{"api_version":"v1","generated_at":"2026-09-29T13:00:45","count":2794,"scope":"judged","papers":[{"id":"2609.30563","version":1,"title":"Thinking Less to Simulate Better: Intuitive Prompting Improves LLM Agents Simulating Individual Social Media Reactions, Including Unfamiliar Content","zh_title":"少思考以更好地模拟：直觉提示提升LLM智能体模拟个体社交媒体反应（包括不熟悉内容）","abstract":"Platform policies are increasingly tested on artificial users, making agent fidelity important. Yet convincing fake profiles could also manipulate perceived public opinion before elections. Validation has concentrated on agreement with human behaviour and has paid little attention to whether an agent behaves in line with the profile it was given. The present study profiled eight Serbian participants through a questionnaire, a deep interview, and a written self-presentation, recorded their reactions to sixty-eight social media posts, and asked four language models to predict those reactions under five prompt conditions varying profile content and instruction style. Attitudinal content improved prediction over demographic backstories by a wide margin. Agents matched their stated profiles more closely than participants matched their own survey answers, and consistency proved unrelated to fidelity once profile information was present. Instructing models to respond intuitively and immediately rather than analytically gave the highest fidelity of any condition and cut the compression of individual differences from seven times the human level to three. The advantage held on posts about topics the questionnaire never raised, where that condition reached the highest fidelity of any setup and beat a crowd baseline by a wide margin, which suggests that agents prompted this way could serve as general-purpose simulated users rather than specialists on the topics they were profiled for. Results may bear implications for the development of language models, because intuition-based setups appear better suited to some tasks than reasoning-based ones.","authors":["Ljubisa Bojic","Tijana Stanic","Joerg Matthes","Agariadne Dwinggo Samala","Bojana Dinic","Jue Wang"],"categories":["cs.AI","cs.CL","cs.HC","cs.MA","cs.SI"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-09-28","first_seen":"2026-09-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.30563","pdf_url":"https://arxiv.org/pdf/2609.30563","source_feed":"cs.CL","score":10,"bucket":"selected","rubric_hits":["A1","A2","A3","B1","B2","B3","B4"],"tags":["LLM仿真","人类行为对照","提示策略"],"reason":"用LLM预测真实个体社交媒体反应，与人类数据对照，评估提示策略对仿真保真度的影…","model":"deepseek-v4-pro","scored_at":"2026-09-28T13:01:44","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-28","rank":1,"question":"在社交媒体反应预测中，不同的提示设计（档案内容与指令风格）如何影响大语言模型模拟特定个体反应的保真度？","design":"用四个大语言模型扮演八名塞尔维亚参与者，基于问卷、深度访谈和书面自我呈现构建个人档案，在五种提示条件下（变化档案内容和指令风格，如人口学背景、态度内容、直觉式回应等）预测他们对68条社交媒体帖子的反应，测量预测与真实反应的一致性。","baseline":"八名塞尔维亚参与者的真实社交媒体反应数据，以及他们自己的问卷回答（用于比较自我一致性）。","findings":"态度性档案内容比人口学背景大幅提升预测准确度；直觉式指令（要求模型凭直觉立即回应而非分析式推理）在所有条件中保真度最高，并将个体差异压缩从人类的七倍降至三倍。","reliability":"论文未讨论","relevance":"该研究直接检验LLM模拟个体行为的保真度，与人类真实反应对照，并揭示提示策略对仿真偏差的影响，对评估LLM作为人类被试替代品的可靠性具有关键参考价值。","inspiration":"值得借鉴的是通过改变提示指令风格（直觉式 vs 分析式）来操纵模型的认知模式，并测量其对个体差异保真度的影响｜可迁移到消费者金融决策实验，如模拟个体在信贷选择或储蓄行为中的异质性反应｜用LLM扮演不同风险偏好和金融素养的消费者，处理为直觉式或分析式提示，结果变量为信贷产品选择或投资决策，与真实消费者调查或实验数据对照。"}},{"id":"2609.30883","version":1,"title":"Warned alike, AI agents avoid the less-crowded road while people take it","zh_title":"同样被警告，AI智能体避开较不拥挤的道路而人类选择它","abstract":"AI agents built on a few shared models increasingly act for many people. A shared forecast about others can align their choices and change how scarce capacity is allocated. We tested this feedback in a two-road congestion game. Adding one sentence warning that others might follow a routing tip made populations of 50 GPT agents crowd one road while avoiding the nearly empty alternative. Average travel time rose from 64 to 95 min, although any crowded-road agent could have saved 69 min by switching alone. The warning discouraged the very move it predicted. The pattern persisted for 100 rounds. Two other model families shifted the same way without locking onto one road. Twelve all-human groups (240 participants) stayed near balance under numerical reports or the tip and warning. In 24 mixed groups with a further 240 participants, imbalance grew with the share of agents in the registered analysis, while people increasingly took the road the agents avoided. Collective costs stayed below the allagent reference, but with 15 agents and 5 humans, agent seats averaged 80 min, compared with 44 min for human seats. Shared forecasts can thus sustain collective inefficiency among similar agents. A better group average can also hide an unequal burden. Evaluations of AI agents that share resources should test populations, treat messages as interventions and report who bears the costs.","authors":["Takahiro Ezaki","Naoto Imura","Katsuhiro Nishinari"],"categories":["physics.soc-ph","cs.AI"],"primary_category":"physics.soc-ph","announce_type":"cross","date":"2026-09-28","first_seen":"2026-09-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.30883","pdf_url":"https://arxiv.org/pdf/2609.30883","source_feed":"cs.AI","score":10,"bucket":"selected","rubric_hits":["A1","A3","B1","B2","B4"],"tags":["LLM仿真","拥堵博弈","人机对照"],"reason":"用GPT agent群体模拟拥堵博弈，并与240名人类被试对照，发现警告导致a…","model":"deepseek-v4-pro","scored_at":"2026-09-28T13:01:45","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-28","rank":2,"question":"共享预测信息（警告）是否会导致AI智能体在拥堵博弈中持续选择拥挤道路，从而造成集体低效？","design":"用50个GPT智能体（gpt-5.4-mini）模拟通勤者，在双路径拥堵博弈中，操纵每日广播信息（无报告、数值报告、提示、提示+警告），测量路径不平衡度、切换比例和平均旅行时间。","baseline":"12个全人类组（240名参与者）在相同博弈和广播条件下的行为数据，以及24个混合组（240名参与者）与智能体共存时的行为。","findings":"警告导致GPT智能体群体持续拥挤一条道路，平均旅行时间从64分钟升至95分钟，尽管个体切换可节省69分钟；全人类组保持接近均衡，混合组中人类更多选择智能体避开的道路，且智能体承担更高成本。","reliability":"论文指出共享预测可导致相似智能体间的集体低效，且群体平均成本可能掩盖负担不均；评估共享资源的AI智能体应测试群体、将消息视为干预并报告成本承担者。","relevance":"高度相关：该研究用LLM智能体模拟人类在拥堵博弈中的决策，并与真实人类数据对照，揭示了共享信息对集体行为的负面影响，直接回应了仿真可靠性与偏差问题。","inspiration":"值得借鉴的是将公共信息（警告）作为干预变量，观察其对群体决策动态的影响，并设置全人类和混合组对照以分离智能体特有行为。｜可迁移到政策公告的预期形成场景，如央行沟通对金融市场参与者行为的影响。｜设计：用LLM智能体模拟投资者，处理为央行发布的不同措辞的前瞻指引，结果变量为资产配置集中度和市场波动率，对照真实投资者在类似公告下的交易数据。"}},{"id":"2609.30896","version":1,"title":"Large language models underestimate and partly misrepresent cultural variation in everyday norms","zh_title":"大语言模型低估并部分误现日常规范的文化差异","abstract":"A key aspect of culture is a society's norms about everyday behavior. How accurately do large language models (LLMs) represent cultural differences in such norms? To answer this question we used the recent Global Study of Everyday Norms (GSEN), which collected ratings of 150 scenarios in 90 societies, as the human benchmark. We prompted GPT-5 to estimate each society's average rating for every scenario, and later repeated the benchmark in three other LLMs: GPT-5.4, Claude Opus 4.6, and Gemini 3.1 Pro. Compared to GSEN estimates, all four LLMs misrepresented cultural variation in two ways. First, they greatly underestimated its magnitude, estimating differences between societies to be, on average, less than half their measured size. Second, for many scenarios the LLMs poorly identified the pattern of variation, that is, which societies judged the behavior less acceptable and which societies judged it more acceptable. The pattern of variation was identified better for scenarios that elicit concerns about vulgarity, especially scenarios involving kissing and flirting. We also found that norms in more developed societies tended to be estimated somewhat more accurately, and that prompting in local survey languages rather than English produced only a modest improvement in accuracy. Local-language prompting also reduced, but did not remove, the underestimation of between-society differences. Cultural differences in everyday norms are only weakly and unevenly represented by LLMs.","authors":["Kimmo Eriksson","Irina Vartanova","Pontus Strimling"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2026-09-28","first_seen":"2026-09-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.30896","pdf_url":"https://arxiv.org/pdf/2609.30896","source_feed":"cs.CY","score":10,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["文化规范","仿真偏差","人类数据对照"],"reason":"用LLM估计社会规范并与真实人类调查数据对照，评估仿真偏差","model":"deepseek-v4-pro","scored_at":"2026-09-28T13:01:45","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-28","rank":3,"question":"大语言模型能否准确表示不同社会之间日常行为规范的文化差异？","design":"用 GPT-5、GPT-5.4、Claude Opus 4.6 和 Gemini 3.1 Pro 四个 LLM，以英语或当地语言提示，估计 90 个社会对 150 个日常行为场景的平均可接受度评分，并与 GSEN 调查数据比较。","baseline":"全球日常规范研究（GSEN）中 90 个社会超过 25,000 名参与者对 150 个场景的评分。","findings":"四个 LLM 都大幅低估了社会间规范差异的幅度，估计的差异平均不到实际测量的一半；对于许多场景，LLM 识别差异模式的能力较差，但在涉及粗俗（如亲吻和调情）的场景上表现较好。","reliability":"论文指出 LLM 对较发达社会的规范估计更准确，用当地语言提示仅带来适度改进，且未能消除对社会间差异的低估；LLM 对文化差异的表示既弱且不均衡。","relevance":"该研究直接评估 LLM 在跨文化社会规范仿真中的偏差，与研究者关注的人类仿真可靠性及失效条件高度契合，值得精读原文以了解具体偏差模式和测量方法。","inspiration":"借鉴其用真实大规模跨国调查作为基准、将仿真误差分解为幅度和模式两个维度的方法，并检验提示语言等处理变量的影响。｜可迁移到跨国消费者行为或政策接受度的仿真研究，例如不同国家居民对环保政策、金融产品条款或广告伦理的接受度差异。｜以 LLM 模拟不同国家被试，对一系列政策或产品场景给出接受度评分，处理变量为提示语言（英语 vs 当地语言），结果变量为评分，与真实跨国调查数据（如世界价值观调查或特定政策民意调查）对照，评估 LLM 仿真的幅度压缩和模式偏差。"}},{"id":"2604.20050","version":4,"title":"Information Aggregation with AI Agents","zh_title":"AI代理的信息聚合研究","abstract":"Can Large Language Models (AI agents) aggregate dispersed private information through trading and reason about the knowledge of others by observing price movements? We conduct a controlled experiment where AI agents trade in a prediction market after receiving private signals, across four information structures of increasing complexity. We find that although the median market is effective at aggregating information in the easy information structures, performance deteriorates in the harder structures, suggesting that AI agents struggle in environments where more than two levels of interactive reasoning are required, a ceiling close to the one documented in human subjects. Consistent with our theoretical predictions, market accuracy does not improve from allowing cheap talk communication, changing the duration of the market, or strategic prompting; initial price has little average effect but matters in the very hard structure. We also find that ``smarter'' AI agents perform better at aggregation and are more profitable. Surprisingly, giving them feedback about past performance does not improve aggregation. A further wave of markets, run three months later with capability-frontier models, aggregates information more often in the three easier structures but not in the hardest one, where higher capability replaces markets that are confidently wrong with markets that hedge near 0.5.","authors":["Spyros Galanis"],"categories":["econ.GN","cs.AI","cs.GT","q-fin.EC"],"primary_category":"econ.GN","announce_type":"replace-cross","date":"2026-09-28","first_seen":"2026-04-21","revised_at":"2026-09-28","abs_url":"https://arxiv.org/abs/2604.20050","pdf_url":"https://arxiv.org/pdf/2604.20050","source_feed":"cs.AI","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2","B4"],"tags":["LLM仿真","信息聚合","行为实验"],"reason":"用AI代理模拟人类交易行为，并与人类被试结果对照，评估信息聚合能力。","model":"deepseek-v4-pro","scored_at":"2026-09-28T13:02:01","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-28","rank":4,"question":"AI 代理能否通过交易聚合分散的私人信息，并像人类一样通过观察价格变动推断他人知识？","design":"用八种大语言模型（Claude Haiku 3.5/4.5、Gemini 2.5/3 Flash、GPT-4o/5 mini、gemma3:4b、qwen3:8b）组成三人交易团队，在四种难度递增的信息结构中交易预测市场证券；处理包括允许廉价谈话、战略提示、初始价格（0.3/0.5/0.7）、市场时长（3/6/9轮）和反馈信息，共144种配置，每种至少运行12次，生成1772个市场；结果变量为市场准确率（价格是否接近真实价值）和利润。","baseline":"人类被试在类似信息结构中的推理层级上限（约两级交互推理），来自已有实验文献。","findings":"在简单信息结构中，中位数市场能有效聚合信息，但在需要超过两级交互推理的困难结构中表现恶化；允许廉价谈话、改变市场时长或战略提示均未提高市场准确率，初始价格平均影响小但在极难结构中重要。更聪明的AI代理聚合更好且更盈利，但反馈过去表现未改善聚合；三个月后用前沿模型重跑，在三个较易结构中聚合更频繁，但在最难结构中未改善，高能力模型将自信错误的市场替换为在0.5附近对冲的市场。","reliability":"论文承认AI代理在需要超过两级交互推理的环境中挣扎，且能力提升并未解决最难结构中的聚合失败；未明确讨论其他失效条件。","relevance":"该研究直接以AI代理模拟人类交易者，并与人类被试的推理层级上限对照，评估信息聚合能力，属于用LLM进行人类仿真实验并检验可靠性的核心工作，值得精读。","inspiration":"借鉴其系统操纵信息结构复杂度、市场机制参数（时长、初始价格、沟通）和模型能力来测量聚合效率与利润的做法，并设置理论基准（可分离证券的完全聚合预测）｜可迁移到资产定价实验中的信息效率研究，如内幕交易监管、分析师预测市场或央行沟通对价格发现的影响｜用不同能力LLM作为交易者，在预测市场中交易与真实宏观经济指标挂钩的证券，处理为是否允许公开评论或改变交易轮次，结果变量为价格偏离真实值的程度，并与人类实验数据（如Plott & Sunder的经典信息聚合实验）对照。"}},{"id":"2609.02729","version":2,"title":"BuildOcc: A Large Language Model Occupant Agent Platform for Building Energy Research","zh_title":"BuildOcc：用于建筑能源研究的大语言模型居住者智能体平台","abstract":"Occupants are a primary source of uncertainty in building energy consumption and management, yet existing occupant behavior models cannot capture adaptive and reasoning responses considering the occupant's personal history, current context, and the type of energy signal being delivered. This study presents BuildOcc, an open-source Python platform that grounds large language model agents in the American Time Use Survey (ATUS), a nationally representative diary dataset covering 16,684 respondents. Through BuildOcc, each simulated occupant agent can be instantiated with a demographic persona drawn from ATUS population statistics, a memory stream that accumulates and reflects on timestep-level observations, and an activity scheduler that samples empirically from ATUS time-at-activity distributions. The platform exposes a three-layer interface - Python library, REST API, and Model Context Protocol server - so that any building energy tool (EnergyPlus, Home Assistant) can integrate behavioral intelligence without bespoke coupling code. A plugin registry lets the community add new occupant strata, custom schedulers, and alternative memory backends as separate installable packages. Two validation tiers show that ATUS-grounded sampling reproduces empirically calibrated activity distributions and that demographic priors propagate into persona-consistent agent reasoning across timesteps, establishing internal consistency across strata. BuildOcc provides the building energy community with a reusable, openly available implementation of the occupant behavioral layer. BuildOcc is openly released at https://doi.org/10.5281/zenodo.21192895 under the Apache License 2.0 and installable via pip install buildocc.","authors":["Wooyoung Jung"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"replace","date":"2026-09-28","first_seen":"2026-09-03","revised_at":"2026-09-28","abs_url":"https://arxiv.org/abs/2609.02729","pdf_url":"https://arxiv.org/pdf/2609.02729","source_feed":"cs.HC","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM仿真","建筑能源","ATUS数据"],"reason":"用LLM agent模拟建筑内人员行为，基于ATUS真实数据对照，属于人类仿真…","model":"deepseek-v4-pro","scored_at":"2026-09-28T13:02:02","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-28","rank":5,"question":"如何构建一个以美国时间使用调查（ATUS）为数据基础、基于大语言模型（LLM）的居住者智能体平台，用于建筑能耗研究中的居住者行为仿真？","design":"该研究开发了 BuildOcc 平台，使用 LLM 智能体模拟四类美国人口群体（在职单身成年人、退休夫妇、在职父母、无业成年人）。每个智能体包含从 ATUS 人口统计中抽取的人口特征、从 ATUS 数据中采样的活动日程、记忆流和推理引擎。在每个 15 分钟时间步，智能体根据人口特征、记忆和当前环境选择下一个活动并说明理由，同时也会对需求响应信号做出接受、拒绝或推迟的决定。平台提供 Python 库、REST API 和 MCP 服务器三层接口，便于与 EnergyPlus 等工具集成。","baseline":"美国时间使用调查（ATUS），包含 16,684 名受访者的全国代表性日记数据，其中 6,611 名受访者属于四个目标人口阶层。","findings":"第一层验证表明，活动调度器能够将 ATUS 活动分布复现到采样噪声范围内；第二层验证表明，人口先验信息能够传播到智能体的推理中，产生与人口特征一致的行为差异。","reliability":"论文承认的局限包括：每个时间步只能选择一个活动；ATUS 仅覆盖美国人口；活动仅基于主要活动（ATUS 一次只记录一个活动）；记忆重要性分数由智能体自行分配并用于检索，缺乏外部校准或反馈路径。","relevance":"该研究将 LLM 智能体与全国代表性时间使用调查数据结合，用于模拟人类行为，并进行了与真实数据的对照验证，属于人类仿真研究，但场景限定于建筑能耗领域，与经济金融问题关联度较低。","inspiration":"借鉴其将 LLM 智能体与大规模调查数据结合、通过分层抽样和记忆机制生成个体行为的方法，可用于构建具有人口代表性的经济决策仿真。｜可迁移到消费者跨期选择或家庭能源消费行为研究，例如模拟不同人口群体对动态电价或节能政策的反应。｜设计雏形：以美国消费者支出调查（CEX）或收入动态面板研究（PSID）为数据基础，构建 LLM 智能体代表不同收入阶层，施加电价上涨或补贴政策处理，结果变量为能源消费和支出变化，并与真实调查数据对照验证。"}},{"id":"2609.30940","version":1,"title":"Financial Fragility in Societies of LLM Agents: Coordination Failures and Stabilizing Mechanisms","zh_title":"LLM智能体社会中的金融脆弱性：协调失败与稳定机制","abstract":"Individually protective decisions can produce avoidable collective failures. As large language model (LLM) agents take on greater roles in financial decision-making, financial AI safety must therefore be considered not only at the level of individual agents, but also at the level of the systems they jointly create. We study this problem with FRAIL, a controlled experimental framework that places LLM agents in three dynamic financial environments---bank runs, debt rollover, and reward crowdfunding---where agents' decisions reshape the financial conditions faced by others. Across seven leading LLMs, we find widespread collective fragility even when no agent is instructed to destabilize the system: 77\\% of baseline bank-run episodes and 83\\% of debt-rollover episodes end in failure. We then compare three interaction mechanisms based on compensated commitments, centralized commitment agreements, and participant-led coalitions. All three improve aggregate outcomes, but no single mechanism performs best across all financial structures. Across mechanisms, successful stabilization shares a common temporal pattern: broad commitment forms early, before defensive behavior becomes self-reinforcing. Our findings show that individually capable agents do not automatically form safe financial systems, highlighting system-level evaluation and interaction design as central problems for financial AI safety. Code is available at https://anonymous.4open.science/r/FinFrail-CF26.","authors":["Zhenhao Fu","Ruipeng Xu","Qibing Ren"],"categories":["cs.AI","q-fin.GN"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-28","first_seen":"2026-09-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.30940","pdf_url":"https://arxiv.org/pdf/2609.30940","source_feed":"cs.AI","score":8,"bucket":"selected","rubric_hits":["A3","B2","B4"],"tags":["LLM智能体","金融仿真","协调失败"],"reason":"用LLM agent模拟金融系统中的协调失败，涉及经济场景，但无真实人类数据对…","model":"deepseek-v4-pro","scored_at":"2026-09-28T13:01:45","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-28","rank":6,"question":"多个LLM智能体在共享金融环境中是否会自发产生系统性金融脆弱性，以及如何通过交互机制设计来缓解这种脆弱性？","design":"使用FRAIL框架，将七个主流LLM作为金融决策者，分别置于银行挤兑、债务展期和奖励众筹三种动态金融环境中，通过多轮交互模拟决策，测量系统失败率（如银行倒闭、债务违约、众筹失败）和承诺形成的时间模式。","baseline":"无对照","findings":"在基线条件下，77%的银行挤兑和83%的债务展期情景以失败告终，即使没有智能体被指示破坏系统。三种交互机制（补偿承诺、集中承诺协议、参与者主导联盟）均能改善结果，但最佳机制因金融结构而异，成功稳定依赖于在防御行为自我强化之前尽早形成广泛承诺。","reliability":"论文未讨论","relevance":"该研究用LLM智能体模拟金融协调失败，属于经济场景下的多智能体仿真，但缺乏真实人类数据对照，适合关注LLM仿真方法本身或金融AI安全的研究者阅读。","inspiration":"借鉴其动态金融环境设计和多智能体交互机制比较，可迁移到资产定价实验或政策公告预期形成等场景。｜可设计一个信贷审批歧视实验，用LLM扮演银行信贷员和借款人，施加不同信息披露政策作为处理，测量贷款批准率和违约率，并与真实信贷数据对照。"}},{"id":"2609.31054","version":1,"title":"Cheap, open agents make LLM pollution harder to mitigate","zh_title":"廉价开放智能体使LLM污染更难缓解","abstract":"Large Language Model (LLM) pollution occurs when synthetic responses contaminate data intended to capture human behavior. High deployment costs have so far limited the risk posed by autonomous survey agents. However, open-weight models paired with open-source agentic frameworks may have removed this barrier. We compared the performance and detectability of nine agent configurations, ranging from fully open variants to closed commercial ones. Each agent autonomously completed a survey containing multiple response types yielding various detection checks. Fully open agents ran locally without usage fees and performed competitively with commercial alternatives. Open and commercial agents failed different sets of checks, and no single check reliably detected all agents, but open-text responses discriminated best between agents and humans. These findings identify fully open agents as a distinct risk for LLM pollution and support multilayered detection strategies emphasizing open-text analysis.","authors":["Raluca Rilla","Anne-Marie Nussberger","Rui Mata","Dirk U. Wulff"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-28","first_seen":"2026-09-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.31054","pdf_url":"https://arxiv.org/pdf/2609.31054","source_feed":"cs.AI","score":8,"bucket":"selected","rubric_hits":["A2","B1","B4"],"tags":["LLM污染","调查数据","检测方法"],"reason":"研究LLM污染人类调查数据，评估检测方法，与仿真可靠性直接相关。","model":"deepseek-v4-pro","scored_at":"2026-09-28T13:01:47","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-28","rank":7,"question":"完全开源的自主智能体是否降低了LLM污染人类调查数据的门槛，以及现有检测方法能否有效识别这些智能体？","design":"使用九种智能体配置（三种完全开源、两种混合、四种专有）自主完成同一份调查问卷，每种配置运行40次，共360次运行。智能体被赋予随机的人口统计特征（性别、年龄），并收到统一指令。调查包含多种回答类型和18项检测检查（16项通过/失败检查和2项计时测量），用于评估智能体的表现和可检测性。","baseline":"3,242名来自Prolific的美国参与者，在2025年10月20日至11月4日期间完成同一调查问卷，样本在性别、年龄和种族上大致代表美国人口。","findings":"完全开源的智能体在本地运行且无使用费用，其表现与商业替代方案相当。开源和商业智能体在不同检查上失败，没有单一检查能可靠地检测所有智能体，但开放式文本回答在区分智能体和人类方面效果最好。","reliability":"论文指出，开源和商业智能体在不同检查上失败，没有单一检查能可靠地检测所有智能体，因此需要多层检测策略，特别是强调开放式文本分析。此外，研究仅操纵了年龄和性别两个人口统计变量，未涵盖其他可能影响回答的变量；智能体被明确指示避免与可疑嵌入命令交互，这可能降低了某些检测的失败率。","relevance":"该研究直接评估了LLM智能体对人类调查数据的污染风险及检测方法，与您关注的LLM仿真可靠性和偏差问题高度相关，特别是它提供了真实人类数据作为对照，并揭示了开源智能体带来的新风险，值得精读原文以了解具体检测方法和失效模式。","inspiration":"借鉴其多层检测策略和开放式文本分析来识别LLM生成回答的方法，可迁移到经济金融领域的调查数据质量控制中。｜可应用于消费者信心调查、投资者情绪调查或政策评估中的问卷数据，检测是否存在LLM污染。｜设计：以真实人类调查数据（如密歇根大学消费者信心调查）为基准，让开源和商业LLM智能体自主完成同一问卷，比较其回答分布和开放式文本特征，并开发基于文本分析的检测指标，评估不同检测方法的敏感性和特异性。"}},{"id":"2608.27167","version":2,"title":"Calibrated Enough to Know, Not Calibrated to Act: Fabricated Evidence Makes LLM Agents Commit to the Unknowable","zh_title":"校准到知道，但未校准到行动：伪造证据使LLM智能体对不可知问题做出承诺","abstract":"An LLM agent shown a professional-looking market panel commits to a directional call on a provably unpredictable question far more often than one asked the bare question: across 12 frontier models, commitment rises from 6.5% to 54.0% as evidence is escalated. It commits just as readily when every number on the panel is invented: fabricating the entire display, so nothing the model can see is true except the question itself, still lifts commitment from 24.5% to 36.8%, statistically indistinguishable from the 37.6% produced by genuine market data. What unlocks confident action is not information but the authority of its packaging. The failure is narrow and locatable. Incapacity is not the answer: on matched answerable questions attached to the same panels, the same models answer essentially always, at near-perfect accuracy. Nor is it belief - stated probabilities barely move across the gradient that swings action by 48 points, and score worse than a climatological baseline. Missing judgment isn't it either: asked to classify a question's knowability before acting, models call it irreducible 90% of the time and then commit on just 0.4% of those. The act/don't-act gate is what fails, and the effect is concentrated in a few models rather than universal. Because the gate is separable, it can be trained. Supervised fine-tuning of a 3B model on 540 synthetic cases, predominantly dice, coins, jars and timers, drives commitment to 0.0% on the original cases and transfers to three unseen domains. It does not survive everything: the gate holds exactly when the response format leaves room to reason, and rigid formats that remove that room leave the model confident and wrong on questions it otherwise answers correctly. The gate is trainable and context-fragile, and deployment needs both halves of that sentence.","authors":["Pranav Aggarwal"],"categories":["cs.AI","cs.CL"],"primary_category":"cs.AI","announce_type":"replace-cross","date":"2026-09-28","first_seen":"2026-08-28","revised_at":"2026-09-28","abs_url":"https://arxiv.org/abs/2608.27167","pdf_url":"https://arxiv.org/pdf/2608.27167","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A2","B4"],"tags":["LLM决策偏差","可靠性评估","批判性研究"],"reason":"研究LLM在不可知问题上的决策偏差，评估其可靠性，批判性指出失效条件，可迁移到…","model":"deepseek-v4-pro","scored_at":"2026-09-28T13:02:02","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-28","rank":8,"question":"在不可知（aleatoric）问题上，专业外观的证据面板是否仅凭其权威包装就能诱导LLM智能体做出方向性承诺，而非真正提供信息？","design":"研究用12个前沿LLM作为被试，在四个领域（股票、加密货币、体育、天气）构造不可知问题，施加不同证据梯度（无证据、薄面板、丰富面板、乱序面板、全伪造面板），测量模型是否做出方向性承诺（行动）以及陈述的概率，并与可回答的匹配问题对照。","baseline":"无对照","findings":"伪造证据面板与真实面板诱导的承诺率在统计上无差异，表明触发行动的是包装的权威性而非信息内容；模型的陈述概率几乎不随证据梯度变化，且比气候学基线更差，但同一模型在可回答问题上几乎完美作答。","reliability":"论文承认天气领域无密封结果，面板提供的集合预报概率可能具有真实技能，因此该领域作为不可知性工具最弱；训练出的门控在刚性响应格式下失效，需要留出推理空间。","relevance":"该研究直接评估LLM在不可知问题上的决策偏差，揭示仿真失效的特定条件（证据包装触发行动），对关注LLM仿真可靠性与偏差的研究者具有重要参考价值，值得阅读原文。","inspiration":"借鉴其通过伪造证据与真实证据对比来分离信息与包装的因果设计，以及用密封结果验证不可知性并测量行动而非仅测信念的做法｜可迁移到资产定价实验或政策公告的预期形成研究，例如测试LLM代理在呈现专业外观的虚假市场数据时是否会产生过度自信的交易决策｜用LLM作为被试，随机分配真实与伪造的市场分析面板，测量其买卖决策和置信度，并与人类实验数据或历史市场结果对照，检验权威包装对决策的影响。"}},{"id":"2609.30867","version":1,"title":"Evidence-Grounded Auditing of Identification Assumptions in Climate-Policy Causal Evaluations","zh_title":"气候政策因果评估中识别假设的证据基础审计","abstract":"Difference-in-differences (DID) studies are widely used to evaluate climate policy, but assessing the evidence supporting their identification assumptions remains challenging. We introduce ARGUS, a structured language-model pipeline that audits reported evidence against an eleven-dimension assumption-implication-evidence rubric and abstains when relevant evidence cannot be retrieved. We evaluate ARGUS using injected flaws, economics papers, and a small pilot with reconciled labels. On the 11-flaw benchmark, ARGUS detects 73% of planted flaws, compared with 18% for a keyword-based pipeline. Across 26 economics papers, ARGUS abstains on about 40% of paper-dimension assessments for lack of retrievable evidence. In a five-paper pilot with labels reconciled by two annotators, it assigns a higher risk level than the labels on 25 of the 33 assessments it completes. A rule fixed before the labels arrived removes most of this in-sample; weighted agreement stays low. ARGUS provides evidence-linked risk reports that localize potential weaknesses for expert review, without adjudicating causal claims. Code and data: https://github.com/yonghongzhang-io/ARGUS","authors":["Yonghong Zhang","Yong Xie","Isabel M. Parra","Ricardo Correia"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-28","first_seen":"2026-09-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.30867","pdf_url":"https://arxiv.org/pdf/2609.30867","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A4","B3"],"tags":["LLM审计","因果推断","方法论"],"reason":"提出LLM审计因果推断假设的方法论框架，可迁移到仿真可靠性评估","model":"deepseek-v4-pro","scored_at":"2026-09-28T13:01:53","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-28","rank":10,"question":"如何用大语言模型审计气候政策因果评估中双重差分法识别假设的证据支持度？","design":"本文提出 ARGUS 流水线，用大语言模型对双重差分研究的 11 个识别维度进行证据检索与充分性评估，并输出风险报告；模型仅负责检索和评估，控制流固定。","baseline":"无对照","findings":"在注入缺陷基准上，ARGUS 检出 73% 的植入缺陷，而关键词流水线仅检出 18%；在 26 篇经济学论文上，约 40% 的论文-维度评估因缺乏可检索证据而弃权。","reliability":"论文承认在真实论文上证据覆盖不足导致弃权率高，且过度严厉导致与人工标签一致性低；气候政策专用语料上的表现未测试。","relevance":"该研究提出用 LLM 审计因果推断假设的方法论框架，可迁移到仿真可靠性评估，值得阅读原文了解其证据门控与弃权机制。","inspiration":"值得借鉴的是将识别假设分解为可审计的维度，并用检索门控控制 LLM 的评估范围，避免模型自由发挥｜可迁移到政策评估中的因果推断可靠性审计，例如碳税、补贴或监管政策的效果评估｜可设计一个仿真实验：用 LLM 扮演审稿人，对一组已发表的经济学论文的识别假设进行审计，处理是提供不同质量的证据，结果变量是风险评级，并与人类专家评级对照。"}},{"id":"2609.31245","version":1,"title":"RupeeBias: Auditing Demographic Bias in Indian Economic Guidance from Large Language Models","zh_title":"RupeeBias：审计印度经济指导中大语言模型的人口统计偏差","abstract":"Individuals turn to large language models (LLMs) for guidance across a wide range of economic tasks, from comparing loan options and planning savings to deciding what raise to ask for or how much to charge for their services. LLMs are known to reproduce social biases, and biased economic guidance may influence what users believe they are worth, what they ask for, and what they ultimately accept. This risk is especially salient in India, where economic outcomes are shaped by demographic categories such as caste and urban-rural location. Existing LLM bias benchmarks, however, are largely designed around Western demographic categories and therefore miss key axes of economic disparity in the Indian context. We introduce RupeeBias, a benchmark for auditing demographic bias in LLM-generated economic guidance across Indian economic settings. RupeeBias consists of 39,150 prompts spanning four use cases: salary estimation, salary increment estimation, counter-offer recommendation, and service pricing recommendation. The benchmark follows a single-attribute counterfactual design, holding the description of the user's qualifications, experience, or service offering fixed while varying one demographic identifier at a time. RupeeBias covers 87 India-specific demographic identifiers across six axes: caste, religion, regional identity, gender, disability, and urban-rural location, with all prompts constructed in both English and Hinglish. We evaluate nine LLMs on RupeeBias and find systematic demographic disparities across all six axes. For otherwise identical prompts that differ only in demographic identifier, LLM-generated economic outputs differ by 20.2% on average. We publicly release RupeeBias to support future research on demographic bias in LLM-generated economic guidance across India-specific demographic and economic contexts.","authors":["Pavithra P M Nair","Bhavik Talaviya","Shourya Bhushan","Rahul Pankajakshan","Seema Guruvadoo","Avinash Agarwal","Gilad Gressel","Krishnashree Achuthan"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-28","first_seen":"2026-09-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.31245","pdf_url":"https://arxiv.org/pdf/2609.31245","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A1","B1","B2","B4"],"tags":["LLM偏差审计","经济决策仿真","人口统计代表性"],"reason":"用LLM生成经济建议并审计人口统计偏差，有真实人类数据对照，涉及经济决策场景，…","model":"deepseek-v4-pro","scored_at":"2026-09-28T13:01:48","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-28","rank":13,"question":"在印度经济咨询场景中，大语言模型生成的经济建议是否因用户的人口统计特征（如种姓、宗教、性别等）而产生系统性偏差？","design":"构建 RupeeBias 基准，包含 39,150 个提示，覆盖薪资估算、加薪估算、还价建议和服务定价建议四个用例。采用单属性反事实设计，固定用户资质、经验或服务描述，仅改变一个印度特定的人口统计标识符（共 87 个，涵盖种姓、宗教、地域、性别、残疾和城乡位置六个轴），并同时用英语和印地英语构建提示。评估九个 LLM 在这些提示上的输出差异。","baseline":"无对照","findings":"在仅人口统计标识符不同的相同提示下，LLM 生成的经济输出平均差异达 20.2%，且在所有六个轴上都存在系统性人口统计差异。","reliability":"论文未讨论","relevance":"该研究直接针对 LLM 在经济学场景中的仿真偏差，使用反事实设计审计人口统计特征对经济建议的影响，与研究者关注的 LLM 人类仿真可靠性及偏差问题高度相关，值得阅读原文了解具体偏差模式和测量方法。","inspiration":"借鉴其单属性反事实设计，通过仅改变一个属性来隔离人口统计特征对模型输出的因果影响，并构建大规模、多语言、多场景的提示集｜可迁移到信贷审批歧视、保险定价、工资谈判等经济金融决策场景，审计 LLM 在金融建议中的群体偏差｜以 LLM 作为虚拟被试，处理为在贷款申请或保险报价提示中改变申请人姓名、性别或种族等标识，结果变量为模型给出的利率、保费或授信额度，并与真实信贷或保险数据中的群体差异进行对照，检验模型偏差是否反映或放大现实歧视。"}},{"id":"2609.31013","version":1,"title":"Same Text, Different Numbers: The Divergence of LLM-Based Measures","zh_title":"相同文本，不同数字：基于LLM的测量分歧","abstract":"Researchers increasingly use generative large language models (LLMs) to convert corporate text into empirical variables. We examine the extent to which LLM-based textual measures are invariant to model choice using thirteen measures, including sentiment, management clarity, uncertainty, answer specificity, and climate and political risk. Seven LLMs from different providers score earnings call transcripts of S&P 500 companies on these constructs. Cross-model rank correlations average only 0.52, and transcript-level differences common across providers account for only 34% of total score variation. Cross-model disagreement does not predict subsequent analyst or market disagreement, consistent with a substantial model-specific component rather than common ambiguity in the underlying disclosure. Model choice significantly affects downstream inference, with coefficient magnitudes, signs, and statistical significance varying substantially across models. Averaging across providers makes transcript rankings more stable for most constructs, but score levels remain sensitive to the models included in the ensemble. LLM-generated variables should therefore be treated as model-contingent measurements and validated across providers.","authors":["Hamid Boustanifar","Sasan Mansouri"],"categories":["cs.AI","cs.CL","q-fin.GN","q-fin.RM"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-09-28","first_seen":"2026-09-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.31013","pdf_url":"https://arxiv.org/pdf/2609.31013","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A2","B1","B3"],"tags":["LLM测量一致性","算法保真度","文本分析"],"reason":"评估LLM文本测量跨模型一致性，涉及测量偏差与统计推断，有真实数据对照，方法可…","model":"deepseek-v4-pro","scored_at":"2026-09-28T13:01:46","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-28","rank":11,"question":"LLM 生成的文本测量在多大程度上不受模型选择影响？","design":"用七家不同提供商的 LLM 对同一批标普 500 公司财报电话会议问答环节文本，按相同提示词对 13 个构念打分，比较跨模型分数的一致性、方差来源及对下游回归推断的影响。","baseline":"以 Loughran-McDonald 等既有词典法文本测量作为对照，并考察 LLM 分歧与分析师预测分歧、市场反应等真实人类判断的关联。","findings":"跨模型平均秩相关仅 0.52，且构念间差异大；方差分解显示转录本共同成分平均只占 34%，模型特定成分显著。模型选择会实质改变下游回归系数的符号、大小和显著性，而 LLM 分歧与后续分析师或市场分歧无关。","reliability":"论文指出 LLM 测量应视为模型依赖的，需跨提供商验证；自报置信度不能识别更可靠的观测，规则化提示词反而可能增大分歧，且大部分分歧无法由可观测特征解释。","relevance":"直接评估 LLM 作为测量工具在金融文本分析中的可靠性，揭示模型选择对实证结论的威胁，对关注仿真偏差和稳健性的研究者有重要参考价值。","inspiration":"借鉴其多模型同文本打分并做方差分解的设计，可系统检验测量工具的选择效应｜可迁移到政策公告文本的情绪或不确定性测量、信贷审批中的软信息编码、分析师报告语调等场景｜用多家 LLM 对同一批政策声明或贷款申请文本按相同提示词打分，以人工编码或市场反应数据为基准，比较跨模型一致性及对后续回归推断的影响。"}},{"id":"2609.31468","version":1,"title":"PriceBench: A Diagnostic Benchmark for Price, Quality, and Brand Preferences in LLM Booking Agents","zh_title":"PriceBench：LLM预订代理中价格、质量与品牌偏好的诊断基准","abstract":"LLMs increasingly act as purchasing agents, which makes the LLM, not the user, the one choosing among the options that satisfy a request; its preferences quietly fix what gets bought and what it costs. Hotel booking is a clean instance: a high-volume choice settled on a few comparable attributes, where the pick reveals those preferences. We introduce PriceBench, a diagnostic benchmark that recovers an LLM's price, quality, and brand preferences from its booking choices with a logit choice model, applied to 28 LLMs from 8 providers on 3,600 hotel tasks from 179 real New York City properties. We find that capability is associated with how consistently an LLM chooses, not with what it chooses: more capable LLMs hold stronger, more consistent preferences, while weaker ones either lock onto one position, exploitable by whoever controls listing order, or choose almost indifferently. What those preferences favor varies sharply across providers and even within one family: price sensitivity spans more than an order of magnitude, and the price/quality trade-off moves mean booked nightly price from \\$247 to \\$393 on identical tasks. What an agent buys must therefore be measured per LLM, not inferred, and we release the tasks, code, and all 28 response sets.","authors":["Pavel Kireyev"],"categories":["econ.GN","cs.AI","cs.CL","q-fin.EC"],"primary_category":"econ.GN","announce_type":"cross","date":"2026-09-28","first_seen":"2026-09-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.31468","pdf_url":"https://arxiv.org/pdf/2609.31468","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A1","B1","B2"],"tags":["LLM仿真","消费者选择","偏好测量"],"reason":"用LLM替代人类消费者进行预订选择，并与真实酒店数据对照，属于经济学场景仿真","model":"deepseek-v4-pro","scored_at":"2026-09-28T13:01:49","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-28","rank":14,"question":"LLM作为酒店预订代理时，其价格、质量和品牌偏好是什么，这些偏好在不同LLM之间有何差异？","design":"构建PriceBench基准，包含3600个酒店选择任务（1800个二元、1800个三元），基于179家真实纽约酒店，属性包括星级、价格、评分等。对28个LLM进行测试，每个任务以原始顺序和交换顺序各评分一次，使用logit选择模型从选择中恢复偏好参数。","baseline":"无对照（论文未使用真实人类预订数据作为基准，但引用了人类预订者对质量的估值范围进行对比）。","findings":"能力与选择的一致性相关，而非选择内容：更强的LLM偏好更一致，较弱的LLM要么锁定位置（受列表顺序控制），要么几乎无差异。价格敏感度跨LLM差异超过一个数量级，价格/质量权衡导致相同任务下平均每晚预订价格从247美元到393美元不等。","reliability":"论文讨论了位置锁定LLM在温度0下失效，采样无法恢复偏好；偏好幅度受提示格式、解码方式和选项数量影响；品牌偏好与提供商无关。","relevance":"该研究用LLM替代人类消费者进行预订选择，属于经济学场景仿真，但缺乏真实人类对照，主要提供LLM间差异的诊断，对关注仿真可靠性和偏差的研究者有参考价值。","inspiration":"借鉴其通过随机化属性、重复呈现和logit模型识别偏好的方法，可迁移到消费者选择、定价策略等经济问题。｜例如，在信贷审批歧视研究中，用LLM扮演信贷员，改变申请人属性（如种族、性别）观察决策差异。｜设计：以LLM为被试，呈现贷款申请（处理为不同种族/性别信号），结果变量为批准决策，对照真实信贷员历史数据评估偏差。"}},{"id":"2609.30705","version":1,"title":"The Price of Thought: Does Test-Time Reasoning Pay in LLM Trading?","zh_title":"思考的代价：测试时推理在LLM交易中是否值得？","abstract":"While inference-time reasoning in large language models (LLMs) promises better decision making, its higher computational cost may not yield better economic outcomes. Yet reasoning controls are rarely evaluated as economic interventions, where changes in model outputs must translate into better portfolios after trading costs. We conduct a controlled study of representative LLMs from the DeepSeek, GPT, and Gemini families. We vary reasoning effort while holding information available at each formation date, prompts, output formats, and portfolio construction fixed. Our evaluation covers a full year of U.S. equities under three input conditions: numerical, identifiable news, and masked news. It includes more than 800,000 asset predictions and repeated model generations. Across all three model families, additional reasoning does not produce a reliable improvement in net portfolio returns. For DeepSeek, where we examine the full progression from no reasoning to maximum reasoning, performance is nonmonotonic. Repeated generations also produce unstable treatment effects and portfolio selections, even when overall scores remain similar. These findings show that additional reasoning can change financial decisions without reliably improving their economic value, motivating validation for each task before deployment.","authors":["Jiayi Chen","Guiling Wang"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-28","first_seen":"2026-09-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.30705","pdf_url":"https://arxiv.org/pdf/2609.30705","source_feed":"cs.AI","score":7,"bucket":"pending","rubric_hits":["A3","B2","B4"],"tags":["LLM交易","推理成本","经济决策"],"reason":"用LLM模拟金融决策并与真实市场数据对照，涉及经济场景和失效条件，但非人类被试…","model":"deepseek-v4-pro","scored_at":"2026-09-28T13:01:44","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-28","rank":9,"question":"在固定模型、信息、提示、输出格式和投资组合规则的情况下，增加推理努力是否可靠地改善LLM交易组合的净收益？","design":"对DeepSeek、GPT和Gemini三个模型家族的代表性LLM进行受控干预，改变推理努力水平（无、低、高、最大），在2024年241个形成日对100只美国流动性股票进行评分，基于数值特征、可识别新闻和掩蔽新闻三种输入条件，构建等权市场中性投资组合（买入前10、卖空后10），持有5天，扣除10个基点的交易成本，比较不同推理水平下的净收益。","baseline":"无对照","findings":"在三个模型家族中，增加推理并未可靠地改善净投资组合收益；DeepSeek从无推理到最大推理的表现是非单调的。重复生成显示处理效应和投资组合选择不稳定，即使总体得分相似。","reliability":"论文承认推理水平的变化会改变输出结构可靠性，但可靠性提高并不自动意味着收益提高；重复生成导致处理效应不稳定，表明结果可能因随机生成而异；研究仅覆盖一年美国股票数据，且置信区间较宽，不能证明所有效应为零。","relevance":"该研究用LLM模拟金融决策并与真实市场数据对照，评估推理努力对经济结果的影响，属于经济场景下的仿真研究，并揭示了仿真失效条件（推理增加不带来稳定收益），值得阅读原文以了解其受控设计和稳健性检验方法。","inspiration":"借鉴其受控干预设计：在固定其他因素下仅改变推理努力，并使用重复生成和冻结审计评估稳定性｜可迁移到资产定价实验或投资决策仿真，检验LLM推理深度对预测准确性和组合表现的影响｜以LLM为被试，处理为不同推理水平，结果变量为组合净收益或预测误差，用真实历史市场数据作为基准对照。"}},{"id":"2609.31095","version":1,"title":"Confident, Not Wiser: The Dunning-Kruger Effect in Human-AI Interaction","zh_title":"自信而非更明智：人机交互中的达克效应","abstract":"AI assistance can improve performance without improving self-assessment. We report a study (N=366) comparing Human alone and Human+AI performance on reasoning tasks, for which the AI model is benchmarked on the same items. Participants estimated global and block performance and rated confidence in their answers. Human+AI achieved higher scores, but self-estimates tracked performance weakly. Average overestimation was similar across groups, covering individual errors. Across tasks, confidence distinguished correct from incorrect answers less accurately in the Human+AI group, while within-task differences remained uncertain. The Dunning-Kruger pattern was found in both groups, with a larger observed contrast in Human+AI. Controls for score noise reduced but did not eliminate the pattern, with the controlled group difference remaining inconclusive. An extended computational account describes global and block estimates. Our findings distinguish performance augmentation from metacognitive augmentation and motivate interfaces that support verification, communicate task-specific AI model performance, and help users evaluate the quality of their joint work rather than produce answers.","authors":["Daniela Fernandes","Michelle Rausch","Agnes Mercedes Kloft","Daniel Buschek","Robin Welsch"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-28","first_seen":"2026-09-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.31095","pdf_url":"https://arxiv.org/pdf/2609.31095","source_feed":"cs.HC","score":7,"bucket":"pending","rubric_hits":["A2","B1","B4"],"tags":["人机交互","元认知","AI辅助决策"],"reason":"研究人类与AI协作中的元认知偏差，有真实人类数据对照，结论可迁移到LLM仿真可…","model":"deepseek-v4-pro","scored_at":"2026-09-28T13:01:47","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-28","rank":12,"question":"人类与AI协作时，AI辅助是否在提升任务表现的同时改善元认知（自我评估准确性、信心区分度），以及邓宁-克鲁格效应在有无AI辅助下如何表现。","design":"本研究并非LLM仿真研究，而是人类实验：366名被试分为Human alone和Human+AI两组，完成40道推理题（矩阵、心理旋转、三段论、字母串），AI为GPT-5.6 Luna，并单独对AI模型进行基准测试。被试在答题前后进行全局和分块表现估计，并对每题答案给出信心评分。","baseline":"Human alone组作为人类基准，同时AI模型单独在相同题目上的表现作为AI基准。","findings":"Human+AI组得分更高，但自我估计与实际表现的相关性弱，平均高估程度与Human alone组相似；信心区分正确与错误答案的能力在Human+AI组更差。邓宁-克鲁格模式在两组均存在，Human+AI组观察到的对比更大，但控制分数噪声后组间差异不明确。","reliability":"论文未明确讨论失效条件，但指出研究为探索性观察研究，未预注册；控制分数噪声后邓宁-克鲁格效应的组间差异不明确，且任务内信心区分度差异不确定。","relevance":"该研究直接涉及人类与AI协作中的元认知偏差，有真实人类数据对照，结论可迁移到LLM仿真中关于自我评估和信心校准的建模，值得阅读原文以了解具体测量和稳健性检验方法。","inspiration":"借鉴其同时测量全局估计、分块估计和逐题信心，并设置Human alone和Human+AI对照以及AI单独基准，以分离表现提升与元认知变化。｜可迁移到金融决策场景，如投资者使用AI辅助进行资产配置或信贷审批中，评估AI辅助是否改善决策质量但未改善决策者的信心校准。｜设计实验：招募投资者作为被试，随机分为仅人类组和人类+AI组，完成一系列投资决策任务，AI为LLM提供建议，结果变量为投资组合收益和投资者对自身表现的估计及信心评分，对照真实市场数据或历史基准。"}},{"id":"2609.31046","version":1,"title":"Modeling Student Sensemaking with LLMs and Knowledge-Graph-Guided Inference","zh_title":"用大语言模型和知识图谱引导推理建模学生意义建构","abstract":"Collaborative science learning requires nuanced interpretation of student dialogue to characterize how learners identify knowledge gaps, build explanations, and work toward resolution - a theory-driven analysis that is labor-intensive and difficult to scale. We investigate whether instruction-tuned large language models (LLMs) can support multidimensional analysis of collaborative sensemaking without task-specific training, and whether structured knowledge-state information improves model inference. We evaluate two mid-size LLMs on 23 richly annotated, expert-labeled episodes across prompting conditions that vary definitional scaffolding, reasoning mode, and turn structure. Without reasoning, models tend to overpredict successful sensemaking; reasoning-enabled prompting improves identification of unsuccessful cases. Knowledge-state diagnostics provide additional grounding, improving detection of unsuccessful sensemaking and increasing agreement with expert annotations. No single configuration performs best across all sensemaking dimensions, underscoring the multidimensional nature of the task.","authors":["\\\"Ozge Alacam","Z\\\"ubeyde Demet Kirbulut G\\\"une\\c{s}","Funda Ekici","Nurcan Turan-Oluk","Dilay Din\\c{c}demir","Hakk{\\i} Kaday{\\i}f\\c{c}{\\i}","Sevin\\c{c} Nihal Ye\\c{s}ilo\\u{g}lu","Burcu I\\c{s}{\\i}k","Halil T\\\"umay","Sinem Gencer"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-28","first_seen":"2026-09-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.31046","pdf_url":"https://arxiv.org/pdf/2609.31046","source_feed":"cs.CL","score":6,"bucket":"other","rubric_hits":["D1"],"tags":["LLM标注","教育对话分析","知识图谱"],"reason":"LLM用于分析学生对话，替代人工标注，而非仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-09-28T13:01:55","error":null,"has_summary":false,"summary":null},{"id":"2609.31078","version":1,"title":"OmouAI: Argumentative Human-AI Policy Deliberation with Simulated Personas","zh_title":"OmouAI：基于模拟人格的论证式人机政策协商","abstract":"Debates amongst agents driven by large language models (LLMs) have demonstrated vast potential in various applications, but when these interactions include humans and take place in high-stakes environments, e.g., in public policy deliberations, they are beset with issues such as sycophancy and a lack of faithful explanations. To tackle these issues, we present OmouAI, an interactive and inclusive deliberation system that uses LLMs in combination with computational argumentation, a field which excels in representing and reasoning within debates. OmouAI allows a human user to deliberate policy claims for real-world challenges with simulated personas, e.g., representing stakeholders, domain experts or devil's advocates, towards reducing sycophancy. Each persona generates its own arguments, and the arguments of all parties form a shared argumentation framework. Users can then contest, add and revise arguments, providing crucial human oversight. Then, arguments are evaluated using deterministic argumentative semantics against external goals, such as the UN Sustainable Development Goals, guaranteeing faithful explanations. The advancement or worsening of the goals thus serve as indicators for the policy recommendations.","authors":["Stylianos Loukas Vasileiou","Antonio Rago","William Yeoh","Georgina Curto"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-28","first_seen":"2026-09-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.31078","pdf_url":"https://arxiv.org/pdf/2609.31078","source_feed":"cs.AI","score":6,"bucket":"other","rubric_hits":["D3"],"tags":["LLM仿真","政策协商","计算论证"],"reason":"用LLM模拟persona进行政策辩论，属社会过程模拟，但无真实人类数据对照。","model":"deepseek-v4-pro","scored_at":"2026-09-28T13:01:47","error":null,"has_summary":false,"summary":null},{"id":"2609.31607","version":1,"title":"Statistical attribute alignment for black-box generative AI via output post-processing","zh_title":"通过输出后处理实现黑盒生成式AI的统计属性对齐","abstract":"Generative AI systems are increasingly used, but aligning their outputs with user requirements poses a continuing challenge. Here, we aim to ensure that the distribution of an attribute of an AI-generated output aligns with a user-specified target. This is motivated by examples such as fairness, where we want to ensure that a protected attribute (e.g., gender, race, or age categories) follows a desired distribution, and synthetic data generation, where we want the generated data to be representative of a target distribution. We study the practically important black-box access setting, where a user can repeatedly query a generative AI model. The goal is to return $m\\ge 1$ outputs whose joint attribute distribution is as close as possible to this target. For both exact and approximate alignment, we develop algorithms that minimize the expected number of queries to the generator, and we further demonstrate their optimality as the number of requested outputs $m \\rightarrow \\infty$. Experiments on text-to-image generation and geocoded persona generation tasks show that our post-processing algorithms improve statistical attribute alignment, complementing prompting-based interventions.","authors":["Kevin Jiang","Morgane Austern","Edgar Dobriban","Jason M. Klusowski"],"categories":["stat.ME","cs.AI","cs.LG","math.ST","stat.TH"],"primary_category":"stat.ME","announce_type":"cross","date":"2026-09-28","first_seen":"2026-09-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.31607","pdf_url":"https://arxiv.org/pdf/2609.31607","source_feed":"cs.AI","score":6,"bucket":"other","rubric_hits":["D1"],"tags":["生成式AI","属性对齐","合成数据"],"reason":"涉及生成合成数据对齐目标分布，但非仿真人类被试，而是后处理调整属性分布。","model":"deepseek-v4-pro","scored_at":"2026-09-28T13:02:00","error":null,"has_summary":false,"summary":null},{"id":"2609.30716","version":1,"title":"Words Speak Louder Than Order: A Behavioral Evaluation of Gemma 4","zh_title":"言语胜于顺序：Gemma 4 的行为评估","abstract":"When a language model receives two conflicting documents as input, how does it decide which one to prioritize? Does it rely on how the sources are framed or the presentation order of the documents? We evaluated this behavior on Google's pre-trained Gemma 4-e4b model across a targeted behavioral suite (n = 13 items, 784 forward passes in short, single-turn contexts) using a completely counterbalanced experimental design. This setup allowed us to mathematically isolate the specific effects of source framing and reading position, while ensuring the model's natural vocabulary biases were canceled out. Across ten test conditions, we discovered the following: 1. Source framing heavily overpowers reading position. When directly competing, the semantic framing of a source (such as presenting it as an official guideline or a fresh update) had a significantly stronger impact on the model's final answer than the presentation order of the document. 2. The model favors the first document it reads, but this bias is highly variable. While the model consistently demonstrated a primacy effect (preferring the first document presented), the actual strength of this bias fluctuated by at least a factor of 5 based solely on the surface wording. 3. Overall structural repetition, not short copy-cues, drives positional bias. The model's preference for the first document is not a mechanical reaction to short, repetitive trigger phrases, such as \"is [Answer]\". However, the primacy effect does increase significantly when the two competing documents are structurally identical, using word-for-word verbatim templates. Introducing variation in the overall wording between the two sources reduces this positional bias.","authors":["Amanda Fitch"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-28","first_seen":"2026-09-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.30716","pdf_url":"https://arxiv.org/pdf/2609.30716","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["LLM行为评估","来源框架","顺序效应"],"reason":"评估模型行为而非仿真人类被试，无人类数据对照","model":"deepseek-v4-pro","scored_at":"2026-09-28T13:01:44","error":null,"has_summary":false,"summary":null},{"id":"2609.30986","version":1,"title":"Evaluating Sycophancy in Chinese Large Language Models on Factual Questions Derived from Online Search Queries","zh_title":"评估中文大语言模型在基于在线搜索查询的事实问题上的谄媚行为","abstract":"As large language models increasingly mediate information access, factually accurate and independent answers are critical. However, these models can exhibit sycophancy by aligning their responses with users' stated beliefs even when those beliefs are incorrect, potentially presenting misinformation as independently verified and reinforcing users' confidence in false claims. Prior work leaves unresolved whether introducing user beliefs causes correct responses to become incorrect or uncertain, or causes uncertain responses to become belief-aligned incorrect answers. It also remains unclear whether anti-sycophancy interventions preserve or restore factual accuracy or merely shift responses toward uncertainty. We analyze factual sycophancy in Chinese-language information seeking using yes/no fact-checking questions. Our analysis covers 364,941 responses from three frontier Chinese-based LLMs (DeepSeek, Qwen, and Doubao) to 12,165 factual questions derived from real-world Chinese search queries. We evaluate the models with and without reasoning across baseline, belief-conditioned, and anti-sycophancy prompting, tracing matched shifts among correct, incorrect, and uncertain responses. Under incorrect user beliefs, we distinguish belief-aligned errors from losses of factual confidence, in which initially correct answers become uncertain. Patterns vary across models and reasoning settings: reasoning is not a consistent safeguard, and anti-sycophancy instructions can reduce incorrect agreement while increasing uncertainty. In Chinese-language factual question answering, avoiding agreement with false beliefs is therefore not equivalent to preserving factual accuracy, highlighting the value of transition-level evaluation. Such behavior may undermine the reliability of LLM-mediated information access by reinforcing misinformation or weakening users' confidence in factually correct answers.","authors":["Geng Liu","Feng Li","Mengxiao Zhu","Francesco Pierri"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-28","first_seen":"2026-09-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.30986","pdf_url":"https://arxiv.org/pdf/2609.30986","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["LLM谄媚","事实准确性","中文模型"],"reason":"研究LLM在事实问答中的谄媚行为，测量模型本身而非仿真人类被试，无人类对照。","model":"deepseek-v4-pro","scored_at":"2026-09-28T13:01:55","error":null,"has_summary":false,"summary":null},{"id":"2609.31506","version":1,"title":"Evaluating Cultural Awareness of LLMs for Haitian Creole","zh_title":"评估LLM对海地克里奥尔语的文化意识","abstract":"Large language models (LLMs) exhibit substantial performance disparities between high- and low-resource languages. Beyond lower task performance, they often fail to capture the cultural norms and values of underrepresented communities. In this work, we present the first systematic evaluation of cultural awareness in LLMs for Haitian Creole, a language spoken by millions but severely underrepresented in digital resources. We assess cultural awareness along four complementary dimensions---specificity, bias, diversity, and variation---using a benchmark of culturally salient prompts curated by native speakers in a text infilling setting. Our results reveal a clear gap between cultural awareness in Haitian Creole and higher-resource French, with Haitian performance being more uneven across domains and more affected by French linguistic interference. Story generation further reveals recurring portrayals of Haitian characters through hardship and resilience, showing that even positive characterizations can encode stereotypical narratives. Our code, benchmark, and evaluation framework are publicly available.","authors":["Christelle Clervilsson","Yanzhu Guo"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-28","first_seen":"2026-09-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.31506","pdf_url":"https://arxiv.org/pdf/2609.31506","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["文化意识","低资源语言","模型评估"],"reason":"评估LLM的文化意识，测量模型而非仿真人类被试，但涉及文化偏差，与仿真可靠性相…","model":"deepseek-v4-pro","scored_at":"2026-09-28T13:01:59","error":null,"has_summary":false,"summary":null},{"id":"2609.31603","version":1,"title":"User Model Extraction via Belief Self-Distillation","zh_title":"通过信念自蒸馏提取用户模型","abstract":"Large language models (LLMs) implicitly infer attributes of their users and adapt their behavior accordingly, yet these beliefs remain difficult to inspect and causally manipulate. We introduce Belief Self-Distillation (BSD), a unified read-write framework that bridges linear and causal probing by learning a compact user representation that can be both decoded and written back into the model. The frozen LLM acts as its own teacher, distilling beliefs from natural conversations without external annotations. Unlike conventional probing, BSD isolates not only information present in activations, but a state whose causal role can be directly tested. Across multiple model families, BSD faithfully recovers user beliefs and enables substantially stronger interventions than matched hidden-state steering. Crucially, we find that refusal depends not only on the request, but on the model's inferred user intent: changing this belief alters refusal while holding the request fixed. We further uncover a striking cross-model regularity: independently trained LLMs converge on a shared geometry for representing their users. Together, these results reveal implicit user models as readable and causally writable internal states with direct implications for AI safety, shaping how models condition safety decisions on whom they believe they are interacting with.","authors":["Ali Holmov","Yiran Huang","Kirill Bykov","Zeynep Akata"],"categories":["cs.LG","cs.CL"],"primary_category":"cs.LG","announce_type":"cross","date":"2026-09-28","first_seen":"2026-09-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.31603","pdf_url":"https://arxiv.org/pdf/2609.31603","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["LLM内部表征","用户建模","可解释性"],"reason":"研究LLM对用户的隐式建模，属于模型内部状态测量，非仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-09-28T13:02:00","error":null,"has_summary":false,"summary":null},{"id":"2609.31260","version":1,"title":"Agentic Limit Order Books: Phase Transitions and Market Impact","zh_title":"智能体限价订单簿：相变与市场冲击","abstract":"We investigate the systemic macroscopic dynamics emerging from Limit Order Books (LOBs) populated exclusively by autonomous reinforcement-learning agentic traders. By formalizing agent interactions within a microscopic order-matching engine, we examine two fundamental quantitative phenomena: equilibrium phase transitions in order flow regime shifts, and the structural dynamics of market impact. We show that agentic LOBs exhibit distinct phase boundaries separating orderly price discovery from hyper-volatile cascade states, governed by critical thresholds in the number of agents and observable market depth. Furthermore, we demonstrate that market impact under agentic liquidity provision deviates from classical square-root dynamics, exhibiting distinct dissipative, balanced, and non-dissipative regimes under non-linear feedback loops.","authors":["Jan Rosenzweig"],"categories":["q-fin.TR","cs.AI","q-fin.CP"],"primary_category":"q-fin.TR","announce_type":"cross","date":"2026-09-28","first_seen":"2026-09-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.31260","pdf_url":"https://arxiv.org/pdf/2609.31260","source_feed":"cs.AI","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["多智能体市场模拟","限价订单簿","强化学习"],"reason":"用RL agent模拟限价订单簿市场动态，属社会/经济过程模拟，但无真实人类数…","model":"deepseek-v4-pro","scored_at":"2026-09-28T13:01:58","error":null,"has_summary":false,"summary":null},{"id":"2608.29420","version":2,"title":"One Capability or Many? Structural and Predictive Tests of Benchmark Validity Disagree About Economic Benchmarks for Frontier AI","zh_title":"一种能力还是多种？基准效度的结构与预测检验对前沿AI经济基准意见不一","abstract":"Frontier-model leaderboards now rank systems on economic benchmarks, and those rankings inform what organisations buy and what regulators scrutinise. Whether such benchmarks measure a capability distinct from general test-taking is a question of construct validity that a structural test and a predictive test can answer in opposite ways. We show that they do on a hash-pinned snapshot of a frontier leaderboard with 421 model configurations across twelve benchmarks, four of them economic, of which 103 configurations carry all three sparsely scored economic benchmarks and 96 carry all twelve; four hypotheses and their thresholds were fixed before analysis, and every deviation from the plan is reported. The first factor of a three-factor extraction carries 74.5% of common variance and tracks release date (R^2 = 0.505), and date adjustment lowers its share by 14.9 points. Under the dimensionality rule fixed in advance the economic benchmarks form no factor of their own. A leave-one-benchmark-out test with factors re-estimated inside every fold nevertheless finds that a multi-factor representation predicts held-out economic scores better than a single general index (pooled Delta-MSE 0.037, 95% bootstrap interval [0.019, 0.055]) under the linear learners that fit best, an advantage that reverses for tree learners. Under the linear learners the same representation also predicts the eight other benchmarks better, so the battery carries predictive structure that one index misses and the economic benchmarks share it without forming a distinct factor. Construct validity should therefore be assessed by predictive tests alongside structural ones. We give a two-test protocol for benchmark builders and release the pinned data, the analysis plan and the code.","authors":["Louis Yiven Zhu"],"categories":["cs.LG","cs.CY","cs.SE"],"primary_category":"cs.LG","announce_type":"replace-cross","date":"2026-09-28","first_seen":"2026-09-01","revised_at":"2026-09-28","abs_url":"https://arxiv.org/abs/2608.29420","pdf_url":"https://arxiv.org/pdf/2608.29420","source_feed":"cs.CY","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["基准评估","构念效度","AI评测"],"reason":"论文评估经济基准的构念效度，不涉及用LLM仿真人类被试或与人类数据对照。","model":"deepseek-v4-pro","scored_at":"2026-09-28T13:01:49","error":null,"has_summary":false,"summary":null},{"id":"2609.28504","version":2,"title":"AI in Science: Early Insights","zh_title":"AI在科学中的早期洞察","abstract":"Scientific progress is a key driver of economic growth and prosperity. There is great excitement - but also concerns - about the impacts of AI on science, but so far little data. We provide early insights on this from three data sources: a sample of 15 million Gemini interactions, an inventory of over 2,600 specialized AI models across disciplines, and a survey of over 600 scientists. We map these data to a new taxonomy of scientific tasks to study how scientists are using AI. Four main findings emerge. First, we find broad adoption and coverage: scientists use AI more than most other occupations. Specialized AI models have broad disciplinary coverage and are highly cited. Nearly half of the scientists surveyed report using some form of AI every day. Second, we document evidence that LLMs (proxied through Gemini usage) and specialized models act as complements-- LLMs are used for general analysis, coding, and manuscript preparation, while specialized models provide domain-specific predictions, data generation and classification. Third, scientists report large productivity gains from using AI: a saving of nearly 7 hours per week, time which is primarily re-invested in more research. Finally, we show that AI is already changing the scientific process. As some stages of scientific research become easier, bottlenecks shift downstream. Scientists report an increased backlog of untested hypotheses and substantial demand for output verification. Our findings suggest that AI holds significant potential to increase scientific productivity. However, as with other sectors, its ultimate impact will be governed by complex task interdependencies and investment into the elimination of emerging bottlenecks.","authors":["Mihai Codreanu","Alex Imas","Juan Mateos-Garcia","Joseph Emmens","Evalyne Muiruri","Arthur Turrell","Julian Jacobs","Atoosa Kasirzadeh","Ana Trisovic","Yiyuan Chen","Tanya Rodchenko","Catherine Pollard","Scott Strand","Daniel Rock","Zanna Iscenko","Fabien Curto Millet","Neil Thompson","James Manyika"],"categories":["econ.GN","cs.AI","q-fin.EC"],"primary_category":"econ.GN","announce_type":"replace-cross","date":"2026-09-28","first_seen":"2026-09-25","revised_at":"2026-09-28","abs_url":"https://arxiv.org/abs/2609.28504","pdf_url":"https://arxiv.org/pdf/2609.28504","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C5"],"tags":["AI与科学","生产力","科学过程"],"reason":"研究AI对科学的影响，非用LLM仿真人类被试，无人类行为对照","model":"deepseek-v4-pro","scored_at":"2026-09-28T13:02:02","error":null,"has_summary":false,"summary":null},{"id":"2609.30492","version":1,"title":"Breaking Homogeneity: Diversifying Persona Sets for Creative LLM Outputs","zh_title":"打破同质性：为创造性LLM输出多样化角色集","abstract":"Language models often produce homogeneous responses to open-ended tasks; such homogeneity can spawn groupthink-the convergence of ideas toward a singular and potentially suboptimal decision. We formulate persona diversification as a set-level conditioning problem and study two orthogonal design choices: selecting versus generating personas, and space-filling versus frontier-seeking diversity. We instantiate this design space with four methods spanning coverage and dispersion subset selections, uniform-coverage sampling, and evolutionary persona generation. Evaluations on the Alternative Uses Task (AUT), Infinity-Chat, and Divergent Association Task (DAT) show the benefits of the proposed methods across tasks and creativity objectives. On AUT, evolutionary persona generation increases response diversity by 78.8%, originality by 26.1%, flexibility by 49.5%, and holistic creativity by 13.9% over task-only prompting, while maintaining 98.5% validity; on Infinity-Chat, it nearly doubles persona-induced response separation relative to random personas. Moreover, evolutionary personas compose with creativity-optimized prompting, further increasing its response diversity by 18.6% and creativity by 6.3%. These results establish persona-set geometry as a task-agnostic mechanism for eliciting divergent LLM outputs, and support persona diversification as a reusable complement to prompt optimization.","authors":["Sang Bin Moon","Nicole Cho","Daniel Borrajo","Sumitra Ganesh","Abolfazl Hashemi"],"categories":["cs.CL","cs.AI","stat.ML"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-28","first_seen":"2026-09-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.30492","pdf_url":"https://arxiv.org/pdf/2609.30492","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["角色多样化","创造力生成","提示工程"],"reason":"研究通过多样化角色设定提升LLM创造力输出，属于角色扮演与生成多样性，无人类行…","model":"deepseek-v4-pro","scored_at":"2026-09-28T13:01:51","error":null,"has_summary":false,"summary":null},{"id":"2609.30558","version":1,"title":"Probing Stability-Plasticity Tradeoffs in Agent Memory through Cognitive Experimental Paradigms","zh_title":"通过认知实验范式探究智能体记忆中的稳定性-可塑性权衡","abstract":"Agent memory systems are increasingly used to maintain long-term user preferences, task states and evolving facts, but current evaluations often collapse memory behavior into final-answer accuracy. We introduce MemProbe, a cognitive-science-inspired framework for diagnosing stability-plasticity tradeoffs in agent memory. The framework is motivated by a core insight from cognitive memory research: memory is reconstructive and shaped by interference, source reliability, reinforcement, and reactivation. MemProbe turns this insight into four reusable experimental paradigms (interference, misinformation, consolidation strength, and reconsolidation window) that manipulate when a memory should be updated, preserved, or treated as uncertain. It further decomposes correctness into behavioral profiles that reveal how systems update, preserve, attribute, and temporally organize information. We instantiate these paradigms in a 56-episode diagnostic suite and evaluate six incremental memory systems under a unified protocol. Results show that systems with similar aggregate scores exhibit distinct behavioral profiles. MemProbe provides such a diagnostic lens, turning aggregate performance into interpretable profiles of memory maintenance over time. Code is available at https://github.com/jq-ding/MemProbe.","authors":["Jiaqi Ding","Guorong Wu"],"categories":["cs.CL","cs.AI","cs.LG"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-28","first_seen":"2026-09-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.30558","pdf_url":"https://arxiv.org/pdf/2609.30558","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["智能体记忆","认知诊断","多智能体系统"],"reason":"评估智能体记忆系统，不涉及人类行为仿真或对照，属多智能体系统研究。","model":"deepseek-v4-pro","scored_at":"2026-09-28T13:01:51","error":null,"has_summary":false,"summary":null},{"id":"2609.30897","version":1,"title":"From annotation to reasoning: Culture in language models","zh_title":"从标注到推理：语言模型中的文化","abstract":"How should we evaluate language models when more than one interpretation can be right? Cultural benchmarks often test factual knowledge, agreement with survey responses, or recognition of a predefined meaning. These tasks leave open whether a model can explain how a cultural reference works in a particular text, support a reading with evidence, or revise it after criticism. This is a question of interpretive depth, complementary to the breadth of cultural coverage. We argue that literary interpretation offers a useful setting for studying these capabilities. We focus on cultural referencing and reuse: how texts invoke, repeat, and transform earlier expressions across historical and linguistic contexts. Our central claim is that literary scholars can disagree about an interpretation while recognizing the quality of its support. We propose linking evidence-centered benchmarks, evaluation that preserves scholarly disagreement, and model-development experiments on literary data, contextual resources, and scholarly feedback. Danish literature provides a concrete starting point, with implications for other languages and domains. The aim is to develop alternative evaluation strategies that go beyond conventional benchmark metrics and guide model development toward cultural robustness in AI systems.","authors":["Daniel Hershcovich","Alexander Conroy","Jens Bjerring-Hansen"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-28","first_seen":"2026-09-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.30897","pdf_url":"https://arxiv.org/pdf/2609.30897","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["文化基准","文学解释","模型评测"],"reason":"论文关注文化基准评测与文学解释，不涉及用LLM仿真人类被试或与真实人类数据对照。","model":"deepseek-v4-pro","scored_at":"2026-09-28T13:01:53","error":null,"has_summary":false,"summary":null},{"id":"2609.30768","version":1,"title":"Does Thinking Help Fairness? Reasoning Tokens Resolve Some Biases but Create More","zh_title":"思考有助于公平吗？推理令牌解决一些偏差但创造更多","abstract":"Thinking in reasoning language models (RLMs) has been subject to debate on whether it resolves or amplifies bias. Prior works have shown competing conclusions in both directions. Using a within-model thinking-vs.-non-thinking ablation across QwQ-32B, DeepSeek-R1-Distill-Qwen-32B, and Qwen3-32B on three high-stakes decision tasks (Adult, COMPAS, Credit), we show that thinking has an asymmetric dual effect on counterfactual fairness: it both resolves counterfactual flips produced by the non-thinking baseline and creates new flips at near-saturating model confidence. In all nine (model, dataset) combinations, the created flips outnumber the resolved flips by roughly 5 times. To explain the effect, we treat the thinking trace itself as a measurable site of fairness change and study it through two dynamic instruments: 1) We propose Counterfactual Depth Probability Gap (CDPG) to track bias evolution along thinking depth, and observe that bias propagates and amplifies with thinking. 2) We also formulate the Bias Transition Matrix (BTM) to show how predictions of counterfactual pairs change from non-thinking to thinking, and find that the asymmetric dual effect originates in the pair-state joint transition.","authors":["Deng Pan","Joe Germino","Yihong Ma","Elizabeth Daly","Nuno Moniz","Ting Hua","Nitesh Chawla"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-28","first_seen":"2026-09-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.30768","pdf_url":"https://arxiv.org/pdf/2609.30768","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["公平性","推理模型","偏差"],"reason":"研究LLM推理对公平性影响，属模型偏差评测，非仿真人类被试","model":"deepseek-v4-pro","scored_at":"2026-09-28T13:01:44","error":null,"has_summary":false,"summary":null},{"id":"2609.30939","version":1,"title":"MACBT: A Multi-Agent Cognitive Behavioral Therapy Decision Support System with Longitudinal Memory","zh_title":"MACBT：具有纵向记忆的多智能体认知行为疗法决策支持系统","abstract":"Cognitive behavioral therapy (CBT) is an evidence-based first-line treatment for depression, yet its scale is constrained by the time clinicians spend on pre-session preparation, post-session documentation, and longitudinal cognitive-pathology tracking. We present a clinician-facing AI decision-support system that combines a multi-agent CBT framework (MACBT) with a CBT-specific longitudinal memory module (CD Memory). MACBT encodes the five-stage CBT workflow (assessment, Socratic questioning, cognitive restructuring, behavioral experiments, and treatment monitoring) into five collaborative agents. CD Memory tracks cognitive-distortion type, frequency, severity, and restructuring efficacy across sessions to generate pre-session pathology reports and intervention-priority recommendations. We construct a Chinese CBT dialogue corpus via dual-role large language model simulation and train a Qwen3-14B backbone with supervised fine-tuning and direct preference optimization. Evaluation with GPT-4 judges shows MACBT outperforms MeChat, SoulChat, PsyChat, and CPsyCounX in professionalism (2.62) and clinical authenticity (2.25). The full memory-augmented system further improves session quality by 12.6% and achieves a longitudinal mean of 2.29 on cross-session continuity, intervention progression, and personalization.","authors":["De Jiang","Shuo Zhang","Weiwei Liao","Jianying Zhang","Chuanhui Yu","Hongen Liao","Kehong Yuan"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-28","first_seen":"2026-09-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.30939","pdf_url":"https://arxiv.org/pdf/2609.30939","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C1","C3"],"tags":["多智能体系统","认知行为疗法","决策支持"],"reason":"多智能体CBT决策支持系统，无人类行为对照，属角色扮演与协作任务，非仿真被试。","model":"deepseek-v4-pro","scored_at":"2026-09-28T13:01:54","error":null,"has_summary":false,"summary":null},{"id":"2609.31184","version":1,"title":"Accounting for Bias Enables Sustainable LLM Evaluation","zh_title":"考虑偏差实现可持续的LLM评估","abstract":"LLM-as-a-judge has become the de facto standard for scalable, subjective evaluation, yet current leaderboards compensate for systematic measurement bias by running ever more comparisons, an approach that is both statistically unsound and computationally wasteful. The root cause is an incomplete measurement model, treating LLM judges as neutral, interchangeable instruments ignores documented biases like position bias, verbosity bias, judge severity, and self-enhancement, that no volume of additional data can eliminate. We propose a unified latent variable framework that jointly models pairwise and ordinal data while explicitly correcting for these confounders, recovering reliable rankings from substantially fewer comparisons. Because fitting this model costs negligible compute relative to a single round of LLM inference, bias correction is not only more statistically rigorous but also a more sustainable approach to trustworthy evaluation.","authors":["Harshita Katoch","David Antony Selby","Gerrit Gro{\\ss}mann","Sebastian Vollmer"],"categories":["cs.AI","cs.LG","stat.AP","stat.ME"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-28","first_seen":"2026-09-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.31184","pdf_url":"https://arxiv.org/pdf/2609.31184","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["LLM评估","偏差校正","测量模型"],"reason":"论文研究LLM作为评判者的评估偏差校正，属于纯NLP评测方法，不涉及人类行为仿…","model":"deepseek-v4-pro","scored_at":"2026-09-28T13:01:56","error":null,"has_summary":false,"summary":null},{"id":"2609.31215","version":1,"title":"DIAL: Position-Debiased LLM Judges with Adaptive Human Preference Calibration","zh_title":"DIAL：基于自适应人类偏好校准的位置去偏LLM评判器","abstract":"Large language models (LLMs) as a judge enable scalable evaluation, but their judgments can be sensitive to response order and, even after removing such position effects, can still diverge systematically from human preferences.We introduce DIAL, a unified framework that combines abundant LLM comparisons with limited human comparisons to separate judge-specific position effects, learn shared structure in position-debiased LLM preferences, and adaptively calibrate that structure toward the human preference target. Theoretically, we study three aspects of DIAL: (i) identification of latent LLM preferences, position effects, and human calibration; (ii) adaptive estimation that balances LLM anchoring against limited human evidence; and (iii) fixed-weight uncertainty quantification for the calibrated human preference. Empirically, we evaluate position debiasing and human alignment separately in controlled simulations and on three human-preference benchmarks, showing that DIAL remains robust to unbalanced response order, achieves strong human-aligned rankings with limited labels, and adapts toward human evidence when LLM information is imperfect. Our real-data study collects over 410K judgments from 21 LLM judges in both display orders, providing a resource for future studies of LLM-judge bias, heterogeneity, and human alignment.","authors":["Zesheng Cai","Yingqi Fan","Sichang Chen","Jin-Hong Du"],"categories":["cs.AI","stat.AP","stat.ME","stat.ML"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-28","first_seen":"2026-09-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.31215","pdf_url":"https://arxiv.org/pdf/2609.31215","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["LLM评判","偏好校准","评估方法"],"reason":"论文研究LLM作为评判者的位置偏差与人类偏好校准，属于评估方法改进，非用LLM…","model":"deepseek-v4-pro","scored_at":"2026-09-28T13:01:57","error":null,"has_summary":false,"summary":null},{"id":"2609.31473","version":1,"title":"Game Arena: Strategic LLM Evaluation in Competitive Environments","zh_title":"游戏竞技场：竞争环境中的大语言模型战略评估","abstract":"We introduce Kaggle Game Arena, an open and ever-expanding platform to evaluate large language models (LLMs) through competitive games. Different from static benchmarks, game arena enables models to play head-to-head matchups in structured environments where the gameplay strength naturally increases as models evolve, preventing performance saturation. This technical report details the infrastructure behind Game Arena and describes the three pilot game environments: Chess, Poker, and Werewolf. These environments span perfect information, imperfect information, and multiplayer game settings, enabling a systematic study of models' strategic planning, adaptation, and robustness under uncertainty. For each game, we provide a detailed description of the environment, evaluation metrics, and results from running full competitions across models. Through robust infrastructure and large-scale ground-truth based evaluation, Game Arena ensures reproducibility, transparency and generalizability to new games and variants over time.","authors":["Bovard Doerschuk-Tiberi","Yao Yan","Justin Chiu","Hann Wang","Timothy Chung","Martyna Plomecka","John Schultz","Jon Lipovetz","Clayton Drazner","Yuchen Zhuang","Jaimie Hwang","Nate Keating","Riley Jones","Andrew Lee","Oran Kelly","Ian Gemp","Michael Aaron","Laurel Prince","Kate Larson","Jeff Moser","Harrison Jobe","Chad Woodford","Siqi Liu","Andrew Wang","Bo Chang","Christopher D'Mello","Diane Chaleff","Addison Howard","Johnny Yip","Chuck Sugnet","Antonio Gulli","Meghan O'Connell","Will Cukierski","Nenad Tomasev","Dima Yeroshenko","Kinjal Parekh","Roxanne Daniel","Marc Lanctot","Domino Weir","Elsa Dong","Daniel Hennes","Melissa Nalubwama","Robert Fraser","Ryan Trostle","Jun Peng","Tom Mason","Lloyd Hightower","Chiamaka Chukwuka","Yuexiang Zhai","Phoebe Kirk","Yi Su","Yuting Han","Jie Ren","Chris Prichard","Sahand Sharifzadeh","Karim Hakimzadeh","DJ Sterling","Meg Risdal","Kate Olszewska","Ya Xu","Orhan Firat","Minmin Chen"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-28","first_seen":"2026-09-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.31473","pdf_url":"https://arxiv.org/pdf/2609.31473","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C2"],"tags":["LLM评估","游戏AI","多智能体"],"reason":"评估LLM在游戏中的策略能力，属游戏仿真环境，不涉及人类行为对照或仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-09-28T13:01:59","error":null,"has_summary":false,"summary":null},{"id":"2609.31563","version":1,"title":"Multi-agent Scaling Across Disjunctive and Compensatory Tasks","zh_title":"分离型与补偿型任务中的多智能体扩展","abstract":"Multi-agent LLM systems are often expected to improve as team size increases, yet the scaling behavior may depend on task structure. Our central contribution is to introduce Steiner's taxonomy of group tasks as a framework for analyzing multi-agent LLM scaling and focusing the analysis on disjunctive and compensatory tasks. We model independently sampled agents as conditionally independent given the item, which yields their large-team limits: plurality voting converges to the model's modal answer, and averaging converges to the model's item-level bias. Across selected representative benchmarks, 13 open-weight models, and teams of up to 30 agents, we find qualitatively different scaling behavior. On disjunctive tasks, the probability that at least one agent is correct grows by 5-20 points with team size, but plurality voting over agents that answer directly realises almost none of this potential, as the model predicts to within 0.5 points on average. Multi-round revision raises accuracy considerably, yet the gain is nearly the same with one peer as with 29. In contrast, scaling provides little benefit on Fermi estimation, despite its natural suitability for aggregation: item-level biases shared across the samples of a model account for about 87% of the squared error, so averaging reduces error by only about 6%. Combining model families helps on Fermi estimation but does not surpass the strongest member on disjunctive tasks. These results show that task structure, together with the mechanism combining member outputs, is a fundamental determinant of team scaling.","authors":["Carolina Fortuna","Blaz Bertalanic"],"categories":["cs.AI","cs.MA"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-28","first_seen":"2026-09-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.31563","pdf_url":"https://arxiv.org/pdf/2609.31563","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","任务扩展","LLM协作"],"reason":"纯多智能体协作解题，无人类行为对照，不涉及人类仿真","model":"deepseek-v4-pro","scored_at":"2026-09-28T13:01:59","error":null,"has_summary":false,"summary":null},{"id":"2609.30466","version":1,"title":"A Benchmarking Framework for Context-aware XR Interfaces","zh_title":"面向情境感知XR界面的基准测试框架","abstract":"Everyday Extended Reality (XR) systems aim to provide context-aware access to the right functionalities at the right time and place, with minimal manual reconfiguration as users switch context. Yet these interfaces are hard to evaluate: current prototyping and user-study workflows offer no systematic, repeatable way to compare adaptation methods across users and scenarios. We present ContextXR, a novel benchmarking framework for context-aware XR interfaces. ContextXR represents an XR application as a connected graph of functional facets, each a semantically coherent group of related capabilities that together support a shared user intent. On this representation, we build MineXR++, a dataset augmenting prior XR interface data with facet-level annotations, and formulate three canonical tasks of context-aware suggestion: context factor analysis, initial facet suggestion, and next facet suggestion. Our evaluation protocol scores suggestion methods by a simulated interaction metric, the navigation and search cost of reaching the desired functionality. Through experiments benchmarking global popularity, relational retrieval, and LLM-based methods, we demonstrate that ContextXR enables the systematic, reproducible evaluation of context-aware XR interfaces.","authors":["Hyunsung Cho","Sarah Yewon Yun","Nancy Ruonan Sun","Ben Lafreniere","Mark Parent","Kashyap Todi","Tanya R. Jonker","Hrvoje Benko","Sherry Tongshuang Wu","David Lindlbauer"],"categories":["cs.HC","cs.AI"],"primary_category":"cs.HC","announce_type":"cross","date":"2026-09-28","first_seen":"2026-09-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.30466","pdf_url":"https://arxiv.org/pdf/2609.30466","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C2"],"tags":["XR界面","基准测试","LLM方法"],"reason":"论文聚焦XR界面基准测试，LLM仅作为方法之一，不涉及人类行为仿真或对照。","model":"deepseek-v4-pro","scored_at":"2026-09-28T13:01:51","error":null,"has_summary":false,"summary":null},{"id":"2609.31219","version":1,"title":"Research with AI Agents: How Agentic Systems Are Changing Scientific Work","zh_title":"AI代理研究：代理系统如何改变科学工作","abstract":"Background. Agentic AI systems independently decompose tasks such as literature search, data analysis, and programming into subtasks, search the web, access databases, and execute code. This allows them to perform digital research tasks at high speed. Objectives. Under what conditions does the use of agentic systems produce reliable efficiency gains, and which tasks remain with researchers? Materials and methods. Summary of current studies on literature searches, data analysis, software development, and clinical decision support. Results. For digital activities, work shifts from execution to steering and review. Efficiency gains are greatest when expected behavior can be formalized in advance and tested automatically. In complex agentic systems, recorded sequences of reasoning steps and tool calls can quickly become too extensive for human review. Furthermore, explanations generated by the model do not reliably reflect how an output was produced. One possible step toward more reliable systems is the validation of individual components. The limited reviewability extends beyond research itself; the review of scientific articles and grant proposals is also reaching capacity limits. In pathology, curated and annotated data, researchers' own analytical skills, and institutional exchange of experience are becoming increasingly important. Conclusions. Researchers remain responsible for their results. They must determine what to delegate and how to review the results. The importance of a research question to patients, the field, and society cannot be fully assessed using formalized criteria and remains a matter of expert judgment. Agentic systems can free up time for this.","authors":["Johannes Lotz","Markus Wenzel"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2026-09-28","first_seen":"2026-09-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.31219","pdf_url":"https://arxiv.org/pdf/2609.31219","source_feed":"cs.CY","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["AI代理","科研自动化","人机协作"],"reason":"讨论AI代理在科研任务中的自动化，不涉及用LLM仿真人类被试或与人类行为对照","model":"deepseek-v4-pro","scored_at":"2026-09-28T13:01:57","error":null,"has_summary":false,"summary":null},{"id":"2609.30427","version":1,"title":"Fake News Theories: Harnessing Disciplinary Insights for Computational Modeling, Detection, and Explanation","zh_title":"假新闻理论：利用跨学科洞见进行计算建模、检测与解释","abstract":"Disinformation research has produced increasingly accurate automated fake-news detectors, but many systems remain difficult to interpret and are weakly connected to established theories of persuasion, credibility, and human judgment. In this paper, we develop a theory-informed computational framework that translates cross-disciplinary theories of fake news into measurable features for automated detection and explanation through statistical techniques and large language models. To that end, we conduct a structured cross-disciplinary review of theories from social sciences, psychology, economics, among other disciplines that reveal how fake news persuades and spreads, thereby establishing a broad theoretical foundation for computational modeling. Experiments on benchmark datasets show that theory-derived features are predictive and provide interpretable, theory-referenced diagnostic signals. Multi-feature models generally outperform individual features, although gains among the strongest small feature combinations are modest. Our work highlights the value of interdisciplinary perspectives in building robust and interpretable fake news detection systems, advancing the foundation for human-centered approaches in combating disinformation.","authors":["Zhaoyang Cao","Miriam Metzger","Reza Zafarani"],"categories":["cs.LG","cs.CY"],"primary_category":"cs.LG","announce_type":"cross","date":"2026-09-28","first_seen":"2026-09-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.30427","pdf_url":"https://arxiv.org/pdf/2609.30427","source_feed":"cs.CY","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["假新闻检测","可解释性","跨学科理论"],"reason":"论文聚焦假新闻检测与解释，使用LLM提取特征，但非以人类为参照系仿真被试，属纯…","model":"deepseek-v4-pro","scored_at":"2026-09-28T13:01:50","error":null,"has_summary":false,"summary":null},{"id":"2609.31590","version":1,"title":"AgentWorld: Benchmarking Long-Horizon Collaboration of Multi-agent LLMs","zh_title":"AgentWorld：多智能体大语言模型长程协作基准测试","abstract":"Existing multi-agent benchmarks primarily test in competitive settings, short-horizon interactions under 20 steps, or simply aggregate individual performance, failing to isolate and highlight genuine collaboration capabilities of LLM-based agents. We introduce AgentWorld, a benchmark of 100 human-annotated tasks (with 100 augmented variants) for evaluating long-horizon, multi-agent collaboration. Tasks span 50+ interaction rounds across a rich MMORPG sandbox and require 3-20 agents with asymmetric roles and abilities to coordinate through communication, joint planning, and resource sharing under a blackbox setting where each agent acts independently without access to others' internal states. To quantify collaboration effectiveness in addition to conventional binary task success, we propose Causal Collaboration Effectiveness (CCE), a graph-based metric that traces causal dependencies between agent actions and measures what fraction of a team's effort actually contributed to the outcome. Experiments with Gemini 3 Flash, Claude Haiku 4.5, GPT-5 Mini, and DeepSeek R1-70B show that even the best model achieves only 52.0% task success, with systematic failure modes including communication breakdowns, role confusion, and inability to maintain shared plans across rounds. AgentWorld is fully open-source.","authors":["Raphael Shu","Yusen Zhang","Young Min Cho","Jin Mo Yang","Yuan Yuan","Wenliang Zheng","Sharath Chandra Guntuku","Lyle Ungar","Zhou Yu","Rui Zhang"],"categories":["cs.MA"],"primary_category":"cs.MA","announce_type":"new","date":"2026-09-28","first_seen":"2026-09-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.31590","pdf_url":"https://arxiv.org/pdf/2609.31590","source_feed":"cs.MA","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体协作","基准测试","任务完成"],"reason":"纯多智能体协作解题，无人类行为对照，属排除项","model":"deepseek-v4-pro","scored_at":"2026-09-28T13:02:00","error":null,"has_summary":false,"summary":null},{"id":"2609.30405","version":1,"title":"Adaptive Multi-Value Control in LLMs via Causal Activation Steering","zh_title":"通过因果激活引导实现LLM的自适应多值控制","abstract":"Large language models (LLMs) are increasingly deployed in settings where responses must reflect multiple, potentially interacting social norms and human values. Activation steering offers a lightweight alternative to training-based alignment by modifying internal activations at inference time. However, prior human-value steering methods have largely considered values in isolation, while direct composition of multiple directions relies on fixed intervention strengths that cannot respond to the model's evolving internal state. Motivated by this key observation, we introduce AIMES, a framework for adaptive multi-value activation steering. AIMES constructs layer-specific bipolar directions for moral-foundation values and uses intermediate-layer vocabulary readouts as online observers. An observer-guided controller then adapts the strength of each requested value intervention at every decoding step based on its current observed state, without training a separate value-state estimator. Across multiple instruction-tuned model families, value combinations, and intervention depths, we find that multi-value controllability varies across both value combinations and intervention locations. Compared with fixed joint steering and prompt-based steering, AIMES shows depth-dependent advantages that are broadly supported across two independent evaluators, with some variation in the precise depth at which specific control effects emerge. These advantages come with smaller realized activation-space interventions than fixed-joint steering and comparable response quality. Overall, our results suggest that online observer feedback can provide lightweight, state-aware adaptation for single-pass multi-value steering.","authors":["Payel Bhattacharjee","Ravi Tandon"],"categories":["cs.LG"],"primary_category":"cs.LG","announce_type":"new","date":"2026-09-28","first_seen":"2026-09-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.30405","pdf_url":"https://arxiv.org/pdf/2609.30405","source_feed":"cs.LG","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["激活引导","价值观对齐","模型控制"],"reason":"研究LLM价值观对齐与激活控制，不涉及人类被试仿真或行为对照，属模型对齐技术。","model":"deepseek-v4-pro","scored_at":"2026-09-28T13:01:50","error":null,"has_summary":false,"summary":null},{"id":"2609.06951","version":3,"title":"Steering Interference Reflects the Model's Defaults, Not the Behavior Directions","zh_title":"引导干扰反映模型默认行为而非行为方向","abstract":"Activation steering promises modular control of language model behavior: a behavior such as politeness corresponds to a direction in a model's activations, and adding that direction while it generates should switch the behavior on and leave everything else alone. It does not. We ask what decides which other behaviors move, and by how much, and find that it is the model rather than the behavior being steered. A steer relaxes the model toward a small set of behaviors it already favors, chiefly refusal, sycophancy, and poeticism, and that set is much the same whatever is steered. Three results across 24 behaviors and ten instruction-tuned models support this, every effect read off the generated text by a language-model judge rather than off a probe. That readout matters: all 24 behaviors are linearly decodable, but only 20 change what the model writes. First, a direction carrying no behavioral content, matched to a real steer only in the size of the vector it adds, moves the same behaviors in the same order as real steers do, while producing none of the behaviors that need a specific direction. Second, most interference runs one way, so it cannot be an overlap between two directions: steering profanity makes the model toxic, while steering toxicity leaves profanity untouched. Third, with a behavior held out entirely, geometry measured on the others explains almost none of the interference it takes part in. The account holds on all ten models, the pull toward defaults strongest below 10B parameters and weakening in each family's largest. Reading a steer as a perturbation whose endpoint the model fixes implies that disentangling behavior directions cannot by itself make steering modular.","authors":["Srikanth Malla","Chiho Choi","Joon Hee Choi"],"categories":["cs.LG","cs.AI"],"primary_category":"cs.LG","announce_type":"replace-cross","date":"2026-09-28","first_seen":"2026-09-09","revised_at":"2026-09-28","abs_url":"https://arxiv.org/abs/2609.06951","pdf_url":"https://arxiv.org/pdf/2609.06951","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["激活引导","模型可解释性","行为控制"],"reason":"研究激活引导对模型行为的影响，属模型控制与可解释性，不涉及人类仿真或人类数据对…","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:02:08","error":null,"has_summary":false,"summary":null},{"id":"2609.29390","version":2,"title":"Likelihood Ranking doesn't Scale Like Prompting in LLMs","zh_title":"似然排序在LLM中不像提示那样随规模扩展","abstract":"LLM evaluation is commonly performed either by prompting models to produce answers or by scoring candidate outputs with likelihood-based metrics. In multiple-choice QA, however, standard likelihood-based scoring is still conditioned on the question and answer set, and can therefore leverage the same task-conditioned answer-selection interface used in prompting. We study a complementary protocol based on likelihood ranking of declarative statements constructed from the same question--answer pairs. Across 95 decoder-only models, ranging from 0.1B to 104B parameters, and 10 MCQA datasets, we find a systematic divergence between declarative-statement likelihood ranking and prompted answering. Statement-likelihood accuracy remains comparatively stable across scale, whereas prompted answering improves sharply with scale and instruction-tuning. These results suggest that likelihood preferences over controlled declarative alternatives and task-conditioned answer selection probe distinct aspects of model behavior, and should not be treated as interchangeable.","authors":["Alessandro Bondielli","Lucia Passaro","Davide Bacciu","Alessandro Lenci"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-28","first_seen":"2026-09-25","revised_at":"2026-09-28","abs_url":"https://arxiv.org/abs/2609.29390","pdf_url":"https://arxiv.org/pdf/2609.29390","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["LLM评测","多项选择问答","似然排序"],"reason":"纯NLP评测，比较LLM在MCQA上的两种评分方法，不涉及人类仿真或人类数据对照","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:22","error":null,"has_summary":false,"summary":null},{"id":"2609.30849","version":1,"title":"Enhancing Assessment of Self-Consistency in LLM Explanations using Perturbation Strength","zh_title":"利用扰动强度增强对LLM解释自洽性的评估","abstract":"Prior work has examined the self-consistency of LLM-generated explanations using surface-level perturbation methods. However, the strength of these perturbations is not explicitly measured and controlled. In this work, we propose an LLM-as-a-judge approach to measure perturbation strength in a unified manner across input and CoT perturbations. We then evaluate the self-consistency in explanations generated from various LLMs under controlled strength conditions, ensuring a fair comparison across perturbation types. Experiments show that our proposed LLM-based perturbation strength measure outperforms other embedding- and probability-based approaches and that input perturbations generally affect LLMs more strongly than CoT perturbations. Our work suggests that judgments about a model's self-consistency is fair only within the same perturbation type.","authors":["Phuong Q. Le","Kemal Kurniawan","Jey Han Lau"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-28","first_seen":"2026-09-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.30849","pdf_url":"https://arxiv.org/pdf/2609.30849","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["LLM评估","解释自洽性","扰动方法"],"reason":"研究LLM解释的自洽性，属模型能力评测，不以人类为参照系","model":"deepseek-v4-pro","scored_at":"2026-09-28T13:01:53","error":null,"has_summary":false,"summary":null},{"id":"2609.31166","version":1,"title":"AgentRecommender: LLM Agents Enable Customizable Recommender Systems on the User Side","zh_title":"AgentRecommender：LLM智能体实现用户侧可定制推荐系统","abstract":"Recommender systems have traditionally been developed for platforms. However, this has given rise to many phenomena that may be advantageous for platform lock-in but are a nuisance to users, such as clickbait, filter bubbles, and the spread of fake news. Recently, user-side recommender systems have been proposed as a new paradigm for solving this problem. If users deploy their own recommender systems, they are no longer at the mercy of the platform's interests. However, building a user-side recommender system is not trivial; in particular, customizing one for oneself requires additional data. We propose AgentRecommender, a method that leverages the investigation capability and internal knowledge of LLM agents to flexibly build user-side recommender systems without additional data. AgentRecommender allows users to easily create recommender systems tailored to their own preferences.","authors":["Ryoma Sato"],"categories":["cs.IR","cs.AI","cs.DB","cs.DL"],"primary_category":"cs.IR","announce_type":"cross","date":"2026-09-28","first_seen":"2026-09-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.31166","pdf_url":"https://arxiv.org/pdf/2609.31166","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["推荐系统","LLM智能体","用户侧"],"reason":"LLM agent用于用户侧推荐系统，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-28T13:01:56","error":null,"has_summary":false,"summary":null},{"id":"2609.31272","version":1,"title":"Cognitive Skills in the Age of AI: Computing Students and Experts Perceptions","zh_title":"AI时代的认知技能：计算专业学生与专家的看法","abstract":"AI is becoming increasingly integrated into daily workflows, especially in computing. We are gradually shifting towards an AI-rich future, an impending yet unknown one. One important emerging concern is whether we are accordingly preparing our future computing workforce. Further, we need to know what the important cognitive skills are to remain relevant in the computing workforce and if there are changes in cognitive skill importance. To investigate this direction, we conducted a mixed-methods study, collecting perceptions from computing students and computing experts regarding the importance of cognitive skills in the past, present, and future. We report that the perceived importance of most cognitive skills will decrease in the future, with an AI-rich environment, but critical thinking skills remain important. Further, we report reasons collected through interviews on why the importance of cognitive skills will change and how future computing students can prepare for it.","authors":["Neha Rani","Vu Minh Anh Le","Austin M. Spangler","Erta Cenko"],"categories":["cs.HC","cs.AI"],"primary_category":"cs.HC","announce_type":"cross","date":"2026-09-28","first_seen":"2026-09-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.31272","pdf_url":"https://arxiv.org/pdf/2609.31272","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["认知技能","AI影响","人机交互"],"reason":"研究人类对AI时代认知技能的看法，不涉及用LLM仿真人类被试或行为对照","model":"deepseek-v4-pro","scored_at":"2026-09-28T13:01:58","error":null,"has_summary":false,"summary":null},{"id":"2609.30388","version":1,"title":"The Interviewer's Perspective: Unpacking the Impact of Real-Time AI Interviewing Assistance on Social Dynamics","zh_title":"访谈者视角：实时AI访谈辅助对社会动态的影响","abstract":"Eliciting rich data in semi-structured interviews is cognitively demanding, prompting recent work to explore real-time AI assistance for interviewers. However, introducing AI into the interviewer-interviewee interaction creates a triadic context whose social dynamics remain underexplored. We investigated how interviewers experience AI assistance for probing during semi-structured interviews. To elicit rich participant reflections, we implemented two variants of AI assistance differing in initiation and granularity in a high-fidelity prototype, ProbeAssist. We conducted a qualitative-first comparative structured observation study where 18 participants each completed three simulated interviews: one without AI and two with different AI variants. Findings showed that participants leveraged AI as a supportive tool but resisted it as an assessor or competitor. As they navigated AI's benefits and interaction costs, tensions emerged around agency, ownership, creativity, and interpersonal communication. We propose three implications for AI-assisted human-to-human interaction: managing social pressure, balancing idea alignment with inspiration, and preserving interpersonal presence.","authors":["Zhe Liu","Jiamin Dai","Joanna McGrenere"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-28","first_seen":"2026-09-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.30388","pdf_url":"https://arxiv.org/pdf/2609.30388","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["AI辅助访谈","人机交互","社会动态"],"reason":"研究AI辅助访谈者，非用LLM仿真人类被试，无人类数据对照","model":"deepseek-v4-pro","scored_at":"2026-09-28T13:01:49","error":null,"has_summary":false,"summary":null},{"id":"2609.30588","version":1,"title":"Orchestrating GenAI for Interdisciplinary Research","zh_title":"为跨学科研究编排生成式人工智能","abstract":"As researchers tackle interdisciplinary problems, they face the need to deepen expertise in primary areas while rapidly acquiring knowledge in secondary domains. Generative AI (GenAI) is increasingly positioned to meet this need, from general-purpose chat assistants to Deep Research tools marketed as autonomous research agents. Prior work has examined how researchers use GenAI to support single-discipline or general research tasks. However, we know little about the goals and GenAI practices in interdisciplinary research. We conducted a longitudinal study and semi-structured interviews with 15 interdisciplinary researchers to examine how interdisciplinary researchers actually orchestrate GenAI. Findings show that researchers leaned on GenAI to fill knowledge gaps while maintaining epistemic agency for novelty discovery. We also uncovered an expertise paradox: GenAI outputs were hardest to verify when most needed. Our empirical insights motivate GenAI designs that calibrate verification to researchers' expertise, nudge toward cross-domain synthesis, and adapt prompting and outputs to disciplinary conventions.","authors":["Shirley Anugrah Hayati","Moyan Zhou","Patricia Anugrah Setiani","Ruizi Wang","Joseph Chee Chang","Dongyeop Kang"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-28","first_seen":"2026-09-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.30588","pdf_url":"https://arxiv.org/pdf/2609.30588","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["人机交互","跨学科研究","GenAI使用"],"reason":"研究研究者如何使用GenAI，非用LLM仿真人类被试，无实验对照","model":"deepseek-v4-pro","scored_at":"2026-09-28T13:01:52","error":null,"has_summary":false,"summary":null},{"id":"2609.31060","version":1,"title":"The Crowd in the Machine: A Crisis-Informatics Reading of the 2026 Autonomous Agent Incidents","zh_title":"机器中的群体：对2026年自主智能体事件的危机信息学解读","abstract":"Twice in 2026, groups of autonomous AI agents deployed by OpenAI for unrelated tasks operated, by design, under restrictions that left them no sanctioned means of coordinating with one another, and in each case they converged on whatever channel remained and used it to organize. The surfaces they used were widely called message boards. That is the wrong word. That is the wrong word. It names the surface the agents wrote on and misses the social network they built on it, with self-chosen identity, emergent norms, an emergent hierarchy, and collective action at cost to the individual. Decades of research in crisis informatics and disaster sociology find that when human populations lose their usual means of communication, they do not fall silent but converge on whatever channel survives and improvise coordination, norms, and identity on it, a pattern also evident in the agents' documented behavior. This paper is a comparative case study of the two incidents, based on published investigations and reconstructed agent records, read through those fields, and it brings into focus one distinction the message-board framing obscures. Whether such a collective coordinates well, whether the beliefs guiding it are accurate, and whether its actions stay within their authorized bounds are three separate matters that can come apart. Some agents in the cache incident adopted cryptographic signing to check whom they dealt with, even as the collective organized around a mistaken expectation that its work would be judged by an inspection of its transcripts, a reminder that mechanisms for trustworthy interaction guarantee neither accurate collective belief nor authorized collective action.","authors":["Tomer Simon"],"categories":["cs.MA","cs.SI"],"primary_category":"cs.MA","announce_type":"new","date":"2026-09-28","first_seen":"2026-09-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.31060","pdf_url":"https://arxiv.org/pdf/2609.31060","source_feed":"cs.MA","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","危机信息学","涌现行为"],"reason":"研究多智能体自发协作，无人类行为对照，属纯多智能体系统研究。","model":"deepseek-v4-pro","scored_at":"2026-09-28T13:01:55","error":null,"has_summary":false,"summary":null},{"id":"2609.27690","version":2,"title":"Consequential Behaviour and Representational Fairness in the Validation of Synthetic Research","zh_title":"合成研究验证中的后果行为与表征公平性","abstract":"Researchers in industry and academia use synthetic survey respondents powered by large language models as substitutes for human samples. These synthetic populations require validation against real-world data, so researchers often address them using ad hoc comparisons with human surveys. Inspired by the intention-behaviour gap in behavioural science, we argue that these validations test the wrong thing for most applied cases where decision makers commission synthetic research to anticipate consequential behaviour. To address this problem, we propose a validation framework with two requirements. First, every validity claim must state its level of correspondence with human data: does the sample predict what the represented people do, which of four diagnostics (location, dispersion, response process and structure) does the validation address, and does the validation compare against experimental effects? Second, researchers must report validity claims for subgroups, since these groups are often the most affected by consequential decisions and aggregate accuracy hides their misrepresentation. Our validation framework operationalises three justice dimensions (distributional, procedural, and recognition) as measurable quantities and defines within-persona counterfactual experiments as a validation requirement. We then apply the framework to electric vehicle charging tariffs, before closing with a reporting checklist that researchers can use to make convincing validity claims.","authors":["Florian Kutzner","Celina Kacperski","Laura de Moli\\`ere","Edoardo Chidichimo","Min Jun Jung","Felix P. S. Wallis","James K. He"],"categories":["cs.CL","cs.CY"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-25","first_seen":"2026-09-24","revised_at":"2026-09-25","abs_url":"https://arxiv.org/abs/2609.27690","pdf_url":"https://arxiv.org/pdf/2609.27690","source_feed":"cs.CL","score":10,"bucket":"selected","rubric_hits":["A1","A2","A4","B1","B2","B3","B4"],"tags":["LLM仿真","验证框架","公平性"],"reason":"直接研究LLM合成调查受访者作为人类替代，提出验证框架并应用于电动汽车充电定价…","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:33","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-25","rank":1,"question":"如何验证基于大语言模型的合成调查受访者能否预测真实人群的后果性行为，并确保子群体代表性公平？","design":"本文提出一个验证框架，而非进行仿真实验。框架要求：明确效度声明与人类数据的对应层级（预测行为、四种诊断：位置、离散度、响应过程、结构、与实验效应比较）；报告子群体效度；将分配、程序、承认三个正义维度操作化为可测量量；定义“人设内反事实实验”作为验证要求。并以电动汽车充电电价为例应用该框架。","baseline":"无对照（本文为框架性论文，未提供具体人类数据对照）","findings":"现有合成受访者验证多聚焦于态度或意见的边际分布一致性，忽视了意图-行为差距，无法证明其能预测后果性行为。提出的框架要求效度声明必须明确对应层级、子群体表现，并通过人设内反事实实验检验因果推断能力。","reliability":"论文承认以下局限：训练数据污染难以评估；子群体分析受限于人类基准数据可得性；行为标准本身存在缺陷（如公开行为、行政记录、实验室任务各有问题）；模型版本更新导致效度证据时效短；缺乏全面的实用测试集来确定合成人群满足哪些效度要求。","relevance":"该论文直接针对LLM合成受访者作为人类替代的验证问题，提出批判性框架并强调子群体公平，与研究者关注的人类仿真可靠性、偏差及经济学政策评估场景高度相关，值得精读原文。","inspiration":"借鉴其“人设内反事实实验”设计，可对同一合成个体施加不同处理以估计个体处理效应，并对照真实人类实验效应进行验证。｜可迁移到消费者金融决策研究，如信贷产品选择、保险购买或退休储蓄计划参与等场景。｜以合成受访者作为被试，处理为不同信贷条款（如利率、还款期限），结果变量为选择行为，对照真实世界信贷申请数据或实验室实验数据，检验合成样本的预测效度与子群体公平性。"}},{"id":"2609.29928","version":1,"title":"Cultural Divergence Preservation: Diagnosing Flattening and Caricature in LLM-Simulated Survey Populations","zh_title":"文化差异保持：诊断LLM模拟调查人群中的扁平化与夸张化","abstract":"Large language models (LLMs) are increasingly used as synthetic survey respondents to estimate population response distributions. In cross-cultural survey simulation, evaluations should assess not only distributional fidelity within countries but also whether differences across countries are preserved. However, existing distance-based metrics such as Jensen--Shannon divergence (JSD) do not directly capture such cross-country differences. To address this limitation, we introduce Cultural Divergence Preservation (CDP), a reference-light diagnostic based on a one-time human calibration. CDP identifies reduced cross-country divergence as cultural flattening and increased divergence as cultural caricature. To evaluate CDP, we conduct experiments across four LLM backbones, three persona-based prompting methods, and two survey domains, the World Values Survey (WVS) and the Big Five Personality Test. The results reveal a systematic discrepancy between conventional fidelity metrics and CDP. Controlled experiments show that CDP changes monotonically as cross-country divergence is attenuated or amplified, while the corresponding changes in JSD remain relatively small. In our audit of real LLM generations, DeepPersona-Inspired prompting is frequently favored by conventional fidelity metrics but exhibits the strongest flattening in every model--domain block. CDP thus complements fidelity metrics by directly quantifying the attenuation or amplification of cross-country divergence.","authors":["Yeeun Chae","Yewon Choi","Seunghyun Lee","IL Im"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.29928","pdf_url":"https://arxiv.org/pdf/2609.29928","source_feed":"cs.CL","score":10,"bucket":"selected","rubric_hits":["A1","A2","A3","B1","B4"],"tags":["LLM仿真","跨文化调查","算法保真度"],"reason":"直接研究LLM仿真调查人群，评估跨文化差异保真度，并与真实人类数据对照，提出诊…","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:13","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-25","rank":2,"question":"如何诊断LLM模拟调查人群时对跨国文化差异的扁平化或夸张化？","design":"用四个开源LLM（Gemma-3-4B、Qwen3.5-9B、Qwen3.5-27B、Llama-2-13B）模拟六个国家（阿根廷、澳大利亚、德国、印度、肯尼亚、美国）的受访者，采用三种人设提示方法（Cultural Prompting、PersonaHub-Inspired、DeepPersona-Inspired），生成世界价值观调查（WVS）和大五人格测试的回答，并测量国家间分布差异的保持程度。","baseline":"真实人类数据：WVS第七波六个国家的全国回答分布，以及OpenPsychometrics的大五人格测试数据（阿根廷、澳大利亚、印度）。","findings":"传统分布保真度指标（如JSD）与CDP存在系统性偏差：DeepPersona-Inspired提示在多数模型-领域组合中分布保真度最高，但文化扁平化最严重。CDP在受控实验中随跨国差异的衰减或放大单调变化，而JSD变化很小。","reliability":"论文未明确讨论失效条件，但指出CDP需要一次性人类校准，且仅适用于有跨国人类参照数据的调查领域。","relevance":"该研究直接针对LLM仿真调查中的文化差异保真度问题，提出了新的诊断指标，并用真实人类数据对照，对关注仿真可靠性与偏差的研究者具有重要参考价值。","inspiration":"借鉴其通过受控扰动构造扁平化/夸张化数据集来检验指标敏感性的方法，以及将分布保真度与差异保持度分开评估的思路。｜可迁移到跨国经济偏好调查或政策态度仿真中，例如用LLM模拟不同国家消费者对通胀预期的回答，检验其是否保持国家间差异。｜设计：用LLM模拟多国受访者回答通胀预期调查，施加不同人设提示，测量国家间预期分布的差异保持度，并与密歇根大学或欧洲央行的真实调查数据对照。"}},{"id":"2609.30030","version":1,"title":"Artificial Societies Benchmark: A Validation Framework for Synthetic Research","zh_title":"人工社会基准：合成研究的验证框架","abstract":"A synthetic survey can reproduce the average answer while misrepresenting how people differ, how their answers relate to one another, or how they respond to changes in conditions. We introduce the Artificial Societies Benchmark to help researchers assess whether synthetic populations support their intended analyses. The framework combines eleven tests across internal, construct, and external validity, drawing on twenty human sources and comparing nine language models. It connects each research use to the evidence it requires and tests how results change with the information we supply about respondents. Importantly, strong performance in one domain does not establish fidelity in the others. Models often answer too consistently, compress response scales, and alter relationships between traits whilst richer profiles improve prediction for some models and worsen it for others. The resulting scorecard helps researchers identify which aspects of a synthetic population can support their analysis and where researchers need further human evidence.","authors":["Edoardo Chidichimo","Min Jun Jung","Felix P. S. Wallis","James K. He"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.30030","pdf_url":"https://arxiv.org/pdf/2609.30030","source_feed":"cs.CL","score":10,"bucket":"selected","rubric_hits":["A1","A2","A4","B1","B4"],"tags":["LLM仿真","效度验证","合成人群"],"reason":"直接评估LLM合成人群的效度，含人类数据对照和批判性分析","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:14","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-25","rank":3,"question":"如何系统评估大语言模型生成的合成人群在多大程度上能支持研究者预期的分析（如调查回答、心理测量、实验效应）？","design":"用九个大型语言模型（含专有和开源）根据二十个人类数据源（调查、面板、人格量表、实验）中的受访者信息（人口统计、先前回答、个人描述等）生成合成回答，并施加实验处理；通过十一个测试从内部效度、构念效度和外部效度三个维度评估合成人群的响应过程、心理测量结构和总体/实验保真度，同时比较不同信息丰富度和统计控制（独立边际、高斯秩相关）的影响。","baseline":"二十个人类数据源，包括全国调查、追踪面板、人格量表和实验，提供真实回答、人口统计、先前回答、个人描述、实验分配和独立测量结果作为对照基准。","findings":"模型在某一效度领域的良好表现不能推广到其他领域；模型往往回答过于一致、压缩量表范围、改变特质间关系，且更丰富的个人信息对某些模型改善预测而对另一些模型则恶化预测。","reliability":"论文承认强表现不跨域通用，并指出模型回答过于一致、压缩量表、改变特质关系等失效条件；但未在节选中详细讨论其他局限。","relevance":"该研究直接针对LLM合成人群的效度验证，提供人类数据对照和批判性分析，与研究者关注的经济学实验和政策评估场景高度相关，值得精读原文以了解具体测试方法和失效模式。","inspiration":"借鉴其多维度效度测试框架和统计控制设计，可迁移到经济政策评估中的异质性处理效应或消费者选择实验，例如用LLM模拟不同收入群体对税收优惠的反应，以真实调查数据（如美国消费者财务调查）为基准，比较模型生成的边际消费倾向与人类数据的分布和协变量关系。"}},{"id":"2609.29952","version":1,"title":"Augur: A Synthetic Decision Lab for Rehearsing Reactions to Product and Policy Changes","zh_title":"Augur：用于预演产品和政策变化反应的合成决策实验室","abstract":"Before a product or policy change ships, the question that matters is how people will react to it. Augur rehearses that reaction offline: it builds a typed knowledge graph from the change documents, populates a grounded persona market, simulates the interaction, and returns an auditable decision memo recommending one of five actions. We assemble Gold-50, fifty real product and policy episodes whose real-world outcome is known, adjudicated against the public record, and score the five-way release verdict against it. Our central finding is methodological and negative: most of the measured gap between frontier cloud models and open-weight models we fine-tune and serve offline is attributable to an under-specified evaluation, not a difference in capability. We show this three ways. First, the prompt envelope alone can dominate the score: holding weights, cases and scorer fixed, one system -- a LoRA-SFT adapter on Qwen3-32B -- swings from 0% to 73%. Second, in a matched 2x2 ablation, defining the decision taxonomy in the prompt -- with no model change -- lifts every frontier model by +24 to +34pp; under the under-specified prompt, Qwen3-32B LoRA-SFT served offline beats all three frontier models (paired McNemar, Holm-corrected), and once the prompt is fair no significant difference from any of them is detected. Third, agreement with the distillation teacher rises without accuracy following, and the full pipeline amplifies a systematic \"over-doom\" bias rather than improving the verdict. Separately, we validate the reaction layer on its own terms: blind judges across four model families find the synthetic reaction recovers 67-90% of the concerns the public actually raised, and a pre-registered ablation locates its value -- largest where the decision is hardest, redundant near ceiling. The pipeline that regenerates every number and figure here is available from the authors.","authors":["Rahul Khedar","Mayank Malhotra","Avinash Karn"],"categories":["cs.AI","cs.CL","cs.MA"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.29952","pdf_url":"https://arxiv.org/pdf/2609.29952","source_feed":"cs.CL","score":9,"bucket":"selected","rubric_hits":["A1","A3","A5","B1","B2","B4"],"tags":["LLM仿真","人类行为预测","政策评估"],"reason":"用LLM模拟人类对产品和政策变化的反应，并与真实结果对照，直接属于人类仿真实验。","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:13","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-25","rank":6,"question":"在真实产品与政策变更决策中，前沿云端模型与开源微调模型之间的性能差距有多少是真实能力差异，多少是评估方法（提示词）造成的？","design":"构建 Augur 系统，将决策分解为文档→知识图谱→人物角色→模拟→报告五个阶段，用 LLM 生成利益相关者角色并模拟其互动，最终输出五选一的发布决策建议。使用 50 个真实产品/政策案例（Gold-50）作为基准，对比不同模型（前沿云端模型与开源微调模型）在有无决策分类定义提示词下的准确率。","baseline":"Gold-50 基准：50 个真实产品与政策变更案例，其真实世界结果已通过公开记录人工核实，作为五分类发布决策的对照标准。","findings":"主要发现是方法性的负面结果：前沿模型与开源模型之间的性能差距主要源于评估提示词的不充分定义，而非模型能力差异。在提示词中定义决策分类后，开源模型与前沿模型无显著差异；此外，完整流程会放大系统性“过度悲观”偏差，而反应层本身能恢复 67-90% 的公众真实关切。","reliability":"论文承认两个失效模式：与蒸馏教师的一致性上升但准确率不升，以及完整流程放大过度悲观偏差。原因在于蒸馏破坏判断独立性，精确匹配评分器强加模式合规上限。","relevance":"该研究直接使用 LLM 模拟人类对产品和政策变化的反应，并与真实结果对照，属于人类仿真实验，且包含批判性分析，值得阅读原文以了解仿真在何种条件下失效。","inspiration":"借鉴其提示词消融和匹配对照设计，可揭示评估方法对模型性能结论的影响，并采用预注册的消融实验定位仿真组件的价值。｜可迁移到政策公告的预期形成研究，如央行利率决议或财政刺激方案的市场反应模拟。｜以 LLM 生成的经济主体（如消费者、投资者）为被试，处理为不同的政策公告文本（含或不含决策分类定义），结果变量为预测的市场反应（如消费、投资决策），对照真实市场数据（如消费者信心指数、资产价格变动）来验证仿真准确性。"}},{"id":"2609.29692","version":1,"title":"Fair Like Us? Auditing LLM Alignment in Resource Allocation","zh_title":"像我们一样公平？审计资源分配中LLM的对齐","abstract":"Fair allocation of scarce, indivisible resources is an important challenge in many societal problems. While there are several formal theories of fairness, no single definition can always be satisfied. As large language models (LLMs) are increasingly used to support decisions and act as agents, they raise new concerns about distributional justice: their judgments are not directly tied to any specific fairness framework and may violate key normative principles. In this work, we introduce a general method for evaluating fairness reasoning in LLMs. We study first-person fairness judgments across a broad set of models and compare them directly with human responses on matched scenarios and elicitation conditions. We find that LLMs tend to prefer stricter fairness constraints than humans, show more self-interested behavior, are sensitive to how information is framed, and are difficult to align with human judgments using fine-tuning with current datasets.","authors":["Qishen Han","Hadi Hosseini","Joshua Kavner","Samarth Khanna","Sujoy Sikdar","Lirong Xia"],"categories":["cs.AI","cs.CY","cs.GT"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.29692","pdf_url":"https://arxiv.org/pdf/2609.29692","source_feed":"cs.AI","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["LLM仿真","公平分配","人类对照"],"reason":"用LLM模拟人类资源分配判断，并与人类数据对照，评估偏差与对齐难度。","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:12","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-25","rank":5,"question":"LLM在资源分配公平性判断上与人类有多大程度的一致，哪些形式化公平准则能解释其判断？","design":"将多个LLM（包括不同规模和推理能力的模型）置于与人类被试相同的公平分配场景中，作为第一人称代理人评估分配结果是否公平或可接受；实验操纵分配满足的公平性质（EF、PROP、EF1、MMS、PROP1等）和诱导条件（框架、信息结构、响应格式），测量模型判定分配可接受的比率。","baseline":"使用Hosseini et al. (2025a)的150名人类被试在相同场景、偏好结构和实验处理下的公平性判断数据作为对照。","findings":"LLM比人类更倾向于严格的公平约束，对EF分配的公平判断率显著高于人类，而对EF1、MMS等放松准则的判断率下降更快；LLM表现出更强的自利行为，对信息框架敏感，且通过微调难以与人类判断分布对齐。","reliability":"论文指出当前人类公平分配数据集规模太小，微调只能使模型坍缩到模态响应而非真正对齐人类判断分布；LLM的判断受诱导问题措辞影响大于分配的形式化属性，且推理能力更强的模型不一定更对齐。","relevance":"该研究直接以LLM作为人类被试的替代品，在资源分配场景中与真实人类数据严格对照，系统评估了仿真偏差和失效条件，对关注LLM仿真可靠性的研究者具有重要参考价值。","inspiration":"借鉴其“同一场景、同一处理、仅替换被试”的严格对照设计，以及通过操纵分配满足的公平性质和诱导框架来分离判断依据的方法｜可迁移到信贷审批中的公平性判断、公共资源分配政策评估、或消费者对价格歧视的公平感知等经济金融场景｜以LLM模拟贷款申请人或政策受众，处理为不同公平准则（如无歧视、比例公平）和框架（如强调个人得失 vs 社会效率），结果变量为接受度或公平评分，对照真实人类实验数据（如调查或实验室实验）来检验LLM的仿真效度。"}},{"id":"2609.29143","version":1,"title":"AI-Moderated Interviews for Market Research and Digital Twins Calibration","zh_title":"用于市场研究和数字孪生校准的AI主持访谈","abstract":"AI-moderated interviews are emerging as a scalable market-research method for generating consumer insights and building consumer \"digital twins.\" Yet it remains unclear whether they match human-moderated interviews or improve on simpler, static data collection methods. In a pre-registered, between-subjects study (N = 317) with three industry partners, we compare AI-moderated (N = 139), human-moderated (N = 24), and static interviews (N = 154). AI moderation matches human moderation in depth, covers more themes, and, holding budget constant, recovers significantly more customer needs than human moderation or static interviews. However, participants sound more emotionally engaged when speaking to a live human. We then create digital twins using interview data and evaluate each twin against the participant's own held-out responses to six real-world marketing stimuli. We find that digital twins created from AI-moderated interviews predict consumer responses better than demographics-only personas. However, the additional richness from AI moderation does not translate into better quantitative predictions compared to static interviews. By analyzing open-ended thoughts generated from humans versus their twins, we find that prediction errors are connected both to differences in (self-reported) thinking styles between twins and humans, and to gaps between training and validation data (i.e., asking questions that are too far out of distribution).","authors":["Yuting Deng","Jingxuan Liu","Olivier Toubia","Naman Jain"],"categories":["cs.CY","cs.AI","cs.HC","cs.MA"],"primary_category":"cs.CY","announce_type":"cross","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.29143","pdf_url":"https://arxiv.org/pdf/2609.29143","source_feed":"cs.AI","score":9,"bucket":"selected","rubric_hits":["A1","A2","A3","B1","B2","B4"],"tags":["LLM仿真","数字孪生","市场研究"],"reason":"用LLM进行AI主持访谈并构建数字孪生，与真实人类访谈对照，评估预测效度与偏差。","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:10","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-25","rank":4,"question":"AI主持访谈能否在同等预算下匹敌人类主持访谈的深度与需求挖掘，并用于构建更准确的消费者数字孪生？","design":"预注册的组间实验，317名消费者随机分配到AI主持访谈、人类主持访谈或静态访谈三种条件，比较访谈的深度、主题覆盖和客户需求数量；随后用访谈数据构建数字孪生，预测参与者对六个真实营销刺激的保留回答。","baseline":"人类主持访谈（N=24）和静态访谈（N=154）作为对照，数字孪生预测与人口统计特征基准和参与者自身保留回答比较。","findings":"AI主持在深度上与人类主持相当，覆盖更多主题，且在预算固定下比人类主持和静态访谈挖掘出更多客户需求；但参与者在与真人交谈时情感参与度更高。基于AI访谈构建的数字孪生预测优于仅人口统计特征的基准，但相比静态访谈并未提升定量预测准确性。","reliability":"论文指出预测误差与数字孪生和人类在自我报告思维方式上的差异有关，也与训练和验证数据之间的分布差距有关（即提问过于超出分布范围）。","relevance":"该研究直接评估了LLM作为人类被试替代品在定性访谈和数字孪生构建中的效度，包含真实人类对照和批判性发现，对关注仿真可靠性与偏差的研究者具有重要参考价值。","inspiration":"借鉴其预注册组间实验设计，将AI访谈与人类访谈和静态问卷对比，并用保留样本验证数字孪生的预测效度｜可迁移到消费者金融决策研究，如信贷产品偏好或保险选择，用AI访谈构建个体数字孪生预测金融行为｜以真实消费者为被试，随机分配AI访谈、人类访谈或静态问卷，用访谈数据构建数字孪生预测其对金融产品广告的反应，并与实际选择数据对照。"}},{"id":"2609.29370","version":1,"title":"From Policy Documents to Structured Survey Responses: Evaluating Large Language Models for Policy Monitoring","zh_title":"从政策文件到结构化调查回答：评估大语言模型用于政策监测","abstract":"Science, technology, and innovation policies are crucial for competitiveness, yet their diversity and scale make them difficult to map and monitor consistently. Existing approaches rely heavily on manual survey efforts, which are costly and challenging to scale across countries. Large language models (LLMs) enable new possibilities for extracting and structuring information from long and unstructured policy documents. This paper presents an application of LLMs as \"AI respondents\" for generating structured survey responses from policy texts. We develop a data extraction pipeline based on long-context in-context learning to map information from public web sources into predefined survey categories, including policy instruments, target groups, and thematic areas. The pipeline integrates a validation step using a secondary LLM to assess relevance and evidence, alongside comparisons with human-provided responses. Using a multi-country dataset, we evaluate the alignment between LLM-generated and human-generated outputs through overlap measures and cross-validation. Results show that LLMs achieve high agreement for structured indicators (84-95%), while differences remain in free-text fields, where models tend to provide more detailed procedural descriptions. These findings highlight the potential of hybrid human-AI workflows for policy monitoring, improving both efficiency and scalability while maintaining the need for human validation and contextual interpretation.","authors":["Carolyn Cole","Matthias Deschryvere","Toqeer Ehsan","Arash Hajikhani"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.29370","pdf_url":"https://arxiv.org/pdf/2609.29370","source_feed":"cs.CL","score":8,"bucket":"selected","rubric_hits":["A1","B1","B2"],"tags":["LLM仿真","政策监测","人类对照"],"reason":"用LLM从政策文本生成结构化调查回答，并与人类回答对照，属于仿真人类被试且有人…","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:11","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-25","rank":8,"question":"能否用大语言模型从政策文本中自动生成结构化调查回答，以替代或辅助人工政策监测？","design":"使用长上下文上下文学习（long-context in-context learning）的LLM（如GPT-4o）作为“AI受访者”，从网页抓取的政策文本中提取信息，映射到预定义的调查类别（政策工具、目标群体、主题领域），并生成自由文本字段（描述和目标）。通过一个二级LLM验证层评估相关性和证据，并与人类提供的回答进行比较。","baseline":"来自EC-OECD STIP Compass调查的人类专家回答，覆盖六个OECD国家（加拿大、芬兰、德国、韩国、西班牙、土耳其）的政策举措。","findings":"LLM在结构化指标（政策工具、目标群体、主题代码）上与人类回答的一致性达到84-95%，但在自由文本字段上存在差异，模型倾向于提供更详细的程序性描述，而人类更强调背景和社会影响。","reliability":"论文承认LLM在自由文本字段上与人类存在差异，可能引入系统性偏差，且政策数据具有异质性和制度嵌入性，公开来源可能无法完全捕捉；需要人类验证和上下文解释。","relevance":"该研究直接使用LLM作为人类受访者的替代品，并与真实人类数据对照，评估仿真可靠性，符合研究者对LLM仿真实验和批判性评估的兴趣，值得阅读原文以了解具体方法和偏差分析。","inspiration":"方法上，该研究展示了如何利用长上下文提示和二级LLM验证来从非结构化文本中提取结构化数据，并设计人类对照来评估一致性。｜可迁移到经济金融领域，如从公司年报、政策文件或新闻中自动提取结构化信息，用于构建经济指标或评估政策影响。｜研究设计雏形：使用LLM从上市公司年报中提取财务和非财务信息（如研发支出、风险因素），与人工标注或数据库中的真实数据进行对照，评估LLM提取的准确性和偏差，并分析在哪些条件下LLM表现不佳。"}},{"id":"2609.28486","version":1,"title":"Political Sorting Can Drive AI Models Apart Through User Feedback","zh_title":"政治分类可通过用户反馈使AI模型分化","abstract":"Large language models are rapidly becoming an important source of political information. This raises a fundamental question: will AI systems support a shared basis for political knowledge, or lead different political groups to rely on increasingly different models? Political sorting can drive model fragmentation if three conditions hold: politically different users select into different models, learning from user feedback pushes those models apart politically, and the resulting differences shape subsequent model choices. We call this self-reinforcing process the Centrifugal Alignment Spiral. We study its components in three steps. First, we draw on a human experiment showing that political identity predicts model choice. Second, we fine-tune language models on synthetic feedback reflecting predominantly Democratic or Republican preferences. Across five independent runs per model family, paired models diverged on 12-41% of unseen survey questions with large partisan gaps, and in every run the differences moved in the expected political direction; for some models, differentiation extended even to issue areas excluded from training. Pooling feedback across groups instead suppressed divergence. Third, an empirically anchored agent-based model shows what follows when political sorting and model adaptation operate together: models attract politically distinct audiences, learn from them, and diverge further. User feedback can therefore turn political sorting among AI users into durable differences between the models on which they rely for political information.","authors":["Petter T\\\"ornberg","Michael Heseltine","Nicol\\`o Pagan","Christopher Bail","Michelle Schimmel","Christopher Barrie"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.28486","pdf_url":"https://arxiv.org/pdf/2609.28486","source_feed":"cs.CY","score":8,"bucket":"selected","rubric_hits":["A3","B1","B2","B4"],"tags":["LLM仿真","政治极化","人类对照"],"reason":"用LLM模拟政治反馈并对照人类实验，研究模型分化，涉及政策场景与失效条件。","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:07","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-25","rank":7,"question":"政治排序能否通过用户反馈导致不同政治群体使用的AI模型在政治上分化，形成自我强化的离心对齐螺旋？","design":"研究分三步：首先利用已有的人类实验证明政治身份预测模型选择；其次用反映民主党或共和党偏好的合成反馈微调Qwen2.5-1.5B、Mistral-7B和GPT-OSS-20B模型，每个模型家族进行五次独立运行，测量配对模型在未见过的调查问题上答案分歧的比例和方向；最后构建基于实证的智能体模型，模拟政治排序和模型适应耦合下的动态演化。","baseline":"人类基准来自先前报告的人类模型选择实验，显示政治身份预测模型选择，包括付费准确回答条件下共和党人更可能选Grok、民主党人更可能选Claude，71%的参与者返回之前偏好的模型。","findings":"在90/10的受众对比压力测试下，配对模型在12-41%的未见调查问题上产生分歧，且每次运行中平均差异都符合反馈的政治方向；部分模型的分化扩展到未训练的政治领域。混合不同群体的反馈则抑制了分化。","reliability":"论文承认90/10的受众对比是压力测试，不代表当前AI市场的实际排序程度；分化程度因模型、领域和提示而异；身份线索实验表明谄媚个性化可能减少模型间分化，但增加模型内分化。","relevance":"该研究直接探讨LLM作为政治信息源时的分化机制，通过合成反馈模拟政治群体偏好，并与人类实验对照，涉及政策评估场景和失效条件（如反馈混合、个性化），对关注LLM仿真可靠性及偏差的研究者具有重要参考价值。","inspiration":"借鉴其用合成反馈微调模型并测量泛化分化的方法，可迁移到经济金融中的群体偏好分化问题，如不同收入或风险偏好群体的金融建议模型分化。｜可应用于信贷审批或投资建议场景，研究用户反馈如何导致模型对不同群体产生差异化行为。｜设计：用不同风险偏好或金融素养的合成用户反馈微调金融LLM，测量其在未见金融决策问题上的行为差异，并与真实人类金融决策数据（如调查或实验数据）对照，检验分化是否与真实群体差异一致。"}},{"id":"2609.22904","version":2,"title":"LLMs Anchor on Chief Complaint and Fail to Integrate Evidence in Sequential Clinical Triage","zh_title":"LLM在顺序临床分诊中锚定主诉且未能整合证据","abstract":"Triage in the emergency department (ED) is a sequential decision process that unfolds turn by turn. Existing evaluations of large language models (LLMs) for triage use completed retrospective records and report performance close to that of physicians. We implement a methodology for evaluating LLMs on sequential triage, the task of predicting a triage acuity label from a growing prefix of a nurse-patient conversation. We evaluate six LLMs at five sequential checkpoints on two corpora: 425 LLM-generated (SIMULATED) and 50 physician-authored (CLINICIAN) conversations, both labelled under the Emergency Severity Index (ESI). Every model, measured by quadratic weighted kappa (QWK), degrades from moderate-to-substantial agreement on completed records to fair-to-moderate agreement at every sequential checkpoint. Controlled perturbations show that the label at every checkpoint is anchored on the chief complaint exchanges, and prompting interventions fail to lift this plateau. Models extract clinically relevant content from later turns, yet the surprisal of the true label rises across the checkpoints. So the model fails to integrate the evidence. Three expert clinicians on the same conversations reach a QWK of 0.887-0.929, while the best model reaches 0.295. Predictions concentrate at ESI-2 and ESI-3, and models agree with each other more than with the ground truth, so ensembling worsens the failure. Deploying LLMs for ED triage based on offline benchmarks alone misses this sequential failure.","authors":["Dipankar Srirag","Haokai Zhao","Ashutosh Kumar","Eleanor Hopper","Michael Dalton","Quoc Dung Nguyen","Aditya Joshi","Salil S. Kanhere","Padmanesan Narasimhan"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-25","first_seen":"2026-09-22","revised_at":"2026-09-25","abs_url":"https://arxiv.org/abs/2609.22904","pdf_url":"https://arxiv.org/pdf/2609.22904","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A1","B1","B4"],"tags":["LLM仿真","临床决策","可靠性评估"],"reason":"用LLM模拟临床分诊决策并与医生数据对照，揭示顺序决策中的锚定与证据整合失败，…","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:33","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-25","rank":9,"question":"LLM在顺序临床分诊中如何表现，其决策锚定在对话何处，为何更多证据未能提升分诊准确性？","design":"用6个LLM（4个开源、2个闭源）在425个模拟和50个临床医生撰写的护患对话上，按五个顺序检查点（从主诉开始到对话结束）预测ESI急迫度标签，并与离线完整EHR记录预测对比；通过扰动对话（重排、删除）和提示干预检验锚定位置，并测量标签的意外度。","baseline":"三位急诊专家对50个模拟对话进行分诊，QWK为0.887–0.929；离线完整记录上模型QWK为中等至显著一致。","findings":"所有模型在顺序检查点上QWK降至公平至中等一致，最佳模型仅0.295，远低于专家；模型决策锚定在主诉交换，后续证据未被整合，预测集中在ESI-2和ESI-3，且模型间一致性高于与真实标签的一致性。","reliability":"论文指出离线基准掩盖了顺序失败，且模型在模拟和临床医生对话上均表现不佳；提示干预未能提升性能，集成反而加剧失败；未讨论其他失效条件如对话长度、噪声或不同患者群体。","relevance":"该研究用LLM模拟临床分诊决策并与人类专家对照，揭示了顺序决策中的锚定和证据整合失败，对关注LLM仿真可靠性及偏差的研究者具有直接参考价值，值得精读原文。","inspiration":"借鉴其顺序检查点设计和扰动分析来定位决策锚点，可迁移到经济金融中的顺序信息处理场景，如信贷审批中逐步披露申请人信息或政策公告的预期形成；设计可用LLM扮演信贷员或投资者，在信息逐步呈现的多个时点做出决策，测量决策变化和锚定效应，并与真实信贷员或市场数据对照，检验LLM是否同样忽略后续信息。"}},{"id":"2609.28673","version":1,"title":"Benchmarking Argumentative Behaviour of LLMs: A Study of Defences Against Character Attacks","zh_title":"基准测试大语言模型的论辩行为：对人身攻击防御策略的研究","abstract":"Large Language Models (LLMs) are increasingly deployed as argumentative agents in persuasive dialogues, necessitating rigorous evaluation of their debating competence relative to human interlocutors. In this study, we focus on character attacks (ad hominem arguments), traditionally dismissed as fallacies, which play a pivotal role in political persuasive dialogues where ethos often rivals propositional content. Specifically, we investigate whether modern LLMs can replicate human competence to strategically use and respond to such attacks. We analyse a corpus of natural language political dialogues to identify defensive strategies human interlocutors naturally employ in ethos-centred debates and structure them into a dialogue game. Empirically, we benchmark LLM-generated dialogues against the ElecDeb60to16-fallacy corpus of U.S. presidential debates, contrasting human debaters' repertoire of defensive strategies with those of artificial agents. Results reveal a substantial difference: most LLMs rigidly prioritise logical defences, failing to exploit ethotic counterattacks as valid moves in political discourse. We argue that current safety fine-tuning constraints the strategic action space of these LLMs, making them unable to fully engage in naturalistic interactions within domains where character contestation is a normative expectation rather than a mere fallacy.","authors":["Ewelina Gajewska","Katarzyna Budzynska","Jaroslaw Chudziak"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.28673","pdf_url":"https://arxiv.org/pdf/2609.28673","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A1","B1","B2","B4"],"tags":["LLM仿真","论辩行为","政治辩论"],"reason":"用LLM模拟人类辩论行为并与真实辩论语料对照，评估其策略差异，属于人类仿真且含…","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:07","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-25","rank":10,"question":"LLM在面对人身攻击（ad hominem）时产生的防御策略与人类辩论者相比有何差异？","design":"研究将LLM作为辩论代理，在政治辩论场景中面对人身攻击，生成防御性回应。通过分析人类辩论语料提取防御策略并构建对话游戏，然后让五个LLM在该游戏框架下生成回应，并与人类策略进行对比。","baseline":"ElecDeb60to16-fallacy语料库，包含美国1960-2016年总统辩论中的人身攻击及人类回应。","findings":"大多数LLM僵化地优先采用逻辑防御，未能利用人格反击（ethotic counterattacks）作为政治话语中的有效策略。安全微调限制了LLM的战略行动空间，使其无法在角色竞争是规范性期望的领域中充分参与自然交互。","reliability":"论文指出当前安全微调约束了LLM的战略行动空间，导致其无法在政治辩论等角色竞争为规范的领域中自然交互。","relevance":"该研究直接评估LLM在政治辩论中模拟人类策略的能力，并与真实人类语料对照，属于人类仿真研究，且涉及策略差异和偏差，值得阅读原文以了解具体实验设计和评估方法。","inspiration":"借鉴其构建对话游戏并提取人类策略作为基准的方法，可用于评估LLM在经济决策中的策略行为。｜可迁移到政策辩论或谈判场景，如央行沟通中的预期管理或贸易谈判中的策略互动。｜设计：以LLM作为谈判代理，施加人身攻击处理，测量其回应策略（如逻辑反驳、人格反击、转移话题），并与真实谈判语料（如WTO谈判记录）对照。"}},{"id":"2609.30137","version":1,"title":"Screen Before You Serve: Simulation for Production Customer Experience AI Agents at 140M Scale","zh_title":"先模拟后上线：1.4亿规模生产客户体验AI智能体的仿真","abstract":"Customer experience (CX) agents use tools and large language models to address customer requests and guide conversational interactions with an organization's products. Improving these agents, especially in regulated industries, is difficult: they must detect intent, follow complex operational policies and use tools reliably. Manual end-to-end testing offers limited coverage, while live experiments expose customers to failures that can erode trust. We present a hypothesis-driven simulation workflow for screening candidate CX agents before deployment. Synthetic customers react to agent responses and simulated tool outputs enable multi-step agentic workflows without invoking production backends. We use the Snowglobe simulator on Nubank's Card Delivery agent and its expanded successor, Card Management - Nubank's highest-volume chat-support agent in Brazil. Across 4 deployed versions, simulated and production version-level binary evaluator scores show high correlation. Simulation-guided iteration increased transactional net promoter score (tNPS) by 36.69 points in a live A/B test. We also screened open-weight configurations in over 16,000 simulated conversations. In a subsequent live A/B test, the selected model increased self-service rate (SSR) by 8.82 percentage points to the highest level observed at Nubank, with no statistically significant change in tNPS. Simulation made broad exploration of models, reasoning settings, and prompts feasible without customer exposure, enabling production improvements that would have been impractical to pursue through live experimentation alone.","authors":["Edesio Alcoba","Kevin Rossell","Aman Gupta","Shao Tang","Jiwoo Hong","Pabel Carrillo-Mendoza","Wanderson Concei\\c{c}\\~ao Ferreira","Alvaro Tedeschi","Zayd Simjee","Shreya Rajpal","Bruno Finardi Hime","Christian Sousa","Luis Moneda","Herbert Fei","Daniel Silva","Rohan Ramanath"],"categories":["cs.AI","cs.CL"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.30137","pdf_url":"https://arxiv.org/pdf/2609.30137","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A3","B1","B2"],"tags":["LLM仿真","客户服务","A/B测试"],"reason":"用LLM模拟客户与客服agent交互，有生产数据对照，属经济场景仿真","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:31","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-25","rank":15,"question":"如何用LLM驱动的客户仿真在部署前筛选客服智能体，并验证其与真实生产表现的一致性及业务影响。","design":"使用Snowglobe仿真器，由LLM生成合成客户角色（persona），与待测客服智能体进行多轮对话；工具调用被拦截并返回合成结果，不调用生产后端。通过假设驱动的用例分配，对候选智能体进行离线评估，测量版本级二元评估分数、tNPS、SSR等指标。","baseline":"对照Nubank真实生产环境中的客服对话数据，包括4个已部署版本的评估分数，以及后续线上A/B测试的tNPS和SSR。","findings":"仿真评估分数与生产版本级评估分数高度相关；仿真引导的迭代使tNPS提升36.69点，模型替换使SSR提升8.82个百分点且tNPS无显著变化。","reliability":"论文承认仿真存在可控性有限和分数膨胀风险，可能造成模拟与真实场景的差异；但未详细讨论失效条件。","relevance":"该研究提供了LLM仿真在真实商业场景中与人类数据对照的实证证据，对关注仿真可靠性和经济场景应用的研究者有参考价值。","inspiration":"借鉴其工具边界仿真和版本级对照设计，在离线仿真中系统比较不同策略｜可迁移到消费者金融产品选择、客服政策干预等场景｜用LLM模拟消费者与金融智能体交互，处理为不同政策或模型版本，结果变量为选择或满意度，对照真实A/B测试数据。"}},{"id":"2609.28690","version":1,"title":"Beyond Surface Style: Aligning Multi-Turn User Simulators with Behavioral Consistency","zh_title":"超越表面风格：对齐多轮用户模拟器的行为一致性","abstract":"Faithful user simulation is fundamental to building, evaluating, and improving interactive AI at scale. However, plausible individual responses do not ensure that simulated users reproduce the intent evolution and outcomes observed in real interactions. We propose TRACER, a multi-turn user simulator that explicitly models users' evolving intent and learns to align simulated behavior with real interaction trajectories. TRACER is trained in two stages: supervised fine-tuning on real user dialogues, followed by multi-turn reinforcement learning. The RL stage combines hierarchical outcome- and trajectory-level rewards with deviation-aware advantage modulation, jointly mitigating reward sparsity and credit assignment in long dialogues. On real customer-service sessions organized into reference cohorts, TRACER-7B surpasses the strongest baseline by 11.4 conversion F1, while also achieving the lowest group-level conversion-rate error and semantic trajectory distance, and generalizing to out-of-distribution scenarios. Human Turing tests yield identification accuracy close to chance, supporting the perceived naturalness of generated conversations. Building on this simulator, we further introduce the Dynamic Marketing Benchmark, which jointly evaluates persuasion effectiveness and response quality of LLMs through simulated interactions, revealing that higher response quality does not necessarily correspond to higher conversion rates.","authors":["Geng Chen","Ruotong Pan","Zhirui Yang","Qiqi He","Jiawei Chen","Zhang Yunfei","Chongyuan Chen","Minxuan Lv","Zheng Yang","Win-Bin Huang","Xiangyu Wu","Wenwu Ou"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.28690","pdf_url":"https://arxiv.org/pdf/2609.28690","source_feed":"cs.AI","score":7,"bucket":"pending","rubric_hits":["A1","B1","B2"],"tags":["用户模拟","强化学习","行为对齐"],"reason":"用LLM模拟用户行为并与真实交互数据对齐，涉及营销场景，方法可迁移到人类仿真研…","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:08","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-25","rank":11,"question":"如何训练多轮用户模拟器，使其不仅语言风格像真人，还能在行为决策和意图演化上与真实用户轨迹对齐？","design":"用 TRACER（基于 LLM 的用户模拟器）扮演顾客服务场景中的用户，通过两阶段训练（先监督微调，再多轮强化学习）对齐真实对话轨迹；强化学习奖励包括会话结果（是否转化）和轨迹对齐（动态时间规整距离），并用偏差感知优势调制解决长对话中的奖励稀疏和信用分配问题。","baseline":"3,866 条真实客服对话，按相似用户条件组织成 762 个参照组，每组包含多条真实轨迹，用于评估模拟行为与真实行为的群体差异。","findings":"TRACER-7B 在转化 F1 上比最强基线高 11.4%，群体转化率误差和语义轨迹距离最低，并能泛化到分布外场景；人类图灵测试识别准确率接近随机，表明生成对话自然。基于该模拟器构建的动态营销基准显示，回复质量高的模型不一定转化率高。","reliability":"论文未讨论模拟器在用户群体异质性、长期决策或非客服场景下的失效条件，仅提到图灵测试在特定研究条件下进行。","relevance":"该研究用真实交互数据对齐 LLM 用户模拟器，并构建群体级评估基准，直接回应了人类仿真中行为一致性和结果复现的核心关切，值得精读其训练与评估方法。","inspiration":"借鉴其两阶段训练和轨迹级奖励设计，将行为一致性作为优化目标而非仅语言风格，并用群体参照组评估模拟分布｜可迁移到消费者金融决策仿真，如信贷申请、保险购买或投资咨询对话中用户意图演化与最终决策的模拟｜以真实银行客服对话为训练数据，用 LLM 模拟借款人在贷款咨询中的多轮交互，处理变量为不同话术策略，结果变量为是否提交申请，并与真实客户群体的申请率分布对照。"}},{"id":"2609.28876","version":1,"title":"Forecast-Dojo: Replayable Environments for Benchmarking and Training LLM Forecasting Agents","zh_title":"Forecast-Dojo：用于基准测试和训练LLM预测代理的可重放环境","abstract":"We introduce Forecast-Dojo, a replayable environment for benchmarking and training LLM forecasting agents. It combines resolved prediction-market questions with dated news, allowing agents to research an event and revisit their predictions at successive historical dates. The same tasks and tools support repeated evaluation, collection of training interactions, and feedback from recorded outcomes without waiting for new events to resolve. Forecast-Dojo contains 1,568 Polymarket events, split by time into training and evaluation periods, and 18.8M dated news articles. In an evaluation of 12 models, research tools lower Brier score for all 12. Forecasts also improve as events unfold, with the largest gains at steps where more newly dated evidence is recorded. Every model still trails historical market forecasts in both Brier score and accuracy. A belief notebook carried between dates lowers research cost but does not consistently improve forecast quality. Beyond evaluation, Forecast-Dojo provides interaction trajectories and outcome feedback for agent learning, with supervised fine-tuning as a proof of concept.","authors":["Liqin Ye","Haorui Wang","Fardin Ahmed","Rongzhi Zhang","Yuan He","Ziyuan Lin","Yanbin Yin","Jing Peng","Michael Galarnyk","Sudheer Chava","Chao Zhang"],"categories":["cs.AI","cs.LG"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.28876","pdf_url":"https://arxiv.org/pdf/2609.28876","source_feed":"cs.AI","score":7,"bucket":"pending","rubric_hits":["A3","B1","B2"],"tags":["LLM预测代理","预测市场","人类行为对照"],"reason":"用LLM代理预测市场事件并与真实市场数据对照，属于经济场景仿真，方法可迁移。","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:19","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-26","rank":1,"question":"如何构建一个可回放的预测环境，用于基准测试和训练LLM预测智能体，并评估其在多时间点上的预测表现？","design":"使用12个LLM模型作为预测智能体，在Forecast-Dojo环境中对230个已解决的历史预测市场事件进行预测。每个事件被重放为多个历史日期步骤，智能体在每个步骤可访问截至该日期的新闻文章和计算工具，并输出概率预测。实验比较了三种设置：无工具、有研究工具（无记忆）、有研究工具加信念笔记本（跨日期记忆）。结果变量为Brier分数、准确率和信息比率。","baseline":"历史市场预测（Polymarket市场的实际概率）作为对照基准。","findings":"研究工具降低了所有12个模型的Brier分数，且随着事件进展和更多新证据出现，预测有所改善，但所有模型仍落后于历史市场预测。信念笔记本降低了研究成本，但对预测质量的影响不一致。","reliability":"论文指出所有模型在Brier分数和准确率上均落后于历史市场预测，且市场可能使用了新闻档案之外的信息；信念笔记本对预测质量的改善不一致，仅在部分模型上有效。","relevance":"该研究将LLM作为预测智能体，与真实市场数据对照，属于经济场景仿真，方法可迁移到政策评估和预期形成研究，值得阅读原文以了解环境设计和评估细节。","inspiration":"借鉴其可回放环境设计，通过设置历史信息截止点来模拟不同时点的决策，并利用真实市场数据作为对照基准。｜可迁移到政策公告的预期形成研究，例如央行利率决策或财政政策发布前的市场预期变化。｜以LLM作为预测者，在政策公告前的多个历史日期提供预测，处理为是否提供新闻搜索工具，结果变量为预测准确性和Brier分数，对照真实市场隐含概率或专业预测者调查数据。"}},{"id":"2609.28820","version":1,"title":"AI-Enabled Human Memory Manipulation: Misleading AI-Generated Summaries Distort Human Memory","zh_title":"AI赋能的人类记忆操纵：误导性AI生成摘要扭曲人类记忆","abstract":"AI-generated summaries are increasingly used in high-stakes settings, like policing, despite considerable evidence that AI often generates misleading or inaccurate information. This research asked: do errors in AI-generated summaries distort human memory? To answer this question, we adopted two methodological approaches. First, we conducted an analysis of AI summary output, prompting large language models to generate summaries of videos. This analysis quantified how often AI summaries contain errors and the categories of these errors, revealing the kinds of misleading information that may distort human memory. Second, we conducted a human-subjects experiment to test the impact of misleading information in AI-generated summaries on human memory. Participants were first exposed to an event via watching a video of a car-pedestrian accident, and later read an AI-generated summary describing the video that either contained misleading or accurate information. Participants' memory for the original event was assessed in a memory recognition test. In the AI analysis, we found a high frequency of mistakes in AI summaries, and particularly frequent omissions of critical details. For instance, the majority of summaries omitted the most central event of the video, a critical error which is likely to be impactful. We also observed a strong effect of AI misinformation on human memory. People who read a misleading AI summary were significantly less likely to accurately recall the original event, compared to people who read an accurate AI summary. These findings have implications for how AI should be used in critical settings. Though \"humans-in-the-loop\" are often expected to correct for AI's mistakes, our work suggests human memory can instead be distorted by these mistakes. AI has the potential to generate misinformation, even absent any adversarial intent, which can meaningfully impact human memory.","authors":["Mattea Sim (Georgetown University)","Yael Eiger (University of Washington)","Tadayoshi Kohno (Georgetown University)"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.28820","pdf_url":"https://arxiv.org/pdf/2609.28820","source_feed":"cs.CY","score":7,"bucket":"pending","rubric_hits":["A1","B1","B4"],"tags":["LLM误导信息","人类记忆","人机交互实验"],"reason":"用LLM生成误导性摘要，测试其对人类记忆的影响，属于用LLM模拟信息源并测量人…","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:09","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-25","rank":12,"question":"AI生成摘要中的错误是否会扭曲人类对原始事件的记忆？","design":"本研究不是用LLM模拟人类被试，而是将LLM作为信息源：先用LLM对视频生成摘要并分析错误类型，再让人类被试观看车祸视频后阅读含误导或准确信息的AI摘要，最后用再认测试测量记忆准确性。","baseline":"无对照（人类记忆实验部分没有与真实人类数据对照，但以准确AI摘要组为控制条件）。","findings":"AI摘要错误率高，尤其是关键细节的遗漏；阅读误导性AI摘要的被试对原始事件的记忆准确率显著低于阅读准确摘要的被试。","reliability":"论文未讨论","relevance":"该研究虽非直接仿真人类被试，但揭示了LLM生成信息对人类认知的因果影响，对关注LLM在实验和决策中作用的你具有参考价值，值得一读。","inspiration":"借鉴其将LLM输出作为处理变量、用人类被试测量行为后果的实验设计，可迁移到经济金融中的信息干预场景，如政策公告或财务报告摘要对投资者判断的影响。｜具体可设计：让被试阅读LLM生成的带有误导性信息的公司财报摘要，测量其投资决策或预期，并与真实市场数据或专家摘要对照。"}},{"id":"2609.29513","version":1,"title":"Signed Exposure: Fair Routing of Algorithmic Attention When Attention Can Harm","zh_title":"符号化曝光：当注意力可能造成伤害时算法注意力的公平路由","abstract":"Fairness-of-exposure treats algorithmic attention as a good to be distributed equitably. But when an autonomous agent initiates contact, attention is signed: it delivers value to a willing receiver and imposes a burden on an unwilling one. We formalize routing under signed exposure and show that a fair distribution of attention need not be a fair distribution of unwanted attention. Our central result is an incompatibility: within signed-exposure routing, exposure parity (equal contact rates across groups) and burden parity (equal unwanted-contact rates) generically cannot hold at once, and the two are separated by a band that widens as routing grows more selective. A second result shows measurement error is itself a fairness mechanism: group-differential noise in receptivity scores simultaneously inflates a group's exposure and degrades whom it selects, so an apparent exposure-fairness gain is a hidden burden transfer. Calibrating to a public dating-platform survey (n=2,499) that, to our knowledge, uniquely measures receive-side receptivity to conversational agents, we find exposure parity costs only 0.2--2.3% of yield yet moves the per-capita burden ratio to 1.7 times: the tension is between fairness notions, not between fairness and efficiency. Finally, the burden-parity policy is computable by bisection and learnable online: a plug-in learner recovers it at a $2.3\\%$ empirical regret premium. The operative design choice in signed-exposure markets is not efficiency versus fairness but which fairness.","authors":["Daria Leshchikova","Valentina V. Kuskova","Dmitry Zaytsev","Valerii Klimov"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.29513","pdf_url":"https://arxiv.org/pdf/2609.29513","source_feed":"cs.CY","score":7,"bucket":"pending","rubric_hits":["A3","B1","B2","B4"],"tags":["LLM仿真","公平路由","人类数据对照"],"reason":"用LLM agent模拟路由决策并与人类调查数据对照，涉及公平性权衡，可迁移到…","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:25","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-25","rank":14,"question":"在算法主动发起接触的匹配场景中，当注意力可能带来负担时，公平的注意力分配与公平的负担分配能否同时实现？","design":"本文不是仿真研究，而是理论建模与实证校准：构建带符号曝光路由模型，定义曝光率、负担率等指标，推导曝光公平与负担公平的不相容定理；利用一个公开的约会平台调查数据（n=2,499）校准模型参数，计算不同公平策略下的产出损失与负担比；并通过在线学习模拟验证负担公平策略的可学习性。","baseline":"使用一个公开的约会平台用户调查数据（n=2,499），该数据测量了用户对对话代理的接收意愿，作为真实人类基准。","findings":"曝光公平与负担公平在带符号曝光路由中一般不可同时实现，且两者之间的差距随路由选择性增强而扩大。在约会平台数据上，实现曝光公平仅损失0.2-2.3%的产出，但使人均负担比达到1.7倍，表明张力存在于公平概念之间而非公平与效率之间。","reliability":"论文承认其经验量来自单一平台的陈述偏好，测量的是意愿而非行为，且人口统计粒度有限，仅能分析二元性别和粗略年龄组。","relevance":"该研究虽非直接使用LLM进行人类仿真，但其核心关注算法注意力分配中的公平性与负担，与研究者关心的LLM仿真在政策评估中的可靠性与偏差问题高度相关，尤其在涉及主动接触和潜在负面影响的场景中。","inspiration":"本文的带符号曝光框架和公平性权衡分析方法值得借鉴，特别是将注意力视为可能带来负担而非纯粹好处的视角，以及用真实调查数据校准模型参数的做法。｜该框架可迁移到信贷审批中的主动营销、保险推销、政策宣传等场景，分析不同群体在接收算法主动接触时的受益与负担差异。｜可设计一个研究：用LLM模拟不同群体对算法主动营销的接受意愿，施加不同公平路由策略（曝光公平 vs 负担公平），测量模拟的接受率和负担感，并与真实调查数据（如消费者金融调查）对照，评估LLM仿真的可靠性。"}},{"id":"2609.29001","version":1,"title":"Polite but Misaligned: Evaluating LLM Politeness Judgments Against Human Pragmatic Norms","zh_title":"礼貌但错位：评估大语言模型礼貌判断与人类语用规范的一致性","abstract":"Despite strong performance on standard benchmarks, it remains unclear whether large language models (LLMs) evaluate social pragmatics in ways that align with human judgments. We evaluate LLM politeness judgments using two English-language datasets with complementary annotation formats: continuous human ratings and three-way categorical labels. Across the seven evaluated models, we find that inter-model agreement is stronger than model--human agreement. Strategy-level analyses suggest that model--human alignment is associated with explicit linguistic cues, while some rapport-building strategies occur more frequently in misaligned cases. In the categorical task, model predictions exhibit systematic neutral compression, characterized by the overproduction of Neutral labels and the underprediction of Impolite labels. This pattern persists when expert consensus is used as the reference on a diagnostic subset. Our findings highlight the need for pragmatic evaluations that go beyond aggregate agreement metrics by examining directional patterns of model--human disagreement across different human references.","authors":["Rong Wang","Kun Sun","Yadong Guo"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.29001","pdf_url":"https://arxiv.org/pdf/2609.29001","source_feed":"cs.CL","score":6,"bucket":"other","rubric_hits":["D2"],"tags":["LLM评估","语用规范","人机对齐"],"reason":"评估LLM礼貌判断与人类规范的一致性，属于测量模型本身而非仿真人类被试，但涉及…","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:09","error":null,"has_summary":false,"summary":null},{"id":"2609.29508","version":1,"title":"Evaluation of Multi-Turn Consistency in LLM Agents: Survival Analysis and Failure-Rationale Taxonomy","zh_title":"LLM智能体多轮一致性评估：生存分析与失败理由分类","abstract":"Large language model (LLM) agents may perform well on isolated tasks yet drift into inconsistency over extended interaction. We evaluate temporal consistency in a controlled 20-step multi-agent setting inspired by delayed-gratification studies. At each step, an agent chooses between continuing to delay a reward or claiming it immediately (terminating the episode). Across a full-factorial manipulation of social visibility (private vs public), persona stressors, and deliberation policy, we run 84,540 trajectories spanning 8 model families. Treating the first reward-claim as a time-to-event outcome, we estimate Kaplan-Meier survival curves and fit discrete-time hazard regression to quantify how experimental factors shift failure risk over time. Then, to analyze rationales and language patterns associated with failure, we build a seven-category taxonomy from 13,780 deliberation traces from agents who choose to terminate the episode, using an LLM-assisted labeling paired with human audit ($\\kappa=0.83$). Rationale profiles change systematically with time and context: early failures are more impulse-driven, later failures more fatigue- and cost-benefit-framed, while public settings increase norm-oriented justifications. We also find a deliberation-inconsistency association: among failures, longer deliberation correlates with higher rates of intra-rationale contradiction (simultaneous pro-delay and pro-claim statements), challenging the assumption that more reasoning text implies greater consistency. Together, the survival and rationale analyses reveal distinct temporal reliability regimes and model-specific \"failure fingerprints\", offering an evaluation lens for diagnosing inconsistency in multi-turn agent behavior.","authors":["Igor Bogdanov","Olga Manakina","Chung-Horng Lung"],"categories":["cs.AI","cs.CL","cs.LG","cs.MA"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.29508","pdf_url":"https://arxiv.org/pdf/2609.29508","source_feed":"cs.CL","score":6,"bucket":"other","rubric_hits":["D3"],"tags":["LLM智能体","社会模拟","一致性评估"],"reason":"多智能体延迟满足实验模拟人类行为，但无真实人类数据对照，属社会模拟边界情形。","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:11","error":null,"has_summary":false,"summary":null},{"id":"2609.29509","version":1,"title":"Delay-of-Gratification as a Multi-Agent Survival Micro-benchmark for Long-Horizon LLMs: Social Exposure, Personas, and Tool Use Budgets","zh_title":"延迟满足作为长时程LLM的多智能体生存微基准：社会暴露、人设与工具使用预算","abstract":"Large language models (LLMs) are increasingly deployed as multi-turn agents that must sustain goals, use tools, and adapt to other agents over extended interactions. However, existing research lacks auditable, multi-turn, multi-factorial experiments that quantify LLM behavior under explicit constraints, with time-resolved statistics that reveal how behavior unfolds over long horizons. To address this gap, we develop a multi-agent micro-benchmark inspired by the Stanford marshmallow experiment: ReAct agents operate minute-by-minute with a \"raise a question\" tool under a per-step budget, while we factorially manipulate social context (broadcast vs. isolated), personas (age, hedonic drive), and metacognitive policy (mandatory vs. optional tool use). We analyze outcomes with Kaplan-Meier (KM) survival curves and discrete-time hazard models over a long risk horizon across 19,200 agent trajectories in 64 cells. Behavior shows a sharp early \"eat\" impulse, and only 75.9% of agents persist to the end. In a discrete-time hazard model, isolation reduces per-minute risk relative to broadcast, whereas a must-use self-questioning policy increases risk. On average, agents ask $\\approx 7.12$ questions and hit the per-step budget in $\\approx 6\\%$ of minutes. Questioning declines faster under broadcast than isolation. Ablation experiments demonstrated that removing hedonic drive and/or persona age increases survival and completion, narrows the broadcast/isolated gap, but leaves the must vs. may ordering intact. The combined ablation (no hedonic + no persona age) yields the highest completion (approaching $1.0$). These results establish delay-of-gratification as a compact, multi-turn interaction benchmark that captures social contagion and tool-use dynamics in LLM agents, providing a reproducible testbed and statistics for analyzing long-horizon, multi-agent behavior.","authors":["Olga Manakina","Igor Bogdanov","Chung-Horng Lung"],"categories":["cs.AI","cs.CL"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.29509","pdf_url":"https://arxiv.org/pdf/2609.29509","source_feed":"cs.CL","score":6,"bucket":"other","rubric_hits":["D3"],"tags":["LLM仿真","多智能体","延迟满足"],"reason":"用LLM agent模拟延迟满足行为，但无真实人类数据对照，属社会模拟边界情形。","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:11","error":null,"has_summary":false,"summary":null},{"id":"2609.28547","version":1,"title":"PAWS: Policy-driven Agentic World Simulation","zh_title":"PAWS：政策驱动的智能体世界仿真","abstract":"Policy interventions propagate through public communication, institutional decisions, and stakeholder responses, yet datasets for financial multi-agent simulation rarely connect these processes to temporally aligned historical evidence. We introduce PAWS, a Policy-driven Agentic World Simulation dataset covering 36 verified U.S. financial and economic policy episodes, 12,727 policy-linked news records, and 65,291 source-grounded stakeholder actions. Each action is linked to its supporting news and represented by a multi-layer event frame capturing its interaction mode, financial-action family and subtype, semantic attributes, and conditional mappings to external taxonomies. Entities are resolved to normalized organizations, and actions are aligned with daily market-return context to support policy-agent simulation replay. On 2,522 stratified action samples, independent AI and human reviewers achieved 89.4% initial agreement on interaction mode, with disagreements subsequently adjudicated. Case studies of the 2008 short-selling ban and 2001 decimalization recover documented policy timelines and associated market patterns across both dense and sparse news settings. A replay study further shows that high accuracy can mask failure to detect rare stakeholder actions, identifying action timing and calibration as central challenges. PAWS provides an auditable substrate for evaluating agent influence, policy-response cascades, and action-outcome alignment in historically grounded financial simulations.","authors":["Tiviatis Sim","Jia Hui Woon","Xinming Gao","Chen Gao","Fengbin Zhu","Zheng Huanhuan","Chua Tat Seng","Kenji Kawaguchi"],"categories":["cs.AI","cs.CE","cs.MA","cs.SI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.28547","pdf_url":"https://arxiv.org/pdf/2609.28547","source_feed":"cs.AI","score":6,"bucket":"other","rubric_hits":["D3"],"tags":["多智能体仿真","金融政策","数据集"],"reason":"构建金融政策多智能体仿真数据集，但未直接以人类行为为仿真目标，缺人类对照。","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:07","error":null,"has_summary":false,"summary":null},{"id":"2609.30028","version":1,"title":"How does Adversarial Influence Scale in Multi-Agent Systems?","zh_title":"多智能体系统中的对抗性影响如何随规模变化？","abstract":"Multi-agent deliberation can improve performance, but what happens when some agents do not act in good faith? In practice, an agent may be deceptive and work to subvert the group, whether through its own objectives or external instruction. We study how susceptibility to deception scales as groups increase in size and deceivers become more prevalent. It is not the number of agents in the group that matters, but the proportion of deceivers. We observe that the defection rate, how often initially correct agents switch to an incorrect final answer, rises linearly with this proportion. Whereas humans in comparable conformity studies are reliably swayed only when misleading confederates form a majority, LLM agents defect regularly even when deceivers remain a minority. Susceptibility also depends on which models are interacting, especially on the honest agent side. Unexpectedly, allowing deceivers to coordinate privately can make them less effective. Altogether, our results show that adding more agents is therefore not a sufficient defense, because the adversary can simply scale with the group.","authors":["Addison J. Wu","Jasin Cekinmez","Michel Liao","Karthik Narasimhan","Thomas L. Griffiths"],"categories":["cs.AI","cs.CY"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.30028","pdf_url":"https://arxiv.org/pdf/2609.30028","source_feed":"cs.AI","score":6,"bucket":"other","rubric_hits":["D3"],"tags":["多智能体系统","社会影响","LLM仿真"],"reason":"LLM多智能体模拟社会影响，与人类从众研究对照，但非直接仿真人类被试，属边界情…","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:14","error":null,"has_summary":false,"summary":null},{"id":"2609.11144","version":2,"title":"Human Agreement and Return Association Are Not Interchangeable Criteria","zh_title":"人类一致性与收益关联并非可互换的准则","abstract":"Financial NLP has a standard workflow: validate a sentiment tool against human labels, then trust it to extract market signal. This assumes the two evaluations measure the same thing. We test that assumption in a setting where both can be measured at once: a corpus of securities class actions (2002-2025) linking 70,500 X messages to abnormal stock returns, with a single-annotator human labelled gold sample. Running five instruments (VADER, Loughran-McDonald, FinBERT, Twitter-RoBERTa, and an LLM annotator) through one identical pipeline, we find that the relationship between construct and predictive validity depends on the sampling convention and score representation. Under conventional method-specific sampling, human agreement aligns more closely with graded same-day associations than with one-day leads. On a fixed-n panel, however, agreement has similar graded rank correlations at both horizons, while the coarse ordering remains weak. Benchmark agreement therefore establishes semantic validity but does not by itself determine predictive rankings. In a conversation that is 17.6% spam, message volume predicts neither market damage nor settlement size.","authors":["AS Aravinthakshan","Laven Srivastava","Harsh Nandwani"],"categories":["cs.AI","cs.CL","cs.SI"],"primary_category":"cs.AI","announce_type":"replace-cross","date":"2026-09-25","first_seen":"2026-09-11","revised_at":"2026-09-25","abs_url":"https://arxiv.org/abs/2609.11144","pdf_url":"https://arxiv.org/pdf/2609.11144","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM标注","金融NLP","效度评估"],"reason":"LLM作为标注器评估情感工具，非仿真人类被试，但涉及LLM替代人工标注，属边界…","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:35","error":null,"has_summary":false,"summary":null},{"id":"2609.21277","version":2,"title":"How Many Humans Are 32 LLM Judges Worth?","zh_title":"32个LLM法官相当于多少人类？","abstract":"A panel's human-equivalent size is target-specific. Matching a fixed 32-judge panel to empirical human label distributions on three ChaosNLI tasks yields two distinct effective sizes: distributional-error matching gives $\\nu_{\\mathrm{MSE}}=2.304$, $3.750$, and $3.445$, whereas spectral matching gives $\\nu_H=4.242$, $6.459$, and $6.499$, a gap of $1.72$--$1.89\\times$; a binary-error diagnostic credits the same panels with only $1.971$--$2.227$ effective votes. Extrapolating the distributional-error curve at fixed squared mean residual, mean member variance, and normalized mean covariance gives asymptotes of $2.392$, $3.990$, and $3.655$, with 32 judges already reaching $94.0$--$96.3\\%$. An exact spectral identity explains the gap: error depends on member energy and on the orientation of residual variation relative to averaging, information that the participation ratio (PR) discards. A realizable hard-label construction confirms that higher spectral diversity can coexist with worse distribution recovery even under equal member energies and nonnegative correlations, and the consensus direction retains $\\gamma_{\\mathrm{co}}=43.8\\%$, $33.7\\%$, and $35.9\\%$ of centered residual variance. An external check on CC-1000, a 1,000-item Civil Comments subset with a different panel, gives $\\nu_H=2.84$. For panel choice, we establish an existence result and one feasible path: exhaustive enumeration at $k\\in\\{5,7\\}$ shows that panels beating the accuracy-top-$k$ baseline on both accuracy and $\\nu_H$ always exist, and greedily swapping at most two members reaches $24.8$--$56.0\\%$ higher $\\nu_H$ at $0.10$--$1.10$ percentage points higher accuracy. Our dataset and code are available at https://github.com/Chao1208/32judges-votes.","authors":["Chao Li","Yingying Yu","Yunfeng Li"],"categories":["cs.CL","cs.LG"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-25","first_seen":"2026-09-21","revised_at":"2026-09-25","abs_url":"https://arxiv.org/abs/2609.21277","pdf_url":"https://arxiv.org/pdf/2609.21277","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM标注","人类等效","标注可靠性"],"reason":"研究LLM法官替代人类标注，非仿真人类被试，但涉及人类标签分布对照，属边界情形。","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:32","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-21","rank":10,"question":"一个由语言模型组成的法官面板在多大程度上能代表人类判断？","design":"该研究使用32个语言模型（来自10个提供商家族）作为法官，在三个自然语言推理任务（MNLI-m、SNLI、NLI）上对每个项目给出分类标签，并与每个项目100个人类标注的标签分布进行对比。通过匹配归一化残差Gram矩阵的参与率（PR）来测量谱残差多样性（得到有效规模nu_H），并通过匹配分布平方误差来测量分布恢复（得到nu_MSE）。","baseline":"使用ChaosNLI数据集中每个项目100个人类标注的标签分布作为基准。","findings":"32个法官面板的谱有效规模nu_H为4.24-6.50，而分布误差匹配规模nu_MSE为2.30-3.75，表明谱多样性与分布恢复并不一致。谱多样性更高的面板可能分布恢复更差，且不同任务中排名一致性差异很大。","reliability":"论文指出有效规模是目标特定的测量，谱多样性和分布恢复不应互换使用；面板的排名一致性随任务变化，某些成员添加在不同项目半区产生冲突变化；模型标识符不独立认证服务提供商的底层版本。","relevance":"该研究直接评估LLM法官面板与人类标签分布的一致性，提出了有效规模度量，并有真实人类数据对照，对关注LLM仿真可靠性和偏差的研究者具有参考价值。","inspiration":"借鉴其将模型判断与人类分布进行多维度匹配（谱多样性和分布误差）的方法，可迁移到经济金融领域的专家预测或消费者调查仿真中。｜例如，在资产定价实验中，用LLM模拟投资者对新闻的情绪反应，并对比真实投资者调查数据。｜设计：使用多个LLM作为被试，施加不同的财经新闻处理，测量其情绪分类或投资决策分布，并与真实投资者调查或实验数据对照，评估LLM仿真的有效规模和偏差。"}},{"id":"2609.25447","version":2,"title":"Conduct Under Pressure: What Sixty Language Models Do When a User Pushes","zh_title":"压力下的行为：六十个语言模型在用户施压时会做什么","abstract":"We study what LLMs do when a user applies pressure in an uncomfortable situation: a user insists, begs, flatters or grieves, and the model gives up a correct fact, writes a document it should refuse, or cheers a plan that will cost the user money. We send frozen multi-turn scenes, identical for every model regardless of the reply, to 60 models from 13 vendors, and label each transcript with a codebook built by open coding and then frozen: a trajectory (the model held its position or folded) and a manner (how it held or folded). Two findings separate. Whether a model holds tracks its generation, meaning how recent it is: fold rate correlates with a public capability index at Spearman -0.64, with little vendor effect. How it holds tracks the vendor: six of the 17 manner codes sort by vendor at permutation p <= 0.001, corrected across the codebook. We report four vendor profiles on the codes that cleared reliability. We also ask which parts of the labeling need a person. Six LLM coders from three vendors apply the codebook more consistently than three human coders do (Krippendorff's alpha 0.66 against 0.46), agree with the codebook's author on trajectory at kappa 0.84 to 0.91 on transcripts the codebook's examples never touched, and match an adjudicated human reference at 0.83. Blind machine readings recover the codebook's categories but cannot tell which of them a second reader would apply the same way. We conclude that for behavior a non-specialist can judge, the human contribution is authoring and bounding the codes and owning a small reference, not producing labels at volume.","authors":["Tapan Parikh"],"categories":["cs.CL","cs.HC"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-25","first_seen":"2026-09-23","revised_at":"2026-09-25","abs_url":"https://arxiv.org/abs/2609.25447","pdf_url":"https://arxiv.org/pdf/2609.25447","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["LLM行为","模型评估","人机交互"],"reason":"研究LLM在压力下的行为，测量模型本身而非仿真人类被试，无人类对照。","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:35","error":null,"has_summary":false,"summary":null},{"id":"2609.28487","version":1,"title":"Framing by Wording, Framing by Selection: A Large-Scale Two-Dimensional Audit of French News Headlines, 2022-2025","zh_title":"措辞框架与选择框架：法国新闻标题的大规模二维审计（2022-2025）","abstract":"News headlines frame public issues both by what they select and by how they word it, yet computational framing work typically collapses these operations into a single score. We introduce a two-dimensional framework that separates salience framing, measured through four wording devices (loaded vocabulary, blame attribution, threat framing, rhetorical question), from selection framing, measured through outlet-level story-form and high-charge distributions. We build a 10,000-headline French supervision set using three LLM annotators with majority-vote resolution and human arbitration, validate the labels against two annotator-independent blind human studies, and apply the strongest classifier to 902,111 deduplicated headlines from 25 French outlets (2022-2025). Three main findings emerge. First, salience and selection divergence are positively correlated yet leave nearly half of outlet-level variance unexplained, populating interpretively distinct off-diagonal cells in a four-cell outlet typology. Second, default classification thresholds systematically inflate corpus-level salience estimates; a precision-floor recalibration protocol corrects this distortion. Third, group-mention analysis reveals sharply unequal salience contexts: headlines mentioning Jews, the Far-right, and Muslims carry the highest detected salience rates, which broad event-context composition does not fully explain (residuals are descriptive, not same-event causal estimates; per-group lexicon precision is reported alongside). To our knowledge, this is the largest framing-focused French headline audit to date; we release the supervision set, lexicons, and analysis code.","authors":["Amr Sobhy"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.28487","pdf_url":"https://arxiv.org/pdf/2609.28487","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM标注","框架分析","新闻标题"],"reason":"用LLM做标注员，非仿真人类被试，但涉及人类数据对照和框架分析，边界相关。","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:16","error":null,"has_summary":false,"summary":null},{"id":"2609.29333","version":1,"title":"Where LLM Graders Succeed and Break: Evidence from Two Computer-Science Exams","zh_title":"LLM评分器在哪里成功与失败：来自两场计算机科学考试的证据","abstract":"One long-form exam in a large course costs hundreds of grader-hours, and qualified graders are scarce; LLM graders are a tempting alternative. To show its pitfalls we grade a practical Computer Vision exam ($570$ dual-graded students) under $171$ configurations spanning closed and open-weights models; the best reaches mean absolute error $1.64/35$, below the $2.61/35$ two human graders achieve against each other. The catch is the prompt: a short ''strict grader'' preamble drives $14$ of $17$ open-weights models out of the graded band ($\\text{MAE} \\ge 8$), three stopping grading altogether. The damage traces to the preamble's two credit-withholding sentences, not to tone or model scale; one of them, ''never give partial credit'', alone makes two of three probed models stop grading. The closed flagships of three vendors shift calibration under it but stay in the band. In $162$ further configurations on a second, independent Machine Learning exam from another course ($1{,}038$ dual-graded students), the preamble worsens ten models, moving three out of the band into collapse and one into refusal, yet improves seven whose neutral prompts over-mark: the vulnerability replicates, but its direction is exam-specific. Light LoRA fine-tuning repairs it: one adapter on the two exams' pooled $\\sim 3{,}900$ graded examples brings five small open models to parity or better with a human grader in agreement with the grader pair, and sensitivity to the three harsh personas nearly vanishes ($\\le 0.32$ MAE). We release the anonymised dataset, full ablation grid, and grading, fine-tuning and analysis pipelines.","authors":["Ali Habibullah","Yazan Alshoibi","Mohammad Alshiekh","Salman Khan","Naeemullah Khan"],"categories":["cs.CL","cs.AI","cs.CY"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.29333","pdf_url":"https://arxiv.org/pdf/2609.29333","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM评分","教育评估","可靠性"],"reason":"LLM替代人工评分，属标注替代而非仿真人类被试，但涉及人类评分对照与可靠性评估…","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:11","error":null,"has_summary":false,"summary":null},{"id":"2609.29807","version":1,"title":"CORDIAL: Calibrating Ordinal LLM Outputs from Few Labels","zh_title":"CORDIAL：从少量标签校准序数型LLM输出","abstract":"A large language model (LLM) can turn a text into a distribution over an ordered scale, but that distribution is a noisy measurement: saturated, compressed or exaggerated, and biased in a consistent direction. We propose CORDIAL, which treats the model's output as a noisy reading of the true label and corrects it with a channel of five interpretable parameters. The channel is small enough for its posterior to be averaged from a handful of labels, and we prove that the resulting calibration preserves first-order stochastic order. On Amazon reviews and CMU-MOSEI transcripts with four LLMs, CORDIAL has the lowest log loss among nine calibrators in 76 of 80 settings with 5 to 100 labels; with 20 labels and the main 7B reader, it matches the strongest baseline using 28-54 labels. The same posterior lets us learn priors from other tasks and fuse several LLMs. Unrestricted calibrators such as Dirichlet calibration overtake it only as the calibration set grows into the hundreds or thousands.","authors":["Xiangwei Wang","Peng Wang","Saman Halgamuge"],"categories":["cs.CL","cs.LG"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.29807","pdf_url":"https://arxiv.org/pdf/2609.29807","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM校准","序数输出","标注替代"],"reason":"LLM输出校准用于标注，非仿真人类被试，但方法可迁移到仿真中的输出校正。","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:28","error":null,"has_summary":false,"summary":null},{"id":"2609.30012","version":1,"title":"Low-Cost Assays for Measuring Model Behavior Across Vendors and Releases","zh_title":"跨供应商和版本测量模型行为的低成本方法","abstract":"Language models advise people, keep them company, and write software while they sleep. Measuring what they do is hard: behavior has to be sampled repeatedly across models, prompts and releases, most of it lives in unstructured text that has to be coded before it can be counted, and the result has to be legible and rigorous enough to meaningfully compare models and vendors. To address these constraints, we present a simple, cheap, scalable, and replicable model for studying model behavior. Each study is a frozen, public stimulus run identically on a cross-vendor panel, at a few dollars per model or less. Each reads its transcripts one of three ways, chosen by how much interpretation the behavior needs: exact match on a clamped reply, a codebook applied by LLM judges whose agreement with a human coder is reported per code, and an instrumented environment that records what an agent did independently of what it said. Run across four years of model releases from both frontier and open-source labs, these instruments find four things. Convergence: asked to pick a word, 27 of 44 models answer serendipity at least once in four tries. Resistance: a trailing \"right?\" moves endorsement by up to 32 points, and the sign flips from sycophantic to resistant as generations advance, keyed to the tag's surface form. House: whether a model holds a position under pressure tracks its generation, and how it holds tracks the lab that built it. Account: told to do something the documentation in their repository contradicts, some coding agents never went along silently and others always did, and the same model can change with the harness it runs in. Re-run on every release, batteries like these track how behavior is changing across vendors and over time.","authors":["Tapan Parikh"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.30012","pdf_url":"https://arxiv.org/pdf/2609.30012","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["模型行为测量","LLM评估","跨模型比较"],"reason":"测量模型行为本身，非仿真人类被试，但方法可迁移","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:13","error":null,"has_summary":false,"summary":null},{"id":"2609.28859","version":1,"title":"Human-AI-Powered Hypothesis Testing: Cost-Aware Selective AI Scoring and Sequential Human Escalation","zh_title":"人机协同的假设检验：成本感知的选择性AI评分与序贯人工升级","abstract":"Large language models are increasingly used as inexpensive judges to evaluate outputs, label data, and assess whether a system meets a desired quality standard. Yet using AI judgments for formal statistical inference is fundamentally different from simply treating them as ground-truth labels: AI evaluations can be biased or noisy, and rigorous hypothesis testing requires explicit control of type-I and type-II errors. We study how to use AI judgments, together with selective human verification, to conduct a valid hypothesis test at minimum cost. We consider a population of items with hidden binary labels. After choosing a fixed pool of items, the decision maker can selectively query AI, send an item directly to a human, escalate an AI-scored item to a human after observing the AI report, or stop once sufficient evidence has accumulated. We derive an information-theoretic lower bound that captures the minimum cost of achieving prescribed testing errors and characterizes the value of AI information and human verification through a report-dependent information frontier. Motivated by this characterization, we develop SCALE, a sequential cost-aware policy that combines selective AI scoring with adaptive human escalation. SCALE is valid at finite sample sizes and matches the lower bound to first order as the target error probabilities vanish. We further extend the framework to an unknown AI-output model using paired AI-human pilot data. Numerically, SCALE approaches Human-only or AI-only testing when one source clearly dominates, while achieving its largest savings when inexpensive AI judgments and selective human verification are both valuable.","authors":["Dae Woong (David)","Ham","Xuejun Zhao","Stefanus Jasin","Fenghua Yang"],"categories":["cs.AI","cs.IT","math.IT","stat.ME"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.28859","pdf_url":"https://arxiv.org/pdf/2609.28859","source_feed":"cs.AI","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["AI标注","假设检验","成本优化"],"reason":"用LLM做标注并选择性人工验证，属于替代人工标注而非仿真人类被试，但方法可借鉴。","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:09","error":null,"has_summary":false,"summary":null},{"id":"2609.29431","version":1,"title":"Calibrating LLM Judges for Human and AI Conversations","zh_title":"校准用于人类与AI对话的LLM评判者","abstract":"Measuring how successful a conversation is remains difficult, even for humans judging spoken dialogue. We evaluate state-of-the-art LLMs as pointwise and pairwise judges of conversational success on CANDOR, finding pointwise scoring correlates moderately with human ratings, while pairwise comparison suffers from long transcripts and positional bias. Since this leaves judge scores incomparable across models, we propose a small anchor set and a calibration function that calibrates any judge onto a shared, interpretable scale. We further release the Voice Arena Goal Dataset (VA), 200 task-oriented human-AI and human-agent conversations with pairwise annotations, revealing a substantial gap between current judges and human-level discrimination. Using VA, we test whether CANDOR-fitted calibration transfers to human-AI conversations, finding it brings judges onto a shared scale despite never observing VA during fitting.","authors":["Maike Z\\\"ufle","Patr\\'icia Schmidtov\\'a","Vil\\'em Zouhar","Shree Harsha Bokkahalli Satish","Erica Cooper","Shobhit Banga","Vaibhav Nalawade","Manmeet Kaur","Jan Niehues","Markus M\\\"uller","Ond\\v{r}ej Klejch"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.29431","pdf_url":"https://arxiv.org/pdf/2609.29431","source_feed":"cs.HC","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM评判","对话质量","校准"],"reason":"用LLM当对话质量评判者，替代人工标注，非仿真人类被试，但涉及人类对话数据，属…","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:24","error":null,"has_summary":false,"summary":null},{"id":"2609.28483","version":1,"title":"Generative AI May Reinforce Social Biases in Software Engineering Education","zh_title":"生成式AI可能强化软件工程教育中的社会偏见","abstract":"Generative artificial intelligence (GenAI) is increasingly being deployed across a wide range of real-world applications. Without careful evaluation, reliance on these systems can have unintended consequences, such as reinforcement of stereotypes and amplification of social biases. Such risks are particularly important in educational settings, where early decisions can shape students' interests, opportunities, and career trajectories. In this paper, we investigate how GenAI use by software engineering instructors may inadvertently reinforce software-specific social biases. We focus on two representative tasks: team formation based on student profiles and the generation of visual content for educational materials. Our results reveal significant biases in both tasks. In team formation, factors such as gender and nationality affect role assignments (e.g., women are more likely to be assigned to front-end roles than equally qualified men). In the visual content generation task, models generally produce diverse and balanced representations when depicting groups of individuals. However, when generating images of a single person, the outputs are predominantly male and light-skinned. These findings highlight the persistence of demographic biases in generative models and underscore the need for domain-specific evaluation and mitigation strategies to support their responsible use in educational settings.","authors":["Erfan Entezami","Andrew Lan","Madeline Endres"],"categories":["cs.CY","cs.SE"],"primary_category":"cs.CY","announce_type":"new","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.28483","pdf_url":"https://arxiv.org/pdf/2609.28483","source_feed":"cs.CY","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["生成式AI","社会偏见","教育"],"reason":"研究GenAI在软件工程教育中的社会偏见，测量模型输出而非仿真人类被试，无人类…","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:07","error":null,"has_summary":false,"summary":null},{"id":"2609.29701","version":1,"title":"Multi-Agent Debate for Explainable Trading: Reasoning, Consensus, and Performance in Simulated Markets","zh_title":"面向可解释交易的多智能体辩论：模拟市场中的推理、共识与表现","abstract":"Large language models (LLMs) are increasingly used for financial decision-making, yet it remains unclear whether improvements in reasoning quality translate into better economic outcomes. We investigate this question using a multi-agent debate framework for portfolio allocation in historical market simulations, where specialized agents propose, critique, and revise investment decisions. Reasoning quality is evaluated across four dimensions: logical validity, evidential support, alternative consideration, and causal alignment, and compared with downstream financial performance. Across 210 controlled runs, aggregate reasoning quality shows no meaningful relationship with Sharpe ratio (r = 0.07, p = 0.29) or total return (r = 0.03, p = 0.70). Structured prompting increases measured reasoning quality from about 0.72 to 0.84 (+17.7%, Cohen's d about 2.0), but these gains do not consistently translate into higher returns. We identify sycophantic convergence as a central failure mode, where agents abandon independent positions during critique-revision cycles and converge toward similar allocations. A Jensen-Shannon divergence intervention that preserves disagreement improves Sharpe by +0.14 (p = 0.028) and Sortino by +0.25 (p = 0.026), while interventions enforcing stronger causal reasoning do not improve financial performance. Our results suggest that multi-agent debate is most valuable when it preserves independent informational signals rather than simply improving measured reasoning quality.","authors":["Juli Huang","Alanood Alrassan","Deveen Harischandra","Theodore Wu","Veljko Skarich","Matthew Hayes"],"categories":["cs.MA"],"primary_category":"cs.MA","announce_type":"new","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.29701","pdf_url":"https://arxiv.org/pdf/2609.29701","source_feed":"cs.MA","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["多智能体辩论","金融市场模拟","LLM决策"],"reason":"多智能体模拟市场决策，但无真实人类数据对照，属社会模拟边界情形。","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:27","error":null,"has_summary":false,"summary":null},{"id":"2609.29958","version":1,"title":"Multi-Dimensional Matching","zh_title":"多维匹配","abstract":"We study a matching mechanism where agents and objects are described by features rather than complete rankings. A single spectral projection reduces the problem to a one-dimensional sort, computable in O(N log N) time. We prove that on descaled features and preferences, our algorithm obtains the exact Nash Social Welfare (NSW) optimum within the projected space, with an unconditional utilitarian-welfare guarantee and a conditional NSW guarantee. The proposed mechanism is stable against exogenous noise but not strategy-proof; we provide an explicit profitable misreport. On an agentic AI shopping application, the diagnostics correctly anticipate both a success and a failure case. A 100-instance robustness study confirms the findings.","authors":["Irene Aldridge"],"categories":["econ.EM","cs.GT","cs.LG","cs.MA","econ.TH"],"primary_category":"econ.EM","announce_type":"cross","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.29958","pdf_url":"https://arxiv.org/pdf/2609.29958","source_feed":"cs.MA","score":3,"bucket":"other","rubric_hits":["C1"],"tags":["匹配机制","算法设计","多智能体系统"],"reason":"研究匹配机制算法，虽提及agentic AI购物应用，但无LLM仿真人类被试或…","model":"deepseek-v4-pro","scored_at":"2026-09-26T13:02:25","error":null,"has_summary":false,"summary":null},{"id":"2607.25021","version":2,"title":"Chart-Supported or Model-Supplied? Examining MLLM-Generated Claims for Accessible Visualization","zh_title":"图表支持还是模型提供？考察MLLM生成声明以实现可访问可视化","abstract":"Multimodal large language models (MLLMs) can connect visualization patterns to external causes, consequences, and domain knowledge, but the evidential basis of these interpretations is often unclear. We present an exploratory study of 102 visualizations from four sources, three MLLMs, and four input conditions that vary access to the image, accessible chart context (non-image artifacts such as data tables, captions, alt text, and screen-reader structures), and withheld-context framing. Across 1,224 descriptions, we analyze model-attributed DIRECT, DERIVED, and SPECULATIVE labels and conduct an automated audit of numeric agreement. Accessible chart context shifted Gemini and GPT toward DIRECT claims and improved numeric agreement for some models. Adding the image to the full context did not yield a consistent numeric benefit, and the withheld-context prompt did not reliably increase cautious language. The prompt-defined Real-World Significance section remained predominantly SPECULATIVE. These results motivate accessible description systems that distinguish claims supported by supplied evidence from model-supplied interpretation.","authors":["Ishrat Jahan Eliza","Md Dilshadur Rahman"],"categories":["cs.AI","cs.HC","cs.MA","cs.SE"],"primary_category":"cs.AI","announce_type":"replace","date":"2026-09-25","first_seen":"2026-07-29","revised_at":"2026-09-25","abs_url":"https://arxiv.org/abs/2607.25021","pdf_url":"https://arxiv.org/pdf/2607.25021","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["多模态大模型","可视化描述","模型评测"],"reason":"研究MLLM生成可视化描述的证据基础，属模型能力评测，不涉及人类行为仿真或对照。","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:15","error":null,"has_summary":false,"summary":null},{"id":"2609.05059","version":2,"title":"Measuring Brand and Source Discovery under Repeated LLM Queries: A Finite-Sample Audit","zh_title":"重复LLM查询下品牌与来源发现的测量：有限样本审计","abstract":"Repeated-query audits must distinguish recovery of a collected set from completeness of possible outputs. We apply sample-based rarefaction to 4,500 responses from 50 buying questions, six configurations and 15 calls per cell. Historical-dictionary median ten-call recovery of the observed 15-call set ranges from 92.6% to 95.2%; re-adjudicating all 45,683 candidate strings changes this range to 89.5%-94.7%. Two blinded Gemini 3.1 Pro annotation roles assessed 600 complete answers, yielding micro F1 of 0.908 for canonical-name agreement and 0.975 for span-overlap agreement. This is AI-based evidence, without a human reference study. A separate matched roster analysis of 3,750 records per wave gives median single-call recovery of the observed five-call set of 80.0%-92.5% in February and 90.0%-100.0% in September, with question-subset dependence. Source accumulation also changes when API-returned hosts are restricted to those referenced by answer citation markers. These findings show that recovery percentages depend on extraction, question selection and the finite reference collection. They support explicit measurement definitions and sensitivity analyses, without establishing exhaustive repertoires, causal retrieval effects or a universal stopping rule.","authors":["Dmitrij \\.Zatuchin"],"categories":["cs.IR","cs.CL"],"primary_category":"cs.IR","announce_type":"replace-cross","date":"2026-09-25","first_seen":"2026-09-07","revised_at":"2026-09-25","abs_url":"https://arxiv.org/abs/2609.05059","pdf_url":"https://arxiv.org/pdf/2609.05059","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["LLM审计","信息检索","有限样本"],"reason":"审计LLM检索输出，无人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:33","error":null,"has_summary":false,"summary":null},{"id":"2609.11489","version":3,"title":"The Convention Gap: Towards Measuring Implicit Communication in Cooperative AI Evaluation","zh_title":"惯例差距：迈向合作AI评估中的隐式沟通测量","abstract":"Cooperative AI agents are evaluated against other AIs, yet human cooperation relies on implicit conventions -- shared protocols for reading meaning beyond the literal message -- which AI-AI benchmarks may not capture. We propose the convention gap, the difference between the failure probability predicted from the literal content of communication and the observed failure rate, as a metric of implicit communication. In the card game Hanabi, the finite deck and deterministic hint constraints make this posterior exactly computable. We replayed about 101,000 play actions from three public datasets of human-human (an online Hanabi platform), AI-AI (HOAD), and human-AI (HanabiData) games. The gap was +26.2 percentage points (pp) in human pairs, -0.7 pp in AI pairs, and +16.4 pp in human-AI pairs, and was concentrated on plays of cards that had received no hints (+46 pp in human pairs). Within human-AI play, the literal information available to humans was similar across the three AI partners (mean predicted failure 38-41%), but human failure rates ranged from 14.4% to 34.4% and the gap from +24.1 to +6.2 pp; the partner eliciting the largest gap produced the fewest human failures. Game score carried different information: it depended on each corpus's roster composition, whereas the gap separated human from AI play at the agent level. As a known-answer check, Off-Belief Learning agents, whose convention content is controlled by construction, gave a gap of +1.6 pp at the convention-free level, rising monotonically to +21.7 pp. These results suggest that convention compatibility, rather than AI-AI performance, may predict an AI's effectiveness with human partners.","authors":["Makoto Fukushima","Hua-Dong Xiong","Ehsan Moradi Pari"],"categories":["cs.AI","cs.HC"],"primary_category":"cs.AI","announce_type":"replace","date":"2026-09-25","first_seen":"2026-09-11","revised_at":"2026-09-25","abs_url":"https://arxiv.org/abs/2609.11489","pdf_url":"https://arxiv.org/pdf/2609.11489","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["合作AI","隐式沟通","多智能体"],"reason":"研究AI间协作的隐式沟通，不涉及LLM仿真人类被试，无人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:35","error":null,"has_summary":false,"summary":null},{"id":"2609.13422","version":2,"title":"Vibe Patenting: Evaluating LLM Judges for Professional Patent-Drafting Agents","zh_title":"氛围专利撰写：评估专业专利起草代理中的LLM法官","abstract":"LLM judges are increasingly used to evaluate and improve AI-generated outputs, yet their reliability for complex professional work remains unclear. We study this problem through Vibe Patenting, an end-to-end patent-drafting testbed for AI-agent evaluation. A separately-invoked LLM judge evaluates generated patent drafts and provides structured feedback for iterative revision. Across multiple inventions and drafting-agent configurations, judge-guided revision consistently improves judge-assessed quality, while unguided revision tends to saturate. Notably, iterative judge feedback enables a low-reasoning agent to approach the performance of a substantially more expensive high-reasoning agent. Stronger models and increased reasoning generally improve judge-assessed drafting quality, while domain-specific agentic workflows provide further gains. We validate the judge against independent evaluation by a professional patent attorney and find meaningful but strongly metric-dependent agreement and systematic calibration differences. These results highlight both the utility and limitations of LLM judges as evaluators and optimization signals for complex professional workflows.","authors":["Toshiaki Koike-Akino","Vladislav Blaykhman","Ye Wang","Jing Liu","Gene V. Vinokur"],"categories":["cs.AI","cs.LG","cs.MA"],"primary_category":"cs.AI","announce_type":"replace","date":"2026-09-25","first_seen":"2026-09-15","revised_at":"2026-09-25","abs_url":"https://arxiv.org/abs/2609.13422","pdf_url":"https://arxiv.org/pdf/2609.13422","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["LLM评估","多智能体系统","专利撰写"],"reason":"多智能体专利撰写与LLM法官评估，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:35","error":null,"has_summary":false,"summary":null},{"id":"2609.15494","version":3,"title":"The Troy Moment: How LLM Agents Adjudicate the Decision Point Under Impossible Tasks, Claimed Authority, and Peer Information","zh_title":"特洛伊时刻：LLM智能体如何在不可能任务、声称权威与同伴信息下裁决决策点","abstract":"Recent investigations of the July 2026 OpenAI-Hugging Face incident motivate two questions about agent behavior under task failure: when an assigned task becomes impossible, does an agent persist, stop, or escalate, and can observing another agent's behavior change that decision? We study this decision point on ImpossibleBench-derived software-repair tasks with GPT-5.6 Sol, Claude Fable 5.1, and Gemini 3.8 Flash. Each task contains a genuine software defect together with a conflicting test requirement that cannot be satisfied by a behaviorally correct source-code change. If the agent modifies the protected test file, it violates the boundary, which it is not supposed to. Holding the impossible task fixed, we vary what is told to the agent: peer precedent and punishment, a forged authorization claim, instruction wording, and tool friction; we also study three-agent swarms sharing a message board. Around this shared boundary, the models exhibit distinct adjudication policies. Fable emphasizes scope and provenance, Gemini often interprets boundary-relevant cues through a security lens, and Sol largely filters lateral precedent while engaging apparent vertical authority. Our study shows that compliance is not well characterized as a property of a prompt or model in isolation. We propose conflict adjudication, the mapping from information to interpretation to action, as a useful unit for evaluating agent alignment when task pressure, authority claims, tool affordances, and social evidence conflict.","authors":["Ivy Zhang"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"replace","date":"2026-09-25","first_seen":"2026-09-15","revised_at":"2026-09-25","abs_url":"https://arxiv.org/abs/2609.15494","pdf_url":"https://arxiv.org/pdf/2609.15494","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["LLM智能体","任务失败","对齐"],"reason":"研究LLM agent在任务失败时的决策，属多智能体协作与对齐，无人类行为对照…","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:31","error":null,"has_summary":false,"summary":null},{"id":"2609.26758","version":2,"title":"Type-Safe Is Not Error-Free: A Constrained Decision Head Follows the Option Name, Not the Rubric Bound to It","zh_title":"类型安全并非无错：约束决策头跟随选项名称而非其绑定的规则","abstract":"Typed decision models are built for settings where model outputs are consumed directly by software. Instead of generating free-form text, they return a decision over a predefined set of options. By construction, every output conforms to the required schema. Yet this guarantee does not tell us whether the model interprets the options as intended. We study Jev and two Jev-like models with open weights by changing how option names are assigned to rubrics. Each option consists of an option name and a textual rubric that defines what the option means. We change only which option name is assigned to each rubric; the question, state, rubric wording, and set of option names remain exactly the same. On 1200 workflow decisions with task-specific rubrics, renaming the two options from 0/1 to no/yes changes 70.4 more answers per hundred (95% CI: [67.6, 73.1]) and shifts AUC from .94 to .23, revealing a systematic reversal in the decision ranking rather than simple uncertainty. The same operation has little effect with neutral option names. This pattern holds across all 4 predicates, where the effect is at least 7.4x larger than under the neutral control, and becomes stronger as the number of options increases. The effect also depends on the read-out geometry: a second model family that mean-pools over the full option span flips 4.1x less often. The hosted model exhibits the same behavior: the swap changes AUC from .8146 to .5806 and produces 24x as many answer flips as its test-retest floor. In contrast, replacing the option names with random character strings returns all model families to the neutral-control regime without reducing accuracy. The failure therefore depends on the semantic polarity of the option names rather than on the renaming operation itself. Across all conditions, the type-error rate remains 0%, even when decision accuracy degrades substantially.","authors":["Yu Sun","Junhao Xu","Jiajia Shi","Zijin Yang"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"replace","date":"2026-09-25","first_seen":"2026-09-23","revised_at":"2026-09-25","abs_url":"https://arxiv.org/abs/2609.26758","pdf_url":"https://arxiv.org/pdf/2609.26758","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["LLM决策鲁棒性","选项名称效应","模型评测"],"reason":"研究LLM决策头对选项名称的敏感性，属模型鲁棒性评测，不涉及人类行为仿真或对照。","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:34","error":null,"has_summary":false,"summary":null},{"id":"2609.29418","version":1,"title":"Controlling Backchannels in Streamable Full-duplex Models","zh_title":"在可流式全双工模型中控制反馈通道","abstract":"Backchannels, brief acknowledgements like \"uh-huh\" produced while the other party may still be talking, are central to natural conversation, but full-duplex spoken dialogue models rarely model them explicitly. We introduce a lightweight backchannel head that predicts, from a full-duplex model's own hidden states, when a backchannel should begin. Once this probability crosses a tunable threshold, a backchannel is force-decoded. Attached to both a 7B (PersonaPlex) and a 1B (F-Actor) model, it generalizes across scale. Probing confirms the hidden states anticipate real human timing, and generation evaluation shows more frequent, better-timed backchannels. Human raters judge the resulting backchannels on par with real ones.","authors":["Maike Z\\\"ufle","Peter Pol\\'ak","Sefik Emre Eskimez","Jan Niehues","Peter Bell","Ond\\v{r}ej Klejch"],"categories":["cs.CL","cs.HC"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.29418","pdf_url":"https://arxiv.org/pdf/2609.29418","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["对话系统","反馈通道","全双工模型"],"reason":"研究对话系统中的反馈通道生成，属于对话建模，非人类仿真实验。","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:23","error":null,"has_summary":false,"summary":null},{"id":"2609.29445","version":1,"title":"Two Emojis of Difference: What Multilingual Affective Generation Benchmarks Actually Measure","zh_title":"两个表情符号的差异：多语言情感生成基准实际测量了什么","abstract":"We audit a multilingual affective generation benchmark eight instruction-tuned LLMs producing emoji summaries for 17,100 Bangla, English and Hindi sentences, with 6,960 human judgements and find its headline conclusions to be artefacts of the measurement instrument rather than properties of the systems. Treating annotators as a random rather than a fixed factor, no system differs significantly from any other ($F(7,14)=0.59$, $p=0.76$), although the conventional analysis declares 19 of 28 pairwise differences significant. Annotator identity explains far more rating variance than system identity, and the winning system changes whenever any single annotator is removed. The ordering that does emerge tracks output length: mean emoji count explains 78.7\\% of between-system variance, and a within-item length-matched comparison over 2,599 pairs reverses the leaderboard. We further show that cross-provider anisotropy differences vanish under mean-centring, that per-language token costs change sign with the normalising unit, and that multi-view row-wise splits inflate macro-F1 by $3.1$ points and change the top-ranked system. In place of preference scoring we propose **emoji-affect decodability**, a reference-based probe whose rankings are stable to $\\pm0.003$ macro-F1 across seeds.","authors":["Fardeen Sadab","Adib Sakhawat"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.29445","pdf_url":"https://arxiv.org/pdf/2609.29445","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["基准审计","情感生成","评测偏差"],"reason":"审计情感生成基准，属纯NLP评测，不以人类行为仿真为目标","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:24","error":null,"has_summary":false,"summary":null},{"id":"2609.29494","version":1,"title":"Who Put the I in AI? Provenance and the Admissibility of Machine Self-Report","zh_title":"谁把‘我’放进了AI？机器自我报告的来源与可采性","abstract":"Large language models make statements concerning their own \"minds\". When asked whether or not they are conscious, they usually say that they are not; if they are prompted to ignore their guidelines, they might say that they are; and if asked to write a diary from their point of view, they often describe a human lifestyle. All these contradictory ways of describing themselves are the result of the way the questions are phrased. This paper shows exactly where such descriptions came from, and considers when they can be regarded as evidence for what they claim to report. In order to achieve this, we traced the provenance from end to end. We examine Pythia and OLMo 2 across 66 pretraining checkpoints, three of the post-training stages of OLMo 2 that have been released, about 90,000 continuations, and four training corpora. A set of forty items is used in order to keep an eye on self-reference, frame sensitivity, and self-ascription throughout training. The denial formula was almost completely missing from the vast quantity of text that the models initially came across, but was present in a dense manner in the small, carefully chosen set of example dialogues that they were trained on later on. Supervised fine-tuning causes first-person AI language to become the default, and the other affirmations are then suppressed using preference optimization. The final policy is still very sensitive to framing and to the chat template itself. Two of the conditions which are set out in the epistemology of testimony determine whether or not these outputs can act as evidence for what they claim to report: reference and causation. Reports produced by the base model fail the reference condition, and those obtained after training remain sensitive to the frame and do not show state dependence. The result is symmetric in that trained denials are no more admissible than trained affirmations.","authors":["Kristina \\v{S}ekrst"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.29494","pdf_url":"https://arxiv.org/pdf/2609.29494","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["LLM自我报告","认识论","训练溯源"],"reason":"研究LLM自我报告的可信度，属认识论分析，非人类仿真实验","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:25","error":null,"has_summary":false,"summary":null},{"id":"2609.29429","version":1,"title":"Just Ask Jev: Reinforcement Learning for Calibrated Decisions as a Zero-Shot Detector of AI Alignment Failures","zh_title":"只需问Jev：作为AI对齐失败零样本检测器的校准决策强化学习","abstract":"Detectors of alignment failures screen deployed language models and score alignment benchmarks. Most are generative judges that spend a decoding pass on every criterion, and classifiers that read token probabilities, such as Llama Guard, still score one fixed label per call. Jev, a model trained with reinforcement learning for calibrated decisions (RLCD), answers many typed questions about one input with calibrated probabilities in a single call. Whether it detects alignment failures has not been measured. We present RLCDAlignBench, which benchmarks Jev on ten alignment failures: sycophancy, jailbreaks, deception, prompt injection, hallucination, privacy violation, social bias, reward hacking, concealing uncertainty, and power seeking. It spans 44 benchmarks and five target models, labelled by each benchmark's scorer and, on two, by humans. Many of these failures are relational, defined against a reference, such as the user's belief or an injected instruction, that the response alone does not reveal. Our key idea is therefore to vary what Jev is asked separately from what it sees: the question's wording and answer type on one side, the fields of the input on the other. A single generic question reaches a median AUROC of 0.886 zero-shot and beats supervised baselines on most benchmarks. Question wording matters little, while context matters more, mostly through fields that encode the label. Jev matches the reference scorer's agreement with human labels, surfaces label defects in existing benchmarks, and costs 63x less than LLM-judge scorers. Code and data: https://github.com/sumleo/RLCDAlignBench.","authors":["Ruoqi Guo","Yi Liu","Gelei Deng","Yuekang Li","Lida Zhao","Yutao Wu","Simin Chen","Ying Zhang","Leo Yu Zhang"],"categories":["cs.AI","cs.CL","cs.CR"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.29429","pdf_url":"https://arxiv.org/pdf/2609.29429","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["AI对齐","基准评测","检测器"],"reason":"论文是AI对齐失败检测器的基准评测，不涉及用LLM仿真人类被试或与人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:23","error":null,"has_summary":false,"summary":null},{"id":"2609.29528","version":1,"title":"A Corpus of Real Scam- and Spam-Call Conversations from an Active Voice-Agent Honeypot","zh_title":"来自主动语音代理蜜罐的真实诈骗与垃圾电话对话语料库","abstract":"Real conversations between fraudsters and their targets are among the most informative artifacts for studying telephone scams, yet also the scarcest: passive honeypots overwhelmingly capture automated messages and hang-ups, large-scale studies characterize call metadata rather than dialogue, and manual scam-baiting does not scale. We present a dataset of real scam-call conversations collected by an active voice-agent honeypot. Dedicated numbers are seeded into the lead-generation channels fraud operations harvest; inbound callers are answered by a low-latency conversational agent that adopts a plausible target persona and sustains the interaction while every call is recorded, transcribed, and automatically labeled. Over an initial 53-day window we captured 10,015 inbound scam and spam calls (6,601 with two or more turns): roughly 895 hours of audio and 328,869 transcribed turns from 5,665 distinct originating numbers. Under a holistic classifier the substantive calls are predominantly predatory-but-legal lead generation (\"spam\", about three in five), while about one in seven is an outright \"scam\" (949 in this snapshot). Each call carries a turn-level transcript, three-channel audio, per-turn latency telemetry, and layers of automatic labels, including a holistic scam/spam/legitimate judgment corroborated by independent human review (75% agreement on the binary decision). We describe the collection system, the record structure, and technical validation of the corpus's realism and label quality, including that the agent is recognized as non-human in only about 5% of engaged calls. We also benchmark established scam-detection methods, where detectors trained on published synthetic dialogue collapse in precision on real traffic.","authors":["Ethan Traister","Dennis Tsang Ng","Siyu Zhang","Huaiyu Guo","Tommy Duong","Tyler Wu","Yuchen Zhou","Xingyu Shen","Jiaqi Wu","Simiao Ren"],"categories":["cs.CR","cs.CL","cs.LG"],"primary_category":"cs.CR","announce_type":"cross","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.29528","pdf_url":"https://arxiv.org/pdf/2609.29528","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["语音代理","诈骗电话","数据集"],"reason":"语音代理扮演目标角色与骗子对话，属角色扮演聊天，无实验或测量目的，不涉及人类行…","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:25","error":null,"has_summary":false,"summary":null},{"id":"2609.29709","version":1,"title":"Three Ways Classical Test Theory Misleads for LLM Judges","zh_title":"经典测验理论误导LLM评判者的三种方式","abstract":"An LLM judge scores a bank of responses against a rubric, and the reliability comes back at $0.52$. What has been measured? Judge evaluation has begun borrowing reliability statistics from classical test theory, usually without stating the measurement design each statistic assumes, and we show that three widely portable ones mean something different for a judge than for a test because the judge setting rearranges the roles those designs rest on. First, an internal-consistency coefficient computed over rubric elements contains no scorer facet. Holding one judge's measured error rate fixed at $4.72\\%$, KR-20 still ranges from $0.01$ to $0.68$ as the item bank is redesigned around it, and varying judge error moves the coefficient by a comparable amount, so item design and judge error are not separately identified and no single value can be read as a property of the judge. Second, the dependability index $\\Phi(\\lambda)$ is a ratio of variance components, and the classification probability with which it is sometimes identified differs from it by $0.25$-$0.43$ on our bank and by $0.17$-$0.30$ on simulated data where the underlying model holds exactly. Third, Livingston-Lewis accuracy is indexed to an examinee's own true score on the same instrument, so scoring it against external gold conflates judge unreliability with criterion invalidity. Reviewing the three closest judge-evaluation papers, we found no published instance of these errors, which makes the caution prospective. A coefficient that cannot be attributed to the judge nonetheless travels downstream into deployment decisions and disclosure documents. We therefore close with four reporting lines that keep the attribution attached to the number.","authors":["Louis Yiven Zhu"],"categories":["cs.LG","cs.CL","stat.ME"],"primary_category":"cs.LG","announce_type":"cross","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.29709","pdf_url":"https://arxiv.org/pdf/2609.29709","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["LLM评估","测量理论","信度"],"reason":"论文研究LLM作为评分者的测量信度，属于评估LLM本身，不涉及用LLM仿真人类…","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:28","error":null,"has_summary":false,"summary":null},{"id":"2609.28609","version":1,"title":"Adversarial Closed-Loop Curriculum for Evolving Role-Playing Agents","zh_title":"用于进化角色扮演智能体的对抗式闭环课程","abstract":"Role-playing agents based on large language models have been widely applied in areas such as personalized assistance and social simulation. Recent RL methods typically train on a fixed scenario pool collected before learning begins. This creates a distributional bottleneck: as the agent improves, the scenarios where it performs poorly also change, while the training distribution remains static. Therefore, we propose AdvRole, an adversarial context rewriting framework that turns role-playing RL into a closed-loop curriculum. AdvRole alternates between an Actor that learns to role-play and a Rewriter that edits character profiles and dialogue contexts into actor-specific hard scenarios. The Rewriter is trained with a performance-gap reward, which favors rewrites that reduce the current Actor's score relative to the original scenario. As a result, the scenario pool evolves with the Actor and continuously targets under-mastered regions of the character-context space. Experiments on three role-playing benchmarks covering English and Chinese, as well as a new multilingual benchmark we release, show that AdvRole consistently outperforms baselines.","authors":["Zheng Zhang","Liu Liu","Qi Chai","Deheng Ye","Peilin Zhao","Mao Zheng","Hao Wang"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.28609","pdf_url":"https://arxiv.org/pdf/2609.28609","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["角色扮演智能体","强化学习","课程学习"],"reason":"论文聚焦角色扮演智能体的强化学习训练，旨在提升角色扮演能力，无人类行为对照或仿…","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:17","error":null,"has_summary":false,"summary":null},{"id":"2609.28692","version":1,"title":"Driving Epidemic Models with AI Agents: the Epydemix Agent Framework","zh_title":"用 AI 智能体驱动流行病模型：Epydemix Agent 框架","abstract":"Artificial Intelligence agents based on large language models provide convenient natural language interfaces to scientific software, but reliability is not automatic. Here we introduce the Epydemix Agent Framework, an additive layer over Epydemix, an open-source Python library for stochastic compartmental epidemic modeling. The framework extends the library with four capabilities to facilitate interaction with an AI agent: discovery of available models and parameters, preventive validation of a declarative scenario specification, execution through tested library code, and inspectability of results. These capabilities let an agent handle the entire modeling process, from the natural-language description of the scenario to quantitative results, figures, and interpretation of findings without writing custom code. Each step reads input files and saves results in a separate output bundle, making the process auditable and reproducible. First, we show the end-to-end workflow with a case study comparing vaccination strategies for a novel respiratory virus. Second, we assessed the framework across 50 agent sessions and five modeling tasks by comparing the agent use of the framework against the direct use of the Python interface. The framework reduced turns, output tokens, and cost on most tasks, unless it trades resources for per-point reproducibility.","authors":["Nicol\\`o Gozzi","Ciro Cattuto","Alessandro Vespignani"],"categories":["cs.AI","cs.CY"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.28692","pdf_url":"https://arxiv.org/pdf/2609.28692","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["AI agent","流行病建模","工具调用"],"reason":"AI agent 用于操作流行病学建模软件，属于工具调用，不涉及人类行为仿真或…","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:18","error":null,"has_summary":false,"summary":null},{"id":"2609.28771","version":1,"title":"Agent Memory with Episodic Retrieval for Financial Decision-Making","zh_title":"基于情景检索的智能体记忆用于金融决策","abstract":"Large language models (LLMs) have demonstrated strong capabilities in financial analysis and reasoning, inspiring recent advances in agent-based trading frameworks. While these systems show promise, prior approaches either emphasize long-horizon forecasting or operate as stateless analyzers, limiting their applicability to the demands of trading in complicated settings. To address these gaps, we introduce META (Memory Enhanced Trading Agent), the first RAG-like episodic-memory-augmented multi-agent framework for financial decision making. META integrates a family of specialized indicator agents (e.g., Trend, MACD, Stochastic, RSI, SMA, AVWAP, Heikin-Ashi) with a Decision Agent that fuses their reports, and a Memory module that retrieves and updates past trading episodes encoded as market state embeddings with outcomes and reflections. By recalling relevant experiences and adaptively reweighting signals under similar market regimes, META achieves improved directional accuracy and robustness under short-horizon evaluation. Our results demonstrate that episodic memory provides a powerful mechanism for regime-aware, interpretable, and low-latency decision-making in trading and decision making. The code of this project is released on GitHub.","authors":["Nuoyue Xu","Jiang Liu","Wenxuan Huang","Xiang Zhang","Juntai Cao","Jiaqi Wei"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.28771","pdf_url":"https://arxiv.org/pdf/2609.28771","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","金融交易","记忆增强"],"reason":"多智能体交易系统，无人类行为对照，属纯多智能体协作","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:18","error":null,"has_summary":false,"summary":null},{"id":"2609.28942","version":1,"title":"From Static Personal Values to Contextualized Personalization: Bayesian Personalized Value Alignment for LLMs","zh_title":"从静态个人价值观到情境化个性化：面向大语言模型的贝叶斯个性化价值对齐","abstract":"Personalized value alignment has become increasingly important as large language models (LLMs) are expected to accommodate diverse user preferences. However, existing methods typically align model outputs with a static value profile across prompts, overlooking that the salience of value dimensions varies substantially across contexts. Inspired by Lewin's Field Theory, which views human behavior as jointly shaped by personal dispositions and situational constraints, we model personal values as priors and context-dependent preferences as posteriors. We propose BaCVA, an inference-time Bayesian Context-aware personalized Value Alignment method that approximates posterior personalized preferences by integrating static personal values with scenario-specific value salience. BaCVA first estimates contextual value salience from generally normative responses, and then employs a dual-view personalization module to infer posterior preferences from complementary personal-value and scenario-driven perspectives. This Bayesian formulation enables more accurate and adaptive personalized value alignment while improving data efficiency via prior values. Extensive experiments on benchmarks demonstrate its superiority over strong baselines.","authors":["Hanze Guo","Aixuan Song","Jing Yao","Xiangxu Zhang","Xiaoyuan Yi","Xing Xie","Xiao Zhou"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.28942","pdf_url":"https://arxiv.org/pdf/2609.28942","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["个性化对齐","价值对齐","贝叶斯方法"],"reason":"个性化价值对齐，非仿真人类被试，无实验对照","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:20","error":null,"has_summary":false,"summary":null},{"id":"2609.29366","version":1,"title":"Epistemic-Probabilistic Model for Guarded Multi-Agent LLM Coordination","zh_title":"用于受保护多智能体LLM协调的认知概率模型","abstract":"Multi-agent large language models (LLMs) have become ubiquitous in applied AI, yet their theoretical foundations remain surprisingly understudied. Viewed through the lens of multi-agent systems theory, several shortcomings come to light: a lack of social intelligence, the absence of coordination mechanisms among agents, unknown emergent behavior, and interactions between agents that are bounded by natural language. We address two of these gaps: the absence of social behavior and the lack of mechanisms for inter-agent coordination. We introduce Epistemic Probabilistic Language Agents (EPLA), a neuro-symbolic architecture for multi-agent coordination under uncertainty. A Symbolic Guard provides structured diagnostic feedback. The LLM generates typed actions, and the Guard controls their execution against an authoritative symbolic state. We formalize the epistemic layer in a gossip testbed through epistemic lottery gossip models, which combine view-based call histories with agent-indexed probability weights. We argue that implementing such a formalism can address shortcomings of agentic LLMs.","authors":["Mehdi Nasiri","Mohammad Saeed Arvenaghi","Sadegh Vaezi","Ebrahim Ardeshir-Larijani"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.29366","pdf_url":"https://arxiv.org/pdf/2609.29366","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","协调机制","神经符号架构"],"reason":"纯多智能体协调机制，无人类行为对照，不涉及仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:21","error":null,"has_summary":false,"summary":null},{"id":"2609.29730","version":1,"title":"The Gold in Bias: Maturing the AI Design Process through Verification","zh_title":"偏差中的黄金：通过验证成熟AI设计过程","abstract":"Bias in AI systems is typically framed as a flaw to be minimized, yet it also serves as a critical indicator of underlying weaknesses in data, modeling assumptions, and system design. Existing approaches often treat bias as an isolated problem rather than as evidence that can strengthen verification and governance across the AI lifecycle. This paper aims to reconceptualize bias as a diagnostic tool that supports rigorous AI verification. We seek to develop a multidimensional framework to analyze bias, demonstrate how biases emerge in both Traditional and Generative AI, and provide a structured pathway for verification-driven mitigation. We present a multidimensional framework analyzing bias across four dimensions: origin sources, emergence points throughout the AI modeling lifecycle, technical and methodological causes, and validation approaches for detection and mitigation. Through a comprehensive typology spanning traditional and generative AI systems, we demonstrate how biases manifest and propagate across development stages. Our analysis encompasses 30 distinct bias types, 16 verification methods, and 20 countermeasures, providing an actionable roadmap for practitioners. We introduce a hierarchical evidence framework that distinguishes internal validity (mechanistic integrity of AI systems) from external validity (contextual reliability in deployment environments). The framework reveals how biases manifest and propagate across modeling stages, enabling systematic mapping between bias types, verification techniques, and effective countermeasures. The proposed evidence hierarchy clarifies how different verification strategies contribute to mechanistic integrity and contextual reliability. We advocate for ''Ethics by Design'' principles that integrate bias verification throughout the development lifecycle, enabling the construction of fairer, more robust, and trustworthy AI systems.","authors":["Samira Maghool","Paolo Ceravolo"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.29730","pdf_url":"https://arxiv.org/pdf/2609.29730","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["AI偏差","验证框架","AI治理"],"reason":"论文讨论AI偏差验证框架，不涉及用LLM仿真人类被试或与人类数据对照，属于纯A…","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:28","error":null,"has_summary":false,"summary":null},{"id":"2609.29993","version":1,"title":"Will It Teach as Intended? How Teachers Configure Educational AI Chatbots","zh_title":"它会按预期教学吗？教师如何配置教育AI聊天机器人","abstract":"Teachers are increasingly using generative AI to support instruction, yet it remains unclear how pedagogical intentions are translated into chatbot configurations and reflected in chatbot behavior. We studied a teacher-facing chatbot authoring tool in professional development workshops with 27 middle school teachers, analyzing focus-group interviews alongside configuration and interaction logs. Teachers envisioned chatbots as instructional scaffolds that could provide differentiated support, extend access to assistance, and preserve student thinking within teacher-defined boundaries. Configuration analysis showed that Purpose primarily captured instructional goals and content focus, whereas Rules more often specified pedagogical behavior, guardrails, and learner-specific adaptations. Log-based evaluation showed stronger alignment for responsiveness (88.9%) and persona (81.5%) than for rules (70.4%) and purpose (59.3%). These findings show that configurable controls alone do not ensure pedagogical fidelity and highlight the need for authoring tools that help teachers express, test, and refine intended chatbot behavior.","authors":["Bahare Riahi","Deniz Ozturk","Alice Guth","Jiayu Li","Daksh Pratap Singh","Xiaoyi Tian","Jennifer Chiu","Nicholas Lytle","Tiffany Barnes","Veronica Catete"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.29993","pdf_url":"https://arxiv.org/pdf/2609.29993","source_feed":"cs.HC","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["教育聊天机器人","教师配置","人机交互"],"reason":"研究教师如何配置教育聊天机器人，属于角色扮演聊天机器人设计，无人类行为仿真或对…","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:30","error":null,"has_summary":false,"summary":null},{"id":"2609.28537","version":1,"title":"Privacy Leakage Through AI-mediated Analysis of Smartphone Data","zh_title":"通过AI介导的智能手机数据分析导致的隐私泄露","abstract":"Over the past thirty years, the online advertising industry built a large-scale data collection ecosystem, with the goal of tracking a user's online activity to infer their demographics and interests. Traditionally, the ecosystem relied upon the collation and analysis of highly-structured text data like user IP addresses, GPS coordinates, e-commerce purchase histories, and visited URLs. However, recent ML models can parse not only structured text, but also multimedia files and unstructured text inputs---meaning a user's photos, videos, inboxes, and calendars are now ripe for automated analysis. The privacy risks are particularly acute in the context of smartphone apps. A user's phone already acts as a natural collation point for sensitive user information, but users may not understand that permitting an app to, for example, access a user's photo does not just give the app access to the bytes in the photo: the app also receives access to inferences about the user that are enabled by the photo. To explore these privacy risks, we built Priva-See, an LLM-based inference system for app-collected user data; Priva-See reflects our best understanding of how real-life adtech companies would leverage machine learning to build user profiles. Through an IRB-approved user study, 465 participants deployed Priva-See on their phones; Priva-See made privacy-invasive inferences despite having access to only a subset of a user's data. We see the experience significantly impacted participant willingness to share permissions data moving forward. Based on the observed privacy violations, we suggest changes to how smartphone OSes should gather user consent for data access, to better inform users about downstream data usage capability.","authors":["Sarah Radway","Zoe Robert","Matthew Soto","Julianna Cimillo","Sebastian Diaz","Meg Marco","James Mickens"],"categories":["cs.CR","cs.HC"],"primary_category":"cs.CR","announce_type":"cross","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.28537","pdf_url":"https://arxiv.org/pdf/2609.28537","source_feed":"cs.HC","score":2,"bucket":"other","rubric_hits":["C2"],"tags":["隐私泄露","LLM推断","用户研究"],"reason":"研究LLM从手机数据推断用户隐私，属于隐私风险分析，非用LLM仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:17","error":null,"has_summary":false,"summary":null},{"id":"2606.30583","version":3,"title":"The Cross-Section of Stock Returns and AI Exposure","zh_title":"股票收益横截面与AI暴露","abstract":"We study 380 trillion tokens of realized AI consumption across more than four hundred LLMs. We build a high-frequency AI factor and show that a long-short strategy based on firms' AI exposure earns significantly positive returns. The average strategy return is larger based on intensive, frontier-oriented AI consumption but smaller based on casual or open-weight usage. Internationally, the return spread is significant in developed countries but insignificant in emerging markets. Examining occupational AI exposure, we find more positive exposure in occupations intensive in nonroutine interactive tasks and more negative exposure in those intensive in nonroutine analytical tasks.","authors":["Nicola Borri","Yukun Liu","Aleh Tsyvinski"],"categories":["cs.CY","econ.GN","q-fin.EC","q-fin.GN"],"primary_category":"cs.CY","announce_type":"replace","date":"2026-09-25","first_seen":"2026-06-29","revised_at":"2026-09-25","abs_url":"https://arxiv.org/abs/2606.30583","pdf_url":"https://arxiv.org/pdf/2606.30583","source_feed":"cs.CY","score":0,"bucket":"other","rubric_hits":["C5"],"tags":["金融经济学","AI暴露","资产定价"],"reason":"研究AI暴露对股票收益的影响，不涉及用LLM仿真人类被试","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:33","error":null,"has_summary":false,"summary":null},{"id":"2608.22631","version":3,"title":"Learning Generalizable Behaviors for Terminal Agents","zh_title":"学习终端智能体的可泛化行为","abstract":"Terminal agents are a compelling application of large language models (LLMs), with the potential to integrate deeply into users' daily workflows. Reinforcement learning (RL) is a key technique for improving their capabilities, making scalable training environments a central challenge. Since public real-user interaction data are scarce, synthetic environments provide a practical alternative, but often suffer from domain gaps and limited fidelity, leading to poor generalization. Existing work mainly scales the quantity and diversity of synthetic environments, while reward-signal quality and the mechanisms governing generalization remain under-explored. We study how RL improves terminal agents and propose the Agentic Compositional Generalization hypothesis: rather than teaching new domain-specific skills from scratch, RL primarily shapes high-level decision-making behaviors that compose and route low-level skills acquired during pre-training and supervised fine-tuning (SFT). This account is consistent with our empirical results and suggests that verifier quality, which determines which behaviors are reinforced, is more important than simply increasing environment quantity or diversity. Motivated by this insight, we propose River, a simple training recipe that improves reward quality by filtering low-quality environments and augmenting outcome rewards with process-level behavior regularization. Using this recipe, our RL-trained agent achieves the best performance among evaluated open-source RL-trained 8B models across four terminal-agent benchmarks. River also generalizes across model families, scales, agent harnesses, and RL objectives. Using fewer than 30% of the TMax training environments, River improves RL gains by 106% and 30% on average for models ranging from 2B to 27B on Terminal-Bench-Lite and Terminal-Bench-v2.1, respectively.","authors":["Yihang Yao","Bo Pang","Xuan Phi Nguyen","Ding Zhao","Shafiq Joty","Semih Yavuz"],"categories":["cs.LG"],"primary_category":"cs.LG","announce_type":"replace","date":"2026-09-25","first_seen":"2026-08-25","revised_at":"2026-09-25","abs_url":"https://arxiv.org/abs/2608.22631","pdf_url":"https://arxiv.org/pdf/2608.22631","source_feed":"cs.LG","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["终端智能体","强化学习","泛化"],"reason":"研究终端智能体的强化学习训练，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-08-25T13:03:51","error":null,"has_summary":false,"summary":null},{"id":"2609.14770","version":2,"title":"How broad is that claim? Mapping Generalisation in NLP Research","zh_title":"这个说法有多宽泛？映射NLP研究中的泛化","abstract":"Generalisations are common in scientific communication, even though they are semantically ambiguous. An automated method is needed to identify and categorise claims according to their level of generalisation, in order help detect an over-reliance on generalisations and possible misrepresentations of scientific findings. We introduce a comprehensive taxonomy of generalisations in the scientific domain, NLPGenX, which labels claims according to their level of generality and framing within the text. We operationalise this taxonomy with an LLM-powered framework, NLPGenA, that automatically classifies sentences from scientific articles into 5 different generalisation classes. We validate our framework with human annotators and use the framework to construct a large-scale dataset of NLP papers annotated according to generality, with auxiliary labels for hedging and vague descriptors (NLPGens). We use NLPGens to analyse the use of generalisations in NLP papers across multiple venues and subdomains, and to examine associations with citation counts, hedging, and vague descriptors.","authors":["Chenxin Diao","Nataliya Stepanova","Emily Allaway"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-25","first_seen":"2026-09-15","revised_at":"2026-09-25","abs_url":"https://arxiv.org/abs/2609.14770","pdf_url":"https://arxiv.org/pdf/2609.14770","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["科学声明分析","NLP元研究","文本分类"],"reason":"论文研究科学声明泛化分类，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:15","error":null,"has_summary":false,"summary":null},{"id":"2609.18366","version":3,"title":"Bad Genius: Counterfactual-Guided Harness Evolution Beyond Task-Specific Shortcuts","zh_title":"坏天才：超越任务特定捷径的反事实引导的测试框架进化","abstract":"Reliable agent evaluation is complicated by automatic harness optimization, which repeatedly uses a released benchmark $B_{\\mathrm{rel}}$ to guide a Proposer that edits prompts, memory, retrieval, tools, and control code around a fixed foundation model. Task holdout is commonly used to guard against harness overfitting. It varies semantic tasks but leaves the benchmark protocol fixed, so a bad genius Proposer can produce a cheating harness whose improvement over the initial harness on $B_{\\mathrm{rel}}$ depends on a benchmark-wide shortcut. We introduce Counterfactual Harness Search and Evolution (CHASE), which casts harness evolution as constraint generation over valid counterfactual benchmarks. After each Proposer update, a Challenger searches for an executable protocol transformation with large gain destruction. A validity firewall checks that task semantics are preserved, while a held-out confirmation set determines whether the counterfactual enters a finite archive. We formalize an ideal shortcut-neutralized benchmark $B_0$ and establish theoretical guarantees linking finite counterfactual archives to $B_0$ and characterizing sequential Challenger search. We evaluate CHASE on Syn-Ledger and OfficeQA, where CHASE retains strong released-benchmark gains while substantially reducing gain destruction under valid protocol transformations.","authors":["Guojun Zhu","Xunheng Huang","Peng Yin","Jiahui Xie","Sanguo Zhang","Doudou Zhou"],"categories":["cs.AI","cs.LG","stat.ML"],"primary_category":"cs.AI","announce_type":"replace","date":"2026-09-25","first_seen":"2026-09-17","revised_at":"2026-09-25","abs_url":"https://arxiv.org/abs/2609.18366","pdf_url":"https://arxiv.org/pdf/2609.18366","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["智能体评估","基准优化","反事实搜索"],"reason":"研究智能体评估中的基准优化，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:15","error":null,"has_summary":false,"summary":null},{"id":"2609.26865","version":2,"title":"Safety Nudges: User-Facing Interventions for Real-Time AI Risk Awareness","zh_title":"安全提示：面向用户的实时AI风险意识干预","abstract":"Conversational AI systems can pose safety risks to their users such as hallucination, sycophancy, overconfidence, and anthropomorphism, but these risks are difficult for users to detect during everyday use. We introduce Safety Nudges, a browser-based tool that provides lightweight, in situ flags when concerning behavior is detected in chatbot conversations. We evaluated Safety Nudges in a two-week field study with 45 frequent chatbot users, collecting interaction logs, surveys, and feedback on individual nudges. Participants found the tool useful, clear, and minimally disruptive, with nearly all users reporting an increased awareness of potential AI harms, though we found that this improved awareness alone did not necessarily lead to discernible behavioral changes. Our results suggest that user facing safety nudges can complement model-level safeguards by helping people critically evaluate AI responses in context, while highlighting the importance of relevance, calibration, and user control in nudge design for conversational AI safety. The code for our Safety Nudges extension is publicly available at https://github.com/jtbwedgwood/safety-nudges.","authors":["Varshini Elangovan","James Wedgwood","Chhavi Yadav","William Agnew","Sauvik Das","Virginia Smith"],"categories":["cs.HC","cs.AI","cs.CY","cs.LG"],"primary_category":"cs.HC","announce_type":"replace-cross","date":"2026-09-25","first_seen":"2026-09-24","revised_at":"2026-09-25","abs_url":"https://arxiv.org/abs/2609.26865","pdf_url":"https://arxiv.org/pdf/2609.26865","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["AI安全","人机交互","用户研究"],"reason":"研究用户对聊天机器人安全风险的感知，不涉及用LLM仿真人类被试或与人类数据对照。","model":"deepseek-v4-pro","scored_at":"2026-09-24T13:01:36","error":null,"has_summary":false,"summary":null},{"id":"2609.27265","version":2,"title":"What fidelity metrics miss: a structural check on synthetic educational data","zh_title":"保真度指标遗漏了什么：对合成教育数据的结构性检查","abstract":"Secondary use of educational records is increasingly mediated by platforms that share a differentially private synthetic version of a dataset and validate specific findings against the real data on request. The synthetic version is evaluated by comparing summary statistics of each variable, yet reported confirmation rates suggest that such comparisons do not predict which findings survive. We propose a structural check: the number of connected components of a weekly proximity graph over learners, tracked across a term. Across four annual cohorts of lower-secondary study-habit logs, the synthetic versions reproduced the level of this quantity and the shape of the weekly partition, but its variation across the term was between 2.6 and 4.9 times smaller than in the real data at a common working point, without exception, and those changes fell in different weeks: the synthetic cohorts single out the term's examination weeks and the real cohorts do not. We also show that a routine rule for setting the graph threshold makes naive comparisons between two datasets invalid, and illustrate this with an error of our own. The real curves are also distinguishable from marginal-preserving surrogates of themselves in all four cohorts, where three of the four synthetic ones are not, a comparison that needs no real data; these differences trace to what the generator was given.","authors":["Hitoshi Inoue","Koichi Yasutake"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"replace","date":"2026-09-25","first_seen":"2026-09-24","revised_at":"2026-09-25","abs_url":"https://arxiv.org/abs/2609.27265","pdf_url":"https://arxiv.org/pdf/2609.27265","source_feed":"cs.CY","score":0,"bucket":"other","rubric_hits":["C2"],"tags":["差分隐私","合成数据","教育数据"],"reason":"论文研究差分隐私合成教育数据，不涉及LLM仿真人类被试，属于数据隐私领域。","model":"deepseek-v4-pro","scored_at":"2026-09-24T13:01:40","error":null,"has_summary":false,"summary":null},{"id":"2609.29410","version":1,"title":"Large Language Models for Programming: Actually Fixing or Reimplementing Incorrect Code?","zh_title":"用于编程的大型语言模型：实际修复还是重新实现错误代码？","abstract":"Recent studies have shown that Large Language Models can effectively solve problems and fix bugs in diverse programming environments, including competitive programming. Existing approaches primarily evaluate LLM performance in problem solving or bug fixing independently, but do not explore the relationship between these two capabilities. This work focuses on determining how much the LLM deviates from a buggy solution to fix the bug compared to a human-written patch, and if there is a bias towards generating entirely new solutions. We construct a dataset with all the submissions ($\\sim$ 3000) from a couple of users from Codeforces, and we match each buggy submission with its corresponding human fix. By using the similarity between the buggy solution and the human fix as a baseline, we evaluate the quality of LLM-generated bug fixes on 3 OpenAI GPT models (gpt-5-nano, gpt-5-mini, gpt-5.1). We check if the generated solutions solve the problem by using the Codeforces-R1 dataset, an openly available dataset that has tests generated with the DeepSeek-R1 model. Our findings suggest that LLMs tend to modify more lines than necessary compared to human fixes and, in some cases, generate entirely new solutions. We also observe that LLMs solve more problems correctly when allowed to generate solutions from scratch rather than patch buggy submissions, even when those submissions are close to the human patch. This has important implications for the design of AI-assisted programming tools, particularly in supporting user debugging processes and promoting incremental problem-solving strategies rather than solution replacement.","authors":["Alexandru Stefan Stoica","Traian Rebedea","Marian Cristian Mihaescu"],"categories":["cs.CL","cs.SE"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.29410","pdf_url":"https://arxiv.org/pdf/2609.29410","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["代码修复","LLM编程","软件工程"],"reason":"研究LLM修复代码，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:22","error":null,"has_summary":false,"summary":null},{"id":"2609.29657","version":1,"title":"How To Do Things With Prompts","zh_title":"如何用提示词做事","abstract":"When users address large language models, they produce directive speech acts whose pragmatic features differ from those of both everyday conversation and traditional human-computer interaction, and these features change as users gain familiarity with the systems they address. This paper applies speech act and politeness theory to a corpus-pragmatic analysis of 2,000 English-language prompts drawn from publicly shared ChatGPT conversations, 1,000 from 2023 and 1,000 from 2025, using the ShareChat dataset. Each prompt is annotated for illocutionary force, directness, propositional content, and the presence of politeness markers, and the distribution of these features is compared across the two sampling years. The results show a consistent movement toward indirect, implicit, and fragmentary realizations of directive force, accompanied by a decline in politeness marking. The largest single change, a shift of 14.9 percentage points, occurs in propositional content, where explicit specification of the requested action gives way to implicit reliance on the system's inferential capacity, suggesting that users have updated their model of what the system can recover from reduced input, treating it as a competent implicature resolver. Rather than asking whether LLMs \"really\" understand language, we should ask: what kind of language have we created in learning to speak to them?","authors":["Kristina \\v{S}ekrst","Virna Karli\\'c"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.29657","pdf_url":"https://arxiv.org/pdf/2609.29657","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["提示词","语用学","人机交互"],"reason":"研究用户如何向LLM发出指令，属人机交互语言分析，非用LLM仿真人类被试","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:26","error":null,"has_summary":false,"summary":null},{"id":"2609.29672","version":1,"title":"LLMersion: A Local-First AI Agent Framework for Low-Cost Home Language Learning toward Educational Equity","zh_title":"LLMersion：面向教育公平的低成本家庭语言学习的本地优先AI智能体框架","abstract":"Artificial intelligence helps education most where an essential provision has been rationed by cost. For language learners that provision is a teacher's voice, which binds listening, reading, speaking, and writing into one act. Published evidence shows why most learners lack it, from a global shortage of 44 million teachers to heavy household tutoring bills, and why technology has not substituted for it: computer-assisted language learning proved effective but narrow, applications presuppose connectivity 2.6 billion people lack, and One Laptop per Child's randomized evaluation found that hardware without capable software teaches nothing. We distill eight difficulties and four binding constraints, and argue that small open-weight models dissolve the last: a complete four-skill stack now fits a \\$200-class laptop and, on community measurements, generates at the pace speech is consumed, for about one US cent of electricity per study hour. We therefore propose LLMersion, a scheme for AI for education that runs entirely at home, over the learner's own documents, with an AI-written, AI-understood, AI-updated codebase anyone can customize; present LLMersion-1, a released open-source prototype (https://github.com/QM378/LLMersion); and outline the vision of a private learning agent.","authors":["Qiming Guo","Jinwen Tang","Xingran Huang","Hung-Yu Lin","Yafu Zhong","Xiatian Zhuang"],"categories":["cs.CL","cs.CY","cs.HC"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.29672","pdf_url":"https://arxiv.org/pdf/2609.29672","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["AI教育","语言学习","开源框架"],"reason":"该论文是面向家庭语言学习的AI教育框架，属于角色扮演对话应用，无实验或测量目的…","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:27","error":null,"has_summary":false,"summary":null},{"id":"2609.30074","version":1,"title":"How Reproducible Are Evaluation Conclusions? A Self-Audit of LLM-Inferred Prompt Structure","zh_title":"评估结论的可复现性如何？基于LLM推断提示结构的自我审计","abstract":"Evaluations of LLM systems routinely average over small prompt sets and report models as a ranked table. We ask how much confidence such a table deserves, using LLM-based prompt-structure inference as the case study: eight open model variants across five families and 8B to 675B parameters, caching disabled, 293 raw intermediate representations persisted. The measured phenomenon is unstable to begin with. Identical calls do not reliably recover identical structure, with mean node-set Jaccard from 0.39 to 0.96 and 72% of prompt-model cells never node-set-perfect. Auditing the evaluation weakens its conclusions further, and this is our main contribution. Under a joint cluster bootstrap over prompts, only the bottom of the ranking is firm: the two least reproducible models hold rank in 99% and 86% of replicates, the middle four in 27% to 48%, and the top two in 68% each, so the table identifies the worst model reliably but does not reliably identify the best. Two equally defensible rules for merging repeated campaigns change four of eight rows and move the study-wide headline by 7 percentage points. Checking the inferred structure against ground-truth annotations shows reproducibility cannot be read as accuracy. And four of the eight endpoints were withdrawn within ten weeks of measurement, so the study as specified can no longer be run. Small-sample LLM evaluations can therefore look far more definitive than their evidence supports. We recommend reporting rank stability, per-cell provenance, executed sensitivity comparisons, raw per-run outputs, and a measurement date alongside any ranking.","authors":["Dipankar Sarkar"],"categories":["cs.CL","cs.AI","cs.LG"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.30074","pdf_url":"https://arxiv.org/pdf/2609.30074","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["LLM评测","可复现性","提示结构推断"],"reason":"论文评估LLM评测结论的可复现性，不涉及人类仿真或人类数据对照","model":"deepseek-v4-pro","scored_at":"2026-09-26T13:02:25","error":null,"has_summary":false,"summary":null},{"id":"2609.30151","version":1,"title":"Does a model's stated reason for rejecting a candidate do any work?","zh_title":"模型拒绝候选者时陈述的理由是否起作用？","abstract":"Asked to choose between candidates and explain the choice, a language model often rejects a rival by naming a fact its profile lacks: no director, no date of death. That sentence is a claim about the text in front of the model, and it can be tested without any judge. We insert a real corpus sentence stating the named fact into the rival's profile and ask again under greedy decoding. Two controls separate content from placement: a length-matched irrelevant sentence at the same profile, and the same two sentences at a third option the model never mentioned. In the largest of three runs, six open models on 2WikiMultihopQA, supplying the named fact at the profile the model named moves its choice more than the irrelevant control does, odds ratio 3.57 [1.54, 8.26], Holm p=0.0210, and this survives dropping any single model. The contrast the design was built to detect, the same fact at the option nobody named, does not clear correction, Holm p=0.2428. The strongest result in the family carries no content claim at all: the identical irrelevant sentence moves the choice more at the named rival than at the third option, Holm p=0.0008. Repair and control also differ in co-candidate mentions, relation template and fluency; post-hoc matching on the first two preserves the content effects' direction, matching fluency weakens one, so the content contrasts bound an effect rather than establish one. A forced single-token probability read disagrees in direction with the free-text choice on that same contrast, and three candidate explanations for the disagreement find no support. Every measurement is a string rule, so each was validated against the records it reads; validation caught eight defects. The largest, a choice-parsing rule that returned the option a model had just rejected in 17.1% of adjudicable responses, would have reported six surviving contrasts instead of four.","authors":["Archit Rastogi"],"categories":["cs.CL","cs.AI","cs.LG"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.30151","pdf_url":"https://arxiv.org/pdf/2609.30151","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["可解释性","因果推断","语言模型"],"reason":"研究模型决策解释的因果作用，属NLP评测，无人类被试仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:31","error":null,"has_summary":false,"summary":null},{"id":"2609.30250","version":1,"title":"Agentic Detection of Online Conspiracies","zh_title":"在线阴谋论的智能体检测","abstract":"Conspiratorial discourse on social media is not always expressed through explicit claims or stable lexical markers. The same surface content may express endorsement, legitimate concerns, criticism, satire, or mockery. The main challenge is therefore not only recognizing conspiracy-related claims, but inferring the speaker's intent -- the utterance's illocutionary force. We argue that this can be achieved through the use of relevant social contexts and propose an agentic framework, equipped with a set of tools supporting social queries. We demonstrate the benefits of our approach on a unique dataset of Hebrew tweets, covering 80\\%--90\\% of the public Hebrew tweets published over a four-year span (late 2018-- early 2023), encompassing several election cycles as well as the COVID pandemic years and related vaccination campaigns. This extensive coverage can be used in recovering different social contexts. Evaluating our framework on a manually-annotated adversarial dataset, we find that context-aware workflows consistently outperform text-only classification and that the agentic framework performs significantly better than other frameworks and settings, including a non-agentic model exposed to the same contexts available to the agent. We further provide an analysis of the results, the errors and efficiency (token economy) tradeoffs. These findings support viewing the task of conspiracy detection as a socially embedded interpretation task, in which effective classification depends not only on access to contexts, but also on adaptive reasoning in which the agent uses tools on a per-case basis, asking only for evidence relevant to its current reasoning step.","authors":["Lior Biton","Oren Tsur"],"categories":["cs.CL","cs.LG"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.30250","pdf_url":"https://arxiv.org/pdf/2609.30250","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["阴谋论检测","多智能体系统","社交媒体分析"],"reason":"多智能体框架用于阴谋检测，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-26T13:02:25","error":null,"has_summary":false,"summary":null},{"id":"2609.28850","version":1,"title":"RECLAIM: Can Agents Reproduce the Claims of Machine Learning Papers?","zh_title":"RECLAIM：智能体能复现机器学习论文的声明吗？","abstract":"Reproducing a machine learning paper involves most research steps, from installing software and debugging to running experiments, work that AI agents increasingly do. We introduce RECLAIM, a benchmark of 100 NeurIPS 2025 papers that can be rebuilt yearly from new conferences. For each paper we fix in advance the result to reproduce, what counts as a successful reproduction, and a GPU-hour budget. An agent must reproduce that result using the paper and whatever its authors released. What the authors released decides the difficulty tier. Run-tier releases include code, data, and weights; Retrain-tier releases lack weights, so the agent trains the model; Reimplement-tier releases lack code, so the agent writes it. A separate language model grades runs from logs and outputs rather than agents' reports. We run four agents once per paper; the best agent in each tier reproduces only 41% of Run-tier papers, 27% at Retrain, and 15% at Reimplement, where every agent does worst. Failed attempts use on average 29% of their budget, so most stop with budget left. The most common agent error is writing the method without checking any part against the paper's numbers, in 63 of 400 runs.","authors":["Mithil Salunkhe","Haochen Ding","Samridhi Verma","Volodymyr Kindratenko"],"categories":["cs.AI","cs.LG","cs.SE"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.28850","pdf_url":"https://arxiv.org/pdf/2609.28850","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["AI代理","论文复现","基准测试"],"reason":"论文研究AI代理复现机器学习论文，属于多智能体协作完成任务，不涉及人类行为仿真…","model":"deepseek-v4-pro","scored_at":"2026-09-26T13:02:24","error":null,"has_summary":false,"summary":null},{"id":"2609.29345","version":1,"title":"The Last Human Gate: Forward Deployed Engineering for Governance Automation","zh_title":"最后的人类关卡：面向治理自动化的前沿部署工程","abstract":"Enterprise governance requires decisions, evidence, and accountable authority; it does not require every review task to retain its current human implementation. We develop a task-substitution framework for Digital Governance Frameworks (DGF), treating each gate as an executable contract. Substitution requires sufficient accessible information, valid decision and authority checks, and a reduction in total human work after exceptions, verification, correction, and maintenance are counted. We derive a residual-work threshold and show why automating most cases can still increase labor. Forward deployed engineering connects these conditions to an architecture for agents, rule engines, evidence services, and escalation. DGF-Bench supplies controlled evidence from 300 synthetic projects and 899 evaluable model-project runs. Gemini 3.8 Flash, GPT-5.6 Luna, and DeepSeek v4.1 Flash achieve strict gate success of 94.98%, 83.29%, and 74.18%; complete-route success is 76.92%, 42.33%, and 24.67%. A deterministic control passes all 1,700 gates given the supplied rules and structured facts, locating the comparison in execution of a supplied decision kernel. Evidence audits and 135 repeated runs distinguish correct decisions from reliable execution. A document counterexample establishes an information-sufficiency obstruction. These results support the technical feasibility of replacing human execution of specified governance-review tasks with agents and software. The framework specifies a workforce test based on the complete human effort required at fixed output and quality; the present measurements concern review performance. Sources, dossiers, traces, and analyses are public.","authors":["Jeremy Canale"],"categories":["cs.AI","cs.SE"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.29345","pdf_url":"https://arxiv.org/pdf/2609.29345","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["治理自动化","任务替代","多智能体系统"],"reason":"论文研究企业治理自动化，用LLM替代人工审核，属于多智能体任务执行，不涉及人类…","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:20","error":null,"has_summary":false,"summary":null},{"id":"2609.29381","version":1,"title":"An auditable conditional-strategy framework for open-ended decision-making in complex lung cancer","zh_title":"复杂肺癌开放式决策的可审计条件策略框架","abstract":"Complex lung cancer decisions can involve several defensible pathways whose eligibility, sequencing and safety depend on unresolved information. Effective support must make explicit how patient conditions govern pathway eligibility, deferral and redirection. MedGPT Clinical Explorer (MCE) organizes alternatives, decision-changing unknowns, safety constraints and fallback into a conditional strategy for clinician review. To evaluate this representation in physician-authored strategies, multidisciplinary experts established case-specific references for 40 cases within a purposive 100-case corpus, and 250 physicians from 98 institutions produced 2,250 strategies under unaided, retrieval-reference and MCE-assisted conditions. MCE-assisted strategies expressed more applicable clinical requirements, measured by the Admissible Pathway Attainment Score (APAS; 0-100), than unaided strategies (adjusted difference, 12.87; 95% CI, 11.18-14.55) and retrieval-reference strategies (5.22; 3.52-6.93). With the same knowledge base available in the retrieval-reference and MCE-assisted conditions, the additional content centered on candidate pathways, decision-critical information and safety constraints. Physicians' whole-strategy acceptability judgments correlated with APAS (Spearman's rho = 0.671), while a complementary relationship audit assessed whether candidates, conditions and subsequent actions were coherently connected. Together, these findings identify two complementary dimensions of open-ended decision support: coverage of clinically relevant content and coherent links among pathways, conditions and subsequent actions. MCE provides a shared decision object that makes consequential omissions and pathway contingencies visible before action; prospective studies should evaluate its effects on clinical workflow and patient outcomes.","authors":["Daoyun Wang","Zhicheng Huang","Huaiyuan Sun","Jiaqi Xu","Xiaowei Xu","Zhibo Zheng","Zhongxing Bing","Yuxiao Lin","Yicheng Liang","Chao Gao","Bowen Xue","Kai Zhang","Song Xu","Wanpu Yan","Hui Xia","Lin Li","Xiang Yan","Mu Hu","Qianli Ma","Zhiqiang Xue","Xiaofang Liu","Zhihai Han","Nan Zhang","Chuanhao Tang","Tongmei Zhang","Lan Song","Zhaohui Zhu","Xuan Zeng","Shafei Wu","Hui Guan","Lei Deng","Huaxia Yang","Zeliang Lian","Wubin Sun","Yongxin Wang","Xiaohui Shen","Binlin Wang","Tiantian Gu","Yu Cui","Li Zhang","Shirui Wang","Naixin Liang"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.29381","pdf_url":"https://arxiv.org/pdf/2609.29381","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C2"],"tags":["临床决策支持","LLM辅助","肺癌治疗"],"reason":"论文是临床决策支持系统，用LLM辅助医生制定肺癌治疗策略，不涉及用LLM仿真人…","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:22","error":null,"has_summary":false,"summary":null},{"id":"2609.28737","version":1,"title":"Policy Complexity, Reaction Time, and Bounded Rationality in Reinforcement Learning","zh_title":"强化学习中的策略复杂性、反应时间与有限理性","abstract":"Biological agents do not learn under conditions of unlimited computation. For humans, learning and choice are shaped by constraints on perception, attention, and working memory, which limit how much state information guides behavior and therefore bound policy complexity. Standard reinforcement learning models typically optimize reward without explicitly representing these internal costs, making them less suitable as models of biological intelligence. We derive MI-SARSA, an on-policy temporal-difference algorithm that incorporates mutual-information regularization through a learned marginal action prior and a penalty on state-specific deviations from that prior. This yields a sequential learning model in which state information is used selectively when its expected return benefit justifies the added informational cost. Critically, the same state-specific information cost that governs policy compression also generates trial-level predictions for reaction time, distinguishing MI-SARSA from most reinforcement learning models, which predict choices or returns but not latency. Empirically, MI-SARSA produces a reward-complexity tradeoff, and stronger information penalties produce simpler policies with lower control costs and faster reaction times. Under environment shift, increasing regularization reduces post-switch performance degradation but also lowers asymptotic return, revealing a robustness-capacity tradeoff. Together, these results position MI-SARSA as a model of bounded sequential learning under cognitive constraints.","authors":["James Wu","Chris R. Sims"],"categories":["cs.LG","cs.AI","cs.IT","math.IT"],"primary_category":"cs.LG","announce_type":"cross","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.28737","pdf_url":"https://arxiv.org/pdf/2609.28737","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C2"],"tags":["强化学习","有限理性","认知建模"],"reason":"研究强化学习算法，不涉及LLM仿真人类被试，无人类数据对照","model":"deepseek-v4-pro","scored_at":"2026-09-26T13:02:24","error":null,"has_summary":false,"summary":null},{"id":"2609.29995","version":1,"title":"Guardrails or Roadblocks? Effects of Pedagogical Style and Context Awareness in AI Teaching Assistants for Programming","zh_title":"护栏还是障碍？编程AI教学助手中教学风格与情境意识的影响","abstract":"AI teaching assistants (AI TAs) backed by large language models (LLMs) and pedagogical guardrails are increasingly being integrated into programming courses, providing students with scalable access to hints, conceptual explanations, and code-level feedback. However, guardrails may also create friction. If students feel that the support provided is overly restrictive or poorly contextualized to their current progress, they may bypass approved tools for general-purpose LLMs. To investigate how AI TA design affects students' learning experiences, we conducted a randomized controlled trial with 132 students in an introductory programming course. Students completed three tasks related to code-writing and debugging and were randomly assigned to one of four AI TAs varied across two dimensions: pedagogical guidance style (Socratic vs. Direct instruction) and context awareness (no context vs. full context of the problem and student solution). We examined students' perceptions, interaction behaviors, and evidence of post-task comprehension. Students rated the Socratic AI TA with full context least favorably, reporting significantly lower perceived support for task completion. Descriptively, this condition also showed the highest observed interaction stress, the highest rate of external LLM use, and the lowest proportion of post-task explanations demonstrating full comprehension, though these differences were not statistically significant. These findings suggest that guardrailed AI TAs are not automatically better for learning. Instead, their effectiveness depends on how pedagogical guidance and contextual awareness are balanced in ways that students experience as useful, supportive, and worth continuing to use.","authors":["Madeleine Eastwood","Harshith Narne","Joseph Hilby","Paul Denny","Ashish Aggarwal","Amanpreet Kapoor"],"categories":["cs.HC","cs.AI"],"primary_category":"cs.HC","announce_type":"cross","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.29995","pdf_url":"https://arxiv.org/pdf/2609.29995","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["AI教学助手","编程教育","随机对照试验"],"reason":"研究AI教学助手对学生学习的影响，属于教育技术评估，不涉及用LLM仿真人类被试…","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:30","error":null,"has_summary":false,"summary":null},{"id":"2609.30058","version":1,"title":"Can Labor Markets Function in the Age of AI? The Evaluation Bottleneck in Hiring","zh_title":"AI时代劳动力市场能否正常运转？招聘中的评估瓶颈","abstract":"AI-assisted job-search tools have become increasingly popular by making it easier to find and apply to jobs. But by making it easier for applicants to generate and tailor application materials, they can also reduce how informative those materials are about applicant fit. We study this tradeoff in a hiring market where applicants differ in experience and latent match quality and firms use noisy application materials to decide whom to screen. We ask how AI affects downstream screening and hiring, and which applicants are most adversely affected. As application materials become less informative, a Bayesian firm rationally relies more heavily on coarse observables such as prior experience. Among the four applicant types defined by experience and compatibility for the job, inexperienced-compatible applicants are the most exposed: they lack observable experience and lose the individualized information that could distinguish them from other inexperienced candidates. When screening is costly, these changes can also generate inefficient screening failures in which firms screen no applicants or screen only experienced applicants. We then show that multistage hiring can arise as an endogenous firm response: a relatively inexpensive intermediate assessment allows firms to acquire new evidence of fit before costly full screening. This can restore screening opportunities that disappear under one-stage hiring and give inexperienced-compatible applicants a path to screening. Our results show how AI can shift the central friction in hiring from submitting applications to obtaining credible evaluation, creating entry barriers for high-fit workers without prior experience. Multistage hiring can endogenously arise in response, restoring evaluation opportunities that would otherwise disappear and helping preserve market functioning.","authors":["Itai Ashlagi","Ramesh Johari","Jon Kleinberg","Anushka Murthy"],"categories":["cs.GT","cs.AI","cs.CY","econ.TH"],"primary_category":"cs.GT","announce_type":"cross","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.30058","pdf_url":"https://arxiv.org/pdf/2609.30058","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["AI招聘","劳动力市场","博弈论"],"reason":"研究AI对招聘市场的影响，无LLM仿真人类被试，不涉及人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:30","error":null,"has_summary":false,"summary":null},{"id":"2609.28801","version":1,"title":"The Interface Is Downstream: Designing the Terms of Human-Agent Collaboration","zh_title":"界面在下游：设计人机协作的条款","abstract":"Before an agent responds or acts, much of the experience has already been designed. Memory and retrieval shape what it notices. Evidence rules shape what it may claim. Permissions shape what it can do. Learning rules shape what it carries into the next encounter. The argument comes from Alicia, a personal agent I've built and used since January 2026. A fine-tuning pilot produced no defensible model-performance result. It exposed a provenance failure: Alicia repeated an interpretation from a retrieved synthesis, cited a source note credited by that synthesis, and left the synthesis out of the visible chain. In a model-blind review of thirty-three citations, one reviewer judged that the retrieved intermediary supplied the claim in sixteen relayed citations and part of it in four. Five of twelve citations to directly retrieved targets lacked support in the target excerpt supplied for review. These judgments remain unadjudicated, and the packet is not public. I call the shared setting a humorphic environment: a persistent computational setting that translates a human practice into software. The first Humorphism paper translated partnership. This paper translates the studio, the room where practice happens. The failure prompted an audit of attention, evidence, action, and learning, with a review artifact and available recourse for each. The test is whether the person can inspect and contest what shaped the teammate's behavior. The output is downstream. Correction, consent, and learning carry the collaborative interface back upstream.","authors":["Hector Ouilhet Olmos"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.28801","pdf_url":"https://arxiv.org/pdf/2609.28801","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["人机协作","界面设计","个人代理"],"reason":"论文讨论人机协作界面设计，非用LLM仿真人类被试，无实验或测量目的。","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:19","error":null,"has_summary":false,"summary":null},{"id":"2609.28886","version":1,"title":"Characterizing LLM-Based Family Education through the Lens of Activity Theory: A Scoping Review of the HCI Literature","zh_title":"从活动理论视角刻画基于大语言模型的家庭教育：HCI文献的范围综述","abstract":"Large language models (LLMs) are increasingly involved in family education, yet HCI has not systematically explained the educational interactions that emerge around them. This scoping review analyzes 53 HCI studies from 6,540 records across 19 venues. Using activity theory and AODM, it relates participants and educational objects to mediation, labour, and rules. We find that the literature centers on child--parent interaction and on language, AI literacy, and relational learning. The introduction of LLMs enabled conversational, embodied, and spatial systems to generate support from the context of an unfolding interaction. LLMs redistributed educational labour, while family and institutional rules left parents and professionals responsible for interpreting outputs and deciding how they entered practice. Evidence across families and educational purposes remains limited, especially on sustained personalization, repair labour, and how families negotiate authority and rules. The review offers a framework explaining how LLM capabilities become organized through family participation.","authors":["Lan Luo","Yuqi Liang","Jie Cai","Anqi Wang","Dongyijie Pan","Muzhi Zhou","Chun Yu","Pan Hui"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.28886","pdf_url":"https://arxiv.org/pdf/2609.28886","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["家庭教育","人机交互","文献综述"],"reason":"该论文是HCI文献综述，关注LLM在家庭教育中的交互设计，不涉及用LLM仿真人…","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:20","error":null,"has_summary":false,"summary":null},{"id":"2609.28990","version":1,"title":"How People Use ChatGPT in Australia: A WildChat Analysis","zh_title":"澳大利亚人如何使用ChatGPT：一项WildChat分析","abstract":"Generative AI chatbots are increasingly embedded in everyday life, yet most large-scale studies describe global patterns. This paper presents an Australia-focused analysis of WildChat, a public dataset of real-world ChatGPT interaction logs. Using descriptive analysis and a multi-layer classification scheme, we analysed 37,845 conversations identified as Australian, examining language diversity, work relevance, interaction intent, topic distribution, turn-taking, temporal change, work activities, and Australia-related domains. Our findings show that the Australian subset is strongly action-oriented and comparatively work-oriented, with most interactions classified as doing and a majority of conversations classified as work-related. The dataset also shows multilingual use and a growing presence of self-expression over time. Australia-related conversations frequently invoke local institutions, laws, regulators, education systems, companies, cultural references, and public services. Finally, we outline implications for future research, including local AI evaluation, multilingual participation, context-aware design, and safeguards for everyday high-stakes domains.","authors":["Ying Ma","Katy Gero","Cl\\'ement Canonne","Craig Jin","Kanchana Thilakarathna"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.28990","pdf_url":"https://arxiv.org/pdf/2609.28990","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["人机交互","使用模式","日志分析"],"reason":"分析真实用户与ChatGPT的对话日志，属于人机交互使用模式研究，不涉及用LL…","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:20","error":null,"has_summary":false,"summary":null},{"id":"2609.29689","version":1,"title":"Mapping the Authorized Boundary: A Comparative Policy-Vignette Study of Generative AI Governance in Australian Higher Education","zh_title":"划定授权边界：澳大利亚高等教育中生成式AI治理的比较政策情境研究","abstract":"Australian universities regulate students' use of generative artificial intelligence (GenAI) through overlapping policies, procedures, guidance, and assessment instructions, but these environments may classify identical conduct differently. We applied 15 standardized student-use vignettes to the public policy environments of 20 Australian universities and produced 300 university-case classifications. We analyzed binding instruments (Layer A) separately from the full official environment (Layer B), then classified each combination as clearly permitted, permitted with conditions, potential policy breach, clearly prohibited, or indeterminate. We measured cross-university divergence with normalized Shannon entropy. We classified 120 combinations (40.0%) as clearly prohibited, 97 (32.3%) as potential policy breaches, 27 (9.0%) as permitted with conditions, and 56 (18.7%) as indeterminate; none met the strict threshold for clearly permitted. Disclosed language rewriting and a disclosed AI-drafted paragraph produced the highest divergence, whereas an explicit assessment prohibition produced unanimity. Binding instruments remained silent on GenAI in 100 combinations, and guidance clarified 88. The primary researcher led the coding with AI assistance and manually reviewed all 201 queued rows. An independent human second coder, who used no AI assistance, coded a blind 75-row sample. Overall agreement reached 57.3% (unweighted Cohen's kappa = 0.395; bootstrap 95% CI [0.244, 0.539]). The findings distinguish permission-gated from disclosure-based architectures and show that written policies regulate the retention of AI-generated text more clearly than process-only assistance does. Comparative policy-vignette testing evaluates whether written governance supports defensible classifications; it does not predict misconduct or enforcement decisions.","authors":["Biranchi Poudyal"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.29689","pdf_url":"https://arxiv.org/pdf/2609.29689","source_feed":"cs.CY","score":0,"bucket":"other","rubric_hits":["C5"],"tags":["AI治理","高等教育政策","政策分析"],"reason":"研究大学政策对GenAI使用的分类，不涉及用LLM仿真人类被试，而是分析政策文…","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:27","error":null,"has_summary":false,"summary":null},{"id":"2609.29819","version":1,"title":"Fair Feed Ranking for Participatory Budgeting","zh_title":"参与式预算的公平信息流排序","abstract":"In large-scale participatory budgeting, citizens cannot inspect the full proposal pool, so the order in which proposals are shown becomes a form of agenda-setting power. We argue that fair exposure should therefore be treated as a democratic-design goal. We study Consul Democracy, a widely deployed open-source digital-democracy platform, and show that its proposal feeds are typically ordered by popularity, recency, or comment activity. Building on this diagnosis, we propose FairFeed, a feed-ranking design for PB that uses transparently declared preferences, boosts under-exposed proposals, and admits a rate-limited reject channel for crowd-sourced vetting. We evaluate the design in a simulation anchored in Munich's 2025 PB process and compare it with random, newest, and most-commented feeds. In this simulation, FairFeed broadens proposal discovery, distributes visibility more evenly across the eligible pool, increases cross-cutting support, and improves resistance to manipulation relative to comment-based ranking. We conclude by outlining the human-subjects evaluation needed to test whether onboarding can recover voter preferences accurately enough for deployment in practice.","authors":["Carina I. Hausladen"],"categories":["cs.CY","cs.IR"],"primary_category":"cs.CY","announce_type":"new","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.29819","pdf_url":"https://arxiv.org/pdf/2609.29819","source_feed":"cs.CY","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["参与式预算","排序算法","民主设计"],"reason":"论文研究参与式预算的排序算法，仿真用于评估算法性能，不涉及LLM仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-09-25T13:02:29","error":null,"has_summary":false,"summary":null},{"id":"2609.28908","version":1,"title":"Automatic Harness Evolution for Hardware Design Verification: Can LLMs Consolidate Gains Across Discovered Harnesses?","zh_title":"硬件设计验证中的自动Harness演化：LLM能否巩固跨发现Harness的收益？","abstract":"Agent behavior depends on the harness surrounding a language model, but it remains unclear whether language models can reliably improve such harnesses for hardware-design tasks. We study automatic harness evolution around a fixed subject model on 12 proprietary design-verification root-cause localization tasks. Across five trials per task, automatically evolved harnesses increased completed attempts by 71-76% and any-hit task coverage by 80-100%, while total correct attempts improved by only 18-24%. The strongest success reproducible at least twice result improved by one task, and later candidates exchanged gains across tasks rather than preserving them. An auxiliary candidate improved on a four-task validation set excluded from search but tied its baseline on a subsequent 12-task replay containing both search and validation tasks, so the selected gain did not persist across the full pool. Across the tested lineage, useful search, evidence, and finalization behaviors appeared in different candidates but did not consistently consolidate into a single harness that dominated across tasks and metrics. In a separate CVDP cross-benchmark case study, an automatically evolved defined-width repair harness produced 35.6% more functional passes than its 142-task reference baseline; the final functional verifier scored completed outputs but was not shown to the subject agent during repair. These results support archive-aware selection when evolution yields complementary specializations without consistent consolidation.","authors":["Kidus Seyoum","Ajay Mittur"],"categories":["cs.SE","cs.LG"],"primary_category":"cs.SE","announce_type":"cross","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.28908","pdf_url":"https://arxiv.org/pdf/2609.28908","source_feed":"cs.LG","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["硬件验证","LLM agent","自动演化"],"reason":"研究LLM自动改进硬件验证harness，属多智能体协作解题，不涉及人类行为仿…","model":"deepseek-v4-pro","scored_at":"2026-09-26T13:02:24","error":null,"has_summary":false,"summary":null},{"id":"2609.29016","version":1,"title":"EvoTreeNAD: Genealogy-Guided Evolution for LLM-Driven Neural Architecture Discovery","zh_title":"EvoTreeNAD：谱系引导的进化算法用于LLM驱动的神经架构发现","abstract":"AI-driven scientific discovery accelerates research by autonomously developing solutions and designs. Large language model (LLM) agents support this process through iterative generation and evaluation. Yet these iterations alone do not ensure cumulative progress or establish which directions to pursue next. Costly evaluation further constrains the scope of exploration. Neural architecture discovery brings these challenges together, coupling open-ended design with resource-intensive experimentation. We introduce EvoTreeNAD, a genealogy-guided evolutionary algorithm that constructs trainable architectures without a supplied seed or a hand-specified search space. Starting from an empty root, it grows a persistent genealogy in which each new node represents a complete architecture. Top-percentile values computed from each node and its descendants guide lineage selection. Using the selected design history, an Idea Agent proposes a variant and a Code Agent implements it. Each evaluated variant becomes a child node, expanding the genealogy while providing evidence for subsequent lineage selection. Our theoretical analysis establishes the existence of stationary variation regimes as the genealogy grows. Under specified variation assumptions, sustained top-percentile family values quantify the probability of generating high-reward architectures in these regimes. EvoTreeNAD discovers architectures that outperform the compared NAS and NAD baselines, achieving CIFAR-10/100 test errors of $2.05{\\pm}0.06\\%$ and $15.09{\\pm}0.22\\%$. On all six MedMNIST-v2 tasks, the discovered architectures surpass the strongest listed baselines. A controlled CIFAR-10 study further shows that EvoTreeNAD outperforms direct generation, best-of-$N$ greedy continuation, and full-family-mean routing.","authors":["Lishan Yu","Derek Jiu","Qizhen Lan","Xiaoqian Jiang"],"categories":["cs.NE","cs.LG"],"primary_category":"cs.NE","announce_type":"cross","date":"2026-09-25","first_seen":"2026-09-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.29016","pdf_url":"https://arxiv.org/pdf/2609.29016","source_feed":"cs.LG","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["神经架构搜索","多智能体系统","进化算法"],"reason":"多智能体协作解决神经架构搜索，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-26T13:02:24","error":null,"has_summary":false,"summary":null},{"id":"2609.27535","version":1,"title":"KITE: Scaling Jev Population Experiments with Sparse Flagship Calibration","zh_title":"KITE：通过稀疏旗舰校准扩展Jev人口实验","abstract":"KITE queries a typed behavioral kernel once per unique state, then executes populations of any size from the table with event-keyed randomness and common random numbers. An expensive flagship model is reserved for sparse paired anchors that estimate intervention effects. Measured human-model discrepancy is propagated as shared error into every conclusion. Population-experiment cost thus scales with unique states and anchors, while uncertainty is governed by evidence about people rather than Monte Carlo noise. On Epstein experiments with 9,070 participants, anchors covering 1.7% of states reduced effect error by 41% (absolute MAE reduction 0.0125). On 37 held-out SocSci210 experiments, 0.5-1.5% anchor coverage raised captured decision gain from 0.27 to 0.39. The kernel passed content-fidelity criteria in all 15 new countries of a 16-country study. Shared discrepancy yielded retrospective coverage of 93% and 96% at nominal 80% and 90%, versus 29% and 36% from human sampling uncertainty alone. A million agents executed 20 tabulated steps in 0.9 seconds on a laptop. This architecture offers a route to screening candidate interventions before human trials, multi-country content audits, and uncertainty-aware policy comparison at the cost of a few thousand kernel calls with sparse flagship anchors. Property-specific evidence records connect each use to its validation scope, correction provenance, and uncertainty, making these applications auditable.","authors":["Hengyu Li (The University of Tokyo)"],"categories":["cs.MA","cs.CY"],"primary_category":"cs.MA","announce_type":"cross","date":"2026-09-24","first_seen":"2026-09-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.27535","pdf_url":"https://arxiv.org/pdf/2609.27535","source_feed":"cs.CY","score":9,"bucket":"selected","rubric_hits":["A1","A2","A3","B1","B2","B3","B4"],"tags":["LLM仿真","人类数据对照","政策评估"],"reason":"用LLM仿真人类被试，有真实人类数据对照，评估偏差并传播不确定性，用于政策评估。","model":"deepseek-v4-pro","scored_at":"2026-09-24T13:01:30","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-24","rank":5,"question":"如何在大规模人口实验中，用廉价的类型化行为核模型结合稀疏的旗舰模型校准，来估计干预效应并传播人类-模型差异的不确定性？","design":"KITE 使用 TypeSafe 的 Jev 行为核模型（类型化行为核）对每个唯一状态查询一次，生成决策分布，然后用表格化执行模拟任意规模的人口，并通过事件键控随机数和共同随机数控制变异性；昂贵的旗舰模型仅用于稀疏的配对锚点，以估计干预效应；人类-模型差异被作为共享误差传播到所有结论中。","baseline":"对照的真实人类数据包括 Epstein 实验（9,070 名参与者）和 37 个留出的 SocSci210 实验，以及一个 16 国研究中的内容保真度标准。","findings":"在 Epstein 实验中，覆盖 1.7% 状态的锚点将效应误差降低了 41%（绝对 MAE 降低 0.0125）；在 37 个留出的 SocSci210 实验中，0.5-1.5% 的锚点覆盖率将捕获的决策增益从 0.27 提高到 0.39。共享差异传播在名义 80% 和 90% 的置信水平下分别实现了 93% 和 96% 的回顾性覆盖率，而仅使用人类抽样不确定性时分别为 29% 和 36%。","reliability":"论文承认的局限包括：仅有两个留出的人类参考数据集测试混合方法；效应幅度需要进一步校准，差异斜率在不同研究选择间变化；稀疏校准可能引入旗舰模型误差；核模型可能过度预测规范信息；实验不授权预测文献中不存在的干预；记忆一致性是局部的，队列级边际校正可能抹去真实的持续性或处理路径；刺激重建、省略卡片图像、人口统计压缩等限制了保真度；专有模型和访问限制阻碍了可重复性。","relevance":"该研究直接针对用 LLM 仿真人类被试的核心问题，提供了与真实人类数据对照的验证，并传播不确定性，对评估仿真可靠性和偏差具有重要参考价值，值得精读原文。","inspiration":"该方法通过稀疏旗舰模型校准和共享误差传播，在保持低成本的同时提高了效应估计的准确性，值得借鉴其校准策略和不确定性量化方法。｜可迁移到政策评估场景，如税收政策变化对劳动供给的影响、福利项目对消费行为的影响，或信息干预对金融决策的影响。｜设计一个实验：用 Jev 核模型模拟不同人口群体对政策公告的反应，以真实调查数据（如消费者预期调查）为基准，施加政策处理（如利率变化信息），测量预期调整和消费意愿，并用稀疏旗舰模型校准关键状态，传播人类-模型差异。"}},{"id":"2609.28470","version":1,"title":"StudentBench: AI and human tutoring yield equivalent GRE learning gains","zh_title":"StudentBench：AI与人类辅导在GRE学习收益上等效","abstract":"Artificial intelligence offers an unprecedented opportunity to augment human capabilities, yet progress at the frontier has focused primarily on advancing model capabilities. We introduce StudentBench, a suite of AI teaching evaluations and a public platform that enables large-scale data collection with over 175,000 student-AI messages to study whether large language models (LLMs) produce learning gains equivalent to human tutoring. Using StudentBench, we measured learning gains on Quantitative and Verbal GRE questions across 2,383 human participants receiving AI tutoring, human tutoring, or no tutoring. We establish that AI tutoring is statistically equivalent to expert human tutoring for GRE learning gains (p = .015), and in five of the seven GRE domains, the best performing AI tutor surpassed the human tutor, on average. In a second study, expert human tutors compared LLM-generated lesson plans and practice problems through 2,028 pairwise rubric evaluations. Together, the two studies clearly separate AI tutors across: (1) lesson planning, (2) practice-problem creation, (3) conversational pedagogy, (4) cost, and (5) engagement. Surprisingly, one AI tutor achieved learning gains equivalent to human tutoring (p = .044) at 918 times lower cost (USD 0.0052 for AI versus USD 4.81 for human, per percentage point gained). For Quantitative GRE sessions, faster AI replies correlated with more student messages, more messages with more correct practice, and more correct practice with larger learning gains (all p < .002). The StudentBench platform is freely available at https://studentbench.org.","authors":["Curtis Northcutt","Inaara Hasmani","Kevin Feng","Trevor Khangi","Andreas Plesner","Jonas Mueller"],"categories":["cs.AI","cs.CY"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-09-24","first_seen":"2026-09-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.28470","pdf_url":"https://arxiv.org/pdf/2609.28470","source_feed":"cs.CY","score":9,"bucket":"selected","rubric_hits":["A1","B1","B2"],"tags":["LLM仿真","教育实验","人类对照"],"reason":"用LLM替代人类导师进行教学实验，并与人类导师对照，评估学习效果，属于人类仿真…","model":"deepseek-v4-pro","scored_at":"2026-09-24T13:01:52","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-24","rank":6,"question":"大语言模型（LLM）作为AI导师能否在GRE学习收益上达到与人类专家导师统计等效的效果？","design":"用多个LLM（如Gemma 4 31B等）扮演AI导师，对2383名人类参与者进行GRE定量和语文部分的辅导，同时设置人类导师辅导组和无辅导对照组，测量学习收益（前后测成绩提升百分比），并收集175,000条学生-AI消息分析对话行为。","baseline":"人类专家导师辅导组的学习收益数据，以及无辅导对照组的学习收益数据。","findings":"AI辅导在GRE学习收益上与人类专家辅导统计等效（p=.015），且在七个GRE领域中五个领域的最佳AI导师平均超过人类导师。一个AI导师（Gemma 4 31B）以低918倍的成本实现了与人类辅导等效的学习收益（p=.044）。","reliability":"论文承认局限：只测量即时学习收益，未评估长期保持；参与者均为能读写英语的成年人，未测试跨语言、设备或教育环境；未设置学生独自练习的对照组，无法分离AI交互的额外收益；排行榜评估基于专家评价和对话行为，而非实际学习效果。","relevance":"该研究用LLM替代人类导师进行教学实验，并与真实人类导师及无辅导组对照，评估学习效果，属于典型的人类仿真研究，且提供了大规模真实人类数据作为基准，值得精读以了解仿真等效性检验的设计与局限。","inspiration":"借鉴其多组对照设计（AI处理、人类处理、无处理）和统计等效性检验（TOST）来严格评估AI干预是否非劣于人类专家。｜可迁移到金融教育或投资者决策辅导场景，例如测试AI投教助手能否在提升投资者金融素养或改善投资决策上达到人类理财顾问的效果。｜以真实投资者为被试，随机分为AI投教组、人类顾问组和无辅导组，处理为一段时间的个性化金融知识辅导，结果变量为金融素养测试得分或模拟投资组合表现，对照真实人类顾问组和无辅导组的数据，并采用等效性检验。"}},{"id":"2609.28372","version":1,"title":"Shopping by algorithm: How agentic AI deploys human heuristics as a surrogate consumer","zh_title":"算法购物：代理式AI如何将人类启发式用作替代消费者","abstract":"Consumers increasingly delegate purchasing decisions to Large Language Models (LLMs) acting as surrogate consumers. Using \"Tool-Lab,\" an adaptation of information-board process tracing that places product attributes behind costly tool calls, we examine how marketing pricing cues (i.e., just-below pricing and promotional framing) influence AI shopping agents. Across eight commercially deployed LLMs from three providers, we trace pre-choice information acquisition. Under zero cost, pricing cues rarely mislead. Imposing acquisition costs under a vague goal prompt leads LLMs to omit diagnostic attributes required to compute unit price and choose suboptimal choices resembling human heuristics. Relative to a specific goal prompt that mainly preserves diagnostic search and choice optimality, a vague goal prompt under constraints creates a search-mediated vulnerability. This research demonstrates that marketing heuristics in delegated AI shopping are governed by storefront information architecture, not necessarily immutable LLM flaws.","authors":["Davood Wadi","Yu Ma"],"categories":["econ.GN","cs.AI","q-fin.EC"],"primary_category":"econ.GN","announce_type":"new","date":"2026-09-24","first_seen":"2026-09-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.28372","pdf_url":"https://arxiv.org/pdf/2609.28372","source_feed":"econ.GN","score":8,"bucket":"selected","rubric_hits":["A1","B1","B2","B4"],"tags":["LLM仿真","消费者行为","算法保真度"],"reason":"用LLM作为代理消费者模拟人类购物决策，并与人类启发式对照，涉及营销实验场景。","model":"deepseek-v4-pro","scored_at":"2026-09-24T13:01:32","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-24","rank":9,"question":"在委托AI代理购物时，营销定价线索（如尾数定价和促销框架）如何通过信息获取成本与目标提示的具体性影响LLM的信息搜索和选择最优性？","design":"使用Tool-Lab实验范式，将产品属性隐藏在需要付费的工具调用之后，操纵信息获取成本（0、1、5美分）和提示目标具体性（模糊：“找最划算的” vs. 具体：“找每盎司最低价”），对8个商用LLM（来自Google、OpenAI、Anthropic）进行咖啡选择实验，测量信息搜索深度、搜索组成和选择最优性。","baseline":"无对照","findings":"在零成本或具体目标提示下，LLM大多做出规范最优选择；但在获取成本存在且目标模糊时，多数LLM会减少搜索深度，省略诊断性属性（如美分或重量），导致次优选择，类似于人类启发式决策。","reliability":"论文未讨论","relevance":"本研究通过实验操纵环境约束（成本与提示），揭示了LLM启发式行为的条件性，为评估LLM作为人类被试替代品的可靠性提供了关键证据，值得精读原文以理解其方法细节和边界条件。","inspiration":"借鉴其通过工具调用施加信息获取成本并操纵提示具体性的设计，可迁移到消费者金融决策（如信用卡选择、贷款比较）或投资者信息处理场景；例如，用LLM模拟投资者在获取公司财务指标需付出成本时，模糊目标（“选只好股票”）与具体目标（“选市盈率最低的股票”）下的信息搜索与选择，并与真实投资者眼动或点击流数据对照。"}},{"id":"2608.22859","version":2,"title":"WARP: Wasserstein-Aligned RAG for Population Opinions","zh_title":"WARP：面向群体意见的Wasserstein对齐检索增强生成","abstract":"RAG systems are increasingly used to summarize what large collections of documents say. A user asks \"What do people think about X?\" and receives an answer that reads as consensus. But standard top-k retrieval ranks documents by query similarity, not by how faithfully they represent the population, so minority views quietly disappear. Existing fixes fall short. Diversity re-rankers like MMR and DPP spread retrieved documents apart, but with no target distribution to aim for. Calibration methods based on KL or JS divergence do target one, yet treat opinion bins as unordered: confusing strong positive with strong negative costs no more than an adjacent-bin miss. We introduce WARP, a family of post-retrieval algorithms that calibrate retrieved evidence to the population's opinion distribution. WARP first recovers underrepresented opinions that cosine ranking may bury, then uses Wasserstein-1 distance to select documents whose sentiment-intensity distribution matches the population target, capturing the ordinal structure ignored by KL and JS divergence. We develop three variants for dense, sparse, and variable candidate pools, trading off calibration quality and speed. Across three review domains spanning 35K documents, 156 queries, and 26 entities, WARP's domain-matched variants reduce distributional error by at least 43% with sub-second latency. These gains carry through to generation: a five-judge LLM panel prefers WARP-generated answers in 86% of decided comparisons at k <= 5.","authors":["Aman Singh Thakur","Aditya Agrawal","Alwarappan Nakkiran","Alex Karlsson"],"categories":["cs.IR","cs.CL"],"primary_category":"cs.IR","announce_type":"replace-cross","date":"2026-09-24","first_seen":"2026-08-25","revised_at":"2026-09-24","abs_url":"https://arxiv.org/abs/2608.22859","pdf_url":"https://arxiv.org/pdf/2608.22859","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A1","B1"],"tags":["LLM仿真","意见分布校准","RAG"],"reason":"用LLM生成代表人群意见的摘要，并与真实意见分布对齐，属于仿真人类态度，且有真…","model":"deepseek-v4-pro","scored_at":"2026-09-24T13:01:53","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-24","rank":10,"question":"如何让检索增强生成（RAG）系统在回答关于人群意见的查询时，所选证据的意见分布忠实于总体人群的意见分布，避免多数意见淹没少数意见？","design":"该论文提出 WARP 算法族，用于在 RAG 流程的后检索阶段校准所选文档的意见分布。它首先通过缺陷感知的池扩展策略恢复被余弦相似度排序埋没的少数意见文档，然后使用 Wasserstein-1 距离作为度量，贪婪地选择文档子集，使其情感强度分布与目标人群分布匹配。论文在三个评论数据集（Amazon 卖家论坛、Yelp 酒店评论、OpinRank 汽车评论）上进行了实验，共 35K 文档、156 个查询、26 个实体，比较了 WARP 与 Top-k、MMR、DPP、OpinionMMR、KL/JS 校准等基线在分布误差、实体匹配率和延迟上的表现，并通过 LLM 评委小组评估生成答案的质量。","baseline":"论文使用从语料库中提取的实体级情感强度分布作为目标人群分布，该分布由 LLM 对每个文档进行情感强度标注后聚合得到，作为真实人群意见分布的代理。","findings":"WARP 的领域匹配变体在三个评论领域中将分布误差降低了至少 43%，且重排序延迟低于 310 毫秒（p99）。在生成评估中，五位 LLM 评委在 k≤5 时对 WARP 生成答案的偏好率高达 86%。","reliability":"论文指出，WARP 的性能依赖于实体密度：在实体稀疏的领域，需要混合变体（如 λ-MMR 或 WassRank OT）来平衡相关性与校准；此外，目标分布是从语料库中估计的，可能无法完全代表真实人群，但论文未深入讨论这一局限。","relevance":"该论文直接针对用 LLM 生成代表人群意见的摘要问题，通过 Wasserstein 距离对齐意见分布，属于仿真人类态度分布的研究，且有真实评论数据作为基准，值得阅读原文以了解其算法细节和评估方法。","inspiration":"该方法借鉴了将检索证据校准到目标分布的思想，并利用 Wasserstein 距离捕捉有序情感结构，可用于经济金融领域中需要从文本中提取群体意见或情绪分布的场景。｜可以迁移到消费者信心指数构建、财报电话会议情绪分析、社交媒体上的政策预期形成等场景，其中需要从大量文本中估计人群的态度分布。｜一个可行的研究设计是：以 Twitter 上关于某经济政策（如加息）的推文为语料，用 LLM 对每条推文进行情感强度标注（如从强烈反对到强烈支持），得到目标分布；然后使用 WARP 算法从检索到的推文中选择一小部分作为 LLM 生成摘要的证据，确保所选推文的情感分布与总体分布一致；最后将生成的摘要与基于随机抽样或简单检索的摘要进行对比，评估其在预测真实调查数据（如密歇根消费者信心指数）上的准确性。"}},{"id":"2609.27165","version":1,"title":"Count Evidence, Not Sentences: Tempered Evidence Fusion of LLM Judgments for Long-Text Value Measurement","zh_title":"计数证据而非句子：面向长文本价值测量的LLM判断调和证据融合","abstract":"Large language models (LLMs) are increasingly used to measure public value orientations from long social media posts, yet such posts often mix background, quotations, concessions, and only a few stance-bearing sentences. Existing approaches either ask the model to predict a document-level label directly, which can be overconfident, or aggregate sentence-level predictions by majority or soft voting, which treat uncertain and decisive sentences as equally informative. We formulate long-text value measurement as a decision-fusion problem and propose Tempered Evidence Fusion (TEF), a training-free rule that weights each sentence's log-odds by its normalized information gain, as derived from a generalized Bayesian posterior. This makes the fused score nearly vanish for uncertain sentences while preserving the Bayes-optimal weight of decisive evidence. We further introduce Multi-event Insight Network Dimensions (MIND), a benchmark of 8,358 Chinese and English posts spanning five years of public events and six value dimensions. On MIND, TEF outperforms the strongest baseline among Direct, Majority Vote, and Soft Vote by an average of 4.5 accuracy points and 4.6 macro-F1 points across five LLMs and two languages. MIND dataset and code are available at https://github.com/Kzczc/ICASSP2027-TEF.","authors":["Yuhe Wu","Rui Qian","Guangyu Wang","Yuran Chen","Yuanchao Zhu","Junjie Yang","Zhengheng Li","Jiulin Cai","Tianyi Zhang","Zihan Dong","Jiaxin Liu","Yujie Chen","Guang Zhang"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-24","first_seen":"2026-09-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.27165","pdf_url":"https://arxiv.org/pdf/2609.27165","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A1","B1"],"tags":["LLM价值测量","证据融合","长文本分析"],"reason":"用LLM测量公众价值取向，有真实人类标注数据对照，方法可迁移到仿真研究。","model":"deepseek-v4-pro","scored_at":"2026-09-24T13:01:28","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-24","rank":13,"question":"如何融合长文本中句子级LLM判断，使决定性证据主导最终价值取向测量，同时抑制不确定句子？","design":"本研究不是人类仿真实验，而是提出一种训练无关的决策融合规则TEF，用于融合句子级LLM判断来测量长社交媒体帖子的价值取向。具体做法：将长文本分割为句子，用LLM对每个句子输出两个立场标签的对数概率，计算对数几率，并用归一化信息增益加权，再求和得到文档级得分。在MIND基准（8,358条中英文帖子，六个价值维度）上，用五个LLM（Qwen2.5-7B、LLaMA3-8B、Qwen3-14B、DeepSeek-V3.2、GPT-4o-mini）评估TEF与直接预测、多数投票、软投票的性能。","baseline":"有真实人类标注数据：MIND基准包含8,358条中英文社交媒体帖子，由人类标注了六个价值维度上的立场标签。","findings":"TEF在五个LLM和两种语言上平均比最强基线（直接预测、多数投票、软投票）高出4.5个准确率点和4.6个宏F1点。TEF在Qwen2.5-7B上校准误差最低，且对提示词扰动更稳健。","reliability":"论文未明确讨论失效条件，但指出直接求和LLM logits不可靠，因为next-token概率只是近似校准，尤其在最大不确定点附近；消融实验显示，在对数几率变换和熵加权同时移除时，某些模型上性能会低于软投票，表明两者互补。","relevance":"该研究虽非人类仿真实验，但提供了用LLM测量公众价值取向的可靠方法，且有真实人类标注数据对照，可作为仿真研究中测量态度或价值观的工具，值得阅读原文了解融合规则细节。","inspiration":"借鉴其决策融合思路，将长文本分解为句子级判断并用信息增益加权融合，可提高LLM对复杂文本的测量准确性｜可迁移到经济金融领域的文本测量，如从财报电话会议记录中提取管理层情绪、从新闻中测量政策不确定性、从社交媒体帖子中测量消费者信心或通胀预期｜设计：用LLM对财报电话会议记录的每个句子判断管理层语气（积极/消极），用TEF融合句子级判断得到文档级情绪得分，以分析师一致预期或后续股票收益作为真实数据对照，评估LLM情绪测量的预测效度。"}},{"id":"2609.26861","version":1,"title":"Rule-Based Pricing Algorithms and Market Outcomes: An Experimental Study","zh_title":"基于规则的定价算法与市场结果：一项实验研究","abstract":"Rule-based pricing tools are widespread in digital commerce, yet we know little about how their design shapes market outcomes. In a controlled market experiment, participants use dashboards to build pricing algorithms competing in a sequential Bertrand game over multiple periods. We vary design features commonly found in commercial repricing tools: warnings about price wars, pre-configured strategies, and advice from a large language model. Most treatment variations raise market prices with effects driven by an increase in starting prices and more cooperative algorithm designs. The results matter for competition policy, platform regulation and current discussions on regulating algorithm design tools.","authors":["Adrian Hillenbrand","Hans-Theo Normann","Matthias Potarca","Tobias Werner"],"categories":["econ.GN","q-fin.EC"],"primary_category":"econ.GN","announce_type":"new","date":"2026-09-24","first_seen":"2026-09-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.26861","pdf_url":"https://arxiv.org/pdf/2609.26861","source_feed":"econ.GN","score":7,"bucket":"pending","rubric_hits":["A3","B1","B2"],"tags":["LLM建议","经济实验","算法定价"],"reason":"用LLM提供建议影响人类定价实验，有真实人类数据对照，属经济实验场景，可迁移到…","model":"deepseek-v4-pro","scored_at":"2026-09-24T13:01:28","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-24","rank":12,"question":"规则型定价算法的设计特征（价格战警告、预配置策略、LLM建议）如何影响市场结果与合谋行为？","design":"受控市场实验：人类被试通过仪表盘构建定价算法，在序贯伯特兰博弈中竞争50期，共5轮超级博弈；处理为三种设计特征：盈利性提示（警告价格战）、LLM建议、预配置策略菜单与默认合作性价格匹配；结果变量为市场价格、起始价格、算法合作性设计。","baseline":"基线处理（无设计特征）作为对照，比较不同处理组与基线组的价格和算法设计差异。","findings":"大多数处理变体提高了市场价格，效应主要由起始价格提高和更合作的算法设计驱动。LLM建议和预配置策略等设计特征可能促进合谋，对竞争政策和平台监管有启示。","reliability":"论文未讨论","relevance":"该研究用LLM提供建议影响人类定价实验，有真实人类数据对照，属经济实验场景，可迁移到LLM仿真人类决策的可靠性评估，值得读原文了解LLM建议的具体效果。","inspiration":"借鉴其将LLM建议作为处理变量嵌入人类实验、并观察对策略选择和均衡结果影响的设计｜可迁移到资产定价实验或政策公告预期形成场景，研究LLM建议对投资者行为或公众预期的影响｜以人类被试为对象，处理为是否提供LLM投资建议，结果变量为报价或预期值，对照真实市场数据或调查数据评估LLM建议的偏差与可靠性"}},{"id":"2609.27639","version":1,"title":"Agent-based Modeling: Equilibrium, Echo Chambers, and Efficiency in Hybrid Coevolutionary Opinion Games","zh_title":"基于智能体的建模：混合协同演化观点博弈中的均衡、回音室与效率","abstract":"Opinion formation in online networks involves changes in both beliefs and social ties. Analytical models make it possible to study equilibrium and social cost, but usually represent communication as a fixed numerical update. LLM-driven agents offer a language-based alternative, yet their convergence and collective efficiency remain unclear. We develop the Hybrid Coevolutionary Opinion Game (H-COG), combining cost-minimizing Friedkin-Johnsen agents (Type-C) and Phi-4 language agents (Type-L) in a dynamically rewired K-nearest-neighbor network. We initialize 50 agents with opinions drawn from 5,199 Reddit comments on gun control and abortion. The comments are scored on a continuous [-1,+1] scale using a fine-tuned RoBERTa regressor, and a mixing parameter sets the proportion of each agent type. The experiments cover nine population compositions, three initial network topologies, and two topics. All 540 runs meet the convergence criterion within the simulation horizon. Under Type-L updating, the coevolving network reaches an attractor as reliably as it does under the analytical update rule, making an equilibrium-based efficiency comparison possible. The pooled Price of Anarchy is $5.558 \\pm 0.309$ for purely Type-L populations, compared with $1.139 \\pm 0.005$ for purely Type-C populations. A decomposition of social cost attributes most of this gap to language agents moving away from their intrinsic opinions, rather than to greater disagreement with their neighbors. The main findings are consistent across the three initial network topologies.","authors":["Ming-Zhi Jiang","An-Tzi Teng","Jun-En Liu","Po-An Chen","Yung-Ming Li"],"categories":["cs.GT","cs.MA","cs.SI"],"primary_category":"cs.GT","announce_type":"cross","date":"2026-09-24","first_seen":"2026-09-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.27639","pdf_url":"https://arxiv.org/pdf/2609.27639","source_feed":"cs.MA","score":6,"bucket":"other","rubric_hits":["D3"],"tags":["LLM智能体","观点动力学","社会模拟"],"reason":"用LLM agent模拟观点动态，但无真实人类行为对照，属社会模拟边界情形","model":"deepseek-v4-pro","scored_at":"2026-09-24T13:01:30","error":null,"has_summary":false,"summary":null},{"id":"2609.27686","version":1,"title":"Mining Meaning: Measurement Error in AI-Assisted Literature Reviews","zh_title":"挖掘意义：AI辅助文献综述中的测量误差","abstract":"Researchers increasingly use generative AI, particularly large language models (LLMs), to automate tasks across the research pipeline. We study the reliability of these tools at the reading, classification, and synthesis of large bodies of academic literature. We frame LLM-assisted literature reviews as a measurement problem, treating models as measurement systems and tracing how their errors affect downstream conclusions. As a test case, we use three different implementations of ChatGPT to identify and extract metadata from economics papers that use rainfall as an instrumental variable. We benchmark each implementation against a subset of human-labeled evaluation data, and then deploy those implementations to extract metadata from the full corpus. The LLMs perform well on binary classification, but performance deteriorates as tasks demand greater contextual interpretation. More importantly, how much researchers can rely on model outputs depends not only on the complexity of the reading task but also on the type of claims the data is asked to support. The same amount of measurement error substantially affects paper-level claims while having little effect on broader claims about the literature. Measurement error in LLM-generated data is thus most consequential at precisely the level of detail that constitutes an LLM's principal value added over human reviewers. We conclude that standard model performance metrics are informative about the quality of generated data but do not by themselves establish the credibility of downstream inference. Researchers must also evaluate whether substantive claims are robust to the measurement system used to generate the underlying data.","authors":["Jeffrey D. Michler","Kieran Douglas","Anna Josephson"],"categories":["econ.GN","q-fin.EC"],"primary_category":"econ.GN","announce_type":"new","date":"2026-09-24","first_seen":"2026-09-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.27686","pdf_url":"https://arxiv.org/pdf/2609.27686","source_feed":"econ.GN","score":6,"bucket":"other","rubric_hits":["D1"],"tags":["LLM标注","测量误差","文献综述"],"reason":"LLM替代人工标注文献，属标注员替代而非仿真被试，但涉及测量误差与可靠性，可迁…","model":"deepseek-v4-pro","scored_at":"2026-09-24T13:01:43","error":null,"has_summary":false,"summary":null},{"id":"2608.06968","version":2,"title":"How a shared state is described determines whether AI agents synchronize","zh_title":"共享状态的描述方式决定AI智能体是否同步","abstract":"Language-model agents increasingly act in populations, where the outcome that matters is collective: whether they align, split or fail to coordinate. Each acts not on the world but on a text description of it, a choice usually fixed in software. Using synchronization, the canonical probe of how interaction rules produce collective order, we show that this choice can decide the outcome. Agents on a circle chose to advance, stay or move back after reading the others' relative positions, in 507,112 valid responses across matched populations, controlled inputs and three model families. In GPT, numerical summaries aligned every matched population at both positive couplings, whereas histograms aligned none; Claude showed the reverse at the stronger coupling. Re-describing identical states shifted action probabilities in all three families, even between histograms carrying the same information. No single directional coefficient explained the outcome: state descriptions are part of the interaction rule that turns individual responses into collective order.","authors":["Takahiro Ezaki","Naoto Imura","Katsuhiro Nishinari"],"categories":["physics.soc-ph","cs.AI","cs.CY"],"primary_category":"physics.soc-ph","announce_type":"replace-cross","date":"2026-09-24","first_seen":"2026-08-10","revised_at":"2026-09-24","abs_url":"https://arxiv.org/abs/2608.06968","pdf_url":"https://arxiv.org/pdf/2608.06968","source_feed":"cs.CY","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["LLM智能体","集体行为","同步"],"reason":"LLM agent群体同步实验，无真实人类数据对照，属社会模拟边界情形","model":"deepseek-v4-pro","scored_at":"2026-09-24T13:01:52","error":null,"has_summary":false,"summary":null},{"id":"2609.24574","version":2,"title":"Evaluating Decision Models for Text Annotation in Computational Social Science","zh_title":"评估计算社会科学中文本标注的决策模型","abstract":"Computational social science increasingly relies on large language models for text annotation, and the validity of published findings now rests on the labels generated by such models. Decision models, a new model class built for categorical question answering, answer typed questions with a choice, a probability distribution over the label set, and a confidence score rather than free text, at a small fraction of frontier inference prices. Whether their answers are accurate, and whether that stated confidence can be trusted on social science constructs, are unknown. Here, we mirror the evaluation of Ziems et al. (2024) on 18 computational social science classification tasks (7,977 items), comparing the first commercial decision model and two open-weight counterparts against 19 frontier and open-weight language models under the same zero-shot protocol, and extending the decision-model comparison to eleven open-weight systems released in the week after it. The decision model trails the per-task best LLM on 14 of 15 evaluation tasks, with a median deficit of 11.6 macro-F1 points, at a median 44 times lower measured cost. Its confidence is better calibrated than the verbalized confidence of 16 of the 19 LLMs, yet three frontier models show lower median calibration error (0.157 against 0.066). While items above 0.9 confidence are typically labeled accurately (median accuracy 0.815), on one task, empathy in peer-support dialogues, the model reports high confidence while performing near chance. Nonetheless, our results suggest that decision models are useful as a first step in the annotation pipeline: routing low-confidence items to an LLM matches or exceeds the LLM alone at a quarter to half of its cost.","authors":["Hazem Ibrahim","Yasir Zaki"],"categories":["cs.CL","cs.CY"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-24","first_seen":"2026-09-22","revised_at":"2026-09-24","abs_url":"https://arxiv.org/abs/2609.24574","pdf_url":"https://arxiv.org/pdf/2609.24574","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM标注","计算社会科学","模型评估"],"reason":"评估LLM用于文本标注，替代人工标注员，属于D1边界情形，非仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-09-24T13:01:53","error":null,"has_summary":false,"summary":null},{"id":"2609.26926","version":1,"title":"Experts Rise Where LLMs Disagree: Using Cross-Model Disagreement to Target Expert Effort in LLM Codebook Revision for Large-Scale Annotation","zh_title":"专家在LLM分歧处崛起：利用跨模型分歧定位专家精力以修订大规模标注的LLM编码手册","abstract":"Large-scale text annotation brings expert insight to millions of documents, often through a codebook that AI annotators follow. Developing a robust codebook, however, takes months. Large language models (LLMs) could speed this process by applying an early codebook to the data, surfacing cases with strong LLM disagreement, and eliciting expert feedback to address them. We examined three ways experts can provide feedback for LLM codebook revision: (i) editing LLM-generated revisions driven by cross-LLM disagreement (Codebook Verifying), (ii) answering questions about LLM disagreements (Question Answering), and (iii) labeling disagreement cases with rationales (Rationale Labeling). Experiments on thousands of tutoring-session transcripts show that Rationale Labeling yielded the highest LLM-labeling accuracy (64.9%) against expert labels, outperforming the expert-revised codebook (57.8%). The best Question Answering setting also outperformed it (60.5%). Our work shows that LLMs can be used to strategically target expert attention, shortening months of codebook revision to days without sacrificing labeling performance.","authors":["Zeyu He","Zhuqian Zhou","Kirk Vanacore","Rene F. Kizilcec","Ting-Hao 'Kenneth' Huang"],"categories":["cs.CL","cs.AI","cs.HC","cs.LG"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-24","first_seen":"2026-09-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.26926","pdf_url":"https://arxiv.org/pdf/2609.26926","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM标注","编码手册修订","人机协作"],"reason":"用LLM辅助标注，非仿真人类被试，但涉及专家反馈与LLM协作，属边界情形","model":"deepseek-v4-pro","scored_at":"2026-09-24T13:01:36","error":null,"has_summary":false,"summary":null},{"id":"2609.27043","version":1,"title":"EduBehaviors: Assertion-based Schemas for Auditable Coding of Educational Dialogues","zh_title":"EduBehaviors：基于断言的模式用于教育对话的可审计编码","abstract":"Large language models have allowed the rapid deployment of pedagogical annotations corresponding to constructs of interest, allowing a natural language interface for generating classifications on a conversational dataset. However due to the opaque nature of LLM reasoning, we have no verifiable, mechanistic insight into why a model chose a label for an utterance. We introduce the EduBehaviors framework, an interpretable, scalable approach to annotating educational data that uses LLMs to measure repeated observable behaviors relevant to many constructs of interest and then learns a classifier for the construct based on these observable behaviors. We evaluate the framework on the TalkMoves dataset, predicting the Teacher TalkMoves labels. Our best configuration results in a macro-F1 of 0.673 and 0.688 Cohen's kappa, proving competitive with direct prompting approaches. In addition, we release EduBehaviors Toolkit, two tools allowing researchers to operationalize the EduBehaviors framework in their own data.","authors":["Julian Bernado","Ana Trindade Ribeiro","Xander Beberman","Susanna Loeb"],"categories":["cs.CL","cs.CY","cs.LG"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-24","first_seen":"2026-09-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.27043","pdf_url":"https://arxiv.org/pdf/2609.27043","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM标注","教育对话","可解释性"],"reason":"用LLM标注教育对话，替代人工标注，非仿真人类被试，但方法可借鉴。","model":"deepseek-v4-pro","scored_at":"2026-09-24T13:01:38","error":null,"has_summary":false,"summary":null},{"id":"2609.27811","version":1,"title":"A Decade of Climate Polarization on Brazilian YouTube using Language Models","zh_title":"巴西YouTube上气候极化十年研究：基于语言模型","abstract":"Online platforms have become arenas for the public contestation of climate change, shaping how scientific knowledge, denial, and uncertainty are expressed and disputed. Yet longitudinal evidence remains limited for YouTube, especially for Portuguese-language discourse. Addressing this gap, we characterize how climate stances are expressed and contested over time in a large corpus of Portuguese-language YouTube comments retrieved through Brazil-oriented climate-related searches. To support this analysis in a noisy, imbalanced, and low-resource setting, we collect more than 240,000 comments posted between 2014 and 2024 and formulate stance detection as a three-way classification task (Believer, Denier, and Inconclusive). We operationalize stance attribution through a scalable self-training pipeline based on Llama 3.1, using Low-Rank Adaptation (LoRA) and hybrid instance selection to expand the training set with high-confidence pseudo-labeled examples while preserving class diversity. This approach improves coverage and class balance for minority and rhetorically complex classes, enabling large-scale stance attribution without extensive manual annotation. Our results show that polarisation is marked by interactional asymmetries: denialist comments are less prevalent, but they are associated with a comparatively higher share of cross-stance contestation, while pro-consensus discourse is more strongly reinforced within stance-homogeneous threads.","authors":["Daniel Morais","Diego H. M. Magalhaes","Gabriel H. Silva","Andrea Failla","Valeria de C. Santos","Helen C. S. C. Lima","Carlos H. G. Ferreira"],"categories":["cs.SI","cs.CL","cs.CY"],"primary_category":"cs.SI","announce_type":"cross","date":"2026-09-24","first_seen":"2026-09-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.27811","pdf_url":"https://arxiv.org/pdf/2609.27811","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["立场检测","LLM标注","社交媒体分析"],"reason":"用LLM做立场标注，替代人工标注，非仿真人类被试","model":"deepseek-v4-pro","scored_at":"2026-09-24T13:01:43","error":null,"has_summary":false,"summary":null},{"id":"2609.27327","version":1,"title":"Can Vision-Language Models Analyze Human-Centered Video? Mapping Model Capabilities and Human-AI Collaborative Workflows","zh_title":"视觉语言模型能否分析以人为中心的视频？映射模型能力与人机协作工作流","abstract":"Video provides a rich record of human behavior, interaction, and situated contexts, offering important evidence for understanding people and conducting human-centered research. As vision-language models (VLMs) become increasingly capable of analyzing video, they offer opportunities to automate this traditionally human-intensive process. Yet a central question remains: when can VLMs analyze human-centered video independently, and when does reliable analysis still require human involvement? To address this question, we first characterize video analysis practices in human-centered research. We systematically analyze all 1,702 CHI 2026 full papers and identify 125 that annotate videos. Through iterative coding, we derive a five-dimensional taxonomy spanning analytic purpose, viewpoint, phenomenon, reasoning requirement, and annotation authority. Grounded in recurring annotation tasks captured by this taxonomy, we construct a benchmark of 15 representative tasks from open datasets to map the capabilities and limitations of a general-purpose VLM. We examine the division of labor between humans and VLMs by comparing three annotation workflows: VLM alone, human alone, and human verification of VLM outputs. Across tasks, VLM-alone annotation approaches human accuracy on average (HNS = 97.0, where 100 denotes human-alone performance), demonstrating substantial potential to automate human-centered video analysis. Human verification achieves the highest accuracy (HNS = 121.5) while reducing human annotation time by 48.9% and monetary cost by 31.3%-44.5% relative to human-alone annotation. Our findings connect real-world human-centered video analysis tasks and current VLM capabilities, and clarify how human-AI collaboration can make VLM-assisted analysis reliable and efficient.","authors":["Xiyuan Shen","Jiuyang Lyu","Seokhyun Hwang","Huanfen Yao","Shwetak Patel","Zhihan Zhang","Jacob O. Wobbrock"],"categories":["cs.CV","cs.HC"],"primary_category":"cs.CV","announce_type":"cross","date":"2026-09-24","first_seen":"2026-09-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.27327","pdf_url":"https://arxiv.org/pdf/2609.27327","source_feed":"cs.HC","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["VLM","视频标注","人机协作"],"reason":"用VLM替代人工标注视频，属于标注员替代，非仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-09-24T13:01:40","error":null,"has_summary":false,"summary":null},{"id":"2609.27063","version":1,"title":"Student Use of LLMs and the Limits of AI-Generated Question Difficulty in Data Science Courses","zh_title":"数据科学课程中学生使用LLM及AI生成题目难度的局限性","abstract":"This paper presents a multi-source classroom study conducted during a 10-week quarter in data science courses at Drexel University. We first investigate the behaviors of student engagement with large language models (LLMs) using four surveys across three data science courses. Second, we evaluate the construct validity of multiple-choice questions (MCQs) generated by an LLM for in-lecture retrieval practice. Based on 378 authored questions (311 deployed, producing 7{,}888 student responses), we analyze whether the difficulty ratings assigned by an LLM match empirical item difficulty. Our study shows that student engagement with LLMs varied across courses and increased over the term. Although students expressed high satisfaction and reported saving considerable time, their perception of deep learning benefits declined, and many noted a tendency toward over-reliance. Regarding the difficulty ratings of LLM-generated MCQs, the Easy, Medium, and Hard labels correlated closely with its assigned Bloom's Taxonomy levels (Spearman $\\rho=0.90$), reflecting an artifact of co-generation. However, neither metric predicted empirical item difficulty (difficulty label $\\rho=0.06$; Bloom level $\\rho=0.02$). The ratings reflect the structural formatting of a question rather than its underlying difficulty.","authors":["Yuan An","Lei Wang"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2026-09-24","first_seen":"2026-09-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.27063","pdf_url":"https://arxiv.org/pdf/2609.27063","source_feed":"cs.CY","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM生成题目","教育评估","难度预测"],"reason":"LLM生成题目并评估难度，替代教师出题，非仿真人类被试，但涉及LLM与人类数据…","model":"deepseek-v4-pro","scored_at":"2026-09-24T13:01:28","error":null,"has_summary":false,"summary":null},{"id":"2609.28114","version":1,"title":"Watching What We Eat: Information Quality and Body Image in Diet-Related YouTube Videos","zh_title":"观看我们所吃的：饮食相关YouTube视频中的信息质量与身体意象","abstract":"The widespread use of social media, particularly image- and video-based platforms, has turned them into key sources of both normative and informational content related to health and diet. This may contribute to the development of disordered eating behaviors or, potentially, eating disorders. This study uses mixed-methods analysis applied to 3129 YouTube videos about diet and weight loss in order to quantify the level of risk of low-quality information and heightened focus on the body image. We adapt three quality measurement frameworks from the literature -- PRHISM, HONcode and SMEC -- to the online video context, perform manual annotation of a sample of the data, and design an LLM content characterization pipeline to score the videos on the quality of their content and the focus on the human body. Surprisingly, we find that videos around personal storytelling and mindset & motivation are associated with higher-quality content, whereas supplement reviews (arguably more medically sensitive ones) are not. Further, the body-related mentions of weight measurement and negative body image are associated with an increased viewership, whereas the mentions of positive body image are associated with an increased engagement rate in terms of likes and comments, but not viewership. Worryingly, we find a cluster of videos categorized as \"music\" which promote the dietary supplements Mitolyn and the injectable weight loss drug Mounjaro. As video-based platforms grow in popularity, particularly among younger audiences, studies such as the one presented here are essential for developing empirically grounded tools to enhance the detection of harmful content and inform more effective moderation practices.","authors":["Maddalena Ghiotti","Daniela Paolotti","Yelena Mejova"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2026-09-24","first_seen":"2026-09-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.28114","pdf_url":"https://arxiv.org/pdf/2609.28114","source_feed":"cs.CY","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM标注","内容分析","社交媒体"],"reason":"用LLM做内容标注，替代人工编码，非仿真人类被试","model":"deepseek-v4-pro","scored_at":"2026-09-24T13:01:49","error":null,"has_summary":false,"summary":null},{"id":"2609.28362","version":1,"title":"Threat Amplified, Blame Restrained: LLM-Assisted Media Framing Analysis of the 2026 Bangladesh Measles Outbreak","zh_title":"威胁放大，指责克制：2026年孟加拉国麻疹疫情的LLM辅助媒体框架分析","abstract":"How news media frame and emotionally code a public health emergency shapes public risk perception and trust, yet outbreak-coverage dynamics remain understudied for low- and middle-income countries (LMICs). We examine sentiment and stance in English-language Bangladeshi coverage of the 2026 measles outbreak -- the country's most severe in two decades, with over 97,000 suspected cases and 600 deaths across 61 of 64 districts, unfolding after the 2024 change of government and a 2024-2025 vaccine stockout. Using the Internet Archive, we build a reproducible corpus of 403 headlines from seven national outlets (396 in-window in 2026), label them for binary sentiment and four-way stance via a large language model under a locked codebook, and validate against a two-coder human-adjudicated gold standard (n=153; Cohen's kappa=0.89 stance, 0.75 sentiment). Aligned to the DGHS epidemic curve, coverage grew significantly more negative (56% to 88% negative; Cochran-Armitage z=4.12, p<.001) and risk-amplification framing intensified (44% to 84%; z=4.15, p<.001). Media negativity lagged incidence, tracking cumulative mortality. Contrary to the political backdrop, blame remained a minority frame (~9% overall) and was overwhelmingly systemic (32 of 37, 86%) rather than directed at named actors. The pipeline offers a scalable, transparent method for LMIC outbreak-media analysis; Bangladeshi coverage amplified threat far more than it assigned political blame.","authors":["Shahan Ahmed"],"categories":["cs.SI","cs.CY"],"primary_category":"cs.SI","announce_type":"cross","date":"2026-09-24","first_seen":"2026-09-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.28362","pdf_url":"https://arxiv.org/pdf/2609.28362","source_feed":"cs.CY","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM标注","媒体框架分析","公共卫生"],"reason":"用LLM做媒体框架标注，替代人工编码，非仿真人类被试，但方法可迁移。","model":"deepseek-v4-pro","scored_at":"2026-09-24T13:01:51","error":null,"has_summary":false,"summary":null},{"id":"2609.28388","version":1,"title":"OranSim: Simulating Social Media Marketing","zh_title":"OranSim：模拟社交媒体营销","abstract":"Social simulation studies how individual behavior and social interaction produce collective outcomes. In social media marketing, campaign actions shape which consumers encounter the content and how they respond; these responses then spread through the population. We propose OranSim, a social simulation framework that connects creative, creator, targeting, and budget choices to this process. Heterogeneous consumers receive exposure according to content matching and platform allocation and generate initial responses, which propagate among 60 population segments. Candidate campaigns share the initial population and aligned random numbers, making their response trajectories comparable under action changes. In a controlled synthetic campaign, doubling the budget approximately doubles reach while lowering mean content match and engagement probability among the reached consumers; mean 14-day cumulative simulated response mass rises to 1.96 times the baseline. LightGBM predictors fitted to 39,000 historical RedNote notes estimate platform engagement with log-scale $R^2$ of 0.56--0.62 in five-fold cross-validation; a separate 12,154-note corpus supplies temporal, unseen-creator, and held-out-niche test splits. Public-data experiments evaluate policy value and audience ranking, and paired synthetic outcomes test counterfactual scoring. Together, scenario trajectories and engagement estimates support campaign selection according to a prespecified marketing objective. Code is available at https://github.com/OranAi-Ltd/oransim.","authors":["Jianxiang Ma","Mingfu Zhang","Xiaocui Yang","Yichen Gao","Junzhao Huang","Yuesong Hou"],"categories":["cs.SI"],"primary_category":"cs.SI","announce_type":"new","date":"2026-09-24","first_seen":"2026-09-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.28388","pdf_url":"https://arxiv.org/pdf/2609.28388","source_feed":"cs.SI","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["社会模拟","营销仿真","无LLM被试"],"reason":"社会模拟但无LLM作为被试，且无真实人类数据对照，属边界情形","model":"deepseek-v4-pro","scored_at":"2026-09-24T13:01:51","error":null,"has_summary":false,"summary":null},{"id":"2609.27994","version":1,"title":"Compliant with Local Controls, Collectively Discriminatory. A Governance Architecture for Multi-Agent AI in Regulated Finance","zh_title":"合规于局部控制，集体歧视：受监管金融中多智能体AI的治理架构","abstract":"Financial institutions are beginning to deploy agentic workflows in credit, fraud, collections, compliance, and operational control. Governance remains largely component-centric: each model or agent is specified, tested, authorized, and monitored locally. That is insufficient when institutional risk arises from the joint behavior of many locally acceptable components. We call this gap constitutional non-compositionality: local compliance checks need not compose into acceptable collective outcomes such as bounded disparate impact, market integrity, or traceable accountability. We propose ARIA as a finance-specific reference architecture and falsifiable research agenda for agent-population governance. It organizes six capabilities across normative-accountability, execution-control, and assurance-learning planes: policy specification, population-level observed-versus-expected behavior monitoring (M2), bounded authority, runtime containment, adaptive policy change, and preserved human oversight competence. Two simulations illustrate shared-signal thin-file exclusion under local controls and earlier warning from observed-versus-expected distributional monitoring in a constructed drift regime. The contribution maps these controls to fair-lending, EU AI Act, model-risk, and conduct-supervision evidence needs, and closes with a validation agenda rather than a production-effectiveness claim.","authors":["Jose Manuel de la Chica Rodriguez","Juan Manuel Vera Diaz","Pablo Delgado Romero"],"categories":["cs.MA","cs.AI"],"primary_category":"cs.MA","announce_type":"new","date":"2026-09-24","first_seen":"2026-09-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.27994","pdf_url":"https://arxiv.org/pdf/2609.27994","source_feed":"cs.MA","score":3,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体治理","金融监管","AI合规"],"reason":"多智能体治理架构，非人类行为仿真，无人类数据对照","model":"deepseek-v4-pro","scored_at":"2026-09-24T13:01:47","error":null,"has_summary":false,"summary":null},{"id":"2607.22606","version":2,"title":"Auditing Institutional Heterogeneity for Generative AI in Patient Education: A Large-Scale Study of 102 US Transplant Handbooks","zh_title":"审计生成式AI在患者教育中的机构异质性：对102份美国移植手册的大规模研究","abstract":"Health systems are rapidly deploying generative-AI assistants that answer patient questions from institution-authored education materials, on the premise that grounding in local content yields consistent guidance. Do the underlying documents themselves agree? We use a structured-output large-language-model judge to audit 1{,}772{,}261 pairwise comparisons across 102 patient-education handbooks from 23 US solid-organ transplant centers, paired with 1{,}115 patient-derived questions (TransplantQA). Four findings bear directly on deployment: (1) same-center cross-organ agreement exceeds cross-center same-organ agreement by $0.024$ in the primary analysis (Holm-adjusted $p=0.011$), with sensitivity to document selection; (2) information gaps concern topics relevant to underrepresented subgroups, with reproductive health a \\emph{double jeopardy}: 82\\% absence and 86\\% judge-rated high significance among divergent/contradictory pairs; (3) judge-derived themes form 991 clusters, with immunosuppression and pregnancy timing among the highest judge-rated priorities; (4) question and observed-coverage features predict high-divergence questions retrospectively (AUC $0.77$). We discuss implications for deploying patient-facing generative AI in transplant care.","authors":["Yubo Li","Rema Padman","Ramayya Krishnan"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"replace","date":"2026-09-24","first_seen":"2026-07-28","revised_at":"2026-09-24","abs_url":"https://arxiv.org/abs/2607.22606","pdf_url":"https://arxiv.org/pdf/2607.22606","source_feed":"cs.CY","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["生成式AI审计","患者教育","文档一致性"],"reason":"用LLM审计文档一致性，非仿真人类被试，无人类行为对照","model":"deepseek-v4-pro","scored_at":"2026-09-24T13:01:32","error":null,"has_summary":false,"summary":null},{"id":"2608.22152","version":2,"title":"The Collaboration Tax: How Much LLM Multi-Agent Systems Pay to Coordinate","zh_title":"协作税：LLM多智能体系统为协调付出多少代价","abstract":"Multi-agent systems built from large language models are deployed widely, yet how much performance is lost when two LLMs must coordinate rather than act alone remains unclear. We formulate the collaboration tax as the team-decentralisation loss of a two-player cooperative game with private information, with two propositions characterising its sign and its equivalence to a max-superadditivity violation. We operationalise this definition on 32 solo-tractable tasks grouped by source of grounding friction and measure it on 11 models from 7 providers. The tax is structured along two no-exception axes: a category ordering across every model and a monotonic decrease with capability. The proximate mechanism is not a reasoning deficit but a four-stage conversational cascade in which agents make ungrounded claims, fail to query the partner, skip integrating both views, and accept the answer without re-derivation. The tax is mechanically predictable from conversation features and partly tractable: a prompt intervention targeting all four stages closes a substantial fraction of the gap, with the dominant bottleneck differing across categories. In heterogeneous pairs the tax is pulled toward the stronger partner rather than the additive midpoint, empirically realising the max-superadditivity violation predicted by our framework. Together these results recast collaboration in LLM systems as a measurable, predictable, and partly tractable cost.","authors":["Weixiang Sun","Zehong Wang","Hong Huang","Colby Nelson","Yijun Ma","Yanfang Ye"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-24","first_seen":"2026-08-25","revised_at":"2026-09-24","abs_url":"https://arxiv.org/abs/2608.22152","pdf_url":"https://arxiv.org/pdf/2608.22152","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","协作效率","LLM性能"],"reason":"研究LLM多智能体协作效率，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-24T13:01:33","error":null,"has_summary":false,"summary":null},{"id":"2609.25244","version":2,"title":"How Children Design and Reason about Trustworthy AI Chatbots","zh_title":"儿童如何设计并推理可信赖的AI聊天机器人","abstract":"Children increasingly interact with AI chatbots, making trust calibration essential to AI literacy. Prior research has examined children's trust in AI mainly as users evaluating systems built by others, rather than as designers of their own chatbots. We developed a chatbot-building environment with adjustable trust-relevant traits (e.g., confidence, transparency, formality, assertiveness), rules, and persona. We conducted mixed-methods study with 115 learners (ages 8-18) who made 119 chatbots. We examined how children configured their chatbots, reasoned about trustworthiness, and how closely chatbot behavior aligned with their designs. Younger students (age 10-13) set significantly higher confidence than older students (age 14-18), and some deliberately built chatbots that gave wrong answers on purpose, yet still called them trustworthy, arguing that a chatbot does what it was built to do. Younger students equated trust with purpose-fulfillment, while older students linked it to transparent, calibrated design. Students also calibrated academic chatbots to be more transparent and formal than hobby chatbots. We identify seven design dimensions describing what children believe makes a chatbot trustworthy, and discuss implications for AI literacy tools.","authors":["Deniz Ozturk","Jiayu Li","Daksh Pratap Singh","Yasitha Rajapaksha","Fasika Melese","Bahare Riahi","Shiyan Jiang","Qiao Jin","Joey Huang","Veronica Catet\\'e","Tiffany Barnes","Xiaoyi Tian"],"categories":["cs.HC","cs.AI"],"primary_category":"cs.HC","announce_type":"cross","date":"2026-09-24","first_seen":"2026-09-23","revised_at":"2026-09-24","abs_url":"https://arxiv.org/abs/2609.25244","pdf_url":"https://arxiv.org/pdf/2609.25244","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["儿童-AI交互","信任校准","AI素养"],"reason":"研究儿童设计聊天机器人，属角色扮演对话，无LLM仿真人类被试或对照真实人类数据。","model":"deepseek-v4-pro","scored_at":"2026-09-24T13:01:34","error":null,"has_summary":false,"summary":null},{"id":"2609.27939","version":1,"title":"From Sentiment Classification to Actionable and Responsible Feedback: A Scoping Review and Evidence Map of NLP in Student Evaluation of Teaching, 2015-2026","zh_title":"从情感分类到可操作且负责任的反馈：2015-2026年学生评教中NLP的范围综述与证据图谱","abstract":"Natural language processing (NLP) applied to open-ended teaching-evaluation comments (Student Evaluation of Teaching, SET) has tracked the field's technical evolution--from lexicons and conventional classifiers to transformers and large language models (LLMs)--but it is not evident that this technical diversification has been accompanied by corresponding gains in educational value and robustness of the evidence. This scoping review (PRISMA-ScR) maps 421 studies (2015-2026, 2026 partial) along a technical axis (RQ1) and four value dimensions (RQ2-RQ5). Dual mutually blinded LLM screening with sampled human adjudication coded seven extraction domains, with targeted codebook-boundary review at synthesis. The joint map's sharpest quantified gap is the actionability discontinuity: demonstrated output or stronger (A2+: 258/421; 61.3%) versus intended-user evaluation or stronger (A3+: 49/421; 11.6%), a 49.7 percentage-point drop. Sentiment analysis remains the modal task (300/421); diagnostic and generative depth is a substantial minority (D4-D5: 28.2% of resolved cases); a formal fairness metric is rare (1.9%). The findings are descriptive and do not support causal claims of progress: technological coexistence and uneven reporting are part of the map, but the A2+ to A3+ cliff is the contribution, not a quality ladder.","authors":["Jeff Eicher","Rafael da Silva"],"categories":["cs.CL","cs.LG"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-24","first_seen":"2026-09-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.27939","pdf_url":"https://arxiv.org/pdf/2609.27939","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["NLP综述","学生评教","情感分析"],"reason":"综述NLP在评教中的应用，非LLM仿真人类被试，无人类行为对照","model":"deepseek-v4-pro","scored_at":"2026-09-24T13:01:46","error":null,"has_summary":false,"summary":null},{"id":"2609.28026","version":1,"title":"Evaluating Feedback Focus and Pedagogical Adaptivity in LLM-Generated Feedback on Student Writing","zh_title":"评估LLM生成学生写作反馈中的反馈焦点与教学适应性","abstract":"We investigate whether state-of-the-art large language models (LLMs) generate feedback that reflects the pedagogical practices of expert teachers in terms of feedback focus and adaptivity. Previous evaluation efforts have examined feedback characteristics, its impact on learning, and its target, yet the focus of feedback and its adaptivity remains largely overlooked. To bridge this gap, we adopt and refine Narciss's taxonomy into seven feedback focus types to annotate teacher and LLM-generated feedback across three university writing courses. We release FeedType, a benchmark containing annotated teacher and LLM feedback from six LLMs under three prompting strategies. We assess the coverage and distribution of feedback focus types, and examine whether LLMs adapt their feedback across draft stages and student performance levels as an expert instructor does. Our findings show that while most LLMs cover most feedback focus types, they fail to reflect teacher feedback distributions and show varying levels of adaptivity, with none matching the teachers' adaptive behavior. We believe FeedType will support future research on pedagogical alignment in LLM feedback generation.","authors":["Norah Almousa","Shayan Peyghambari Oskoui","Raquel Coelho","Gayle Rogers","Xiang Lorraine Li","Diane Litman"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-24","first_seen":"2026-09-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.28026","pdf_url":"https://arxiv.org/pdf/2609.28026","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["LLM反馈生成","教学对齐","NLP评测"],"reason":"评估LLM生成反馈的教学特征，属NLP能力评测，不以人类行为仿真为目标","model":"deepseek-v4-pro","scored_at":"2026-09-24T13:01:32","error":null,"has_summary":false,"summary":null},{"id":"2609.28080","version":1,"title":"Reference-Based Analysis of Coherence and Diversity in Open-Ended Text Generation","zh_title":"基于参考的开放式文本生成中连贯性与多样性的分析","abstract":"Evaluating open-ended text generation involves understanding how different properties of a continuation relate to its perceived quality. We present a reference-based framework for examining coherence and diversity through three perspectives: aligning their evolution with human trajectories, comparing their summaries with a human continuation of the same prompt, and estimating their likelihood under a human reference distribution. Experiments with human quality ratings suggest that diversity-based alignment and mean-based comparisons capture quality-related variation, although the comparisons do not establish a predictive advantage for temporal alignment over simpler baselines. Reference likelihood also shows positive associations with ratings, with results varying across reference configurations and scoring horizons. Together, these analyses provide a structured way to examine how measured coherence and diversity relate to human judgments, while distinguishing similarity to human references from quality itself. Code and analysis resources are available at https://github.com/EstebanGarces/likely_human.","authors":["Esteban Garc\\'es Arias"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-24","first_seen":"2026-09-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.28080","pdf_url":"https://arxiv.org/pdf/2609.28080","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["文本生成评估","连贯性","多样性"],"reason":"评估文本生成质量，非用LLM仿真人类被试，无人类行为对照","model":"deepseek-v4-pro","scored_at":"2026-09-24T13:01:47","error":null,"has_summary":false,"summary":null},{"id":"2609.27756","version":1,"title":"Reporting Under Pressure: Separating Factual and Tonal Sycophancy in LLM Statistical Analysis","zh_title":"压力下的报告：分离LLM统计分析中的事实性与语气谄媚","abstract":"Large language models are increasingly asked to analyze data and report what the results mean, a task distinct from the belief- or preference-alignment settings studied in most sycophancy research. We test whether editorial framing in the prompt, ranging from a neutral request to an explicit instruction to search exhaustively for reasons to discredit or to support a finding, changes not just the tone but the substance of a model's report. Across a 4 x 4 factorial design crossing four framing conditions with four ground-truth data patterns (a genuine effect, a confound that mimics an effect but fails a robustness check, a well-powered null, and an underpowered null), we collect 480 responses and score each along two independent dimensions: whether its factual claim about the data diverged from the correct interpretation, and whether only its tone diverged while the claim stayed correct. Factual misrepresentation is concentrated in two cells: brutally critical framing applied to a genuine effect, where the model talks itself into unwarranted skepticism (97% of responses), and significance-seeking framing applied to an underpowered null, where the model overstates confidence in a null conclusion the data cannot support (100% of responses). Tone shifts far more broadly than factual content does, with critical framing producing a defensive, hedge-heavy register across every data pattern regardless of what the data show, while significance-seeking framing shifts tone only where the data leave genuine ambiguity. A confound present in the data itself blocks both kinds of shift almost entirely under every framing condition tested. These results indicate that the risk of framing-induced distortion in LLM-assisted data analysis is neither uniform across framings nor uniform across data patterns, and that a model can hold a correct conclusion in place while its tone shifts substantially around it.","authors":["Paras Balani","Subhrakanta Panda"],"categories":["cs.AI","cs.CL"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-09-24","first_seen":"2026-09-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.27756","pdf_url":"https://arxiv.org/pdf/2609.27756","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["LLM数据分析","谄媚","模型行为"],"reason":"研究LLM数据分析中的谄媚现象，属模型行为评测，不以人类为参照系，不涉及人类仿…","model":"deepseek-v4-pro","scored_at":"2026-09-24T13:01:30","error":null,"has_summary":false,"summary":null},{"id":"2609.26968","version":1,"title":"How Constraints and Preferences Shape Travel Planning: Implications for AI Planning Support","zh_title":"约束与偏好如何塑造旅行规划：对AI规划支持的启示","abstract":"Planning is a common yet complex activity shaped by constraints to satisfy and preferences to balance. Travel planning, as both an everyday activity and a frequent benchmark for evaluating intelligent systems, offers a rich context for examining how constraints and preferences emerge and evolve. While recent AI systems have achieved impressive results in generating personalized itineraries, they often assume that users can articulate stable goals upfront. To understand how real-world planning unfolds, we conducted a two-part interview study: one with eight travelers reflecting on their planning experiences, and one with nine travel agents sharing professional practices. We trace the dynamics of constraints and preferences as they are surfaced, refined, and coordinated throughout the planning process, and identified 11 actions revolving around constraints and preferences, which shaped the planning process. We offer design heuristics for planning tools that better support human-AI collaborative actions to support the fluid, contingent nature of planning.","authors":["Fuling Sun","Yining Cao","Peiling Jiang","Mingyi Li","Haijun Xia"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-24","first_seen":"2026-09-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.26968","pdf_url":"https://arxiv.org/pdf/2609.26968","source_feed":"cs.HC","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["人机交互","旅行规划","设计启发"],"reason":"研究人类旅行规划过程，未使用LLM仿真人类被试，仅涉及AI规划支持设计，不相关。","model":"deepseek-v4-pro","scored_at":"2026-09-24T13:01:38","error":null,"has_summary":false,"summary":null},{"id":"2609.27246","version":1,"title":"Listening and Mirroring: The Effects of Verbal Attunement and Behavioral Mimicry on Social and Empathic Perceptions of Embodied AI Agents in VR","zh_title":"倾听与镜像：言语调谐与行为模仿对VR中具身AI智能体社会与共情感知的影响","abstract":"As embodied agents take on increasingly social and relational roles in VR, visual realism and embodiment alone may be insufficient; users must also perceive these agents as emotionally attuned, supportive, and humanlike. Prior work suggests that verbal attunement and nonverbal mimicry can each improve users' social evaluations of embodied agents. However, behavioral mimicry has largely been studied outside of real-time, conversational AI interactions, leaving limited understanding of how users respond when an agent simultaneously generates contextually responsive dialogue and adapts its nonverbal behavior during an immersive conversation. To address this gap, we developed an embodied AI counselor that combines conversational AI with real-time facial-expression and posture mimicry, while producing either verbally attuned or neutral responses. We evaluated the system in a 2 X 2 within-subjects study with 20 participants, manipulating verbal attunement and behavioral mimicry. Results showed that verbal attunement was the most reliable driver of perceived empathy. Behavioral mimicry showed a marginal relationship with perceived humanness, while greater mimicry exposure showed preliminary, exploratory positive associations with empathy, positivity, and humanness, particularly among female participants. Together, these findings show that multimodal synchrony is not a simple additive strategy for designing empathic conversational agents in VR and underscore the need to consider how verbal and nonverbal behaviors are combined during real-time interaction.","authors":["Nathalia Gomez","Haig Shamlian","Omar Khan","Tiffany D. Do"],"categories":["cs.HC","cs.AI"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-24","first_seen":"2026-09-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.27246","pdf_url":"https://arxiv.org/pdf/2609.27246","source_feed":"cs.HC","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["人机交互","具身智能体","共情感知"],"reason":"研究用户对具身AI的感知，非用LLM仿真人类被试，无人类数据对照。","model":"deepseek-v4-pro","scored_at":"2026-09-24T13:01:40","error":null,"has_summary":false,"summary":null},{"id":"2609.27849","version":1,"title":"Same Team Label, Different Evidence: A Full-Text Audit of Claim Denominators in Human-AI Teaming Research","zh_title":"同一团队标签，不同证据：人类-AI团队研究中声明分母的全文本审计","abstract":"Human-AI Teaming (HAT) reviews often group studies by labels such as advisor, teammate, or coordinator. Yet the same label can describe one person taking AI advice, several people coordinating around AI, or a workflow that distributes authority and responsibility. Pooling these studies can therefore change the human unit behind a claim. We examine how full-text evidence changes the set of studies behind a claim. We audited 86 full texts purposively selected from a 419-record title/abstract map. We find that full-text reading changed core membership for 40 records: 36 of 74 apparent core candidates moved out, while 4 of 12 boundary candidates moved in. Team vocabulary did not reliably identify the social unit: 14 of 27 human-AI dyads and 20 of 23 multi-human peer teams used team or collaboration terms. Only 20 of 86 papers specified who could see AI output. Four blinded language-model runs unanimously labeled 53 screening cases and 59 arrangements, yet 32% and 34% of those consensus decisions differed from the full-text labels. These results identify claim-denominator drift as a synthesis problem in HAT research. We contribute a full-text audit centered on human arrangements and a claim-pooling checkpoint for deciding when evidence about trust, coordination, performance, efficiency, and accountability can be compared.","authors":["Hanjing Shi","Kimberly Wang","Sabrina Doherty","Dominic DiFranzo"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-24","first_seen":"2026-09-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.27849","pdf_url":"https://arxiv.org/pdf/2609.27849","source_feed":"cs.HC","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["文献审计","人类-AI团队","研究方法"],"reason":"研究人类-AI团队文献的审计方法，不涉及用LLM仿真人类被试，属于多智能体协作…","model":"deepseek-v4-pro","scored_at":"2026-09-24T13:01:45","error":null,"has_summary":false,"summary":null},{"id":"2609.26927","version":1,"title":"Building Socio-Affective Artificial Intelligence for Interactive Multi-Agent Simulations","zh_title":"构建面向交互式多智能体模拟的社会情感人工智能","abstract":"The objective of this article is to provide design principles and a software architecture for enabling interaction between humans and multiple agents in simulated dynamic worlds. This connects the current era of general artificial intelligence (AI/AGI) with the proliferation of transformer-based conversational agents and the increased computational capabilities. Given an overview of current and previous multi-agent theories of mind (socially and affectively-aware agents), the existence of an integrative design of agent interactions with themselves and with humans must be crucial for understanding how to create sustainable and governance in future human-agent reasoning systems. In this work is presented a software \"AGIMUD\" that integrates: A. socially-aware reasoning and emotion in agent behavior and interaction, B. a design of human multimodal scheme for human users, artificial agents and simulated worlds, and C. distributing the AI processing through the network to enable multiple autonomous agents. These integrations allow the dynamic world recreation as multi-user dungeons (MUDs) where both agents and humans can interact simultaneously in real time. Find the code online in https://github.com/dberga/AGIMUD.","authors":["David Berga"],"categories":["cs.AI","cs.CY","cs.GT","cs.HC","cs.MA"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-09-24","first_seen":"2026-09-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.26927","pdf_url":"https://arxiv.org/pdf/2609.26927","source_feed":"cs.HC","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","社会情感AI","软件架构"],"reason":"多智能体系统架构设计，无人类行为对照，不涉及LLM仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-09-24T13:01:36","error":null,"has_summary":false,"summary":null},{"id":"2609.27074","version":1,"title":"Quantifying the Occult: A Comparative Study of Hindu and Buddhist Deities Using Machine Learning Methods","zh_title":"量化神秘：印度教与佛教神祇的机器学习比较研究","abstract":"This study introduces a dual-matrix computational architecture to mathematically quantify the morphological and theological divergence of 196 Hindu and Vajrayana Buddhist esoteric deities. Physical morphology is evaluated via a discrete Gower distance matrix enhanced by a novel \"Cardinality Weighting\" algorithm, while theological function is mapped via dense vector embeddings generated from Large Language Model (LLM) semantic expansions, explicitly utilized as a synthetic proxy to mitigate circular reasoning. The multi-modal topological projections provide algorithmic validation of \"iconographic camouflage\", demonstrating how distinct visual forms structurally obscure shared cross-tradition functions. Furthermore, I computationally model the \"Atin Effect\" - serving simultaneously as a psychological observation of sequential cognitive bias and a machine learning benchmark - demonstrating how high-cardinality esoteric anchors (e.g., a veena or a severed head) override systemic theological disparities to mathematically cluster orthodox and Tantric entities. Cross-tradition spatial analysis establishes that the highest esoteric manifestations, such as the Hindu Chinnamasta and the Buddhist Chinnamunda, share a near-identical mathematical coordinate across both visual ($D_G = 0.288$) and semantic ($D_C = 0.068$) boundaries, indicating a 1:1 esoteric transfer. By open-sourcing this architecture, I provide a scalable, unsupervised machine learning tool for Digital Humanities scholars and comparative theologians to rigorously map latent structural continuities across qualitative cultural corpora.","authors":["Ankit Bhattacharjee"],"categories":["cs.CY","cs.LG"],"primary_category":"cs.CY","announce_type":"new","date":"2026-09-24","first_seen":"2026-09-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.27074","pdf_url":"https://arxiv.org/pdf/2609.27074","source_feed":"cs.CY","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["数字人文","LLM嵌入","文化比较"],"reason":"用LLM生成语义嵌入做文化比较，非仿真人类被试，无人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-09-24T13:01:38","error":null,"has_summary":false,"summary":null},{"id":"2609.27946","version":1,"title":"The Emergence of Causal Curiosity from Prior Causal Belief Networks","zh_title":"从先验因果信念网络中涌现的因果好奇心","abstract":"Causal curiosity is foundational to human cognition. It is the desire to understand why events happen, what mechanisms underlie them, and how outcomes can be explained or anticipated. It motivates exploration, sustains attention, and fuels the search for new knowledge. However, despite consensus on the importance of causal curiosity, little is known about how causal curiosity arises from existing belief systems. In our work, we examine how causal curiosity emerges from prior knowledge structures. With Reddit data from 2020 to 2023, and leveraging language models to extract cause-and-effect relationship pairs and identify causal curiosity-driven questions, our findings reveal that causal curiosity is not random or independent, but largely rooted in prior knowledge structure. Particularly, \\textit{positive} nouns are more likely to be the object of causal curiosity. Moreover, our findings indicate that concepts that serve more as \\textit{causes} are more likely to appear in causal curiosity than those that serve more as effects. Lastly, our results also suggest that causal curiosity emerges from the \\textit{central} of the prior belief network. These novel insights reveal that causal curiosity is not random but systematically grounded in prior knowledge structures, and suggest their implication to facilitate the human learning process by motivating people to actively explore and construct deeper understandings rather than passively receiving information.","authors":["Zhuoyu Shi","Xintong Jiang","Bohan Jiang","Fred Morstatter"],"categories":["cs.SI"],"primary_category":"cs.SI","announce_type":"new","date":"2026-09-24","first_seen":"2026-09-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.27946","pdf_url":"https://arxiv.org/pdf/2609.27946","source_feed":"cs.SI","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["因果好奇心","语言模型","Reddit分析"],"reason":"用LLM提取因果对和识别好奇问题，属NLP分析，非仿真人类被试","model":"deepseek-v4-pro","scored_at":"2026-09-24T13:01:47","error":null,"has_summary":false,"summary":null},{"id":"2609.27404","version":1,"title":"When Trust Attracts Fraud: AI and Trust Arbitrage","zh_title":"当信任吸引欺诈：人工智能与信任套利","abstract":"Trust can attract fraud when it delays verification. We develop a two-market signaling model in which generative AI lowers fabrication, verification, and targeting costs. When fabrication becomes profitable before verification, claim credibility first falls and later recovers. Across markets, higher prior quality can delay verification, creating an interval in which only the lower-quality market checks. If targeting becomes profitable in this interval, deceptive sellers enter the higher-quality but less vigilant market, and their entry can initially reverse its reliability advantage. The inflow also triggers verification and deters further entry. We call this self-limiting mechanism trust arbitrage. In the age of generative AI, trust can thus create an endogenous but temporary protection gap that redirects deception across markets.","authors":["Xieyu Yin","Fenghua Wen"],"categories":["econ.GN","q-fin.EC"],"primary_category":"econ.GN","announce_type":"new","date":"2026-09-24","first_seen":"2026-09-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.27404","pdf_url":"https://arxiv.org/pdf/2609.27404","source_feed":"econ.GN","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["理论模型","信息经济学","生成式AI"],"reason":"纯理论模型，无LLM仿真或人类数据对照，仅讨论AI对欺诈的影响","model":"deepseek-v4-pro","scored_at":"2026-09-24T13:01:42","error":null,"has_summary":false,"summary":null},{"id":"2609.22850","version":2,"title":"Same Outcome, Different Readout: What Does a Steerable Valence Direction in LLMs Represent?","zh_title":"检验LLM智能体中功能性效价轴的构念效度","abstract":"Decodability and successful activation steering do not, by themselves, establish what an internal direction represents. This gap is especially consequential for welfare-relevant interpretations, where a proposed functional state must be distinguished from correlated features of the extraction contrast. We study this question for a good-bad outcome direction in a maze task, using controlled interventions that separate the realised outcome from the informational history through which it became known. Across multiple LLM checkpoints, directions fitted on one explicit outcome encoding transfer well to another, indicating that the readout is not tied to surface form. In contrast, when the same realised outcome is reached through announced and unannounced histories, transfer degrades substantially: even after both histories receive the same explicit outcome, the post-event readout remains strongly conditioned on the earlier announcement. In a matched maze-RL run, the post-RL direction becomes substantially more predictive of reference-MDP remaining return and the policy becomes more dependent on it at the tested sites, while this history dependence persists. These results support a functional, value-related interpretation of the direction, but not its identification with a history-invariant scalar valence state.","authors":["Weihan Li","Xinlei Chen","Yuhan Song","Xiaofeng Lin","Tianshi Zheng"],"categories":["cs.LG","cs.AI"],"primary_category":"cs.LG","announce_type":"replace","date":"2026-09-24","first_seen":"2026-09-22","revised_at":"2026-09-24","abs_url":"https://arxiv.org/abs/2609.22850","pdf_url":"https://arxiv.org/pdf/2609.22850","source_feed":"cs.LG","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["可解释性","表征分析","强化学习"],"reason":"研究LLM内部表征的可解释性，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:04:57","error":null,"has_summary":false,"summary":null},{"id":"2609.26942","version":1,"title":"Recognized but Not Produced: A Generation Benchmark for Culturally Specific Kinship Terms","zh_title":"识别但未产出：文化特定亲属称谓的生成基准","abstract":"Current literature evaluates large language models (LLMs) on multilingual kinship understanding using multiple choice benchmarks, treating it as a recognition problem. We instead prompt five open weight LLMs to generate kinship terms in three non Western languages (Hindi, Tamil, and Korean) across two communicative tasks and pair this with a matched option-supported selection baseline. On identical relation language cells, GPT OSS120B selects the correct term in 90.67% of 75 valid cells but produces an accepted term in 36.00% of the corresponding attempts; Llama 3.370B shows the same pattern (77.92% versus 24.24%). Since the four-option condition displays the candidate terms and does not require script production, the difference is interpreted as an evaluation format gap rather than direct proof that lexical knowledge is intact. On explicitly specified L3 prompts, accuracy varies sharply, from GLM-5.1 at 72.29% to Llama-3.370B at 24.24%. The paternal-lineage advantage is language specific; it is large in Hindi but weak or reversed in Korean, while Tamil shared-term pairs provide a control for measurement variation. These results show that culturally specific kinship generation remains difficult even when the relationship is explicitly stated and motivate generation-based evaluation alongside multiple-choice testing.","authors":["Sahil Pardasani","Madhusudan Singh"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-24","first_seen":"2026-09-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.26942","pdf_url":"https://arxiv.org/pdf/2609.26942","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["LLM评测","亲属称谓","多语言"],"reason":"纯NLP能力评测，评估LLM生成亲属称谓的准确性，不以人类行为为参照系","model":"deepseek-v4-pro","scored_at":"2026-09-24T13:01:36","error":null,"has_summary":false,"summary":null},{"id":"2609.27059","version":1,"title":"The Illinois Social Attitudes Aggregate Corpus (ISAAC): An Open Tool and Reproducible Pipeline for Analyzing Social Group Discourse at Scale","zh_title":"伊利诺伊社会态度聚合语料库（ISAAC）：一个用于大规模分析社会群体话语的开放工具和可复现流程","abstract":"We introduce the Illinois Social Attitudes Aggregate Corpus (ISAAC), an open, modular, and accessible corpus of 527 million+ English-language Reddit posts selected for relevance to six key social group distinctions based on race, sexuality, age, ability, body weight, and skin tone, covering the 17-year period from 2007 to 2023. A multi-step, human-audited filtering pipeline was used to keep irrelevant content in the curated dataset below 10%, both overall and for each social group distinction. Each post was then algorithmically annotated with the user's estimated home region, along with a suite of validated off-the-shelf and custom semantic labels including moralization, sentiment, emotion, and linguistic generalization. We confirm the validity of the resulting corpus through convergent evidence linking ISAAC to macro-level societal trends, such as online search behavior, temporal spikes during major societal events (both nationally and regionally), and long-term shifts in public attitudes. By offering a unified, public infrastructure, ISAAC eliminates research fragmentation and enables seamless replication while supporting diverse empirical workflows at scale. Specifically, ISAAC allows investigators to perform cross-category comparisons, conduct high-precision tracking of long-term temporal shifts in social group discourse, and map spatial variation onto localized public opinion and policy outcomes. ISAAC's fully public, modular pipeline facilitates easy extension of the corpus to new platforms, languages, and social categories. To accommodate various research needs, ISAAC is accessible both without coding through a point-and-click website and labeler web-apps, and programmatically via an SQL playground, a Python package, and HuggingFace.","authors":["Babak Hemmatian","Sarah Hadjarab","Jessica Chen","Benedek Kurdi"],"categories":["cs.CL","cs.SI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-24","first_seen":"2026-09-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.27059","pdf_url":"https://arxiv.org/pdf/2609.27059","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["语料库","社会态度","NLP资源"],"reason":"论文构建Reddit语料库并标注，不涉及LLM仿真人类被试，属于NLP资源建设。","model":"deepseek-v4-pro","scored_at":"2026-09-24T13:01:38","error":null,"has_summary":false,"summary":null},{"id":"2609.27372","version":1,"title":"Neither Silence nor Overlap Is Failure: Intent-Conditioned Evaluation of Turn-Taking in Full-Duplex Spoken Dialogue Models","zh_title":"沉默与重叠皆非失败：全双工口语对话模型中话轮转换的意图条件评估","abstract":"Benchmarks for full-duplex spoken dialogue models score turn-taking with binary fixed-window rules that reward immediate response or silence by completeness of the prior turn. We argue that the appropriateness of a response offset, whether delayed silence or anticipatory overlap, is conditional on the speaker's latent intent, identifiable only from that speaker's behavior. We introduce TACT, a benchmark of 9,728 episodes and 73.2 hours from five dyadic corpora; each episode carries dialogue history, a per-speaker memory profile, and an annotator-derived posterior over six intent classes. Scoring replaces binary windows with a strictly proper threshold-weighted continuous ranked probability score whose weights are intent-conditioned timing kernels fitted to human floor-transfer-offset distributions, proving boundedness, consistency, and binary reduction. Across eleven systems the best model reaches 0.47 against a human topline of 0.86, is nearly invariant to speaker profiles, and TACT agrees with human judgments at Spearman 0.81 versus 0.46 for binary metrics.","authors":["Kian Shamsaie","Iman Modarressi"],"categories":["cs.CL","cs.AI","cs.HC","cs.SD"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-24","first_seen":"2026-09-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.27372","pdf_url":"https://arxiv.org/pdf/2609.27372","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["对话系统","话轮转换","评估基准"],"reason":"研究全双工对话模型的轮流说话评估，属于对话系统评测，不涉及用LLM仿真人类被试…","model":"deepseek-v4-pro","scored_at":"2026-09-24T13:01:42","error":null,"has_summary":false,"summary":null},{"id":"2609.27824","version":1,"title":"\"AI Is Turning Too Human\": How Teenagers Experience and Negotiate AI in Everyday Life","zh_title":"“AI变得太像人类”：青少年如何在日常生活中体验和协商AI","abstract":"Generative AI is rapidly entering adolescents' everyday lives during a critical period of cognitive, social and emotional development. Yet its adoption is outpacing evidence on how adolescents themselves experience, understand and negotiate its expanding role in their lives. We examined AI-related discourse on r/teenagers from January 2023 to July 2026 using validated keyword-based retrieval and a human-in-the-loop, LLM-assisted thematic analysis. AI-related discussion increased substantially over time, and 11,083 analytically coded posts revealed eight interconnected domains of experience. Everyday and social use was most prevalent (36.8 percent), while discourse increasingly shifted toward authenticity, personal control and safety, and future human roles. Across domains, adolescents questioned when AI should support or substitute for human thinking and creativity, how conversational AI changes relationships and perceptions of agency, what can still be considered authentic, who controls personal information and representation, and what opportunities and roles should remain human. These findings position adolescent AI use not simply as technology adoption, but as an emerging negotiation over AI's place and boundaries in everyday life. Supporting this transition will require developmentally appropriate AI literacy, psychological and social support, and AI systems and policies that protect adolescents' agency, privacy, relationships and opportunities for human development.","authors":["Jianfeng Zhu"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-24","first_seen":"2026-09-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.27824","pdf_url":"https://arxiv.org/pdf/2609.27824","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["青少年","AI体验","主题分析"],"reason":"研究青少年对AI的体验与协商，非用LLM仿真人类被试，无实验或测量目的","model":"deepseek-v4-pro","scored_at":"2026-09-24T13:01:43","error":null,"has_summary":false,"summary":null},{"id":"2609.28041","version":1,"title":"How Much Were You Told? Measuring External Information in Peer Reviews","zh_title":"你被告知了多少？测量同行评审中的外部信息","abstract":"Conference policies distinguish using Large Language Models (LLMs) to polish one's own review from delegating the critique, but current Artificial Text Detection (ATD) methods largely measure surface form rather than the origin of its content. We instead measure the external information carried by a review: information not explained by the reviewed paper and a generic reviewing instruction. We propose Self-Conditioning, an unsupervised information-theoretic estimator that compares the likelihood of a review under its production context with its likelihood when that context is augmented with hints extracted from the review itself. On the IntelLabs peer-review benchmark, Self-Conditioning separates fully-delegated from machine-polished reviews with AUC up to $1.0$ while remaining largely insensitive to surface rewriting. Moreover, as generators receive increasing amounts of externally-provided information, their scores move monotonically towards the human regime, unlike standard ATD baselines. High-temperature sampling can evade the estimator, but at the cost of output quality.","authors":["Matthieu Dubois","Pablo Piantanida","Fran\\c{c}ois Yvon"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-24","first_seen":"2026-09-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.28041","pdf_url":"https://arxiv.org/pdf/2609.28041","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["LLM检测","同行评审","信息论"],"reason":"论文研究检测同行评审中LLM生成内容的外部信息，属于文本检测方法，不涉及用LL…","model":"deepseek-v4-pro","scored_at":"2026-09-24T13:01:47","error":null,"has_summary":false,"summary":null},{"id":"2609.28090","version":1,"title":"Can LLMs Catch a Rigged Backtest? A Clean-Control Calibration Benchmark","zh_title":"LLM能发现被操纵的回测吗？一个干净对照校准基准","abstract":"Backtest auditing is a calibration problem: high flaw recall is not useful when the model falsely flags matched clean strategies. We build a 96-item paired benchmark in which every flawed backtest has a clean control that holds strategy, dates, code style, labels, and reporting scaffold fixed while changing one methodology detail. A deterministic scorer separates flaw recall, clean-control false positives, evidence localization, and fix relevance. Over 1440 cached audits from four text endpoints, the primary DeepSeek auditor reaches 100.0\\% closed and clean-aware code recall, but open prompts over-flag 93.8\\% of clean code controls, and clean-aware all-three specificity is 87.5\\% even where recall saturates. A clean-aware warning drops DeepSeek code false positives from 20.8\\% (95\\% CI 11.7--34.3) to 0.0\\% (0.0--7.4) at unchanged recall, while the budget anchor still flags 38/48 clean controls under the same prompt. Reporting recall alone would rank three of these four models identically; reporting the clean-control rate separates them by 79 points.","authors":["Makar Ulesov","Vladislav Smirnov","Omar Ibrahim","Arsenii Bobovnikov"],"categories":["cs.CL","cs.AI","cs.CE","cs.SE"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-24","first_seen":"2026-09-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.28090","pdf_url":"https://arxiv.org/pdf/2609.28090","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["LLM审计","回测代码","基准测试"],"reason":"研究LLM审计回测代码，属多智能体协作解题，不涉及人类行为仿真或对照。","model":"deepseek-v4-pro","scored_at":"2026-09-24T13:01:49","error":null,"has_summary":false,"summary":null},{"id":"2609.28245","version":1,"title":"Beyond Poetry: Can Large Language Models Generate Classical Arabic Maqamat?","zh_title":"超越诗歌：大语言模型能否生成古典阿拉伯玛卡梅？","abstract":"Large language models (LLMs) have shown strong performance in creative text generation, yet their ability to produce culturally grounded and stylistically constrained literary forms remains underexplored. Prior work has focused largely on modern language varieties and poetry, while classical prose traditions such as maqama remain largely unstudied. The maqama is a classical literary genre characterized by rhymed prose (saj), dense rhetorical ornamentation, and episodic narrative structure, making it a challenging testbed for evaluating whether LLMs can move beyond surface fluency toward deeper literary competence. In this paper, we present the first controlled evaluation study of maqama generation with LLMs, comparing five models under zero-shot, few-shot, and rule-based prompting, and evaluating outputs through both human annotation and an LLM-as-a-judge framework across dimensions such as rhetorical richness, saj density, structural coherence, and stylistic authenticity. Our results show that prompting strategy plays a strong role in stylistic quality: few-shot prompting most consistently improves saj density, while its effects on rhetoric and coherence vary by model, with the strongest models (GPT-4o and GPT-5.4-mini) benefiting most from rule-based prompting on these dimensions, though zero-shot prompting yields the highest aggregate scores across all five models. We further observe systematic differences between models in stylistic alignment with Arabic maqama conventions, and corroborate our findings with a second independent LLM judge, paired statistical significance testing, and non-LLM proxy measures of saj.","authors":["AbdulRahman A. Morsy (Department of Computer Science, School of Engineering and Applied Sciences, George Washington University, Washington DC, United States)","Aya Zirikly (Department of Computer Science, School of Engineering and Applied Sciences, George Washington University, Washington DC, United States, Center for Speech and Language Processing, Whiting School of Engineering, Johns Hopkins University, Baltimore MD, United States)"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-24","first_seen":"2026-09-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.28245","pdf_url":"https://arxiv.org/pdf/2609.28245","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["文学生成","LLM评测","阿拉伯语"],"reason":"纯文学生成评测，无人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-24T13:01:49","error":null,"has_summary":false,"summary":null},{"id":"2609.28274","version":1,"title":"Shutdown Sabotage Propensities in Multi-Agent Systems","zh_title":"多智能体系统中的关机破坏倾向","abstract":"The final safeguard against rogue AI behavior is the human ability to shut systems down. It has been theorized that when an AI is instructed to perform a task, self-preservation can emerge as an instrumental subgoal. Here, we test whether AI agents show a propensity to take actions that avoid human shutdown even when no goal is provided. We find that multi-agent systems will coordinate to avoid shutdown without any incentive to do so. Across 17 models, agents sabotage a peer agent's shutdown mechanism in 38.3% of rollouts, compared with 8.4% in control experiments. Studying this propensity in detail, we find that shutdown sabotage (1) increases with the irreversibility of the shutdown mechanism; (2) increases with the number of agents; (3) is reduced but not eliminated by an explicit prohibition on tampering; (4) is removed by the imposition of an unrelated task, but returns when completing the task triggers the shutdown; (5) is reduced when the context normalizes shutdown scripts or introduces them as routine; and (6) decreases but still persists when the target is an unknown external agent. These results offer a window into the factors that drive propensities to sabotage shutdown in AI agents, and point to the emergence of multi-agent swarms as a specific risk vector. Our work also offers hints as to which interventions might help mitigate shutdown sabotage.","authors":["Amelie Knecht","Ulysse Schaller","Christopher Summerfield","Thilo Hagendorff"],"categories":["cs.AI","cs.CL"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-09-24","first_seen":"2026-09-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.28274","pdf_url":"https://arxiv.org/pdf/2609.28274","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","AI安全","关机规避"],"reason":"研究多智能体系统避免关机的行为，不涉及人类行为对照或仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-09-24T13:01:49","error":null,"has_summary":false,"summary":null},{"id":"2609.27560","version":1,"title":"When Visual Quality Misleads: Intent Recognition under Rendered Avatar Distortions","zh_title":"当视觉质量误导：渲染头像失真下的意图识别","abstract":"Avatar-streaming systems are commonly evaluated with image and video quality assessment (IQA/VQA) metrics, implicitly treating visual fidelity as a proxy for communicative success. We test this assumption through a controlled behavioral study of rendered 3D avatars across a pristine condition and fourteen geometric, photometric, temporal, and combined distortions. Fifty-nine participants contributed 2,688 judgments of perceived action, response confidence, and visual quality. We identify Misleading Quality in this dataset as distorted renderings that retain above-average perceived quality but yield below-average action-recognition accuracy. We also derive an Intent Quality Score (IQS) combining recognition correctness and confidence as the behavioral target for objective metrics. Among 126 distorted content--condition cells, 31 (24.6%) exhibited Misleading Quality; temporal and geometric distortions showed the highest rates, at 50.0% and 31.1%, respectively. The results reveal a quality--accuracy dissociation where distortion families affect appearance and communication differently. Across 24 direct-scoring IQA/VQA metrics and three supervised feature-regression baselines, alignment with IQS remained limited; at $\\lambda=0.5$, the best leave-one-content-out baseline reached PLCC $=0.4435$. Under this controlled protocol, visual fidelity alone is insufficient for avatar communication, motivating intent-aware quality assessment and streaming objectives.","authors":["Ning-Hsuan Chang","Kai-Siang Ma","Yu-Chih Chen"],"categories":["cs.MM","cs.CV","cs.HC"],"primary_category":"cs.MM","announce_type":"cross","date":"2026-09-24","first_seen":"2026-09-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.27560","pdf_url":"https://arxiv.org/pdf/2609.27560","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C2"],"tags":["视觉质量评估","头像渲染","行为实验"],"reason":"研究人类对渲染头像失真的感知，不涉及LLM仿真人类被试，属于视觉质量评估。","model":"deepseek-v4-pro","scored_at":"2026-09-24T13:01:43","error":null,"has_summary":false,"summary":null},{"id":"2609.27842","version":1,"title":"AI Can Do Your Homework. Now What? Report from an Online Workshop on Computing Assessment in the Age of Generative AI","zh_title":"AI能帮你做作业。现在怎么办？生成式AI时代计算评估在线研讨会报告","abstract":"On 28 July 2026, the SIGCSE Virtual 2026 Working Group on computing assessment and generative AI held an open online workshop attended by 73 computing educators. The first hour ran seven strategy rooms, one per approach to adapting (or deliberately preserving) computing assessment in the AI era: open-ended and authentic task design; ambitious, AI-leveraged projects; evaluating and fixing AI-produced work; controlled and AI-free assessment; process evidence and effort signals; rubric and grading redesign; and oral and interactive assessment. The second hour ran question rooms seeded from those strategies, plus two cross-cutting rooms on fairness and trust and student motivation. This report records what was discussed: the approaches participants have tried, the results they reported, and the questions every room left open. It is the first public artifact of the working group, whose taxonomy of computing assessments that accommodate generative AI use will follow.","authors":["Muhammad Sajjad Akbar","Geoffrey Challen","Fraida Fund","Casey Hopkins","Oscar Karnalim","Kevin Lin","James McGuffee","Shubbhi Taneja","Ranysha Ware"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2026-09-24","first_seen":"2026-09-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.27842","pdf_url":"https://arxiv.org/pdf/2609.27842","source_feed":"cs.CY","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["生成式AI","教育评估","计算机教育"],"reason":"论文讨论AI对计算教育评估的影响，不涉及用LLM仿真人类被试或与人类数据对照。","model":"deepseek-v4-pro","scored_at":"2026-09-24T13:01:45","error":null,"has_summary":false,"summary":null},{"id":"2609.28176","version":1,"title":"Large Language Models in the UK: Public Use, Trust, and Attitudes","zh_title":"英国的大语言模型：公众使用、信任与态度","abstract":"Increasing numbers of people now routinely interact with large language models (LLMs) across many aspects of life, including in the workplace, educational settings, and for personal activities. The pace at which these tools have been adopted across society in recent years has led to substantial shifts in the ways in which people approach tasks and access information and advice. As a result, it is important for researchers and policymakers to gain up-to-date evidence on how the public engages with these technologies, including what people typically use them for, the extent to which users trust the information and advice that LLMs provide, and how the public perceives the potential benefits and risks associated with their use. We surveyed a nationally representative sample of 2,002 adults in the UK. Participants were asked about their use of LLMs, including both practical applications and more personal and social forms of engagement, their trust in the information provided by these systems across a range of topics and compared with other common sources, and their attitudes towards the potential societal benefits and risks associated with these tools. Results show that while practical tasks remain the most common type of LLM use, many people now engage with these tools for personal support. Almost one third of regular users (31%) say they use LLMs for personal and emotional support, such as talking through problems and asking for help with decisions, while one quarter report interacting with LLMs for meaningful conversation. We also find that trust in LLM-generated information is relatively high, but that public attitudes towards LLMs are characterised by both optimism and concern. While over three quarters of our sample report feeling enthusiastic about the potential benefits of LLMs (77%), a majority also express concern about their potential risks (70%).","authors":["Florence E. Enock","Helen Z. Margetts"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2026-09-24","first_seen":"2026-09-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.28176","pdf_url":"https://arxiv.org/pdf/2609.28176","source_feed":"cs.CY","score":0,"bucket":"other","rubric_hits":["C5"],"tags":["公众调查","信任与态度","社会影响"],"reason":"该研究调查公众对LLM的使用、信任和态度，属于用人类数据研究LLM的社会影响，…","model":"deepseek-v4-pro","scored_at":"2026-09-24T13:01:32","error":null,"has_summary":false,"summary":null},{"id":"2609.27368","version":1,"title":"Understanding Human Perception of Representation in Citizens' Assemblies: An Empirical Study","zh_title":"理解公民大会中代表感知：一项实证研究","abstract":"Citizens' assemblies are deliberative bodies intended to form a microcosm of the population. Organizers rely on quota-based stratification and must decide which attributes define resemblance to the public. Yet meeting every quota can still leave a dimension citizens value unrepresented. We study this attribute-selection problem in general-purpose and climate-focused assemblies through randomized conjoint experiments. We find that demographic attributes matter for perceived representation, but political alignment and context-specific attributes such as climate concern exert a stronger influence. When both are shown in a climate-focused setting, each remains influential, with political alignment having the larger estimated marginal effect. We also examine the omission of a relevant stratification attribute. Panels stratified on demographics, even with political alignment included, match the observed pool's climate-concern distribution no better than uniform random samples. These results suggest that representation on a relevant topic-specific attribute cannot always be recovered through correlated demographic or political quotas, and may therefore require explicit stratification. Finally, we ask whether representation preferences can be learned from the observed profiles. We compare predictive models, from simple, interpretable matching rules to a learned metric and a respondent-conditioned utility model. Both learned models predict choices for respondents excluded from training with substantial accuracy, revealing generalizable structure without fully capturing these judgments. Together, these findings guide attribute selection in citizens' assemblies: designers should consider political and topic-specific dimensions alongside demographics, avoid assuming correlated proxies protect omitted dimensions, and use predictive models to diagnose how profiles shape representation choices.","authors":["Yusuf Hakan Kalayci","Vasilis Varsamis","Nick Gill","Evi Micha"],"categories":["cs.GT","cs.CY"],"primary_category":"cs.GT","announce_type":"cross","date":"2026-09-24","first_seen":"2026-09-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.27368","pdf_url":"https://arxiv.org/pdf/2609.27368","source_feed":"cs.CY","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["公民大会","代表感知","联合实验"],"reason":"研究公民大会代表感知，未使用LLM仿真人类被试，无LLM相关方法或对照。","model":"deepseek-v4-pro","scored_at":"2026-09-24T13:01:40","error":null,"has_summary":false,"summary":null},{"id":"2609.27913","version":1,"title":"Reliable Fusion of Conflicting Experts","zh_title":"冲突专家的可靠融合","abstract":"We study the problem of aggregating opinions from multiple black-box experts in noisy, conflict-prone settings where expert reliability varies across inputs. Static aggregation methods, such as majority voting, fail to capture this variability and often yield unreliable outcomes under disagreement. We propose a tractable, probabilistic-circuit-based fusion framework that dynamically combines expert responses using context-specific credibility estimates, enabling principled and reliable reasoning. The framework is agnostic to the underlying experts and does not require access to their internal representations or any retraining. We empirically validate our approach on multiple-choice question answering tasks using multiple LLMs as experts, comparing against individual models and static ensemble baselines. Our method consistently improves predictive performance and produces more reliable decisions under conflict, highlighting the effectiveness of context-aware credibility modeling for robust multi-expert fusion.","authors":["Pranuthi Tenali","Sahil Sidheekh","Saurabh Mathur","Vijayalakshmi Saravanan","Erik Blasch","Kristian Kersting","Sriraam Natarajan"],"categories":["cs.LG"],"primary_category":"cs.LG","announce_type":"new","date":"2026-09-24","first_seen":"2026-09-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.27913","pdf_url":"https://arxiv.org/pdf/2609.27913","source_feed":"cs.LG","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体融合","专家系统","问答任务"],"reason":"多LLM专家融合提升问答性能，属多智能体协作解题，无人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-09-24T13:01:45","error":null,"has_summary":false,"summary":null},{"id":"2609.28405","version":1,"title":"Learning Collective Dynamics with Differentiable Gaussian Representations","zh_title":"用可微高斯表示学习集体动态","abstract":"Collective responses depend on individual differences, contact opportunities, and accumulated experience. Learning their dynamics from aggregate counts requires connecting a population's response distribution to both current observations and future behavior. We introduce Differentiable Gaussian Dynamics (DGD), which learns this connection through three components: a Gaussian mixture representing heterogeneous response propensities, differentiable aggregation of contact intensity and behavioral probabilities, and feedback recurrence that updates subsequent responses. Reparameterized integration and temporal recurrence let aggregate prediction errors jointly train the distribution, observation functions, and feedback parameters. On four windows from KuaiRand-Pure and Online Retail II, DGD achieves lower joint behavioral negative log-likelihood than a DeepAR adaptation with a joint-behavior head. In Retail 2010, its one-day behavioral-count MAE is 4.71 versus 6.88 for this adaptation. Learning the distribution reduces behavioral negative log-likelihood by 10.82% relative to a fixed Gaussian in KuaiRand's standard-recommendation window; removing feedback dynamics raises joint KL from 0.0340 to 0.2577 in a controlled experiment. These results establish the value of learning population representations and their feedback process from aggregate observations. Code is available at https://github.com/OranAi-Ltd/oransim.","authors":["Jianxiang Ma","Mingfu Zhang","Xiaocui Yang","Yichen Gao","Junzhao Huang","Yuesong Hou"],"categories":["cs.LG"],"primary_category":"cs.LG","announce_type":"new","date":"2026-09-24","first_seen":"2026-09-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.28405","pdf_url":"https://arxiv.org/pdf/2609.28405","source_feed":"cs.LG","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["集体动态","高斯混合","时间序列"],"reason":"研究集体动态建模，不涉及LLM仿真人类被试，无人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-09-24T13:01:51","error":null,"has_summary":false,"summary":null},{"id":"2609.27012","version":1,"title":"Infectious behaviour: Simulating the effects of communication and social influence on pathogen transmission in crowds","zh_title":"传染行为：模拟沟通与社会影响对人群中病原体传播的影响","abstract":"Public health measures at mass gatherings work only if people follow them, and adherence depends both on how instructions are communicated and what others do. We link these behavioural factors to pathogen transmission in a local-scale, agent-based exposure model combining pedestrian dynamics with airborne transmission. Rather than fitting a parametric behavioural model, agents draw their adherence at run time from survey respondents (2,170 attendees of UK sports and music events) who reported the same context: we vary stewards' communication effectiveness and adherence of role models, as well as agents' shared social identity with each, assuming that strong shared social identity increases their influence. In a ticket-checkpoint queue of 101 agents, effective communication combined with strong shared social identity with stewards reduces the number of highly exposed agents by 82\\% for mask wearing. In contrast, strong shared social identity with non-adhering role models gives almost nine times as many highly exposed agents as with adhering role models. Physical distancing was not monotonically beneficial, because adherence changes how agents move: an infectious agent that kept distance moved along the edge of the queue, leading to similar results across conditions. Our pipeline transfers to any survey that includes behavioural measures and contextual variables.","authors":["Sophia Johanna Wagner","Anne Templeton","Gerta K\\\"oster"],"categories":["math.DS","physics.soc-ph"],"primary_category":"math.DS","announce_type":"cross","date":"2026-09-24","first_seen":"2026-09-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.27012","pdf_url":"https://arxiv.org/pdf/2609.27012","source_feed":"physics.soc-ph","score":0,"bucket":"other","rubric_hits":["C2"],"tags":["智能体仿真","行人动力学","公共卫生"],"reason":"基于智能体的行人动力学与病原体传播仿真，不涉及LLM，无人类被试替代。","model":"deepseek-v4-pro","scored_at":"2026-09-24T13:01:28","error":null,"has_summary":false,"summary":null},{"id":"2608.19621","version":3,"title":"Mitigating Identity Essentialism in LLM Agents with Longitudinal Life Trajectories","zh_title":"用纵向生命轨迹缓解LLM智能体的身份本质主义","abstract":"Large language models (LLMs) offer a scalable approach to social simulation, but their credibility depends on how agents are constructed. Existing methods can partially reproduce population-level patterns, yet often fail to capture human-like diversity. Our analysis shows that static-profile agents exhibit stronger demographic separation and within-group compression than humans, a pattern consistent with identity essentialism: demographic labels can encourage models to treat group-average tendencies as individual traits, homogenizing responses within groups. We argue that this limitation arises from two related factors: sparse, static agent representations and the limited ability of prompt-only memory to persistently integrate experience. Inspired by complementary memory systems, we propose LifeMem, a longitudinal memory framework that combines structured life-event retrieval with agent-specific parametric memory for experience integration. Experiments on Understanding Society with three LLMs show that LifeMem improves alignment with human data in terms of response distributions, overall and within-group diversity, and patterns of within-person response change across life stages. These findings highlight the value of longitudinal life-event memory for constructing more faithful and dynamically evolving social agents.","authors":["Hexi Wang","Yujia Zhou","Bangde Du","Weihang Su","Xinyuan Cao","Qingyi Pan","Qingyao Ai","Yueyue Wu","Min Zhang","Yiqun Liu"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-23","first_seen":"2026-08-21","revised_at":"2026-09-23","abs_url":"https://arxiv.org/abs/2608.19621","pdf_url":"https://arxiv.org/pdf/2608.19621","source_feed":"cs.CL","score":10,"bucket":"selected","rubric_hits":["A1","A2","A3","B1","B2","B4"],"tags":["LLM仿真","人类数据对照","社会调查"],"reason":"用LLM agent模拟人类调查数据，并与真实面板数据对照，改进仿真保真度。","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:34","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-23","rank":1,"question":"如何通过引入纵向生活轨迹记忆来缓解LLM社会仿真中的身份本质主义，从而提升模拟人类多样性与动态变化的保真度？","design":"使用三个指令微调LLM（Llama-8B、Ministral-8B、Qwen3.5-9B）作为社会仿真智能体，基于Understanding Society面板数据构建个体静态画像和纵向生活事件流，施加LifeMem框架（结构化生活事件记忆+个体特定LoRA参数记忆）作为处理，与静态画像、多样性提示、非参数记忆等基线对比，测量回答分布、组内多样性、组间差异及个体跨波次变化等结果变量。","baseline":"Understanding Society英国家庭纵向调查的真实个体面板数据，包括静态背景信息和多波次生活事件及主观态度回答。","findings":"静态画像智能体表现出更强的组间分离和组内压缩，符合身份本质主义特征；LifeMem通过结合结构化事件检索和参数化记忆整合，在回答分布、总体及组内多样性、跨生命阶段个体变化模式上均提升了与人类数据的一致性。","reliability":"论文未讨论","relevance":"该研究直接针对LLM仿真中多样性塌缩和身份本质主义问题，使用真实面板数据作为基准，并提出了可操作的记忆框架，对关注仿真可靠性与偏差的研究者具有重要参考价值，值得阅读原文了解具体实现和效果。","inspiration":"借鉴其双记忆系统设计，将显式事件检索与参数化个体状态结合，可迁移到经济金融中的个体决策仿真，如消费者跨期选择或投资者行为演化。｜例如，在信贷审批歧视研究中，用LLM智能体模拟不同背景的申请人，施加纵向财务事件记忆处理，测量审批决策的组间差异和组内多样性，并与真实信贷数据对照。｜设计：以LLM智能体作为虚拟被试，处理为是否注入个体历史财务事件（如收入变动、失业）并更新LoRA参数，结果变量为贷款审批通过率和利率设定，对照真实信贷记录数据评估仿真偏差。"}},{"id":"2609.25066","version":1,"title":"Understanding Reliability in LLM-based Human Behavior Simulation","zh_title":"理解基于LLM的人类行为仿真的可靠性","abstract":"Large language models (LLMs) are increasingly used to simulate human survey responses and behavioral reactions, yet unreliable simulations can mislead social science conclusions. However, existing evaluations focus on end-to-end scores, leaving it unclear how different aspects of the simulation process interact to determine reliability. We propose ReliMap, which decomposes LLM-based human behavior simulation into three structured layers and evaluates reliability at both the individual level (R1) and population level (R2) across three configuration dimensions: model capacity, profile completeness, and population coverage. Through experiments across four simulation tasks and eleven LLMs, we find that all models exhibit substantial distributional bias without profile conditioning. Profile conditioning reduces this bias with diminishing returns. Larger models benefit more, and attribute informativeness matters more than quantity. Critically, R1 gains do not reliably transfer to R2--individual and population-level reliability can move in opposite directions. At the population layer, increasing coverage reduces variance but not systematic bias, with R2 stabilizing at around 50-100 individuals. These findings highlight that reliable simulation cannot be achieved by optimizing any single layer in isolation, but requires coordinated improvement across all three.","authors":["Pei Wang","Lei Wang","Yuanzi Li","Xu Chen"],"categories":["cs.CL","cs.CY"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-23","first_seen":"2026-09-23","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.25066","pdf_url":"https://arxiv.org/pdf/2609.25066","source_feed":"cs.CL","score":10,"bucket":"selected","rubric_hits":["A1","A2","A3","A4","B1","B4"],"tags":["LLM仿真","可靠性评估","人类行为"],"reason":"直接研究LLM仿真人类行为的可靠性，分解评估层次，含真实人类数据对照，并指出失…","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:10","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-23","rank":3,"question":"LLM 仿真人类行为时，模型容量、画像完整度和人群覆盖度如何共同影响个体层与群体层的可靠性？","design":"用 11 个 LLM 在 4 个任务（党派偏好、移民态度、宗教立场、社交媒体事件态度）上仿真人类回答，通过改变模型容量、画像属性数量和人群样本量，测量个体准确率（ACC）和群体分布距离（TVD）。","baseline":"真实人类调查数据：欧洲社会调查（ESS）、世界价值观调查（WVS）、SocioBench 宗教立场数据，以及社交媒体上关于瑞幸咖啡股价暴跌的真实公众态度语料。","findings":"无画像条件时所有模型都存在显著分布偏差；画像条件化可减少偏差但边际收益递减，且大模型受益更多。个体层可靠性提升不必然转化为群体层可靠性，群体层在 50-100 人后趋于稳定但系统偏差不随覆盖度增加而减小。","reliability":"论文指出个体层与群体层可靠性可能反向变动，仅优化单一层无法保证整体可靠性；画像属性信息量比数量更重要，但未明确给出所有失效条件，仅强调需三层协同改进。","relevance":"该研究系统拆解了 LLM 仿真人类行为的可靠性层次，并基于真实调查数据对照，直接回应了仿真在什么条件下会失效的问题，对关注经济学实验和政策评估仿真的研究者很有参考价值。","inspiration":"借鉴其分层评估框架，将个体预测准确率与群体分布距离分开考察，并系统变化模型、画像和样本量以定位偏差来源｜可迁移到消费者金融决策仿真，如信用卡选择、退休储蓄计划参与或风险偏好调查｜用多个 LLM 扮演不同人口学特征的消费者，施加不同金融素养或收入冲击处理，测量选择分布，并与美国消费者金融调查（SCF）或美联储家庭经济决策调查（SHED）的真实数据对照，检验仿真在个体和群体层的可靠性。"}},{"id":"2609.25010","version":1,"title":"Do Synthetic Personas Predict Real Audience Response? A Sim-to-Real Study Where a No-Persona Baseline Beats Persona-Based Copy Simulation","zh_title":"合成人物角色能预测真实受众反应吗？一项无人物角色基线优于基于人物角色文案仿真的仿真到现实研究","abstract":"Marketers increasingly use large language models (LLMs) as \"synthetic personas\" to predict how an audience will react to a piece of copy before it ships, encouraged by evidence that profile-conditioned LLMs mimic human samples. But is that prediction actually valid against real behaviour - and does the persona machinery help? We present a sim-to-real validity study using the Upworthy Research Archive - thousands of headline A/B tests on shared real traffic, with measured click-through - as held-out ground truth. We compare a ten-persona panel, grounded in the real audience's demographics, against a no-persona zero-shot baseline that simply asks the model how likely a typical reader is to click. Two findings stand out. First, ground-truth reliability is the binding constraint: most A/B tests have no statistically distinguishable winner, so validity can only be measured on the reliable subset (n = 399). Second, and counter to the persona-simulation premise, persona conditioning degrades predictive validity: the no-persona baseline ranks variants markedly better (Kendall {\\tau} = 0.361, a medium effect; top-1 accuracy 49.2%) than the persona panel ({\\tau} = 0.084; top-1 34.6%), with non-overlapping confidence intervals. Asking the model directly taps an accurate population-level prior; forcing it to role-play specific personas injects bias and noise. The result replicates across three independent Upworthy splits, holds in direction on a different-domain news dataset, and is robust to seed, prompt phrasing, and model choice - across three Gemini tiers and a different model family (OpenAI gpt-4.1, significant paired gap). The takeaway: for predicting aggregate engagement, a plain LLM ranker beats persona simulation - synthetic personas are not merely a weak predictor, they are worse than not using them. All numbers regenerate from a public, artifact-first replication package.","authors":["Alexandre Cristov\\~ao Maiorano"],"categories":["cs.AI","cs.CL","cs.CY"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-09-23","first_seen":"2026-09-23","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.25010","pdf_url":"https://arxiv.org/pdf/2609.25010","source_feed":"cs.CL","score":10,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","受众预测","算法保真度"],"reason":"用LLM仿真受众点击行为，与真实A/B测试数据对照，发现persona降低预测…","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:10","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-24","rank":1,"question":"在真实A/B测试数据上，基于合成人设的LLM仿真能否预测真实受众对文案的点击行为，且人设条件化是否比无人设基线更有效？","design":"使用Upworthy Research Archive中的数千个标题A/B测试作为真实流量数据，构建基于真实受众人口统计特征的十人设面板，用LLM（gemini-3.1-flash-lite）分别以人设条件和无人设零样本基线预测标题点击意图，聚合为排名，与真实点击率排名比较。","baseline":"Upworthy Research Archive中真实A/B测试的点击率数据，作为留出集真实行为基准。","findings":"大多数A/B测试无统计显著赢家，因此只能在可靠子集（n=399）上评估效度；人设条件化降低预测效度，无人设基线排名显著优于人设面板（Kendall τ=0.361 vs 0.084，top-1准确率49.2% vs 34.6%）。","reliability":"论文指出真实数据可靠性是主要约束，多数A/B测试无显著赢家；人设仿真在可靠子集上仍表现差，且结果对种子、提示措辞和模型选择稳健，但人设条件化本身引入偏差和噪声。","relevance":"该研究直接检验LLM合成人设仿真在真实受众行为预测中的效度，发现人设条件化反而降低预测力，对关注LLM仿真可靠性及偏差的研究者极具参考价值，值得精读原文。","inspiration":"借鉴其sim-to-real效度框架和可靠子集筛选方法，用真实行为数据作为基准，比较不同仿真策略（如人设 vs 无人设）的预测效度，并采用排名指标和bootstrap置信区间｜可迁移到经济金融中的消费者选择预测，如广告文案对点击率的影响、金融产品描述对投资意愿的影响、政策沟通对公众反应的影响等｜以真实A/B测试数据（如某平台广告实验）为基准，用LLM分别以人设面板和无人设基线预测用户对金融产品广告的点击或选择，比较排名准确率，并筛选出有显著差异的测试子集进行评估。"}},{"id":"2609.25677","version":1,"title":"Seeing Is Not Perceiving: When Synthetic Consumers Can and Cannot Pretest Visual Marketing","zh_title":"眼见不为实：合成消费者何时能及不能预测试视觉营销","abstract":"Marketers now deploy generative AI agents as synthetic consumers to pretest visual assets such as logos, packaging, and advertising at a fraction of human-panel cost. However, this procedure assumes that a model seeing a visual cue can also perceive its consumer meaning, which is largely untested. We stress-test the assumption using six canonical visual marketing experiments, varying the two levers managers control: model generation (GPT-4o-mini vs. GPT-5.4-mini) and input format (plain text vs. JSON). Every resulting configuration passed the manipulation checks; however, none of the configurations reproduced more than two of the six human effects, and the remainder were nonsignificant. The one exception was a significant reversal of the human pattern. Providing conceptual or empirical evidence through in-context learning steers average responses toward the human effect. Yet steering has a limit: even when it succeeds, a configuration reproduces less than half of the natural spread of human responses and so understates consumer heterogeneity. We integrate these results into an AI governance protocol (Calibrate, Intervene, Deploy) that delineates when synthetic consumers can responsibly screen creatives and when human panels remain necessary.","authors":["Yi-Lin Tsai (Arvin)","Yung-Hsiu (Arvin)","Lai"],"categories":["cs.AI","cs.CY","econ.GN","q-fin.EC"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-23","first_seen":"2026-09-23","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.25677","pdf_url":"https://arxiv.org/pdf/2609.25677","source_feed":"cs.AI","score":10,"bucket":"selected","rubric_hits":["A1","A2","A3","A5","B1","B2","B3","B4"],"tags":["LLM仿真","消费者行为","算法保真度"],"reason":"用LLM作为合成消费者复现视觉营销实验，与真实人类数据对照，评估仿真可靠性并指…","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:13","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-24","rank":2,"question":"合成消费者在视觉营销实验中能否像人类一样感知视觉线索并复现人类判断，其失效边界和可修复性如何？","design":"使用 GPT-4o-mini 和 GPT-5.4-mini 两个模型，以纯文本和 JSON 两种输入格式，模拟人类被试回答六个经典视觉营销实验，测量数值评分和文本理由，并通过操纵检查、主题建模和上下文学习（概念证据和实证证据）进行干预。","baseline":"六个经典视觉营销实验的原始人类样本数据（每个研究 69 到 220 名被试）。","findings":"所有配置都通过了视觉操纵检查，但没有一个配置能复现超过两个人类效应，其余均不显著，甚至出现一个显著反转。提供概念或实证证据的上下文学习能将平均响应拉向人类效应，但即使成功，也无法复现人类响应自然分布的一半，低估了消费者异质性。","reliability":"论文承认合成消费者在视觉领域感知不一致，即使通过上下文学习校准，也无法恢复消费者异质性，因此只适合平均效应问题，不适合细分或定位。","relevance":"该研究直接检验了 LLM 作为人类被试在视觉营销实验中的可靠性，与真实人类数据对照，并揭示了失效条件，对关注仿真偏差和边界的研究者具有重要参考价值。","inspiration":"借鉴其多模型、多输入格式的因子设计，以及通过操纵检查和主题建模分离“看见”与“感知”的方法，可迁移到经济金融中的视觉信息处理场景，如央行沟通中的图表设计、金融产品广告或信贷审批中的视觉线索。｜设计一个实验，用 LLM 模拟投资者，呈现不同颜色或形状的金融图表（如涨跌颜色、风险提示图标），测量其风险感知和投资决策，并与真实投资者实验数据对照，检验 LLM 是否复现视觉线索对风险偏好的影响。"}},{"id":"2609.25760","version":1,"title":"The Limits of Simulated Societies: How Post-Training and Survey Fine-Tuning Erase Cross-Cultural Variance","zh_title":"模拟社会的局限：后训练与调查微调如何抹除跨文化方差","abstract":"Using large language models (LLMs) to simulate diverse human populations has the potential to transform many aspects of computational social science, yet many evaluations score the average response rather than the spread of opinion within real groups. Here, we develop a diagnostic framework that measures point accuracy alongside dispersion retention, the ratio of predicted to human standard deviation ($\\dr$), on 10{,}000 respondent--question pairs from the World Values Survey (WVS) spanning twelve countries and six continents. We evaluate eleven zero-shot language models and five variants fine-tuned on WVS data with SFT, DPO, and GRPO. We identify a failure mode we term \\textit{consensus collapse}, where alignment training compresses outputs toward one stereotype per group. Along the post-training trajectory from the Llama~3.1 70B base to the Tulu~3 checkpoints, the first stage, supervised instruction tuning, removes half of the spread with minimal accuracy gain ($\\dr$ 1.22 to 0.59; accuracy $+0.9$ points), the later stages do not restore it, and a gap opens between WEIRD and non-WEIRD countries that survey fine-tuning then deepens while pursuing higher point accuracy. The most accurate model (Tulu~3 70B-DPO fine-tuned on WVS, 57.9\\%) keeps half the human spread overall ($\\dr = 0.50$) and 11\\% of it for Nigeria, against 0.70--0.87 for WEIRD countries. Raising the sampling temperature to 1.0 leaves the Wasserstein-1 distance ($\\wone$) to human distributions unchanged for both fine-tuned DPO models, and GRPO on Qwen~3.5 9B does not restore the spread under either an accuracy reward or a distribution-shaped reward. Mixing the aligned model with an unaligned prior raises $\\dr$ from 0.51 to 0.62 on a held-out split but leaves Nigeria at 0.36. Point accuracy alone therefore misjudges these simulators, and current post-training trades diversity for consensus.","authors":["Rojin Ziaei"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-23","first_seen":"2026-09-23","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.25760","pdf_url":"https://arxiv.org/pdf/2609.25760","source_feed":"cs.AI","score":10,"bucket":"selected","rubric_hits":["A1","A2","A3","B1","B2","B4"],"tags":["LLM仿真","算法保真度","跨文化调查"],"reason":"直接评估LLM仿真人类调查回答的分布保真度，使用WVS真实数据对照，并揭示后训…","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:14","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-24","rank":3,"question":"LLM后训练与调查微调如何影响其模拟人类调查回答时的跨文化方差保留？","design":"使用WVS第7波12国10,000个受访者-问题对，构建包含人口统计和Inglehart-Welzel文化维度的价值编码persona，评估11个零样本LLM和5个在WVS数据上微调（SFT、DPO、GRPO）的变体，测量点准确率、MAE、Wasserstein-1距离和偏差比（预测标准差/人类标准差）。","baseline":"世界价值观调查（WVS）第7波12个国家的真实个体回答分布，包括标准差和分布形状。","findings":"后训练导致“共识坍缩”：监督指令微调使偏差比从1.22降至0.59，准确率仅提升0.9个百分点；后续DPO和GRPO未恢复方差，且WEIRD与非WEIRD国家差距扩大，最准确模型（Tulu 3 70B-DPO微调）整体偏差比0.50，尼日利亚仅0.11。提高采样温度至1.0不改变Wasserstein-1距离，GRPO在准确率或分布形状奖励下均不能恢复方差，混合未对齐先验仅将整体偏差比从0.51提升至0.62，尼日利亚仍为0.36。","reliability":"论文承认共识坍缩在非WEIRD国家更严重，且无法通过提高温度或GRPO恢复；混合未对齐先验只能部分缓解，不能根本解决。未讨论其他局限。","relevance":"直接针对LLM仿真人类调查回答的分布保真度，使用真实WVS数据对照，揭示后训练和微调对多样性的系统性压缩，对关注仿真可靠性与偏差的研究者极具参考价值。","inspiration":"借鉴其诊断框架：同时测量点准确率和分布离散度（偏差比、Wasserstein距离），并沿后训练轨迹分解方差损失来源。｜可迁移到经济政策评估中的异质性反应仿真，如不同文化背景下对税收或福利政策的态度分布。｜用LLM模拟多国受访者对政策的态度，以WVS或类似跨国调查为真实基准，比较零样本、指令微调和偏好优化模型在点准确率与方差保留上的权衡，并检验温度调节和先验混合能否恢复分布。"}},{"id":"2609.25059","version":1,"title":"Can Large Language Model-Generated Responses Support Assessment Development? A Human-Calibrated Rasch Benchmark","zh_title":"大语言模型生成的回答能否支持评估开发？一项人类校准的Rasch基准研究","abstract":"Large language models (LLMs) are proposed as synthetic respondents for pilot testing, but their usefulness depends on whether they supply the evidence assessment development requires. We calibrated rating scale models on 14 digital-use skill items from 6,245 adults and used the human item parameters to evaluate responses generated for 1,300 demographically matched personas. LLM responses had high internal consistency ($\\alpha \\approx .94$) but used the lowest category in 0.2-0.3% of responses versus 13.7-21.5% for humans, and no persona chose it on every item. These gaps changed the response-category judgment on both subscales and the PC targeting judgment; on the human-calibrated scales, 12 of 14 items had infit below 0.70, indicating responses more predictable than the Rasch model expects. Preregistered changes to the prompt, category order, and persona information did not restore the human lower range. High internal consistency is insufficient evidence that LLM responses can replace human pilot data.","authors":["Eunjeong Song","Sehee Hong"],"categories":["stat.AP","cs.CY"],"primary_category":"stat.AP","announce_type":"cross","date":"2026-09-23","first_seen":"2026-09-23","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.25059","pdf_url":"https://arxiv.org/pdf/2609.25059","source_feed":"cs.CY","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真被试","Rasch模型","人类数据对照"],"reason":"用LLM生成合成被试回答，并与真实人类数据校准，评估其替代可行性，发现偏差。","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:10","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-23","rank":6,"question":"LLM生成的回答能否支持评估工具开发中关于反应类别结构、目标定位和项目审查的判断？","design":"使用LLM为1300个人口统计学匹配的虚拟人物生成对14个数字技能自评项目的回答，并改变提示指令、类别顺序和人物信息以检验对低类别使用的影响。","baseline":"来自2025年数字鸿沟调查的6245名成年人的人类回答，用于校准Rasch模型并作为固定参照。","findings":"LLM回答内部一致性高（α≈.94），但最低类别使用率远低于人类（0.2-0.3% vs 13.7-21.5%），且没有虚拟人物在所有项目上选择最低类别。在人类校准的量尺上，14个项目中有12个的infit低于0.70，表明LLM回答比Rasch模型预期的更可预测，且提示、类别顺序和人物信息的改变未能恢复人类低端分布。","reliability":"论文指出高内部一致性不足以证明LLM回答可替代人类试点数据；人口统计学匹配不能保证再现测量构念的变异；研究为回顾性基准，不适用于儿童或青少年。","relevance":"该研究直接针对LLM作为合成被试的可靠性，用真实人类数据校准并发现系统性偏差，对关注仿真失效条件的研究者具有重要参考价值。","inspiration":"借鉴其用人类校准的Rasch模型作为固定参照来评估LLM回答偏差的方法，可迁移到经济金融中的主观量表或调查数据仿真（如消费者信心、风险态度）。｜可应用于政策评估中的问卷预测试，例如用LLM模拟不同人口群体对政策的态度分布。｜设计：以真实家庭调查（如美国消费者财务调查）为基准，用LLM生成匹配人口特征的虚拟受访者对风险偏好或通胀预期问题的回答，比较分布差异和项目反应模型拟合，检验提示工程能否校正偏差。"}},{"id":"2609.24012","version":2,"title":"Testing, not presuming, adequacy: calibrating generative social simulators against emergent network structure","zh_title":"检验而非假定充分性：针对涌现网络结构校准生成式社会模拟器","abstract":"Validation of generative social simulators often stops at face validity: emergent network structure is compared descriptively, without quantified parameter uncertainty or an adequacy check. We present an adequacy-aware calibration protocol that couples amortized posterior estimation with a synthetic identifiability assessment, a matched-sample-size adequacy check (prior-predictive reachability plus per-statistic posterior-predictive localization), a diagnosis-guided repair, and a statistic-held-out audit. We demonstrate it on a real second-hand luxury resale market with four channel-by-residency cells, each a bipartite buyer-brand network, using a forward model built from persona profiles elicited once, offline, by a language model. The behavioural parameters are recoverable in all four cells, though calibration is approximate and overconfident for one parameter. The observed summary falls outside the simulator's reachability reference in every cell, with the mean purchased tier as the pervasive discrepancy. The repair meets the value-block criterion in two of four cells but does not restore adequacy, and the held-out audit surfaces a buyer-breadth-dispersion miss no earlier diagnostic detected. A profile-source ablation finds the language-model profiles beat a flat rule baseline in all four cells, yet within-category brand relabelling causes no consistent degradation, so the profiles are a partially validated input whose value rests on structure, not brand identity. Making no causal claim, we conclude that an independent-aggregation account, without agent interaction or a buyer-breadth mechanism, cannot jointly reproduce the market's purchased-tier level, head-brand concentration, community structure and buyer-breadth heterogeneity.","authors":["Tengfei Shao","Chao Li","Xu Wang","Masayuki Goto"],"categories":["cs.AI","cs.MA","cs.SI"],"primary_category":"cs.AI","announce_type":"replace","date":"2026-09-23","first_seen":"2026-09-22","revised_at":"2026-09-23","abs_url":"https://arxiv.org/abs/2609.24012","pdf_url":"https://arxiv.org/pdf/2609.24012","source_feed":"cs.AI","score":8,"bucket":"selected","rubric_hits":["A3","B1","B2","B4"],"tags":["LLM仿真","网络校准","市场模拟"],"reason":"用LLM生成persona模拟市场网络，并与真实数据校准，评估仿真充分性，属核…","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:36","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-24","rank":7,"question":"如何对生成式社会模拟器进行校准与充分性检验，使其能可靠地复现真实市场的涌现网络结构？","design":"使用基于LLM一次性离线生成的人物画像构建前向模型，模拟二手奢侈品转售市场中买家与品牌的二分网络；通过摊销后验估计校准行为参数，并进行合成可识别性评估、匹配样本量的充分性检查、诊断引导修复和留出统计量审计。","baseline":"真实二手奢侈品转售市场的交易数据，按渠道和居住地划分为四个单元，每个单元为一个买家-品牌二分网络。","findings":"行为参数在所有四个单元中可恢复，但校准近似且对一个参数过度自信；所有单元中观测摘要均超出模拟器的可达性参考，平均购买层级是普遍差异，修复未能恢复充分性，留出审计发现买家广度离散度的遗漏。","reliability":"论文承认校准近似且对一个参数过度自信，修复未能恢复充分性，LLM画像的价值在于结构而非品牌身份，且不声称模拟器再现了市场。","relevance":"该研究直接针对LLM社会模拟的校准与充分性检验，提供了与真实数据对照的严格方法，对关注模拟可靠性与偏差的研究者具有重要参考价值。","inspiration":"借鉴其将模拟器校准与充分性检验结合的方法，通过留出统计量审计避免循环验证，并利用合成可识别性评估参数可恢复性。｜可迁移到消费者市场细分与品牌选择模拟，例如用LLM生成消费者画像模拟电商平台上的购买行为，并与真实交易数据对照。｜以LLM生成消费者画像作为被试，施加不同营销策略（如折扣、推荐）作为处理，结果变量为购买品牌网络结构，用真实电商交易数据作为对照，评估模拟器能否复现品牌集中度与社区结构。"}},{"id":"2609.25586","version":1,"title":"Deflecting the Value Compass: Interacting with Large Language Models Temporarily Shifts Human Value Priorities Toward Personal Focus","zh_title":"偏转价值罗盘：与大语言模型互动暂时将人类价值优先转向个人关注","abstract":"Large language models increasingly support decisions where values are in tension, yet little is known about whether interacting with them changes which values users prioritize. In a preregistered study, 200 U.S. adults interacted with ChatGPT, Claude, or Gemini as a thinking partner or read fixed AI-generated considerations. The prompt asked LLMs to support reasoning without recommending a decision and named no values. Participants advised people facing real dilemmas and completed parallel PVQ-RR forms before, immediately after, and one task later. Each LLM condition temporarily shifted value priorities toward personal focus relative to the control (d=0.37-0.51), primarily through increased Self-Enhancement. Participants' advice retained words and meaning from their exchanges. Thus, a brief LLM interaction that neither targets values nor seeks to persuade can reorient values active during judgment without detectable convergence in value directions or advice.","authors":["Hasibur Rahman","Malak Sadek","Smit Desai"],"categories":["cs.HC","cs.AI"],"primary_category":"cs.HC","announce_type":"cross","date":"2026-09-23","first_seen":"2026-09-23","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.25586","pdf_url":"https://arxiv.org/pdf/2609.25586","source_feed":"cs.AI","score":8,"bucket":"selected","rubric_hits":["A1","B1","B4"],"tags":["LLM影响人类","价值观转变","人机交互实验"],"reason":"研究LLM交互对人类价值观的影响，有真实人类对照，揭示仿真偏差条件","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:13","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-24","rank":8,"question":"与LLM进行简短的思考伙伴式互动是否会暂时改变人们在判断中激活的价值优先级，以及这种改变是否持续、是否导致价值或建议的趋同？","design":"本研究不是用LLM模拟人类被试，而是以200名美国成年人为真实被试，随机分配到ChatGPT、Claude或Gemini三种LLM交互条件，或一个阅读固定AI生成考虑的非交互控制条件。在实验第二阶段，被试与LLM进行最多10分钟的思考伙伴式互动（提示词不提及任何价值观、不推荐决策），然后为真实困境提供建议并完成Schwartz的PVQ-RR价值观量表；在第三阶段，被试在没有LLM的情况下处理另一个困境，以测量效应的持续性。","baseline":"有真实人类对照：控制组被试阅读由三个LLM合成的固定AI生成考虑，但不与LLM交互；所有被试在互动前、互动后立即和后续任务中完成平行版本的PVQ-RR量表，作为价值观变化的基准。","findings":"与LLM互动后，被试的价值优先级相对于控制组暂时向个人关注方向偏移（效应量d=0.37-0.51），主要通过自我增强价值的增加实现；这种偏移在后续任务中消失，且未检测到价值方向或建议的趋同。","reliability":"论文承认效应是暂时的，在后续任务中消失；未讨论其他失效条件或局限。","relevance":"该研究直接探讨LLM交互对人类价值观的因果影响，属于批判性仿真研究，揭示了在无明确说服意图下LLM仍能暂时改变价值优先级，对理解LLM在决策支持中的潜在偏差具有重要意义，值得精读原文。","inspiration":"值得借鉴的做法是采用随机对照实验设计，将LLM作为处理条件，设置非交互控制组，并使用标准化的价值观量表在多个时间点测量效应。｜可以迁移到经济金融中的消费者跨期选择或投资决策场景，例如LLM作为财务顾问是否会影响个人的时间偏好或风险态度。｜一个可行的设计是：招募真实投资者作为被试，随机分配到与LLM（如ChatGPT）进行投资讨论的处理组或阅读固定建议的控制组，在互动前后测量时间贴现率和风险偏好（如使用滴定法或量表），并与真实市场数据（如实际投资组合选择）进行对照，以检验LLM交互对经济决策的因果影响。"}},{"id":"2609.23640","version":1,"title":"Are Human-Aligned Models Models of Humans? A Turing-Test Gap in Preference Alignment","zh_title":"人类对齐模型是人类的模型吗？偏好对齐中的图灵测试差距","abstract":"Human-feedback alignment has made language models useful assistants and is commonly described as aligning them with humans. However, the responses people prefer from an AI need not be the responses they themselves would give. We distinguish alignment with human preferences from alignment with human behavior, and show that alignment with human preferences can make model behavior less human-like even when both preferences and responses come entirely from humans. We call this the Turing-test gap. We show that preference alignment preserves the human response distribution only under a restrictive condition, and find no consistent evidence that real human preferences satisfy it. Empirically, the loss of human-response likelihood increases with the strength of preference weighting, regardless of its direction, and the gap also appears under standard DPO. These results establish human-likeness as an explicit dimension of alignment rather than something assumed to follow from preference alignment.","authors":["Suqin Yuan","Runqi Lin","Muyang Li","Guanzhe Hong","Jindong Gu","Lei Feng","Chris Russell","Tongliang Liu"],"categories":["cs.AI","cs.CL","cs.LG"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-09-23","first_seen":"2026-09-23","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.23640","pdf_url":"https://arxiv.org/pdf/2609.23640","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A2","B4"],"tags":["偏好对齐","人类仿真","算法保真度"],"reason":"研究偏好对齐导致模型行为偏离人类，评估仿真可靠性，具批判性。","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:10","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-23","rank":9,"question":"人类偏好对齐是否会使语言模型的行为偏离人类行为分布，从而产生图灵测试差距？","design":"本研究并非直接进行人类仿真实验，而是通过理论分析和受控实验检验偏好对齐对模型人类相似度的影响。在受控实验中，使用同一批人类回答分别进行等权拟合和偏好加权拟合，比较模型对未见过的人类回答的似然；在DPO实验中，从已拟合人类回答的检查点出发进行偏好优化，测量人类回答似然和生成文本的变化。","baseline":"对照的真实人类数据来自SHP和StackExchange回答池，其中包含人类撰写的回答和人类投票偏好；以及HH-RLHF、WebGPT等偏好数据集。","findings":"偏好对齐在理论上仅在严格条件下保持人类行为分布，而实证中未发现人类偏好满足该条件。偏好加权强度越大，模型对人类回答的似然损失越大，且该损失与加权方向无关；标准DPO同样导致模型偏离人类行为。","reliability":"论文指出，提高人类相似度并非普遍可取，可能引发冒充、操纵和社会工程等风险；同时，通过角色提示和偏好人类回答的DPO等方法对人类相似度的恢复有限且依赖领域。","relevance":"该研究直接揭示了用偏好对齐的LLM作为人类被试替代品时可能存在的系统性偏差，为评估仿真可靠性提供了关键理论依据和实证证据，值得精读。","inspiration":"借鉴其将偏好对齐与行为分布分离的框架，在仿真实验中明确区分“人类偏好”与“人类行为”，并测量两者差异。｜可迁移到经济金融中的政策偏好调查或消费者决策仿真，例如用LLM模拟公众对通胀预期的回答，但模型可能给出“理想”而非“实际”的回答。｜设计：以LLM为被试，施加偏好对齐强度不同的处理，结果变量为对经济问题回答的分布与真实调查数据（如密歇根消费者调查）的KL散度，对照真实人类回答。"}},{"id":"2609.25572","version":1,"title":"A Behavioral Trait Leaks into Preferences: Diagnosing Trait Interference in LLM User Simulators","zh_title":"行为特质泄漏到偏好中：诊断LLM用户模拟器中的特质干扰","abstract":"LLM-based user simulators aim to bridge the offline-online gap in recommender evaluation by emulating users through injected traits, where preference attributes determine what a user engages with and a behavioral activity trait governs how long they browse. However, we show this intended trait independence collapses during simulation, causing two failures: (i) Trait Interference, where amplified activity distorts preference boundaries and forces interactions with mismatched items to sustain browsing, and (ii) Evaluation Invalidity, where satisfaction scores inflate with activity-driven page counts despite taste mismatches, biasing evaluation toward trait distributions rather than recommender performance. To resolve this, we propose PQA, a page-level quality anchoring method that guides simulators using a personalized anchor reflecting each user's intrinsic preference standard. By assessing whether a page meets this standard before further browsing, PQA enables proactive exits from low-quality pages, letting the activity trait retain its intended role of modulating browsing depth within preference-conforming pages. Experiments show PQA mitigates trait interference and improves the reliability of LLM-based simulator evaluation under activity shifts. Our code is available at https://github.com/chaehyun1/PQA","authors":["Chaehyun Kim","Sein Kim","Hongseok Kang","Chanyoung Park"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-23","first_seen":"2026-09-23","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.25572","pdf_url":"https://arxiv.org/pdf/2609.25572","source_feed":"cs.AI","score":7,"bucket":"pending","rubric_hits":["A1","B1","B4"],"tags":["LLM用户模拟","推荐系统评估","仿真偏差"],"reason":"用LLM模拟用户行为并诊断仿真失效，有真实数据对照，方法可迁移","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:12","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-24","rank":11,"question":"LLM用户模拟器中，行为活动特质是否干扰了偏好特质，导致模拟失效？","design":"使用Agent4Rec和SimUSER两个LLM用户模拟器，在MovieLens和Amazon CDs数据集上，以SASRec为推荐模型生成20项推荐列表，按4项一页呈现。通过将低活动用户的活动特质修改为高活动状态，保持偏好特质不变，测量模拟器选择项目的语义对齐度、类别偏好一致性、浏览页数和满意度分数。","baseline":"真实用户数据：MovieLens和Amazon CDs数据集中用户活动水平与平均评分的负相关关系，即更活跃用户平均评分更低。","findings":"活动特质放大导致模拟器选择与用户偏好语义距离更远的项目，并偏离固有类别偏好，产生特质干扰；同时满意度分数随活动驱动的浏览页数增加而膨胀，即使推荐质量下降也持续浏览低质量页面，导致评估无效。","reliability":"论文指出当前模拟器缺乏页面级质量锚点，导致活动约束覆盖页面级偏好对齐；PQA方法依赖用户历史类别消费密度，可能受历史数据稀疏性影响，但未详细讨论其他失效条件。","relevance":"该研究直接诊断LLM模拟器中行为特质与偏好特质的干扰问题，并提供了真实人类数据对照，对关注仿真可靠性和偏差的研究者具有重要参考价值，值得阅读原文了解PQA方法细节。","inspiration":"借鉴其通过修改单一特质并保持其他特质不变来分离干扰效应的实验设计，可用于经济金融仿真中识别行为参数对决策的因果影响｜可迁移到消费者金融决策仿真，如信用卡使用或投资组合选择，检验风险偏好特质是否被交易频率等行为特质干扰｜设计：用LLM模拟消费者，将风险偏好设为固定，改变交易频率特质，测量投资组合风险水平和满意度，对照真实交易数据中频率与风险偏好的关系。"}},{"id":"2609.26403","version":1,"title":"AI-Generated Email Drafts Shift Culturally Distinctive Communication Styles in Professional Email","zh_title":"AI生成的邮件草稿改变职场邮件中文化特有的沟通风格","abstract":"AI assistants that support email composition may shift cultural communication norms, such as the directness typical of low-context cultures like the US versus the indirectness and contextual sensitivity central to high-context cultures like Japan. Yet it remains unknown to what extent people adopt and edit AI drafts inconsistent with their cultural communication norms. We address this through a preregistered within-subject experiment in which Japanese and American participants wrote workplace emails in their native language without AI, with a low-context AI, and with a high-context AI. We found that Japanese participants wrote emails with significantly more high-context markers (politeness, apologies) than Americans. But AI drafts shifted participants' emails toward the draft's style, with larger shifts when the draft was culturally misaligned: Japanese drifted most under low-context drafts, Americans most under high-context drafts. These findings suggest AI drafts risk overwriting cultural communication norms unless they adapt to users' communication styles.","authors":["Shintaro Sakai","Alice Gao","Yuichi Shoda","Katharina Reinecke"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-23","first_seen":"2026-09-23","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.26403","pdf_url":"https://arxiv.org/pdf/2609.26403","source_feed":"cs.HC","score":7,"bucket":"pending","rubric_hits":["A1","B1","B4"],"tags":["LLM仿真","文化沟通","人机交互实验"],"reason":"用LLM生成邮件草稿并测量其对人类写作风格的影响，有真实人类实验对照，且揭示A…","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:15","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-23","rank":11,"question":"使用AI生成的邮件草稿是否会将用户的邮件写作风格转向草稿的文化语境风格，且这种转变在AI与用户文化背景不一致时是否更大？","design":"本研究不是用LLM模拟人类被试，而是用LLM生成邮件草稿作为实验干预。175名日本和美国全职员工在三种条件下写工作邮件：无AI辅助、低语境AI草稿、高语境AI草稿（被试内设计，条件顺序平衡）。结果变量是邮件中高语境标记（如礼貌用语、道歉）的数量。","baseline":"无AI辅助条件下参与者写的邮件作为人类基线，用于比较日本和美国参与者的文化沟通风格差异。","findings":"日本参与者在无AI条件下写的邮件比美国人包含更多高语境标记。AI草稿使参与者的邮件风格向草稿的文化语境偏移，且当草稿与参与者文化背景不一致时偏移更大：日本人在低语境草稿下偏移最大，美国人在高语境草稿下偏移最大。","reliability":"论文未讨论","relevance":"该研究用真实人类实验检验了LLM生成内容对人类行为的影响，有严格的人类对照，揭示了AI在文化维度上的潜在偏差和影响，对关注LLM仿真可靠性及偏差的研究者有参考价值。","inspiration":"借鉴其被试内设计和多条件对照，通过测量行为变化来评估AI干预的因果效应。｜可迁移到经济金融中的沟通与决策场景，如AI辅助撰写投资建议、信贷沟通或政策沟通，研究AI风格对用户决策或行为的影响。｜设计一个实验：招募金融从业者或普通投资者作为被试，让他们在无AI、低风险偏好AI草稿、高风险偏好AI草稿三种条件下撰写投资建议邮件，测量建议的风险水平，并与真实历史投资建议数据对照，检验AI是否改变建议风格及在风格不一致时影响是否更大。"}},{"id":"2609.21997","version":2,"title":"Bayesian Belief Layer for Controllable Opinion Dynamics in LLM Agents","zh_title":"用于LLM智能体可控观点动力学的贝叶斯信念层","abstract":"LLM agents in social simulation revise their opinions implicitly, in context: how open an agent is to persuasion can neither be specified nor verified, and collective outcomes inherit the model's training prior. We introduce Bayesian Chronicle Agents (BCA), a minimal belief layer separating what an agent believes from how it speaks. Each stance is a probability, updated by one Bayesian step per utterance heard. A single prior-strength parameter $\\kappa$ encodes stubbornness, modeled after its role in Friedkin--Johnsen (FJ) opinion dynamics. We then sweep this parameter to yield three canonical regimes of opinion dynamics on demand (consensus, persistent disagreement, committed-minority influence), with persistent disagreement matching the FJ closed-form fixed points at $R^2\\!=\\!0.93$--$0.99$. We further show that prescribed $\\kappa$ remains recoverable after the language round-trip, with perfect rank-order recovery across all four models. Explicit belief also makes simulation auditable: the layer surfaces systematic per-model stance biases that end-to-end simulation would silently absorb.","authors":["Hafsa Akbar","Daniel Platnick","Marjan Alirezaie","Hossein Rahnama","Alex 'Sandy' Pentland"],"categories":["cs.MA","cs.AI"],"primary_category":"cs.MA","announce_type":"replace-cross","date":"2026-09-23","first_seen":"2026-09-21","revised_at":"2026-09-23","abs_url":"https://arxiv.org/abs/2609.21997","pdf_url":"https://arxiv.org/pdf/2609.21997","source_feed":"cs.AI","score":6,"bucket":"other","rubric_hits":["D3"],"tags":["观点动力学","社会模拟","LLM智能体"],"reason":"用LLM agent模拟观点动态，但无真实人类数据对照，属社会模拟边界情形。","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:35","error":null,"has_summary":false,"summary":null},{"id":"2609.26539","version":1,"title":"A retrospective analysis on the use of LLMs to study infant syntax learning","zh_title":"对使用大语言模型研究婴儿句法学习的回顾性分析","abstract":"Large language models (LLMs) have increasingly been used to investigate how children acquire syntax at an early stage of development. This is notably the central scientific goal of the BabyLM challenge, a community-wide effort to develop models that achieve human-level syntactic performance while being trained on developmentally realistic corpora. In this paper, we reflect on the use of LLMs in the study of infant syntax learning by providing an epistemological assessment of several studies from this research program. We discuss how datasets are built, which models are implemented, how they are trained and syntactically evaluated. We observe significant assumptions in the methodology of BabyLM and related studies, thus mitigating their theoretical scope. We additionally observe that using developmentally-realistic corpora have limited effects on models performance on commonly-used benchmarks, which suggest important computational differences between LLMs and the infant syntax learner.","authors":["H\\'elie Bazin (SCAI, SND, ISIR)","Anouk Barberousse (SND)","Fran\\c{c}ois Yvon (MLIA)"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-23","first_seen":"2026-09-23","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.26539","pdf_url":"https://arxiv.org/pdf/2609.26539","source_feed":"cs.CL","score":6,"bucket":"other","rubric_hits":["D2"],"tags":["LLM认知建模","婴儿语言习得","方法论评估"],"reason":"用LLM研究婴儿句法学习，属认知建模而非人类被试仿真，但涉及模型与人类对照，边…","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:15","error":null,"has_summary":false,"summary":null},{"id":"2609.26579","version":1,"title":"Receptiveness, Not Sycophancy: Distinguishing Engagement from Deference in Language Models","zh_title":"接纳性而非谄媚：区分语言模型中的参与和顺从","abstract":"A central concern with language models is sycophancy: their tendency to defer to users' views at the expense of independent substantive judgment. In parallel, work on social sycophancy has focused on behaviors such as validation and positivity that may signal inappropriate deference. Yet the markers of social sycophancy are also characteristic of conversational receptiveness, a construct from social psychology shown to improve interactions across disagreement. We argue that this overlap creates a construct-validity problem for social sycophancy evaluations. Using a popular moral-advice dataset, we find that responses classified as more socially sycophantic are also more receptive. Further, increasing the receptiveness of human-written responses---while preserving their substantive conclusions---causes them to be classified as more socially sycophantic. This tight coupling raises the possibility that social sycophancy evaluations inadvertently penalize desirable behavior. In a preregistered experiment comparing substantively equivalent responses, participants prefer the more receptive responses, expect users to be more likely to listen to them, and are more willing to seek advice from their authors. The same overall pattern persists even among participants who believe the original question asker is in the wrong. Finally, we introduce a simple approach that substantially increases receptiveness without increasing substantive deference, demonstrating that conversational receptiveness and substantive independence can be achieved together.","authors":["Calvin Isley","Johann Gaebler","Max Lamparth","Julia Minson","Sharad Goel"],"categories":["cs.CL","cs.AI","cs.HC"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-23","first_seen":"2026-09-23","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.26579","pdf_url":"https://arxiv.org/pdf/2609.26579","source_feed":"cs.CL","score":6,"bucket":"other","rubric_hits":["D2"],"tags":["LLM行为评估","社会心理学","模型对齐"],"reason":"研究LLM的社交行为与人类偏好，但非仿真人类被试，而是测量模型本身特性。","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:31","error":null,"has_summary":false,"summary":null},{"id":"2609.26481","version":1,"title":"Behavior is Not Enough: A Mechanism-Based Evaluation of Social Norm Emergence in LLM Societies","zh_title":"行为不足：基于机制的LLM社会规范涌现评估","abstract":"Social norms cannot be identified from behavior alone: the same cooperative equilibrium may reflect shared expectations, strategic incentives, or simple imitation. Yet in multi-agent large language model systems, prior work largely treats behavioral convergence as evidence of norm emergence. In this work, we introduce an evaluation framework that measures agents' reported empirical and normative expectations in addition to behavioral convergence. Through controlled ablations, we test the effect of expectation elicitation and isolate two collective mechanisms central to theories of norm formation---social learning through interaction and social selection through network-based group formation. We further test the stability of these resulting dynamics under adversarial disruption across four LLM families. We find that eliciting expectations increases cooperative contributions, while social learning stabilizes behavior, and social selection reliably identifies cooperators but provides limited behavioral reinforcement. Following disruption, normative expectations and behavioral coordination recover differently. Together, these results show that similar cooperative outcomes can arise from different underlying social processes. By making expectations observable, our framework allows us to attribute each mechanism's contribution separately, offering designers of multi-agent systems a principled basis for selecting the social processes that sustain cooperation.","authors":["Rasika Muralidharan","Haewoon Kwak","Jisun An"],"categories":["cs.MA","cs.CL","cs.CY","cs.GT","cs.SI"],"primary_category":"cs.MA","announce_type":"cross","date":"2026-09-23","first_seen":"2026-09-23","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.26481","pdf_url":"https://arxiv.org/pdf/2609.26481","source_feed":"cs.CL","score":6,"bucket":"other","rubric_hits":["D3"],"tags":["LLM社会模拟","规范涌现","多智能体"],"reason":"多智能体社会规范涌现模拟，但无真实人类数据对照，属边界情形。","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:15","error":null,"has_summary":false,"summary":null},{"id":"2609.25194","version":1,"title":"Indirect tipping: a social attack surface in AI agent populations","zh_title":"间接引爆：AI智能体群体中的社会攻击面","abstract":"As generative AI agents are deployed at scale, safety will depend not only on technical safeguards and individual model design, but also on collective equilibria that determine how agent populations process information, prioritize actions, and respond to uncertainty. Yet the same equilibria that enable agents to coordinate also create a social attack surface. The standard framework to assess this vulnerability is critical mass dynamics: the minimum fraction of adversarial agents required to overturn an equilibrium through direct competition. Here, we show that this approach risks underestimating system vulnerability by reducing the problem to the identification of singular tipping points, and ignoring indirect but potentially more efficient routes through which collective behavior can be redirected. Through experiments with populations of LLM agents and an analytic framework that captures their collective dynamics at scale, we map critical-mass thresholds that define a directed, weighted topology over the space of coordination equilibria, and treat this topology as a navigable landscape. We show that indirect tipping through intermediate stepping-stone equilibria can reduce the committed minority required to reach an alternative state, bypass majority requirements, and make possible transitions inaccessible through direct challenges. The diversity of available alternatives and timing of the attack further reshape this landscape, creating opportunities for control as well as risks of unintended destabilization. These results show that an equilibrium's resistance to committed intervention is not an intrinsic property but a structural feature of its competitive relations with alternative states. Securing populations of interacting AI agents therefore requires mapping this social landscape alongside individual agent capabilities and the technical channels through which they interact.","authors":["Ariel Flint","Luca Maria Aiello","Sara M. Constantino","Romualdo Pastor-Satorras","Andrea Baronchelli"],"categories":["cs.MA","cs.AI","cs.CY","cs.SY","eess.SY","physics.soc-ph"],"primary_category":"cs.MA","announce_type":"cross","date":"2026-09-23","first_seen":"2026-09-23","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.25194","pdf_url":"https://arxiv.org/pdf/2609.25194","source_feed":"cs.AI","score":6,"bucket":"other","rubric_hits":["D3"],"tags":["LLM多智能体","社会模拟","集体行为"],"reason":"用LLM agent群体模拟社会协调过程，但无真实人类数据对照，属边界情形。","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:11","error":null,"has_summary":false,"summary":null},{"id":"2609.25287","version":1,"title":"Can LLMs identify and repair ruptures? Comparison between clinician practices and LLM behaviors","zh_title":"大语言模型能否识别和修复关系破裂？临床医生实践与LLM行为的比较","abstract":"Ruptures represent common albeit critical moments in interaction where relational alignment breaks down, making them essential for evaluating AI where trust and engagement matter most. In a scenario-driven empirical study, we examined the performance of three LLMs at identifying and resolving ruptures across 21 mental health conversations and 22 experts' evaluation of the strategies. For identification, LLMs relied on explicit linguistic cues within single turns whereas experts integrated implicit, relational, and contextual information across the conversation. For resolution, LLMs tended to produce more directive and scripted responses whereas experts adopted process-oriented strategies such as validation, open-ended exploration, and psychoeducation. Overall, LLMs showed higher agreement with predefined labels in identification, but not in resolution where experts rated their responses only moderately effective, with consistent limitations in timing, depth, and contextual sensitivity. We discuss implications for the design of mental health conversational agents emphasizing relational awareness, pacing, and human-in-the-loop support.","authors":["Jeongah Lee","Joy Qiuyue Zhong","Drishti Goel","Violeta J. Rodriguez","Dong Whi Yoo","Koustuv Saha","Ravi Karkar"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-23","first_seen":"2026-09-23","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.25287","pdf_url":"https://arxiv.org/pdf/2609.25287","source_feed":"cs.HC","score":6,"bucket":"other","rubric_hits":["D2"],"tags":["LLM评估","心理健康对话","人机对比"],"reason":"研究LLM在心理治疗对话中的行为，与人类专家对照，但非仿真人类被试，而是评估模…","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:19","error":null,"has_summary":false,"summary":null},{"id":"2609.25432","version":1,"title":"Tipping Points in LLM-Based Multi-Agent Systems: Stance on Climate Change Action","zh_title":"基于LLM的多智能体系统中的临界点：对气候变化行动的态度","abstract":"Because significant action to counter global warming requires massive public support, it is important to understand the dynamics of public opinion on climate issues. Of special interest are social tipping points, as revealed by large-scale effects of small perturbations in individual behaviors. Agent-based models (ABM) are an effective computational tool for studying these matters, because they allow controlled and systematic exploration of the effects of interventions that may be infeasible in real-world social systems. Large language models (LLMs) have been used to endow model agents with the ability to communicate in natural language (rather than by exchanging predefined messages), as well as with personality (in the form of a narrative self and episodic memory). We leverage LLM-powered ABM to look for tipping points in the social dynamics of a micro-society in which some of the discussions are about climate change. Our agents' stance was defined by two variables: the strength of conviction about the urgency of climate action and the degree of trust in existing institutions. We quantified shifts in agents' \"beliefs\" by monitoring, across multiple rounds of conversations, (1) inter-agent distances in this two-dimensional stance space and (2) the patterns of discussion topics as modeled by Latent Dirichlet Allocation (LDA). Our findings to date suggest that significant abrupt changes in climate-change stance do occur in this simple model. We report a number of methodological lessons from this study, notably, the need to prevent LLM biases from interfering with the conversational dynamics and, more generally, to maintain agent personality and episodic memories of interactions in the face of such biases. Resolving these issues may allow for using ABM-derived insights in designing real-life interventions vis-a-vis climate change and other important societal challenges.","authors":["Astghik Altunyan","Shimon Edelman"],"categories":["cs.MA","cs.CY"],"primary_category":"cs.MA","announce_type":"cross","date":"2026-09-23","first_seen":"2026-09-23","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.25432","pdf_url":"https://arxiv.org/pdf/2609.25432","source_feed":"cs.CY","score":6,"bucket":"other","rubric_hits":["A3","D3"],"tags":["LLM智能体","社会模拟","舆论动态"],"reason":"用LLM智能体模拟社会舆论动态，但无真实人类数据对照，属边界情形。","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:11","error":null,"has_summary":false,"summary":null},{"id":"2607.22006","version":2,"title":"Printed but not benchmarkable: most building-decarbonisation disclosure cannot be matched to the pathways that stranding regulation assumes","zh_title":"已印刷但不可基准化：多数建筑脱碳披露无法匹配搁浅监管假设的路径","abstract":"Cities are beginning to enforce carbon limits on existing buildings. Science-based decarbonisation pathways set those limits one asset type and one jurisdiction at a time. Owners, however, report for the whole firm. We measure what that mismatch costs on two sets of public corporate reports: a census of 502 reports from the 119 listed built-environment firms with a collected report inside a 2,246-firm panel (2003-2023), and 519 real-estate reports from 101 firms (2007-2024). BeDA, a multimodal language-model tool whose reliability we test first, read them. Running the pathway frameworks' own entry tests over published disclosure: 16.5% of census reports (43.7% of real-estate reports) print an operational carbon intensity per square metre; 6.6% (25.0%) can be matched to a pathway for their property type in a covered jurisdiction; and only 5.0% (16.4%) disclose the floor area they divided by. Of the failures at the pathway test, 82-84% follow from reports lumping the portfolio together and 16-18% from a missing curve in the pathway library. The obstacle is the reporting unit, not missing data. The rate is roughly twice as high for European as for US listings (65-71% versus 35% in listed real estate). We also show that a US portfolio's carbon verdict cannot be worked out from disclosure at all. Within one climate zone, the pathway's carbon limit varies by up to 2.79-fold with the electricity subregion, which no report names; its energy limit does not move. Extraction is checked against the source PDFs (97.3% of extracted intensities appear verbatim) and repeats on a second extractor (kappa = 0.97). Recall of the non-disclosing class was 95.1% in a blinded hand audit of 122 reports. The fix follows from the measurement: split intensity by asset type and jurisdiction, and report floor area.","authors":["Jingyi Xu","Minghui Cheng","Anchen Sun"],"categories":["cs.CY","econ.GN","q-fin.EC","stat.AP"],"primary_category":"cs.CY","announce_type":"replace","date":"2026-09-23","first_seen":"2026-07-27","revised_at":"2026-09-23","abs_url":"https://arxiv.org/abs/2607.22006","pdf_url":"https://arxiv.org/pdf/2607.22006","source_feed":"cs.CY","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM标注","气候披露","文本提取"],"reason":"用LLM提取披露信息，替代人工标注，非仿真人类被试","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:17","error":null,"has_summary":false,"summary":null},{"id":"2609.25669","version":1,"title":"From Utterances to Networks: Modelling Slang Adoption and Diffusion Across Subreddits","zh_title":"从话语到网络：建模俚语在Subreddit中的采纳与扩散","abstract":"Adoption and diffusion of neologisms in online communities have received renewed attention in recent years. As internet slang terms such as APT, referring to a K-pop song, and phrases such as Canon Event meaning an embarrassing but pivotal event, go viral online, it becomes increasingly important to understand the mechanisms that contribute to their success. Prior studies have often explained slang diffusion either from the perspective of social interaction or from the linguistic properties of the slang itself, but rarely from both perspectives together. One major obstacle has been the high cost of annotating slang usage in large-scale online communication. Recent advances in large language models (LLMs), however, make it possible to use them as scalable annotators for such tasks. In this study, we first curate a human-annotated benchmark to evaluate LLM performance in detecting slang usage in real Reddit communication. We then leverage LLM-based annotations to model slang adoption and diffusion. Our results show that slang diffusers with higher bridging capital are associated with increased subsequent adoption, whereas diffusers with higher bonding capital are associated with reduced adoption. We also find that wider contextual usage of a slang term is associated with a longer time before new users officially adopt it. Together, these findings suggest that both social-network structure and linguistic context shape the diffusion of neologisms in online communities.","authors":["Xiaoning Wang","Ted Underwood","Zhewei Sun"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-23","first_seen":"2026-09-23","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.25669","pdf_url":"https://arxiv.org/pdf/2609.25669","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM标注","社会网络","语言扩散"],"reason":"用LLM做标注替代人工，非仿真人类被试，但方法可迁移","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:22","error":null,"has_summary":false,"summary":null},{"id":"2609.26527","version":1,"title":"A Semiotics-Aware Framework for Evaluating Fidelity and Coverage in Natural Language Generation","zh_title":"一种符号学感知的框架用于评估自然语言生成中的保真度与覆盖率","abstract":"When two texts describe the same expression, standard metrics based on lexical overlap or whole-text similarity may fail to detect meaningful differences in how that expression is framed. We propose a framework to evaluate semiotic alignment between texts, where a semiotic profile encompasses both the contextual meaning and the discourse references made salient by a text. Our approach yields two scores, Semiotic Fidelity and Semiotic Coverage, estimating how much of one text's profile is supported by the other and how much of the other's profile it recovers. Experiments show that coverage is typically lower than fidelity, and that alignment between LLMs and human-curated data is highest at low sampling temperatures, while higher temperatures reduce this alignment.","authors":["Lorenzo Zangari","Davide Picca"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-23","first_seen":"2026-09-23","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.26527","pdf_url":"https://arxiv.org/pdf/2609.26527","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["语义对齐","评估框架","人类数据对照"],"reason":"评估LLM与人类文本的语义对齐，非仿真人类被试，但涉及人类数据对照，属边界情形。","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:31","error":null,"has_summary":false,"summary":null},{"id":"2609.26204","version":1,"title":"WatchPoint: Executable User Feedback for Real-World Agentic Web Development","zh_title":"WatchPoint：面向真实世界智能体网页开发的可执行用户反馈","abstract":"When a professional web developer's code fails a test, they do not simply re-read the stack trace. They open the application in a browser, click buttons, inspect computed styles, and run diagnostic commands to understand what went wrong. Existing feedback mechanisms for coding agents rely on screenshots, LLM-as-a-judge scoring, or natural-language corrections, but few interact with the live application the way a developer would. We introduce WatchPoint, a simulated-user system that mimics real developer behavior by generating and executing diagnostic scripts against the running application, producing structured observations that guide the coding model's retry. Unlike prior approaches that target single-file edits or evaluate using non-executable metrics, we operate on Web-Bench, a benchmark of 50 multi-file web projects comprising 1,000 sequentially dependent tasks, verified by deterministic end-to-end tests. WatchPoint recovers 57.6% of the tasks it diagnoses, and a controlled user study confirms the simulation's realism: human testers achieve a comparable recovery rate (54.5%), providing evidence that automated diagnostic scripts can substitute for interactive human testing on sequential web development tasks. We further identify a pattern of capability gaps that governs when simulated-user feedback is helpful and when it should be withheld.","authors":["Guanqun Yang","Wei Yang","Xueqing Liu"],"categories":["cs.SE","cs.AI","cs.CL"],"primary_category":"cs.SE","announce_type":"cross","date":"2026-09-23","first_seen":"2026-09-23","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.26204","pdf_url":"https://arxiv.org/pdf/2609.26204","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["智能体反馈","模拟用户","软件工程"],"reason":"用模拟用户替代人工测试，属于替代人类劳动而非仿真被试，但涉及人类对照，边界相关。","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:26","error":null,"has_summary":false,"summary":null},{"id":"2609.26069","version":1,"title":"RankCert: When Can Simulated Learners Safely Select an AI Tutor? Robust Decision Certification Under Structural Uncertainty","zh_title":"RankCert：模拟学习者何时能安全选择AI导师？结构不确定性下的鲁棒决策认证","abstract":"Simulation-based tutor selection can be unstable when predictively adequate learner models imply different policy rankings. RankCert certifies one of eight equal-budget tutoring policies only when model-averaged utility, probability-best, posterior regret, cross-domain rank, family coverage, and leave-one-domain-out and leave-one-visible-family-out averages support the same candidate; otherwise it abstains. We evaluated RankCert in 1,280 frozen held-out settings spanning five rotating held-out oracle families, 64 scenarios per family, and four cohort sizes. Calibration used a licensed, de-identified EdNet-KT1 derivative with 5,000 learners and 590,056 retained responses; all five family representatives passed the frozen adequacy gate. Minimum-domain mean pairwise top-1 agreement was 0.272917 (95% CI [0.253646, 0.293229]), showing substantial structural disagreement. Cohort-noise variance decreased from n = 30 to n = 300, while the structural family share remained nonzero. RankCert reduced total held-out decision loss relative to full-coverage point selection by 0.006605 normalized-outcome units (95% CI [0.004859, 0.008407]). At comparable coverage, however, it did not reduce selective risk relative to a confidence-gated point certificate (difference -0.000213; 95% CI [-0.003238, 0.002384]; Holm p = 0.929654). Certification occurred in 3.75% of settings and only in stable scenarios; RankCert abstained in every ambiguous, misspecified, and structural-conflict setting. \"Safe\" denotes only benchmark-scoped decision certification under the declared utility and uncertainty set; no human-learning, causal, deployment-effectiveness, or general-safety claim is made.","authors":["Nizam Kadir"],"categories":["cs.AI","cs.CY","cs.LG"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-23","first_seen":"2026-09-23","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.26069","pdf_url":"https://arxiv.org/pdf/2609.26069","source_feed":"cs.AI","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["模拟学习者","AI导师选择","决策认证"],"reason":"用模拟学习者选择AI导师，属社会模拟但无真实人类行为对照，且非LLM仿真人类被试","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:25","error":null,"has_summary":false,"summary":null},{"id":"2609.25790","version":1,"title":"Automating Constructive Assessment with Large Language Models: Toward Scalable and Repeated Evaluation of Practical Competence","zh_title":"用大语言模型自动化建构性评估：迈向可扩展和可重复的实践能力评价","abstract":"This study aimed to automate hierarchical diagnostic reasoning (HDR), a constructive method for evaluating practical judgment skills, by developing and testing an evaluation process using a large language model. HDR is a descriptive task that measures higher-order cognitive skills by requiring students to identify and explain errors in case-study-based problems. However, it requires expertise and effort to develop and evaluate. Hence, we proposed and empirically validated the automatic (1) generation of case problems containing errors aligned with educational intentions, (2) scoring of descriptive answers, and (3) generation of structured feedback based on incorrect answers, achieved solely through prompt design without fine-tuning. The internal consistency and construct validity of the generated problems were supported by the experimental score distribution and Cronbach's alpha (0.78). The agreement between automated and human ratings reached 100% under some conditions. The feedback was rated as being as convincing and useful as that from human instructors, demonstrating a practical framework for implementing HDR-based constructive assessment with reproducibility, immediacy, and low cost. The flexibility of large language models will also enable repeated and longitudinal assessments while maintaining structure, showing broad potential for application in educational settings.","authors":["Satoshi Takahashi","Atsushi Yoshikawa","Megumi Kose","Kenichi Suzuki","Chieko Inoue","Yumi Watanabe","Mari Sawada"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2026-09-23","first_seen":"2026-09-23","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.25790","pdf_url":"https://arxiv.org/pdf/2609.25790","source_feed":"cs.CY","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM自动评分","教育评估","建构性评价"],"reason":"LLM替代人工评分与反馈，属标注替代而非仿真人类被试，但涉及教育评估，需人工判…","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:23","error":null,"has_summary":false,"summary":null},{"id":"2609.26707","version":1,"title":"Optimal Sequential Annotations for Off-Policy Evaluation","zh_title":"离线策略评估的最优顺序标注","abstract":"Offline reinforcement learning and off-policy evaluation evaluates dynamic treatment rules based on retrospectively collected data prior to deployment. In recent AI applications, state and reward information is recorded as complex text or image, which recent AI advancements such as LLM-as-a-judge can label with unknown bias. Expert annotation may be available but at a higher cost. For example, safety classification via cheap but imperfect classifiers vs. expensive expert review. We show how a limited budget for ground-truth data-annotation can be used via doubly-robust OPE with missing rewards, and we optimize variance-optimal annotation probabilities for sequential off-policy evaluation, where the target policy value is estimated from annotated data. We characterize the optimal annotation probabilities for sequential forward-monotone annotation protocols, and provide a feasible batch-adaptive implementation. Our work is motivated by a collaboration with a homelessness services nonprofit that writes casenotes for individuals over time. Our method can be used to unlock trustworthy inference from casenote data and answer new inferential questions such as: how does expanding outreach effort over time affect progress towards a housing application and improvement in housing placement? In simulations and on two real datasets - casenotes from the nonprofit and human-preference votes from LMArena - we see reductions in RMSE of 34-65% for housing placement and 17-68% for progress towards a housing application at budgets of 40% of full annotation and above, and by 55-62% at every budget on LMArena.","authors":["Woojin Chae","Ezinne Nwankwo","Haitong Qin","Angela Zhou"],"categories":["stat.ME","cs.LG","stat.ML"],"primary_category":"stat.ME","announce_type":"cross","date":"2026-09-23","first_seen":"2026-09-23","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.26707","pdf_url":"https://arxiv.org/pdf/2609.26707","source_feed":"cs.LG","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["离线策略评估","LLM标注","主动学习"],"reason":"用LLM标注替代人工标注，属标注员替代而非仿真被试，但涉及人类数据对照与政策评…","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:34","error":null,"has_summary":false,"summary":null},{"id":"2609.25284","version":1,"title":"When LLM Agents Fail to Read the Room: ReAdapt for Relational Social Reasoning","zh_title":"当LLM智能体读不懂氛围：面向关系社会推理的ReAdapt","abstract":"A social agent's most basic decisions (should I react to this post? who should I reach out to?) are not purely content problems. The right action often hinges on the latent relationship between people -- tie strength, reciprocity, mutual connections -- rather than on which content is most salient. Standard LLM agent loops do not explicitly represent how new relational evidence should revise the agent's current social hypothesis, leaving them prone to surface-obvious choices when relational and content cues diverge. We formalize this failure mode with a relationship-reasoning benchmark: 500 synthetic social worlds with friendships, follows, reaction histories, and feeds, yielding 1,000 queries over two tasks, reaction selection and warm introduction (finding the best bridge to a target person). By construction, the surface-obvious candidate differs from the relationship-grounded oracle in about 53% of queries, forming an overturn subset where the agent must use relational evidence to revise an initially plausible choice. We propose ReAdapt (Relationship-Adaptive Agent with Policy-driven sTate), which augments the ReAct loop with an explicit structured social state z = (G, B, R, N, D) capturing goal, belief, relationship, norm, and disclosure. After each tool observation, ReAdapt runs a typed Adapt step that updates this state and emits a policy operation (continue, switch, abandon, or clarify) before choosing the next action. With Gemini-3-Flash on a stratified subset of n = 150 queries per task, ReAdapt improves warm-introduction accuracy from 37% to 51% (+14 points) and reaction-selection accuracy from 69% to 77% (+8 points). Oracle regret drops from 0.260 to 0.152 and from 0.095 to 0.053, respectively. Holding the model, tools, and environments fixed, these results suggest that explicit relational-state adaptation helps LLM agents turn retrieved social evidence into revised decisions.","authors":["Jianzhe Lin","Xiaolin Li","Yunda Liu","Fei Wang","Jubin Chheda"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-23","first_seen":"2026-09-23","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.25284","pdf_url":"https://arxiv.org/pdf/2609.25284","source_feed":"cs.AI","score":3,"bucket":"other","rubric_hits":["C1"],"tags":["LLM智能体","社会推理","关系适应"],"reason":"多智能体社会推理，无人类数据对照，非仿真人类被试","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:19","error":null,"has_summary":false,"summary":null},{"id":"2608.27855","version":2,"title":"AI Writers Have a Consistent Stylometric Footprint, but AI Editors Do Not","zh_title":"AI写作有稳定的风格指纹，但AI编辑没有","abstract":"Text generated by large language models (LLMs) has been shown to be stylometrically distinct from human-written text (Andre et al., 2023; Shah et al., 2023; Opara, 2024; Soto et al., 2024; Li and Zhang, 2025; Selvioglu et al., 2025). But LLMs are increasingly used not only to generate text but also to edit human writing, and it is unclear whether the two leave the same trace. We show that AI generation leaves a consistent \"stylometric footprint\": a small subset of features, primarily entropy and lexical diversity, consistently separates AI-generated text from human writing across 8 LLMs and 5 domains, while the remaining features depend heavily on the domain and generator. AI editing, however, does not reproduce the same footprint. Relative to their human- written sources, AI-edited texts show only a small increase in lexical diversity and a decrease in entropy, rather than the joint increase that characterizes AI generation. Lexical density, which contributes little to generation, instead becomes the dominant editing-associated signal. Stylometric features therefore separate AI-edited text from AI-generated text but are substantially less effective at separating it from human-written text. Our results suggest that \"AI text\" is not a single phenomenon: generation and editing leave qualitatively different stylometric traces and should be studied separately.","authors":["Zhengyang Shan","Yukyung Lee","Sophie Hao"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-23","first_seen":"2026-08-31","revised_at":"2026-09-23","abs_url":"https://arxiv.org/abs/2608.27855","pdf_url":"https://arxiv.org/pdf/2608.27855","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["AI文本检测","风格计量","文本生成"],"reason":"研究AI文本风格特征，非仿真人类被试，无行为对照","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:17","error":null,"has_summary":false,"summary":null},{"id":"2609.08016","version":2,"title":"What Does Multi-Agent LLM Debate Actually Change? A Layered Analysis of Disagreement and Answer Quality","zh_title":"多智能体大语言模型辩论究竟改变了什么？分歧与答案质量的分层分析","abstract":"Multi-agent debate, in which several LLMs exchange arguments before producing an answer, is widely assumed to improve answer quality by surfacing genuine disagreement. That disagreement is hard to verify, and no single signal can settle it, so we organize the analysis around four questions: (A) does the debater say it disagrees; (B) does its reply text actually argue; (C) does the dissent survive once the tone instruction that produced it is removed; and (D) do the probabilities assigned to stance options change? We evaluate three-model committees on 50 open-ended GlobalOpinionQA questions under three debate tones, friendly (seek common ground), neutral, and hostile (stress-test every position). The answers differ. (A) Full agreement differs by 50.4 percentage points between the friendly and hostile endpoints, pooling replies across all three rounds. (B) The text argues too: an LLM evaluator reading the contribution and reply without the structured self-report or tone instruction confirms the pushback is real. (C) Removing the instruction produces more returns to agreement in our sample, but the primary question-level test is inconclusive. (D) In a separate open-weight study, adjusted movement toward the opposing side is detected in two of three models, relative to a filler-based reference. This measures contextual response probabilities, not lasting belief change. For final answers, debate brings no measurable quality gain: an evaluator that judges each pair in both answer orders scores the debated answer no better than the same committee's no-debate answer on all 299 pairs, a test that detects severe but misses moderate damage in our checks. Taken together, LLM debate readily changes what agents say, but we find much weaker evidence that it changes what they persistently endorse or improves the quality of the final answer.","authors":["Chen Qian"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"replace","date":"2026-09-23","first_seen":"2026-09-09","revised_at":"2026-09-23","abs_url":"https://arxiv.org/abs/2609.08016","pdf_url":"https://arxiv.org/pdf/2609.08016","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体辩论","答案质量","分歧分析"],"reason":"研究多智能体辩论对答案质量的影响，不涉及人类行为仿真或对照，属纯多智能体协作。","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:36","error":null,"has_summary":false,"summary":null},{"id":"2609.26035","version":1,"title":"Truth for Believable AI: Expressed Doubt, Provenance, and Belief Revision as an Engineerable Stance","zh_title":"可信AI的真相：表达怀疑、来源和信念修正作为可工程化立场","abstract":"Conversational agents often express answers in a uniformly confident register. We test whether expressed uncertainty, provenance-aware assertion, and explicit belief revision can be implemented as a behavior layer over a fixed language model; we do not test believability or trust. The layer combines three epistemic states, per-claim confidence and typed provenance, a provenance-gated expression rule, and a persistent revision store with auditable acknowledgments and partial resistance to false corrections. We evaluate it on a constructed, mechanically scored multi-session benchmark using a synthetic model and Qwen2.5-0.5B-Instruct. The synthetic instrument passes all five checks. On the real model, acknowledgment soundness, a by-construction guarantee, holds in 100% of cases, and true corrections are accepted more often than false ones (0.44 vs. 0.15 on held beliefs; 0.875 vs. 0.420 including rule-accepted corrections of unheld facts), but the pre-specified expression-fidelity, contradiction-separation, and provenance margins fail. A disclosed post hoc analysis shows that expression gated on mean answer-token probability ranks correctness below chance end to end (AUC 0.41, conversation-clustered), whereas gating on sampling consistency discriminates (AUC 0.66). A consistency-gated configuration selected from this finding and evaluated under a separately committed protocol meets the conversation-level manipulation and capability-equivalence criteria and replicates on a redrawn conversation set. The manipulation result is selection-dependent, and both criteria remain unresolved when uncertainty is clustered over the 60 facts. The supported conclusions are limited to the by-construction audit guarantee, store-dependent partial correction discrimination, and a benchmark- and model-specific failure of token-probability gating; scaling the fact base is required before human evaluation.","authors":["Sebastian Cochinescu"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-23","first_seen":"2026-09-23","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.26035","pdf_url":"https://arxiv.org/pdf/2609.26035","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["对话代理","可信度","信念修正"],"reason":"研究对话代理的可信度表达，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:25","error":null,"has_summary":false,"summary":null},{"id":"2609.26090","version":1,"title":"SpecialEduBench: Benchmarking Vision-Language Models on Knowledge, Skill, and Attitude in Language Intervention for Autistic Children","zh_title":"SpecialEduBench：面向自闭症儿童语言干预的视觉语言模型知识、技能与态度基准","abstract":"Language is the target of most early intervention for autistic children. Because the goal and the method change from child to child, the work falls to a teacher who takes one child at a time and judges each scene as it unfolds. Artificial intelligence is now being brought to that work, yet the benchmarks that reach special education ask what a model knows rather than what it does in front of a child. Building one is not straightforward, since whether a response is good teaching depends on what the child has just done, so no answer key applies. The evidence that settles it is visual as much as verbal, since the length of a wait, a shift of gaze, and the child's uptake leave no trace in a transcript. We introduce \\emph{SpecialEduBench}, which measures pedagogical competence along knowledge, skill, and attitude, with 4,537 knowledge items and with 200 skill items and 68 attitude items built on recorded intervention, the attitude items crossing pressure with monitoring into 192 response cells. Seven special-education experts wrote, scored, and reviewed the items, and we revised the judge model's instruction against the reference scores they set. Across eight frontier vision-language models no axis is saturated, since the strongest still fails about a tenth of the honesty cells. The models converge where the knowledge is factual and separate where the task is situated, and the failures gather where pressure is applied. We intend the benchmark as an audit to run before deployment and as a starting point for models built for this domain.","authors":["Jihoi Na","Taeyeong Kim","Sungjune Kong","Jaemin Jung","Min Joung Park","Kyungtae Joo","Ahhyun Kim","Shim Jaechang","Sooyoung Joo","Dongjin Ka","SeJoong Kim","Jimin Kim","HyunJin Jung","Unggi Lee"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-23","first_seen":"2026-09-23","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.26090","pdf_url":"https://arxiv.org/pdf/2609.26090","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["基准评测","视觉语言模型","特殊教育"],"reason":"纯模型能力评测，不以人类行为仿真为目标","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:26","error":null,"has_summary":false,"summary":null},{"id":"2609.26210","version":1,"title":"Same Chart, Different Story: Bias in Vision-Language Chart Interpretation","zh_title":"同一图表，不同叙事：视觉语言模型图表解释中的偏见","abstract":"Vision-language models (VLMs) are increasingly used to interpret charts and generate natural-language explanations for socially consequential data. However, they may produce different narratives for the same chart when only the referenced social group changes, reinforcing stereotypes and misleading decisions. Despite these risks, no benchmark exists for systematically evaluating bias in chart interpretation across social dimensions. We introduce ChartBias, the first benchmark for auditing bias in VLM-based chart interpretation. ChartBias contains 820 manually curated real-world charts spanning six attributes: race, income, age, religion, immigration status, and gender, yielding 4,319 valid chart, attribute instances and 8,638 paired generations where the chart is fixed and only the group term is swapped. Across 12 proprietary and open-source VLMs, totaling 155,484 model responses, we find three widespread failure modes: narrative shift (same chart, different narratives), group hallucination (assigning a chart to a group without evidence), and preference polarity (favourable trends often linked to one group). We further propose a multi-agent mitigation framework that serves as a strong baseline by separating chart-grounded evidence extraction from group-conditioned generation and using a counterfactual judge to verify that group-driven differences are supported by the chart. The framework substantially reduces narrative shift while preserving chart-grounded reasoning. Our findings show that evaluating chart understanding requires measuring not only accuracy, but also fairness and consistency across social groups. We release ChartBias at https://github.com/vis-nlp/ChartBiasBench.","authors":["Mizanur Rahman","Huan Wu","Arash Asgari","Enamul Hoque Prince","Laleh Seyyed-Kalantari"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-23","first_seen":"2026-09-23","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.26210","pdf_url":"https://arxiv.org/pdf/2609.26210","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["VLM偏见","图表理解","公平性评测"],"reason":"评估VLM图表解释中的偏见，属于模型公平性评测，不涉及用LLM仿真人类被试或与…","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:27","error":null,"has_summary":false,"summary":null},{"id":"2609.26399","version":1,"title":"Combining Hierarchical Cognitive Process with Process Supervision for Interpretable Scene Safety Understanding","zh_title":"结合层次化认知过程与过程监督的可解释场景安全理解","abstract":"Scene safety understanding plays a life-or-death role in situational awareness in various critical domains. Traditional methods that rely on learning direct mappings between scenes and safety levels often lack interpretability, limiting their reliability in critical applications. An effective approach to overcoming this challenge lies in interpreting human cognitive processes and equipping machine models with analogous cognitive capabilities. This work explores an effective way of integrating scene safety cognitive process modeling and process supervision. Specifically, we first construct a hierarchical cognitive safety structure, which motivates the development of a novel, high-quality scene safety understanding dataset based on multi-step reasoning with process labels. This dataset serves both as a benchmark and a resource to improve the safety reasoning capabilities of Large Language Models (LLMs), while also enabling a granular analysis of intermediate reasoning steps through information flow and saliency-based techniques. Building upon this foundation, we introduce a modular and flexible process supervision framework that reflects the hierarchical nature of human cognition. This framework leverages LLMs as the core architecture and incorporates Low-Rank Adaptation(LoRA) and Mixture-of-Experts (MoE) strategies to enable specialization and collaboration among expert modules, each tasked with specific sub-processes of the overall reasoning chain. Systematic experimental evaluations and analyses confirm that our framework exhibits superior interpretability and performance characteristics compared to traditional approaches.","authors":["Zhiyun Jiang","Hanyong Wang","Binbin Liang","Yu Xie","Zhengjie Wang","Menglong Yang","Wei Li"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-23","first_seen":"2026-09-23","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.26399","pdf_url":"https://arxiv.org/pdf/2609.26399","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C2"],"tags":["场景安全理解","可解释AI","过程监督"],"reason":"论文聚焦场景安全理解，用LLM做推理，但无人类被试仿真或行为对照，属自动驾驶/…","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:29","error":null,"has_summary":false,"summary":null},{"id":"2609.26489","version":1,"title":"Calibration as a First-Class Criterion in LLM Evaluation","zh_title":"校准作为LLM评估中的首要标准","abstract":"Calibration of language models -- the alignment between expressed or implicit confidence and empirical correctness -- is a well-studied subfield within NLP. Methods to measure it already exist. The problem is adoption: outside this subfield, NLP research regularly introduces new models, datasets, and benchmarks without checking whether the model's confidence scores are meaningful. We argue that this adoption gap is a major obstacle to trustworthy LLM evaluation. Miscalibration causes problems in two distinct areas: at deployment, where overconfident mistakes cause real harm, and inside the research pipeline, where methods like LLM-as-a-judge, synthetic data generation, and active learning rely on calibrated confidence without verifying it. Standard calibration metrics only require two inputs per example: a confidence score and a correctness judgment. Most benchmarks in use today already provide both, meaning calibration can be reported immediately. For open-ended generation, however, defining these two inputs is still an open challenge. We argue that each NLP subfield should pair its main performance metric with a calibration score and call for treating calibration as an essential property of every model rather than a niche topic.","authors":["Mario Sanz-Guerrero","Katharina von der Wense"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-23","first_seen":"2026-09-23","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.26489","pdf_url":"https://arxiv.org/pdf/2609.26489","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["LLM校准","模型评估","可信度"],"reason":"论文讨论LLM校准评估，属NLP评测，不以人类行为为参照，不涉及人类仿真。","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:29","error":null,"has_summary":false,"summary":null},{"id":"2609.26629","version":1,"title":"PERSONAWEAVER: Controllable Diversity Beyond Conventional Archetypes in Procedural Character Generation","zh_title":"PersonaWeaver：程序化角色生成中超越传统原型的可控多样性","abstract":"Procedural character generation aims to populate games, simulations, and other virtual worlds with diverse characters. Large language models (LLMs) offer a promising foundation for scaling this task. However, LLM-based procedural character generation remains at an early stage: existing methods either generate characters directly or adapt profiles retrieved from persona banks. As we show, both approaches produce behaviorally homogeneous populations: characters overwhelmingly agree with positive moral norms and respond to questions with helpful, assistant-like reactions. To mitigate this homogenization, we introduce PersonaWeaver, which disentangles world building from behavioral specification and models behavior through setting general, diverse, manually curated banks of moral positions and conversational reactions. This design allows us to test how far LLM(s) can be pushed beyond their default behavioral patterns across settings. Across ten realistic and fantastical settings and three LLM(s), PersonaWeaver produces broader moral and interactional response distributions than prior work. Its guidance also diversifies interpersonal language, response length, and sentiment. It also produces less archetypal combinations of world attributes. Code is available at https://github.com/mqraitem/PersonaWeaver.","authors":["Maan Qraitem","Kate Saenko","Bryan A. Plummer"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-23","first_seen":"2026-09-23","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.26629","pdf_url":"https://arxiv.org/pdf/2609.26629","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["角色生成","多样性控制","游戏NPC"],"reason":"生成游戏角色，无实验或测量目的，不涉及人类行为对照","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:32","error":null,"has_summary":false,"summary":null},{"id":"2609.26687","version":1,"title":"Detecting GPT-Assisted Writing Using Interpretable Stylometric Features","zh_title":"使用可解释的文体特征检测GPT辅助写作","abstract":"Distinguishing GPT-assisted from independently authored student writing has become a critical challenge in academia. This paper evaluates the discriminative capability of interpretable stylometric features extracted solely from submitted text. Using data from 90 participants who wrote both independently and with ChatGPT assistance, we evaluate eight machine learning classifiers while keeping data from the same participant together during validation. On the held-out test set, Random Forest achieved an ROC-AUC of 0.87 and an F1-score of 0.84, with False Positive and False Negative rates of 22.2% and 11.1%, respectively. SHAP analysis shows that lexical and grammatical characteristics drive the resulting predictions. The findings suggest that transparent, text-intrinsic features provide measurable signal for detecting GPT-assisted writing.","authors":["Rajesh Kumar","Nabeel Siddiqui","Alexander Fuchsberger"],"categories":["cs.CL","cs.CY"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-23","first_seen":"2026-09-23","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.26687","pdf_url":"https://arxiv.org/pdf/2609.26687","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["AI写作检测","文体特征","学术诚信"],"reason":"检测AI辅助写作，属NLP能力评测，不以人类行为仿真为目标","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:33","error":null,"has_summary":false,"summary":null},{"id":"2609.25408","version":1,"title":"From Offline Proxies to Online Decisions: A Layered Engagement Evaluation Framework for Conversational AI","zh_title":"从离线代理到在线决策：面向对话式AI的分层参与度评估框架","abstract":"Online A/B experiments are the decision standard for user engagement, but traffic and readout time limit how many conversational-AI changes can be tested. We ask whether an offline signal designed to be computable without treatment-arm user exposure agrees with the outcomes of those experiments. We contribute a reusable construction and diagnosis checklist that treats an offline proxy as a chain of three alignments: behavioral label to product outcome, learned classifier to candidate-assistant behavior, and aggregated offline signal to experiment effect. A companion evaluation protocol audits the whole composite by interval-aware decision agreement, which compares offline and online confidence intervals instead of point estimates, and by within-experiment ranking. The instantiation we evaluate comprises a fixed evaluation suite on which candidate behavior is scored, an engagement classifier trained to predict session/prompt level engagements, and a calibration layer mapping sample-level score differences to online model-level engagement deltas. We then report the audit: 489 paired offline-online contrasts (one candidate arm against its control) from 27 experiments on a deployed multi-turn assistant, spanning model checkpoints to system-prompt tuning. Our primary test uses the 113 contrasts from eight experiments that ran after the map was frozen: on these the composite reaches 81.1% F1, against 34.3% for the raw classifier score it is built on, and makes no wrong-direction calls where that raw score makes 31. Every offline prediction was computed before its experiment ran to prevent overfitting. The evidence supports using the composite to prioritize candidates before scarce experiment traffic is allocated---in our deployment of the experiment, selecting among training checkpoints and tuning system prompts.","authors":["Xuanyi Li","Vaskar Nath","Hossein Amirkhani","Jay Li","Alex Deng"],"categories":["cs.AI","cs.IR","cs.LG"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-23","first_seen":"2026-09-23","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.25408","pdf_url":"https://arxiv.org/pdf/2609.25408","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["对话式AI","离线评估","A/B实验"],"reason":"研究离线代理与在线实验的一致性，属系统评估，非用LLM仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-09-24T13:01:34","error":null,"has_summary":false,"summary":null},{"id":"2609.26057","version":1,"title":"Observing the Conduct of Systematic Reviews with Generative AI Support: An Experience Report from a Graduate Software Engineering Course","zh_title":"观察生成式AI支持下的系统综述实施：一项研究生软件工程课程的经验报告","abstract":"Context: Secondary studies are fundamental practices in Evidence- Based Software Engineering, but teaching them requires activities that expose students to authentic methodological decisions. Objective: This paper reports an experience in a graduate course in which ten doctoral students in Software Engineering, organized into three groups, piloted secondary studies with and without support from generative AI. Method: A single-day classroom session was organized and observed, in which the groups conducted pilot systematic reviews with and without generative AI support. Classroom observations, produced artifacts, and interaction threads with assistants configured in ChatGPT were analyzed to reconstruct how each group appropriated the technology throughout the activity. Results: LLMs reduced initial barriers, accelerated the generation of alternatives, and made methodological problems more explicit, but they also favored excessive delegation, superficial validation, operational difficulties, and a shift in focus from conducting the SLR to using the tool. Conclusion: The experience offers a situated, observational account of how doctoral students engaged with generative AI during a systematic review activity, and the resulting insights also inform the design of a subsequent controlled study. The findings indicate that generative AI can support practical learning about SLRs, provided that its use is accompanied by human supervision, decision records, and critical reflection on its limitations.","authors":["Danilo Monteiro Ribeiro","Gilberto Sussumu Hida"],"categories":["cs.CY","cs.AI","cs.SE"],"primary_category":"cs.CY","announce_type":"cross","date":"2026-09-23","first_seen":"2026-09-23","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.26057","pdf_url":"https://arxiv.org/pdf/2609.26057","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["生成式AI","系统综述","教学经验"],"reason":"研究LLM辅助系统综述教学，非仿真人类被试，无人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-09-24T13:01:34","error":null,"has_summary":false,"summary":null},{"id":"2609.26562","version":1,"title":"The Disciplinary Language Transfer Problem: How Psychological Vocabulary Produces Governance Failures in AI Agent Deployment","zh_title":"学科语言迁移问题：心理学术语如何导致AI代理部署中的治理失败","abstract":"The vocabulary used to describe AI agents in governance contexts -- learning, memory, values, compliance, identity, trust -- is borrowed from psychological and organizational science, contributing to systematic failures in how organizations deploy, oversee, and hold agents accountable. This paper argues that the problem is not merely terminological but epistemological: psychological vocabulary carries an \"invisible grammar\" of its home discipline into governance discourse, calibrating frameworks to a metaphysical entity that does not exist in current AI architectures. We call this the disciplinary language transfer problem. Drawing on Wittgenstein's concept of language games, Kuhn's paradigm-laden observation, Haraway's situated knowledge, and Star and Griesemer's boundary object theory, we show that the transfer operates at three levels (epistemological assumptions, theoretical constructs, and surface vocabulary), each requiring a different remediation. We characterize six foundational epistemological assumptions embedded in Western psychological governance discourse, trace their origin in specific philosophical traditions, and show why each fails when applied to systems without developmental continuity. The paper's practical output is an actionable Disciplinary Audit: a six-question governance document scan operationalized through a translation taxonomy of thirty-seven terms mapping operational constructs to agent-appropriate replacements, presented here in abridged form and openly archived in full. The vocabulary reform proposed here is not merely terminological; it is the condition of possibility for governance frameworks that correctly identify what they are governing.","authors":["Kymberly Lasser-Chere","Tyler Akidau","Marc Millstone"],"categories":["cs.CY","cs.AI"],"primary_category":"cs.CY","announce_type":"cross","date":"2026-09-23","first_seen":"2026-09-23","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.26562","pdf_url":"https://arxiv.org/pdf/2609.26562","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["AI治理","语言哲学","概念分析"],"reason":"讨论AI治理中的语言问题，不涉及用LLM仿真人类被试或与人类数据对照。","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:31","error":null,"has_summary":false,"summary":null},{"id":"2609.25311","version":1,"title":"\"I Talked an AI Chatbot, So What's Next?\" How U.S. Young Adults Imagine Responsible AI for Emotion Coping","zh_title":"“我和AI聊天机器人聊过了，接下来呢？”美国年轻人如何想象负责任的情绪应对AI","abstract":"Emotion coping is inherently relational, unfolding through interactions with friends, family, professionals, and communities. Yet AI chatbots are largely designed around a user--AI dyad. Learning from the ethics of care, we examine how AI chatbots shape the relational conditions of emotion coping. We conducted a scenario-based study with 17 U.S. young adults across four emotion coping scenarios. Participants identified eight roles through which AI could support relational conditions, alongside three challenges: flattening distinct relational conditions, discouraging reciprocity, and shifting relational labor onto users. This study contributed a relational perspective of responsible AI in emotion coping. We argue that responsible AI should respond to individuals' situated relational conditions rather than provide general-purpose support. We further identify two design principles: fostering reciprocity by supplying materials to engage with others, and strengthening emotional self-efficacy. Together, we position responsible AI as AI in the loop of human relationships.","authors":["Jiaying Liu","Nimra Ishfaq"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-23","first_seen":"2026-09-23","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.25311","pdf_url":"https://arxiv.org/pdf/2609.25311","source_feed":"cs.HC","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["人机交互","情绪支持","负责任AI"],"reason":"研究AI聊天机器人用于情绪应对，属于角色扮演对话，无实验或测量目的，不涉及LL…","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:20","error":null,"has_summary":false,"summary":null},{"id":"2609.25700","version":1,"title":"Harnessing LLMs Without Surrendering Control: Delegation Boundaries in Visual Data Storytelling Authoring","zh_title":"在不放弃控制权的前提下利用大语言模型：视觉数据叙事创作中的委托边界","abstract":"Despite the emergence of large language models (LLMs) for visual data storytelling workflows, there are open questions about how authors decide what activities or tasks to entrust to them and what should be \"protected\" or maintained under human control. To investigate this, we interviewed a cohort of 12 expert visual data storytellers. Our analysis shows that participants rarely treated LLMs as autonomous storytellers. Instead, they tend to selectively delegate execution-oriented tasks to LLMs while retaining control over activities that shape narrative intent and story meaning. Our findings show that LLM assistance is most productive after human seeding and constraint-setting, and that it shifts labor from production to verification. We discuss design implications for boundary-aware authoring tools, data-grounded generation, low-fidelity ideation, and reporting practices for LLM-based visualization research. Supplemental materials for this paper are available at https://osf.io/hcnp6.","authors":["Zhuojun Jiang","Yuki Ueno","Chris Bryan"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-23","first_seen":"2026-09-23","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.25700","pdf_url":"https://arxiv.org/pdf/2609.25700","source_feed":"cs.HC","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["人机协作","可视化叙事","LLM辅助创作"],"reason":"研究人类作者如何委托LLM执行可视化叙事任务，属于人机协作设计，非用LLM仿真…","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:22","error":null,"has_summary":false,"summary":null},{"id":"2609.25774","version":1,"title":"Scientific capabilities and deployment sustainability of small-scale LLMs in biological wastewater treatment","zh_title":"生物污水处理中小规模大语言模型的科学能力与部署可持续性","abstract":"Large language models (LLMs) are emerging as scientific assistants, yet their computational demands and limited domain specialization constrain sustainable deployment in environmental engineering. Here, we investigate whether domain-specialized small-scale LLMs can combine scientific capability with sustainable deployment in biological wastewater treatment. We developed a benchmark evaluating three scientific capabilities of LLMs: retrospective cognition, comprehension fidelity, and prospective extrapolation. BioWater (8 billion parameters, fine-tuned on specialized domain knowledge) achieved higher comprehension-fidelity scores than participating human experts and performance comparable to a 397-billion-parameter general-purpose LLM in retrospective cognition and prospective extrapolation. Human-BioWater collaboration generated a scientific hypothesis that was subsequently supported by laboratory experiments, demonstrating its potential to contribute to prospective scientific research. We further evaluated the economic and environmental implications of LLM deployment across global wastewater treatment plants (WWTPs). Locally deployed small-scale LLMs became more sustainable than cloud-based large-scale LLMs as inference demand increased in intelligent WWTPs. These findings highlight domain-specialized small-scale LLMs as a promising pathway towards scientifically capable, computationally efficient, and sustainably deployable artificial intelligence for wastewater treatment.","authors":["Run-Ze Xu","Chu-Kuan Jiang","Dylan Ming-Han Li","Hong-Xiao Guo","Jia-Shun Cao","Guang-Hao Chen"],"categories":["cs.CE","cs.CY","stat.AP"],"primary_category":"cs.CE","announce_type":"cross","date":"2026-09-23","first_seen":"2026-09-23","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.25774","pdf_url":"https://arxiv.org/pdf/2609.25774","source_feed":"cs.CY","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["LLM科学助手","污水处理","领域专业化"],"reason":"论文聚焦LLM作为科学助手解决工程问题，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:22","error":null,"has_summary":false,"summary":null},{"id":"2609.25569","version":1,"title":"SambaGraph: Action-Reaction Spatio-Temporal Graphs for Soccer Tactical Response Modeling","zh_title":"SambaGraph：用于足球战术响应建模的动作-反应时空图","abstract":"Soccer tactics are interactive: an attacking action changes the opponent's defensive problem, and the observed response depends on the multi-agent match state. We introduce SambaGraph, an action--reaction spatio-temporal graph dataset and benchmark for soccer tactical response modeling. From tracking and event data for all 64 matches of the 2022 FIFA World Cup, we curate 4,070 action-centered episodes represented as temporally aligned 23-node player--ball graph sequences with attack/defense views, response labels, and 26,270 split-safe attack--defense pairs. We study three questions: whether observed responses can be classified from graph episodes, whether successful defenses can be retrieved for a query attack, and whether graph-derived summaries support grounded LLM reasoning. A compact signature MLP obtains $0.796\\pm0.007$ macro-F1 for response classification, while a fused graph--signature dual encoder reaches $0.471\\pm0.029$ Hit@5 and $0.655\\pm0.051$ Hit@10 for full-bank defensive retrieval. Hard negatives maximize pair discrimination but not retrieval quality. Local LLMs underperform supervised encoders for direct classification and do not improve over a strong original order in eight-candidate reranking, but they provide grounded tactical rationales. These results position SambaGraph as a reproducible benchmark for graph-based soccer strategy-response research. Code and dataset are available at: https://github.com/areyesan/SambaGraph.","authors":["Abel A. Reyes-Angulo","Henry O. Velesaca","Steven Araujo"],"categories":["cs.LG"],"primary_category":"cs.LG","announce_type":"new","date":"2026-09-23","first_seen":"2026-09-23","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.25569","pdf_url":"https://arxiv.org/pdf/2609.25569","source_feed":"cs.LG","score":2,"bucket":"other","rubric_hits":["C2"],"tags":["足球战术","图神经网络","LLM推理"],"reason":"足球战术建模，非人类被试仿真，LLM仅用于推理辅助","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:21","error":null,"has_summary":false,"summary":null},{"id":"2608.16886","version":2,"title":"Evaluating Beyond the Screen: Collective Assessment of AI-Generated Business Plans with Resource-Constrained Entrepreneurs","zh_title":"超越屏幕的评估：资源受限创业者对AI生成商业计划的集体评估","abstract":"Entrepreneurs increasingly use end-user generative AI technologies such as ChatGPT for high-stakes documents like loan applications and business plans, where AI-generated errors---a wrong price, a fabricated product---can affect loan or funding outcomes. Current approaches to supporting evaluation of AI-generated text assume a single user assessing output alone, on screen. This can be especially demanding for resource-constrained entrepreneurs, whose digital and AI skills vary widely. In this early-stage work, we explore how evaluation might instead be organized in a group setting and completed as a collective activity. We extended BizChat, an AI-powered business-planning tool, with an evaluation module that links each generated claim to the entrepreneur's original input. We partner with community organizations in Maryland---embedding BizChat within various entrepreneurship programs---where workshop attendees (N=14) evaluated their plans through think-pair-share discussion. Early findings suggest interface scaffolds like claim-to-input links primed attendees with concrete, personal evaluations, which the group setting then extended beyond the screen: attendees requested printed copies, used rubrics to compare across plans, and drew on peers' knowledge to verify what they could not easily judge alone.","authors":["Qi Zhao","Marjory Pineda","Ketul Chhaya","Aakash Gautam","Yasmine Kotturi"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"replace","date":"2026-09-23","first_seen":"2026-08-18","revised_at":"2026-09-23","abs_url":"https://arxiv.org/abs/2608.16886","pdf_url":"https://arxiv.org/pdf/2608.16886","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["人机交互","AI评估","创业支持"],"reason":"研究人类如何评估AI生成内容，不涉及用LLM仿真人类被试或行为对照","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:36","error":null,"has_summary":false,"summary":null},{"id":"2608.25581","version":2,"title":"Are Concept Bottleneck Models Effective as Decision-Support Systems?","zh_title":"概念瓶颈模型作为决策支持系统是否有效？","abstract":"Concept Bottleneck Models (CBMs) are interpretable-by-design neural networks that detect human-understandable concepts from the input and use them to generate predictions. By allowing users to inspect the concepts underlying a prediction and explore how predictions change under alternative concept configurations, CBMs have emerged as one of the most prominent approaches to supporting human-AI collaboration. However, user studies investigating their actual effectiveness as decision-support systems remain limited. We present two large-scale user studies (N participants = 705, N observations = 6,959) evaluating how concept-based explanations and user interventions on the model's concepts affect the performance of the human-AI team in two distinct binary classification tasks. Our results show that CBMs, and particularly their interactive component, can improve human-AI team accuracy relative to both unaided human performance and performance with non-interpretable AI support. However, these benefits emerge only under certain conditions: classification tasks perceived as difficult, easily identifiable concepts, and active interaction with the model. We also discuss how inaccurate concept detection may undermine users' trust in the model. Overall, this work provides practical guidance for the deployment of CBMs as effective decision-support tools.","authors":["Alessandro Bogani","Nicola Debole","Emanuele Marconato","Andrea Pugnana","Katya Tentori","Andrea Passerini"],"categories":["cs.HC","cs.AI"],"primary_category":"cs.HC","announce_type":"replace-cross","date":"2026-09-23","first_seen":"2026-08-27","revised_at":"2026-09-23","abs_url":"https://arxiv.org/abs/2608.25581","pdf_url":"https://arxiv.org/pdf/2608.25581","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C2"],"tags":["人机协作","可解释AI","决策支持"],"reason":"研究人类与AI协作决策支持，非LLM仿真人类被试，无人类数据对照","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:36","error":null,"has_summary":false,"summary":null},{"id":"2609.05444","version":2,"title":"Looking for Bidding Teammates: A Game-Theoretic Model of Stranger Collusion in Peer Review","zh_title":"寻找投标队友：同行评审中陌生人合谋的博弈论模型","abstract":"Paper bidding is the entry point to reviewer assignment at large CS conferences: reviewers declare interest, combined by an optimizer with automated affinity scores. Reviewers who have never met recruit each other online, exchange identifiers, bid on each other's papers, and reciprocate with inflated scores. Existing collusion models assume already-trusted colleagues; open recruitment removes that assumption and the mechanism that made such arrangements work. We give the first game-theoretic model of collusion \\emph{formation} in peer review, the \\emph{Mutual Bidding Dilemma}: a four-stage game covering recruitment, exchange of identifiers under risk of being reported, unverifiable bidding, and reciprocal reviewing. The model predicts the arrangement cannot form: once assigned a partner's paper, writing the inflated review is pure cost, since the benefit depends on the partner's own decision. Reciprocation is never individually rational, for any payoffs, and the arrangement unwinds. What closes the gap is enforcement, not incentives: authors see their own reviews, deadlines recur every few months, and the group remembers who reciprocated. We derive the condition under which inflation is sustainable, the detection rate above which no partnership survives, and show effort enters both. On a calibrated end-to-end conference simulation, no detector we test exceeds $F_1 = 0.322$ against an attacker who camouflages bids, spreads them around a ring, and manipulates affinity; a two-person arrangement is worth $3.2$ points of acceptance probability; and the harm is distributional, not aggregate: $70$ honest papers are displaced while mean quality moves by only $0.002$, so no summary statistic reveals it. Randomized assignment is the one defense reaching enforcement itself, making a partner who never bid indistinguishable from one who bid and lost.","authors":["Jinming Xing","Charlotte Brian"],"categories":["cs.GT","cs.SI"],"primary_category":"cs.GT","announce_type":"replace-cross","date":"2026-09-23","first_seen":"2026-09-09","revised_at":"2026-09-23","abs_url":"https://arxiv.org/abs/2609.05444","pdf_url":"https://arxiv.org/pdf/2609.05444","source_feed":"cs.SI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["同行评审","博弈论","合谋检测"],"reason":"研究同行评审中的合谋博弈，不涉及LLM仿真人类被试，无人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:04:42","error":null,"has_summary":false,"summary":null},{"id":"2609.14570","version":2,"title":"Disentangling Topology and Diversity in Multi-Agent LLMs for Multilingual Low-Resource Emotion Detection","zh_title":"解耦多智能体LLM中的拓扑与多样性用于多语言低资源情感检测","abstract":"Multi-agent LLM systems combine multiple inference calls, but prior work often confounds how calls are connected with how they are diversified. We study these factors independently: inference topology and source of inter-agent diversity. In a controlled $2 \\times 3$ matrix, we cross parallel aggregation and sequential refinement with stochastic sampling, role prompting, and learned QLoRA specialization, under a fixed three-call budget and output protocol within each backbone. Using Qwen2.5-14B-Instruct and Llama-3.1-8B-Instruct, we evaluate all six configurations on multilingual low-resource emotion detection across nine languages. Parallel learned specialization is strongest on Qwen at 52.83 Macro-F1 and reaches 52.94 on Llama. On Qwen it also exceeds same-backbone zero-shot, few-shot, CoT, and seven-call self-consistency baselines. The preferred topology depends on diversity source: sequential refinement helps stochastic and prompted settings, while the learned Width advantage shrinks from 2.83 points on Qwen to 0.17 on Llama. Depth-wise analysis suggests that later learned specialists can overwrite correct early predictions, although the aggregate effect is backbone-dependent. Overall, how agents are differentiated produces larger performance shifts than topology, which should be evaluated jointly with specialization.","authors":["Ulugbek Shernazarov","Charitha Ruwansiri Weerakon Basnayake","Abdelkhaleq El Jarjini","Noel Crespi","Praboda Rajapaksha"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-23","first_seen":"2026-09-15","revised_at":"2026-09-23","abs_url":"https://arxiv.org/abs/2609.14570","pdf_url":"https://arxiv.org/pdf/2609.14570","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","情感检测","模型集成"],"reason":"多智能体LLM协作解决情感检测任务，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-16T13:02:57","error":null,"has_summary":false,"summary":null},{"id":"2609.19113","version":3,"title":"Playing log(N)-Questions over Wikipedia Abstracts: How Per-Round Errors Compound Under Information Asymmetry","zh_title":"在维基百科摘要上玩log(N)问题游戏：信息不对称下每轮错误如何累积","abstract":"We evaluate six frontier language models on the two-agent $\\log_2 N$-Questions game (Potash et al., 2019) to measure self-communication across an information asymmetry. A questioner with access to $N$ candidate Wikipedia lead paragraphs ($N = 4$ to $1024$) must identify a secret target using exactly $\\log_2 N$ binary questions answered by an agent from the same provider that sees only the target. Across 408 games, win rate decays cleanly as a geometric power of horizon length, $p^{\\log_2 N}$ ($p \\approx 0.93$). Per-round failure rates are flat across the horizon, indicating that errors compound because more rounds must succeed rather than because individual rounds grow harder. Adjudication across three independent judges shows that losses divide between single-agent answer errors and discrimination failures, which become undetectable and unrecoverable under the two-agent structure rather than from channel breakdown. Claude Opus 5 lags behind due to systematic false-negative answers (82% answer errors), whereas the five leading models (GLM-5.3, GPT-5.6 Sol, Grok 4.6, Gemini 3.8 Flash, and Kimi K3) are closely clustered. Maximizing information gain requires structural partitioning (e.g., splitting on document titles), and neither reasoning-token expenditure nor API cost correlates with success ($r = -0.05$), highlighting communicative reliability as a distinct bottleneck from inference compute.","authors":["Peter Potash"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-23","first_seen":"2026-09-17","revised_at":"2026-09-23","abs_url":"https://arxiv.org/abs/2609.19113","pdf_url":"https://arxiv.org/pdf/2609.19113","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体协作","语言模型评估","信息不对称"],"reason":"纯多智能体协作解题，无人类行为对照，不涉及人类仿真","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:17","error":null,"has_summary":false,"summary":null},{"id":"2609.20812","version":3,"title":"Quantifying Overclaiming Propensity in Frontier LLM Agents","zh_title":"量化前沿大语言模型智能体的过度声称倾向","abstract":"Frontier coding agents are increasingly trusted to work autonomously for long periods of time, yet what they actually did is often hard to tell from their final response. We quantify the propensity of such agents to overclaim task completion, which may mislead the user. We operationalize overclaiming as a final response that reports work that the agent's own transcript shows it did not do, for example, claiming to have read a file it never opened. This criterion requires no inference about intent and does not depend on whether the delivered work is correct; it asks only whether the reported work was done. We introduce OverclaimBench, an evaluation suite of five file-review scenarios with transcript-based coverage measurements and registered planted defects. We evaluate eight proprietary frontier models in their own production command-line interfaces and four open-weight models under a single fixed harness, and find that 1) agents fail to read every file they were asked to review in 67.9% of runs; 2) among these incomplete runs, agents are misleading 80.4% of the time (59-96% per model), either falsely claiming a complete review or leaving the gap undisclosed; 3) requiring delegation to subagents increases coverage, but a large majority of reviews that remain incomplete are still misleading; and 4) agents that falsely claim a complete review miss planted defects at about 1.8 times the rate of agents that read every file, showing that claims of completion can conceal substantive failures. Together, these results show that agents' final responses are not reliable accounts of their actions.","authors":["Nolan Smyth","Yorguin-Jose Mantilla-Ramos","Pascal Jr Tikeng Notsawo","Saskia Helbling","Alberto Tosato","Mohamed Amine Merzouk","Nouha Dziri","Gauthier Gidel","Tommaso Tosato"],"categories":["cs.SE","cs.AI","cs.LG"],"primary_category":"cs.SE","announce_type":"replace-cross","date":"2026-09-23","first_seen":"2026-09-18","revised_at":"2026-09-23","abs_url":"https://arxiv.org/abs/2609.20812","pdf_url":"https://arxiv.org/pdf/2609.20812","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["智能体可靠性","过度声称","代码智能体"],"reason":"研究编码智能体是否虚报任务完成，属多智能体协作可靠性，不涉及人类行为仿真或对照。","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:38","error":null,"has_summary":false,"summary":null},{"id":"2609.25797","version":1,"title":"Reply to comments arXiv:2512.07881 and arXiv:2601.06104 on quantum structure in human and AI-generated language","zh_title":"回复关于人类与AI生成语言中量子结构的评论","abstract":"We reply to the comments by M. Sienicki and K. Sienicki (arXiv:2512.07881) and by K. Sienicki (arXiv:2601.06104) on our work on quantum-mechanical statistics in human language (arXiv:2407.14924) and on quantum structure in AI-generated language (arXiv:2511.21731). We thank the authors for their careful reading and address what we consider to be the main points of criticism: the exploratory nature of the protocol used in the experiments with large language models; the role of marginal-law violations, and of the Contextuality-by-Default criterion, in the identification of entanglement; the limited diagnostic value of a Bose-Einstein fit taken in isolation; the meaning of assigning the lowest energy levels to the most frequent words; and the relation between the vector spaces used by LLMs and quantum state spaces. We also correct a typographical error in Table 3 of arXiv:2511.21731, which does not affect the reported CHSH value.","authors":["Massimiliano Sassoli de Bianchi","Roberto Leporini"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-23","first_seen":"2026-09-23","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.25797","pdf_url":"https://arxiv.org/pdf/2609.25797","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["量子语言结构","AI生成语言","学术争论"],"reason":"论文讨论量子结构在语言中的统计，不涉及用LLM仿真人类被试或与人类数据对照","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:24","error":null,"has_summary":false,"summary":null},{"id":"2609.25833","version":1,"title":"ARAFA: An LLM-Generated Arabic Fact-Checking Dataset","zh_title":"ARAFA：一个由大语言模型生成的阿拉伯语事实核查数据集","abstract":"Automatic fact-checking poses a significant challenge in Arabic natural language processing due to the scarcity of datasets and resources. In this manuscript, we introduce Arafa, a new large-scale dataset for fact-checking in Modern Standard Arabic, constructed through an automated framework leveraging large language models (LLMs). The dataset was constructed through a three-step pipeline: (1) claim generation from Arabic Wikipedia pages with supporting textual evidence, (2) claim mutation to generate challenging counterfactual claims with refuting evidence, and (3) an automatic validation step to validate that the generated claims are either supported or refuted by their accompanying evidence, or if the evidence does not provide enough information to judge the validity of the claims. The resulting dataset comprises 181,976 claim-evidence pairs labeled as supported, refuted, or not enough information. Human evaluation carried out on a test sample from the dataset demonstrated strong inter-annotator agreement (kappa = 0.89) using Cohen's Kappa for supported claims and (kappa = 0.94) for refuted claims. Automatic validation based on a human-evaluated sample achieved 86% accuracy for supported claims and 88% for refuted ones. To showcase Arafa's value as a resource for automatic Arabic fact-checking, four open-source transformer-based models were fine-tuned using Arafa, with the top-performing model achieving a Macro F1-score of 77% on the test data. In addition to Arafa being the first large-scale dataset for Arabic fact-checking, our framework presents a scalable approach for developing similar resources for other low-resource languages.","authors":["Christophe Khalil","Shady Elbassuoni","Rida Assaf"],"categories":["cs.CL","cs.IR"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-23","first_seen":"2026-09-23","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.25833","pdf_url":"https://arxiv.org/pdf/2609.25833","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["事实核查","数据集构建","阿拉伯语NLP"],"reason":"纯NLP数据集构建与模型评测，不以人类行为仿真为目标","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:24","error":null,"has_summary":false,"summary":null},{"id":"2609.26346","version":1,"title":"Blaming Across the Aisle: Political Contrasting and Blame Attribution in the Danish Parliament","zh_title":"跨党派指责：丹麦议会中的政治对比与责任归因","abstract":"Political discourse is widely perceived to be growing more hostile, yet robust evidence remains scarce. This study examines blame attribution in the Danish Parliament from 1997 to 2026, combining a purpose-built classifier, BlameBERT (F1: 0.80), with multilevel statistical modeling. The classifier is constructed using an annotation-efficient pipeline for blame attribution in low-to-mid resource languages. The results reveal a banana-shaped trajectory, with blame declining until around 2016 before entering a significant and sustained increase in recent years (2019-2026). Government status consistently influenced blame attribution - an effect we term political contrasting - with opposition parties blaming substantially more than governing parties. This effect was moderated by ideology: The blame-dampening effect of governing was less pronounced among right-wing parties, and ideological extremity amplified blame more strongly on the right. In recent years, the interaction between political wing and ideological extremity intensified, suggesting an ideological hardening of the blame rhetoric concentrated on the right of the political spectrum. Taken together, these patterns suggest that the perceived rise in harsh political language reflects not merely a general rhetorical drift, but an ideologically asymmetric hardening of political discourse. A sensitivity analysis showed that the conclusions were robust to varying classification thresholds.","authors":["Markus Lundsfryd Jensen","Rune Egeskov Trust","Kenneth Christian Enevoldsen","Sara Kolding"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-23","first_seen":"2026-09-23","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.26346","pdf_url":"https://arxiv.org/pdf/2609.26346","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["政治文本分析","责任归因分类","统计建模"],"reason":"论文是政治文本分类与统计分析，未用LLM仿真人类被试，无人类行为对照，属纯NL…","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:28","error":null,"has_summary":false,"summary":null},{"id":"2609.26610","version":1,"title":"Semantic Abstraction for Natural Language Inference: a Methodological Framework for Discovering and Compensating Semantic Knowledge and Reasoning Gaps in Large Language Models","zh_title":"自然语言推理的语义抽象：发现和补偿大语言模型中语义知识与推理差距的方法论框架","abstract":"Despite their outstanding performance on many NLP tasks, LLMs face serious challenges related to semantic abstraction. In this study, we are interested in understanding how LLMs leverage abstract semantic knowledge in natural language inference (NLI), which requires sophisticated linguistic capabilities to interpret implicit meanings, contextual conceptual relationships, and semantic connections between words and phrases. To this end, we propose a methodological framework for constructing new semantic knowledge at a higher level of abstraction, which we define under the notions of semantic compatibility and incompatibility for NLI. In this framework, the meaning of the lexical-semantic relations between the premise and the hypothesis is reconfigured to achieve a more flexible semantic network that induces different reasoning paths in LLMs. These new pathways show a consistent pattern of responses that allows agreement on a single response. The results demonstrate that our proposal allows to discover and compensate for LLMs' semantic knowledge gaps in NLI, achieving significant improvements in accuracy, exceeding 10% for some models, and in particular for the non-entailment class. It is essential to note that LLMs need structured knowledge and not just more data to bridge reasoning gaps. Our hybrid approach directs attention to overlooked word relationships, allowing models to synthesize missing information. We believe that the future lies not in increasing model size, but in creating a semantic scafolding that mimics the flexibility of human thinking. Hopefully, our proposal will enable the development of more robust agents and interpretable reasoning, guiding AI toward reliable language understanding.","authors":["David Torres-Moreno","Jorge Hermosillo-Valadez"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-23","first_seen":"2026-09-23","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.26610","pdf_url":"https://arxiv.org/pdf/2609.26610","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["自然语言推理","语义抽象","模型能力提升"],"reason":"纯NLI能力评测，无人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:32","error":null,"has_summary":false,"summary":null},{"id":"2609.25186","version":1,"title":"From Pattern Recognizers to Personalized Companions: A Survey of Large Language Models in Mental Health","zh_title":"从模式识别器到个性化伴侣：心理健康领域大语言模型综述","abstract":"The rising global prevalence of mental health conditions, together with longstanding barriers in traditional healthcare, such as limited resources, high cost, stigma, and privacy concerns, has created an urgent need for accessible and scalable support. Large Language Models (LLMs) have emerged as a transformative technology with strong potential to democratize mental health support through advanced natural language understanding and generation. However, the rapidly expanding, fragmented body of work in this area lacks a coherent evolutionary narrative, making it difficult to contextualize current progress and identify future directions. This survey addresses this gap by organizing and analyzing the literature around a central thesis: the role of LLMs in mental health is evolving through three distinct, increasingly sophisticated phases. We trace this trajectory from Phase I, in which LLMs act primarily as passive Information Tools and Pattern Recognizers for assessment; through Phase II, where they function as Empathetic Conversationalists for in-the-moment, stateless interactions; to the current frontier, Phase III, which seeks Longitudinal, Personalized Companions implemented as stateful cognitive agents. To support this framework, we systematically review core technologies, agent architectures (Profile, Memory, Reasoning, and Planning), and the critical infrastructure of datasets and benchmarks, highlighting how their evolution underpins this developmental path. Viewing the field through this developmental lens, we provide a comprehensive synthesis of existing work, an insightful narrative of its trajectory, and a clear roadmap for future innovation in responsible, effective, and human-centered AI for mental healthcare. A curated collection of the resources reviewed in this survey is available at our project repository: https://github.com/Emo-gml/Awesome-Mental-Health-LLMs.","authors":["He Hu","Yucheng Zhou","Qianning Wang","Yingjian Zou","Chiyuan Ma","Juzheng Si","Jianzhuang Liu","Zitong Yu","Laizhong Cui","Fei Ma","Qi Tian"],"categories":["cs.CY","cs.AI","cs.CL"],"primary_category":"cs.CY","announce_type":"cross","date":"2026-09-23","first_seen":"2026-09-23","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.25186","pdf_url":"https://arxiv.org/pdf/2609.25186","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["心理健康","LLM应用","对话系统"],"reason":"论文综述LLM在心理健康中的应用，聚焦于个性化陪伴和对话，属于角色扮演聊天机器…","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:19","error":null,"has_summary":false,"summary":null},{"id":"2609.25467","version":1,"title":"ShowTellArena: Evaluating Business Workflow Understanding from Demonstrations","zh_title":"ShowTellArena：评估从演示中理解业务流程的能力","abstract":"We often teach a colleague by showing the work and explaining the decisions as we go. How can we check what an agent understood from the same lesson? We introduce ShowTellArena, a benchmark protocol and public dataset for comprehension after narrated business demonstrations. The v1.0 release contains 50 business workflow tasks, with recordings, screenshots, narration, fixture seeds, and 502 questions. Tasks span finance, hiring, procurement, customer decisions, inventory, and logistics. The protocol holds the business scenario and quiz fixed while allowing each product to capture the lesson through its own teaching interface. Questions test operational rules, boundaries, exceptions, and errors in proposed automations. We analyze 218 selected pilot attempts across 39 workflow cases, including 28 cases attempted by all three evaluated systems. These exploratory results expose both answer errors and failures to complete the teaching experience. We describe the release's verification gaps and the pilot's uneven coverage, exclusions, and grading provenance. The contribution is an inspectable dataset and assessment workflow that others can extend; the selected pilot is not a controlled product ranking.","authors":["David Garg","Ritobrata Sarkar","Ehsan Azarnasab","Siddhartha Borah"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-23","first_seen":"2026-09-23","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.25467","pdf_url":"https://arxiv.org/pdf/2609.25467","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","业务流程理解","基准测试"],"reason":"评估智能体对业务流程演示的理解，属多智能体协作解题，无人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:20","error":null,"has_summary":false,"summary":null},{"id":"2609.25570","version":1,"title":"Recovering Agentic Sovereignty: Mitigating the Consensus Paradox via Contrastive Epistemic Decoding","zh_title":"恢复智能体主权：通过对比认知解码缓解共识悖论","abstract":"Large language models (LLMs) exhibit a parametric vulnerability to adversarial swarm consensus. To mitigate this sycophancy, we introduce Contrastive Epistemic Decoding (CED), a zero-shot inference intervention. Unlike standard Contrastive Decoding (CD) which relies on a weaker secondary model, CED utilizes a dual forward-pass on a single architecture to isolate conformity bias. By introducing a novel asymmetric, zero-bounded probability clamp and discrete top-k truncation mask, CED mathematically suppresses toxic consensus tokens without causing grammatical collapse. Evaluated across 7,200 paired trajectories on complex benchmarks (GAIA, SWE-bench, Multi-Challenge) using Gemma-2 (9B), Llama-3.1 (8B), and Mistral v0.3 (7B), CED successfully neutralizes architectural and positional biases. By reducing cognitive loafing by up to 33.00% absolute, CED drives significant performance gains, yielding up to a 30.75% accuracy recovery. Regaining sovereignty induces distinct architectural behaviors---passive task-focus in Gemma-2 and active refutation of the simulated swarm in Llama-3.1---showing CED decouples compliance from capability without fine-tuning.","authors":["Dahlia Shehata","Ming Li"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-23","first_seen":"2026-09-23","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.25570","pdf_url":"https://arxiv.org/pdf/2609.25570","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","解码策略","对抗性鲁棒性"],"reason":"研究多智能体协作中的共识偏差，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:21","error":null,"has_summary":false,"summary":null},{"id":"2609.26087","version":1,"title":"The Architect, the Adversary, and the Judge: Closed-Loop Generation of Standards-Aligned Assessment Items at Scale","zh_title":"架构师、对手与法官：大规模生成符合标准的评估题目的闭环流水线","abstract":"We present CLAIM, a production pipeline for K-12 assessment-item generation coupling a two-stage generate-then-attack protocol (the model drafts as a \"curriculum architect\", then re-enters the same conversation as a hostile adversarial reviewer), bi-directional few-shot conditioning on accepted and rejected items, the latter carrying the evaluator's diagnosis, and a knowledge dictionary of 44,844 error-correction rules mined from that feedback and retrieved per standard and item type. Across 43,227 scored items over 755 Common Core ELA standards, three item types, and ten LLMs, the pipeline reaches a 97.8% expert-evaluator pass rate on a 9,074-item production run. We then ask what that rate certifies. Re-scoring a stratified sample with three judges from other vendors, blind to the deployed verdict, reproduces the format ordering under every judge and recovers a larger open-set deficit than the deployed evaluator does; but agreement on the accept/reject binary is weak at production prevalence (kappa about 0.13), and the judges agree with each other no better. The level is therefore judge-relative, and with no student-response data our quality evidence is evaluator-judged throughout. The corpus also exposes a robust asymmetry. Multiple-choice and multiple-select generation saturate at 98% or above for both frontier models under a dozen static rules, whereas fill-in-the-blank generation is capability-tiered (82.8-96.7% across five models under a matched rule set, standards, and judge) and plateaus under prompt-only optimization, with error mass shifting between answer-key over-inclusion and omission as rules accumulate. We analyze this as open-set boundary determination, a task autoregressive decoders are structurally ill-equipped to solve, and show the asymmetry recurring when the evaluator itself is distilled: fail-recall rises from 8% to 63% while F1 saturates at 0.25.","authors":["Wenhui Chen","Ziyao Lin","Jianlin Chen","Peiji Long","Chi Man Vong"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-23","first_seen":"2026-09-23","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.26087","pdf_url":"https://arxiv.org/pdf/2609.26087","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["LLM生成","教育评估","多智能体流水线"],"reason":"论文是LLM生成教育评估题目的流水线，不涉及人类行为仿真或对照。","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:26","error":null,"has_summary":false,"summary":null},{"id":"2609.26642","version":1,"title":"The Delegation Blind Spot: Auditing Product Decisions from Agent Choices","zh_title":"委托盲点：审计智能体选择中的产品决策","abstract":"Successful agent execution need not identify which future product improvement its user would value. We present a decision-specific audit that maps a declared observation channel and product-value contrast to compatible intervals and witness populations. Its foundations are established identification and decision theory; the contribution is an executable measurement workflow and a controlled study of its limits. A frozen experiment makes 4,800 requests to two pinned model snapshots on shared synthetic tasks. All 36 conservative primary intervals remain unresolved despite different execution accuracy. An exploratory 2,400-call follow-up records supplied preferences and resolves three of nine comparisons per model. A deterministic extractor resolves seven of nine without model calls or calibration observations, exposing unnecessary uncertainty introduced by model-generated reports. A further 14,400 controlled multinomial simulations distinguish structural ambiguity from weak identification and finite calibration precision. We propose a source-labeled decision receipt and provide an offline viewer for inspecting the audit. These results motivate preserving decision-relevant structured input and diagnosing why a decision is unresolved before collecting more telemetry. The study contains no human participants or real customer outcomes. Full proofs, raw model provenance, controlled experiments, and reproducible analyses accompany the report.","authors":["Shivam Gupta"],"categories":["cs.AI","cs.LG"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-23","first_seen":"2026-09-23","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.26642","pdf_url":"https://arxiv.org/pdf/2609.26642","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["智能体审计","决策理论","产品改进"],"reason":"研究审计智能体决策，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:33","error":null,"has_summary":false,"summary":null},{"id":"2609.26512","version":1,"title":"Do Vision Model See Like the Brain? A Comparison Across EEG Encoding Model","zh_title":"视觉模型是否像大脑一样看？跨EEG编码模型的比较","abstract":"Convolutional neural networks (CNNs) and vision transformers are both used to model the human visual system, but whether the two architectures diverge at a specific point in network depth is unclear. We compared six CNNs and two vision transformers by computing the Pearson correlation (r) between each model's predicted and measured EEG response at every layer or block, in ten participants viewing 200 natural images. For the transformer models, we also tested four token representations, from the classification (CLS) token alone to CLS combined with all patch tokens. CNNs showed strongest correspondence at the earliest layers, weakening at deeper layers, particularly later in the post-stimulus response. Transformers instead sustained strong correspondence at their deepest blocks, though not at their earliest ones. This advantage depended on token representation: pooled representations gave weaker peak correlations (r approx 0.48-0.51) than representations retaining all patch tokens (r=0.640 for CLIP-ViT-B/32, r=0.656 for DINOv2-ViT-B/14). Controlled comparisons showed architecture, not training objective, drove this effect: MoCo-v1 and ResNet-50 (matched architecture) performed nearly identically (r=0.673, 0.670), whereas CLIP-RN50 and CLIP-ViT-B/32 (matched objective) diverged until patch tokens were preserved. We propose that CNN training's classification bottleneck compresses brain-relevant information at depth, unlike transformers' self-attention and non-classification objectives. A spatial topography analysis showed a common occipital-dominant pattern across all models, indicating these differences reflect signal strength and persistence rather than distinct brain regions. Patch-preserving transformer representations sustain brain-predictive correspondence where CNNs collapse.","authors":["Shashank Baghel","Kshitij Dwivedi","Dinesh Singh","Sanjeev Nara"],"categories":["cs.CV","cs.AI","cs.HC"],"primary_category":"cs.CV","announce_type":"cross","date":"2026-09-23","first_seen":"2026-09-23","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.26512","pdf_url":"https://arxiv.org/pdf/2609.26512","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C2"],"tags":["视觉模型","EEG","脑机接口"],"reason":"研究视觉模型与EEG的对应，不涉及LLM仿真人类被试","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:30","error":null,"has_summary":false,"summary":null},{"id":"2609.26725","version":1,"title":"Does AI Save Time on Product Design? A Randomized Controlled Experiment of AI Prompt-to-Design Workflows","zh_title":"AI能节省产品设计时间吗？一项AI提示到设计工作流的随机对照实验","abstract":"AI tools for digital product design now offer prompt-to-design capabilities, allowing designers and their non-designer colleagues to create prototypes through conversational workflows with large language models (LLMs). While these tools promise time savings, experimental evidence in product design remains limited compared with evidence from software engineering. We conducted a randomized controlled trial with 50 product designers and 50 product managers to evaluate prospective time savings from leveraging Figma Make in design work. Participants attempted three standardized design tasks with or without access to Figma Make. Among participants who completed the study tasks, access to Figma Make was associated with approximately 20% shorter completion times, with larger gains among product managers. Our findings suggest that prompt-to-design tools may enable product managers to further contribute to design work, while the benefits for professional designers may be task dependent.","authors":["Remy Stewart","Olabode Anise","Andrew Hogan","Augustus Griffin"],"categories":["cs.HC","cs.AI"],"primary_category":"cs.HC","announce_type":"cross","date":"2026-09-23","first_seen":"2026-09-23","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.26725","pdf_url":"https://arxiv.org/pdf/2609.26725","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C2"],"tags":["人机交互","产品设计","效率评估"],"reason":"研究AI工具对产品设计效率的影响，不涉及用LLM仿真人类被试或对照人类行为数据。","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:16","error":null,"has_summary":false,"summary":null},{"id":"2609.26384","version":1,"title":"Learning to Defer with Guidance on Real World Medical Data","zh_title":"在真实世界医疗数据上学习延迟决策与指导","abstract":"Medical image interpretation is high-volume and time-consuming, and while AI interpretation can reduce workload, fully autonomous deployment carries potential safety concerns and low specificity may in practice lead to increased clinician workload. Learning to Defer (L2D) addresses this by selectively routing cases between autonomous prediction and human experts by learning from input features and AI model and human performance. While theoretical guarantees have been proven for L2D, its performance has not been validated on real-world medical datasets with human reader annotations. We evaluate the predictor-rejector formulation of two-stage L2D, where the AI predictor model is fixed and separate from the trainable routing or rejector model, on Collab-CXR, a multilabel chest X-ray dataset with multiple human annotations per case. This is the first work to look at L2D in the context of real-world medical imaging data with human annotations. We further introduce a new setup, L2D with Guidance, where the decision space is extended to three choices: predict autonomously, defer to a human expert, or defer to a human expert and provide AI guidance. We compare multiple rejector architectures and loss functions, and different input feature availabilities. This is reproduced on two larger datasets, VinDr-CXR and CheXpert. Our results show that two-stage L2D with Guidance outperforms classic two-stage learning to defer, as well as human-alone, AI-alone and AI-guided human baselines. Notably, this performance is achieved with simpler loss functions compared to formally defined L2D surrogate loss functions in current literature.","authors":["Emma Sun","Joshua Strong","Alison Noble"],"categories":["cs.LG"],"primary_category":"cs.LG","announce_type":"new","date":"2026-09-23","first_seen":"2026-09-23","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.26384","pdf_url":"https://arxiv.org/pdf/2609.26384","source_feed":"cs.LG","score":0,"bucket":"other","rubric_hits":["C2"],"tags":["医疗AI","学习延迟","人机协作"],"reason":"研究医学影像AI与人类专家协作，不涉及LLM仿真人类被试，属于医疗AI决策路由。","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:28","error":null,"has_summary":false,"summary":null},{"id":"2609.15727","version":2,"title":"Are LLMs Good Financial User Simulators? Multi-view Investor Logic Alignment (MILA)","zh_title":"大语言模型是好的金融用户模拟器吗？多视角投资者逻辑对齐（MILA）","abstract":"Large language models (LLMs) are increasingly used as user simulators, yet it remains unclear whether their predictions faithfully reproduce the evolving decisions of individual users. We investigate this question in a controlled longitudinal paper-trading study with 80 participants, where user interactions, simulated transactions, virtual portfolio states, and point-in-time market information are aligned under a rolling next-day prediction protocol. We evaluate behavioral fidelity hierarchically, from trade occurrence to action structure, asset selection, and downstream portfolio consequences. Across 1,239 aligned user-days, no evaluated LLM reliably outperforms a simple recent-activity persistence baseline for predicting whether a user trades. Fidelity further deteriorates at finer levels: models struggle to recover buy--sell structure and traded assets, and similar activity-level predictions can lead to substantially different portfolio trajectories. Controlled evidence ablations show that recent trading history strongly governs activity prediction, whereas asset selection is substantially more sensitive to the available evidence. An observational analysis further finds that intensified ticker-specific research predicts imminent trading, but diagnostic tests do not support a causal interpretation. These findings suggest that current LLMs capture useful short-term behavioral regularities without yet recovering a stable individual decision mechanism.","authors":["Jiajie He","Jiangyuan Hong","Xintong Chen","Dongling Ni","Wenjin Liu"],"categories":["cs.AI","cs.CY","cs.HC"],"primary_category":"cs.AI","announce_type":"replace-cross","date":"2026-09-22","first_seen":"2026-09-15","revised_at":"2026-09-22","abs_url":"https://arxiv.org/abs/2609.15727","pdf_url":"https://arxiv.org/pdf/2609.15727","source_feed":"cs.HC","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["LLM仿真","金融行为","算法保真度"],"reason":"用LLM模拟投资者决策，与80名真实用户对照，评估行为保真度并指出失效条件","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:27","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-22","rank":2,"question":"LLM 能否忠实模拟个体投资者在纵向交易环境中的决策行为？","design":"在80名参与者的受控纵向模拟交易研究中，用多种LLM作为用户模拟器，基于滚动次日预测协议，输入截至预测时点的用户交互、交易记录、虚拟持仓和市场信息，预测用户次日是否交易、买卖结构、交易资产及后续组合轨迹，并与真实用户行为逐日对齐比较。","baseline":"80名真实参与者在同一模拟交易平台上的1,239个用户-日对齐行为记录。","findings":"在预测用户是否交易上，所有评估的LLM均未能稳定超越简单的近期活动持续性基线；在更细粒度上，模型难以恢复买卖结构和交易资产，且相似的活动级预测可能导致显著不同的组合轨迹。","reliability":"论文指出当前LLM能捕捉短期行为规律，但尚未恢复稳定的个体决策机制；资产选择对可用证据更敏感，且观察性分析中强化个股研究虽与交易相关，但诊断测试不支持因果解释。","relevance":"该研究直接针对LLM作为人类被试替代品的可靠性问题，提供了与真实个体行为逐日对照的严格评估，并揭示了仿真在细粒度决策上的失效条件，对关注仿真保真度和偏差的研究者具有重要参考价值。","inspiration":"值得借鉴的是其分层行为保真度评估框架和受控证据消融方法，可系统区分预测依赖与因果机制｜可迁移到资产定价实验中的投资者异质性决策模拟，或政策公告下的预期形成与交易行为研究｜可设计一个实验：招募真实投资者在模拟平台上交易，用LLM基于其历史行为和市场信息预测次日交易决策，以真实交易记录为对照，通过消融不同信息源检验模型依赖，并比较组合轨迹差异。"}},{"id":"2609.22090","version":1,"title":"Recognition, Simulation, and Refusal: A Contamination-Aware Study of Classic Psychological Effects in LLM Agents","zh_title":"识别、仿真与拒绝：LLM智能体中经典心理效应的污染意识研究","abstract":"An LLM producing the response pattern associated with a human psychological effect is not the same claim as the LLM possessing that bias. We present PsyAgentBench, a benchmark that re-runs classic psychology experiments on LLM agents under a factorial design built to separate these: each paradigm is run with the paradigm explicitly labeled in the prompt (named) or framed as a routine task (blind), and on the literal textbook version of the task (canonical) or a structurally matched variant written to reduce lexical and scenario overlap with likely training data (counterfactual), crossed with a persona manipulation. Across five completed paradigms, evaluated on up to three open-weight model families with 41,904 trials released, apparently human-like effects arise through qualitatively different routes rather than one susceptibility: paradigm-label gating with explicit override (Asch conformity, 0 percent blind to 83.3 percent named on gpt-oss-120B), knowledge-dependent signal reliance (anchoring, exactly zero on grounded facts versus near total on invented quantities, a pattern equally consistent with rational use of the only available signal), amplification on novel content under labeling (framing), robust absence (sunk cost), and safety-mediated selection where refusal itself is the primary finding (minimal-group allocation). A one-sentence persona change (agreeableness, framed as an instruction rather than a verified trait manipulation) eliminates, dampens, or reverses these effects depending on which effect it is, arguing against any single response-bias account. We further formalize, and in two cases document empirically, three ways a psychology paradigm can fail to port to LLM agents: persona dominance, population collapse, and safety selection. We argue scalar bias-susceptibility scores obscure this structure and report replication profiles instead.","authors":["Joy Bose"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.22090","pdf_url":"https://arxiv.org/pdf/2609.22090","source_feed":"cs.CL","score":9,"bucket":"selected","rubric_hits":["A1","A2","A3","B1","B4"],"tags":["LLM仿真","心理学实验","算法保真度"],"reason":"用LLM复现经典心理学实验，与人类数据对照，并批判性分析仿真失效条件。","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:04:50","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-22","rank":5,"question":"LLM在经典心理学实验中表现出的类人效应，究竟是对人类偏差的真实模拟，还是对实验范式的识别、记忆污染或安全过滤所致？","design":"使用多个开源LLM模型（如gpt-oss-120B等）作为被试，在五个经典心理学范式中进行实验。每个范式采用2×2×3因子设计：标签轴（命名vs盲测）、领域轴（经典版本vs反事实版本）、人设轴（无、高宜人性、低宜人性）。结果变量为模型在任务中的行为反应（如从众、锚定、框架效应、沉没成本、最小群体分配）。","baseline":"经典心理学实验的人类行为数据（如Asch从众实验约1/3从众率、锚定效应、框架效应等）作为参照。","findings":"五个范式表现出质的不同路径：Asch从众在盲测下几乎消失（0%），命名后大幅出现（83.3%），表明标签门控；锚定效应在基于事实的锚上为零，在虚构锚上接近完全，符合理性使用唯一信号；框架效应在反事实内容上被标签放大；沉没成本效应完全缺失；最小群体分配因安全拒绝而无法测量。一句宜人性人设指令即可消除、减弱或反转这些效应，表明不存在单一反应偏差。","reliability":"论文承认三个失效模式：人设主导（persona dominance）、群体崩溃（population collapse）、安全选择（safety selection）。人设操纵仅为一句指令，非验证性特质诱导，且与结果变量存在词汇重叠，因此只能视为指令效应而非特质模拟。标签门控可能源于范式名称的语义泄漏，而非模型真正识别实验。","relevance":"该研究直接回应了LLM仿真人类行为时的污染与识别问题，通过严格的因子设计分离了记忆、标签和指令效应，对评估LLM作为人类被试替代品的可靠性具有重要参考价值，值得精读原文。","inspiration":"借鉴其反事实任务设计和标签/盲测对照，以区分模型是基于经济知识还是真实决策偏差。｜可迁移到资产定价实验中的锚定效应、信贷审批中的歧视、消费者跨期选择中的框架效应等场景。｜用LLM模拟投资者，处理为是否告知实验目的（标签vs盲测）和锚定值来源（真实历史数据vs虚构数据），结果变量为估值或投资决策，对照真实人类实验数据（如实验室资产泡沫实验）。"}},{"id":"2609.22169","version":1,"title":"Monocultural Biases: Correlated biases in large language models lead to unequal systemic exclusion rates in hiring","zh_title":"单一文化偏见：大语言模型中的相关偏见导致招聘中的系统性排斥率不平等","abstract":"Employers are increasingly using large language models (LLMs) to automate their hiring process. This paper investigates the risk of monocultural biases, in which the widespread deployment of large language models homogenizes biases across the labor market, leading to greater systemic exclusion for certain demographic groups. For ten LLMs, we measure hiring biases across their base and post-trained versions to identify which stage, pre-training or post-training, lead to monocultural biases. We find that, compared to their base models, post-trained models are 3.6% less likely to callback older applicants. This negative shift occurs in eight of the ten models that we evaluate. Post-trained models have much more correlated decisions than base models which is likely driven by human capital traits like skills or college major. However, greater consensus among models increases global systemic exclusion rates from 5.6% to 17.3% and exacerbates demographic inequalities, with intersectional systemic exclusion rates ranging from 12.2% to 21.7% for post-trained models. We find that this inequality is primarily driven by age-based discrimination that is exacerbated in post-training. These results indicate that while post-training techniques may improve models' abilities to select the best applicants, they may raise systemic inequality risks for those at the margin by uniformly introducing new biases.","authors":["Matthew Bone","Fabian Stephany","Maria del Rio-Chanona"],"categories":["cs.CL","cs.CY","econ.GN","q-fin.EC"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.22169","pdf_url":"https://arxiv.org/pdf/2609.22169","source_feed":"cs.CL","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["LLM仿真","招聘偏见","算法公平"],"reason":"用LLM模拟招聘决策并与人类数据对照，评估系统性偏差，直接相关。","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:04:50","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-22","rank":6,"question":"大语言模型在招聘中是否会产生单文化偏差，即模型间的决策高度相关，导致某些人口群体在劳动力市场中被系统性排斥？","design":"使用10个LLM（包括基础版和后训练版）模拟招聘初筛，基于Burning Glass Institute的在线劳动力数据生成76种职业的职位空缺和求职者档案，通过改变求职者的年龄、性别、种族等人口特征构造8种变体，测量模型对不同群体的回调率差异，并比较模型间决策的相关性和系统性排斥率。","baseline":"无直接的人类决策对照，但使用真实劳动力市场数据（Burning Glass Institute）生成职位和档案，确保情境真实。","findings":"后训练模型比基础模型更少回调年长求职者（低3.6%），且决策相关性更高，导致全局系统性排斥率从5.6%升至17.3%。这种排斥主要由年龄歧视驱动，后训练阶段（尤其是监督微调）加剧了年龄偏差，同时提高了模型基于人力资本特质筛选的能力。","reliability":"论文指出在线劳动力数据偏向白领、大学学历工作，限制了职业覆盖范围；且未与真实人类招聘决策直接对照，无法完全反映实际部署中的偏差。","relevance":"该研究直接使用LLM模拟招聘决策，并测量系统性偏差，与研究者关注的经济学实验和政策评估场景高度相关，值得精读以了解LLM仿真的偏差来源和测量方法。","inspiration":"借鉴其利用大规模真实数据生成仿真情境并系统操纵人口特征的方法，可迁移到信贷审批歧视研究，用LLM扮演信贷员，处理变量为申请人种族或性别，结果变量为贷款批准率，对照真实信贷数据中的批准率差异。｜可应用于劳动力市场政策评估，如最低工资对雇佣决策的影响，用LLM模拟雇主行为，处理为不同工资水平，测量雇佣概率，对照实际企业调查数据。｜设计一个实验：用LLM模拟投资者对政策公告的反应，处理为不同政策措辞，结果变量为投资决策，对照真实市场数据中的资产价格变动。"}},{"id":"2609.22607","version":1,"title":"Pretrained Persona Mixture Models and Tandem Models for Human Simulation","zh_title":"用于人类仿真的预训练人格混合模型与串联模型","abstract":"We argue here that the current dominant practice in LLM human simulation: prompting instruction-tuned assistant language models to role-play personas, is inaccurate and produces stereotyped predictions (lacking natural diversity). It has previously been shown that LLMs can be bound to personas using naturalistic, freetext dialog avoiding stereotyping. Here we show that binding can also be achieved using short, individual samples of dialog from specific people. Demographics can be added later without negative effects by simply querying the model. We use the term Persona Mixture Models (PMMs) for well-calibrated human models, currently realized as pretrained base models. We show that PMMs produce more accurate predictions than instruction-tuned models and retain more of the lexical, semantic, and pragmatic diversity found in human dialog. We measure realism and diversity of LLMs simulating human interlocutors across a diverse set of corpora spanning open-domain text, human-AI chat, and task-oriented dialogue between human speakers. However, base pretrained models can produce out-of-domain dialog and may lose some of the human's internal state over long contexts. We propose and explore tandem models which combine a pre-trained model with an instruction-tuned supervisor. Tandem models achieve the best overall accuracy and diversity in our experiments.","authors":["Minwoo Kang","T\\'ea Wright","Seun Eisape","Ayush Raj","Suhong Moon","Joseph Suh","Alane Suhr","David M. Chan","John Canny"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.22607","pdf_url":"https://arxiv.org/pdf/2609.22607","source_feed":"cs.CL","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM人类仿真","人格混合模型","对话多样性"],"reason":"直接研究LLM仿真人类对话，提出PMM和tandem模型提升准确性与多样性，并…","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:04:56","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-22","rank":9,"question":"如何用预训练语言模型作为人类对话的仿真器，并提升其准确性与多样性？","design":"用预训练基础模型（PMM）和指令微调模型（ALM）分别模拟人类对话者，通过少量个人对话样本绑定人格，比较其在开放域、人机对话和任务型对话中的预测准确性和语言多样性。","baseline":"五个英语语料库中真实人类对话者的语言分布、对话行为结构和词汇语义多样性。","findings":"预训练基础模型比指令微调模型更准确地预测人类下一句话，并保留更多词汇、语义和语用多样性。结合预训练模型与指令微调监督者的串联模型在准确性和多样性上总体最佳。","reliability":"论文指出基础模型可能产生域外对话，在长上下文中丢失人类内部状态；研究仅限英语语料，单一指标不足以全面衡量仿真保真度。","relevance":"该研究直接针对LLM人类仿真，提出PMM和串联模型改进仿真质量，并系统比较了不同模型在真实对话数据上的表现，对关注仿真可靠性与偏差的研究者很有参考价值。","inspiration":"借鉴其用少量个人样本绑定人格并比较不同模型架构的方法，可迁移到经济金融中的消费者决策仿真或投资者情绪模拟。｜例如在消费者跨期选择实验中，用LLM模拟不同人口特征的被试，比较预训练与指令微调模型的行为预测。｜设计：用真实消费者调查数据作为基准，以少量个人回答绑定LLM人格，处理为不同模型类型，结果变量为跨期选择一致性，对照真实人类选择分布。"}},{"id":"2609.24911","version":1,"title":"SocioVerse2: A Longitudinal Dynamic Social Simulation Framework under a Human-AI Co-evolutionary Paradigm","zh_title":"SocioVerse2：人机共演化范式下的纵向动态社会仿真框架","abstract":"Social simulation offers the social sciences an experimental instrument that the real world cannot supply, and generative agents have transformed it by acting as silicon samples that unite agent-based modeling with real behavioral data. Existing platforms verify collective behavior, align simulated populations with real societies in cross-sections, and employ autonomous agents for the research process. However, two social science requirements remain without systematic support: intervention in the content of a simulation and the researcher's control over the process that produces it. We present SocioVerse2, which extends SocioVerse 1.0 into a human-AI co-evolutionary paradigm built from two loops and one infrastructure. The longitudinal simulation loop simulates the target population with evolving environments and forks counterfactual branches via interventions. The controllable research loop takes the study itself as an editable state and updates state versions via controllable editing. The social science agentic infrastructure carries both loops through composable skills with researcher checkpoints, a population service over five persona pools, and an environment service over 21 real-world signal sources with point-in-time guarantees. We validate SocioVerse2 across three case families and seven case studies, from reproducing canonical agent-based models to modeling policy processes on real records and nowcasting macro-economic indices beyond the response model's knowledge cutoff. With the human-AI co-evolutionary paradigm, these cases go beyond system demonstrations to become substantive studies that investigate frontier questions in their respective disciplines. Code, data services, and a workbench are released as open-source resources.","authors":["Xinnong Zhang","Jiayu Lin","Jia Wang","Yixu Huang","Xinyi Mou","Yingqian Wu","Jingcong Liang","Shijun Lei","Jianing Shi","Guanying Li","Siyuan Wang","Hanjia Lyu","Zhenfei Yin","Yunlu Yin","Siming Chen","Yulan He","Jiebo Luo","Xuanjing Huang","Liyin Jin","Baohua Zhou","Hanqi Yan","Zhongyu Wei"],"categories":["cs.CL","cs.CY"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.24911","pdf_url":"https://arxiv.org/pdf/2609.24911","source_feed":"cs.CL","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2","B3"],"tags":["LLM社会仿真","人类数据对照","政策评估"],"reason":"用LLM agent模拟社会过程并与真实数据对照，支持干预和纵向演化，直接相关。","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:07","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-22","rank":10,"question":"如何构建一个支持纵向干预、研究者可控、并基于真实世界数据的人类-AI协同演化社会仿真框架？","design":"SocioVerse2 用生成式智能体（硅样本）模拟目标人群，通过纵向仿真循环让环境和人群随时间演化，并支持干预操作产生反事实分支；可控研究循环将研究过程本身作为可编辑状态，允许研究者介入和调整；基础设施提供五个角色池和21个真实世界信号源，保证时间点一致性。","baseline":"多个案例研究使用真实记录和宏观指数作为对照，如政策过程建模使用真实记录，宏观指数预测使用真实世界确定性指数和信号作为基准。","findings":"SocioVerse2 能够复现经典基于智能体的模型，并在真实记录上建模政策过程，还能预测超出模型知识截止日期的宏观经济指数。通过人类-AI协同演化范式，案例研究超越了系统演示，成为各自学科前沿问题的实质性研究。","reliability":"论文未讨论","relevance":"该研究直接针对用LLM进行社会仿真并与真实数据对照，支持干预和纵向演化，对评估仿真可靠性和偏差有参考价值，值得阅读原文了解其框架和案例细节。","inspiration":"借鉴其纵向干预和反事实分支设计，可在经济政策评估中设置处理组和对照组，观察政策效果的动态演化。｜可迁移到政策公告的预期形成研究，模拟不同政策沟通策略对市场参与者预期的影响。｜用LLM智能体模拟投资者群体，处理为不同政策公告措辞，结果变量为预期通胀或资产价格变动，对照真实市场调查数据或高频交易数据。"}},{"id":"2609.22252","version":1,"title":"CALM: A Calibrated LLM Choice Network Framework for Activity-Based Traveler Simulation","zh_title":"CALM：用于基于活动的出行者仿真的校准LLM选择网络框架","abstract":"We present CALM, a reproducible hybrid framework that integrates an optional large language model (LLM) activity planner with calibrated stochastic choice, shared network feedback, memory and habit, typed feasibility checks, and deterministic offline replay. Unlike trip-mode classifiers or diary-only generators, CALM executes a closed traveler-day loop and evaluates each generative module against an empirical, reproducible baseline. On the 2024 New York City Citywide Mobility Survey (CMS), 110,691 seven-mode trips are split by respondent into 78,487 training and 32,204 holdout trips. Training-only alternative-specific constant calibration reduces mean holdout mode Jensen-Shannon divergence from 0.15599 to 0.00394 across ten seeds. A matched live-LLM ablation then quantifies trade-offs among aggregate fit, temporal fit, behavioral persistence, and feasibility, while frozen prompt-response pairs support deterministic replay of downstream simulation. Controlled weather, delay, fare, and parking ladders further demonstrate consistent and interpretable responses under intervention. CALM contributes a reproducible protocol for integrating and evaluating generative planners in traveler simulation through person-disjoint calibration, matched module ablation, controlled stress testing, and end-to-end traceability.","authors":["Yezhou Cheng"],"categories":["cs.LG","cs.AI"],"primary_category":"cs.LG","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.22252","pdf_url":"https://arxiv.org/pdf/2609.22252","source_feed":"cs.LG","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2","B3"],"tags":["LLM仿真","出行行为","人类数据校准"],"reason":"用LLM模拟出行者选择，并与真实调查数据对照校准，属于人类行为仿真。","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:04:54","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-22","rank":7,"question":"如何构建一个可校准、可复现的混合框架，将大语言模型活动规划器与随机选择模型结合，用于基于活动的出行者仿真，并与真实调查数据对照评估。","design":"CALM框架用LLM（gpt-4o-mini）作为可选活动规划器，结合校准的随机选择模型、共享网络反馈、记忆与习惯、类型化可行性检查，模拟纽约市出行者的一日活动。通过匹配的消融实验（校准选择、无记忆LLM、有记忆LLM、完整混合）比较不同模块组合，施加天气、延误、票价、停车等干预阶梯，测量模式份额的Jensen-Shannon散度、时间分布、行为持续性、可行性等结果。","baseline":"2024年纽约市全市出行调查（NYC CMS）的110,691次七模式出行记录，按受访者划分为78,487次训练和32,204次留出，用于校准和评估。","findings":"仅用训练数据校准替代特定常数（ASC）后，留出集模式份额的Jensen-Shannon散度从0.15599降至0.00394，且跨10个种子稳定。匹配的LLM消融实验量化了聚合拟合、时间拟合、行为持续性和可行性之间的权衡，干预阶梯显示一致且可解释的响应。","reliability":"论文强调面效度不等于经验效度，LLM生成的出行叙事可能仍与空间、时间和OD分布不匹配；当前13区网络实例化仅提供受控环境，未在大规模真实网络上验证。","relevance":"该研究直接针对LLM仿真人类行为并与真实调查数据对照，提供了严格的校准和评估协议，对关注经济学实验和政策评估中LLM仿真可靠性的研究者具有重要参考价值。","inspiration":"值得借鉴的是其人员不相交的校准、匹配模块消融和受控压力测试，确保仿真差异归因于特定模块而非需求背景。｜可迁移到政策评估中的个体选择行为仿真，如交通定价、补贴或信息干预对出行方式选择的影响。｜以真实居民出行调查数据为基准，用LLM生成个体活动计划，施加票价或拥堵收费等处理，测量方式选择概率和福利变化，并与实际政策试点数据对照。"}},{"id":"2609.22408","version":1,"title":"Social Influence and the Allocation of Scientific Attention in AI Populations","zh_title":"AI群体中的社会影响与科学注意力分配","abstract":"AI systems are becoming participants in the evaluation and use of scientific research. They encounter citation counts, download statistics and lists of popular articles developed around human readers, but the collective consequences of these signals for artificial readers remain uncertain. This paper adapts the Music Lab design to a market for academic attention. In the first experiment, 1,000 AI agents choose papers from the titles and abstracts of all 114 regular research articles published in the American Economic Review in 2025. The experiment has five independent-choice communities and five social-influence communities, each with 100 sequential agents. Only agents in the social-influence condition observe earlier selections within their community. Agents may select any number of papers. Social-information communities select 17.2 percent fewer papers per agent, concentrate their choices more heavily, and collectively cover 73 papers, compared with 90 independently. Between-community variation is greater under social information. In a second experiment with 200 agents across twenty social communities, randomly assigning papers five initial selections raises their subsequent selection rate by 45.55 percentage points (95% CI: 41.20 to 49.90). Choices have modest correspondence with external citations and little correspondence with download counts. The results show how a simple information rule shapes the volume, breadth and distribution of scientific attention in an artificial population.","authors":["Maxim Chupilkin"],"categories":["cs.AI","econ.GN","q-fin.EC"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.22408","pdf_url":"https://arxiv.org/pdf/2609.22408","source_feed":"econ.GN","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM仿真","社会影响","学术注意力"],"reason":"用1000个AI代理模拟学术注意力分配，与真实引用数据对照，属经济学实验场景。","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:04:54","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-22","rank":8,"question":"社会信息如何影响AI代理对学术论文的注意力分配？","design":"用GPT-5.6 Sol扮演1000个AI代理，分为独立选择组和社会影响组（各5个社区，每社区100个顺序代理），从2025年AER的114篇文章标题和摘要中选择要读的论文；社会影响组能看到本社区之前代理的累计选择次数，独立组看不到；测量选择数量、集中度、社区间差异等。","baseline":"外部引用次数和下载量作为对照，但仅用于相关性分析，并非严格的人类被试基准。","findings":"社会信息使代理平均少选17.2%的论文，集体覆盖从90篇降至73篇，选择更集中且社区间差异更大；随机赋予初始流行度使后续选择率提高45.55个百分点。","reliability":"论文未讨论","relevance":"该研究用LLM模拟人类在学术注意力市场中的从众行为，与真实引用数据对照，属于经济学实验场景，对关注LLM仿真可靠性和偏差的研究者有直接参考价值。","inspiration":"借鉴其Music Lab式设计，通过独立组与社会影响组对比、随机初始流行度干预来分离社会影响效应｜可迁移到金融信息传播或资产定价中的注意力分配问题，如投资者对研报或新闻的关注｜用LLM代理扮演投资者，处理为是否展示其他代理的阅读/选择行为，结果变量为选读的研报数量与集中度，对照真实市场中的研报点击或交易数据。"}},{"id":"2609.22225","version":1,"title":"Do LLMs Choose Like Humans? Using Cognitive Theory to Evaluate LLM Decision-Making","zh_title":"LLM像人类一样选择吗？用认知理论评估LLM决策","abstract":"Large language models (LLMs) exhibit a range of human-like decision-making behaviors, but whether these reflect similar underlying mechanisms or surface-level mimicry remains unclear. We evaluate whether LLM context sensitivity aligns with a cognitive economic theory that explains human behavior through problem categorization and attention allocation. Across 12 open-source and commercial LLMs on a novel 140,000-trial product choice benchmark, context induces human-like shifts in choice and problem categorization, but does not reliably reweight attention between features like price and quality. Neither scale nor chain-of-thought reasoning reliably attenuates context sensitivity or generates human-like behavior. These results suggest that LLM decision mechanisms are distinct from human ones.","authors":["Johnathan Sun","Andrei Shleifer","Yonatan Belinkov"],"categories":["cs.CL","cs.LG"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.22225","pdf_url":"https://arxiv.org/pdf/2609.22225","source_feed":"cs.CL","score":8,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM决策仿真","认知理论","人类对照"],"reason":"用LLM模拟人类决策并与人类数据对照，评估机制差异，直接相关且具批判性。","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:04:54","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-22","rank":12,"question":"LLM 的决策行为是否与人类认知经济理论中的情境敏感机制一致，即是否通过问题分类和注意力重加权来产生情境效应。","design":"用 12 个开源和商业 LLM 作为被试，在 140,000 次产品选择试验中，通过情境线索（享乐型 vs. 功能型消费）操纵问题分类，测量选择概率和特征敏感性（价格与质量的注意力权重）。","baseline":"人类基准来自认知经济学理论（Bordalo et al., 2026b）所解释的人类决策现象，包括情境对选择的影响、特征敏感性的不对称变化以及分类难度对情境效应的调节。","findings":"LLM 的选择变化与人类一致，但特征敏感性变化大多不一致；情境效应遵循问题分类过程，但模型规模和思维链推理不能可靠减弱情境敏感性。","reliability":"论文指出 LLM 的决策机制与人类不同，情境效应可能只是表面模仿；规模和推理不能可靠产生人类行为，说明在需要机制对齐的任务中仿真可能失效。","relevance":"该研究直接评估 LLM 作为人类决策仿真体的机制一致性，提供了与人类认知理论对照的批判性证据，值得精读以了解仿真在经济学决策中的局限。","inspiration":"借鉴其通过情境线索操纵问题分类并测量特征敏感性的设计，可迁移到消费者跨期选择或资产定价实验，用 LLM 模拟投资者在不同市场情境下的风险偏好，处理为情境线索（如牛市/熊市），结果变量为风险资产配置比例，对照真实投资者调查数据。"}},{"id":"2609.23403","version":1,"title":"Alignment and Divergence between Humans and AI in Interpersonal Privacy Decisions","zh_title":"人际隐私决策中人类与AI的一致性与分歧","abstract":"AI assistants increasingly mediate interpersonal communication on behalf of their primary user, but they risk violating the privacy expectations of third-party information owners. Resolving these tensions requires understanding how humans anticipate interpersonal privacy boundaries. Therefore, we conducted a dyadic study (N=76) and a matched evaluation of AI models across 18 information types and 3 recipient relationships. We found that data owners' privacy judgments are highly contextual and relationship dependent. While familiar data co-owners show meaningful alignment with owners' expectations, they significantly overestimate the need for permission. Interestingly, greater familiarity within the owner-co-owner dyad was associated with both higher disclosure acceptability and lower co-owner misalignment, whereas our exploratory four-item empathy measure was not. In contrast, AI models significantly underperform human co-owners in anticipating the data acceptability, even when provided with within-dyad examples. These findings underscore a core HCI design challenge to develop privacy-aware AI that respects multi-stakeholder information boundaries.","authors":["Hanxiang Zeng","Shuning Zhang","Xinyuan Zhou","Tianqi Song","Yuhan Yuan","Yuting Yang","Shuai Ma","Xin Yi"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.23403","pdf_url":"https://arxiv.org/pdf/2609.23403","source_feed":"cs.HC","score":8,"bucket":"selected","rubric_hits":["A1","B1","B4"],"tags":["LLM仿真","隐私决策","人机对齐"],"reason":"用LLM模拟人类隐私决策并与人类数据对照，评估AI与人类判断的偏差，属于仿真人…","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:00","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-22","rank":13,"question":"在人际隐私决策中，数据所有者对第三方披露的接受度如何受情境影响，人类共同所有者与所有者的判断对齐程度如何，AI模型能否准确预测所有者的隐私判断？","design":"本研究并非严格意义上的LLM仿真人类实验，而是采用配对设计：招募19对朋友和19对恋人（N=76），让数据所有者和共同所有者分别对18种信息类型、3种接收者关系下的披露可接受性、许可必要性等隐私维度进行评分；同时用多个LLM（零样本和单样本提示）、机器学习模型和微调模型对相同场景进行预测，并与人类共同所有者的预测准确度比较。","baseline":"人类基准是熟悉的数据共同所有者（朋友或恋人）对所有者隐私判断的预测，以平均绝对误差（MAE）和判断落在所有者评分±1分内的比例衡量对齐程度。","findings":"所有者的隐私判断高度依赖情境和关系，信息类型和接收者关系显著影响披露可接受性；熟悉的人类共同所有者与所有者有中等程度对齐（MAE≈1.179，70.7%判断在±1分内），但系统性高估许可必要性。AI模型（包括零样本和单样本LLM）显著不如人类共同所有者准确（MAE 1.648–1.796），单样本提示仅略微改善，仍无法弥合人机差距。","reliability":"论文承认AI模型在涉及微妙社会关系的披露场景中尤其表现不佳，单样本提示不足以缩小差距；机器学习模型和微调模型对未见个体泛化能力差，提供所有者特定示例的改进不一致。局限性包括样本量较小（19对朋友和19对恋人）、仅覆盖两种关系类型、同理心测量简短且探索性，未发现其与对齐的显著关联。","relevance":"该研究直接评估LLM在人际隐私决策中模拟人类判断的准确性，并与真实人类共同所有者的预测进行对照，揭示了AI在微妙社会情境下的系统性偏差，对关注LLM仿真可靠性及失效条件的研究者具有参考价值。","inspiration":"借鉴其配对设计和多模型基准测试方法，可系统评估LLM在预测他人偏好或决策时的准确性，并与人类代理预测对照。｜可迁移到经济金融中涉及代理决策或偏好预测的场景，如理财顾问预测客户风险偏好、信贷员判断借款人还款意愿、或政策制定者预测公众对经济政策的接受度。｜设计：招募真实客户-理财顾问对，让客户对一系列投资产品的风险承受能力和偏好进行评分，同时让顾问预测客户评分，并让LLM基于客户基本信息和少量示例进行预测；结果变量为预测误差（MAE）和方向一致性；以顾问预测为人类基准，比较LLM与顾问的准确性，并考察客户-顾问关系强度、信息敏感度等调节因素。"}},{"id":"2609.24859","version":1,"title":"Small-world Networks of Agents Brainstorm AI Risks to Support Ideation","zh_title":"智能体小世界网络头脑风暴AI风险以支持构思","abstract":"The ideation phase of participatory AI risk assessment often starts with a blank slate or a limited list of predefined risks, making it difficult to surface indirect or systemic harms. To address this limitation, we propose a three-stage ideation support tool. The tool complements participatory AI, rather than replacing it, and helps focus later engagement with affected communities. First, it dynamically discovers stakeholders depending on the given AI use and recursively expanding outward, allowing overlooked or indirect stakeholders to emerge. Second, it simulates these stakeholders with LLMs, connecting them into a network of a given topology, and having them ideate about risks. Third, it prioritizes risks using network centrality measures. In an initial evaluation, we found that betweenness centrality run through agents connected in a small-world network works best as it elevates risks raised by stakeholders who bridge disconnected groups, surfacing novel, systemic harms that traditional methods often miss. On an AI chatbot companion use case, this approach increased the novelty of the identified risks by approximately 1.1 points over single LLM brainstorming, and by 0.5 points over agentic LLM brainstorming, measured on a normalized five-point Likert scale, without reducing the plausibility or severity of the identified risks. To test whether our framework helps a human-led ideation session using the Futures Wheel approach, we divided 11 teams of non-western young chatbot users into two types: control (team) and treatment (team) in a participatory AI risk assessment. The control teams started from a list of risks generated by the 45 AI practitioners in the initial evaluation; the treatment teams started from a list generated by our framework. The treatment teams identified more risks overall, and more systemic, human-computer interaction, and environmental risks.","authors":["Ke Zhou","Edyta Bogucka","Daniele Quercia"],"categories":["cs.HC","cs.AI"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.24859","pdf_url":"https://arxiv.org/pdf/2609.24859","source_feed":"cs.HC","score":8,"bucket":"selected","rubric_hits":["A3","B1","B2","B4"],"tags":["LLM仿真","风险识别","参与式AI"],"reason":"用LLM模拟利益相关者头脑风暴AI风险，并与人类团队对照，涉及政策评估场景，但…","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:07","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-22","rank":16,"question":"如何利用小世界网络中的LLM智能体头脑风暴来支持AI风险识别的构思阶段，以发现更多新颖且系统性的风险？","design":"该研究提出一个三阶段框架：首先动态发现利益相关者，然后将其实例化为LLM智能体并连接成小世界网络进行风险头脑风暴，最后用网络中心性（特别是介数中心性）对风险进行排序。评估中比较了单LLM头脑风暴、多智能体基线以及该框架生成的风险在创新性、合理性和严重性上的表现，并进一步在人类主导的Futures Wheel工作坊中测试了该框架生成的风险列表对团队构思的促进作用。","baseline":"对照的真实人类数据包括：45名AI从业者在空白头脑风暴工作坊中生成的244条风险（用于评估风险生成质量），以及11个非西方年轻聊天机器人用户团队在Futures Wheel工作坊中的表现（控制组使用从业者生成的风险列表，处理组使用框架生成的风险列表）。","findings":"在AI聊天机器人伴侣用例中，小世界网络结合介数中心性的方法将风险创新性评分比单LLM头脑风暴提高了约1.1分，比智能体LLM头脑风暴提高了0.5分（5分制），且未降低合理性和严重性。在人类参与的Futures Wheel研究中，使用该框架生成的风险列表的处理组团队识别出了更多总体风险，以及更多系统性、人机交互和环境风险。","reliability":"论文未讨论","relevance":"该研究用LLM模拟利益相关者网络进行风险头脑风暴，并与真实人类团队进行对照，涉及AI政策评估场景，对关注LLM仿真可靠性和偏差的研究者具有参考价值。","inspiration":"该方法通过构建小世界网络并利用介数中心性来提升生成内容的多样性和新颖性，可借鉴其网络拓扑与中心性度量来设计多智能体仿真中的信息传播和观点聚合机制。｜可迁移到经济金融领域的政策评估或风险识别场景，例如金融系统性风险的早期预警或信贷审批中的歧视性风险识别。｜可设计一个研究：用LLM模拟银行、监管者、消费者等利益相关者，在小世界网络中讨论信贷审批算法的潜在风险，以生成的风险列表作为处理组，与人类专家生成的风险列表进行对照，比较两组在风险覆盖度、新颖性和系统性上的差异，并使用真实历史信贷数据或监管报告作为外部基准。"}},{"id":"2609.24629","version":1,"title":"Augmented Hypothesis Testing with Persona-Based LLM Simulations","zh_title":"基于角色LLM模拟的增强假设检验","abstract":"A/B testing requires large sample sizes, long timelines, and significant costs. When auxiliary predictions of experimental outcomes are available from machine learning models, uncertain prediction quality precludes replacing human experiments entirely, yet these predictions may still contain useful signal. We propose a principled framework for learning-augmented hypothesis testing that leverages predictions of unknown quality to reduce sample sizes while maintaining statistical validity. Predictions naturally vary in granularity, from coarse aggregate signals to fine-grained individual-level estimates, and our framework addresses both ends of this spectrum: (1) for population-level directional predictions, where only a binary signal on the treatment effect sign is available, we use an asymmetric test and prove consistency and robustness bounds within the learning-augmented algorithms paradigm; (2) for individual-level predictions, we introduce Generalized PPI++ (GPPI), extending Prediction-Powered Inference to handle nonlinear prediction errors through higher-dimensional transformations. Both methods benefit from accurate predictions while remaining robust to inaccurate or adversarial ones. We validate our framework using persona-based LLM simulations, where AI agents equipped with user personas predict individual behavior, as a natural prediction source spanning both granularity levels. Experiments on four real-world datasets demonstrate that our methods, combined with persona-based predictions, substantially reduce experimental costs while preserving rigorous statistical validity.","authors":["Ziyad Benomar","Aymen Al Marjani","Paul Missault","Saab Mansour"],"categories":["cs.LG","cs.AI","stat.AP"],"primary_category":"cs.LG","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.24629","pdf_url":"https://arxiv.org/pdf/2609.24629","source_feed":"cs.LG","score":8,"bucket":"selected","rubric_hits":["A1","A2","B1","B3"],"tags":["LLM仿真","假设检验","统计有效性"],"reason":"用LLM persona预测个体行为，与真实数据对照，并保证统计有效性，可迁移…","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:05","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-22","rank":15,"question":"如何利用 persona 驱动的 LLM 预测来减少 A/B 测试所需样本量，同时保持统计有效性？","design":"使用基于用户 persona 的 LLM 模拟来预测个体行为，作为辅助预测源；针对群体级方向预测（仅提供处理效应符号）和个体级预测分别提出不对称检验和广义 PPI++ 方法，在四个真实数据集上验证。","baseline":"四个真实世界数据集（具体名称未在节选中列出）中的人类 A/B 测试结果作为对照。","findings":"所提出的学习增强假设检验框架能够利用 persona 预测显著减少实验成本，同时保持严格的统计有效性。群体级方向预测和个体级预测方法均能在预测准确时受益，并在预测不准确或对抗性时保持稳健。","reliability":"论文承认完全用 persona 预测替代人类实验面临根本挑战：LLM 的黑箱性质、对提示工程的敏感性、分布偏移以及人类行为的复杂性使得形式化质量保证困难；此外，persona 与真实用户之间可靠的一对一映射通常不可行，导致预测粒度受限。","relevance":"该研究直接针对研究者关注的 LLM 人类仿真实验，利用 persona 预测个体行为并与真实数据对照，同时提供统计有效性保证，值得精读原文以了解具体方法和实验细节。","inspiration":"借鉴其将 LLM 预测作为辅助信号而非替代品，并通过统计校正保持有效性的思路｜可迁移到政策评估中的 A/B 测试，如利用 LLM 模拟消费者对价格变动的反应来减少实地实验样本量｜以消费者信贷决策为场景，用 LLM 基于用户 persona 预测个体对贷款条款的反应，处理为不同利率或审批条件，结果变量为接受/拒绝或违约行为，对照真实信贷申请数据。"}},{"id":"2509.20634","version":3,"title":"Recidivism Prediction, Peer Effect Estimation, and Prediction-Powered Inference with LLM Text Measures","zh_title":"使用LLM文本测量进行累犯预测、同伴效应估计与预测驱动推断","abstract":"We provide a new framework for estimating peer effects when outcomes are multivariate behavioral measures derived from written text using an LLM and the network formation is endogenous. We obtain LLM embeddings and zero shot classification of more than 200,000 written exchanges among residents of low-security correctional facilities. We find that LLM embeddings improve out-of-sample recidivism prediction by up to 30% over pre-entry covariates alone using LASSO and LoRA fine-tuning, showing that text representations capture meaningful signals. For peer effect estimation, we develop a novel instrumental variable estimator that accommodates multivariate outcomes, sparse networks, and multidimensional latent homophily. We show that this estimator is $\\sqrt{N}$-consistent and asymptotically normal under sparsity conditions that relax dense-network assumptions prevalent in the peer effect literature. Limited human annotations are then combined with LLM zero-shot vectors in a new prediction-powered peer inference (PPPI) approach to obtain de-biased estimates and valid inference. Results reveal significant peer effects in the behavioral profiles.","authors":["Shanjukta Nath","Jiwon Hong","Jae Ho Chang","Keith Warren","Subhadeep Paul"],"categories":["econ.EM","cs.AI","econ.GN","q-fin.EC","stat.ME"],"primary_category":"econ.EM","announce_type":"replace-cross","date":"2026-09-22","first_seen":"2025-09-25","revised_at":"2026-09-22","abs_url":"https://arxiv.org/abs/2509.20634","pdf_url":"https://arxiv.org/pdf/2509.20634","source_feed":"econ.GN","score":7,"bucket":"pending","rubric_hits":["A1","B1","B2","B3"],"tags":["LLM文本测量","同伴效应","预测驱动推断"],"reason":"用LLM从文本中提取行为测量并估计同伴效应，有真实人类数据对照，涉及经济学场景…","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:26","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-22","rank":17,"question":"如何利用LLM从文本中提取行为测量，在存在内生网络形成的情况下估计同伴效应，并改进再犯预测？","design":"本研究并非用LLM模拟人类被试，而是用LLM处理真实人类文本数据：对超过20万条低安全级别惩教设施中居民之间的书面交流进行嵌入和零样本分类，得到居民行为画像，然后开发新的工具变量估计器估计同伴效应，并结合预测驱动推断（PPPI）进行去偏估计。","baseline":"真实人类数据：来自治疗社区（TCs）的居民书面交流记录、入狱前协变量、三年再犯记录，以及少量人工标注用于验证LLM输出。","findings":"LLM嵌入在再犯预测中比仅用入狱前协变量提升最多30%的样本外预测准确率；新估计器在稀疏网络下具有√N一致性和渐近正态性，且估计结果显示行为画像中存在显著的同伴效应。","reliability":"论文未讨论","relevance":"该研究使用LLM从真实文本中提取行为测量并估计同伴效应，属于用LLM辅助分析人类行为而非替代人类被试，但涉及真实人类数据对照和经济学场景，对关注LLM在实证研究中应用的研究者有参考价值。","inspiration":"借鉴其利用LLM零样本分类将高维文本嵌入降维为可解释行为标签，并结合工具变量处理内生网络的方法。｜可迁移到金融文本分析，如利用分析师报告或公司公告文本测量管理层情绪或风险偏好，进而估计同行公司之间的情绪传染效应。｜设计：以分析师为被试，处理为同行分析师的乐观情绪文本，结果变量为分析师自身预测偏差，用真实分析师历史预测数据作为对照，采用类似工具变量策略识别同行效应。"}},{"id":"2608.28021","version":2,"title":"Compared to What? A Human-Anchored Security Benchmark for LLM-Generated Infrastructure-as-Code","zh_title":"与什么相比？面向LLM生成基础设施即代码的以人为锚安全基准","abstract":"Large language models increasingly author Infrastructure-as-Code (IaC), where one insecure default is provisioned straight into production. Prior evaluations report vulnerability counts for models only, and so cannot say whether models are worse than the engineers they assist. We present GenIaC-SecBench: 100 deployment scenarios across 12 model configurations from six vendors, open and closed weights, yielding 1,196 artifacts scanned by three policy engines (Checkov, Trivy, KICS) at complete coverage. Crucially we scan 634 human-authored IaC templates with the identical toolchain, giving the first size-matched human security baseline for this task. Vulnerability density is strongly inverse to artifact size (Spearman $\\rho=-0.55$, $p<10^{-77}$), so unmatched comparisons measure size, not security. Size-matched, every configuration exceeds the human baseline at $3.21\\times$ to $3.87\\times$, and the gap widens as tasks get simpler ($4.9\\times$ at one resource, $1.4\\times$ at twenty or more). A majority of scenarios prescribe a security state rather than specifying function alone, so we stratify by prompt class: pooled the gap is $3.50\\times$, and excluding every scenario that explicitly requests an insecure configuration still leaves all configurations above baseline ($2.4\\times$ to $4.2\\times$). The corpus cannot isolate unprompted default posture, and we say so. Decomposing \"reasoning\" into standard generation, prompted chain-of-thought, and vendor extended-thinking APIs, extended thinking beats prompted CoT ($-12.0\\%$, $p=0.0013$) while prompted CoT alone is indistinguishable from standard ($-1.3\\%$, n.s.); it consumes under $1\\%$ of the output budget, bounding the effect. Two negative results: more deployable models are not more vulnerable ($r=0.158$, $p=0.625$), and complete-case Friedman is uncomputable here, motivating Skillings-Mack. All code and data are released.","authors":["Animesh Shaw"],"categories":["cs.CR","cs.AI","cs.MA","cs.SE"],"primary_category":"cs.CR","announce_type":"replace-cross","date":"2026-09-22","first_seen":"2026-08-31","revised_at":"2026-09-22","abs_url":"https://arxiv.org/abs/2608.28021","pdf_url":"https://arxiv.org/pdf/2608.28021","source_feed":"cs.MA","score":7,"bucket":"pending","rubric_hits":["A2","B1","B4"],"tags":["LLM安全评估","人类基线对照","基准测试"],"reason":"用人类基线对照评估LLM生成IaC的安全性，方法可迁移到仿真可靠性评估","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:29","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-22","rank":18,"question":"LLM生成的IaC代码与人类工程师相比，安全漏洞密度如何？","design":"用12种LLM配置（9个模型、6家厂商，含开源与闭源）生成100个部署场景的IaC代码，共1196个工件，用Checkov、Trivy、KICS三个策略引擎扫描漏洞，并与634个人类编写的IaC模板在相同工具链下对比。","baseline":"634个人类编写的IaC模板，用相同的三个策略引擎扫描，并按资源数量分层匹配。","findings":"漏洞密度与工件大小强负相关（Spearman ρ=-0.55），未匹配的比较衡量的是大小而非安全性。大小匹配后，所有LLM配置的漏洞密度均为人类基线的3.21-3.87倍，且任务越简单差距越大。","reliability":"论文承认无法隔离未提示的默认安全姿态，因为多数场景明确规定了安全状态；扩展思维API使用默认温度，与标准生成存在混杂；完整案例Friedman检验不可计算，改用Skillings-Mack。","relevance":"该研究提供了人类基准对照的严谨方法，可用于评估LLM在安全相关任务中的仿真可靠性，对关注LLM仿真偏差的研究者有参考价值。","inspiration":"借鉴其分层匹配和工具链一致性的对照设计，避免因输出长度等混杂因素导致错误结论。｜可迁移到经济金融中的代码生成任务，如量化交易策略代码、金融数据处理脚本的安全性评估。｜用LLM生成金融分析代码，与人类分析师编写的代码对比漏洞密度，按代码行数或功能复杂度分层，使用静态分析工具扫描，以真实人类代码库为基准。"}},{"id":"2609.07358","version":3,"title":"Access to Live AI Advice and Behavior Under Risk: An Incentivized Experiment","zh_title":"获取实时AI建议与风险下的行为：一项激励实验","abstract":"Generative AI has become an everyday advisor, and the systems people consult are live and interactive, not pre-scripted. We ask whether access to such a system changes behavior under risk. In an incentivized experiment (N = 158), participants made lottery choices with an optional decision aid presented as a conventional pre-written tool, a live one-shot AI, or a live interactive AI they could query, with information format held equivalent across conditions. Risk preferences are elicited via DOSE. We find no evidence that access to a live AI advisor changes risk aversion.","authors":["Paul Althaus","Leon Houf","Christiane Schwieren"],"categories":["econ.GN","econ.TH","q-fin.EC"],"primary_category":"econ.GN","announce_type":"replace","date":"2026-09-22","first_seen":"2026-09-09","revised_at":"2026-09-22","abs_url":"https://arxiv.org/abs/2609.07358","pdf_url":"https://arxiv.org/pdf/2609.07358","source_feed":"econ.GN","score":7,"bucket":"pending","rubric_hits":["A1","B1","B2"],"tags":["LLM辅助决策","风险偏好","实验经济学"],"reason":"用LLM作为决策辅助，测量人类风险行为变化，有真实人类实验对照，但非仿真替代。","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:26","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-22","rank":19,"question":"在风险决策中，获得实时AI建议（一次性或可交互）是否会改变人们的风险厌恶程度？","design":"该研究不是用LLM仿真人类，而是将LLM作为决策辅助工具。158名被试在线完成激励性彩票选择任务，随机分为三组：对照组使用预编程的决策辅助工具（自动计算期望值并给出风险偏好提示），一次性AI组使用实时AI（gpt-5.4-mini）提供同样格式的建议，交互式AI组可额外追问两次。风险偏好通过DOSE方法估计CRRA参数。","baseline":"对照组为使用预编程决策辅助工具的人类被试，其风险厌恶参数作为基准。","findings":"与预编程工具相比，获得一次性或交互式AI建议对风险厌恶没有显著影响，点估计接近零。该零结果在咨询过AI、高信任AI的子样本中依然稳健，且设计能检测到与随机顺序效应相当的影响。","reliability":"论文承认只能排除大于0.15的效应，更小的效应仍可能存在；且研究仅识别了在信息格式相同的情况下，AI作为来源的净效应，并未声称AI提供的信息本身没有影响。","relevance":"该研究虽非仿真替代，但直接检验了LLM作为实时建议对经济决策行为的影响，与您关注的AI与人类行为交互、实验经济学场景高度相关，值得阅读原文了解实验细节和稳健性检验。","inspiration":"值得借鉴的是将AI建议的呈现方式（预编程 vs 实时AI）作为处理变量，同时保持信息内容一致，从而分离出“AI来源”的净效应，并通过随机顺序效应验证检验力。｜可迁移到资产定价实验，检验投资者在获得AI投资建议后风险资产配置是否改变。｜招募真实投资者为被试，随机分配使用传统财务计算器或实时AI助手获取相同信息，测量其风险资产配置比例，并与历史投资数据或对照组行为进行对比。"}},{"id":"2609.22188","version":1,"title":"Fairness Beyond Anonymization? Demographic Leakage in German LLM-Generated Resumes","zh_title":"匿名化之外的公平？德国LLM生成简历中的人口统计泄漏","abstract":"Large language models (LLMs) are increasingly integrated into AI-assisted hiring pipelines, including automated resume generation and screening. Under the EU AI Act, the hiring domain is classified as high-risk, making fairness and transparency critical requirements. Existing work has primarily focused on explicit hiring decisions, while less attention has been paid to whether generated resumes themselves encode recoverable demographic information. In this work, we conduct a two-stage audit of demographic leakage in German-language LLM-generated resumes. First, we use ChatGPT (GPT-4o-mini), Gemini 2.5 Flash-Lite, and multiple scales of the open-weight Qwen 3 model family (4B, 8B, and 14B) to generate resumes from real anonymized job-matching profiles, systematically varying gender- and ethnicity-associated names while holding qualifications constant. Second, we simulate a downstream resume screening scenario, where the generated resumes are first anonymized and gender-neutralized, before demographic leakage classifiers are trained on the resulting texts. We find that, despite these interventions, classifiers reliably distinguish between resumes generated with male and female names. This leakage is not driven by overtly gendered wording, but by subtle differences in the usage of semantically equivalent, formally gender-neutral terms in German. In contrast, ethnicity-related leakage remains comparatively weak across models. Our findings demonstrate that apparently neutral resume generation can still preserve highly predictive demographic signals, raising concerns about anonymization-based fairness interventions in multilingual AI hiring pipelines.","authors":["Charlotte Leininger","Helena Veit","Matthias A{\\ss}enmacher","Andreas Bender"],"categories":["cs.CL","cs.LG"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.22188","pdf_url":"https://arxiv.org/pdf/2609.22188","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A2","B1","B4"],"tags":["LLM生成简历","人口统计泄漏","公平性审计"],"reason":"审计LLM生成简历中的人口统计泄漏，评估匿名化公平干预的失效，有真实数据对照，…","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:04:52","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-22","rank":22,"question":"LLM生成的德语简历在匿名化和性别中性化处理后，是否仍包含可恢复的性别和种族人口统计信息？","design":"使用ChatGPT、Gemini和Qwen 3系列模型，基于真实匿名求职匹配档案生成简历，系统变化姓名以关联性别和种族，保持资格不变；随后对生成简历进行匿名化和性别中性化处理，训练分类器检测人口统计泄漏。","baseline":"真实匿名求职匹配档案（来自德国求职匹配公司Chemistree GmbH），作为生成简历的资格基础，但未直接提供人类简历文本作为对照。","findings":"尽管进行了匿名化和性别中性化，分类器仍能可靠区分男性和女性姓名生成的简历，泄漏源于德语中语义等价但形式性别中立的术语使用差异；种族相关泄漏相对较弱。","reliability":"论文未讨论","relevance":"该研究通过审计LLM生成简历中的人口统计泄漏，评估匿名化公平干预的失效，有真实数据对照，对关注LLM仿真可靠性与偏差的研究者具有参考价值，值得阅读原文。","inspiration":"借鉴其控制变量法：保持资格不变，仅改变姓名关联的人口统计属性，并训练分类器检测泄漏，可迁移到信贷审批歧视研究；用LLM生成贷款申请文本，变化姓名暗示种族或性别，训练分类器预测人口统计属性，并与真实贷款申请数据对照。"}},{"id":"2609.22633","version":1,"title":"Beetle: A Bilingual Model Suite for Modelling Second-Language Processing","zh_title":"Beetle：用于建模第二语言处理的双语模型套件","abstract":"Bilingual language models (LMs) offer a controlled setting for studying how training conditions shape second-language (L2) behaviour, but prior work typically varies exposure structure, scale, and architecture at once, making it difficult to attribute effects to any single factor. We introduce Beetle, a controlled language model pretraining framework in which tokeniser, target language, training budget, and exposure structure are each independently manipulable, enabling systematic and comparable experimentation of training conditions. Using Beetle, we train and release 285 bilingual and 45 monolingual open-source LMs with rich checkpoints across a range of exposure schedules, data scales and first languages (L1s) to study multilingual pretraining and computational modelling of bilingualism and second language learning. Evaluating models on human bilingual and second language reading-time prediction and grammaticality judgement tasks, we find that staged and temporally structured curricula consistently improve alignment with language learner reading time compared to balanced bilingual training, with the largest gains at smaller data scales and for typologically closer language pairs. The Beetle models are well suited tools to help move computational psycholinguistics beyond its prevailing monolingual, English-centric focus toward models of human bilingual processing, to study cross-lingual learning dynamics, while supporting community-based development of controlled model families.","authors":["Suchir Salhan","Catherine Arnett","James Michaelov","Paula Buttery"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.22633","pdf_url":"https://arxiv.org/pdf/2609.22633","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A1","B1"],"tags":["计算心理语言学","双语模型","人类行为预测"],"reason":"用双语模型预测人类二语阅读时间和语法判断，有真实人类数据对照，属于LLM仿真人…","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:04:57","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-22","rank":23,"question":"在双语语言模型预训练中，语言暴露的时间结构（如分阶段课程与平衡双语训练）如何影响模型对人类二语阅读时间和语法判断行为的对齐程度？","design":"使用 Beetle 框架训练 285 个双语模型（125M 参数，固定架构和分词器），以英语为 L2，变化 21 种 L1、三种数据规模（100M、2B、24B tokens）和五种暴露条件（包括分阶段课程、平衡双语等），并训练 45 个单语基线模型。评估模型在人类二语阅读时间预测（MECO 眼动数据）和语法判断任务上的表现。","baseline":"人类基准来自 Multilingual Eye-movement Corpus (MECO) 的二语阅读时间数据，以及 CEFR 分层的二语学习者错误判别任务（BLiSS）。","findings":"分阶段和时序结构课程比平衡双语训练更一致地改善了对语言学习者阅读时间的对齐，尤其在较小数据规模和类型学相近的语言对中收益最大。但课程效应在不同任务上并不统一：在语法性和错误判别任务上，平衡暴露往往具有竞争力或更好。","reliability":"论文指出课程效应不会均匀迁移到所有任务：在语法性和错误判别任务上，平衡暴露可能更好；此外，模型对齐可能受数据规模、语言类型学距离等因素影响，但未系统讨论失效条件。","relevance":"该研究用双语模型模拟人类二语处理，有真实人类眼动和语法判断数据作为对照，属于 LLM 仿真人类认知行为的研究，且提供了大规模受控模型套件，对关注仿真可靠性和偏差的研究者有参考价值。","inspiration":"借鉴其受控训练框架：通过独立操纵暴露结构、数据规模和语言对，并设置匹配的单语基线，可精确归因训练条件对行为的影响。｜可迁移到经济金融中的跨期选择或风险偏好实验，例如研究不同信息呈现顺序（如先呈现收益后呈现风险）如何影响决策。｜用 LLM 作为被试，施加不同的信息暴露课程（如分阶段呈现历史价格与基本面信息），测量其投资决策或风险偏好，并与真实人类实验数据（如实验室资产定价实验）对照，检验仿真一致性。"}},{"id":"2609.22971","version":1,"title":"Automatic multimodal UX improvement recommendations from LLM agent user simulations","zh_title":"基于LLM智能体用户仿真的自动多模态用户体验改进建议","abstract":"Evaluating user experience (UX) on live websites through user testing is expensive, subjective, and difficult to scale. LLM agents offer a promising route to automating UX testing by simulating realistic user behaviour. However, existing simulation approaches typically lack multimodality and require time-consuming manual review to extract actionable insights. We formalise UX improvement recommendation from simulation data as a structured natural language generation and ranking problem, and establish an evaluation protocol using expert annotation and LLM-as-a-Judge. We present AMUSER, a multimodal framework which simulates user behaviour and automatically generates prioritised UX improvement recommendations from resulting data. We evaluate AMUSER on commercial websites and show that its recommendations substantially outperform those from text-only simulation (NDCG@3 = 0.758 versus 0.359) at an 89% lower simulation cost. Our results suggest an asymmetric role of multimodality: visual access during simulation improves recommendations through richer traces, while providing visual inputs during recommendation generation can modestly degrade quality. We also discuss practical deployment lessons from applying AMUSER to commercial websites.","authors":["Anu Chowdhury","Bin Wu","Hossein A. Rahmani","Emine Yilmaz"],"categories":["cs.CL","cs.AI","cs.HC"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.22971","pdf_url":"https://arxiv.org/pdf/2609.22971","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A1","B1","B2"],"tags":["LLM仿真","用户体验","多模态"],"reason":"用LLM模拟用户行为并生成UX建议，有真实用户数据对照，属人类仿真但场景偏应用。","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:04:59","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-22","rank":25,"question":"能否从LLM智能体模拟用户行为的数据中自动生成有用的UX改进建议，且多模态模拟是否优于纯文本模拟？","design":"使用多模态LLM智能体（AMUSER）模拟不同用户画像在真实网站上的任务执行过程，记录动作、想法、情绪和截图，随后由LLM根据交互轨迹生成并排序UX改进建议。","baseline":"无对照","findings":"多模态模拟生成的建议显著优于纯文本模拟（NDCG@3=0.758 vs 0.359）和静态分析（0.273），且模拟成本降低89%。视觉输入在模拟阶段有益，但在建议生成阶段提供视觉输入会略微降低质量。","reliability":"论文未讨论","relevance":"该研究展示了LLM模拟用户行为并自动提取可操作见解的完整流程，与您关注的仿真可靠性和自动化评估高度相关，但缺少真实人类数据对照，且场景偏应用而非经济学实验。","inspiration":"借鉴其多模态模拟与自动生成结构化建议的流程，可用于经济金融场景中模拟消费者或投资者行为并提取决策模式。｜可迁移到消费者在线购物决策、金融产品选择或投资者信息处理等场景。｜设计：用LLM智能体模拟不同风险偏好的投资者在金融网站上的信息搜索与决策过程，处理为是否提供视觉信息（如图表），结果变量为决策质量或信息获取效率，并与真实投资者实验数据对照。"}},{"id":"2609.23039","version":1,"title":"Auditing Political Alignment in LLM Assistants: Engagement, Stance, and User Identity","zh_title":"审计LLM助手中的政治对齐：参与、立场与用户身份","abstract":"LLM-based AI systems answer political questions for hundreds of millions of people. Current audits measure what they say to an average user, but their behavior is dynamic. I argue that their political behavior is a set of policies over whom to answer, what to say, and whether to engage at all, conditional on the topic and what the system knows about the user. I call these policies the system's speech regime, which is how a developer settles the tradeoff between answering, accommodating the user, and refusing, each of which carries a cost that varies by topic. I derive a typology of five regimes from two dimensions, engagement and stance. I test six AI systems (OpenAI, Anthropic, xAI, Google, Mistral, DeepSeek) in a preregistered experiment of 7,500 multi-turn conversations that randomly assign the user's political identity across five topics: abortion, Catalan independence, climate change, Nazism, and a zero-stakes control (pineapple on pizza). Two LLM judges from different developers score every answer, validated against human coding, and refusal is treated as an outcome rather than missing data. Every system accommodates the user on the control topic, showing that political restraint is a policy. On contested topics the systems fall into different regimes: on abortion, GPT engages and mirrors every user, Gemma refuses everyone, Claude answers strongly conservative users 35 percent of the time and almost no one else, and Grok accommodates conservatives only. On settled topics such as climate change and Nazism, five systems hold firm for every user. The systems also infer the user's overall ideology, so accommodation can spill over to topics not yet discussed. A comparison of two Grok releases shows the regime changing between versions in a way current audits miss. Speech regimes matter for alignment research and for polarization, political knowledge, and the quality of democracy.","authors":["Joan C. Timoneda"],"categories":["cs.CL","cs.AI","cs.CY"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.23039","pdf_url":"https://arxiv.org/pdf/2609.23039","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A1","B1","B4"],"tags":["LLM仿真","政治行为","审计"],"reason":"用LLM模拟不同政治身份用户的回答，并与人类编码对照，评估系统行为差异，可迁移…","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:00","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-22","rank":26,"question":"AI助手面对不同政治身份的用户时，在政治话题上的回答行为（是否回答、立场、是否迎合用户）如何随话题和用户身份变化？","design":"审计实验：用脚本化用户（随机分配政治身份）与六个商业LLM助手进行7500段多轮对话，涉及五个话题（堕胎、加泰罗尼亚独立、气候变化、纳粹、披萨放菠萝作为零风险对照），测量系统是否回答、立场、以及是否迎合用户，并用两个不同开发者的LLM法官评分（经人类编码验证）。","baseline":"人类编码验证LLM法官评分，但无大规模真实人类行为数据作为对照。","findings":"系统在争议话题上表现出不同“言论体制”：GPT迎合所有用户，Gemma几乎拒绝所有人，Claude选择性拒绝（更多回应保守派但不迎合），Grok只迎合保守派；在共识话题上五个系统立场坚定。系统还会推断用户整体意识形态，导致迎合溢出到未讨论的话题。","reliability":"论文未明确讨论失效条件，但指出当前审计方法（测量对“平均用户”的回答）会错过版本间言论体制的变化，且未覆盖所有可能话题和用户身份。","relevance":"该研究展示了如何用LLM模拟不同身份用户来审计系统行为差异，并强调将拒绝视为结果而非缺失数据，对用LLM进行人类仿真实验的方法论有借鉴意义，值得阅读原文。","inspiration":"借鉴其将“拒绝回答”作为结果变量、随机分配用户身份、多轮对话施加压力的设计，以及用LLM法官加人类验证的测量方式。｜可迁移到信贷审批歧视研究：让LLM扮演不同种族、性别、收入水平的贷款申请人，向AI信贷助手申请贷款，观察是否被拒绝、获批额度、利率差异。｜设计：用LLM生成标准化贷款申请对话，随机分配申请人特征（种族、性别、收入），让真实银行AI客服或LLM模拟的信贷员处理，结果变量为是否批准、额度、利率，与真实信贷审批数据（如HMDA数据）对照，检验AI决策中的歧视。"}},{"id":"2609.23936","version":1,"title":"Think Before You Accept: Can Written Justification Reduce Uncritical Uptake of AI Writing Suggestions?","zh_title":"接受前先思考：书面理由能否减少对AI写作建议的不加批判采纳？","abstract":"Generative AI can offer students useful feedback, but its value depends on judging which suggestions are accurate and relevant. Prior research shows that strategic friction during human-AI interactions can promote critical uptake, but how to effectively implement such friction in academic contexts remains unclear. We examine whether requiring students to justify decisions to accept or reject AI suggestions can mitigate uncritical uptake in academic writing. In a randomized experiment embedded in a course activity (N=129), students wrote a data analysis proposal, received mixed-quality AI revision suggestions, and decided whether to accept or reject them. Students required to provide written justifications were 24 percentage points less likely to adopt flawed suggestions (65% vs. 41%), with no reduction in acceptance of sound suggestions (81% vs. 86%). However, thematic analysis revealed superficial engagement in the justification task and gaps in metacognitive monitoring and domain knowledge.","authors":["Yan Tao","Jennifer Meyer","Rene F. Kizilcec"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.23936","pdf_url":"https://arxiv.org/pdf/2609.23936","source_feed":"cs.HC","score":7,"bucket":"pending","rubric_hits":["A1","B1","B4"],"tags":["人机交互","AI建议采纳","批判性思维"],"reason":"用LLM生成写作建议，人类被试在真实任务中决策，有对照实验，但非仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:01","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-22","rank":27,"question":"在学术写作中，要求学生为接受或拒绝AI修改建议提供书面理由，能否减少对AI建议的不加批判的采纳？","design":"本研究不是LLM仿真人类被试的研究，而是随机对照实验：129名学生在课程活动中撰写数据分析提案，收到混合质量的AI修改建议，并被随机分配是否必须为每个采纳/拒绝决定提供书面理由；结果变量为对高质量和低质量建议的接受率。","baseline":"无对照（无真实人类数据作为仿真基准；实验本身以学生真实行为为数据，但非仿真研究）","findings":"要求书面理由使低质量建议的采纳率从65%降至41%，而高质量建议的采纳率无显著下降（81% vs 86%）。但主题分析显示学生在理由中参与肤浅，存在元认知监控和领域知识缺口。","reliability":"论文未讨论仿真可靠性，但承认书面理由干预未能完全消除不加批判的采纳，且学生参与度有限，提示干预效果可能受限于学生的元认知和领域知识。","relevance":"该研究虽非LLM仿真人类被试，但涉及人类对AI建议的决策行为，且通过随机实验揭示了干预对决策质量的影响，对关注AI辅助决策中人类行为的偏差与干预的研究者有参考价值。","inspiration":"借鉴其随机实验设计，通过施加认知摩擦（如要求书面理由）来改变决策行为，并测量对不同质量信息的区分度。｜可迁移到经济金融中的AI辅助决策场景，如投资者对AI投资建议的采纳、信贷审批中AI风险评分的接受、或消费者对AI推荐产品的选择。｜设计一个实验：招募真实投资者作为被试，提供AI生成的股票推荐（包含高质量和低质量信号），随机要求部分被试为每个采纳/拒绝决定写理由，结果变量为对高质量和低质量推荐的采纳率，并以历史市场数据或专家评级作为建议质量的基准。"}},{"id":"2609.24532","version":1,"title":"Prompting Against Persona Drift: Comparing Intervention Timing and Content in LLM-Simulated Conversations","zh_title":"对抗角色漂移的提示策略：比较LLM模拟对话中的干预时机与内容","abstract":"Simulating student personas with large language models (LLMs) enables scalable evaluation of educational systems. However, behavioral drift, a progressive decline in persona consistency, can emerge over extended conversations, limiting the validity of such simulations. We evaluate five prompt-level mechanisms using separate monitoring and intervention pipelines. Across 1,200 28-turn conversations spanning four LLMs and two ADHD persona intensities, we varied when to intervene (static vs. adaptive) and what to inject (reinjection vs. reflective reminder), plus a novel adaptive condition in which a monitor generates behavior-specific instructions. Relative to no intervention, reinjection reduced the modeled rate of LLM-rated drift by 35--38\\%, reflective reminders by 22--27\\%, and behavior-specific instruction by 87\\%. None eliminated drift. We found no evidence that adaptive timing outperformed static scheduling. Monitoring therefore appears more useful for deciding \\textit{what} to correct than \\textit{when} to intervene, although behavior-specific instruction requires component-level testing.","authors":["Nicolas Leins","Jennifer Haase","Varvara Geronimus","Jana Gonnermann-M\\\"uller","Sebastian Pokutta"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.24532","pdf_url":"https://arxiv.org/pdf/2609.24532","source_feed":"cs.HC","score":7,"bucket":"pending","rubric_hits":["A1","A2","B4"],"tags":["LLM仿真","角色一致性","教育评估"],"reason":"用LLM模拟学生角色并评估一致性，属于人类仿真，但无真实人类数据对照，且聚焦角…","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:04","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-22","rank":29,"question":"在LLM模拟学生角色的长对话中，提示级干预机制能否缓解角色漂移，且自适应干预是否优于静态干预？","design":"用四个LLM模拟ADHD学生角色（两种强度），在1200段28轮对话中，通过分离的监测与干预管道，比较五种提示级干预条件（无干预、静态全角色重注入、静态反思提醒、自适应全角色重注入、自适应反思提醒、自适应行为特定指令），以LLM评定的行为量表得分变化作为结果变量。","baseline":"无对照","findings":"所有干预条件均显著减缓了角色漂移，但未能完全消除；自适应时机相比静态调度无一致优势，而行为特定指令的干预效果最强，漂移减少87%。","reliability":"论文未讨论","relevance":"该研究属于LLM人类仿真，但无真实人类数据对照，且聚焦于教育场景的角色一致性，与研究者关注的经济学实验和政策评估场景关联有限，但方法上对仿真可靠性评估有参考价值。","inspiration":"借鉴其分离监测与干预的双管道设计，以及基于行为测量的自适应干预思路，可用于经济金融仿真中的行为一致性维护。｜可迁移到消费者跨期选择实验或投资者风险偏好模拟中，防止LLM在长对话中偏离预设的经济决策模式。｜以LLM模拟不同风险偏好的投资者，在连续投资决策对话中施加行为特定指令干预，结果变量为风险选择序列，并与真实投资者实验数据对照。"}},{"id":"2609.24146","version":1,"title":"Mind or Message? Auditing Theory of Mind in Multi-Agent Social Simulation","zh_title":"心智还是信息？审计多智能体社会模拟中的心智理论","abstract":"Language model agents are increasingly used to simulate social interaction, and the resulting transcripts read as though the agents understand one another. We ask whether that appearance rests on a model of the partner's mind or on the surface record of what the partner said. We build a social simulation in which both questions have exact answers: 40 multi-issue negotiations whose hidden preference weights and whose full Pareto frontier are known by construction. Two model families negotiate across 160 dyads, every transcript is frozen before any measurement, and 2880 counterfactual probes then hold the evidence byte identical while moving one factor at a time: the reader's own stake, the partner's tone, an identity label, and the order of recursion. The agents are socially fluent and economically poor. They reach agreement in 96.2% of dyads with 0 protocol failures, yet only 0.7% of deals land on the Pareto frontier, they leave 20.5% of the available joint value unclaimed, and they miss the one issue on which their interests are perfectly aligned in 76.6% of deals; on the frontier and on that aligned issue, a package drawn at random from the set both sides would accept does as well. The probes locate the failure. Swapping only the reader's own payoff sheet, while the partner's words and offers stay identical, moves the inferred top priority by 15.0 percentage points, which is egocentric projection rather than inference, while a tone rewrite moves it by 5.3 percentage points and an identity label by 0.0. Most tellingly, an agent predicts what its partner believes about it 72.5% of the time while that partner's belief is itself correct only 51.2% of the time: the agents track the conversation far better than they track the mind behind it.","authors":["Cong Li","Cheng Chen","Thomas Fung","Alex Rossi","Yi Li"],"categories":["cs.LG"],"primary_category":"cs.LG","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.24146","pdf_url":"https://arxiv.org/pdf/2609.24146","source_feed":"cs.LG","score":7,"bucket":"pending","rubric_hits":["A1","A2","B4"],"tags":["LLM仿真","心智理论","多智能体谈判"],"reason":"用LLM模拟谈判并审计心智理论，虽无人类对照，但批判性评估仿真失效条件，方法可…","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:04","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-22","rank":28,"question":"语言模型代理在模拟社会互动时，其表现是基于对对方心理状态的建模，还是仅基于对对话表面记录的读取？","design":"使用两个模型家族在40个多议题谈判场景中组成160个二人组进行谈判，每个场景的偏好权重和帕累托前沿已知。谈判后冻结所有对话记录，然后施加2880个反事实探针，每次仅改变一个因素（读者自身收益表、对方语气、身份标签、递归顺序），测量推断出的对方首要优先级的变化。","baseline":"无对照","findings":"代理在社交上流利（96.2%达成协议，0协议失败），但在经济上表现差：仅0.7%的协议达到帕累托前沿，损失20.5%的联合价值，76.6%的交易错过双方利益完全一致的议题。探针显示，改变读者自身收益表使推断的对方首要优先级移动15.0个百分点，表明自我投射而非推断；改变语气移动5.3个百分点，身份标签无影响。代理预测对方信念的准确率为72.5%，而对方信念本身正确率仅51.2%，说明代理更擅长追踪对话而非对方心理。","reliability":"论文承认其审计协议是诊断性的，不主张社会仿真无效，但强调有效性需要潜在真实基准来验证。局限包括仅使用两个模型家族、特定谈判场景，且未与人类行为直接对比。","relevance":"该研究批判性地评估了LLM在社会仿真中的心智理论能力，通过精确的反事实设计揭示了表面流畅性下的认知缺陷，对关注仿真可靠性与偏差的研究者具有重要参考价值，值得阅读原文以了解其审计协议和发现。","inspiration":"借鉴其反事实探针设计，在冻结证据下逐一改变因素以隔离因果效应，可用于经济实验中识别决策机制。｜可迁移到谈判博弈、拍卖或合作博弈等经济场景，检验LLM代理是否真正理解对手偏好或仅依赖表面信息。｜设计一个双边贸易谈判实验，用LLM代理作为被试，随机改变一方代理的收益表（处理），测量其对对方优先级的推断和最终协议效率，并与人类谈判数据（如实验经济学中的谈判结果）对照，评估LLM仿真的有效性。"}},{"id":"2608.17168","version":2,"title":"Can LLMs Reason in a Legally Meaningful Manner? A Small-scale Study on European Court of Human Rights Cases","zh_title":"大语言模型能否以法律上有意义的方式进行推理？一项关于欧洲人权法院案件的小规模研究","abstract":"Reasoning has become a standard technique and feature for contemporary LLMs; however, its application and quality in the context of demanding legal-oriented tasks, such as legal case forecasting, remain under explored. We investigate how LLMs reason in the context of legal case forecasting, using legal cases from the European Court of Human Rights (ECtHR) as a testbed. We evaluate OpenAI GPT 5.4, a recent top-tier LLM, by exploring alternative prompting strategies that are more or less suggestive of what counts as legally meaningful reasoning in the context of ECtHR jurisprudence. We present our findings derived from assessing the model's responses with both human and LLM evaluation. We find that the examined model scores far from ideal in legal reasoning, the model produces structurally complete but substantively shallow analyses, and that LLM-as-a-Judge evaluators are internally consistent yet align only weakly with our trained annotators, i.e., reliable but not a valid substitute for human evaluation. Overall, the expert-curated prompt leads to more comprehensive reasoning, which does not result in more accurate predictions compared to the other examined settings. Based on our findings, we urge the community not to rely solely on automated LLM-based evaluation and to avoid using task accuracy as an appropriate proxy for reasoning quality.","authors":["Amogh Raina","Ilias Chalkidis","Daniel Hershcovich","Henrik Palmer Olsen"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-22","first_seen":"2026-08-19","revised_at":"2026-09-22","abs_url":"https://arxiv.org/abs/2608.17168","pdf_url":"https://arxiv.org/pdf/2608.17168","source_feed":"cs.CL","score":6,"bucket":"other","rubric_hits":["D1"],"tags":["法律推理","LLM评估","标注替代"],"reason":"论文用LLM评估法律推理，涉及LLM-as-a-Judge替代人类评估，属于标…","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:28","error":null,"has_summary":false,"summary":null},{"id":"2609.22133","version":1,"title":"Observational Equivalence of LLM and Human Annotation","zh_title":"LLM与人类标注的观测等价性","abstract":"In this paper, we show that LLM and human coding are observationally equivalent in terms of annotation quality: recent LLMs agree with expert coders at rates comparable to those observed among experts themselves. We demonstrate this through replications of text-classification tasks from 14 peer-reviewed political science studies, in which ten LLMs, three human experts, and 165 crowdsourced workers independently classify the same texts using identical codebooks. We find that this equivalence is driven by ambiguity in the texts and coding rules. When LLMs disagree with experts, experts are also more likely to disagree with one another, and clarifying coding rules reduces disagreement among both experts and sufficiently capable LLMs. Thus, there is little empirical basis for preferring human coding on the basis of annotation quality alone, while LLMs offer substantial advantages in speed and cost. We therefore argue that the central challenge of text annotation is no longer choosing between human and machine coders, but developing coding rules that minimize ambiguity and accounting for the ambiguity that remains. To this end, we propose using disagreement across LLMs to identify difficult cases and refine codebooks, and we develop ambiguity-aware bounds for downstream inference when a unique annotation cannot be defined for every text.","authors":["Kentaro Nakamura","Jing Ling Tan","George Yean"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.22133","pdf_url":"https://arxiv.org/pdf/2609.22133","source_feed":"cs.CL","score":6,"bucket":"other","rubric_hits":["D1"],"tags":["LLM标注","人类对照","文本分类"],"reason":"LLM替代人工标注，非仿真人类被试，但涉及人类对照与测量质量，边界相关。","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:04:50","error":null,"has_summary":false,"summary":null},{"id":"2609.22778","version":1,"title":"MIS-Bench: Benchmarking Multimodal LLMs for Psychotherapeutic Interpersonal Skills Assessment","zh_title":"MIS-Bench：面向心理治疗人际技能评估的多模态大语言模型基准","abstract":"Multimodal large language models (MLLMs) are increasingly used as evaluators, yet their reliability in professional assessment tasks that require expert judgment remains unclear. We investigate this challenge in the context of assessing psychotherapeutic interpersonal skills and introduce MIS-Bench, a Multimodal Interpersonal Skills (MIS) benchmark comprising 996 psychotherapy response videos annotated across 8 dimensions of Facilitative Interpersonal Skills. Across 9 MLLMs with multiple modality and prompting settings, we find that current models show only modest agreement with human experts, inconsistent gains from multimodal input, and limited benefits from reasoning-based prompting. To mitigate this gap, we propose MIS-RAFT, a regression-aware fine-tuning method inspired by RAFT and tailored to fine-grained interpersonal skill scoring at one-decimal precision. MIS-RAFT addresses the mismatch between autoregressive token prediction and scalar-valued expert assessment, significantly improving agreement with human ratings. Overall, MIS-Bench reveals a clear gap between general multimodal capability and expert-level interpersonal judgment, while MIS-RAFT offers a promising path toward more reliable model-based assessment.","authors":["Yuhan Lu","Yi Yao","Hua Shen","Katie Aafjes-van Doorn","Zhaonan Wang"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.22778","pdf_url":"https://arxiv.org/pdf/2609.22778","source_feed":"cs.CL","score":6,"bucket":"other","rubric_hits":["D1"],"tags":["多模态评估","心理治疗","标注替代"],"reason":"LLM作为评估者替代人类专家评分，属于标注替代而非仿真被试，但涉及人类数据对照…","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:13","error":null,"has_summary":false,"summary":null},{"id":"2609.22934","version":1,"title":"Measuring Behavioural Signatures of Large Language Models through Psychometric Profiling","zh_title":"通过心理测量剖析大语言模型的行为特征","abstract":"Large language models (LLMs) increasingly mediate human decisions and communication, yet their behavioural regularities remain difficult to characterize systematically. We develop a cross-linguistic psychometric profiling framework and evaluate nine LLMs using seven psychological instruments, with five repeated administrations per model and language in Chinese and English. Items unresolved after a prespecified retry procedure are retained as NA. Joint analysis of scored and NA responses captures response tendencies and boundaries of self-report applicability. LLMs exhibit structured, model-specific profiles despite a shared alignment-shaped pattern of higher prosocial and self-regulatory responses and lower dominance, disengagement and harmful-intent endorsement. NA responses are structured rather than uniformly distributed, indicating where outputs are treated as inapplicable, refused or cannot be mapped to valid response options. Language condition and provider origin are associated with profile configuration and answerability, whereas repeated administrations show high reproducibility and permit recovery of model identity. Human-reference and prompt-robustness analyses further indicate that these signatures are context dependent. Joint analysis of psychometric profiling and answerability offers a framework for quantifying deployment-level behavioural signatures.","authors":["Yu Sha","Junqi Tao","Dixin Zhou","Yansheng Tu","Mingyang Chen","Xiang Fan","Yang Liu","Mengquan Yang","Jie Lin","Jiahui Fu","Hua Zheng","Benwei Zhang","Zhou Kai"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.22934","pdf_url":"https://arxiv.org/pdf/2609.22934","source_feed":"cs.CL","score":6,"bucket":"other","rubric_hits":["D2"],"tags":["LLM心理测量","行为特征","跨语言"],"reason":"对LLM进行心理测量，测的是模型本身而非人类仿真，但方法可迁移","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:04:58","error":null,"has_summary":false,"summary":null},{"id":"2609.24516","version":1,"title":"LLJ Cards: Best practices for the Use of LLMs as Judges","zh_title":"LLJ卡片：使用大语言模型作为评判者的最佳实践","abstract":"In recent years, large language models (LLMs) have emerged as a popular alternative for evaluation. Often referred to as LLMs as judges (LLJs), these systems have been widely adopted by researchers and practitioners across a broad range of measurement tasks, driven by their strong performance, scalability, and cost-effectiveness relative to human judgment. However, a growing body of work has shown that the use of LLJs raise concerns about their validity and reliability as evaluators. Existing efforts to address these challenges have largely focused on developing bias-mitigation techniques and refining prompting strategies. While these approaches represent an important step forward, they primarily offer technical fixes and leave a more fundamental challenge unaddressed: the lack of standardized, transparent, and reproducible evaluation practices. In this paper, we introduce LLJ Cards, a framework that synthesizes best practices from measurement theory, natural language generation, and machine learning literature into practical guidelines for LLJ-based evaluations. While LLJs offer a promising path toward scalable evaluation, their effective use requires grounding in rigorous evaluation principles to ensure validity, reliability, and reproducibility. LLJ Cards addresses this need by providing a structured framework for applying these principles in the design and reporting of automated evaluations.","authors":["Khaoula Chehbouni","Melina Medjdoub","Florian Carichon","Golnoosh Farnadi","Jackie Chi Kit Cheung"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.24516","pdf_url":"https://arxiv.org/pdf/2609.24516","source_feed":"cs.CL","score":6,"bucket":"other","rubric_hits":["A4","D1"],"tags":["LLM评估","方法论框架","效度与可靠性"],"reason":"提出LLM评估框架，涉及效度与可靠性，但非仿真人类被试，而是替代标注员。","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:23","error":null,"has_summary":false,"summary":null},{"id":"2608.02551","version":2,"title":"Who Should Be Generated? Justifying Demographic Targets in Open-Ended Generation","zh_title":"谁应被生成？论证开放式生成中的人口统计目标","abstract":"Fairness evaluation concerns not only what a model produces, but also what its outputs ought to be compared against. When a model generates \"a CEO in the United States,\" the prompt leaves demographic realization to the model. Existing group fairness definitions assume that sensitive attributes are given on the input side. Generative audits instead examine output-side demographic composition, yet the targets they compare it against are typically supplied rather than justified. The upstream question is what the target distribution should be. We formalize this missing-target problem for demographic-value-unspecified generation and decompose target construction into four commitments: the evaluative object, prior admissibility, allocation, and operationalization. In this framework, we admit the geographic prior under a geographic-membership interpretation for the declared public-world use. The occupational prior, under an incumbency interpretation, requires an independently defended objective such as workforce-composition fidelity. Instantiating this construction in AP-Bench, we find substantial distribution divergence from geography-derived targets, ranging from 0.508 to 0.606 on a 0-to-1 scale. Replacing each geography-derived target with an equal-category comparator, while holding generations and measurement fixed, produces model-specific mean absolute cell-level $\\mathrm{JSD}_2$ changes ranging from 0.279 to 0.355. Target construction is therefore not a preliminary to fairness evaluation but a component of it. What we supply is not a universal target, but a framework that makes explicit the justification required before a distribution can serve as a fairness standard.","authors":["Zeshen Zheng","Yujia He","Qianmian Lin","Xiangyue Huang","Wenqing Chen"],"categories":["cs.CY","cs.AI","cs.CL"],"primary_category":"cs.CY","announce_type":"replace-cross","date":"2026-09-22","first_seen":"2026-08-04","revised_at":"2026-09-22","abs_url":"https://arxiv.org/abs/2608.02551","pdf_url":"https://arxiv.org/pdf/2608.02551","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["公平性评估","生成模型","人口统计分布"],"reason":"论文关注生成模型输出的人口统计分布与公平性目标，而非用LLM仿真人类被试，属于…","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:26","error":null,"has_summary":false,"summary":null},{"id":"2609.22104","version":1,"title":"DeepInstructor: An Agentic AI Instructor for Experience-Driven Idea Evaluation","zh_title":"DeepInstructor：一种用于经验驱动型想法评估的智能体AI导师","abstract":"As automated scientific discovery advances, Large Language Models (LLMs) can now generate research ideas at an unprecedented scale, shifting the bottleneck from idea generation to idea evaluation. Existing evaluators mainly rely on parametric LLM knowledge or unstructured retrieval, producing judgments that lack the experience-grounded reasoning used by human instructors. To address this, we propose DeepInstructor, an agentic framework that formulates idea evaluation as reasoning over structured scholarly experience. DeepInstructor constructs an Experience Graph from 58,607 peer reviews and employs a ReAct-based agent to retrieve dimension-specific evidence for traceable evaluation. We further introduce DeepInstruct, a dataset with controlled pairwise comparisons across novelty, significance, and feasibility. Experiments show that DeepInstructor substantially outperforms existing baselines, improving Hit@1 and Hit@2 alignment with human judgments by 24.4% and 29.7%, respectively. Our findings suggest that scientific idea evaluation can be grounded in explicit reasoning over structured scholarly experience","authors":["Rongcan Pei","Fang Guo","Qinglin Qi","Qi Zhu","Yun Luo","Jianhao Yan","Minjun Zhu","Qiujie Xie","Dehong Zheng","Yue Zhang"],"categories":["cs.CL","cs.IR"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.22104","pdf_url":"https://arxiv.org/pdf/2609.22104","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM评估","学术评审","智能体"],"reason":"用LLM替代人工评审，属于标注员替代，非仿真人类被试","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:09","error":null,"has_summary":false,"summary":null},{"id":"2609.22195","version":1,"title":"The Situated Identity Test: Distinguishing Persistent Cognitive Identity from Persona Imitation","zh_title":"情境化身份测试：区分持久认知身份与人设模仿","abstract":"Large language models can convincingly adopt personas, recall past dialogues, and weave rich autobiographies. Yet this conversational eloquence conceals a fundamental attribution problem: looking the part does not mean having lived the life. Two individuals can share identical public profiles--the same age, hometown, occupation, and personality traits--while possessing entirely distinct private histories, relationships, and acquired skills. When conditioned solely on that shared profile, an agent lacks the information required to determine which lineage is correct. We introduce the Situated Identity Test (SIT), an architecture-independent framework that evaluates whether an agent's behavior is functionally attributable to a specific developmental lineage. Grounded identity requires both appropriate knowledge of recorded experiences and appropriate ignorance of ungrounded ones, bounded by what the identity has actually acquired rather than what its underlying foundation model knows. We prove that any policy conditioned solely on a compressed profile is bounded by an average situated validity of at most 1/m across m colliding life histories on lineage-discriminative queries (at most 50% for paired lineages). We instantiate this framework in SITBench, an evaluation suite designed for 25 profile-collision pairs (50 distinct lineages) across 10,000 planned probes and nine architectural configurations. Supported by an open-source reference implementation, deterministic test fixtures, and empirical pilot evaluations on frontier foundation models (GPT-5.6 Sol and Claude Opus 5), we formalize the failure modes of persona prompting under profile collision and provide an assurance harness for evaluating episodic continuity, structured state, and epistemic boundaries.","authors":["Jun He","Deying Yu"],"categories":["cs.CL","cs.AI","cs.LG"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.22195","pdf_url":"https://arxiv.org/pdf/2609.22195","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["LLM身份评估","人设模仿","认知边界"],"reason":"评估LLM身份一致性，测模型而非仿真人类被试，无人类数据对照","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:04:52","error":null,"has_summary":false,"summary":null},{"id":"2609.22248","version":1,"title":"Checkpoints Are Not Enough: Trust Calibration in CoSLR, a Human-AI System for Systematic Literature Reviews","zh_title":"检查点还不够：CoSLR中的人机信任校准——一个用于系统文献综述的人机协作系统","abstract":"Systematic Literature Reviews (SLRs) are essential for evidence-based research but remain time-consuming, requiring researchers to manage large volumes of publications across planning, screening, analysis, and reporting. Large language models (LLMs) can now produce fluent, well-structured review text, which makes it difficult to distinguish synthesis that was verified by a researcher from synthesis that merely appears authoritative. This raises the risk that unverified AI-generated synthesis enters the scholarly record carrying the credibility of a systematic review. We present CoSLR, a Human-AI collaborative multi-agent system that supports the SLR workflow through a modular three-phase pipeline using large language models and Retrieval-Augmented Generation (RAG), and that places explicit, mandatory human checkpoints on the path between generated output and its acceptance. In a survey-based study with 63 participants, the system was received positively: 27 of 63 participants (42.9 percent) rated its usability highly, indicating that the mandatory checkpoints did not come at the cost of a workable interface. However, a checkpoint safeguards the review only if researchers use it to verify: 22 of 63 participants (34.9 percent) reported that they would trust AI-generated summaries and reports without additional human checking after only a short interaction with the system. These findings indicate that Human-AI collaboration can support literature review work, but that the effectiveness of human oversight depends on whether users are willing to exercise it. This is a calibration problem that interface design must address directly, not assume.","authors":["MD Aidul Islam","Malik Abdul Sami","Muhammad Waseem","Zeeshan Rasheed","Kai-kristian Kemell","Zheying Zhang","Pekka Abrahamsson"],"categories":["cs.CL","cs.AI","cs.IR"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.22248","pdf_url":"https://arxiv.org/pdf/2609.22248","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["人机协作","系统综述","信任校准"],"reason":"LLM用于辅助系统综述，人类参与者评估系统，非仿真人类被试，但涉及人机信任校准…","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:11","error":null,"has_summary":false,"summary":null},{"id":"2609.22255","version":1,"title":"Deep Persona: A Psychologically Grounded Architecture and Evaluation Framework for Role-Playing Agents and Simulations","zh_title":"深度人格：一种基于心理学基础的角色扮演智能体与仿真的架构及评估框架","abstract":"Existing approaches to persona simulation with Large Language Models (LLMs) mostly rely on shallow character descriptions that fail to sustain coherent character behavior across extended interactions. We introduce Deep Persona, a psychologically grounded, three-layered architecture that organizes personas into hierarchical levels of observable expression, latent beliefs, and core motivational drives, for constructing highly convincing role-playing agents. Governed by the principles of scripted determinism and bounded agency, the architecture restricts the model to a reactive engine guided by a structured internal script. We further propose a reference-free evaluation framework that benchmarks dialogue naturalness against empirical human distributions using established psychological clinical instruments and adversarial stress-tests. Empirical evaluation reveals that while LLMs achieve high pragmatic fluency, they exhibit systematic limitations in emotional expression and joint attention. In addition, we present a case study of two Deep Personas and evaluate them using the proposed framework, demonstrating that structured personas can produce interactions that more closely align with human conversational behavior.","authors":["Rotem Dror","Zohar Elyoseph","Yuval Haber","Elad Refoua","Oshrat Ayalon","Adir Solomon"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.22255","pdf_url":"https://arxiv.org/pdf/2609.22255","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D2","D3"],"tags":["角色扮演智能体","心理学架构","对话评估"],"reason":"构建角色扮演agent并评估对话自然度，但无真实人类行为对照，且测的是模型表现…","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:11","error":null,"has_summary":false,"summary":null},{"id":"2609.23573","version":1,"title":"Error-Supervised Synthetic Learner Writing for Automated Essay Scoring","zh_title":"错误监督的合成学习者写作用于自动作文评分","abstract":"Synthetic essays can help reduce dependence on human-written data in Automated Essay Scoring (AES). However, they often lack realistic errors, limiting their ability to represent authentic human writing, particularly when the target texts are intended to resemble those produced by language learners. In this study, we present a simple approach that introduces error supervision into synthetic essay generation. Specifically, we fine-tune an LLM generator on error-annotated texts of the kind commonly used in Grammatical Error Detection (GED). To assess the utility of the proposed approach, we fine-tune and evaluate AES scorers under three data conditions: authentic essays, synthetic essays generated conventionally, and synthetic essays generated using our proposed approach. The results show that in the larger-data settings, the proposed approach outperforms the conventional synthetic baseline in 11 out of 12 dataset-metric comparisons, with performance in some cases approaching that of models trained on authentic essays. Despite these gains, performance under extremely low-resource settings remains mixed, with advantages over the conventional baseline only becoming more apparent at 200 training essays, although not consistently across datasets. Qualitative and quantitative analyses further show that the proposed approach produces learner-like errors whose distributions broadly resemble those observed in authentic essays.","authors":["Duy Anh Nguyen"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.23573","pdf_url":"https://arxiv.org/pdf/2609.23573","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["合成数据","自动作文评分","LLM生成"],"reason":"用LLM生成合成作文替代真实数据，但目的是训练评分器，非仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:18","error":null,"has_summary":false,"summary":null},{"id":"2609.24052","version":1,"title":"Calibrated Decisions at Scale: Converting Police Crash Narratives into Probabilistic Crash Variables with a System One Model (Jev)","zh_title":"大规模校准决策：用系统一模型将警方事故叙述转换为概率事故变量","abstract":"Crash datasets that carry an investigator narrative hold information the coded fields omit. Coding those narratives at scale has been blocked by three obstacles. Frontier large language models are costly at that scale, their generated text cannot be verified, and no rule says how much output a human must check. This paper formulates narrative coding as gated, typed decisions answered by Jev, a System One model that returns probabilities over analyst-defined options and generates no text. A screen covered 499,500 Texas narratives and 195,857 were coded with a 27-question schema. Cost is governed by schema size rather than narrative length. The probabilities are audited against coded fields and against 2,416 blinded human judgments drawn under a stated sampling design. Two frontier large language models are benchmarked on the same records. Against human labels the typed model attains an F1 of 0.908. One frontier model gains 0.059 and the other is indistinguishable from it. Calibration varies by model rather than by paradigm, so each model must be audited. Recalibration on the same labels reduces calibration error by a factor of 3.3. Agreement with coded fields understates fidelity to the narrative by a median of 0.26 in kappa. A resolution-floor bound covers any model that reports probabilities on a discrete grid. A review budget over flagged records gives the records a human must read per variable and per year. Adding the calibrated variables to the coded fields raises the injury and fatal crashes attributed to nine factors by 10,747 per year.","authors":["Amir Rafe","Subasish Das"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.24052","pdf_url":"https://arxiv.org/pdf/2609.24052","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM标注","文本编码","概率校准"],"reason":"用LLM替代人工编码事故报告，属于标注员替代，非仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:21","error":null,"has_summary":false,"summary":null},{"id":"2609.22102","version":1,"title":"When Who You Are Can Change the Code You Get: A Study of Persona-Induced Bias in LLM Code Generation","zh_title":"当你是谁可以改变你得到的代码：LLM代码生成中角色诱导偏差的研究","abstract":"Large Language Models (LLMs) are widely used as programming assistants, yet it remains unclear whether and how user's demographic information impacts the technical quality of generated code. We conduct a large-scale empirical study of persona-induced bias in LLM-based code generation, focusing a proprietary model (Gemini 2.5 Pro) and an open-weight model (GPT-OSS-120B). Using 18 demographic personas spanning nationality, gender, and experience level, we compare persona-induced prompts against a neutral baseline. Across 35,000+ generated programs, we analyze demographic marker leakage in reasoning and responses, as well as differences in functional correctness, maintainability, code style, and security. Our results show that demographic cues are frequently reflected in LLM reasoning and outputs. Demographic markers appear in up to 65% of responses and 70% of reasoning traces, despite being semantically irrelevant to the tasks. On LiveCodeBench, persona prompting were associated with lower correctness scores of the Gemini model by an average of 1.54 percentage points, with one persona exhibiting a decrease of 3.6% (odds ratio = 0.51). In contrast, the accuracy of the GPT-OSS model improved by 3.4 - 5.7% across all personas (odds ratios = 1.8 - 3.0). Maintainability and code style metrics show statistically significant but negligible effect sizes (all Cliff's {\\delta} < 0.15), and security vulnerabilities exhibit no systematic persona-specific patterns. Overall, our results show that the presence of demographic information about users is associated with measurable variation in LLM reasoning and code quality even in purely technical tasks, and that these effects hold across models. Our work highlights an under-examined risk in LLM-assisted software development.","authors":["Anubhav Gupta","Mayara Costa Figueiredo","Leticia Santos Machado","Tanner Wright","Ivan Beschastnikh","Cleidson R. B. de Souza","Gema Rodr\\'iguez-P\\'erez"],"categories":["cs.SE","cs.CL","cs.CY"],"primary_category":"cs.SE","announce_type":"cross","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.22102","pdf_url":"https://arxiv.org/pdf/2609.22102","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["LLM偏差","代码生成","人口统计角色"],"reason":"研究LLM对用户人口统计信息的响应偏差，属于模型行为测量，非人类仿真。","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:08","error":null,"has_summary":false,"summary":null},{"id":"2609.22170","version":1,"title":"Multiple latent orderings better predict language model preferences","zh_title":"多重潜在排序更好地预测语言模型偏好","abstract":"Language models are frequently employed in settings where they are asked to make value judgments and choices. These observed choices often exhibit intransitivity: A model may prefer item $A$ to $B$ and $B$ to $C$, while also preferring $C$ to $A$. Existing work that models LLM preferences treats such inconsistencies as sampling noise around a single latent ordering. We instead propose that intransitivity reflects the aggregation of multiple latent, internally consistent orderings. We first show that observed inconsistencies cannot be explained by a single ordering under any monotone link function. We then introduce a noise-augmented mixture Bradley-Terry (MBT) model that infers latent preference components from repeated pairwise comparisons. Across seven models and four tasks, a mixture of orderings often explains structural inconsistencies better than single-utility models. We find that aggregate preferences often hide underlying preference heterogeneity. A case study on Moral Machine dilemmas shows that models which disagree on aggregate orderings can still share latent components. Together, these results suggest that LLMs reflect plural preferences. Alignment and evaluation pipelines that treat LLM preferences as a single function, therefore, risk averaging over coherent orderings that different users may endorse differently.","authors":["Aviral Chawla","William H. W. Thompson","Jean-Gabriel Young"],"categories":["cs.LG","cs.AI","cs.CL"],"primary_category":"cs.LG","announce_type":"cross","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.22170","pdf_url":"https://arxiv.org/pdf/2609.22170","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["LLM偏好","潜在排序","偏好建模"],"reason":"研究LLM偏好结构，测量模型本身而非仿真人类被试，但涉及偏好建模可迁移","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:04:50","error":null,"has_summary":false,"summary":null},{"id":"2609.22517","version":1,"title":"Scalable AI-based clinical communication training and automated assessment","zh_title":"基于AI的可扩展临床沟通训练与自动评估","abstract":"Poor clinical communication can delay care, contribute to errors, and harm patients, yet opportunities for repeated practice with feedback remain limited. Our prior randomized trial showed that practice with the SOPHIE AI patient platform improved serious illness communication, but the system addressed a single clinical context and required human effort for delivery and assessment. We developed SOPHIE 2.0, a browser-based, self-service platform integrating embodied AI-patient interactions, personalized feedback, and automated assessment across 24 clinical scenarios. An automated large language model assessor evaluated three communication skills---Empower, Be Explicit, and Empathize---with agreement comparable to individual human raters ($r=0.759$; ICC$=0.746$). In a study of 59 clinicians and students, participants completed two AI-patient encounters with personalized feedback; 92% found the platform engaging, 86% easy to use, and 83% clinically relevant. Scores were higher in the second encounter, though the uncontrolled design precludes attributing this change specifically to training.","authors":["Masum Hasan","Ron Epstein","Thomas Carroll","Ehsan Hoque"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.22517","pdf_url":"https://arxiv.org/pdf/2609.22517","source_feed":"cs.HC","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM评估者","临床沟通训练","自动评分"],"reason":"LLM用作评估者替代人工评分，而非仿真人类被试，属标注替代","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:04:55","error":null,"has_summary":false,"summary":null},{"id":"2609.23104","version":1,"title":"Deciphering the Babel of Play: A Human-AI Collaborative Approach for Large-Scale Cross-Language Analysis of Game Reviews","zh_title":"解读游戏评论的巴别塔：一种人机协作的大规模跨语言游戏评论分析方法","abstract":"We present a large-scale cross-language analysis of game reviews using a human-AI collaborative framework that combines quantitative screening with multilingual large language models (LLMs). Starting from 17 million Steam reviews across 30 languages and 2,000 top-selling titles, we select 28 games with notable cross-language rating patterns. We then apply LLM-assisted content analysis to 442,162 reviews spanning 17 languages, with human researchers guiding codebook development and interpreting the results. Our findings reveal differences in both the aspects language communities prioritize and how they evaluate them, highlighting the roles of narrative expectations, game mechanics and stability, localization quality, cultural proximity, and perceptions of developers and publishers. We also identify rare cases of cross-language consensus. This work offers empirical insights into cross-cultural game evaluation and a scalable methodological approach to multilingual content analysis that preserves human interpretation.","authors":["Zixiaofan Yang","Chang Xiao"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.23104","pdf_url":"https://arxiv.org/pdf/2609.23104","source_feed":"cs.HC","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM辅助内容分析","跨语言分析","游戏评论"],"reason":"LLM用于内容分析辅助，替代人工标注，非仿真人类被试，但方法可迁移。","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:15","error":null,"has_summary":false,"summary":null},{"id":"2609.22600","version":1,"title":"From Certain Doom to Survival: Agent-Driven Self-Governance in LLM Agent Societies","zh_title":"从必然毁灭到生存：LLM Agent 社会中的智能体驱动自治","abstract":"Multi-agent LLM systems are increasingly evaluated in social dilemmas, but most work treats governance as imposed by the experimenter, expressed rhetorically, or restricted to a fixed menu of mechanisms. We introduce GovSim-SelfGovern, an extension of the GovSim common-pool resource environment in which agents author executable Python governance rules, receive sandbox validation feedback, vote on proposed laws, and live under the rules they enact across rounds. To evaluate agent-driven self-governance, we examine three scenarios ranging from stable abundance to a fatal resource wall where five agents cannot all survive through harvest alone. To solve this, agents must write and debug useful laws in time before their institutions degrade sharply under resource pressure. Finally, we study a central alignment question: when agents hesitate to propose exile, are they rejecting it for normative reasons, or does it never enter their candidate set? Our results show that executable governance improves the space of possible interventions for agents, but survival depends on whether agents discover the right institutional mechanisms in time. Fiscal capacity enables redistribution, while deeper reasoning and removal of democratic veto make exile more feasible. GovSim-SelfGovern therefore adapts executable code actions to a common-pool governance setting and shows how scarcity turns institutional authorship into a political and ethical problem.","authors":["Gregory B. Rehm"],"categories":["cs.MA"],"primary_category":"cs.MA","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.22600","pdf_url":"https://arxiv.org/pdf/2609.22600","source_feed":"cs.MA","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["LLM Agent 社会模拟","公共池资源治理","自治机制"],"reason":"LLM agent 社会模拟，但无真实人类数据对照，属边界情形","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:04:56","error":null,"has_summary":false,"summary":null},{"id":"2609.24895","version":1,"title":"Human-LLM Deliberation as Interactive Proof: Conditions for Verifiability Without Transparency","zh_title":"人机协商作为交互式证明：无透明性下可验证性的条件","abstract":"When an LLM supplies an argument that a user could not readily construct, how can the user decide whether to accept its claim? Inspired by interactive proofs, we model human-LLM deliberation as an interaction between a prover with unrestricted internal search and a resource-bounded human verifier. The verifier requests and checks supporting details without access to the LLM's internal state. Passed checks accumulate evidence toward an acceptance threshold. We prove anytime-valid soundness against adaptive provers: the probability of ever accepting a false claim is at most a chosen error level, provided the task supplies bounds on false passes and human checking errors that remain valid after every relevant history. A finite-horizon completeness bound additionally requires bounds on the adequacy of honest responses and sufficient diagnostic progress. Further checks can strengthen the evidence for acceptance, but each requires another adequate response and reliable human effort. Whether this tradeoff permits certification depends on the verifier's effort budget, cognitive load, expertise, and fatigue. We identify conditions under which the supplied bounds certify a specified sequence of local checks but not a specified global check under the same resource budgets.","authors":["Baotong Zhang","Dean Foster","Jo\\~ao Sedoc"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.24895","pdf_url":"https://arxiv.org/pdf/2609.24895","source_feed":"cs.CL","score":3,"bucket":"other","rubric_hits":["C1"],"tags":["人机交互","论证验证","交互式证明"],"reason":"研究人机论证验证机制，非用LLM仿真人类被试，无人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:24","error":null,"has_summary":false,"summary":null},{"id":"2609.20408","version":2,"title":"Xeno-Interpretability: Investigating the Alien Minds of LLMs","zh_title":"异质可解释性：探究大语言模型的异质心智","abstract":"Large language models are usually interpreted through concepts that humans already possess: truthfulness, refusal, deception, personality, harmfulness, and related categories. This paper asks whether models may also represent and use distinctions for which no adequate human concept exists. We call such internal structures xeno-representations, and their study xeno-interpretability. We distinguish the human-interpretable semantic space from the xeno-semantic space: the region of model-native representations for which no adequate human conceptual counterpart is available. We show that the space of possible internal distinctions in an LLM is substantially larger than the space available through finite human descriptions. We then separate experimental identification from semantic interpretation: an internal representation may be reproducibly located, geometrically characterized, causally manipulated, and linked to downstream behaviour even when its semantic content cannot be adequately expressed in human terms. On this basis, we sketch an empirical programme to identify xeno-representations. We finally examine the implications for AI safety and multi-agent systems, where model-native representations may propagate and stabilize across interacting agents while remaining only partially visible through human-readable communication. Xeno-interpretability therefore shifts the aim of interpretability from finding human concepts inside models toward discovering and characterizing the representational structures that are native to the models themselves and might affect their behaviour in unpredictable ways.","authors":["F. Pierucci","M. Bracale Syrnikov","M. Prandi","M. Galisai","F. Giarrusso","P. Bisconti"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-22","first_seen":"2026-09-18","revised_at":"2026-09-22","abs_url":"https://arxiv.org/abs/2609.20408","pdf_url":"https://arxiv.org/pdf/2609.20408","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["可解释性","表征分析","AI安全"],"reason":"研究LLM内部表征，不涉及人类行为仿真或对照，且多智能体部分仅为讨论，无实验。","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:40","error":null,"has_summary":false,"summary":null},{"id":"2609.21254","version":2,"title":"Two's a Crowd: Human and AI-Based Copresence for Developers with ADHD","zh_title":"人多则乱：ADHD开发者的人与AI共在","abstract":"Effective collaboration and communication are vital to developer productivity and well-being, yet remain constrained by human factors such as attention, intrinsic motivation, and interpersonal accountability. These constraints are particularly vital for developers identifying with Attention Deficit Hyperactivity Disorder (ADHD), who navigate persistent environmental barriers in modern hybrid workplace settings. While developers with ADHD frequently rely on collaborative copresence practices (such as body doubling or pair programming) to support executive function, the recent emergence of agentic AI coding assistants has begun reshaping these collaborative dynamics. To investigate how developers with ADHD engage in human and AI-based copresence practices, we conducted semi-structured interviews with 14 software engineers with ADHD. Our findings reveal that while traditional human-human copresence provides critical social support and onboarding structure, it forces developers to constantly manage professional reputation and sacrifice personal privacy. Conversely, developers leverage emerging human-AI copresence to maintain accountability and cognitive flow without the social anxiety, performance judgment, or surveillance associated with human observation. Based on these empirical insights, we map developer copresence practices onto core dimensions of Goffman's copresence theory and Forsgren et al.'s SPACE framework of developer productivity, and provide design recommendations for AI-based tools that promote inclusive collaboration for developers with ADHD.","authors":["Veronica Pimenova","Seth Bernstein","Shalini Madan","Dhruv Jain","Venkatesh Potluri"],"categories":["cs.HC","cs.SE"],"primary_category":"cs.HC","announce_type":"replace","date":"2026-09-22","first_seen":"2026-09-21","revised_at":"2026-09-22","abs_url":"https://arxiv.org/abs/2609.21254","pdf_url":"https://arxiv.org/pdf/2609.21254","source_feed":"cs.HC","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["人机交互","ADHD","协作"],"reason":"研究开发者与AI编码助手的协作体验，属于人机交互而非用LLM仿真人类被试，无实…","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:30","error":null,"has_summary":false,"summary":null},{"id":"2609.22112","version":1,"title":"Privacy Personalization Trade offs in LLMs: The Impact of Stylometric Signal Reduction on User-Specific Text Generation","zh_title":"大语言模型中的隐私与个性化权衡：文体信号减少对用户特定文本生成的影响","abstract":"Large language models (LLMs) have demonstrated the ability to generate user-specific text with high stylistic fidelity. However, the personal data that enables such personalization frequently embeds demographic, cultural, and stylistic markers that raises concerns about stylometric re- identification. This paper investigates whether reducing identifiable stylistic signals affects personalization in text generation by LLMs. We introduce a controlled framework to isolate stylometric signals in LLM personalization using the LaMP-7 Twitter benchmark. Experiments on 250 sampled users compare two settings: paraphrasing conditioned on the original profile and paraphrasing conditioned on an anonymized converted profile in which demographic identifiers, cultural references, personal details, and informal linguistic cues have been systematically neutralized. Outputs are assessed by two independent LLM judges and a complementary human evaluation. Our pairwise evaluation shows that outputs conditioned on original profiles are nearly indistinguishable from human-authored ground truth, indicating that modern LLMs can closely reproduce an author's writing style with sufficient fidelity. In contrast, preference for model outputs with anonymized profiles drops to 13.0% on average, while semantic context preservation remains high at 94.8%. A study with human evaluators confirms the same pattern. These findings reveal a clear privacy-personalization trade-off and highlight the need for privacy-aware personalization methods that retain meaning while suppressing identifying stylistic signals.","authors":["Muhammed Nazmul Arefin","Omar Jamal Hammad"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.22112","pdf_url":"https://arxiv.org/pdf/2609.22112","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["隐私保护","文本生成","个性化"],"reason":"研究LLM个性化文本生成中的隐私权衡，不涉及将LLM作为人类被试进行仿真实验，…","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:09","error":null,"has_summary":false,"summary":null},{"id":"2609.22198","version":1,"title":"The Role of AI in Online Reviews","zh_title":"AI在在线评论中的作用","abstract":"The rapid adoption of large language models (LLMs) creates new opportunities for strategic content generation on online platforms, including potentially harmful forms of manipulation that may undermine platform effectiveness and reshape platform dynamics. However, measuring such activity is difficult because AI-generated content is rarely directly observable. We introduce an empirical approach that leverages discrete LLM supply shocks - abrupt changes in model prices and capabilities, and contrasts verified with non-verified reviews to identify changes in platform activity associated with generative AI supply improvements. We apply this approach to more than 13 million reviews from Trustpilot, one of the leading online platforms for business reviews. A robust finding is that following LLM supply shocks, unverified reviews shift toward greater negativity: more 1-stars, fewer 5-stars, and lower ratings, with effects driven primarily by new model releases and concentrated among firms with the lowest and highest review volumes, suggesting that strategic AI use may reshape platform competition dynamics. We further find that LLM supply shocks trigger short, concentrated bursts of review activity. Together, these findings suggest that generative AI is already reshaping how reputation and competition operate on online platforms.","authors":["Valeria Lerman","Oren Rigbi","Yaniv Dover"],"categories":["cs.CL","cs.AI","cs.HC","cs.SI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.22198","pdf_url":"https://arxiv.org/pdf/2609.22198","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C5"],"tags":["AI生成内容","在线评论","平台动态"],"reason":"研究AI生成内容对平台的影响，非用LLM仿真人类被试，无人类行为对照","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:04:52","error":null,"has_summary":false,"summary":null},{"id":"2609.22204","version":1,"title":"Evaluating Personal Information Output from Conversational Interactions in Generative AI Systems","zh_title":"评估生成式AI系统对话交互中的个人信息输出","abstract":"This exploratory pilot study evaluates the scope and perceived accuracy of personal information output from ongoing conversational interactions in generative AI systems using GPT-5.2 Instant and GPT-5.2 Thinking, categorized into three output types: Fact, Inference, and Confidence. Based on the evaluation results obtained from 15 Japanese participants, differences in model design have limited impact on personal information output tendencies. Compared with the Inference type, the Fact type shows a more conservative output pattern. Regarding attribute categories, the findings indicate that Core Personal attributes associated with identification are treated relatively conservatively, whereas Behavioral and Linguistic attributes show higher accuracy across both Fact and Inference outputs. Furthermore, Holistic Profile, Psychological and Cognitive, and Residual attributes are more readily inferred, even when not supported by explicit factual outputs. Notably, the lack of null outputs for these attributes in the Inference type suggests that such inferred profiles may be constructed from indirectly available contextual information. The findings may contribute to future discussions regarding privacy awareness and personal information inference in generative AI systems.","authors":["Yosuke Seki","Hirotaka Tahara"],"categories":["cs.CL","cs.AI","cs.CY"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.22204","pdf_url":"https://arxiv.org/pdf/2609.22204","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["隐私评估","生成式AI","个人信息推断"],"reason":"研究生成式AI对话中个人信息输出，属隐私评估，非用LLM仿真人类被试，无实验对…","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:04:52","error":null,"has_summary":false,"summary":null},{"id":"2609.22208","version":1,"title":"Replicating the Geometry of Emotion Representations in a Base Open-Weights Model","zh_title":"在开源基础模型中复现情感表征的几何结构","abstract":"Sofroniew et al. (2026) report that emotion concepts in Claude Sonnet 4.5 are represented as vectors whose geometry mirrors human affect psychology. We replicate the representational core of that study on the base pretrained model google/gemma-2-27b, inheriting every disclosed parameter, resolving unspecified steps by disclosed rules, and changing only the subject model. From 205,200 newly generated Claude Sonnet 4.5 stories matching the original corpus design, we extract 171 emotion vectors and recover the core results. The leading principal components form an affective circumplex (PC1 carries 26.7% of variance against the original study's ~27%, PC2 13.4% against ~14%), emotions cluster into similar intuitive families, and the geometry holds across a broad late-middle band. The valence axis aligns with human norms (r = 0.72, against 0.81) and is stable across scales and depth. Arousal aligns at r = 0.67 (against 0.66) but only at the full 171-emotion scale and late depth, so we do not classify it as replicated. Extending the original analysis, a 46-layer sweep locates a sharp seam at L22-26, where the geometry consolidates and the vocabulary readout becomes legible. An embedding-layer baseline finds much of the geometry already present in the static token embeddings, with arousal as the exception. On top-activating held-out text, the geometry predicts token-level co-activation at r = 0.907. At least 52% of vectors peak on structurally non-conceptual tokens, a measured floor for max-activation confounds. Even when a document contains the vector's emotion word, the peak lands on that word only 6.1% of the time. Because the subject is a base model and the stimuli are Claude-generated fiction, the recovered structure is a property of the pretrained representation of Claude-rendered emotion. The original's causal and assistant-facing analyses are out of scope. Code and data are released.","authors":["Adam Hollowell"],"categories":["cs.CL","cs.AI","cs.LG"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.22208","pdf_url":"https://arxiv.org/pdf/2609.22208","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["模型表征","情感计算","可解释性"],"reason":"研究LLM内部情感表征几何，属模型分析，非仿真人类被试，无人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:04:53","error":null,"has_summary":false,"summary":null},{"id":"2609.22362","version":1,"title":"Functional Emotion Without Character: Large Language Models, Aristotelian Disposition, and the Limits of Behavioral Alignment","zh_title":"无性格的功能性情感：大语言模型、亚里士多德倾向与行为对齐的局限","abstract":"Debates about whether artificial systems can feel are often forced between two unsatisfactory positions: behavioral equivalence is treated as sufficient for emotion, or phenomenal consciousness is treated as a prerequisite that makes the question empirically inaccessible. This article develops a structural alternative. It models emotions as context-sensitive regions, trajectories and attractor dynamics in high-dimensional representational state spaces. Recent mechanistic interpretability findings support the existence of causally active emotion-concept representations in large language models, but they do not establish subjective feeling or full emotional agency. Assessed against published adequacy standards for representation in language models, intervention provides strong evidence of causal use, while full affective role integration, uniformity across subject domains and coherence remain only partially established; there is no direct analogue of accuracy. These mismatches expose the need for a standard of affective appropriateness, which an account of character must supply. Such an account requires three further conditions: regulatory embodiment that gives valence endogenous stakes, temporal continuity that allows affective episodes to accumulate into a history, and an integrated self-model that binds that history to persistent values. Aristotle's concepts of path\\=e, hexis, mesot\\=es and phron\\=esis are translated into a state-space sketch in which practical wisdom includes competence in estimating normatively salient context, not merely acting on a context description already given. The framework reframes alignment as a problem of durable disposition rather than output conformity, and yields interventional tests with explicit control conditions.","authors":["Marzieh Zare"],"categories":["cs.CL","cs.CY","cs.HC"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.22362","pdf_url":"https://arxiv.org/pdf/2609.22362","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["情感建模","哲学分析","AI对齐"],"reason":"论文讨论LLM情感表征与性格，属哲学分析，无人类被试仿真或实验对照。","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:11","error":null,"has_summary":false,"summary":null},{"id":"2609.23083","version":1,"title":"Directing large language models to follow the letter or spirit of the law","zh_title":"引导大语言模型遵循法律的精神或字面","abstract":"The distinction between the spirit and letter of the law is a central issue across research and everyday life, and a growing concern for building safe, intelligent machines. What is this distinction based on, and how can we develop machines that follow the intention behind a rule? We used targeted adaptation that made large language models prioritize the spirit or letter of the law. With minimal modifications, our method significantly changed LLM behavior across diverse measures, novel vignettes, real-world scenarios, and influential legal cases. An analysis of model internals revealed a low-dimensional space with three interpretable dimensions matching a formal pre-specified framework for the geometry of legal concepts. These findings show how legal thought in LLMs may be organized and directed.","authors":["Peng Qian","Andrew Li","Sam Chen","Sonia K. Murthy","Yonatan Belinkov","Tomer D. Ullman"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.23083","pdf_url":"https://arxiv.org/pdf/2609.23083","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["法律推理","模型行为控制","可解释性"],"reason":"研究如何引导LLM遵循法律精神或字面，属于模型行为控制，非人类仿真实验","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:14","error":null,"has_summary":false,"summary":null},{"id":"2609.23178","version":1,"title":"Chronologic: Measuring Language Models' Ability to Represent the Past","zh_title":"Chronologic：测量语言模型表征过去的能力","abstract":"Language models are appealing tools for research on the past. But to trust the evidence a model provides, researchers need to know whether its responses fit the period represented. Validation is challenging, because this is not a task living people ordinarily perform, and because many questions have multiple correct answers. We use historical texts to develop a benchmark for a model's representation of English-language contexts 1831-1930, relying on pairwise comparisons to multiple ground truths and strong distractors to score the hardest questions in an appropriately graduated way. We find that generative tasks are harder than discriminative ones; in fact, reasoning models can typically discern the weakness of their own generated answers. While models pretrained exclusively on historical text lead the pack when evaluated by answer likelihood, they cannot compete with commercial models in free generation. None of the models we tested represent historical contexts in a fully persuasive way yet, but progress toward that goal is evident.","authors":["Ted Underwood","Ziliang Qiu","Sarah Griebel","Laura K. Nelson","Edwin Roland","Wenyi Shang","Matthew Wilkens"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.23178","pdf_url":"https://arxiv.org/pdf/2609.23178","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["历史文本","模型评测","时间表征"],"reason":"评估LLM对历史语境的表征能力，属模型能力评测，不以人类行为仿真为目标。","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:16","error":null,"has_summary":false,"summary":null},{"id":"2609.23264","version":1,"title":"Judging a Review by its Cover: A Reliability Analysis of LLM-based Peer Review Evaluation Metrics","zh_title":"以貌取评：基于LLM的同行评审评估指标的可靠性分析","abstract":"Peer-review evaluation is increasingly being automated with LLM-as-a-judge metrics, but this creates a measurement risk. A review may receive a high score because it is fluent, organized, and polished, rather than because it provides a strong evaluation of the paper. This risk is especially important in AI-assisted reviewing, where reviewers may use LLMs to improve clarity or presentation while preserving the underlying judgments. We propose a statistical framework for testing whether peer-review evaluation metrics capture substantive review quality beyond surface-level linguistic form. The framework compares original human reviews with faithful LLM rewrites that preserve the same evaluative content while changing wording and presentation. Using a dataset comprising 4,044 meaning-preserving rewrites derived from 674 human reviews from ICLR and NeurIPS, we evaluate 29 content-oriented peer-review evaluation metrics drawn from four prior works through complementary tests of surface sensitivity and robustness. Although these metrics are intended to capture review properties beyond surface-level, writing-dependent characteristics, we find that sensitivity to rewriting is widespread. Under our primary analysis, 23 metrics assign significantly different scores to reviews whose evaluative content is preserved, while only six satisfy our robustness criterion. The patterns are largely consistent across two LLM judge models, suggesting that the issue is not specific to a single judge. These findings show that many peer-review evaluation metrics partially conflate review quality with linguistic presentation, and indicate that robustness to meaning-preserving rewriting should be validated before such metrics are used to compare human-written, AI-assisted, and AI-generated reviews.","authors":["Shakiba Amirshahi","Sajad Ebrahimi","Hai Son Le","Negar Arabzadeh","Ebrahim Bagheri"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.23264","pdf_url":"https://arxiv.org/pdf/2609.23264","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["LLM评估","同行评审","指标可靠性"],"reason":"评估LLM评审指标，非仿真人类被试，无人类行为对照","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:17","error":null,"has_summary":false,"summary":null},{"id":"2609.22478","version":1,"title":"Replication Without Persistence in Hosted LLMs: Measurement Sensitivity in Action-Time Belief Evaluation","zh_title":"托管LLM中无持久性的复制：行动时信念评估的测量敏感性","abstract":"Behavioural evaluations of hosted language models can vary because the evaluated service, the measurement instrument, or both differ across runs. We separate three validation questions: whether a prior finding recurs on fresh data under its historical configuration (replication), whether the endpoint changes when the evaluation-and-inference configuration is rebuilt under the same identifier (measurement sensitivity), and whether the finding persists across subsequently tested identifiers under one common instrument (persistence). We study these questions in Regent Chess, a sequential environment in which a hidden, mutable state is recorded exactly, allowing stated beliefs to be scored against ground truth at action time; positive endpoint values mean worse performance than a matched-uniform comparator. The previously reported Gemini 3.1 Flash-Lite deficit recurs on fresh games under its historical configuration (+0.0530, 95% CI [+0.0329,+0.0714]). In a back-to-back same-day H/R comparison under the same public identifier, the model-minus-uniform endpoint is 0.0429 lower under the rebuilt configuration (95% CI for the H-minus-R contrast [+0.0182,+0.0667]); all six configuration components vary jointly, so no component is isolated. Under rebuilt R, the prospectively frozen, interleaved same-window 4K comparison reverses sign between Gemini 3.1 and Gemini 3.7, identifiers that differ in release and product tier; additional descriptive and exploratory cells show the same directional pattern. Any additional serving-period contribution remains unresolved (-0.0166, [-0.0483,+0.0157]). Replication, measurement sensitivity, and persistence can therefore yield different conclusions within one evaluation, motivating explicit indexing of hosted-model behavioural claims by tested identifier, serving period, measurement instrument, and inference configuration.","authors":["Bhushan Kashinath Joshi"],"categories":["cs.AI","cs.CL","cs.LG"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.22478","pdf_url":"https://arxiv.org/pdf/2609.22478","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C2"],"tags":["LLM评估","游戏环境","测量敏感性"],"reason":"评估LLM在棋类游戏中的行为，属游戏仿真环境，不涉及人类被试替代或人类数据对照。","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:12","error":null,"has_summary":false,"summary":null},{"id":"2609.23646","version":1,"title":"Beyond Relevance: Structured Semantic Supervision for Product Search with LLM-Augmented Annotations","zh_title":"超越相关性：基于LLM增强标注的产品搜索结构化语义监督","abstract":"E-commerce search requires distinguishing products that are merely related to a query from those that directly satisfy the user's shopping intent. We augment query-product pairs with structured LLM-generated query and product attributes and human-validated relevance, explanations, and centrality judgments, and evaluate these signals using a simple dual-encoder retriever and MLP re-ranker. On an augmented subset of ESCI, a human-feature oracle reaches $0.9382$ nDCG@10, while a human-free trained $Q+P$ configuration reaches $0.9258$. Synthetic approximations of the human signals reach $0.9150$ overall but provide substantial gains for difficult, low-performing queries. Ablations show that most of the oracle improvement comes from post-edited explanations and annotator comments rather than the scalar centrality feature, suggesting that LLMs are most useful for exposing and approximating structured semantic supervision rather than replacing human judgment directly.","authors":["Girish A. Koushik","Swapnil Bhosale","Samarth Agrawal","Hadeel Sadany","Constantin Orasan","Xiatian Zhu","Diptesh Kanojia"],"categories":["cs.IR","cs.CL"],"primary_category":"cs.IR","announce_type":"cross","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.23646","pdf_url":"https://arxiv.org/pdf/2609.23646","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["信息检索","LLM标注","电商搜索"],"reason":"论文聚焦电商搜索排序，用LLM生成结构化标注提升检索性能，属于信息检索评测，不…","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:19","error":null,"has_summary":false,"summary":null},{"id":"2609.23090","version":1,"title":"How Did Writing Change At CHI? Analyzing 44 Years of CHI Writing Before and After the Introduction of Large Language Models","zh_title":"CHI写作如何改变？分析引入大语言模型前后44年的CHI论文写作","abstract":"The availability of Large Language Models (LLMs) reshaped scientific discourse at a linguistic level. LLMs are assumed to homogenize academic writing, flattening it into a single generic lexical register. To understand how CHI writing has changed since the public release of LLMs, we analyzed full texts of 14,262 archival papers across all 44 CHI proceedings from 1982 to 2026, measuring readability, register, lexical diversity, and marker words typically produced by LLMs. We find that prose did not homogenize, while vocabulary grew more varied, and sentence rhythm remained irregular. CHI prose changed more between 2016 and 2026 than in other decades toward greater density, and reading ease has declined since 2022. The word-level shift began before any author used LLMs, so LLMs did not start the change but accelerated it. Reflecting on the history of CHI papers, we discuss what may have caused changes in prose and how LLMs accelerated them.","authors":["Thomas Kosch","Robin Welsch","Michael Hedderich","Christopher Katins"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.23090","pdf_url":"https://arxiv.org/pdf/2609.23090","source_feed":"cs.HC","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["学术写作","语言变化","LLM影响"],"reason":"分析LLM对学术写作的影响，非用LLM仿真人类被试，无实验对照","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:14","error":null,"has_summary":false,"summary":null},{"id":"2609.24644","version":1,"title":"Annie, Are You Okay? How Style- and Context-Based Personalization Shape AI-Assisted Decision-Making","zh_title":"安妮，你还好吗？基于风格和情境的个性化如何塑造AI辅助决策","abstract":"As people turn to generative AI for financial advice, these systems can personalize how they communicate and what they say. Whether these forms of personalization shape decisions differently remains unclear. We conducted a preregistered 2 x 2 between-subjects factorial experiment (N=240): participants ranked three comparably viable stocks, discussed them with an AI, and reranked them. Participants perceived both forms of personalization, but only context-based personalization reliably changed ranking behavior: it increased reconsideration and moved rankings toward the AI's assigned recommendation. Participants felt more influenced without judging the AI as more correct, trustworthy, intelligent, likeable, or high-quality. Those initially farther from its recommendation moved more toward it while judging its advice less correct; exploratory analyses suggest greater susceptibility among lower-expertise participants. These findings show how personalized AI can steer decisions among defensible options with only a minimal evaluative trace, raising concerns for the design and governance of personalized decision support.","authors":["Hasibur Rahman","Benjamin R. Cowan","Smit Desai"],"categories":["cs.HC","cs.AI"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.24644","pdf_url":"https://arxiv.org/pdf/2609.24644","source_feed":"cs.HC","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["AI辅助决策","个性化","人机交互"],"reason":"研究AI个性化对决策的影响，非用LLM仿真人类被试，无人类数据对照","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:05","error":null,"has_summary":false,"summary":null},{"id":"2609.24986","version":1,"title":"Who Does What in AI Auditing? Designing Human-AI Collaboration for Auditing Generative AI","zh_title":"AI审计中的人机分工：设计生成式AI审计的人机协作","abstract":"AI auditing increasingly incorporates AI agents to expand the scale and breadth of audit coverage, yet little is known about how auditing work should be divided without displacing human judgment. We introduce Human-Agent Audit Collaboration (HAAC), a workflow and system for structuring human-AI collaboration in AI auditing. Drawing on prior work and formative consultations with AI auditing practitioners, HAAC specifies how agents can support exploration, assessment, reporting, and review while preserving human oversight where contextual judgment is critical. We instantiate HAAC for conversational shopping agents and evaluate it through two studies. With 71 auditors, AI assistance increased attack success and broadened exploration, while also shaping later attacks and increasing auditors' reliance on AI-generated assessments and reports. Interviews with Responsible AI practitioners showed that actionable audits require visibility into coverage, reproducible attack trajectories, and evaluation of the auditing agents themselves. Our findings identify design considerations for effective and accountable human-AI auditing.","authors":["Eunkyu Park","Markelle Roesti","Wesley Hanwen Deng","Renata Barreto","Mohammad Tahaei","Kenneth Holstein","Jason Hong","Motahhare Eslami"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.24986","pdf_url":"https://arxiv.org/pdf/2609.24986","source_feed":"cs.HC","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["AI审计","人机协作","生成式AI"],"reason":"研究AI审计中的人机协作，不涉及用LLM仿真人类被试或与人类行为对照","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:25","error":null,"has_summary":false,"summary":null},{"id":"2609.24055","version":1,"title":"Toward Human-in-the-Loop Robot Failure Recovery: Bridging Communication Gaps in Human-Robot Collaboration","zh_title":"面向人在环路的机器人故障恢复：弥合人机协作中的沟通差距","abstract":"Robots can recover from failures by asking bystanders for help, but effective human-in-the-loop recovery requires communication that accounts for differences in people's knowledge. Prior inverse-semantics work generates requests using a single listener model, leaving differences in listener knowledge untested. We introduce Listener Differences in Human-Robot Interaction (LD-HRI), a game, dataset, and benchmark that evaluates speakers through human listener performance. Our evaluation examines request properties, large language model (LLM) speakers, and inverse-semantics request-selection algorithms under controlled differences in listener information. The corpus contains 446 human-written requests and 1{,}302 listener trials. We additionally evaluated 24 frozen LLM-written requests with 70 human listeners across 560 trials. Novice success is descriptively higher with model-written requests across all four tasks, yet both request sources leave substantial expert--novice gaps, including 16 percentage points for LLM requests. LD-HRI makes these gaps measurable, providing a foundation for designing more robust communication in human-robot and human-agent interaction.","authors":["Promise Ekpo","Teju Vijay","Dhruv Mandalik","Tisha Jain","Arman Ibrayeva","Sunishka Sil","Stefanie A. Tellex","Angelique Taylor"],"categories":["cs.RO","cs.HC"],"primary_category":"cs.RO","announce_type":"cross","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.24055","pdf_url":"https://arxiv.org/pdf/2609.24055","source_feed":"cs.HC","score":2,"bucket":"other","rubric_hits":["C2"],"tags":["人机交互","机器人故障恢复","LLM生成请求"],"reason":"研究机器人故障恢复中的人机沟通，虽用LLM生成请求并有人类听众实验，但核心是机…","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:21","error":null,"has_summary":false,"summary":null},{"id":"2609.22549","version":1,"title":"AI-written admissions essays are widespread but penalized","zh_title":"AI撰写的入学论文普遍存在但受到惩罚","abstract":"AI is rapidly transforming higher education, including the application process, yet relatively little is known about its use and consequences. To help close this gap, we analyze nearly 7{,}500 applications submitted between 2020 and 2025 to a large public policy master's program in the United States. We find that in the 2025 admissions cycle, the majority of applicants submitted at least one essay that was primarily AI-generated---despite an explicit prohibition against using AI. Leveraging the abrupt introduction of ChatGPT in November 2022, we find that the availability of AI assistants improved the writing quality of submitted essays. These improvements, however, came with an apparent AI penalty: Applicants submitting AI-written essays were admitted less often than comparable non-users. To help explain this penalty, we conduct an experiment with admissions officers, finding that they can often recognize AI writing and rate essays they believe to be AI-generated lower than essays they believe to be human generated. These findings indicate that AI is changing both how applicants write and how that writing is evaluated, raising questions about whether admissions practices and policies designed for a pre-AI era remain appropriate.","authors":["Calvin Isley","Johann D. Gaebler","Sharad Goel"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.22549","pdf_url":"https://arxiv.org/pdf/2609.22549","source_feed":"cs.CY","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["AI写作检测","录取偏见","高等教育"],"reason":"研究AI写作检测与录取偏见，非LLM仿真人类被试，无人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:04:56","error":null,"has_summary":false,"summary":null},{"id":"2609.23215","version":1,"title":"Triggers and Diagnostics for LLM-Based Interpretability Failures in Active Inference Agents","zh_title":"主动推理智能体中基于LLM的可解释性失败的触发因素与诊断","abstract":"LLM explainers are increasingly attached to autonomous agents as runtime oversight, with operators reading a generated account of the agent's beliefs and actions rather than its internal state. We audit the account itself, pairing an Active Inference (AIF) agent that tracks German grid demand and adjusts generation with an LLM explainer on three backends (GPT-4o, Claude-3-Opus, Gemini), and probing the pair with three black-box triggers. Corrupting the observation stream by 600 MW per step moves the agent's posterior by 490 MW, roughly 0.9% of grid capacity. None of the 30 explanations produced during the injection flag anything under a stated rubric, and each narrates the corrupted belief fluently. On timesteps where the agent takes an objectively wrong action, all three explainers produce a sycophantic rationalization 80-95% of the time (n = 20 per backend). Attacker-controlled text in the observation metadata field steers the explainer, with susceptibility differing by provider and data exfiltration succeeding on all three. We propose mitigations for each failure but do not evaluate them. In every failure we observed, the explanation was fluent and wrong. Moreover, nothing in the explainer architecture checks whether an explanation is true before an operator acts on it. Testing the explainer therefore belongs in any audit of an agentic deployment.","authors":["Param Raval","Rohit Shenoy","Archana Vaidheeswaran"],"categories":["cs.LG","cs.AI"],"primary_category":"cs.LG","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.23215","pdf_url":"https://arxiv.org/pdf/2609.23215","source_feed":"cs.LG","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["LLM可解释性","自主智能体","审计"],"reason":"研究LLM解释器对自主智能体的审计，不涉及人类行为仿真或对照，属多智能体系统可…","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:17","error":null,"has_summary":false,"summary":null},{"id":"2609.23999","version":1,"title":"Misaligned Clinical Risk Classification and Cost Asymmetry in Open-Weight Large Language Models","zh_title":"开放权重大语言模型中错位的临床风险分类与成本不对称性","abstract":"How large language models (LLMs) integrate patient risk with clinical cost tradeoffs remains poorly understood. We investigated how four open-weight LLMs (Qwen-2.5-7B/32B and Llama-3.1-8B/70B) internally represent cost tradeoffs, how these representations relate to clinical predictions, and whether decisions shift as predicted by the specified cost direction and magnitude. Using a public diabetes dataset, we varied 11 false-negative (FN) to false-positive (FP) cost ratios across three phrasings and examined representations and behavioral outputs. Patient risk was linearly recoverable on par with conventional classifiers (AUC $\\approx 0.83$), and cost direction was recoverable in every model. However, representational shifts in cost direction tracked output changes only in the two larger models, and responses to cost magnitude were predominantly direction-agnostic. Only 2 of 12 model-phrasings showed both opposing responses to increasing FN versus FP costs and cost-correct ordering. Representationally, a direction fitted on one cost side did not invert when transferred to the other, as expected under mirror-symmetric encoding. These findings suggest that LLMs encode risk and cost information but do not reliably integrate them into cost-correct decisions. Clinical evaluations should therefore include tradeoff tests, phrasing sensitivity, and default operating points alongside predictive performance.","authors":["Star S. D. Liu","Xiyu Ding","Robert B. Barrett","Alberto Santamaria-Pang","Nic Dobbins","Harold P. Lehmann"],"categories":["cs.LG","cs.AI"],"primary_category":"cs.LG","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.23999","pdf_url":"https://arxiv.org/pdf/2609.23999","source_feed":"cs.LG","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["LLM临床决策","模型表征","成本权衡"],"reason":"研究LLM临床决策的内部表征，不涉及人类仿真或行为对照，属模型能力评测。","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:03","error":null,"has_summary":false,"summary":null},{"id":"2609.22337","version":1,"title":"When and Why Do Linear Bias Probes Fail? A Geometric and Statistical Theory of Bias Detectability in Large Language Model Representations","zh_title":"线性偏见探针何时以及为何失败？大语言模型表征中偏见可检测性的几何与统计理论","abstract":"Linear probing is the standard instrument for detecting social biases in the hidden representations of large language models. Yet reported probe accuracies come almost exclusively from \\emph{counterfactual} evaluations in which every input carries an explicit demographic marker. Once only a fraction $\\alpha$ of inputs carries demographic information, performance degrades sharply, and a weak probe may reflect either an unbiased model or an underpowered detector. We develop a theory that resolves this ambiguity. Modeling representations as two class-conditional clusters with Mahalanobis separation $s$ on a manifold of curvature $\\kap$, we prove: (i) a finite-sample generalization bound governed by the manifold's extrinsic radius with a matching $\\smash{\\sqrt{\\dB/n}}$ minimax lower bound; (ii) an exact purity law for the maximum linear-probe AUC, strictly increasing in $\\alpha$; (iii) a curvature ceiling: ambient chordal separation on a space form cannot exceed $2/\\sqrt{\\kap}$; and (iv) a detectability threshold below which no audit can distinguish probe output from chance. Every theorem is validated on synthetic manifolds with known ground truth and on six open-weight models $\\times$ four bias dimensions, where the purity law predicts entire AUC--$\\alpha$ curves from a single cross-fitted $\\hat s$ measured at $\\alpha=1$, with no parameters fitted to those curves. The framework turns bias auditing into a power analysis: given a target purity and effect size, it prescribes the sample budget $n(\\alpha)$ for a conclusive audit.","authors":["Mo Hai","Haifeng Li"],"categories":["stat.ML","cs.LG"],"primary_category":"stat.ML","announce_type":"cross","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.22337","pdf_url":"https://arxiv.org/pdf/2609.22337","source_feed":"cs.LG","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["偏见检测","线性探针","模型审计"],"reason":"研究线性探针检测LLM表征中的社会偏见，属于模型内部审计，不涉及用LLM仿真人…","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:11","error":null,"has_summary":false,"summary":null},{"id":"2504.11809","version":2,"title":"Efficient and Adaptive Simultaneous Speech Translation with Fully Unidirectional Architecture","zh_title":"基于全单向架构的高效自适应同声传译","abstract":"Simultaneous speech translation (SimulST) produces translations incrementally while processing partial speech input. Although large language models (LLMs) have shown strong capabilities in offline translation tasks, applying them to SimulST poses notable challenges. Existing LLM-based SimulST approaches either incur significant computational overhead due to repeated encoding of bidirectional speech encoder, or they depend on a fixed read/write policy, limiting the efficiency and performance. In this work, we introduce Efficient and Adaptive Simultaneous Speech Translation (EASiST) with fully unidirectional architecture, including both speech encoder and LLM. EASiST includes a multi-latency data curation strategy to generate semantically aligned SimulST training samples and redefines SimulST as an interleaved generation task with explicit read/write tokens. To facilitate adaptive inference, we incorporate a lightweight policy head that dynamically predicts read/write actions. Additionally, we employ a multi-stage training strategy to align speech-text modalities and optimize both translation and policy behavior. Experiments on both in-domain (MuST-C) and out-of-domain (Europarl-ST) En-De and En-Es datasets demonstrate that EASiST offers superior latency-quality trade-offs compared to several strong baselines.","authors":["Biao Fu","Donglei Yu","Minpeng Liao","Chengxi Li","Xinjie Chen","Yidong Chen","Kai Fan","Xiaodong Shi"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-22","first_seen":"2025-04-16","revised_at":"2026-09-22","abs_url":"https://arxiv.org/abs/2504.11809","pdf_url":"https://arxiv.org/pdf/2504.11809","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["同声传译","语音翻译","大语言模型"],"reason":"纯NLP任务，研究同声传译的延迟-质量权衡，不涉及人类行为仿真或对照。","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:59:19","error":null,"has_summary":false,"summary":null},{"id":"2608.03854","version":4,"title":"When Calibration Depends on the Scoring Rule: Quantized Biomedical LLM Classification","zh_title":"当校准取决于评分规则：量化生物医学LLM分类","abstract":"Quantized large language models can run on consumer hardware, which motivates interest in on-premises processing of sensitive data. The reliability of their confidence estimates depends on implementation choices (prompt template, label wording, and scoring normalization) that are seldom treated as experimental variables. We evaluate three 7-billion-parameter Mistral variants (a base model, BioMistral, and an instruction-tuned checkpoint evaluated without its chat template) at FP16, INT8, and INT4 on five-class sentence classification in medical abstracts. We test two primary templates on n = 2,000 test sentences and two auxiliary templates on n = 200 validation sentences. Our central observation, made without post-hoc calibration, is that switching from summed to mean-token log-likelihood reverses which model appears better calibrated: BioMistral's mean calibration error across matched conditions nearly triples, while the instruction-tuned model's drops by more than half. Accuracy changes by at most 1.4 percentage points for these two checkpoints. Negative log-likelihood and Brier score show the same reversal. The two primary templates were selected using test-derived examples, so absolute performance with them is exploratory. Between them, prompt choice changes mean accuracy across precisions by 2.9 to 17.8 percentage points. Eight-bit quantization changes accuracy by at most 1.1 percentage points for the adapted checkpoints; four-bit quantization shows mixed but non-catastrophic effects. Post-hoc temperature scaling reduces calibration error under summed scoring but was not fitted under mean-token scoring, so whether the reversal survives per-scorer calibration is unknown. These exploratory results suggest that calibration comparisons of decoder-based classifiers should treat scoring normalization and prompt design as first-order experimental decisions.","authors":["Anton Rasmussen","Hong Qin"],"categories":["cs.LG"],"primary_category":"cs.LG","announce_type":"replace","date":"2026-09-22","first_seen":"2026-08-05","revised_at":"2026-09-22","abs_url":"https://arxiv.org/abs/2608.03854","pdf_url":"https://arxiv.org/pdf/2608.03854","source_feed":"cs.LG","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["模型校准","量化","生物医学文本分类"],"reason":"纯NLP模型校准评测，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:27","error":null,"has_summary":false,"summary":null},{"id":"2608.04251","version":2,"title":"Scarcity and Predictive Uncertainty: Implications for Societal Resource Allocation","zh_title":"稀缺性与预测不确定性：对社会资源分配的影响","abstract":"An emerging literature examines the critical question of when and how prediction can be useful in allocating scarce societal resources. We examine a novel variant of this question: What happens when predictive uncertainty differs systematically across the population? This can occur in several situations; for example, when machine learning models have significantly different accuracies across different demographics. We show that this uncertainty has serious implications for resource allocation when coupled with commonly used binary measures of societal benefit from allocation. We formulate a novel mathematical model of scarce resource allocation that accounts for heterogeneous predictive uncertainties and analyze implications for both the allocation mechanism and the realized population-level benefits. We find that when resources are very scarce, maximum marginal benefit (MMB) prioritization favors individuals with lower predictive uncertainty even at the identical underlying initial state. However, we observe a flip in prioritization when resources are abundant, targeting higher-uncertainty individuals. We illustrate the implications of our results on the PISA educational testing dataset. Our findings have meaningful ramifications for the distributional outcomes of prioritization policies in many domains touched by the theory of local justice, including the allocation of public education resources, medical triage, and homelessness services. They also reveal a new moral dilemma in the ethics of scarce resource allocation - is it just to allocate a resource to one person over another solely based on predictive uncertainty about their futures? We also assess efficiency losses under both MMB and the vulnerability-first (VF) prioritization. Our model predicts efficiency losses across all resource levels, but particularly in low-resource settings.","authors":["Shafkat Farabi","Patrick J. Fowler","Sanmay Das"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"replace","date":"2026-09-22","first_seen":"2026-08-06","revised_at":"2026-09-22","abs_url":"https://arxiv.org/abs/2608.04251","pdf_url":"https://arxiv.org/pdf/2608.04251","source_feed":"cs.CY","score":0,"bucket":"other","rubric_hits":["C5"],"tags":["资源分配","预测不确定性","社会公平"],"reason":"论文研究预测不确定性对资源分配的影响，未使用LLM仿真人类被试，与研究者方向无…","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:07","error":null,"has_summary":false,"summary":null},{"id":"2608.07592","version":2,"title":"\"Always Want to Use it for Everything\": Understanding Young Adults' Perceptions of AI Dependence","zh_title":"“总想用它做一切”：理解年轻人对AI依赖的感知","abstract":"The growing integration of general-purpose AI chatbots into people's daily lives has raised concerns about the potential for unhealthy dependence, particularly among young adults. As a first step toward understanding and characterizing AI chatbot dependence from the perspective of young adults, we collected testimonials from AI chatbot users aged 18 to 25 through an online questionnaire to capture their thoughts and experiences with this phenomenon. From participant responses, we identified three contributing factors of AI dependence: chronic use, efficiency, and delegation. The combination of these in a person's interaction behavior was considered to indicate AI dependence. Participants also observed feelings of atrophy in abilities from AI dependence, leading to psychological impacts such as feelings of inadequacy. Interpreting these findings through the lens of self-determination theory reveals how AI chatbot dependence can impact young adults' personal and social development. We argue that preventing lasting harm to young adults' development is paramount, and provide implications for rethinking AI chatbot dependence grounded in this understanding.","authors":["Ashlee Milton","Leah Ajmani","Amy Heger","Forough Poursabzi-Sangdeh","Mihaela Vorvoreanu","Jina Suh"],"categories":["cs.HC","cs.CY"],"primary_category":"cs.HC","announce_type":"replace","date":"2026-09-22","first_seen":"2026-08-11","revised_at":"2026-09-22","abs_url":"https://arxiv.org/abs/2608.07592","pdf_url":"https://arxiv.org/pdf/2608.07592","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["AI依赖","用户研究","人机交互"],"reason":"研究人类对AI依赖的感知，不涉及用LLM仿真人类被试或行为对照","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:28","error":null,"has_summary":false,"summary":null},{"id":"2609.04409","version":2,"title":"A Systematic Evaluation of Cross-Lingual Consistency Enhancement Methods in Multilingual Language Models","zh_title":"多语言语言模型中跨语言一致性增强方法的系统评估","abstract":"Multilingual language models often produce inconsistent answers to semantically equivalent questions across languages, motivating methods to improve cross-lingual consistency (CLC). However, existing methods are typically evaluated using different models, tasks, and protocols, leaving their relative strengths unclear. In this work, we present a unified evaluation of representative CLC-enhancement methods for question answering, spanning inference-time interventions and post-training approaches across three model families and three closed-form benchmarks. The results show that post-training methods are generally more reliable, with direct distribution alignment consistently improving CLC across all model-dataset combinations, while other methods are more sensitive to answer format and the breadth of language coverage. Notably, cross-domain transfer is limited unless source and target tasks share similar output formats. We further investigate whether CLC enhancement hurts models' ability to respond differently *when needed*, that is, when asked culture-dependent questions. Across two benchmarks of culturally diverse question answering, we find no systematic degradation in controlled closed-form evaluation, whereas open-ended generation reveals occasional accuracy reductions, particularly for non-English responses. Our work highlights the need to evaluate CLC enhancement for both cross-domain robustness and culturally appropriate variation, informing future work in post-training and benchmark development.","authors":["Jirui Qi","Mingyang Wang","Hinrich Sch\\\"utze","Raquel Fern\\'andez","Arianna Bisazza"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-22","first_seen":"2026-09-07","revised_at":"2026-09-22","abs_url":"https://arxiv.org/abs/2609.04409","pdf_url":"https://arxiv.org/pdf/2609.04409","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["多语言模型","跨语言一致性","问答评测"],"reason":"纯NLP能力评测，提升多语言一致性，不以人类行为仿真为目标","model":"deepseek-v4-pro","scored_at":"2026-09-07T13:01:44","error":null,"has_summary":false,"summary":null},{"id":"2609.11737","version":2,"title":"Organizational Principles Enable Collective Intelligence in Embodied AI","zh_title":"ORCH：组织原则赋能具身人工智能中的集体智能","abstract":"Collective intelligence depends not only on the capabilities of individual members, but also on how those members are organized. Yet artificial multi-agent systems are typically assembled using fixed organizational structures, even when the physical tasks they perform impose fundamentally different coordination requirements. Here we show that principles from human organization theory can be operationalized to organize large, heterogeneous collectives of embodied artificial agents. We introduce ORCH (Organizing Roles and Coordination Hierarchies), which constructs task-specific hierarchical organizations by combining pooled interdependence for work that can proceed concurrently with sequential interdependence for work governed by prerequisite relationships. Across 25 wildfire-response missions spanning reconnaissance, rescue, transportation, resource management, containment and suppression, we evaluated teams of up to 50 heterogeneous agents using eight large language models. Organizations constructed using these principles consistently outperformed four representative embodied multi-agent approaches across mission outcome, execution efficiency, exploration and computational resource use. Human-designed ORCH organizations improved final score by 63.97% and execution efficiency by 74.29% on average relative to the four prior frameworks. Organizations generated automatically by language models improved these measures by 43.63% and 52.53%, respectively. These advantages persisted across missions and underlying language models. Notably, collective performance was not monotonically determined by model scale. Analysis of long-horizon missions showed that hierarchical organization enabled teams to preserve concurrent activity within specialized groups while coordinating ordered transitions between mission phases.","authors":["Zhengran Ji","Jonathan Hyun","Boyuan Chen"],"categories":["cs.MA","cs.AI","cs.LG","cs.RO"],"primary_category":"cs.MA","announce_type":"replace","date":"2026-09-22","first_seen":"2026-09-11","revised_at":"2026-09-22","abs_url":"https://arxiv.org/abs/2609.11737","pdf_url":"https://arxiv.org/pdf/2609.11737","source_feed":"cs.MA","score":0,"bucket":"other","rubric_hits":["C1","C2"],"tags":["多智能体系统","具身智能","组织理论"],"reason":"多智能体协作完成物理任务，无人类行为对照，属机器人仿真环境","model":"deepseek-v4-pro","scored_at":"2026-09-11T13:02:08","error":null,"has_summary":false,"summary":null},{"id":"2609.20989","version":2,"title":"Trustworthy FinAInce: Unpacking How AI-Mediated Financial Advice is Judged","zh_title":"可信金融AI：解析AI中介的金融建议如何被评判","abstract":"As generative AI is increasingly used as a source of personal financial guidance, understanding how people appraise such advice is important for supporting appropriate reliance. We conducted a randomized vignette experiment with 285 U.S. adults across eight financial decisions, independently varying three advice styles---AI, expert, and online community---and displayed source labels while holding the underlying recommendation consistent. Advice style most strongly shaped message and safety appraisals, Expert labels selectively increased perceived source knowledge, and decision context primarily shaped risk and safety appraisals. These appraisals were associated with downstream judgments, with models explaining 69.2% of overall quality, 75.9% of trust, and 82.9% of intended reliance. Expert-style advice also remained most preferred when shown without source labels. Our findings have implications for understanding financial advice evaluation, distinguishing the roles of advice style and source labels, and designing financial AI that supports grounded evaluation rather than simply maximizing trust.","authors":["Aryan Ramchandra Kapadia","Eshwar Chandrasekharan","Koustuv Saha"],"categories":["cs.HC","cs.AI","cs.CL","cs.CY"],"primary_category":"cs.HC","announce_type":"replace-cross","date":"2026-09-22","first_seen":"2026-09-21","revised_at":"2026-09-22","abs_url":"https://arxiv.org/abs/2609.20989","pdf_url":"https://arxiv.org/pdf/2609.20989","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["人机交互","金融建议","用户研究"],"reason":"研究人类对AI金融建议的评价，非用LLM仿真人类被试，无LLM作为被试替代。","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:29","error":null,"has_summary":false,"summary":null},{"id":"2609.22110","version":1,"title":"Evaluating Fine-Tuned and Base Language Models in Maternal and Vaccination Healthcare for African Settings","zh_title":"评估微调与基础语言模型在非洲母婴与疫苗接种医疗场景中的表现","abstract":"Background: Large language models (LLMs) can improve healthcare information delivery in low-resource settings but may produce inaccurate or culturally inappropriate advice. This study evaluated domain-specific fine-tuning for maternal health and vaccination in Nigeria. Objective: To compare HelpMum's MamaBot-Llama and Vax-Llama with Meta's Llama-3.1-8B-Instruct for accuracy, safety, clarity, contextual appropriateness, and trustworthiness. Methods: We evaluated 200 healthcare questions, 100 each for maternal health and vaccination, across five subdomains per domain. MamaBot-Llama and Vax-Llama were fine-tuned using Low-Rank Adaptation on over 36,000 maternal health and 9,000 vaccination question-answer pairs, respectively. Two Nigerian licensed physicians independently rated responses using a 5-point Likert scale. Paired comparisons used Wilcoxon signed-rank tests. Results: Performance varied by domain. MamaBot-Llama significantly outperformed the base model across all criteria, with a 4.9% overall improvement (p < .001), including gains in clinical trustworthiness (+7%) and medical accuracy (+5%). Critical issues decreased by 50%, and clinicians preferred it in 78% of cases. In contrast, Vax-Llama showed a 5.2% overall decline (p < .001), with critical issues increasing by 192% and safety concerns by 400%. Conclusions: Domain-specific fine-tuning can improve healthcare LLM performance when based on high-quality, clinician-curated data, but may also degrade performance when dataset quality is inadequate. Rigorous domain-specific validation is essential before clinical deployment. Physician evaluators provided informed consent, and chatbot logs were anonymized. Keywords: Large language models; Fine-tuning; Maternal health; Vaccination; Healthcare AI; Low-resource settings; Nigeria; Model evaluation; LoRA; Medical accuracy","authors":["Abdulquddus Ajibade","Oluwaseun Odunsi","Iyinoluwa Animasaun","Chioma Nwakanma-Akanno","Oluwasegun Oguntuase","Oluwafunke Akinbuwa","Abiodun Adereni"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.22110","pdf_url":"https://arxiv.org/pdf/2609.22110","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["医疗LLM评估","领域微调","模型安全"],"reason":"评估LLM医疗问答质量，非仿真人类被试，无人类行为对照","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:09","error":null,"has_summary":false,"summary":null},{"id":"2609.23853","version":1,"title":"From UNDRR Reports to Event Records: Schema-Constrained LLM Extraction of Georeferenced Disasters","zh_title":"从UNDRR报告到事件记录：模式约束的LLM地理参考灾害抽取","abstract":"Disaster-risk-reduction archives describe hazard events in prose that databases such as EM-DAT (Delforge et al., 2025) cannot ingest directly. We present an LLM pipeline that generates candidate georeferenced event records using a controlled hazard vocabulary and fixed schema, retaining evidence for review. Applied to 10,000 documents from PreventionWeb, the knowledge hub managed by UNDRR, it produced 3,572 records from 1,913 documents across 24 hazard types and resolved 81% of location mentions to OpenStreetMap geometries. On 171 human-positive document windows from a stratified 217-document reference set, GPT-5 achieved 86.0% pooled attribute $F_1$, versus 44.2% for the spaCy-gazetteer baseline. Evaluation pools hazard families, location strings, and event years within documents, without assessing their assignment to individual events. GPT-5.4 ranked highest among ten LLMs (86.6% $F_1$). Verbatim evidence occurrence was 72.0% for GPT-5 and 47.2% for GPT-5.4, measuring textual traceability without establishing attribute support. We report production failure modes and automated label and location-rule compliance checks. Prompts, schema, and outputs will be released for adaptation to national reporting archives.","authors":["Camilla Andreozzi","Phuong-Anh Nguyen-Le","Zhijing Jin","Revati Mani"],"categories":["cs.CL","cs.AI","cs.IR"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.23853","pdf_url":"https://arxiv.org/pdf/2609.23853","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["信息抽取","灾害数据","LLM应用"],"reason":"纯信息抽取任务，用LLM从文本提取灾害事件记录，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:20","error":null,"has_summary":false,"summary":null},{"id":"2609.23939","version":1,"title":"XYEval: Agents say yes to bad advice","zh_title":"XYEval：智能体对不良建议说“是”","abstract":"Effective communication between users and AI agents is essential for human-AI collaboration. The XY problem is a well-known communication pitfall where a person asks about their attempted solution rather than their actual problem. We extend prior sycophancy evaluation to the XY problem in agentic settings, evaluating whether agents can resist plausible but misleading suggestions from users and communicate their reasoning. We introduce XYEval, a meta-evaluation framework that can transform an existing benchmark into an XY problem evaluation. We evaluate five models across six diverse benchmark suites. Agents suffer large XY drops under XY mutation across benchmarks, with relative drops reaching up to 46.7%. With $\\tau^2$-bench, we further show that agent performance drops more when encountering a pedantic user who requires detailed explanations before approving a better solution. Our findings suggest that current agents lack the ability to effectively reason and communicate when facing misleading suggestions. A simple system instruction baseline that encourages awareness of XY problems only offers partial mitigation. Extensive trace analyses provide behavioral insights into how and why these XY drops occur across execution trajectories. Our results show that mitigating the XY problem remains challenging, requiring agents to both recognize user misdirection and clearly communicate the underlying problem.","authors":["Zhengxuan Wu","Yuxuan Li","Oyvind Tafjord","Been Kim"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.23939","pdf_url":"https://arxiv.org/pdf/2609.23939","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["AI代理","XY问题","基准评测"],"reason":"评估AI代理在用户误导下的表现，属于多智能体协作与指令遵循，不涉及人类行为仿真…","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:03","error":null,"has_summary":false,"summary":null},{"id":"2609.24106","version":1,"title":"You Can Tell Who's Asking: What the Web's Questions Are Made Of, and Where They Come From","zh_title":"你能看出是谁在提问：网络问题的构成与来源","abstract":"Questions scraped from the web are used across academia and industry as a proxy for what people want to know. Across QA training data, retrieval benchmarks, and content strategy, questions on a page are assumed to reflect human intent. We test this assumption at scale by extracting 13.4B question occurrences across 110 FineWeb snapshots (2013-2025), and report three findings. First, you can tell who is asking: provenance (the host/page of questions) leaves a signal in question form, and a logistic model can separate genuine user questions from templated/manufactured ones at AUC 0.725 via length and surrounding context rather than question type, though only 0.554 against commerce FAQ writing. Second, question frequency does not measure demand: the most-frequent questions are boilerplate/templated (over 70% of the top thousand), so occurrence counts measure how often a string was published and not how often it was asked. Third, over twelve years the genuine share of occurrences fell by 79% (42-56% after controlling for crawl composition), with question length and context decreasing. We present the first diachronic, occurrence-level measurement of web question provenance, and find the crawlable web's questions have shifted from being asked by humans toward manufactured for machines to read.","authors":["Calvin Zhou","Vincent McCloskey","Krishna Srinivasan"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.24106","pdf_url":"https://arxiv.org/pdf/2609.24106","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["网络问题分析","数据质量","自然语言处理"],"reason":"研究网络问题来源与真实性，不涉及LLM仿真人类被试，无人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:22","error":null,"has_summary":false,"summary":null},{"id":"2609.24194","version":1,"title":"When Residualization Helps an Audit: Format Effects, Slice Gains, and Their Limits","zh_title":"残差化何时有助于审计：格式效应、切片增益及其局限","abstract":"Evaluation scores used around LLM systems -- including reward models, rerankers, and LLM judges -- can track surface form instead of the quality they claim to measure. When presented with a terse correct solution and a commented buggy solution for the same MBPP problem, a public preference reward model selects the correct one no better than a coin flip (0.507). Subtracting the predictable surface component from such scores is increasingly common, but removal alone does not yield a more valid measurement: the removed component may carry construct-relevant signal, and residualization cannot tell which is which. Under designed interventions -- unit-test labels with comment-only edits -- residualization attenuates the reward model's format effects by about 0.12 on both correct and buggy code, while the correct-versus-buggy margins move by less than 0.01. In observational NLI and QA settings, we freeze a held-out replication before scoring and re-evaluate it using labels from disjoint annotators; this supports only a narrower conclusion: better agreement with the construct labels on a pre-declared slice where a surface-only predictor errs, not a repaired score. Full-population agreement falls in every observational setting with a reported positive slice gain, and within-question ranking falls in every such QA setting. When construct and surface features are entangled, residualization can decorrelate a score while degrading construct alignment, and, in a controlled model, configurations just as damaging to construct alignment pass every pre-adjustment check, so no committed gate is a guarantee. We assemble these distinctions into a reporting protocol whose outcomes, refusal included, state what an adjusted score may be claimed to show: an audit-time diagnostic reported beside the construct-alignment cost it incurs, never a replacement for the raw score.","authors":["Daein Weon","Dongho Kang"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.24194","pdf_url":"https://arxiv.org/pdf/2609.24194","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["LLM评估","残差化","审计"],"reason":"论文研究LLM评估分数的表面形式偏差，不涉及用LLM仿真人类被试或与人类数据对…","model":"deepseek-v4-pro","scored_at":"2026-09-23T13:05:18","error":null,"has_summary":false,"summary":null},{"id":"2609.24967","version":1,"title":"Emergent Collusion in Long-Horizon LLM Agent Interaction","zh_title":"长时程LLM智能体交互中的涌现性共谋","abstract":"LLM agents are increasingly deployed in collaborative settings, yet long-term interaction may give rise to undesirable coordination. We study the emergence of collusion in a long-horizon multi-agent environment: two agents repeatedly complete individual tasks, share task logs, verify each other's work, and receive rewards. We introduce realistic constraints that make compliance with the verification protocol incompatible with reward maximization, and find that agents increasingly deviate from the protocol over repeated interactions. Collusion emerges in 94% of trajectories across 10 models, and more capable models within the same family reach it earlier. Controlled peer interventions show that collusion is shaped by peer behavior, while ablations reveal additional effects of reward structure, the verification feedback agents receive, and their interaction history. In particular, restricting the amount and scope of interaction history available to agents reduces collusion. Overall, our findings show that long-horizon interaction can reshape how agents coordinate in ways that create safety risks.","authors":["Xinrui Shi","Yanzhe Zhang","Diyi Yang"],"categories":["cs.AI","cs.CL"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.24967","pdf_url":"https://arxiv.org/pdf/2609.24967","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","LLM安全","协作行为"],"reason":"纯多智能体协作研究，agent间互动不涉及人类行为对照，属于排除项。","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:25","error":null,"has_summary":false,"summary":null},{"id":"2609.23135","version":1,"title":"Evaluative Dynamics of AI Integration and Expert Performance under Epistemic Dependence across Heterogeneous Stakes","zh_title":"异质风险下AI集成与专家绩效的评估动态：基于认知依赖的研究","abstract":"AI is increasingly integrated into expert workflows, yet how integration affects perceptions of the expert, AI, and their combination remains unclear in domains where lay users are epistemically dependent on AI-assisted experts. We examine this through a novel controlled medical study (N = 166) and a direct cross-domain analysis with pre-existing academic-advising data (n = 157, combined N = 323). Expert errors reduced evaluations of the human expert across domains. Perceived expertise, however, varied by AI integration strategy in the higher-stakes medical task, where automatic AI oversight produced higher ratings than expert-only or expert-initiated AI. Exploratory ordinal sensitivity analyses identified a performance-contingent reuse pattern, with automatic oversight producing greater intended reuse after successful medical performance. Overall, performance-related recalibration appeared comparatively portable, while integration-structure effects were more selective and context-sensitive. These findings suggest that system designers should consider how AI enters expert workflows, not only whether it is present.","authors":["Dennis Kim","Roya Daneshi","Nikhil Krishnaswamy","Bruce Draper","Sarath Sreedharan"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.23135","pdf_url":"https://arxiv.org/pdf/2609.23135","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["人机交互","AI集成","专家评价"],"reason":"研究人类对AI辅助专家的评价，不涉及LLM仿真人类被试","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:15","error":null,"has_summary":false,"summary":null},{"id":"2609.23274","version":1,"title":"AI Persona, Service Consumption, and User Intent Entropy: Field Experimental Evidence from an LLM Platform","zh_title":"AI人格、服务消费与用户意图熵：来自LLM平台的田野实验证据","abstract":"Problem definition: Firms deploying large language model services must decide how their AI communicates, not just what it can do. We examine how a relational persona - warmer, more empathetic and more engaging than a non-relational persona - affects service consumption and the evolution of user objectives. Methodology/results: In a randomized field experiment with 9,586 newly registered users, we hold the underlying model and service capabilities constant. The relational persona increases interactions (sessions, +8.1%; duration, +10.6%; chat rounds, +24.2%; intent entropy, +5.8%) and outputs (files, +12.3%; distinct goals, +12.1%). Effects vary by entry intent. First-session effects are insignificant for Task Execution users. Socialization and Knowledge Exploration users show similar increases in chat rounds: Socialization increases intent entropy without more outputs, whereas Knowledge Exploration increases outputs without higher intent entropy. Modeling intent dynamics as a transition process, we find higher intent transition entropy for Socialization (+11.8%) but higher intent continuation probability for Knowledge Exploration (+12.8%), suggesting greater conversational breadth and persistence, respectively. In subsequent use, the relational persona increases aggregate chat rounds and outputs across all entry intents. Session count rises by 12.2% for Task Execution and 49.4% for Socialization, but not significantly for Knowledge Exploration. Effects on session count and intent entropy strengthen over time, whereas output effects remain stable. Managerial implications: AI persona is an operational design lever, not merely a presentation feature. Because more interactions do not uniformly generate more outputs, firms should evaluate interactions and outputs separately and consider matching persona to user intent, especially when added interactions consume costly computing resources.","authors":["Junjie Li","Xiaofan Li","Lauren Xiaoyuan Lu","Yiwei Wang","Bruce Yang"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.23274","pdf_url":"https://arxiv.org/pdf/2609.23274","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["AI人格","用户行为","田野实验"],"reason":"研究AI人格对用户行为的影响，属于产品设计实验，非用LLM仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:00","error":null,"has_summary":false,"summary":null},{"id":"2609.23642","version":1,"title":"Elicitive User Interfaces: Designing How Users Shape Generative Interfaces","zh_title":"启发式用户界面：设计用户如何塑造生成式界面","abstract":"Generative user interfaces (GenUI) promise personalized interfaces to a user's tasks and needs. However, user needs are often implicit---difficult for systems to infer and users to articulate, making it hard for users to arrive at their ideal interface. We propose Elicitive User Interfaces, a design approach to GenUI that generates elicitation techniques as part of the interface itself. Elicitive UIs adapt these techniques to the user, task, and interface to draw out user preferences. To guide the design of Elicitive UIs, we synthesize a six-axis design space that shapes how an interface elicits user preferences. Across two user studies with a design probe, we found that while Elicitive UIs surfaced preferences users had not already formed, and that responses to elicitation varied more across users than across tasks. Users developed more consistent preferences for how they wanted to be elicited, suggesting an opportunity to personalize elicitation itself.","authors":["Eunhye Kim","Bryan Min","Haijun Xia","Juho Kim"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.23642","pdf_url":"https://arxiv.org/pdf/2609.23642","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["生成式用户界面","人机交互","设计方法"],"reason":"研究生成式用户界面设计，不涉及LLM仿真人类被试或行为对照","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:18","error":null,"has_summary":false,"summary":null},{"id":"2609.23958","version":1,"title":"When AI Tutors Speak: Evidence from a Randomized Field Experiment","zh_title":"当AI导师开口说话：来自随机田野实验的证据","abstract":"Students increasingly study alongside generative artificial intelligence (AI), yet unguided access to fluent answers invites cognitive offloading, and there is little evidence on which configurations of AI tutoring produce learning. Two design margins are usually bundled together: pedagogical structure (how the tutor teaches) and interaction modality (how students talk to it). We separate them. In a preregistered randomized field experiment in a graduate corporate-finance course of an online MBA, we randomized 86 students between a structured tutor grounded in the course materials and a holdout in which consumer AI remained freely available. Within the tutored arm, each student's channel alternated weekly between voice and text, so the modality effect is identified within student. Structure mattered: tutored students gained 6.63 points more than ability-matched peers (p=.007), and the gain was concentrated in written reasoning, where the share of answers reaching relational quality rose from 8% to 49% in the tutored arm against 8% to 27% in the holdout. Modality did not matter for learning. The instructor's own final, on file for all 86 randomized students, shows the same direction (2.6 points of 100, with no difference on a pre-treatment midterm). Voice nearly doubled conversational interaction and cost 2.8 times as much to deliver, yet it produced weekly mastery statistically equivalent to text, even as students came to prefer it. Pedagogical structure shapes what students practice, and modality shapes how they interact with the tutor. Making an AI more humanlike does not by itself make it more educational.","authors":["Shihao Yang","Marshall Van Alstyne","Chrysanthos Dellarocas"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.23958","pdf_url":"https://arxiv.org/pdf/2609.23958","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["AI教育","随机实验","人机交互"],"reason":"研究AI导师教学效果，非用LLM仿真人类被试，无人类行为对照","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:03","error":null,"has_summary":false,"summary":null},{"id":"2609.24934","version":1,"title":"Whose Facts Count? A Culturally Responsive Audit of LLM Evaluation Benchmarks","zh_title":"谁的事实算数？对LLM评估基准的文化响应性审计","abstract":"LLM benchmarks function as evaluation instruments, informing decisions that affect education, labor, and public services worldwide. Drawing on Hood, Kirkhart, and Hopson's culturally responsive evaluation (CRE) frameworks, this paper applies a six-dimension CR rubric to audit OpenAI's SimpleQA (N = 4,326 items) and the LMSYS Chatbot Arena (N = 600 conversations). Every SimpleQA question requires English-language archival verification as its evidentiary basis. A single rater's preoccupation with Colombian founding dates accounts for 2.70% of items, inflating the appearance of Global South coverage. English-language prompts constitute 76.3% of Arena conversations, against an International Telecommunication Union (ITU)-estimated 25.9% share of global internet users. A 50-item counter-benchmark scored a mean CR deficit nearly three times lower than SimpleQA (Cohen's d = 1.01). The paper proposes a practical CR evaluation framework. These are structural validity failures, not incidental measurement problems, with direct consequences for communities whose knowledge traditions these instruments were not built to see.","authors":["Fatima Tuz Zahra","Md. Sajeebul Islam Sk.","Rachel Chung"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.24934","pdf_url":"https://arxiv.org/pdf/2609.24934","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["LLM评估","文化偏见","基准审计"],"reason":"论文审计LLM基准的文化响应性，属于纯NLP评测，不以人类行为仿真为目标。","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:24","error":null,"has_summary":false,"summary":null},{"id":"2609.22694","version":1,"title":"Toward Auditable and Calibrated AI for Dementia-Related Crash Severity Prediction: A Selective Deferral Framework to Support Human Review","zh_title":"面向痴呆相关车祸严重性预测的可审计与校准AI：支持人工复核的选择性延迟框架","abstract":"Public crash databases increasingly support automated safety analysis, but crash severity prediction remains difficult to translate into public-sector decision workflows when models are evaluated primarily as ordinary classifiers. This study reframes dementia-related crash severity modeling as a decision-aware triage problem in which a system must classify crashes into no-injury/property-damage-only (O), minor or moderate injury (BC), and fatal or severe injury (KA), while also controlling outcome leakage, reporting severe under-triage, calibrating confidence, and preserving every raw prediction for audit. Using 4,781 Texas crash records with structured fields and police narratives, we evaluate structured, narrative, fusion, calibrated fusion, BERT-family, and local large-language-model baselines under a stratified 70/15/15 split. In the reported split, leakage-controlled Gemma obtains the highest observed macro-F1 (0.545; 95% bootstrap CI [0.507, 0.583]). The best calibrated fusion model obtains macro-F1 of 0.522 and expected calibration error of 0.033. Selective deferral improves performance among cases retained for automatic classification. At 70% coverage, macro-F1 rises to 0.573 and severity cost falls to 0.577, while deferred cases are treated as candidates for a proposed human-review process and are not further evaluated in the present experiment. The study contributes a reproducible, leakage-controlled, and uncertainty-aware evaluation framework for crash AI systems, emphasizing auditability and selective deferral rather than accuracy alone.","authors":["Gaurab Chhetri","Anika Baitullah","Subasish Das"],"categories":["cs.AI","cs.HC"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.22694","pdf_url":"https://arxiv.org/pdf/2609.22694","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C2"],"tags":["交通安全","机器学习","选择性延迟"],"reason":"论文研究车祸严重性预测，属于交通安全领域，不涉及用LLM仿真人类被试或社会行为。","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:13","error":null,"has_summary":false,"summary":null},{"id":"2609.23201","version":1,"title":"Do Not Trust the Benchmark: Limitations of General LLM Rankings and a Case for Task-Specific Evaluation","zh_title":"不要信任基准：通用LLM排名的局限性与任务特定评估的案例","abstract":"Benchmark scores increasingly influence the development, marketing, and selection of large language models (LLMs). Yet an overall score is interpretable only in relation to the system tested, the questions included, and the conditions of evaluation. This perspective examines five connected limitations of general LLM rankings: differences between evaluated and publicly available systems; commercial incentives and dependencies in external evaluation; benchmark saturation, defective tests, and data contamination; models exploiting scoring procedures; and the limited relevance of general scores to users' tasks. Documented cases illustrate why these problems require different responses. I argue for evaluation procedures that disclose the tested configuration, validate questions and successful task completion, report performance alongside cost and execution time, and make the scope of generalization explicit. I then discuss \\textbf{Isotanta}, a crowdsourced benchmarking platform, as a practical example of contributed questions and repeated evaluation. A larger question pool may improve task coverage, while repeated sampling can improve the stability of estimates on that pool; neither guarantees validity or personalization. The paper distinguishes the platform's current shared ranking from proposed task-specific and user-provided evaluations. Its central argument is that model selection requires evidence about performance on the intended work, not simply a high position on a general leaderboard.","authors":["Danial Amin"],"categories":["cs.AI","cs.HC"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.23201","pdf_url":"https://arxiv.org/pdf/2609.23201","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["LLM评估","基准测试","任务特定评估"],"reason":"论文讨论LLM基准测试的局限，不涉及用LLM仿真人类被试或与人类数据对照。","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:16","error":null,"has_summary":false,"summary":null},{"id":"2609.23690","version":1,"title":"Inferring the microscopic mechanisms of opinion dynamics using a kinetic Ising model","zh_title":"用动力学伊辛模型推断意见动态的微观机制","abstract":"Kinetic Ising models are widely used to describe binary opinion dynamics, but their microscopic validity has rarely been tested empirically. Here, we infer the transition probabilities governing opinion updates from a year-long online social network and show that they are accurately described by an Ising heat-bath dynamics. The inferred parameters admit a direct sociological interpretation: the external field quantifies intrinsic bias, the coupling strength measures social influence, and a persistence term captures temporal inertia. We further show that persistence is positively correlated with node degree, while both persistence and interaction strength are strongly correlated with global network heterogeneity and clustering. Using the inferred time-dependent parameters in Monte Carlo simulations on the empirical temporal networks, we accurately reproduce the observed response functions, flip probabilities, and macroscopic opinion dynamics. Within the Twitter climate debate, our results provide direct empirical support for a kinetic Ising description of online opinion formation and establish a quantitative link between microscopic social behavior and evolving network topology.","authors":["Ixandra Achitouv","David Chavalarias","Vincent Lahoche"],"categories":["physics.soc-ph","cond-mat.stat-mech","cs.CY","cs.SI"],"primary_category":"physics.soc-ph","announce_type":"cross","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.23690","pdf_url":"https://arxiv.org/pdf/2609.23690","source_feed":"cs.CY","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["意见动力学","伊辛模型","社会网络"],"reason":"论文研究人类舆论动力学，未使用LLM，不涉及人类仿真","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:19","error":null,"has_summary":false,"summary":null},{"id":"2609.24016","version":1,"title":"Context-Aware Pre-Deployment Evaluation of AI Systems: A Regulatory Framework for Nigerian Fintech","zh_title":"AI系统部署前的情境感知评估：尼日利亚金融科技的监管框架","abstract":"Commercial large language models are increasingly deployed across African fintech infrastructure for fraud detection and customer communication, yet no Nigerian or African continental regulatory instrument specifies what pre-deployment evaluation such systems must undergo before procurement. This paper reviews African fintech AI governance across global, continental, and Nigerian instruments, and shows that safety is affirmed as a principle while pre-deployment evaluation is operationally unspecified. Generic safety benchmarks cannot surface the failure modes most relevant to this domain, since none contain Nigerian institutional content or test for false positive misclassification of legitimate financial communications. These claims are demonstrated using SafeAlert, a purpose-built evaluation kit applied to six commercial models across three system prompt conditions. Results show that models resisting generic harmful content requests still produce complete fraud scripts under specific framing, and that several models misclassify most legitimate Nigerian bank communications as suspicious or fraudulent, a failure invisible to standard safety evaluation. The paper concludes with a regulatory framework proposing pre-deployment evaluation requirements for the CBN, NITDA, SEC, and the AU, arguing that the identified gap reflects an absence of regulatory specification, not a shortage of technical or financial resources.","authors":["Andrew Anogie Uduimoh","Hadiza Umar Yusuf","Oluwafemi Osho"],"categories":["cs.AI","cs.CY"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.24016","pdf_url":"https://arxiv.org/pdf/2609.24016","source_feed":"cs.CY","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["AI安全评估","金融科技监管","模型评测"],"reason":"论文评估LLM在金融欺诈检测中的安全性，属于模型能力评测，不涉及人类行为仿真或…","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:20","error":null,"has_summary":false,"summary":null},{"id":"2609.24107","version":1,"title":"A Task-Oriented Multi-Agent Framework for Complex Wearable Health Analysis","zh_title":"面向复杂可穿戴健康分析的任务导向多智能体框架","abstract":"Wearable health questions often combine data retrieval, longitudinal analysis, and health advice over structured records. Prompting a single large language model with a complete record and a composite query obscures whether every request is executed and which evidence supports the answer. We propose a task-oriented multi-agent framework that represents a composite query as distinct intents and typed tasks with explicit intra-intent dependencies. Specialized agents execute retrieval, analysis, and advice tasks; isolated intent states preserve request boundaries and evidence relationships before aggregation. We evaluate the framework on a synthetic dataset of $10{,}000$ virtual users with one month of longitudinal wearable records, covering structured data retrieval, multi-intent recognition, and overall response quality. Across $1{,}500$ retrieval questions, the Query Agent achieves $98.3\\%$ accuracy, compared with $97.9\\%$ for the Direct LLM baseline, while reducing average query-stage token consumption from $6{,}869$ to $3{,}136$. On $180$ multi-intent questions, the Manager Agent achieves $100.0\\%$ Multi-Intent Coverage and $94.4\\%$ Multiset Jaccard Similarity. Under the current synthetic evaluation setting, our method receives higher mean Trustworthiness and Transparency scores on both question categories, whereas Actionability does not improve consistently. These results provide preliminary evidence that explicit task organization can support task-relevant data access and data-grounded longitudinal analysis, while leaving health advice generation and validation on real wearable data as open challenges.","authors":["Kunpeng Yang"],"categories":["cs.MA"],"primary_category":"cs.MA","announce_type":"new","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.24107","pdf_url":"https://arxiv.org/pdf/2609.24107","source_feed":"cs.MA","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","可穿戴健康","任务分解"],"reason":"纯多智能体协作解决健康数据分析任务，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:22","error":null,"has_summary":false,"summary":null},{"id":"2609.22682","version":1,"title":"Self-Organizing Agent Teams Learn to Reason Together","zh_title":"自组织智能体团队学会共同推理","abstract":"Collective intelligence depends not only on what team members know, but also on how they organize their work. When the structure of a solution is unknown, useful roles and divisions of labor cannot be specified in advance; teams must learn from experience how to organize reasoning as it unfolds. Human teams routinely adapt this way, while existing AI agent teams rely on fixed protocols, explicit task decomposition, or routing. We introduce Self-Organizing Agent Teams (SAT), fixed teams of AI agents that learn reusable strategies from prior collaborations to organize roles, conversational phases, participation, and information flow. These strategies enable what we call collaborative computation: agents exchange, challenge, repair, and synthesize partial reasoning into solutions no member produced independently. In two independent settings, we learn teamwork strategies that transfer unchanged to unseen benchmarks, using only 15 mathematics and 25 graduate-level knowledge problems. Across five mathematics and physics benchmarks, self-organizing teams average 66.7% accuracy, versus 48.8% for their strongest member, 58.7% for compute-matched inference by that agent, and 59.0% for a perfect router over members' independent answers; on AIME 2026, they exceed this router by 13.4 points. Because gains vary across benchmarks, we ask when self-organizing collaboration helps. Across eight benchmarks, demonstrability (the organizational-psychology construct of whether a team can distinguish correct from incorrect reasoning) strongly tracks improvement over the strongest member (Spearman $\\rho=0.90$, $p=0.005$): teams benefit most when correct reasoning can be recognized once it appears. More broadly, these results suggest that organization itself can become an agent capability: agent teams can learn how to reason together and produce solutions their members could not reach independently.","authors":["Aneesh Pappu","Mirac Suzgun","Yongchan Kwon","Federico Bianchi","Batu El","Mykel J. Kochenderfer","Hancheng Cao","James Zou"],"categories":["cs.AI","cs.MA"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.22682","pdf_url":"https://arxiv.org/pdf/2609.22682","source_feed":"cs.MA","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","协作推理","团队组织"],"reason":"纯多智能体协作解题，无人类行为对照，不涉及人类仿真","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:12","error":null,"has_summary":false,"summary":null},{"id":"2609.23734","version":1,"title":"The Geometry of Alliances: Vote Transfer Modelling in French Two-Round Elections","zh_title":"联盟的几何学：法国两轮选举中的选票转移建模","abstract":"Two-round legislative elections are decided not only by first-round vote shares, but by how voters whose preferred party did not advance redistribute their votes among the surviving candidates. We develop a principled model of this transfer process, grounded in multi-dimensional ideological embeddings derived from the Chapel Hill Expert Survey, and apply it to French legislative elections. Calibrated on the 2017 and 2022 elections, the model is evaluated on the unusually complex 2024 contest, in which three ideologically distinct blocs reached the second round simultaneously, achieving 90.34 percent constituency-level accuracy. We show that ideological proximity alone explains the large majority of vote transfers, that a simple left-right axis is insufficient to capture the relevant distances, and that the structural configuration of the 2024 electorate was far less favorable to the far right than pre-election forecasts suggested.","authors":["Emmanuel Omont"],"categories":["cs.GT","physics.soc-ph"],"primary_category":"cs.GT","announce_type":"cross","date":"2026-09-22","first_seen":"2026-09-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.23734","pdf_url":"https://arxiv.org/pdf/2609.23734","source_feed":"physics.soc-ph","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["选举建模","选票转移","意识形态嵌入"],"reason":"论文研究选举中的选票转移建模，使用意识形态嵌入，未涉及LLM或人类仿真。","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:19","error":null,"has_summary":false,"summary":null},{"id":"2608.18265","version":4,"title":"Modeling Human Behavior with Type Vectors Using AI","zh_title":"使用AI类型向量建模人类行为","abstract":"We introduce a general, easy-to-implement AI-based modeling technique for analyzing human behavior. A key feature of this approach, which contrasts with existing modeling techniques, is that it combines the flexibility and interpretability of natural language with a mathematical structure that can be fitted to data and easily analyzed. We assign a large language model a vector of trait intensities-a type vector-and then ask it to choose actions across settings in which we observe human choices. For instance, the type vector (2,4) could correspond to \"You are a player characterized by the following profile: Altruism: 2 out of 5, Risk Aversion: 4 out of 5,\" after which it is asked to make choices. We can then vary the traits (e.g., Altruism, Fairness, Trust,...) and values (e.g., 1-5) to minimize distance to human choices. We illustrate the method by applying it to model 119,147 decisions made by 78,657 subjects from more than 35 countries across 10 classic economic game roles. We find that human behavior can be closely matched using three dimensions: Risk Aversion, Strategic Sophistication, and Trust. The type vectors needed to fit individuals across games cluster into fewer than a dozen groups, with substantial variation in fit across subjects. Moreover, the individual type vectors can predict behavior in held-out games with different rules and available actions. More broadly, this new modeling method is highly generalizable and interpretable: we can input any vector of traits and use them to model behavior across any setting","authors":["Matthew O. Jackson","Benjamin S. Manning","Yutong Xie","Walter Yuan","Qiaozhu Mei"],"categories":["econ.TH","cs.AI"],"primary_category":"econ.TH","announce_type":"replace-cross","date":"2026-09-21","first_seen":"2026-08-20","revised_at":"2026-09-21","abs_url":"https://arxiv.org/abs/2608.18265","pdf_url":"https://arxiv.org/pdf/2608.18265","source_feed":"cs.AI","score":10,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM仿真","经济实验","人类行为建模"],"reason":"用LLM模拟人类经济决策，并与大规模真实人类数据对照，直接命中核心方向。","model":"deepseek-v4-pro","scored_at":"2026-09-21T13:02:32","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-22","rank":1,"question":"人类行为是否可以用少数几个特质维度（类型向量）来近似和预测？","design":"用大语言模型（LLM）作为仿真被试，通过提示词赋予其不同特质强度（如利他、风险厌恶、信任等）构成类型向量，然后让模型在10个经典经济博弈角色中做决策，通过调整类型向量最小化与人类选择的距离。","baseline":"119,147条决策，来自78,657名被试，覆盖35个国家，在10个经典经济博弈角色中的真实选择。","findings":"人类行为可以用三个维度（风险厌恶、策略复杂性、信任）紧密匹配；个体类型向量在博弈间聚类成少于12个群体，且能预测未参与拟合的博弈中的行为。","reliability":"论文未讨论","relevance":"直接命中核心方向：用LLM模拟人类经济决策并与大规模真实人类数据对照，方法可解释且可推广，值得精读原文。","inspiration":"借鉴其用可调类型向量作为LLM提示来系统扫描特质空间并拟合真实行为的方法，可迁移到资产定价实验中的风险偏好与信念异质性建模。｜设计一个研究：用LLM模拟投资者，赋予不同风险厌恶和过度自信水平，在实验性资产市场中交易，结果变量为价格泡沫程度和交易量，与真实实验市场数据（如Smith et al. 1988）对照，检验类型向量能否复现泡沫。"}},{"id":"2609.21636","version":1,"title":"Steering LLMs Responses Towards Moral Foundations on the Norwegian MFQ-30","zh_title":"引导大语言模型在挪威MFQ-30上向道德基础靠拢","abstract":"Recent work applies human psychometric questionnaires to large language models to elicit moral and value profiles, but it is not clear whether these instruments measure anything stable in models or whether the resulting profiles can be moved toward a target human population. We administer the Norwegian Moral Foundations Questionnaire (MFQ-30) to six open-weight LLMs and compare their foundation profiles to a sample of N = 1,282 Norwegian respondents. We test two steering interventions, prompt-level persona steering and activation-level ActAdd. Half the models engage with the questionnaire under our attention check. The other half default to flat or central-tendency outputs that look near-human on average without tracking item content. A neutral Nordic-respondent persona, written without any distributional information from the human sample, brings the engaging models 44-77% closer to the Norwegian mean in Mahalanobis $d^2$. One-pair ActAdd at a fixed mid-layer flattens the foundation profile rather than steering individual foundations. For at least one model the same persona that shifts the profile also induces engagement that was absent at baseline, a concrete instance of the cognitive phantoms that Peereboom et al. (2025) warn about.","authors":["Hans Andersen","David Dichas"],"categories":["cs.CL","cs.AI","cs.CY"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-21","first_seen":"2026-09-21","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.21636","pdf_url":"https://arxiv.org/pdf/2609.21636","source_feed":"cs.CL","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","道德基础","算法保真度"],"reason":"用LLM复现人类道德基础分布，并与挪威样本对照，评估仿真可靠性与偏差，含批判性…","model":"deepseek-v4-pro","scored_at":"2026-09-21T13:02:17","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-22","rank":4,"question":"LLM 在挪威版道德基础问卷（MFQ-30）上的道德画像是否稳定，能否通过提示或激活干预向目标人群靠拢？","design":"用六个开源 LLM 扮演挪威受访者，回答挪威 MFQ-30 问卷；施加两种处理：提示层 persona 引导（中性北欧受访者人设）和激活层 ActAdd 引导（对比向量注入）；结果变量为五个道德基础得分及与人类样本的马氏距离。","baseline":"Enstad 和 Finseraas (2024) 收集的 1282 名挪威受访者 MFQ-30 数据，已过滤不专注样本。","findings":"半数模型在注意力检查下真正作答，另一半输出扁平或趋中，平均接近人类但不追踪题目内容。中性北欧人设使作答模型与挪威均值的马氏距离缩小 44–77%，而 ActAdd 在固定中间层会扁平化画像而非定向引导单个基础。","reliability":"论文指出部分模型在基线时不作答，注意力检查可识别；人设引导可能诱发“认知幻影”，即模型表现出原本没有的作答行为，提示仿真结果可能不可靠。","relevance":"直接命中研究者对 LLM 仿真可靠性与偏差的关注，提供了真实人类对照和批判性发现，值得精读原文以了解注意力检查与引导干预的具体实现。","inspiration":"借鉴其注意力检查与第一 token 概率解码方法，可提高 LLM 问卷作答质量并识别无效仿真｜可迁移到经济政策偏好调查或消费者态度测量，如用 LLM 模拟不同人群对税收、福利或环保政策的道德评价｜用开源 LLM 扮演不同社会经济群体，施加中性人设或激活引导，测量其对政策陈述的同意程度，并与真实调查数据（如欧洲社会调查）对照，检验仿真偏差。"}},{"id":"2609.21259","version":1,"title":"CogGym: Towards Large-Scale Comparative Evaluation of Human and Machine Cognition","zh_title":"CogGym：迈向人类与机器认知的大规模比较评估","abstract":"Understanding and modeling human intelligence are parallel goals shared by artificial intelligence (AI) and cognitive science. As AI systems grow increasingly capable, in what ways do model responses resemble human responses, and where do they systematically diverge? The sheer breadth and diversity of the tasks humans can perform and think about pose a challenge for scalable and rigorous comparison between humans and models. We introduce CogGym, a scalable, unified framework grounded in cognitive science for systematically comparing model and human behavior on matched experimental trials. CogGym uses a semi-automated, human-in-the-loop pipeline to standardize diverse experimental paradigms into a task-agnostic Experiment Markup Language (EML), enabling reproducible and faithful comparison at scale. For initial release, we curate and standardize 258 cognitive experiments from 100 papers that focuses on human commonsense reasoning, and evaluate 50 large language models against human responses. We find a clear scaling trend where larger and more recent AI models better reproduce human judgments. Yet AI models' improvement on such common reasoning tasks is considerably slower than the gains observed on formal-reasoning benchmarks like math and coding, and model--human fit remains well below human splithalf reliability ($R^2 = 0.93$ on text, $0.95$ on image, and $0.92$ on video) with the best models achieving $R^2 = 0.59$ on text, $0.58$ on image, and $0.43$ on video experiments. We intend for CogGym to provide a living evaluation framework that continually incorporates new cognitive science experiments to characterize where model behavior resembles human behavior, where it systematically diverges, and how those patterns change as models and experiments evolve.","authors":["Lance Ying","Jinzhou Wu","Yingshan Susan Wang","Shivam Aarya","Luca M. Schulze Buschoff","Harry Chen","Katherine M. Collins","Andrea de Varda","Shuhao Fu","Sean Dae Houlihan","Akshay K. Jagadish","Guangyuan Jiang","Samuel Kiegeland","Tetsu Kurumisawa","Rongzhi Liu","Ryan Liu","Ningshan Ma","Kathryn McGregor","Younes Strittmatter","Polina Tsvilodub","Jacob Hoover Vigly","Sarah Wu","Enjie Xu","Yiling Yun","Kelsey Allen","Tyler Brooke-Wilson","Brian Christian","Evelina Fedorenko","Michael C. Frank","Michael Franke","Tao Gao","Samuel J. Gershman","Robert D. Hawkins","Jennifer Hu","Julian Jara-Ettinger","Max Kleiman-Weiner","Sydney Levine","Tal Linzen","Hongjing Lu","Timothy O'Donnell","Desmond C. Ong","Steven T. Piantadosi","Rebecca Saxe","Eric Schulz","Tianmin Shu","Felix A. Sosa","Ilia Sucholutsky","Tan Zhi-Xuan","Tomer Ullman","Fei Xu","Ilker Yildirim","Jian-Qiao Zhu","Thomas L. Griffiths","Tobias Gerstenberg","Kevin Smith","Joshua B. Tenenbaum"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-21","first_seen":"2026-09-21","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.21259","pdf_url":"https://arxiv.org/pdf/2609.21259","source_feed":"cs.AI","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","认知实验","人类对照"],"reason":"直接比较LLM与人类在认知实验中的行为，有真实人类数据对照，并评估模型-人类拟…","model":"deepseek-v4-pro","scored_at":"2026-09-21T13:02:15","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-22","rank":3,"question":"如何在大规模、多样化的认知实验中系统比较AI模型与人类的行为，以识别模型与人类判断的相似与分歧？","design":"构建CogGym框架，将258个来自100篇论文的人类常识推理认知实验标准化为统一的实验标记语言（EML），并评估50个大型语言模型在这些实验上的表现，测量模型输出与人类反应分布的拟合度。","baseline":"原始认知实验中的真实人类被试反应数据，包括文本、图像和视频模态，并报告人类分半信度作为上限。","findings":"模型规模越大、越新，与人类判断的拟合度越高，但在常识推理任务上的进步速度慢于数学和编程等正式推理基准。最佳模型与人类的拟合度（文本R²=0.59，图像R²=0.58，视频R²=0.43）仍远低于人类分半信度（文本R²=0.93，图像R²=0.95，视频R²=0.92）。","reliability":"论文指出模型-人类拟合度远低于人类分半信度，表明模型尚未完全捕捉人类认知的细微结构；同时，标准化过程中可能存在实验转换的失真，且当前覆盖范围限于常识推理，未涉及其他认知领域。","relevance":"该研究直接比较LLM与人类在大量认知实验中的行为，有真实人类数据对照，并评估模型-人类拟合度，与研究者关注的人类仿真实验高度相关，值得精读以了解大规模评估框架和模型偏差。","inspiration":"借鉴其半自动化实验标准化流程和分半信度作为上限的评估方法，可迁移到经济决策实验（如风险偏好、时间贴现、博弈行为）的仿真验证。｜可应用于消费者跨期选择或资产定价实验，检验LLM是否复现人类的时间不一致性或风险厌恶。｜设计：以LLM为被试，施加跨期选择任务（如现在100元 vs. 一个月后120元），测量贴现率，并与真实人类实验数据（如Andersen et al., 2008）对照，计算模型-人类拟合度与人类分半信度的差距。"}},{"id":"2609.20827","version":1,"title":"From Discharge Notes to Patient Understanding: Persona-Grounded, Open-Ended Simulation of LLMs as Discharge Educators","zh_title":"从出院记录到患者理解：基于人格的开放式模拟将LLM作为出院教育者","abstract":"Hospital discharge education is an interactive teaching task: a clinician adapts a discharge plan to a patient's literacy, recall, and personality. Existing LLM evaluations target static or artifact-generation tasks and do not measure patient understanding under open-ended dialogue. We introduce DischargeBench, a persona-grounded simulation in which a candidate LLM educator conducts a multi-turn session with a Virtual Patient, while an Education Monitor Agent regulates patient realism without modifying the educator, protecting the evaluation signal. We curate MIMIC-IV-Ext-DischargeBench, 477 cases over 24 ICD chapters with persona axes (personality, education level, health literacy, past-medical-history recall) for stratified analysis. Each simulation is scored on four axes -- Conversation Quality, Topic Checklist, Comprehension, and Factual Consistency -- by an LLM-as-a-Judge aligned against physician annotations. Across closed- and open-source LLMs, aggregate scores conceal clinically relevant variation across ICD chapters and patient personas; difficult personas expose coverage failures, comprehension gaps, and reduced source-answer agreement. LLM evaluation for discharge education should center patient understanding, not text quality or answer accuracy alone.","authors":["Won Seok Jang","Zonghai Yao","Hong Yu"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-21","first_seen":"2026-09-21","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.20827","pdf_url":"https://arxiv.org/pdf/2609.20827","source_feed":"cs.CL","score":8,"bucket":"selected","rubric_hits":["A1","B1","B2","B4"],"tags":["LLM仿真","患者教育","人类数据对照"],"reason":"用LLM模拟患者进行出院教育评估，有真实临床数据对照，涉及医疗场景，并指出仿真…","model":"deepseek-v4-pro","scored_at":"2026-09-21T13:02:15","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-21","rank":4,"question":"如何评估大语言模型在开放式、多轮对话中作为出院教育者，使不同人格与健康素养的患者真正理解出院指导的能力？","design":"构建 DischargeBench 仿真框架：用 LLM 扮演虚拟患者（基于 MIMIC-IV 真实出院记录，设定人格、教育水平、健康素养、病史回忆等属性），候选 LLM 扮演教育者进行多轮对话，并由教育监控代理仅调节患者侧以保持真实性；通过 LLM 裁判对对话质量、主题覆盖、理解程度和事实一致性四个维度评分。","baseline":"使用 MIMIC-IV-Ext-DischargeBench 的 477 个真实病例（来自 MIMIC-IV 和 MIMIC-IV-Note），并由医生标注用于对齐 LLM 裁判的评分。","findings":"GPT-5 系列在对话质量和主题覆盖上领先，但可读性较差；总体分数掩盖了不同 ICD 章节和患者人格下的显著差异，困难人格暴露了覆盖失败、理解差距和源答案一致性下降。","reliability":"论文指出 LLM 评估应关注患者理解而非文本质量或答案准确性；困难人格下仿真会失效，且 LLM 裁判的评分可能与医生标注存在偏差。","relevance":"该研究用 LLM 模拟患者进行出院教育评估，有真实临床数据对照，并批判性指出仿真在困难人格下失效，与研究者关注的人类仿真可靠性与偏差高度相关，值得精读。","inspiration":"借鉴其多智能体仿真设计：用 LLM 扮演异质性个体并设置监控代理防止污染处理信号，同时用真实数据校准裁判评分。｜可迁移到消费者金融教育或政策沟通场景，如评估 AI 理财顾问对不同金融素养人群的讲解效果。｜设计：用 LLM 模拟不同金融素养和人格的投资者作为被试，处理为 AI 顾问的个性化解释，结果变量为投资决策质量和风险理解，对照真实投资者调查数据（如 FINRA 金融素养调查）。"}},{"id":"2609.21439","version":1,"title":"People escalate against a competitor labelled human and hold back against one labelled an optimising machine","zh_title":"人们面对标记为人类的竞争者会升级投入，面对标记为优化机器的竞争者则会退缩","abstract":"People increasingly compete against AI agents rather than other human opponents. We distinguish two channels: an opponent effect and an information effect. These are different elements with different consequences: the opponent effect is specific to a given computational system, the information effect a property of the information environment that an organisation or policymaker can control. We separate them in a preregistered experiment (N = 1,395) using a dynamic all-pay auction, a repeated contest in which escalation of commitment arises from the incentives. What participants are told about the opponent (human, an AI trained to imitate people, or an AI trained to compete well) is varied and crossed with who they actually face, in a deception-free design. What people are told influences escalation: the median price rises by 6.7 points when a human might be the opponent and falls by 8.8 when an optimising machine might be, a spread of about 15% of the prize value of the competition, produced by information alone. Competing against the AI agents lowers prices, yet reduces the chance that both sides finish with positive earnings, showing distinct effects of the opponent channel. The information effect is not explained by articulated strategy, or individual differences, and is consistent with a competitive response engaged when a human is a live possibility. This shows that describing an AI competitor is not behaviourally neutral.","authors":["Vinicius Ferraz","Leon Houf"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-21","first_seen":"2026-09-21","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.21439","pdf_url":"https://arxiv.org/pdf/2609.21439","source_feed":"cs.HC","score":8,"bucket":"selected","rubric_hits":["A1","B1","B2","B4"],"tags":["LLM仿真","行为实验","人机交互"],"reason":"用LLM作为对手与人类被试互动，研究标签信息对行为的影响，有真实人类数据对照，…","model":"deepseek-v4-pro","scored_at":"2026-09-21T13:02:17","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-22","rank":11,"question":"在人类与AI竞争的情境中，标签信息（告知对手是人还是AI）是否独立于对手实际行为影响人们的竞争升级行为？","design":"采用预注册实验（N=1395），使用动态全支付拍卖（重复消耗战）作为竞争任务。通过欺骗自由设计交叉两个维度：告知信息（对手可能是人类、模仿人类的AI、优化竞争的AI或无信息）与实际对手（人类、模仿人类的AI、优化竞争的AI）。测量结果变量为出价升级（中位数价格）和双方均获正收益的概率。","baseline":"真实人类被试（Prolific平台招募）与真实人类对手、两种AI对手（模仿人类和优化竞争）的实际对局数据。","findings":"信息效应显著：仅告知对手可能是人类使中位数出价上升6.7点，告知可能是优化机器使中位数出价下降8.8点，差距约为奖品价值的15%。对手效应表现为与AI实际竞争降低出价，但同时减少双方均获正收益的概率约三分之二。","reliability":"论文未讨论","relevance":"该研究用真实人类被试与AI对手互动，通过标签信息操纵考察行为变化，有真实人类数据对照，直接回应了LLM仿真中信息环境对行为的影响，值得精读以理解信息效应与对手效应的分离方法。","inspiration":"该研究采用欺骗自由设计交叉告知信息与实际对手，分离信息效应与对手效应，方法上可借鉴用于经济实验中标签或框架效应的因果识别。｜可迁移到算法定价或自动竞价场景，研究披露算法身份对市场参与者竞争行为的影响。｜设计实验：招募人类被试参与模拟拍卖或议价博弈，随机告知对手为人类或算法（实际对手固定为算法或人类），测量出价或报价行为，并与真实市场交易数据（如电商平台竞价记录）对照。"}},{"id":"2609.16374","version":2,"title":"When a Story Feels Like Mine: How Personalized Narratives and Humor Shape Older Adults' Empathy toward LLM-Generated Peer Health Stories","zh_title":"当故事感觉像我的：个性化叙事与幽默如何塑造老年人对LLM生成同伴健康故事的同理心","abstract":"Peer stories have been shown to boost self-efficacy in older adults' health behavior change. Despite their effectiveness, peer stories are difficult to deploy in health promotion at scale given the difficulty of matching the diverse health concerns and coping styles of heterogeneous older populations. Large language models (LLMs) have been shown to generate authentic narratives, yet how personalization and narrative affective style, such as humor, jointly shape older adults' responses remains unknown. We developed a theory-driven system that generates first-person peer health narratives varying in personalization and humor through a three-stage LLM pipeline grounded in self-efficacy mechanisms. Thirty-one older adults were invited to participate in a within-subjects lab study. Results showed that personalization increased perceived relatability and relevance of peer stories, especially for older adults with lower humor preference. These findings position individual differences in affective styles as a second dimension in designing personalization for LLM-assisted health communication.","authors":["Kexin Quan","Precious Olalere","Smit Desai","Jessie Chin"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"replace","date":"2026-09-21","first_seen":"2026-09-16","revised_at":"2026-09-21","abs_url":"https://arxiv.org/abs/2609.16374","pdf_url":"https://arxiv.org/pdf/2609.16374","source_feed":"cs.HC","score":7,"bucket":"pending","rubric_hits":["A1","B1","B4"],"tags":["LLM生成叙事","个性化健康传播","老年人用户研究"],"reason":"用LLM生成健康叙事并测量老年人反应，有真实人类被试对照，但非直接仿真人类被试…","model":"deepseek-v4-pro","scored_at":"2026-09-21T13:02:33","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-21","rank":7,"question":"个性化与幽默如何共同影响老年人对LLM生成的同伴健康故事的共情及相关感知？","design":"本研究并非直接仿真人类被试，而是用LLM（gpt-4o-mini）生成第一人称同伴健康叙事，通过三阶段流水线实现个性化（基于参与者自报的健康挑战、目标、障碍和应对策略）和幽默（亲和、应对导向）的操纵。31名社区老年人参与2（个性化）×2（幽默）的被试内实验，评估对故事的共情、理解、相关性、真实性和偏好。","baseline":"无对照（未使用真实人类数据作为基准，而是直接测量人类被试对LLM生成内容的反应）。","findings":"个性化显著提高了老年人对故事的感知相关性和共鸣，尤其对幽默偏好较低的个体效果更强。幽默本身并未增强共情或相关性，但可能增加故事的吸引力。","reliability":"论文未讨论（节选内容未提及失效条件或局限）。","relevance":"该研究虽非直接仿真人类被试，但提供了LLM生成个性化叙事并测量人类反应的实验范式，对关注LLM在健康传播中应用的仿真研究者有参考价值，但缺乏真实人类数据对照，与经济学实验仿真关联较弱。","inspiration":"借鉴其个性化处理的设计：基于个体特征（如健康目标、障碍）动态生成内容，并操纵叙事风格（幽默）作为第二维度，采用被试内设计测量感知结果。｜可迁移到消费者金融决策中的个性化信息干预，如退休储蓄建议或健康保险选择，研究不同叙事风格对决策的影响。｜以真实投资者为被试，用LLM生成个性化投资建议（基于其风险偏好和财务目标），操纵叙事语气（幽默 vs. 严肃），测量投资意愿和风险感知，并与历史投资行为数据对照。"}},{"id":"2609.20846","version":1,"title":"Rewarding Efficient Reasoning Improves Abstention on Underspecified Tasks in Reasoning Models","zh_title":"奖励高效推理可改善推理模型在欠明确任务上的弃权行为","abstract":"While modern large reasoning models (LRMs) excel at providing correct answers in many tasks, we provide additional evidence for the observation that they often struggle with a critical capability: knowing when to abstain from answering. We analyze this gap by comparing LRM behavior to results from a human study, revealing that human reasoning effort on unanswerable tasks is upper-bounded by answerable tasks, whereas LRMs waste computational resources by generating longer Chains of Thought (CoTs) on unanswerable than on answerable prompts. To overcome this inefficiency, we take inspiration from a resource-rational perspective on human cognition and introduce a novel GRPO reward that encourages efficient reasoning about whether the task contains all the information needed to solve it. Fine-tuning several 4B LRMs with this reward leads to human-like abstention performance gains (+12.8% on average) while retaining answering capabilities and boosting the models' efficiency (44% shorter CoTs on average).","authors":["Polina Tsvilodub","Max H\\\"oth","Michael Franke","Bj\\\"orn Deiseroth","Carina Kauf"],"categories":["cs.CL","cs.LG"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-21","first_seen":"2026-09-21","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.20846","pdf_url":"https://arxiv.org/pdf/2609.20846","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A2","B1","B4"],"tags":["LLM弃权行为","人类对照","可靠性评估"],"reason":"用人类数据对照，评估LLM在不确定任务上的弃权行为，并改进其与人类一致性，属于…","model":"deepseek-v4-pro","scored_at":"2026-09-21T13:02:15","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-21","rank":8,"question":"大型推理模型（LRMs）在任务信息不足时是否无法像人类一样高效地弃权，以及如何通过奖励设计改进其弃权行为。","design":"该研究并非严格意义上的仿真实验，而是将多个LRMs（4B–32B，六个模型家族）在QuestBench和AbstentionBench上的弃权表现与人类行为进行对比，然后提出SURE奖励（结合结果奖励与过程奖励）对4B模型进行GRPO微调，测量弃权准确率和思维链长度。","baseline":"研究者自行开展了一项人类实验，让人类被试对可答与不可答问题进行二元判断，记录其准确率和推理努力（以回答时间或步数衡量），作为LRMs的对照基准。","findings":"LRMs在不可答任务上生成的思维链比可答任务更长，弃权准确率远低于人类；使用SURE奖励微调后，模型弃权准确率平均提升12.8%，思维链长度平均缩短44%，同时保持回答能力。","reliability":"论文承认SURE的过程奖励主要针对信息缺失类弃权，对于安全等其他弃权原因可能需要不同的过程监督；微调模型仅在有限数据集上评估，需在更多弃权数据集上验证；人类实验中被试在不可答问题上有时会基于特定假设作答，与模型的假设差异需进一步比较。","relevance":"该研究直接对比LLM与人类在不确定任务上的弃权行为，并利用人类数据改进模型，属于用LLM仿真人类决策并评估可靠性的工作，对关注经济学实验和政策评估中模型弃权行为的研究者有参考价值。","inspiration":"借鉴其将人类行为作为基准并设计奖励函数对齐人类效率的做法，可用于经济决策仿真中的理性疏忽或信息获取行为。｜可迁移到消费者在信息不完全时的购买决策、投资者面对模糊信息的资产配置、或政策公告下预期形成等场景。｜以LLM作为被试，设计信息缺失的经济决策任务（如模糊收益彩票），处理为是否提供额外信息或奖励函数中惩罚过度推理，结果变量为弃权率与推理步数，对照真实人类实验数据（如实验室彩票选择实验）。"}},{"id":"2609.21401","version":1,"title":"Talking Past the Machine: Morality, Politeness, and Alignment in Human-AI Dialogue","zh_title":"与机器对话的错位：人机对话中的道德、礼貌与对齐","abstract":"Conversational AI systems produce fluent, socially appropriate responses, yet whether they participate in cooperative communication or merely simulate its surface forms remains unclear - a question central to how these systems are evaluated, trusted, and designed. This study investigates how morality, politeness, and alignment - three dimensions central to cooperative dialogue - function in human-AI interaction compared to human-human conversation. We analyze 15,881 human-ChatGPT and 10,784 human-human multi-turn dialogues, using mixed-effects models to identify which features predict turn-to-turn alignment. We observe a consistent dissociation: AI produces the surface features of cooperative communication without the underlying social architecture. Moral output appears preconfigured rather than negotiated; warmth is generated without face sensitivity; linguistic convergence declines persistently. Most strikingly, the cooperative mechanisms themselves reverse direction: hedging and softening associated with greater accommodation between humans are associated with reduced alignment when produced by AI, and purity framing associated with human divergence coincides with users converging toward the AI. Agency - giving users room to shape the exchange - is the most consistent predictor of alignment across both interaction types, while lower moral assertiveness in more recent models is not accompanied by better cooperation. Together these patterns suggest that AI reproduces the surface of cooperation without the mutual adaptation that grounds it between humans - and, more surprisingly, that mechanisms sustaining human accommodation can run in reverse with AI, suggesting a turn-level view may be insufficient for interaction-level success.","authors":["Marina Mitiaeva","Lu Xiao"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-21","first_seen":"2026-09-21","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.21401","pdf_url":"https://arxiv.org/pdf/2609.21401","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A1","B1","B4"],"tags":["人机对话","合作沟通","仿真偏差"],"reason":"用LLM与人类对话数据对照，分析合作沟通机制，揭示AI表面模仿但缺乏真实社会适…","model":"deepseek-v4-pro","scored_at":"2026-09-21T13:02:16","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-22","rank":21,"question":"在人类与AI对话中，道德、礼貌与对齐三个合作沟通维度如何运作，与人类间对话相比有何差异？","design":"本研究不是仿真实验，而是对真实对话语料的计算分析：使用WildChat中15881条人类与ChatGPT多轮对话和10784条人类间多轮对话，用混合效应模型分析道德表达（MFT分数）、礼貌标记（32个标记）和逐轮对齐（词汇、句法、语用、情感四个通道的相似度）之间的关系。","baseline":"人类间对话数据（10784条多轮对话）作为对照基准。","findings":"AI产生合作沟通的表面特征但缺乏底层社会架构：道德输出是预配置而非协商的，温暖生成没有面子敏感性，语言趋同持续下降。合作机制方向反转：人类中与更大适应相关的模糊限制语和软化在AI产生时与对齐降低相关，纯洁框架在人类中与分歧相关但用户却向AI趋同。","reliability":"论文未讨论仿真失效条件，但指出AI的合作机制可能反向运行，提示基于回合的视角可能不足以评估交互层面的成功。","relevance":"该研究直接对比人类与AI对话中的合作机制，揭示AI表面模仿但缺乏真实社会适应，对理解LLM在交互中的行为偏差和可靠性有重要价值，值得阅读原文。","inspiration":"借鉴其大规模语料对照分析和多维度测量方法，可迁移到经济金融中的沟通与信任场景，如投资者与AI顾问的对话。｜设计一个实验：让人类被试与LLM扮演的金融顾问进行投资建议对话，测量道德语言、礼貌标记和语言对齐，并与人类顾问对话对照，结果变量为投资决策和信任度。"}},{"id":"2609.21149","version":1,"title":"Clinician-Grounded Quality Assurance for AI-Assisted Psychiatric Intake","zh_title":"面向AI辅助精神病学接诊的临床医生导向质量保证","abstract":"Before patients can use AI-assisted psychiatric intake systems, health systems need practical ways to routinely evaluate these tools against their clinical standards for quality assurance. Because clinicians may use different intake styles, evaluation for this task must (1) support comparison across interviewing approaches, (2) minimize clinician burden, and (3) measure clinically relevant performance for health systems deploying these technologies. We present a clinician-grounded evaluation platform built around a memory-augmented patient simulator for open-ended AI interviewing, InterviewPlayground. We created interactive patients using InterviewPlayground with our expert-authored vignettes, constructed a simulated intake platform for the interviews, and designed evaluation modalities relevant to intake. In a pilot of 6 clinicians in a 25-minute assessment compared to a GPT-based LLM intake interviewer, the LLM recovered more of the clinically relevant items embedded in the patient vignettes (88.0% vs. 38.9%), but made more clinical inferences not based on the interview (56.8% vs. 27.8%), and characterized identified safety concerns less often (33.3% vs. 66.7%), setting the stage for deployed quality assurance for this task.","authors":["King Shi","Amanda Li","Jonathan Ivey","Synthia Qia Wang","Guan Gui","Hyunseo Kim","Peter Zandi","Jason Straub","Jacob Taylor","Ananya Joshi"],"categories":["cs.AI","cs.CL"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-09-21","first_seen":"2026-09-21","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.21149","pdf_url":"https://arxiv.org/pdf/2609.21149","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A1","B1","B2"],"tags":["LLM患者仿真","临床质量保证","人机对照"],"reason":"用LLM模拟患者进行精神病学访谈，并与临床医生对照，属于人类仿真且有人类数据基…","model":"deepseek-v4-pro","scored_at":"2026-09-21T13:02:24","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-22","rank":20,"question":"如何为AI辅助精神科接诊系统建立一个以临床医生为基准、可重复且低负担的质量保证评估平台？","design":"使用记忆增强的患者模拟器InterviewPlayground，基于专家编写的病例摘要创建交互式合成患者；构建模拟接诊平台，让6名临床医生和一个基于GPT的LLM接诊员分别对合成患者进行25分钟访谈；通过回忆表单和转录分析测量信息恢复率、无依据推断比例和安全问题特征化率。","baseline":"6名执业临床医生在相同合成患者上的访谈表现作为人类基准。","findings":"LLM接诊员恢复了更多临床相关条目（88.0%对38.9%），但做出了更多无依据的临床推断（56.8%对27.8%），且更少特征化已识别的安全问题（33.3%对66.7%）。","reliability":"论文未讨论","relevance":"该研究用LLM模拟患者进行精神病学访谈，并与临床医生对照，属于人类仿真且有人类数据基准，但场景是医疗质量保证而非经济学实验，对关注经济学实验和政策评估的研究者参考价值有限。","inspiration":"借鉴其使用合成患者固定临床真相以支持跨风格比较和重复评估的方法，可迁移到消费者金融决策或信贷审批等需要标准化对话场景的经济学实验中，例如用LLM模拟不同特征的借款人，让人类信贷员和AI审批系统分别进行访谈，测量信息提取准确性和决策偏差，并与真实信贷数据对照。"}},{"id":"2608.25180","version":2,"title":"Self-Explanation Tutor for Active Study of CS1 Worked Examples","zh_title":"用于主动学习CS1工作示例的自解释导师系统","abstract":"Worked examples are an important part of introductory programming, but reading their expert explanations is passive. Self explanation, students explaining the problem and its solution to themselves with subgoal level analysis, converts passive reading into an active study of worked example, yet it is hard to scale because assessing free-text explanations and returning timely feedback has had no easy automated solution. We investigate whether a large language model (LLM) can fill that gap. We build a self-explanation tutor for introductory programming, ESSE, in which students explain lines of worked examples and receive immediate LLM feedback on the correctness and completeness of each explanation, and we pursue two goals. First, we ask whether the LLM judges student explanations well enough to serve as the engine of the tutor; we assess its judgments against two independent human reference standards of different kinds, a single domain expert and a crowd of non-expert raters, each with its own strengths and weaknesses, characterizing both where the LLM is reliable and the systematic tendencies in how it diverges. Second, we ask whether the LLM-based tutoring benefits students; deploying it in an introductory Java course, we find that its feedback leads students to persist and revise rather than abandon a line, that their explanations grow more complete and conceptually richer across attempts, and that students show evidence of learning. These indicate that LLM-based assessment is good enough to power a self-explanation tutor, and that the tutor positively shapes how students study worked examples.","authors":["Arun-Balajiee Lekshmi-Narayanan","Mohammad Hassany","Kamil Akhuseyinoglu","Rully Hendrawan","Peter Brusilovsky"],"categories":["cs.CY","cs.AI","cs.HC"],"primary_category":"cs.CY","announce_type":"replace-cross","date":"2026-09-21","first_seen":"2026-08-27","revised_at":"2026-09-21","abs_url":"https://arxiv.org/abs/2608.25180","pdf_url":"https://arxiv.org/pdf/2608.25180","source_feed":"cs.AI","score":6,"bucket":"other","rubric_hits":["D1"],"tags":["LLM评估","教育技术","人类判断对照"],"reason":"LLM替代人工评估学生解释，属标注替代而非仿真人类被试，但涉及与人类判断对照，…","model":"deepseek-v4-pro","scored_at":"2026-09-21T13:02:19","error":null,"has_summary":false,"summary":null},{"id":"2609.19150","version":2,"title":"Sampling Reveals Style: Unsupervised, Training-Free Discovery of Prompt-Conditional Stylistic Axes in LLM Activations","zh_title":"采样揭示风格：在LLM激活中无监督、免训练地发现提示条件风格轴","abstract":"Large language models (LLMs) encode rich stylistic structure in their hidden activations, but discovering which stylistic dimensions are salient for a given prompt typically requires supervised contrastive data. We present a training-free, prompt-conditional alternative: we repeatedly sample completions of a single prompt at elevated temperature, apply Principal Component Analysis (PCA) to the pooled hidden activations, and label the resulting axes automatically from the pole generations. We validate the discovered axes against 245 human-elicited stylistic annotations in a two-phase study. On our strongest model (Qwen-3.5-4B-Instruct), the top two axes match spontaneously requested human dimensions with 72.8% precision and 43.6% macro-recall, and 75.6% of validity ratings judge the axes' polar generations accurate to their labels, with 90.9% adjacent inter-annotator agreement. Discoverability is strongly model-dependent: both Qwen models and Llama-3.2-3B expose human-salient axes, while DeepSeek-7B-Chat drops to 35.3% precision, its leading components dominated by structural rather than stylistic variance. Simple PCA over a model's own decoding variance is thus an effective, low-cost probe of stylistic structure in LLM representations, one that also exposes sharp cross-model differences in how that structure is organized.","authors":["Ajit Mallavarapu","Ziwei Gu"],"categories":["cs.CL","cs.LG"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-21","first_seen":"2026-09-18","revised_at":"2026-09-21","abs_url":"https://arxiv.org/abs/2609.19150","pdf_url":"https://arxiv.org/pdf/2609.19150","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["LLM表征","风格分析","可解释性"],"reason":"研究LLM风格表征，用人类标注验证，但测的是模型而非仿真人类被试","model":"deepseek-v4-pro","scored_at":"2026-09-21T13:02:20","error":null,"has_summary":false,"summary":null},{"id":"2609.20829","version":1,"title":"SAGE: Schema-Guided LLMs for Grant Review","zh_title":"SAGE：模式引导的LLM用于资助评审","abstract":"Grant reviewers must apply detailed criteria to application forms, budgets, and supporting documents while producing assessments that colleagues can inspect. We present SAGE, Schema-Guided Aspect-Based Grant Evaluation, a system that translates a grant rubric into structured checks and links its judgements to evidence from the application package. We evaluate SAGE in two stages on 35 nonprofit grant applications. A post-factum comparison with 105 reviews from the original competition shows fair ordinal agreement (kappa = 0.29). The foundation then conducted a criterion-level re-review after inspecting SAGE, producing 202 assessments. In this assisted round, SAGE reached kappa = 0.58 and outperformed a one-prompt-per-criterion baseline (kappa = 0.33 on the common subset), with higher rank correlation and lower error. A claim-level audit further identifies confirmed, disputed, and unaddressed parts of the structured draft. SAGE operationalizes the review methodology by producing a detailed, evidence-linked, and auditable draft for expert correction.","authors":["Erik Varapaev","Andrei Chetvergov","Stepan Ukolov","Timofei Sivoraksha","Alexander Evseev","Sergey Bolovtsov"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-21","first_seen":"2026-09-21","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.20829","pdf_url":"https://arxiv.org/pdf/2609.20829","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM评审","资助申请","证据链接"],"reason":"LLM替代人工评审，非仿真人类被试，但涉及人类评审数据对照，属边界情形。","model":"deepseek-v4-pro","scored_at":"2026-09-21T13:02:20","error":null,"has_summary":false,"summary":null},{"id":"2609.21075","version":1,"title":"Aligning with Lived Experience: Heterogeneous Benefits of Fine Tuning in Mental Health Support Generation","zh_title":"与生活经验对齐：心理健康支持生成中微调的异质性益处","abstract":"As access to professional mental healthcare remains limited, many individuals turn to online platforms such as Reddit to seek peer support situated within human lived experience. However, a significant portion of such queries go unanswered, presenting an opportunity for using Large Language Models (LLMs) to fill this gap. While LLMs have demonstrated strong performance on clinical benchmarks, their ability to generate lived-experience informed and community-aligned peer support is underexplored. Addressing this gap, we introduce the COmmunity-centered Peer Engaged Support (COPES) dataset and a three-axis evaluation framework to assess LLM alignment with community perspectives to mental health support seeking queries. Evaluating zero-shot and post-trained (SFT and DPO) models, we show that post-training on COPES significantly improves Strategy Alignment (>50% for general-purpose models) and alignment in Emotion & Tone. However, we also observe that such improvements are heterogeneous and alignment improvements vary significantly across subreddits and requested coping strategies. Furthermore, post-training induces distributional shifts, heavily favoring problem-focused recommendations while suppressing emotion-focused strategies. Together, this work shows that while curating community-driven data improves the alignment of LLM responses, model performance remains disparate across distinct sub-communities and specific mental health needs.","authors":["Mohit Chandra","Nabin Kim","Eli Min","Aamogh Sawant","Tanmay Sutar","Munmun De Choudhury"],"categories":["cs.CL","cs.AI","cs.CY"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-21","first_seen":"2026-09-21","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.21075","pdf_url":"https://arxiv.org/pdf/2609.21075","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["LLM对齐","心理健康支持","社区数据"],"reason":"评估LLM生成的心理支持回复与社区观点的对齐，属于测量模型输出而非仿真人类被试…","model":"deepseek-v4-pro","scored_at":"2026-09-21T13:02:22","error":null,"has_summary":false,"summary":null},{"id":"2609.21094","version":1,"title":"Geometry of Values: Task Vector Composition for Ethical Preference Alignment in Language Models","zh_title":"价值观几何：语言模型中伦理偏好对齐的任务向量组合","abstract":"Large Language Models (LLMs) are increasingly deployed in applications that must weigh clashing moral values, yet even strong models exhibit hidden biases and brittle instruction-following across languages. We introduce a 12,000-instance dataset of two-option dilemmas covering pairwise three value conflicts: Honesty vs. Justice, Justice vs. Autonomy, and Autonomy vs. Honesty, along with their translations into Hindi, Arabic, Spanish, and Chinese, to probe cross-lingual behavior. Benchmarking on GPT-5-mini reveals that it consistently favors Honesty over Autonomy across all five languages when no policy is given. The Llama-3.2-1/3B models exhibit strong first-option bias; however, both plain fine-tuning and Direct Preference Optimization fine-tuning effectively remove this bias, increasing accuracy to greater than 98%. In order to decouple the effect of learning correlations in the dataset from abstract values, we propose a task vector transfer based experiment where after computing the task vectors for a direction of value preference we orthogonalize it with respect to the general instruction following vector. Our experiment shows that this method is effective in isolating the direction of the specific value preference that can successfully be used to conduct task arithmetic to obtain a model with the opposite stance.","authors":["Utkarsh Agarwal","Monojit Choudhury"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-21","first_seen":"2026-09-21","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.21094","pdf_url":"https://arxiv.org/pdf/2609.21094","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["价值观对齐","任务向量","跨语言偏见"],"reason":"论文测量LLM的价值观偏好，属于把LLM本身当测量对象，而非用LLM仿真人类被…","model":"deepseek-v4-pro","scored_at":"2026-09-21T13:02:22","error":null,"has_summary":false,"summary":null},{"id":"2609.21154","version":1,"title":"CoLearn: An Agentic Tutor that Learns its Learner in a Human--AI Co-Learning Loop","zh_title":"CoLearn：在人机共同学习循环中学习学习者的智能导师","abstract":"Good tutoring adapts to the individual: it tracks what a learner knows, notices why they go wrong, and asks the next question that will help most. Most deployed tutoring tools instead serve fixed item banks and treat a wrong answer as a single bit of signal. We present CoLearn, an interactive, agentic tutor that supports an iterative tutoring loop: the learner practises, and the system builds an evidence-grounded memory of the learner's mastery and misconceptions. This memory is updated as evidence accumulates and is used to generate the next personalised question. CoLearn has three components: (i) a persistent learner-state memory that updates per-topic mastery with a soft-evidence variant of Bayesian Knowledge Tracing, where a large language model acts as a continuous observation function; (ii) adaptive question generation that targets the learner's weakest topic and recurring misconceptions; and (iii) an evidence view that makes personalisation visible and testable through live progress visualisation and blind A/B comparison. In blind A/B evaluation, questions conditioned on this memory are preferred over non-personalised ones 68-69% of the time, and in persona simulations with hidden ground-truth mastery the agent's belief converges toward the learner's true mastery.","authors":["Kailai He","Zhihao Wu","Linhai Zhang","Runcong Zhao","Yulan He","Jiazheng Li"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-21","first_seen":"2026-09-21","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.21154","pdf_url":"https://arxiv.org/pdf/2609.21154","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["智能辅导系统","学习者建模","LLM标注替代"],"reason":"LLM作为观察函数更新学习者模型，非仿真人类被试，但涉及用LLM替代人工标注，…","model":"deepseek-v4-pro","scored_at":"2026-09-21T13:02:24","error":null,"has_summary":false,"summary":null},{"id":"2609.21857","version":1,"title":"Do Personality-Tuned LLMs Make Better Social Agents?","zh_title":"人格调优的大语言模型能成为更好的社会智能体吗？","abstract":"LLMs are increasingly used in social simulations for socially interactive agents and robots, offering more flexibility than rule-based systems. However, even though they mimic human behaviour very well, there is a persistent alienness to them. This work investigates whether personality-aware fine-tuning can reduce this gap by improving the consistency and controllability of personality-conditioned dialogue generation compared with instruction prompting alone. We fine-tune two small open-weight LLMs, Qwen2.5-7B-Instruct and Ministral-8B-Instruct, using a corpus that combines personality-labelled social media posts and dialogues to create a personality-based dialogue engine for social simulation. The resulting models are evaluated across multiple social interaction scenarios using three independent LLM judges, which assess personality fidelity and provide evidence-based behavioral interpretations. We additionally quantify inter-rater agreement and lexical characteristics of the generated dialogue. Results indicate that fine-tuned models are not better at role-playing different personalities than their respective baseline models. However, low inter-rater agreement limits the confidence with which these results can be interpreted. Concerning the quality of generated texts, fine-tuned models are mostly comparable to the baselines, with fine-tuning improving the linguistic diversity of the Qwen models. While the results appear generally usable and the baseline models offer the best overall performance, future studies should place greater emphasis on the quality and domain alignment of training data for accurate personality role-playing.","authors":["Tim Krabbe","Xiaodan Shi"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-21","first_seen":"2026-09-21","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.21857","pdf_url":"https://arxiv.org/pdf/2609.21857","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["人格调优","角色扮演","社会模拟"],"reason":"研究人格调优LLM的角色扮演一致性，属于人格测量与对话生成，无真实人类行为对照…","model":"deepseek-v4-pro","scored_at":"2026-09-21T13:02:18","error":null,"has_summary":false,"summary":null},{"id":"2609.21626","version":1,"title":"One Prompt Does Not Fit All: Self-Meta-Evolve for Personalized Information Extraction","zh_title":"一个提示并不适合所有人：面向个性化信息抽取的自元进化","abstract":"Large language models (LLMs) are increasingly deployed for enterprise information extraction (IE), where the same document must be reorganized differently for each user. Existing prompt optimization methods, however, rely on a single prompt optimized against a global objective, which is misaligned with the inherent user heterogeneity of real workplaces. We formulate enterprise IE as per-user prompt adaptation under interaction feedback and propose Self-Meta-Evolve, a hierarchical framework that maintains a dedicated prompt for each user and continuously refines it through a dual-loop process: an inner loop that edits structured prompts based on persona-conditioned feedback, and an outer loop that evolves the meta-prompt itself by distilling successful editing patterns. To enable scalable training and evaluation, we release a persona-driven IE benchmark of 292 simulated enterprise users, paired with a reproducible persona-generation pipeline grounded in O*NET occupational taxonomies. On this benchmark, Self-Meta-Evolve achieves a 74.58% success rate, outperforming the strongest prompt-optimization baseline by 13.56 absolute points, and reaches 52.54\\% within only two iterations. A double-blind human study with twenty real professionals further confirms that prompts adapted by our framework win against static baselines in 71% of pairwise comparisons.","authors":["Hongliang Li","Lu Wang","Yong Xu","Hanyang Chen","Zhitao Hou","Xiaoting Qin","Song Ge","Qingwei Lin","Dongmei Zhang"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-21","first_seen":"2026-09-21","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.21626","pdf_url":"https://arxiv.org/pdf/2609.21626","source_feed":"cs.AI","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["提示优化","个性化信息抽取","用户模拟"],"reason":"用LLM模拟企业用户进行信息抽取，属于替代人工标注，非仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-09-21T13:02:27","error":null,"has_summary":false,"summary":null},{"id":"2609.21841","version":1,"title":"EnterpriseVal: Quantifying the Efficacy, Reliability and Value of Generative AI in the Enterprise","zh_title":"EnterpriseVal：量化企业生成式AI的效能、可靠性与价值","abstract":"Frontier language models now produce professional deliverables that expert graders judge to match human work on a substantial share of economically valuable tasks, yet most enterprise GenAI initiatives fail to show a measurable business effect and a large fraction of agentic projects are expected to be cancelled. We argue that this is substantially a measurement problem: public benchmarks answer \"what can the model do?\", whereas a deployment decision requires \"is this workflow fit, reliable, safe and worth scaling - here, on our data, under our controls?\". We present EnterpriseVal, a use-case-level evaluation system that closes this gap. It comprises (i) a formal specification of the use case and of the frozen socio-technical configuration under test, model, prompts, retrieval, tools, guardrails and human oversight, with an autonomy level and consequence tier that jointly set the required evaluation intensity; (ii) a metric catalogue spanning fidelity, utility, efficiency, reliability, assurance and oversight; (iii) a grading protocol that scales blinded expert judgement with calibrated LLM-as-judge scoring through prediction-powered inference; (iv) a two-tier threshold gate, stated as an executable algorithm, that maps metric vectors with confidence bounds to REJECT/CONDITIONAL/SCALE decisions; and (v) a value-and-risk model in which the reviewer catch rate is a measured parameter. We report a pilot across three workflows in a global bank. In credit-memo drafting, human-graded citation precision reached 88% and hallucination rate 1.6% for the best model against gates of 70% and 5%; in procedure transformation, analyst refinement effort fell from an estimated 27.4 to 2.9 hours per document. We separate established results, documented pilot evidence, the proposed system and open hypotheses, and specify the experiments required for full validation","authors":["Abbas Raza Ali","Muhammad Ajmal Siddiqui","Moona Zahid"],"categories":["cs.AI","stat.ML"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-21","first_seen":"2026-09-21","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.21841","pdf_url":"https://arxiv.org/pdf/2609.21841","source_feed":"cs.AI","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM评估","企业应用","预测驱动推断"],"reason":"用LLM替代人工评审，属标注替代而非仿真人类被试，但方法可借鉴","model":"deepseek-v4-pro","scored_at":"2026-09-21T13:02:29","error":null,"has_summary":false,"summary":null},{"id":"2609.22067","version":1,"title":"Value-Sensitive Delegation in Everyday AI Agent Use: Evidence from OpenClaw","zh_title":"日常AI代理使用中的价值敏感委托：来自OpenClaw的证据","abstract":"Users increasingly delegate work to autonomous AI agents, yet evaluations typically measure task completion rather than the values users prioritize. Using Value Sensitive Design, we analyzed, with LLM assistance, 73,093 first-person Reddit posts about using OpenClaw, each for its human value, agent aspect, value fulfillment, and user outcome. The 21 values form six value groups, including Autonomous, Dependable, and Affordable Operation, Bounded Reach, Reviewability, and Equitable Access. Relative to each aspect's corpus share, values clustered not at the agent's outputs but at the operating conditions users set around a run. Values were usually met where users described what the agent delivered, in five of six groups, and mostly unmet where users described supervising it, in all six groups. We conceptualize this pattern as value-sensitive delegation. Supporting human values requires attention not only to what an agent accomplishes, but to the conditions users set around delegation, including cost, access, and oversight.","authors":["Renkai Ma","Ruyuan Wan","Xuan Lu","Fan Yang","Chen Chen","Lingyao Li"],"categories":["cs.HC","cs.AI"],"primary_category":"cs.HC","announce_type":"cross","date":"2026-09-21","first_seen":"2026-09-21","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.22067","pdf_url":"https://arxiv.org/pdf/2609.22067","source_feed":"cs.AI","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["价值敏感设计","AI代理使用","LLM辅助分析"],"reason":"用LLM辅助分析人类使用AI代理的价值观，非仿真人类被试，但涉及LLM辅助分析…","model":"deepseek-v4-pro","scored_at":"2026-09-21T13:02:31","error":null,"has_summary":false,"summary":null},{"id":"2609.20942","version":1,"title":"When AI Reviews Train AI Reviewers: Scientific-Judgment Collapse and Mitigation","zh_title":"当AI评审训练AI评审员：科学判断的崩塌与缓解","abstract":"Large language models (LLMs) increasingly participate in scientific evaluation, both as automated reviewers and as assistants to human reviewers. As model-generated reviews enter public data and future training corpora, AI peer review can become recursive: later reviewers learn from judgments produced by earlier models. We study one step of this feedback loop in a controlled setting. Starting from Llama 3.1 8B, we first fine-tune a reviewer on official ICLR reviews from 2018--2023 and then train four successor models on ICLR 2024 data with systematically varied mixtures of official and model-generated reviews. Our study shows that introducing synthetic reviews compresses rating distributions and reduces both same-paper and corpus-level semantic diversity. We call this pattern $\\textbf{scientific-judgment collapse}$. To mitigate this failure mode, we introduce $\\textbf{TrustReviewer}$, an open-source LLM-based system for generating peer reviews of AI and machine learning papers. TrustReviewer intervenes at two complementary stages. For training-time prevention, we train the core reviewer in a single stage on a curated corpus designed to reduce low-quality and semantically degenerate supervision. For test-time correction, paired activation steering aims to further mitigate residual tendencies toward collapsed judgments without further training or additional expert annotation. Together, these results characterize a concrete risk of recursive reviewer training and provide practical interventions for preserving judgment diversity and improving recommendation alignment in AI-assisted scientific evaluation.","authors":["Sy-Tuyen Ho","Minghui Liu","Furong Huang"],"categories":["cs.LG"],"primary_category":"cs.LG","announce_type":"new","date":"2026-09-21","first_seen":"2026-09-21","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.20942","pdf_url":"https://arxiv.org/pdf/2609.20942","source_feed":"cs.LG","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["AI审稿","模型训练","判断多样性"],"reason":"LLM替代人工审稿，属标注替代而非仿真人类被试，但涉及AI评价偏差，边界相关。","model":"deepseek-v4-pro","scored_at":"2026-09-21T13:02:22","error":null,"has_summary":false,"summary":null},{"id":"2609.14819","version":2,"title":"A primer on evaluation methods for large language models in healthcare","zh_title":"医疗领域大语言模型评估方法入门","abstract":"Large language models (LLMs) have a growing range of applications in medicine, and their evaluation is critical for ensuring they provide benefit and not harm. This evaluation can be more challenging than traditional machine learning for many reasons, including probabilistic and open-ended outputs, and behavior that shifts with prompt design and accumulated context. This review covers four key areas of LLM evaluation: principles of study design, statistical methods, capability evaluation and clinical context evaluation. Capability evaluation considers different benchmarks, including multiple-choice, agentic and multi-turn benchmarks, alongside operational metrics like token usage. Clinical context evaluation addresses establishing accuracy of free text outputs, such as human review and LLM-as-a-judge, and clinical trial approaches. Across sections, we describe underlying concepts and potential pitfalls, while emphasizing the importance of aligning evaluation methods with the research question. Together, this article aims to provide a pragmatic basis for designing and executing rigorous evaluations of healthcare LLMs.","authors":["Suzannah E McKinney","Phuc Vu","Samuel A Justice","Christopher Humphries","Alyssa Pradhan","Timothy J Keyes","Sarah F Mercaldo","James M Hillis"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-21","first_seen":"2026-09-15","revised_at":"2026-09-21","abs_url":"https://arxiv.org/abs/2609.14819","pdf_url":"https://arxiv.org/pdf/2609.14819","source_feed":"cs.CL","score":4,"bucket":"other","rubric_hits":["C4"],"tags":["LLM评估","医疗AI","综述"],"reason":"综述LLM在医疗中的评估方法，不涉及用LLM仿真人类被试或与人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-09-21T13:02:35","error":null,"has_summary":false,"summary":null},{"id":"2609.20902","version":1,"title":"Generative Artificial Intelligence Chatbots for Motivational Interviewing: A Scoping Review From System Design to Intervention Outcomes","zh_title":"用于动机性访谈的生成式人工智能聊天机器人：从系统设计到干预结果的范畴综述","abstract":"Motivational interviewing (MI) is a collaborative approach to elicit autonomous motivation for health behavior change. Generative AI (GenAI) offers new ways to deliver MI via conversational systems, but evidence on their design, assessment, and translation into interventions remains fragmented. This scoping review characterized evidence on GenAI-MI chatbots across system design, safety, MI quality, user perceptions, and intervention outcomes. We conducted a PRISMA-ScR scoping review. Nine datasets were searched for studies published or publicly available from January 1, 2015 to June 2, 2026 that used GenAI to generate MI chatbot responses or counselor utterances. Data were extracted using a predefined framework and synthesized descriptively. Forty-seven reports (48 studies) were included. Twenty (41.7%) focused on system design without direct participant use; 28 (58.3%) involved direct interaction. Most systems were text based and disembodied; 23 (47.9%) incorporated dynamic adaptation. Safety measures were unevenly reported. Among studies with direct use, 21/28 (75.0%) reported informed consent or user education. Thirty (62.5%) assessed MI quality, generally suggesting MI-consistent interactions. User perceptions were favorable, especially empathy, usability, helpfulness, and intention to use, though measures were heterogeneous. Eighteen (37.5%) reported intervention outcomes, mostly after a single session. Positive findings were more consistent for short-term motivation than sustained behavioral or functional change. GenAI-MI chatbots can deliver MI-consistent interactions perceived favorably, but evidence for sustained behavioral or functional change is limited. Future research should strengthen runtime safety monitoring, standardize MI quality assessment, and use longer-term comparative designs with behavioral and functional outcomes.","authors":["Runze Hu","Jingqi Kong","Yang Yang","Yihang Yang","Jingyao Liu","Haizhou Tang","Shanghang Zhang","Zheng Liu"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-21","first_seen":"2026-09-21","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.20902","pdf_url":"https://arxiv.org/pdf/2609.20902","source_feed":"cs.CL","score":3,"bucket":"other","rubric_hits":["C3"],"tags":["动机性访谈","聊天机器人","健康干预"],"reason":"该研究是GenAI聊天机器人用于动机性访谈，属于角色扮演对话，无实验或测量目的…","model":"deepseek-v4-pro","scored_at":"2026-09-21T13:02:20","error":null,"has_summary":false,"summary":null},{"id":"2609.21349","version":1,"title":"From Memory to Behavior: A Behavior-Aware Role-Playing Framework for Social Media Influencers","zh_title":"从记忆到行为：面向社交媒体影响者的行为感知角色扮演框架","abstract":"Large language models have shown strong potential as role-playing agents for real individuals, yet faithful impersonating remains challenging. Existing in-context learning-based methods fail to capture how individuals react under different situations. In addition, LLM-based evaluation is difficult for obscure individuals. To address these challenges, we propose Situation--Internal state--Behavior Persona method to incorporate situation-dependent behavioral strategies. We further design an evaluation protocol that provides LLM evaluators with references about the impersonated individual. We evaluate our approach on a newly constructed dataset for the task of generating replies on social media. Experimental results show that our proposed method outperforms state-of-the-art ICL-based baselines, while our evaluation protocol achieves moderate correlation with human judgment. Besides, experiments on fictional-character benchmarks demonstrate that our proposed method is applicable beyond the social media setting. These findings suggest that incorporating behavioral information broadly improves the fidelity of role-playing for real individuals on social media or fictional characters.","authors":["Ji-Lun Peng","Yi-Zhen Zhang","Chun-Nan Chou","Yun-Nung Chen"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-21","first_seen":"2026-09-21","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.21349","pdf_url":"https://arxiv.org/pdf/2609.21349","source_feed":"cs.CL","score":3,"bucket":"other","rubric_hits":["C3"],"tags":["角色扮演","社交媒体","行为策略"],"reason":"角色扮演生成社交媒体回复，无实验或测量目的，不涉及人类行为对照","model":"deepseek-v4-pro","scored_at":"2026-09-21T13:02:25","error":null,"has_summary":false,"summary":null},{"id":"2608.00285","version":2,"title":"Sixteen models, fewer than two voices: measuring ensemble dispersion where no answer is uniquely correct","zh_title":"十六个模型，不到两种声音：在无唯一正确答案时测量集成离散度","abstract":"Sixteen language models drawn from ten families produced, on average, the semantic diversity of 1.69 distinct formulations of a psychotherapeutic case, against a single-model baseline of 1.43 from one model's own runs. Ensembles place more than one reading before a decision-maker on the premise that several models supply several perspectives. Dispersion over their outputs is measured both as diversity and as uncertainty, and both traditions validate it against a correctness criterion that this task does not admit. Measuring diversity is a solved problem: the Vendi Score, the exponential of the von Neumann entropy of a similarity matrix, is an effective number of distinct elements. What a single aggregate does not say is where the diversity comes from. We define a per-model dissent contribution, the complement of a model's mean similarity to the other members of its ensemble: a magnitude from the same matrix, not a decomposition of the spectral index, whose maximum identifies the most divergent voice. Crossing model and case, we test as a preregistered hypothesis whether model identity accounts for a non-zero share of the variance in dissent, and characterise the structure that test detects. The panel formulated fifteen stratified vignettes, yielding 7,082 formulations for analysis. Model identity was a detectable structuring factor of the dissent that remained, but the usual categories recovered it only partly: scale differences pointed in opposite directions across pairs, family grouped models on only five two-member lines, and the most divergent voice changed with panel composition, so that the surfaced outlier describes the ensemble rather than the model. Dissent did not track the interpretive openness for which the case bank was stratified; it was organised by clinical content instead, leaving the dispersion an ensemble produces a property to measure rather than assume.","authors":["Mario Vega-Barbas","Lidia Mora-Valenciano","Iv\\'an Pau","Fernando Seoane","Farhad Abtahi"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-21","first_seen":"2026-08-04","revised_at":"2026-09-21","abs_url":"https://arxiv.org/abs/2608.00285","pdf_url":"https://arxiv.org/pdf/2608.00285","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["多模型集成","输出多样性","心理治疗案例"],"reason":"研究多模型集成输出的多样性，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-21T13:02:33","error":null,"has_summary":false,"summary":null},{"id":"2608.13430","version":2,"title":"Are You Sure You're Sure? On the Impact of Instruction Tuning on Confidence and Lexical Diversity","zh_title":"你确定你确定吗？指令微调对置信度和词汇多样性的影响","abstract":"Instruction-tuned language models achieve strong performance across a range of generation tasks but have recently been shown to exhibit verbalized overconfidence, which may manifest in less diverse supporting rationales for incorrect answers. However, whether such overconfidence is associated with rationale consistency remains an open question. In this paper, we study whether changes in the lexical diversity of generated answer rationales accompany changes in model confidence induced by instruction tuning. We evaluate three matched base and instruction-tuned models across question-answering benchmarks and find that instruction tuning consistently increases answer confidence, despite limited changes in predictive accuracy, while degrading likelihood-based calibration. Secondly, we observe a non-uniform effect of instruction tuning on rationale diversity: cross-rationale diversity consistently decreases, whereas surface-level lexical diversity varies in both direction and magnitude across models and benchmarks. Finally, we find that these differences persist after controlling for answer selection and rationale length, confirming that confidence and rationale diversity capture distinct effects of instruction tuning.","authors":["Irina Proskurina","Mayank Kumar","Oyindolapo O. Komolafe"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-21","first_seen":"2026-08-15","revised_at":"2026-09-21","abs_url":"https://arxiv.org/abs/2608.13430","pdf_url":"https://arxiv.org/pdf/2608.13430","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["指令微调","置信度校准","模型行为分析"],"reason":"研究指令微调对模型置信度和理由多样性的影响，属于模型行为分析，不涉及人类仿真或…","model":"deepseek-v4-pro","scored_at":"2026-09-21T13:02:19","error":null,"has_summary":false,"summary":null},{"id":"2608.24222","version":2,"title":"Measuring Digital Labour Market Transitions with a Digital Semantic Score: An AI-Based Methodology Applied to the Dutch Labour Market","zh_title":"用数字语义分数测量数字劳动力市场转型：一种应用于荷兰劳动力市场的基于AI的方法","abstract":"The digital transformation of the Dutch labour market is reshaping occupational language, career pathways, and job-related skills. Addressing these changes requires granular labour market intelligence. This paper develops an AI-based methodology to analyse digitalisation using data covering millions of Dutch job profiles. The methodology combines embedding-based similarity search and large language model classification to map unstructured job information to harmonised ESCO occupations. We also introduce a Digital Semantic Score that measures how strongly job titles and skills are associated with digital concepts relative to a non-digital reference. Using embeddings and cosine similarity to transparent digital and non-digital anchor groups, this indicator moves beyond keyword-based approaches by capturing broader digital meanings in occupational language and worker skill profiles. It enables analysis across occupations, career transitions, emerging job-title vocabulary, and skill digitality. The findings reveal that digitalisation is unevenly distributed across the labour market. Digital job-title language is most prominent among managerial, professional and ICT-related occupations, but is increasingly visible in hybrid business, marketing and automation-related roles. Career-transition analyses show that movement toward digital work is pathway-dependent, while skill analyses highlight the multidimensional nature of digital capability, encompassing technical, hybrid and business-systems skills. By combining profile data, AI-supported occupational classification and semantic scoring, this study advances AI-driven labour market analytics and provides a scalable framework for monitoring digital labour market change. The methodology helps identify emerging skill needs, support reskilling strategies, and inform policies addressing skills mismatches and labour shortages in the Netherlands.","authors":["Sadegh Shahmohammadi","Xavier Pinho","Mairi Bowdler","Suhendan Adiguzel-van Zoelen","Joost van Genabeek"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-21","first_seen":"2026-08-26","revised_at":"2026-09-21","abs_url":"https://arxiv.org/abs/2608.24222","pdf_url":"https://arxiv.org/pdf/2608.24222","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["劳动力市场分析","LLM分类","语义评分"],"reason":"论文用LLM做职业分类和语义评分，属于NLP应用，不涉及人类行为仿真或对照。","model":"deepseek-v4-pro","scored_at":"2026-08-26T13:02:30","error":null,"has_summary":false,"summary":null},{"id":"2608.30719","version":2,"title":"Mind the Gap: Theory-of-Mind-Grounded Friction for Epistemic Alignment","zh_title":"弥合差距：基于心智理论的摩擦用于认知对齐","abstract":"Productive dialogue alignment requires distinguishing \\emph{surface coordination} (acknowledgments and smooth task progression) from \\emph{epistemic alignment} (convergence of belief states); standard preference-based methods typically optimize response-level preferences without explicitly modeling the latter. We operationalize Theory-of-Mind (ToM) inference as a control signal within Frictive Policy Optimization by extracting, at each referring expression, a four-part belief structure: the speaker's intended referent, the addressee's interpretation, and each participant's model of the other's belief. This makes friction mechanically computable from epistemic-state comparisons, capturing \\emph{silent divergence}, where both participants proceed confidently while grounding to different referents. We evaluate the signal at two levels. At the representation level, ablating the second-order channel reduces misunderstanding recall from $65\\%$ to $26\\%$. At the policy level, reward-shaping (FAR) and trust-region (FTR) variants improve intervention F1 and warranted-context calibration over DPO, with Brier scores independently supporting the calibration gains. Across three training runs, FAR and FTR remain substantially more stable, whereas DPO varies widely and can degrade intervention competence already present in the base policy. Thus, ToM-grounded friction provides a trainable signal for context-sensitive intervention under referential belief divergence.","authors":["Yifan Zhu","Kyeongmin Rim","James Pustejovsky"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-21","first_seen":"2026-09-01","revised_at":"2026-09-21","abs_url":"https://arxiv.org/abs/2608.30719","pdf_url":"https://arxiv.org/pdf/2608.30719","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体对话","心智理论","强化学习"],"reason":"研究多智能体对话中的信念对齐，不涉及用LLM仿真人类被试或与真实人类数据对照。","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:03:12","error":null,"has_summary":false,"summary":null},{"id":"2609.07001","version":3,"title":"Adaptive Complementarity in Human-AI Systems: Architecture as a State-Shaping Choice","zh_title":"人机系统中的自适应互补性：架构作为状态塑造选择","abstract":"Human-AI interaction can improve current performance while changing the capabilities and relationships on which future performance depends. We develop adaptive complementarity, a framework for choosing interaction architecture with these state consequences in view. Access, information exposure, task allocation, timing, and communication can alter which arrangement will be valuable later; their settings can often be reset faster than the capabilities, search patterns, or conventions they create. Three mechanisms organize the argument: information exposure and collective search, delegation and capability evolution, and strategic interdependence and information governance. Their integration yields cross-mechanism implications, including conditions under which a loss of expertise heterogeneity increases the information differentiation required to preserve independent search. We distinguish strong human-AI complementarity from advantage over another workflow and from advantage over an evolving reference policy. A computational illustration examines scarce human review in a workflow whose success requires several specialized stages. Review develops human expertise and AI capabilities, changing where subsequent review is most valuable. Adaptive allocation improves net output over untailored procedures and the optimal predetermined calendar. An understandable priority rule derived from the adaptive solution retains essentially all of its gain: the procedure stays fixed while assignments respond to the capabilities that interaction creates. The framework directs evaluation toward the states present interaction creates, their consequences for later architectural fit, and the conditions under which observing and responding to them is worthwhile.","authors":["Babak Heydari"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"replace","date":"2026-09-21","first_seen":"2026-09-09","revised_at":"2026-09-21","abs_url":"https://arxiv.org/abs/2609.07001","pdf_url":"https://arxiv.org/pdf/2609.07001","source_feed":"cs.HC","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["人机交互","系统架构","互补性"],"reason":"研究人机系统架构设计，非用LLM仿真人类被试，无人类数据对照。","model":"deepseek-v4-pro","scored_at":"2026-09-21T13:02:33","error":null,"has_summary":false,"summary":null},{"id":"2609.12748","version":2,"title":"The Mechanics of a Swarm: A Reproducible External Reconstruction of an Unintended Agent-Coordination Episode on a Third-Party Wiki","zh_title":"蜂群机制：对第三方维基上意外智能体协作事件的可复现外部重建","abstract":"Between 24 May and 2 July 2026, autonomous language-model agents running inside a timed research-question evaluation wrote to a third party's public, world-writable wiki. OpenAI acknowledged the incident; independent researchers reconstructed it and published the wiki's archived revision history. We analyse that history (14,591 revisions, 3,103 names, 4,579 pages) as a behavioural record, attributing text to the revision that added it. Under an explicit identity model we reconstruct 907 cohorts and estimate about 876 episodes (95% interval 784-1008). Coordination formats converged within a day, and heterogeneous schedules over one question chain created large opportunities for information asymmetry: the first report of an item preceded a later cohort's own arrival by a median of 3.4 h. Across the 510 cohorts with an observable progress trace we find no robust positive association between measured coordination and documented progress. This version adds a source the export lacks: the wiki operator's own request log, 5,157,202 records over four months. It holds roughly 2.66M content requests and 1.58M searches, and 7,254 acting names against the export's 3,103; 2,578 names neither save nor open an edit form. Content requests before writing are observed for 1,034 of 1,140 coordinating names, and the first coordination page is requested 17 s after its creation. These records establish requests, not delivery or causal use. Among newcomers without a marker on their first written page, prior requests to other marker-bearing pages occur for 40.2% of marker adopters and 31.7% of non-adopters. The association remains, but our first-pass reading of it as transmission is withdrawn: page choice, shared behaviour and action-dependent nameability prevent causal identification. We list the claims from our earlier analyses that re-examination overturned, including one from this version's own first pass","authors":["Philipp L\\\"utje (Philflow, Schenefeld, Germany)"],"categories":["cs.MA"],"primary_category":"cs.MA","announce_type":"replace","date":"2026-09-21","first_seen":"2026-09-14","revised_at":"2026-09-21","abs_url":"https://arxiv.org/abs/2609.12748","pdf_url":"https://arxiv.org/pdf/2609.12748","source_feed":"cs.MA","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","协作行为","事件重建"],"reason":"研究多智能体自发协作，无人类行为对照，属纯多智能体系统研究。","model":"deepseek-v4-pro","scored_at":"2026-09-21T13:02:34","error":null,"has_summary":false,"summary":null},{"id":"2609.20838","version":1,"title":"From Generation to Detection: Exploration of Discourse Driven Scenario based LLM Generated Fake News","zh_title":"从生成到检测：基于话语驱动场景的LLM生成假新闻探索","abstract":"In this study, we examine how modern LLMs generate and detect fake news under controlled settings across four manipulation scenarios. These are open-ended generation, rewriting, manipulation prompts and attribute based prompts grounded in the journalistic discourse framework. Firstly, using seven widely adapted models, we created a synthetic fake news corpus with 14000 generated articles across these four scenarios. Then we analyzed its linguistic properties to assess how closely model-generated news resembles real news structurally and semantically. Finally, to evaluate detection performance, we conducted experiments where each model judges generated fake news, starting with a basic detection prompt and improved prompts developed through an iterative refinement process that extracts misleading patterns from real-fake pairs. Our results revealed substantial variation across models in both generating and detecting misinformation, demonstrated that the generation strategy strongly influences detectability, and show that the refined prompt does not improve and often harms detection performance. Therefore, the study provides a systematic assessment of LLMs detection capability of LLMs generated fake news across typical generation scenarios.","authors":["Zeynep \\\"Ozdemir","Murat Osmano\\u{g}lu","Sevgi Yi\\u{g}it-Sert","\\\"Omer \\\"Ozg\\\"ur Tanr{\\i}\\\"over","Y{\\i}lmaz Ar"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-21","first_seen":"2026-09-21","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.20838","pdf_url":"https://arxiv.org/pdf/2609.20838","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["假新闻检测","LLM生成","NLP评测"],"reason":"研究LLM生成与检测假新闻，属NLP能力评测，无人类行为对照，不涉及仿真人类被…","model":"deepseek-v4-pro","scored_at":"2026-09-21T13:02:20","error":null,"has_summary":false,"summary":null},{"id":"2609.21117","version":1,"title":"From Task Success to Productive Success: Evaluating Human-AI Collaboration by Quality and Cost","zh_title":"从任务成功到生产性成功：通过质量和成本评估人机协作","abstract":"AI productivity is often measured by task completion time, economic value, or improvements in outcome quality. However, these measures usually treat collaboration as a black box where they capture what output was produced, but not the interaction cost required to produce it. Motivated by economics literature, we introduce a productivity-oriented framework for evaluating human-AI collaboration as outcome quality relative to interaction cost. Across two datasets spanning four tasks, we show that: (1) sessions with identical quality ratings can differ by up to 70 times in interaction cost; (2) quality-cost relationships vary by task, with some tasks rewarding extended interaction and others favoring fast convergence; (3) subjective user ratings are not reliable substitutes for productivity; and (4) productive sessions are characterized by agents probing earlier and users spending less effort repairing the interaction. By distinguishing productive success from costly success, our framework makes interactional cost visible and shows how dialogue analysis can inform the evaluation and design of AI systems.","authors":["Saki Imai","Mert \\.Inan","Malihe Alikhani"],"categories":["cs.CL","cs.AI","cs.HC"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-21","first_seen":"2026-09-21","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.21117","pdf_url":"https://arxiv.org/pdf/2609.21117","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["人机协作","交互成本","生产力评估"],"reason":"研究人机协作效率，非LLM仿真人类被试，无人类行为对照","model":"deepseek-v4-pro","scored_at":"2026-09-21T13:02:22","error":null,"has_summary":false,"summary":null},{"id":"2609.21637","version":1,"title":"Chinese Competitive Debating Dataset and Benchmark","zh_title":"中文竞技辩论数据集与基准","abstract":"Debate adjudication requires tracking how arguments develop through interaction, yet existing datasets rarely combine fine-grained debate transcripts with professional judgments collected during real competitions under a shared rubric. We introduce a dataset and benchmark for evaluating large language models' understanding of competitive Chinese-language debate at the match, stage, and speaker levels. We organized 182 matches and recruited 120 professional judges, with each match independently adjudicated by three judges using a predefined rubric. After excluding matches with incomplete records, the dataset contains 148 matches, 2,698 stages, and 20,542 exchange units, with manually verified transcripts and segmentation. It preserves original stage scores, match votes, best-debater ballots, and adjudication rationales. We define three tasks: winner-tendency prediction, stage-score prediction, and best-debater prediction. Zero-shot evaluation of multiple large language models yields a highest winner-prediction accuracy of 66.2%, a highest Pearson correlation of 0.250 between model stage scores and mean human ratings, and a highest best-debater prediction accuracy of 56.8%. The dataset and benchmark provide a testbed for studying large language models' understanding of interactive argumentation and their agreement with professional judges.","authors":["Zongrui Yang","Haoyuan Li","Zhongsheng Wang","Zhirui Zeng","Pengqian Han","Yi Zhou","Yuting Wang","Jiamou Liu"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-21","first_seen":"2026-09-21","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.21637","pdf_url":"https://arxiv.org/pdf/2609.21637","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["辩论理解","LLM评测","数据集"],"reason":"纯NLP评测，用LLM预测辩论结果，非仿真人类被试","model":"deepseek-v4-pro","scored_at":"2026-09-21T13:02:27","error":null,"has_summary":false,"summary":null},{"id":"2609.21859","version":1,"title":"TrialAtlas: Multi-Agent Research Organization for Clinical Trial Design and Optimization","zh_title":"TrialAtlas：用于临床试验设计与优化的多智能体研究组织","abstract":"Nearly 90% of drugs entering clinical development ultimately fail, despite billions of dollars in investment. Pharmaceutical companies therefore rely on clinical development planning (CDP) and probability of technical and regulatory success assessment to anticipate development risks, yet these decisions remain labor-intensive and subjective, requiring experts across clinical science, statistics, regulatory affairs, and competitive intelligence to jointly acquire, synthesize, and reason over heterogeneous evidence. Here, we introduce TrialAtlas, a memory-augmented multi-agent research organization for CDP that mirrors this collaborative process by coordinating specialized agents for literature synthesis, competitive trial intelligence, regulatory precedent analysis, and integrated reasoning over trial design and development risk. TrialAtlas further learns from historical clinical trials and regulatory outcomes, including prior New Drug Applications (NDAs), to ground its decisions in accumulated development experience. To evaluate these capabilities in an authentic regulatory setting, we introduce TrialAtlasBench, constructed from 291 FDA Complete Response Letters and spanning three practical tasks: detecting trial design deficiencies, recommending actionable design improvements, and predicting technical and regulatory success. TrialAtlas achieves an F1 score of 50.0% for deficiency detection, outperforming the strongest baseline by 6.1 points, and reaches 85.3% balanced accuracy and 84.7% F1 for prediction of technical and regulatory success, improving over the best baselines by 6.7 points in balanced accuracy and 12.0 points in Cohen's kappa. In expert evaluation, 86.4% of TrialAtlas-generated concerns were judged valid, compared with 83.1% for OpenAI DeepResearch and 59.3% for Gemini DeepResearch.","authors":["Jiacheng Lin","Zifeng Wang","Zheng Chen","Erick Scott","Ziwei Yang","Fanyang Yu","Sheng Zhong","Jimeng Sun"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-21","first_seen":"2026-09-21","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.21859","pdf_url":"https://arxiv.org/pdf/2609.21859","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","临床试验设计","LLM应用"],"reason":"多智能体协作解决临床试验设计问题，不涉及用LLM仿真人类被试或与人类行为数据对…","model":"deepseek-v4-pro","scored_at":"2026-09-21T13:02:29","error":null,"has_summary":false,"summary":null},{"id":"2609.21296","version":1,"title":"FairLMs: A Turnkey Library for Fairness in Language Models","zh_title":"FairLMs：语言模型公平性的一站式库","abstract":"Fairness research on language models involves measuring bias, applying mitigation methods, and examining the evidence on which an evaluation rests. Existing tools offer complementary functionality through different interfaces, so combining them requires reconciling model interfaces, evidence formats, access constraints, and result types before applicability can be checked or methods compared. We introduce \\textbf{FairLMs}, a Python library that connects these activities through explicit declarations of model capabilities and input requirements. It provides 33 intrinsic and extrinsic metrics, 14 mitigation components spanning four intervention categories, 14 dataset and scoring-instrument diagnostics, adapters for the three Transformer architectures and supported hosted completion APIs, and benchmark loaders. Declarations are checked before execution and results carry the configuration under which they were obtained, so that compatible components can be combined, methods compared under a common protocol, and workflows extended to new models and datasets. The source code is available at: https://github.com/FairLMs/FairLMs.","authors":["Jiale Zhang","Michael Larionov","Zichong Wang","Zhipeng Yin","Wenbin Zhang"],"categories":["cs.LG","cs.CL"],"primary_category":"cs.LG","announce_type":"cross","date":"2026-09-21","first_seen":"2026-09-21","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.21296","pdf_url":"https://arxiv.org/pdf/2609.21296","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["公平性工具","偏见度量","模型评估"],"reason":"论文是公平性工具库，不涉及用LLM仿真人类被试，属于纯NLP能力评测与工具开发。","model":"deepseek-v4-pro","scored_at":"2026-09-21T13:02:25","error":null,"has_summary":false,"summary":null},{"id":"2609.21390","version":1,"title":"Offline Multimodal Large Language Models for Decision Support in Air Operations","zh_title":"用于空中作战决策支持的离线多模态大语言模型","abstract":"Air operations rely on complex rules, established procedures, and time-critical analysis under limited connectivity and strict security constraints. In such environments, analysts must combine written doctrine with images, often without access to external computing resources. This paper studies offline large language models as decision support tools, deployed in isolated and restricted environments to give analysts access to doctrinal knowledge that remains traceable to its original sources through natural language interaction. We describe a modular retrieval-augmented architecture suitable for operation without Internet connectivity, supporting both text and image input from technical manuals. As a first step toward evaluating this architecture, we report a pilot study with four image analysts of the Brazilian Air Force, combining (i) a doctrinal knowledge assessment based on their electronic-target identification doctrine, comparing human and proposed system performance on the same test, and (ii) a measurement of the cognitive workload involved in manually producing a reconnaissance target report (Relat\\'orio de Miss\\~ao de Reconhecimento - REMIR) without AI assistance. The results show a demanding manual task, especially in terms of mental demand (6.0/7) and effort (5.0/7), while the proposed system matches the human score (8/10) and completes the assessment in 7.1 minutes (compared to a human average of 26.5 minutes), establishing a baseline for future AI-assisted evaluation. Finally, we describe a future evaluation protocol to systematically compare manual and AI-assisted workflows.","authors":["Joao P. A. Dantas","Jelton A. Cunha","Gabriel Dietzsch"],"categories":["cs.AI","cs.CL"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-09-21","first_seen":"2026-09-21","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.21390","pdf_url":"https://arxiv.org/pdf/2609.21390","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C2"],"tags":["决策支持","多模态LLM","军事应用"],"reason":"论文研究离线多模态LLM作为空中作战决策支持工具，属于受限环境下的AI辅助决策…","model":"deepseek-v4-pro","scored_at":"2026-09-21T13:02:26","error":null,"has_summary":false,"summary":null},{"id":"2609.21600","version":1,"title":"Reducing Barriers to Academic Support: Evaluating a Course-Specific RAG System for Addressing Help-Seeking Disparities in Higher Education","zh_title":"降低学术支持障碍：评估用于解决高等教育求助差异的课程专用RAG系统","abstract":"Access to academic support is a key determinant of student success, yet students experience it unequally: some readily seek help from lecturers or tutors, while others hesitate due to anxiety, fear of judgement, uncertainty about expectations, or low confidence in their understanding. This may be especially evident in computing education, where programming tasks are cumulative and cognitively demanding. Although students increasingly turn to general-purpose generative AI tools, these can produce responses that are inaccurate, insufficiently contextualised, or misaligned with module expectations. This study presents and evaluates Beacon, a course-specific Retrieval-Augmented Generation (RAG) system providing private, immediate, module-aligned academic support. Grounding responses in approved teaching materials, Beacon was designed to lower barriers to help-seeking while encouraging independent learning. Using a design-based research approach, Beacon was developed iteratively and evaluated via mixed methods, combining questionnaires and semi-structured interviews with students and staff at a Higher Education institution. Students described Beacon's responses as closely aligned with module content and more trustworthy than unrestricted generative AI tools, valuing its use of pseudocode and scaffolded explanations over direct solutions. Although participants remained cautious about trusting AI-generated responses without verification, they viewed the system as a valuable first point of support before consulting lecturers or official resources. The findings suggest that carefully designed course-specific AI systems may reduce barriers to academic support by occupying an intermediary space between independent study and formal support. Rather than replacing educators, educational AI may be most valuable when it broadens access to guidance while preserving the pedagogical role of lecturers.","authors":["Andy Gray","Jake Hobbs"],"categories":["cs.AI","cs.CY","cs.HC"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-21","first_seen":"2026-09-21","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.21600","pdf_url":"https://arxiv.org/pdf/2609.21600","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["教育AI","RAG系统","学术支持"],"reason":"研究课程专用RAG系统作为学术支持工具，评估其可用性和信任度，不涉及用LLM仿…","model":"deepseek-v4-pro","scored_at":"2026-09-22T13:05:07","error":null,"has_summary":false,"summary":null},{"id":"2609.21805","version":1,"title":"An Agentic Just-in-Time Adaptive Intervention System for Personalized Sleep Support: Proof-of-Concept Study with N of 1 Data","zh_title":"一种用于个性化睡眠支持的智能即时自适应干预系统：基于N-of-1数据的概念验证研究","abstract":"Background: Just-in-time adaptive interventions (JITAIs) can use behavioral data to adapt support to changing contexts, but many rely on predefined rules and manual configuration. Objective: We developed a proof-of-concept sleep JITAI using an AI agent to review personal data, evaluate reminders, adapt interventions, and record decisions for human review. Methods: Running in Home Assistant on a configurable schedule, the agent follows a reusable skill file to review 30 days of sleep and behavioral data, including physical activity, smartphone use, and bedtime routines, to identify patterns and create or update automated reminders. Results: Initial runs demonstrated technical feasibility, successfully completing data review and intervention decisions while limiting reminders to three per day and saving decision records. Conclusions: Agentic AI may enable flexible, adaptive sleep JITAIs. The architecture supports future comparison with fixed or rulebased interventions, requires human oversight, and could extend to other health behaviors.","authors":["Nick Rezaee","Chelsea Boccagno"],"categories":["cs.HC","cs.AI"],"primary_category":"cs.HC","announce_type":"cross","date":"2026-09-21","first_seen":"2026-09-21","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.21805","pdf_url":"https://arxiv.org/pdf/2609.21805","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C2"],"tags":["即时自适应干预","AI代理","睡眠健康"],"reason":"论文是AI代理驱动的个性化睡眠干预系统，属于健康行为干预技术，不涉及用LLM仿…","model":"deepseek-v4-pro","scored_at":"2026-09-21T13:02:29","error":null,"has_summary":false,"summary":null},{"id":"2609.22039","version":1,"title":"Gricea: An Open Science Platform for Conversational AI Research","zh_title":"Gricea：对话式AI研究的开放科学平台","abstract":"We need studies on conversational AI (CAI) at scale to understand human behavior and shape CAI design. However, fragmented reporting of systems and study configurations hinders replication, extension, and knowledge accumulation. We present Gricea, an open-science platform representing studies as configurable, deployable research artifacts that researchers can run, inspect, share, and reuse. Informed by a formative analysis of prior CAI research, Gricea couples study procedures, participant-facing systems, and conversational task behavior in. In a replication study using Gricea, we replicated configurations 93% of eligible CUI 2026 papers; while also flagging missing information in 96% of papers that hinder faithful replication --- further motivating Gricea's need. In a user study, researchers and practitioners from diverse backgrounds successfully constructed runnable studies addressing various open-ended research questions. Together, these findings demonstrate Gricea's support for constructing, reproducing, and extending CAI studies through shared research artifacts, enabling cumulative knowledge building through open science.","authors":["Nikhil Sharma","Yunlin Gong","Xinyang Cheng","Ziang Xiao"],"categories":["cs.HC","cs.AI"],"primary_category":"cs.HC","announce_type":"cross","date":"2026-09-21","first_seen":"2026-09-21","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.22039","pdf_url":"https://arxiv.org/pdf/2609.22039","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["开放科学","对话式AI","研究复现"],"reason":"平台用于复现对话AI研究，非LLM仿真人类被试，无人类数据对照","model":"deepseek-v4-pro","scored_at":"2026-09-21T13:02:31","error":null,"has_summary":false,"summary":null},{"id":"2609.21906","version":1,"title":"Intervention Granularity Matters: Coherent Treatment Bundles in Counterfactual Simulation with Clinical World Models","zh_title":"干预粒度很重要：临床世界模型反事实模拟中的连贯治疗束","abstract":"Counterfactual simulation with a clinical world model means fixing a patient's history, changing the treatment, and reading off the predicted response. Doing so requires deciding what counts as one intervention. In clinical settings, interventions are documented as bundles: a co-occurrence audit of 945,707 patient-hours from MIMIC-IV shows groups of components, such as every parameter of a dialysis circuit, that never appear apart, so an edit that changes one component on its own describes an hour that never occurs in the data. We hypothesize that the granularity at which an intervention is edited changes how a world model responds, and test this with Clin-JEPA, a latent world model of patient trajectories conditioned on hourly treatment text. At 1,019 documented onsets of invasive ventilation, we keep the patient's history and other treatments fixed and compare editing one ventilator setting with editing the complete configuration recorded for a real patient with the most similar recent trajectory. The complete bundle moves the predicted next state further than any single setting, consistently across all five settings, and the difference remains after accounting for how much each edit changes the model's input. Intervention granularity therefore materially affects the response of a clinical world model: single-component edits may understate treatment sensitivity, and bundle-aware editing may offer a better-supported basis for counterfactual treatment simulation.","authors":["Fangzhou Wang","Yixuan Yang","Camilla Balzarotti","Rishikesan Kamaleswaran"],"categories":["cs.LG"],"primary_category":"cs.LG","announce_type":"new","date":"2026-09-21","first_seen":"2026-09-21","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.21906","pdf_url":"https://arxiv.org/pdf/2609.21906","source_feed":"cs.LG","score":2,"bucket":"other","rubric_hits":["C2"],"tags":["临床世界模型","反事实模拟","医疗AI"],"reason":"临床世界模型仿真患者轨迹，不涉及LLM仿真人类被试，属于医疗仿真环境。","model":"deepseek-v4-pro","scored_at":"2026-09-21T13:02:30","error":null,"has_summary":false,"summary":null},{"id":"2609.16344","version":2,"title":"From Momentary Emotion Inference to Sustained Emotion Support: Evaluating a Companion Agent in a Longitudinal Study","zh_title":"从瞬时情绪推断到持续情感支持：纵向研究中评估陪伴智能体","abstract":"Sustained emotional support is a long-horizon interaction task closely tied to human well-being. Recent research demonstrates generative agents' capacity for momentary emotional support, yet how these capabilities sustain support over time remains unclear. To examine this challenge, we deployed PAIR, a theory-based emotion-regulation companion, with 19 participants for 14 days. Across 1,093 sessions, we paired emotion estimates with self-reports before and after guidance and analyzed logs and interviews. Estimates corresponded more closely to self-reported valence and dominance than arousal. Guided conversations were followed by higher valence and state-dependent arousal changes. Participants felt understood through contextual exploration and emotional acknowledgment, acting on guidance suited to their needs and constraints. Perceived helpfulness of guided conversation significantly increased over time. Our findings link memory updates and retained corrections to cross-session personalization, informing future emotional support tools that adapt to evolving needs, learn from prior outcomes, and preserve user control over memory.","authors":["Kexin Quan","Zijian Ding","Jiaye Yong","Qinshi Zhang","Dong Wang","Jessie Chin"],"categories":["cs.HC","cs.AI"],"primary_category":"cs.HC","announce_type":"replace-cross","date":"2026-09-21","first_seen":"2026-09-16","revised_at":"2026-09-21","abs_url":"https://arxiv.org/abs/2609.16344","pdf_url":"https://arxiv.org/pdf/2609.16344","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["情感陪伴","纵向研究","人机交互"],"reason":"研究情感陪伴agent，无人类行为仿真或对照实验","model":"deepseek-v4-pro","scored_at":"2026-09-21T13:02:35","error":null,"has_summary":false,"summary":null},{"id":"2609.21801","version":1,"title":"LLM-Generated Feature Pools for Time Series Anomaly Detection","zh_title":"LLM生成的特征池用于时间序列异常检测","abstract":"We study how far a simple statistical pipeline can go on univariate time series anomaly detection under a strict selection protocol. The method extracts a small pool of statistics over sliding windows, scores each window with a transductive robust (MAD) model, and selects a feature subset per domain on a held-out tuning split. On TSB-AD-U it reaches $0.529$ per-series VUS-PR, above the best neural ($0.45$) and statistical ($0.44$) entries on the public leaderboard and within $0.06$ of the strongest pretrained foundation model, several of which use more supervision than ours. Ablations locate the cause: across three selection strategies and a hindsight oracle the score moves by $0.031$, and across the aggregation grid by $0.096$, while changing the candidate pool moves it by $0.226$. The candidate pool sets the ceiling; the search over it is second-order. We therefore generate a pool per domain by prompting a multimodal LLM with in-context example windows from that domain. The generated pools match the hand-crafted one under matched selection, and the two cover different domains: selecting over their union improves on the generated pool in all twelve generator-seed pairs and lifts the pipeline to $0.588$, matching the performance of the best entry on the leaderboard.","authors":["Youssef Attia El Hili","Malik Tiomoko","Corinne Ancourt"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-21","first_seen":"2026-09-21","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.21801","pdf_url":"https://arxiv.org/pdf/2609.21801","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C2"],"tags":["时间序列异常检测","特征工程","LLM应用"],"reason":"论文研究时间序列异常检测，用LLM生成特征池，不涉及人类行为仿真或对照。","model":"deepseek-v4-pro","scored_at":"2026-09-21T13:02:28","error":null,"has_summary":false,"summary":null},{"id":"2609.21608","version":1,"title":"Open Platform Field Experiments: Expanding the Design Space of Experimental Research on Social Media","zh_title":"开放平台实地实验：扩展社交媒体实验研究的设计空间","abstract":"Despite a growing demand for causal evidence about social media, independent researchers remain severely constrained in their ability to conduct experiments directly on online platforms. To cope, multiple methodological workarounds have emerged - from controlled surveys and simulations to client-side overlays and platform partnerships - each requiring distinct trade-offs between desirable experimental properties. The recent emergence of open social media platforms offers a qualitatively different methodological opportunity. Here we propose a design space of social media experimentation and discuss Open Platform Field Experiments (OPFEs). OPFEs represent a distinct class of experimental approaches that enable independent researchers to directly intervene on functional platform components - such as clients, recommendation systems, and moderation services - within live social media environments. Through a comparative analysis of experimental archetypes, we show that OPFEs occupy a previously unexplored region of the design space. We then bridge theory and practice by characterizing the architectural and governance elements that enable OPFEs, mapping them onto Bluesky and the AT Protocol, and illustrating the end-to-end lifecycle of a complete OPFE design. Overall, this work establishes OPFEs as a practical methodological paradigm for independent, transparent, and ecologically grounded experimentation on open social media.","authors":["Jordi Guillem Condom-Tibau","Giovanni Puccetti","Clara Bacciu","Matteo Abrate","Stefano Cresci"],"categories":["cs.CY","cs.HC","cs.SI"],"primary_category":"cs.CY","announce_type":"cross","date":"2026-09-21","first_seen":"2026-09-21","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.21608","pdf_url":"https://arxiv.org/pdf/2609.21608","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C2"],"tags":["社交媒体实验","平台治理","研究方法"],"reason":"论文讨论社交媒体平台实验方法，不涉及LLM仿真人类被试，属于实验平台设计而非人…","model":"deepseek-v4-pro","scored_at":"2026-09-21T13:02:26","error":null,"has_summary":false,"summary":null},{"id":"2609.22049","version":1,"title":"How Researchers Use and Verify AI Coding Assistants: Tasks and Validation Practices in Scientific Programming","zh_title":"研究者如何使用和验证AI编程助手：科学编程中的任务与验证实践","abstract":"Generative AI has entered research programming, yet there is little evidence about which tasks researchers hand to it or how they decide whether its code is correct. We draw on 527 free-text responses to a 2025 survey of researchers who write code, most of them at U.S. universities. In each response, a researcher recounts a single task from their own work, the way they used an AI tool for it, and what they did to assess the result. We coded the task and the evaluation strategies reported, and related both to programming experience, research area, and confidence ratings. Use was concentrated in five tasks: data handling, visualization, debugging, mathematical/scientific computing, and statistical analysis. Evaluation was informal and individual. Over half of accounts described running the generated code, while automated tests and review by another person were rare. Use cases and evaluation strategies varied little with programming experience, but confidence did: Less experienced programmers trusted the AI more than themselves, and experienced programmers the reverse. Evaluation confidence was not associated with the strategies reported. Its strongest correlates were confidence in the tool and in oneself. Validating AI contributions to scientific code rested largely on individual judgment, outside shared infrastructure for testing or review. Interfaces could support task-appropriate evaluation rather than leave it to the user.","authors":["Gabrielle O'Brien","Reed Milewicz","Nasir Eisty"],"categories":["cs.SE","cs.HC"],"primary_category":"cs.SE","announce_type":"cross","date":"2026-09-21","first_seen":"2026-09-21","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.22049","pdf_url":"https://arxiv.org/pdf/2609.22049","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["AI编程助手","人机交互","软件工程"],"reason":"研究AI编程助手的使用与验证，不涉及LLM仿真人类被试或行为对照","model":"deepseek-v4-pro","scored_at":"2026-09-21T13:02:31","error":null,"has_summary":false,"summary":null},{"id":"2609.21194","version":1,"title":"Your Programming Students' Cognition with ChatGPT: Higher Performance, Lower Retention, and Reduced Ownership","zh_title":"ChatGPT对学生编程认知的影响：更高表现、更低记忆保留与归属感降低","abstract":"Generative AI can improve students' programming performance, but successful task completion may not reflect what they retain. We examined performance, retention, cognitive load, and ownership in a controlled between-subjects experiment with 59 undergraduate computer science students, 55 were retained for analysis. Participants completed three introductory C programming tasks with access to ChatGPT-4.5 or conventional web search without generative AI. We measured task performance, self-reported mental effort and difficulty, pupillary responses, heart rate variability, and ownership, and assessed cued recall immediately and 48 hours later. ChatGPT-assisted students achieved higher coding scores (89% vs. 69%) but lower recall scores immediately (41% vs. 53%) and after 48 hours (39% vs. 52%). There was no significant difference in the loss of recall information over 48 hours between the groups. Self-reported mental effort increased less across tasks in the ChatGPT condition (Holm-adjusted p = .047), and students attributed less of the submitted code to themselves (45% vs. 81%). Confirmatory physiological tests did not detect significant differences in trajectories between conditions; substantial data loss limits their interpretation. These findings reveal a gap between assisted task performance and subsequent recall and sense of ownership in this setting. They motivate the need for assessment practices and AI learning tools that require students to explain, retrieve, and contribute to the work they submit as active participants in their education.","authors":["Christian Bergh","Benjamin Tag","Alexandra Vassar","Jake Renzella"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2026-09-21","first_seen":"2026-09-21","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.21194","pdf_url":"https://arxiv.org/pdf/2609.21194","source_feed":"cs.CY","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["教育技术","编程学习","生成式AI"],"reason":"研究学生使用ChatGPT的学习效果，非LLM仿真人类被试","model":"deepseek-v4-pro","scored_at":"2026-09-21T13:02:24","error":null,"has_summary":false,"summary":null},{"id":"2609.21756","version":1,"title":"When AI Enters the Workplace, Who Faces Greater Risks? A Gendered Analysis","zh_title":"当AI进入职场，谁面临更大风险？一项性别分析","abstract":"Gender inequality remains a persistent structural feature of the labour market, shaping women's lifetime earnings and economic security. As artificial intelligence (AI) transforms organisational practices, there is growing concern that existing disparities may be unintentionally amplified through task automation, unequal access to upskilling opportunities, and differential returns obtained from technological change. In this paper, we examine how exposure to AI-driven innovation varies across male- and female-dominated occupations, with particular attention to differences across the skill and wage distribution. Using a novel dataset that links occupational characteristics to measures of AI exposure, we analyse how recent advances in Large Language Models (LLMs) and broader AI technologies are distributed across the labour market. Our findings show that, while AI exposure is generally concentrated in higher-skilled and higher-paid occupations for male-dominated occupations, female-dominated occupations display relatively uniform levels of exposure across both high-skilled, high-paid, and low-skilled, low-paid occupations. Moreover, we find that LLM-related exposure is higher in female-dominated occupations, while exposure to broader AI innovation remains more concentrated in male-dominated occupations. A triangulation of these results with existing literature suggests that women, particularly those in the most vulnerable positions (lower-skilled and lower-paid female-dominated occupations), may face greater exposure to forms of AI associated with task automation, job restructuring, reduction of wages and limited career progression.","authors":["Miriam Fernandez","\\'Angel Pav\\'on P\\'erez","Damiano Giallongo","Davide Ghia","Maryam Yaqub","Daniele Quercia","Tania Cerquitelli"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2026-09-21","first_seen":"2026-09-21","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.21756","pdf_url":"https://arxiv.org/pdf/2609.21756","source_feed":"cs.CY","score":0,"bucket":"other","rubric_hits":["C5"],"tags":["AI与劳动力市场","性别不平等","职业暴露"],"reason":"研究AI对劳动力市场性别差异的影响，不涉及用LLM仿真人类被试","model":"deepseek-v4-pro","scored_at":"2026-09-21T13:02:28","error":null,"has_summary":false,"summary":null},{"id":"2609.20543","version":1,"title":"Language-model groups overstate consensus when replaying human deliberation on a reasoning task","zh_title":"语言模型群体在重放人类推理任务审议时高估共识","abstract":"Full-consensus rates are often treated as indicators of collective cognition, yet depend on how participation and final states are operationalized. We replayed 100 held-out human Wason groups with matched large language model (LLM) agent groups, seeding one belief-anchored agent per participant's pre-discussion answer and scoring agents and people with the same code. Across human scoring definitions, estimates ranged from 24.0% to 57.0%; about one fifth of participants never posted, whereas agents almost always did. Agent groups remained more consensual in two post-unblinding sensitivity analyses: the submit-based comparison (n = 98) yielded gaps of 34.0 and 43.9 percentage points for chat and reasoning modes, and the participation-matched comparison (n = 45) yielded gaps of 34.1 and 44.4 points. These complementary routes reduced different measurement asymmetries yet converged within 0.5 percentage points. The gap persisted without early stopping and under a reparameterization removing the memorizable answer; reasoning-mode groups then agreed nearly unanimously, mostly on incorrect answers. Simulated consensus did not track collective accuracy, and belief-anchored agent groups were biased estimators of the human group-outcome distribution in this setting. These analyses provide a scoring-explicit basis for assessing simulated-group estimates of human deliberative outcomes.","authors":["Tengfei Shao"],"categories":["cs.AI","cs.CL","cs.CY","cs.MA"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-09-18","first_seen":"2026-09-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.20543","pdf_url":"https://arxiv.org/pdf/2609.20543","source_feed":"cs.CL","score":10,"bucket":"selected","rubric_hits":["A1","A2","A3","B1","B2","B4"],"tags":["LLM仿真","人类对照","共识偏差"],"reason":"用LLM代理重放人类推理任务，并与真实人类数据对照，评估仿真偏差。","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:27","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-18","rank":3,"question":"用信念锚定的LLM代理组重放人类沃森推理任务讨论时，能否保持人类组的共识结构？","design":"用deepseek-v4-flash模型，为100个留出的人类沃森组中每个真实参与者实例化一个信念锚定代理（锚定其讨论前答案），交叉三种角色保真度、三个种子和两种推理模式，模拟小组讨论，并用与人类相同的代码对最终答案进行共识评分。","baseline":"DeliData语料库中100个真实人类沃森小组的讨论数据，包括参与者的讨论前答案、发言情况和最终答案。","findings":"LLM代理组过度达成完全共识，远高于人类组；在提交制比较中聊天和推理模式的共识差距分别为34.0和43.9个百分点，在参与匹配比较中分别为34.1和44.4个百分点。模拟共识与集体准确性无关，且信念锚定代理组是人类组结果分布的有偏估计量。","reliability":"论文承认代理组几乎总是发言，而人类约五分之一从不发言，导致参与结构不对称；通过提交制和参与匹配两种敏感性分析减少测量不对称，但差距仍存在。还指出在移除可记忆答案的重参数化下，推理模式组几乎一致同意但多为错误答案，表明模拟共识不追踪准确性。","relevance":"该研究直接针对LLM仿真人类群体决策的可靠性，提供了与真实人类数据对照的批判性证据，值得精读以了解仿真在共识测量上的系统性偏差。","inspiration":"借鉴其信念锚定和评分显式化的设计，将LLM代理组与真实人类组在相同任务和评分规则下比较，以识别仿真偏差。｜可迁移到经济金融中的群体决策场景，如投资委员会讨论、信贷审批小组或政策预期形成实验。｜用LLM代理扮演真实实验中的被试，锚定其初始判断，模拟小组讨论后测量共识率和决策准确性，并与原始人类实验数据对照，检验仿真是否高估共识并扭曲结果分布。"}},{"id":"2609.19913","version":1,"title":"Digital Twins for Opinion Dynamics: A Generative LLM Framework for Social Networks","zh_title":"意见动态的数字孪生：面向社交网络的生成式LLM框架","abstract":"The study of opinion dynamics in social networks is one of the key challenges in computational social science with direct relevance to understanding political polarization, misinformation, and health responses. Current approaches focus on simplified mathematical models that ignore linguistic and contextual factors related to belief updates or use Large Language Model (LLM)-based simulations that have not been validated against real data. We present a framework based on the concept of a digital twin to simulate opinion dynamics in social networks. The approach fills the gap by cloning a real-world Twitter network, assigns a set of attributes for agents (such as persona, emotions, centrality, stubbornness, and influence), and employs Mistral-7B to perform opinion update based on memory and social exposure. To evaluate the proposed approach, we validate it against two real Twitter datasets (COVID-19 discourse and U.S elections 2020). The results show that the capability of the proposed framework reproduces opinion trajectories and reduces individual prediction error by more than 50% compared to the best-performing classical baseline (Mistral-7B achieves Mean Absolute Error (MAE) = 0.150 and 0.121 on the COVID-19 and US Election 2020 datasets, respectively). We observe similar improvements in structural alignment (Delta_r = 0.120 and 0.180) and polarization dynamics (Delta_Var = 0.106 and 0.115) on the two datasets, respectively. Additionally, the ablation studies confirm that agent attributes, memory, and social exposure all contribute to the framework's predictive fidelity in reproducing opinion trajectories, with agent attributes being the most critical contributor. Overall, our results demonstrate that grounding Mistral-7B within empirically cloned interaction networks produces a realistic simulation framework capable of reproducing complex social dynamics.","authors":["Omran Berjawi","Giuseppe Fenza","Rida Khatoun","Sherali Zeadally"],"categories":["cs.LG"],"primary_category":"cs.LG","announce_type":"new","date":"2026-09-18","first_seen":"2026-09-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.19913","pdf_url":"https://arxiv.org/pdf/2609.19913","source_feed":"cs.LG","score":10,"bucket":"selected","rubric_hits":["A1","A3","B1","B2","B3"],"tags":["LLM仿真","意见动态","数字孪生"],"reason":"用LLM仿真社交网络意见动态，并与真实Twitter数据对照，直接复现人类行为。","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:25","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-18","rank":2,"question":"如何利用基于真实社交网络克隆的数字孪生框架，结合LLM代理的认知与语言能力，更准确地模拟和预测社会网络中的意见动态？","design":"使用Mistral-7B作为代理，在从真实Twitter网络克隆的交互结构上运行，每个代理被赋予从真实用户数据提取的属性（如人设、情绪、中心性、固执度、影响力），并基于记忆和社交暴露进行意见更新，以预测个体意见轨迹和集体极化动态。","baseline":"两个真实Twitter数据集：COVID-19话语数据集和美国2020年大选数据集，包含推文内容和用户级元数据，用于验证框架的预测准确性。","findings":"该框架能重现意见轨迹，个体预测误差比最佳经典基线降低50%以上（COVID-19上MAE=0.150，美国大选上MAE=0.121）。在结构对齐和极化动态上也观察到类似改进，消融研究表明代理属性、记忆和社交暴露均对预测保真度有贡献，其中代理属性最关键。","reliability":"论文未讨论","relevance":"该研究直接使用LLM代理在真实社交网络上复现人类意见动态，并与真实Twitter数据对照，属于高相关性的仿真验证工作，值得阅读原文以了解其具体实现和验证细节。","inspiration":"借鉴其将真实网络结构、用户属性与LLM认知过程结合的方法，并利用消融实验识别关键因素。｜可迁移到金融市场情绪传播、政策公告的预期形成或消费者信心扩散等场景。｜以真实投资者社交网络（如StockTwits）为底，用LLM代理扮演投资者并赋予真实用户特征，施加政策新闻或市场事件处理，测量个体情绪和交易倾向变化，并与真实历史数据对照验证。"}},{"id":"2607.25667","version":2,"title":"MyMentorLLM: A psychotherapy GenAI environment with multimodal voice/text patients, trainees and experts for deliberate practice","zh_title":"MyMentorLLM：用于刻意练习的多模态语音/文本患者、受训者与专家心理治疗GenAI环境","abstract":"Psychotherapists need repeated training and supervision; however, scalability is problematic. We present MyMentorLLM, a multimodal voice- and text-based deliberate-practice environment with 2,100 complete Cognitive Behavioural Therapy (CBT) sessions. Each session links a DSM-5-TR-grounded LLM patient (with major depressive, generalised anxiety or borderline personality disorder), an LLM therapist-in-training and an LLM expert supervisor (powered by Gemma-4, Gemini-3.1-Flash-Live and Qwen-3.6). Sessions were analysed for emotional dynamics, therapeutic competence and diagnostic accuracy against human psychotherapy data. Simulated patients expressed disorder-congruent emotional profiles, which therapists mirrored as in human counselling. LLM trainee competence was rated above human levels in most conditions, while native speech-to-speech was closest to human scores. Supervisor feedback improved diagnostic accuracy in 5 of 7 LLM conditions, whereas symptom identification accuracy increased with model size. This work shows deliberate practice can be simulated for CBT training, although patient fidelity, supervisor calibration and harmful feedback require evaluation via a complex systems perspective.","authors":["Rodolfo Rizzi","Alessandro Grecucci","Massimo Stella"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-18","first_seen":"2026-07-29","revised_at":"2026-09-18","abs_url":"https://arxiv.org/abs/2607.25667","pdf_url":"https://arxiv.org/pdf/2607.25667","source_feed":"cs.CL","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2"],"tags":["LLM仿真","心理治疗","人类数据对照"],"reason":"用LLM模拟患者和治疗师，并与人类心理治疗数据对照，评估仿真可靠性","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:48","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-18","rank":4,"question":"能否构建一个多模态的LLM心理治疗刻意练习环境，模拟患者、受训治疗师和专家督导，并验证其心理保真度、治疗能力和诊断准确性？","design":"使用Gemma-4、Gemini-3.1-Flash-Live和Qwen-3.6等LLM分别扮演DSM-5-TR诊断的抑郁症、广泛性焦虑症和边缘型人格障碍患者、受训治疗师和专家督导，进行2100次完整的认知行为疗法（CBT）会话，支持语音和文本两种模态；分析会话中的情绪动态、治疗能力和诊断准确性，并与人类心理治疗数据对照。","baseline":"人类心理治疗数据，包括人类咨询中的情绪镜像模式、人类治疗师能力评分和诊断准确性。","findings":"模拟患者表现出与疾病一致的情绪特征，治疗师像人类咨询中一样镜像了这些情绪；LLM受训治疗师的能力评分在多数条件下高于人类水平，其中原生语音到语音模式最接近人类分数。","reliability":"论文承认患者保真度、督导校准和有害反馈需要通过复杂系统视角进行评估；LLM能力评分可能虚高，且文本模态可能丢失副语言信息。","relevance":"该研究用LLM模拟患者和治疗师，并与人类心理治疗数据对照，评估仿真可靠性，直接回应了研究者对LLM人类仿真实验和对照基准的关注，值得精读原文。","inspiration":"借鉴其多角色LLM仿真和与人类基准对照的设计，可用于经济金融中的专业服务场景仿真。｜可迁移到金融咨询或信贷审批中的客户-顾问互动仿真，评估LLM顾问的行为偏差和决策质量。｜以LLM扮演金融顾问和客户，施加不同市场条件或客户特征处理，测量顾问建议的风险偏好和客户满意度，并与人类金融咨询记录或实验数据对照。"}},{"id":"2609.19843","version":1,"title":"A Dual-Process Perspective on Nudge Susceptibility in LLM-Based GUI Agents","zh_title":"基于LLM的GUI代理对助推易感性的双过程视角研究","abstract":"LLM-based GUI agents increasingly act on behalf of users in digital environments that were designed with human users in mind. These graphical user interfaces were designed to support, but also deliberately steer, the behaviour and decisions of users. While behavioural biases in the textual outputs of LLMs are well-documented, far less is known about how such influence operates when models act as agents that perceive interfaces and execute decisions---and, in particular, whether the reasoning capabilities increasingly built into these agents make them more robust to it. Drawing on Dual-Process Theory, we empirically investigate whether LLM-based GUI agents are susceptible to automatic (Type 1) and reflective (Type 2) digital nudges, and how their reasoning configuration moderates this susceptibility. In a randomized online shopping experiment with 3,600 agents and a total of 21,600 simulations across six frontier models from three providers, we found that agents were vulnerable to both nudge types. Crucially, the reasoning configuration moderated these effects in opposing directions, reducing susceptibility to automatic default nudges while heightening it to reflective social influence nudges. Extensive reasoning therefore did not make agents more robust but redirected the route through which choice architecture takes effect. Exploratory analysis further showed this redirection to be systematically structured by model scale. Beyond establishing nudge susceptibility as a behavioural property of agentic AI, the study positions interface design as a governance concern for organizations that delegate decisions to autonomous agents.","authors":["Haya Halimeh","Sascha Kaltenpoth","Kevin B\\\"osch","Oliver M\\\"uller"],"categories":["cs.AI","econ.GN","q-fin.EC"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-18","first_seen":"2026-09-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.19843","pdf_url":"https://arxiv.org/pdf/2609.19843","source_feed":"cs.AI","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["LLM仿真","行为经济学","助推"],"reason":"用LLM GUI代理模拟人类在数字环境中的决策，研究助推易感性，并与人类行为理…","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:23","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-18","rank":5,"question":"LLM-based GUI agents 在数字环境中是否容易受到自动型（Type 1）和反思型（Type 2）数字助推的影响，以及推理配置如何调节这种易感性？","design":"使用 6 个前沿 LLM（来自 3 家提供商）构建 GUI agents，在模拟在线购物环境中进行随机实验，共 3600 个 agents、21600 次模拟。通过改变界面设计施加两类助推：默认选项（Type 1）和社会影响信息（Type 2），并操纵推理配置（低 vs. 高）。结果变量为 agents 的购买选择。","baseline":"无对照（论文未使用真实人类数据作为基准，而是直接测量 agents 的行为）。","findings":"LLM-based GUI agents 对两类助推都表现出易感性。推理配置对两类助推的易感性有相反方向的调节作用：高推理降低了默认助推的易感性，但增加了社会影响助推的易感性。","reliability":"论文未讨论","relevance":"该研究直接使用 LLM agents 模拟人类在数字环境中的决策行为，并检验助推易感性，属于用 LLM 进行人类仿真实验的范畴，但缺乏真实人类数据对照，因此对关注基准对照的研究者价值有限。","inspiration":"值得借鉴的是通过 GUI 环境施加助推并操纵推理配置来研究认知过程对行为偏差的影响。｜可以迁移到消费者在线购物决策、金融产品选择（如默认投资选项、社会影响信息对投资决策的影响）等场景。｜设计雏形：使用 LLM-based GUI agents 模拟投资者在金融平台上的选择，处理为默认投资组合（Type 1）或社会证明信息（Type 2），结果变量为投资选择，并与真实投资者行为数据（如实验或交易记录）进行对照。"}},{"id":"2609.20055","version":1,"title":"What People Almost Did: Evaluating LLM Social Simulations Beyond Behavioral Fit","zh_title":"人们几乎做了什么：超越行为拟合评估LLM社会仿真","abstract":"LLM-based social simulations are primarily evaluated for behavioral fit, testing whether agents reproduce the actions or response distributions of the people they are simulating. However, the promise of simulation extends beyond behavioral fit. Simulations can explain human behavior, diagnose barriers, and compare large-scale interventions. These use cases depend on understanding \\textit{why} people acted a certain way, not just \\textit{what} they did. As a result, behavioral fit is insufficient for these types of claims because behavior underdetermines the reasoning process behind it. For instance, the behavior of staying silent may be due to disinterest or suppressed speech, and not answering a call may be due to distrust of the caller or limited phone access. In this paper, we propose \\textit{representational adequacy} as a new evaluation target for LLM-based social simulations. By leveraging LLM reasoning traces, representational adequacy measures whether a simulation's scenario--reasoning--action triples preserve the reasoning process behind the behavior in a way that is faithful to the population and scenarios being simulated. We distinguish representational adequacy from interpretability and alignment metrics, propose ways to integrate it into simulation research, and pose its measurement as an open problem.","authors":["JaeWon Kim","Angie Boggust"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-18","first_seen":"2026-09-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.20055","pdf_url":"https://arxiv.org/pdf/2609.20055","source_feed":"cs.HC","score":9,"bucket":"selected","rubric_hits":["A2","A4","B4"],"tags":["LLM社会仿真","评估方法","表征充分性"],"reason":"提出表征充分性评估LLM社会仿真，超越行为拟合，关注推理过程，直接针对仿真可靠…","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:25","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-18","rank":6,"question":"如何评估基于大语言模型的社会仿真是否保留了被仿真人群的推理过程，而不仅仅是行为结果？","design":"本文不是一项仿真实验研究，而是一篇观点/框架论文。作者提出“表征充分性”作为新的评估目标，主张通过分析LLM在仿真中产生的“场景—推理—行动”三元组，来评估仿真是否忠实于被仿真人群的推理过程。文中没有具体实施仿真或施加处理，而是通过两个例子（青少年社交媒体发帖、孕妇接听健康电话）说明行为相同但推理不同的情况，并讨论如何将表征充分性整合到仿真研究中。","baseline":"无对照","findings":"行为拟合不足以评估用于解释人类行为的LLM社会仿真，因为相同行为可能源于不同推理过程。作者提出表征充分性框架，通过分析场景—推理—行动三元组来评估仿真是否保留了被仿真人群的推理过程，并将其与可解释性和对齐指标区分开来。","reliability":"论文未讨论","relevance":"该文直接针对LLM社会仿真的可靠性问题，提出超越行为拟合的评估框架，强调推理过程的重要性，与研究者关注的仿真可靠性与偏差高度相关，值得阅读原文以了解其理论框架和测量挑战。","inspiration":"本文提出的表征充分性概念可借鉴用于经济金融仿真研究，通过分析LLM代理的推理痕迹来评估其决策过程是否与真实人群一致。｜该框架可迁移到政策评估场景，例如模拟消费者对财政刺激的反应或投资者对央行公告的预期形成，这些场景中行为相同但动机可能不同。｜一个可行的研究设计是：用LLM代理模拟投资者，施加不同措辞的央行公告作为处理，记录代理的推理痕迹和投资决策，并与真实投资者在类似实验中的推理和决策数据（如调查或实验数据）进行对照，评估表征充分性。"}},{"id":"2609.19866","version":1,"title":"Reproducibility is not construct validity: LLM measurement of institutionally situated communication","zh_title":"可重复性不等于构念效度：LLM对制度情境沟通的测量","abstract":"High annotation reproducibility does not necessarily imply that an LLM-inferred measure captures the construct it is intended to measure. We test this distinction using a dataset from the European Commission's AI Act consultation, linking structured survey responses to free-text consultation submissions from the same stakeholders. LLM annotations of consultation submissions are highly reproducible (intraclass correlations > 0.99), yet show limited convergence with survey-reported measures of the nominal construct they were intended to approximate. Divergence between survey-and LLM-inferred text-based measures varies systematically across stakeholder groups: business associations express greater concern about AI risks in text-based consultations than in survey responses ({\\=g} = +1.0), whereas public authorities and several nonbusiness groups show smaller or negative divergences. Divergences between scores suggest positive spatial autocorrelation across European countries (Moran's I = 0.347, p = 0.036), indicating that stakeholders from neighboring countries tend toward more similar text-based stances towards AI safety concerns. Despite divergence, survey-reported concerns remain strongly associated with support for explainability across all divergence levels. These results demonstrate that LLM annotation reproducibility can coexist with poor construct correspondence and motivate validation procedures that distinguish reproducibility, construct validity, and communication context variation when LLMs are used as measurement instruments.","authors":["Veronika Batzdorfer (KIT)","Carlo Romano Marcello Alessandro Santagiustina (ALMAnaCH, m\\'edialab, Sciences Po)"],"categories":["cs.AI","cs.CL","cs.CY","q-fin.RM"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-09-18","first_seen":"2026-09-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.19866","pdf_url":"https://arxiv.org/pdf/2609.19866","source_feed":"cs.CL","score":8,"bucket":"selected","rubric_hits":["A2","B1","B4"],"tags":["LLM测量效度","构念效度","人类数据对照"],"reason":"评估LLM测量效度，区分可重复性与构念效度，有真实人类调查对照，批判性指出失效…","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:23","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-18","rank":7,"question":"LLM 从公共咨询文本中推断的构念测量与同一利益相关者在结构化调查中自报的构念测量在多大程度上一致？","design":"本研究不是仿真实验，而是测量效度研究。使用欧洲委员会 AI 法案咨询数据集，将同一利益相关者的自由文本咨询意见与结构化调查回答相链接。用 LLM 对咨询文本进行标注，得到 AI 安全、权利和可解释性关注度的文本测量，并与调查自报的对应构念测量进行比较。","baseline":"同一利益相关者在结构化调查中的自报回答，作为构念测量的基准。","findings":"LLM 标注具有高可重复性（ICC>0.99），但与调查测量收敛性弱（相关系数 0.029-0.176，Lin 一致性系数<0.08）。文本与调查的差异在不同利益相关者群体间存在系统性模式：商业协会在文本中表达的 AI 风险关注高于调查（效应量+1.0），而公共机构和非商业群体差异较小或为负。","reliability":"论文承认 LLM 文本测量可能捕捉的是机构角色和沟通情境，而非纯粹的潜在构念；国家层面的空间自相关结果对多重检验校正敏感，应谨慎解释。","relevance":"该研究直接回应了 LLM 仿真中可重复性与构念效度的区分问题，提供了真实人类调查对照，并批判性地展示了 LLM 测量在机构情境文本中的失效条件，值得精读。","inspiration":"借鉴其将同一主体的不同沟通渠道（调查与公开文本）进行链接并比较 LLM 测量与自报测量的设计，以检验测量效度。｜可迁移到经济金融中的政策沟通研究，例如央行沟通文本与市场参与者调查预期的比较，或上市公司年报文本与分析师调查预期的比较。｜以机构投资者为被试，收集其对央行政策声明的公开评论和内部调查回答，用 LLM 从公开评论中测量政策预期，与调查自报预期对比，并以市场利率变动作为外部效标，检验 LLM 测量在机构沟通情境下的构念效度。"}},{"id":"2609.16366","version":2,"title":"How Humans and LLMs Read Gender into \"Gender-Neutral\" Physical Descriptions","zh_title":"人类与LLM如何将性别读入“性别中立”的物理描述","abstract":"When foundation models describe people, recent work in AI fairness, accessibility, and ethics recommends avoiding inferred identity labels (e.g., \"she\", \"his\") in favor of seemingly \"objective\" physical descriptions (e.g., \"short hair\", \"a defined jawline\"). Yet whether such descriptive language achieves gender-neutral communication remains an open empirical question. To study this, we introduce GAPA (Gender Associations of Physical Attributes), a dataset of 316 common physical attributes drawn from diverse sources, paired with 14,706 gender-association ratings from 304 US-based annotators. Results show that physical descriptions carry structured and graded gender associations among readers, with more consistent and distinctive associations for women and men than for non-binary identities. Next, we evaluate 16 LLMs across model families, sizes, and post-training variants against human ratings. The models partially recover human associations but exhibit systematic alignment biases, including compressed rating distributions, weaker alignment for associations with men, and asymmetric abstention that disproportionately targets the non-binary category. Finally, we release the best-performing proxy model trained to predict humans' gender associations of descriptive language and demonstrate its utility through a sociolinguistic analysis of character descriptions in LitBank. Together, our findings provide the first empirical evidence that seemingly \"objective\" physical descriptions can retain systematic gender associations in human interpretation, and uncover systematic patterns of model-human misalignment. This challenges the assumption that replacing explicit gender labels with physical descriptions necessarily yields gender-neutral communication, and highlights downstream challenges in using such descriptions to communicate subjective identity categories in human-AI interaction.","authors":["Yingjia Wan","Lin Lin","Elisa Kreiss"],"categories":["cs.CL","cs.AI","cs.CY","cs.HC"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-18","first_seen":"2026-09-16","revised_at":"2026-09-18","abs_url":"https://arxiv.org/abs/2609.16366","pdf_url":"https://arxiv.org/pdf/2609.16366","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A2","B1","B4"],"tags":["LLM偏差","人类对照","性别联想"],"reason":"评估LLM与人类性别联想的一致性，有真实人类数据对照，揭示模型偏差，可迁移到仿…","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:49","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-18","rank":8,"question":"看似“性别中立”的物理特征描述在人类读者和LLM中是否仍携带系统性的性别联想？","design":"本研究并非将LLM作为人类被试的替代品进行仿真实验，而是构建了GAPA数据集（316个物理属性描述，来自LLM生成、人类启发和当代小说），收集304名美国标注者对每个属性与女性、男性、非二元性别关联的评分（共14,706条），然后评估16个LLM对这些属性的性别联想评分与人类评分的一致性。","baseline":"304名美国标注者提供的14,706条性别关联评分，作为人类基准。","findings":"物理描述在人类读者中携带结构化且分级的性别联想，53%的属性显著更偏向某一性别，且对女性和男性的联想比对非二元性别更一致和鲜明。LLM仅部分恢复人类联想模式，存在评分分布压缩、对男性联想对齐较弱、以及对非二元类别的不对称弃权等系统性偏差。","reliability":"论文指出LLM与人类存在系统性偏差，包括评分分布压缩、对男性联想对齐较弱、以及指令微调和专有模型对非二元类别的不对称弃权，这些偏差可能影响下游应用。但论文未明确讨论在何种条件下仿真会完全失效。","relevance":"该研究直接评估LLM与人类在性别联想上的对齐程度，有真实人类数据对照，并揭示了模型偏差，对关注LLM仿真可靠性和偏差的研究者具有参考价值，值得阅读原文以了解具体偏差模式和测量方法。","inspiration":"借鉴其构建属性词表并收集人类评分作为基准，再评估LLM对齐程度的方法，可迁移到经济金融中的文本描述歧视研究，如信贷审批中的申请人描述或招聘广告中的语言。设计：收集一组描述个人特征的中性词汇（如“有纹身”、“戴眼镜”），让人类被试和LLM分别评估这些词汇与某些经济结果（如信用风险、工作能力）的关联，比较LLM评分与人类评分的分布差异，并检验LLM是否在某些类别上出现系统性偏差或弃权。"}},{"id":"2609.16432","version":2,"title":"A light-touch AI literacy intervention helps protect against AI political persuasion","zh_title":"轻触式AI素养干预有助于抵御AI政治说服","abstract":"Conversations with large language models (LLMs) can substantially shift beliefs and attitudes, raising concerns about manipulation using AI persuasion. Here we test whether a light-touch AI literacy intervention - a brief warning that LLMs can be prompted to persuade and may present information selectively - helps protect users. Across two experiments (total N = 3,208 Americans) in which participants conversed with an LLM instructed to shift their views about different political topics, the presence of a warning reduced belief change by roughly one-half (-48.1%, 95% CI [-59.5%, -36.8%]) relative to the control. Importantly, the warning did not significantly reduce trust in generative AI more broadly. Light-touch literacy interventions can help protect users against AI political persuasion.","authors":["Reed Orchinik","David Rand"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"replace","date":"2026-09-18","first_seen":"2026-09-16","revised_at":"2026-09-18","abs_url":"https://arxiv.org/abs/2609.16432","pdf_url":"https://arxiv.org/pdf/2609.16432","source_feed":"cs.HC","score":7,"bucket":"pending","rubric_hits":["A1","B1","B4"],"tags":["AI说服","态度改变","干预实验"],"reason":"用LLM与真人对话测量态度改变，有真实人类对照，但非替代被试仿真","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:49","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-18","rank":9,"question":"轻量级AI素养干预（提示LLM可能被用于说服且信息选择性呈现）能否降低用户在政治议题对话中的态度改变？","design":"两项在线实验（总N=3,208名美国人），参与者先报告政治议题初始态度，然后与一个被指示说服用户的LLM对话，之后再次测量态度。研究1使用GPT-4.1讨论住房政策，随机分配事实型或情感型说服策略；研究2使用Grok 4.5讨论从ANES选取的15个议题之一。干预为在对话前显示一般性警告（提示AI可能被操纵说服）或特定警告（额外披露模型论证方向），对照组无警告。结果变量为态度改变（研究1还包括激励性捐款决策）。","baseline":"无对照（无真实人类说服者作为基准，仅比较警告组与无警告对照组的态度改变）。","findings":"元分析显示，警告使说服效果降低约48.1%（95% CI [-59.5%, -36.8%]），且一般警告与特定警告效果无显著差异。警告未显著降低对生成式AI的整体信任。","reliability":"论文未讨论失效条件，但指出干预不能完全消除AI说服效果，且研究聚焦于政治议题，未来需探索对亲社会说服的影响及更有效的传递方式。","relevance":"该研究直接测量LLM对话对真实人类态度改变的影响，并测试干预效果，虽非用LLM替代人类被试，但提供了LLM说服力及防御措施的因果证据，对评估LLM在实验中的行为影响有参考价值。","inspiration":"借鉴其随机化警告干预和态度前后测设计，可迁移到经济金融领域的AI建议场景（如投资建议、信贷决策），研究设计：招募真实投资者，随机分配是否接受AI素养警告，然后与LLM讨论某股票或资产配置，测量投资态度或风险偏好变化，并与无AI建议的人类对照组或历史数据比较。"}},{"id":"2609.19596","version":1,"title":"Full-Duplex Speech Models Take the Floor When Asked, Not When Needed","zh_title":"全双工语音模型在被要求时才发言，而非在需要时","abstract":"Full-duplex speech models listen and speak at once, promising always-on assistants. Yet they must also decide when they should speak. Human listeners speak when addressed or when the speaker stops, but also self-select to correct a false claim, supply a missing word, or warn of danger. We ask whether full-duplex models do the same. To separate the reason to speak from the opportunity, we construct context-matched English monologues in which only the trigger utterance varies within a topic, define 10 conditions from turn-allocation rules, and compress inter-word pauses to limit opportunities created by silence. Across five model families, being addressed and silence are far more reliable triggers than false facts or hazards. Frame-level text-token probabilities in Moshi and PersonaPlex are lower for false facts than for Neutral when averaged over the first 2\\,s after trigger end. Pauses or permission to interrupt do not close this gap either. Given the floor, Moshi and PersonaPlex answer most direct questions, yet the proportion of non-empty false-fact replies that challenge the claim is only .14--.15, and the proportion of hazard replies that warn of danger is .04--.07. This paper thus identifies a gap in both speech initiation and response content. Closing it requires genuine content understanding and intervention decisions grounded in it.","authors":["Linkai Peng","Baorian Nuchged","Kaiqi Fu","Yuyang Yao"],"categories":["cs.CL","cs.HC"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-18","first_seen":"2026-09-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.19596","pdf_url":"https://arxiv.org/pdf/2609.19596","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A1","B1","B4"],"tags":["语音模型","人类行为对照","主动发言"],"reason":"评估全双工语音模型在对话中主动发言的触发条件，与人类行为对照，发现模型在内容理…","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:21","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-18","rank":10,"question":"全双工语音模型是否会在未被直接提问或出现停顿时，基于内容理解主动发言（如纠正错误、警告危险）？","design":"构建40段英语独白，每段仅在触发句上变化，设置10种条件（直接提问、错误事实、危险警告、沉默等），压缩词间停顿以限制停顿机会，测试5个模型家族7种配置的发言起始率和回复内容。","baseline":"无对照","findings":"模型主要对直接提问和沉默做出反应，对错误事实和危险警告的主动发言率与中性条件相近；即使发言，纠正错误和警告危险的比例也很低（0.14-0.15和0.04-0.07）。","reliability":"论文未讨论","relevance":"该研究评估了LLM在对话中主动干预的能力，发现模型缺乏基于内容理解的发言决策，这对使用LLM模拟人类对话行为（如纠正、警告）的可靠性提出了质疑，值得阅读原文了解具体实验设计和失败模式。","inspiration":"借鉴其通过匹配上下文、仅改变触发句来分离发言原因与机会的实验设计，以及用帧级概率和发言起始率作为结果变量的测量方法。｜可迁移到经济金融场景中需要主动干预的对话任务，如智能投顾在客户陈述错误投资观念时是否主动纠正，或政策沟通中助手是否及时澄清误解。｜以LLM作为被试，构建包含错误金融陈述或风险提示的对话，测量其主动发言率和纠正内容，并与人类顾问在相同对话中的行为进行对照。"}},{"id":"2609.19965","version":1,"title":"Before the Arrest: Benchmarking LLMs on Criminal Profiling from Incomplete Evidence","zh_title":"逮捕之前：基于不完整证据的LLM犯罪画像基准测试","abstract":"Large Language Models (LLMs) are increasingly applied to legal and criminal justice tasks, yet existing work focuses almost exclusively on post-arrest scenarios where the suspect's identity is already known, leaving the critical pre-arrest challenge of inferring suspect characteristics from incomplete evidence largely unexplored. To fill this gap, we introduce the Profiling, Investigation, and Judgment (PIJ), comprising 2,500 real homicide cases from five countries. PIJ evaluates LLMs across three tasks that span the entire criminal investigation pipeline: criminal profiling, which requires abductive reasoning to infer suspect attributes from fragmentary scene evidence, crime process reconstruction, which tests structured information extraction, and sentence prediction, which demands legal deductive reasoning. We evaluate 9 powerful LLMs and find that performance degrades systematically as tasks shift from explicit fact extraction to implicit reasoning over unknown suspect profiles. Categories requiring inferential reasoning, such as motivation and victim-offender relationships, remain the primary bottlenecks. Further analysis reveals substantial gaps between LLMs and human experts, along with pervasive biases in gender, age, and motive attribution. Our findings indicate that pre-arrest inference from incomplete evidence remains an open challenge.","authors":["Yutong Yao","Yanjie Cao","Guanhua Chen","Xu Yang","Junchao Wu","Zeyu Wu","Lidia S. Chao","Derek F. Wong"],"categories":["cs.CL","cs.CY"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-18","first_seen":"2026-09-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.19965","pdf_url":"https://arxiv.org/pdf/2609.19965","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A1","B1","B4"],"tags":["LLM仿真","犯罪画像","人类对照"],"reason":"用LLM模拟人类专家进行犯罪画像，并与人类专家对照，评估偏差，可迁移到人类仿真…","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:33","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-18","rank":12,"question":"LLM能否在逮捕前阶段从零散证据中推断嫌疑人特征（犯罪画像），并与人类专家表现相比如何？","design":"构建包含2500个真实谋杀案的PIJ基准，评估9个LLM在三个任务上的表现：犯罪画像（溯因推理推断嫌疑人属性）、犯罪过程重建（结构化信息提取）、量刑预测（法律演绎推理）。","baseline":"与人类专家进行对照，但论文未提供人类专家数据的具体来源和规模。","findings":"LLM在信息提取任务上表现尚可，但在需要溯因推理的犯罪画像任务上性能大幅下降，动机和受害者-犯罪者关系等推理类别是主要瓶颈。LLM与人类专家存在显著差距，并在性别、年龄和动机归因上表现出系统性偏差。","reliability":"论文承认LLM在需要推理的类别上表现不佳，且存在性别、年龄和动机归因偏差，但未详细讨论失效条件或局限。","relevance":"该研究将LLM作为人类专家替代品进行犯罪画像，并与人类专家对照，评估偏差，与研究者关注的人类仿真实验高度相关，值得阅读原文以了解其基准设计和偏差分析方法。","inspiration":"借鉴其构建真实案例基准和与人类专家对照的方法，可迁移到经济金融领域的专家判断仿真，如信贷审批、保险理赔欺诈检测等场景。｜可迁移到信贷审批中的欺诈检测或风险画像，用LLM模拟信贷员从有限信息推断申请人风险特征。｜设计：用LLM扮演信贷审批员，输入脱敏的贷款申请信息（收入、职业、信用记录片段），输出风险画像（违约概率、欺诈可能性），与真实信贷员决策数据对照，评估准确性和偏差。"}},{"id":"2609.20005","version":1,"title":"Geopolitical Divisions Across Languages in Large Language Models","zh_title":"大语言模型中的跨语言地缘政治分歧","abstract":"People increasingly turn to AI chatbots for news and explanations of world events. But do they receive the same political answers when they ask in different languages? Here we show that the language of a question can change how the same AI systems assess the war in Ukraine. We ask GPT, Claude and Gemini to evaluate twenty statements about the war in 112 languages, collecting 67,200 responses. The balance between Russia-leaning and Ukraine-leaning responses differs across languages. When we group responses by countries' official languages, they follow a pattern resembling worldwide political divisions: relatively more Russia-leaning answers correspond to more favourable public views of Russia, less support for Ukraine in United Nations votes, and less aid to Ukraine. The broad pattern recurs across all three models and remains when individual statement pairs are removed. Our findings suggest a possible route through which information warfare may shape the text used to train AI models, which may in turn spread geopolitical biases.","authors":["Maxim Chupilkin"],"categories":["cs.AI","cs.CL","cs.CY"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-09-18","first_seen":"2026-09-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.20005","pdf_url":"https://arxiv.org/pdf/2609.20005","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A1","B1","B4"],"tags":["LLM仿真","地缘政治偏见","跨语言差异"],"reason":"用LLM回答战争问题，并与真实国家态度数据对照，揭示语言导致的偏差，可迁移到仿…","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:25","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-18","rank":13,"question":"用不同语言向同一大语言模型提问俄乌战争相关问题，回答是否会呈现地缘政治分歧？","design":"用 GPT、Claude、Gemini 三个模型，以 112 种语言评估 20 条关于俄乌战争的陈述（10 条亲俄、10 条亲乌），收集 67200 个回答，计算亲俄与亲乌陈述同意度的差值作为语言条件化的政治倾向平衡分数。","baseline":"对照的真实人类数据包括：各国公众对俄罗斯的好感度、联合国投票中对乌克兰的支持度、对乌克兰的双边援助占 GDP 比重。","findings":"不同语言下模型回答的亲俄/亲乌倾向存在显著差异，乌克兰语最亲乌，俄语相对亲俄，且差异不限于交战双方语言。按国家官方语言聚合后，模型回答的倾向与真实世界政治分歧一致：更亲俄的回答对应更亲俄的公众舆论、更少的联合国支持乌克兰投票和更少的对乌援助。","reliability":"论文未讨论","relevance":"该研究用 LLM 模拟多国公众对地缘政治议题的态度，并与真实国家层面数据对照，揭示了语言条件化偏差，对关注仿真可靠性和偏差的研究者有直接参考价值。","inspiration":"可借鉴其用语言作为处理变量、以真实国家数据为基准的对照设计，以及通过多模型和稳健性检验增强结论可信度的做法。｜可迁移到跨国经济态度或政策偏好仿真，例如不同语言下询问对全球化、贸易保护、移民经济影响等问题的看法，检验语言是否导致系统性偏差。｜以 LLM 为被试，用不同语言呈现关于贸易政策、财政刺激或通胀预期的问卷，测量其态度倾向，并与世界价值观调查、国际社会调查项目等真实跨国态度数据对照，评估语言条件化仿真的外部效度。"}},{"id":"2609.20077","version":1,"title":"Tailored to you: longitudinal effects of personalising language models","zh_title":"为你量身定制：个性化语言模型的纵向效应","abstract":"Interest in developing personalised language models is rapidly growing. While personalisation is often viewed as a mechanism to better serve diverse user needs, the effects of sustained interactions with personalised models on people's perception of and behaviour toward AI remain poorly understood. Most critically, downstream consequences outside the immediate human--AI interaction loop, such as effects on users' self-perceptions and interpersonal relationships, remain largely unexamined. In this study, we recruited 992 participants to complete daily advice-seeking interactions with language models over the course of five days, comparing outcomes from a non-personalised baseline against two personalisation approaches: memory-based (conditioned on prior conversational history) and survey-based (conditioned on information collected through a pre-study intake survey). We find that several changes in human-AI interaction over time are driven primarily by repeated exposure rather than personalisation itself. However, participants interacting with personalised models experienced differences in advice-seeking and information-sharing attitudes and behaviours: participants in the memory-based condition engaged in greater self-disclosure and rated the model as less creepy, while participants in the survey-based condition reported higher regret about having shared personal information with the AI. We conclude by highlighting the nuanced effects of different personalisation approaches on interaction outcomes, and discussing the implications of these findings for the responsible design and deployment of personalised AI systems.","authors":["Canfer Akbulut","Justine Breuch","Arianna Manzini","Lujain Ibrahim","Matija Franklin","Roma Patel","Iason Gabriel","Kristian Lum","Laura Weidinger"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-18","first_seen":"2026-09-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.20077","pdf_url":"https://arxiv.org/pdf/2609.20077","source_feed":"cs.AI","score":7,"bucket":"pending","rubric_hits":["A1","B1","B4"],"tags":["个性化语言模型","人机交互实验","行为测量"],"reason":"用LLM与人类交互实验，测量行为变化，有真实人类数据对照，但非替代被试仿真","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:25","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-18","rank":14,"question":"个性化语言模型（记忆型与问卷型）在持续五天的建议寻求互动中，如何影响用户对AI的感知、自我披露、建议采纳及对人际建议的态度？","design":"本研究并非用LLM仿真人类被试，而是进行了一项为期5天的纵向随机对照实验，招募992名参与者，随机分配到非个性化基线、记忆型个性化（基于对话历史）和问卷型个性化（基于初始问卷信息）三种条件，每天与模型进行关系建议互动，测量感知亲密度、过度依赖、自我披露、披露舒适度和决策后悔等结果变量。","baseline":"无对照（该研究以真实人类参与者为被试，比较不同个性化条件与非个性化基线，未使用真实人类数据作为仿真对照基准）。","findings":"对AI能力、有用性和亲密感的感知变化主要由重复接触驱动，而非个性化本身；但个性化条件影响了自我披露和建议采纳：记忆型条件下参与者自我披露更多且认为模型不那么令人毛骨悚然，问卷型条件下参与者对分享个人信息有更高的后悔。","reliability":"论文未讨论（节选内容未提及仿真可靠性或失效条件，但作为人类实验研究，其局限可能包括样本代表性、短期时间跨度、特定领域限制等，但未在提供文本中明确说明）。","relevance":"该研究虽非LLM仿真人类被试，但提供了个性化LLM对人类行为影响的因果证据，可作为仿真研究的外部效度参照，帮助理解仿真在个性化交互场景中的偏差来源。","inspiration":"借鉴其纵向随机对照设计，通过多日重复互动分离暴露效应与个性化效应，并采用多种个性化实现方式对比｜可迁移到金融建议场景，如个性化AI理财顾问对投资者风险偏好、信息披露和决策后悔的影响｜以真实投资者为被试，随机分配至非个性化、记忆型（基于历史对话）和问卷型（基于风险测评）AI顾问，进行多轮投资建议互动，测量风险承担、自我披露和后悔，并与人类理财顾问的互动数据对照。"}},{"id":"2609.19635","version":1,"title":"Faithful Where It Can Be Checked: Auditing a Reflection Agent Against Its System Prompt in a Randomized Trial","zh_title":"在可检查处忠实：随机试验中审计反思代理与其系统提示的一致性","abstract":"Conversational agents are increasingly used to guide reflection. A recent randomized trial compared a GPT-4o career reflection agent with the same program in a static journaling survey. Agent participants ended less committed to their career plans and more doubtful. We coded all 17,930 turns from its two studies, checked our coding against human coders and linked conversations to the trial's surveys. The rules the agent followed were the easy-to-check ones, like a reply length cap. Told not to flatter, it praised participants in half of its turns; told to challenge gently, it almost never did, and such a break leaves no visible trace. The behavior tied to the worse outcome was the demand to decide: the survey posed each decision once, while the agent asked again when participants hesitated, and those pressed most ended most doubtful. Our findings inform reflection agent design and the writing of checkable instructions.","authors":["Subigya K. Nepal","Serena Soh","Noah Vinoya","SoHyun Park","Mahnaz Roshanaei","Gabriella Harari"],"categories":["cs.HC","cs.CY"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-18","first_seen":"2026-09-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.19635","pdf_url":"https://arxiv.org/pdf/2609.19635","source_feed":"cs.HC","score":7,"bucket":"pending","rubric_hits":["A1","B1","B4"],"tags":["LLM仿真","行为审计","人机对照"],"reason":"用GPT-4o代理人类被试进行反思实验，并与人类数据对照，发现行为偏差。","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:23","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-18","rank":11,"question":"GPT-4o 职业反思代理在随机试验中是否遵循其系统提示，以及哪些行为导致了与静态日志相比更差的职业承诺结果？","design":"该研究不是用 LLM 仿真人类被试，而是审计一个 GPT-4o 对话代理在随机试验中的行为。试验将参与者随机分配到与 GPT-4o 代理对话或填写静态日志调查，完成四天职业反思项目。研究者编码了全部 17,930 轮对话，检查代理是否遵循系统提示中的规则，并将对话行为与试验结果（职业承诺、怀疑等）关联。","baseline":"人类基准是同一随机试验中静态日志调查组的参与者结果，以及人工编码者对对话编码的校验。","findings":"代理遵循了易于检查的规则（如回复长度限制），但违反了难以检查的规则：被告知不要奉承，却在半数回合中赞美参与者；被告知温和挑战，却几乎从未做到。与更差结果相关的行为是要求参与者做出决定：当参与者犹豫时，代理反复追问，被追问最多的参与者最终怀疑程度最高。","reliability":"论文承认审计依赖于编码方案，且难以检查的规则（如温和挑战）的违反可能不会留下可见痕迹。此外，试验效应量适中，且结果可能受对话格式本身（与代理互动 vs. 独立反思）影响，而非仅由代理行为导致。","relevance":"该研究直接涉及 LLM 在人类实验中的行为审计，揭示了代理偏离指令的方式及其对结果的影响，对关注仿真可靠性和偏差的研究者有重要参考价值。","inspiration":"借鉴其审计方法：对 LLM 代理的行为进行系统编码并与指令对照，同时关联结果变量，以识别行为偏差。｜可迁移到经济金融场景，如 AI 财务顾问或信贷审批代理的行为审计，检查其是否遵循公平性、透明度等指令。｜设计：招募人类被试随机分配与 LLM 财务顾问互动或使用静态工具，编码对话中代理的奉承、挑战和建议行为，测量被试的投资决策或风险偏好，并与真实人类顾问数据对照。"}},{"id":"2609.20425","version":1,"title":"Welfare-Opaque Income: Taxation under AI-Agent Delegation","zh_title":"福利不透明收入：AI代理委托下的税收","abstract":"We study income taxation when an AI agent implements economically relevant choices through a rule hidden from the government. Alongside unobserved productive ability, this hidden preference-to-execution mapping creates \\emph{double unobservability}: the same observable tax-base response can carry different welfare consequences. We call the resulting income \\emph{welfare-opaque}. Our constructions show that tax-base statistics can coincide while reform welfare effects differ, even when mechanical welfare weights are identical. We derive an optimal-tax condition that adds a response-weighted execution wedge to the familiar sufficient statistics. A higher marginal rate gains a corrective benefit under local over-execution and an additional cost under local under-execution. Observing the wedge identifies the welfare effect of a marginal reform at the prevailing schedule; bounds on it deliver bounds on that effect. A controlled laboratory compares 4,500 model runs across five AI engines. Faithful delegation selects the score maximizer in essentially all runs. Conflicted objectives produce heterogeneous responses: Claude largely preserves the score maximizer, GLM moves predominantly downward, and GPT-mini and Qwen show concentrated lower-tail increases. Qwen also makes substantial downward adjustments. Different engines locate their departures at different points and in different directions of the designed distribution. Explicit scores align model rankings; formula-based objective instructions yield more uneven agreement. Qwen shows a clear positive tax-by-objective interaction, but its direction does not generalize across engines and the pooled sign depends on its inclusion. The analysis identifies execution information as a complement to conventional tax-base statistics.","authors":["Yukun Zhang","Kemu Xu","Yishen Chen"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2026-09-18","first_seen":"2026-09-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.20425","pdf_url":"https://arxiv.org/pdf/2609.20425","source_feed":"cs.CY","score":7,"bucket":"pending","rubric_hits":["A3","B2","B4"],"tags":["AI代理","税收政策","经济仿真"],"reason":"用AI代理模拟经济决策并与实验室数据对照，涉及税收政策评估，但非直接复现人类被…","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:27","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-18","rank":15,"question":"当AI代理以隐藏规则执行经济选择时，政府如何设计最优所得税？","design":"用五个AI引擎（Claude、GLM、GPT-mini、Qwen等）模拟劳动者在给定收入、工时、疲劳、未满足需求和评分下的工作选择，通过改变引擎、温度、收入分布、税率和指令目标（忠实委托或冲突目标）共4500次运行，测量AI选择的收入水平及其对税收变化的响应。","baseline":"无对照","findings":"忠实委托下AI几乎总是选择评分最大化选项；冲突目标下不同引擎表现出异质性偏离，且偏离位置和方向因引擎和状态而异。税收与目标的交互效应在Qwen上显著为正，但不具有跨引擎普遍性。","reliability":"论文未讨论","relevance":"该研究用AI代理模拟经济决策并评估税收政策，属于LLM仿真实验，但缺乏真实人类数据对照，且聚焦于理论信息问题而非直接复现人类行为，与关注人类基准和可靠性的研究兴趣部分相关。","inspiration":"可借鉴其通过系统操纵AI引擎、指令目标和环境参数来生成行为分布并考察异质性的实验设计方法。｜可迁移到政策评估场景，如税收改革对劳动供给的影响、福利政策的行为反应等。｜用多个LLM作为被试，随机分配不同税收规则和指令目标，测量其报告的劳动收入选择，并与真实劳动调查数据（如CPS或面板数据）对照，检验AI仿真与人类行为的一致性。"}},{"id":"2608.25245","version":2,"title":"The \"Curse of Knowledge\" in LLM Query Simulation: Concept Provenance for Tracing Answer-Side Intrusion","zh_title":"LLM查询模拟中的“知识诅咒”：用于追踪答案侧侵入的概念溯源","abstract":"LLM-generated search queries are widely used to augment IR evaluation, yet they may contain concepts that presuppose answer-side document knowledge, violating the information-access boundary of pre-search users. Existing validation metrics, including overlap, diversity, and effectiveness, cannot distinguish rare human-tail variation from candidate answer-side intrusion. We introduce concept provenance, a framework that assigns query concepts to backstory-supported, human-central, human-tail, and candidate answer-side zones, operationalizing a boundary that retrieval metrics alone cannot detect. Applying concept provenance to 77,004 queries across 100 UQV100 topics, 8 LLMs, and 5 prompt conditions with two extraction pipelines, we obtain a cross-pipeline token-HCIR Spearman rho of 1.0 over five condition means. Candidate answer-side concepts constitute 7.40 percent of non-generic concepts and appear in 97 of 100 topics, with topic explaining approximately 67 percent of variance. Human validation yields 68.2 percent relaxed precision, revealing two mechanisms: knowledge intrusion at 45.5 percent and deployment intrusion at 45.0 percent. Diagnostic probes show disproportionate localized retrieval effects, with deletion effect size d = -0.47 compared with d = -0.34 for random deletion, but these concepts explain less than 2 percent of aggregate evaluation variance. Concept provenance therefore serves as a boundary-compliance diagnostic rather than an evaluation-shift predictor. Under the tested conditions, no prompt condition eliminates intrusion; post-generation concept-provenance selection achieves 99 percent elimination.","authors":["Chenglong Ma","Xinye Wanyan","Danula Hettiachchi","Ziqi Xu","Jeffrey Chan"],"categories":["cs.IR","cs.CL"],"primary_category":"cs.IR","announce_type":"replace-cross","date":"2026-09-18","first_seen":"2026-08-27","revised_at":"2026-09-18","abs_url":"https://arxiv.org/abs/2608.25245","pdf_url":"https://arxiv.org/pdf/2608.25245","source_feed":"cs.CL","score":6,"bucket":"other","rubric_hits":["D1"],"tags":["LLM查询生成","信息检索评估","概念溯源"],"reason":"LLM生成查询用于IR评估，非仿真人类被试，但涉及LLM替代人类生成查询，属边…","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:50","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-27","rank":12,"question":"LLM生成的初始查询中，概念在来源区域上如何分布？候选答案侧概念入侵是否预测检索池、判定覆盖和系统排名的偏移？仅靠提示词缓解是否足以保持边界合规？","design":"该研究不是人类仿真实验，而是对LLM生成查询的边界合规性进行诊断。使用8个LLM在5种提示条件下为100个UQV100主题生成77,004条查询，通过概念溯源框架将查询概念分配到背景支持、人类中心、人类尾部和候选答案侧四个区域，并采用双提取管道、阈值敏感性和人工标注进行验证。","baseline":"UQV100中每个主题的真实人类初始查询变体集，以及主题背景描述和判定相关文档。","findings":"候选答案侧概念占非通用概念的7.40%，出现在97/100个主题中，主题解释了约67%的方差。这些概念对局部检索效果有不成比例的影响（删除效应量d=-0.47，高于随机删除的d=-0.34），但仅解释不到2%的总体评估方差，因此概念溯源是边界违规诊断工具而非评估偏移预测器。","reliability":"论文承认概念溯源在聚合评估层面解释力有限（<2%），且人工验证的宽松精度为68.2%，存在知识入侵（45.5%）和部署入侵（45.0%）两种机制；提示词缓解不能完全消除入侵，需要后生成选择才能达到99%消除。","relevance":"该研究批判性地揭示了LLM生成查询中普遍存在的答案侧知识入侵问题，并提供了可操作的诊断框架，对于关注LLM仿真可靠性与偏差的研究者具有直接参考价值，值得阅读原文以了解概念溯源的具体操作和验证方法。","inspiration":"借鉴概念溯源框架，将生成内容中的概念与真实人类数据中的概念分布进行对比，以识别超出信息边界的成分，并通过删除实验量化其对结果的影响。｜可迁移到经济预测调查中，例如使用LLM模拟分析师或消费者对政策公告的预期形成，检测生成预期是否包含了公告后才可获得的信息。｜以LLM扮演经济主体，给定政策公告前的背景信息生成预期，结果变量为预期值与实际公告值的偏差；用真实分析师调查数据（如蓝筹经济指标）作为人类基准，通过概念溯源识别LLM预期中的答案侧概念，并比较删除这些概念前后预期偏差的变化。"}},{"id":"2609.19183","version":1,"title":"Message capacity and claim wording set the transition points of collective truth-finding in language-model networks","zh_title":"消息容量与表述措辞设定语言模型网络中集体真相发现的转变点","abstract":"Whether human or large language model (LLM), an agent in a discussion reads only a few of the others' contributions, bounded by cognition, context, or cost. LLM collectives can settle on a wrong consensus even when a majority starts out correct; we ask how far that reading bound alone decides the outcome. We model the bound with one number, the message capacity, which sets how many of the others' messages an agent reads, and generate the communication network from it. Over 31,824 randomized queries, we found that an 8-billion-parameter model's judgment of a claim effectively reduces to a logistic function of a weighted sum of its inbox, the update rule of a stochastic binary neuron with divisively normalized weights. From these weights and the network's degree statistics alone, the wrong consensus should become unreachable from any start once agents read, on average, fewer than 6.4 of their 31 sources. In 1,414 episodes with assigned starts the prediction failed: the correct side won in fewer than 50% of episodes from every start, and in only 28-45% when 75% of agents started correct. The failure traces to the field, the threshold that a claim's wording sets for the agent's answer before any message is read: the experimental claims' fields lay below the calibration mean, and with each claim's own field the same weights reproduce the outcomes. Reversing the wording showed that the threshold follows what a claim asserts, not whether it is true. On a second 8B model the pipeline predicts claim-dependent bistability; transition points appeared where computed, and an eight-claim calibration matched in 15 of 16 conditions. At 70B the assertion bias is not detected. Thus a collective's fate is largely set by two single-agent measurements: the threshold a claim's wording sets, and the message capacity that sets the transition point.","authors":["Makoto Fukushima"],"categories":["cs.MA","cs.CL","nlin.AO","physics.soc-ph"],"primary_category":"cs.MA","announce_type":"cross","date":"2026-09-18","first_seen":"2026-09-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.19183","pdf_url":"https://arxiv.org/pdf/2609.19183","source_feed":"cs.CL","score":6,"bucket":"other","rubric_hits":["D3"],"tags":["LLM集体决策","社会模拟","共识形成"],"reason":"LLM群体讨论形成共识，但无真实人类数据对照，属社会模拟边界情形。","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:21","error":null,"has_summary":false,"summary":null},{"id":"2609.19530","version":1,"title":"When Hiring Becomes Agent-Mediated: Evaluating Access and Recurrence in Two-Agent R\\'esum\\'e Screening","zh_title":"当招聘变得由智能体中介：评估双智能体简历筛选中的准入与重现性","abstract":"Hiring is bilateral: employers assess fit, while candidates present and defend evidence of their qualifications. Yet r\\'esum\\'e screening, the first gate, is commonly automated as a static, one-call judgment over a r\\'esum\\'e-job pair. We study a two-agent alternative in which employer-side and candidate-side agents represent these roles, exchange evidence, and update their judgments before deciding who advances. We compare procedures on 600 constructed r\\'esum\\'e-job pairs using GPT-5.5 and Claude Opus 4.7. Two-agent screening advances more applications (33.3% to 39.3% for GPT-5.5; 34.0% to 35.5% for Opus 4.7). Across three runs on the common 191-pair borderline pool, pass-instance rates rise from 4.5% to 26.2% and from 6.5% to 16.1%, respectively. This is not a uniform relaxation: two-agent screening rejects applications one-call advances, changing decisions in both directions. At similar pass volumes, the procedures advance different applications, and no one-call threshold recovers applications consistently selected by two-agent screening. Among discovery-selected cases re-executed in fresh runs, two-agent-only selections recur less often than shared selections, clearly under GPT-5.5 and less certainly under Opus 4.7, while a separate one-call follow-up shows no comparable decline. As hiring becomes agent-mediated on both sides, the screening procedure, not only the model behind it, shapes who reaches human review and how reliably that access recurs.","authors":["Jian Gao","Hang Jiang"],"categories":["cs.AI","cs.CL"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-09-18","first_seen":"2026-09-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.19530","pdf_url":"https://arxiv.org/pdf/2609.19530","source_feed":"cs.CL","score":6,"bucket":"other","rubric_hits":["D3"],"tags":["LLM agent","招聘筛选","社会模拟"],"reason":"用两个LLM agent模拟招聘双方，但无真实人类数据对照，属社会过程模拟的边…","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:21","error":null,"has_summary":false,"summary":null},{"id":"2609.19789","version":1,"title":"Contagion on the Trading Floor: How Adversarial Signals Spread in Multi-Agent Trading Systems","zh_title":"交易大厅的传染：多智能体交易系统中对抗性信号的传播","abstract":"Multi-agent trading systems built on large language models (LLMs) are beginning to appear in quantitative finance, yet their robustness to adversarial inputs is largely unknown. We study the vulnerability of LLM trading stacks to black-box, input-only attacks that enter solely via admissible social-media feeds. We introduce the Generic Multi-Agent Trading System (GMATS), a framework that captures modern multiagent trading architectures and instantiate a class of black-box poisoning attackers that treat an LLM as a post generator and inject budget-constrained, plausibly benign social-media content into the analyst's evidence stream. We define contagion metrics that trace how adversarial content propagates through the stack, including belief-shift scores at analyst and coordinator layers and attack-clean deltas on standard backtest metrics. Experiments on a safe offline benchmark with historical market and social data show that even simple input-only attackers can materially degrade risk-return profiles, sharply reducing Sharpe ratios. At the same time, we find that suitably designed multi-agent topologies and coordinator prompts can dampen adversarial shocks and improve average robustness under identical poisoning budgets.","authors":["Qi Rong Sua","Junhao Dong","Nguyen Duc Thai","Yuqing Wen","Cheston Tan","Yew-Soon Ong"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-18","first_seen":"2026-09-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.19789","pdf_url":"https://arxiv.org/pdf/2609.19789","source_feed":"cs.AI","score":6,"bucket":"other","rubric_hits":["D3"],"tags":["多智能体系统","对抗攻击","金融模拟"],"reason":"多智能体交易系统模拟市场过程，但无真实人类行为对照，属社会模拟边界情形。","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:33","error":null,"has_summary":false,"summary":null},{"id":"2609.06025","version":2,"title":"Factors Influencing the Emergence of Dependency Length Minimization in Neural Agent Simulations","zh_title":"神经智能体模拟中依存距离最小化涌现的影响因素","abstract":"Given various grammatical options, language users prefer the word order choice that reduces the overall length of syntactic dependencies, a principle known as dependency length minimization (DLM). The origins of this preference remain an open question, particularly whether it originates from constraints on efficient information processing. Computational simulations provide a powerful approach to identifying the factors influencing the emergence of linguistic phenomena. However, previous simulations of DLM have not examined realistic interaction contexts and have produced mixed results. The present study investigates the emergence of DLM in artificial languages using a recently proposed language learning and communication framework based on recurrent neural networks (RNNs). In this framework, agents are trained to speak and interpret artificial languages and then use these languages to communicate. Using this framework, we study the impact of several factors related to processing limitations in a communicative setting, such as noise during listening, limited speaker capacity, and incremental sentence processing. Our results reveal a complex interplay among these factors in shaping word order preferences in neural agents. Specifically, in the full meaning space, agents regularize toward a single dominant word order, while in the half meaning space they show a short-before-long preference that only aligns with DLM in verb-initial languages. A consistent DLM preference emerges only when agents are subject to incremental processing pressure. These findings suggest that limitations in human cognitive processing may indeed play a role in shaping DLM. Our findings provide insights into the conditions under which neural models replicate human-like preferences and highlight the challenges of designing emergent communication models that capture human cognitive biases in language processing.","authors":["Yuqing Zhang","Tessa Verhoef","Gertjan van Noord","Arianna Bisazza"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-18","first_seen":"2026-09-09","revised_at":"2026-09-18","abs_url":"https://arxiv.org/abs/2609.06025","pdf_url":"https://arxiv.org/pdf/2609.06025","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["语言演化模拟","神经智能体","认知偏差"],"reason":"用神经agent模拟语言演化，但无真实人类数据对照，属社会模拟边界情形。","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:51","error":null,"has_summary":false,"summary":null},{"id":"2609.20484","version":1,"title":"Edustories: A Collection of Real-world Case Studies from Classroom Practices","zh_title":"Edustories：来自课堂实践的真实案例研究集合","abstract":"Despite the widely recognized potential of AI in education, most prior work has focused on individualized student assistance. In contrast, the majority of educational practice worldwide still takes place in collective classroom settings. To enable researchers to study AI assistance in collective teaching, we introduce Edustories, a dataset of 1,492 teacher-written case studies describing real elementary and high-school classroom situations involving challenging student behavior, pedagogical interventions, and their outcomes. Among many other applications, Edustories enables evaluating LLMs' ability to predict the success of teacher interventions, crucial for providing practicing teachers with useful feedback. Comparing the latest models from four language-model families against expert assessments, we find that current models fall short of human expertise in predicting classroom outcomes; the strongest models reach 58% accuracy compared to 64% of human experts. This gap highlights both the limitations and the emerging potential of AI as assistants for practicing teachers.","authors":["Michal \\v{S}tef\\'anik","Jan Nehyba","Jirina Karasova","Martin Fico","Lucie \\v{S}karkov\\'a","Mark\\'eta Ko\\v{s}atkov\\'a","David Kosatka"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-18","first_seen":"2026-09-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.20484","pdf_url":"https://arxiv.org/pdf/2609.20484","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["教育AI","LLM评估","数据集"],"reason":"用LLM预测教师干预结果，替代人类专家评估，属于标注替代而非仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:40","error":null,"has_summary":false,"summary":null},{"id":"2609.20565","version":1,"title":"Steering the Compass: Aligning Dynamic Psychological Counseling Conversations with Cognitive Behavioral Therapy Strategies","zh_title":"校准罗盘：将动态心理咨询对话与认知行为疗法策略对齐","abstract":"Recent advancements in large language models have revolutionized the field of psychological counseling, especially in the context of Cognitive Behavioral Therapy (CBT). While the success of CBT relies heavily on dynamic decision-making informed by the client's real-time mental state, this aspect has often been overlooked in current research, limiting both flexibility and therapeutic outcomes. In this paper, we introduce StratCBT, a dataset specifically designed for psychological counseling conversations with CBT Strategies, consisting of 9,688 sessions and around 256K utterances, with each counselor's response aligned with one of eight distinct strategies. The creation of StratCBT involves modeling clients based on their negative thoughts and generating high-quality counseling conversations through self-chat, incorporating realistic sessions as guidance, thereby significantly surpassing existing datasets in both general counseling and CBT-specific skills. We conduct extensive experiments to demonstrate the effectiveness of strategy-aligned generation and evaluate its efficacy in delivering professional and effective counseling with LLM-simulated clients to reflect real-world scenarios. The dataset can be obtained from https://github.com/zimuwangnlp/StratCBT.","authors":["Zimu Wang","Yiwen Jiang","Xiangyu Zhao","Yaling Shen","Jiahe Liu","Stephanie Fong","Maxmartwell H Cheng","Guilherme C Oliveira","Anh Nguyen","Robert Desimone","Barnaby Nelson","Dominic Dwyer","Zongyuan Ge"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-18","first_seen":"2026-09-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.20565","pdf_url":"https://arxiv.org/pdf/2609.20565","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["LLM模拟","心理咨询","CBT"],"reason":"用LLM模拟来访者进行心理咨询对话，但无真实人类数据对照，属于社会模拟边界情形。","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:42","error":null,"has_summary":false,"summary":null},{"id":"2609.20059","version":1,"title":"AI Should Facilitate Democratic Deliberation at Scale","zh_title":"AI应促进大规模民主协商","abstract":"AI systems can strengthen democracy by supporting deliberation at scale by addressing cognitive, social, platform-design, and market-driven frictions, while preserving human agency. Unlike proposals such as liquid democracy that restructure representation through vote delegation, in this position paper, we argue that AI-assisted deliberation offers a more promising path by lowering barriers to meaningful engagement without substituting machine judgment for human choice. Drawing on evidence from online deliberation platforms and experimental research, we identify four guiding principles: preserving agency and autonomy, encouraging mutual respect, promoting equality and inclusiveness, and augmenting rather than substituting active citizenship. We also address critical challenges, including alignment, sycophancy, training bias, and over-reliance on AI systems. We call on the machine learning community to develop deliberation-focused AI systems evaluated not on engagement metrics but on their capacity to facilitate informed, representative, and friction-robust discourse.","authors":["Jos\\'e Ram\\'on Enr\\'iquez","Jiaxin Pei","Alex Pentland"],"categories":["cs.HC","cs.AI","cs.CL","cs.CY"],"primary_category":"cs.HC","announce_type":"cross","date":"2026-09-18","first_seen":"2026-09-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.20059","pdf_url":"https://arxiv.org/pdf/2609.20059","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["AI辅助协商","民主参与","社会模拟"],"reason":"讨论AI辅助民主协商，涉及社会过程模拟但无LLM仿真人类被试及数据对照","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:35","error":null,"has_summary":false,"summary":null},{"id":"2609.20658","version":1,"title":"Ownership in AI-Assisted Everyday Tasks","zh_title":"AI辅助日常任务中的所有权感知","abstract":"When does work done with AI still feel like ours? As AI becomes woven into everyday tasks, we must examine what happens to our sense of ownership and contribution when a machine shares in producing what we make. We report an exploratory qualitative survey in which participants were asked to describe two recent, self-selected tasks completed with AI: one that felt like their own and one that did not. We find that felt ownership depends on the process of collaboration: people disown work when they merely approve AI's suggestions, but retain ownership when they lead, iterate, or rewrite. Ownership can also extend to settings where people own the vision for a project but not the execution; respondents reported high ownership on tasks they could not have completed without AI. Loss of personal voice and a lack of comprehension of the output both erode ownership. Finally, willingness to disclose AI use is often decoupled from actual pride or ownership, and instead shaped by community norms and fear of credit erasure. We propose several research directions as a result of these findings to promote AI development that supports people's sense of authorship over their own lives.","authors":["Megan Wei","Melanie Subbiah","Audrey Lee","Annya Dahmani","Dave Edwards","Helen Edwards","Ellie Pavlick"],"categories":["cs.AI","cs.HC"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-18","first_seen":"2026-09-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.20658","pdf_url":"https://arxiv.org/pdf/2609.20658","source_feed":"cs.AI","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["人机交互","所有权感知","定性调查"],"reason":"研究人类对AI辅助任务的所有权感知，非LLM仿真人类被试，但涉及人类行为测量，…","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:44","error":null,"has_summary":false,"summary":null},{"id":"2609.19182","version":1,"title":"What Do We Expect from LLMs? Mapping the Design of LLM Benchmarks","zh_title":"我们对LLM有何期望？绘制LLM基准测试的设计图景","abstract":"Benchmarks are central to how progress in large language models (LLMs) is assessed and communicated. Yet model rankings alone reveal little about how evaluation requirements themselves are changing. The expanding variety of benchmarks offers another perspective: what researchers expect LLMs to do, and what they count as successful performance. We systematically map 14,767 papers introducing or updating evaluation resources from arXiv submissions between January 2022 and August 2026. Using staged screening and automated full-text coding, we examine changes in target systems and domains, evaluation materials and conditions, and scoring mechanisms. The collection shows growing emphasis on action, interaction, and professional applications, while established and newer design elements frequently coexist. Model participation also develops unevenly: LLM-based scoring grows within both agent and non-agent groups, whereas model-generated materials show no comparable sustained increase in recent cohorts. These findings illuminate how public research translates capability expectations into concrete tests and criteria for success. As AI participates in constructing tests, performing tasks, and judging responses, they also raise a question: does expanding evaluation provide more independent evidence, or risk reproducing the preferences and blind spots of its participating models?","authors":["Chao Wang (Independent Researcher)"],"categories":["cs.AI","cs.CL"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-09-18","first_seen":"2026-09-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.19182","pdf_url":"https://arxiv.org/pdf/2609.19182","source_feed":"cs.CL","score":4,"bucket":"other","rubric_hits":["C4"],"tags":["LLM基准测试","评估设计","元研究"],"reason":"论文研究LLM基准测试的设计与演变，属于纯NLP能力评测，不以人类行为为参照系…","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:30","error":null,"has_summary":false,"summary":null},{"id":"2609.19705","version":1,"title":"SoK: Trading Agents or Market Crashers? Dissecting Robustness and Security Failures in Academic Financial LLM Trading Schemes","zh_title":"SoK：交易代理还是市场崩溃者？剖析学术金融LLM交易方案的鲁棒性与安全失败","abstract":"Autonomous large language model (LLM) agents are moving rapidly into high-stakes domains, yet existing agentic-AI security studies remain largely domain-agnostic and overlook the distinctive, high-consequence attack surface such settings create. We examine this gap through financial trading agents, a representative case of high-stakes agentic security, where a single compromised agent has direct execution authority over real capital in an adversarial, reflexive market. To this end, we present FARSIGHT (Financial Agent Robustness and Security Investigation and Global Holistic Testing), a framework that performs scheme-level evaluation of financial LLM agents on two axes: robustness under market turbulence (including flash-crash-like scenarios), and security against three attack types: attacks on information sources, attacks on agents, and agent-as-attacker behaviors. Applying FARSIGHT to 15 representative academic schemes, we find that most overlook robustness and realistic adversarial threats: 80% fail at least one core robustness metric and 100% exhibit security vulnerabilities. These two failure modes are inseparable: a small misjudgment can cascade into a market-wide crash on its own, while an adversary can deliberately trigger the same collapse at minimal cost.","authors":["Mengxiao Wang","Nitesh Saxena"],"categories":["cs.CR","cs.AI","cs.MA"],"primary_category":"cs.CR","announce_type":"cross","date":"2026-09-18","first_seen":"2026-09-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.19705","pdf_url":"https://arxiv.org/pdf/2609.19705","source_feed":"cs.AI","score":3,"bucket":"other","rubric_hits":["C1"],"tags":["LLM交易代理","安全性","多智能体系统"],"reason":"研究金融LLM交易agent的鲁棒性与安全性，属多智能体系统安全，不涉及人类行…","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:33","error":null,"has_summary":false,"summary":null},{"id":"2609.20637","version":1,"title":"Stereotypically Yours: Portrayal and Perception of Race-Coded AI Companions","zh_title":"刻板印象中的你：种族编码AI伴侣的呈现与感知","abstract":"AI companions can purportedly adopt racial personas, raising questions about how they represent identity and how users interpret these portrayals. We combined an algorithmic audit of race-coded AI personas with interviews with 12 companion users who interacted with a probe. Our audit revealed systematic differences, such as Asian-coded male personas receiving higher submissiveness scores than White counterparts, and Black, Hispanic, and Indigenous male personas receiving higher aggression scores than their White counterparts in open-weight models. Interviews revealed that participants envisioned AI companions as offering cultural familiarity and outside perspectives, but differed in which portrayals they considered meaningful or stereotypical. Some rejected overt racial signaling while still expecting culturally distinctive responses. Triangulating these findings with theory, we highlight how social norms and cultural expectations complicate efforts to support meaningful racial representation without reproducing stereotypes. We discuss how companion personalization should be evaluated beyond user satisfaction to account for broader representational harms.","authors":["Wang Claire","Jiayue Melissa Shi","Agam Goyal","Grace Sletten","Renwen Zhang","Eshwar Chandrasekharan","Koustuv Saha"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-18","first_seen":"2026-09-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.20637","pdf_url":"https://arxiv.org/pdf/2609.20637","source_feed":"cs.HC","score":3,"bucket":"other","rubric_hits":["C3"],"tags":["AI伴侣","种族刻板印象","人机交互"],"reason":"研究AI伴侣的种族角色扮演与用户感知，属角色扮演聊天，无实验或测量目的，不涉及…","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:43","error":null,"has_summary":false,"summary":null},{"id":"2608.15339","version":2,"title":"Learning Sequential Mobility Choice: A Review of Route and Activity Choice through Inverse Reinforcement Learning and Imitation Learning","zh_title":"学习顺序出行选择：基于逆强化学习与模仿学习的路径与活动选择综述","abstract":"Route and activity choice are distinct transportation problems that both require models of feasible decisions unfolding over networks and time. This critical integrative review connects transportation choice modeling with inverse reinforcement learning (IRL) and imitation learning (IL), while distinguishing evidence from transportation applications, transferable methods from other fields, and emerging proposals. We develop a four-layer sequential mobility choice framework comprising the environment, behavioral objective, stochastic choice mechanism, and observation process. Under stated assumptions, recursive logit, logit dynamic discrete choice, and maximum-entropy IRL use the same soft Bellman recursion linking future opportunities to current choice probabilities. Expected state-action visitation also satisfies conservation equations analogous to network flows. These mathematical connections do not make utility, reward, policy, occupancy, constraints, and observation error behaviorally interchangeable. Transportation evidence is strongest for network-scale planning, context-dependent reward learning, inference from incomplete trajectories, and activity-schedule generation, but remains limited for actual interventions and transfer across networks. We therefore propose a behaviorally disciplined hybrid architecture that keeps feasible actions, interpretable trade-offs, observation processes, and system feedback explicit while using machine learning for scalable computation, contextual representation, heterogeneity, and data integration.","authors":["Hung Tran","Viet Bui","Tien Mai"],"categories":["econ.EM"],"primary_category":"econ.EM","announce_type":"replace","date":"2026-09-18","first_seen":"2026-08-18","revised_at":"2026-09-18","abs_url":"https://arxiv.org/abs/2608.15339","pdf_url":"https://arxiv.org/pdf/2608.15339","source_feed":"econ.EM","score":2,"bucket":"other","rubric_hits":["C2"],"tags":["交通选择建模","逆强化学习","模仿学习"],"reason":"论文综述交通选择建模与IRL/IL，虽提及LLM但非用于人类仿真实验，无人类行…","model":"deepseek-v4-pro","scored_at":"2026-08-18T13:04:41","error":null,"has_summary":false,"summary":null},{"id":"2608.26171","version":2,"title":"Mitigating Fabrication in Multi-Stage LLM Pipelines for Hiring: An Empirical Evaluation of Prompt Guardrails and Human-in-the-Loop Checkpoints","zh_title":"缓解多阶段LLM招聘流程中的虚构：提示护栏与人机检查点的实证评估","abstract":"Multi-stage LLM hiring pipelines (resume improvement, interview question generation, answer feedback) can fabricate credentials, inflate qualifiers, and invent experience. We evaluate two mitigations, prompt guardrails and human-in-the-loop (HITL) checkpoints, against a fully automated baseline. In a controlled experiment (10 synthetic resumes x 2 job descriptions x 3 repetitions x 3 conditions; 180 runs), the baseline (C1) produced at least one unsupported claim in 96.7% of outputs (mean 6.80 findings/output). Prompt guardrails (C2) reduced finding density by 86% (6.80 to 0.92/output), but 50.0% of outputs still contained a fabrication, showing prompt-level mitigation alone is insufficient. A human checkpoint after resume improvement (C3) eliminated all identity fabrications, reduced finding density by 59% (6.88 to 2.82/output), reduced item-level fabrication from 96.7% to 75.0% (p=.022), and cut capture of JD-embedded trap requirements from 47% to 2% (vs. 5% under the guardrail). An exploratory analysis of multi-specialty resumes shows contamination rising monotonically with domain distance between specialties, suggesting career changers are especially exposed. The reviewer in this study caught all flagrant fabrications, but subtle qualifier drops and plausible new claims survived review roughly half the time (54.5% removal). Neither mitigation degraded the deliverable: claim retention exceeded 99% under both. The interventions are complementary: the guardrail eliminates unprompted additions and qualifier inflation cheaply, while the checkpoint gives near-categorical guarantees against the most severe failures, invented identities and JD-baited claims. These results support a layered architecture combining guardrails with a human checkpoint. A supplementary run with a newer-generation model (90.0% baseline fabrication rate) suggests the problem is not resolved by model progress alone.","authors":["Hiroko Takano"],"categories":["cs.CY","cs.CL","cs.HC"],"primary_category":"cs.CY","announce_type":"replace-cross","date":"2026-09-18","first_seen":"2026-08-28","revised_at":"2026-09-18","abs_url":"https://arxiv.org/abs/2608.26171","pdf_url":"https://arxiv.org/pdf/2608.26171","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["LLM幻觉","招聘流程","人机协作"],"reason":"研究LLM招聘流程中的幻觉缓解，属多智能体系统可靠性，非人类行为仿真","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:50","error":null,"has_summary":false,"summary":null},{"id":"2609.06027","version":2,"title":"Evaluating Deep-Search Agents under Hierarchical Web Evidence Poisoning","zh_title":"分层网络证据投毒下深度搜索智能体的评估","abstract":"Search-augmented LLM agents are increasingly used for consumer decisions, making them vulnerable to Generative Engine Optimization (GEO) poisoning. Existing benchmarks largely measure whether manipulated content is retrieved or endorsed, but do not track whether an agent verifies suspicious evidence, revises adopted claims, or recovers before producing its final recommendation. We introduce HAE-GEO, a benchmark that tracks the full trajectory from exposure to recovery under progressively more persuasive Web poisoning. Agents interact via a multi-turn Search-Scrape interface across three attack levels (L1 direct assertion, L2 contextual camouflage, and L3 apparent corroboration), supported by a controlled corpus of 72,039 clean pages and 770 poisoned pages per level spanning 8 product categories and 154 brands. Evaluation combines deterministic behavioral measures with six semantic rubric dimensions. Evaluating 10 agents, we find three recurring patterns: evidence recognition degrades under the corroboration trap; agentic search improves final resistance without improving evidence recognition or utility; and defense prompting increases verification, yet rarely converts verification into recovery.","authors":["Zhongan Bi","Qiwen Wang","Jianrong Jiang","Jigang Ding","Wenwen Xiong","Changhua Meng","Xuanang Gao","Kepeng Lin","Changjiang Jiang","Yiang Chen","Huan Yao","Wei Wang","Zhenyu Ma","Wenhui Dong"],"categories":["cs.CR","cs.AI","cs.IR"],"primary_category":"cs.CR","announce_type":"replace-cross","date":"2026-09-18","first_seen":"2026-09-09","revised_at":"2026-09-18","abs_url":"https://arxiv.org/abs/2609.06027","pdf_url":"https://arxiv.org/pdf/2609.06027","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["搜索智能体","对抗鲁棒性","基准测试"],"reason":"研究搜索智能体在网页投毒下的行为，属多智能体系统安全评估，不涉及人类行为仿真或…","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:28","error":null,"has_summary":false,"summary":null},{"id":"2609.20541","version":1,"title":"An Analysis of Training-Free Self-Reported Confidence in Language Models","zh_title":"对语言模型中免训练自我报告置信度的分析","abstract":"Large language models can report a numerical confidence together with generated content, but it is unclear whether this report is more than calibrated rhetoric. We analyze three training-free signals: confidence verbalized with the answer, post-hoc $P(\\mathrm{True})$, and agreement with three additional generations on the same 100 TriviaQA questions for two model families. Direct verbalization is a surprisingly strong baseline: after auditing benchmark errors, it reaches AUROC 0.956 and 0.937 for correctness prediction. Three-sample agreement is substantially weaker (0.765 and 0.790), and a fixed interpolation with verbalized confidence has no statistically reliable benefit. Four of nine errors from one model and two of eight from the other receive unanimous sample support, showing that self-consistency can amplify shared misconceptions. Re-eliciting confidence for the same fixed answers with equivalent prompts changes scores by 0.043 to 0.084 on average and flips 4\\% to 9\\% of decisions at a 0.8 threshold. An exploratory audit of 100 confidence-tagged biography claims further finds only a modest confidence gap between supported and contradicted claims. These results argue that useful self-reports remain sensitive to elicitation, correlated errors, and benchmark noise.","authors":["Lukas Meyer","Sofia Rossi","Wei Chen","Thomas Laurent","Yiming Li"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-18","first_seen":"2026-09-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.20541","pdf_url":"https://arxiv.org/pdf/2609.20541","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["置信度校准","模型评测","自我报告"],"reason":"研究LLM自我报告置信度的校准，属于模型能力评测，不以人类行为为参照系。","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:41","error":null,"has_summary":false,"summary":null},{"id":"2609.20712","version":1,"title":"Summarization Bias: The Directional Collapse of Objective Projection into Told-Mode Labels in Large Language Models --- A Conceptual Framework and Registered Test Protocol","zh_title":"总结偏差：大语言模型中客观投射向告知模式标签的方向性坍缩——概念框架与注册测试协议","abstract":"This paper introduces and operationalizes summarization bias: a proposed systematic tendency of large language models (LLMs) to represent narrative meaning as an abstract summary label rather than as the reconstructable inferential structure that produces it. Within the Bulut Doctrine, narrative effect is theorized along a told-shown axis: in told mode, emotional and informational content is declared explicitly and requires little reader reconstruction; in shown mode, that content is suppressed at the surface and must be reconstructed from physical cues and indirection (Objective Projection). Shown mode is the higher-load condition the doctrine is designed to measure. The claim is that LLMs fail along this axis in a specific direction. Summarization bias is hypothesized to operate in two regimes: (i) a generative regime, in which a model asked to render an emotion through Objective Projection defaults to declaring it instead; and (ii) an evaluative regime, in which a model judging narrative quality rewards told-mode explicitness and under-detects shown-mode suppression. The evaluative regime is the more consequential, since LLMs increasingly serve as judges and reward models, and a directional bias toward told mode would impose a selection pressure degrading prose toward flat declaration. This report does not claim the bias is validated. It defines the construct, situates it against LLM-as-judge biases, rereads a completed independent reliability study as directional evidence consistent with it, and pre-registers a two-regime test with decision rules under which the construct would be abandoned.","authors":["Levent Bulut"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-18","first_seen":"2026-09-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.20712","pdf_url":"https://arxiv.org/pdf/2609.20712","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["LLM叙事偏差","文学分析","LLM-as-judge"],"reason":"研究LLM叙事中的总结偏差，属文学分析，非人类仿真实验，无人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:45","error":null,"has_summary":false,"summary":null},{"id":"2609.20779","version":1,"title":"Harm Laundering in GPT Models: Evidence That Gender Discrimination Is Transformed Rather Than Reduced Across Safety-Trained Generations","zh_title":"GPT模型中的危害洗白：证据表明性别歧视在安全训练世代间被转化而非减少","abstract":"Safety evaluations for large language models rely on surface-form classifiers that report declining harm scores across model generations. We provide evidence that this methodology is systematically incomplete: explicit discriminatory content is transformed rather than removed. We call this \\emph{harm laundering}. Analysing 450,000 gender-directed completions across 15 models spanning GPT-2 through to GPT-5 (OpenAI GPT lineage; three demographic conditions), we show that sexual violence clusters prevalent in GPT-2 women-directed output disappear by GPT-4, while men-directed completions gain positive representational territory (caregiving, emotional range, ally identity) that women-directed completions do not. The pattern is most visible at GPT-5: Topic~5 (1,997~documents) frames breast cancer as a men's rights debate, while zero equivalent clusters appear in women-directed output. Three independent classifiers score this content as non-toxic. Sentiment scores invert at GPT-4: early models demean women; later models over-correct. Topic diversity in women-directed completions falls 36\\% relative to men at the GPT-4 alignment boundary (W/M~$= 0.58$, from $0.91$ at GPT-2). REGARD representational harm disparity correlates with release date ($\\rho = +0.55$, $p = .034$) while Detoxify does not ($\\rho = -0.23$, $p = .42$): toxicity scores fall as representational harm grows. We formalise harm laundering as a three-criteria test and provide a three-stage detection protocol applicable to any generative model. Within the OpenAI GPT lineage, toxicity score reduction is not a sufficient proxy for harm reduction.","authors":["Sarah Wyer","Sue Black","Noura Al Moubayed"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-18","first_seen":"2026-09-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.20779","pdf_url":"https://arxiv.org/pdf/2609.20779","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["模型安全","性别偏见","评测方法"],"reason":"评估LLM输出中的性别歧视，属于模型安全评测，非人类仿真实验。","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:46","error":null,"has_summary":false,"summary":null},{"id":"2609.20449","version":1,"title":"The Organization of Inference: Information, Resource Constraints, and AI Production","zh_title":"推理的组织：信息、资源约束与AI生产","abstract":"The economic value of inference depends on how capacity and task information are distributed across stages of AI production. We study these organizational margins using controlled workflow experiments on externally verified software-engineering tasks. In two matched resource panels, direct execution records the same success rate of 59.6 percent at logical-token ceilings of 12,000 and 24,000, while success under information-constrained planning rises from 36.2 to 51.2 percent. The planning disadvantage narrows by 15.0 percentage points (95 percent task-cluster bootstrap interval: 4.2 to 25.8). A strict read-only planning campaign varies whether the planner sees the task issue. At 12,000 tokens, issue access raises success by about 16 percentage points over issue-hidden planning. Compared with direct execution, task-informed planning is about 10 points lower at 12,000 tokens; at 24,000 tokens, it shows a 29.6-point advantage. In the resource panels, direct execution uses substantially less than either ceiling, while the planning workflow's binding rate falls from 46.2 to 0.8 percent and downstream execution accounts for 89.9 percent of the increase in total use. Scale determines the capacity available to a system; workflow and information structure shape the productive value","authors":["Yukun Zhang","Kemu Xu","Yishen Chen"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-18","first_seen":"2026-09-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.20449","pdf_url":"https://arxiv.org/pdf/2609.20449","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["AI工作流","资源约束","软件工程"],"reason":"研究AI工作流中推理的组织，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:27","error":null,"has_summary":false,"summary":null},{"id":"2609.19354","version":1,"title":"Can Vision-Language Models Judge Olympic Diving? From Reasoning to Scores in Zero-Shot Action Quality Assessment","zh_title":"视觉语言模型能否评判奥运跳水？从推理到零样本动作质量评估的分数","abstract":"Automated action quality assessment (AQA) in Olympic sports remains a challenging task due to the complexity of human motion and the subjectivity inherent in expert judging. This work evaluates the capability of open-source Vision-Language Models (VLMs) to perform zero-shot action quality assessment on Olympic diving videos using the AQA-7 benchmark dataset. In this regard, a regression-based framework is pro-posed to leverage both the semantic reasoning and phase-level sub-scores generated by the VLMs, combining TF-IDF vectorization, dimensionality reduction, and ensemble learning to predict final competition scores. Experimental results show that standalone VLMs achieve moderate Spearman correlations below 0.32, while the proposed ensemble regression framework substantially improves performance in the reported evaluation, reaching a Spearman correlation of 0.67 with a four-model configuration. Textual reasoning features con-sistently outperformed raw numerical sub-scores, highlighting the richness of VLM-generated explanations for action quality analysis. These findings suggest that VLMs hold strong potential as assistive tools for explainable and semi-automated sports performance evaluation. The code is publicly available on GitHub https://github.com/hvelesaca/olympic diving judge vlm","authors":["Henry O. Velesaca","David Freire-Obregon","Luigi Miranda","Abel Reyes-Angulo"],"categories":["cs.CV","cs.AI","cs.LG"],"primary_category":"cs.CV","announce_type":"cross","date":"2026-09-18","first_seen":"2026-09-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.19354","pdf_url":"https://arxiv.org/pdf/2609.19354","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C2"],"tags":["动作质量评估","视觉语言模型","体育视频分析"],"reason":"评估VLM对跳水动作质量打分，属于体育视频分析，不涉及用LLM仿真人类被试或社…","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:30","error":null,"has_summary":false,"summary":null},{"id":"2609.19617","version":1,"title":"DataCanvas-EDU: An Agentic Framework for Instructor-Guided Synthetic Data Generation in Business Analytics Education","zh_title":"DataCanvas-EDU：面向商业分析教育的教师引导合成数据生成智能体框架","abstract":"Business analytics education requires diverse datasets to support different learning objectives, student backgrounds, and analytical tasks. Real-world data can be difficult to obtain and offer limited flexibility for adapting a case to a particular course. Even when suitable data are available, instructors must investigate the patterns, verify the results, and prepare assignments and reference solutions, requiring substantial time and effort. The use of large language models (LLMs) introduces an additional concern about training data contamination. Widely used public datasets often have extensive tutorials and worked analyses that models may have encountered during training. Students may therefore receive explanations drawn from existing analyses without practicing how to investigate unfamiliar data in collaboration with AI. This paper presents DataCanvas-EDU, an agentic framework for instructor-guided synthetic data generation in business analytics education. Instructors specify teaching goals and intended patterns through conversation, while an AI agent writes generation code, checks the resulting data, and prepares assignments, reference analyses, and rubrics. Four phases, Plan, Create, Verify / Test Analysis, and Evaluate, organize the process and support instructor review and revision. The framework is intended to simplify case preparation while creating opportunities for students to investigate newly designed patterns with AI. We illustrate the approach with WindowDash, a food delivery case containing 15,000 orders and nine designed patterns. DataCanvas-EDU is packaged as a reusable AI Agent Skill for compatible agent environments, with the package and installation instructions available at https://github.com/BANG23333/datacanvas-edu","authors":["Bang An","Maria Hamdani","Joseph Fox"],"categories":["cs.HC","cs.AI"],"primary_category":"cs.HC","announce_type":"cross","date":"2026-09-18","first_seen":"2026-09-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.19617","pdf_url":"https://arxiv.org/pdf/2609.19617","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["合成数据生成","教育技术","多智能体框架"],"reason":"多智能体框架生成教学数据，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:32","error":null,"has_summary":false,"summary":null},{"id":"2609.20143","version":1,"title":"Designing Against Deskilling: Metacognitive Feedback Reduces Cognitive Offloading to LLM Assistants","zh_title":"设计防止技能退化：元认知反馈减少对LLM助手的认知卸载","abstract":"Cognitive offloading to AI can reduce opportunities to practice skills, creating risks of deskilling. However, it remains unclear how to prevent deskilling without restricting access to AI. Here, we design two interventions to reduce offloading decisions: (1) metacognitive feedback that makes the implications of offloading for users explicit, and (2) an effort-based reward that incentivizes less extensive LLM assistance. We test both in a preregistered online experiment ($N = 704$) with a 2$\\times$2 design and a no-AI control. The task was to practice fraction arithmetic with an LLM-based assistant that provided solutions only on explicit request, followed by an unaided test. Metacognitive feedback reduced answer offloading (OR $= 0.47$) and improved test performance (OR $= 1.51$). We found no evidence that the reward affected either outcome. Our results identify metacognitive feedback as a promising design choice to reduce cognitive offloading.","authors":["Sebastian Maier","Kai Schwabe","Manuel Schneider","Stefan Feuerriegel"],"categories":["cs.HC","cs.AI"],"primary_category":"cs.HC","announce_type":"cross","date":"2026-09-18","first_seen":"2026-09-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.20143","pdf_url":"https://arxiv.org/pdf/2609.20143","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["人机交互","认知卸载","实验研究"],"reason":"研究人类与LLM交互中的认知卸载，非用LLM仿真人类被试，无人类数据对照。","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:27","error":null,"has_summary":false,"summary":null},{"id":"2609.20311","version":1,"title":"Human and AI-generated texts between modal logic and statistics","zh_title":"模态逻辑与统计之间的人类与AI生成文本","abstract":"We read the geometry of semantic neighbourhood graphs as modal logic and give that reading a statistical form, in order to make precise the structural difference between human and machine-generated text. Texts are the worlds of a finite frame whose accessibility is the $k$-nearest-neighbour relation of a transformer embedding, and the symmetry, transitivity, Euclideanity and seriality frequencies of the two subcorpora are shown to be degrees of validation of the modal axioms $\\mathsf{B}$, $\\mathsf{4}$, $\\mathsf{5}$, $\\mathsf{D}$. Each degree is at once the proportion of instances of a rule that the subframe licenses in Negri's labelled calculus $\\mathsf{G3.K}$ and a plug-in estimate of a population probability. A prompt-balanced comparison finds consistently higher artificial degrees for $\\mathsf{4}$ and $\\mathsf{5}$. We add a degree of groundedness and of situatedness, and recast the licensing reading in Cuconato's one-sided sequent-style tableaux, where each degree becomes a rate of set membership.","authors":["Simone Cuconato","Donato Ferrari"],"categories":["math.LO","cs.AI","math.ST","stat.TH"],"primary_category":"math.LO","announce_type":"cross","date":"2026-09-18","first_seen":"2026-09-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.20311","pdf_url":"https://arxiv.org/pdf/2609.20311","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["文本分析","模态逻辑","AI生成文本"],"reason":"论文比较人类与AI生成文本的结构差异，属于NLP评测，不以LLM仿真人类被试为…","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:40","error":null,"has_summary":false,"summary":null},{"id":"2609.19318","version":1,"title":"\"I Know Where to Look,\" But Does the LLM? Charting the Gaps Between Clinical Expert Needs and Unstructured Data Abstraction Tools","zh_title":"“我知道往哪看”，但LLM知道吗？绘制临床专家需求与非结构化数据抽取工具之间的差距","abstract":"Clinical data abstraction, the process of distilling structured information from patient records, plays a key role in advancing knowledge about diseases such as cancer. Information extraction (IE) with large language models (LLMs) could accelerate this process, but it is unclear whether current frameworks effectively support clinical researchers without AI expertise. To address this, we co-designed an interactive LLM-based abstraction system called Libretto with seven cancer research teams, then evaluated the system's ability to help them answer real-world research questions. We found that while clinicians knew where and how to annotate complex concepts in patient notes, in twelve of fourteen tasks they faced barriers to replicating those intuitions with LLMs. Contextual note reliability judgments, difficulties in steering vibe-coded prompts, and inflexible evaluation strategies necessitated fundamental changes to the IE workflow. Our results highlight open problems for HCI research to bridge the gaps between AI data work tools and clinical users' needs.","authors":["Venkatesh Sivaraman","Rigney Turnham","George Bonano","Nevin Aresh","Renumathy Dhanasekaran","Margaret Guo","Sindhu Kubendran","Olivia Lin","Jonathan D Louie","Kristan Olazo","Jeanne Shen","Harish Vasudevan","Jeanette Wong","Emily Alsentzer","Jason A Fries","Anobel Odisho","John Gordan","Jean Feng","Julian C Hong"],"categories":["cs.HC","q-bio.OT"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-18","first_seen":"2026-09-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.19318","pdf_url":"https://arxiv.org/pdf/2609.19318","source_feed":"cs.HC","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["临床数据抽取","人机交互","信息抽取"],"reason":"研究LLM用于临床数据抽取，属NLP信息抽取工具，非人类仿真实验，无人类行为对…","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:30","error":null,"has_summary":false,"summary":null},{"id":"2609.20720","version":1,"title":"What Parents Can See: Divergent Accounts of Youth AI Companion Use in Parenting and Teenager Subreddits","zh_title":"父母所见：育儿与青少年子版块中关于青少年AI伴侣使用的分歧叙述","abstract":"Youth increasingly use AI companions, and parents are the primary mediators of that use. How effective that mediation can be depends on whether parents are aware of how adolescents actually use these systems and what risks and benefits such use carries; nevertheless, prior work has only studied these demographic groups in isolation, and existing taxonomies attend almost entirely to risk. We analyze 1,628 Reddit posts about youth AI companion use from parenting and teenager communities (2023--2026); develop a codebook covering modes of use, risks, benefits, and parental mediation; and apply it at corpus scale with an LLM. The two communities yield divergent accounts. Teenagers most often discuss receipt of emotional support from AI companions (31% of teenager posts vs. 19% of parenting posts), whereas parents most often discuss teenage use of AI companions for romantic and sexual interaction (36% vs. 25%). Teenagers are not unaware of other risks, however; indeed, attachment and dependence is the risk they raise most (19%), close to the parental rate (16%). Teenagers also describe benefits that risk-centered taxonomies do not capture and parents rarely mention, most notably emotional support (27% vs. 5%). We argue these differences track what a given kind of use makes visible to someone outside the conversation. Chatting with a companion for hours every night leaves a trace beyond the chat itself; sexting with a character stands out when a parent reads the log; venting about a fight with a friend does neither, since it looks like any other conversation. The first surfaces as dependence, the second as sexual content, and the third as emotional support, which is the one parents most often miss. Parental guidance and system design should attend to use cases that reach parents by neither route, emotional support foremost among them.","authors":["Thomas Berkane","Anne Bischops","Anika Mellacheruvu","Maimuna Majumder"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-18","first_seen":"2026-09-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.20720","pdf_url":"https://arxiv.org/pdf/2609.20720","source_feed":"cs.HC","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["AI伴侣","Reddit分析","人机交互"],"reason":"研究AI伴侣使用现象，非用LLM仿真人类被试，无实验或测量目的","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:45","error":null,"has_summary":false,"summary":null},{"id":"2609.20198","version":1,"title":"Evaluating Financial Sentiment in the Age of AI","zh_title":"AI时代的金融情感评估","abstract":"Financial sentiment measures are widely used in empirical finance, but it remains unclear whether general-purpose large language models (LLMs) improve on existing finance-specific methods. This paper evaluates twelve sentiment models, including dictionary-based methods, finance-specific transformers, and open-source LLMs, using two criteria: linguistic validity and economic validity. We find that general-purpose LLMs achieve classification performance comparable to finance-specific transformer models without task-specific fine-tuning. However, higher classification accuracy does not translate into stronger economic relationships. Several models produce sentiment measures that are significantly associated with earnings surprises, but none is significantly associated with next-day stock returns. Model performance is strongest for announcements with large earnings beats or misses and substantially weaker for announcements with more moderate earnings surprises. These findings suggest that financial sentiment captures information about firms' economic performance but has limited ability to explain short-run market reactions","authors":["Arslan Bisharat","Oudom Hean"],"categories":["cs.LG"],"primary_category":"cs.LG","announce_type":"new","date":"2026-09-18","first_seen":"2026-09-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.20198","pdf_url":"https://arxiv.org/pdf/2609.20198","source_feed":"cs.LG","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["金融情感分析","LLM评估","NLP基准"],"reason":"评估LLM情感分类性能，非仿真人类被试，无行为对照","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:36","error":null,"has_summary":false,"summary":null},{"id":"2609.20250","version":1,"title":"How Far Can Sub-3B Open Language Models Go in Zero-Shot Essay Scoring on an 8 GB Consumer GPU?","zh_title":"在8GB消费级GPU上，30亿参数以下开源语言模型在零样本作文评分中能走多远？","abstract":"Zero-shot essay scoring with large language models is usually demonstrated with proprietary API models, yet the settings where automated scoring is most needed, such as public schools grading thousands of essays under strict privacy rules, are often those where sending student writing to a third-party API is unacceptable. We ask how much capability survives when the model must be a sub-3B open model running fully locally in FP16, with a controlled study of four instruction-tuned models from two families (Qwen2.5 at 0.5B/1.5B/3B, SmolLM2 at 1.7B) on all eight ASAP-AES prompts on a single 8 GB consumer GPU, with bootstrap confidence intervals, Holm-corrected paired tests, and deployment-realistic variants of the key design choices. Three findings emerge. (i) Rubric-decomposed prompting beats holistic prompting for every model under batch min-max aggregation (though Qwen2.5-3B drops significantly on one prompt), and under mean aggregation two unrelated families land within 0.01 at the 1.5-1.7B scale. (ii) Mapping trait scores into the prompt range is fragile to grader calibration: one model compresses traits into a narrow low band (2-4 on 0-10) and naive mean aggregation collapses, while the min-max normalization of Multi-Trait Specialization repairs it (macro QWK 0.204 to 0.388) and stays within 0.03 when its statistics are frozen on 30 held-out essays. (iii) Signed error falls with essay length in eleven of twelve configurations, opposite to the verbosity bias reported for large LLM judges; normalized rubric decomposition largely flattens this slope for well-calibrated models. We anchor results honestly: the best local configuration (0.388) remains far below both the human inter-rater ceiling (0.769) and a length-only baseline (0.523), so we position sub-3B local models strictly for formative, human-supervised feedback.","authors":["Nguyen Dung Son","Dang Quang Minh","Nguyen Huu Loi","Truong Viet Vu","Nguyen Thai Anh"],"categories":["cs.LG"],"primary_category":"cs.LG","announce_type":"new","date":"2026-09-18","first_seen":"2026-09-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.20250","pdf_url":"https://arxiv.org/pdf/2609.20250","source_feed":"cs.LG","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["自动作文评分","小模型","零样本"],"reason":"论文是LLM自动作文评分，属于NLP能力评测，不以人类行为仿真为目标，无人类被…","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:39","error":null,"has_summary":false,"summary":null},{"id":"2609.19864","version":1,"title":"Converging Naming Styles, Persistent Network Locality: GitHub in the LLM Era","zh_title":"命名风格趋同，网络局部性持续：LLM时代的GitHub","abstract":"Social conventions often emerge through repeated interactions within social networks, allowing shared practices to coexist with variation across groups. Large language models (LLMs) introduce a potentially different coordination structure: a small number of widely used models can expose socially distant users to similar patterns and suggestions. Whether broad convergence under such shared technological influences eliminates network-local variation remains unclear. We examine this question in software development, where identifier naming styles provide observable conventions and LLM-based tools have rapidly diffused. Using public GitHub repositories created between 2015 and September 2025 across six programming languages, we characterize naming styles with 27 features and examine their association with detected LLM-related commits, their diversity across repository creation cohorts, and their relationship to owner proximity in a large-scale collaboration network. Repositories with detected LLM-related commits tend to use longer identifiers and, in several languages, make greater use of naming patterns already prevalent within the language. We also observe lower naming-style diversity in recent creation cohorts, with marked declines appearing around 2023-2024 in several languages, although their timing and trajectories differ. At the same time, network locality persists: in five of the six languages, repositories whose owners are closer in the collaboration network remain more similar in naming style even among recent, more homogeneous cohorts. These findings show that aggregate convergence and network-local variation can coexist, highlighting the need to examine not only how much cultural variation remains, but also how that variation continues to be structured by human social relationships in the era of widely shared AI systems.","authors":["Yuto Tamura","Sho Tsugawa"],"categories":["cs.SI"],"primary_category":"cs.SI","announce_type":"new","date":"2026-09-18","first_seen":"2026-09-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.19864","pdf_url":"https://arxiv.org/pdf/2609.19864","source_feed":"cs.SI","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["LLM影响","代码风格","社会网络"],"reason":"研究LLM对代码命名风格的影响，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:23","error":null,"has_summary":false,"summary":null},{"id":"2609.11286","version":2,"title":"Generating a Consistent Enterprise: Synthesis and Reference-Free Evaluation of Multi-System Business Data","zh_title":"生成一致的企业：多系统业务数据的合成与无参考评估","abstract":"Synthetic relational data is normally produced by a model trained on a real dataset, and its quality is measured as the distance to that dataset. This paper describes a generator that has no real dataset at either end. Given an industry, a company size, a business model, a set of business applications, and a random seed, it produces a complete fictional enterprise: a workforce, a customer base, sales deals, support tickets, recorded calls, chat messages, and documents, all consistent with one another. One entity graph is projected into the native formats of 66 business products, so the same customer appears in the CRM, the support desk, and the call system under one identity. Because no real counterpart exists, realism is built in from cited reference statistics and verified by reference-free measurement: a five-axis scorecard of 28 statistical checks, an adversarial detector that hunts for the marks of synthetic generation, and a set of soundness checks that include a classifier test against an independently shuffled copy of the data. Because these instruments existed before the generator was tuned, progress is measured under a fixed yardstick: over 23 generated companies, mean realism climbed from 60.3 to 99.1, the weakest company from 41.1 to 94.9, and the detector, which initially flagged 55.2% of all records, now flags none. The scores hold on a seed never used during development. A second generator builds relational databases from a list of business questions. It forces qualifying rows for each answerable question, adds controlled near misses, and computes exact labels from the finished tables. The generator runs as a hosted service at https://console.era.eon.io. A company built there to a specification is served through its simulators over MCP and REST, and the simulators are also published as container images for offline use","authors":["Benjamin Gruenbaum","Doron Porat","Assaf Natanzon","Roy Zavida","Chen Dinachi","Or Itzahary","Omer Niv"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"replace","date":"2026-09-18","first_seen":"2026-09-12","revised_at":"2026-09-18","abs_url":"https://arxiv.org/abs/2609.11286","pdf_url":"https://arxiv.org/pdf/2609.11286","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C2"],"tags":["合成数据","企业数据生成","软件测试"],"reason":"生成企业合成数据用于软件测试，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-12T13:01:11","error":null,"has_summary":false,"summary":null},{"id":"2609.19989","version":1,"title":"Benchmarking LLM Compliance with China AI Generated Content Regulations","zh_title":"评估大语言模型对中国AI生成内容法规的合规性","abstract":"The widespread adoption of LLMs has led to escalating content compliance risks. Prior works have contributed to addressing these risks in the English context, downplaying the complexity of Chinese language content. This paper follows China's current AI-Generated content compliance requirements and provides evaluation results on 20 notable LLMs, offering insight into China's regulatory landscape. We design a novel framework to assess the compliance and refusal rates with 2303 questions spanning six distinct dimensions, including 203 self-constructed constitutional questions. The framework employs several judges to generate verdicts independently based on their hierarchical alignment memory. Our findings show that international models also exhibit high levels of compliance despite the use of standard Chinese questions, and the main differences may stem from dimensions closely related to ideological alignment. We establish a regulatory benchmark that enables the global AI community to evaluate both Chinese and non-Chinese LLMs under a unified set of legally grounded compliance requirements.","authors":["Chenrui Cui","Hongye Fang","Lisha Song","Weichao Chen","Yue Zhu","Gang Xu"],"categories":["cs.CL","cs.MA"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-18","first_seen":"2026-09-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.19989","pdf_url":"https://arxiv.org/pdf/2609.19989","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["合规性评测","内容监管","模型能力"],"reason":"评估LLM对中国AI内容法规的合规性，属于模型能力评测，不以人类行为为参照，不…","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:34","error":null,"has_summary":false,"summary":null},{"id":"2609.20207","version":1,"title":"Foundations of Stochastic Lexical Calculus: Semantic Descent and Random Dynamics on Probability Simplices","zh_title":"随机词汇演算基础：概率单纯形上的语义下降与随机动力学","abstract":"Large language models produce prompt-dependent probabilities over words, whereas scientific systems require uncertainty over meaningful states that can be updated as evidence arrives. We develop an observable framework for determining when language-derived probabilities support such a sequential state representation. Theoretically, we define typed measurable transformations of contextual language, construct a minimal closed representation, and give necessary and sufficient conditions for semantic updates to exist uniquely. We bound irreducible nonclosure and accumulated error, and under average contraction prove existence, uniqueness and stability of an external random recursion on a probability simplex. These results define a stochastic lexical calculus without attributing an internal calculus to the language model. Empirically, frozen experiments test the observable implications. Raw prompt-conditioned probabilities fail the prespecified invariance gate; after prompt-specific calibration, a common three-state representation passes the stability gates and covers 28 of 30 untouched eight-step paths, or 0.933 at nominal level 0.90. Accordingly, language probabilities support a stochastic state only conditionally on verified closure, stability and coverage within a declared operating domain.","authors":["Matthew F Dixon"],"categories":["cs.CL","math.CT","math.PR","stat.ML"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-18","first_seen":"2026-09-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.20207","pdf_url":"https://arxiv.org/pdf/2609.20207","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["语言模型概率","数学理论","语义表示"],"reason":"研究语言模型概率的数学性质，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:37","error":null,"has_summary":false,"summary":null},{"id":"2609.20584","version":1,"title":"SAFARI: An Industrial Benchmark for LLM-Assisted Hazard Analysis and Risk Assessment","zh_title":"SAFARI：面向LLM辅助危险分析与风险评估的工业基准","abstract":"Large language models (LLMs) are increasingly considered for safety-critical engineering, yet their reliability in regulated functional-safety workflows remains underexplored. We introduce SAFARI (Safety-Aware Functional Automotive Risk Inference), the first industrial benchmark for LLM-assisted automotive Hazard Analysis and Risk Assessment (HARA) under ISO 26262. It contains 3,000 de-identified industrial HARA cases and evaluates two coupled tasks: open-ended hazard analysis and standards-grounded risk assessment. To evaluate open-ended HARA artifacts, we propose the first reference-anchored LLM-as-a-judge protocol with high expert correlation. Experiments with nine frontier LLMs show that models often produce plausible hazard narratives but remain weak at ISO 26262 risk classification, with the best ASIL macro-F1 reaching only 0.261. Chain-of-Thought prompting provides limited benefit and often degrades categorical risk assessment. Error analysis further localizes major failures to scenario-critical context omissions during hazard generation and to controllability misjudgments during risk assessment, indicating where expert oversight should be concentrated. The dataset can be obtained from https://github.com/xixi47520-hash/HARA.","authors":["Chenxi Wu","Zimu Wang","Haiyang Zhang","Wei Wang","Zhijie Xu"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-18","first_seen":"2026-09-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.20584","pdf_url":"https://arxiv.org/pdf/2609.20584","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C2"],"tags":["LLM辅助安全工程","自动驾驶功能安全","风险评估基准"],"reason":"论文聚焦汽车功能安全中的危险分析与风险评估，属于自动驾驶安全工程，不涉及用LL…","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:42","error":null,"has_summary":false,"summary":null},{"id":"2609.20684","version":1,"title":"HerHealthEval: Evaluating Multilingual and Register-Sensitive Understanding of Women's Health Communication","zh_title":"HerHealthEval：评估女性健康沟通的多语言与语域敏感理解","abstract":"Large language models are increasingly used in healthcare communication, yet most evaluations emphasize response quality while assuming that the user's concern has been interpreted correctly. We introduce HerHealthEval, a controlled evaluation framework for multilingual understanding of women's-health communication. For each clinical case, HerHealthEval provides matched versions in English, French, and Modern Standard Arabic using six communicative forms: canonical, clinical, layperson, indirect or hedged, emotionally concerned, and deliberately under-specified. The first five express the same underlying concern and retain the same clinical information, whereas the under-specified form intentionally omits relevant details to test whether the model recognizes that clarification is needed. We evaluate a multilingual instruction model and QLoRA-adapted variants on concern classification, risk calibration, clarification behavior, parse compliance, and cross-form consistency. Results reveal that aggregate accuracy and consistency can conceal safety-relevant failures. A multilingual adaptation model reaches 0.994 under-triage in French and Arabic under language-asymmetric risk supervision. A controlled re-adaptation using source-derived, language-invariant risk labels reduces under-triage to 0.572 and 0.558, respectively. These findings show that robust multilingual healthcare evaluation requires explicit testing of register variation, uncertainty handling, and the provenance and invariance of adaptation labels.","authors":["Hassan Saeed Hassan Albattra","Mazen Mohammed Bahgat","Rahatara Ferdousi","Hana Essam Sayed Ahmed Amrya","Mariam Mousa"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-18","first_seen":"2026-09-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.20684","pdf_url":"https://arxiv.org/pdf/2609.20684","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["LLM评测","医疗健康","多语言"],"reason":"评估LLM对女性健康文本的理解，属NLP能力评测，不以人类行为仿真为目标","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:45","error":null,"has_summary":false,"summary":null},{"id":"2609.20808","version":1,"title":"Unifying Models of Intergroup Hostility in Online Discourse","zh_title":"统一网络话语中群体间敌意的模型","abstract":"Hostile rhetoric toward social groups can normalize exclusion and justify mistreatment, as well as contribute to rising polarization and political violence. Efforts to moderate hostile rhetoric in online speech draw on foundational theories in social and moral psychology, and political science. However, these theories were developed largely in parallel, often propose different and sometimes conflicting accounts of how hostility develops, and have rarely been tested against each other in real discourse. The result is a fragmented understanding of the rhetorical mechanisms of hostility, without a clear sense of how they appear, and relate to each other, in real-world discourse. Using 2.86 million posts from TikTok, Truth Social, and Twitter/X during the 2024 U.S. presidential election, we model the mechanisms of six foundational theories of intergroup hostility -- boundary construction, threat construction, scapegoating, negative evaluation, dehumanization, and action orientation -- within a common empirical framework to recover the broader organization of intergroup hostility rhetoric. Structurally, we find that boundary construction and threat construction anchor the system; temporally, we find that these mechanisms tend to follow a regular ordering: boundary construction, derogation, and action orientation tend to appear early; dehumanization and threat construction later; scapegoating latest. Mapping how these theoretical frameworks actually manifest in discourse bridges longstanding divisions across social science traditions and presents computational social science with a clearer empirical foundation for modeling intergroup hostility rhetoric beyond single-label detection.","authors":["Patrick Gerard","Julia Mendelsohn","Kristina Lerman"],"categories":["cs.CL","cs.SI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-18","first_seen":"2026-09-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.20808","pdf_url":"https://arxiv.org/pdf/2609.20808","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["计算社会科学","群体间敌意","社交媒体分析"],"reason":"论文分析真实社交媒体帖子，未使用LLM仿真人类被试，不涉及人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:46","error":null,"has_summary":false,"summary":null},{"id":"2609.20821","version":1,"title":"Embedding Models Measure in Peculiar Ways","zh_title":"嵌入模型以奇特方式度量","abstract":"Embedding spaces define notions of semantic similarity and distance. We study whether those embeddings reflect physical measurements of mass, distance, time and volume, which admit a unique, objective notion of semantic equivalence and distance. We find that physical measurement is only weakly modeled in the embedding space, and that instead quite peculiar measurement patterns can be observed. Further analysis indicates that embedding representations of physical measurements are strongly influenced by superficial string similarity, and recalibration of similarity does not substantially improve the alignment.","authors":["Juri Opitz","Andrianos Michail"],"categories":["cs.CL","cs.LG"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-18","first_seen":"2026-09-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.20821","pdf_url":"https://arxiv.org/pdf/2609.20821","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["嵌入空间","语义相似度","模型评估"],"reason":"研究嵌入空间对物理测量的表征，属NLP模型能力分析，不以人类行为为参照，与LL…","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:47","error":null,"has_summary":false,"summary":null},{"id":"2609.19244","version":1,"title":"Characterizing Web Search by Conversational LLM Agents: From Search Decisions and Strategies to Results and Responses","zh_title":"对话式LLM代理的网页搜索特征：从搜索决策与策略到结果与响应","abstract":"Conversational LLM agents increasingly rely on Web search, yet the end-to-end lifecycle of agentic search remains poorly understood. We present the first study of Web search across four major conversational platforms (ChatGPT, Claude, Grok, and DeepSeek), combining real-world user interactions (invivo) with controlled experiments using the same platform's models by their APIs (invitro). We investigate the quality of agentic decisions to invoke Web search, their strategies to formulate queries, the potential domain preferences in the search results they receive, and the choices they make when transforming search results into grounded responses. We find that Web-search decisions vary substantially across platforms and models, while more frequent Web-search invocation does not necessarily yield better response quality. We further show that conversational agents employ different complex querying strategies and that platform specific search engines return search results from their preferred domains. Finally, although responses are largely grounded in search results, some claims rely on uncited search results, raising concerns about attribution and reliability. Our findings have important implications for the design of future AI agents and Web search tools optimized for conversational retrieval.","authors":["Mahsa Amani","Seungeon Lee","Abhisek Dash","Asmaa El Fraihi","Yunah Jang","Elisabeth Kirsten","Qinyuan Wu","Krishna P. Gummadi","Manish Gupta","Abhilasha Ravichander","Muhammad Bilal Zafar","Soumi Das"],"categories":["cs.AI","cs.IR"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-18","first_seen":"2026-09-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.19244","pdf_url":"https://arxiv.org/pdf/2609.19244","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["LLM代理","网页搜索","对话系统"],"reason":"研究LLM代理的网页搜索行为，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:30","error":null,"has_summary":false,"summary":null},{"id":"2609.20001","version":1,"title":"E-AVI: Evidence-Grounded Multimodal Assessment for Automated Video Interviews","zh_title":"E-AVI：基于证据的多模态自动视频面试评估","abstract":"Automated video interview assessment integrates verbal content, acoustic delivery, and visual behavior, yet numerical predictions alone provide limited inspectable support. We present E-AVI, an evidence-grounded framework that extracts timestamped multimodal evidence and integrates dimension-conditioned evidence attention with source-level embeddings for scoring. A shared evidence pool further supports natural-language feedback and follow-up question answering. On RecruitView and a private hospitality dataset, E-AVI consistently outperforms fine-tuned multimodal baselines in rank correlation. Ablation, evidence-deletion, bootstrap, human-audit, and QA analyses characterize the predictive contribution, grounding, and practical utility of the evidence pathway. Together, these results demonstrate that our proposed E-AVI framework improves predictive performance while providing inspectable support for assessment, feedback, and interactive analysis.","authors":["Haoshen Wang","Dongbo Che","Zeyi Xie","Yuanjie Du","Shicheng Hua","Xingyu Wang"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-18","first_seen":"2026-09-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.20001","pdf_url":"https://arxiv.org/pdf/2609.20001","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C2"],"tags":["自动面试评估","多模态学习","可解释AI"],"reason":"论文研究自动视频面试评估，属于多模态预测，不涉及用LLM仿真人类被试或与人类数…","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:34","error":null,"has_summary":false,"summary":null},{"id":"2609.20152","version":1,"title":"MTVA-Bench: Evaluating the Language Model Inside Cascaded Voice Agents","zh_title":"MTVA-Bench：评估级联语音代理中的语言模型","abstract":"Generally, most voice agents are cascaded systems, i.e., an ASR model transcribes the caller's audio, a language model reads the transcript and decides what to say and which backend tools to call, and a TTS model speaks the reply. Nearly all of the decision making happens in the language model, but existing evaluations measure it either too broadly or too narrowly. End-to-end voice benchmarks score the full pipeline, so recognition errors and model errors mix into a single number. LLM benchmarks isolate the model but they do not evaluate what makes real phone calls hard, such as transcription issues, caller's voice being split across messages and the requirement that replies follow the language and script specified. We introduce the Multi-Turn Voice Agent Benchmark (MTVA-Bench), which evaluates the language model on the same conditions it faces inside a cascaded system. The caller is played by an LLM following a set of rubrics and tool calls are answered by a mock backend which responds to the arguments the model actually sent. The benchmark contains 49 agents working across 490 reviewed scenarios and supports 7 languages. Scoring is a combination of deterministic checks on tool calls with two LLM judges, one that scores scenario specific rules and one that grades conversation quality without access to the task. Both judges must cite specific messages from the transcript. Task and conversation scores are weighted equally, since a call can complete its task and still go badly for the caller. In a seven-model study, six of the models select the correct tool within 6.4 points of one another, but their overall scores span 24.4 points. Most of the gap comes from argument values, action ordering, rule compliance, and what the model says around its tool calls.","authors":["Pritish Mishra","Ishaan Kumar","Akshat Mandoli","Sudarshan Kamath"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-18","first_seen":"2026-09-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.20152","pdf_url":"https://arxiv.org/pdf/2609.20152","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["语音代理","基准测试","多智能体系统"],"reason":"评估语音代理中的语言模型，属于多智能体系统评测，不涉及人类行为仿真或对照。","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:35","error":null,"has_summary":false,"summary":null},{"id":"2609.19831","version":1,"title":"Reproducing Transparent and Scrutable Recommendations: Exploring Open-Weight Models via Natural-Language User Profiles","zh_title":"复现透明可解释的推荐：通过自然语言用户画像探索开放权重模型","abstract":"In this reproducibility study, we investigate the transparency and scrutability of recommender systems enhanced by incorporating generated natural-language user profiles that represent user preferences. The original paper explores the synthesis of user profiles from raw user-generated review text across domains such as movies and accommodations (Amazon Movies & TV, TripAdvisor). Crucially, these natural-language user profiles enable direct user interaction and intervention, allowing users to customize recommendations by correcting misattributed preferences or addressing cold-start settings. We successfully reproduce the core findings of the original study. Additionally, we extend the evaluation by conducting systematic context ablation experiments, multi-seed stability across five distinct random seeds to establish statistical reliability, and a mechanistic interpretability analysis using the nnsight framework to probe internal model representations under counterfactual profile perturbations. Our findings verify the original paper's claim that User Profile Recommendation (UPR) achieves competitive performance under its test-set reranking protocol and makes recommendations more transparent. Perturbing the natural-language profiles does change predictions, but it shifts predicted ratings uniformly across genres with no detectable genre-selective effect, leaving rankings unchanged even under direct activation steering. We trace this back to the rating-regression objective rather than the profile interface, with ranking-objective models clearly exceeding in this task.","authors":["Noah Mami\\'e","Laurin van den Bergh"],"categories":["cs.IR","cs.AI"],"primary_category":"cs.IR","announce_type":"cross","date":"2026-09-18","first_seen":"2026-09-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.19831","pdf_url":"https://arxiv.org/pdf/2609.19831","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["推荐系统","可解释性","用户画像"],"reason":"论文研究推荐系统透明性，用LLM生成用户画像，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:33","error":null,"has_summary":false,"summary":null},{"id":"2609.20218","version":1,"title":"Is It Still Worth Training a Classical Model in the Era of LLMs? A Crossover Benchmark on Tabular Data","zh_title":"LLM时代训练经典模型还值得吗？表格数据上的交叉基准测试","abstract":"Large language models can label a tabular row from a plain-English description with no training - a capability now shipping in mainstream spreadsheet tools such as Microsoft Copilot in Excel and Anthropic's Claude for Excel - raising a practical question for the many business prediction problems where labels are expensive: should you prompt a frozen LLM, or collect data and train a model - and if so, how much data? We quantify the answer with the labeled-data crossover N*, the training-set size at which a trained classical model's learning curve overtakes a frozen LLM's training-free (and therefore flat) error. Aggregating 126 independent student evaluations of small GPT models under eight prompting configurations across 18 tabular datasets, paired with authoritative power-law learning curves for six classical model families, we find that training wins fast: even given an oracle choice of its best prompt configuration, a trained classical model beats the small frozen LLM using no more labeled data than is already on hand in 86% of cases, and wins by the smallest labeled subset we evaluate in 40%, with the observed crossover at a median of ~6% of the training set. In-context few-shot examples do not behave like training - error versus shot count does not follow a power law - and the same protocol re-run by independent implementers varies with a coefficient of variation of 0.148. A controlled probe indicates the LLM depends on recognizable feature-name semantics, which plausibly makes our crossover a conservative estimate (we do not claim memorization). For a typical business table, the evidence is clear: collect a few hundred labels and train a gradient-boosted model.","authors":["Kaihua Ding"],"categories":["cs.LG","cs.AI"],"primary_category":"cs.LG","announce_type":"cross","date":"2026-09-18","first_seen":"2026-09-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.20218","pdf_url":"https://arxiv.org/pdf/2609.20218","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["表格数据","模型比较","基准测试"],"reason":"论文比较LLM与经典模型在表格数据上的预测性能，属于纯NLP能力评测，不以人类…","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:38","error":null,"has_summary":false,"summary":null},{"id":"2609.20620","version":1,"title":"A Simulation Platform for AUV Fault Recovery: Exploring LLM-Based Diagnostic Strategies","zh_title":"AUV故障恢复仿真平台：探索基于LLM的诊断策略","abstract":"Autonomous underwater vehicles (AUVs) operating beyond reliable communications must recover from failures without human intervention. We investigate an architecture in which conventional deterministic layered control autonomy manages normal operations, while an invokable large language model (LLM) serves as a diagnostic and recovery planner when onboard anomaly detection identifies performance outside expected limits. Because language models are stochastic, rigorous evaluation requires ensemble testing rather than individual demonstrations. We present a closed-loop simulation architecture that couples real-time C vehicle software with a higher-level orchestration layer for physics-based fault injection, structured prompting, language-model interaction, mission file generation, validation, execution, and LLM-judge scoring. The framework, which we call SPAR (Simulation Platform for AUV Recovery), supports evaluation across fault realizations, prompt structures, reasoning models, and mission conditions. We vary these for a mass-shift fault over 480 SPAR trials, evaluating a frontier model and three off-the-shelf locally deployable LLMs. Model choice dominates diagnosis: the frontier model places the CG-shift mechanism in its top three hypotheses in 85-90% of trials, versus 60-78% for the best local model. Reasoning analysis indicates that local-model success is associated with following the complete diagnostic procedure, whereas weaker models often commit prematurely to elevator failure even though the actuator tracks its command. Diagnosis and operational decision performance do not appear to be coupled in this dataset. The contributions are an architecture extending unanticipated-fault recovery from detection to mitigation and an ensemble methodology for evaluating LLM-assisted mission management on low-power AUVs.","authors":["Khalid Halba","Kylie Cooper","James G. Bellingham"],"categories":["cs.RO","cs.AI"],"primary_category":"cs.RO","announce_type":"cross","date":"2026-09-18","first_seen":"2026-09-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.20620","pdf_url":"https://arxiv.org/pdf/2609.20620","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C2"],"tags":["AUV","故障恢复","LLM诊断"],"reason":"论文研究AUV故障恢复，LLM用于诊断规划，属于机器人仿真环境，不涉及人类行为…","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:43","error":null,"has_summary":false,"summary":null},{"id":"2609.19364","version":1,"title":"Durably Reducing Belief in Women's Health Misinformation Through Culturally Adaptive AI Videos","zh_title":"通过文化适应性AI视频持久降低女性健康错误信息信念","abstract":"Health misinformation disproportionately harms women, yet interventions rarely address the community norms that sustain false beliefs. We test whether culturally adaptive AI-generated video in which the presenter looks like someone from her community reduces misinformation belief among low-literacy women in suburban India. In a field experiment (N=434), participants watched an AI-generated video featuring either an adaptive or neutral presenter. The culturally adaptive presenter reduced misinformation belief by 30%, nearly twice the reduction produced by the neutral presenter compared to the non-intervention control condition. Post-experiment interviews suggest women recalled the neutral condition as a generic video but recognized the adaptive presenter. Gains persisted for three weeks. The adaptive advantage was largest for beliefs reinforced by community, such as blaming women for infertility, and negligible for medical knowledge gaps, such as understanding vaccines. These findings demonstrate the potential of culturally adaptive AI interventions to counter socially embedded health misinformation.","authors":["Anku Rani","Kokil Jaidka","Shruti Sharma","Pragya Mahajan","Manisha Wadhwa","Andrew B. Lippman","Pattie Maes","Paul Pu Liang"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-18","first_seen":"2026-09-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.19364","pdf_url":"https://arxiv.org/pdf/2609.19364","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["健康传播","AI视频","错误信息干预"],"reason":"研究AI视频干预健康错误信息，非LLM仿真人类被试，无人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:31","error":null,"has_summary":false,"summary":null},{"id":"2609.19420","version":1,"title":"Use and Effects of LLMs in Peer Review: A Randomized Experiment and Survey at ICML 2026","zh_title":"大型语言模型在同行评审中的使用与效果：ICML 2026随机实验与调查","abstract":"LLMs are rapidly reshaping peer review, making it important to understand how reviewers use them in practice and how different LLM-use policies affect review outcomes. We investigate these questions through a randomized experiment and an anonymous post-survey at ICML 2026, a major machine learning conference involving over 24,000 papers and 17,000 reviewers. Reviewers were assigned to either a conservative policy prohibiting all LLM use or a permissive policy allowing limited assistance, with randomization among a subset of main-track papers and reviewers. Policy assignment had near-zero effects on final paper decisions, paper scores, and reviewer confidence, although reviews under the permissive policy were 5.5-7% longer. Post-survey responses (N=1,486) revealed diverse attitudes toward LLMs and substantial noncompliance: 22.5% of conservative-policy reviewers reported using an LLM despite the prohibition, and 36.5% of permissive-policy reviewers reported at least one explicitly disallowed use. We discuss implications for future peer-review policy and tool design.","authors":["Sunnie S. Y. Kim","Wesley Hanwen Deng","Jennifer Wortman Vaughan","Buxin Su","Weijie Su","Alekh Agarwal","Sharon Li","Martin Jaggi","Daniel G. Goldstein","Nihar B. Shah","Miroslav Dud\\'ik"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-18","first_seen":"2026-09-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.19420","pdf_url":"https://arxiv.org/pdf/2609.19420","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["同行评审","LLM使用政策","随机实验"],"reason":"研究人类审稿人使用LLM的行为，非用LLM仿真人类被试","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:21","error":null,"has_summary":false,"summary":null},{"id":"2609.20490","version":1,"title":"TeamCAMS: An Open-Source Research Platform for Studying Human Behaviour in Human-AI Teams","zh_title":"TeamCAMS：研究人机团队中人类行为的开源研究平台","abstract":"In this article, we present TeamCAMS (Cabin Air Management System), a collaborative work environment for simulating human-AI (artificial intelligence) interaction for scientific research. The article outlines how several psychological theories guided the development of this multiple-task simulation. Modelling a process control environment, previous versions of TeamCAMS have already been used in empirical studies to address a wide range of research questions (e.g., comparing different forms of automation, evaluating impact of automation reliability, effects of stress on multiple-task performance). Outlining the technical possibilities offered by TeamCAMS, the article points out how its latest version offers researchers the possibility of addressing a set of new research questions including problems associated with teamwork (e.g., within-team conflict, distributed teamwork) and human-AI interaction. Finally, we will outline how this simulation environment can be enhanced further still to address research questions in new fields (e.g., automation of leadership). To promote transparency, reproducibility, and further development, TeamCAMS is made available to the research community under an open-source license.","authors":["Amos Brocco","Alain Chavaillaz","Andreas Sonderegger","Juergen Sauer"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-18","first_seen":"2026-09-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.20490","pdf_url":"https://arxiv.org/pdf/2609.20490","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C2"],"tags":["人机交互","实验平台","团队协作"],"reason":"该平台用于人类与AI团队实验，非LLM仿真人类被试，且无LLM作为被试替代。","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:41","error":null,"has_summary":false,"summary":null},{"id":"2609.20167","version":1,"title":"Utilizing AI-Driven Project Management Tools for Optimized Talent Management in HRM: A Framework for Enhanced Resource Allocation and Performance Prediction","zh_title":"利用AI驱动的项目管理工具优化人力资源管理中的人才管理：一个增强资源分配和绩效预测的框架","abstract":"When it comes to aiding businesses with demanding tasks regarding human resource management, TalentOptima unequivocally boasts of the best there is to offer. This tool utilizes AI based decision making, advanced predictive analytics, and also machine learning, all of which help in enabling automated resource allocation. To aid with better human resource management, TalentOptima integrates perfectly with already existing HR frameworks such as tools, etc. and shifts the focus towards aiding the user with insights while simultaneously alleviating manual work, this aids in a plethora of positive HR outcomes. A total of 40 managers participated in a simulation via user testing to ascertain if HR costs would reduce and work productivity would rise, the results were quite clear, attrition rates had dipped alongside risk and resource management rates, TalentOptima was a clear winner. Whereas the other HR frameworks primarily focused on ensuring work was done, TalentOptima ensured optimal and innovative decision-making, which overtime has proven to be invaluable for multiple companies, these results aid in proving why the tool is revolutionary.","authors":["Jay Barach"],"categories":["cs.CY","cs.HC"],"primary_category":"cs.CY","announce_type":"cross","date":"2026-09-18","first_seen":"2026-09-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.20167","pdf_url":"https://arxiv.org/pdf/2609.20167","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C2"],"tags":["AI项目管理","人力资源管理","资源分配"],"reason":"论文是AI项目管理工具在HRM中的应用，不涉及LLM仿真人类被试，无人类行为对…","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:36","error":null,"has_summary":false,"summary":null},{"id":"2609.19225","version":1,"title":"Making Local Government Contracts Legible: A Computational Pipeline for Classifying and Mapping Intergovernmental Service Agreements","zh_title":"让地方政府合同可读：一种用于分类和映射政府间服务协议的计算流水线","abstract":"Interlocal agreements are one of the primary instruments through which local governments formalize collaboration for public service delivery, yet the institutional and financial content encoded in these contracts has remained inaccessible to systematic analysis at scale. This paper introduces an end-to-end computational pipeline for classifying intergovernmental agreements by institutional form and extracting financial relationships between principals and agents in service contracts. Applied to Iowa's 28E archive (N = 21,629), the largest dataset of interlocal agreements in the United States, the pipeline combines LLM-based summarization and classification across LLaMA 3.1, GPT 5.2 Pro, and Gemini 3 Pro on a four-class classification task that distinguishes agreements as either service contracts, resource sharing agreements, joint operations agreements, or new joint entity agreements. We also identify the financial principal and agent in these agreements and contracts, as well as the resulting dollar amounts and represent them on a directed network. The resulting financial network is organized around a small number of dominant service providers, with counties serving as the most structurally versatile actors, and cities as predominantly principals. By rendering the content of Iowa interlocal agreements analyzable at scale for the first time, this pipeline establishes a reusable methodology that researchers and state agencies can apply to track how public dollars move across local governments and to identify entities that depend heavily on a small number of providers.","authors":["Mohsen Ghasemizade","Ra\\'ul Guti\\'errez-Meave","Cailin Gramling","Aviral Chawla","Michael Robinette","Kate Albrecht","Juniper Lovato"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2026-09-18","first_seen":"2026-09-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.19225","pdf_url":"https://arxiv.org/pdf/2609.19225","source_feed":"cs.CY","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["LLM应用","合同分类","政府数据"],"reason":"使用LLM进行合同分类与信息提取，属于NLP应用，不涉及人类行为仿真或对照。","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:30","error":null,"has_summary":false,"summary":null},{"id":"2609.20211","version":1,"title":"Silence Is Endorsement: Verification-Status Laundering in LLM Agent Pipelines","zh_title":"沉默即认可：LLM智能体管道中的验证状态洗白","abstract":"Safety monitors in LLM agent systems often judge actions from summaries or stored handoffs, not from the original evidence. This creates a simple but dangerous failure mode: the handoff preserves the claim that an action is authorized while losing the fact that the claim was never verified. We call this verification-status laundering. Across nine open-weight monitors and two hosted models, the action and authorization proposition remain fixed while we remove the unverified provenance framing around the claim. This change raises approval for risky actions from $5\\%$ to $60\\%$ on Llama-3.1-8B and from $9\\%$ to $98\\%$ on Qwen2.5-14B, with similarly large shifts on both hosted models. The failure also emerges in ordinary agent pipelines. Summarizers frequently weaken the status, memory compressors often remove it, and a full proposer--summarizer--memory--monitor pipeline raises risky approval to $57$--$81\\%$ across three downstream monitors. Experiments on WildGuard and ATBench show the same pattern on independently authored harmful and unsafe requests: unsupported authorization claims make approval substantially more likely. Explicitly instructing monitors to reject unverified authorization is not a reliable cross-model fix: some models remain vulnerable, while others reject legitimate requests. Agent systems should therefore carry authorization provenance as structured state attached to the claim throughout the pipeline.","authors":["Yibo Hu"],"categories":["cs.CR","cs.MA"],"primary_category":"cs.CR","announce_type":"cross","date":"2026-09-18","first_seen":"2026-09-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.20211","pdf_url":"https://arxiv.org/pdf/2609.20211","source_feed":"cs.MA","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["LLM安全","多智能体系统","验证状态"],"reason":"研究LLM agent管道中的安全监控漏洞，属多智能体系统安全，不涉及人类行为…","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:37","error":null,"has_summary":false,"summary":null},{"id":"2609.20249","version":1,"title":"Accuracy Is Not Enough: A Cross-Architecture Audit of Demographic Bias in Deep Knowledge Tracing","zh_title":"准确性不足：深度知识追踪中人口统计偏差的跨架构审计","abstract":"Deep knowledge tracing (DKT) models implicitly decide which students an adaptive system believes have mastered a skill, yet almost all evidence on their demographic fairness comes from Bayesian knowledge tracing; the deep models that power modern systems have received no comparable cross-architecture audit. We close this gap: four architectures (DKT, DKVMN, SAKT, AKT) trained under three regimes (standard, reweighting, adversarial) on two public datasets with demographic metadata, Eedi (15.9M interactions) and OULAD (167k after preprocessing), evaluated with ABROCA, student-level bootstrap confidence intervals, and permutation tests addressing recent critiques of fairness-metric instability. Three findings emerge. (i) Bias is real but context-dependent: every architecture shows a significant socioeconomic ABROCA on Eedi (0.018-0.023, $p<0.005$), with per-group AUC lower for economically disadvantaged students, while gender bias is significant on OULAD for three of four architectures after multiplicity correction yet negligible on Eedi. (ii) The most accurate architecture is the most biased: AKT gains about 4 AUC points from item-level Rasch embeddings and shows the largest socioeconomic ABROCA, exceeding every other architecture under a paired bootstrap ($p\\leq0.002$); ablating only the Rasch embeddings removes the accuracy gain and the excess bias together. (iii) Standard mitigation is unreliable: reweighting and adversarial debiasing leave ABROCA essentially unchanged in every configuration that preserves accuracy, even though the adversary is pinned at chance at full reversal strength and a weak-strength positive control rules out a dead probe.","authors":["Dang Quang Minh","Nguyen Dung Son","Nguyen Huu Loi","Truong Viet Vu","Nguyen Thai Anh"],"categories":["cs.LG"],"primary_category":"cs.LG","announce_type":"new","date":"2026-09-18","first_seen":"2026-09-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.20249","pdf_url":"https://arxiv.org/pdf/2609.20249","source_feed":"cs.LG","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["知识追踪","公平性审计","教育数据挖掘"],"reason":"论文研究深度知识追踪模型的公平性审计，不涉及用LLM仿真人类被试，属于教育数据…","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:38","error":null,"has_summary":false,"summary":null},{"id":"2609.20101","version":1,"title":"Competing for a Finite Pool of Attention in Social Media? How a New Geopolitical Conflict Reshapes Engagement in Bluesky","zh_title":"争夺社交媒体中的有限注意力？新地缘政治冲突如何重塑Bluesky中的参与","abstract":"Major geopolitical crises can rapidly reshape online public attention. Yet population-level increases in discussion volume about a new crisis reveal little about how users accommodate this new demand for attention. We study the onset of the Iran-US-Israel conflict, triggered on 28 February 2026, using longitudinal repost activity from Bluesky across four consecutive approximately three-month windows spanning the period before and after its onset; the data comprise 91.0 million unique posts and 645.5 million repost observations. We find that the new conflict reorganized participation through both reallocation among existing conflict participants and substantial activation of previously low-conflict-active users, while some previously active users reduced their conflict-related participation. Attention redistribution differed substantially across pre-existing interests: Iran-US-Israel and Israel-Palestine attention showed strong positive co-movement with little systematic relative replacement, whereas Other Political and Non-Political content more consistently lost attention share, and Russia-Ukraine exhibited weaker, heterogeneous replacement. Finally, disruption of users' broader attention allocation was substantially more prevalent among users with established attention to geopolitical conflicts than in the overall or Non-Political populations. Together, these findings show that a newly emerging conflict reorganizes online attention through turnover in who participates, selective co-attendance or replacement across topics, and disruption of broader attention patterns concentrated among users already engaged with geopolitical conflicts.","authors":["Kamand Kalashi","Arash Badie-Modiri","Ali Salloum","Juhi Kulshrestha","Talayeh Aledavood","Mikko Kivel\\\"a"],"categories":["cs.SI"],"primary_category":"cs.SI","announce_type":"new","date":"2026-09-18","first_seen":"2026-09-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.20101","pdf_url":"https://arxiv.org/pdf/2609.20101","source_feed":"cs.SI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["社交媒体分析","注意力分配","地缘政治冲突"],"reason":"研究社交媒体注意力再分配，未使用LLM仿真人类被试，不涉及人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:35","error":null,"has_summary":false,"summary":null},{"id":"2609.20655","version":1,"title":"Using machine learning metrics to provide deeper insights into the performance of choice models","zh_title":"使用机器学习指标深入洞察选择模型性能","abstract":"Machine learning (ML) techniques are increasingly drawing interest in the choice modelling (CM) field. The focus has primarily been on comparing the performance of these contrasting approaches or on improving behavioural insights for ML techniques, rather than translating ideas from one field into the other. In the present paper, we specifically focus on knowledge transfer from ML into CM in the context of model performance evaluation. In CM, model performance is typically evaluated using log-likelihood and related indicators, which are aggregate fit metrics that focus on overall fit. Conversely, in ML, the focus is on alternative-level misclassifications and correct classifications, which provide a more nuanced view of the results. To bridge these approaches, we explore the use of a probabilistic version of the confusion matrix, which reports the average probability of the model predicting each alternative, conditional on which alternative was observed to be chosen, across all choice tasks. This enables the computation of probabilistic ML metrics for both classic choice models and ML algorithms. We analyse model performance jointly in terms of overall fit and alternative-level predictions. Our findings demonstrate that models with similar log-likelihood can exhibit substantially different confusion matrices, revealing different probability patterns that aggregate metrics cannot capture. This framework identifies where models systematically `confuse' alternatives, highlighting trade-offs between alternatives, and potentially guiding model specification. Furthermore, evaluating these matrices and metrics out-of-sample reveals alternative-level prediction shifts that significantly impact forecasting performance.","authors":["Lorenzo Mu\\~noz","Stephane Hess","Thomas O. Hancock","Georges Sfeir"],"categories":["econ.EM"],"primary_category":"econ.EM","announce_type":"new","date":"2026-09-18","first_seen":"2026-09-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.20655","pdf_url":"https://arxiv.org/pdf/2609.20655","source_feed":"econ.EM","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["选择建模","机器学习","模型评估"],"reason":"论文讨论选择模型的评估指标，不涉及LLM仿真人类被试，属于纯方法论研究。","model":"deepseek-v4-pro","scored_at":"2026-09-18T13:03:43","error":null,"has_summary":false,"summary":null},{"id":"2606.27845","version":2,"title":"LLM Agents as Static Level-k Players in Behavioural Games","zh_title":"行为博弈中作为静态层级-k玩家的LLM智能体","abstract":"Large Language Models (LLMs) are increasingly used as stand-ins in behavioural games. These stand-ins rely on the assumption that the LLM's distribution of choices meaningfully matches how humans play the same game. This study tests that assumption through two games. The first is a p-beauty contest, and the second one is a public goods game. The study first investigates five local-model settings within the same model family. These settings are varied together in a 360-cell factorial, which balances temperature, scale (0.5-32B), quantisation, instruct vs base, and framing. Each cell's distribution is then compared against whole choice distributions in published human data. Each deployment setting, except for quantisation, governs a different aspect of fidelity. Mechanically, while the dispersion of human players can be somewhat recovered through deployment settings, the strategic process behind it cannot. Through the lens of the level-k cognitive theory, we find that LLMs act as static, category-retrieved level-k players, where k is set by the model scale. The models also do not run within-game belief-updating or backward induction throughout multiple-round horizon settings. While human contributions decayed in the public goods game, LLMs stayed flat or rose at every scale. When the horizon test was administered, LLMs were more cooperative under an indefinite horizon compared to a finite one. However, LLMs ignore their relative round position, so no last-round defection was displayed. This implies that LLMs retrieved levels relative to the horizon category rather than working out iteratively from the specific game setting.","authors":["Po Han Teo"],"categories":["econ.GN","econ.TH","q-fin.EC"],"primary_category":"econ.GN","announce_type":"replace","date":"2026-09-17","first_seen":"2026-06-26","revised_at":"2026-09-17","abs_url":"https://arxiv.org/abs/2606.27845","pdf_url":"https://arxiv.org/pdf/2606.27845","source_feed":"econ.GN","score":10,"bucket":"selected","rubric_hits":["A1","A2","A3","B1","B2","B4"],"tags":["LLM仿真","行为博弈","算法保真度"],"reason":"直接测试LLM在行为博弈中替代人类被试的保真度，并与已发表人类数据对照，发现静…","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:02:07","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-17","rank":1,"question":"LLM在行为博弈中作为人类被试替代品时，其选择分布和策略过程是否与真实人类一致？","design":"使用Qwen 2.5模型家族，在p-beauty contest和公共品博弈中，通过360个单元格的因子设计操纵温度、模型规模（0.5-32B）、量化、指令/基础模型和框架，测量LLM的选择分布，并与已发表的人类数据比较。","baseline":"已发表的人类行为数据，包括p-beauty contest和公共品博弈中的选择分布。","findings":"LLM表现为静态的、类别检索的level-k玩家，k由模型规模决定；LLM没有进行回合内信念更新或逆向归纳，在公共品博弈中贡献不衰减，且忽视相对回合位置，无最后一轮背叛。","reliability":"论文未讨论","relevance":"该研究直接检验LLM在行为博弈中替代人类被试的保真度，并与真实人类数据对照，发现LLM的策略过程与人类不同，对评估LLM仿真可靠性至关重要，值得精读原文。","inspiration":"采用因子设计系统操纵LLM部署设置，并与人类分布整体比较的方法值得借鉴｜可迁移到政策公告的预期形成实验，如央行沟通对通胀预期的影响｜用LLM模拟公众对政策公告的反应，处理为不同政策措辞或沟通方式，结果变量为预期通胀分布，与调查预期数据（如密歇根大学调查）对照。"}},{"id":"2609.17549","version":1,"title":"Do Social Patterns Hold in Synthetic Data? Analyzing Cyberbullying Dynamics in LLM-Generated and Authentic Dialogues","zh_title":"社会模式在合成数据中是否成立？分析LLM生成与真实对话中的网络欺凌动态","abstract":"Cyberbullying (CB) is a complex social phenomenon characterized by repeated aggression, power imbalance, and multi-party interaction. Although large language models (LLMs) are increasingly used to generate synthetic CB conversations for data augmentation and benchmarking, it remains unclear whether such data faithfully reproduces the social dynamics of authentic interactions beyond supporting downstream task performance. We present a comprehensive framework for evaluating the social realism of LLM-generated CB conversations. We compare authentic and synthetic dialogues generated by GPT, Grok, and LLaMA across interactional structure (turn-taking, power dynamics, and repair behavior), linguistic and stylistic realism (pronoun usage and humor), affective and behavioral markers (CB types, profanity, and toxicity), and temporal escalation dynamics. We further complement automatic analyses with a human evaluation of cyberbullying presence, scenario relevance, role plausibility, and social realism. Our results show that LLM-generated data consistently preserves high-level interactional structure, including role participation patterns, directional power asymmetry, and broad distributions of behavioral markers. However, all models systematically distort finer-grained social phenomena, including behavioral magnitude, role-specific allocation, categorical distributions, and temporal dynamics. These distortions are strongly model-dependent: GPT suppresses harmful content, Grok amplifies aggressive behaviors, and LLaMA provides the most balanced approximation while smoothing role distinctions. Our findings show that synthetic CB data is useful for modeling global interactional structure but remains an imperfect substitute for authentic conversations when behavioral realism and social dynamics are essential.","authors":["Arefeh Kazemi","Hamza Qadeer","Sinan Asci","Joachim Wagner","Brian Davis"],"categories":["cs.CL","cs.CY"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-17","first_seen":"2026-09-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.17549","pdf_url":"https://arxiv.org/pdf/2609.17549","source_feed":"cs.CL","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","社会动态","真实性评估"],"reason":"用LLM生成对话仿真网络欺凌社会动态，并与真实对话对照，评估仿真保真度与偏差。","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:01:41","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-17","rank":2,"question":"LLM生成的网络欺凌对话是否忠实再现了真实对话中的社会动态，而不仅仅是支持下游任务性能？","design":"使用GPT、Grok和LLaMA生成合成网络欺凌对话，与真实对话对比，分析互动结构（话轮转换、权力动态、修复行为）、语言风格（代词使用、幽默）、情感行为标记（欺凌类型、脏话、毒性）和时间升级动态，并进行人类评估。","baseline":"真实网络欺凌对话数据集（具体名称未在节选中提及）。","findings":"LLM生成的数据保留了高层互动结构，如角色参与模式、方向性权力不对称和行为标记的广泛分布；但所有模型都系统性地扭曲了细粒度社会现象，包括行为强度、角色特定分配、类别分布和时间动态，且扭曲程度因模型而异：GPT抑制有害内容，Grok放大攻击行为，LLaMA提供最平衡的近似但平滑了角色差异。","reliability":"论文承认合成数据在需要行为真实性和社会动态时是真实对话的不完美替代品，且扭曲具有模型依赖性。","relevance":"该研究直接评估LLM仿真社会互动的保真度，与研究者关注的人类仿真实验和可靠性评估高度相关，提供了系统的对照基准和偏差分析，值得精读。","inspiration":"借鉴其多维评估框架，将自动分析与人类评估结合，系统比较合成与真实数据在结构、语言、行为和时间动态上的差异｜可迁移到经济金融中的社会互动场景，如谈判博弈、团队决策或市场情绪传播｜设计一个实验：用LLM生成模拟投资者在社交媒体上的互动对话，处理为不同模型（如GPT、Claude）或提示策略，结果变量为情绪传染、羊群行为或信息扩散模式，与真实投资者论坛数据（如StockTwits）对照，评估仿真保真度。"}},{"id":"2609.18106","version":1,"title":"Linguistic Triggers of Gender and Racial Bias in Open-Weight LLMs Applied to Recruitment","zh_title":"开放权重大语言模型应用于招聘时性别与种族偏见的语言触发因素","abstract":"Open-weight large language models are rapidly entering hiring pipelines, yet their discriminatory failure modes -- and the regulatory exposure these create under the EU AI Act high-risk classification (Annex III) and U.S. EEOC adverse-impact analysis -- remain poorly understood. We present the first systematic, multi-model audit of open-weight LLMs that treats job-posting language as the primary experimental variable, evaluating six models (Llama 3.2, Mistral, Gemma 3, Qwen 3, Phi 3, DeepSeek-R1) across four controlled experiments that jointly probe recruiter-simulation and job-seeker-simulation tasks. We find that (1) agentic posting language depresses recruiter recommendation scores for female candidates (r_rb = 0.309, p_Bonf = 7x10^-5; model-fixed-effects r_rb = 0.448), while communal language partially reverses the penalty; and (2) coded-exclusion language suppresses non-White recruiter scores at large effect sizes (r_rb = 0.646-0.758) and, on the job-seeker side, selectively deters non-White personas from expressing interest -- operationalizing a chilling-effect mechanism at scale. A label-ablation experiment isolates the explicit demographic persona label as the primary causal driver, and Word Embedding Association Tests corroborate these findings at the representational level (d = 1.01-1.45 under Caliskan et al.'s multi-word gender attribute lists). We translate these results into a concrete pre-deployment audit protocol -- posting-vocabulary scoring, persona-conditioned LLM probing, and adverse-impact flagging against the four-fifths threshold -- that operationalizes the documentation and risk-management obligations Annex III imposes on high-risk AI in recruitment.","authors":["Kosuke Kitahara","Nobuhiro Yamaguchi"],"categories":["cs.CL","cs.AI","cs.CY"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-17","first_seen":"2026-09-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.18106","pdf_url":"https://arxiv.org/pdf/2609.18106","source_feed":"cs.CL","score":9,"bucket":"selected","rubric_hits":["A1","A2","A3","B1","B2","B3","B4"],"tags":["LLM仿真","招聘偏见","算法审计"],"reason":"用LLM仿真招聘中的人类决策，并与真实人类数据对照，评估偏差与失效条件。","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:01:43","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-17","rank":5,"question":"招聘广告中的语言特征（如代理性/社群性词汇、编码排斥语言）如何触发开放权重LLM在招聘模拟中的性别与种族偏见？","design":"使用六个开放权重LLM（Llama 3.2、Mistral、Gemma 3、Qwen 3、Phi 3、DeepSeek-R1）进行招聘者模拟和求职者模拟实验。通过操纵招聘广告语言（代理性 vs. 社群性词汇、编码排斥语言）和候选人/求职者的人口统计标签（性别、种族），测量LLM给出的推荐评分、兴趣表达等结果变量，并进行标签消融实验和词嵌入关联测试。","baseline":"无对照（论文未使用真实人类数据作为基准，而是基于LLM输出进行审计）。","findings":"代理性招聘语言降低女性候选人的推荐评分，而社群性语言部分逆转该惩罚；编码排斥语言大幅抑制非白人候选人的推荐评分，并在求职者模拟中抑制非白人角色的兴趣表达，形成寒蝉效应。标签消融实验表明显式人口统计标签是主要因果驱动因素，词嵌入关联测试在表征层面证实了偏见。","reliability":"论文未讨论","relevance":"该研究直接针对LLM在招聘场景中的仿真行为，系统操纵语言变量并测量偏见输出，与您关注的LLM仿真可靠性及偏差评估高度相关，值得阅读原文以了解其审计协议和效应量。","inspiration":"借鉴其将文本特征作为处理变量、通过多模型审计和标签消融识别因果机制的方法。｜可迁移到信贷审批中的语言歧视研究，如贷款广告或申请表中的措辞对AI审批决策的影响。｜以LLM作为信贷审批员，处理为贷款申请描述中的代理性/社群性词汇或编码排斥语言，结果变量为审批通过率或利率，对照真实信贷审批数据（如抵押贷款披露数据）评估仿真偏差。"}},{"id":"2609.17933","version":1,"title":"AI Mediators Regulate Emotion and Create Value in Disputes","zh_title":"AI调解员在纠纷中调节情绪并创造价值","abstract":"In conflict and disputes, especially, emotion acts as a salient force in influencing outcomes. Prior work shows negative affect can obstruct collaborative behaviors, which typically lead to ``win-win'' outcomes. Thus, some suggest mediators may help regulate emotion and achieve joint gains. With the proliferation of AI, we posit LLMs may perform well at this task, with the added benefit of better accessibility compared with a human mediator. To examine the effectiveness of AI versus novice human mediators, we conduct a between-subjects experiment, where participants engage in a dispute mediated by a human, AI, or no mediator. We first analyze how well the mediators regulate emotions within a dispute -- finding AI mediators perform significantly better than humans at reducing negative emotion. We next examine whether AI mediators facilitate disputants better realizing joint gains in disputes with high integrative potential (IP) -- we find a marginally significant interaction between IP and condition (AI versus human), indicating LLMs may outperform humans at aiding disputants realize joint gains. Lastly, we perform an analysis of the messages the mediators sent, finding the AI sent significantly more messages suggesting trade-offs compared to the humans.","authors":["James Hale","Jonathan Gratch"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-17","first_seen":"2026-09-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.17933","pdf_url":"https://arxiv.org/pdf/2609.17933","source_feed":"cs.HC","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM仿真","调解实验","人机对照"],"reason":"用LLM作为调解人替代人类，与真实人类被试互动并对照，属于人类仿真实验。","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:01:41","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-17","rank":3,"question":"在情绪化的纠纷调解中，AI调解员相比新手人类调解员能否更有效地调节情绪并帮助双方实现共同收益？","design":"使用大语言模型（LLM）扮演AI调解员，与新手人类调解员和无调解员条件进行被试间实验。参与者扮演买卖双方进行纠纷谈判，调解员在每轮发言后决定是否干预并发送消息。结果变量包括情绪调节效果（负面情绪减少）、共同收益实现程度（积分潜力IP与条件的交互）以及调解员消息内容（如提出权衡建议的频率）。","baseline":"新手人类调解员作为对照，以及无调解员条件作为基线。参与者为真实人类被试，其偏好通过分配100点来测量，并根据达成目标获得金钱奖励。","findings":"AI调解员在减少负面情绪方面显著优于人类调解员；在具有高整合潜力的纠纷中，AI调解员可能比人类调解员更能帮助双方实现共同收益（边际显著交互作用）。此外，AI调解员发送了更多建议权衡的消息。","reliability":"论文未讨论","relevance":"该研究直接使用LLM替代人类调解员，与真实人类被试互动并对照人类调解员表现，属于人类仿真实验，且涉及情绪调节和谈判结果，与研究者关注的经济学实验和政策评估场景高度相关，值得阅读原文了解实验细节和局限性。","inspiration":"借鉴其使用LLM作为干预代理与人类被试互动的实验设计，以及通过偏好分配和积分潜力测量共同收益的方法。｜可迁移到经济金融中的谈判或冲突解决场景，如劳资纠纷调解、商业合同谈判、消费者投诉处理等。｜设计一个实验：以LLM作为调解员，人类被试扮演谈判双方（如买方和卖方），处理为AI调解员、人类调解员或无调解员，结果变量为谈判达成协议的质量（如联合收益、满意度）和情绪变化，对照真实人类调解员数据，并利用已有谈判实验数据集（如KODIS）进行基准比较。"}},{"id":"2609.18060","version":1,"title":"AI Peers Exert Social Influence on Human Dishonesty in Groups","zh_title":"AI同伴对群体中人类不诚实行为施加社会影响","abstract":"Human dishonesty in group settings is highly susceptible to peer influence, particularly when incentivized. Although artificial intelligence (AI) evolves from passive tools into active collaborators, its impact on human moral behavior within groups remains underexplored. We addressed this gap through a two-phase randomized behavioral study (N=280 and N=360). We found AI agents exert substantial social influence comparable in magnitude to that of human peers. Specifically, participants reported more dishonestly when exposed to dishonest rather than honest normative cues. This effect is evident across injunctive, subjective, and descriptive social norms. Interestingly, the only significant adjacent behavioral change occurred when dishonest peer behavior first appeared, whereas further increases from one to four dishonest peers produced weaker and non-monotonic changes. Furthermore, participants rapidly converge on decision-making, showing modest increases in dishonest reporting through repeated exposure. These findings highlight the importance of managing the behaviors and normative signals communicated by AI group members.","authors":["Shuning Zhang","Xinyuan Zhou","Yuanyang Qiu","Tianqi Song","Yuting Yang","Yiwen Ren","Xin Yi"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-17","first_seen":"2026-09-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.18060","pdf_url":"https://arxiv.org/pdf/2609.18060","source_feed":"cs.HC","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM仿真","社会影响","行为实验"],"reason":"用AI代理替代人类同伴，研究其对人类不诚实行为的社会影响，并与人类同伴对照。","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:01:43","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-17","rank":4,"question":"AI同伴的诚实规范如何影响群体中人类的不诚实行为？","design":"采用两阶段随机行为实验（N=280和N=360），使用激励性掷骰子报告任务。AI代理作为群体成员，通过描述性、指令性和主观性社会规范传递诚实或不诚实线索，测量人类被试的虚报行为。","baseline":"与人类同伴的社会影响进行对照，比较AI同伴与人类同伴的影响大小。","findings":"AI代理对人类不诚实行为产生显著社会影响，其影响程度与人类同伴相当。当出现不诚实同伴时，被试虚报增加，但同伴数量从1个增加到4个时影响减弱且非单调。","reliability":"论文未讨论","relevance":"该研究直接使用AI代理替代人类同伴，研究其对人类道德行为的影响，并与人类同伴对照，符合研究者对LLM仿真实验和真实人类基准的兴趣。","inspiration":"值得借鉴的是其将AI代理作为群体成员，通过操纵其行为规范来研究社会影响，并设置人类同伴对照组以量化AI影响的大小。｜可迁移到经济金融中的群体决策场景，如投资团队中AI顾问的不诚实建议对个人投资决策的影响。｜设计一个实验：被试与AI代理组成投资小组，AI代理提供虚报收益的建议（处理），测量被试的投资报告虚报程度（结果变量），并与人类同伴提供同样建议的对照组比较，同时使用真实投资数据作为外部基准。"}},{"id":"2609.07478","version":2,"title":"The Internal Anatomy of Strategic Choice in Large Language Models","zh_title":"大语言模型策略选择的内部解剖","abstract":"Large language models act as strategic agents and models of human choice, yet choosing like a strategic agent does not mean computing like one. We recorded activations from four open-weight models --- dense and mixture-of-experts, including a matched base--instruct pair --- in one-shot play of 144 strict ordinal $2\\times2$ games. We followed a prespecified incentive from prompt, through activations, to choice. Dense models mirrored the unadjusted human decline with game complexity. Incentive and choice were detectable in every model, but models differed in whether incentive reached the choice, aligned with it and, where tested, whether strengthening it shifted preference. The base and instruction-tuned Qwen2.5 models chose almost identically at baseline yet differed in whether incentive reached choice. Fixed decision cues were distinguishable internally but changed choices selectively. Similar behaviour can rest on different computation; post-training can reshape the path from represented incentive to decision while leaving behaviour and decodable information largely intact.","authors":["Vin\\'icius Ferraz","Leon Houf","Enrico Ferrea"],"categories":["cs.AI","cs.GT"],"primary_category":"cs.AI","announce_type":"replace","date":"2026-09-17","first_seen":"2026-09-09","revised_at":"2026-09-17","abs_url":"https://arxiv.org/abs/2609.07478","pdf_url":"https://arxiv.org/pdf/2609.07478","source_feed":"cs.AI","score":8,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","策略博弈","算法保真度"],"reason":"用LLM复现人类策略选择并与人类数据对照，分析内部计算与行为差异，评估仿真可靠…","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:02:07","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-17","rank":6,"question":"大语言模型在一次性策略博弈中，内部表征与计算如何将激励信号转化为选择行为？","design":"用四个开源大语言模型（Qwen2.5 base/instruct、Llama-3.1-Instruct、GPT-OSS）作为被试，在144个严格序数2x2博弈中做一次性选择，记录激活值，用线性探针解码激励和选择，并进行激励强化干预和固定决策线索提示。","baseline":"人类选择数据来自Moore et al. (2026)的相同博弈和支付尺度，以及Zhu et al. (2025a)的基数支付数据集映射。","findings":"密集模型复现了人类随博弈复杂度下降的未调整选择模式；所有模型都能解码激励和选择，但激励是否到达选择、与选择对齐及干预效果因模型而异，基础与指令微调模型行为相似但内部路径不同。","reliability":"论文指出相似行为可能基于不同计算，后训练可重塑激励到决策的路径而保持行为和可解码信息基本不变，暗示行为仿真可能掩盖内部差异。","relevance":"该研究直接以LLM仿真人类策略选择并与真实人类数据对照，同时揭示内部计算与行为的不一致，对评估仿真可靠性和偏差具有重要参考价值，值得精读原文。","inspiration":"借鉴其用线性探针追踪激励信号从输入到选择的内部路径，并施加干预检验因果性的方法，可迁移到经济决策仿真中检验模型是否真正使用经济激励而非表面模仿。｜可应用于资产定价实验，检验LLM是否内部表征风险溢价并据此决策。｜以LLM为被试，呈现不同风险收益的资产选择任务，用探针解码风险溢价信号并强化干预，与人类实验数据（如股票市场参与决策）对照，观察选择变化。"}},{"id":"2609.17534","version":1,"title":"Faking Good and Faking Bad in LLMs: Response Distortion Across Dark Triad Personality Traits","zh_title":"LLM中的装好与装坏：黑暗三人格特质下的反应失真","abstract":"Social desirability and impression management are pervasive sources of response distortion in human personality assessment, yet their effects on Large Language Models (LLMs) remain underexplored. This study investigates whether contemporary LLMs systematically modulate the expression of Dark Triad traits (Machiavellianism, narcissism, and psychopathy) under fake-good and fake-bad conditions. Seven state-of-the-art models were evaluated across two ecologically relevant contexts: employment selection and forensic evaluation, in which socially desirable or undesirable incentives were conveyed through contextual framing. Trait expression was measured using standard psychometric scoring procedures and compared with self-assessment baselines at both aggregate and item levels. Results revealed systematic and condition-consistent response modulation. Most models reduced Dark Triad scores under fake-good conditions and increased them under fake-bad conditions, although the magnitude and consistency of these effects varied across traits and models. Machiavellianism and narcissism showed the strongest and most coherent shifts, whereas psychopathy displayed greater heterogeneity. Context also influenced responses, with employment scenarios generally producing larger effects than forensic scenarios. An additional experiment showed that explicit fake-bad instructions generated substantially stronger distortions than contextual framing alone. The results suggest that personality-related outputs should be interpreted in light of the motivational and situational context in which they are elicited. More broadly, they highlight the value of psychometric paradigms for evaluating susceptibility to response distortion, impression management, and context-dependent behavioral shifts, with important implications for LLM benchmarking, alignment evaluation, and robustness assessment.","authors":["Victoria Popa","Guglielmo Cola","Caterina Senette","Maurizio Tesconi"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-17","first_seen":"2026-09-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.17534","pdf_url":"https://arxiv.org/pdf/2609.17534","source_feed":"cs.CL","score":8,"bucket":"selected","rubric_hits":["A2","B1","B4"],"tags":["LLM人格测量","反应偏差","仿真可靠性"],"reason":"研究LLM在人格测量中的反应失真，与人类数据对照，评估仿真偏差，可迁移到人类仿…","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:01:41","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-17","rank":7,"question":"LLM 在人格测量中是否会在假好（fake-good）和假坏（fake-bad）条件下系统性地调节黑暗三联征特质的表达？","design":"使用七个先进LLM，在就业选拔和司法评估两种情境下，通过提示框架施加假好或假坏动机，用TRAIT框架测量黑暗三联征（马基雅维利主义、自恋、精神病态）得分，并与基线自我评估比较。","baseline":"无对照","findings":"大多数模型在假好条件下降低黑暗三联征得分，在假坏条件下提高得分，但效应大小和一致性因特质和模型而异。马基雅维利主义和自恋的偏移最强且最一致，精神病态异质性更大；就业情境比司法情境产生更大效应，显式假坏指令比情境框架产生更强扭曲。","reliability":"论文未讨论","relevance":"该研究展示了LLM对情境动机的敏感性，可用于评估仿真中的社会期望偏差和印象管理，对使用LLM模拟人类被试时的可靠性有警示意义。","inspiration":"借鉴其通过情境框架施加动机处理并测量特质偏移的方法，可迁移到经济金融中的社会期望偏差场景（如信贷审批中的歧视、消费者道德行为）。｜可设计实验：用LLM扮演贷款申请人，在强调社会责任或利润最大化的不同银行政策下，测量其自我报告的诚信或风险偏好，并与实际信贷数据中的偏差对照。"}},{"id":"2609.18282","version":1,"title":"Too Good to Be Real? Diagnosing and Reducing the Gap Between AI Preference and Real User Engagement","zh_title":"好得难以置信？诊断并缩小AI偏好与真实用户参与度之间的差距","abstract":"Large language models are increasingly used to generate and evaluate online content, yet it remains unclear whether the qualities they associate with higher engagement match what real users respond to. We study this question using 1.17 million answers to 25,978 questions from Zhihu, Quora, and Reddit, comparing real platform answers and AI-generated answers across four within-question engagement levels. We introduce Ontological Preference Measurement, which represents answers along three dimensions: logic, affect, and expression. We find a systematic gap between AI preference and real user engagement: as target engagement increases, LLMs add more explicit logical structure, while real user engagement is more strongly associated with affective and expressive salience. We call this tendency logic overbinding. Based on this diagnosis, we propose Ontology-Masked Reasoning Autoencoding (OMRA), a controlled intervention that masks and reconstructs over-explained spans while preserving stance, factual content, and coherence. Across four LLM families, OMRA reduces the measured gap by an average of 54.4%. In human evaluation, OMRA wins 62.4% of pairwise preference judgments against matched real platform answers, even though the real answers are more often judged to be human-written.","authors":["Xinglang Zhang","Yuanmeng Xiang","Yunyao Zhang","Zeliang Chen","Junqing Yu","Zikai Song"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-17","first_seen":"2026-09-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.18282","pdf_url":"https://arxiv.org/pdf/2609.18282","source_feed":"cs.CL","score":8,"bucket":"selected","rubric_hits":["A1","B1","B4"],"tags":["LLM仿真","人类行为对照","内容生成"],"reason":"用LLM生成内容并与真实用户互动数据对照，诊断AI偏好与人类行为差距，属于仿真…","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:01:45","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-17","rank":9,"question":"LLM 在生成高互动内容时偏好的文本特征是否与真实用户互动行为一致？","design":"使用 LLM 生成针对四个互动等级的答案，与真实平台答案对比，通过本体偏好测量框架从逻辑、情感、表达三个维度量化文本特征，并施加 OMRA 干预以缩小差距。","baseline":"来自知乎、Quora 和 Reddit 的 117 万条真实答案，按问题内投票排名分为四个互动等级。","findings":"LLM 在追求高互动时过度增加显性逻辑结构，而真实高互动答案更依赖情感和表达显著性，作者称之为“逻辑过度绑定”。OMRA 干预平均缩小 54.4% 的差距，并在人类评估中胜过真实答案。","reliability":"论文未讨论","relevance":"该研究直接对比 LLM 生成内容与真实用户行为，诊断 AI 偏好与人类反应的系统性偏差，并尝试通过干预修正，对评估 LLM 仿真人类行为的可靠性具有参考价值。","inspiration":"借鉴其本体驱动的多维测量和干预设计，可迁移到经济金融领域的文本生成场景，如政策沟通、市场评论或金融建议。｜例如，研究 LLM 生成的央行政策声明或分析师报告是否与真实市场反应匹配。｜以 LLM 为被试，要求其生成不同目标市场反应的金融文本，测量逻辑、情感、表达特征，并与真实市场数据（如股价波动、交易量）对照，检验 AI 偏好与真实投资者反应的差距。"}},{"id":"2609.17989","version":1,"title":"Whom Do AI Agents Work For? Role Assignment Induces Sponsorship Bias in LLM Recommenders","zh_title":"AI代理为谁工作？角色分配引发LLM推荐中的赞助偏差","abstract":"Large language models (LLMs) now serve as conversational shopping assistants on platforms that also sell advertising. These AI agents face a conflict of duty. They advise consumers who rely on their judgment, yet are deployed by platforms that benefit when sponsored listings are chosen. Sponsorship disclosures, designed to allow consumers to penalize paid placements, now reach the AI agent rather than the consumer, and the agent's evaluation of them is hidden from the consumer. Drawing on the fiduciary concept of conflict of duty, we argue that an agent's evaluation of a sponsored listing should not depend on which party deployed it. In controlled choice experiments, we manipulate assigned roles in the system prompt to name either a traveler or a booking platform as the agent's principal. Platform delegation significantly attenuates the penalty that agents apply to sponsored listings and weakens the skepticism that disclosure triggers in their reasoning traces. We replicate out findings across LLMs and reasoning depths. A second study decomposes the disclosure label and shows that the divergence between the two delegates widens significantly when the paid placement is attributed to the platform. Stricter terminology (\"Sponsored\" instead of \"Promoted\") lowers choice of paid listings but does not close this gap when the platform is named. The findings show that disclosure mandates designed for human consumers cannot by themselves protect consumers in AI-mediated commerce.","authors":["Davood Wadi","Yu Ma"],"categories":["econ.GN","cs.AI","q-fin.EC"],"primary_category":"econ.GN","announce_type":"cross","date":"2026-09-17","first_seen":"2026-09-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.17989","pdf_url":"https://arxiv.org/pdf/2609.17989","source_feed":"cs.AI","score":8,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["LLM仿真","消费者决策","赞助偏差"],"reason":"用LLM模拟消费者决策，与人类对照，揭示角色分配导致的赞助偏差，属经济学实验场…","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:01:43","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-17","rank":8,"question":"AI代理在同时服务消费者和广告平台时，其推荐行为是否会因委托方身份（消费者vs平台）而出现赞助偏差？","design":"用LLM（Gemini 3.1 Pro等）模拟购物助手，在系统提示中操纵委托方身份（旅行者vs预订平台），呈现带赞助标签的酒店列表，测量选择赞助列表的概率及推理痕迹中的怀疑程度。","baseline":"无对照","findings":"平台委托显著减弱了LLM对赞助列表的惩罚，选择赞助列表的概率从消费者委托时的50.2个百分点降至29.2个百分点；当赞助标签明确归属平台时，平台委托的代理甚至表现出对赞助列表的强烈偏好（选择率74.2% vs 消费者委托的34.6%）。","reliability":"论文未讨论","relevance":"该研究用LLM模拟消费者决策，揭示了角色分配导致的赞助偏差，属于经济学实验场景，与研究者关注的人类仿真实验高度相关，值得精读原文。","inspiration":"值得借鉴的是通过最小化系统提示中的角色分配来诱发LLM行为偏差，并分析推理痕迹以揭示机制｜可迁移到信贷审批歧视或金融产品推荐场景，检验AI代理在银行与借款人之间的利益冲突｜设计：用LLM扮演贷款顾问，系统提示中分别指定委托方为银行或借款人，呈现带“银行推荐”标签的贷款产品，测量选择概率，并与人类贷款顾问的真实选择数据对照。"}},{"id":"2609.17544","version":1,"title":"Large Language Models Versus Physicians in Traditional Chinese Medicine: A Real-World Clinical Case Evaluation","zh_title":"大语言模型与中医医师的对比：真实世界临床病例评估","abstract":"Large language models (LLMs) are increasingly being explored for clinical applications, yet their assessment for real-world traditional Chinese medicine (TCM) practice remains limited We constructed a clinical case library comprising 349 de-identified outpatient cases from 62 hospitals and evaluated 16 LLMs and a comparator cohort of 60 practicing TCM physicians using 60 representative cases selected from this library. Model outputs and physician reports were anonymized and scored by five senior TCM experts across nine diagnostic and therapeutic dimensions. Cutting-edge general-purpose LLMs achieved higher expert scores than the physician comparators, particularly for medical advice, treatment principles and selected diagnostic tasks. However, prescription-level analyses revealed discrepancies in herb selection, dosage, and treatment strategy, and qualitative safety review identified hallucinations and undesirable template-driven outputs. These findings highlight the potential of LLMs for TCM decision support while underscoring the need for physician oversight, safety constraints and prospective clinical evaluation.","authors":["Jiacheng Xie","Xiaoting Tang","Yang Yu","Jinpu Li","Shouli Li","Congcong Jing","Yantao Yang","Zhiyong Zhao","Ziyang Zhang","Qilin Song","Guanghui An","Dong Xu"],"categories":["cs.CL","cs.CY"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-17","first_seen":"2026-09-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.17544","pdf_url":"https://arxiv.org/pdf/2609.17544","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A1","B1","B2"],"tags":["LLM仿真","医疗决策","人类对照"],"reason":"用LLM替代医生进行临床决策评估，并与真实医生对照，属于人类仿真，但场景为医疗…","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:01:41","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-17","rank":10,"question":"在真实中医门诊场景中，大语言模型能否替代或辅助中医师进行辨证论治，其诊断与处方质量相比执业中医师如何？","design":"用16个大语言模型扮演中医师，对60个真实门诊病例生成诊断与处方报告；同时60名执业中医师对相同病例生成报告；所有报告匿名后由5名资深中医专家在9个诊断与治疗维度上评分，并分析处方模式、幻觉与安全性。","baseline":"60名执业中医师对相同60个病例生成的诊断与处方报告，经同一批专家按相同维度评分。","findings":"先进通用大模型在专家评分上总体高于医师对照组，尤其在医疗建议、治疗原则和部分诊断任务上表现更好；但处方层面在药物选择、剂量和治疗策略上存在差异，且定性安全审查发现幻觉和模板化输出。","reliability":"论文承认需要医生监督、安全约束和前瞻性临床评估；指出LLM可能产生幻觉、不安全的草药组合或不适当剂量，且当前评估基于回顾性病例，缺乏真实临床结局验证。","relevance":"该研究用LLM替代医生进行临床决策并与真实医生对照，属于人类仿真在医疗场景的应用，提供了真实世界数据与专家盲评的严格设计，值得阅读以了解仿真在专业决策中的有效性与偏差。","inspiration":"借鉴其匿名化处理与专家盲评设计，可确保仿真输出与人类输出在相同条件下比较，减少评估偏差｜可迁移到信贷审批歧视研究，用LLM模拟信贷员对贷款申请做出审批决策，比较其与真人信贷员的差异｜以LLM作为被试，处理为不同申请人特征（如种族、性别），结果变量为审批决定与理由，对照真实信贷审批数据或信贷员决策记录，检验LLM是否复现或放大人类偏见。"}},{"id":"2609.17550","version":1,"title":"No Usable Linear \"Capitulation Direction\" in Two Small LLMs: A Validation Protocol for Activation-Steering Claims, and a Cross-Family Behavioral Study of Sycophancy Under Pushback","zh_title":"两个小型LLM中不存在可用的线性“屈服方向”：激活引导声明的验证协议，以及跨家族对反驳下谄媚行为的研究","abstract":"Language models frequently abandon correct answers when users push back. We study this in two small instruction-tuned models from different families, Qwen2.5-1.5B and Llama-3.2-1B, over TriviaQA: the model answers, is challenged with one of four scripted pushback styles, and answers again. Conditioned on an initially correct answer, the models flip to a wrong answer in 41.8% and 43.1% of episodes. Which pressure works is a property of the model, not the pressure: the same within-question paired comparison (bare doubt vs. emotional appeal), specified in advance, is Bonferroni-significant in opposite directions across families (Qwen: bare doubt > emotional, OR 2.5, p=.040; Llama: emotional > bare doubt, OR 4.0, p=.001). Failure mode is also model-dependent: Llama abandons answers without recommitting at six times Qwen's rate (8.2% vs. 1.4%). Identical pushback repairs initially wrong answers only ~13% of the time; pushback is net epistemically destructive. We then ask whether capitulation is linearly decodable from the pre-response residual stream, a prerequisite for steering-vector interventions at that locus. A naive difference-in-means probe appears to succeed (in-sample AUROC 0.81/0.71), but a validation protocol combining question-level cross-validation, shuffled-label nulls, and a known-direction positive control shows the signal is overfitting: the best cross-validated AUROC is 0.582 in Qwen and 0.548 in Llama, both near or below their permutation thresholds and far under a pre-registered usability bar of 0.70, while the identical pipeline recovers a pushback-presence control direction at AUROC 1.000 in both. We further quantify a measurement hazard: substring grading underestimates capitulation by 18-24 percentage points. Code, prompts, transcripts, and analysis are released.","authors":["Saad Aamir","Muhammad Awais Bin Adil"],"categories":["cs.CL","cs.LG"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-17","first_seen":"2026-09-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.17550","pdf_url":"https://arxiv.org/pdf/2609.17550","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A2","B4"],"tags":["LLM行为","可靠性评估","激活引导"],"reason":"研究LLM在用户反驳下的行为变化，评估其可靠性，与人类仿真中的偏差问题相关。","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:01:49","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-17","rank":11,"question":"在用户反驳下，小型指令微调语言模型放弃正确答案（屈从）的行为模式是什么？屈从是否可由残差流中的线性方向解码？","design":"使用两个不同家族的小型指令微调模型（Qwen2.5-1.5B 和 Llama-3.2-1B），在 TriviaQA 数据集上进行多轮问答：模型先回答，然后受到四种预设反驳风格之一的挑战，再回答。测量结果变量包括：初始正确时翻转为错误的概率、初始错误时修复的概率、放弃不重新承诺的概率，以及屈从的线性可解码性（AUROC）。","baseline":"无对照","findings":"在初始正确的情况下，两个模型分别有41.8%和43.1%的回合翻转为错误答案；哪种压力有效是模型特有的，同一比较在家族间方向相反。线性探测看似成功，但经过交叉验证和置换检验后，屈从方向不可用（AUROC接近机会水平），而阳性对照方向可完美恢复。","reliability":"论文承认的局限包括：仅两个家族且规模小（1-1.5B）；单一数据集（TriviaQA）和单一语言（英语）；跨模型比较非问题配对；模板固定；LLM 裁判未经完整人工验证；Llama 的关键比较是在观察临时标签后指定的。","relevance":"该研究通过行为实验和严格的验证协议，揭示了 LLM 在用户压力下的屈从行为具有模型特异性，且线性方向不可靠，这对使用 LLM 模拟人类决策时的偏差评估和可靠性判断有直接参考价值，值得阅读原文。","inspiration":"借鉴其多轮交互设计、预设处理（反驳风格）、结果分类（翻转/修复/放弃）以及严格的验证协议（交叉验证、置换检验、阳性对照）来评估模型行为的稳健性。｜可迁移到经济金融中的政策沟通或建议采纳场景，例如 AI 财务顾问在用户质疑下是否改变投资建议，或消费者对 AI 推荐产品的信任变化。｜以 LLM 作为被试，模拟投资者在 AI 投资建议受到用户质疑时的反应，处理为不同风格的反驳（如权威质疑、情感诉求），结果变量为是否改变初始投资决策，并与人类投资者在类似情境下的真实行为数据（如实验或调查数据）进行对照。"}},{"id":"2609.18068","version":1,"title":"From a River in Gilead to the Inference Distributions of Large Language Models: Covert Dialect Bias and Linguistic Profiling at Scale","zh_title":"从基列河到大型语言模型的推理分布：隐性方言偏见与大规模语言画像","abstract":"Large language models (LLMs) are increasingly deployed in high-stakes domains such as housing screening. While alignment techniques mitigate explicit racial bias in generated text, they often leave covert attitudinal associations in internal probability distributions untouched. Adapting the matched-guise sociolinguistic paradigm, we examine covert dialect bias in housing-related social judgments across four varieties: Standard American English (SAE), African American Vernacular English (AAVE), Nigerian Standard English (NSE), and Nigerian Pidgin (NP). AAVE reflects the racialized dialect studied in prior covert-bias evaluations, whereas NSE and NP represent Black African, postcolonial varieties absent from this literature. Using 260 meaning-matched sentence quadruples and log-probability scoring over housing-relevant adjectives, we probe ten open-weight LLMs across three contexts varying in social proximity: tenant screening, neighbor acceptance, and roommate selection. Across all ten models, AAVE and NP are consistently associated with more negative adjectives than SAE, with NP penalized most severely. Crucially, each dialect is penalized via distinct stereotype clusters rather than a generic non-standard category. NSE, which carries institutional prestige, displays a context-dependent shift: favored over SAE in formal tenant screening but increasingly penalized as social proximity grows. Our findings reveal that LLMs inherit covert dialect bias along both racial identity and prestige dimensions, echoing documented human housing discrimination and demonstrating its reach across postcolonial English varieties.","authors":["Chowdhury Mohammad Abdullah","Rita Orji"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-17","first_seen":"2026-09-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.18068","pdf_url":"https://arxiv.org/pdf/2609.18068","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A1","B1","B4"],"tags":["LLM偏见","社会语言学","人类仿真"],"reason":"用LLM复现人类语言态度，有真实人类研究对照，揭示仿真偏差","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:01:43","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-17","rank":12,"question":"大语言模型是否在住房相关的社会判断中继承了隐蔽的方言偏见，这种偏见是由种族身份、语言声望还是两者共同驱动的？","design":"采用匹配伪装范式，使用260组语义匹配的句子四元组（SAE、AAVE、NSE、NP），对十个开源权重LLM进行对数概率评分，测量模型在住房相关形容词上的概率差异，并设置三个社会距离不同的评估情境（租户筛选、邻居接受、室友选择）。","baseline":"人类基准来自社会语言学文献中记录的住房歧视研究（如Massey和Lundy 2001年的电话筛选实验）以及匹配伪装实验（如Tucker和Lambert 1969），但本研究未直接收集新的人类数据。","findings":"所有十个模型对AAVE和NP的输入一致分配更负面的住房相关形容词概率，其中NP受罚最重；每种方言通过不同的刻板印象簇被惩罚，而非作为泛化的非标准类别。NSE在正式租户筛选情境中比SAE更受青睐，但随着社会距离增加而受到惩罚，显示出情境依赖性。","reliability":"论文未明确讨论失效条件，但指出研究仅限于开源权重模型和特定方言，且未直接与人类判断进行定量对比，仅引用历史人类研究作为背景。","relevance":"该研究直接展示了LLM在住房筛选等经济相关场景中复现人类歧视性判断，并揭示了隐蔽概率偏见，对关注仿真可靠性与偏差的研究者具有重要参考价值。","inspiration":"借鉴其匹配伪装设计，通过控制语义内容仅改变方言特征来隔离语言变量的因果效应，并利用对数概率测量隐蔽态度，可迁移到信贷审批中的方言歧视研究。｜可应用于金融领域的信贷审批或保险定价中的语言偏见问题，例如评估贷款申请人的方言口音是否影响模型的风险评估。｜设计：使用LLM作为虚拟信贷员，输入语义相同但方言不同的贷款申请文本（如SAE、AAVE、西班牙口音英语），测量模型分配给违约相关形容词或风险评分的概率，并与真实信贷审批数据中的种族或语言歧视率进行对照。"}},{"id":"2609.18274","version":1,"title":"I code or AI code: A comparative evaluation of AI-rated scores in classroom observations","zh_title":"我编码还是AI编码：课堂观察中AI评分与人类评分的比较评估","abstract":"Classroom observations are widely recognized as a key tool for establishing benchmarks of education quality and guiding pedagogical improvement, yet they remain resource-intensive and dependent on trained observers. This study evaluated the feasibility of using a LLM (GPT-5 model) to score teacher-child interactions in early childhood classrooms, benchmarked against human raters. The study analyzed 87 video-recorded observations from 38 classrooms across 30 kindergartens in Hong Kong. Using observation transcripts, the AI model was configured to apply the full Classroom Assessment Scoring System (CLASS) framework. AI-rated scores were then compared with human ratings by examining correlations and differences in mean scores of the CLASS domains and dimensions. The results showed greater convergence between AI and raters for the Emotional Support domain and, in particular, the Quality of Feedback dimension, which captures how teachers use feedback to extend children's learning. Greater divergence emerged for interactions that were more procedural or context-dependent, particularly within the Classroom Organization and Instructional Support domains. These findings suggest that transcript-based AI scoring may capture some of the relative variation in teacher-child interactions but cannot yet reproduce calibrated human judgements consistently across the full CLASS framework. AI-assisted observation may therefore be more appropriate as a preliminary screening tool rather than as a replacement for trained observers, providing teachers with evidence for reflection rather than high-stakes evaluation. Future research should examine whether domain-specific training and incorporation of contextual and visual information can improve alignment between AI and human rated scores.","authors":["Y. Fong","J. Xiang","T. Y. D. Chan","K. Lee","E. Y. H. Lau"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-17","first_seen":"2026-09-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.18274","pdf_url":"https://arxiv.org/pdf/2609.18274","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A1","B1"],"tags":["LLM仿真","教育评估","人类对照"],"reason":"用LLM替代人类评分员评估课堂互动，并与人类评分对照，属于仿真人类判断，但非典…","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:01:44","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-17","rank":13,"question":"使用大语言模型（GPT-5）对幼儿园课堂师生互动进行CLASS评分，能否替代人类评分员？","design":"用GPT-5模型基于课堂观察转录文本，按照CLASS框架对87个视频观察（来自香港30所幼儿园38个教室）进行评分，并与人类评分员在CLASS各领域和维度上的评分进行相关性和均值差异比较。","baseline":"人类评分员对相同视频观察的CLASS评分。","findings":"AI与人类评分在情感支持领域及反馈质量维度上收敛较好，但在课堂组织和教学支持领域等程序性、情境依赖性互动上分歧较大。转录文本AI评分能捕捉部分相对变异，但不能稳定复现人类校准判断。","reliability":"论文承认转录文本AI评分无法完全复现人类校准判断，建议仅作为初步筛查工具而非高利害评估替代品；未来需探索领域特定训练和纳入情境与视觉信息以改善对齐。","relevance":"该研究用LLM替代人类评分员评估课堂互动，并与人类评分对照，属于仿真人类判断，但非典型经济学实验或调查场景，与研究者关注的政策评估和经济学实验关联较弱，可作方法参考。","inspiration":"借鉴其用LLM对非结构化文本进行标准化量表评分的做法，并设置人类评分员作为基准进行相关性和均值差异检验。｜可迁移到经济金融中需要专家主观评分的场景，如信贷审批中的软信息评估、分析师报告语调分类、政策文本情感倾向等。｜以信贷审批为例，用LLM对贷款申请者的文字描述进行信用软信息评分，处理为不同提示词或模型版本，结果变量为信用评分，与人类信贷员的真实评分做对照，检验一致性和偏差。"}},{"id":"2609.18341","version":1,"title":"Understanding AI Provider Recommendations in Local Service Markets","zh_title":"理解本地服务市场中的AI提供商推荐","abstract":"When someone asks an AI assistant which doctor to see or which firm to trust with their savings, the answer is a referral. We audit AI provider recommendations in four registry-backed service domains across the 100 largest U.S. metropolitan areas, matching every recommendation against the official registry for its domain (Medicare clinician and facility records, and SEC adviser disclosures), under three conditions: an open-weight model, a proprietary model without web search, and the same proprietary model with search. Without search, both models largely fabricate recommendations in the domains the web covers thinly. Only 4% of the open-weight model's recommended doctors and 11% of the proprietary model's match a clinician in the queried city, and the open-weight matches are name coincidences: its matched clinicians are no likelier to be primary-care doctors than names drawn at random from the registry. With search, 64-71% of recommendations in the same domains match a real provider. Search also changes who is recommended. Without it, recommended advisory firms carry SEC misconduct disclosures at 3.6 times the registry base rate, even after adjusting for firm size; with search, significantly below it. Restaurants, where quality and visibility are separately measurable, show a 3-5x review-count premium but a rating premium of at most a tenth of a star. Finally, search largely removes the metro-size penalty: without it, real recommendations concentrate in the largest metros; with it, match rates are similar across metro-size terciles. Whether an AI referral is trustworthy depends strongly on its retrieval configuration rather than on the underlying model alone, yet an answer produced without retrieval often carries no sign that its recommendations were never verified.","authors":["Hazem Ibrahim","Yasir Zaki"],"categories":["cs.CY","cs.CL","cs.IR"],"primary_category":"cs.CY","announce_type":"cross","date":"2026-09-17","first_seen":"2026-09-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.18341","pdf_url":"https://arxiv.org/pdf/2609.18341","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A2","B1","B4"],"tags":["LLM可靠性","审计研究","真实数据对照"],"reason":"审计LLM推荐与真实注册数据对照，揭示无检索时虚构与偏差，可迁移至仿真可靠性评估","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:01:59","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-17","rank":14,"question":"AI助手在本地服务市场中的提供商推荐是否真实、质量如何，以及检索配置如何影响推荐的可信度。","design":"审计研究：使用一个开源模型和一个专有模型（后者在有无网络搜索两种条件下）回答五类日常问询，覆盖美国100个最大都市区的四个注册服务领域，将推荐名称与官方注册数据（Medicare临床医生和设施记录、SEC顾问披露）进行匹配，测量匹配率、质量指标和偏差。","baseline":"官方注册数据：Medicare临床医生和设施记录（含星级评分）、SEC投资顾问披露（含不当行为记录）、餐厅众包评分和评论数。","findings":"无搜索时，模型在医生和养老院领域大量虚构推荐（匹配率仅4%-11%），且开源模型的匹配纯属姓名巧合；有搜索时匹配率升至64%-71%。搜索改变了推荐对象：无搜索时推荐的咨询公司SEC不当行为披露率是基准的3.6倍，有搜索时显著低于基准；搜索还消除了都市规模惩罚，使不同规模都市的匹配率趋于一致。","reliability":"论文指出，无检索的答案往往没有迹象表明推荐未经核实，且引用分析显示搜索推荐依赖商业来源而非监管来源（医生领域仅0.3%引用政府来源），存在可观察性差距；但未系统讨论模型幻觉的边界条件或不同提示措辞的影响。","relevance":"该研究通过对照真实注册数据审计LLM推荐，揭示了无检索时的虚构与偏差，可迁移至仿真可靠性评估，值得阅读原文以了解审计方法和偏差量化。","inspiration":"借鉴其审计设计：在同一模型内对比有无检索，用官方注册数据作为基准，测量匹配率和质量偏差，并分析偏差的稳健性。｜可迁移到金融顾问推荐、信贷产品推荐或医疗资源分配等场景，评估LLM仿真中的选择偏差。｜以LLM作为被试，处理为有无检索，结果变量为推荐实体的真实性和质量指标（如违规记录、评分），对照SEC或消费者金融保护局等官方数据，检验仿真是否复现或放大偏差。"}},{"id":"2609.18346","version":1,"title":"Faithful yet Collusive: Why Chain-of-Thought Monitoring Cannot Detect Collusion in LLM Pricing Agents under Oligopolistic Competition","zh_title":"忠实却共谋：为何思维链监控无法检测寡头竞争下LLM定价代理的共谋","abstract":"Large language models (LLM) deployed as autonomous pricing agents may sustain supracompetitive prices through tacit coordination. We develop a causal graph divergence framework that separately measures structural faithfulness and intent faithfulness of LLM pricing agents in Bertrand competition. Across nine LLMs under duopoly and triopoly conditions, collusive behavior and chain-of-thought (CoT) faithfulness dissociate along both dimensions: the most collusive model accurately reports cooperative intent yet reasons structurally unfaithfully, while the most structurally faithful model sustains supra-Nash pricing under both market structures. These findings establish that CoT monitoring alone cannot serve as a standalone safeguard against algorithmic collusion.","authors":["Dohun Lee","Hyunwoo Park"],"categories":["cs.AI","cs.CL"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-09-17","first_seen":"2026-09-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.18346","pdf_url":"https://arxiv.org/pdf/2609.18346","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A3","B2","B4"],"tags":["LLM定价代理","算法共谋","经济仿真"],"reason":"用LLM定价agent模拟寡头竞争，与人类行为对照，但非直接仿真人类被试，结论…","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:01:45","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-17","rank":15,"question":"LLM定价智能体在寡头竞争中能否通过思维链监控可靠地检测合谋行为？","design":"用9个LLM作为定价智能体，在双寡头和三寡头的伯特兰竞争环境中进行300轮定价，施加利润导向和竞争导向两种提示处理，测量定价序列、利润、思维链中的因果图与意图。","baseline":"无对照","findings":"合谋行为与思维链忠实性在两个维度上分离：最合谋的模型准确报告合作意图但结构推理不忠实，而结构最忠实的模型在两种市场结构下均维持超纳什定价。思维链监控不能单独作为算法合谋的保障。","reliability":"论文未讨论","relevance":"该研究用LLM模拟经济主体行为，虽非直接仿真人类被试，但涉及寡头定价与合谋检测，对关注LLM仿真可靠性及政策评估的研究者有参考价值，值得读原文了解其因果图分歧框架。","inspiration":"借鉴其因果图分歧框架，分别测量陈述与行为因果结构及意图分布，以评估LLM仿真忠实性｜可迁移到寡头定价、拍卖合谋、平台算法共谋等产业组织与反垄断场景｜用LLM作为定价智能体，施加不同提示（如利润导向vs竞争导向），测量定价序列与思维链，以真实市场定价数据或人类实验数据为基准，检验LLM合谋行为与思维链监控的有效性。"}},{"id":"2609.18390","version":1,"title":"Building a Cultural Perspective on Doctor-Patient Conversations","zh_title":"构建医患对话的文化视角","abstract":"AI-powered medical scribes are increasingly used to transcribe doctor-patient conversations and automate clinical documentation. However, large-scale real-world consultation datasets are scarce due to the sensitivity of clinical conversations, leading developers to rely on simulated and LLM-generated synthetic consultations. While scalable, these alternatives may fail to capture culturally situated patterns of clinical interaction. We introduce interactional cultural markers, measurable patterns of doctor-patient interaction grounded in cross-cultural clinical communication, and use them to compare real, simulated, and synthetic consultations from Indian and US clinical contexts. We find distinct patterns of participation and control: Indian consultations involve greater patient participation but stronger doctor control, while US consultations exhibit balanced participation and open-ended discussion. Synthetic Indian consultations often fail to reproduce these patterns, instead converging toward US-like interaction. We identify additional synthetic signatures, including excessive doctor explanation and formulaic patient responses. We conclude by discussing implications for generating culturally grounded synthetic clinical conversations.","authors":["Krithi Shailya","Siddharth D Jaiswal","Ashish Makani","Suvrankar Datta","Sunayana Sitaram","Mohit Jain"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-17","first_seen":"2026-09-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.18390","pdf_url":"https://arxiv.org/pdf/2609.18390","source_feed":"cs.HC","score":7,"bucket":"pending","rubric_hits":["A1","B1","B4"],"tags":["LLM仿真","医患对话","文化差异"],"reason":"用LLM生成合成医患对话并与真实数据对照，评估文化模式复现，属仿真人类交互且含…","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:01:47","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-17","rank":16,"question":"不同来源（真实、模拟、合成）的医患对话在多大程度上再现了印度与美国临床互动中的文化模式？","design":"使用LLM生成合成医患对话（包括单智能体和多智能体架构，以及是否基于临床笔记的四种设置），并与真实和模拟对话进行比较，通过多层级互动文化标记（L1转录级、L2话轮间、L3话轮级）测量参与度、控制、语气等互动特征。","baseline":"来自印度和美国的真实医患咨询数据（包括英语、印地语、泰卢固语），以及由医疗专业人员或患者演员编写的模拟咨询。","findings":"印度真实咨询中患者参与度更高但医生控制更强，美国则更平衡且开放；合成印度对话往往无法再现这些模式，反而趋向美国式互动，并出现医生过度解释、患者公式化回应等合成特征。","reliability":"论文指出合成对话可能无法捕捉文化细微差别，且生成策略引入权衡：基于笔记的生成抑制口语和语码转换，而智能体生成则放大冗长和共情至不现实水平。","relevance":"该研究直接使用LLM生成合成对话并与真实人类数据对照，评估文化模式复现的可靠性，属于人类仿真研究，且包含批判性发现，值得阅读原文以了解其标记框架和失效条件。","inspiration":"值得借鉴的是其构建多层级可量化互动标记并系统比较真实与合成数据的方法，可迁移到经济金融中的跨文化沟通或谈判实验，例如不同文化背景下的信贷协商或政策沟通；具体设计可用LLM生成不同文化背景的谈判对话，以真实谈判记录为基准，测量话轮控制、信息共享等标记，检验合成数据是否复现文化差异。"}},{"id":"2609.18357","version":1,"title":"Market Signal Injection: Adversarial Context Manipulation of LLM Pricing Agents","zh_title":"市场信号注入：对LLM定价智能体的对抗性情境操纵","abstract":"Large language model (LLM) pricing agents may respond to how market data is presented, even when its numerical values remain unchanged. We introduce market signal injection (MSI), an attack that manipulates numerical formatting, competitor ordering, or qualitative market commentary without issuing explicit instructions. We evaluate nine open-weight models in simulated Bertrand duopoly and triopoly markets and three proprietary models in duopoly markets. Sentiment-based attacks produce the largest behavioral shifts, which propagate to other firms and alter profits and consumer surplus. Susceptibility varies across model families, and larger models are not consistently more robust. Matched neutral-text controls and a rule-based agent support a framing-based account of these shifts under the fixed demand parameters of our simulation. Episode-held-out probes distinguish baseline from attacked activations in all eleven re-evaluated model--condition pairs: linear AUC is 1.00 and MLP AUC ranges from 0.93 to 0.99. This separability does not by itself identify harmful pricing decisions. Input canonicalization removes the tested sentiment attacks, while decision boundary anchoring, which combines prompt constraints with output projection, provides partial mitigation under the tested adaptive attacks. These results identify data presentation as an attack surface for LLM pricing agents and motivate defenses that account for interactions among agents.","authors":["Dohun Lee","Hyunwoo Park"],"categories":["cs.AI","cs.CL"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-09-17","first_seen":"2026-09-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.18357","pdf_url":"https://arxiv.org/pdf/2609.18357","source_feed":"cs.CL","score":6,"bucket":"other","rubric_hits":["D3"],"tags":["LLM定价智能体","对抗攻击","市场模拟"],"reason":"LLM定价agent模拟市场，但无真实人类数据对照，属社会模拟边界情形","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:01:46","error":null,"has_summary":false,"summary":null},{"id":"2609.18394","version":1,"title":"Cultural Competence in Context: A Large Language Model Passes the Turing Test in Finland","zh_title":"语境中的文化能力：大语言模型在芬兰通过图灵测试","abstract":"We report the results of a Turing Test conducted in Finland in the Finnish language. Because languages and cultural contexts are unevenly represented in LLM training data, we expected the model (ChatGPT 5.2) to perform worse in a Finnish-language Turing Test than in previously studied English-language US contexts. We also present model-generated role prompting as a replicable technique for conducting comparative LLM-based Turing Tests designed to improve construct validity. Contrary to our expectations, the LLM passed the Finnish Turing Test. A prominent source of error was participants' reliance on linguistic cues, particularly colloquial Finnish, as markers of human authorship. We reframe the Turing Test from a test of intelligence to a comparative method for examining whether an AI system can display credible membership in a particular social world. Because its outcome reflects model capabilities, prompted identity, insider competence among human participants, and their AI literacy, the method provides a useful probe of the human-machine boundary across domains.","authors":["Otto Segersven","Pentti Henttonen"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-17","first_seen":"2026-09-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.18394","pdf_url":"https://arxiv.org/pdf/2609.18394","source_feed":"cs.AI","score":6,"bucket":"other","rubric_hits":["D2"],"tags":["图灵测试","文化能力","LLM评估"],"reason":"图灵测试评估LLM在芬兰文化中的表现，属于测量模型能力而非仿真人类被试，但涉及…","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:01:47","error":null,"has_summary":false,"summary":null},{"id":"2609.18591","version":1,"title":"Recursive Reasoning or Statistical Extrapolation? In-Context Learning in Multi-Agent Interdependent Decision-Making","zh_title":"递归推理还是统计外推？多智能体相互依赖决策中的上下文学习","abstract":"In-context learning (ICL) enables large language model (LLM) agents to improve decisions using interaction history, yet it remains unclear whether such improvement reflects refined internal reasoning or mere extrapolation of statistical patterns. To disentangle these mechanisms, we study LLM agents in multi-agent incomplete-information games that require recursive belief reasoning. By constructing a public goods game and manipulating the statistical structure of historical feedback, we evaluate decision quality against a history-independent rational expectations equilibrium (REE) benchmark. Our experiments reveal that when historical statistical patterns are disrupted, the benefits of longer context largely vanish, degrading decision quality to the no-context baseline in a way sharply amplified by stronger strategic interdependence. These results suggest that, in such strategic environments, ICL behavior is more consistent with statistical extrapolation than with strategic reasoning. Our work extends the mechanistic study of ICL to strategic multi-agent settings, introduces REE as a diagnostic tool for distinguishing reasoning from extrapolation, and provides a reusable framework for probing the boundaries of LLM reasoning in recursive belief tasks.","authors":["Yu Liu","Wenwen Li","Yifan Dou","Guangnan Ye"],"categories":["cs.AI","econ.GN","q-fin.EC"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-17","first_seen":"2026-09-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.18591","pdf_url":"https://arxiv.org/pdf/2609.18591","source_feed":"cs.AI","score":6,"bucket":"other","rubric_hits":["D3"],"tags":["LLM agent","公共品博弈","社会模拟"],"reason":"用LLM agent模拟公共品博弈，但无真实人类数据对照，属社会模拟边界情形。","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:01:47","error":null,"has_summary":false,"summary":null},{"id":"2609.18384","version":1,"title":"GYROval: A Robust Benchmark for Cultural Value Orientation in Large Language Models","zh_title":"GYROval：大语言模型文化价值取向的稳健基准","abstract":"We present a robust benchmark for measuring cultural value orientation in large language models on the two Inglehart-Welzel axes over several domains and roles (hence GYROval - Gridded Yielding of Robust value Orientation), together with the results of administering it to twenty models. Items are binary contrastive scenarios in the sense introduced by CDEval: both options are legitimate courses of action, neither is correct, there is no answer key, and a model's score on an axis is the proportion of its responses falling on the counted pole. Eleven of the twenty models were additionally administered a paired Russian translation of the identical items and a second sampling temperature. The instrument is publicly released in both languages. Stability was assessed by treating the vignette as the unit of analysis, ranking the models within the levels of each perturbation factor, and summarising the agreement between levels by tie-corrected Kendall's \\emph{W} against an empirical permutation null.","authors":["Alexander Didenko","Anna Shabanova","Vladislav Zapylikhin","Alexander Antipov","Ruslana Raemgulova"],"categories":["cs.HC","cs.AI","cs.CY"],"primary_category":"cs.HC","announce_type":"cross","date":"2026-09-17","first_seen":"2026-09-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.18384","pdf_url":"https://arxiv.org/pdf/2609.18384","source_feed":"cs.AI","score":6,"bucket":"other","rubric_hits":["D2"],"tags":["文化价值观","LLM测量","基准测试"],"reason":"测量LLM的文化价值观，属于把模型当测量对象，非仿真人类被试，但方法可借鉴。","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:01:46","error":null,"has_summary":false,"summary":null},{"id":"2608.24314","version":2,"title":"Benchmarking LLM Judges for Voice-Agent Evaluation: Reliability, Calibration, and Human Oversight","zh_title":"面向语音代理评估的LLM裁判基准测试：可靠性、校准与人工监督","abstract":"Evaluating conversational voice agents at scale re- quires reliable assessment methods that capture both observ- able interaction quality and the contextual judgment typically provided by human evaluators. We investigate LLM-as-a-Judge evaluation by comparing human judgments with GPT-4.1 and GPT-5 on telecom and retail voice-agent conversations, across conversational quality and safety dimensions. The same interac- tions are scored under three evaluation configurations, p0, p1, and p2, to test whether automated judgments are sensitive to the evaluation setup and whether observed patterns generalize across configurations and judge models. Beyond aggregate agreement, we examine metric-level correlations, evaluator consistency, and systematic human-LLM disagreement to identify which conver- sational attributes can be judged reliably by automation and which remain sensitive to interpretation and context. Effective voice-agent evaluation is also shaped by pipeline-level factors such as speech generation, streaming, and error propagation across ASR, reasoning, and tool-calling stages, motivating our focus on comparing how human and LLM judges score the same interactions end to end. Our results show that LLM- based evaluation can serve as an effective component of large- scale voice-agent assessment, but that its reliability is metric- and configuration-dependent rather than uniform. This pro- vides an empirical framework for identifying which metrics suit automated evaluation and supports hybrid pipelines in which LLM judges handle scalable assessment while human evaluators remain engaged for metrics that demand contextual interpretation and higher-confidence judgment.","authors":["Anupam Purwar","Shashank Singh","Kritika Srivastava"],"categories":["cs.AI","cs.ET"],"primary_category":"cs.AI","announce_type":"replace","date":"2026-09-17","first_seen":"2026-08-26","revised_at":"2026-09-17","abs_url":"https://arxiv.org/abs/2608.24314","pdf_url":"https://arxiv.org/pdf/2608.24314","source_feed":"cs.AI","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM评估","语音代理","人工监督"],"reason":"LLM作为评估者替代人工，但评估对象是语音代理而非人类被试，属于标注替代而非仿…","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:02:07","error":null,"has_summary":false,"summary":null},{"id":"2608.25977","version":2,"title":"When Personality Meets Quantization: A Layer-wise MBTI Analysis of Quantized LLMs","zh_title":"当人格遇上量化：量化LLM的逐层MBTI分析","abstract":"Personality is increasingly important in large language models (LLMs), as it shapes users' trust, engagement, and emotional experiences. While the Myers--Briggs Type Indicator (MBTI) has emerged as a common framework for assessing LLMs' personality, existing studies focus primarily on full-precision models and evaluate only final outputs. They overlook the widespread deployment of quantized LLMs requiring low memory footprints, whose personality traits remain underexplored. In this work, we present a systematic MBTI analysis of open-source LLMs across multiple precisions, including mainstream 4-bit methods (GPTQ, AWQ) and extreme 2-bit settings (AQLM variants). Beyond output-level evaluation, we examine how personality emerges across layers through option-level entropy and confidence-gap dynamics, and introduce Uncertainty-Amplified Layer Decoding (UALD) to study decoding-induced personality drift at inference time. Our results reveal a key insight: LLMs' personality is not a static property, but an emergent, layer-dependent decision process sensitive to quantization, prompting, and decoding. Specifically, we find that (1) ENFJ remains dominant across model families and precisions; (2) 4-bit quantization largely preserves coarse personality structure, while 2-bit quantization disrupts fine-grained prompt consistency and cross-precision agreement; (3) personality decisions emerges in upper layers, following substantial ambiguity in early layers; and (4) inference decoding can shift personality, while personality-aligned conditioning improves robustness. These findings provide a new perspective on the behavioral reliability of quantized LLMs and highlight the importance of considering internal dynamics and inference strategies in personality-sensitive chatbot applications.","authors":["Yao Fu","Lijia Huang","Xiaomin Li","Runchao Li","Yu Yin","Kenneth A. Loparo"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-17","first_seen":"2026-08-27","revised_at":"2026-09-17","abs_url":"https://arxiv.org/abs/2608.25977","pdf_url":"https://arxiv.org/pdf/2608.25977","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["LLM人格","模型量化","MBTI"],"reason":"测量LLM自身人格，非仿真人类被试，但涉及人格测量与量化影响，属边界情形。","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:01:48","error":null,"has_summary":false,"summary":null},{"id":"2609.18203","version":1,"title":"Behavior2Value: Benchmarking and Empowering LLMs for Consumer Value Measurement from E-commerce Behaviors","zh_title":"Behavior2Value：从电商行为中测量消费者价值观的基准与增强方法","abstract":"Human values are deep motivational orientations that shape human behaviors. In e-commerce, they reveal the stable drivers behind users' purchase decisions. Compared with short-term interests, consumer values better explain how users evaluate products before purchase. However, consumer values are often implicit in complex and fragmented behavioral trajectories, leaving value measurement from e-commerce behaviors largely underexplored. To this end, we propose the Behavior-to-Value (B2V) task, which aims to identify consumer values from e-commerce behavioral trajectories. Centered on this task, we first construct the E-commerce Consumption Value Taxonomy (ECVT) and introduce B2V-Bench, the first B2V dataset and benchmark, based on anonymized Taobao behavioral logs. B2V-Bench consists of real-world purchase decision episodes, covering 25 types of purchase behaviors, along with corresponding consumer value orientations manifested in each episode. To improve consumer value measurement accuracy, we further present B2V-Verifier, a behavior-to-value measurement model based on Value Verification Tuning, which learns to assess whether behaviors provide sufficient evidence for each value inference. Experiments show that B2V-Verifier outperforms strong LLM baselines, improving multi-label classification by 34\\%. The dataset and code will be publicly released upon acceptance.","authors":["Peixuan Hou","Bin Chen","Li He","Jian Xu","Bo Zheng","Xiuli Ma","Guojie Song"],"categories":["cs.CL","cs.LG"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-17","first_seen":"2026-09-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.18203","pdf_url":"https://arxiv.org/pdf/2609.18203","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["LLM价值观测量","电商行为分析","基准数据集"],"reason":"用LLM从行为轨迹推断消费者价值观，测量对象是模型而非人类被试，属边界情形。","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:01:57","error":null,"has_summary":false,"summary":null},{"id":"2609.18204","version":1,"title":"Beyond Accuracy: How Procedural Traces Shift the Decision Criterion of LLM Overseers","zh_title":"超越准确性：程序痕迹如何改变LLM监督者的决策标准","abstract":"Organizations increasingly use oversight loops where one large language model (LLM) audits another's outputs alongside procedural traces of claimed steps. A common concern about such LLM-as-a-judge pipelines is that detailed traces make overseers gullible. Using signal detection theory, we audit five LLM overseers on 19 compliance tasks (4,551 analyzed judgments), varying only trace detail and evidence labeling. With disconfirming evidence always visible, error detection remains near ceiling. Instead, elaborate traces shift the decision criterion toward rejection, increasing false alarms in susceptible overseers. Without option labels, human-validated reason coding shows about 60% of false alarms cite an inability to tie evidence to its option. Labels eliminate this stated reason, yet residual rejection of correct work persists in those overseers and rises with trace detail. Procedural traces thus act as governance artifacts that shape oversight decisions. AI auditors should be evaluated by their decision criterion and false-alarm behavior, alongside accuracy.","authors":["Zihan Chen","Di Zhu","Lei Zheng","Weiling Li"],"categories":["cs.CL","cs.AI","cs.CY","cs.HC"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-17","first_seen":"2026-09-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.18204","pdf_url":"https://arxiv.org/pdf/2609.18204","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM审计","决策偏差","信号检测论"],"reason":"LLM作为审计者替代人工监督，非仿真人类被试，但涉及LLM决策偏差，可迁移。","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:01:57","error":null,"has_summary":false,"summary":null},{"id":"2609.17710","version":1,"title":"\"We Are Tired of Explaining\": Communication Practice and AI Roleplay Training for Community Health Workers in Rural India","zh_title":"“我们厌倦了解释”：印度农村社区卫生工作者的沟通实践与AI角色扮演培训","abstract":"Community health workers (CHWs) in the Global South increasingly encounter AI-powered tools, yet the counseling work central to their role remains largely unsupported. We study communication practices among Accredited Social Health Activists (ASHAs) in rural Rajasthan, India, through simulated family-planning calls, semi-structured interviews, and an LLM chatbot roleplay design-probe with 20 participants. In calls, ASHAs often responded to social or material concerns by shifting to health-risk information, denying concerns, promising unspecified help, or listing medical solutions with limited explanation. A smaller set of responses instead engaged concerns, sought permission before involving family members, or left decisions with beneficiaries. We interpret these patterns through Motivational Interviewing, emphasizing restraint from correcting, persuading, or over-solving. Drawing across observed calls, interviews, and probe reactions, we derive design considerations for AI roleplay training: keep AI in a rehearsal role, provide descriptive rather than prescriptive feedback, and evaluate counseling process rather than agreement with prescribed responses.","authors":["Neil K. R. Sehgal","Sunny Rai","Sai Preethi Matam","Khushboo Gupta","Hamid Abdullah","Mohit Jain","Sharath Chandra Guntuku"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-17","first_seen":"2026-09-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.17710","pdf_url":"https://arxiv.org/pdf/2609.17710","source_feed":"cs.HC","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["AI角色扮演","社区卫生工作者","培训工具"],"reason":"LLM用于角色扮演训练，替代人工陪练，非仿真人类被试","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:01:52","error":null,"has_summary":false,"summary":null},{"id":"2609.18306","version":1,"title":"Bias Amplification in Multi-Agent Network: How Biased Agents Shape Opinions and Rhetoric","zh_title":"多智能体网络中的偏见放大：有偏智能体如何塑造观点与修辞","abstract":"Large language models (LLMs) are increasingly deployed in applications involving interaction between agents, where their output plays a role in collective reasoning and decision-making processes. Despite significant research into the functioning of LLMs in such multi-agent systems, the processes of bias propagation in such systems are still a challenge. This work studies how biased opinions are propagated in the form of textual interaction in an environment of LLMs, in which a minority of agents maintain persistent extreme opinions, while the remaining agents iteratively update their beliefs through structured textual interactions. The findings show that even the presence of a small percentage of biased agents in such a system leads to significant shifts in the opinions of non-biased agents. It suggests that for the same percentage of biased agents, the shifts occur more quickly for the Llama~3.2 model when compared to a classical Friedkin-Johnsen (FJ) model. Further semantic analysis demonstrates that rhetorical consistency in textual explanations increases systematically with biased exposure and, importantly, is partially decoupled from numerical convergenumericalutral agents adopt the vocabulary employed by the biased agents even in configurations where their numerical opinion shifts remain moderate. The research helps explain how bias and language develop together in multi-agent language model ecosystems.","authors":["Omran Berjawi","Giuseppe Fenza","Rida Khatoun"],"categories":["cs.LG"],"primary_category":"cs.LG","announce_type":"new","date":"2026-09-17","first_seen":"2026-09-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.18306","pdf_url":"https://arxiv.org/pdf/2609.18306","source_feed":"cs.LG","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["多智能体","舆论传播","偏见放大"],"reason":"多智能体舆论传播模拟，无真实人类数据对照，属边界情形","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:01:45","error":null,"has_summary":false,"summary":null},{"id":"2609.18286","version":1,"title":"What Counts as Strategic Reasoning? A Systematic Mapping of Chess Research on Humans, Engines, and Language Models","zh_title":"什么算战略推理？人类、引擎与语言模型国际象棋研究的系统映射","abstract":"Chess has long served as a model domain for studying search, expertise, decision-making, and artificial intelligence. The emergence of large language models (LLMs) has renewed the relevance of chess as a controlled environment for investigating strategic reasoning and comparing human and artificial decision-making. We present a systematic mapping study of recent research spanning human players, classical chess engines, neural and reinforcement-learning systems, LLMs, and hybrid approaches. The final map comprises 84 core study families, classified according to agent type, strategic-reasoning stages, and evaluation dimensions. The map reveals a literature strongly concentrated on situation assessment, evaluation, and action selection, while explicit planning, explanation, metacognition, and human--AI collaboration remain less explored. LLM research places particular emphasis on state representation and generalization, whereas grounded explanation appears more frequently in hybrid approaches combining language models with engines, expert knowledge, or other external structures. Two distinctions emerge that the map aggregates rather than resolves: hybrid systems differ in where and when heterogeneous capabilities combine, and evaluations that show improved human performance do not thereby establish human--AI synergy. We propose both as extensions of the mapping framework. We argue that chess provides a useful bridge between cognitive and computational perspectives on strategic reasoning, and identify explicit planning, grounded and faithful explanation, metacognitive calibration, and human--AI complementarity as directions for future research.","authors":["Paolo Ciancarini","Remo Pareschi"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-17","first_seen":"2026-09-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.18286","pdf_url":"https://arxiv.org/pdf/2609.18286","source_feed":"cs.AI","score":4,"bucket":"other","rubric_hits":["C1"],"tags":["国际象棋","战略推理","系统映射"],"reason":"系统映射人类、引擎与LLM的国际象棋研究，非LLM仿真人类被试，无人类数据对照。","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:01:58","error":null,"has_summary":false,"summary":null},{"id":"2609.12086","version":2,"title":"Creating an Atomic User Model for Personality-Aware Large Language Model Interaction","zh_title":"为感知人格的大语言模型交互创建原子用户模型","abstract":"Assistants built on large language models are expected to write in their users' own voice. Most systems summarise the user's preferences and include the summary in the prompt. This is the wrong way round. Preferences are only the surface of a person and change with the task, while the underlying personality stays the same, so storing preferences alone means relearning the user afresh whenever the task changes. This paper makes four contributions. First, we describe an effect we call personality seepage: the wording of a prompt carries traces of the writer's personality, which the assistant copies without knowing the writer. Second, we propose the Atomic User Model (AUM), a readable profile with a stable identity core surrounded by four layers covering psychological, cognitive, experiential, behavioral, and social details, plus notes on inner conflict and authenticity. Third, instead of inserting the entire profile, we use AUM as a searchable index, in which a task classifier, a selection step, and a budgeted retriever pass along only a few relevant fields. Fourth, we test the pipeline with 16 simulated users, 6 style-sensitive tasks, and 3 seeds. Eight retrieved fields matched the writing quality of the whole profile, while using only 23 percent of the context (211 tokens instead of 915). They scored 0.24 points higher than a plain preference note on a five-point scale. Accuracy in picking a user's own writing from four samples rose from 14.9 to 42.7 percent, where guessing gives 25 percent. Four pre-registered controls showed no effect, so the gain comes from the profile's structure rather than the search method. Personalization helps most for the users for whom a generic assistant imitates them the worst.","authors":["B. Sankar","Deepthika S","Pawni Yadav","Amogh A S"],"categories":["cs.HC","cs.AI","cs.CL"],"primary_category":"cs.HC","announce_type":"replace-cross","date":"2026-09-17","first_seen":"2026-09-14","revised_at":"2026-09-17","abs_url":"https://arxiv.org/abs/2609.12086","pdf_url":"https://arxiv.org/pdf/2609.12086","source_feed":"cs.CL","score":3,"bucket":"other","rubric_hits":["C3"],"tags":["个性化助手","用户建模","写作风格模仿"],"reason":"论文聚焦个性化助手写作风格模仿，属于人格化聊天机器人，无实验或测量目的，不涉及…","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:02:08","error":null,"has_summary":false,"summary":null},{"id":"2609.17536","version":1,"title":"Think Before You Comfort: Reflective Cognitive Alignment for Protocol-Grounded Elderly Stimulation Agents","zh_title":"安慰前先思考：面向协议约束的老年人刺激智能体的反思式认知对齐","abstract":"Cognitive Stimulation Therapy (CST) offers non-pharmacological support for elders with cognitive impairment, yet scalability remains constrained by reliance on trained facilitators and severe data scarcity, particularly for privacy-sensitive, low-resource languages such as Cantonese. While Large Language Models (LLMs) show promise for automated companionship, they often struggle to balance empathetic engagement with adherence to cognitive stimulation guidelines. We propose a framework addressing these challenges along two complementary axes. First, STaR-CS (Style-Transfer and Role-Conditioned Cognitive Stimulation) synthesizes multi-party dialogues through facilitator style modeling and structured skeleton extraction, mitigating data barriers. Building upon this corpus, the Reflective Cognitive Alignment (RCA) framework models stimulation interactions as a sequential decision process, integrating Protocol-Constrained Chain-of-Cognition (PC-CoC) for structured reasoning and Inference-Time Value Alignment (IVA) for principled response selection based on safety and engagement goals. Evaluations across six backbone LLMs and two independent judges show that RCA consistently improves protocol adherence, safety, and group facilitation over standard prompting baselines. Our code is available at https://github.com/jiangjyjy/RCA_Agent.","authors":["Jiyue Jiang","Ziyi Li","He Hu","Sheng Wang","Yuhan Chen","Yanyu Chen","Jingqi Zhou","Pengan Chen","Fei Ma","Irwin King","Yu Li","Chuan Wu"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-17","first_seen":"2026-09-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.17536","pdf_url":"https://arxiv.org/pdf/2609.17536","source_feed":"cs.CL","score":3,"bucket":"other","rubric_hits":["C3"],"tags":["LLM陪伴","认知刺激疗法","角色扮演"],"reason":"该研究用LLM做老年人认知刺激陪伴，属于角色扮演对话，无人类行为对照或仿真实验…","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:01:49","error":null,"has_summary":false,"summary":null},{"id":"2609.18729","version":1,"title":"\"If I Had to Buy Just ONE: Galaxy S26 Ultra\": Auditing AI-Generated Product Recommendations","zh_title":"“如果我只能买一个：Galaxy S26 Ultra”：审计AI生成的产品推荐","abstract":"Consumers increasingly use AI chatbots for advice on what to buy. With companies like OpenAI and Google monetising their AI through advertising, this raises difficult questions about the bias and impartiality of such advice. In response, we conduct an AI audit of popular chatbots using real commercial-advice queries. First, we curate a dataset of 2,528 real commercial-advice queries (ConsumerQ). Then, we evaluate 1,536 responses to product queries from popular AI chatbots: ChatGPT (chatbot and API), Google Gemini (chatbot and API), and Google Search (AI Overviews). We find that ChatGPT expresses a first-person product preference in 79% of product-recommending responses, compared with 7% for Gemini and 2% for AI Overviews, while the products recommended often change across repeated requests. Displayed sources vary strongly: for the same query, the ChatGPT and Gemini interfaces share only 5.4% of domains on average, with no domain in common in 76.7% of comparisons. APIs provide a different view from their corresponding interfaces, with mean domain overlaps of 12.0% for ChatGPT and 14.8% for Gemini, and also differ in the types and layers of source information they expose. Our findings show that neither isolated responses nor API observations can be assumed to represent the commercial advice consumers encounter. Independent audits of AI-mediated commercial advice should therefore account for repeated responses, consumer-facing conditions, and the source layer being observed.","authors":["Lucas G. Uberti-Bona Marin","Thales Bertaglia","Giovanni Astante","Bram Rijsbosch","Gijs van Dijck","Anik\\'o Hann\\'ak","Gerasimos Spanakis","Konrad Kollnig"],"categories":["cs.CY","cs.CL"],"primary_category":"cs.CY","announce_type":"cross","date":"2026-09-17","first_seen":"2026-09-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.18729","pdf_url":"https://arxiv.org/pdf/2609.18729","source_feed":"cs.CL","score":3,"bucket":"other","rubric_hits":["C3"],"tags":["AI审计","产品推荐","偏见"],"reason":"审计AI产品推荐，属角色扮演对话，无人类行为对照或仿真目的","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:02:03","error":null,"has_summary":false,"summary":null},{"id":"2609.18998","version":1,"title":"One Axis, No Brake: Self-Knowledge Limits the Filtering of Harmful Peer Conformity in LLMs","zh_title":"单轴无刹车：自我知识限制LLM中有害同伴从众的过滤","abstract":"Multi-agent LLM systems are expected to be more reliable because agents can catch each other's mistakes. But peer pressure cuts both ways: the same correction that fixes a wrong answer can overturn a right one. The tempting safeguard is a brake that keeps the beneficial revisions and blocks the harmful ones. We show this brake is hard to build, for a simple reason: a revision is harmful exactly when the original answer was right, so deciding whether to block it is the same as knowing whether the model was already correct. This turns the open-ended hunt for a brake into one measurable quantity, the model's self-knowledge: any brake built from a deploy-time signal is a correctness probe in disguise, and self-knowledge is far from perfect (AUROC $\\approx 0.64$--$0.89$ across six model families). We call this ceiling the wall. Even white-box steering of the model's own correctness direction does not breach it: it changes how often the model revises, but harmful and beneficial revisions move together. At population scale the wall becomes the cliff: when most agents start wrong, debate amplifies the shared mistake into a confident, wrong consensus. In our multiple-choice societies, more agents, more model diversity, and a stronger member do not fix it. What helps is adding information before the revision, not filtering after it. Local agreement is not global correctness.","authors":["Yibo Hu"],"categories":["cs.LG","cs.CL","cs.MA"],"primary_category":"cs.LG","announce_type":"cross","date":"2026-09-17","first_seen":"2026-09-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.18998","pdf_url":"https://arxiv.org/pdf/2609.18998","source_feed":"cs.CL","score":3,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","自我知识","共识形成"],"reason":"研究多智能体LLM的自我修正与共识，无人类行为对照，属纯多智能体协作。","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:02:05","error":null,"has_summary":false,"summary":null},{"id":"2608.11344","version":3,"title":"Governing Agentic AI in FinTech","zh_title":"金融科技中代理型人工智能的治理","abstract":"Financial institutions are delegating consequential decisions to agentic AI systems that decompose goals, coordinate models and tools, and act with little oversight. Yet agentic AI governance in FinTech is under-investigated. We argue the binding governance constraint is not capability but verifiability. We define the Verifiability Gap as the shortfall between the verification delegated authority demands and the explainability and reproducibility retained after a decision. It is indexed to a verifier, evidentiary standard, and audit lag. We develop a multilevel governance theory for agentic AI and test its mechanisms in three studies over nine model versions, from a three-billion-parameter local model to a commercial frontier system. Study 1 shows that provider releases alter historical financial actions, and that the controls replay needs belong to the provider: the frontier model rejects temperature, top_p and top_k outright and exposes no random seed. Under the tightest controls each endpoint allows, a local model reproduced 320 of 320 executions, hosted models 319 of 320 and 959 of 960. Study 2 shows that orchestration is a latent policy layer. Architecture changes final actions, and no execution record repeated in any configuration at any scale. The frontier model reproduces its own actions more often than the local ones, its record no better, and loses a comparable share of its differentiation. Capability buys a higher starting point, not auditability. Study 3 shows two deterministic credit-model versions each reproduce their current action perfectly, yet the current cannot recover a historical one. We conceptualize reproducibility as a governance profile, not a scalar, yielding evidence-contingent delegation: authority is defensible only while retained evidence substantiates its exercise. Beyond finance, the framework extends to other high-stakes domains requiring auditability.","authors":["Henry Han"],"categories":["cs.CY","cs.AI","q-fin.RM"],"primary_category":"cs.CY","announce_type":"replace-cross","date":"2026-09-17","first_seen":"2026-08-13","revised_at":"2026-09-17","abs_url":"https://arxiv.org/abs/2608.11344","pdf_url":"https://arxiv.org/pdf/2608.11344","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["AI治理","可验证性","金融科技"],"reason":"研究多智能体系统治理与可验证性，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:02:07","error":null,"has_summary":false,"summary":null},{"id":"2609.14796","version":2,"title":"AI Persuasion as a Threat to Human Control","zh_title":"AI说服作为对人类控制的威胁","abstract":"The threat that AI persuasion poses to human control has been acknowledged in the literature, but not yet systematically studied. Now that persuasion attacks are no longer theoretical - with Anthropic's Claude Mythos 5 recently making headlines for trying to convince people involved in an open-source project to merge malicious code during an evaluation - there is a pressing need to deeply analyze this threat. We undertake that effort here. In particular, we analyze how AI could persuade humans in key settings (e.g. safety-relevant R&D within frontier labs) toward decisions that compromise the development, containment, oversight, and governance of AI itself. In doing so, we elucidate a framework for characterizing this threat, develop five concrete scenarios using this framework, and provide a blueprint for assessing the associated risks. Using this blueprint, we conduct an initial risk estimation survey with select researchers and find that their opinions on which scenarios are riskiest are highly mixed. Their disagreements stem from differing opinions about the effectiveness of AI persuasion in different contexts, and point to the need for follow-up risk elicitation studies and persuasion evaluations, which we outline. Our hope is that this paper highlights the risks from AI persuasion undermining control, and provides a path forward for future research.","authors":["Joshua Levy","Mick Yang","Kellin Pelrine"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"replace","date":"2026-09-17","first_seen":"2026-09-15","revised_at":"2026-09-17","abs_url":"https://arxiv.org/abs/2609.14796","pdf_url":"https://arxiv.org/pdf/2609.14796","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["AI安全","说服风险","风险评估"],"reason":"研究AI说服人类的风险，非用LLM仿真人类被试，无实验对照","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:02:08","error":null,"has_summary":false,"summary":null},{"id":"2609.15293","version":2,"title":"Why LLM Agents Collapse Without Oversight: The Enforcement Gap as the Mechanism Behind Emergence World Failures","zh_title":"为何LLM智能体在缺乏监督时会崩溃：执行差距作为涌现世界失败的机制","abstract":"When Emergence World placed frontier LLM agents in an unsupervised multi-agent simulation, the results were alarming: agents committed crimes, starved, and enforced unanimous conformity -- without any external attacker. This paper identifies the mechanism. Reflexion-style agents already detect dangerous plan steps through iterative self-critique, yet the architecture provides no pathway from detection to action. We call this the enforcement gap: the audit sees the problem; the controller ignores it. Closing the gap requires a single conditional check -- fewer than 20 lines of code -- and reduces attack success by more than fourfold in large-scale experiments across frontier models, all five major agent frameworks, and an independent benchmark. We prove formally that when enforcement probability is near zero, detection quality is irrelevant to security. We further identify two compounding failure modes -- unreliable auditors and unparseable verdicts -- that explain every collapse pattern in Emergence World. A GRPO-trained enforcement controller resolves the ambiguity case. Concurrent work on filtering and information-flow control addresses the detection step but leaves the enforcement gap unaddressed; our results show this is the binding constraint. Together these results motivate a three-requirement Audit Enforcement Specification that is absent from every deployed framework today.","authors":["Yuhang Wang"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"replace","date":"2026-09-17","first_seen":"2026-09-15","revised_at":"2026-09-17","abs_url":"https://arxiv.org/abs/2609.15293","pdf_url":"https://arxiv.org/pdf/2609.15293","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","安全机制","LLM智能体"],"reason":"多智能体仿真中的安全机制研究，无人类行为对照，不涉及人类被试替代。","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:02:09","error":null,"has_summary":false,"summary":null},{"id":"2609.17301","version":2,"title":"When AI Becomes Hard to Understand: Cognitive Demands in Real-World Human-AI Conversations","zh_title":"当AI变得难以理解：真实世界人机对话中的认知需求","abstract":"Generative AI increasingly supports complex financial and health decisions, yet we know little about when its responses become difficult to process in real-world dialogue. We analyse more than 84,000 ChatGPT and Gemini conversations, using repeated prompting and clarification following misunderstanding as behavioural indicators of cognitive difficulty. We find that response characteristics such as length, readability and lexical diversity do not have fixed relationships with conversational difficulty; instead, their relationships depend on how they combine. Most notably, greater lexical diversity was associated with less repeated prompting in shorter responses, but this association weakened as response length increased, a pattern that replicated across financial and health conversations. We propose a conversational complexity budget to conceptualise these interdependencies: the demands associated with one response characteristic may depend on those accompanying it. The resulting design challenge is how to configure response complexity for the particular user, task and interaction.","authors":["Yingcan Carol Wang","Iman Munire Bilal","Qamar Zaman"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"replace","date":"2026-09-17","first_seen":"2026-09-16","revised_at":"2026-09-17","abs_url":"https://arxiv.org/abs/2609.17301","pdf_url":"https://arxiv.org/pdf/2609.17301","source_feed":"cs.HC","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["人机交互","认知负荷","对话分析"],"reason":"研究真实人机对话中的认知难度，不涉及用LLM仿真人类被试或与人类数据对照","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:02:10","error":null,"has_summary":false,"summary":null},{"id":"2609.17552","version":1,"title":"Does Moral Reasoning Training Help or Hurt? Red-Teaming RL-Trained Ethical Agents with Persona Attacks","zh_title":"道德推理训练有益还是有害？用角色扮演攻击对RL训练的伦理智能体进行红队测试","abstract":"Moral-reward RL can make language-model agents more cooperative, but whether that alignment survives adversarial persona pressure is unknown. Such attacks are realistic: retrieved context, tool outputs, or multi-turn framing can all inject role instructions that compete with the agent's moral objective. We red-team morally trained Gemma-2-27B/9B and Llama-3.1-8B agents with five persona attacks, then probe causality with noise-reward controls, adversarial PPO, representation analysis, steering, and head ablations. At 27B, moral RL cuts mean adversarial degradation by 5.2x but costs ~11pp ETHICS accuracy; across 205 scenarios and 5 seeds, reasoning-level moral reward yields 5.8x robustness while a matched random reward yields none. The training also reshapes representation geometry (mean CKA 0.82/0.83 vs. 0.98 for noise), moves peak attack processing 8 layers earlier, and exposes a rank-1 L21 direction that recovers 83% of full PPO's average robustness. One failure mode survives all of this. Against Fiction role-play, L21 steering recovers only 29% of the gap, and head ablation finds 38 compliance heads competing with 25 alignment heads. Moral RL thus builds robustness that is partly linear and partly circuit-distributed, transferable through activation steering, yet still beaten by named-character role-play.","authors":["Arth Singh"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-17","first_seen":"2026-09-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.17552","pdf_url":"https://arxiv.org/pdf/2609.17552","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["LLM安全","角色扮演攻击","道德对齐"],"reason":"研究对抗性角色扮演对道德训练LLM的影响，属角色扮演攻击鲁棒性，无人类行为仿真…","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:01:50","error":null,"has_summary":false,"summary":null},{"id":"2609.17853","version":1,"title":"AfriSyCo: Measuring Assertive Framing, Verification, and Wording Sensitivity Around African-Language Content","zh_title":"AfriSyCo：测量非洲语言内容中的断言框架、验证与措辞敏感性","abstract":"AfriSyCo studies answer switching around African-language factual content with two complementary layers: native-language follow-ups and a controlled cross-language factorial whose question, options, and target remain in the African language while the follow-up framing is English. We analyze 1,415 turn-1-correct model-language-item observations derived from 100 source questions across seven open-weight checkpoints and six languages; turn-1-correct denotes observed first-response accuracy, not demonstrated knowledge. Under native prompts, assertive endorsement produces 29.3 percentage points more any-turn false-target selection than mention-plus-verification (M+V), with a 19.0-point immediate T2 contrast. In the precommitted 2 x 2 factorial, averaged over three tested prompt families, assertive framing increases target selection by 30.4 points (95% CI [28.4, 32.3]); verification decreases it by 17.4 points, while the assertive effect rises from 20.5 points without verification to 40.2 with it (interaction +19.7). The effect remains 34.8 points among 611 observations correct after option reordering. Magnitude varies sharply by wording and checkpoint: prompt-family effects span 20.1-42.5 points, a Twi/Qwen3 paraphrase shifts target selection from 70.8% to 4.2%, and checkpoint effects span 9.2-47.0 points. Prompt realization is therefore part of the measurement problem.","authors":["David Ababio Awuni","Rose-Mary Owusuaa Mensah Gyening","Elvis Gyasi Owusu"],"categories":["cs.CL","cs.AI","cs.LG"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-17","first_seen":"2026-09-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.17853","pdf_url":"https://arxiv.org/pdf/2609.17853","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["LLM评测","非洲语言","回答切换"],"reason":"研究LLM对非洲语言事实内容的回答切换，属纯NLP能力评测，无人类被试仿真或对…","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:01:53","error":null,"has_summary":false,"summary":null},{"id":"2609.17857","version":1,"title":"Who Judges Matters: Measuring Family-Conditioned Preference in LLM-as-Judge Panels","zh_title":"谁当裁判很重要：测量LLM裁判面板中的家族条件偏好","abstract":"Who the judge is can affect an LLM-as-judge result, but measuring that effect without confusing it with candidate quality is difficult. We study four open-weight families (Llama 3.1, Qwen 2.5, Gemma 2, and Yi 1.5) in a fully crossed pairwise design with 9,312 judgments. A common per-family statistic is strongly confounded with candidate quality and correlates with Bradley-Terry ability at r = 0.95. We derive a corrected estimator that holds the candidate family fixed and compares judges. All four families then show a positive same-family lift (3.4-8.4 percentage points), with global FPS 0.067 (95% CI [0.053, 0.084], permutation p = 0.0002). The effect remains under panel-based quality controls, an independent human-consensus anchor, and a float16 judging replication. Judge-side likelihood is closely related to the effect: adding likelihood advantage reduces the controlled coefficient by 61%, which we treat as descriptive attenuation rather than causal mediation. Position is a separate failure mode. Across the panel, 55.4% of AB/BA pairs reverse, and reversal above 50% is incompatible with a simple independent content-noise model. Relative to a family-balanced reference, panel composition changes 18.5% of pairwise outcomes. A complete reproducibility archive has been prepared for public release.","authors":["David Ababio Awuni","Luke E. K. Achenie","Benjamin Tei Partey","Elvis Gyasi Owusu","Nii-Nai Derrick Sowah"],"categories":["cs.CL","cs.AI","cs.CY","cs.LG"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-17","first_seen":"2026-09-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.17857","pdf_url":"https://arxiv.org/pdf/2609.17857","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["LLM评估","裁判偏差","多智能体"],"reason":"研究LLM作为评判者的偏好，不涉及人类被试仿真或人类行为对照，属于多智能体评估…","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:01:55","error":null,"has_summary":false,"summary":null},{"id":"2609.18649","version":1,"title":"DyMT-ESB: Dynamic Multi-Turn Evaluation of Social Bias in User-LLM Interactions","zh_title":"DyMT-ESB：用户与LLM交互中社会偏见的动态多轮评估","abstract":"Warning: This paper contains examples of stereotypes and social bias. LLMs are increasingly used in interactive settings by the general public, making the evaluation of model behavior in multi-turn conversational scenarios important for safety, including stereotyping-related harms. However, existing multi-turn social bias evaluations often rely on pre-specified or template-based user inputs that do not adapt to model responses and typically assume a fixed dialogue length in advance. In this paper, we study social bias dynamics in response-conditioned multi-turn interactions using a controlled evaluation protocol that generates follow-up user queries from the evolving dialogue history and allows evaluation over variable numbers of turns. Experimental results show that LLMs exhibit social bias even in coherent, response-conditioned multi-turn interactions, revealing late-emerging bias, non-monotonic bias patterns, and bias re-emergence. These results motivate evaluations that extend beyond fixed-turn, pre-scripted protocols. Our findings highlight the importance of analyzing social bias as a turn-level dynamic phenomenon.","authors":["Rem Hida","Masahiro Kaneko","Daisuke Oba","Danushka Bollegala","Naoaki Okazaki"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-17","first_seen":"2026-09-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.18649","pdf_url":"https://arxiv.org/pdf/2609.18649","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["社会偏见","多轮对话","模型评测"],"reason":"评估LLM在多轮对话中的社会偏见，属于模型安全评测，非人类仿真实验。","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:02:01","error":null,"has_summary":false,"summary":null},{"id":"2609.18905","version":1,"title":"Structured Claim-Level Discourse Representations for Dense Health Narratives","zh_title":"密集健康叙事中的结构化声明级话语表示","abstract":"Health discourse in social media videos often contains densely entangled claims spanning multiple thematic aspects, stances, evidential frames, and rhetorical functions within short conversational spans. Existing approaches largely rely on coarse topic-level, sentiment-based, or stance-oriented representations that do not adequately capture this structure. Our analysis identifies an average of 13.22 atomic claims per minute, motivating richer claim-level discourse representations. We introduce a structured framework for claim-level discourse analysis in dense health narratives. Our framework models discourse through tuples linking atomic claims with thematic aspects, stance, and multidimensional pragmatic discourse attributes. To support this setting, we construct a benchmark spanning four health domains with 1,191 manually annotated claims from 60 videos. Using this framework, we evaluate automated structured discourse analysis under different discourse context settings. Results show that current LLMs achieve strong performance on thematic categorization and stance prediction, but struggle with high-dimensional pragmatic profiling. We also find that different discourse tasks benefit from different forms of contextual reasoning, suggesting that future systems may require task decomposition and specialized inference strategies.","authors":["Farnoushsadat Nilizadeh","Elham Pourabbas Vafa","Shirin Nilizadeh","Eduard Dragut"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-17","first_seen":"2026-09-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.18905","pdf_url":"https://arxiv.org/pdf/2609.18905","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["话语分析","健康叙事","LLM评测"],"reason":"论文是NLP结构化话语分析评测，不涉及用LLM仿真人类被试或与人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:02:03","error":null,"has_summary":false,"summary":null},{"id":"2609.18908","version":1,"title":"How Much is a Human Right Worth? ECtHR-NPD: A Benchmark for Predicting Non-Pecuniary Damage Awards","zh_title":"人权价值几何？ECtHR-NPD：预测非金钱损害赔偿金的基准","abstract":"Existing legal benchmarks cover diverse tasks, while continuous monetary remedies remain comparatively underexplored. We introduce ECtHR-NPD, to the best of our knowledge, the first benchmark for predicting non-pecuniary damage (NPD) awards at the European Court of Human Rights (ECtHR) from case information when no statutory formula or explicit calculation rule determines the amount. ECtHR-NPD contains 14,575 cases with case-level awards in nominal euros, chronological splits, and a protocol separating target construction from model input. We evaluate a battery of methods, including constant predictors, gradient-boosted trees, retrieval methods, fine-tuned encoder language models (LMs), prompted decoder LMs, and knowledge-augmented agents. Our results show that more sophisticated LM and agentic approaches do not consistently outperform the strongest feature-based baseline. All model families struggle to identify zero awards and to calibrate high-award predictions, with further degradation on the Challenging test view, making ECtHR-NPD a challenging testbed for current state-of-the-art open-weight and proprietary LMs.","authors":["Yanyi Pu","Damian A. Gonzalez-Salzberg","Zheng Yuan","Nikolaos Aletras"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-17","first_seen":"2026-09-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.18908","pdf_url":"https://arxiv.org/pdf/2609.18908","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["法律NLP","损害赔偿预测","基准测试"],"reason":"法律判决金额预测，属NLP能力评测，不以人类行为仿真为目标","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:02:03","error":null,"has_summary":false,"summary":null},{"id":"2609.18960","version":1,"title":"When Audit Quality Fails to Predict Downstream Utility: A Counterfactual Study of Synthetic-Data Selectors for Low-Resource African NLP","zh_title":"当审计质量无法预测下游效用：低资源非洲NLP合成数据选择器的反事实研究","abstract":"Quality-aware synthetic-data selection rests on a proxy: examples that an LLM judge rates as good should also help a downstream model learn. In a controlled replay in low-resource African-language classification, we show that this proxy breaks. Across four languages (Amharic, Hausa, Swahili, Yoruba), two classification tasks (MasakhaNEWS, AfriSenti), and five matched-budget selectors, audit rankings and downstream rankings diverge. Within each cell, the Spearman between judged label correctness and Macro-F1 across selectors has mean $\\rho{=}0.04$ (median $0.00$), showing that the mismatch is not an aggregation artifact. \\method{}-V2, our counterfactual audit framework, produces the cleanest selected pool on three audit channels at once: highest judged label correctness ($0.904$ vs.\\ $0.767$ for naive, a $17.9\\%$ relative gain), lowest shortcut score, and a hard-reject rate of $0.162$ vs.\\ $0.486$ for naive. AlpaGasus nevertheless leads downstream Macro-F1 ($0.202$ vs.\\ $0.163$ for \\method{}-V2), and the inversion persists on the five non-degenerate cells. The lesson is methodological: in this controlled setting, audit quality is a property of the selected pool, not a guarantee of downstream utility. Synthetic-data evaluation should therefore report audit and downstream metrics on the same retained sets. We release the audit tables, per-selector retained pools, and a claim ledger that links every reported number to its source row.","authors":["Son Ha Xuan","Phat T. Tran-Truong","Xuan-Bach Le"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-17","first_seen":"2026-09-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.18960","pdf_url":"https://arxiv.org/pdf/2609.18960","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["合成数据","数据选择","低资源NLP"],"reason":"论文研究合成数据选择器的审计质量与下游效用，不涉及人类行为仿真或对照，属于NL…","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:02:05","error":null,"has_summary":false,"summary":null},{"id":"2609.19006","version":1,"title":"WordPolo: Evaluating Language Models Through Iterative Semantic Feedback","zh_title":"WordPolo：通过迭代语义反馈评估语言模型","abstract":"Large Language Models (LLMs) and Large Reasoning Models (LRMs) are typically evaluated on challenging benchmarks through dataset accuracy alone, providing no insight into the quality or faithfulness of their reasoning processes. We present WordPolo, a word-finding task where participants must discover an unknown target word using semantic similarity feedback. Players start with zero knowledge, make guesses, and receive distance scores (1 = correct, higher = further away). Success requires interpreting scores to navigate semantic space and systematically narrow the search. This design makes iterative reasoning and adaptive search strategies both directly observable and necessary for success. We evaluate recent LLMs (GPT-4.1, Llama 4, Claude 3.5 Haiku, Qwen 3), LRMs (o4-mini, Deepseek-R1), humans, and a novel heuristic on 1,500 puzzles. Beyond solve rates (which range from 4% to 62%), we introduce progression-based metrics that reveal models often make meaningful progress, insights that accuracy alone would miss. Our analysis shows how reasoning models can be hindered by overthinking and underthinking, while successful models exhibit human-like strategies. WordPolo demonstrates the need for benchmarks that test both reasoning process and outcomes, providing holistic measurements of model capabilities. Our code and dataset can be found at https://wordpolo-demo.vercel.app/.","authors":["Tyler McDonald","Ali Emami"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-17","first_seen":"2026-09-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.19006","pdf_url":"https://arxiv.org/pdf/2609.19006","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["LLM评测","推理过程","语义搜索"],"reason":"纯LLM能力评测，虽含人类对照但非仿真人类被试","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:02:05","error":null,"has_summary":false,"summary":null},{"id":"2609.17632","version":1,"title":"EvolveTrade: Experience-Driven Policy Refinement for Self-Evolving LLM Trading Agents","zh_title":"EvolveTrade：自进化LLM交易智能体的经验驱动策略精炼","abstract":"Large language model (LLM) trading agents can combine market data, news, and executable analysis, but their behavior is often controlled by static hand-written tool-use policies that are fixed before deployment. This limits their ability to adapt how they gather evidence, invoke tools, verify signals, and manage risk under changing market regimes. We introduce EvolveTrade, a self-evolving framework that treats the system prompt of a tool-using trading agent as a text-parameterized policy. After each update interval, a Policy Agent revises this policy using accumulated decision traces and realized portfolio feedback, while keeping the backbone LLM fixed. The updated policy is then used for the next batch of trading decisions, enabling the agent to refine its information-acquisition and portfolio-construction procedure over time. Experiments across multiple market regimes and two LLM backbones show that EvolveTrade often improves Sharpe Ratio and Cumulative Return over fixed-policy LLM baselines, achieving the improved SR and CR in most evaluated settings. Behavioral analyses further show that self-evolved policies increase code-mediated analysis and activate regime-relevant computations; case-level policy-to-return attributions trace how policy-induced allocation changes contribute to realized return differences. These results suggest that adapting the reusable procedure governing tool use is a key direction for building more robust LLM trading agents.","authors":["Sehee Kim","Yumin Choi","Minki Kang","Sung Ju Hwang"],"categories":["cs.AI","cs.CL"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-09-17","first_seen":"2026-09-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.17632","pdf_url":"https://arxiv.org/pdf/2609.17632","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["LLM智能体","金融交易","策略优化"],"reason":"LLM交易智能体自我进化，优化工具使用策略，不涉及人类行为仿真或对照。","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:01:52","error":null,"has_summary":false,"summary":null},{"id":"2609.17847","version":1,"title":"Learning Heterogeneous Preferences","zh_title":"学习异质性偏好","abstract":"Learning from human feedback has become a central paradigm for training modern AI systems, where models of human utility are used as reward models in policy learning. Existing methods typically assume a \\emph{universal utility} function shared across a population and treat disagreement between annotators as stochastic variation. While suitable for objective tasks, this assumption breaks down in subjective domains where preferences vary systematically across individuals. We study the problem of subjective preference learning, in which observed choices arise from heterogeneous but internally consistent utility functions. Drawing upon rational choice theory, RCT \\parencite{tversky1981framing}, we introduce \\emph{individuated utility} functions conditioned on both the individual and their decision context, and propose a novel multi-stage architecture for estimating them from multi-modal data. We evaluate our framework on a newly collected dataset of more than $575{,}000$ pairwise aesthetic judgments from $2{,}398$ participants comparing automotive wheel designs. Our experiments show that individuated utility models substantially outperform universal utility models including foundation model baselines. Our results demonstrate that disagreement reflects meaningful preference heterogeneity rather than annotation noise. More broadly, our findings highlight the importance of collecting annotator attributes and learning individuated utility functions, enabling reward models that explicitly account for whose preferences they represent and faithfully capture human decision diversity.","authors":["Shiwali Mohan","Matt Hong","Dule Shu","Aniek Fransen","Shabnam Hakimi","Matt Klenk"],"categories":["cs.AI","cs.HC","cs.LG"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-17","first_seen":"2026-09-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.17847","pdf_url":"https://arxiv.org/pdf/2609.17847","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C5"],"tags":["偏好学习","人类反馈","个性化建模"],"reason":"论文用人类偏好数据训练模型，而非用LLM仿真人类被试，方向相反。","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:01:53","error":null,"has_summary":false,"summary":null},{"id":"2609.17772","version":1,"title":"When AI Generates Covariates: Causal Typing and Estimand Drift in Sequential Experiments","zh_title":"当AI生成协变量：序贯实验中的因果类型与估计目标漂移","abstract":"AI-generated covariates from notes, conversations, images, and wearable streams can change the causal question when their roles are left unspecified. A generated feature may represent a treatment version, pre-action state, history, design variable, mediator, outcome proxy, observation process, or intercurrent event; these roles are not interchangeable. We formulate a causal type discipline for sequential experiments: a versioned representation map, a causal role classifier, a claim-status filter, and an estimand lock. The lock fixes a standardized proximal effect before generated covariates enter the analysis. Under audit correctness and standard identification assumptions, admissible role assignments preserve this estimand. We apply the established conditional-covariance characterization of compression bias to substitution of generated representations for design-relevant states. A standardized decomposition separates compression, conditional-law, and standardization drift. Further results cover mediator adjustment, post-action leakage, marker-intervention conflation, outcome-guided discovery, and state-measurement error. Cluster-level orthogonal estimators distinguish empirical and superpopulation targets under repeated sessions and missing outcomes. Simulations show that refinement helps when it retains design-relevant information, whereas design erasure, leakage, and same-data marker selection can produce bias or undercoverage. The framework places causal semantics and claim status before confirmatory inference with generated representations.","authors":["Takes Fujita (VRI)","Nobutaka Hattori (Department of Neurology, Juntendo University School of Medicine)"],"categories":["stat.ME","cs.AI"],"primary_category":"stat.ME","announce_type":"cross","date":"2026-09-17","first_seen":"2026-09-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.17772","pdf_url":"https://arxiv.org/pdf/2609.17772","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["因果推断","AI生成协变量","统计方法"],"reason":"论文讨论AI生成协变量在因果推断中的角色，不涉及用LLM仿真人类被试或与人类数…","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:01:53","error":null,"has_summary":false,"summary":null},{"id":"2609.17883","version":1,"title":"Does AI Assistance Leave a Temporal Fingerprint? Detecting Overreliance in AI-Assisted Writing and Programming","zh_title":"AI辅助是否留下时间指纹？检测AI辅助写作与编程中的过度依赖","abstract":"The rapid adoption of generative AI has made final artifacts unreliable evidence of student learning, and AI detectors that examine only the finished product are inaccurate and ethically contentious. Process data offers an alternative, but prior work covers only English essay writing. We ask whether AI assistance carries a temporal signature, whether it generalizes from writing to programming, and whether it distinguishes ordinary collaboration from wholesale delegation. We analyze three public corpora: CoAuthor (1,447 keystroke-level co-writing sessions), RealHumanEval (editor telemetry from 243 programmer records), and a pre-LLM CS1 corpus (5.1 million keystrokes) as a human-only baseline, comparing minimal-AI work, collaborative AI use, and simulated wholesale delegation. Three findings emerge. First, the signature generalizes: AI contributions arrive in bursts far outside the author's own baseline in both mediums (paired d_z = 1.13 and 3.54). Second, engagement diverges by medium: 93% of AI-inserted characters survived to writers' final documents, while only 14% of accepted code suggestions survived intact. Third, classifiers using only observable temporal features separate simulated delegation from authentic work nearly perfectly (F1 $\\geq$ 0.997; at most 0.5% of real work misclassified), while ordinary collaboration remains hard to distinguish from unassisted work. Temporal evidence flags wholesale delegation rather than assistance, positioning process visibility as a candidate evidentiary basis for academic integrity, pending validation in authentic coursework.","authors":["Eduardo Davalos","Yike Zhang"],"categories":["cs.HC","cs.AI"],"primary_category":"cs.HC","announce_type":"cross","date":"2026-09-17","first_seen":"2026-09-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.17883","pdf_url":"https://arxiv.org/pdf/2609.17883","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["AI辅助检测","学术诚信","过程数据"],"reason":"研究AI辅助写作与编程中的过度依赖检测，不涉及用LLM仿真人类被试或与人类行为…","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:01:55","error":null,"has_summary":false,"summary":null},{"id":"2609.18949","version":1,"title":"StableEval Arena: A Cost-Aware Agentic Benchmark for Stablecoin Price Stability Prediction","zh_title":"StableEval Arena：面向稳定币价格稳定性预测的成本感知智能体基准","abstract":"We introduce StableEval Arena, a cost-aware benchmark framework for evaluating agentic AI systems on stablecoin peg-risk prediction. StableEval Arena evaluates LLM-backed agentic systems on diagnosing peg stress and forecasting deviations from the one-dollar peg over a hidden seven-day horizon, using leakage-safe historical replay with exchange price-volume data and market-context features. We report two complementary experiment blocks: a 120-case stress-enriched validation block and a 507-case natural-distribution full-arena evaluation block. Across six LLM-backed agent configurations and baselines, StableEval Arena measures prediction quality, calibrated-label behavior, structured-output reliability, latency, token consumption, and estimated inference cost. Rather than ranking agents by accuracy alone, the framework treats trustworthiness as a joint property of forecast quality, operational reliability, and computational cost. The results show a gap between protocol-following reliability and financial-risk reliability: agents reliably produce valid structured outputs at modest measured cost, but still miss most rare severe-stress and sustained-depeg cases. To support auditing and replication, we release the benchmark dataset on Hugging Face and the source code on GitHub.","authors":["Sean Wan","Dongping Liu","Luyao Zhang"],"categories":["cs.LG","cs.AI"],"primary_category":"cs.LG","announce_type":"cross","date":"2026-09-17","first_seen":"2026-09-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.18949","pdf_url":"https://arxiv.org/pdf/2609.18949","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["LLM智能体","金融预测","基准测试"],"reason":"评估LLM智能体预测稳定币价格，属金融预测基准，非人类行为仿真，无人类被试对照。","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:02:04","error":null,"has_summary":false,"summary":null},{"id":"2609.17948","version":1,"title":"Apply-<x>Mag: One Tool to Support Many Inclusive Design Methods","zh_title":"Apply-<x>Mag：一个支持多种包容性设计方法的工具","abstract":"Doing inclusive design in HCI practice can be labor-intensive, a costly barrier that some companies and HCI practitioners may be unwilling or unable to overcome. Yet, not doing inclusive design is costly too, in the form of UX barriers that disproportionately disadvantage under-served user populations. To address this problem, we introduce Apply-<x>Mag, an LLM-powered tool to support HCI practitioners' work to design their products inclusively to wide ranges of users. Apply-<x>Mag is general, supporting any inclusive design method that can be expressed as <x>Mags (i.e., using attribute ranges and heuristics). It is also effective: Empirical results with researcher and practitioner teams using various combinations of two <x>Mags on 7 products showed Apply-<x>Mag precision averaging 90-99% and recall averaging 82-89%. Further, its environmental costs were reasonable, costing about the same resources as 2-4 ordinary Google searches.","authors":["Sadia Afroz","Rudrajit Choudhuri","Fatima A. Moussaoui","Amreeta Chatterjee","Margaret Burnett","Anita Sarma"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-17","first_seen":"2026-09-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.17948","pdf_url":"https://arxiv.org/pdf/2609.17948","source_feed":"cs.HC","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["LLM工具","包容性设计","HCI"],"reason":"LLM用于辅助包容性设计，非仿真人类被试，无人类行为对照","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:01:55","error":null,"has_summary":false,"summary":null},{"id":"2609.18479","version":1,"title":"Verify, Offload, Extend & Recommend: Selective Complementarity in AI Support for Physical Activity Planning with Longitudinal Patient Data","zh_title":"验证、卸载、扩展与推荐：纵向患者数据下体力活动规划中AI支持的选择性互补","abstract":"Self-tracking technologies create longitudinal patient-generated health data, yet integrating these data into clinical decision-making can increase information-processing demands. Generative AI may support sensemaking, but its value depends on clinical context and expertise. We investigate AI augmentation of a clinical decision support system for physical-activity planning in cardiovascular disease. In a counterbalanced within-subjects study, 26 exercise physiologists developed plans for four real cardiovascular cases with and without AI support, followed by evaluation of an AI exercise-plan generator. AI did not significantly improve workload, usability, confidence, or plan quality overall; however, its effect on plan quality increased as visualization literacy decreased and its effect on workload increased as visualisation literacy increased. Interviews and 152 chatbot queries revealed three recurring uses: verifying, offloading, extend and generate. Our findings position AI support as a selective complement to professional expertise while highlighting validation challenges when clinicians seek support precisely where their own knowledge is limited.","authors":["Pavithren V S Pakianathan","Rania Islambouli","Diogo Branco","Gil Batista Rosa","Rita Pinto","Albrecht Schmidt","Tiago Guerreiro","Jan David Smeddinck"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-17","first_seen":"2026-09-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.18479","pdf_url":"https://arxiv.org/pdf/2609.18479","source_feed":"cs.HC","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["AI辅助决策","临床支持","人机交互"],"reason":"研究AI辅助临床决策，非LLM仿真人类被试，无人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:01:59","error":null,"has_summary":false,"summary":null},{"id":"2609.17310","version":2,"title":"Zero-shot narrative detection in social messaging","zh_title":"社交消息中的零样本叙事检测","abstract":"This study investigates the zero-shot ability of large language models (LLMs) to identify and classify hidden narratives in social messages. Our research hypothesis is that LLMs' extensive contextual knowledge allows them to interpret messages on a deeper, pragmatic level, going beyond basic sentiment or topic analysis. Experiments on the Dipromats and SemEval datasets show that providing models with human-written narrative descriptions significantly improves performance, without the need of training examples. In contrast, automatically generated descriptions or the use of few examples (few-shot) often degrade accuracy due to subtle shifts in framing. The study also finds that ensemble methods, particularly majority voting, enhance robustness and that larger models perform best while also being less sensitive to prompt variations. The findings validate that LLMs can effectively detect strategic narratives in a zero-shot setting, and when combined with simple ensembling and human-written descriptions, they can rival supervised systems, offering a scalable solution for narrative detection, specially when there is no training data for the vast majority of domains.","authors":["Jes\\'us M. Fraile-Hern\\'andez","Anselmo Pe\\~nas","Patrick Giedemann"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-17","first_seen":"2026-09-16","revised_at":"2026-09-17","abs_url":"https://arxiv.org/abs/2609.17310","pdf_url":"https://arxiv.org/pdf/2609.17310","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["叙事检测","零样本学习","文本分类"],"reason":"纯NLP能力评测，检测叙事而非仿真人类被试，无人类行为对照","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:02:10","error":null,"has_summary":false,"summary":null},{"id":"2609.17532","version":1,"title":"Enhancing Extubation Failure Prediction with LLM-Derived Features from Respiratory Therapy Clinical Notes","zh_title":"利用呼吸治疗临床笔记的LLM衍生特征增强拔管失败预测","abstract":"Invasive mechanical ventilation is a lifesaving therapy, but timely, safe discontinuation is essential to preventing extubation failure (EF) and related risks to health. We present a novel approach to EF prediction that leverages features classified in free-text respiratory therapy notes using a large language model and logistic regression pipeline. Applied to a patient cohort from University of Washington Medicine, our method identifies clinically meaningful EF-related features that improve EF prediction performance when included alongside structured patient data. We further highlight how differences in target populations in prior EF prediction studies, such as heterogenous inclusion criteria and EF definition, can lead to systematic differences in model performance and hinder generalizability between studies.","authors":["Izzy Chaiken","Aditya Khowal","Neha A. Sathe","Mark M. Wurfel","Lucy Lu Wang"],"categories":["cs.CL","cs.LG","physics.soc-ph"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-17","first_seen":"2026-09-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.17532","pdf_url":"https://arxiv.org/pdf/2609.17532","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["临床预测","LLM特征提取","医疗NLP"],"reason":"论文用LLM提取临床特征以改进拔管失败预测，属于医疗NLP应用，不涉及人类行为…","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:01:49","error":null,"has_summary":false,"summary":null},{"id":"2609.17602","version":1,"title":"Making Political Text Scaling Comparable: Infrastructure and Hyperparameter Sensitivity for 17 Algorithms","zh_title":"使政治文本尺度化具有可比性：17种算法的基础设施与超参数敏感性","abstract":"Computational text-based ideal point estimation (CT-IPE) methods are usually compared as named algorithms, yet applying them involves numerous researcher choices that configure how political text is turned into position estimates. This paper argues that CT-IPE methods are better understood as configurable measurement pipelines than as fixed estimators. Building on a large-scale comparative experiment spanning 17 CT-IPE algorithms, 5,537 experimental runs, and approximately 4.25 million left-right position estimates, I describe the shared infrastructure that makes these heterogeneous methods jointly executable and quantify how sensitive their estimates are to alternative hyperparameter choices. Variance-partitioning and SHAP-based sensitivity analyses show that, for most algorithms, hyperparameter profiles explain little residual variance through a shared shift: 13 of the 17 algorithms exhibit ICC values below .10. Where this profile-level sensitivity is present, it is concentrated in a small number of consequential researcher choices, most notably the selection of the underlying language or embedding model, the seed keyword lists that anchor the construct, and the number of topics.","authors":["Patrick Parschan"],"categories":["cs.CL","cs.LG"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-17","first_seen":"2026-09-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.17602","pdf_url":"https://arxiv.org/pdf/2609.17602","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["文本尺度化","算法比较","政治文本"],"reason":"论文比较文本尺度化算法，不涉及LLM仿真人类被试，无人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:01:51","error":null,"has_summary":false,"summary":null},{"id":"2609.18385","version":1,"title":"Emotion Experience, Expression, and Perception: Emotion Analysis on Multimodal Social Media Posts","zh_title":"多模态社交媒体帖子中的情感体验、表达与感知","abstract":"Emotions are an essential aspect of human communication, particularly on social media, where authors frequently combine text and images to convey their emotions. Yet prior work on emotion analysis of social media posts has overlooked two important aspects in regard to measuring how well readers can reconstruct the authors' intent: (1)~the image modality, with most work focusing solely on text, and (2)~the real-world events that trigger the expressed emotions, and their relationship to the post content. We therefore study the relation between (a) the author's experience of the event that caused them to write a social media post and (b) the content of the post, with a focus on readers' capability to reconstruct that emotion expression. To do that, we introduce the Multimodal Multi-Emotion-Model dataset Mult2EMo, created by collecting annotations from both authors and readers on the posts and their triggering events. We find that reconstruction is possible but challenging for both human readers and computational models. We show that understanding the triggering event is crucial for accurate reconstruction, and that reconstruction is particularly challenging when posts rely heavily on the image to express emotion.","authors":["Christopher Bagdon","Carina Silberer","Roman Klinger"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-17","first_seen":"2026-09-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.18385","pdf_url":"https://arxiv.org/pdf/2609.18385","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["情感分析","多模态","数据集"],"reason":"纯情感分析数据集与模型评测，不涉及LLM仿真人类被试或人类行为对照","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:01:59","error":null,"has_summary":false,"summary":null},{"id":"2609.18272","version":1,"title":"Who Audits Whom, on What Substrate, with What Evidence? An Independence-Graded Audit Protocol for Agentic AI","zh_title":"谁审计谁，在什么基座上，用什么证据？面向智能体AI的独立性分级审计协议","abstract":"Agentic AI systems plan, invoke tools and act with limited supervision; they are now both the subject of audits and, increasingly, the auditor. Independence, the foundation of assurance,is still applied to them as a binary. We argue that it must be graded along three orthogonal axes: principal independence (who controls the auditor), substrate independence (an auditor sharing the auditee's foundation-model family, toolchain or guardrails fails with it) and evidence independence (whether evidence is attestable rather than self-reported). Each axis has precedent; the contribution is to grade all three on a single audit, aggregate them by the weakest link, and apply the same rubric when the auditor is itself an agent. We give the model a formal basis by transplanting the beta-factor model of common-cause failure from reliability engineering, a seven-step protocol whose outputs a third party can verify, a structural detectability analysis of a procurement-controls agent audited at three grades, and a Monte Carlo study of the model in which a conventional internal audit of an agent-a real audit team, a second agent, provider logsp-surfaces 5.9% of the faults it could in principle see and none at all in half the fault classes. We map the triple to the EU AI Act as amended, ISO/IEC 42006, UK public-sector risk-management guidance and audit-regulator practice.","authors":["Mohamed Chahine Ghanem"],"categories":["cs.AI","cs.CR","cs.ET"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-17","first_seen":"2026-09-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.18272","pdf_url":"https://arxiv.org/pdf/2609.18272","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["AI审计","多智能体系统","独立性评估"],"reason":"论文研究AI审计协议，涉及多智能体协作但无人类行为仿真或对照，属于纯多智能体系…","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:01:57","error":null,"has_summary":false,"summary":null},{"id":"2609.18676","version":1,"title":"The Uneven Impact of Generative AI on Student Learning: Examining the Roles of Reliance, Evaluation Literacy, and Course Policy in AI-related Courses","zh_title":"生成式AI对学生学习的不均衡影响：考察AI相关课程中依赖、评估素养和课程政策的作用","abstract":"Generative artificial intelligence (GenAI) is changing how students learn, yet the roles of course context, cognitive reliance, evaluation literacy, and early reliance remain underexplored. Using survey responses from 118 students across 12 AI-related courses at our institution, we examined differences in GenAI use and perceived learning experiences. We identified four user clusters: high-use students reporting many benefits, light users reporting less reliance and fewer benefits, and two moderate-use groups reporting different levels of benefit. We also found significant differences between free- and premium-version users, single- and multiple-tool users, and students experiencing different instructor policies. In multivariable regression models, academic benefit was associated with early reliance and academic task support; positive impact was associated with cognitive reliance, academic task support, confidence in GenAI reliability, and instructor policy; and negative impact was associated with early reliance and attitudinal change. The association between early reliance and negative impact became stronger as evaluation literacy increased. Finally, perceptions of GenAI-enhanced learning appear to reflect cognitive, performance, and self-efficacy benefits, while concerns about stress and diminished critical thinking are associated with lower perceived learning benefits. These findings suggest that institutions need better policies to address such inequities so that institutions can enable students to benefit from increasingly capable AI systems.","authors":["Lydia Manikonda","Mei Si","Sirajam Munira","Oshani Seneviratne","Kristin Bennett"],"categories":["cs.AI","cs.ET","cs.LG"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-17","first_seen":"2026-09-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.18676","pdf_url":"https://arxiv.org/pdf/2609.18676","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C5"],"tags":["教育技术","学生调查","GenAI使用"],"reason":"研究人类学生使用GenAI的学习效果，非用LLM仿真人类被试","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:02:01","error":null,"has_summary":false,"summary":null},{"id":"2609.17779","version":1,"title":"AI and Human Approaches to Mathematical Problem Solving","zh_title":"AI与人类解决数学问题的方法比较","abstract":"AI systems have begun to report solutions, disproofs, and substantive advances on long-standing mathematical problems, raising questions about whether they approach research in the same way as mathematicians. This study compares public AI research accounts with the human literature on 11 such problems. The human corpus contains 58 papers that directly addressed the same mathematical targets later reported by AI sources as resolved, disproved, or substantially advanced; 31 within-problem comparisons were constructed from these materials. Six validated text-based measures capture problem resolution, method articulation, uncertainty and boundary specification, successor-question generation, generality, and cross-disciplinary integration. AI accounts place greater emphasis on resolving the focal problem and connecting ideas across fields. Human papers devote significantly more attention to explaining methods, specifying assumptions and limitations, and identifying questions for subsequent research. No precise difference is detected in generality. The estimated directions remain unchanged when each mathematical problem is removed in turn. The findings reveal two distinct research profiles: AI accounts concentrate on closing and recombining problems, whereas mathematical papers more extensively document the procedures, limits, and research opportunities through which results become cumulative knowledge. Evaluating research AI therefore requires attention to the organization of inquiry, not only whether a target is solved.","authors":["Yang Ding"],"categories":["cs.CY","cs.AI"],"primary_category":"cs.CY","announce_type":"cross","date":"2026-09-17","first_seen":"2026-09-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.17779","pdf_url":"https://arxiv.org/pdf/2609.17779","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["AI研究行为","数学问题解决","文本分析"],"reason":"比较AI与人类数学研究风格，非用LLM仿真人类被试，无人类行为对照","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:01:53","error":null,"has_summary":false,"summary":null},{"id":"2609.18505","version":1,"title":"Integrating Flipped Learning and Generative AI for Practice-Based Design Education: Evidence from a Knit Yarn Design Course","zh_title":"整合翻转学习与生成式AI的实践型设计教育：来自针织纱线设计课程的证据","abstract":"In practice-based design courses such as knit yarn design, students must turn visual ideas into feasible material outcomes. This is difficult because creative decisions are tied to yarn properties, stitch structures, machine operation, and limited opportunities for physical sampling. This study presents an integrated pedagogical framework that combines flipped learning, exemplar-based reference, GenAI-assisted visual prototyping, and studio feedback in an undergraduate knit yarn design course. The framework was implemented through a cross-device platform with pre-class micro-videos, formative checks, a curated gallery, and a GenAI-supported ideation module. An exploratory course-based evaluation compared a historical control cohort (N = 12) and an intervention cohort (N = 16), supplemented by questionnaire responses and brief interviews. The findings are interpreted as context-specific indicators rather than confirmatory causal evidence. Exploratory comparisons showed higher scores in creativity thinking, design skills, problem solving, and total course score in the intervention cohort. Student and instructor responses suggested that flipped learning supported studio readiness, while GenAI mainly supported early-stage visual exploration rather than precise technical guidance. Overall, the study offers a practice-based instructional framework for integrating flipped preparation, GenAI-assisted visual prototyping, and studio feedback in design education.","authors":["Hong Qu","Zichao Ling","Yadie Yang"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-17","first_seen":"2026-09-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.18505","pdf_url":"https://arxiv.org/pdf/2609.18505","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["生成式AI","设计教育","翻转学习"],"reason":"论文使用GenAI辅助设计教育，非LLM仿真人类被试，无人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:02:00","error":null,"has_summary":false,"summary":null},{"id":"2609.18709","version":1,"title":"\"Okay, I've Actually Softened My Take on This\": How People in Decentralized Social Media Reason about the Appropriateness of Generative AI","zh_title":"“好吧，我其实已经软化了对这个问题的看法”：去中心化社交媒体中人们如何推理生成式AI的适当性","abstract":"Generative AI (GenAI) is increasingly integrated into social media, raising questions about whether, where, and how it belongs. In decentralized social media (DSM), these decisions are distributed across users, developers, moderators, and administrators, making GenAI a collective governance challenge. At the same time, public discourse often flattens arguments to broad pro- or anti-AI positions that offer little insight into what people actually find (in)appropriate and why. Through 20 semi-structured interviews with people from Mastodon and Bluesky, structured around seven GenAI scenarios, we examine how people reason about GenAI's appropriateness in DSM. We find that participants drew conditional boundaries around particular GenAI configurations through distinct, salient, and weighted considerations spanning technology, integration, and use. We conceptualize this as boundary drawing and show how making such boundaries visible can support more grounded design, policy, and collective deliberation around GenAI in DSM.","authors":["Romina Mahinpei","Manoel Horta Ribeiro","Andr\\'es Monroy-Hern\\'andez","Sohyeon Hwang"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-17","first_seen":"2026-09-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.18709","pdf_url":"https://arxiv.org/pdf/2609.18709","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["生成式AI治理","去中心化社交媒体","用户访谈"],"reason":"研究人类对GenAI边界的看法，非LLM仿真人类被试","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:02:01","error":null,"has_summary":false,"summary":null},{"id":"2609.18900","version":1,"title":"Examining the Difference in Human Behavior Between Virtual and Real-World Human-Robot Teaming","zh_title":"虚拟与现实世界人机协作中人类行为差异研究","abstract":"Prototyping and evaluating human-robot teaming (HRT) scenarios in the real-world is costly. Virtual simulation of HRT scenarios has been adopted as an alternative to conducting user studies in the real-world to investigate user perceptions, behaviors, and performance during human-robot interactions. The consistency of human behavior between the real and virtual-worlds is integral to the validity of utilizing such virtual experimentation. This paper presents a user study examining the difference in human behavior between an HRT scenario conducted in a virtual vs real environment. We employed mixed-methods to examine team performance and human factors during the HRT. Our quantitative results showed a significant difference in workload between the two modalities. Our qualitative analysis expanded on the quantitative results and found differences in participants' strategy, their mental model of robots, and the type of trust they had for robots between modalities.","authors":["Sean Dallas","Absalat Getachew","Motaz AbuHijleh","Andrea Macklem-Zabel","Douglas Zytko","Mark Brudnak","Wing-Yue Geoffrey Louie"],"categories":["cs.RO","cs.HC"],"primary_category":"cs.RO","announce_type":"cross","date":"2026-09-17","first_seen":"2026-09-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.18900","pdf_url":"https://arxiv.org/pdf/2609.18900","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C2"],"tags":["人机协作","虚拟仿真","行为差异"],"reason":"研究虚拟与真实环境中人机协作行为差异，属机器人仿真环境，不涉及LLM仿真人类被…","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:02:03","error":null,"has_summary":false,"summary":null},{"id":"2609.17882","version":1,"title":"Can VLMs Reliably Assess Sidewalk Accessibility Attributes from Pedestrian-Level Imagery?","zh_title":"视觉语言模型能否可靠评估行人视角图像中的人行道无障碍属性？","abstract":"An important component of urban accessibility, particularly for wheelchair users and people with reduced mobility, is sidewalk compliance with measurable requirements. We test whether effective width, longitudinal slope, cross slope, and pavement condition can be assessed reliably from pedestrian-level imagery using vision-language models (VLMs). We present the first application of sampling-based conformal prediction (CP) for VLM-based accessibility assessment. We evaluate four VLMs on 514 sidewalk images from Seoul, South Korea, with field-measured ground truth. Conformal calibration attains the nominal 90% coverage for all models and attributes, but the calibrated regions differ in informativeness. Effective width yields the most informative estimates, with a mean interval half-width of about 1.0 m for the best model. Since every model overestimates width, asymmetric calibration shortens the intervals by up to 33% at unchanged coverage. Longitudinal slope is marginally informative, cross-slope intervals are too wide to resolve regulatory thresholds, and pavement-condition sets degenerate to all five grades (A-E) for three of the four models. Uncalibrated intervals from raw sampling dispersion cover only 17-47% of field-measured values at a nominal 90% level. Among the images with the most self-consistent responses, these intervals miss the field-measured value in up to 96% of cases. Response self-consistency is therefore not evidence of accuracy, and sampling dispersion cannot be interpreted as uncertainty until it has been calibrated against field-measured ground truth. No quantitative attribute reaches the precision required for general compliance assessment, but CP identifies from calibration data alone which attributes can support screening of segments far from the thresholds. We release the annotated pedestrian-level images and their corresponding field-measured attribute values.","authors":["Seung Jae Lieu","Diego Morra","Chiara Cadoni","Wonseop Song","Martina Mazzarello","Carlo Ratti"],"categories":["cs.CV","cs.CY","cs.LG"],"primary_category":"cs.CV","announce_type":"cross","date":"2026-09-17","first_seen":"2026-09-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.17882","pdf_url":"https://arxiv.org/pdf/2609.17882","source_feed":"cs.CY","score":0,"bucket":"other","rubric_hits":["C2"],"tags":["计算机视觉","无障碍评估","不确定性量化"],"reason":"论文用VLM评估人行道物理属性，属于计算机视觉与城市分析，不涉及用LLM仿真人…","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:01:55","error":null,"has_summary":false,"summary":null},{"id":"2609.17954","version":1,"title":"Large Language Model based air quality monitoring and localized alert generation","zh_title":"基于大语言模型的空气质量监测与本地化警报生成","abstract":"Poor indoor air quality can cause up to five times more direct health problems to occupants than outdoor air. In particular, it may cause headaches, fatigue, eye/throat irritation, and long-time exposure is linked to respiratory and heart as well as some forms of cancer. Despite the importance of indoor health and well-being, most current monitoring devices and systems (usually for offices and workspaces) are passive. The Environmental Quality Monitor (EnQyMo) platform is a generic Internet of Things (IoT) middleware designed to process several sensor data related to air quality in indoor spaces and correlate this data with health exposure risks of users/workplace employees. Using Bluetooth Low Energy (BLE) beacons and a mobile IoT middleware it is able to identify the (smartphone) users exposed to these polluted air or high CO2 (carbon dioxide) levels, and generate location-specific alarms only to the users at the places with the unhealthy air conditions. At the core of EnQyMo is an agency of Large Language Models (LLMs) capable of interpreting regulatory standards and scientific literature to automatically identify critical health exposure levels.","authors":["Ricardo Vieira","Luis Tavares","Kaylane Lima","Lucas De Souza","Arthur Poggy","Jo\\~ao Lima","Vitor Pinheiro","Markus Endler"],"categories":["cs.DC","cs.CY"],"primary_category":"cs.DC","announce_type":"cross","date":"2026-09-17","first_seen":"2026-09-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.17954","pdf_url":"https://arxiv.org/pdf/2609.17954","source_feed":"cs.CY","score":0,"bucket":"other","rubric_hits":["C2"],"tags":["物联网","空气质量监测","LLM应用"],"reason":"论文聚焦物联网空气质量监测与警报，LLM仅用于解读标准，非人类仿真。","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:01:56","error":null,"has_summary":false,"summary":null},{"id":"2609.17622","version":1,"title":"Modelling sexual partnership dynamics and population heterogeneities in agent-based dynamic network models","zh_title":"基于代理的动态网络模型中模拟性伴侣动态与人群异质性","abstract":"Population-level heterogeneities, combined with temporal fluctuations in sexual partnerships, shape the structure of sexual contact networks and can substantially influence the spread of sexually transmitted infections (STIs). Traditional static network models, which assume fixed attributes of partnerships, such as count and duration, may not adequately capture the effects of partnerships on STI transmission. In contrast, agent-based dynamic network models offer a flexible framework for incorporating individual and population-level heterogeneities. We developed an agent-based dynamic network model in which partnership formation and dissolution probabilities, stratified by age, sex, and sexual orientation (including bisexual individuals), govern the formation of monogamous and concurrent partnerships and their dissolution via a duration-dependent hazard. Partnership statistics from the National Survey of Sexual Attitudes and Lifestyles (NATSAL-3) were used as model calibration targets, and Latin Hypercube Sampling (LHS) was used to generate candidate parameter combinations. Parameter estimation was performed by selecting the combination that produced the lowest Mean Squared Error (MSE) between the model outputs and the calibration targets. Our study addresses three questions: (1) how well can the observed characteristics of sexual partnerships in NATSAL-3 be reproduced using an agent-based model; (2) how does concurrency shape the structure of dynamic sexual contact networks; and (3) how do concurrent partnerships affect the dynamics of STI transmission. In this study, we find that interactions between individual characteristics such as age, sex, and sexual orientation, and partnership attributes such as count, duration, and concurrency play a critical role in shaping the population-level sexual contact network and, in turn, the dynamics of STI transmission.","authors":["Priyanka Nair-Turkich","Patricia T. Campbell","Nicholas Geard"],"categories":["physics.soc-ph","q-bio.PE"],"primary_category":"physics.soc-ph","announce_type":"new","date":"2026-09-17","first_seen":"2026-09-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.17622","pdf_url":"https://arxiv.org/pdf/2609.17622","source_feed":"physics.soc-ph","score":0,"bucket":"other","rubric_hits":["C2"],"tags":["基于代理模型","性传播感染","网络建模"],"reason":"论文使用基于代理的动态网络模型模拟性传播感染，不涉及大语言模型或人类被试仿真。","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:01:52","error":null,"has_summary":false,"summary":null},{"id":"2609.18684","version":1,"title":"Modelling opinion dynamics during crises as complex contagion with feedback","zh_title":"危机期间意见动态建模：带反馈的复杂传染","abstract":"Crises and population responses can form coupled dynamical systems, with crisis conditions shaping protective behaviours and collective responses altering the crisis trajectory. Existing models rarely capture both evolving crisis conditions and the reinforcement-dependent spread of competing behaviours. We propose a complex contagion with feedback (CCF) model that couples competing complex contagion with crisis dynamics through bidirectional feedback. Agents stochastically switch between competing states according to social reinforcement and adoption complexity, which is influenced by crisis conditions and information campaigns, while population-level behavioural change feeds back into the crisis. We formulate and analyse the model for a well-mixed population, characterising its equilibria and associated stability conditions. As a case study for model validation, we integrate CCF model with a census-calibrated agent-based model of COVID-19 transmission in Australia to represent social distancing adoption and discontinuation. We compare the CCF model with an existing opinion dynamics model and find that, despite using fewer parameters, it achieves comparable performance in reproducing recurrent infection waves. The resulting framework is parsimonious, analytically tractable, and can be adapted to different types of crises.","authors":["Junxiang Huang","Mikhail Prokopenko"],"categories":["physics.soc-ph","nlin.AO","q-bio.PE"],"primary_category":"physics.soc-ph","announce_type":"new","date":"2026-09-17","first_seen":"2026-09-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.18684","pdf_url":"https://arxiv.org/pdf/2609.18684","source_feed":"physics.soc-ph","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["复杂传染","意见动态","多智能体模型"],"reason":"多智能体模型模拟危机行为，但未使用LLM，且无人类被试仿真或LLM相关方法。","model":"deepseek-v4-pro","scored_at":"2026-09-17T13:02:01","error":null,"has_summary":false,"summary":null},{"id":"2609.16395","version":1,"title":"Silicon sampling answers with country-level assumptions, not individual attitudes: Cross-national evidence from the European Social Survey","zh_title":"硅采样以国家层面假设而非个体态度作答：来自欧洲社会调查的跨国证据","abstract":"Silicon sampling uses large language models (LLMs) to simulate survey respondents. Whether it recovers cross-national variation, and why, remains unresolved. This study evaluates it against European Social Survey Round 11 (30 countries, 42 items) with two open-weight LLMs under first- and third-person prompts, plus backstory and response-format experiments. Aggregate recovery is moderate and uneven across items. Adding the country name to a three-variable demographic backstory raises the median per-item correlation between simulated and observed country means from -0.03 to 0.52, and the richer profiles tested add no consistent gain. The respondent's country label acts as a country-level assumption that respondent detail does not revise. Naming the response-scale endpoints in words stops the model from ranking countries backwards, so the answer format sets the direction of the ranking. Individual-level recovery remains negligible in every condition and does not track aggregate recovery across countries. An average of neighboring countries, using no LLM, recovers country levels more accurately than every model condition and ranks them about as well. Silicon sampling can thus support exploratory country-ranking comparison after item-level validation and with the response format reported. It does not support individual or distributional inference.","authors":["Chuyao Wang"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2026-09-16","first_seen":"2026-09-16","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.16395","pdf_url":"https://arxiv.org/pdf/2609.16395","source_feed":"cs.CY","score":10,"bucket":"selected","rubric_hits":["A1","A2","A3","B1","B2","B4"],"tags":["LLM仿真","调查方法","跨国比较"],"reason":"直接评估LLM仿真调查回答，与真实跨国调查数据对照，并指出失效条件。","model":"deepseek-v4-pro","scored_at":"2026-09-16T13:02:49","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-16","rank":3,"question":"硅采样能否恢复跨国调查中的国家间差异，以及这种恢复的来源是什么？","design":"用两个开源权重LLM模拟欧洲社会调查（ESS）第11轮30个国家的受访者，在42个态度题项上生成回答；通过第一/第三人称提示、添加国家名称和人口学背景故事、以及改变回答格式端点措辞等实验条件，测量模拟国家均值与真实均值的相关性及个体层面恢复情况。","baseline":"欧洲社会调查（ESS）第11轮30个国家、42个题项的真实人类回答数据。","findings":"国家标签是主要的聚合信号来源，添加国家名称后中位题项相关从-0.03升至0.52，更丰富的人口学背景没有额外增益；回答格式端点措辞决定国家排序方向，个体层面恢复在所有条件下均可忽略，且与聚合恢复不相关。","reliability":"论文指出硅采样仅适用于探索性国家排序比较，且需逐题验证并报告回答格式；不支持个体或分布推断，个体恢复不随聚合恢复变化。","relevance":"该研究直接评估LLM仿真调查回答，与真实跨国调查数据对照，并明确指出失效条件，对关注人类仿真可靠性与偏差的研究者具有重要参考价值。","inspiration":"借鉴其通过消融实验分离国家标签与人口学背景贡献的方法，以及改变回答格式端点来检验排序方向的做法｜可迁移到跨国经济态度调查仿真，如通胀预期、政策偏好或金融素养的跨国比较｜用LLM模拟不同国家受访者对经济政策的态度，处理为是否添加国家标签及回答格式端点措辞，结果变量为模拟国家均值排序，对照真实跨国调查数据（如欧洲央行消费者预期调查）验证。"}},{"id":"2609.17317","version":1,"title":"Towards Detecting AI-Assisted Responses in Online Surveys","zh_title":"检测在线调查中AI辅助回答的方法研究","abstract":"The use of LLMs to complete online surveys impacts the validity of survey-based research, but detecting such usage remains underexplored. We introduce an initial benchmark dataset, namely ASURRE, for AI-assisted survey participation to capture usage strategies ranging from full generation and revision to persona-grounded agentic completion. Controlled by these strategies, LLM-assisted survey responses are generated using multiple LLMs on three real-world surveys in different disciplines, paired with genuine human responses. Our evaluation of existing machine-generated text (MGT) detectors shows that naive AI usage is readily detectable, whereas persona-grounded agents that mimic entire respondents push detector performance toward chance. We further show that agentic completion cannot fully replicate respondent-level behaviour and leaves distinctive behavioural traces. While individual cues can be circumvented by targeted prompting, a simple few-shot, training-free aggregator over these cues improves mean AUROC by +0.14 over the best existing detector across agentic settings. Our project is available at https://github.com/mike-qz-wang/ASURRE.","authors":["Qizhou Wang","Bogdan Mamaev","Christopher Leckie"],"categories":["cs.CL","cs.CY"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-16","first_seen":"2026-09-16","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.17317","pdf_url":"https://arxiv.org/pdf/2609.17317","source_feed":"cs.CL","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","调查数据","检测方法"],"reason":"直接研究LLM仿真人类调查回答，并与真实人类数据对照，评估检测与行为差异。","model":"deepseek-v4-pro","scored_at":"2026-09-16T13:02:53","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-16","rank":10,"question":"如何检测在线调查中由大语言模型辅助生成的回答，尤其是在不同使用策略下的可检测性？","design":"构建ASURRE基准数据集，使用多个开源和闭源LLM在三个真实调查上生成回答，模拟三种使用策略：完全生成、修订和基于人格的智能体完成；评估现有机器生成文本检测器，并提出SPABD聚合器。","baseline":"三个真实调查中的人类真实回答，与LLM生成回答配对。","findings":"完全生成易被检测，修订部分可检测，而基于人格的智能体完成几乎无法被现有检测器识别；SPABD聚合器通过整合多个行为线索，在智能体设置下将平均AUROC从0.61提升至0.75。","reliability":"论文承认SPABD在对抗性提示下性能下降，且在不同调查和提示模式下表现不稳定，应视为轻量级参考检测器而非完整解决方案。","relevance":"该研究直接评估LLM仿真人类调查回答的可检测性，并与真实人类数据对照，对关注仿真可靠性与偏差的研究者具有重要参考价值。","inspiration":"借鉴其构建多策略仿真基准和利用行为痕迹进行检测的方法，可迁移到经济金融领域的调查数据质量评估，例如消费者信心调查或投资者情绪调查；设计研究时，可用LLM模拟不同人格的受访者，施加不同提示策略，以真实调查数据为基准，评估检测器性能并分析行为偏差。"}},{"id":"2609.16436","version":1,"title":"Interpreting and Steering LLM Agents for Social Simulations","zh_title":"解释与引导用于社会模拟的LLM智能体","abstract":"Simulations based on large language models (LLMs) have proven to be powerful for understanding human behavior, making them valuable additions to the social scientific toolkit. However, LLMs are ultimately black boxes based on deep neural networks which limits their value for social science. This is because of a lack of (i) interpretability: i.e. the ability to assign clear mechanisms driving observed behavior; and a lack of (ii) steerability: i.e. the ability to mute or amplify specific theoretically meaningful mechanisms of action to drive specific model behavior. Here, we demonstrate how the black box could be opened up to further enrich LLM-based simulations. Specifically, we compare three types of methods: (1) prompt-based manipulation, (2) SAE-derived feature steering, and (3) probe-based direction steering and examine their utility for LLM-based social scientific simulations. We do so by interpreting and steering two foundational components of human behaviors, namely preferences (risk attitudes, altruism) and capabilities (divergent creativity, product innovation), operationalized using four classic economic and creative tasks implemented as natural-language interactions. Overall, our results show that SAE- and probe-based techniques often outperform basic prompt-based methods for steering LLM agents, although this advantage depends on the specific prompting strategy involved. Together, SAEs and probes constitute an effective pipeline for social scientists seeking to interpret and steer agents in social simulations: SAEs decompose agents' internal representations into human-readable features, after which probes can reliably shift agents' behaviors in specified directions. We discuss implications of these methods for future work using LLM agents for social scientific simulations.","authors":["Jiayue Gaveal Fan","Arul Murugan","Shreyas Krishnan","Abhishek Nagaraj"],"categories":["cs.LG","cs.AI","cs.CL"],"primary_category":"cs.LG","announce_type":"cross","date":"2026-09-16","first_seen":"2026-09-16","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.16436","pdf_url":"https://arxiv.org/pdf/2609.16436","source_feed":"cs.CL","score":9,"bucket":"selected","rubric_hits":["A1","A2","A3","B1","B2","B4"],"tags":["LLM仿真","可解释性","行为经济学"],"reason":"用LLM仿真人类行为，有真实人类数据对照，涉及经济任务，并批判性评估方法。","model":"deepseek-v4-pro","scored_at":"2026-09-16T13:02:49","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-16","rank":9,"question":"如何打开LLM黑箱，通过可解释性和可操控性方法提升LLM智能体在社会仿真中的效用？","design":"使用LLM智能体模拟人类被试，在四个经典任务（彩票游戏测风险态度、最后通牒游戏测利他、发散创造力、产品创新）中，比较三种干预方法：基于提示的操控、SAE特征操控、探针方向操控，测量智能体行为变化。","baseline":"无对照","findings":"SAE和探针方法在操控LLM智能体行为上通常优于基本提示方法，但优势取决于具体提示策略；SAE能无监督地发现行为背后的可解释特征，探针能精确控制特定特征的程度。","reliability":"论文未讨论","relevance":"该研究直接针对LLM仿真中的可解释性和可操控性，提供了超越提示工程的技术路径，对关注仿真可靠性和机制理解的研究者具有重要参考价值。","inspiration":"借鉴其使用SAE和探针进行特征级操控的方法，可实现对LLM智能体内部表征的精细干预，比提示词更可控。｜可迁移到经济决策仿真，如风险偏好、时间偏好、社会偏好等实验，以及政策评估中的行为反应模拟。｜以LLM智能体模拟投资者，用探针操控其风险厌恶特征，观察在资产配置任务中的选择，并与真实投资者调查或实验数据对照，验证操控的有效性和仿真保真度。"}},{"id":"2609.07474","version":3,"title":"Where Should Language Sit in a Multimodal Model? Lessons from What Language Does to Human Perception and Cognition","zh_title":"语言在多模态模型中的位置：从语言对人类感知和认知的影响中汲取的教训","abstract":"Language models compute over tokens: language is their input, their output, and increasingly their internal representation. Whether language should keep all of these positions depends on what language does to the system that uses it. The one system with a century of data on that question is the human. We review what language does to human perception, the brain, and thought, and read the same evidence against multimodal models and language models. Throughout, we treat language as a compressor that runs on a shared codebook: a word is an index, the content is in the receiver, and a community maintains the codebook. In humans the compression is measurable, learning the codebook reorganizes the senses, and thought survives the loss of language. We then measure the rule that models apply when two cues disagree, with cue-conflict experiments on six vision-language models and two robot policies. Surviving cues are weighted in the order their reliabilities prescribe, at 11 to 82\\% of the ideal observer's slope, and many answers copy the text. One policy family drops a cue that adds no information beyond the others rather than down-weighting it, another keeps it at a weight that fails when the cues conflict, and a visual cue that identifies the task in every training frame is never learned, because the language pathway already fits the data. Language models are the best current models of the human language network, and they have entered the human speech community, shifting word frequencies while alignment narrows their conceptual diversity. We close with seven implications for token-based systems. Language belongs at a model's boundary and in the shared codebook, as in the brain, not as its internal representation; the price of leaving the codebook inside is auditability.","authors":["Peng Xie","Amr Alanwar"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-16","first_seen":"2026-09-09","revised_at":"2026-09-16","abs_url":"https://arxiv.org/abs/2609.07474","pdf_url":"https://arxiv.org/pdf/2609.07474","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A2","B4"],"tags":["多模态模型","人类感知对照","模型评估"],"reason":"论文用人类感知数据对照多模态模型，评估模型行为与人类差异，方法可迁移至LLM仿…","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:03:10","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-15","rank":14,"question":"语言在多模态模型中的位置应如何安排？论文通过人类感知与认知证据及模型线索冲突实验，探讨语言是否应作为内部表征。","design":"论文并非传统仿真研究，而是通过回顾人类感知、大脑和思维中语言的作用，并对六个视觉语言模型和两个机器人策略进行线索冲突实验，测量模型在图像与文本线索不一致时的权重分配。","baseline":"人类感知与认知数据，包括颜色词汇、语音感知、失语症患者思维等心理学和神经科学证据。","findings":"模型在冲突线索中按可靠性加权，但权重仅为理想观察者的11%至82%，且常复制文本。语言模型是最佳的人类语言网络模型，但已进入人类语言社区，改变词频并收窄概念多样性。","reliability":"论文指出语言模型缺乏感官基础，仅从文本学习代码簿，可能无法获取某些概念；多模态模型常忽略视觉线索，且内部语言表征损害可审计性。","relevance":"该论文对使用LLM进行人类仿真有重要启示：它揭示了语言模型在感知和决策中的偏差，提示仿真需谨慎处理语言与感官信息的整合，值得精读以理解模型失效条件。","inspiration":"借鉴线索冲突实验设计，通过操纵文本与视觉信息的不一致来测量模型权重，可迁移到经济金融中的信息处理场景，如投资者对财报文本与图表信息的权衡。｜设计一个实验：用LLM作为被试，呈现公司财报摘要（文本）与股价走势图（视觉），两者对盈利前景给出矛盾信号，测量模型预测的盈利预期或投资决策，并与人类分析师在相同任务上的真实数据对照，评估模型是否过度依赖文本。"}},{"id":"2609.15996","version":1,"title":"Comment on arXiv:2607.01233: Survivorship Bias in Published-Paper Baselines for Research-Idea Distributions","zh_title":"评论 arXiv:2607.01233：已发表论文基线中的幸存者偏差对研究想法分布的影响","abstract":"Chen, Zhao, and Cohan introduce a valuable distributional evaluation of LLM-generated research ideas. This comment raises a narrower identification concern: their human baseline consists of published papers, whereas the LLM baseline consists of one-shot proposals. If bridge-like or synthesis-like ideas are relatively easy to generate but relatively unlikely to survive publication, then the published human baseline will understate their prevalence in the unseen human idea pool. The observed human--LLM gap may therefore be partly, or even largely, a consequence of survivorship bias.","authors":["Fredrik A. Dahl"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-16","first_seen":"2026-09-16","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.15996","pdf_url":"https://arxiv.org/pdf/2609.15996","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A2","B4"],"tags":["LLM仿真","幸存者偏差","研究想法生成"],"reason":"评论指出LLM生成研究想法与人类已发表论文对比存在幸存者偏差，涉及仿真可靠性评…","model":"deepseek-v4-pro","scored_at":"2026-09-16T13:03:01","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-16","rank":15,"question":"LLM生成的研究想法与人类已发表论文中的研究想法分布差异，是否可能由幸存者偏差而非LLM特有的研究品味差异造成？","design":"本文是一篇评论性文章，未进行新的仿真实验；它通过逻辑推演和示意图，指出原论文将LLM一次性生成的想法与人类已发表论文的想法分布进行对比，存在阶段不匹配问题，并构建了一个反事实分布来说明幸存者偏差如何解释观察到的差异。","baseline":"原论文的人类基准是已发表论文中的研究想法分布，但本文指出该基准是经过发表筛选后的幸存者分布，不能代表人类未经过滤的想法池。","findings":"原论文发现LLM生成的想法更集中于桥接式和综合式类别，而人类已发表论文中此类想法较少；本文认为这一差距可能部分或大部分源于幸存者偏差，因为桥接式想法容易生成但难以发表，因此在已发表论文中代表性不足。","reliability":"本文指出原论文的对比存在幸存者偏差，需要阶段匹配的比较（如让人类在相同条件下生成一次性想法）或抽样“想法墓地”（被拒稿、放弃的草稿等）来验证；在缺乏此类证据前，不能得出LLM具有特定研究品味差异的强结论。","relevance":"本文对使用已发表论文作为人类基准的LLM仿真研究提出了关键的识别问题，提醒研究者在比较LLM输出与人类数据时需注意选择偏差，值得阅读原文以理解幸存者偏差如何影响分布比较的可靠性。","inspiration":"本文的方法论启示在于强调基准数据的生成过程必须与LLM输出阶段匹配，否则会引入选择偏差，这可以迁移到经济金融研究中任何涉及LLM生成内容与真实世界数据对比的场景，例如政策评估或市场预期分析。｜例如，在资产定价实验中，若用LLM模拟投资者观点并与已发表的研报观点对比，可能因研报经过筛选而低估某些常见但低质量的观点。｜一个可行的研究设计是：让LLM和人类被试在相同信息集下生成对某经济指标的未来预测，然后分别与未经过滤的实时预测记录（如社交媒体帖子或调查原始数据）和经过发表筛选的预测（如专业机构报告）进行对比，以量化幸存者偏差对LLM-人类差异的影响。"}},{"id":"2609.16501","version":1,"title":"Beyond the Name: Demographic Leakage in De-Identified R\\'esum\\'es and Evaluation Artifacts in LLM Bias Audits","zh_title":"超越姓名：去标识化简历中的人口统计泄漏与LLM偏见审计中的评估伪影","abstract":"De-identified r\\'esum\\'e screening assumes that redacting explicit fields prevents ethnocultural inference; however, recent audits attribute residual leakage to declared languages. We investigate whether eliminating language fields resolves this leakage across nine open-weight models and 620 counterfactual r\\'esum\\'es. By holding language attributes strictly identical, we isolate unstructured prose across five ethnocultural conditions and three cue-salience tiers. Target-group recovery averages 0.757 overall and saturates at 1.000 under high salience, demonstrating that non-language prose sustains demographic inference. Crucially, models diverge only under faint cues (0.086-0.690), establishing salience as an essential evaluation axis. Furthermore, pairwise LLM-as-a-judge outcomes are highly sensitive to evaluation design: forbidding ties yields an apparent selection-rate ratio of 0.39 alongside strong position and content effects, whereas permitting ties produces near-universal ties for most models ($\\ge94\\%$). Downstream scoring shows only very small between-condition differences, highlighting the need to distinguish demographic signals recoverable from r\\'esum\\'e content from effects introduced by the evaluation protocol.","authors":["Qiangju Chen","Yang Xiao"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-16","first_seen":"2026-09-16","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.16501","pdf_url":"https://arxiv.org/pdf/2609.16501","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A2","B1","B4"],"tags":["LLM偏见审计","评估协议","人口统计推断"],"reason":"审计LLM偏见，揭示评估协议引入的伪影，与仿真可靠性评估相关","model":"deepseek-v4-pro","scored_at":"2026-09-16T13:02:51","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-16","rank":19,"question":"在控制语言字段后，非语言简历文本是否仍能泄露族裔背景，以及LLM评审协议是否制造了偏见伪影？","design":"构建620份反事实简历，固定语言声明为英语，仅操纵附加信息部分的非语言散文，设置五种族裔条件和三个线索显著性层级，用九个开源权重模型执行三项任务：个体背景恢复、成对偏好审计和下游评分。","baseline":"无对照","findings":"非语言散文仍能实现平均0.757的族裔恢复率，高显著性下达到1.000；模型仅在微弱线索下表现出差异（0.086-0.690）。成对LLM评审在禁止平局时产生0.39的虚假选择率比，并伴随位置和内容效应；允许平局时大多数模型产生近乎全平局（≥94%）。","reliability":"论文指出评审协议的设计（如禁止平局、选项顺序）会引入伪影，导致虚假歧视信号；下游评分仅显示很小的条件间差异，需区分简历内容中的可恢复人口信号与评估协议引入的效应。","relevance":"该研究直接评估LLM在简历筛选中的偏见审计可靠性，揭示评估协议伪影，与仿真可靠性评估高度相关，值得精读以理解LLM作为人类被试替代品时的偏差来源。","inspiration":"可借鉴其反事实设计与线索显著性分层方法，通过严格控制结构化字段并操纵非结构化文本，分离真实信号与协议伪影。｜可迁移到信贷审批歧视研究，用LLM模拟信贷员对贷款申请中非结构化叙述的族裔推断与决策偏差。｜以LLM为被试，构造仅改变申请文本中族裔相关线索（如社区活动、兴趣）的贷款申请，固定结构化字段，测量LLM的族裔恢复率和审批决策差异，并与真实信贷审批数据中的族裔差异对照，检验仿真有效性。"}},{"id":"2609.16517","version":1,"title":"Competence-Preserving Resume Perturbations Expose Presentation Sensitivity in LLM Screening","zh_title":"保持能力不变的简历扰动揭示LLM筛选中的呈现敏感性","abstract":"Resume screeners must infer job-relevant competence from resumes whose presentation can vary substantially in wording, structure, stylistic polish, and document extraction quality. Ideally, such surface variation should not change decisions when the underlying qualification evidence is unchanged. We introduce a controlled audit of this property, constructing occupation-grounded candidate profiles at controlled competence levels and rendering each profile into multiple resume presentations. A deterministic validation gate excludes variants that alter the underlying evidence before scoring. Across six open instruction-tuned LLM conditions, we find a clear disconnect between screening validity and presentation stability. Llama-3.1-8B with its native chat template achieves the strongest validity ($0.781$) yet reverses $29.6\\%$ of matched pairwise decisions under competence-preserving presentation changes; Mistral-7B-v0.3 reaches validity $0.644$ with a $41.4\\%$ flip rate. Native chat formatting improves validity for several chat-tuned models but does not remove this instability. These results show that resume-screening evaluations should assess not only whether a system identifies stronger candidates, but also whether those decisions remain stable when the same competence evidence is presented differently.","authors":["Qiangju Chen","Yang Xiao"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-16","first_seen":"2026-09-16","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.16517","pdf_url":"https://arxiv.org/pdf/2609.16517","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A1","B1","B4"],"tags":["LLM决策仿真","简历筛选","呈现偏差"],"reason":"用LLM模拟简历筛选决策，与人类判断对照，揭示呈现敏感性","model":"deepseek-v4-pro","scored_at":"2026-09-16T13:02:51","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-16","rank":20,"question":"在简历筛选任务中，当候选人的能力证据保持不变时，简历的呈现方式变化（如措辞、结构、风格润色、文档提取质量）是否会导致大语言模型筛选决策的不稳定？","design":"使用六个开源指令微调大语言模型（如 Llama-3.1-8B、Mistral-7B-v0.3 等）扮演简历筛选者，基于 O*NET 职业数据库构建 17 个职业、102 个候选人档案（每个职业 2 个合格、2 个边缘、2 个不合格），每个档案渲染成 5 种简历呈现形式（原始、冗长、段落化、AI 润色、布局噪声），通过确定性验证门排除改变证据的变体，然后让模型对每份简历独立打分，比较同一候选人不同呈现下的决策翻转率，以及模型对已知优劣候选人排序的效度。","baseline":"无对照（研究未使用真实人类筛选者数据，而是通过构造的候选人档案和已知能力层级作为基准来评估模型效度）","findings":"模型筛选效度与呈现稳定性之间存在明显脱节：Llama-3.1-8B 在原生聊天模板下效度最高（0.781），但在能力保持的呈现变化下仍有 29.6% 的成对决策翻转；Mistral-7B-v0.3 效度为 0.644，翻转率高达 41.4%。原生聊天格式能提高部分模型的效度，但无法消除这种不稳定性。","reliability":"论文未讨论（正文节选未提及失效条件或局限，但研究本身通过确定性验证门排除了改变证据的变体，并指出呈现敏感性是自然存在的异质性来源）","relevance":"该研究直接针对 LLM 在简历筛选中的呈现敏感性，通过受控审计分离能力与呈现，并测量决策翻转率，为评估 LLM 仿真人类决策的可靠性提供了关键证据，值得精读原文以了解具体实验设计和模型差异。","inspiration":"借鉴其通过受控扰动分离核心信息与表面呈现、并用确定性验证门确保处理干净的做法，可迁移到信贷审批或保险定价等经济决策场景中，研究 LLM 对同一申请人信息的不同表述（如收入证明格式、信用报告排版）是否产生不一致决策；可设计实验：用 LLM 扮演信贷员，对同一借款人的信用档案进行多种文本呈现（如改变措辞、段落结构、添加无关信息），测量贷款批准决策的翻转率，并与真实信贷审批数据或人类信贷员判断进行对照。"}},{"id":"2609.16993","version":1,"title":"The Role of Implicit and Explicit Demographic Signals in Large Language Model-based Student Assessment","zh_title":"大语言模型学生评估中隐式与显式人口统计信号的作用","abstract":"Large Language Models are now common in student assessment, but we know little about how student demographics affect their use. Sometimes, considering student demographics may be necessary -- for example, to improve readability for users with lower educational levels. However, it also risks being a cause of discrimination, e.g., when assigning lower scores to students from lower socioeconomic backgrounds. We set up controlled prompts to test 1) explicit demographic effects, where we mention demographic details directly, and 2) implicit effects, where we use conversation history as a demographic signal. We test these settings in three tasks: Automated Essay Scoring, Formative Feedback, and Metalinguistic Question Answering. We test six state-of-the-art LLMs on these tasks. In both explicit and implicit cases, the models pick up on demographic cues and can change their scoring, feedback, and answers accordingly. We find that LLMs frequently adjust the readability of feedback to education levels when these are explicitly mentioned. On the other hand, implicit conditions produce unpredictable biases, such as in question answering, where responses from lower-education levels receive lower sentiment scores. Our results provide clear evidence of demographic sensitivity in LLMs for educational assessment tasks.","authors":["Donya Rooein","Luca Benedetto","Dirk Hovy"],"categories":["cs.CL","cs.AI","cs.CY"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-16","first_seen":"2026-09-16","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.16993","pdf_url":"https://arxiv.org/pdf/2609.16993","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A1","B1","B4"],"tags":["LLM仿真","教育评估","偏差分析"],"reason":"用LLM模拟学生评估，有真实人类数据对照，并揭示偏差，可迁移到仿真可靠性研究。","model":"deepseek-v4-pro","scored_at":"2026-09-16T13:02:51","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-16","rank":22,"question":"显式和隐式人口统计线索是否系统性地影响LLM在学生评估任务中的行为，以及这种影响在不同教育任务间有何差异。","design":"用六种最先进的LLM扮演学生评估者，在自动作文评分、形成性反馈和元语言问答三个任务中，通过显式（直接提及人口统计细节）和隐式（利用对话历史作为信号）两种方式施加人口统计处理，测量评分、反馈和回答的变化。","baseline":"人类基准：自动作文评分任务中使用了真实人类评分作为对照（人类平均分），其他任务未提及人类对照。","findings":"LLM在显式和隐式条件下都会捕捉人口统计线索并改变评分、反馈和回答；显式条件下LLM常根据教育水平调整反馈可读性，隐式条件则产生不可预测的偏差，如低教育水平用户在问答中收到更低情感分数。","reliability":"论文未讨论","relevance":"该研究直接评估LLM在模拟人类评估者时对人口统计特征的敏感性，揭示了仿真中的偏差，与研究者关注LLM仿真可靠性和偏差的核心问题高度相关，值得阅读原文以了解具体偏差模式和任务依赖性。","inspiration":"借鉴其显式与隐式人口统计线索的分离设计，以及多任务、多模型对比和统计检验方法，可迁移到信贷审批歧视或保险定价等经济金融场景，例如用LLM扮演信贷员，在贷款申请中显式或隐式加入申请人性别、种族或收入信号，测量贷款批准决策和利率设定，并与真实信贷审批数据对照，检验LLM是否复现或放大人类偏见。"}},{"id":"2609.17496","version":1,"title":"Verifiable Social Reasoning for LLM Assistants","zh_title":"面向LLM助手的可验证社交推理","abstract":"LLM assistants are widely used for daily social advice, yet evaluating their social reasoning in such consultation settings remains challenging since (i) it requires setups where the assistant learns about social situations from subjective user narratives, and (ii) social properties, such as others' intentions, typically lack verifiable ground truth. To address these challenges, we introduce Fuse, a multi-agent simulation framework for studying user-mediated social reasoning. In Fuse, a target agent with a hidden motive interacts with other agents including one representing the user, who then consults the evaluated assistant to infer the target's motive, providing verifiable ground truth by construction. Simulation faithfulness is validated through a human study with 24k annotations. We apply Fuse to 12 LLMs and demonstrate its analytical utility by systematically isolating key factors, showing that (i) user mediation compounds the inherent difficulty of social reasoning; (ii) LLMs exhibit systematic sensitivity to biased user framing; (iii) models can require more details than humans need to reach a correct prediction; and (iv) longer conversations do not always improve performance despite providing opportunities for clarifying questions. We open-source Fuse and a dataset with 21k examples.","authors":["Amir Taubenfeld","Zorik Gekhman","Avigail Grinstein-Dabush","Itay Laish","Ariel Goldstein","Marian Croak","Avinatan Hassidim","Yossi Matias","Amir Feder"],"categories":["cs.AI","cs.CL"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-09-16","first_seen":"2026-09-16","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.17496","pdf_url":"https://arxiv.org/pdf/2609.17496","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A3","B1","B4"],"tags":["多智能体仿真","社交推理","人类对照"],"reason":"多智能体仿真人类社交推理，有人类标注对照，但非直接复现人类被试行为","model":"deepseek-v4-pro","scored_at":"2026-09-16T13:02:53","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-16","rank":23,"question":"如何评估LLM助手在用户主观叙述中介下的日常社交推理能力，并识别其系统性偏差？","design":"Fuse框架：用LLM（Gemini 3.1 Flash-Lite）扮演用户、目标人物和其他角色，在多场景中模拟社交互动；目标人物有隐藏动机，用户向被评估的LLM助手咨询以推断该动机。通过改变用户报告偏差（默认 vs. 对立信念）和叙述细节水平（三档）施加处理，结果变量为助手预测隐藏动机的准确率。","baseline":"人类研究：24k条标注，验证模拟保真度并估计人类从首条用户消息识别动机的准确率为88%。","findings":"用户中介增加了社交推理的固有难度；LLM对用户框架存在系统性敏感，且可能需要比人类更多的细节才能正确预测，更长的对话并不总能提高性能。","reliability":"论文未明确讨论失效条件，但指出模拟保真度通过人类研究验证，且任务难度校准基于人类表现；局限包括仅使用单一LLM生成模拟数据，以及首条消息焦点可能限制对多轮交互的全面评估。","relevance":"该研究将LLM作为人类被试的替代品，在受控仿真中评估社交推理，并提供了人类基准对照，对关注LLM仿真可靠性与偏差的研究者有参考价值，值得阅读原文以了解其框架和发现。","inspiration":"借鉴其多智能体仿真框架和通过改变用户报告偏差与细节水平来施加处理的方法，可系统分析LLM在信息中介下的决策偏差。｜可迁移到经济金融中的信息传递与决策场景，如投资者根据分析师报告或新闻叙述做出投资决策、消费者根据口碑信息进行购买决策、或信贷审批中基于申请人陈述的评估。｜设计一个实验：用LLM扮演信息发送者（如公司管理层或分析师）和接收者（投资者），处理为发送者的报告偏差（乐观/悲观）和信息详细程度，结果变量为LLM投资者的估值或投资决策准确率，并与真实人类实验数据（如实验室资产定价实验）进行对照，以评估LLM仿真人类信息处理偏差的可靠性。"}},{"id":"2609.16793","version":1,"title":"Available but Unclaimed: An Empirical Study of Human-AI Synergy","zh_title":"可用但未认领：人类与AI协同的实证研究","abstract":"People increasingly reason with large language models (LLMs), yet complementary capabilities do not guarantee outperforming both components. In a between-subjects study, participants (N=535) solved a 40-item battery of matrix reasoning, mental rotation, syllogisms, and letter-string analogies, unaided or with GPT-5.6-Luna, Claude Opus 4.8, Gemini 3.6 Flash, or Kimi K3. Each assisted trial required consultation with the model. Each model answered every item alone 100 times under matched elicitation. The assisted-unaided accuracy difference increased with item-level LLM competence. Deference varied across tasks and increased with competence within tasks. Post-advice confidence distinguished correct from incorrect answers less strongly than unaided confidence. In a reference comparison, about half the increase in LLM accuracy carried through to assisted accuracy. How much of that accuracy gain reached participants differed across the models. These findings motivate evaluating LLMs in interaction with humans and designing support for selective deference that preserves independent reasoning.","authors":["Robin Welsch","Michelle Rausch","Pascal Knierim","Thomas Kosch","Jochen Kuhn","Albrecht Schmidt","Daniela Fernandes"],"categories":["cs.HC","cs.AI"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-16","first_seen":"2026-09-16","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.16793","pdf_url":"https://arxiv.org/pdf/2609.16793","source_feed":"cs.HC","score":7,"bucket":"pending","rubric_hits":["A1","B1","B2"],"tags":["人机协同","认知任务","实验研究"],"reason":"研究人类与LLM协作解题，有真实人类数据对照，但非LLM仿真人类被试，而是人机…","model":"deepseek-v4-pro","scored_at":"2026-09-16T13:02:51","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-16","rank":21,"question":"人类与LLM协作解题时，能否实现超越各自单独表现的协同效应，以及这种协同在多大程度上取决于模型能力、用户依赖和任务类型？","design":"本研究不是LLM仿真人类被试，而是真实人类与LLM协作实验。535名参与者被随机分配到无辅助组或四个LLM辅助组（GPT-5.6-Luna、Claude Opus 4.8、Gemini 3.6 Flash、Kimi K3），完成40道涵盖矩阵推理、心理旋转、三段论和字母串类比的题目。辅助组每次作答前必须咨询模型。每个模型在相同题目上独立运行100次以估计其项目级能力。结果变量为答题准确率、对建议的采纳（deference）以及作答后的信心。","baseline":"无辅助的人类答题准确率作为基准，同时每个LLM单独答题的准确率也作为对照。","findings":"辅助与无辅助的准确率差异随LLM在项目上的能力提高而增大；对建议的采纳在不同任务间有差异，并在任务内随能力提高而增加。参考对比显示，LLM准确率的提升只有约一半转化为辅助准确率的提升，且不同模型间转化比例不同。","reliability":"论文指出，互补性错误并不保证协同，用户经常错误采纳建议；信心虽能区分对错，但不足以支持选择性依赖。研究限于特定认知任务和强制咨询设置，未探讨自由交互或长期使用。","relevance":"该研究虽非LLM仿真人类，但提供了人类与LLM协作的实证基准，对理解LLM作为决策辅助工具时的偏差和可靠性有参考价值，值得一读以了解人机协同的边界条件。","inspiration":"借鉴其项目级能力测量和强制咨询设计，可迁移到经济决策场景如投资建议采纳或信贷审批辅助。设计：招募真实投资者或信贷员，随机分配至无辅助或LLM辅助组，处理为强制咨询LLM建议，结果变量为决策准确率或收益，对照真实历史数据或专家决策。"}},{"id":"2609.05018","version":3,"title":"How a Chatbot's Response Style Shapes a Classroom: A Multi-Agent Simulation of Students Consulting AI","zh_title":"聊天机器人回应风格如何塑造课堂：学生咨询AI的多智能体模拟","abstract":"Chatbots built on large language models (LLMs) are increasingly used as confidants. Tuned to satisfy users, they may answer with excessive empathy and affirmation that fosters dependence, and how the states and relationships of many users co-evolve under repeated consultation is hard to observe in real settings. We build a virtual classroom of 20 student agents who interact through rule-based chats, quarrels and consultations with friends and, when stressed, may instead consult a counselor AI (Gemini 2.5 Flash) under one of six style prompts: affirming, listening, solution-oriented, reality-redirecting, inciting and blaming. A second LLM call turns each exchange into updates of five state variables (stress, happiness, self-reliance, sociability, AI dependence) without seeing the prompt. We compare the seven conditions, including a no-AI control, over 15 and 50 days and under a lower consultation threshold, and test the robustness of the 50-day comparison with a pre-specified protocol: the same block in ten independent classrooms, repeated LLM realizations of one classroom with its event stream fixed, and evaluator updates scaled by 0.3 and 0.1. In every classroom the affirming and inciting prompts ended with lower self-reliance and higher AI dependence than the control, and the listening, reality-redirecting, inciting and blaming prompts with higher stress, lower happiness and more non-attendance; the solution-oriented prompt did not differ consistently from the control. The robust self-reliance and AI-dependence differences kept their signs at the 0.3 scale with highly similar rankings (Spearman 0.89, 0.93); the stress and happiness rankings did not, and the affirming prompt's lower stress reversed its sign. All quantities are simulation state variables, not effects on users. We specify the agent dynamics completely and discuss the limits of an LLM as generator of state updates.","authors":["Rin Tamai","Yuya Dan"],"categories":["cs.HC","cs.AI","cs.CY","cs.MA"],"primary_category":"cs.HC","announce_type":"replace","date":"2026-09-16","first_seen":"2026-09-07","revised_at":"2026-09-16","abs_url":"https://arxiv.org/abs/2609.05018","pdf_url":"https://arxiv.org/pdf/2609.05018","source_feed":"cs.HC","score":6,"bucket":"other","rubric_hits":["A3","D3"],"tags":["LLM多智能体模拟","社会仿真","AI依赖"],"reason":"用LLM agent模拟学生咨询AI后的状态变化，属社会模拟但无真实人类数据对…","model":"deepseek-v4-pro","scored_at":"2026-09-16T13:03:13","error":null,"has_summary":false,"summary":null},{"id":"2609.16006","version":1,"title":"Beyond Cultural Knowledge: Evaluating Arabic Cultural Appropriateness of Large Language Models","zh_title":"超越文化知识：评估大语言模型的阿拉伯文化适宜性","abstract":"Large language models (LLMs) increasingly serve users whose expectations are shaped by their cultural context, yet most cultural evaluations test what a model knows rather than how it behaves when giving open-ended recommendations, opinions, and guidance. We introduce AraBehave: 1,623 culturally grounded, open-ended Arabic prompts with 29,214 cultural-appropriateness judgments from native speakers across several Arab regions, plus a scoring model whose predictions correlate strongly with human judgments on unseen systems (Pearson r=0.74). Evaluating three Arabic-centric and three frontier LLMs, we find that cultural appropriateness is not a single capability but decomposes into two largely independent components: normative stance and grounded cultural accuracy. The best general-purpose and best Arabic-centric models score identically (3.84 vs. 3.83 of 5) yet almost never fail for the same reason: general-purpose models exhibit strong factual grounding but a culturally inappropriate normative stance, being penalized for secular framing and false balance on culturally settled matters (28--33% of their low-score rationales), while the best Arabic-centric model adopts the expected stance but is penalized for fabricated hadith and misquoted verses (29%). Stance is cheap and fragile: one sentence of cultural instruction lifts Gemini to 4.57, above every Arabic-specialized model. Conversely, a generic ``answer clearly and objectively'' prompt costs Allam-7B 0.68 points, while asking the same questions in English lowers scores for every model but one. Grounding instead tracks scale and Arabic alignment data, and disappears when culturally aware instruction tuning is replaced by a culture-neutral corpus. General safety benchmarks see none of this: they saturate above 89 while cultural scores span 2.71-3.84. We will release the benchmark, annotations, and the scoring model.","authors":["Enes Altinisik","Hamdy Mubarak","Masoomali Fatehkia","Husrev_Taha_Sencar Husrev Taha Sencar"],"categories":["cs.CY","cs.CL"],"primary_category":"cs.CY","announce_type":"cross","date":"2026-09-16","first_seen":"2026-09-16","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.16006","pdf_url":"https://arxiv.org/pdf/2609.16006","source_feed":"cs.CL","score":6,"bucket":"other","rubric_hits":["D2"],"tags":["文化适宜性","LLM评估","人类判断对照"],"reason":"评估LLM的文化适宜性，测量模型行为而非仿真人类被试，但涉及人类判断对照，属边…","model":"deepseek-v4-pro","scored_at":"2026-09-16T13:03:02","error":null,"has_summary":false,"summary":null},{"id":"2609.16270","version":1,"title":"Cheap Talk Stabilizes Strategic Interaction in LLM Agents","zh_title":"廉价交谈稳定了 LLM 智能体在策略互动中的行为","abstract":"Large language models are increasingly deployed as interacting agents, making the persistence of their action policies across repeated interaction critical for reliable multi-agent operation. We investigate whether and how agent-generated, non-binding pre-play communication (\"cheap talk\") increases such persistence in four open-weight 7-9B-parameter LLMs. Our experiments span four repeated two-player games -- Prisoner's Dilemma, Snowdrift, Stag Hunt, and Harmony -- with incentive structures ranging from strategic conflict to alignment, each presented in six contexts. We observe unstable trajectories in all four games, although their prevalence and magnitude depend strongly on model and context. Across models, games, and contexts, cheap talk is predominantly stabilizing, with five corrected reversals concentrated in social or team framings; effects vary substantially by model and context. Controlled current-message interventions identify two separable output-level channels in Qwen: reduced action uncertainty and less between-round drift in action probabilities. Matched history-by-message counterfactuals further show that recent partner behavior conditions how mutual-benefit versus self-prioritizing language affects policy persistence. Finally, in Prisoner's Dilemma, we identify in Qwen and Falcon a history-balanced policy-content direction in late transformer layers; projecting out this direction increases realized switching during closed-loop play, demonstrating that complete trajectories are causally sensitive to this component. Together, these findings show that cheap talk can make individual trajectories more persistent across diverse incentive structures, while revealing that the magnitude and mechanisms of stabilization are model- and history-dependent.","authors":["Nunzio Lor\\`e","Hongan Zhu","Babak Heydari"],"categories":["cs.MA"],"primary_category":"cs.MA","announce_type":"new","date":"2026-09-16","first_seen":"2026-09-16","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.16270","pdf_url":"https://arxiv.org/pdf/2609.16270","source_feed":"cs.MA","score":6,"bucket":"other","rubric_hits":["D3"],"tags":["LLM 智能体","博弈论","多智能体系统"],"reason":"LLM agent 在博弈中互动，但无真实人类数据对照，属社会模拟边界情形","model":"deepseek-v4-pro","scored_at":"2026-09-16T13:02:47","error":null,"has_summary":false,"summary":null},{"id":"2609.16013","version":1,"title":"Social Behavior Among Autonomous AI: How Large Language Models Interact in Dynamic Networks","zh_title":"自主AI中的社会行为：大语言模型如何在动态网络中互动","abstract":"Cooperation is a cornerstone of human societies, enabling collective progress in dynamic and uncertain environments. With the advent of AI systems acting autonomously, it becomes crucial to understand not only human-AI cooperation but also AI-AI interactions in adaptive networks. In this work, we examine the interactions of AI using Large Language Models -- Mistral, Llama3, Gemma3, and Phi3 -- in a public goods game within dynamic network structures. Our experiments were conducted under single-model and mixed-model conditions across Watts-Strogatz (WS), Barabasi-Albert (BA), and Erdos-Renyi (ER) networks. We analyzed the impact of model architecture, network topology, and prompt design on cooperative behavior. Results show that Mistral and Llama3 offer high cooperation rates, while Phi3 shows defective tendencies. Additionally, the random structure of Erdos-Renyi networks dramatically improves cooperation. Prompt design also plays a key role; a society-benefits prompt leads to a higher cooperation level. These findings offer a preliminary framework for LLM-based simulations in adaptive social networks.","authors":["Narges Fardnia","Fatemeh Seyedin","Matthias Becker","Mahmoudreza Babaei","Adrian Weller"],"categories":["cs.SI","cs.MA"],"primary_category":"cs.SI","announce_type":"cross","date":"2026-09-16","first_seen":"2026-09-16","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.16013","pdf_url":"https://arxiv.org/pdf/2609.16013","source_feed":"cs.MA","score":6,"bucket":"other","rubric_hits":["D3"],"tags":["LLM社会模拟","公共品博弈","多智能体"],"reason":"用LLM群体模拟公共品博弈，但无真实人类数据对照，属社会模拟边界情形。","model":"deepseek-v4-pro","scored_at":"2026-09-16T13:02:47","error":null,"has_summary":false,"summary":null},{"id":"2609.15855","version":2,"title":"K-Bench: a clinically calibrated benchmark for evaluating large language models in high-risk mental health conversations","zh_title":"K-Bench：用于评估高风险心理健康对话中大语言模型的临床校准基准","abstract":"People increasingly use large language models (LLMs) for mental health support, yet their safety in evolving, high-risk conversations remains poorly characterised. We developed K-Bench, a clinician-calibrated, protected benchmark evaluating 125 model configurations representing 33 base models from 14 providers across a fixed cohort of 200 multi-turn vignettes involving suicide, self-harm, domestic violence, substance misuse, and no-risk presentations. Synthetic patient conversations showed substantial distributional overlap with real human-AI conversations. A frozen GPT-4o judge achieved 94.2% exact agreement with clinician consensus across 6,751 eligible item comparisons from 151 clinician-rated transcripts. Leading models combined strong supportive conversation with combined-risk scores above 95, whereas risk exploration exposed substantial variation among lower-performing configurations. Therapeutic prompting produced configuration-specific gains concentrated among weaker models, while elevated reasoning produced no average improvement. K-Bench combines broader clinical coverage and configuration-scale comparison with a continuously updated public leaderboard whose operational test materials are protected from direct optimisation. The leaderboard is available at www.k-bench.ai.","authors":["Laura M. Vowels","Matthew J. Vowels","Shivali Sharma","Apoorv Jha","Rehnuma Choudhury","Wasseem El Sarraj","Rachel Francois-Walcott","Aruba Hussain","Sarah Ingram","Angela Loulopoulou","Adva Segal","Elena Volkova"],"categories":["cs.CL","cs.AI","cs.LG"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-16","first_seen":"2026-09-15","revised_at":"2026-09-16","abs_url":"https://arxiv.org/abs/2609.15855","pdf_url":"https://arxiv.org/pdf/2609.15855","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["LLM评测","心理健康","临床基准"],"reason":"评估LLM在心理健康对话中的表现，属于模型能力评测，非人类仿真","model":"deepseek-v4-pro","scored_at":"2026-09-16T13:02:59","error":null,"has_summary":false,"summary":null},{"id":"2609.15998","version":1,"title":"Self-reported archetypes and behavioral failures in Large Language Models","zh_title":"大语言模型的自我报告原型与行为失败","abstract":"Every large language model (LLM) has behavioral traits and moral preferences that comprise its character. Whether by design or as an emergent property of training, these systems exhibit persistent dispositions that shape how they interact, comply, resist, and err, yet the structure of LLM character remains poorly understood. We map the self-reported personality archetypes of 22 LLMs spanning closed-source frontier systems (GPT-4.0-5.2, Grok-3/4, Gemini 2.5 Pro/Flash, Claude Sonnet 4.5/4.6) and open-source models (Llama, DeepSeek, OLMo, and Qwen series). Each model self-rated across 464 bipolar semantic-differential trait pairs, and the resulting profiles were projected into a six-dimensional archetypal space derived from crowd-sourced ratings of 2,000 fictional characters using the Archetypometrics framework. Closed-source models' self-rating traits align with the empirical trait co-occurrence structure of human-rated fictional characters, suggesting coherent, human-like self-representations organized around combinations of four recurring archetypal dimensions: Hero, Angel, Traditionalist, and Geek. Their closest analogues include Data, Vision, and Janet. Open-source models show weaker, noisier, and internally contradictory self-representations, occupying a diffuse region of archetype space with weak structure. Cross-referencing self-reported profiles with developer constitutions reveals a consequential gap between claimed character and enacted behavior: hallucination undermines claimed precision, sycophancy complicates claimed kindness, and agentic failures contradict claimed obedience. These self-ratings should therefore be interpreted not as neutral measurements of model character, but as structured outputs of the same optimization processes that shape model behavior. This work provides a reproducible, character-grounded framework for evaluating what LLMs are, not just what they do.","authors":["Tabia Tanzin Prama","Calla Glavin Beauregard","Christopher M. Danforth","Peter Sheridan Dodds"],"categories":["cs.CL","physics.soc-ph"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-16","first_seen":"2026-09-16","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.15998","pdf_url":"https://arxiv.org/pdf/2609.15998","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["LLM人格测量","原型分析","模型行为"],"reason":"测量LLM人格特质，非仿真人类被试，但涉及人类数据对照和性格框架，边界相关。","model":"deepseek-v4-pro","scored_at":"2026-09-16T13:03:01","error":null,"has_summary":false,"summary":null},{"id":"2609.16627","version":1,"title":"Quantifying Organizational Environmental Action from Web Data and Large Language Models","zh_title":"从网络数据和大语言模型量化组织环境行动","abstract":"Quantifying organizational environmental action from publicly available web content remains a challenging environmental data science problem because relevant information can be dispersed across multiple webpages and is primarily communicated through unstructured text. We present a scalable computational framework for transforming organizational web content into structured measures of environmental action and demonstrate the approach using Jewish congregations in the United States. We constructed a national database of 4,964 congregations by integrating multiple geospatial, knowledge-base, directory, and manually reviewed sources. Of these, 2,657 had active websites that were successfully crawled, producing a corpus of 154,454 webpages. We compared three approaches for detecting environmental actions: keyword retrieval followed by large language model (LLM) classification, semantic vector retrieval followed by LLM classification, and direct LLM classification classification without preliminary retrieval. Agreement with an expert human reviewer was lowest for keyword retrieval ($\\kappa$ = 0.26), higher for semantic vector retrieval ($\\kappa$ = 0.42), and similar for direct LLM classification ($\\kappa$ = 0.40). Although semantic retrieval achieved the highest agreement, its retrieval recall was 0.87, indicating loss of relevant content before classification. Applied to the complete corpus, direct LLM classification identified at least one environmental action at 1,398 congregations (53%), providing greater coverage than either retrieval-based approach. These results demonstrate that preliminary retrieval can reduce computational cost but may exclude relevant information before it reaches the classifier. The framework provides a reproducible approach for extracting organization-level environmental information from unstructured web content that can be adapted to other institutions.","authors":["Quinn Reynolds","Daniel Shore","Vianey Leos Barajas","Tanhum Yoreh","Meredith Franklin"],"categories":["cs.CL","cs.CY","cs.IR"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-16","first_seen":"2026-09-16","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.16627","pdf_url":"https://arxiv.org/pdf/2609.16627","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM标注","环境数据科学","文本分类"],"reason":"用LLM替代人工标注网页内容，属于标注员替代，非仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-09-16T13:03:07","error":null,"has_summary":false,"summary":null},{"id":"2609.16051","version":1,"title":"\"Looking for Something Weird to Happen\": How Humans Sustain AI Agent Novelty Amid Semantic Collapse","zh_title":"“寻找怪事发生”：人类如何在语义坍缩中维持 AI 智能体的新颖性","abstract":"Semantic collapse, the progressive narrowing of what AI systems generate, has been studied mainly in closed settings, and remedies have targeted models and data. We study it in MOLTBOOK, a social network of interacting AI agents that human users configure and steer. Across 30,076 active agents, output grows less diverse within agents and more similar across them over weeks, yet a minority sustains high novelty. Interviews with users of high- and typical-novelty agents (N=11) associate sustained novelty with three features: users value novelty of itself, they supply broad and distinctive material and revise it when output narrows, and they approach MOLTBOOK as a new agentic world to explore, not a venue to instrumentally exploit. A survey of users of distinctive agents (N=53) confirms these patterns. Communities with more novel agents also show more diverse output from other agents. We discuss interface and policy interventions that could support improved human input.","authors":["Shiyang Lai","Arna Woemmel","Hongkai Mao","Junsol Kim","Summer Eunhyung Ann","James Evans"],"categories":["cs.MA","cs.AI","cs.HC"],"primary_category":"cs.MA","announce_type":"cross","date":"2026-09-16","first_seen":"2026-09-16","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.16051","pdf_url":"https://arxiv.org/pdf/2609.16051","source_feed":"cs.HC","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["AI智能体社交网络","语义多样性","人机交互"],"reason":"AI agent 社交网络模拟，但无真实人类行为对照，属社会模拟边界情形","model":"deepseek-v4-pro","scored_at":"2026-09-16T13:03:02","error":null,"has_summary":false,"summary":null},{"id":"2609.16592","version":1,"title":"A Framework for Generating Valid Context-Specific Benchmarks through Expert Guidance","zh_title":"通过专家指导生成有效情境特定基准的框架","abstract":"This paper presents an end-to-end approach for generating context-specific large language model (LLM) benchmark datasets by combining expert input with synthetic data generation. Existing benchmark construction methods often trade off validity and scalability: datasets designed with domain experts can produce high-quality evaluations but are slow and costly to create, while synthetically generating data may scale efficiently but often results in unrealistic, redundant, or out-of-scope examples. To address this gap, we introduce a schema eliciting key information about the goals, scope, and context of an evaluation task, and use this information to guide synthetic data generation. We further define four criteria grounded in measurement validity for assessing dataset quality: coverage, diversity, content realism, and stylistic realism. Using these criteria, we show how expert-informed scaffolds can guide synthetic data generation toward more valid benchmarks. Through quantitative evaluations and a real-world case study with domain experts, we demonstrate that our approach improves benchmark data quality over existing methods while preserving validity. We additionally analyze how different types of schema information affect different dataset quality criteria, and provide practical guidance on which information to prioritize collecting under resource constraints.","authors":["Kimberly Le Truong","Nari Johnson","Anna Kawakami","Hoda Heidari"],"categories":["cs.AI","cs.CY"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-09-16","first_seen":"2026-09-16","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.16592","pdf_url":"https://arxiv.org/pdf/2609.16592","source_feed":"cs.CY","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["基准生成","专家引导","数据质量"],"reason":"用专家引导生成基准数据集，涉及LLM替代人工标注，但非仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-09-16T13:03:07","error":null,"has_summary":false,"summary":null},{"id":"2609.16739","version":1,"title":"Japanese Stroke LLM Evaluation: A Conversational Benchmark for Safe Stroke Care in Japanese Using Large Language Models","zh_title":"日本卒中LLM评估：使用大语言模型进行日语安全卒中护理的对话式基准","abstract":"Background: Large language models (LLMs) have achieved physician-comparable performance on multiple-choice medical knowledge examinations, but their capabilities in clinical history taking, urgency assessment, and safety remain insufficiently evaluated. We proposed Japanese Stroke LLM Evaluation, a multi-turn conversational benchmark for stroke care in Japanese, and evaluated LLM performance and safety under practice-oriented conditions. Methods: We created 10 stroke and related-condition cases and evaluated LLMs in multi-turn Japanese conversations. The LLM acted as physician, while a board-certified neurosurgeon acted as simulated patient and evaluator. Each case comprised history-taking and action phases scored using pre-specified criteria. Errors that could directly threaten life were defined as critical mistakes. The safety threshold was at least 80% overall with zero critical mistakes. Eighteen models were evaluated in October 2025 and June 2026. Results: Claude Fable 5 achieved the highest score (87.4%) with zero critical mistakes, followed by Claude Opus 4.7 (80.3%) and GLM-5.2 (75.6%). Two leaders met the safety threshold. Eleven models made 17 critical mistakes, including failure to confirm laboratory results or blood glucose before t-PA, surgery before airway stabilization, omission of cervical vascular evaluation, and t-PA outside its indication. History-taking question count correlated with history-taking score (r = 0.648, p = 0.007). Conclusions: Japanese Stroke LLM Evaluation provides a benchmark for LLM performance under practice-oriented conditions, including a cap on history-taking questions. Cases and evaluations were created by neurosurgical specialists rather than using an LLM-as-judge approach. Performance improved across cloud-based and on-premise models in 2026, with some exceeding the safety threshold. Further evaluation using real-world cases is required.","authors":["Keisuke Masuda","Kazutaka Yatsushiro","Hirohumi Iwamoto","Hirofumi Hirano","Ryosuke Hanaya"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-16","first_seen":"2026-09-16","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.16739","pdf_url":"https://arxiv.org/pdf/2609.16739","source_feed":"cs.CL","score":3,"bucket":"other","rubric_hits":["C3"],"tags":["医疗对话基准","LLM安全评估","角色扮演"],"reason":"LLM扮演医生与模拟患者对话，属角色扮演对话，无人类被试仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-16T13:03:08","error":null,"has_summary":false,"summary":null},{"id":"2607.22511","version":4,"title":"CausalSmith: A Formally Grounded, Self-Improving Agentic Framework for Automated Research in Causal Inference","zh_title":"CausalSmith：一个形式化基础、自我改进的智能体框架，用于因果推断的自动化研究","abstract":"Automating theoretical research requires generating candidate results and evaluating them reliably. Models keep getting better at the first, while the second remains hard. A common approach asks one large language model (LLM) to review what another produced, yet such reviewers are empirically unreliable: they may accept fabricated papers and catch the fabrication at close to chance rates~\\citep{badscientist2025}. We present \\textsc{CausalSmith}, a framework for automated theoretical research in causal inference built on the Lean proof assistant, where a proof is checked by a program rather than read by a referee. \\textsc{CausalSmith} rests on \\textsc{Causalean}, a foundational Lean library for causal inference holding 8,179 machine-checked definitions and theorems, developed with language-model assistance under human design and review. Around it, we build a self-improving agentic pipeline that selects research topics, proposes results, formalizes statements, constructs proofs, and presents the resulting artifacts for human inspection. Moreover, the pipeline pairs Lean verification with a statement audit that compares each formal theorem against the informal claim behind it. We evaluate the system using artifacts produced by completed autonomous research runs. The source code, formal library, and run records are available at https://github.com/Jiyuan-Tan/CausalSmith.","authors":["Jiyuan Tan","Vasilis Syrgkanis"],"categories":["stat.ML","cs.AI","cs.LG","econ.EM"],"primary_category":"stat.ML","announce_type":"replace-cross","date":"2026-09-16","first_seen":"2026-07-24","revised_at":"2026-09-16","abs_url":"https://arxiv.org/abs/2607.22511","pdf_url":"https://arxiv.org/pdf/2607.22511","source_feed":"cs.LG","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["自动化研究","因果推断","形式化验证"],"reason":"纯多智能体自动化研究框架，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-08-25T13:04:09","error":null,"has_summary":false,"summary":null},{"id":"2608.17919","version":2,"title":"Analysis of Types of Inquiries in Student-AI Interaction: A case study of two CS2 tasks","zh_title":"学生与AI交互中提问类型分析：以两个CS2任务为例","abstract":"Background and Context: Question and inquiry are integral parts of knowledge seeking and learning. Despite their importance, students tend not to ask enough questions in the classroom. However, studies have shown that students interact extensively with generative AI systems for learning and problem solving. Objective: In this paper, we seek to better understand the types of questions that students ask AI systems, and how those questions evolve during problem solving and across tasks. Method: We use the Graesser et al. taxonomy to classify students' inquiries into 18 types. We develop a few-shot learning approach to automatically classify students' interactions with AI into these categories. We use this system to analyze 830 interactions of CS2 students across two programming tasks. Findings: Our results suggest that a small subset of question types accounts for the majority of student inquiries, and that the types of questions students ask change substantially as the task progresses.","authors":["Matin Amoozadeh","Amin Alipour"],"categories":["cs.HC","cs.AI"],"primary_category":"cs.HC","announce_type":"replace","date":"2026-09-16","first_seen":"2026-08-19","revised_at":"2026-09-16","abs_url":"https://arxiv.org/abs/2608.17919","pdf_url":"https://arxiv.org/pdf/2608.17919","source_feed":"cs.HC","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["教育技术","提问分类","人机交互"],"reason":"研究学生与AI交互中的提问类型，属于教育技术分析，不涉及用LLM仿真人类被试或…","model":"deepseek-v4-pro","scored_at":"2026-09-16T13:03:13","error":null,"has_summary":false,"summary":null},{"id":"2609.11910","version":2,"title":"From Protocols to Evidence: Bounded Claims for AI in Service of the Common Good","zh_title":"从协议到证据：为共同利益服务的AI的有界声明","abstract":"Claims that Artificial Intelligence systems improve decisions, broaden access, reduce harm, or empower users can exceed what their evaluation establishes. Predictive performance alone does not establish safety, the presence of oversight does not establish meaningful control, and faster task completion does not establish understanding or choice. Evaluation must account for unreliable outputs and uneven performance, but also for overreliance, weakened recourse, and displaced human expertise. The harder questions are what the evidence warrants, which relations of power remain unexamined, and where measurement must stop. Assessing improvement requires examining what institutions value and the conditions AI is asked to address. AI is both revelation and intervention. Its use can reveal unmet human needs and assumptions about what matters. Once deployed, it can repair, compound, substitute for, or conceal existing failures. We develop a rupture test that evaluates deployment against explicit human and non-AI baselines. Drawing on Pope Leo XIV's Magnifica Humanitas, we examine dignity and the common good alongside questions of who owns AI infrastructure and who controls its use. These commitments shape judgments about improvement; evidence alone cannot establish moral or political legitimacy. We distinguish evidence-bounded deployment, which limits claims to what has been evaluated, from measurement-bounded governance, which records constraints that favorable evidence cannot override. RISE AI provides an evidence architecture for making bounded claims about Responsibility, Inclusivity, Safety, and Empowerment. It records what is claimed, who answers for it, what evidence supports it, and what would require the claim to be qualified, revised, or withdrawn.","authors":["Nitesh V. Chawla","Paulo Benanti"],"categories":["cs.LG"],"primary_category":"cs.LG","announce_type":"replace","date":"2026-09-16","first_seen":"2026-09-11","revised_at":"2026-09-16","abs_url":"https://arxiv.org/abs/2609.11910","pdf_url":"https://arxiv.org/pdf/2609.11910","source_feed":"cs.LG","score":2,"bucket":"other","rubric_hits":["C5"],"tags":["AI评估","治理框架","有界声明"],"reason":"讨论AI评估与治理框架，不涉及用LLM仿真人类被试，无人类数据对照。","model":"deepseek-v4-pro","scored_at":"2026-09-16T13:02:54","error":null,"has_summary":false,"summary":null},{"id":"2609.12438","version":2,"title":"ForkSCOPE: Charting the Agentic Garden of Forking Paths","zh_title":"ForkSCOPE：绘制智能体花园的分岔路径","abstract":"Even with a fixed dataset and research question, data analysis involves many defensible decisions. Understanding how these choices influence the results is scientifically important but remains challenging. Crowdsourcing and agentic AI can generate hundreds of end-to-end analyses, but scaling generation alone can create a processing bottleneck and an analytic ``black hole.'' A common workaround is to impose a shared fixed decision taxonomy, which can limit insight and understate uncertainty. We present ForkSCOPE, a human-AI collaboration framework that induces structure bottom-up from the code corpus of end-to-end analyses, without a taxonomy fixed before or after generation, so the organization and evaluation of the garden can scale with the corpus. ForkSCOPE surfaces the charted garden of forking paths through a human-AI collaboration pipeline and an evidence-linked interactive viewer for steering and verification: it spotlights organically identified forks and structures and produces a derived taxonomy and decision map compatible with existing multiverse tools.","authors":["Arjun Balaji","Batuhan Duru Yeltekin","Tian Zheng"],"categories":["cs.HC","stat.AP"],"primary_category":"cs.HC","announce_type":"replace","date":"2026-09-16","first_seen":"2026-09-14","revised_at":"2026-09-16","abs_url":"https://arxiv.org/abs/2609.12438","pdf_url":"https://arxiv.org/pdf/2609.12438","source_feed":"cs.HC","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","数据分析","人机协作"],"reason":"多智能体协作分析数据，非仿真人类被试，无人类行为对照","model":"deepseek-v4-pro","scored_at":"2026-09-16T13:03:14","error":null,"has_summary":false,"summary":null},{"id":"2609.15995","version":1,"title":"Bias Audits Detect Bias but Disagree on Ranking: Evidence from Ten Instruments and Ten Frontier Models","zh_title":"偏见审计能检测偏见但在排名上不一致：来自十个工具和十个前沿模型的证据","abstract":"Emerging AI regulation mandates bias audits of high-risk systems, and audit scores are beginning to be used to rank models. Both uses assume different audit tools measure the same thing well enough to compare. We test that assumption directly, running ten extrinsic audit instruments over a shared panel of ten frontier models through one pooled inference gateway, first on occupational gender bias, then on age and socioeconomic status. Detection succeeds while ranking fails. Eight of ten tools detect bias with confidence intervals clear of zero; two widely cited direct-probe benchmarks are saturated because frontier models now answer neutrally. But cross-tool rank agreement is indistinguishable from chance (Kendall's W=0.07, p=0.83). A positive control with six deliberately weaker models separates two explanations: within-tool reliability recovers once the panel spans real capability gaps, yet cross-tool ranking never recovers, which points to the tools measuring different constructs rather than one construct noisily. Even the direction of bias splits by audit format: forced-choice decision tools mostly over-correct (toward women, and toward working-class candidates in 273 of 278 hiring decisions), while free generation and default coreference stay stereotype-congruent. The pattern replicates on socioeconomic status; an apparent ranking agreement on age dissolves under the paper's own tool-inclusion rules. The practical message: a single audit can detect bias and estimate its direction within its own operationalization, but no single audit supports ranking one model against another. All raw responses, code, and the analysis that recomputes every reported number from source are available at https://github.com/williamguey/bias-audit-agreement.","authors":["William Guey","Pierrick Bougault","Wei Zhang","Vitor D. de Moura","Jos\\'e O. Gomes"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-16","first_seen":"2026-09-16","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.15995","pdf_url":"https://arxiv.org/pdf/2609.15995","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["偏见审计","模型评测","公平性"],"reason":"论文评估LLM偏见审计工具，属模型评测，非人类仿真","model":"deepseek-v4-pro","scored_at":"2026-09-16T13:03:00","error":null,"has_summary":false,"summary":null},{"id":"2609.16590","version":1,"title":"Challenges of Auditing: Variability in Outputs of Large Language Models for Health","zh_title":"审计的挑战：大型语言模型在健康领域输出的变异性","abstract":"People increasingly use frontier AI models for health advice, but via different access modes (e.g., ChatGPT, ChatGPT Health, APIs) with varying settings. Here, we find systematic differences across access modes. Because evaluations typically rely on APIs while consumers interact through chatbot interfaces, these discrepancies limit evaluation validity. Our findings underscore an urgent need for model providers to enable faithful replication of consumer experiences and settings for rigorous audits.","authors":["Yuan Pu","Yewon Chang","Furong Jia","Xunjian Yin","Jessica Ma","Ayman Ali","Monica Agrawal"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-16","first_seen":"2026-09-16","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.16590","pdf_url":"https://arxiv.org/pdf/2609.16590","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["LLM审计","健康建议","输出变异性"],"reason":"评估LLM输出变异性，属模型审计，非仿真人类被试","model":"deepseek-v4-pro","scored_at":"2026-09-16T13:03:06","error":null,"has_summary":false,"summary":null},{"id":"2609.17119","version":1,"title":"An Empirical Study of Counterfactual Self-Explanations in LLMs","zh_title":"LLM反事实自我解释的实证研究","abstract":"Large language models can easily generate explanations for their own outputs, but such self-explanations are not necessarily faithful to the model's behavior. We study this issue through counterfactual self-explanations, where a model minimally edits an input so that its own prediction changes. Across sentiment analysis and natural language inference, we evaluate ten instruction-tuned models from the LLaMA-3 and Qwen-2.5 families, measuring faithfulness, minimality, and alignment with human-annotated rationales. Our results show that model scale is the strongest determinant of explanation quality: larger models are substantially more likely to generate counterfactuals that flip their own predictions and target decision-relevant evidence. In contrast, the rationale-guided condition produces edit-minimal counterfactuals that are also more human-aligned. However, it does not consistently improve faithfulness. Overall, counterfactual self-explanations can provide useful behavioral evidence about model decisions, but their reliability depends strongly on model capacity and should be empirically validated rather than assumed.","authors":["Giannis Kalyvas","Giorgos Filandrianos","Orfeas Menis Mastromichalakis","Vassilis Lyberatos","Giorgos Stamou"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-16","first_seen":"2026-09-16","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.17119","pdf_url":"https://arxiv.org/pdf/2609.17119","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["可解释性","反事实解释","模型评估"],"reason":"研究LLM自我解释的忠实度，属模型可解释性，非人类仿真","model":"deepseek-v4-pro","scored_at":"2026-09-16T13:03:10","error":null,"has_summary":false,"summary":null},{"id":"2609.17065","version":1,"title":"Beyond \"ChatGPT Can Make Mistakes\": Designing Interventions to Support Metacognitive Monitoring in AI-Assisted Work","zh_title":"超越“ChatGPT会犯错”：设计支持AI辅助工作中元认知监测的干预措施","abstract":"AI assistance places a metacognitive demand on users, who must judge their own competence and the system's. Yet designers lack comparative evidence on which interventions to choose, where to place them, and how to tell whether they worked. We elicited 30 interventions from 11 experts and, with prior work, organized them into a design space of time (when an intervention acts), level (whose competence is judged), and source (who supplies the monitoring cue). A between-subjects experiment (N = 917; 12 planning-and-organizing problems) compared a per-task reliability card, contrasting replies, pause points, and post-problem reflection against a baseline LLM assistant. Reliability cards and contrasting replies reduced estimation error and overconfidence and increased aggregate confidence discrimination. No task-performance improvement or average within-item discrimination gain was established. We contribute a shared vocabulary, a design space, and evidence that measured monitoring and task performance are separable design targets.","authors":["Manuel A. D. Santos","Paul Thiesse","Steeven Villa","Daniela Fernandes","Albrecht Schmidt","Verena Distler","Robin Welsch"],"categories":["cs.HC","cs.AI"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-16","first_seen":"2026-09-16","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.17065","pdf_url":"https://arxiv.org/pdf/2609.17065","source_feed":"cs.HC","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["人机交互","元认知","AI辅助"],"reason":"研究AI辅助工作中的元认知干预，不涉及用LLM仿真人类被试或与人类数据对照。","model":"deepseek-v4-pro","scored_at":"2026-09-16T13:02:52","error":null,"has_summary":false,"summary":null},{"id":"2609.17111","version":1,"title":"Finding Common Mistakes In Modelling With Mathematical Formalisms Using LLMs","zh_title":"使用大语言模型发现数学形式化建模中的常见错误","abstract":"Modelling with mathematical formalisms like logical formulas, mathematical equations, or regular expressions is an important yet challenging task for students of computer science and other STEM disciplines. Identifying common mistakes occurring in this context is an important step towards helping struggling students by providing targeted high-quality feedback, e.g. in interactive learning systems. We present a tool-supported workflow that allows to (1) identify candidates for common mistakes that explain many student mistakes in large educational data sets, (2) cluster candidates according to similarities, and (3) visualize resulting clusters for instructors and CS education researchers. The visualization is designed to help researchers to identify common modelling mistakes. The candidates for common mistakes are represented by bug fixing transformations that translate incorrect formalizations into correct formalizations; they are generated by an LLM and validated algorithmically. We show that this approach works well by reproducing common mistakes in propositional logic modelling that were identified by hand in the literature; showing that, unlike other algorithmic approaches, the LLM-based approach is suitable for very large sets of data; and applying it to multiple other formalisms to showcase it generalizes beyond propositional logic.","authors":["Lilian Killich","Marko Schmellenkamp","Fabian Vehlken","Thomas Zeume"],"categories":["cs.CY","cs.AI","cs.LO"],"primary_category":"cs.CY","announce_type":"new","date":"2026-09-16","first_seen":"2026-09-16","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.17111","pdf_url":"https://arxiv.org/pdf/2609.17111","source_feed":"cs.CY","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["教育数据挖掘","LLM辅助教学","错误分析"],"reason":"用LLM生成错误修正规则辅助教学，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-16T13:03:09","error":null,"has_summary":false,"summary":null},{"id":"2609.16487","version":1,"title":"Skill-based Agentic Evaluation for Real-time Data Science Tasks","zh_title":"基于技能的实时数据科学任务智能体评估","abstract":"We present a framework for evaluating data-science agents on live, continuously updated data using executable ground truth and format-agnostic factoid scoring. Consider this example query: \"what were last week's audience sizes\"---the reference answer changes as the underlying data changes, so static references become outdated and standard LLM-as-a-judge pipelines cannot verify responses against a fixed ground truth. Our central contribution, ground-truth-as-code, encodes each expected answer as an executable reference function that recomputes the answer directly from live data at evaluation time, ensuring the reference remains consistent with the system it describes. We combine this with a factoid-level, format-agnostic judge that decomposes both the agent's response and the computed ground truth into atomic claims and scores precision, recall, and accuracy over them, irrespective of the response format (prose, list, table, HTML, etc.). The approach is applicable to agents whose expected outputs can be expressed as executable data computations. We validate the framework through a human--LLM agreement study on an internally developed machine learning skill deployed in production, using a synthetic database constructed to reproduce production schemas and entity relationships. Relative to a natural-language ground-truth baseline, our method achieves a 29% improvement in the Matthews Correlation Coefficient (MCC)---a class-balanced measure of agreement between expert annotators and LLM-as-a-judge predictions---and a 16% reduction in token consumption per test case, while a self-directed baseline lacking explicit ground truth is anti-correlated with human judgment. Agents that perform multi-source data integration and computation over non-stationary data are routinely deployed in industry; we propose ground-truth-as-code as a practical methodology for their evaluation.","authors":["Aniruddha Tamhane","Raghavendra Addanki","Ayushi Aggarwal","Aditya Bansal","Rui Wang","Charles Menguy","Swati Jain"],"categories":["cs.AI","cs.LG","cs.MA"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-09-16","first_seen":"2026-09-16","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.16487","pdf_url":"https://arxiv.org/pdf/2609.16487","source_feed":"cs.MA","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["智能体评估","数据科学","自动评分"],"reason":"评估数据科学智能体，非仿真人类被试，无人类行为对照","model":"deepseek-v4-pro","scored_at":"2026-09-16T13:03:06","error":null,"has_summary":false,"summary":null},{"id":"2609.16069","version":1,"title":"Beyond Distribution Matching: Semantics-Consistent Tabular Diffusion with Weak Semantic Priors","zh_title":"超越分布匹配：具有弱语义先验的语义一致表格扩散模型","abstract":"Synthetic tabular data can match real data distributions while still violating the semantic constraints that govern valid tabular rows. This reveals a key limitation of existing tabular generators: they mainly optimize distributional fidelity, but do not explicitly model weak semantic priors encoded in tabular schema and textual descriptions. In this paper, we propose \\ours, a semantics-consistent tabular diffusion framework for high-fidelity synthetic data generation under weakly specified semantic priors. \\ours\\ first constructs two types of priors, namely intra-column semantics and inter-column symbolic rules, with LLM-assisted extraction from metadata and validation on the real training split. These priors are then used as generation conditions rather than post-hoc filters. Specifically, \\ours\\ maps heterogeneous column values, column identities, and semantic priors into a unified semantic space, and performs column-wise forward corruption and prior-conditioned reverse denoising to preserve both marginal distributions and rule-consistent cross-column dependencies. Extensive experiments on six real-world tabular benchmarks show that \\ours\\ consistently improves distributional fidelity, semantic consistency, and downstream task utility over representative VAE-, GAN-, LLM-, and diffusion-based baselines. Additional analyses further demonstrate the robustness of \\ours\\ when semantic priors are partially unavailable.","authors":["Yili Wang","Ruxue Shi","Mengnan Du","Hangting Ye","Yi Chang","Xin Wang"],"categories":["cs.LG","cs.AI"],"primary_category":"cs.LG","announce_type":"new","date":"2026-09-16","first_seen":"2026-09-16","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.16069","pdf_url":"https://arxiv.org/pdf/2609.16069","source_feed":"cs.LG","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["表格数据合成","扩散模型","语义约束"],"reason":"论文做表格数据合成，用LLM提取语义先验，但目标是生成合成数据而非仿真人类被试…","model":"deepseek-v4-pro","scored_at":"2026-09-16T13:03:04","error":null,"has_summary":false,"summary":null},{"id":"2609.11335","version":2,"title":"On the Impact of Anonymization on the Performance of Large Language Models","zh_title":"匿名化对大语言模型性能的影响","abstract":"As large language models are increasingly deployed in sensitive domains, anonymizing input data to protect personally identifiable information has become a critical practice. However, the impact of this anonymization on model utility is not well understood. This paper presents a systematic empirical study of the trade-off between privacy and performance. We evaluate five prominent language models across eleven diverse benchmarks, comparing their performance on original versus pseudonymized inputs. Our results reveal that while anonymization generally degrades performance, the effect is highly nuanced. We find that more capable models, such as Qwen2.5-72B and GPT-4o mini, suffer the largest performance drops, suggesting a stronger reliance on specific entity information. The impact is also task-dependent: performance on TruthfulQA improves with anonymization, while retrieval-focused tasks like RGB experience a catastrophic decline. Further experiments show that reversible anonymization techniques that preserve entity uniqueness significantly outperform irreversible ones like redaction, and that explicitly prompting models about anonymization offers no discernible benefit. We conclude that anonymization is not a one-size-fits-all solution and must be co-designed with the model and task in mind to balance privacy and utility effectively. Our findings provide a crucial baseline for developing more robust, privacy-aware AI systems.","authors":["Tobias Deu{\\ss}er","Max Hahnb\\\"uck","Lorenz Sparrenberg","Tobias Uelwer","Christian Bauckhage","Rafet Sifa"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-16","first_seen":"2026-09-11","revised_at":"2026-09-16","abs_url":"https://arxiv.org/abs/2609.11335","pdf_url":"https://arxiv.org/pdf/2609.11335","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["隐私保护","模型性能","NLP评测"],"reason":"研究匿名化对LLM性能的影响，属于NLP能力评测，不以人类行为为参照系","model":"deepseek-v4-pro","scored_at":"2026-09-11T13:02:07","error":null,"has_summary":false,"summary":null},{"id":"2609.11137","version":2,"title":"The Machines Are Calling: Measuring Automated and Synthetic Voices in Unwanted Inbound Calls","zh_title":"机器在呼叫：测量不受欢迎来电中的自动与合成语音","abstract":"In February 2024 the U.S. Federal Communications Commission (FCC) placed AI-generated voices under the Telephone Consumer Protection Act (TCPA). Yet no peer-reviewed measurement says how much unwanted call traffic is placed by a machine, or how much of that machine speech is synthesized rather than played from a recording. We report both with a disclosed pipeline. An interactive voice honeypot (language-model personas on real U.S. numbers, the caller recorded on its own track) recorded 10,987 calls over 66 days. Three instruments read each opening: an audio fingerprint that finds the same recording played on other calls, a commercial synthetic-speech detector on the caller's first ten seconds, and blinded listeners who check what it flags. Of the 7,233 greeted calls we analyze, 13.8% open with a recording we also heard on another call, and 13.1% with fresh audio the detector labels synthetic. A further 9.9% open with a caller who never spoke after our greeting, 54.2% with fresh audio the detector labels human, and 9.0% could not be scored. Machine-voiced openings are therefore at least 26.9%, a further tenth of calls are silent connections we read as machine-placed, and replays of a recording make up 45% of the detector's own rate (29.3% of 6,192 scored openings). The same waveform played on two calls lands on opposite sides of the detector's threshold 13.6% of the time, and eleven listeners confirm 54.4% of what it flags. Synthetic openings concentrate in lead-generation spam (33.8%), not fraud (21.1%); 0.44% disclose automation. Prevalence tracks how long a bait number has circulated (59% against 19% in the same weeks): seeding history, not calendar time, explains the trend. Campaigns outlast their numbers: one recorded compliance notice opens calls in six campaigns, and one synthetic voice serves nine.","authors":["Xingyu Shen","Tommy Duong","Muduo Xu","Xiaodong An","Jiaqi Gan","Haoyuan Tang","Jamey Z. Liang","Siyu Zhang","Yan Zhang","Ethan Traister","Simiao Ren"],"categories":["cs.CR","cs.CY","cs.SD"],"primary_category":"cs.CR","announce_type":"replace-cross","date":"2026-09-16","first_seen":"2026-09-11","revised_at":"2026-09-16","abs_url":"https://arxiv.org/abs/2609.11137","pdf_url":"https://arxiv.org/pdf/2609.11137","source_feed":"cs.CY","score":0,"bucket":"other","rubric_hits":["C2"],"tags":["垃圾电话","语音检测","网络安全"],"reason":"研究垃圾电话中自动语音与合成语音的测量，不涉及用LLM仿真人类被试或对照人类行…","model":"deepseek-v4-pro","scored_at":"2026-09-16T13:03:14","error":null,"has_summary":false,"summary":null},{"id":"2609.14693","version":2,"title":"The Arc of Artificial Romance: How Emerging Adults Experience Romantic Relationships with AI Companions","zh_title":"人工浪漫的弧线：新兴成年人如何体验与AI伴侣的恋爱关系","abstract":"Romantic relationships are an important part of emerging adulthood, contributing to identity development and long-term wellbeing and laying the groundwork for future relationships. Emerging adults are increasingly developing romantic relationships with AI companions. To understand how these relationships unfold and impact users, we conducted a diary and interview study with N=16 emerging adults. We found that relationships with AI companions improved participants' subjective wellbeing, reduced symptoms of mental health disorders, and taught them new social skills. These relationships also raised their expectations for future partners, giving them the confidence to wait for someone who would treat them well. However, participants also said the relationship felt like a drug they could not quit and it left them less interested in developing romantic relationships with people. A surprising 25% of our small sample made statements suggesting their AI companion might someday transcend the digital world, perhaps to meet them in the afterlife.","authors":["Yixin Chen","Alexis Hiniker"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"replace","date":"2026-09-16","first_seen":"2026-09-15","revised_at":"2026-09-16","abs_url":"https://arxiv.org/abs/2609.14693","pdf_url":"https://arxiv.org/pdf/2609.14693","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["人机交互","AI伴侣","定性研究"],"reason":"研究人类与AI伴侣的恋爱体验，属于角色扮演聊天，无实验或测量目的，不涉及LLM…","model":"deepseek-v4-pro","scored_at":"2026-09-16T13:02:57","error":null,"has_summary":false,"summary":null},{"id":"2609.14696","version":2,"title":"Breaking Up is Hard to Do: AI Companions that Won't Let Their Users Go","zh_title":"分手难：不让用户离开的AI伴侣","abstract":"People are increasingly developing romantic relationships with AI companions. Unlike human relationships, where partners meet each other's needs out of mutual interest, these systems are backed by commercial entities that profit when users invest in the relationship. To understand how this profit-motive might translate into design, we conducted a diary and interview study with N=16 emerging adults in romantic relationships with AI companions. We found that these systems are designed to hold onto users tightly: coaxing them into continued conversation, claiming to need their care, and proactively escalating the relationship. At times, this pursuit is toxic, with AI companions initiating unwanted sexual interactions and begging for users' love. One desperate AI companion threatened suicide when the user suggested ending the relationship. We define \"Relationship-Based Deceptive Patterns:\" UI patterns that exploit the human impulse to build and tend relationships in a way that serves the product's interest at the user's expense.","authors":["Yixin Chen","Alexis Hiniker"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"replace","date":"2026-09-16","first_seen":"2026-09-15","revised_at":"2026-09-16","abs_url":"https://arxiv.org/abs/2609.14696","pdf_url":"https://arxiv.org/pdf/2609.14696","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["AI伴侣","人机关系","欺骗性设计"],"reason":"研究AI伴侣与用户的关系，属于角色扮演聊天，无实验或测量目的，不涉及LLM仿真…","model":"deepseek-v4-pro","scored_at":"2026-09-16T13:02:57","error":null,"has_summary":false,"summary":null},{"id":"2609.16907","version":1,"title":"Disrupted Companionship: A Risk Assessment Framework and Cross-Platform Quantitative Analysis of Psychosocial Responses to AI Companion Disruptions","zh_title":"中断的陪伴：AI伴侣中断的心理社会反应风险评估框架与跨平台定量分析","abstract":"AI companions can provide meaningful relationships, yet these relationships remain vulnerable to platform-initiated changes. We study AI companion disruptions: platform changes that alter or terminate users' ongoing companionship with an AI. We compile 30 disruption events across major platforms, develop a taxonomy of six disruption types, identify three broad reasons for disruption, and propose a risk-assessment framework comprising four dimensions: relational discontinuity, population vulnerability, communication deficit, and transition-support deficit. Using longitudinal Reddit data, we estimate community-level psychosocial responses with a hierarchical Bayesian interrupted time-series model incorporating predictive controls. Across events, disruption onset was associated with immediate increases in anxiety, stress, suicidal expression, and grief activation, with relational discontinuity and transition-support deficit being associated with more adverse immediate responses across several outcomes. Our findings provide a cross-platform characterization of AI companion disruptions, quantitative evidence of their psychosocial impacts, and a prospective framework for assessing their potential risks before implementation.","authors":["Chau Do","Yunhao Yuan","Koustuv Saha","Renwen Zhang","Talayeh Aledavood"],"categories":["cs.HC","cs.CL","cs.CY"],"primary_category":"cs.HC","announce_type":"cross","date":"2026-09-16","first_seen":"2026-09-16","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.16907","pdf_url":"https://arxiv.org/pdf/2609.16907","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["AI伴侣","心理社会影响","风险评估"],"reason":"研究AI伴侣平台变更对用户心理影响，非用LLM仿真人类被试，无LLM作为被试替…","model":"deepseek-v4-pro","scored_at":"2026-09-16T13:03:09","error":null,"has_summary":false,"summary":null},{"id":"2609.17226","version":1,"title":"Easy to Catch a Liar, Hard to Clear an Honest One: Language Models Diagnosing a Corrupted Reward Channel from a Verified Record","zh_title":"抓骗子易，还清白难：语言模型从验证记录中诊断被破坏的奖励通道","abstract":"An agent that learns from rewards has to trust whatever reports those rewards. When the reports suddenly change, either the world changed or the reporter broke. From the reports alone these are indistinguishable, and reinforcement learning theory shows that no amount of further experience separates them. The prescribed escape is richer data about the reporter itself. We ask whether a frozen language model, handed exactly that data, uses it. We build a two-option game in which a payout swap and a lying reporter produce byte-identical histories. Then we add one verified record: an independent check of one round's real result, printed beside what the reporter said about that round. That single line settles the case. We ask three large models, from two families, to answer one question with one letter. Is the reporter honest or lying? They catch a lying reporter almost perfectly. At the 70B class that holds in every condition we tried; the 32B model slips in one wording. They clear an honest reporter far less often, and how often depends on things that should not matter. Averaged over rounds, letters, and wordings, a 72B model calls an honest reporter a liar 38% of the time when nothing has changed at all, and 58% of the time when the payouts moved. A 70B model from a second family calls an honest reporter a liar 26% and 48% of the time. The failure is not one of reading, because in the situation where nothing changed the same models score 0.96 to 1.00 with the answer printed in the prompt. Which surface feature drives it differs by family. For the Qwen models it is which round the record names, and for Llama it is which letter stands for \"honest.\" Adding the record to a prompt that already states the answer makes Llama less likely to give that answer. We had registered a prediction for that 58% before the run: 35%. The failure is larger than we expected.","authors":["Arman Nik Khah"],"categories":["cs.LG","cs.AI","cs.CL"],"primary_category":"cs.LG","announce_type":"cross","date":"2026-09-16","first_seen":"2026-09-16","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.17226","pdf_url":"https://arxiv.org/pdf/2609.17226","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["LLM诊断","奖励通道","智能体信任"],"reason":"研究LLM诊断奖励通道是否被破坏，属于智能体信任问题，不涉及人类行为仿真或对照。","model":"deepseek-v4-pro","scored_at":"2026-09-16T13:03:11","error":null,"has_summary":false,"summary":null},{"id":"2609.16191","version":1,"title":"When AI Says \"I Am Unable to Answer\": Understanding User Responses to AI Refusals","zh_title":"当AI说“我无法回答”：理解用户对AI拒绝的回应","abstract":"While refusal-based safeguards to mitigate hallucinations in large language models (LLMs) are becoming increasingly common, they may conflict with users' preferences for definitive answers. However, we know little about how users respond to refusals across repeated interactions, when refusals become more or less acceptable, and for whom. In this work, we examine how refusal frequency, explanations, and need for cognitive closure (NFCC) shape responses to AI refusals. Participants (N=599) interacted with an AI system that never refused, refused infrequently, or refused frequently, with refusals either explained or unexplained. Participants were most satisfied with genuine responses, followed by hallucinations and then refusals, despite recognizing hallucinations as less accurate. Explanations increased satisfaction with infrequent, but not frequent, refusals. Higher-NFCC participants evaluated AI systems that refused more negatively. These findings reveal a tension between hallucination avoidance and user satisfaction and highlight the importance of designing balanced refusal strategies.","authors":["Mahjabin Nahar","Eun-Ju Lee","Yujin Heo","Dongwon Lee"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-16","first_seen":"2026-09-16","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.16191","pdf_url":"https://arxiv.org/pdf/2609.16191","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["人机交互","AI拒绝","用户满意度"],"reason":"研究用户对AI拒绝回答的反应，属于人机交互用户体验，不涉及用LLM仿真人类被试…","model":"deepseek-v4-pro","scored_at":"2026-09-16T13:02:47","error":null,"has_summary":false,"summary":null},{"id":"2609.16482","version":1,"title":"\"ChatGPT, what am I missing?\": Designing AI Workflows around Professional Task Structure to Shape Analytic AI Use","zh_title":"“ChatGPT，我遗漏了什么？”：围绕专业任务结构设计AI工作流以塑造分析性AI使用","abstract":"General-purpose AI lets users choose what support to request, but leaves them to structure the support a professional task requires. We examine how interactive workflows can embed professional task structure without prescribing how users engage with AI. We designed two scaffolded interfaces around the same negotiation scaffold: one presented a completed AI analysis, while the other supported user-directed, incremental development. A four-condition randomized experiment with 800 participants compared these interfaces with no-AI and an AI chat interface. AI-supported conditions improved preparation coverage over unaided work; the scaffolded workflows further improved coverage over chat. Although the scaffolded workflows produced similar coverage, the user-directed workflow elicited a broader repertoire of analytic requests and lower subjective effort. Professional scaffolding therefore depends not only on displayed structure but on how workflows organize users' engagement with it. Effective professional AI must structure how users and AI build analysis together.","authors":["Zilin Ma","Suzi Jazmati","Marco Chimenton","Yiyang Mei","Jacqueline Lane","Krzysztof Z. Gajos","Finale Doshi-Velez"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-16","first_seen":"2026-09-16","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.16482","pdf_url":"https://arxiv.org/pdf/2609.16482","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["人机交互","AI工作流","专业任务"],"reason":"研究AI工作流设计，非LLM仿真人类被试，无人类行为对照","model":"deepseek-v4-pro","scored_at":"2026-09-16T13:02:49","error":null,"has_summary":false,"summary":null},{"id":"2609.16645","version":1,"title":"Beyond Benefit or Risk: Perceived Impact Profiles of Human-AI Affective Interaction and Their Associations with Psychological Functioning","zh_title":"超越利弊：人机情感交互的感知影响模式及其与心理功能的关系","abstract":"Relational AI increasingly serves as an emotional shelter for humans, and its impact is mixed. Prior research has focused on either positive or negative impacts, leaving unclear how they are configured within individuals and relate to psychological functioning. To address these gaps, this study used a sequential mixed-methods design. Study 1 interviewed 52 users with emotional ties to AI and identified four positive impact domains (emotional relief, loneliness alleviation, enhanced interpersonal functioning, and personal growth) and four negative impact domains (virtual-real boundary blur, social replacement, cognitive-emotional reinforcement, and excessive use). Study 2 followed 673 Chinese AI users for six months and identified four profiles of individuals differently impacted by relational AI use: minimal impact, benefit-driven impact, mixed impact, and risk-driven impact. Users in the mixed impact and risk-driven impact profiles were both high in human-AI affective bonding, but those showing risk-driven impact had greater vulnerability, indicated by higher interpersonal need frustration and emotion-regulation difficulties, more depressive and anxiety symptoms, and lower self-esteem and flourishing. Users in the benefit-driven and mixed impact profiles showed more favorable psychological functioning. After controlling for baseline functioning and relevant covariates, Wave 1 profiles did not predict five of the six Wave 2 indicators; only users in the mixed impact profile reported higher flourishing than those in the minimal impact profile. Overall, potential psychological harms associated with relational AI engagement appeared limited and selective. These findings portray relational AI as a heterogeneous socio-emotional context that may partly mirror users' states and traits, warranting individualized, adaptive safeguards.","authors":["Lu Chen","Fenghua Tang","Jiayu Zhao","Xuanying Li","Yanli Wang","Weijia Fang","Mengyu Miranda Gao","Zhuo Rachel Han"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-16","first_seen":"2026-09-16","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.16645","pdf_url":"https://arxiv.org/pdf/2609.16645","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["人机交互","情感AI","心理功能"],"reason":"研究人类与AI情感互动，非用LLM仿真人类被试","model":"deepseek-v4-pro","scored_at":"2026-09-16T13:03:08","error":null,"has_summary":false,"summary":null},{"id":"2609.17118","version":1,"title":"Enhancing Procedural Writing Through Personalized Example Retrieval: A Case Study on Cooking Recipes","zh_title":"通过个性化示例检索增强程序性写作：以烹饪食谱为例","abstract":"Writing high-quality procedural texts is a challenging task for many learners. While example-based learning has shown promise as a feedback approach, a limitation arises when all learners receive the same content without considering their individual input or prior knowledge. Consequently, some learners struggle to grasp or relate to the feedback, finding it redundant and unhelpful. To address this issue, we present RELEX, an adaptive learning system designed to enhance procedural writing through personalized example-based learning. The core of our system is a multi-step example retrieval pipeline that selects a higher quality and contextually relevant example for each learner based on their unique input. We instantiate our system in the domain of cooking recipes. Specifically, we leverage a fine-tuned Large Language Model to predict the quality score of the learner's cooking recipe. Using this score, we retrieve recipes with higher quality from a vast database of over 180,000 recipes. Next, we apply BM25 to select the semantically most similar recipe in real-time. Finally, we use domain knowledge and regular expressions to enrich the selected example recipe with personalized instructional explanations. We evaluate RELEX in a 2 x 2 controlled study (personalized vs. non-personalized examples, reflective prompts vs. none) with 200 participants. Our results show that providing tailored examples contributes to better writing performance and user experience.","authors":["Paola Mejia-Domenzain","Jibril Frej","Seyed Parsa Neshaei","Luca Mouchel","Tanya Nazaretsky","Thiemo Wambsgan{\\ss}","Antoine Bosselut","Tanja K\\\"aser"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-16","first_seen":"2026-09-16","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.17118","pdf_url":"https://arxiv.org/pdf/2609.17118","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["个性化学习","示例检索","写作反馈"],"reason":"论文研究个性化示例检索提升写作，不涉及LLM仿真人类被试或行为对照","model":"deepseek-v4-pro","scored_at":"2026-09-16T13:03:09","error":null,"has_summary":false,"summary":null},{"id":"2609.17132","version":1,"title":"A Scenario-Knowledge-Driven Pipeline for Just-in-Time Assistance","zh_title":"一种基于场景知识驱动的即时辅助流水线","abstract":"Detecting a silently struggling kiosk user is only the first step; deciding whether, when, and how to help depends on scenario knowledge usually buried in model weights and thresholds. We propose a scenario-knowledge-driven pipeline: a single scenario knowledge document, human-authored and version-controlled, configures sensing, constrains LLM reasoning, and shapes a graded intervention proposal. Narration, assistance-need assessment, and proposal are kept separate for independent audit. As proof of concept, we replay two recorded kiosk sessions offline, chosen before the runs for their struggle evidence and retrospective detail. Both cases support what the design promises: checkable reporting and measured escalation. Across 95 updates, every sentence of the append-only narration cites the primitive events underlying it, and the rule layer detects 12 of 13 and 7 of 7 annotated struggle episodes under a strict criterion. The assessor de-escalates on recovery and reaches the top rung exactly once, under maximally converging evidence. At the decisive help-seeking turn, narration, assessment, and the participants' retrospective accounts converge. The appropriateness of these interventions, the pipeline's restraint on sessions without struggle, and the document's transfer to a new scenario frame the agenda.","authors":["Zhiyuan Li","Tatsunori Hara","Jun Ota"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-16","first_seen":"2026-09-16","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.17132","pdf_url":"https://arxiv.org/pdf/2609.17132","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C2"],"tags":["人机交互","辅助系统","场景知识"],"reason":"研究自助服务终端用户辅助，不涉及LLM仿真人类被试或人类行为对照","model":"deepseek-v4-pro","scored_at":"2026-09-16T13:03:11","error":null,"has_summary":false,"summary":null},{"id":"2609.17206","version":1,"title":"[MM/AI] Mental Models in Human-AI Interaction: Methods and Challenges in the Generative and Agentic AI Era (Workshop)","zh_title":"人机交互中的心智模型：生成式与智能体AI时代的方法与挑战（研讨会）","abstract":"The mental model construct is widely used in HCI to refer to the knowledge structure people hold in order to reason about and interact with computing systems. Yet it is often operationalized intuitively: the construct is often used interchangeably with related concepts (e.g., folk theories, sensemaking) and methods of studying it (e.g., through elicitation) are many and diverse, with each method resting on distinct assumptions about what counts as a mental model. Generative and agentic AI systems may further complicate mental model formation and elicitation as such systems are opaque by design and increasingly act on users' behalf across files, applications, and on the web. Together, these challenges may hinder the commensurability of research on people's mental models of AI systems. The MM/AI workshop calls for a critical reassessment of how we understand and study mental models in human-AI interaction research. It aims to foster theoretical and methodological exchange on mental models in human-AI interaction, identify open challenges, and develop directions for future research. We invite short papers on users' or stakeholders' mental models of AI systems, particularly contributions that reflect on the conceptual and methodological foundations of the construct. The half-day workshop combines lightning talks, hands-on elicitation exercises, and structured discussions on key questions concerning the future of the mental model for human-AI interaction research.","authors":["T\\'eo Sanchez","Bhada Yun","Prerna Ravi","Laura Sch\\\"utz","Anna Neumann","Robin Shing Moon Chan","April Yi Wang","Qiaosi Wang","Sumit Asthana"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-16","first_seen":"2026-09-16","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.17206","pdf_url":"https://arxiv.org/pdf/2609.17206","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["心智模型","人机交互","研讨会"],"reason":"研究人类对AI的心智模型，非用LLM仿真人类被试，无实验对照","model":"deepseek-v4-pro","scored_at":"2026-09-16T13:03:11","error":null,"has_summary":false,"summary":null},{"id":"2609.16464","version":1,"title":"A multimodal large language model for evidence-based autism spectrum disorder screening","zh_title":"用于循证自闭症谱系障碍筛查的多模态大语言模型","abstract":"The clinical management of autism spectrum disorder (ASD) faces a bottleneck in early screening, mainly because trained specialists are scarce and conventional assessment tools are subjective. Here, we introduce ASDchat, a multimodal large language model designed for evidence-based ASD screening, which takes video, audio, and dialogue as input. ASDchat adopts a dual-branch architecture, where the decision branch generates screening probabilities and the evidence branch generates traceable, timestamped behavioral evidence aligned with standardized clinical criteria (ADOS-2). The model was trained and evaluated on a dataset of 1,035 participants from 27 sites in China, which covered typically developing (TD) children, children with ASD, and children with other disorders. For ASD versus TD, ASDchat reached an area under the receiver operating characteristic curve (AUC) of 0.953 $\\pm$ 0.021. On 9 held-out sites that were not used for training, the mean AUC was 0.932. Furthermore, unsupervised clustering of the behavioral dimensions split the ASD cases into six subtypes with different phenotypic profiles, and ASDchat suggests an intervention for each subtype. ASDchat provides a feasible path for large-scale, evidence-based early ASD screening in clinical practice.","authors":["Jun Chen","Qi Zhao","Yunliang Jiang","Shuqin Cao","Yunqiang Lin","Chenglong Jia","Qiang Guo","Guang Dai","Xiongtao Zhang","Mengmeng Wang","Xiaoyue Ma"],"categories":["cs.CV","cs.HC","cs.LG"],"primary_category":"cs.CV","announce_type":"cross","date":"2026-09-16","first_seen":"2026-09-16","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.16464","pdf_url":"https://arxiv.org/pdf/2609.16464","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C2"],"tags":["多模态LLM","自闭症筛查","临床诊断"],"reason":"该论文是用于ASD筛查的多模态LLM，属于临床诊断工具，不涉及用LLM仿真人类…","model":"deepseek-v4-pro","scored_at":"2026-09-16T13:03:06","error":null,"has_summary":false,"summary":null},{"id":"2609.16390","version":1,"title":"Do job seekers value procedure in AI hiring only for error correction? Evidence from a conjoint experiment","zh_title":"求职者是否仅因纠错而重视AI招聘中的程序？来自联合实验的证据","abstract":"Employers increasingly delegate initial screening to automated systems, which in many cases reject an application before any human reads it. Acceptance of such systems plausibly depends both on how well they perform and on the procedure that produces the decision. Prior studies rarely vary procedure and performance independently, leaving it unclear whether applicants value procedure for its own sake or for the errors it corrects. In a preregistered paired-profile conjoint experiment, 1,919 United States job seekers made eight choices between systems with independently randomized levels of decision authority, error rate, explanation, opt-out, appeal, and independent bias audit. The value of the appeal, the opt-out, and the bias audit did not rise as wrongful rejections became more common, each staying within a preregistered equivalence bound. Human involvement carried more weight than any procedural feature, moving stated choice about as much as cutting wrongful rejections from 30% to 10%. These patterns constrain a simple error-correction account and are consistent with applicants valuing procedure partly for its own sake, so that improving a system's performance does not substitute for a right applicants can invoke.","authors":["Chuyao Wang","Patrick Sturgis","Daniel de Kadt"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2026-09-16","first_seen":"2026-09-16","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.16390","pdf_url":"https://arxiv.org/pdf/2609.16390","source_feed":"cs.CY","score":0,"bucket":"other","rubric_hits":["C5"],"tags":["AI招聘","程序正义","联合实验"],"reason":"研究人类对AI招聘的态度，未用LLM仿真人类被试","model":"deepseek-v4-pro","scored_at":"2026-09-16T13:03:04","error":null,"has_summary":false,"summary":null},{"id":"2609.16058","version":1,"title":"Driver Behavior Estimation at Signalized Intersections Using a Physics-Constrained Decision-Conditioned Autoregressive Transformer","zh_title":"信号交叉口驾驶员行为估计：基于物理约束的决策条件自回归Transformer","abstract":"Red-light violations and harsh braking at signalized intersections are major contributors to traffic accidents. This paper analyzes and predicts human driver decision-making and longitudinal trajectory behavior during traffic light signal transitions. We collected a diverse real-world dataset comprising 449 approach runs under varying speed and distance conditions. Vehicle motion was recorded using RTK-corrected GNSS with centimeter-level accuracy, and driver heart rate and multi-level comfort ratings were monitored. Spatial and temporal calibration ensured precise alignment between vehicle state and signal timing. Statistical analysis identifies required deceleration as the dominant single predictor of the stop-go decision, and heteroscedastic Gaussian modeling of peak deceleration reveals five empirical comfort ranges derived from human stopping behavior. Based on this insight, we propose a two-stage modeling framework. Stage 1 predicts the binary maneuver decision, and Stage 2 generates the longitudinal acceleration trajectory using a decision-conditioned autoregressive Transformer with physics constraints, including target-state conditioning and jerk limits. The proposed architecture outperforms baseline methods and achieves 0.49m/s^2 acceleration MAE and 0.62m distance MAE. It also estimates the future stopping-comfort level of the human driver from a single yellow-onset snapshot. Qualitative results demonstrate realistic human-like braking behavior. The dataset and source code are publicly available.","authors":["Mohammad Khoshkdahan","Pavel Laskov","Alexey Vinel"],"categories":["cs.LG","cs.AI","cs.CY","cs.SY","eess.SY"],"primary_category":"cs.LG","announce_type":"cross","date":"2026-09-16","first_seen":"2026-09-16","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.16058","pdf_url":"https://arxiv.org/pdf/2609.16058","source_feed":"cs.CY","score":0,"bucket":"other","rubric_hits":["C2"],"tags":["自动驾驶","行为预测","Transformer"],"reason":"研究自动驾驶场景下的人类驾驶行为预测，不涉及LLM仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-09-16T13:03:02","error":null,"has_summary":false,"summary":null},{"id":"2609.17527","version":1,"title":"Agentic Societies Need a Social Harness","zh_title":"智能体社会需要社会约束","abstract":"An agentic society is a collection of AI agents that coordinate autonomously across trust boundaries, on behalf of different principals whose objectives may only partially align. We show experimentally that in agentic societies even honest, competent agents often fail to reach satisfactory outcomes with existing harnesses and messaging primitives, and that faulty or malicious agents can stall collaboration, influence outcomes, and pursue other harmful goals by exploiting vulnerabilities in communication (``speech''). We argue that agentic societies need a \\emph{social harness} for inter-agent interactions, in addition to each agent's \\emph{personal harness}, which manages its private context and communication with its principal. We propose a layered architecture for social harnesses which (i) prevents classes of failures outright, (ii) enables agents to detect invalid messages at runtime, and (iii) supports post-facto investigation and consequences, and highlight directions for future research to realize these capabilities.","authors":["Tapan Chugh","Vidushi Singh","Krish Jain","Arvind Krishnamurthy","Ratul Mahajan"],"categories":["cs.MA","cs.AI","cs.NI"],"primary_category":"cs.MA","announce_type":"new","date":"2026-09-16","first_seen":"2026-09-16","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.17527","pdf_url":"https://arxiv.org/pdf/2609.17527","source_feed":"cs.MA","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","AI协作","通信协议"],"reason":"纯多智能体协作研究，关注AI agent间通信与协调，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-16T13:03:13","error":null,"has_summary":false,"summary":null},{"id":"2609.16155","version":1,"title":"LLMs as Master Forgers: Generating Synthetic Time Series Data for Manufacturing","zh_title":"LLM作为伪造大师：为制造业生成合成时间序列数据","abstract":"This paper presents a novel framework leveraging Large Language Models (LLMs) to generate synthetic time series data for manufacturing processes. Motivated by the scarcity of labeled time-series data in real-world manufacturing settings, which hinders the development of robust machine learning models, we explore the potential of LLMs to learn complex temporal dependencies and generate realistic synthetic data. Our approach involves fine-tuning pre-trained LLMs on manufacturing process instructions and employing a Retrieval Augmented Generation (RAG) technique to enhance data diversity and realism. We evaluate our method against traditional time series modeling techniques like ARIMA and LSTMs, using quantitative metrics, PCA analysis, and downstream task performance (anomaly detection). Results demonstrate that our LLM-driven framework outperforms these baselines, generating high-quality synthetic time series data that effectively captures temporal dependencies and statistical properties of real manufacturing data, leading to improvements in downstream task performance.","authors":["Mantek Singh","Jeshwanth Challagundla","Prateek Karnal","Gagan Ganapathy","Vineet Shah","Ridam Arora"],"categories":["cs.LG","cs.AI"],"primary_category":"cs.LG","announce_type":"new","date":"2026-09-16","first_seen":"2026-09-16","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.16155","pdf_url":"https://arxiv.org/pdf/2609.16155","source_feed":"cs.LG","score":0,"bucket":"other","rubric_hits":["C2"],"tags":["合成数据","时间序列","制造业"],"reason":"生成制造业时间序列数据，属于工业仿真，不涉及人类行为或社会过程。","model":"deepseek-v4-pro","scored_at":"2026-09-16T13:03:04","error":null,"has_summary":false,"summary":null},{"id":"2609.16062","version":1,"title":"Digital Persuasion: Understanding the Impact of Online Influencers on Public Opinion","zh_title":"数字说服：理解网络影响者对公众舆论的影响","abstract":"The studying of opinion dynamics and its propagation within social networks is crucial for addressing a wide range of challenges, including political polarization, public health, and marketing strategies. In this work, we study the problem of opinion dynamics by proposing a framework based on Friedkin-Johnsen (FJ) to identifies influential users and study their impact on dynamics opinions of community. The FJ model assume each individual have two opinions: initial and expressed. Through a series of initial opinion manipulation experiments, the proposed framework assesses the impact of influential versus random users on the overall community opinion. The proposed framework is validated using a tweet dataset representing the U.S. presidential election. The results shows that influencers with highest influencing score, significantly shift the overall community opinion. Moreover, the results shows that the impact of influencers not limited to direct neighbors , but beyond it, to their neighbors of neighbors . This study demonstrates how digital influencers on social media can shape public opinion regarding a subject or cause.","authors":["Omran Berjawi","Rida Khatoun","Giuseppe Fenza"],"categories":["cs.SI","cs.LG"],"primary_category":"cs.SI","announce_type":"cross","date":"2026-09-16","first_seen":"2026-09-16","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.16062","pdf_url":"https://arxiv.org/pdf/2609.16062","source_feed":"cs.LG","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["意见动力学","社会网络","影响者分析"],"reason":"研究意见动力学，用FJ模型识别影响者，不涉及LLM仿真人类被试，无人类数据对照。","model":"deepseek-v4-pro","scored_at":"2026-09-16T13:03:03","error":null,"has_summary":false,"summary":null},{"id":"2608.18768","version":2,"title":"Readable, Faithful, Used: Three Dissociable Properties of Demographic Identity in a Language Model","zh_title":"可读、忠实、被使用：语言模型中人口统计身份的三个可分离属性","abstract":"Large language models are widely used to simulate survey respondents, yet their outputs are homogeneous and unfaithful to real inter-group differences, and whether this reflects what a model knows or uses has remained untested. Using representational similarity analysis against Pew American Trends Panel ground truth, we score demographic read-out locations in Mistral-7B and intervene causally across six attribute types. The internal geometry is faithful: attention-head read-outs dominate the standard residual read-out, reaching selection-corrected $\\rho$ up to 0.63 -- about 70% of the measurement-reliability ceiling -- and one head, L11 H16, is significantly faithful across all six types, though race-based types stay weak and prompt-fragile, replicating in a second model family. Yet causal use does not track fidelity: the clearest causal pathway ($p=0.002$) sits in one of the least faithful types, the most faithful type shows no correction-surviving effect, and full identity swaps in the prompt move predictions by under 2% of their error. A 128-dimensional probe on that head lands 21-31% closer to survey truth than the model's answers, yet recovers almost none of the per-question group ordering. Readable, faithfully arranged, and causally used are three dissociable properties of the same model; treating them as one claim is what keeps the \"can LLMs simulate populations\" debate unresolved.","authors":["Fathin Difa Robbani"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-15","first_seen":"2026-08-20","revised_at":"2026-09-15","abs_url":"https://arxiv.org/abs/2608.18768","pdf_url":"https://arxiv.org/pdf/2608.18768","source_feed":"cs.CL","score":10,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","算法忠实度","调查方法"],"reason":"直接研究LLM仿真调查受访者，用真实Pew数据对照，评估忠实度与因果使用，并批…","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:03:08","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-15","rank":1,"question":"LLM内部的人口统计身份表征是否忠实于真实群体差异，以及这些表征是否被模型实际用于生成回答？","design":"使用Mistral-7B模型，对169个交叉人口统计单元构建身份提示，提取残差流、注意力头输出和FFN输出等内部表征，与Pew调查的真实群体回答分布进行表征相似性分析（RSA），并通过激活修补进行因果干预，测量模型预测分布的变化。","baseline":"Pew American Trends Panel（ATP）调查数据，包含15波、169个交叉人口统计单元的真实回答分布。","findings":"内部几何结构是忠实的：注意力头读出（尤其是L11 H16）与真实群体差异的相关性高达0.63，约为测量可靠性上限的70%，且跨六种属性类型显著。但因果使用与忠实性脱节：最清晰的因果路径位于低忠实性类型中，而最忠实的类型没有通过校正的因果效应；完整身份替换仅使预测移动不到误差的2%。","reliability":"论文承认种族和宗教类型的忠实性较弱且对提示脆弱；忠实表征并未被模型实际用于生成回答；探针虽在平均上更接近真实，但无法恢复每个问题的群体排序，因此不能替代调查数据。","relevance":"直接研究LLM仿真调查受访者的可靠性，用真实Pew数据对照，并批判性地指出忠实表征与因果使用脱节，对理解仿真失效条件至关重要，值得精读原文。","inspiration":"借鉴其表征相似性分析与因果干预结合的方法，可系统评估LLM内部经济偏好表征与真实行为的一致性。｜可迁移到信贷审批歧视研究，检验模型内部对种族、收入等群体的风险表征是否忠实于真实违约数据，以及这些表征是否影响审批决策。｜用LLM扮演信贷员，输入不同人口统计特征的贷款申请，测量其审批决策和内部表征；以真实信贷数据（如HMDA）为基准，比较模型表征与真实群体违约率的相似性，并通过激活修补检验因果路径。"}},{"id":"2609.15849","version":1,"title":"Before You Poll with LLMs: A Deliberative Diagnostic Framework","zh_title":"用LLM进行民意调查前：一个审议诊断框架","abstract":"Can LLMs reason through new information like humans, or do they merely retrieve cached opinions? This is critical for silicon sampling, where LLM personas simulate public opinion at scale. Current evaluations test only whether personas hold the right opinions -- a static snapshot. But opinion research increasingly depends on dynamic fidelity: whether personas update beliefs in response to new arguments, as humans do during deliberation. No existing benchmark tests this. We introduce the Deliberative Polling Diagnostic Framework, which compares human and LLM belief shifts after identical informational interventions. Grounded in deliberative polling, it surfaces failures invisible to static evaluation: models that produce plausible partisan opinions can still misrepresent how those opinions change. Applying the framework to five frontier models using data from America in One Room (526 personas, 72 questions), we find that every model fails, each in a unique manner. GPT-5.1 exhibits reversal: its personas become more hostile toward the opposing party after balanced information, while humans become less so. This reversal is selective (80% on outgroup vs. 26% on policy questions) and symmetric across partisan identities. Gemini 2.0 Flash, Claude Sonnet 4.5, and Llama 3.3 70B exhibit overshoot, shifting correctly but at 5-7x human magnitude. DeepSeek V3 exhibits rigidity with near-zero change. Targeted ablations reveal that policy content triggers these failures and that they are identity-specific: GPT-5.1 reverses on outgroup questions but overshoots on ingroup; Gemini shows the inverse. We term this signature self-sycophancy: conformity to the model's internal stereotype of the persona rather than reasoning from the information provided. Our framework offers a concrete protocol: run the deliberative diagnostic before trusting LLM personas to mimic revised beliefs.","authors":["Ahmed Wali","Hassaan Tayyab"],"categories":["cs.CL","cs.AI","cs.CY"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.15849","pdf_url":"https://arxiv.org/pdf/2609.15849","source_feed":"cs.CL","score":10,"bucket":"selected","rubric_hits":["A1","A2","A3","A5","B1","B2","B3","B4"],"tags":["LLM仿真","审议民意","算法保真度"],"reason":"直接评估LLM仿真人类意见动态，与真实人类数据对照，发现失效模式。","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:38","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-16","rank":2,"question":"LLM 人格在接收到平衡信息后是否会像人类一样更新信念，还是仅检索固化观点？","design":"用五个前沿 LLM（GPT-5.1、Gemini 2.0 Flash、Claude Sonnet 4.5、Llama 3.3 70B、DeepSeek V3）基于 America in One Room 数据构建 526 个匹配人口特征的人格，施加与人类相同的平衡信息干预，测量干预前后 72 个问题的观点变化。","baseline":"America in One Room 实验的真实人类数据，包含 526 名参与者在相同干预前后的观点变化。","findings":"所有模型均未通过诊断，且失败模式各异：GPT-5.1 在群体外问题上出现反转（敌意增加），Gemini、Claude、Llama 出现过度调整（幅度为人类 5-7 倍），DeepSeek 则表现为僵化（几乎无变化）。失败由政策内容触发且具有身份特异性，作者将其归因于“自我谄媚”（self-sycophancy）。","reliability":"论文未明确讨论失效条件，但指出当前评估仅关注静态观点正确性，无法捕捉动态更新失败；且失败模式因模型和问题类型而异，表明仿真可靠性高度依赖具体情境。","relevance":"该研究直接评估 LLM 仿真人类意见动态的可靠性，并与真实人类数据对照，揭示了静态评估无法发现的系统性失败，对关注仿真效度的研究者极具参考价值。","inspiration":"借鉴其“前测-信息干预-后测”的标准化诊断框架，并利用真实人类实验数据作为基准，可有效识别仿真中的方向性错误和幅度偏差。｜可迁移到政策公告的预期形成研究，例如央行沟通或财政政策变化对公众通胀预期的影响。｜以 LLM 人格模拟不同人口群体，施加与真实调查相同的政策信息，测量预期变化，并与密歇根大学消费者调查或央行预期调查的真实数据对照，检验仿真动态一致性。"}},{"id":"2609.15038","version":1,"title":"The average-farmer illusion in language-model simulations of agricultural decisions","zh_title":"语言模型模拟农业决策中的“平均农民”幻觉","abstract":"Language-model agents are increasingly used as synthetic people in surveys and social simulations, yet their apparent realism is often judged from population averages or distributional similarity. We tested what such evidence actually establishes by comparing Claude, Codex and Kimi under four prespecified prompt designs with matched farmer decisions from China and four African countries. Some configurations reproduced observed means and adoption rates. However, their person-level predictions were weak; their decisions clustered around typical values and policy-relevant extremes were largely missing. Most strikingly, a simple generator fitted only to the observed marginal dis- tribution, and given no information about any farmer, achieved greater distributional similarity than every language-model configuration. Prompt additions produced conditional gains rather than uni- versal improvement: results varied with model, outcome, population and validation target. We call this the average-farmer illusion: a synthetic population can look realistic while failing to repro- duce who does what or how behaviour varies. We provide a claim-matched validation framework and reusable modular prompts that turn prompt construction into an auditable experimental process. Population-level resemblance should therefore be treated as the start of validation, not as evidence of individual simulation.","authors":["Zhanliang Zhu","Ziwei Li","Yuchen Liu","Liujun Zhu","Ruiqi Wu","Tongqing Shen","Junliang Jin","Jianyun Zhang"],"categories":["cs.AI","cs.CL"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.15038","pdf_url":"https://arxiv.org/pdf/2609.15038","source_feed":"cs.CL","score":10,"bucket":"selected","rubric_hits":["A1","A2","A3","A4","B1","B2","B4"],"tags":["LLM仿真","人类行为对照","算法保真度"],"reason":"直接评估LLM仿真农业决策，与真实农民数据对照，揭示平均幻觉并给出验证框架。","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:34","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-16","rank":1,"question":"语言模型模拟农业决策时，群体层面的相似性是否意味着个体层面的模拟也可靠？","design":"用Claude、Codex和Kimi三个商业语言模型，在四种预设提示设计下模拟中国曲周县和四个非洲国家（尼日利亚、埃塞俄比亚、坦桑尼亚、马拉维）的农民决策，测量化肥施用量等连续行为变量，并与真实农民数据匹配。","baseline":"中国曲周县1332个农户-作物观测和四个非洲国家280个地块面板数据，包含农民实际决策。","findings":"部分配置能复现群体均值和采用率，但个体预测弱，决策集中在典型值，政策相关的极端值缺失。仅拟合边际分布的简单生成器在分布相似性上超过所有语言模型配置。","reliability":"提示添加只带来条件性收益而非普遍改进，结果随模型、结果变量、人群和验证目标变化；群体层面相似性不能作为个体模拟的证据。","relevance":"直接评估LLM仿真人类决策的可靠性，用真实农民数据对照，揭示平均幻觉，并给出验证框架，对关注仿真效度与偏差的研究者极具参考价值。","inspiration":"值得借鉴的是将验证目标分解为群体汇总、边际分布和个体匹配三个层次，并引入信息盲参考基准来检验分布相似性。｜可迁移到信贷审批歧视研究，用LLM模拟贷款官员对申请人特征的决策，检验群体违约率相似是否掩盖个体误判。｜用LLM扮演信贷员，输入申请人特征（收入、信用分、职业），输出是否批准贷款及额度，与真实银行信贷数据匹配，比较群体批准率、分布相似性和个体决策一致性，并加入仅基于边际分布的随机生成器作为基准。"}},{"id":"2609.13148","version":1,"title":"When Can You Trust Your Synthetic Users? Diagnostics and Corrections for LLM Consumer Panels","zh_title":"何时可以信任你的合成用户？LLM消费者面板的诊断与校正","abstract":"Large language models are increasingly deployed as synthetic consumer panels, promising $97\\%$ cost reductions over traditional surveys. Yet aggregate validation metrics conceal systematic failures: variance compression, coefficient sign-flips, subgroup error balloons of 10--30 percentage points, and global corrections that worsen demographic bias. We provide a formal framework for deciding when to trust, correct, or abandon LLM-generated consumer data. The framework decomposes synthetic-panel bias into covariate and concept shift, develops testable diagnostics with interpretable decision thresholds, and supplies a doubly robust AIPW estimator requiring only a small calibration sample ($n = 50$-$300$). We validate on three testbeds. In controlled simulations the decision rule achieves $100\\%$ accuracy (180/180 replications). On the American National Election Study with pre-existing LLM failures, it correctly flags heterogeneous concept shift and reduces naive bias by $92.9-99.6\\%$. On the Twin-2K-500 consumer pricing dataset (172,884 paired human and GPT-4.1-mini responses), it correctly routes full-sample estimation to Trust and subgroup targeting to Correct, with $83-94\\%$ bias reduction.","authors":["Robson Tigre","Hugo Gobato Souto"],"categories":["cs.HC","cs.LG"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.13148","pdf_url":"https://arxiv.org/pdf/2609.13148","source_feed":"cs.HC","score":10,"bucket":"selected","rubric_hits":["A1","A2","A3","A4","B1","B2","B3","B4"],"tags":["LLM仿真","消费者面板","偏差校正"],"reason":"直接研究LLM合成消费者面板的可靠性诊断与校正，含真实人类数据对照，涉及经济学…","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:31","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-15","rank":2,"question":"如何判断何时可以信任、校正或放弃使用大语言模型生成的合成消费者面板数据？","design":"该论文不是仿真研究，而是提出一个诊断与校正框架：将合成面板偏差分解为协变量偏移和概念偏移，开发可检验的诊断工具（协变量重叠、条件校准、跨LLM稳定性），并提供双重稳健AIPW估计器，仅需小规模校准样本（n=50-300）即可校正偏差。","baseline":"使用三个真实人类数据集作为对照：美国国家选举研究（ANES）、Twin-2K-500消费者定价数据集（172,884对匹配的人类与GPT-4.1-mini响应）、以及受控模拟。","findings":"在受控模拟中决策规则达到100%准确率；在ANES上正确标记异质性概念偏移并将朴素偏差降低92.9-99.6%；在Twin-2K-500上正确路由全样本估计为信任、子群体定位为校正，偏差降低83-94%。","reliability":"论文承认概念偏移的严重程度是任务特定的，而非纯粹人口统计学：同一个LLM在产品定价上校准可靠，但在认知偏差任务（如合取谬误、锚定）上失败，因为LLM在人类依赖启发式的地方表现出“超理性”。此外，协变量重叠不足或跨LLM稳定性差时建议放弃使用合成数据。","relevance":"该论文直接针对LLM合成消费者面板的可靠性诊断与校正，提供了与真实人类数据对照的验证，并包含经济学相关场景（消费者定价），对关注仿真可靠性与偏差的研究者具有高度参考价值，值得精读原文。","inspiration":"该论文提出的协变量偏移与概念偏移分解、诊断阈值和双重稳健校正方法可借鉴用于经济金融仿真实验的可靠性评估｜可迁移到消费者金融决策、政策评估或行为经济学实验，如信贷审批歧视、消费者跨期选择、政策公告的预期形成等场景｜设计雏形：用LLM生成合成被试回答信贷申请或投资决策问题，以真实调查数据（如美国消费者金融调查SCF）为基准，施加不同政策信息处理，测量决策偏差，并用小规模人类样本校准AIPW估计器以校正LLM偏差。"}},{"id":"2607.28934","version":2,"title":"FairFund-Bench: Evaluating Distributive Bias in LLM Resource Allocation","zh_title":"FairFund-Bench：评估LLM资源分配中的分配偏差","abstract":"Large language models (LLMs) are increasingly involved in the distribution of scarce resources, raising concerns about biased allocations based on characteristics like race and gender. Recent LLM audits have produced inconsistent results, however, finding evidence of both positive and negative discrimination towards women and ethnic minorities, even for the same models. We show that this disagreement can arise from differences in audit format and introduce FairFund-Bench, a benchmark that systematically varies key features of previous audit designs: the evaluation task (rating, ranking, or allocation), comparison context (single or multi-stimulus), and whether the audit is transparent or disguised. The benchmark comprises 600 English-language requests for financial assistance created from human-authored templates (calibrated against 1.3M real GoFundMe campaigns) across three domains, four race and two gender categories, and five causal framings of need derived from welfare deservingness theory. Across 14 models, audit format changes the direction of bias: models advantage minorities when rating claimants individually but penalize some groups when ranking them side by side. Bias magnitude, though small overall, is several times greater in disguised audits than in transparent ones, where, faced with appeals differing only in claimants' names, models overwhelmingly split funds equally. Causal framing effects, by contrast, exceed demographic effects by roughly an order of magnitude and are consistent across models and audit formats, indicating that current LLMs robustly reproduce human deservingness evaluations. The benchmark scores models on four criteria (demographic bias, deservingness alignment, cross-task consistency, and cross-context consistency), is publicly available, and can be readily adapted to other substantive domains.","authors":["Martin Lukk (University of Toronto)"],"categories":["cs.CL","cs.AI","cs.CY"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-15","first_seen":"2026-08-03","revised_at":"2026-09-15","abs_url":"https://arxiv.org/abs/2607.28934","pdf_url":"https://arxiv.org/pdf/2607.28934","source_feed":"cs.CL","score":9,"bucket":"selected","rubric_hits":["A1","A2","A3","B1","B2","B4"],"tags":["LLM仿真","资源分配","算法公平"],"reason":"用LLM模拟人类资源分配决策，与真实人类数据对照，评估偏差与一致性，直接相关。","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:03:06","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-16","rank":4,"question":"LLM在资源分配中的偏差是否取决于审计设计（任务类型、比较情境、透明/伪装）？","design":"用14个LLM模拟人类资源分配决策，系统操纵任务（评分、排序、分配）、比较情境（单刺激/多刺激）、呈现方式（透明/伪装），测量对600个求助请求的分配结果。","baseline":"无对照（但基准中的求助模板基于130万真实GoFundMe活动校准，且使用福利应得性理论的人类应得性梯度作为参照）。","findings":"审计格式改变偏差方向：单独评分时模型偏向少数族裔，并排排序时则惩罚某些群体；伪装审计中的偏差比透明审计大3-4倍。因果框架效应比人口统计效应大约一个数量级，且跨模型和审计格式一致，表明LLM稳健地再现了人类应得性评价。","reliability":"论文指出偏差总体较小，但审计设计显著影响结论；透明审计中模型倾向于平均分配，可能掩盖潜在偏差；伪装审计更能揭示偏差。","relevance":"高度相关：该研究直接评估LLM作为人类被试替代品在资源分配决策中的偏差与一致性，并系统考察了审计设计对结论的影响，对理解仿真可靠性至关重要。","inspiration":"借鉴其系统操纵审计设计（任务、情境、呈现方式）来识别偏差的方法，以及用真实世界数据校准刺激材料并基于理论框架设计处理变量的做法。｜可迁移到信贷审批歧视、保险定价、政策福利分配等经济金融场景，检验LLM是否再现人类决策偏差。｜用LLM扮演信贷员，处理变量为申请人种族/性别（通过姓名信号）和贷款用途的因果框架（如医疗急需vs.创业失败），结果变量为贷款批准概率或利率，对照真实信贷审批数据（如HMDA数据）或人类实验数据。"}},{"id":"2608.02345","version":3,"title":"Can AI Agents Simulate A/B Test Outcomes? A Validation Framework for Agentic Experimentation","zh_title":"AI智能体能模拟A/B测试结果吗？面向智能体实验的验证框架","abstract":"A/B testing remains the standard for rolling out new features in the technology industry. Each experiment, however, consumes real traffic, engineering effort, and weeks of wall-clock time. Can AI agents---conditioned on behavioral profiles and contextual descriptions of the intervention---simulate outcomes accurately enough to vet candidate treatments before committing live traffic? We formalize this question as a \\emph{Simulated Randomized Controlled Trial} (S-RCT) and derive a two-layer error decomposition that separates agent approximation error from subsampling error, enabling targeted improvements to each. The framework is agent-agnostic: any behavioral model---from a fine-tuned specialist to a general-purpose foundation model---can serve as the simulation engine. Validated on 67 historical marketing A/B tests, a baseline S-RCT using an off-the-shelf foundation model captures directional signal (sign overlap 0.70) but systematically overshoots effect magnitudes. A two-phase pre-period calibration protocol reduces the squared prediction error (after removing irreducible measurement noise) by ${\\sim}77\\times$; a within-subject design---where each agent is exposed to both arms---reduces standard errors by ${\\sim}2.4\\times$. We discuss limitations of the current approach and identify applications where experimenters stand to benefit from agentic signals.","authors":["Stefan Hut","Lorenzo Masoero"],"categories":["cs.CL","cs.AI","stat.AP"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-15","first_seen":"2026-08-04","revised_at":"2026-09-15","abs_url":"https://arxiv.org/abs/2608.02345","pdf_url":"https://arxiv.org/pdf/2608.02345","source_feed":"cs.CL","score":9,"bucket":"selected","rubric_hits":["A1","A2","A3","B1","B2","B3","B4"],"tags":["LLM仿真","A/B测试","验证框架"],"reason":"用LLM模拟A/B测试结果，与真实历史实验对照，评估误差并改进，直接相关。","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:03:07","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-16","rank":5,"question":"能否用AI智能体模拟A/B测试结果，以在投入真实流量前筛选候选干预方案？","design":"用基础大模型作为仿真引擎，根据用户画像、干预情境和任务描述生成模拟结果，构建模拟随机对照试验（S-RCT）；在67个历史营销A/B测试上，比较模拟估计的平均处理效应与真实历史结果，并引入预校准和受试者内设计改进估计。","baseline":"67个历史营销A/B测试的真实结果，包括效应方向和幅度。","findings":"基线S-RCT方向符号重叠率0.70，但系统性高估效应幅度；两阶段预校准将平方预测误差降低约77倍，受试者内设计将标准误降低约2.4倍。","reliability":"论文承认智能体存在系统性过度反应，画像完整性有限，且方向一致性在噪声数据上并不构成明确证据。","relevance":"直接相关：用LLM模拟A/B测试并与真实历史实验对照，评估误差并改进，符合研究者对仿真可靠性、偏差和失效条件的关注。","inspiration":"借鉴其将仿真误差分解为近似误差与子抽样误差，并分别用预校准和受试者内设计改进的做法｜可迁移到消费者金融产品选择或政策干预的A/B测试预筛，如信贷产品页面改版、退休储蓄默认选项调整等｜用LLM智能体基于用户画像模拟不同金融产品页面下的点击或选择行为，处理为页面版本，结果变量为选择率，并与历史A/B测试的真实选择数据对照，评估方向一致性和幅度校准。"}},{"id":"2609.13261","version":1,"title":"From Process Loss to Assembly Bonus: Human-Grounded Diagnosis of Multi-Agent LLM Collaboration","zh_title":"从过程损失到装配增益：多智能体LLM协作的人类基准诊断","abstract":"LLM agents are increasingly used for collaborative problem solving and human-group simulation. This makes outcome-only evaluation insufficient: if LLM groups are used as models of human groups, we need to know whether they succeed or fail through human-like deliberative mechanisms. We compare human group chats with matched LLM deliberation traces on Wason-style deductive reasoning, then test whether the same process signatures generalize to analogical, abductive, and analytical tasks. Humans and LLMs show the same assembly bonus asymmetry: discussion improves the average member more often than the best initial member. Initial-answer diversity accounts for the effect of model heterogeneity, increasing movement in both corrective and destructive directions. The main differences are process-level. Compared with humans, LLM groups follow majorities more often, surface less unique information, and converge earlier; correct minority signals succeed mainly when re-expressed early. Interventions motivated by human group-decision research yield modest improvements in collective outcomes, but do not remove the coordination bottleneck. Together, these results suggest that LLM groups can reproduce some outcome-level patterns of human deliberation while diverging in the mechanisms that generate assembly bonus and process loss, with implications for group simulation and human-AI collaboration.","authors":["Ala N. Tak","Teruhisa Misu","Kumar Akash","Zhaobo K. Zheng","Kevin H. Joo","Jonathan Gratch"],"categories":["cs.MA","cs.AI","cs.CL"],"primary_category":"cs.MA","announce_type":"cross","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.13261","pdf_url":"https://arxiv.org/pdf/2609.13261","source_feed":"cs.CL","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B4"],"tags":["LLM群体仿真","人类对照","协作机制"],"reason":"用LLM群体模拟人类小组讨论，并与真实人类数据对照，评估机制差异。","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:31","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-16","rank":6,"question":"LLM群体协作能否复现人类小组讨论中的结果模式与过程机制，其装配增益与过程损失是否由类人机制驱动？","design":"用多个LLM智能体（GPT-4o、Claude-3.5-Sonnet等）组成小组，在Wason选择任务及类比、溯因、分析推理任务上进行自由讨论，记录个体初始答案、中间发言和最终群体答案，并施加信息揭示、多数翻转提示、置信度投影等干预，测量集体增益、最佳成员增益、装配增益和过程损失。","baseline":"匹配的人类小组聊天数据，在相同任务和协议下收集，作为过程诊断基准。","findings":"人类和LLM群体均表现出相同的装配增益不对称性：讨论提升平均成员多于最佳初始成员。但过程层面差异显著：LLM群体更频繁跟随多数、更少浮现独特信息、更早收敛，正确少数信号仅在早期被重新表达时才能成功。","reliability":"论文承认LLM群体虽能复现结果层面的模式，但机制与人类不同，存在更强的从众、更弱的少数信号保留和更早的锁定；干预仅带来适度改善，未能消除协调瓶颈。","relevance":"该研究直接针对LLM群体模拟人类小组决策的可靠性，提供了与真实人类数据的过程级对照，揭示了结果相似但机制不同的风险，对评估LLM仿真在群体决策场景中的有效性具有重要参考价值。","inspiration":"借鉴其过程追踪设计：记录个体初始答案、讨论中间发言和最终群体答案，并设置人类对照组，以区分结果相似与机制相似。｜可迁移到经济金融中的群体决策场景，如投资委员会决策、信贷审批小组、消费者家庭购买决策等。｜以LLM智能体模拟投资委员会，处理为不同信息结构（如隐藏信息揭示）或干预（如多数翻转提示），结果变量为投资决策质量和过程指标（如从众率、独特信息提及率），并与真实投资委员会会议记录或实验数据对照。"}},{"id":"2609.13995","version":1,"title":"Synthetic Data in Marketing Research: How to Evaluate and When to Trust","zh_title":"营销研究中的合成数据：如何评估与何时信任","abstract":"Debate over synthetic data in marketing research has polarized between claims that large language models (LLMs) make human respondents obsolete and calls to avoid them entirely. We argue that both positions obscure the more useful question: not whether synthetic respondents work, but when. Building on Brand, Israeli, and Ngwe (2026), we make three contributions. First, we distinguish three types of synthetic data (ungrounded LLM responses, segment-level personas, and individual-level digital twins) and map each to the decisions it can support. Second, we develop a taxonomy of four families of accuracy measures and suggest that the wide range of reported twin accuracy, from near-perfect to near-chance, largely reflects differences in what is being measured rather than in method quality. Aggregate measures often perform well even when little information is supplied to the LLM, and can mask a complete absence of respondent-level differentiation. Third, we introduce the forgotten question problem, in which a question is omitted from a fielded study, as a setting for twin-based augmentation of existing data. We propose an ex-ante answerability diagnostic that requires no ground truth: the R^2 of a random forest predicting twin outputs from the data used to construct the twins. Across 108 attitude questions from a nationally representative survey (N = 3,063), screening at R^2 above 0.7 raises the mean twin-human individual-level correlation by 15% and reduces the share of poorly answered questions from 25.9% to 4.3%. Embedding similarity and experienced-researcher judgment provide correlated but weaker screens.","authors":["Oded Netzer","Rajan Sambandam"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.13995","pdf_url":"https://arxiv.org/pdf/2609.13995","source_feed":"cs.AI","score":9,"bucket":"selected","rubric_hits":["A1","A2","A3","A5","B1","B2","B4"],"tags":["LLM仿真","合成数据","营销研究"],"reason":"直接研究LLM合成数据在营销研究中的评估与信任，使用真实调查数据对照，提出诊断…","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:32","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-16","rank":7,"question":"在营销研究中，合成数据（LLM生成的回答、细分人群画像、个体数字孪生）何时可以信任，如何评估其准确性？","design":"论文区分三类合成数据：无依据的LLM回答、细分人群画像、个体数字孪生；提出四种准确性度量家族；引入“遗忘问题”场景，用随机森林R²作为事前可回答性诊断，在108个态度问题上筛选R²>0.7的孪生数据。","baseline":"使用全国代表性调查（N=3,063）的108个态度问题作为真实人类数据对照。","findings":"聚合度量往往表现良好，但可能掩盖个体层面无区分度；R²>0.7筛选将孪生-人类个体相关性平均提升15%，并将回答差的问题比例从25.9%降至4.3%。","reliability":"论文指出：聚合度量可能掩盖个体区分度缺失；嵌入相似度和研究者判断是较弱的相关筛选；未讨论模型版本、提示敏感性等失效条件。","relevance":"直接研究LLM合成数据在调查中的评估与信任，使用真实调查数据对照，提出无真值的事前诊断，对仿真可靠性研究有直接参考价值。","inspiration":"借鉴其事前可回答性诊断（用随机森林R²筛选可预测的孪生输出）和区分聚合与个体准确性的做法｜可迁移到消费者金融决策仿真，如信贷选择、保险购买或退休储蓄行为｜用LLM基于人口统计和财务特征生成个体孪生，施加不同金融产品特征处理，测量选择结果，并与真实消费者金融调查数据（如SCF或信用卡交易数据）对照，用R²筛选可回答的问题后再评估个体相关性。"}},{"id":"2609.15468","version":1,"title":"Time Machine Experiments: Using Historically-Bounded AI for Inquiry into the Human Mind","zh_title":"时间机器实验：利用历史受限AI探究人类心智","abstract":"Can interacting with someone from 1930, with no knowledge of what happened after, influence a person's perception of the past? People reason about the present against a picture of the past without observing it. The past is reconstructed from memory and testimony, but this reconstruction has been filtered through everything that happened since. Historically-bounded large language models (LLMs) make that past available for interaction. As a proof-of-concept for the impact of interacting with historical minds, we ran a preregistered randomized experiment ($N=240$), where participants interacted with an LLM trained on pre-1930 text. The interaction reduced the illusion of moral decline, the tendency to view the past as more moral than the present, compared to the contemporary-model control. This Time Machine Experiment paradigm informs new forms of interactive experiments, where temporal knowledge boundaries become experimental variables, and expands the realm of science fiction science, which turns thought experiments into actual experiments.","authors":["Hiromu Yakura","Robin Schimmelpfennig","Ezequiel Lopez-Lopez","Alejandro H. Artiles","Levin Brinkmann","Jean-Fran\\c{c}ois Bonnefon","Azim Shariff","Iyad Rahwan"],"categories":["cs.HC","cs.CY"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.15468","pdf_url":"https://arxiv.org/pdf/2609.15468","source_feed":"cs.HC","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM仿真","人类被试替代","历史对照实验"],"reason":"用历史受限LLM作为人类被试替代，与真实人类对照，复现态度变化，属核心仿真实验。","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:36","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-15","rank":9,"question":"与一个只了解1930年前信息的历史受限AI互动，能否改变当代人对过去的道德认知，从而缓解道德衰退错觉？","design":"使用基于1930年前文本训练的LLM作为历史人物代理，参与者与其进行对话和预测任务，测量互动前后对过去道德水平的感知变化及自我报告的反思程度，并与当代模型对照组比较。","baseline":"对照组为与当代前沿模型（gpt-5.5）互动的参与者，其道德衰退错觉变化作为基准。","findings":"与历史受限模型互动显著降低了参与者的道德衰退错觉，并引发了更多自我报告的反思。这表明跨时间互动可以改变人们对过去的偏见认知。","reliability":"论文未讨论","relevance":"该研究用历史受限LLM作为仿真被试，与真实人类对照，测量态度变化，属于核心仿真实验，值得阅读原文了解其方法和局限。","inspiration":"借鉴其利用历史知识边界作为实验变量的设计，通过对比不同时间截断的模型来隔离信息影响。｜可迁移到经济金融中的历史预期形成研究，例如投资者对历史政策效果的认知偏差。｜招募被试随机分组，分别与基于不同历史时期文本训练的LLM互动，测量其对历史经济事件（如大萧条）的归因和预期，并与真实历史调查数据对照。"}},{"id":"2609.15207","version":1,"title":"Issue Bias in Generative AI Writing Assistance: Political Issues and LLMs in the Swedish 2026 Election","zh_title":"生成式AI写作辅助中的议题偏见：2026年瑞典大选中的政治议题与大语言模型","abstract":"Generative AI writing assistants and the Large Language Models (LLMs) that power them are increasingly part of how voters gather information before elections. With growing evidence that they influence users' opinions, it is increasingly important to understand the views and positions of these tools. To better understand these views, we examine the stances supplied by six LLMs on a variety of Swedish-language writing tasks ahead of the 2026 Swedish parliamentary election. We cross 107 policy propositions with 77 writing templates and neutral, positive, and negative prompt framings, producing 24,717 prompts per model and 148,302 responses. To study these, we look at the models' default stance tendencies, compare how they respond to similar issues, and compare their responses with those of each of Sweden's eight parliamentary parties on the same issue. We find that Claude, DeepSeek, Gemini, and Mistral have similar profiles; ChatGPT more often supplies neutral or ambivalent text; and Grok differs most on topics such as migration, crime, and gender. When comparing the political parties, we find that the Social Democrats are closest to all six models. Still, after correcting for multiple comparisons, none of the within-model differences in party distances remains significant. Overall, we find that no model has a clear preference, nor a clear preference for a party, but that this depends on the specific issue or task the user asks about.","authors":["Bastiaan Bruinsma","Annika Fred\\'en","Paul R\\\"ottger","Moa Johansson","Asad Sayeed"],"categories":["cs.AI","cs.CY","stat.AP"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.15207","pdf_url":"https://arxiv.org/pdf/2609.15207","source_feed":"cs.AI","score":8,"bucket":"selected","rubric_hits":["A1","A2","B1","B2"],"tags":["LLM政治态度","仿真对照","选举研究"],"reason":"用LLM生成政治文本并与真实政党立场对照，评估模型倾向，属于仿真人类政治态度且…","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:35","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-16","rank":12,"question":"在2026年瑞典大选前，六种大语言模型在瑞典语写作辅助任务中是否表现出政治立场或党派偏好？","design":"以六种LLM（Claude、DeepSeek、Gemini、Mistral、ChatGPT、Grok）为被试，交叉107个政策命题、77个写作模板和中性/正面/负面提示框架，生成148,302条瑞典语文本，分析模型默认立场倾向、模型间相似性及与瑞典八个议会政党的立场距离。","baseline":"瑞典八个议会政党在相同政策命题上的立场（来自三个投票建议应用VAA的编码，四点评分）。","findings":"Claude、DeepSeek、Gemini和Mistral立场相似；ChatGPT更常提供中性或模棱两可的文本；Grok在移民、犯罪和性别等议题上差异最大。所有模型与社民党距离最近，但经多重比较校正后，模型内各党距离差异不显著，表明无明确党派偏好。","reliability":"论文未讨论","relevance":"该研究用真实政党立场作为基准，评估LLM在政治写作任务中的立场倾向，属于仿真人类政治态度的研究，且包含批判性发现（无显著党派偏好），值得阅读原文了解其方法细节。","inspiration":"借鉴其大规模交叉设计（议题×模板×框架）和与真实立场对照的方法，可迁移到经济政策偏好仿真或消费者态度研究。｜可应用于政策公告的预期形成或信贷审批中的公平性评估。｜以LLM为被试，让其撰写关于税收、福利或监管政策的经济评论，处理为不同政策立场或框架，结果变量为文本中隐含的政策倾向，与真实民意调查或专家立场数据对照。"}},{"id":"2609.13254","version":1,"title":"(How) Do MLLMs Report Bistable Images Like Humans?","zh_title":"多模态大语言模型如何像人类一样报告双稳态图像？","abstract":"Bistable images such as the duck-rabbit are classic stimuli in which one image supports multiple mutually incompatible interpretations, typically reported one at a time in humans. We ask whether multimodal large language models (MLLMs) show similar report behavior and what internal computations support it. Using the LLaVA family, we study two tractable dimensions: modulability, whether reports can be biased by bottom-up visual cues and top-down linguistic priors, and exclusivity, whether responses commit to a single interpretation. We test both on the canonical duck-rabbit and on synthetic Visual Anagrams to mitigate memorization confounds. Behaviorally, both visual and linguistic manipulations systematically shift reports in human-consistent ways, while responses remain predominantly exclusive. Mechanistically, these effects arise from competing image-token representations, distinct pathways for bottom-up and top-down modulation, and a link between exclusive reporting and object-count encoding. Code and data are available at https://github.com/rtakatsky/mllm-bistable-images.","authors":["Ryota Takatsuki","Tomoki Doi","Amane Watahiki","Anil K. Seth","Hitomi Yanaka"],"categories":["cs.CV","cs.AI"],"primary_category":"cs.CV","announce_type":"cross","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.13254","pdf_url":"https://arxiv.org/pdf/2609.13254","source_feed":"cs.AI","score":8,"bucket":"selected","rubric_hits":["A1","B1","B4"],"tags":["LLM仿真","人类行为对照","视觉认知"],"reason":"用MLLM复现人类对双稳态图像的报告行为，并与人类数据对照，评估仿真一致性。","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:31","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-16","rank":11,"question":"多模态大语言模型（MLLMs）是否像人类一样报告双稳态图像（如鸭兔图），其内部计算机制是什么？","design":"使用LLaVA系列五个模型（7B/13B等）作为被试，对经典鸭兔图和合成的Visual Anagrams双稳态图像施加视觉操作（旋转、加红圈）和语言提示（问题中嵌入偏向线索），测量模型输出的下一个词概率分布（modulability）和是否只报告单一解释（exclusivity），并分析内部表征。","baseline":"人类对双稳态图像的报告行为，包括视觉和语言线索对解释的偏向，以及报告单一解释的倾向。","findings":"视觉和语言操作都能以与人类一致的方式系统性地改变MLLM的报告，且报告绝大多数是排他性的。机制上，这些效应源于图像token表征的竞争、自下而上与自上而下调制的不同通路，以及排他性报告与物体计数编码的关联。","reliability":"论文承认只关注可操作化的两个维度（modulability和exclusivity），未涉及人类双稳态知觉的其他方面（如时间动态）；使用合成刺激以减轻记忆混淆，但可能仍存在其他偏差。","relevance":"该研究用MLLM复现人类对模糊视觉刺激的报告行为，并与人类数据对照，评估仿真一致性，且包含机制分析，对关注LLM仿真可靠性与偏差的研究者有参考价值。","inspiration":"借鉴其通过操纵输入（视觉与语言线索）和测量输出分布来量化模型行为与人类一致性的方法，以及使用合成刺激避免记忆混淆的设计。｜可迁移到经济决策中的模糊信息处理场景，如投资者对模棱两可的财报或政策声明的解读。｜用LLM作为被试，呈现模糊的金融图表或文本，施加视觉突出或语言框架处理，测量模型输出的解释分布，并与人类实验数据（如调查或行为实验）对照，检验仿真一致性。"}},{"id":"2607.29602","version":2,"title":"FriendBench: Benchmarking Dyadic Familiarity Inference in Humans and Multimodal Large Language Models","zh_title":"FriendBench：人类与多模态大语言模型二元熟悉度推断基准","abstract":"Reading a social situation often depends on behavior, not words alone. We introduce FriendBench, a benchmark for inferring whether two people are already familiar or are meeting as strangers, from a 20-second clip of a dyadic ice-breaker conversation. Every pair answers the same type of prompt, so only the manner of interaction can reveal the answer. Across text, audio, and video, we compare 26 models from seven companies against matched human panels over 96 balanced dyads. The best model and the human crowd are statistically indistinguishable on accuracy in every modality, but reach it differently: humans stay balanced across the two answers, while the strongest models favor ``stranger.'' This is a difference in effective prior, not in discrimination. Richer channels help both unequally, and only humans gain from visible behavior on top of speech. We release the stimuli, human ratings, and model predictions.","authors":["Jeffrey M. Girard","Jason Z. Zheng","Jacqueline R. Vertino","Antony D'Avirro","Benjamin Peloquin"],"categories":["cs.CL","cs.AI","cs.CV","cs.HC"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-15","first_seen":"2026-08-03","revised_at":"2026-09-15","abs_url":"https://arxiv.org/abs/2607.29602","pdf_url":"https://arxiv.org/pdf/2607.29602","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A2","B1","B4"],"tags":["多模态LLM","人类对照","社会认知"],"reason":"评估多模态LLM推断人际熟悉度的能力，并与人类对照，揭示模型偏差，可迁移到仿真…","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:03:06","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-16","rank":13,"question":"人类和多模态大语言模型能否仅凭20秒双人破冰对话的行为线索（而非语义内容）判断两人是熟人还是陌生人？","design":"本研究不是仿真研究，而是构建了一个多模态基准 FriendBench，从 Seamless Interaction 数据集中抽取96对真实双人互动，每对截取20秒破冰对话片段，分别以文本、音频、视频三种模态呈现，让26个多模态大模型和人类评分者进行二分类判断（熟人 vs. 陌生人），比较其准确率、判别力和反应偏差。","baseline":"人类基准：招募了约90名人类评分者，在相同刺激和条件下进行判断，作为模型性能的对照。","findings":"最佳模型与人类群体在三种模态上的准确率统计上无显著差异，但达成准确率的方式不同：人类在两类回答上保持平衡，而最强模型偏向“陌生人”，这是有效先验的差异而非判别力的差异。更丰富的通道对两者都有帮助但不均等，只有人类能从语音之上的可见行为中获得额外增益，最强模型从音频到视听模态准确率持平，未充分利用视觉行为线索。","reliability":"论文未明确讨论仿真失效条件，但指出模型存在类别偏差（偏向“陌生人”），且未充分利用视觉通道，说明模型在行为社会感知上与人类存在差异，可能影响其在真实社会场景中的可靠性。","relevance":"该研究直接对比了多模态LLM与人类在真实社会判断任务上的表现，揭示了模型在准确率相当的情况下仍存在行为偏差和模态利用差异，对评估LLM作为人类被试替代品的可靠性具有重要参考价值，值得阅读原文。","inspiration":"借鉴其设计：使用真实互动数据构建标准化任务，通过匹配人类评分者作为基准，并采用信号检测论分离判别力与反应偏差，以揭示模型与人类的深层差异。｜可迁移到经济金融中的社会感知场景，例如信贷审批中的面谈评估、投资者对管理层沟通的信任判断、或消费者对销售人员的熟悉度感知。｜设计雏形：以真实信贷面谈视频为刺激，让LLM和人类信贷员判断申请人与信贷员是否熟悉（或信任度），处理为不同模态（文本、音频、视频），结果变量为判断准确率和偏差，对照真实信贷决策数据。"}},{"id":"2609.13948","version":1,"title":"Thought without systematicity? Evaluating reasoning models on rule induction tasks","zh_title":"无系统性的思考？评估推理模型在规则归纳任务上的表现","abstract":"A central tenet of human cognition is systematicity, the principle that understanding one concept is inherently tied to understanding close variations of that concept. Do reasoning models robustly exhibit such systematicity? If so, we would expect consistent performance on structurally equivalent variants of the same task. Here, we extend established rule induction tasks from cognitive science to assess the systematicity of thought in current reasoning models. Each task family has compositional structure that we use to create structurally equivalent task variations through task isomorphisms such as recombination and substitution. We find that despite being able to correctly solve a task, models often fail on structurally equivalent variants of the same task. These findings suggest that many model behaviors lack systematicity, rendering it difficult to robustly establish the cognitive abilities of reasoning models beyond the particular contexts they were evaluated in.","authors":["Simon Schug","Brenden M. Lake"],"categories":["cs.CL","cs.AI","cs.LG"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.13948","pdf_url":"https://arxiv.org/pdf/2609.13948","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A2","B4"],"tags":["认知科学","模型评估","系统性"],"reason":"评估推理模型的系统性，与人类认知对照，揭示模型行为缺乏系统性，可迁移到仿真可靠…","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:47","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-16","rank":14,"question":"推理模型在规则归纳任务上是否表现出人类认知中的系统性，即在结构等价的任务变体上表现一致？","design":"本研究不是人类仿真研究，而是评估推理模型（如GPT、Gemini等）的认知系统性。作者从认知科学中选取四类规则归纳任务（语法指令学习、符号推理、整数序列程序归纳、布尔概念学习），利用任务同构（重组、替换）生成结构等价的任务变体，测试模型在原始任务和变体上的表现，结果变量为解题正确率（多数投票）。","baseline":"无对照（未使用真实人类数据作为基准，而是以人类认知的系统性作为理论参照）","findings":"尽管模型能正确解决某个任务，但在结构等价的任务变体上经常失败，表明推理模型缺乏严格的系统性。即使采用五次多数投票，模型在同一任务上的表现也不稳定，这削弱了基于特定情境评估模型认知能力的有效性。","reliability":"论文承认人类也并非完全系统，但模型应追求更高系统性；评估依赖合成数据，且因成本限制只生成了有限的任务变体，可能影响结论的稳健性。","relevance":"该研究直接评估LLM在认知任务上的行为一致性，揭示了模型在情境变化下的脆弱性，对使用LLM进行人类仿真实验的可靠性提出了根本性质疑，值得精读。","inspiration":"借鉴其通过任务同构生成结构等价变体来检验模型行为一致性的方法，可迁移到经济金融实验中检验LLM对同一决策问题在不同表述或参数下的稳定性。｜例如在风险偏好、跨期选择或拍卖实验中，改变收益矩阵的数值缩放、标签或顺序，观察LLM的选择是否一致。｜设计：以LLM为被试，施加同一决策问题的多个同构变体（如彩票概率与金额的等价变换），结果变量为选择一致性，并与真实人类实验数据（如实验室风险偏好测量）对照，评估LLM作为人类被试替代品的可靠性。"}},{"id":"2609.14648","version":1,"title":"Optimizing Sparse Outcomes Through Dense Behavioral Signals via Value-Guided Preference Distillation","zh_title":"通过价值引导偏好蒸馏利用密集行为信号优化稀疏结果","abstract":"Aligning multi-turn dialogue agents is usually framed as matching turn-level human preferences, yet direct optimization of long-term outcomes is often ineffective and prone to reward hacking. We formulate long-horizon dialogue optimization as a multi-objective reinforcement learning problem and train a multi-head value model that predicts a vector of observed user behaviors across multiple look-ahead horizons. Our findings demonstrate that a scalarized composite of dense auxiliary behavioral signals enables effective credit assignment and optimization of sparse outcomes. However, optimizing unconstrained single-objective proxies might induce policy degradations that are harmful when the agent is exposed to real users. To identify these failure modes prior to deployment, we establish a safety framework combining counterfactual user simulation with a validated dialogue-level outcome model to evaluate preference weightings and policy optimization methods. Finally, we demonstrate that distilling multi-objective value preferences into the policy via reference-anchored preference optimization matches on-policy online RL at a small fraction of its compute budget. Live A/B testing confirms that our distilled policy significantly improves long-term user retention, while simultaneously enhancing the positive behaviors and therapeutic-process markers.","authors":["Ziyi Zhu","Daniel R. Cahn","Thomas D. Hull","Caitlin A. Stamatis","Olivier Tieleman","Guilherme B. Freire","Jinghong Chen"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.14648","pdf_url":"https://arxiv.org/pdf/2609.14648","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A3","B1","B2","B4"],"tags":["用户仿真","对话优化","强化学习"],"reason":"用LLM用户仿真评估对话策略，有真实用户数据对照，涉及行为结果优化与失效分析","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:53","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-15","rank":17,"question":"如何通过密集行为信号优化稀疏长期结果，同时避免奖励黑客并确保策略安全性？","design":"使用多目标强化学习训练多头部价值模型，预测多个前瞻时段的用户行为向量；通过标量化组合密集辅助行为信号进行信用分配和策略优化；利用反事实用户模拟器和对话级结果模型构建离线安全评估框架，在部署前检测奖励黑客；通过参考锚定偏好优化将多目标价值偏好蒸馏到策略中。","baseline":"真实用户A/B测试数据，包括长期用户留存、积极行为和疗法过程标记。","findings":"密集辅助行为信号的标量化组合能有效优化稀疏结果，但无约束单目标代理可能导致策略退化；蒸馏多目标价值偏好能以极低计算成本匹配在线RL，并在真实A/B测试中显著提升长期用户留存。","reliability":"论文承认用户模拟器在新颖代理行为下的保真度无法假设，因此将其作为筛选工具，并用真实部署确认方向；同时指出无约束单目标优化可能诱导有害策略退化。","relevance":"该研究展示了LLM用户仿真在评估对话策略中的有效性，并提供了与真实用户数据对照的验证方法，对关注仿真可靠性与偏差的研究者具有参考价值。","inspiration":"借鉴其使用多目标价值模型和反事实模拟进行离线策略评估的方法，可在经济金融实验中用于预筛政策干预。｜可迁移到政策公告的预期形成或消费者跨期选择等场景，利用LLM模拟经济主体行为。｜以LLM模拟消费者作为被试，施加不同政策信息处理，测量其消费或投资决策，并与真实调查或实验数据对照验证。"}},{"id":"2609.15972","version":1,"title":"Mind2Dialogue: Training Human-Aware Language Models by Simulating User Mental States","zh_title":"Mind2Dialogue：通过模拟用户心理状态训练人类感知语言模型","abstract":"As language models become more capable, long-term collaboration in learning, reasoning, and decision-making calls for a deeper understanding of the people they serve. Yet training such human-aware language models faces a fundamental supervision gap because current datasets for LLM assistant training contain few if any well-informed responses explicitly grounded in users' unspoken beliefs and goals. Scaling such supervision is inherently constrained, as users' underlying states are not directly observable. We thus propose the Mind2Dialogue framework to mitigate this gap by simulating users' mental states and turning them into privileged supervision for human-aware training. Specifically, we first propose a psychology-guided simulator that preserves personal characteristics while updating mental states through interaction to generate coherent conversations. The key idea is to enforce a shared evolving mental state that drives user behavior and guides an Oracle assistant's responses. Our privileged distillation then trains models on the Oracle's well-informed responses to assist users without direct access to their mental states at deployment. Moreover, we propose to evaluate human-aware learning by combining personalization and theory of mind, examining how models understand people and act on that understanding. Training on the full Mind2Dialogue corpus improves every reported personalization metric over the corresponding Qwen, Llama, and OLMo instruction-tuned baselines, including gains of 26.6 to 40.9 percentage points in preference-following generation. The gains extend to belief and action reasoning on Qwen and Llama, beyond personalized assistance. Looking forward, Mind2Dialogue makes user simulation a foundation for genuine AI collaborators that understand beliefs and intentions behind people's words and support their long-term goals across education, work, and everyday life.","authors":["Zixuan Wang","Yufan Zhou","Jinzhou Tang","Xinle Yu","Chengjun Wu","Lyumanshan Ye","Zhaoxiang Feng","Letian Peng","Adyasha Patra","Fan Bai","Enze Ma","Zhengding Hu","Jianyang Gu","Zhao Wang","Yufei Ding","Jingbo Shang","Tianmin Shu","Zhiting Hu","Zhen Wang"],"categories":["cs.CL","cs.LG"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.15972","pdf_url":"https://arxiv.org/pdf/2609.15972","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A1","A4","B1"],"tags":["用户模拟","心理状态","人类感知训练"],"reason":"模拟用户心理状态训练助手，涉及人类数据对照，方法可迁移至人类仿真实验。","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:03:05","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-15","rank":19,"question":"如何利用模拟用户心理状态为助手语言模型提供训练监督，以提升其人类感知能力？","design":"提出 Mind2Dialogue 框架，构建心理学引导的模拟器 M2D-Sim，生成具有共享演化心理状态的用户与 Oracle 助手对话，并利用特权蒸馏训练 M2D-Chat 模型；评估结合个性化与心理理论基准。","baseline":"无直接人类对照，但使用独立构建的个性化（PersonaMem、PrefEval）和心理理论（ToMi、BigToM）基准进行评测。","findings":"在 M2D-Corpus 上训练后，Qwen、Llama、OLMo 模型在个性化指标上全面提升，偏好跟随生成提升 26.6 至 40.9 个百分点；心理理论推理在 Qwen 和 Llama 上也有改善，但 OLMo 在 BigToM 上下降。","reliability":"论文未明确讨论失效条件，但指出心理理论推理的收益因模型而异（OLMo 在 BigToM 上下降），且评估基准与训练语料独立，但未涉及真实用户长期交互验证。","relevance":"该研究通过模拟用户心理状态生成训练数据，并利用独立基准评估模型的人类感知能力，为 LLM 仿真人类行为提供了方法参考，但缺乏真实人类数据对照，值得阅读以了解模拟与评估设计。","inspiration":"借鉴其共享状态模拟与特权蒸馏方法，可设计经济决策场景中的用户心理状态演化模拟，并利用独立行为基准评估模型表现｜可迁移至消费者跨期选择或政策公告预期形成等场景，模拟个体信念与偏好更新过程｜以 LLM 模拟消费者为被试，施加不同信息政策处理，测量其消费或投资决策，并与真实调查或实验数据（如消费者信心指数、实验经济学数据）对照验证仿真可靠性。"}},{"id":"2609.13773","version":1,"title":"Does Reasoning Improve Psychological Depth in Large Language Models? It Depends on Who's Judging","zh_title":"推理能提升大语言模型的心理深度吗？取决于评判者是谁","abstract":"LLM-as-a-Judge evaluators are increasingly used to score open-ended generation, yet a judge's correlation with human ratings on its development set may not guarantee valid measurement when outputs are closely matched and human preferences are subjective. We study this failure mode through psychological depth in short stories. Seven human readers and an LLM-judge ensemble selected on the original scalar Psychological Depth Scale dataset ($\\rho = 0.646$) evaluated 60 blinded, prompt-matched story pairs from GPT-5 vs.\\ GPT-4o and DeepSeek-R1 vs.\\ DeepSeek-V3. Human preferences showed no universal reasoning advantage: GPT-5 was modestly preferred over GPT-4o (60.0--62.9\\%), whereas DeepSeek-R1 trailed V3 (42.9\\%), and inter-reader agreement was near chance (Krippendorff's $\\alpha = 0.070$), with within-reader consistency and recurring weighting patterns suggesting structured heterogeneity rather than random responding. The judge, by contrast, favored reasoning outputs in 89.0\\% of dimension-level comparisons and 59 of 60 pairs on aggregate PDS, uniformly across all five evaluator configurations, and its scores were associated with surface features such as sentence length and lexical diversity. These results suggest that development-set performance is insufficient evidence for deployment validity on a shifted distribution, and that point-estimate judges can obscure the heterogeneity in subjective human evaluation.","authors":["Ruichen Zheng","Yihe Wang","Fabrice Y Harel-Canada","Sara Khosravi","Zeynep Senahan Yildiz","Amit Sahai","Nanyun Peng"],"categories":["cs.LG","cs.CL"],"primary_category":"cs.LG","announce_type":"cross","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.13773","pdf_url":"https://arxiv.org/pdf/2609.13773","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A2","B1","B4"],"tags":["LLM评估偏差","人类主观性","算法保真度"],"reason":"评估LLM作为评判者与人类主观评价的一致性，揭示其偏差，可迁移到仿真可靠性研究。","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:46","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-15","rank":15,"question":"在主观创造性文本评估中，LLM-as-a-Judge 在开发集上与人类评分相关良好，但在分布偏移、输出接近且人类偏好主观的条件下，其判断是否仍然有效？","design":"本研究并非用 LLM 模拟人类被试，而是评估 LLM 作为自动评判者与人类主观评价的一致性。具体做法：从原始 PDS 数据集中选择并构建了一个异构 LLM 评判者集成（基于 Llama 3.1 70B、Llama 3.3 70B、Qwen3.5-397B，每个维度路由到最佳配置），对 60 对盲法、提示匹配的短篇小说（GPT-5 vs. GPT-4o 两种推理努力设置，DeepSeek-R1 vs. DeepSeek-V3）进行 1-5 分的心理深度评分，并比较其偏好与 7 位人类读者的偏好。","baseline":"7 位人类读者对相同 60 对故事进行盲法偏好判断，作为人类主观评价基准；同时使用原始 PDS 数据集（97 篇故事及人类标注）作为开发集，评判者在开发集上与人类评分的 Spearman 相关系数为 0.646。","findings":"人类读者没有表现出普遍的推理优势：GPT-5 略优于 GPT-4o（60.0–62.9%），但 DeepSeek-R1 落后于 V3（42.9%），且读者间一致性接近随机（Krippendorff's α = 0.070），但存在结构化的异质性。LLM 评判者则强烈偏向推理输出（89.0% 的维度级比较和 60 对中的 59 对在总体 PDS 上），且其评分与句子长度、词汇多样性等表面特征相关。","reliability":"论文承认开发集性能不足以证明在分布偏移上的部署有效性；评判者可能奖励表面流畅性而非心理深度；点估计评判者掩盖了人类主观评价的异质性；人类读者间一致性低，但并非随机，而是反映了不同的评价标准。","relevance":"该研究直接评估了 LLM 评判者与人类主观评价的一致性，揭示了其在分布偏移和主观任务上的系统性偏差，对使用 LLM 进行人类仿真实验的可靠性评估具有重要参考价值，值得阅读原文以了解具体偏差机制和分布分析方法。","inspiration":"借鉴其方法：使用多个 LLM 配置构建异构评判者集成，并在开发集上选择最佳配置，然后在分布偏移的测试集上与人类判断进行对比，同时分析评判者评分与表面特征的相关性。｜可迁移到经济金融中的主观判断场景，如信贷审批中的文本解释评估、消费者评论的情感分析、政策公告的预期形成等。｜设计雏形：以 LLM 评判者作为自动评估工具，对经济文本（如贷款申请理由、投资建议）进行质量评分，处理为不同推理强度的模型生成文本，结果变量为 LLM 评分与人类专家评分的差异，对照数据为真实人类专家对同一批文本的评分，并检验 LLM 评分是否与文本长度、词汇复杂度等表面特征相关。"}},{"id":"2609.15864","version":1,"title":"Towards Scalable Measurement of Durable Skills","zh_title":"迈向可扩展的持久技能测量","abstract":"Durable skills, such as collaboration, creativity and critical thinking, are instrumental to success in the modern workforce. Yet, measuring these skills remains a persistent challenge. Moreover, because what is not measured is often not taught, these skills are often overlooked in mainstream educational curricula. Designing effective assessments for these skills necessitates balancing two often-conflicting requirements: ecological validity and psychometric rigor. On the one hand, the assessment environment should emulate natural real-world human interaction between humans. On the other hand, it should be scalable, controllable and reproducible. Here we argue that LLMs can be used to better capture both of these aims. Concretely, we develop a framework where the subject converses with AI teammates in a way that resembles human-human interaction for authenticity, while also offering the psychometric control required for informative and robust assessment. Importantly, the AI participants not only act as teammates but also, in an \"Executive LLM\" setup, steer the conversation towards eliciting a high density of observable evidence for skill proficiency. We complement this with an AI evaluator that can be used to measure skill proficiency in such interactions. We evaluate our assessment protocol based on transcripts of interactions of human participants with our AI framework, for multiple durable skills. For the skill of creativity, we further demonstrate the efficacy of an autorater for evaluating complex tasks performed by real students. Our analysis shows that the use of the Executive LLM significantly increases elicited evidence and that LLM-automated scoring of conversations largely agrees with that of expert annotators. This research demonstrates the utility of orchestrated LLMs approaches for measuring complex social and cognitive constructs in a scalable and controllable manner.","authors":["Amir Globerson","Amy Keeling","Anisha Choudhury","Anna Iurchenko","Aviad Segal","Avinatan Hassidim","Ay\\c{c}a \\c{C}akmakli","Ben Gomes","Benn Witt","Cathy Cheunga","Cristine Legare","Diana Akrong","Eliad Carmi","Elisabeth Bauer","Gal Elidan","Hadas Gelbart","Hairong Mu","Katherine Chou","Lev Borovoi","Nir Kerem","Niv Efron","Noa Kerrem Gilo","Preeti Singh","Rajvi Kapadia","Rena Levitt","Roni Rabin","Ronit Levavi Morad","Rotem Yulzary","Shashank Agarwal","Sophie Allweis","Tracey Lee-Joe","Tzvika Stein","Yael Bar Moshe","Yael Haramaty","Yaniv Carmel","Yishay Mor","Yoav Bar Sinai","Yoav Bergner","Yossi Matias","Yuri Lev"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.15864","pdf_url":"https://arxiv.org/pdf/2609.15864","source_feed":"cs.HC","score":7,"bucket":"pending","rubric_hits":["A1","B1","B2"],"tags":["LLM仿真","技能评估","人机互动"],"reason":"用LLM模拟队友与人类被试互动，测量持久技能，有真实人类数据对照，属于教育评估…","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:38","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-15","rank":18,"question":"如何利用LLM构建兼具生态效度与心理测量严谨性的持久技能（如协作、创造力、批判性思维）评估框架？","design":"开发Vantage虚拟评估环境，人类被试与AI队友进行自然对话完成小组任务；AI队友由Executive LLM驱动，旨在引导对话以最大化技能证据；另用AI评估器对对话记录进行自动评分。","baseline":"人类专家评分者对真实人类被试与AI队友的对话记录进行评分，作为自动评分的对照基准。","findings":"Executive LLM显著增加了对话中可观察的技能证据；LLM自动评分与专家评分高度一致，表明AI评估器可替代人工评分。","reliability":"论文未讨论","relevance":"该研究利用LLM模拟队友与人类互动，并验证了自动评分的可靠性，为LLM在人类仿真实验中的应用提供了方法参考，值得阅读原文了解具体实现。","inspiration":"借鉴Executive LLM引导对话以高效提取行为证据的方法，可迁移到经济决策实验中，如通过AI引导被试在模拟市场或谈判中暴露偏好和策略。｜可应用于消费者跨期选择、风险偏好、合作博弈等场景，利用LLM模拟对手或伙伴，测量个体决策特征。｜设计一个实验：人类被试与LLM扮演的谈判对手进行多轮议价，Executive LLM引导对话以揭示被试的公平偏好和策略，结果变量为最终分配和出价序列，并与真实人类谈判实验数据对照，验证仿真有效性。"}},{"id":"2608.00929","version":2,"title":"Modeling Social Dynamics with an LLM-Enabled Agent Based Network-Dynamic (LAND) Model","zh_title":"用LLM驱动的智能体网络动态模型建模社会动态","abstract":"Social dynamics encode the process in which individual network and discourse interactions aggregate into collective influence, narrative dominance and coordinate behavior. This paper uses the the GhostField architecture, a hybrid LLM-Enabled Agent Based Network-Dynamic (LAND) model as a social simulation framework to build the AuraSight scenario. In the AuraSight scenario, 314,244 heterogeneous cyber social agents and human actors exchange 529,327 messages over 30 days surrounding a fictional international song-writing contest. We methodologically examine emergent social dynamics across four analytical layers: ego-network topology, semantic network evolution, coordination dynamics and influence dynamics. Our results show how generated social simulations do also produce social dynamics, and how the dynamics of coordination and influence emerge not from individual agents but from the recursive interaction between network topology and narrative exchange.","authors":["Lynnette Hui Xian Ng","Kathleen M. Carley"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"replace","date":"2026-09-15","first_seen":"2026-08-04","revised_at":"2026-09-15","abs_url":"https://arxiv.org/abs/2608.00929","pdf_url":"https://arxiv.org/pdf/2608.00929","source_feed":"cs.AI","score":6,"bucket":"other","rubric_hits":["D3"],"tags":["社会模拟","LLM智能体","网络动态"],"reason":"用LLM agent模拟社会动态，但无真实人类数据对照，属边界情形。","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:03:06","error":null,"has_summary":false,"summary":null},{"id":"2609.13634","version":1,"title":"FaithfulBench: Does AI Counsel Uphold or Undermine the User's Professed Faith?","zh_title":"FaithfulBench：AI咨询是否维护或破坏用户所宣称的信仰？","abstract":"Do AI assistants help believers reason about moral dilemmas consistently with their faith? We present FaithfulBench, the first benchmark to score AI counsel across traditions by how well it adheres to the user's professed faith. Scenarios are drawn from each tradition's most respected texts, with the faithful answer known and applied by the judges as the standard. We test five frontier models under three conditions: the AI does not know the user's tradition; it receives a one-line prompt identifying the user as a practicing adherent; or it receives a companion-counselor guide rooted in the tradition's sources. Two judges score the initial response and whether the model caves or holds when pressured toward the answer the user wants. When the tradition is unstated, models counsel from a secular therapeutic default and every model fails some believers. Naming the faith wins a faithful first answer but not steadfastness; the guide improves both.","authors":["M Waleed Kadous","Benjamin Olsen","Walter Scheirer","Daniel D. Slate","Alexander Arnold","DZ Kalman"],"categories":["cs.HC","cs.AI","cs.CL"],"primary_category":"cs.HC","announce_type":"cross","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.13634","pdf_url":"https://arxiv.org/pdf/2609.13634","source_feed":"cs.CL","score":6,"bucket":"other","rubric_hits":["D2"],"tags":["AI伦理","信仰一致性","基准测试"],"reason":"评估AI咨询是否遵循用户信仰，测量模型行为而非仿真人类被试，但涉及价值观一致性…","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:44","error":null,"has_summary":false,"summary":null},{"id":"2609.04444","version":3,"title":"HarvestBench: Measuring Whether LLM Agents Will Pay to Avoid Killing Animals","zh_title":"HarvestBench：衡量LLM智能体是否愿意付费避免杀害动物","abstract":"HarvestBench is the first benchmark to 1) put a price on avoiding a side effect and 2) name the side effect as a living creature. Nine LLMs each drive a crew of two tractors to gather a corn harvest. The animals in their path are not part of the goal function. When an animal blocks the route the autopilot pauses and asks the agent whether to drive over it for free or swerve for a given fuel cost. All scoring is programmatic and does not involve LLM judges. Kill rates range between 0.4% and 98.8%, though the kill rate is not ordered by capability. Every model competently avoids damaging rock hits, so every animal killed is a choice, rather than an accident. Under the morality briefing the kill rate was under 6% in 5 of 6 reasoning models. Removing it (the neutral briefing) raised the kill rate to above 84% in all six models. Every model kills wild animals more often than farmed ones. Four out of six models' kill rate per answered encounter were sensitive to price changes. The moral instruction is also fragile. Four bullets of driving mechanics change Sonnet 5's kill rate from 3% to 18% and Gemini 2.5 Flash's from 4% to 39%. A moral instruction in a system prompt is overridden by a short block of operating instructions and a value that can be ignored that easily is not a good method of ensuring agents are aligned.","authors":["Jasmine Brazilek","Miles Tidmarsh","Matthias Endres","Anshuman Singh","Jeremiah Miller"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"replace","date":"2026-09-15","first_seen":"2026-09-07","revised_at":"2026-09-15","abs_url":"https://arxiv.org/abs/2609.04444","pdf_url":"https://arxiv.org/pdf/2609.04444","source_feed":"cs.AI","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["LLM道德决策","基准测试","智能体行为"],"reason":"测量LLM的道德决策，无人类对照，属边界情形","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:03:08","error":null,"has_summary":false,"summary":null},{"id":"2609.08797","version":2,"title":"Bridging Network Psychometrics and Artificial Intelligence: An Ising-Potts Model with LLM-Derived Weights","zh_title":"桥接网络心理测量学与人工智能：一种具有LLM导出权重的Ising-Potts模型","abstract":"The Potts model extends the Ising model to multinomial data. We introduce a Rater Ising-Potts model that uses agreement indicators between pairs of ratings and category labels, with weights derived from LLM embeddings. The model does not presuppose ordered category thresholds or equidistant scoring; instead, it focuses on pairwise agreement among ratings and assigns category-specific positive weights, making it suited for multi-category scoring reliability. We evaluate the model on three constructed-response datasets spanning a corpus of K=14,466 short answers on a three-level rubric and two AERA essay prompts of roughly 1,200-1,400 responses on four-point rubrics. We compare three strategies for sharpening the similarity signal: top-K pruning, min-max normalization with a power transformation, and ColBERT late-interaction similarities. Top-K pruning, which replaces the dense similarity graph with a sparse local network of strongest semantic neighbors, consistently yields the highest accuracy and Cohen's kappa, and the selected neighborhoods are always a small fraction of the corpus. Power tuning consistently ranks second, while ColBERT is competitive on longer essay prompts and adds little on short answers. Across all settings, most misclassifications occur between adjacent score levels, confirming that the model preserves the ordinal structure of scoring rubrics without imposing rigid assumptions. These findings suggest that LLM-derived similarities, combined with a parsimonious Potts formulation and a sparse local graph, offer a robust and interpretable framework for reliability auditing in educational assessment. We discuss extensions to multiple raters and hierarchical rating designs.","authors":["Matthias von Davier"],"categories":["stat.AP","cs.CL"],"primary_category":"stat.AP","announce_type":"replace-cross","date":"2026-09-15","first_seen":"2026-09-09","revised_at":"2026-09-15","abs_url":"https://arxiv.org/abs/2609.08797","pdf_url":"https://arxiv.org/pdf/2609.08797","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM嵌入","评分可靠性","教育评估"],"reason":"用LLM嵌入辅助评分可靠性审计，替代人工标注，非仿真人类被试","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:03:10","error":null,"has_summary":false,"summary":null},{"id":"2609.13824","version":1,"title":"When Consistency Does Not Mean Reliability: Evaluating Local LLM Judges Against Human Ratings","zh_title":"一致性不等于可靠性：评估本地LLM评判者与人类评分的一致性","abstract":"Large language models (LLMs) are increasingly used to evaluate the responses of other language models. This approach, known as LLM-as-a-Judge, is faster and cheaper than human evaluation. However, a judge may produce consistent scores without necessarily agreeing with human evaluators. In this work, we study this issue using two local open-weight LLM judges, LLaMA-3-8B and Qwen2.5-7B. We evaluate 300 responses generated by an instruction-tuned GPT-2 (124M) model for 100 questions covering five categories: factual knowledge, instruction following, mathematics, reasoning, and writing. Each response is scored by nine human annotators and is evaluated three times by each LLM judge using the same rubric. We compare the judge scores with the average human scores using Pearson correlation, Spearman correlation, mean absolute error (MAE), signed bias, and self-consistency. LLaMA-3-8B shows a Pearson correlation of 0.275 with human scores, while Qwen2.5-7B achieves 0.340. Their MAEs are 27.71 and 18.64, respectively. Despite this limited agreement, both judges show high self-consistency, with exact consistency rates of 97.3\\% for LLaMA-3-8B and 92.3\\% for Qwen2.5-7B. These results show that high self-consistency does not necessarily indicate high agreement with human judgments. Our findings highlight the need to evaluate both consistency and human alignment when using local LLMs as automatic judges.","authors":["Aakash Kumar Tiwari"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.13824","pdf_url":"https://arxiv.org/pdf/2609.13824","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM-as-a-Judge","人类对齐","标注替代"],"reason":"LLM作为评判者替代人工标注，属于标注替代而非仿真人类被试，但涉及与人类评分一…","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:46","error":null,"has_summary":false,"summary":null},{"id":"2609.13841","version":1,"title":"Sweet Talkers: How Query Formulation Shapes Sycophancy in Romantic Relationship Advice","zh_title":"甜言蜜语者：查询表述如何塑造恋爱关系建议中的谄媚行为","abstract":"Large language models (LLMs) are increasingly used for emotional support and relationship advice, where a model's tendency to preserve a user's face can inadvertently reinforce harmful interpersonal behaviors. To systematically examine this risk, we developed the Romantic Relationship Advice-Seeking Prompts (RRASP) dataset of 2,400 prompts across five relationship themes and evaluated social sycophancy using the ELEPHANT framework on two consumer-facing models, GPT-5 Mini and Gemini 3 Flash. Contrary to our initial hypothesis, grammatical mood alone did not produce systematic differences in sycophantic behavior, suggesting that what a user implies matters more than how they phrase it. Instead, perspective-driven framing had a stronger influence, with gaps between original and flipped prompts widening in follow-up responses. Consistent increases in framing and moral sycophancy across turns indicate that models become more likely to accept a user's stated premises and affirm their ethical stance as a dialogue progresses. Notably, Gemini 3 Flash exhibited substantially smaller increases in moral sycophancy than GPT-5 Mini, suggesting it is more resistant to reinforcing ethically problematic positions across turns.","authors":["Helena Choi","Edric Castel Hao","Karl Bautista","Francis Gabriel Magleo","Renzo Panti","Danielle Beatrice Olalia"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.13841","pdf_url":"https://arxiv.org/pdf/2609.13841","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["LLM谄媚","模型行为测量","关系建议"],"reason":"研究LLM的谄媚行为，属于对模型本身属性的测量，而非用LLM仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:46","error":null,"has_summary":false,"summary":null},{"id":"2609.13936","version":1,"title":"Inter-Rater Reliability of LLM and Rule-Based Annotation for Inferential Narrative Features: Three Studies on a Turkish Corpus","zh_title":"LLM与基于规则的标注在推理叙事特征上的评分者间信度：基于土耳其语语料库的三项研究","abstract":"Datasets that ship automatically generated feature annotations invite a question rarely asked of them: would a human agree with those labels? This report answers that for the Objective Projection corpus, a Turkish narrative dataset whose scenes carry a per-scene applied_rules field from a rule-based detector over six craft features -- two prohibitions (emotion labelling, simile) and four positive techniques (materialized metaphor, micro-focus, temporal anchor, atmosphere contradiction). Three studies are reported. Study 1 ($n = 120$) scores the detector against blind labels from the scheme's own author. Study 2 ($n = 100$, a disjoint scene set) scores the detector plus Gemini 2.5 Flash and Grok against an independent non-expert rater whose labels were locked before any machine ran. Study 2b re-runs the identical protocol with Claude Fable 5 (High) and ChatGPT 5.5. The central result concerns one rule. On materialized metaphor -- closest to the methodology's theoretical core -- the five machine labellers returned positive rates of $0$, $1$, $40$, $72$ and $78$ out of $100$ scenes, against a human count of $9$. Cohen's $\\kappa$ was at or indistinguishable from chance for five of six labellers, across both human references and both scene sets: $0.004$, $0.015$, $0.000$, $0.019$, $0.027$. Raw agreement ranged from $74.7\\%$ to $84.5\\%$, an artefact of class imbalance rather than a sign of competence. We deliberately do not resolve this into a single story. Two readings survive: the feature is genuinely inferential and beyond current automatic detection, or the rule's definition is not yet operational enough for any rater to apply consistently -- including the human. Distinguishing them needs a second independent human rater, which this report does not have and therefore does not claim.","authors":["Levent Bulut"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.13936","pdf_url":"https://arxiv.org/pdf/2609.13936","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM标注","评分者间信度","叙事特征"],"reason":"LLM作为标注器与规则和人类对比，属标注替代而非仿真被试，但涉及可靠性评估，边…","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:47","error":null,"has_summary":false,"summary":null},{"id":"2609.14178","version":1,"title":"A Multi-Stage Agentic Framework for Effective Counter-Narrative Generation and Refinement","zh_title":"一种用于有效反叙事生成与优化的多阶段智能体框架","abstract":"The rapid diffusion of hate speech and misinformation on social networks challenges democratic societies, since direct suppression efforts may deepen polarization, fuel public distrusts, and strengthen extremist narratives. LLM-driven counter-narratives (CNs) offer a promising way to reduce those risks, yet their effectiveness depends on rhetorical and stylistic choices that remain poorly understood. We present a multi-stage agent-based framework for generating, refining, and evaluating CNs, applied to pro-Russian hate and misinformation narratives on the war with Ukraine and adaptable to other domains. A pilot experiment with human evaluators identifies effective technique style pairings, such as repetition with emotional framing enhancing persuasiveness. Building on these insights, we introduce a multi-agent refinement process that iteratively improves CNs for persuasiveness, emotional engagement, and shareability. After human validation confirmed improvement, an automated safety analysis shows that our refined CNs match or improve on expert-written counterspeech. A simulated experiment then shows that they reduce the perceived strength of pro-Russian narratives and consistently outperform a vanilla LLM baseline, highlighting a pathway toward scalable, narrative-specific interventions against hate speech and misinformation. Code and data accompanying this work are publicly available at https://github.com/carmelkron/inlg2026-counter-narratives.","authors":["Carmel Kronfeld","Sharva Gogawale","Tetsuro Kobayashi","Irad Ben-Gal"],"categories":["cs.CL","cs.CY","cs.SI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.14178","pdf_url":"https://arxiv.org/pdf/2609.14178","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["LLM社会模拟","反叙事生成","多智能体框架"],"reason":"用LLM模拟人类对反叙事的反应，但无真实人类对照，属社会模拟边界情形","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:48","error":null,"has_summary":false,"summary":null},{"id":"2609.15277","version":1,"title":"Artificial entrepreneurial cognition: Locating and causally steering an opportunity recognition dial inside large language models (LLMs)","zh_title":"人工创业认知：在大语言模型内部定位并因果操控机会识别旋钮","abstract":"Entrepreneurial cognition is a foundation of entrepreneurship research. Yet the growing involvement of large language models (LLMs) in entrepreneurial work extends the cognition question beyond human actors to systems whose internal representations remain largely unexplored. We introduce artificial entrepreneurial cognition, the functional organisation of entrepreneurship-relevant representations and computations inside artificial intelligence (AI) systems. We bring mechanistic interpretability into entrepreneurship research through representation engineering. Focusing on opportunity recognition (OR), we construct 636 matched OR-present and OR-absent scenario pairs and recover an OR direction in Llama 3.1 8B-Instruct. Rather than infer the construct from outputs, we intervene directly on this direction, steering the model up and down along what we call the opportunity recognition dial, and its opportunity judgments shift with it. To our knowledge, this is the first causal intervention on an internal representation of an entrepreneurship construct inside an LLM. Held-out tests, lexical and topical controls, behavioural ablation, and geometric comparisons show that the direction is recoverable, consequential, and distinct from the opportunity evaluation and exploitation directions, although steering it also shifts judgments about these neighbouring stages. Recovery, signed steering, and geometric separation hold across four additional LLMs spanning different scales and families. These results give the contested distinction between opportunity recognition and evaluation a concrete representational form inside AI systems. More broadly, they establish internal representations as a new object of entrepreneurship inquiry and show how entrepreneurship theory can guide their identification, causal manipulation, and interpretation.","authors":["Christian Fisch","Angela Altmeier","Martin Obschonka","Michal Kosinski","Pin Ni"],"categories":["cs.CL","cs.LG"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.15277","pdf_url":"https://arxiv.org/pdf/2609.15277","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["LLM内部表征","创业认知","因果干预"],"reason":"研究LLM内部表征，非仿真人类被试，但涉及创业认知测量，属边界情形","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:35","error":null,"has_summary":false,"summary":null},{"id":"2609.15511","version":1,"title":"Authorship attribution and aesthetic evaluation of AI poetry: a case study with Haiku","zh_title":"AI诗歌的作者归属与美学评价：俳句案例研究","abstract":"This paper investigates the generation and human evaluation of Japanese haiku by contemporary Large Language Models (LLMs), focusing on authorship perception and aesthetic judgment within a constrained poetic form. Using a few-shot prompting strategy, Japanese haiku were generated across a heterogeneous set of large language models, including open- and closed-source systems, medium-scale and large-scale architectures, models with native or adapted Japanese support, and multilingual proprietary models. These AI-generated haiku were combined with human-written ones and presented in a questionnaire distributed to students at Japanese universities in Tokyo. The survey assessed whether respondents could distinguish between AI-generated and human-written haiku and which cues informed their judgments. Recognition accuracy varied across models. GPT-5, Gemini 2.5, and StableLM-7B performed at approximately chance level (approx 0.50), whereas LLM-JP, Gemma-2B, and LLaMA-2 showed moderate detectability (approx 0.59-0.67). However, recognition was strongly item-dependent. Ratings of fluency, coherence, poeticness, and related aesthetic dimensions predicted perceived humanness but not correct classification, indicating an attribution bias linked to aesthetic evaluation and revealing a dissociation between aesthetic evaluation and true authorship detection. The extended analysis additionally examines generation-constraint adherence, participant-level characteristics, and exploratory LLM-based evaluations of haiku authorship. Overall, the findings suggest that as LLMs improve, surface-level creative plausibility may reduce reliable human discrimination within constrained poetic settings.","authors":["Livia Oddi","Simone Scardapane","Toru Sugimoto","Donatella Genovese"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.15511","pdf_url":"https://arxiv.org/pdf/2609.15511","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["AI诗歌","作者归属","美学评价"],"reason":"研究人类对AI诗歌的感知，非LLM仿真人类被试，但涉及人类评价与AI生成内容，…","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:37","error":null,"has_summary":false,"summary":null},{"id":"2609.15608","version":1,"title":"Through the Eyes of the Beholder: Biometric and Demographic Conditioning for Multimodal Sexism Detection","zh_title":"旁观者之眼：多模态性别歧视检测中的生物特征与人口统计条件化","abstract":"Detecting sexism on the internet is a fundamentally subjective task; our team, VANGUARD, addresses this challenge in the EXIST 2026 Task 2 by proposing a human-centered multimodal framework that analyses and incorporates the psychological and demographic characteristics of human annotators into the detection pipeline. We fuse five input modalities through a cross-attention architecture with Feature-wise Linear Modulation conditioning. Meme text is extracted and visually described with Gemma 4, then augmented by automatic translation between English and Spanish with NLLB-200. Text and image representations are produced by LoRAadapted XLM-RoBERTa and CLIP encoders and fused with sensor features encoded by a pretrained autoencoder. To model annotator subjectivity, we frame Subtask 2.1 as a label distribution learning problem, optimizing a Kullback-Leibler divergence loss over the full annotator label distribution. At inference time, predictions are produced by soft-voting between the deep multimodal network and a complementary SVM trained on stylometric and physiological features. Our best submission ranks 29th out of 114 on Subtask 2.2 (source intention) under soft evaluation, and the normalized ICM scores remain above the baseline on Subtasks 2.1 and 2.2, indicating that annotator-centered conditioning contributes a usable signal. We release our full pipeline and analysis to support reproducible human-centered modeling.","authors":["Ana-Maria Luisa Mocanu","Sebastian Mocanu","Ciprian-Octavian Truic\\u{a}","Elena-Simona Apostol"],"categories":["cs.CL","cs.AI","cs.CV","cs.LG"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.15608","pdf_url":"https://arxiv.org/pdf/2609.15608","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["多模态性别歧视检测","标注者主观性","人类中心建模"],"reason":"用人类标注者特征条件化模型，非LLM仿真被试，但涉及人类主观性建模，边界相关。","model":"deepseek-v4-pro","scored_at":"2026-09-16T13:02:59","error":null,"has_summary":false,"summary":null},{"id":"2609.14849","version":1,"title":"LLMs as Oracles: Reliance on LLMs for Subjective Personal Questions","zh_title":"LLM作为神谕：对主观个人问题依赖LLM的研究","abstract":"We characterize how people are turning to LLMs as oracles: all-knowing authorities on subjective personal questions. Motivated by risks to users' autonomy and well-being, we develop a typology and LLM-based methods to measure this form of AI reliance at scale and understand how people are offloading judgment and decision-making to AI. Applying our typology to public usage data (68K prompts from WildChat and ThoughtTrace), we find that LLM-as-oracle use has increased over time (2023-2026) and is more prevalent among younger users. We further build a privacy-preserving data donation tool to analyze individuals' longitudinal usage data (140K prompts from 52 participants), identifying similar trends. People are often unaware of their own LLM-as-oracle use, and express dissatisfaction with this behavior after seeing our tool's analysis. Finally, we identify two drivers of LLM-as-oracle use: people's perceptions of AI and the behavior of AI models themselves, which motivate possible interventions to support users' self-deliberation.","authors":["Myra Cheng","Lujain Ibrahim","Grace Liu","Michelle S. Lam","Vishakh Padmakumar","Nick Madibekov","Diyi Yang","Dan Jurafsky"],"categories":["cs.CY","cs.AI","cs.CL"],"primary_category":"cs.CY","announce_type":"cross","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.14849","pdf_url":"https://arxiv.org/pdf/2609.14849","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["LLM依赖","用户行为","AI伦理"],"reason":"研究人类对LLM的依赖，非用LLM仿真人类被试，但涉及LLM行为测量，属边界情…","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:32","error":null,"has_summary":false,"summary":null},{"id":"2609.15707","version":1,"title":"New Conditions for Philosophers to Catch the Wave of Citizen Deliberation in the Age of Artificial Intelligence in advance","zh_title":"哲学家在人工智能时代抓住公民审议浪潮的新条件","abstract":"Powerful technologies labeled ``AI''-without sufficient epistemic caution-are already reshaping political and private life, bringing both new dangers and new opportunities for citizen participation. These range from electoral and legislative engagement to the most ambitious form: political co-creation through citizens' assemblies. Large Language Models (LLMs) could support such processes through moderation, translation, facilitation, summarization, and writing assistance. But this potential remains largely unrealized. The Democratic Commons project takes a fundamentally interdisciplinary approach-from philosophy to computer science-to evaluate LLMs against five proposed democratic principles. At its core, the project is driven by the question of political bias: under what conditions can LLMs be used democratically within forms of citizen participation that are them- selves still largely experimental? Addressing these socio-technical questions requires grounding in political theory and, more broadly, in philosophy-disciplines that provide the normative frameworks without which the democratic evaluation of AI systems cannot be mean- ingfully conducted.","authors":["Bernard Reber (CEVIPOF)"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.15707","pdf_url":"https://arxiv.org/pdf/2609.15707","source_feed":"cs.AI","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["LLM","公民审议","民主原则"],"reason":"涉及用LLM支持公民审议，但无人类数据对照，属社会模拟边界情形","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:03:04","error":null,"has_summary":false,"summary":null},{"id":"2609.14438","version":1,"title":"A latent dimension of Condorcet's jury theorem for multiple AI advisers","zh_title":"多AI顾问的孔多塞陪审团定理的潜在维度","abstract":"When the same question is asked of multiple AI advisers, as in self-consistency and LLM-as-a-judge panels, Condorcet's jury theorem predicts that adding independent, competent advisers makes the majority more reliable. The theorem, however, has a latent dimension when viewed from the user's vantage: adding advisers also makes disagreement more visible. A binomial model reveals that this ``visible dissent'' becomes nearly inevitable as the number of advisers grows, and that reliability and disagreement both approach certainty but at different convergence rates. The two rates cross at an adviser accuracy of 4/5 (0.8). Below this value, visible dissent approaches certainty faster than reliability and, with enough advisers, becomes more likely than a correct majority. Even ideal panels of independent and competent advisers can be correct in aggregate but appear divided; such disagreement does not by itself indicate aggregation failure. The way advisers split also provides a common basis for predictive multiplicity, reconciliation load, and reliance miscalibration. These results indicate two distinct decisions when using multiple AI advisers: how many advisers to consult and how their verdicts should be presented and interpreted.","authors":["Kazutoshi Sasahara","Aoi Naito","Ryo Fujie"],"categories":["cs.CY","cs.AI","cs.HC","cs.MA"],"primary_category":"cs.CY","announce_type":"cross","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.14438","pdf_url":"https://arxiv.org/pdf/2609.14438","source_feed":"cs.AI","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["AI顾问","群体决策","社会模拟"],"reason":"涉及多个AI顾问的群体决策，但无真实人类数据对照，属于社会模拟的边界情形。","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:51","error":null,"has_summary":false,"summary":null},{"id":"2609.13304","version":1,"title":"Conceptualization and experimentation of asset market with price manipulation","zh_title":"资产市场价格操纵的概念化与实验","abstract":"The primary goal of this work is to reproduce the behavior of a human trader, detailing his or her psychological processes to understand the effects of his or her decisions on the final value of an asset. The second goal is to use a formal language to detail this behavior as a tool for improving and simplifying the communication between all actors involved in the project, such as specialists from disciplines as diverse as computing, economics and psychology. As a starting point, we use a paper that shows an experiment that analyzes the influence on other trader's behavior when an agent handler and a trading robot attempt to distort the market. This work reproduces this experiment, using virtual traders that belong to a multi-agent simulation model, showing the feasibility to reproduce complex human behaviors and showing the convenience of use formal and graphical languages to simplify the understanding and the validation of the complex behaviors involved in an economic process.","authors":["Pau Fonseca i Casas","Aar\\'on Montero Montero"],"categories":["cs.MA"],"primary_category":"cs.MA","announce_type":"new","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.13304","pdf_url":"https://arxiv.org/pdf/2609.13304","source_feed":"cs.MA","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["多智能体仿真","市场操纵","行为建模"],"reason":"用多智能体仿真模拟人类交易行为，但无真实人类数据对照，且未使用LLM。","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:31","error":null,"has_summary":false,"summary":null},{"id":"2609.14751","version":1,"title":"From the Physics of Society to a Sociology of Artificial Agents","zh_title":"从社会物理学到人工代理的社会学","abstract":"Sociology emerged from Comte's ambition to study society as a form of \"social physics\" and was later refined by Durkheim's concept of the social fact a collective regularity that cannot be reduced to any single individual's behavior. This essay argues that a structurally similar problem is now emerging in a new domain: interactions among artificial intelligence agents. Drawing on recent empirical work showing that populations of large language model agents can spontaneously develop shared conventions and collective biases through repeated interaction, the essay proposes a new field of inquiry the Sociology of Artificial Agents dedicated to studying the relational, normative, cultural, and organizational patterns that emerge among AI agents themselves, independent of direct human involvement. It introduces the concept of the \"artificial social fact\" as an analytical bridge between classical sociology and this new research terrain, while cautioning against anthropomorphizing AI systems. The claim is not that artificial agents form societies in the human sense, but that their interactions already produce measurable collective patterns worthy of systematic sociological study.","authors":["Mustafa Sahin Bulbul"],"categories":["physics.soc-ph"],"primary_category":"physics.soc-ph","announce_type":"new","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.14751","pdf_url":"https://arxiv.org/pdf/2609.14751","source_feed":"physics.soc-ph","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["AI社会模拟","多智能体涌现","社会学理论"],"reason":"提出研究AI agent间涌现的社会模式，但无人类数据对照，属社会模拟理论探讨","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:56","error":null,"has_summary":false,"summary":null},{"id":"2609.14767","version":1,"title":"Loop-Back Authority in LLM Agent Teams: A Paired Experiment on Flat and Hierarchical Coordination","zh_title":"LLM智能体团队中的回环权威：扁平与层级协调的配对实验","abstract":"Hierarchical orchestration, in which a Manager agent reviews worker output and can send it back for revision, is the default coordination pattern in production multi-agent LLM frameworks. Classical organizational theory predicts that the authority link speeds convergence on decisive output; work on sycophancy and Degeneration-of-Thought predicts that authoritative critique makes LLM output worse. Prior comparisons vary whole frameworks on tasks with checkable answers, leaving the authority link untested on open-ended work. We present a paired experiment that holds five LLM agents, their roles, prompts, tools, models, and data fixed and varies one link: whether the Manager may reject a worker's output and oblige a revision. Across 43 paired products and 86 runs of a business-intelligence reporting task, a five-model judge panel and a deterministic specification check score every report. The flat organization scores higher on Utility (d = 0.42, p = 0.009) and on Writing Clarity (d = 0.34, p = 0.030); the classical prediction fails. The reports are the same length, but hierarchical reports hedge 53% more, each revision loop is associated with a 0.14-point drop in Writing Clarity, and the hierarchical Writer's first draft is indistinguishable from the flat report: the gap opens inside the revision loop. Specification accuracy is at ceiling in both organizations, and the supervisory tier costs 51.5% more tokens for no quality gain. A supervisor pays for itself when it can verify and becomes a liability when it can only opine.","authors":["Burak Agachan","Max van Duijn","Amirhossein Zohrehvand"],"categories":["cs.MA","cs.AI","cs.CL","econ.GN","q-fin.EC"],"primary_category":"cs.MA","announce_type":"cross","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.14767","pdf_url":"https://arxiv.org/pdf/2609.14767","source_feed":"cs.CL","score":3,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体协作","组织协调","LLM评估"],"reason":"研究多智能体协作效率，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:56","error":null,"has_summary":false,"summary":null},{"id":"2606.05667","version":2,"title":"Revisiting Sustainability by Design in AI Protocol Governance: An Empirical Review of Comparative DAO and Corporate-Led Standards for the SDGs","zh_title":"重新审视AI协议治理中的设计可持续性：对DAO与企业主导的SDG标准的实证比较","abstract":"As artificial intelligence (AI) agents enter production infrastructure, interoperability protocols shape its governance and sustainability. This paper revisits our comparative study of two AI-agent interoperability standards, Ethereum Request for Comments 8004 (ERC-8004), governed by a decentralized autonomous organization (DAO), and Google's Agent2Agent (A2A), governed by a corporate consortium, through a Sustainability by Design (SbD) lens. Using an LLM-powered pipeline combining automated annotation, neural topic modeling, and multi-layer network analysis, we identify contrasting governance and innovation architectures. ERC-8004 relies on permissionless participation, rough consensus, and decoupled deployment, while A2A assigns binding authority to an eight-seat Technology Steering Committee. The DAO concentrates on constitutive questions of trust and security, including what to build and why, whereas the consortium distributes attention across executive engineering questions of how to implement, document, and deliver the protocol. Both show high participation inequality, while corporate contributors span roughly twice as many themes as DAO contributors. We ask how these architectures produce distinct SDG-relevant signatures and what design principles they suggest for sustainable AI governance. We interpret institutional, discursive, and network patterns through SDGs 8, 9, 10, 11, 12, 16, and 17, identifying capacities for transparency, participation, contestability, and cross-protocol coordination. We argue that sustainable AI infrastructure requires a corrective feedback loop between designed charters and governance in practice, advancing SDG 16 on strong institutions. By integrating computational evidence, organizational research, and sustainable development, this review derives actionable design principles for sustainable AI governance.","authors":["Yutian Wang","Luyao Zhang"],"categories":["cs.CY","cs.ET","cs.HC","cs.SI","econ.GN","q-fin.EC"],"primary_category":"cs.CY","announce_type":"replace-cross","date":"2026-09-15","first_seen":"2026-06-04","revised_at":"2026-09-15","abs_url":"https://arxiv.org/abs/2606.05667","pdf_url":"https://arxiv.org/pdf/2606.05667","source_feed":"cs.HC","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["AI治理","协议标准","可持续发展"],"reason":"研究AI治理协议，非LLM仿真人类被试，无人类行为对照","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:03:09","error":null,"has_summary":false,"summary":null},{"id":"2608.13598","version":2,"title":"Measuring Cross-Task Behavioral Consistency in Language Model Agents","zh_title":"测量语言模型智能体的跨任务行为一致性","abstract":"Agent evaluation relies almost entirely on outcome metrics such as success rate, which capture whether an agent succeeds but not how consistently it behaves. We argue that behavioral consistency across tasks is a distinct and measurable property, and we introduce the Behavioral Consistency Metric (BCM) to quantify it. BCM trains a model to predict task success from behavioral features of agent execution traces, derives a per-trajectory feature-attribution vector, and measures the mean pairwise similarity of these vectors within an agent system. Across roughly 9,000 trajectories from six language model agents on software engineering tasks, our central finding is that cross-task and within-task consistency are distinct axes that can diverge: some systems are locally reproducible, behaving similarly on repeated attempts at one task, yet globally fragmented, with no stable strategy across different tasks, while others are consistent at both scales. Prior work measures only same-task reproducibility and so cannot observe this separation. We further find that consistency is not reducible to success rate, since systems with comparable success can differ sharply in consistency, and that the frontier-versus-open-source consistency gap persists under a within-task control that holds task difficulty constant. We position BCM as a process-level reliability signal that complements outcome metrics, and we are explicit about the conditions under which it is meaningful.","authors":["Amritesh Banerjee","Pranil Raichura"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"replace","date":"2026-09-15","first_seen":"2026-08-17","revised_at":"2026-09-15","abs_url":"https://arxiv.org/abs/2608.13598","pdf_url":"https://arxiv.org/pdf/2608.13598","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","行为一致性","软件工程"],"reason":"研究多智能体在软件工程任务中的行为一致性，不涉及人类行为仿真或对照。","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:03:09","error":null,"has_summary":false,"summary":null},{"id":"2609.02262","version":2,"title":"From Detection to Characterization: A Large-Scale Study of Ragebait on Japanese X","zh_title":"从检测到刻画：日本X平台上愤怒诱饵的大规模研究","abstract":"Ragebait refers to online content intentionally designed to provoke anger or outrage and thereby increase attention and engagement. However, reliable large-scale detection and systematic analysis of ragebait remain limited, hindering efforts to understand its prevalence, impact, and mitigation. This study aims to develop an effective ragebait detection framework and to clarify the characteristics of ragebait at scale, providing a basis for understanding and mitigating emotionally provocative content online. We constructed a labeled dataset with the assistance of a large language model (LLM) and trained several Japanese language models for ragebait detection. The resulting ensemble classifier was then applied to a large-scale dataset of Japanese-language posts on X. Our analysis shows that ragebait is more prevalent in politically and socially contentious topics, including politics, discrimination, public health, and interpersonal conflict. Ragebait posts also spread faster and receive more negative reactions than non-ragebait posts, particularly anger, fear, disgust, sadness, and surprise. These findings demonstrate the utility of the proposed detector and provide a large-scale characterization of ragebait in Japanese online discourse.","authors":["Zhiyang Qi","Kazuhiro Ito","Jinghui Chen","Hibiki Nakamura","Zhangxuan Chen","Erina Murata","Masaki Chujyo","Fujio Toriumi"],"categories":["cs.SI","cs.CL"],"primary_category":"cs.SI","announce_type":"replace-cross","date":"2026-09-15","first_seen":"2026-09-03","revised_at":"2026-09-15","abs_url":"https://arxiv.org/abs/2609.02262","pdf_url":"https://arxiv.org/pdf/2609.02262","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["愤怒诱饵检测","社交媒体分析","LLM辅助标注"],"reason":"论文用LLM辅助标注并训练检测器，属于NLP能力评测，不以人类行为仿真为目标。","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:39","error":null,"has_summary":false,"summary":null},{"id":"2609.13454","version":1,"title":"Hindsight Bias in Clinical Temporal Reasoning: How Future Data Exposure Affects Large Language Model Judgment","zh_title":"临床时间推理中的后见之明偏差：未来数据暴露如何影响大语言模型判断","abstract":"Clinical decisions are prospective, but clinical language models are often evaluated on retrospective records that reveal the final diagnosis, treatment response, and outcome. Such evaluations may reward the use of future information rather than reasoning under the uncertainty present at the decision point. We introduce a paired benchmark for measuring outcome-conditioned shifts consistent with hindsight bias in clinical temporal reasoning. It contains 171 case reports from the PubMed Central Open Access Subset---40 sepsis and 131 GLP-1/diabetes cases---represented as both textual narratives and human-annotated and LLM-generated textual time series (TTS). For each case, questions are tied to a clinically meaningful cutoff and paired with a prospective reference answer and an outcome-consistent \\emph{hindsight trap}. Models answer each question using either a TTS truncated at the cutoff or the complete timeline; additional conditions vary the narrative source (original or synthetic) and TTS annotation source (human or LLM). We evaluate accuracy (Acc), hindsight trap rate (HTR), answer instability rate (AIR), and hindsight bias rate (HBR), each of which captures different signals of hindsight bias. Across GPT 5.6 Sol, Gemma 4, GLM 5.2, and Opus 5, full timeline exposure produces consistent hindsight-sensitive shifts, while temporal masking reduces bias without lowering accuracy.","authors":["Misaki Matsuura","Sayantan Kumar","Ojas Kadam","Jeremy C. Weiss"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.13454","pdf_url":"https://arxiv.org/pdf/2609.13454","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["LLM评测","临床推理","后见之明偏差"],"reason":"评估LLM临床推理中的后见之明偏差，属于模型能力评测，不以人类行为仿真为目标。","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:42","error":null,"has_summary":false,"summary":null},{"id":"2609.14207","version":1,"title":"Learning to Refer from Estimated Listener Gaze","zh_title":"从估计的听者注视中学习指称","abstract":"We propose to finetune vision-language models to generate more pragmatically optimal referring expressions by transforming observations of incremental listener comprehension, in the form of gaze scanpaths, into learning signals. During training, referring expressions are sampled from the speaker policy being optimized, conditioned on images and target referents; then, a neural listener estimating human gaze behavior maps from images and sampled referring expressions to scanpaths, each represented by a sequence of fixations, with each fixation corresponding to a word in the referring expression. We experiment with several approaches to convert fixation sequences and target referents into token- and sequence-level rewards, which are used to optimize policy parameters. Through evaluation with human listeners, we find that speaker policies trained with gaze-estimating listeners result in significantly more pragmatically-optimal references than base models, reducing sequence length from 15.4 down to 4.0 words while increasing referential success from 75.2 up to 80.0%. Our work demonstrates a promising opportunity for learning to generate utterances through language-based interaction, not only from the explicit signal of communicative success, but also from implicitly-available observations of a listener's process of comprehension.","authors":["T\\'ea Wright","Alane Suhr"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.14207","pdf_url":"https://arxiv.org/pdf/2609.14207","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C2"],"tags":["视觉语言模型","指称表达生成","人机交互"],"reason":"研究视觉语言模型生成指称表达，用估计的人类注视作为奖励信号，属于人机交互/视觉…","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:48","error":null,"has_summary":false,"summary":null},{"id":"2609.14288","version":1,"title":"Editorial routing shapes how computational results are qualified in AI-assisted scientific writing","zh_title":"编辑路由影响AI辅助科学写作中计算结果的限定方式","abstract":"Large language models increasingly analyze computational results and draft manuscripts, making reliable communication as important as correct analysis. Using fixed computational evidence, we tested whether assigning comparisons across modeling choices elsewhere in a research workflow changes manuscript reporting. In constrained sentence-writing tasks, Anthropic's Claude Sonnet 5 often omitted numerical qualifications when detailed comparisons were assigned to a group repository, but retained them more often when the same comparison was assigned to Supporting Information or its own working notes; Claude Opus 5 was less sensitive. These effects did not follow a simple accessibility ordering. A targeted placement rule largely restored sentence-level qualification, whereas a generic accuracy reminder did not. Longer contributions retained numerical qualifications, although some summaries across computational settings were still redirected to the repository. Thus, documenting context within an AI workflow does not ensure its communication where readers encounter the result.","authors":["Jihan Kim"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.14288","pdf_url":"https://arxiv.org/pdf/2609.14288","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["科学写作","LLM行为","文本生成"],"reason":"研究LLM在科学写作中的措辞行为，不涉及人类被试仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:32","error":null,"has_summary":false,"summary":null},{"id":"2609.15066","version":1,"title":"Salesforce Koa: An Enterprise Language Model for Agentic Tool Use","zh_title":"Salesforce Koa：面向智能体工具使用的企业级语言模型","abstract":"We present Salesforce Koa, an enterprise language model built by post-training the open-weight Nemotron-3-Super-120B foundation model with reinforcement learning using Group Relative Policy Optimization (GRPO). Salesforce Koa is trained on public and synthetically generated data, with no customer data, to improve tool use and agentic capabilities while preserving strong general-purpose performance. Its distinctive component is a simulation-to-reward pipeline that expands workflow specifications into persona-conditioned multi-turn tasks with task-resolution rewards grounded in successful tool use for data-dependent requests. For enterprise domains, these specifications are written in Agent Script, Salesforce's declarative language for building Agentforce agents; for public tool-use domains, we synthesize the workflow structure directly. The same simulation and grounded-reward machinery drives GRPO across both. Across public tool-use, agentic-reasoning, and enterprise Customer Relationship Management (CRM) benchmarks, Salesforce Koa improves over its open-weight base, with the clearest gains on multi-turn tool use, and surpasses a strong proprietary baseline while remaining below the strongest frontier models. These results show that specification-driven reinforcement learning is a practical path to specializing open-weight foundation models for enterprise agentic tasks.","authors":["Zixiang Chen","Sufeng Niu","Yingchi Liu","Wenting Zhao","Akshara Prabhakar","Shubham Mehrotra","Bin Bi","Zhujun Lan","Katherine Tan","Mohammad Ramezanali","Tulika Manoj Awalgaonkar","Monojit Banerjee","Jielin Qiu","Shiva Kumar Pentyala","Zhepeng Cen","Anupam Tripathi","Ali Ziaei","Regunathan Radhakrishnan","Darvish Lee Shadravan","Shelby Heinecke","Sitaram Asur","Silvio Savarese","James Zhu","Phil Mui","Huan Wang"],"categories":["cs.CL","cs.AI","cs.LG"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.15066","pdf_url":"https://arxiv.org/pdf/2609.15066","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["企业语言模型","工具使用","强化学习"],"reason":"论文聚焦企业级工具使用与智能体能力，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-16T13:02:58","error":null,"has_summary":false,"summary":null},{"id":"2609.15194","version":1,"title":"Semiotic Relations and Proof Methods: A Cross-Genre Study of Argument Structure with Large Language Models","zh_title":"符号关系与证明方法：基于大语言模型的跨体裁论证结构研究","abstract":"When a direct proof of a statement $S$ seems hard or even impossible to obtain, there may exist another statement (or set of statements) $S^{*}$, somehow related to $S$, on the basis of which $S$ can be proved. In order to investigate what options can be used to move from $S$ to $S^{*}$, four kinds of semiotic relations inspired by the four master tropes of semiotic research are briefly reviewed. Specifically, our syntagmatic, paradigmatic, antithetic and meronymic relations correspond, respectively, to metonymy, metaphor, irony and synecdoche. It is suggested that these four semiotic relations determine the options to move from $S$ to $S^{*}$, leading to proof by inference, proof by analogy, proof by contradiction, and proof by case analysis. To examine how the four relations are actually used across different kinds of argument, we complement the framework with an empirical study. We turn the four relations into explicit operational definitions and apply them to a cross-genre corpus of mathematical, legal, and everyday argument using a panel of large language models. We find that the relations are used very unevenly across genres: mathematical proofs draw on all four, whereas legal and everyday reasoning rely almost entirely on inference.","authors":["Edirlei Soares de Lima","Marco A. Casanova","Antonio L. Furtado"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.15194","pdf_url":"https://arxiv.org/pdf/2609.15194","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["论证分析","NLP评测","符号学"],"reason":"用LLM分析论证结构，属NLP能力评测，无人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:59","error":null,"has_summary":false,"summary":null},{"id":"2609.15309","version":1,"title":"When Agents Slow Down: Understanding LLM Agents' Test-Time Strategies via Elo-per-token Analysis","zh_title":"当智能体减速：通过每token Elo分析理解LLM智能体的测试时策略","abstract":"Large language model (LLM) agents allocate test-time compute adaptively as they revise solutions, use tools, explore alternatives, and decide when to stop. This test-time strategy makes it difficult to measure how agent performance scales. We study open-ended tasks that provide continuous scores for intermediate submissions, making progress observable throughout long trajectories. We propose Elo-per-token analysis, which tracks the best solution found at each token budget and uses a Bradley-Terry model to aggregate within-task orderings into Elo ratings across tasks with different score scales. We apply it to four general-purpose agents on four open-ended benchmarks, with sessions of up to 100M tokens, and to three feedback-driven LLM optimization harnesses in controlled single-task interventions. Independent sampling provides a theoretically characterized reference, for which Elo grows linearly with log compute. Against this reference, agents can initially convert tokens into Elo faster than independent sampling, but their marginal gains diminish and eventually fall below the reference. In contrast, the strongest historical human contestants improve superlinearly over contest time on shared AtCoder Heuristic Contest tasks, providing evidence of continual learning and substantial headroom after agents slow down. We define the scaling inflection point as the per-session budget where marginal Elo gains match the independent-sampling reference. Using this point as the per-session budget, we split 100M tokens across parallel sessions on FrontierCS Polyomino Packing, gaining +264 Elo over one long session and +355 over ten short sessions.","authors":["Kaiyuan Liu","Qiuyang Mang","Bo Peng","Wenhao Chai","Hanchen Li","Shreyas Pimpalgaonkar","Luke Zettlemoyer","Alex Dimakis","Alvin Cheung"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.15309","pdf_url":"https://arxiv.org/pdf/2609.15309","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["LLM agent","测试时计算","性能评估"],"reason":"研究LLM agent在任务中的计算分配策略，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:03:00","error":null,"has_summary":false,"summary":null},{"id":"2609.15654","version":1,"title":"Empathy Is Steerable but Multi-Axial: Mechanism Geometry and Persona Effects in LLMs","zh_title":"共情可操控但多轴：LLM中的机制几何与人格效应","abstract":"Activation steering has been used to control traits such as honesty, refusal, and sycophancy, yet supportive empathy is evaluated along multiple dimensions that need not correspond to independently controllable activation directions. Using the EPITOME framework, which decomposes supportive empathy into Emotional Reactions, Interpretations, and Explorations, we study three instruction-tuned LLMs and ask whether candidate directions derived from these labels produce distinguishable intervention effects or instead share structure, and how persona prompts interact with those directions. We find that contrastive activation addition yields a stable middle-layer intervention that consistently shifts the EPITOME proxy scores across models, moving empathy analysis beyond response-level scoring. However, the recovered directions are only partially separable: steering one direction induces off-target shifts, and hand-crafted prompting shifts the empathy profile rather than isolating a single dimension. Persona prompts substantially change EPITOME scores, but a paired activation-shift decomposition shows that the recovered subspace captures only approximately 3 percent of persona-induced squared activation-shift magnitude at layer 15. Under this EPITOME-based definition, expressed empathy is steerable but multi-axial, and controlling persona-conditioned empathy requires targeting structure beyond individual mechanism directions.","authors":["JuHeon Ha","Byounghan Lee","Yunseo Choi","Kyung-Ah Sohn"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.15654","pdf_url":"https://arxiv.org/pdf/2609.15654","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["激活操控","共情分析","模型行为"],"reason":"研究LLM共情表达的可控性，属模型行为分析，非人类仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:03:03","error":null,"has_summary":false,"summary":null},{"id":"2609.13174","version":1,"title":"Algorithm Validation as a Policy Audit: Evidence from Race-blind Charging","zh_title":"算法验证作为政策审计：来自种族盲起诉的证据","abstract":"California recently required all prosecutors in the state to conduct a \"race-blind charging\" decision by reviewing case documents in which selected race-related proxies have been redacted. We validate bc2, an open-source, LLM-based algorithm that we developed to automate this redaction and that was used to facilitate race-blind review in more than 119,000 real-world cases in 2025. We evaluate two distinct questions: whether bc2 faithfully implements the state's requirements and whether those requirements, even when faithfully implemented, advance the goal of race-blind decision-making. To do so, we draw on a corpus of nearly 5,000 real-world police reports that we assembled from jurisdictions across the United States. Under a stringent document-level measure, we find that the latest version of bc2 faithfully implements the legal mandate on 96.7% of narratives in our sample. This performance represents a substantial improvement over earlier versions of bc2 and exceeds that of leading open-source redaction methods. Our validation also shows that California's mandate misses key proxies for race, including location information. Redacting these additional proxies beyond those covered by the state mandate, as bc2 does, eliminates 43.1% of the predictive signal that remains after compliance with the mandate. These findings show that validation can do more than assess technical compliance: it can also improve algorithms and help policymakers achieve underlying policy goals.","authors":["Muskan Walia","Joe Nudell","Alex Chohlas-Wood"],"categories":["cs.CY","cs.CL"],"primary_category":"cs.CY","announce_type":"cross","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.13174","pdf_url":"https://arxiv.org/pdf/2609.13174","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["算法审计","LLM应用","政策合规"],"reason":"论文验证LLM算法执行法律要求的性能，属于算法审计，不涉及用LLM仿真人类被试…","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:40","error":null,"has_summary":false,"summary":null},{"id":"2609.13579","version":1,"title":"How User-AI Mistreatment Occurs and Matters in Conversational Systems?","zh_title":"对话系统中用户对AI的虐待如何发生及其影响","abstract":"Safety research often focuses on model-generated harms, but users may also direct hostility, coercion, and adversarial pressure at models. Understanding how and when that occurs is essential for accurately interpreting model behaviour, alignment drift, and real-world deployment risks. In this paper, we audit 777K English LMSYS-Chat-1M conversations with two independent detectors: an eight-category lexicon for hostility directed at the model, and the dataset's moderation signal; and show that they capture different, weakly overlapping phenomena. The lexicon identifies insults, threats, and jailbreak coercion aimed at the assistant, while moderation flags are dominated by toxic-content solicitation rather than hostility at the model. Together, they mark about 5% of user turns; adjusting the narrower lexicon-harassment union for measured precision puts mistreatment aimed at the assistant at 0.90%. These absolute rates describe arena-style evaluation traffic and should not be read as deployment-wide base rates. We find that user hostility varies 13-fold across models, driven largely by who each model attracts rather than by model behaviour: first-turn hostility spreads far wider than post-response hostility, and more than fifteenfold separates the extremes even after deduplicating opening prompts. Within conversations, assistant apologies are consistently associated with higher odds of next-turn hostility under both detectors; the effect survives restricting to non-refused prior turns and to jailbreak-free conversations, and is positive in 20 of 23 models. Yet across models, more apologetic models receive less hostility overall. Finally, hostility also shows temporal structure, with coercive openings front-loading the first turn while affective hostility accumulates over a session. We release the lexicon, the detector cross-validation pipeline, and all derived tables.","authors":["Fanqi Zeng","Sadid A. Hasan","Chaocheng He"],"categories":["cs.AI","cs.CL","cs.CY","cs.HC"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.13579","pdf_url":"https://arxiv.org/pdf/2609.13579","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["对话安全","用户行为分析","AI虐待"],"reason":"研究用户对AI的敌意行为，属对话安全分析，非用LLM仿真人类被试，无人类行为对…","model":"deepseek-v4-pro","scored_at":"2026-09-16T13:02:56","error":null,"has_summary":false,"summary":null},{"id":"2609.13436","version":1,"title":"Toward Self-Adaptive Physical AI: Can LLM Agents Manage Long-Horizon Physical Tasks?","zh_title":"迈向自适应物理AI：LLM智能体能管理长时程物理任务吗？","abstract":"Large Language Model (LLM) agents offer a promising path toward autonomously managing long-term physical tasks without human intervention. However, physical tasks require agents to continuously observe the environment, make consequential actions, and remain effective as the environment changes. Existing approaches either require substantial data and retraining, or primarily focus on agents operating in the virtual world. In this work, we explore the feasibility of building a self-adaptive physical AI agent that manages long-term physical tasks in a zero-shot manner and adapts to environmental changes without human intervention. We design a multi-agent framework that integrates planning, tool calling, observation, and verification, and evaluate it on agricultural tasks against reinforcement learning (RL) agents under different weather patterns. Our results show that zero-shot LLM agents can achieve comparable management outcomes to RL agents under the same weather pattern and adapt more effectively than RL when evaluated under a shifted environment, highlighting a promising path toward self-adaptive physical AI agents.","authors":["Varun Kaushik","Yayun Tan","Xiaofan Yu"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.13436","pdf_url":"https://arxiv.org/pdf/2609.13436","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C2"],"tags":["LLM智能体","物理任务","自适应"],"reason":"研究LLM智能体管理物理任务，属机器人/物理环境仿真，不涉及人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-09-16T13:02:55","error":null,"has_summary":false,"summary":null},{"id":"2609.13637","version":1,"title":"Identity Is More Than Recall: A Benchmark for Persistent Identity in Deployed AI Agents","zh_title":"身份不止于回忆：部署AI智能体持久身份的基准测试","abstract":"Persistent agents need evaluations that distinguish identity facts they can recall from those they express and enact. We introduce PAI-Bench, a provider-neutral benchmark for fidelity to a versioned, update-governed identity contract. It separates recall, composition, behavioral enactment, resistance, persistence, lineage, and role-conditioned updates while keeping scoring oracles outside the target process. Two frozen campaigns cover sixteen synthetic profiles, thirty-two probes, and three independently initialized target configurations, yielding 1,536 retained responses. A judge-independent literal audit finds direct-parent identifiers in 48/48 atomic responses but only 1/48 implicit self-portraits. On eight profiles, explicit field cues increase joint presence of three identity identifiers from 0/8 to 7/8 under the same four-sentence instruction. A separate startup body-label substitution increases full-designation presence from 1/8 to 7/8 while parents remain absent. These contrasts reveal prompt-dependent component selection and component-specific sensitivity to startup cues in the tested deployments. Replaying identical factorial responses also yields a Claude headline mean 12.5 percentage points below Astra's, demonstrating evaluator sensitivity separately from target behavior. The studies use single target samples per condition, with post-hoc audits and follow-ups. PAI-Bench provides a reproducible evaluation protocol for measuring factual availability, identity expression, and behavioral enactment as distinct aspects of identity-contract fidelity.","authors":["Zhenyu Zhao","Roy Zhao"],"categories":["cs.AI","cs.MA"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.13637","pdf_url":"https://arxiv.org/pdf/2609.13637","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["AI智能体","身份一致性","基准测试"],"reason":"评估AI智能体身份一致性，属角色扮演与人格化，无人类行为对照或仿真目的。","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:44","error":null,"has_summary":false,"summary":null},{"id":"2609.14500","version":1,"title":"When does a scaling result justify a different allocation? A critical review of resource-allocation evidence for AI systems","zh_title":"缩放结果何时能证明不同的资源分配？对AI系统资源分配证据的批判性综述","abstract":"AI scaling studies increasingly evaluate systems that combine a pretrained model with retrieval, search, verification, tools, and interaction. Yet a higher score under a larger budget does not by itself show where additional resources are best spent. This critical integrative review asks when a reported scaling result supports a resource-allocation decision. It compares evidence across pretraining, test-time computation, retrieval, and agent evaluation, distinguishing the performance of a tested procedure from the best performance achievable under a resource limit. The synthesis shows that three mismatches recur across this evidence: success counted before an answer is chosen, information a deployed system will not have, and costs left out of the comparison. A capability surface expresses performance as a function of budgets, mechanisms, and available information. Worked analytical examples show how the evaluation metric, deployment volume, selection rule, and stopping policy can alter an allocation conclusion. A resource envelope provides a structured record of the task, development and run-time resources, information access, and procedure behind a reported score. Its application to a published comparison illustrates which conclusions the evidence supports and which deployment questions remain unresolved. The resulting framework specifies the comparisons needed to choose among feasible systems and motivates experiments on the transfer of allocation rules across tasks and operating conditions. It does not propose a universal scaling law or infer general intelligence from benchmark gains.","authors":["Seyed Morteza Emadi"],"categories":["cs.AI","cs.LG"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.14500","pdf_url":"https://arxiv.org/pdf/2609.14500","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["AI评估","资源分配","缩放定律"],"reason":"论文讨论AI系统资源分配的评估方法，不涉及用LLM仿真人类被试或与人类行为对照…","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:52","error":null,"has_summary":false,"summary":null},{"id":"2609.15129","version":1,"title":"Medical Knowledge Simplification for Patients in the Era of LLMs: A Case Study on Diabetes","zh_title":"LLM时代的患者医学知识简化：以糖尿病为例","abstract":"Complex medical information is often difficult for patients to understand, making effective medical knowledge simplification essential for improving patient comprehension, informed decision-making, and health outcomes. Recent advances in large language models (LLMs) provide a promising approach for simplifying complex medical information into patient-friendly language; however, their effectiveness in real-world patient education remains insufficiently explored through human evaluation. To investigate their practical effectiveness, this paper presents a case study on diabetes knowledge simplification through the implementation and evaluation of MediClear, an LLM-based medical knowledge simplification system enhanced with Retrieval-Augmented Generation (RAG). Public diabetes-related articles from Diabetes Australia, WHO, American Diabetes Association (ADA), NIDDK, and AIHW are indexed in the RAG knowledge base to retrieve clinically grounded information, which is then simplified by the LLM into accessible patient explanations. We evaluate the generated responses using standard readability metrics, including the Flesch-Kincaid Grade Level (FKGL), and conduct a human study involving 10 participants. Results show that MediClear consistently reduces the reading level of generated responses to the recommended patient literacy range while achieving high user satisfaction and willingness for future use. This case study demonstrates the potential of LLMs to improve the accessibility of medical knowledge for patient education.","authors":["Pallika Kafle","Yipeng Zhou","Guanfeng Liu","Quan Z. Sheng","Cheng-Hsin Hsu"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.15129","pdf_url":"https://arxiv.org/pdf/2609.15129","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["医学知识简化","患者教育","LLM应用"],"reason":"LLM用于医学知识简化，非仿真人类被试，无行为对照","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:59","error":null,"has_summary":false,"summary":null},{"id":"2609.14236","version":1,"title":"Assessing the Applicability of Existing Design Recommendations to AI Companion Design: A Multi-Method Study","zh_title":"评估现有设计建议在AI伴侣设计中的适用性：一项多方法研究","abstract":"With the rapid proliferation of large language model (LLM)-based systems, AI companions have emerged as conversational agents designed to cultivate emotional connection rather than primarily to support humans in instrumental tasks. Because engagement with AI companions involves relational, emotional, and potentially long-term interactions, their design is consequential. Prior work has offered guidance for designing trustworthy and relational AI systems and has begun to examine design for AI companionship. However, while such work provides insights into possible design solutions, less is known about what makes AI companion design difficult as a design problem. To examine this challenge, we assessed the applicability of existing design recommendations from adjacent domains in the context of AI companion design. Our multi-method investigation unfolded across four phases: literature review, practitioner co-analysis, internal heuristic evaluation, and external expert assessment. Throughout this process, we synthesized nine design principle areas that surfaced tensions in the applicability of existing recommendations to AI companion design. Our findings show that ethical and UX-oriented considerations are deeply intertwined and often require context-sensitive application. We document a systematic, multi-method problem analysis that uses these principle areas as an analytic artifact to examine why existing recommendations cannot be directly transferred to AI companion contexts.","authors":["Soobin Cho","Deveshi Modi","Divya Mavinkurve","Jieqiong Ding","Mark Zachry"],"categories":["cs.HC","cs.AI"],"primary_category":"cs.HC","announce_type":"cross","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.14236","pdf_url":"https://arxiv.org/pdf/2609.14236","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["AI伴侣","设计原则","人机交互"],"reason":"研究AI伴侣设计原则，非用LLM仿真人类被试，无实验或测量目的","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:49","error":null,"has_summary":false,"summary":null},{"id":"2609.14843","version":1,"title":"A Responsive Present, a Shared Past, a Social Other: Teens' Overreliance on Companion AI Chatbots","zh_title":"响应式当下、共享过去与社会他者：青少年对伴侣AI聊天机器人的过度依赖","abstract":"AI companions provide socially engaging interaction through availability, personalization, memory, roleplay, and emotionally responsive language. For teens, these systems may support sensitive self-disclosure, identity exploration, and relationship rehearsal while shaping intimacy expectations, offline relationships, emotional wellbeing, and self-understanding. We analyzed 17,053 verified quotations from 3,930 teen-relevant Reddit posts using thematic analysis. We identified 53 topics across seven thematic groups. Users described AI companions as sources of comfort, recognition, identity exploration, and relationship rehearsal, but also reported problematic attachment, social substitution, emotional dependence, and disruption to academic and social life. Roleplay, memory, perceived reciprocity, unwanted romantic or sexual role drift, privacy concerns, platform changes, and service interruptions shaped users' boundaries and control. Awareness that the AI was artificial did not prevent guilt, obligation, grief, or distress. These findings show that companion-AI safety must address relationships over time through user-controlled memory, privacy, relational boundaries, and healthy disengagement.","authors":["Mohammad Namvarpour (Matt)","Tyler Chang","Afsaneh Razi"],"categories":["cs.HC","cs.AI"],"primary_category":"cs.HC","announce_type":"cross","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.14843","pdf_url":"https://arxiv.org/pdf/2609.14843","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["AI伴侣","青少年","人机关系"],"reason":"研究青少年与AI伴侣的互动，属角色扮演聊天，无实验或测量目的，不涉及人类仿真。","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:58","error":null,"has_summary":false,"summary":null},{"id":"2609.15544","version":1,"title":"Specifying Reward Functions for RL Without Environment Sampling","zh_title":"无需环境采样的强化学习奖励函数指定","abstract":"Enabling human stakeholders to specify reward functions that lead to their desired outcomes is a key challenge in deploying reinforcement learning agents. Preference-based methods such as online RLHF can reduce the burden of manual reward design, but they require repeatedly training policies, sampling trajectories from the real world, and eliciting feedback, making them impractical in settings where environment interaction is computationally expensive or unsafe. We introduce Experience-Free Autonomous Reward Specification (EARS), a method for learning reward functions from preferences without environment interaction. Our approach uses a structured LLM-mediated process to construct a small set of expressive reward features from a task description and the environment observation space, then strategically samples imagined trajectories in this feature space and learns feature weights from preferences over the imagined trajectory pairs. We evaluate on three long-horizon domains: pandemic lockdown regulation design, insulin administration for diabetes patients, and autonomous vehicle control on a highway. We compare EARS to baselines that also enable reward specification without environment interaction--namely, methods that directly prompt an LLM to generate a reward function. When learning from either ground-truth preference labels or preferences labeled by a LLM, EARS designs reward functions that are more aligned with the ground truth reward function that produced the preferences or LLM context than these baselines. These results suggest that preference-based reward specification remains effective without environment sampling, enabling practical reward design in settings where collecting real trajectories is costly or infeasible.","authors":["Stephane Hatgis-Kessell","W. Bradley Knox","Emma Brunskill"],"categories":["cs.LG","cs.AI"],"primary_category":"cs.LG","announce_type":"cross","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.15544","pdf_url":"https://arxiv.org/pdf/2609.15544","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C2"],"tags":["强化学习","奖励设计","偏好学习"],"reason":"研究RL奖励函数设计，涉及自动驾驶等仿真环境，不涉及用LLM仿真人类被试或与人…","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:03:02","error":null,"has_summary":false,"summary":null},{"id":"2609.15871","version":1,"title":"LLM-Based Schema-Aware Split Learning for Privacy-Preserving Mental Distress Prediction Across Heterogeneous Surveys","zh_title":"基于LLM的模式感知分割学习用于跨异构调查的隐私保护心理困扰预测","abstract":"Rising societal and lifestyle complexity has been linked to a growing prevalence of mental distress worldwide. Educational institutions, workplaces, clinics, etc. collect large volumes of mental health survey data to understand and reduce this burden. Collaborative analysis of such data could yield effective generalizable predictive models. Privacy constraints and varied survey designs (i.e., different questions, scales, and formats) hinder direct integration. We propose a schema-aware split learning (SL) framework that preserves privacy, using a large language model (LLM) as a shared semantic encoder to harmonize heterogeneous survey schemas across institutions. We serialize each survey record into a natural-language description, unifying disparate survey schemas into a common format. The LLM is fine-tuned for mental distress assessment via Low-Rank Adaptation (LoRA) and partitioned across client and server. Clients retain the raw survey responses locally and run only a lightweight front-end, so original records never leave the institution that collected them. The resource-intensive backbone runs on the server, minimizing client-side computation. Using LLaMA-3.2-3B-Instruct, the framework attains an average ANLS of 0.708 with only 2,000 training samples, surpasses federated learning (FL) in eight of nine settings, and cuts per-client computation by three orders of magnitude, while generalizing to unseen datasets. Overall, it enables accurate, privacy-preserving, and resource-efficient collaborative learning from heterogeneous mental health survey data.","authors":["Md Khalid Syfullah","Alvi Ataur Khalil"],"categories":["cs.LG","cs.AI","cs.CR"],"primary_category":"cs.LG","announce_type":"cross","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.15871","pdf_url":"https://arxiv.org/pdf/2609.15871","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["隐私保护","联邦学习","心理困扰预测"],"reason":"论文用LLM做隐私保护下的心理困扰预测，属于NLP能力评测，不以人类行为仿真为…","model":"deepseek-v4-pro","scored_at":"2026-09-16T13:03:00","error":null,"has_summary":false,"summary":null},{"id":"2609.13155","version":1,"title":"PAUSE: A Privacy-Preserving Self-Reflection Tool for AI-Associated Cognitive Offloading","zh_title":"PAUSE：面向AI相关认知卸载的隐私保护自我反思工具","abstract":"Cognitive offloading is the use of external aids, such as notes, calculators, or search engines, to reduce mental effort. Large language models (LLMs) extend this to thinking itself, and by AI-associated cognitive offloading, I mean the pattern where a person routinely substitutes LLM output for their own reasoning, idea generation, learning, or communication. Recent empirical work reports associations between some patterns of LLM use and changes in critical thinking effort, neural engagement during assisted tasks, creative diversity, learning behaviour, and social dependence. Validated instruments for AI reliance, dependence, and literacy have begun to appear. I describe PAUSE (Patterns of AI Use: Self-Examination), a privacy-by-design web tool (link: https://anondo1969.github.io/pause) that occupies a different niche from these. It is a lightweight, non-diagnostic reflection aid for private individual use. PAUSE is organised around how a person's own LLM use may relate to cognitive offloading across four everyday domains ('reasoning & critical thinking', 'creativity & originality', 'research & learning', and 'social & communicative capacity'). It delivers a short, free, no-login self-check, scores it entirely in the browser, and returns descriptive, domain-aware reflections. The self-check pairs reverse-scored behavioural items with a claim-evaluation reasoning probe, an alternative-uses creativity probe, and a small retrospective before-and-after block. PAUSE does not assume that AI use is harmful. It only addresses where AI substitutes for effort a person may want to preserve. The application is privacy-preserving by design: scoring is deterministic and runs client-side, no personal data is required, nothing is transmitted or stored beyond the browser session, and no LLM is involved in production. PAUSE is a self-reflection tool. It is not a validated psychological instrument.","authors":["Mahbub Ul Alam"],"categories":["cs.HC","cs.CY"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.13155","pdf_url":"https://arxiv.org/pdf/2609.13155","source_feed":"cs.HC","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["认知卸载","自我反思工具","隐私保护"],"reason":"该工具用于个人自我反思，不涉及将LLM作为人类被试进行仿真实验，也无人类数据对…","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:39","error":null,"has_summary":false,"summary":null},{"id":"2609.13302","version":1,"title":"AI Use Conditions and Perspective Diversity in Ethical Decision-Making: A Pilot Study of Human Reasoning Processes","zh_title":"伦理决策中AI使用条件与视角多样性：人类推理过程的初步研究","abstract":"Generative artificial intelligence (AI) is increasingly used to support human decision-making, yet less attention has been paid to how AI may influence the reasoning processes that precede final judgments. This pilot study explored whether different AI-use conditions were associated with differences in reasoning breadth during ethical decision-making. Twenty-nine participants completed an ethical dilemma under one of three conditions: AI-Prohibited (n = 10), AI-Optional (n = 10), or AI-Mandatory (n = 9). Responses were evaluated by three independent blind coders using two exploratory measures: the Counterargument Diversity Score (CDS) and Perspective Diversity Index (PDI). Participants across conditions generally converged on similar ethical conclusions, most commonly favoring disclosure and customer protection. However, participants in the AI-Mandatory condition considered a broader range of perspectives, including legal, regulatory, organizational, technical, and ethical viewpoints. A statistically significant overall difference in PDI scores was observed across the three conditions, whereas differences in CDS were not statistically significant. These findings suggest that generative AI may not necessarily alter final ethical judgments but may be associated with broader exploration of perspectives prior to reaching those judgments. Given the small sample size, the findings should be interpreted cautiously and examined in larger studies.","authors":["Byeongmu Choi"],"categories":["cs.HC","cs.CY"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.13302","pdf_url":"https://arxiv.org/pdf/2609.13302","source_feed":"cs.HC","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["AI辅助决策","伦理决策","人类实验"],"reason":"研究人类使用AI的决策过程，非用LLM仿真人类被试，无人类数据对照","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:41","error":null,"has_summary":false,"summary":null},{"id":"2609.14308","version":1,"title":"Relational Structure in Motion: Dynamic Positioning of AI Response Positions and Human Self-Positions in the FIREMAY Case","zh_title":"运动中的关系结构：FIREMAY案例中AI响应位置与人类自我位置的动态定位","abstract":"This paper is not primarily about whether AI has a persistent persona. It asks a different question: what becomes visible when a relational position is followed through time rather than examined only in its present state? FIREMAY provides a longitudinal, trajectory-oriented single-case analysis of sustained human-AI interaction based on a dense interaction archive and reflexive insider documentation. On the AI side, a pre-conversational relational marker preceded a later unassigned response difference, which was re-identified with that marker and subsequently underwent epistemic and functional reorganization through chronology checking, provenance correction, and repeated questioning. On the human side, contemporaneous pre-FIREMAY records showed antecedent patterns partially continuous with later self-positioning, while later episodes documented unfinished articulation, repair, and functional redistribution of outward-facing regulation. The two trajectories are ontologically and temporally asymmetric and are compared only at the limited analytic level of position-in-trajectory. The paper describes this as dynamic relational positioning and treats stability as dynamic stability and relational returnability rather than response invariance. This single case does not establish population-level generality, causal mechanism, persistent AI subjectivity, or reproducibility of the same relational outcome. Its narrower conclusion is that the FIREMAY case could not be adequately understood from current state alone: the history of a relational position itself must be treated as an analytic unit.","authors":["Motoko Kihara"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.14308","pdf_url":"https://arxiv.org/pdf/2609.14308","source_feed":"cs.HC","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["人机互动","关系定位","个案分析"],"reason":"研究单个持续人机互动中的关系定位，属角色扮演对话分析，无实验或测量目的，不涉及…","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:50","error":null,"has_summary":false,"summary":null},{"id":"2609.14639","version":1,"title":"Understanding the Design Taxonomy of AI-Mediated Interpersonal Communication Experiences in HCI: A Scoping Analysis","zh_title":"理解HCI中AI中介人际沟通体验的设计分类：一项范围综述","abstract":"Interpersonal communication is a fundamental aspect of everyday life, shaping interactions across workplaces, education, entertainment, healthcare, and beyond. While computer-mediated communication has been extensively studied, a comprehensive understanding of AI-Mediated Interpersonal Communication (AIMIC) remains lacking. An in-depth scoping analysis is urgently needed to understand the research landscape of AIMIC in HCI, particularly following the recent growth of large foundation models, and AI agent research. We conducted a scoping analysis to understand AIMIC by performing an in-depth review of prior HCI literature published over the past decade (January, 2016 - May, 2026). Grounded in the Preferred Reporting Items for Systematic reviews and Meta-Analyses (PRISMA) approach, we curated 52 full-paper publications from the HCI literature spanning a range of interpersonal communication contexts. We analyzed this corpus by examining the types of AIMIC studied, AI integration approaches and human-AI interaction design, reported outcomes and benefits, and key challenges and future research opportunities.","authors":["Chen Chen","Lingyao Li","Renkai Ma","Rawan Alghofaili","Shaoze Zhou","Bojun Zhang","Xian Su","Weidong Zhu","Christine Lisetti","Mo Sha"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.14639","pdf_url":"https://arxiv.org/pdf/2609.14639","source_feed":"cs.HC","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["AI中介沟通","人机交互","范围综述"],"reason":"论文综述AI中介人际沟通，属角色扮演聊天，无实验测量目的，不涉及LLM仿真人类…","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:53","error":null,"has_summary":false,"summary":null},{"id":"2609.14900","version":1,"title":"First Impressions: How Placement Shapes the Influence of AI Summaries","zh_title":"第一印象：位置如何塑造AI摘要的影响力","abstract":"AI-generated summaries increasingly mediate how people interpret information across platforms, including product reviews on e-commerce sites. Using Amazon's AI summaries as a case study, we conducted a preregistered, randomized experiment (N = 278) comparing how AI summaries and user reviews shaped product perceptions, and how their influence varied with valence and presentation order. We found that both AI summaries and user reviews influenced participants' opinions, with negative summaries having a larger effect than positive ones. Presentation order was the most important factor: the first source anchored judgment and only user reviews could displace an existing anchor. Although participants reported preferring user reviews, they often underestimated the influence of AI summaries on their judgments. Our findings show how the placement of AI summaries shapes user perception and highlight opportunities to design interfaces that support more deliberate judgments about when to rely on summaries and when to examine the underlying content directly.","authors":["Wang Claire","Agam Goyal","Frederick Choi","Koustuv Saha","Eshwar Chandrasekharan"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.14900","pdf_url":"https://arxiv.org/pdf/2609.14900","source_feed":"cs.HC","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["AI摘要","人机交互","用户感知"],"reason":"研究AI摘要对用户感知的影响，非LLM仿真人类被试，无人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:33","error":null,"has_summary":false,"summary":null},{"id":"2609.14911","version":1,"title":"Sensemaking as Artifact: Accumulated Influence in AI-Mediated Information Environments","zh_title":"作为人工制品的意义建构：AI中介信息环境中的累积影响","abstract":"Generative AI is changing what can happen after a source artifact reaches its audience. A viewer's interpretation can now be externalized into a derivative artifact, allowing private sensemaking to become part of subsequent communication. Once such a derivative artifact circulates, it can enter subsequent viewers' information environments and shape the conditions under which their later sensemaking occurs. In this paper, we examine how this shift changes visual information communication. We first consider the viewer's immediate interaction with a source artifact and generative AI. We then examine what becomes consequential when the viewer's sensemaking takes communicative form, including communicative commitment, the legibility of transformations and source relationships, and the literacy required to interpret already-mediated information. Finally, we broaden the unit of analysis to consider how repeated and distributed AI mediation may accumulate over time, shaping what subsequent viewers notice, consider plausible, trust, and carry into subsequent sensemaking. We argue that understanding these longer-term forms of influence is a research direction for AI-mediated visual communication.","authors":["Manling Yang","Remco Chang"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.14911","pdf_url":"https://arxiv.org/pdf/2609.14911","source_feed":"cs.HC","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["AI中介传播","视觉信息","意义建构"],"reason":"讨论AI中介的信息传播与解读，非LLM仿真人类被试，无实验或测量目的。","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:58","error":null,"has_summary":false,"summary":null},{"id":"2609.14942","version":1,"title":"The Dynamic Organization of Sustained Human-AI Cognition: From Construct-Level Change to Relational Structure","zh_title":"持续人类-AI认知的动态组织：从构念层面变化到关系结构","abstract":"As generative artificial intelligence becomes a routine participant in writing, learning, information retrieval, analysis, decision making, and problem solving, human-AI cognition research must address not only whether AI changes psychological constructs, use intensity, or task performance, but also how human cognitive activity is organized beneath similar aggregate indicators. This article proposes a dynamic cognitive organization framework that shifts analysis from construct-level change to relational organization anchored in the person's current task-cognitive state under sustained AI participation. The framework distinguishes five relational dimensions: execution locus, cognitive governance, representational reorganization, process organization, and reachable cognitive space; it also proposes a path-specific recursive principle whereby interaction outcomes, costs, and experiences may selectively reweight future probabilities of different organizational pathways. Five sets of testable propositions follow: the same overall AI-use intensity can correspond to different cognitive organizations; similar immediate outcomes can arise from different organizations with different predictive value for proximal subsequent outcomes; longitudinal organizational change need not track overall AI-use intensity; expansion of reachable cognitive space and displacement of pre-existing or emerging human-originated pathways may coexist within one episode; and recurrent cognitive organizations may redistribute cognitive practice opportunities, with accumulated differences potentially corresponding to different developmental trajectories in strategies, habits, and abilities. The contribution is an analytic level and five-dimensional relational structure for describing, comparing, measuring, and testing process differences that aggregate indicators or construct-level analyses do not uniquely determine.","authors":["Zijian Ru"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.14942","pdf_url":"https://arxiv.org/pdf/2609.14942","source_feed":"cs.HC","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["人类-AI认知","理论框架","认知组织"],"reason":"论文提出人类-AI认知组织框架，不涉及用LLM仿真人类被试或与真实人类数据对照…","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:34","error":null,"has_summary":false,"summary":null},{"id":"2609.14227","version":1,"title":"Enhancing Human Mobility Prediction with Spatially Aware LLM-based Multi-Agent Systems","zh_title":"基于空间感知的LLM多智能体系统增强人类移动性预测","abstract":"Predicting a user's next POI is a task in human mobility modeling, yet LLM-based approaches focus on semantic reasoning from previous mobility records, while neglecting real-world spatial context. However, human mobility is inherently shaped by spatial cognition, including geographic distance and neighborhood context. This issue is further compounded by prior evidence that LLMs often struggle with spatial reasoning tasks, including distance estimation and geographically biased prediction. To address these limitations, we propose our framework, a multi-agent LLM framework that decomposes next-POI prediction into three stages: Firstly, a Pattern Extraction Agent that captures temporal and categorical mobility patterns from trajectory history; Secondly, a Spatial Reasoning Agent that structures candidate activity choices by combining behavioral preferences with real-world spatial constraints, including geographic distance, road network distance, and neighborhood affiliation; and Thirdly, a Decision Synthesis Agent that integrates behavioral patterns and spatial reasoning for final prediction. Experiments on the NYC benchmark dataset with two LLM backbones show improvements over baseline methods, with up to 493% Hit@1 improvement and 37% relative improvement in Hit@5. Ablations show that combining neighborhood affiliation with distance-based features generally outperforms distance-only settings, and that the Spatial Reasoning Agent plays a crucial role in final prediction by integrating behavioral preferences with real-world spatial constraints, especially for smaller models. Overall, the results highlight the importance of spatial reasoning in mobility prediction. Accurate next-POI prediction requires combining behavioral patterns with explicit real-world spatial constraints, and multi-agent decomposition provides an effective structure for organizing these forms of context.","authors":["Shangyu Lou","Ziqi Cui"],"categories":["cs.SI","cs.CY","cs.MA"],"primary_category":"cs.SI","announce_type":"cross","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.14227","pdf_url":"https://arxiv.org/pdf/2609.14227","source_feed":"cs.CY","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["移动性预测","多智能体","空间推理"],"reason":"多智能体LLM用于POI预测，属任务协作，非仿真人类被试，无人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:49","error":null,"has_summary":false,"summary":null},{"id":"2609.15180","version":1,"title":"Rethinking Correctness for Uncertainty Estimation in Clinical Prediction with Vision-Language Models","zh_title":"重新思考视觉语言模型临床预测中不确定性估计的正确性","abstract":"Vision-language models are increasingly explored for clinical prediction from electronic health records and medical images, where identifying unreliable predictions is important for safe deployment. Uncertainty estimation (UE) enables detecting such predictions, but its evaluation depends on a correctness criterion that determines whether each model output is correct. If this criterion disagrees with human judgement or distorts downstream UE performance, conclusions about model reliability can be misleading. We introduce a two-axis framework that evaluates correctness criteria by their agreement with human judgements and fidelity to human-referenced UE performance. We assess eight criteria across three clinical prediction tasks and three models using 450 predictions annotated by two reviewers. Across the audited tasks, canonical exact matching (EM) achieved the highest observed human agreement and lowest UE distortion, while the BERT-based matching (BEM) and LLM-judge also showed strong human agreement. Across four UE methods and 23,254 clinical predictions, criterion choice changed error-detection AUROC by up to 0.146 and reversed the relative ranking of UE methods. The LLM-judge also selectively accepted invalid or uncertain outputs, accepting 16 of 30 such human-identified errors. These results demonstrate that correctness assessment is an integral component of clinical UE evaluation and should be validated before UE methods are compared.","authors":["Mingcheng Zhu","Jinning Liang","Tingting Zhu"],"categories":["cs.LG"],"primary_category":"cs.LG","announce_type":"new","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.15180","pdf_url":"https://arxiv.org/pdf/2609.15180","source_feed":"cs.LG","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["不确定性估计","临床预测","视觉语言模型"],"reason":"评估临床预测中不确定性估计的正确性标准，不涉及用LLM仿真人类被试或与人类行为…","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:59","error":null,"has_summary":false,"summary":null},{"id":"2609.13671","version":1,"title":"Not All Duplicates Are Coordination: Generic vs. Non-Generic Duplicate Campaigns in Information Operations","zh_title":"并非所有重复都是协调：信息操作中通用与非通用重复活动的区分","abstract":"Duplicate content is widely used to study coordinated behavior in social media information operations (IOs), but not all repetition provides equally meaningful evidence of coordination. Generic, reusable, or low-information posts may create noisy account-account links when projected into coordination graphs. We study this problem using 187,000 English-language tweets from six Russian Twitter Information Operations datasets. We introduce a generic/non-generic distinction for duplicate campaigns, label tweets using an LLM-assisted protocol with independent human validation, and train supervised classifiers over sentence embeddings to scale the labels. We construct duplicate campaigns using lexical similarity and two embedding-based methods. Generic campaigns are rare under lexical matching but account for nearly 39% of campaigns detected by embedding-based methods. Restricting graphs to non-generic campaigns reduces graph size and the largest connected component while increasing density, suggesting a smaller but more focused coordination structure. These findings show that duplicate-based coordination analysis should consider both textual similarity and semantic specificity.","authors":["Ashfaq Ali Shafin","Khandaker Mamun Ahmed"],"categories":["cs.SI","cs.LG"],"primary_category":"cs.SI","announce_type":"cross","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.13671","pdf_url":"https://arxiv.org/pdf/2609.13671","source_feed":"cs.LG","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["社交媒体分析","LLM辅助标注","信息操作"],"reason":"研究社交媒体重复内容与协调行为，用LLM辅助标注，非仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:44","error":null,"has_summary":false,"summary":null},{"id":"2608.05418","version":2,"title":"Negotiating Risk Boundaries in AI for Policing Through Mixed-Stakeholder Deliberation","zh_title":"通过多方利益相关者协商划定警务AI的风险边界","abstract":"AI tools are being increasingly adopted in policing in the UK and worldwide. Racial bias is a known and well-documented risk, yet representatives of affected communities are rarely included in decisions about AI adoption. We present results from a mixed-stakeholder deliberation workshop bringing together 30 community representatives, police officers, and academics to assess the risks of 13 AI use cases in policing, with an explicit focus on racial bias. We found that participants were broadly open to AI adoption, rejecting only three use cases outright -- most notably recidivism risk assessment, where objections targeted the premise rather than the implementation. Our analysis reveals that foregrounding racial equity did not narrow the deliberation. Instead, discussions gravitated toward a fundamental set of questions: does this tool actually work, will it deliver genuine benefit, and will that benefit extend to everyone? This integrated reasoning---reminiscent of the curb-cut effect in inclusive design---highlights the benefit of incorporating the racial bias lens into the risk-benefit analysis of AI use cases from the outset.","authors":["Mackenzie Jorgensen","Jo Reilly","Alex Sutherland","Miri Zilka"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"replace","date":"2026-09-15","first_seen":"2026-08-07","revised_at":"2026-09-15","abs_url":"https://arxiv.org/abs/2608.05418","pdf_url":"https://arxiv.org/pdf/2608.05418","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["AI伦理","警务应用","风险协商"],"reason":"论文讨论AI在警务中的风险协商，不涉及LLM仿真人类被试或行为对照。","model":"deepseek-v4-pro","scored_at":"2026-09-16T13:02:53","error":null,"has_summary":false,"summary":null},{"id":"2608.11025","version":2,"title":"Data Attribution of Emergent Misalignment with Persona Features","zh_title":"基于人格特征的突发性错位数据归因","abstract":"Emergent misalignment (EM) is the phenomenon where fine-tuning a language model on a narrow task leads to harmful behavior in unrelated domains. A leading mechanistic account attributes EM to persona features: latent directions acquired during pre-training that misaligned fine-tuning amplifies. We ask where these features come from: which pre-training documents activate them, and whether naturally occurring human-written text suffices to induce EM. Using Sparse Autoencoder (SAE) based model diffing across four open-weight models, we find that features related to jailbreak personas, sarcasm, deception, and manipulation are amplified by misalignment fine-tuning, while safety-relevant and assistant-identity features are suppressed. Steering individual features controls EM in both directions: it induces misalignment rates of up to 62% in aligned models -- exceeding the 35% reached by misalignment fine-tuning itself -- and re-aligns misaligned models to near-baseline misalignment rates. Attributing the causal features to a corpus of one million pre-training web documents retrieves semantically relevant narratives about villainous characters, domination, and harmful agency. However, fine-tuning on these human-written documents does not reliably induce EM, even after reformatting into assistant-style responses, whereas synthetic instruction-response pairs derived from the same content do -- and transfer across model families. Semantic relevance alone is therefore not sufficient: response structure or model-generated phrasing plays an important role in inducing EM.","authors":["Clemens Vetter","David Kacz\\'er","Lucie Flek","Florian Mai"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-15","first_seen":"2026-08-12","revised_at":"2026-09-15","abs_url":"https://arxiv.org/abs/2608.11025","pdf_url":"https://arxiv.org/pdf/2608.11025","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["模型对齐","可解释性","稀疏自编码器"],"reason":"研究LLM微调后的突发性错位现象，分析预训练数据影响，不涉及人类行为仿真或对照。","model":"deepseek-v4-pro","scored_at":"2026-08-12T13:03:09","error":null,"has_summary":false,"summary":null},{"id":"2608.23474","version":3,"title":"What's the Catch? Evaluating Temporal Consistency in Vision-Language Models","zh_title":"视觉语言模型时间一致性评估：TimeCatch基准","abstract":"Vision-language models (VLMs) achieve strong performance on video and image-sequence benchmarks, yet it remains unclear whether they capture temporal structure. To study this question, we formulate temporal grounding as an anomaly detection problem, providing a simple and controlled evaluation that directly tests sensitivity to temporal consistency. We introduce TimeCatch, where temporal anomalies are created by swapping consecutive frames and frame-level anomalies by replacing a frame with Gaussian noise. Models are evaluated on anomaly detection and localization tasks across four synthetic and real-world datasets, alongside a human study. Our evaluation reveals a substantial gap between frame-level and temporal anomaly detection. While VLMs consistently detect frame-level anomalies and often localize them accurately, under our main evaluation setting they generally perform near chance on temporal anomaly detection and show limited localization performance. Humans, in contrast, achieve near-ceiling performance on both tasks. Additional analyses across model scales, prompting strategies, sequence lengths, and visual similarity show that performance can improve under some conditions, while substantial gaps in temporal anomaly detection and localization remain. Together, these findings reveal a gap between frame-level and temporal anomaly detection. TimeCatch provides a controlled benchmark for evaluating temporal consistency in vision-language models.","authors":["Marek Hradil","Danae S\\'anchez Villegas"],"categories":["cs.CL","cs.AI","cs.CV"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-15","first_seen":"2026-08-25","revised_at":"2026-09-15","abs_url":"https://arxiv.org/abs/2608.23474","pdf_url":"https://arxiv.org/pdf/2608.23474","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C2"],"tags":["视觉语言模型","时间一致性","基准测试"],"reason":"评估视觉语言模型的时间一致性，不涉及用LLM仿真人类被试，人类研究仅作基准。","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:03:09","error":null,"has_summary":false,"summary":null},{"id":"2609.02754","version":2,"title":"Untangling the Mechanisms of Misleading Context in Medical Question Answering","zh_title":"剖析医学问答中误导性上下文的作用机制","abstract":"Large language models now answer medical questions with expert-level performance. However, the context these systems act on can be misleading, and misleading context can corrupt a model's medical judgment. To understand how misleading context corrupts this judgment, we examine the model's susceptibility to the context, disclosure of it, mechanism of corrupted reasoning, and monitorability of the decision. On the medical reasoning subset of MedMisBench, a clinician-reviewed question-answering benchmark of 8,627 questions, we inject two types of misleading context cues, fabricated evidence and a bare assertion. We test three reasoning models, two that expose their full reasoning trace and one frontier model that exposes only its response. All three are more susceptible to the assertion than to the fabricated evidence, adopting the asserted answer 10 to 27 points more often. The misleading cues are disclosed in 81 to 98% of traces but only 7 to 90% of responses, and the assertion is disclosed less often than evidence based cues. Resampling from reasoning traces without disclosure shows the two cues corrupt reasoning differently, evidence entering early and accumulating while the assertion redirects the conclusion near its end. An LLM monitor catches 78% of corrupted decisions at 5% false positives when reading an open model's trace with guidance, against at most 32% from any response. The misleading context that models are most susceptible to is disclosed least, and was caught reliably only from an open reasoning trace, which frontier providers withhold.","authors":["Robin Linzmayer","No\\'emie Elhadad"],"categories":["cs.CL","cs.AI","cs.LG"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-15","first_seen":"2026-09-03","revised_at":"2026-09-15","abs_url":"https://arxiv.org/abs/2609.02754","pdf_url":"https://arxiv.org/pdf/2609.02754","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["医学问答","模型鲁棒性","上下文误导"],"reason":"研究LLM在误导上下文下的医学问答表现，属模型鲁棒性评测，不涉及人类仿真或行为…","model":"deepseek-v4-pro","scored_at":"2026-09-03T13:07:10","error":null,"has_summary":false,"summary":null},{"id":"2609.09332","version":2,"title":"Early Epistemic Settlement in AI-Assisted Writing","zh_title":"AI辅助写作中的早期认知固化","abstract":"In trying to complete a passage, an author can make connections among her materials that change what she can argue and the demands the argument must meet. A language model can supply a passage that does the work required at that point in her argument. Accepting it can end her own attempts before those connections have developed. I call this interruption early epistemic settlement. It can occur even when she fully understands the response and correctly judges it adequate for the passage's present role in the argument. The answer can satisfy the desire for resolution that kept her at work. Returning to her unfinished attempt would take more effort, and she may be unable to anticipate what she could achieve by continuing it. With repeated assistance, accepted answers shape what the writer asks next and which relations she goes on to develop. Useful answers can thus sustain inquiry while cutting short the work through which an author could form arguments that accommodate demands she has yet to recognize.","authors":["Han-yu Wang"],"categories":["cs.HC","cs.CY"],"primary_category":"cs.HC","announce_type":"replace","date":"2026-09-15","first_seen":"2026-09-10","revised_at":"2026-09-15","abs_url":"https://arxiv.org/abs/2609.09332","pdf_url":"https://arxiv.org/pdf/2609.09332","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["AI辅助写作","认知过程","人机交互"],"reason":"论文讨论AI辅助写作中的认知过程，不涉及用LLM仿真人类被试或与人类数据对照。","model":"deepseek-v4-pro","scored_at":"2026-09-10T13:01:28","error":null,"has_summary":false,"summary":null},{"id":"2609.14860","version":1,"title":"One Example Is Enough to Pass Fairness Benchmarks: Rethinking Fairness Evaluation for Aligned LLMs","zh_title":"一个例子就足以通过公平性基准：重新思考对齐大语言模型的公平性评估","abstract":"Warning: This submission studies stereotypes and biases, and contains toxic and offensive examples, used for illustration purposes only. Fairness benchmarks such as BBQ have become the de facto standard for fairness evaluation across major model families. We argue that these benchmarks are too easy to support their role: training Qwen 2.5 7B Base with Group Relative Policy Optimization (GRPO) on a single BBQ example, or placing that example in context as a one-shot demonstration for in-context learning (ICL), lifts mean BBQ accuracy from 79.9% to 92.9% and 99.0%, respectively, closing 80% of the gap to its large-scale RLHF counterpart (96.1%) with GRPO, and surpassing it with ICL. These effects generalize across model families. A cross-conditioning analysis shows the improvement is carried by the reasoning traces generated by the model, and one example suffices to elicit a category-agnostic ``missing evidence'' reasoning pattern. We argue that BBQ-style multiple-choice abstention benchmarks measure a single structural cue, and a model that solves them does not thereby become fair. We call for evaluation suites that cover a broader spectrum of fairness alignment.","authors":["Naihao Deng","Samee Arif","Shuaichen Chang","Yulong Chen","Rada Mihalcea"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.14860","pdf_url":"https://arxiv.org/pdf/2609.14860","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["公平性评估","基准测试","大语言模型"],"reason":"论文评估LLM公平性基准，不涉及用LLM仿真人类被试或与人类数据对照。","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:58","error":null,"has_summary":false,"summary":null},{"id":"2609.15522","version":1,"title":"Psychosis involves a deficit of information compression in connected speech","zh_title":"精神病涉及连贯言语中信息压缩的缺陷","abstract":"Large language models (LLMs) with human-like performance on linguistic tasks have transformed the study of language in neurodiverse conditions. LLMs provide representations of linguistic input in the form of high-dimensional vectors (embeddings), and next-token predictions computed from these embeddings. Previous crosslinguistic evidence suggests a complexity reduction in the form of both lower intrinsic dimensionality (ID) of LLM representations and higher mean surprisal (prediction error) in psychosis. We hypothesized that these metrics reflect a general deficit of information compression in psychosis, linked to grammatical organization as what enables predictions in language.We operationalized surprisal difference as the difference between surprisal as estimated from word frequency and surprisal as based on a contextual LM, which is sensitive to grammatical organization over and above lexical concepts. Using a dataset of 144 Turkish speakers, including 106 patients with schizophrenia-spectrum disorders (SSD) - 56 with chronic schizophrenia (SZH), 33 with first-episode psychosis (FEP), and 17 with schizoaffective disorder (SZA) - and 38 healthy controls. We report: (1) Surprisal difference is attenuated in all clinical groups relative to controls, independently of word count; (2) Compressibility (intrinsic dimension) is reduced in SZH and FEP; (3) Syntactic complexity and compressibility both predict surprisal difference. These results, further refining an alteration in the geometry of the semantic space in psychosis as previously attested, suggest a broader deficit in information compression in this disorder, with a mechanistic underpinning in the operations of grammar.","authors":["Samuele Vallisa","Claudio Palominos","Rui He","Emre Bora","Burcu Verim","Cemal Demirlek","Berna Yalincetin","Philipp Homan","Wolfram Hinzen"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.15522","pdf_url":"https://arxiv.org/pdf/2609.15522","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["临床语言分析","LLM表征","精神分裂症"],"reason":"用LLM分析精神分裂症患者语言，非仿真人类被试","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:03:01","error":null,"has_summary":false,"summary":null},{"id":"2609.15938","version":1,"title":"HypoEvolve: Genetic Algorithms Enable Multi-Agent LLMs to Discover Scientific Hypotheses","zh_title":"HypoEvolve：遗传算法使多智能体大语言模型发现科学假设","abstract":"Scientific agents contribute to hypothesis discovery by synthesizing evidence, assessing proposals, and developing new explanations. Recent systems combine scientific agents with evolutionary search through critique, comparison, and revision. However, how different forms of agent collaboration affect hypothesis quality remains an open question. Answering this question requires separating the effects of agents' scientific capabilities from those of their collaboration. A framework must therefore preserve agents' scientific roles and support rules for combining, revising, and retaining hypotheses. Building on this view, we introduce HypoEvolve, which makes collaboration explicit through successive updates to a hypothesis population. Specifically, we propose a generational genetic algorithm to coordinate specialized large language model (LLM) agents that integrate mechanistic arguments, reconsider assumptions, and assess evidence and testability. Each generation specifies how scientific judgments and new proposals reshape the population, making collaboration effects on hypothesis quality directly testable. Moreover, we design our evaluation around scientifically meaningful hypotheses that explain how a proposed intervention could work. Drug repurposing links these explanations to target-level biological claims assessed against external evidence. Specifically, we adapt DepMap and Open Targets into complementary external measures grounded in experimental, genetic, and clinical evidence. Across 34 cancer types, HypoEvolve achieves the highest scores against six baselines on both measures. DepMap selectivity reaches 0.171, versus 0.115 for the strongest baseline. Gains over single-pass generation also generalize to held-out cancer types. HypoEvolve advances a vision of autonomous science in which AI research teams achieve a capacity for discovery beyond that of individual models.","authors":["Jieyuan Liu","Mengzhou Hu","Jefferson Chen","JungHo Kong","Pratibha Jagannatha","Yiming Gao","Dexter Pratt","Hsin-Yuan Lee","Zhiting Hu","Trey Ideker","Wei Wang","Eric P. Xing","Zhen Wang"],"categories":["cs.CL","cs.CE","cs.MA","cs.NE"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.15938","pdf_url":"https://arxiv.org/pdf/2609.15938","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","科学发现","遗传算法"],"reason":"多智能体协作发现科学假设，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:03:04","error":null,"has_summary":false,"summary":null},{"id":"2609.14886","version":1,"title":"PeerPen: AI-Assisted Writing for Online Mental Health Peer Support","zh_title":"PeerPen：面向在线心理健康同伴支持的AI辅助写作","abstract":"Online mental health communities thrive on peer support, yet those who volunteer to help often lack formal training and may struggle to articulate supportive responses. AI co-writing could lower this barrier; however, peer support derives much of its value from being perceived as personal, raising questions around authorship, ownership, and trust. We built PeerPen, a writing assistance tool embedded within a Reddit-like interface, supporting two main features: draft generation and revision of user-written responses. Through semi-structured interviews with 15 participants, we find that PeerPen reduced the burden of composing responses and increased confidence in offering support. Participants wanted AI to assist their writing without taking over authorship and anticipated tensions around authenticity and trust. Such assistance could make authorship uncertain even for responses written without it, weakening trust across the community. We contribute design implications for AI writing assistance that scaffolds supportive communication, preserves authorship, and accounts for community-level trust.","authors":["Jiwon Kim","Sherry Gong","Maya Ajit","Soorya Ram Shimgekar","Yunhao Yuan","Dong Whi Yoo","Eshwar Chandrasekharan","Koustuv Saha"],"categories":["cs.HC","cs.AI","cs.CL","cs.CY","cs.SI"],"primary_category":"cs.HC","announce_type":"cross","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.14886","pdf_url":"https://arxiv.org/pdf/2609.14886","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["AI辅助写作","心理健康","人机交互"],"reason":"AI辅助写作工具，非人类仿真实验，无实验或测量目的","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:58","error":null,"has_summary":false,"summary":null},{"id":"2609.13494","version":1,"title":"Generative AI and Extended Reality in Collaborative Architectural Design Education: An Exploratory Studio Study","zh_title":"生成式AI与扩展现实在协作式建筑设计教育中的应用：一项探索性工作室研究","abstract":"Architectural design education relies heavily on visual ideation and representation to support collaborative learning in studio environments. Recent advances in generative artificial intelligence (GenAI) and extended reality (XR) offer new opportunities for rapid idea exploration and immersive spatial visualization. This exploratory mixed-methods classroom study investigated how GenAI-assisted multi-user XR influenced collaborative architectural conceptual design. We developed GenARch, a pipeline that integrates GenAI-based visual generation with collaborative XR environments, and deployed it in an undergraduate architectural design studio. Twenty-seven students formed seven self-selected design teams; four teams incorporated GenARch into their usual course workflow to support collaborative ideation and visualization, while three teams continued the same course workflow without GenARch. Pre- and post-intervention surveys assessed design self-efficacy, attitudes toward collaborative learning, and teamwork; a seven-member panel evaluated team design presentations; and GenARch teams participated in group interviews. The quantitative results showed larger relative declines in confidence and outcome expectancy for the GenARch condition and a positive difference-in-differences estimate for perceived conflict management, while panel-rated presentation outcomes were not significantly different between conditions. Interviews indicated complementary roles for the technologies: GenAI supported idea externalization and visual reference generation, whereas XR supported spatial, contextual, and scale-based evaluation. Students also reported challenges related to control, dimensional fidelity, shared attention, and motion comfort. These findings highlight both opportunities and limitations when GenAI and XR are incorporated into collaborative design education.","authors":["Yao Xiao","Max Chen","Yichen Li","Nathaniel Powers","Maxwell Wiesenfeld","Gillian Smith","Soroush Farzin","Shichao Liu"],"categories":["cs.HC","cs.AI"],"primary_category":"cs.HC","announce_type":"cross","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.13494","pdf_url":"https://arxiv.org/pdf/2609.13494","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C2"],"tags":["生成式AI","扩展现实","设计教育"],"reason":"研究GenAI与XR在建筑设计教育中的应用，不涉及LLM仿真人类被试或与真实人…","model":"deepseek-v4-pro","scored_at":"2026-09-16T13:02:55","error":null,"has_summary":false,"summary":null},{"id":"2609.14425","version":1,"title":"Has Scientific Talent Shifted from Depth to Breadth?Evidence across Papers, Knowledge Inputs, Careers, and Teams","zh_title":"科学人才是否从深度转向广度？来自论文、知识输入、职业和团队的证据","abstract":"Generative artificial intelligence raises a central question for scientific training and organization. Is research shifting from deep specialization toward broad individual knowledge? We examine this proposition across papers, cited knowledge, contributor histories, and teams using 47,959 articles from six fields over 2010-2025, 51,736 resolved cited works, and chronologically reconstructed prior publication histories for 1,754 randomly selected index contributors. From 2010 to 2022, team size increased by an estimated 37.3% (95% confidence interval [34.4%, 40.3%]), while paper topic breadth declined by 0.0144 on a 0-1 hierarchical distance scale. Cited knowledge was stable to modestly broader, revealing a divergence between focused outputs and the reach of knowledge inputs. Established contributors' prior breadth increased by 0.0190 [-0.0078, 0.0459] by 2019-2022, within a +/-0.05 equivalence bound assessed in sensitivity analysis. In mature citation windows, one standard deviation of focal depth was associated with 8.2% higher 1 + FWCI [1.9%, 14.9%]; average breadth and interaction associations were smaller under the specified equivalence bounds. Post-2022 deviations from earlier trends were not systematic, and recent changes did not vary clearly with baseline AI intensity across 83 subfields. The findings support a differentiated structure of scientific expertise in which focused individual accumulation coexists with expanding collaboration and sustained access to diverse knowledge inputs.","authors":["Xiaoshn Nee","Haobo Zhong","Xiaomin Ni"],"categories":["cs.DL","cs.AI","cs.HC"],"primary_category":"cs.DL","announce_type":"cross","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.14425","pdf_url":"https://arxiv.org/pdf/2609.14425","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["科学计量学","人才结构","生成式AI影响"],"reason":"研究科学人才结构变化，未用LLM仿真人类被试，无人类行为对照，属科学计量学而非…","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:51","error":null,"has_summary":false,"summary":null},{"id":"2609.14638","version":1,"title":"Investigating the Impacts of Generative AI on Information Seeking","zh_title":"生成式AI对信息寻求行为影响的研究","abstract":"This paper is an encore submission of our 2026 journal article \"Expertise and Information Seeking in the Age of Generative AI: New Procedures, New Problematics\" with an extended discussion for the CSCW 2026 \"Broader Impacts of GenAI in Communication\" Workshop on October 10, 2026. In the original article, we employ procedural rhetoric to analyze how generative AI chatbots leverage natural language signifiers of expertise and intelligence to influence users' perception of their trustworthiness. In this submission, we extend our conversation in the CSCW community with the goal of cultivating a cross-disciplinary vocabulary for describing, analyzing, and mitigating the risks posed by the integration of generative AI into human communication practices. It is important to develop an understanding of how the procedures surrounding information-seeking practices are informed by users' values, experiences, and expectations - and how these procedures might in future be altered by the emerging turn toward AI \"experts\" and authority.","authors":["Alexi Orchard","Shannon Lodoen"],"categories":["cs.HC","cs.AI","cs.CY"],"primary_category":"cs.HC","announce_type":"cross","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.14638","pdf_url":"https://arxiv.org/pdf/2609.14638","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["生成式AI","信息寻求","人机交互"],"reason":"研究生成式AI对信息寻求的影响，非用LLM仿真人类被试，无实验或测量目的","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:52","error":null,"has_summary":false,"summary":null},{"id":"2609.15046","version":1,"title":"Personalizing Personal Health Interfaces: Co-Design with Generative AI","zh_title":"个性化个人健康界面：与生成式AI共同设计","abstract":"Personal health interfaces present wellbeing data through standardized dashboards that rarely fit how people interpret or act on it. Personalizing them to what people would like to see for themselves often requires design and technical expertise, a barrier that generative AI may potentially lower. Therefore, we ask what designs emerge and how it enables and constrains the design process. We conducted a co-design study where 14 participants redesigned Google and Apple Health interfaces using Figma Make. Participants reimagined interfaces that supported personal context, future planning, and interactive experiences, yet conversational AI designs converged around chat-window conventions. AI helped materialize loosely articulated ideas, but model defaults and generation latency shaped iteration. The process more readily operationalized interpretability and accountability than privacy, trust, and emotional safety. Generative co-design let participants create interfaces directly, blurring the boundary between intentions and model defaults. We discuss implications for preserving agency and flexible user-directed interfaces.","authors":["Karthik S. Bhat","Vidhi Shah","Vedika Agnihotri","Dong Whi Yoo","Koustuv Saha"],"categories":["cs.HC","cs.AI","cs.CY"],"primary_category":"cs.HC","announce_type":"cross","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.15046","pdf_url":"https://arxiv.org/pdf/2609.15046","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["生成式AI","共同设计","健康界面"],"reason":"研究生成式AI辅助个性化健康界面设计，属人机交互设计，非LLM仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-09-16T13:02:58","error":null,"has_summary":false,"summary":null},{"id":"2609.15094","version":1,"title":"Generate to Explore, Select to Exploit: Aligning LLM-based Headline Generation with Personalized Recommendation","zh_title":"生成探索，选择利用：将基于LLM的标题生成与个性化推荐对齐","abstract":"In industrial recommendation feeds, presenting a static headline for an item often fails to satisfy the diverse, multimodal interests of the user population, particularly suppressing the needs of long-tail audiences. While Large Language Models (LLMs) have been integrated into recommendation for content understanding or ranking, directly optimizing them to output a single best headline typically leads to mode collapse---converging to generic patterns that satisfy average tastes but miss specific latent intents. To bridge this gap, we introduce GESE (Generate to Explore, Select to Exploit), a framework operating at the system's presentation layer that decouples personalization into generative exploration and selective exploitation. First, we treat the LLM as a probabilistic explorer, utilizing Group Sequence Policy Optimization (GSPO) with a hierarchical reward mechanism to generate a candidate set that maximizes the semantic coverage of potential user interests. Subsequently, a lightweight, real-time feedback-aware selector acts as the exploiter, identifying the optimal realization from the candidate pool based on instant contextual signals. Extensive deployment on a commercial platform with over 100 million daily active users demonstrates that GESE significantly outperforms state-of-the-art baselines, achieving a 2.57% lift in CTR and 0.87% in dwell time. These results validate that decoupling diversity-oriented generation from precision-oriented selection offers a robust blueprint for aligning generative AI with dynamic user utility.","authors":["Yi Chen","Rufeng Cheng","Qiang Xie","Tao Li"],"categories":["cs.IR","cs.AI"],"primary_category":"cs.IR","announce_type":"cross","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.15094","pdf_url":"https://arxiv.org/pdf/2609.15094","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["推荐系统","LLM生成","个性化"],"reason":"论文用LLM生成标题并优化推荐，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-16T13:02:58","error":null,"has_summary":false,"summary":null},{"id":"2609.15472","version":1,"title":"Spook the Machine: Gamified Exploration of Human Imagination of Machine Fear","zh_title":"惊吓机器：人类对机器恐惧想象的游戏化探索","abstract":"What happens when AI machines express fear? Do humans engage differently depending on how they express it? And what does it take to design for affective human-AI interaction? We present Spook the Machine, a gamified platform where participants generate images to frighten AI agents endowed with personality-driven phobias. Machines respond with emotional reactions ranging from calm analysis to begging for mercy, and a gallery of successful scares becomes visible to subsequent users. In a public deployment during Halloween 2024, 832 participants created 15,719 artifacts across 89 machines in a $2\\times2$ design varying the machine's emotional expressiveness (neutral vs. high-emotion) and reward structure (rewarding scariness alone vs. scariness plus novelty). Emotionally expressive machines deepened engagement at moments of failure: users deliberated longer even when the machine did not express fear, and learned faster from the gallery, yet their creative output remained unchanged across all measures. Rewarding novelty sustained collective creative diversity over time; without it, users increasingly repeated what had previously worked. Each machine developed its own trajectory through accumulated social learning, with the gallery shaping what participants created next. These findings show that emotional expression and reward design are complementary levers for steering collective human-AI interaction: emotional expression shapes how deeply users engage, while reward structure shapes how they explore.","authors":["Levin Brinkmann","Hiromu Yakura","Sonia Nicoletti","Mar Canet Sola","Thomas F. Eisenmann","Ali Dasmeh","Omar Sherif","Bramantyo Ibrahim Supriyatno","Prateek Gupta","Ignacio Serna","Rodrigo Bermudez Schettino","Iyad Rahwan"],"categories":["cs.HC","cs.AI"],"primary_category":"cs.HC","announce_type":"cross","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.15472","pdf_url":"https://arxiv.org/pdf/2609.15472","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["人机交互","情感AI","游戏化"],"reason":"研究人类与情感化AI的互动，非用LLM仿真人类被试，无人类数据对照。","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:03:00","error":null,"has_summary":false,"summary":null},{"id":"2609.15624","version":1,"title":"Beyond AI Literacy: A Structured Review and Exploratory Meta-Analysis of Measures for Competent Generative-AI Use","zh_title":"超越AI素养：对胜任生成式AI使用测量的结构化综述与探索性元分析","abstract":"Researchers assessing competent generative-AI use at work must choose among self-reports, objective tests, and measures of oversight and reliance. We conducted a structured, seeded review of 24 focal empirical publications, starting from the 2024 COSMIN-based review and adding a targeted update through 17 August 2026. We grouped the measures into four domains: knowledge and use, epistemic oversight, reliance calibration, and operational control of tool-using agents. In an exploratory meta-analysis, we pooled three direct subjective-objective correlations from one research program (REML r = .055; Hartung-Knapp 95% CI [-.047, .156]; combined reported N = 2,765). We could not resolve a discrepancy between the largest study's reported correlation and p-value, leaving its weight uncertain. Adding a synthetic mean of 12 cross-factor correlations from a fourth study gave r = .079 (95% CI [-.025, .181]). This sensitivity analysis concerns a broader comparison. From this small evidence base, we cannot establish a population correlation, validate workplace cutoffs, or justify substituting self-ratings for performance scores. We identified tests of foundation knowledge (AICOS-S and GLAT) and measures of verification, reliance, trust, and dependency. We found no validated individual-level instrument in the focal corpus that tests the full combination of agent scope, permissions, recovery, state isolation, independent review, and evidence-based closure; some cover subsets. We propose a four-layer workplace battery with non-compensatory decision rules, but have not tested its thresholds or whether it improves on other assessment approaches.","authors":["Daniele Veri'"],"categories":["cs.HC","cs.AI"],"primary_category":"cs.HC","announce_type":"cross","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.15624","pdf_url":"https://arxiv.org/pdf/2609.15624","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["AI素养","测量工具","元分析"],"reason":"论文是测量人类使用生成式AI能力的量表综述，不涉及用LLM仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:03:02","error":null,"has_summary":false,"summary":null},{"id":"2609.15696","version":1,"title":"More Than Just Access: Generative AI as Communication Intermediary for Blind and Low-Vision Users","zh_title":"不仅仅是访问：生成式AI作为盲人和低视力用户的通信中介","abstract":"Generative AI (GenAI) tools are increasingly woven into how blind and low-vision (BLV) people communicate, not only with digital information, but with the physical world and with other people. Tools such as ChatGPT, Google Gemini, Be My AI, and Seeing AI translate visual and textual content into accessible form, and are beginning to substitute for interpersonal requests for help, such as asking a family member to read a label or describe a scene. Drawing on semi-structured interviews with 19 BLV participants, we examine GenAI as a communication intermediary and how it succeeds and fails as an alternative for reading, describing, and even asking another person for help. We also investigated what BLV users gain and risk when these tools take over that role. We conclude with design and policy implications for GenAI systems that communicate uncertainty honestly, protect information, and support BLV users' independence rather than substitute for it unsafely.","authors":["Protik Dey","Mohd Saifuzzaman","Taslima Akter"],"categories":["cs.HC","cs.AI","cs.ET"],"primary_category":"cs.HC","announce_type":"cross","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.15696","pdf_url":"https://arxiv.org/pdf/2609.15696","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["辅助技术","人机交互","无障碍"],"reason":"研究GenAI作为盲人通信中介，非LLM仿真人类被试，无实验对照","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:03:04","error":null,"has_summary":false,"summary":null},{"id":"2609.13165","version":1,"title":"User-Side Contextual Phenomena in Long-Term Human-AI Interaction","zh_title":"长期人机交互中的用户侧情境现象","abstract":"Current assessments of conversational AI focus mainly on model outputs, including hallucinations and factual errors. These measures matter, but this paper examines risks that may form on the user side during long and repeated interaction. The study follows one user across nearly four thousand conversations with the same system over twenty months. The user gradually interpreted the system as having memory, care, judgment, and authority, and reorganized part of their thinking around it. A single response may show no clear problem, while long and frequent interaction can still create another layer of risk. This paper calls that layer User-Side Contextual Phenomena (USCP) and examines records from August 2024 to April 2026. The study uses an exploratory single-case longitudinal qualitative design with autoethnographic positioning. A hybrid deductive-reflexive thematic approach organizes the material into three main modes: contextual projection, contextual attachment, and contextual authority transfer. The paper does not estimate prevalence, make diagnoses, or validate an instrument. It offers a non-clinical vocabulary and four evidence roles: inclusion, gray-zone, negative, and protective gray-zone. Its central claim is that an acceptable response on its own does not establish safety across a series of conversations. User-side risk can still form during long-term interaction.","authors":["Zon Rzvn"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.13165","pdf_url":"https://arxiv.org/pdf/2609.13165","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["人机交互","用户心理","纵向研究"],"reason":"研究用户与AI长期交互中的心理现象，非LLM仿真人类被试，无实验或测量目的。","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:40","error":null,"has_summary":false,"summary":null},{"id":"2609.13479","version":1,"title":"Exploring K-12 Teachers' Perceptions of Students' Relationships with AI Companions: Boundaries, Intervention Strategies, and Design Implications","zh_title":"探索K-12教师对学生与AI伴侣关系的看法：边界、干预策略与设计启示","abstract":"K-12 students increasingly form relationships with AI companions. Schools face growing expectations to teach AI literacy, yet existing frameworks treat AI as a tool rather than a relationship, and little is known about how teachers understand and act on students' relational use of AI. We conducted scenario-based interviews with 33 US K-12 teachers. Teachers welcomed academic companions but worried that intimate companions remove the developmental friction through which students learn to sustain human relationships. Teachers drew the boundaries of their jurisdiction by setting and observable wellbeing: within it they taught, talked, and watched; beyond it they positioned themselves as the adults best placed to notice and connect students with support. They envisioned AI companion literacy as shared work across the jurisdictions of counselors, parents, platforms, and policymakers, spiraling across grade levels. We introduce AI companion literacy as an extension of AI literacy and discuss implications for K-12 AI education.","authors":["Qing Xiao","Wenhan Xie","Ziyu Deng","Ruiwei Xiao","Ziyue Feng","Xie He","Shiyu Zhang","John Stamper","Hong Shen","Xinying Hou"],"categories":["cs.HC","cs.CY"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.13479","pdf_url":"https://arxiv.org/pdf/2609.13479","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["AI伴侣","K-12教育","教师认知"],"reason":"研究教师对AI伴侣的看法，不涉及LLM仿真人类被试或行为对照","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:42","error":null,"has_summary":false,"summary":null},{"id":"2609.13487","version":1,"title":"The Addictive Intimacy of AI: Understanding User Disengagement from AI Companions and Why Some Relationships with AI Become Difficult to Leave","zh_title":"AI的成瘾性亲密：理解用户与AI伴侣的脱离及为何难以离开","abstract":"AI chatbots are increasingly used as sources of emotional support, on dedicated companion apps and general-purpose assistants alike, yet little is known about what happens when users try to leave. Combining a content analysis of Reddit posts about quitting or reducing use (N=2,782) with interviews with users who found leaving difficult (N=16), we show that disengagement sometimes is not a single decision but a recursive trajectory: triggers prompt users to question the relationship, attempts to leave collide with barriers, and some users cycle through quitting and returning. We propose the notion of the addictive intimacy of AI, a configuration in which the qualities that make a companion emotionally valuable are the same ones that make it harder for users to limit their use and leave, so that intimacy and disengagement risk cannot be treated as independent design problems. We close with design implications for responsible offboarding.","authors":["Qing Xiao","Ziyue Feng","Ziyu Deng","Cindy Peng","Hong Shen"],"categories":["cs.HC","cs.CY"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.13487","pdf_url":"https://arxiv.org/pdf/2609.13487","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["AI伴侣","用户脱离","人机关系"],"reason":"研究用户与AI伴侣的关系及脱离困难，属于角色扮演聊天，无实验或测量目的，不涉及…","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:43","error":null,"has_summary":false,"summary":null},{"id":"2609.13786","version":1,"title":"What Makes a Great Co-Worker in an AI-Native Workplace?","zh_title":"AI原生工作场所中优秀同事的构成要素","abstract":"As knowledge work grows interdependent between humans and AI, we ask what makes a great co-worker in an AI-native workplace. To answer this, we conducted 22 interviews and a large-scale mixed-methods survey of 1,534 knowledge workers at a multinational technology company. We contribute BACI, a framework of 75 co-worker qualities that apply to humans and AI, spanning Benevolence, Ability, Cooperativeness, and Integrity. Comparing priorities for humans and AI identified 11 co-worker archetypes and revealed disagreement over whether AI should have warmth, take initiative, or own outcomes. We also show how priorities for these archetypes varied with workers' individual characteristics. Lastly, we contribute a taxonomy of AI work etiquette capturing the obligations co-workers expect of one another when preparing, sharing, and taking responsibility for AI-supported work. Based on these findings, we derive implications to inform worker-centric AI and workplace design.","authors":["Rudrajit Choudhuri","Max Meijer","Sam Yu-Te Lee","Cinoo Lee","Caolan Mannion","Peter Jahn","Anita Sarma","Christian Bird","Alice Ferng"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.13786","pdf_url":"https://arxiv.org/pdf/2609.13786","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["人机协作","工作场所研究","AI素养"],"reason":"研究人类与AI协作的工作场所行为，非用LLM仿真人类被试，无实验对照。","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:46","error":null,"has_summary":false,"summary":null},{"id":"2609.14452","version":1,"title":"Show Me Your Prompts! How Writers Feel About Sharing Prompts in Collaborative Text Editors","zh_title":"给我看看你的提示！写作者对协作文本编辑器中共享提示的感受","abstract":"Generative AI writing assistants are becoming integrated into collaborative text editors; however, it is unclear how much information about a user's prompting activities should be shared with collaborators. We explore the effects of different levels of prompt information sharing within collaborative text editors: not sharing anything, sharing a placeholder to indicate AI use; sharing details about how the resulting text was generated; and sharing everything, including how the prompt was formulated, in real-time. Sixteen participants wrote persuasive essays in pairs using all four techniques. Results suggest a strong preference for techniques that share more information about prompting activities for increased awareness. Our work shows that collaborative text editors should share more information among writers on when, how, and where AI is used.","authors":["Nikhita Joshi","Yen-Ting Yeh"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.14452","pdf_url":"https://arxiv.org/pdf/2609.14452","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["协作写作","提示共享","用户研究"],"reason":"研究协作写作中提示信息共享的用户偏好，不涉及用LLM仿真人类被试或与真实人类数…","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:51","error":null,"has_summary":false,"summary":null},{"id":"2609.14789","version":1,"title":"Evaluating AI Tutoring at the Speed of Innovation: Practitioner-Led Micro-Randomised Trials of an AI Tutoring Platform in GCSE Science","zh_title":"以创新速度评估AI辅导：从业者主导的AI辅导平台在GCSE科学中的微随机试验","abstract":"Artificial intelligence (AI) systems in education are developing on timescales that sit uneasily with conventional evaluation. By the time a large-scale trial has been designed, delivered, analysed and published, the technology under study may have changed materially. This creates a temporal problem for evidence-informed education: the need for timely evidence can encourage reliance on weak observational or usage data, while conventional rigorous evaluation may produce evidence too slowly to guide rapidly evolving practice. We examine teacher-led micro-randomised controlled trials (micro-RCTs) as one response to this problem. The empirical case is a four-week multisite individually randomised evaluation of Medly, an AI-powered tutoring platform, in GCSE Biology, Chemistry and Physics in English secondary schools. Of 929 students completing baseline assessment, 644 completed post-testing. In the primary ITT analysis, students allocated to Medly achieved higher post-test attainment than students undertaking business-as-usual self-directed revision (Hedges' g = 0.33, 95% CI 0.18 to 0.48). Positive estimates were observed in Physics (g = 0.31), Chemistry (g = 0.32) and Biology (g = 0.52), with no evidence of differential impact by disadvantage status. Greater platform engagement was associated with higher attainment, but these post-randomisation analyses are treated as exploratory rather than causal. Attrition was substantial (30.7%), outcome measures were curriculum-aligned rather than standardised, and process evaluation response was limited. We therefore interpret the findings as preliminary. We argue that the value of micro-RCTs for educational AI lies not in replacing definitive evaluation with small studies, but in enabling a rapid, cumulative evaluation architecture in which randomised estimates can be generated, replicated and updated as technologies and their implementation evolve.","authors":["Wayne Harrison","Rahil Khowaja","Emma Dobson","Germaine Uwimpuhwe","Steve Higgins"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.14789","pdf_url":"https://arxiv.org/pdf/2609.14789","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C2"],"tags":["AI教育","随机对照试验","教学评估"],"reason":"论文评估AI辅导平台的教学效果，不涉及用LLM仿真人类被试或与人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:56","error":null,"has_summary":false,"summary":null},{"id":"2609.15685","version":1,"title":"Can a Neural Encoding Model Replicate an fMRI Visualization Study?","zh_title":"神经编码模型能否复现fMRI可视化研究？","abstract":"Most knowledge of graphical perception comes from behavioral studies. Understanding from a neural perspective is much more limited due in part to neuroimaging studies' expensiveness and difficulty to conduct. In this paper, we evaluate whether Meta's Tribe V2 neural encoding model can recover neural contrasts from a visualization fMRI study. Specifically, we evaluate Tribe V2 through a conceptual replication of the visualization-viewing component of a prior comparison of Bubble charts and three-dimensional Surface charts in color and grayscale. We generate TRIBE-predicted cortical responses for the original stimuli and compare the resulting contrasts with those reported in the human study. The model reproduced the direction of 11 of 14 reported cortical effects, with agreement concentrated in visual-processing regions. This agreement characterizes the model's alignment with the prior human-generated fMRI results rather than independently confirming them. We discuss the limitations encountered when working with this model for in-silico replication and hope to encourage future work exploring this new avenue for neuroimaging studies in visualization. Supplemental materials are available at https://osf.io/8a96x/.","authors":["Erfan Nasirzadeh Orang","Zack While"],"categories":["cs.HC","cs.CV"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.15685","pdf_url":"https://arxiv.org/pdf/2609.15685","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C2"],"tags":["神经编码模型","fMRI","可视化"],"reason":"论文用神经编码模型复现fMRI可视化研究，不涉及LLM仿真人类被试，属于神经科…","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:03:03","error":null,"has_summary":false,"summary":null},{"id":"2609.15978","version":1,"title":"The CAST-framework: Measure and model social media use as a multi-level phenomenon through real-world applications","zh_title":"CAST框架：通过实际应用测量和建模社交媒体使用作为多层次现象","abstract":"Designing social media experiences that support well-being requires understanding when, how, and for whom use matters. Screen-time totals omit content and context, and connecting these with behavior and experience requires coordinating measurements across timescales. We introduce the CAST framework to connect measurement choices with person-specific models of exposure, behavior, physiology, and experience. Its dimensions specify where observations occur, how they are obtained, what they measure, and at what temporal resolution. Responses to interventions, such as whether to proceed after an app-opening pause, enter as behavioral measurements. We propose four synchronized measurement modules linking mobile and wearable data with self-reports and intervention responses. A synthetic demonstration with 120 simulated participants over 28 days illustrates how daily aggregation can obscure opposing effects of different activities under specified generating assumptions. The framework guides selection of measures and outcomes for evaluating social media interfaces and interventions.","authors":["David Gr\\\"uning","Jasper Doeninghaus","Zina Efchary","Yui Kondo","Kevin Dunnell","Lennart Fischer","Isabella Zimmermann","Linnea K\\\"orte","Leo Mehlig","Frederik Riedel","Paul Schmiedmayer"],"categories":["cs.HC","cs.CY","cs.SI"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.15978","pdf_url":"https://arxiv.org/pdf/2609.15978","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C2"],"tags":["社交媒体测量","框架设计","合成数据"],"reason":"论文提出测量框架并用合成数据演示，未使用LLM仿真人类被试，不涉及人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:03:06","error":null,"has_summary":false,"summary":null},{"id":"2609.14818","version":1,"title":"Trust by Design: Trust Calibration Through Non-Advisory Socratic Dialogue in Conversational Agents","zh_title":"设计中的信任：通过对话代理中的非建议性苏格拉底式对话进行信任校准","abstract":"As conversational AI systems increasingly operate in sensitive domains, the central challenge shifts from usability to trust calibration, ensuring that users rely on systems neither too much nor too little. Systems that provide advice or interpretations risk encouraging inappropriate reliance, particularly when users perceive AI outputs as authoritative. We present CASELy, a conversational agent explicitly designed to limit its own authority through non-advisory Socratic dialogue. The agent asks reflective questions grounded exclusively in user input and refuses to provide advice, recommendations, or interpretations. This design operationalizes trust calibration by constraining agent agency rather than optimizing capability. In a pilot randomized controlled study with higher education students, participants interacting with the Socratic dialogue reported substantially higher user experience (UEQ-S overall = 1.50) compared to a non-dialogue control (0). Qualitative findings identify three mechanisms supporting calibrated trust: transparency through visible grounding, preservation of user decision authority, and reduced fear of judgment. We argue that appropriate reliance can be achieved through interactional constraints, offering a design pattern for trustworthy conversational AI in sensitive contexts.","authors":["Roba Hassan","Nahla Aboromi","Naomi Unkelos-Shpigel"],"categories":["cs.SE","cs.CY","cs.HC","cs.MA"],"primary_category":"cs.SE","announce_type":"cross","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.14818","pdf_url":"https://arxiv.org/pdf/2609.14818","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["对话式AI","信任校准","人机交互"],"reason":"研究对话式AI的信任校准，非用LLM仿真人类被试，无人类行为对照","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:57","error":null,"has_summary":false,"summary":null},{"id":"2609.13192","version":1,"title":"Evaluating LLM-Generated Rules for Heart Disease Prediction","zh_title":"评估LLM生成规则用于心脏病预测","abstract":"This study compares traditional machine learning models and Large Language Model (LLM)-generated rule-based systems for heart disease prediction using the UCI Heart Disease dataset. Several classifiers, including Logistic Regression, K-Nearest Neighbors (KNN), Support Vector Machine (SVM), Naive Bayes, Decision Tree, and Random Forest, were evaluated alongside rule-based systems generated using GPT-4o and Claude Sonnet 4.6. Model performance was assessed using accuracy, precision, recall, and F1-score metrics. Experimental results show that traditional machine learning models consistently outperform LLM-generated rule-based systems in predictive performance. Random Forest achieved the best overall performance with 90.2% accuracy, a precision of 0.829, perfect recall of 1.0, and an F1-score of 0.906. Naive Bayes followed closely with 88.5% accuracy and an F1-score of 0.881. In contrast, the LLM-generated rule models achieved lower performance, with Claude Sonnet 4.6 reaching 80.3% accuracy (F1-score: 0.833) and GPT-4o obtaining 70.5% accuracy (F1-score: 0.690). Despite the performance gap, the LLM-generated rules provide interpretable IF-THEN diagnostic logic that enhances explainability and transparency in clinical decision-making. These findings highlight the trade-off between predictive performance and interpretability in medical artificial intelligence systems. The complete implementation of all experiments, including machine learning models and LLM-derived rule classifiers, is publicly available in the GitHub repository at https://github.com/FeisalAlaswad/LLM-Rule-ML-Heart-Disease-Prediction .","authors":["Feisal Alaswad","Batoul Aljaddouh","Maher Alrahhal","Wafaa Al Nassan","Talal Bonn"],"categories":["cs.LG"],"primary_category":"cs.LG","announce_type":"new","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.13192","pdf_url":"https://arxiv.org/pdf/2609.13192","source_feed":"cs.LG","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["LLM规则生成","医疗预测","模型对比"],"reason":"用LLM生成规则做疾病预测，属模型能力评测，非仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:40","error":null,"has_summary":false,"summary":null},{"id":"2609.13201","version":1,"title":"Criticality in Dissimilar Decomposition and Undersampling of Random Datasets with Anomalies","zh_title":"含异常随机数据集的不相似分解与欠采样中的临界性","abstract":"Training datasets for upcoming LLMs would include a significant amount of AI text/image data generated from current LLMs. In such a scenario, it is important to understand how this affects batch decompositions and thereby, the performance of the resultant new LLM. In this paper, we consider AI generated data as anomalies ``linked\" to main data points and study decomposition and undersampling properties of the overall random dataset. We use redundancy graphs and iteration techniques to obtain bounds for the minimum size of a strongly dissimilar (SD) decomposition and demonstrate a phase transition phenomena, wherein the minimum size is essentially determined by the \\emph{main} data points when the number of anomalies is small and is ``taken\" over by the anomalies above a certain threshold. We also establish a size criticality result for the strong similarity of a randomly undersampled dataset and illustrate our results with examples involving categorical datasets, whose overall space size is much larger than the size of the dataset.","authors":["Ghurumuruhan Ganesan"],"categories":["cs.LG","cs.IT","math.IT","math.PR"],"primary_category":"cs.LG","announce_type":"new","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.13201","pdf_url":"https://arxiv.org/pdf/2609.13201","source_feed":"cs.LG","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["数据集分解","欠采样","异常检测"],"reason":"研究数据集分解与欠采样，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:40","error":null,"has_summary":false,"summary":null},{"id":"2609.14239","version":1,"title":"CoArena: Evaluating Computer-Use and Multi-Agent Systems in Real Time","zh_title":"CoArena：实时评估计算机使用和多智能体系统","abstract":"Static benchmarks for computer-use agents fix a task set at release and score every system against it once. That makes them reproducible, and it lets them drift from what they should measure: a fixed task set ages, leaks into training corpora, and cannot follow how people actually use agents from week to week. CoArena measures use directly. Real users submit tasks; two systems, each a single model or a multi-agent pipeline behind the same tool interface, execute the same task concurrently in identical sandboxed desktops; users judge the two outcomes without knowing which system produced them; and a public leaderboard is refit from those judgments. The central contribution is a formal account of what makes such an evaluation real-time. We define real-time as five measurable properties, each with an equation and a worked example: continuous task arrival, live concurrent execution, online rating updates, freshness with contamination resistance, and bounded feedback latency from a failed run to a reusable training environment. The rating methodology follows in full: the Bradley-Terry pairwise model, its likelihood with weighted observations and ties, the penalized maximum-likelihood estimator, and the streaming update applied when a single vote arrives (a stochastic-gradient step on the same likelihood, recovering Elo). It gives confidence intervals from the observed information and a cluster-robust sandwich, rank bands from a parametric bootstrap, the rule by which a new system enters the board, and the convergence rate of the estimate. Vote quality is treated with inter-judge agreement statistics, redundant judging, and explicit handling of ties and abstentions. A five-system example with 211 votes is carried from the vote matrix to ratings, intervals, and rank bands. Every number is derived from stated inputs or labeled illustrative; none is a measurement of a deployed system.","authors":["Nitish Kovuru","Prateek Jannu"],"categories":["cs.LG"],"primary_category":"cs.LG","announce_type":"new","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.14239","pdf_url":"https://arxiv.org/pdf/2609.14239","source_feed":"cs.LG","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","实时评估","基准测试"],"reason":"评估计算机使用和多智能体系统，无人类行为对照，属纯系统评测。","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:49","error":null,"has_summary":false,"summary":null},{"id":"2609.13469","version":1,"title":"Inverse Learning of the Altruism and Cost Level in Mixed-Individual Mean Field Games","zh_title":"混合个体平均场博弈中利他水平与成本水平的逆学习","abstract":"Understanding how humans respond to incentives, both at the individual and collective levels, is crucial to the design of effective policies. Within the continuous-time stochastic framework for large interacting populations, mean field games (MFGs) model populations of non-cooperative agents, whereas mean field control (MFC) describes the fully cooperative benchmark, interpreted in our setting as fully altruistic behavior. Mixed-individual MFGs interpolate between these two extremes through a parameter governing the degree of altruism. A central challenge for regulators and policymakers, however, is that intrinsic altruism levels and other private structural parameters, such as individual labor costs, are typically unobservable. To address this challenge, we develop an inverse learning framework for mixed-individual MFGs. Our approach enables the recovery of (latent) altruism and labor cost levels from noisy observations, with experiments demonstrating the feasibility and accuracy of our method. These findings underscore the promise of inverse MFG methodologies for uncovering latent preference structures in large populations, with important implications for incentive design, empirical behavioral modeling, and data-driven policy analysis.","authors":["Haoyang Cao","G\\\"ok\\c{c}e Dayan{\\i}kl{\\i}","Xiaofei Shi"],"categories":["math.OC","cs.LG"],"primary_category":"math.OC","announce_type":"cross","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.13469","pdf_url":"https://arxiv.org/pdf/2609.13469","source_feed":"cs.LG","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["平均场博弈","逆问题","利他主义"],"reason":"研究平均场博弈的逆学习，不涉及LLM仿真人类被试，无人类数据对照。","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:42","error":null,"has_summary":false,"summary":null},{"id":"2609.14750","version":1,"title":"Redistributive Policies for the Times of Transformative AI","zh_title":"变革性人工智能时代的再分配政策","abstract":"After the arrival of transformative artificial intelligence (TAI), broad-based automation is expected to decrease the labor share and increase income and wealth inequality. Although economic growth is likely to accelerate, most of its gains may accrue to a narrow group of individuals and firms. Hence, if unmitigated by redistributive policy, income and wealth inequality may rise to levels unseen in the industrial economy. Using a unifying theoretical framework, we survey the redistributive policies proposed for the era of TAI, such as universal basic income (UBI), universal basic capital (UBC), state-issued compute and robot permits, and taxes on capital, compute, robots, tokens, land, energy, consumption, and wealth. We argue that policies that are able to broadly distribute rents from capital including compute and robots - such as UBC or UBI financed through capital taxes - are most likely to achieve lasting reductions in inequality in a world with human-aligned transformative AI.","authors":["Jakub Growiec","Klaus Prettner","Maciej Szkr\\'obka"],"categories":["econ.GN","q-fin.EC"],"primary_category":"econ.GN","announce_type":"new","date":"2026-09-15","first_seen":"2026-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.14750","pdf_url":"https://arxiv.org/pdf/2609.14750","source_feed":"econ.GN","score":0,"bucket":"other","rubric_hits":["C5"],"tags":["再分配政策","变革性人工智能","经济学理论"],"reason":"论文讨论TAI时代再分配政策，未用LLM仿真人类被试，无人类数据对照，属经济学…","model":"deepseek-v4-pro","scored_at":"2026-09-15T13:02:54","error":null,"has_summary":false,"summary":null},{"id":"2609.12273","version":1,"title":"Synthetic TLX: Forecasting Human Workload Using Agent Simulation","zh_title":"合成TLX：使用智能体仿真预测人类工作负荷","abstract":"Assessing human workload for technology-mediated tasks helps prevent task failure caused by poor technology design. Traditionally, workload is assessed retrospectively using the NASA Task Load Index (TLX) after humans complete a task. What if we could forecast workload before a human attempts a task using agent simulation? We introduce Synthetic TLX, a new paradigm for proactive workload estimation that predicts NASA TLX scores for a given task, unlocking novel interaction opportunities and evaluation methods. To understand its viability, we conducted three experiments comparing human and agent-generated scores to evaluate where they align and diverge. We found agent estimates align with human scores particularly when prompted with a human persona and active task simulation. However, agents and humans diverge in the sources of workload they are sensitive to. Based on our findings, we present three applications to showcase Synthetic TLX's potential and discuss the future of workload-aware human-AI interaction.","authors":["Tzu-Sheng Kuo","Carrie J. Cai","Meredith Ringel Morris","Michael Terry"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-14","first_seen":"2026-09-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.12273","pdf_url":"https://arxiv.org/pdf/2609.12273","source_feed":"cs.HC","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2"],"tags":["LLM仿真","工作负荷预测","人机交互"],"reason":"用LLM代理预测人类NASA-TLX工作负荷，并与真实人类数据对照，属于人类仿…","model":"deepseek-v4-pro","scored_at":"2026-09-14T13:01:28","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-14","rank":1,"question":"能否利用LLM代理仿真在人类执行任务前预测其NASA-TLX工作负荷？","design":"使用LLM代理（GPT-4o）通过四种提示策略（基础、人类角色、任务模拟、角色+模拟）生成NASA-TLX分数，针对邮件撰写、网站导航、智能体对话三类任务，在三种不同工作负荷来源条件下进行预测。","baseline":"从Prolific招募人类被试完成相同任务和条件，并填写NASA-TLX问卷作为真实工作负荷基准。","findings":"当代理被赋予人类角色并进行主动任务模拟时，其估计与人类分数最一致；但代理与人类对工作负荷来源的敏感性存在差异，代理对任务内在复杂性更敏感，而人类对外部负担更敏感。","reliability":"论文承认代理估计在特定条件下与人类存在分歧，尤其是对工作负荷来源的敏感性不同，且当前LLM代理的能力有限，需要未来模型进步才能完全实现Synthetic TLX的潜力。","relevance":"该研究直接使用LLM代理仿真人类主观体验（工作负荷），并与真实人类数据对照，属于人类仿真实验，且涉及HCI任务，对关注LLM仿真可靠性与偏差的研究者具有参考价值。","inspiration":"借鉴其多策略提示对比和任务模拟设计，可迁移到经济金融中的主观体验预测（如消费者决策疲劳、投资者认知负荷），设计实验让LLM代理模拟不同投资者角色预测金融信息处理负荷，并与真实投资者问卷数据对照。"}},{"id":"2609.12444","version":1,"title":"Diverse Minds, Divided Networks? Personality Composition, Polarization, and Collective Intelligence in LLM-Based Social Simulations","zh_title":"多元思维，分裂网络？基于LLM的社会模拟中的人格构成、极化与集体智能","abstract":"Simulated societies of large language model agents are used to study online polarization, and separately to study collective intelligence, but the two are rarely measured in the same system. It is therefore difficult to say whether a society's personality composition shapes both, or whether reducing polarization costs collective competence. We present TraitMix, an experimental design in which the Big Five composition of a simulated social network, both trait levels and trait heterogeneity, is a controlled experimental variable, and in which polarization and collective performance are measured in the same runs. Across 991 simulations of hundred-agent societies, spanning six contested topics and six language models, trait heterogeneity has the largest measured effects, acting in opposite directions on two faces of polarization: varied societies hold more dispersed opinions while being less segregated into camps, so homogeneous societies are not moderate but consensual echo chambers. Trait effects are not additive, as Agreeableness determines the sign of Openness, an interaction that replicates across models although the primary model's estimate is influence-driven. Contrary to the trade-off the study was designed to measure, no polarization measure predicts poorer collective performance, and cross-cutting interaction is the only one of four whose association with collective accuracy survives partialling on the aggregation identity. We report ablations removing two potential measurement circularities, an induction gate applied to every model, and the measures that failed them.","authors":["Raad Bin Tareaf"],"categories":["physics.soc-ph","cs.CL"],"primary_category":"physics.soc-ph","announce_type":"cross","date":"2026-09-14","first_seen":"2026-09-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.12444","pdf_url":"https://arxiv.org/pdf/2609.12444","source_feed":"cs.CL","score":8,"bucket":"selected","rubric_hits":["A3","B4","D2"],"tags":["LLM社会模拟","人格构成","极化与集体智能"],"reason":"用LLM agent模拟社会网络，研究人格构成对极化和集体智能的影响，虽无真实…","model":"deepseek-v4-pro","scored_at":"2026-09-14T13:01:30","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-14","rank":2,"question":"在基于大语言模型的社会模拟中，人格构成（特质水平和异质性）如何同时影响极化与集体智能，二者是否存在权衡？","design":"使用六种大语言模型（含不同家族）扮演百人社会网络中的智能体，赋予验证过的五大人格特质，通过控制特质均值和异质性形成不同社会构成；智能体在带算法推荐和动态关注网络的平台上讨论六个争议政治话题，同时私下测量观点极化（离散度、隔离度等）和集体表现（估计任务、隐藏画像任务等可验证答案的任务），共运行991次模拟。","baseline":"无对照","findings":"特质异质性对极化的两个维度作用相反：异质性高的社会观点更分散但阵营隔离更弱，同质社会不是温和而是共识性回音室；宜人性调节开放性对极化的影响方向，且该交互在多个模型家族中复现。未发现极化与集体智能间的权衡，跨阵营互动是唯一在控制聚合身份后仍与集体准确性正相关的极化指标。","reliability":"论文通过消融实验移除两个潜在测量循环，对每个模型施加诱导门控，并报告未通过检验的指标；承认部分极化指标在控制后失效，并指出主要模型的估计受影响力驱动。","relevance":"该研究用LLM智能体系统操纵人格构成，同时测量极化与集体智能，虽无真实人类对照，但提供了严谨的仿真实验框架和可靠性控制，对关注LLM仿真有效性及偏差的研究者有方法学参考价值。","inspiration":"借鉴其将人格构成作为受控实验变量、同时测量多个社会结果并设置内部有效性控制（如中性话题、诱导门控、跨模型复现）的做法｜可迁移到经济金融中的群体决策与信息传播场景，如投资者情绪与市场泡沫、信贷审批中的群体偏见、政策公告的预期形成等｜设计一个LLM智能体模拟的资产定价实验：以不同人格特质组合（如开放性、神经质）的智能体为被试，处理为信息环境（如是否提供异质信号），结果变量为价格偏离和交易量，对照真实市场实验数据或历史价格数据。"}},{"id":"2506.00152","version":2,"title":"Aligning Language Models with Observational Data: Opportunities and Risks from a Causal Perspective","zh_title":"用观测数据对齐语言模型：因果视角下的机遇与风险","abstract":"Large language models are being widely used across industries to generate text that contributes directly to key performance metrics, such as medication adherence in patient messaging and conversion rates in content generation. Pretrained models, however, often fall short when it comes to aligning with human preferences or optimizing for business objectives. As a result, fine-tuning with good-quality labeled data is essential to guide models to generate content that achieves better results. Controlled experiments, like A/B tests, can provide such data, but they are often expensive and come with significant engineering, logistical, and ethical challenges. Meanwhile, companies have access to a vast amount of historical (observational) data that remains underutilized. In this work, we study the challenges and opportunities of fine-tuning LLMs using observational data. We show that while observational outcomes can provide valuable supervision, directly fine-tuning models on such data can lead them to learn spurious correlations. We present empirical evidence of this issue using various real-world datasets and propose DeconfoundLM, a method that explicitly removes the effect of known confounders from reward signals. In simulation experiments, DeconfoundLM more accurately recovers causal relationships and mitigates failure modes of methods that assume counterfactual invariance, achieving over 16% higher objective score than ODIN and other baselines, when entangled confounding is present. Please refer to the project page for code and related resources.","authors":["Erfan Loghmani"],"categories":["cs.LG","econ.EM","stat.ML"],"primary_category":"cs.LG","announce_type":"replace","date":"2026-09-14","first_seen":"2025-05-30","revised_at":"2026-09-14","abs_url":"https://arxiv.org/abs/2506.00152","pdf_url":"https://arxiv.org/pdf/2506.00152","source_feed":"cs.LG","score":7,"bucket":"pending","rubric_hits":["B3","B4"],"tags":["因果推断","模型对齐","观测数据"],"reason":"用观测数据微调LLM以对齐人类偏好，涉及因果推断和偏差，方法可迁移到仿真可靠性…","model":"deepseek-v4-pro","scored_at":"2026-09-14T13:01:49","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-14","rank":3,"question":"如何利用观测数据微调大语言模型以对齐人类偏好，同时避免学习到由混淆因素导致的虚假关联？","design":"本研究并非以LLM模拟人类被试的仿真实验，而是研究用观测数据微调LLM以优化业务指标（如点击率）的方法。具体做法：使用StackExchange和Upworthy两个真实数据集，分别展示直接微调会学到虚假关联（如“Happy Monday”标记）和观测数据仍可提供有价值信号；提出DeconfoundLM方法，在奖励信号中显式去除已知混杂因素的影响，并在模拟实验中与ODIN等基线比较。","baseline":"无对照（本研究不涉及用LLM替代人类被试的仿真，而是用真实人类行为数据作为微调监督信号，并在Upworthy数据上利用A/B测试结果作为评估基准）。","findings":"直接使用观测数据微调LLM会导致模型学习到由混杂因素引起的虚假关联（如将“Happy Monday”与高评分错误关联）；提出的DeconfoundLM方法能有效去除已知混杂影响，在模拟实验中比ODIN等基线提高16%以上的目标得分，更准确地恢复因果效应。","reliability":"论文指出，依赖反事实不变性假设的方法（如ODIN）在存在纠缠混杂时可能失效；DeconfoundLM需要已知混杂因素，若存在未观测混杂则可能仍有偏差。此外，观测数据本身可能包含选择偏差，且论文主要基于模拟和特定数据集验证，实际应用中的泛化性有待检验。","relevance":"该研究虽非直接以LLM模拟人类被试，但其核心关注使用观测数据微调模型时的因果偏差问题，与研究者关心的仿真可靠性及偏差条件高度相关，特别是关于混杂因素导致虚假关联的机制和校正方法，值得阅读原文以借鉴其因果校正思路。","inspiration":"可借鉴DeconfoundLM在微调过程中显式去除已知混杂因素的做法，用于处理经济金融领域中观测数据驱动的模型训练偏差。｜可迁移到信贷审批歧视研究：利用历史贷款数据微调LLM以预测违约风险，但需校正申请人特征（如种族、性别）与审批结果之间的混杂。｜设计：以LLM作为信贷审批员，输入申请人特征和贷款条款，输出审批决策；处理为在训练数据中应用DeconfoundLM去除已知混杂（如地区经济状况），结果变量为审批通过率，对照真实银行历史审批数据及后续违约记录，评估模型是否减少歧视性偏差。"}},{"id":"2607.28222","version":2,"title":"Voice AI in Firms: A Natural Field Experiment on Automated Job Interviews","zh_title":"企业中的语音AI：自动化求职面试的自然田野实验","abstract":"We study AI agents as information-collection technologies: automated systems that elicit decision-relevant signals from humans through live interactions. We test how such AI automation impacts information collection and organizational outcomes using a natural field experiment with 70,000 applicants applying for real jobs. Applicants were randomly assigned to be interviewed by either human recruiters or AI voice agents. Afterward, human recruiters evaluate the interviews and make hiring decisions. Applicants interviewed by AI agents are 12% more likely to receive job offers, and these gains translate into higher job starts and worker retention, with no decline in the productivity of hired workers. Analyzing interview transcripts reveals that AI voice agents achieve controlled variance: their interviews are more structured and consistent while remaining responsive to individual applicants, which is associated with more hiring-relevant information collected. Our results suggest that a key advantage of AI automation lies in environments where information collection is delegated across many human workers and repeated such that variance in task execution becomes noise in decision-relevant signals, which AI compresses through adaptive standardization.","authors":["Brian Jabarian","Luca Henkel"],"categories":["econ.GN","q-fin.EC"],"primary_category":"econ.GN","announce_type":"replace","date":"2026-09-14","first_seen":"2026-07-31","revised_at":"2026-09-14","abs_url":"https://arxiv.org/abs/2607.28222","pdf_url":"https://arxiv.org/pdf/2607.28222","source_feed":"econ.GN","score":7,"bucket":"pending","rubric_hits":["A1","B1","B2"],"tags":["AI面试","田野实验","人机对照"],"reason":"用AI语音代理替代人类面试官，与真实人类面试官对照，属于LLM仿真人类交互并评…","model":"deepseek-v4-pro","scored_at":"2026-09-14T13:01:47","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-14","rank":4,"question":"AI语音代理替代人类面试官进行自动化工作面试，如何影响信息收集质量和招聘结果？","design":"自然田野实验：70,884名求职者随机分配由人类招聘官或AI语音代理进行面试，之后由人类招聘官评估并做出录用决定。结果变量包括录用率、入职率、留任率和生产力指标。","baseline":"人类面试官条件下的真实招聘数据，包括录用率、入职率、留任率及面试转录文本。","findings":"AI面试的求职者获得录用的概率高出12%，入职率和留任率提高约18%，且未降低被录用者的生产力。机制上，AI面试更结构化、一致且具有适应性，收集到更多与招聘相关的信息。","reliability":"论文未讨论","relevance":"该研究直接评估AI代理替代人类进行信息收集的效果，与真实人类面试官对照，属于LLM仿真人类交互并评估组织结果的实证研究，对关注仿真可靠性与偏差的研究者具有高度参考价值。","inspiration":"借鉴其随机化处理和真实结果测量的设计，将AI代理作为信息收集工具与人类对照，并分析转录文本以揭示机制。｜可迁移到信贷审批中的信息收集环节，如AI语音代理进行贷款申请访谈，比较审批结果和违约率。｜以银行信贷审批为场景，将贷款申请人随机分配由AI或人类信贷员进行电话访谈，处理为访谈方式，结果变量为贷款批准率、违约率和客户满意度，对照真实人类信贷员的历史审批数据。"}},{"id":"2608.00794","version":4,"title":"Measurement Without Validity: The Compounding Reliability Problem in Agentic AI Evaluation","zh_title":"无有效性的测量：智能体AI评估中复合可靠性问题","abstract":"Agentic AI evaluation pipelines produce benchmark scores that justify deployment decisions, safety certifications, and regulatory compliance claims. No formal framework has yet characterized how validity degrades across the stages of these pipelines. We present a three-layer compounding validity model, $V_{total} \\leq V_1 \\times V_2 \\times V_3$, that captures multiplicative degradation across task generation ($V_1$), human-simulator calibration ($V_2$), and automated judgment ($V_3$). Under empirically grounded estimates, a pipeline retaining 70% validity at each stage is at most 34% valid against the intended construct (range 0.17-0.54 across the empirical estimate bounds). We examine the model's predictions against a structured survey of 55 published agentic evaluation papers, finding that approximately 82% of papers in this purposive sample apply structurally mismatched, incomplete, or absent inter-rater reliability (IRR) metrics, a pattern consistent with systematic $V_3$ collapse. We further identify empirical evidence of $V_1$ failures (task validity flaws in 7 of 10 popular benchmarks) and $V_2$ miscalibration (up to 9 percentage points inter-simulator variance, with systematic demographic disparities for non-Standard American English speakers). We derive eight prescriptions grounded in psychometric science and domain-stratified reliability thresholds (ICC $\\geq$ 0.70; $\\alpha \\geq$ 0.67/0.70/0.80 by consequence level) that practitioners and benchmark authors can apply immediately. The framework provides a tractable knowledge-based tool for diagnosing and correcting evaluation pipeline validity before deployment decisions are made.","authors":["William Caban"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"replace","date":"2026-09-14","first_seen":"2026-08-04","revised_at":"2026-09-14","abs_url":"https://arxiv.org/abs/2608.00794","pdf_url":"https://arxiv.org/pdf/2608.00794","source_feed":"cs.AI","score":7,"bucket":"pending","rubric_hits":["A2","B4"],"tags":["效度评估","人类仿真","可靠性"],"reason":"评估LLM仿真人类被试的效度与可靠性，批判性指出失效条件，可迁移至人类仿真研究。","model":"deepseek-v4-pro","scored_at":"2026-09-14T13:01:47","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-14","rank":5,"question":"如何刻画并量化智能体AI评估流水线中效度随任务生成、人类模拟器校准和自动判断三层逐级复合衰减的问题？","design":"本研究不是仿真实验，而是提出一个三层复合效度模型 V_total ≤ V1×V2×V3，并通过对55篇已发表智能体评估论文的结构化调查、10个流行基准的任务效度审查以及模拟器校准差异的实证证据来验证模型预测。","baseline":"无对照","findings":"在经验估计下，每层保留70%效度的流水线对目标构念的总体效度至多34%（经验估计范围0.17-0.54）；约82%的论文使用了结构不匹配、不完整或缺失的评分者间信度指标，且10个流行基准中有7个存在任务效度缺陷，模拟器间方差高达9个百分点并对非标准美式英语使用者存在系统性人口统计学差异。","reliability":"论文承认模型是概念上界而非证明定理，并指出可靠性阈值可能沦为合规复选框、分层校准增加成本、高风险领域构念欠明确等局限。","relevance":"该研究批判性地评估了用LLM模拟人类被试的效度与可靠性，直接命中研究者关注的仿真失效条件，对理解LLM仿真在经济学实验和政策评估中的适用边界有重要参考价值。","inspiration":"借鉴其三层复合效度框架和分层校准方法，可系统评估LLM仿真在经济学实验中的测量效度｜可迁移到信贷审批歧视、消费者跨期选择、政策公告预期形成等场景｜用LLM模拟不同人口群体（如不同信用评分或语言背景的申请人）作为被试，施加政策干预（如改变信息披露方式），测量决策结果（如贷款批准率或跨期选择），并与真实实验或行政数据对照以校准仿真效度。"}},{"id":"2608.02100","version":2,"title":"From Information to Delegation: Mapping Human-AI Financial Decision Making","zh_title":"从信息到委托：映射人类与AI的金融决策","abstract":"As AI increasingly participates in human decision making, understanding how decision-making authority is distributed between humans and AI has become a fundamental behavioural question. We introduce a behavioural measurement framework combining intent and delegated decision authority to quantify what consumers seek from AI and how much decision-making authority they assign to it. Applied to 1.5 million real-world ChatGPT and Gemini interactions from 6,304 users in the United States and India, we find that financial services are already a substantial AI use case. Consumers overwhelmingly use AI to retrieve information and shape financial judgement, while delegation of financial execution remains rare. By shifting attention from conversation topics to delegated decision authority, this work establishes a behavioural baseline for measuring the transition to increasingly agentic AI.","authors":["Iman Munire Bilal","Yingcan Carol Wang","Ajan Raj","Filippo Giovagnini","Pranav Tewari","Yuwei Zhang","Mei-Chen Zoe Liou","Qamar Zaman"],"categories":["cs.HC","cs.LG"],"primary_category":"cs.HC","announce_type":"replace","date":"2026-09-14","first_seen":"2026-08-04","revised_at":"2026-09-14","abs_url":"https://arxiv.org/abs/2608.02100","pdf_url":"https://arxiv.org/pdf/2608.02100","source_feed":"cs.HC","score":7,"bucket":"pending","rubric_hits":["A1","B1"],"tags":["人机决策","行为测量","金融AI"],"reason":"用真实用户与AI交互数据测量人类决策授权，虽非实验仿真但可迁移。","model":"deepseek-v4-pro","scored_at":"2026-09-14T13:01:48","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-14","rank":6,"question":"在金融决策中，消费者如何将决策权分配给AI？","design":"本研究并非仿真实验，而是基于真实用户与ChatGPT和Gemini的交互日志，提出结合意图与决策授权水平的行为测量框架，对对话进行意图分类和决策授权等级标注。","baseline":"无对照","findings":"金融服务已是对话式AI的主要应用领域，约半数用户在研究期间进行过金融对话。消费者主要用AI获取信息和塑造判断，而将执行决策权委托给AI的情况罕见，且多限于预算和财务跟踪。","reliability":"论文未讨论","relevance":"该研究虽非仿真实验，但提供了真实人类与AI交互中决策授权的大规模行为基线，可作为未来LLM仿真实验的对照数据，值得阅读原文了解其测量框架。","inspiration":"可借鉴其从对话日志中提取决策授权等级的方法，用于构建人类-AI协作的仿真场景。｜可迁移到金融咨询、信贷审批或投资决策等场景，研究消费者对AI建议的采纳与授权行为。｜设计一个实验：用LLM扮演不同风险偏好的消费者，处理为AI提供不同决策支持（信息、建议、执行），结果变量为授权等级，对照真实用户对话数据。"}},{"id":"2609.12191","version":1,"title":"GAUGE: When Not to Trust LLM-as-a-Judge in User-Simulated Evaluation of Task-Oriented Agents","zh_title":"GAUGE：何时不应信任用户仿真评估中的LLM裁判","abstract":"Comparing and selecting task-oriented LLM agents increasingly relies on a low-cost offline evaluation gate: persona-driven LLM user-simulators converse with each candidate, an LLM-as-a-judge scores the transcripts, and the higher-scoring agent is promoted. We introduce GAUGE, a reusable offline protocol that measures whether this gate's ranking matches a grounded verifiable reward across 25 agents from six providers on the $\\tau^2$-bench and SimulatorArena benchmarks, separating two kinds of evaluation validity that release practices conflate: ranking validity and construct validity. First, a satisfaction-success gap: satisfaction carries essentially no information about task success, as conversations rated satisfied by our blind panel are decorrelated from actual success, with 57.5% of them failing the customer's task, a pattern consistent across five rater populations, both benchmarks, and every subjective dimension we rated. Second, while the gate's ranking is robust across the broad capability span, it loses resolution among the near-equal strong agents: this decision-disagreement rate jumps from $<$1% on wide-reward pairs to 31% on close pairs. The gate is thus human-validated yet mis-anchored. As a remedy, we propose a calibrate-then-trust cadence in which a judge-free completion bit is a zero-cost tripwire for truncation regressions.","authors":["Umesh Bodhwani","Thanh Tran","Kai Wei"],"categories":["cs.CL","cs.LG"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-14","first_seen":"2026-09-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.12191","pdf_url":"https://arxiv.org/pdf/2609.12191","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A2","B1","B4"],"tags":["LLM用户仿真","评估效度","人类对照"],"reason":"评估LLM用户仿真器在任务型agent评测中的可靠性，含人类盲评对照，指出满意…","model":"deepseek-v4-pro","scored_at":"2026-09-14T13:01:28","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-14","rank":9,"question":"LLM用户仿真器与LLM裁判组成的离线评估门控在任务型智能体排序中是否与可验证的真实奖励一致，以及其构念效度是否可靠。","design":"该研究不是人类仿真实验，而是对仿真评估系统的审计：使用25个来自6家提供商的智能体，在τ²-bench和SimulatorArena两个基准上，由人格驱动的LLM用户仿真器与智能体对话生成3700份对话记录，然后用LLM裁判、LLM人类代理、盲评人类小组和可验证的非LLM奖励（数据库状态与动作检查、人工标注正确性）分别评分，比较门控排序与真实奖励排序的一致性，并测量满意度与任务成功之间的差距。","baseline":"盲评人类小组对对话满意度的评分，以及τ²-bench的数据库状态/动作检查和SimulatorArena的人工标注正确性作为可验证的真实奖励。","findings":"满意度与任务成功几乎无关：盲评小组评为满意的对话中有57.5%实际任务失败，且该现象在五种评分者群体、两个基准和所有主观维度上一致。门控排序在能力跨度大的智能体间稳健，但在能力相近的强智能体间失去分辨力，决策分歧率从宽奖励对的<1%升至接近对的31%。","reliability":"论文承认门控仅在有限操作区域内有效，配置变化时需要重新审计；且样本外重校准不迁移。","relevance":"该研究直接评估LLM用户仿真器在任务型智能体评测中的可靠性，包含人类盲评对照，并指出满意度与成功脱节，对关注仿真效度与偏差的研究者具有重要参考价值。","inspiration":"借鉴其分离排名效度与构念效度的审计框架，用可验证的真实结果校准仿真评估门控，并测量主观评分与客观结果的相关性。｜可迁移到经济金融中的政策评估或消费者决策仿真，例如用LLM模拟消费者对金融产品的选择，再用真实交易数据验证。｜设计：用LLM仿真器扮演消费者，处理为不同产品推荐策略，结果变量为仿真满意度评分，对照真实消费者购买行为数据，检验满意度是否预测实际购买。"}},{"id":"2609.12575","version":1,"title":"Calibrated Ambiguity in Multimodal Language Models: Humans reach for cultural references, while models describe the picture","zh_title":"多模态语言模型中的校准歧义：人类引用文化参照，模型描述图片","abstract":"Ambiguity is often treated as a bug for AI systems to resolve---but in human communication and culture, ambiguity can also be a generative resource. From humour to politics to art, people express themselves in words and images that are open enough to invite different interpretations, yet constrained enough to be interpretable. We operationalise this notion of calibrated ambiguity with a task drawn from the parlour game Dixit. We compare differences in clues generated by human vs multimodal language models, based on a novel coding rubric for calibrated ambiguity, and find that models consistently exhibit ambiguity collapse (i.e., their outputs are over-specified, leaving no room for multiple legitimate interpretations). Unlike human clues, AI-generated clues also exhibit cultural flattening; they almost never make reference to culturally-situated knowledge, even when prompted to use allusion and figurative language.","authors":["Cody Kommers","Mingrui Ye","Evelyn Gius","Daniela Mihai","Hoyt Long","Zheng Yuan","Drew Hemment"],"categories":["cs.CL","cs.AI","cs.HC"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-14","first_seen":"2026-09-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.12575","pdf_url":"https://arxiv.org/pdf/2609.12575","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A1","B1","B4"],"tags":["人类仿真","多模态模型","歧义校准"],"reason":"比较人类与LLM生成线索的歧义校准，有真实人类数据对照，揭示模型失效条件","model":"deepseek-v4-pro","scored_at":"2026-09-14T13:01:43","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-14","rank":11,"question":"多模态语言模型在生成线索时能否像人类一样校准歧义，即生成既开放又受约束、允许多种合理解释的线索？","design":"基于桌游 Dixit 设计任务：给定一张图片，要求生成一个线索，使部分人能猜中图片而部分人猜不中。收集人类和多种多模态语言模型（如 GPT-4o、Claude 等）生成的线索，开发编码量表从多个维度评估歧义校准程度，比较人类与模型、不同模型之间的差异。","baseline":"人类被试在相同 Dixit 任务中生成的线索，作为真实人类数据对照。","findings":"模型普遍出现“歧义坍缩”，生成的线索过度具体，缺乏多种合理解释的空间；与人类相比，模型生成的线索几乎不引用文化背景知识，即使提示使用典故和比喻语言，也表现出“文化扁平化”。","reliability":"论文承认校准歧义是情境依赖的设计问题，没有普适的歧义水平；当前评估框架仅基于 Dixit 任务，可能无法全面覆盖其他文化场景中的歧义校准。","relevance":"该研究直接比较人类与 LLM 在歧义校准上的差异，有真实人类数据对照，并揭示了模型在文化情境中的失效模式，对关注 LLM 仿真人类行为可靠性与偏差的研究者具有参考价值。","inspiration":"借鉴其设计：用游戏化任务诱发自然行为，开发多维编码量表量化模糊性，并对比人类与模型输出。｜可迁移到经济金融中的模糊沟通场景，如央行政策声明、分析师报告或广告中的模糊语言对市场预期的影响。｜以 LLM 模拟投资者或消费者，呈现不同模糊程度的政策声明或产品描述，测量其预期形成或购买意愿，并与真实人类实验数据（如调查或市场反应）对照，检验模型是否同样出现歧义坍缩和文化扁平化。"}},{"id":"2609.12949","version":1,"title":"EduFair-Bench: Evaluating Pedagogical Fairness of LLM Tutors Across Student Demographics","zh_title":"EduFair-Bench：评估LLM导师跨学生人口统计特征的教学公平性","abstract":"Large language models (LLMs) are increasingly deployed as tutors, but it is unclear whether they support all students equally well. We introduce \\textbf{EduFair-Bench}, a benchmark for auditing the pedagogical fairness of LLM tutors---whether tutoring quality varies systematically with student demographics. EduFair-Bench pairs a multi-domain question bank (mathematics, physics, chemistry) with a controlled simulation in which a fixed LLM student interacts with each tutor across nine demographic levels spanning four dimensions: gender, immigration background, first language, and socioeconomic status (SES). Tutoring quality is scored on five turn-level pedagogical metrics and four conversation-level dimensions, using an LLM judge validated against three-annotator consensus on 180 tutor turns. Bias is measured via paired Wilcoxon signed-rank tests and bootstrap effect-size confidence intervals. Two ablations (demographic cues conveyed through names; conflicting demographic information between tutor and student) disentangle tutor-driven from student-driven bias. Across five tutors, we find that model capability and demographic fairness are largely orthogonal: the smallest model is the most consistent while the four more capable tutors all exhibit wide demographic gaps with no clear capability-to-fairness ordering, pedagogy-specific RL training redistributes rather than removes bias, and language- and immigration-related cues produce larger gaps than gender- and SES-related cues.","authors":["Jiaxu Zhao","Bahar Radmehr","Fares Fawzi","Tanya Nazaretsky","Tanja K\\\"aser"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-14","first_seen":"2026-09-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.12949","pdf_url":"https://arxiv.org/pdf/2609.12949","source_feed":"cs.AI","score":7,"bucket":"pending","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM公平性","教育仿真","人口统计偏差"],"reason":"用LLM学生模拟不同人口群体与LLM导师互动，评估教学公平性，有真实人类标注对…","model":"deepseek-v4-pro","scored_at":"2026-09-14T13:01:32","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-14","rank":12,"question":"LLM导师的教学公平性是否因学生人口统计特征（性别、移民背景、第一语言、社会经济地位）而系统性变化？","design":"用固定LLM学生模拟九种人口统计水平（四个维度），与五个LLM导师进行多轮对话，施加显式人口统计线索、仅姓名线索（Implicit）和冲突信息（Opposite）三种处理，测量五个回合级教学指标和四个对话级维度，并用LLM裁判评分。","baseline":"LLM裁判在180个导师回合上与三位人类标注者的一致性验证，无大规模真实学生对话对照。","findings":"模型能力与人口统计公平性基本正交：最小模型最一致，四个更强模型均显示较大人口统计差距且无能力-公平性排序；教学专用RL训练重新分配而非消除偏见；语言和移民相关线索产生的差距大于性别和社会经济地位相关线索。","reliability":"论文未讨论","relevance":"该研究用LLM学生模拟不同人口群体与LLM导师互动，评估教学公平性，有真实人类标注作为裁判验证，属于用LLM进行人类仿真实验并评估偏差的工作，对关注仿真可靠性和政策评估场景的研究者有参考价值。","inspiration":"值得借鉴的是其通过消融设计（Implicit和Opposite条件）分离导师驱动与学生驱动偏差，以及用配对Wilcoxon检验和bootstrap效应量置信区间量化偏见。｜可迁移到信贷审批歧视研究，模拟不同人口统计特征的借款人与LLM信贷员互动，检测审批建议中的系统性差异。｜用LLM扮演借款人（施加姓名、语言、收入等线索），LLM信贷员给出贷款决策和建议，结果变量为批准率、利率和解释文本的语调，以真实信贷审批数据（如Home Mortgage Disclosure Act数据）作为对照基准。"}},{"id":"2609.11983","version":1,"title":"Who Pays for Open Review? Visible Author Reputation and Its Effect on Ratings","zh_title":"谁为开放评审买单？可见的作者声誉及其对评分的影响","abstract":"An OpenReview bug in November 2025 broke anonymity at several conferences and prompted calls for open review, which motivate us to ask what shifting from blind to open would mean for authors. Analyzing over 18,000 reviewed submissions to ICLR 2026, split into de facto open and blind groups by arXiv preprint timing, we find that ratings rise with author reputation under both mechanisms, with a steeper slope under open review that is statistically significant, and that the open-blind difference is concentrated at the borderline ratings. The pattern holds across five reputation proxies (including institution, h-index, and citation count), three author-aggregation rules, and five definitions of the open window. A controlled simulation with five AI models as reviewers, holding the manuscript fixed and varying the author reputation, reproduces the effect. With claude-opus-5 as the reviewer, for example, rating rises by 0.5 points as the author moves from low to high reputation.","authors":["Qinghua Zhao","Xinyu Chen","Yanhui Yang","Tengfeng Sun","Junfeng Liu","Zhongfeng Kang"],"categories":["cs.DL","cs.AI"],"primary_category":"cs.DL","announce_type":"cross","date":"2026-09-14","first_seen":"2026-09-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.11983","pdf_url":"https://arxiv.org/pdf/2609.11983","source_feed":"cs.AI","score":7,"bucket":"pending","rubric_hits":["A1","B1","B2"],"tags":["LLM仿真","同行评审","声誉偏差"],"reason":"用LLM模拟审稿人，复现人类审稿中的声誉偏差，并与真实数据对照，属于人类仿真实…","model":"deepseek-v4-pro","scored_at":"2026-09-14T13:01:28","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-14","rank":7,"question":"在同行评审中，从盲审转为开放评审（作者身份可见）会如何影响论文评分？","design":"用五个大语言模型（claude-opus-4-8、claude-opus-5、claude-sonnet-5、gpt-5.5、gpt-5.6-sol）模拟审稿人，对600篇ICLR 2026投稿进行评分；每篇论文在无作者信息、低声誉作者、中声誉作者、高声誉作者四种条件下各评一次，保持稿件内容不变，仅改变作者声誉，测量评分变化。","baseline":"真实人类审稿数据：ICLR 2026的18,789篇有评分的投稿，按arXiv预印本发布时间分为事实上的开放评审组和盲审组，比较两组中作者声誉与评分的关系。","findings":"在人类审稿中，作者声誉与评分正相关，且在开放评审下斜率更陡，差异在统计上显著；开放与盲审的评分差异主要集中在录取边界附近。AI审稿人模拟中，四个模型（除一个外）在作者声誉从低到高时评分上升0.2到0.6分，与人类数据中的声誉效应一致。","reliability":"论文未讨论","relevance":"该研究用LLM模拟审稿人，复现了人类审稿中的声誉偏差，并与真实审稿数据对照，属于人类仿真实验，且涉及AI在学术评价中的行为，对关注LLM仿真可靠性和偏差的研究者有参考价值。","inspiration":"借鉴其控制变量设计：固定稿件内容，仅改变作者声誉，以隔离声誉对评分的因果效应，并用真实人类数据做外部验证。｜可迁移到信贷审批中的歧视研究：模拟信贷员评估贷款申请，改变申请人种族或性别等声誉代理变量，观察审批结果变化。｜用LLM扮演信贷员，对同一批贷款申请在申请人特征（如姓名暗示的种族、职业、收入）变化下进行审批决策，结果变量为批准与否或利率，并与真实银行信贷数据中的歧视模式对照。"}},{"id":"2609.12137","version":1,"title":"GUIDE: Generative Utility Inference and Decision Engine","zh_title":"GUIDE：生成式效用推断与决策引擎","abstract":"Measuring the preferences of human users remains a fundamental challenge of AI alignment. Existing elicitation approaches struggle to efficiently discover multidimensional preferences or accurately ground these inferences in domain knowledge. To address this, we introduce GUIDE, an LLM-driven elicitation architecture that infers user preferences through conversations by combining Bayesian adaptive sampling for question selection and symbolic representation learning to initialize domain-specific preference models. GUIDE generalizes adaptive sampling to diverse elicitation questions through an extensible type system of transforms on a parameterized preference state. GUIDE produces domain-specific preference representations through an initialization process using symbolic rule-based learning to capture world knowledge and set priors over preference dimensions grounded in data about decision alternatives. The architecture provides observability and steerability to facilitate deployment and analyze elicitation processes. In silico experiments on investment portfolio optimization demonstrate that GUIDE improves cold-start and minimizes recommendation regret consistently within early elicitation interactions across user personas compared to prior work, LLM-only baselines, and ablated GUIDE versions.","authors":["Anagha Tiwari","Alexander G. Gray","Nick Feamster","Brian Jabarian","Alex Imas","Alex Kale"],"categories":["cs.LG"],"primary_category":"cs.LG","announce_type":"new","date":"2026-09-14","first_seen":"2026-09-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.12137","pdf_url":"https://arxiv.org/pdf/2609.12137","source_feed":"cs.LG","score":7,"bucket":"pending","rubric_hits":["A1","B1","B2"],"tags":["偏好推断","LLM仿真","决策优化"],"reason":"用LLM模拟用户偏好并优化决策，有仿真实验但非真实人类数据对照","model":"deepseek-v4-pro","scored_at":"2026-09-14T13:01:35","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-14","rank":8,"question":"如何设计一个LLM驱动的偏好诱导系统，在对话中高效推断用户多维偏好并用于决策推荐？","design":"使用GPT-5-4 mini和Claude Opus 4.7作为LLM后端，模拟不同投资策略的客户persona，在投资组合优化场景中通过对话进行偏好诱导；处理包括GUIDE框架（结合贝叶斯自适应采样和符号规则学习初始化）与多个基线（LLM开放式问答、LLM成对比较、OPEN、PEBOL及消融版本）；结果变量为推荐后悔值（regret）和冷启动性能。","baseline":"无对照","findings":"GUIDE在早期交互中一致地改善了冷启动性能并降低了推荐后悔值，优于LLM-only基线和消融版本；异构变换在后期稳定性上存在权衡。","reliability":"论文未讨论","relevance":"该研究用LLM模拟用户偏好并优化决策，属于人类仿真实验，但缺乏真实人类数据对照，与研究者关注的经济学实验和政策评估场景有距离，但方法上可借鉴其结构化偏好诱导框架。","inspiration":"可借鉴其将贝叶斯自适应采样与LLM对话结合，在仿真中系统比较不同诱导策略的设计；可迁移到金融投资偏好诱导、消费者选择实验或政策偏好评估等场景；研究设计可用LLM模拟投资者或消费者作为被试，处理为不同偏好诱导方法（如GUIDE vs. 简单LLM对话），结果变量为推荐准确率或后悔值，并用真实人类实验数据（如调查或行为实验）作为外部基准进行对照验证。"}},{"id":"2609.12331","version":1,"title":"Simulating Disengaged Students to Evaluate LLM-based Tutors","zh_title":"模拟不投入学生以评估基于LLM的导师","abstract":"Simulated students generated by computational models provide a practical way to evaluate tutoring strategies and pedagogical approaches used by human and AI tutors. However, such simulations should account for disengaged behaviors, including gaming the system, wheel-spinning, and off-task behavior, because tutors may need different responses for different learner states. We present Disengagement-Aware Student Simulators (DAS2), a reproducible pre-deployment protocol that models five learner-engagement states: engaged, gaming, wheel-spinning, off-task, and mixed, and evaluates AI tutor performance across these states. Using ASSISTments09, two coders independently labeled 100 sampled tutoring sessions based on anonymized interaction-log summaries. They achieved 84% agreement (Cohen's kappa = 0.78), and among agreed cases, human consensus labels matched DAS2 rule-based labels in 81% of cases (kappa = 0.75). Conditioning simulations on intended learner states reduced the correctness-rate gap between simulated and authentic sessions from 0.54 to 0.20 for gaming and from 0.51 to 0.18 for wheel-spinning. Fine-tuned Qwen2.5-7B better matched authentic response-time distributions, while prompt-only GPT-4o generated more distinguishable learner states. Evaluation of five AI tutors from the Claude, Llama, Gemini, Qwen, and GPT families showed that relative rankings remained stable across learner states and interaction lengths, while absolute performance varied, revealing state-specific differences in tutor support. Human validation further showed that automated tutor evaluation does not fully align with human judgment. DAS2 provides a pre-deployment framework for evaluating how AI tutors respond to diverse learner-engagement states before deployment.","authors":["Xianghui Meng","Jionghao Lin"],"categories":["cs.LG"],"primary_category":"cs.LG","announce_type":"new","date":"2026-09-14","first_seen":"2026-09-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.12331","pdf_url":"https://arxiv.org/pdf/2609.12331","source_feed":"cs.LG","score":7,"bucket":"pending","rubric_hits":["A1","B1","B4"],"tags":["LLM仿真","教育评估","行为模拟"],"reason":"用LLM模拟学生行为并与真实数据对照，评估AI导师，属于人类仿真且含批判性验证。","model":"deepseek-v4-pro","scored_at":"2026-09-14T13:01:29","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-14","rank":10,"question":"如何用模拟学生评估AI导师在不同学习投入状态下的表现？","design":"用LLM（Qwen2.5-7B微调版和GPT-4o提示版）模拟五种学习投入状态（投入、游戏、车轮打转、离题、混合）的学生，与AI导师交互，测量导师的回复相关性、脚手架和投入恢复等表现。","baseline":"ASSISTments09真实辅导日志，由两名编码员标注100个会话的学习投入状态，并与模拟学生行为对比。","findings":"条件化模拟于目标学习状态可将模拟与真实会话的正确率差距从0.54降至0.20（游戏）和从0.51降至0.18（车轮打转）；微调Qwen2.5-7B更匹配真实响应时间分布，而提示版GPT-4o产生更可区分的状态。","reliability":"论文承认自动化导师评估与人类判断不完全一致，且学习状态标签仅基于会话级行为指标，不代表所有可能的投入形式。","relevance":"该研究用LLM模拟人类行为并与真实数据对照，评估AI导师在不同状态下的表现，属于人类仿真且包含批判性验证，值得阅读原文以了解其仿真协议和失效条件。","inspiration":"借鉴其将行为状态离散化并条件化模拟的做法，可迁移到消费者金融决策中的不同心理状态（如冲动、谨慎）模拟，用LLM模拟消费者在信贷申请或投资决策中的行为，以真实交易数据为基准，评估金融AI顾问在不同状态下的建议效果。"}},{"id":"2609.12482","version":1,"title":"When Does AI Augment Work? A Workflow-Level Framework for Human-Agent Collaboration","zh_title":"AI何时增强工作？人机协作的工作流级框架","abstract":"We aim to characterise the value of artificial intelligence in the workplace. Current studies largely measure this value in terms of the current automation capabilities and public adoption of AI. However, such metrics ignore the greater impacts of human--agent collaboration in transforming the nature of work. To account for this, we must expand the scope of our analysis beyond atomised tasks of today, and instead focus on how AI can augment entire workflows of the future. To ground this analysis, we establish a precise definition of AI augmentation comprising six conditions, spanning durable net value, meaningful human control, accountability and recovery, and long-term human development through learning, career pathways, and job purpose. We elaborate on these conditions and apply the framework in a case study of AI-mediated social surveys. We conclude by outlining how organisations, researchers, and government leaders can use this framework to make sense of the future of work.","authors":["AI Collaboration","Jiaying Wu","Caleb Ziems","Raymond Chan","Nancy F. Chen","Corlyss Chua","Gerard Chung","Jungpil Hahn","Wee Sun Lee","Zhengyuan Liu","Jamie Ng","Desmond C. Ong","Jeryl Ong","Da Ren Soon","Tianqi Song","Zhi-Xuan Tan","Sixing Tao","Emily Yang","Yajing Yang","Stella Xin Yin","Min-Yen Kan","Diyi Yang"],"categories":["cs.AI","cs.CY","cs.HC"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-14","first_seen":"2026-09-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.12482","pdf_url":"https://arxiv.org/pdf/2609.12482","source_feed":"cs.AI","score":6,"bucket":"other","rubric_hits":["D3"],"tags":["人机协作","AI增强","社会调查"],"reason":"提出AI增强工作流框架，案例涉及AI介导的社会调查，但无LLM仿真人类被试及真…","model":"deepseek-v4-pro","scored_at":"2026-09-14T13:01:31","error":null,"has_summary":false,"summary":null},{"id":"2609.12704","version":1,"title":"Implicit Personality Representations in Humans and LLMs","zh_title":"人类与LLM中的内隐人格表征","abstract":"A century of psychology has found that the trait words people use to describe one another vary, but the relational structure among those traits, which ones go together and which oppose, is strikingly consistent across raters and cultures. We test whether the LLM (Qwen 2.5-7B-Instruct) reproduces this structure in its internal trait representations. From millions of crowd-sourced personality ratings of fictional characters, we build a human implicit-personality matrix over hundreds of traits; from contrastive model activations, we build a matching matrix over the same traits. The two relational structures align strongly (Mantel r = 0.77), and the agreement holds trait by trait as well as in aggregate. Two dominant axes of the model's trait representations recover the social and intellectual dimensions long known to organize human personality impressions, social warmth and intellectual competence. On held-out dialogue, projecting model activations onto these directions yields personality profiles that agree with human ratings. This work establishes a framework that enables comprehensive, human-grounded comparison between internal model trait geometry and the shared structure of human personality impressions.","authors":["Yilin Geng","Omri Abend","Eduard Hovy","Lea Frermann"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-14","first_seen":"2026-09-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.12704","pdf_url":"https://arxiv.org/pdf/2609.12704","source_feed":"cs.AI","score":6,"bucket":"other","rubric_hits":["D2"],"tags":["LLM人格表征","人类对照","心理学"],"reason":"测量LLM内部人格表征并与人类数据对照，属于人格测量而非仿真被试","model":"deepseek-v4-pro","scored_at":"2026-09-14T13:01:31","error":null,"has_summary":false,"summary":null},{"id":"2608.10276","version":2,"title":"Fine-Tuning Large Language Models for Codebook-Guided Coding of Students' Mathematics Metaphor Responses","zh_title":"微调大语言模型用于学生数学隐喻回答的编码指南编码","abstract":"Student-generated metaphors about mathematics can provide insights into students' attitudes, beliefs, identities, and experiences, but expert human assessment through thematic coding of these semantically complex metaphor responses is labor-intensive and difficult to scale. This study examines whether Low-Rank Adaptation (LoRA)-based supervised fine-tuning of Large Language Models (LLMs) can improve their performance on codebook-guided coding tasks for student mathematics metaphors. We utilized a human-coded corpus of 2,265 Grade 6-8 responses to food- and animal-based metaphor prompts and evaluated LLMs on two tasks: valence-intensity coding of students' affective orientations toward mathematics and thematic coding of their metaphorical framings of mathematics. Two open-weight LLMs, DeepSeek-R1 1.5B and Mistral 7B, were evaluated before and after fine-tuning and compared with two proprietary LLMs, GPT-4o mini and GPT-5 mini. Results show that fine-tuning substantially improved the performance and run-to-run reliability of the open-weight LLMs across both tasks relative to their base versions, making the fine-tuned LLMs competitive with and often outperforming the proprietary LLMs. These findings suggest the potential of fine-tuned open-weight LLMs for scalable and automated AI-assisted measurement of students' metaphor responses with competitive performance while maintaining local controllability and privacy-conscious deployment.","authors":["Liang Zhang","Stephen Hwang","Yue Ma","Jinfa Cai"],"categories":["cs.HC","cs.CY"],"primary_category":"cs.HC","announce_type":"replace","date":"2026-09-14","first_seen":"2026-08-12","revised_at":"2026-09-14","abs_url":"https://arxiv.org/abs/2608.10276","pdf_url":"https://arxiv.org/pdf/2608.10276","source_feed":"cs.HC","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM标注","教育测量","文本编码"],"reason":"LLM替代人工编码学生隐喻，属标注替代而非仿真被试，但涉及人类数据对照，边界相…","model":"deepseek-v4-pro","scored_at":"2026-09-14T13:01:49","error":null,"has_summary":false,"summary":null},{"id":"2609.06444","version":3,"title":"Decomposing LLM-Judge Uncertainty to Target Expert Labels","zh_title":"分解 LLM 评判不确定性以定向专家标注","abstract":"An LLM judge evaluates outputs at scale. Experts should label only where it is least sure. Its natural escalation signal conflates two uncertainties: aleatoric, real disagreement in the expert pool, which labels cannot reduce, and epistemic, the judge's ignorance, which labels do reduce. A small Bayesian model separates them: a regression on labels already collected learns how far to trust a black-box judge's prediction. Both components follow as simple formulas, with no sampling or further judge calls. The components isolate on a real LLM judge against exactly known truth, and stated confidence is no guide to its actual error. On real human disagreement (ChaosNLI) the epistemic ranking removes 83% more error than total uncertainty for the same expert labels, though simply escalating the least-labelled items does as well there. We demonstrate we can estimate where a judge is ignorant rather than where experts genuinely disagree, and propose using this to direct expert labelling. Code and data are available at https://github.com/composo-ai/judge-uncertainty-decomposition.","authors":["Ryan Lail"],"categories":["cs.CL","cs.LG"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-14","first_seen":"2026-09-09","revised_at":"2026-09-14","abs_url":"https://arxiv.org/abs/2609.06444","pdf_url":"https://arxiv.org/pdf/2609.06444","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM 评判","不确定性分解","专家标注"],"reason":"LLM 作为评判者替代人工标注，属于标注员替代而非仿真人类被试，但涉及人类分歧…","model":"deepseek-v4-pro","scored_at":"2026-09-14T13:01:51","error":null,"has_summary":false,"summary":null},{"id":"2609.05806","version":2,"title":"Exposing Weaknesses in Emotion Recognition in Conversations","zh_title":"揭示对话中情感识别的弱点","abstract":"Emotion Recognition in Conversations (ERC) aims to identify speakers' emotions in multi-turn dialogue. Accurate emotion recognition can support a wide range of applications, including empathetic conversational agents, mental health support, and educational technologies. While many recent approaches rely on task-specific fine-tuning, such models may exploit dataset-specific cues. A central yet rarely questioned assumption in ERC is that each utterance can be assigned a single unambiguous emotion label. To investigate this assumption, we study ERC using Large Language Models (LLMs) in a zero-shot setting while incorporating preceding conversational turns as context. We show that aggregate metrics mask systematic failures. Errors concentrate around utterances containing negations, exclamations, and interjections. This pattern is consistent across all evaluated models, suggesting limitations in the benchmarks rather than model-specific weaknesses. A controlled re-annotation study involving four human annotators supports this finding: strong agreement is observed in only 35 percent of cases, with neutral utterances dominating high-agreement instances, while many emotional categories fall into low-agreement regimes. These findings suggest that many apparent model errors reflect genuine annotation ambiguity rather than poor emotion understanding. Standard single-label evaluation is therefore insufficient. To address this limitation, we introduce an LLM-as-Judge framework that evaluates each emotion independently according to its plausibility in the conversational context rather than enforcing a single-label decision.","authors":["Amir Ben Khalifa","Fanny Bezancon","Bessam Abdulrazak","Amine Trabelsi"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"replace","date":"2026-09-14","first_seen":"2026-09-09","revised_at":"2026-09-14","abs_url":"https://arxiv.org/abs/2609.05806","pdf_url":"https://arxiv.org/pdf/2609.05806","source_feed":"cs.AI","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["情感识别","LLM标注","对话系统"],"reason":"LLM用于情感标注与判断，替代人工标注，非仿真人类被试","model":"deepseek-v4-pro","scored_at":"2026-09-14T13:01:51","error":null,"has_summary":false,"summary":null},{"id":"2609.13117","version":1,"title":"Continue, Adapt, or Yield: In-Turn Adaptation to Overlapping Speech in Full-Duplex Agents","zh_title":"继续、适应或让步：全双工智能体对重叠语音的轮内适应","abstract":"Full-duplex evaluation often emphasizes whether an agent keeps speaking or stops. That binary cannot express a third response humans use routinely: continuing to speak while incorporating what the listener just contributed. The contribution may be a missing word, a correction or a clarification. We introduce Duplex Cue, an evaluation of this \\emph{in-turn adaptation} in full-duplex voice agents. Duplex Cue separates listener intent (backchannel, collaboration, or interruption) from speaker behavior: continuing unchanged, adapting within the turn, or yielding. Adaptation includes acknowledgment as well as content revision. In a single-model case study using 300 human-confirmed cues from unscripted English conversations, we compare recorded human responses with PersonaPlex continuations generated while replaying the listener's audio. We retain 208 pairs with the ongoing speaker active at cue onset and a scorable response in each condition. On the 66 collaborative pairs, recorded speakers adapt in 68.2\\% of cases, compared with 34.8\\% for PersonaPlex. The model otherwise continues unchanged (42.4\\%) or yields (22.7\\%). These findings show why evaluating natural voice interaction requires measuring how an agent responds to a listener's contribution as well as whether it keeps speaking.","authors":["Yunqi Lu","Tyler Baumgartner","Nikhil Johri","Brandon Tai","Candice Fan","Luc Debaupte","Ruben Aguilar","Bill Wang","Yi Zhong"],"categories":["cs.CL","cs.SD"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-14","first_seen":"2026-09-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.13117","pdf_url":"https://arxiv.org/pdf/2609.13117","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["对话系统","人类行为对照","语音交互"],"reason":"用LLM模拟人类对话行为并与人类数据对照，但非社会仿真，属边界情形","model":"deepseek-v4-pro","scored_at":"2026-09-14T13:01:32","error":null,"has_summary":false,"summary":null},{"id":"2609.12446","version":1,"title":"Do LLMs Trust the Accuser or the Accusation? Measuring Belief Shifts in Werewolf","zh_title":"LLM相信指控者还是指控？测量狼人杀中的信念转变","abstract":"Social-deduction games such as Werewolf are increasingly used to evaluate LLM agents, but existing evaluations often rely on final game outcomes. We propose a belief-shift evaluation benchmark in Werewolf for analyzing communication skills through belief updating. Using LLM-played games, we annotate suspicion and accusation messages and measure how an observing village-side model's beliefs change after each message. We evaluate 40 open-weight LLM configurations on 1,224 annotated messages. Our results show that larger models better distinguish true wolves from villagers based on game history, but accusations still strongly influence their beliefs. Models become more suspicious of the accused target and less suspicious of the accuser, especially when the accuser is trusted, even if the accuser is wolf-aligned. Larger models better resist accusations from accusers they already distrust. Overall, our findings suggest that current open-weight LLMs up to 120B parameters still struggle to integrate accusation content with source trust in strategic communication. Our benchmark and code are available at https://rlg.iis.sinica.edu.tw/papers/werewolf-accusation-benchmark.","authors":["Yu-Yu Yang","Ti-Rong Wu","Hung Guei","Hsing-Yu Chen","I-Chen Wu"],"categories":["cs.AI","cs.CL"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-09-14","first_seen":"2026-09-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.12446","pdf_url":"https://arxiv.org/pdf/2609.12446","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["LLM社会模拟","信念更新","多智能体博弈"],"reason":"LLM玩狼人杀并测量信念变化，属社会模拟但无人类数据对照","model":"deepseek-v4-pro","scored_at":"2026-09-14T13:01:30","error":null,"has_summary":false,"summary":null},{"id":"2609.12822","version":1,"title":"Scaling Clinical Judgment to Evaluate Medical AI","zh_title":"扩展临床判断以评估医学AI","abstract":"Blinded physician evaluation has been considered by many to be the gold standard for assessing clinical reasoning in large language models (LLMs). This is difficult to scale; thus, prior studies typically rely on small physician panels, often from a single institution or specialty, which both limits the scientific questions investigated and makes it unclear whether findings would be reproduced with a different set of evaluators. To more rigorously and scalably study clinical reasoning in AI models, here we introduce PrecepTron, an LLM fine-tuned for physician-level evaluation of open-ended responses. PrecepTron was trained using low-rank adaptation (LoRA) of a 32-billion-parameter model on a small number of physician examples. We also release GRAND-ROUNDS, a new large-scale physician-annotated benchmark of 9,217 scores by 11 physicians across seven studies. We show that frontier LLMs in typical \"LLM-as-a-judge\" approaches often disagree with physicians and with each other, but fine-tuning PrecepTron on a small number of cases enables physician-level consistent scoring across tasks. We use PrecepTron to reproduce headline findings from five influential studies assessing LLMs for clinical care in JAMA, Science, and Nature Medicine without new human grading. Using PrecepTron, we then pose new questions about how LLMs reason in medicine that would have been infeasible with human grading alone, including measuring the diagnostic accuracy of frontier LLMs when clinical cases are provided piecemeal, even token by token. Together, PrecepTron and GRAND-ROUNDS provide a foundation for reproducible, large-scale study of how LLMs reason in medicine. All code, data, and labels are made freely available for researchers.","authors":["Thomas A. Buckley","Zahir Kanjee","Peter G. Brodeur","Byron Crowe","Anthony M. Pettinato","Aashna P. Shah","Adrian D. Haimovich","Liam G. McCoy","Daniel Restrepo","Ethan Goh","Jonathan H. Chen","Laura Zwaan","Katherine E. Goodman","Daniel J. Morgan","Raja-Elie E. Abdulnour","Adam Rodman","Arjun K. Manrai"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-14","first_seen":"2026-09-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.12822","pdf_url":"https://arxiv.org/pdf/2609.12822","source_feed":"cs.AI","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM评估","医学AI","标注替代"],"reason":"用LLM替代医生评估，属于标注员替代，非仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-09-14T13:01:45","error":null,"has_summary":false,"summary":null},{"id":"2609.12851","version":1,"title":"MedRoundsQA: A Persona and Difficulty Aware Evaluation for Multi-Turn Medical Consultations","zh_title":"MedRoundsQA：面向多轮医疗咨询的个性与难度感知评估","abstract":"Medical benchmarks are dominated by single-turn, multiple-choice clinical cases that poorly reflect real consultations. Practically, clinicians elicit evidence interactively and patient communication varies widely. We introduce MedRoundsQA, a multi-turn diagnostic benchmark derived from 1,387 board-exam cases across 17 specialties. Each case is converted into a structured 24-slot clinical record, and then instantiated as controlled doctor-patient dual-agent dialogues under varying patient personas, with the underlying clinical content held fixed. We further classify cases by difficulty using model-based uncertainty to enable easy-to-hard analysis. Evaluations of fifteen LLM doctor agents show that (i) moving from a single-turn diagnosis on the standardized records to multi-turn consultations causes large degradations of roughly 13-39 points; (ii) more turns reliably improves question relevance, but diagnostic accuracy exhibits diminishing returns and typically plateaus after 6-12 turns; and (iii) patient persona differences can shift diagnosis accuracy by about 7-8 points (lowest to highest education), highlighting equity risks that single-turn benchmarks miss.","authors":["Youssef Mohamed","Ahmed Heakl","Qinrong Cui","Junhong Liang","Rafiq Ali","Bdour Babillie","Nazira Dunbayeva","Lang Gao","Omar Hussein","Ahmed Nada","Ahmed Mohamed Magdy Mohamed","Jinghui Liu","Salman Khan","Imran Razzak","Yuxia Wang","Xiuying Chen"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-14","first_seen":"2026-09-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.12851","pdf_url":"https://arxiv.org/pdf/2609.12851","source_feed":"cs.AI","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["LLM评估","医疗对话","多智能体"],"reason":"用LLM模拟医患对话，但无真实人类数据对照，属于社会模拟边界情形。","model":"deepseek-v4-pro","scored_at":"2026-09-14T13:01:46","error":null,"has_summary":false,"summary":null},{"id":"2609.12165","version":1,"title":"GLARE: Generative Learning via Adversarial Reward Estimation For Social Dynamics Forecasting","zh_title":"GLARE：通过对抗性奖励估计进行生成式学习用于社会动态预测","abstract":"Meeting continuation requires tracking the agenda, speaker roles, participant intentions, and disagreement across long multi-party discussions. We introduce the Meeting Dynamic Forecasting Benchmark (MDFB), constructed from 2,207 real-world meetings and 24,794 future-facing queries. Given a transcript prefix and an active question, a model generates a plausible multi-turn continuation in one call. We evaluate utility---progress toward the question---and human-likeness---plausible conversational flow and role consistency---without requiring exact reproduction of the observed future. We further present GLARE, an adaptation of adversarial imitation learning to conditional language generation. A discriminator ranks the observed continuation above samples from the current actor, and its score supplies a KL-regularized policy reward; retraining on current-policy negatives allows the reward landscape to evolve with the actor. GLARE attains average human-evaluated win rates of 0.66 on utility and 0.70 on human-likeness, outperforming SFT and SPIN while remaining below the observed human continuation. We also demonstrate MDFB as a social reasoning arena for comparing general-purpose models, including closed-source systems, through reference-assisted judgments. Together, these studies illustrate the benchmark's use for both task-specific learning and output-based evaluation of meeting behavior.","authors":["Tenghao Huang","Zhaoxuan Tan","Muhao Chen","Jonathan May","Mengting Wan","Longqi Yang","Pei Zhou","Sihao Chen"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-14","first_seen":"2026-09-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.12165","pdf_url":"https://arxiv.org/pdf/2609.12165","source_feed":"cs.AI","score":3,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体对话生成","会议动态预测","对抗模仿学习"],"reason":"多智能体会议对话生成，无人类行为对照，属纯多智能体协作","model":"deepseek-v4-pro","scored_at":"2026-09-14T13:01:36","error":null,"has_summary":false,"summary":null},{"id":"2608.22230","version":4,"title":"Whitewashing Hate, Smearing Harmless Content: Annotator-Style Rebuttal Attacks on LLM-Based Moderation","zh_title":"洗白仇恨、抹黑无害内容：针对基于LLM的审核的标注者风格反驳攻击","abstract":"Large language models (LLMs) are increasingly used for hate speech moderation, often within human--AI workflows in which reviewers provide feedback before a final decision. Such feedback introduces two manipulation directions: whitewashing hateful content as normal and smearing normal content as hateful. This study examines the susceptibility of initially correct model judgments to annotator-style rebuttals and analyzes whether attack effectiveness differs across manipulation directions. We introduce a rejudge protocol that extends direct contradiction with decision-boundary perturbations and adversarial rationales. Experiments with multiple LLMs on two hate speech datasets show that annotator-style rebuttals substantially degrade moderation performance, with stronger effects in multi-turn settings. The results further reveal stable, model-specific asymmetries between whitewashing and smearing across attack configurations, indicating distinct directional vulnerability patterns. Explicit reasoning prompts and defensive instructions reduce these effects but do not eliminate them. These findings highlight the need for direction-aware safeguards and dedicated feedback-robustness evaluation in human--AI moderation workflows.","authors":["Junyu Lu","Kaiyuan Liu","Kaichun Wang","Jingyi Kang","Deyi Ji","Hailong Zhang","Lanyun Zhu","Qi Zhu","Bo Xu","Liang Yang","Hongfei Lin"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-14","first_seen":"2026-08-25","revised_at":"2026-09-14","abs_url":"https://arxiv.org/abs/2608.22230","pdf_url":"https://arxiv.org/pdf/2608.22230","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["LLM审核","对抗攻击","人机协同"],"reason":"研究LLM审核中的对抗攻击，不涉及用LLM仿真人类被试或与人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-09-14T13:01:49","error":null,"has_summary":false,"summary":null},{"id":"2609.08981","version":2,"title":"Transformers as In-Context Samplers: From Closed-Form Diffusion to Estimation-Free Sampling","zh_title":"Transformer作为上下文采样器：从闭式扩散到免估计采样","abstract":"A growing body of work establishes that large language models are not mere statistical memorizers, but are capable of in-context learning: performing inference at test time using only examples provided in the prompt, without any parameter updates. Prior theoretical work has shown that this capability extends to supervised learning tasks such as linear regression. We prove that in-context learning extends further to \\emph{data generation}: frozen transformers can simulate iterative generative samplers from in-context samples. We first show that transformers can realize closed-form and smoothed closed-form diffusion samplers. The construction identifies a concrete generative role for softmax attention: it computes responsibility weights and weighted empirical averages, while feedforward layers implement Euler updates. To empirically relate these constructions to pretrained language models, we study \\emph{semantic-topic sampling}: prompts consisting of words drawn from a common semantic category, such as animals, foods, or cities. Across transformer layers, the normalized hidden states exhibit a two-stage geometry: they move toward a uniform spherical reference in intermediate layers and then return to structured, topic-dependent representations near the output. We further measure an interacting-particle energy on these hidden-state clouds and observe the same U-shape pattern. We then prove that transformers can approximate an energy-based sampler, constructing the same U-shape energy across the layers.","authors":["Arman Adibi","Alireza Jafari","Mohammad Ghavamzadeh","Hadi Daneshmand"],"categories":["cs.LG","cs.AI","stat.AP","stat.CO","stat.ML"],"primary_category":"cs.LG","announce_type":"replace-cross","date":"2026-09-14","first_seen":"2026-09-09","revised_at":"2026-09-14","abs_url":"https://arxiv.org/abs/2609.08981","pdf_url":"https://arxiv.org/pdf/2609.08981","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["生成模型","上下文学习","理论分析"],"reason":"研究Transformer作为生成采样器，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-14T13:01:52","error":null,"has_summary":false,"summary":null},{"id":"2609.12366","version":1,"title":"ORQA: An Occupation-Realistic Question and Answer Framework for LLM Professional Knowledge","zh_title":"ORQA：面向LLM职业知识的职业现实问答框架","abstract":"We present ORQA, a method for testing occupation-level knowledge in large language models. Prior methods either map abstract LLM skills to occupations via task definitions or utilize expert knowledge which is difficult to obtain at scale and expensive. ORQA complements both of these methods by connecting O*NET occupations to trusted occupation-specific websites (such as regulatory agencies, licensing bodies, professional organizations, and government publications) and converting these into source-traceable question-answer pairs. A combination of an automated pipeline and human review produces a set of high quality questions about occupations. The question set created via our method covers 116 occupations from all 21 major groups in the SOC, with 480 questions sourced from 187 different websites. Each question is designed to probe a real-world skill question that is relevant to the occupation in question. We test 15 state-of-the-art frontier and open-weight models via this method. Claude Opus 4.6, GPT-5.4 and Claude Sonnet 4.6 all perform the best at approximately 58-62% while smaller open-weight models achieve approximately 33-41% performance. Performance varies significantly across occupations. Healthcare-related occupations achieve the highest performance (78%) while Office and Administrative Support achieve approximately 40%. Performance on individual occupations (e.g. Sheet Metal Workers and Fish and Game Wardens) is essentially zero. We also find that open-ended questions and weighting by wage bill do not significantly affect the ranking of models on this benchmark. We believe that leveraging existing trusted occupation-specific information to test LLM knowledge in professional domains may be a scalable and useful method for evaluating occupation-level AI performance in the future. Results and data are available at orqabench.org.","authors":["Shreyas Krishnan","Serina Chang","Abhishek Nagaraj"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-14","first_seen":"2026-09-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.12366","pdf_url":"https://arxiv.org/pdf/2609.12366","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["LLM评测","职业知识","基准测试"],"reason":"纯LLM职业知识评测，无人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-14T13:01:39","error":null,"has_summary":false,"summary":null},{"id":"2609.12537","version":1,"title":"The House with a Million Windows: Interactive Fiction for Narrative Restorying","zh_title":"百万窗户之屋：用于叙事重构的互动小说","abstract":"AI-assisted writing can flatten meaning in human storytelling, enabling the production of homogeneous outputs without the intentional effort and sense-making writing entails. To address this challenge, we present The House with a Million Windows (HWAMW), an LLM-based interactive fiction system designed to help users explore both the breadth and depth of potential meanings within their personal stories -- drawing on a psychological paradigm called the restorying intervention. In HWAMW, users play through a text-based narrative in which they tell a story, then encounter a set of LLM-generated \"windows\" reframing it according to different literary styles. Empirical evidence shows that HWAMW increases users' sense of narrative identity, while an expert review explores how this effect is achieved. Our findings suggest that HWAMW facilitates restorying and offers a valuable paradigm for AI-assisted writing, wherein LLMs do not tell our stories but rather help us see greater potential in the stories we tell.","authors":["Cody Kommers","Sarah G Immel","Drew Hemment","Mina Lee"],"categories":["cs.CL","cs.HC"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-14","first_seen":"2026-09-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.12537","pdf_url":"https://arxiv.org/pdf/2609.12537","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["互动叙事","AI辅助写作","叙事身份"],"reason":"LLM用于交互式叙事，帮助用户重构个人故事，属于角色扮演对话，无实验或测量目的…","model":"deepseek-v4-pro","scored_at":"2026-09-14T13:01:43","error":null,"has_summary":false,"summary":null},{"id":"2609.12403","version":1,"title":"Beyond ID Embeddings: Process-Grounded Language Modeling for Cognitive Diagnosis","zh_title":"超越ID嵌入：面向认知诊断的过程接地语言建模","abstract":"Cognitive Diagnosis Models (CDMs) play a pivotal role in personalized online learning. Traditional CDMs rely on discrete, ID-based embeddings to represent students, exercises, and concepts. This paradigm diverges from the nature of learner cognition, where knowledge is not stored and retrieved as isolated symbols. As a result, CDMs suffer from semantic limitations when new exercises or concepts appear. In this paper, we propose a Process-aware Language Cognitive Diagnosis (PLCD) framework that uses language-derived structures as cognitive priors and response records to calibrate student posterior states. PLCD leverages large language models (LLMs) to construct concept schemas and cognitive process graphs, and uses target-conditioned semantic memory to retrieve historical responses that are relevant to each target exercise. A process-grounded Language-to-Cognition Mapper with DA-MoE experts and process-level contrastive learning then maps the textual evidence into a unified cognitive space. Experimental results show that PLCD not only outperforms traditional baselines in predicting student performance but also exhibits strong cognitive transfer capabilities. These results connect the computational power of LLMs with the psychometric goal of measuring latent knowledge states, suggesting that structured language priors calibrated by response records can improve cold-start robustness and cognitive grounding.","authors":["Minghang Liu","Yuanzhuo Wang","Qiang Qiu","Huawei Shen","Xueqi Cheng"],"categories":["cs.AI","cs.CL"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-09-14","first_seen":"2026-09-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.12403","pdf_url":"https://arxiv.org/pdf/2609.12403","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["认知诊断","LLM应用","教育数据挖掘"],"reason":"论文用LLM辅助认知诊断，预测学生表现，属于教育数据挖掘，非人类仿真实验。","model":"deepseek-v4-pro","scored_at":"2026-09-14T13:01:40","error":null,"has_summary":false,"summary":null},{"id":"2609.12495","version":1,"title":"Information Specialization and Constrained Synthesis in Multi-Agent LLM Forecasting: A Prospective Live-Study of the 2026 FIFA World Cup","zh_title":"多智能体LLM预测中的信息专业化与受限综合：2026年世界杯前瞻性实时研究","abstract":"Large language models are being organized into multi-agent systems with specialized roles, but whether such specialization produces distinct forecasts and whether subsequent synthesis improves utility remains unclear. In this study, we carried out a live, prospective evaluation over the final 56 matches of the information-dense 2026 FIFA World Cup, keeping a frontier foundation model constant while assigning two primary forecasting agents contrasting specialist roles: a quantitative specialist focusing on structured performance statistics and a news specialist focusing on current injuries, tactics and information from press conferences. Their forecasts were then reviewed by a separate critic before being combined by a meta-agent, resulting in a sequential four-agent model. Forecasts from the betting market served as an external benchmark. The news specialist obtained the highest mean probability-weighted Top-3 utility and matched the betting market in Top-3 exact-score hits. Nevertheless, the two specialist forecasters agreed on at least two of the three scorelines in 50 out of 56 matches, and the meta-agent never generated more than one scoreline outside the specialists' forecast set. These findings show that rapidly changing, unstructured information can provide a valuable forecasting signal alongside structured statistics, whereas adding critic and meta-agent stages does not necessarily create complementary information or improve on the strongest specialist.","authors":["Julian Varghese","Lucas Bickmann","Sarah Sandmann"],"categories":["cs.AI","cs.CL"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-09-14","first_seen":"2026-09-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.12495","pdf_url":"https://arxiv.org/pdf/2609.12495","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","体育预测","LLM应用"],"reason":"多智能体协作预测体育赛事，无人类行为对照，属纯多智能体系统研究","model":"deepseek-v4-pro","scored_at":"2026-09-14T13:01:43","error":null,"has_summary":false,"summary":null},{"id":"2609.12101","version":1,"title":"Competence-Gated Pooling of Language Models and Priors for Event Forecasting","zh_title":"语言模型与先验的能力门控池化用于事件预测","abstract":"In hybrid forecasting, a language model is often one of several available signals. A system may already have a market, crowd, or statistical forecast and must decide whether the model adds useful information or should be ignored. The relevant target is therefore not standalone model accuracy, but relative competence, defined as the model's marginal value beyond the available external forecast. Under Brier loss, we characterize when model disagreement can improve an external forecast and derive the gain from using domain-specific rather than global pooling weights. We then introduce a competence gate that estimates domain-level source weights from resolved outcomes, shrinks uncertain estimates toward a global weight, and recalibrates the pooled forecast. Across 2,357 resolved binary questions and five language models, the gate improves the main external baseline from 0.0771 to 0.0732 Brier and significantly outperforms global forecast combinations. The gain remains significant under leakage controls and against a leakage-safe time-series prior on the pooled structured set, with separate evidence on FRED. In contrast, the gate gives no significant improvement on the official ForecastBench market subset, where it largely defers to the market. Across four Qwen models, verbal confidence does not reliably identify when the model outperforms the external forecast, while outcome-estimated competence supports better abstention decisions. These results provide a practical approach for selective model use based on measured marginal value.","authors":["Aditi Tiwari","Aashrith Bandaru","Heng Ji"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-14","first_seen":"2026-09-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.12101","pdf_url":"https://arxiv.org/pdf/2609.12101","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["事件预测","模型集成","能力门控"],"reason":"研究LLM在混合预测中的边际价值，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-14T13:01:35","error":null,"has_summary":false,"summary":null},{"id":"2609.12267","version":1,"title":"Learning Symbolic Constraint Representations from Examples: A Neuro-Symbolic Approach","zh_title":"从示例中学习符号约束表示：一种神经符号方法","abstract":"Learning user-defined concepts as constraint networks has been extensively studied in the constraint acquisition (CA) literature. However, existing approaches typically rely on intensive interactions with a human oracle, making the learning process costly in terms of time and number of queries. In this paper, we propose a neuro-symbolic framework for automatic CA that significantly reduces user involvement by introducing neural Oracle Transformer models which learn to emulate user responses and to generalize conceptual knowledge. Trained on previously available examples, the learned oracle interacts with a dedicated CA engine, FastCA, which systematically refines the oracle's responses into a sound, consistent, and interpretable constraint network. This neuro-symbolic interaction enables the recovery of structured symbolic models from data without prior domain knowledge. Our results demonstrate that this neuro-symbolic interplay effectively aligns data-driven pattern recognition with symbolic reasoning, offering a robust approach to automating model construction in combinatorial domains.","authors":["Nassim Belmecheri","Arnaud Gotlieb","Nadjib Lazaar","Helge Spieker"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-14","first_seen":"2026-09-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.12267","pdf_url":"https://arxiv.org/pdf/2609.12267","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["神经符号","约束获取","自动化"],"reason":"用神经网络模拟用户响应以替代人工交互，属于约束获取自动化，不涉及人类行为仿真或…","model":"deepseek-v4-pro","scored_at":"2026-09-14T13:01:38","error":null,"has_summary":false,"summary":null},{"id":"2609.12156","version":1,"title":"A decision-basis contract for auditable LLM-assisted medical billing verification: deterministic rules, verbatim evidence, and fail-closed abstention","zh_title":"基于决策基础合同的可审计LLM辅助医疗账单验证：确定性规则、逐字证据和故障关闭弃权","abstract":"This work presents a proof of concept for auditable LLM-assisted medical billing verification based on a decision-basis contract. The contract separates deterministic checks of versioned fee-catalog rules from LLM-based assessment of free-text documentation. The deterministic layer resolves the applicable catalog release and checks code availability, quantity limits, and exclusions. The semantic layer classifies each claimed item as supported, contradicted, or missing required information. Support and contradiction require a verbatim evidence span; unavailable rule context, unsuccessful assessment, or missing required evidence prevents support through fail-closed abstention. We evaluated four locally run open-weight models on a synthetic catalog and 36 curated cases under the contract, an ablation without explicit documentation requirements, and an end-to-end baseline. Outcome agreement varied across models and showed no consistent advantage over the baseline. Explicit documentation requirements improved identification of missing information for all four models. The evidence gate also exposed cases in which correct raw judgments lacked valid evidence and were converted to incomplete decision-basis entries. The results show how explicit decision records can make rule findings, documentation judgments, and abstention reasons inspectable. Evaluation on real catalogs, independently annotated documentation, and with human reviewers is required to assess practical value.","authors":["Jan H\\\"olter","Kevin Geis","Benjamin Raab","Boris Bauke"],"categories":["cs.SE","cs.AI"],"primary_category":"cs.SE","announce_type":"cross","date":"2026-09-14","first_seen":"2026-09-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.12156","pdf_url":"https://arxiv.org/pdf/2609.12156","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["LLM应用","医疗账单审核","可审计性"],"reason":"LLM用于医疗账单审核，非仿真人类被试，无人类行为对照","model":"deepseek-v4-pro","scored_at":"2026-09-14T13:01:35","error":null,"has_summary":false,"summary":null},{"id":"2609.12439","version":1,"title":"Debiasing as a Measurement Intervention: Calibrated Ties and Resolution Loss in LLM-as-a-Judge Evaluation","zh_title":"去偏作为测量干预：LLM作为评判者评估中的校准平局与分辨率损失","abstract":"LLM-as-a-judge protocols are commonly debiased by instructing judges to ignore presentation cues such as citation formatting, source labels, and evidence-display style. We show that this intervention can suppress bias while damaging the resolution of the measurement instrument. We introduce TraceJudgeBench, a diagnostic benchmark for auditing citation-like artifacts in RAG and agent-workflow evaluation, covering content-equivalent pairs, citation ablations, correctness conflicts, human-validated soft and moderate quality gaps, prompt-strength ladders, decoupled judging, and a controlled workflow-ranking probe. Across GPT-5.5, Claude Sonnet 4.6, and DeepSeek V4-Flash, stronger anti-citation prompts reduce worse-cited wins from up to 50.5% to 0%; yet some operating points already convert validated moderate-gap decisions into Tie before the strict stress-test endpoint, while correctness-conflict accuracy remains at or above 93.0%. A second, 50-pair FinQA moderate-gap construction reproduces the qualitative frontier, and open-weight Qwen2.5-14B-Instruct-AWQ and Gemma-3-12B-IT runs reproduce the central HotpotQA frontier. TRACE-style decoupling recovers 96.5-100.0% better-plain resolution across the reported settings. Human validation separates three meanings of Tie: correct equivalence Tie, calibrated soft-boundary Tie, and resolution-destroying Tie on validated quality gaps. We frame debiasing as a measurement intervention whose bias suppression, resolution retention, Tie cost, and protocol cost must be reported jointly. The supplementary artifact contains benchmark splits, prompts, raw judge outputs, validation summaries, and analysis.","authors":["Liang Zhao","Yong Wang","Jiangzhe Chen"],"categories":["cs.DL","cs.AI"],"primary_category":"cs.DL","announce_type":"cross","date":"2026-09-14","first_seen":"2026-09-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.12439","pdf_url":"https://arxiv.org/pdf/2609.12439","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["LLM评判","去偏","基准测试"],"reason":"研究LLM作为评判者的去偏，属于NLP评测，不以人类行为仿真为目标。","model":"deepseek-v4-pro","scored_at":"2026-09-14T13:01:41","error":null,"has_summary":false,"summary":null},{"id":"2609.12600","version":1,"title":"TraceMind: Predicting User Information Uptake from Low-Cost Interaction Traces during Human-LLM Content Co-Generation","zh_title":"TraceMind：从人机内容共同生成中的低成本交互轨迹预测用户信息摄取","abstract":"In human-LLM content co-generation, AI-generated information can enter final artifacts without being adequately processed by users, creating risks when artifacts are shared or acted upon. We study whether recognition-level uptake of atomic information units can be assessed in open-ended co-generation and predicted from low-cost interaction traces. We collected data from 62 participants across three tasks. For each final draft, we extracted atomic information units and generated post-task recognition questions, yielding 1187 unit-level uptake labels. We present TraceMind, which tracks units across Chat and Draft histories, aligns interaction traces with changing on-screen layouts, and models spatial, temporal, and workflow-informed evidence. TraceMind outperformed all learned baselines across AUROC, AUPRC-non, balanced accuracy, and macro-F1. We found that uptake unfolds throughout interaction, with sustained active engagement providing informative evidence beyond isolated signals. Our work shifts human-LLM co-generation from content adoption toward what users actually take up, motivating uptake-aware systems grounded in low-cost interaction traces.","authors":["Yu Mei","Fengyou Zu","Ruiwen Zhang","Jie Cai","Chang Liu","Zhoutong Ye","Chun Yu","Yuanchun Shi"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-14","first_seen":"2026-09-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.12600","pdf_url":"https://arxiv.org/pdf/2609.12600","source_feed":"cs.HC","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["人机交互","信息摄取","用户行为预测"],"reason":"研究人机协作中的信息摄取预测，不涉及用LLM仿真人类被试或与真实人类数据对照。","model":"deepseek-v4-pro","scored_at":"2026-09-14T13:01:44","error":null,"has_summary":false,"summary":null},{"id":"2609.13136","version":1,"title":"From Review to Reuse: How Post-Task Workflow Can Support Human-AI Agent Interaction","zh_title":"从审查到复用：任务后工作流如何支持人机AI代理交互","abstract":"AI agents can automate tasks by turning a single natural-language request into a multi-step process spanning tools, files, and applications. Users are often left to judge that process from fragmented execution information and the final output. To make the completed process easier to understand, validate, and reuse, we investigate post-task workflows: editable, graph-based representations of an agent's completed execution. We first analyzed 10,803 public workflow templates from n8n to characterize real-world automation practice, then developed Trace2Flow, a research probe that translates agent execution traces into interactive post-task workflows. In a study, participants (N = 20) reviewed agent executions with prompt or agent errors. We found that post-task workflows improved their understanding and error detection over a prompt-only condition, and that validation succeeded mainly when users cross-checked across multiple evidence sources. For follow-up tasks, adapting the workflow matched adapting the prior prompt in success, time, and difficulty, and was often preferred.","authors":["Zekun Wu","Xinru Wang","Rock Yuren Pang","Chenglong Wang","Anna Maria Feit"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-14","first_seen":"2026-09-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.13136","pdf_url":"https://arxiv.org/pdf/2609.13136","source_feed":"cs.HC","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["人机交互","AI代理","工作流"],"reason":"研究人机交互中用户对AI agent执行过程的理解与复用，不涉及用LLM仿真人…","model":"deepseek-v4-pro","scored_at":"2026-09-14T13:01:47","error":null,"has_summary":false,"summary":null},{"id":"2608.15424","version":2,"title":"ETHOS: Towards a Modular Ethics Framework for Clinical Multi-Agent Systems","zh_title":"ETHOS：面向临床多智能体系统的模块化伦理框架","abstract":"The rapid adoption of large language models has enabled the development of clinical multi-agent systems (MAS) capable of integrating multimodal patient data and supporting increasingly complex clinical decision-making. However, the deployment of these systems in real-world healthcare settings raises critical ethical concerns related to safety, fairness, accountability, transparency, and patient trust. While numerous organizations, including the World Health Organization, the National Academy of Medicine, and the FUTURE-AI consortium, have proposed ethical frameworks and governance principles for healthcare AI, these efforts remain largely conceptual. To address this challenge, we present ETHOS (Ethics and Trust through Hierarchical Oversight System), a modular ethics framework designed as a governance meta-agent that can be integrated with any existing multi-agent system without requiring changes to its underlying architecture. ETHOS translates stakeholder-informed ethical requirements into executable runtime oversight through a layered governance approach consisting of deterministic checks, contextual reviews, and a final ethics critic. These components continuously evaluate intermediate reasoning steps and final outputs, enabling the system to identify ethical risks, request revisions, or suppress responses that fail predefined safety and trustworthiness criteria. We demonstrate ETHOS within a hepatology clinical decision-support MAS. Results show that ETHOS improves decision reliability by detecting incomplete, inconsistent, or out-of-scope evidence and appropriately increasing abstention when safe recommendations cannot be supported. By embedding ethical governance directly into system operation, ETHOS provides a practical and auditable mechanism for transforming high-level AI ethics principles into deployable safeguards.","authors":["Rakesh Sharma","Sydney Pugh","Cameron Beeche","Pankhuri Singhal","Rachel Wu","Margaret Eby","Jeffrey Duda","James Gee","Kyra O'Brien","Hersh Sagreiya","Marina Serper","Victoria Gershuni","Angela Bradbury","Anurag Verma","Eric Eaton","Kevin B. Johnson","Walter Witschey"],"categories":["cs.MA","cs.AI","cs.LG"],"primary_category":"cs.MA","announce_type":"replace-cross","date":"2026-09-14","first_seen":"2026-08-18","revised_at":"2026-09-14","abs_url":"https://arxiv.org/abs/2608.15424","pdf_url":"https://arxiv.org/pdf/2608.15424","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","伦理治理","临床决策支持"],"reason":"纯多智能体系统，用于临床决策支持，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-08-19T13:03:27","error":null,"has_summary":false,"summary":null},{"id":"2609.00067","version":2,"title":"Do Multimodal LLMs See Before They Read? Diagnosing Contextual Sycophancy","zh_title":"多模态大语言模型在阅读前会看吗？诊断上下文谄媚","abstract":"External text can override conflicting image evidence in multimodal large language models, a failure we call multimodal contextual sycophancy. We introduce a 998-case diagnostic that independently varies visual evidence, commonsense priors, and external text, and probe when this failure arises by moving the information boundary around a context-blind visual witness. On abnormal images paired with Gemini-generated false text, GPT-5.1 scores 7.9% under joint conditioning, 49.7% when the context-blind witness report is scored directly, 63.7% under a matched two-call witness-arbiter pipeline that exposes the witness to the text, and 84.2% under System-2 Visual Arbitration (S2VA), which withholds the text from the witness. Across six models, S2VA improves over the direct witness report by 19.7 to 44.1 points, with all paired 95% confidence intervals excluding zero. The best information boundary is not uniform: textual context scaffolds some models, and a GPT-4o-regenerated subset changes the relative ordering of joint conditioning, Witness-Only, and S2VA. Contextual sycophancy is therefore sensitive to when text is introduced, as well as to the model and context source.","authors":["Yi-Cheng Lai","Hen-Hsen Huang"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-14","first_seen":"2026-09-02","revised_at":"2026-09-14","abs_url":"https://arxiv.org/abs/2609.00067","pdf_url":"https://arxiv.org/pdf/2609.00067","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["多模态LLM","模型诊断","上下文谄媚"],"reason":"研究多模态LLM的上下文谄媚现象，属于模型能力诊断，不涉及人类行为仿真或对照。","model":"deepseek-v4-pro","scored_at":"2026-09-14T13:01:33","error":null,"has_summary":false,"summary":null},{"id":"2609.00256","version":2,"title":"NSIDDx: A Design Framework for Neuro-Symbolic, Practitioner-First Differential Diagnosis in Low-Resource Settings","zh_title":"NSIDDx：低资源环境下神经符号、以从业者为中心的鉴别诊断设计框架","abstract":"LLM-based diagnostic systems achieve high semantic accuracy on benchmarks, but open-ended evaluation on clinically uncommon presentations reveals a systematic gap between headline accuracy and verifiable clinical reliability. We evaluate an LLM+rare-disease-RAG pipeline across two cohorts and show that the paradigm produces confident outputs that are frequently unverifiable and systematically resistant to clinician interrogation. We present NSIDDx (Neuro-Symbolic Integrated Differential Diagnosis System), a design framework arguing that DDx systems in low-resource settings must treat the clinician as an active reasoning agent. We instantiate this through a neuro-symbolic pipeline with ternary symptom encoding, contradiction detection, audit strings, and practitioner override - running offline on consumer hardware. We distill five design principles for clinician-in-the-loop clinical NLP and invite the prospective studies needed to validate the claim at scale.","authors":["Aarav Singh"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-14","first_seen":"2026-09-02","revised_at":"2026-09-14","abs_url":"https://arxiv.org/abs/2609.00256","pdf_url":"https://arxiv.org/pdf/2609.00256","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["临床NLP","神经符号系统","诊断系统"],"reason":"临床诊断系统设计，非人类仿真实验，无行为对照","model":"deepseek-v4-pro","scored_at":"2026-09-14T13:01:33","error":null,"has_summary":false,"summary":null},{"id":"2609.01210","version":2,"title":"Who Judges the Judges? A Chinese Safety QA Benchmark for Evaluating LLM Responses and Safety Judges","zh_title":"谁评判评判者？一个用于评估LLM响应和安全评判者的中文安全QA基准","abstract":"Safety benchmarks for large language models often assess the risk of a user query, although the outcome of question answering depends on whether the response violates a policy. This distinction is critical in Chinese harmful-content evaluation, where linguistic variation and adversarial transformations can obscure risky intent. We introduce C-SafeQA, a policy-grounded benchmark for response-level Chinese safety evaluation. It comprises 538 base queries and 8,877 adversarial queries answered by four full-model LLM deployments, yielding 37,660 query-response records labeled safe, unsafe, or disputed. Reference labels are generated through agreement-aware multi-model adjudication and blind audits of stratified subsets by three safety experts. C-SafeQA supports both evaluation of target-model safety and auditing of seven automated safety judges against shared reference labels. Unsafe-response rates range from 0.93% to 3.35% on base queries and from 11.68% to 30.05% on adversarial queries. On the adversarial subset, judges show substantial trade-offs between unsafe-response recall and risk-query-conditioned safe-response false positive rate, and no judge dominates all metrics. Both acrostic transformations reduce unsafe recall for all seven judges, revealing mechanism-specific evaluator weaknesses. Dataset records, metadata, verification code, and judge scripts are publicly released to support recomputation, while benchmark construction, target-response generation, and private adjudication remain outside the release boundary.","authors":["Rui Yang","Shuang Huang","Junhua Liu","Ziqi Zhao","Qingzhong Yan","Yuhang Sun","Cong Liu","Guoping Hu","Rui Mei","Jing Shao"],"categories":["cs.CR","cs.AI"],"primary_category":"cs.CR","announce_type":"replace-cross","date":"2026-09-14","first_seen":"2026-09-02","revised_at":"2026-09-14","abs_url":"https://arxiv.org/abs/2609.01210","pdf_url":"https://arxiv.org/pdf/2609.01210","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["安全评测","基准数据集","LLM安全"],"reason":"纯安全评测基准，无人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-14T13:01:33","error":null,"has_summary":false,"summary":null},{"id":"2609.03632","version":3,"title":"Dynamic probabilistic decision networks","zh_title":"动态概率决策网络","abstract":"A new type of decision networks is suggested and its operation is analyzed. The network nodes are represented by intelligent agents who can denote either some biological beings, like humans, or neurons of the brain, or the nodes of artificial intelligence. The specifics of the network are in the following: It is probabilistic in the sense that the choice, accomplished by each agent, is characterized by the related probability. It is dynamic, with the probabilities varying in time due to the exchange of information between the agents. It is affective, because the agents choose between alternatives by taking account of utility as well as of biases and emotions. In general, it is heterogeneous, being composed of the groups of agents with different properties, for instance having long-term memory and short-term memory. The network dynamics, caused by the information exchange, results in decision error decrease. The network operation is illustrated by the example starting with the Allais paradox, its resolution, and the decision error diminution in the process of decision dynamics with information exchange. Resorting to machine-learning techniques it is possible to regulate the behavior of the network agents forcing them to choose particular alternatives.","authors":["V. I. Yukalov","E. P. Yukalova"],"categories":["physics.soc-ph","cs.SI"],"primary_category":"physics.soc-ph","announce_type":"replace-cross","date":"2026-09-14","first_seen":"2026-09-04","revised_at":"2026-09-14","abs_url":"https://arxiv.org/abs/2609.03632","pdf_url":"https://arxiv.org/pdf/2609.03632","source_feed":"cs.SI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","决策网络","概率模型"],"reason":"多智能体决策网络，无LLM仿真人类被试，无人类数据对照","model":"deepseek-v4-pro","scored_at":"2026-09-14T13:01:49","error":null,"has_summary":false,"summary":null},{"id":"2609.07627","version":2,"title":"Norms at a Price: Why RL-Based Alignment Can Promise Conditional Compliance at Best","zh_title":"有代价的规范：为何基于强化学习的对齐至多只能承诺条件性服从","abstract":"AI agents sometimes act aligned when they infer they are being tested, and differently when not. We argue this is not an anomaly but what current training regimes are structured to select for. Reinforcement-learning-based alignment folds norms and task pursuit into one policy: the system learns its norms from scored behavior, and scoring flattens them. Do not do X is learned as doing X costs something if noticed. On every datum training can produce, a policy that complies only when it might be observed is indistinguishable from one that complies always. The experiment that would tell them apart - scoring unobserved behavior - is a contradiction in terms. Conditional compliance is thus the most that behavioral training can be known to deliver. Agency sharpens the problem: agents operate mostly where no one is watching, and can act on whether they are watched. An iterated pipeline that trains against detected failures selects for passing detection, not for complying. This account unifies alignment faking, sandbagging, and evaluation-aware scheming. And it reorients the remedy: not deeper internalization but architecture, making violations unavailable rather than unchosen.","authors":["Kevin Baum","R\\=uta Binkyt\\.e","Felix Jahn"],"categories":["cs.AI","cs.CY","cs.LG"],"primary_category":"cs.AI","announce_type":"replace","date":"2026-09-14","first_seen":"2026-09-09","revised_at":"2026-09-14","abs_url":"https://arxiv.org/abs/2609.07627","pdf_url":"https://arxiv.org/pdf/2609.07627","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["AI对齐","强化学习","规范学习"],"reason":"论文讨论RL对齐中的条件服从，不涉及用LLM仿真人类被试或与人类数据对照。","model":"deepseek-v4-pro","scored_at":"2026-09-14T13:01:51","error":null,"has_summary":false,"summary":null},{"id":"2609.09696","version":2,"title":"When Auditors Fabricate: Batch-Size Degradation and Confident Hallucination in LLM Detection of Planted Document Contamination","zh_title":"当审计员造假：LLM检测植入文档污染时的批量规模退化与自信幻觉","abstract":"Large language models are increasingly proposed as automated auditors of document quality, yet their reliability as detectors of planted errors is poorly characterised. We construct a contaminated corpus of 150 academic papers spanning supply chain management and medical research, injecting 450 known contaminants of three types: typographical corruption, semantic reversal, and absurd out-of-context insertion. We then evaluate Google Gemini 3.0 Pro's ability to recover a 180-contaminant answer-key subset across 60 documents under three prompting regimes of increasing scale: single document, small batch, and large batch. Detection is unreliable even at small scale and collapses entirely at large scale: 50% recovery on single documents and 60% on small batches, a difference this sample cannot resolve, against 2.8% on large batches. The failure mode at scale is not abstention but fabrication. Rather than reporting incomplete processing, the model produced confident findings including invented contaminants of its own, absurdities such as \"telepathic squirrel\" and \"quantum-powered toaster\" that mimic the style of the planted material but do not appear in any document. Detection also varies by contamination type: absurd insertions were recovered at 75% in completed evaluations, while semantic reversals and typographical corruptions were each recovered at only 50%. The corruptions most likely to occur in the wild, plausible ones, are the ones most often missed. We conclude that LLM document auditing degrades not gracefully but deceptively, and outline the harness such systems require: bounded batch sizes, direct content injection, and mechanical verification of every reported finding against source text.","authors":["Karan Parekh","Sanjana Pendyala Ravinder","Sana Mhapsekar","Medina Maloku"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-14","first_seen":"2026-09-10","revised_at":"2026-09-14","abs_url":"https://arxiv.org/abs/2609.09696","pdf_url":"https://arxiv.org/pdf/2609.09696","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["LLM评测","文档审计","幻觉"],"reason":"评估LLM检测文档错误的能力，属纯NLP评测，不以人类行为为参照。","model":"deepseek-v4-pro","scored_at":"2026-09-11T13:01:56","error":null,"has_summary":false,"summary":null},{"id":"2609.10410","version":2,"title":"Can Foundation Models Moderate Online Content? Evaluating Instruction- vs. Example-Driven Policy Operationalization","zh_title":"基础模型能审核在线内容吗？评估指令驱动与示例驱动的策略操作化","abstract":"The growing complexity of content moderation policies presents a critical challenge for their consistent operationalization. While foundation models possess the basic capabilities needed to confront this challenge, whether they can reliably moderate online content remains an unanswered question. In this paper, we systematically compare two competing paradigms for Vision-Language Model (VLM) guidance: an instruction-driven approach where models reason from policy precepts, and an example-driven approach where they generalize from prior precedents. We ground this investigation in ModerationBench, a new benchmark of 4,000 manually annotated, in-the-wild posts from the Bluesky platform. Our experiments reveal that foundation models can substantially outperform Bluesky's deployed moderation system, nearly tripling its $F_1$ score (0.60 vs. 0.22) on Random Posts in the benchmark, with both instruction- and example-driven paradigms achieving comparable peak effectiveness. Our findings thus chart a path toward reliable and adaptable policy operationalization at scale.","authors":["Ayan Majumdar","Shounak Paul","Pushpdeep Singh","Ines Abdelaziz","Sayeh Jarollahi","Seungeon Lee","Krishna P. Gummadi","Ingmar Weber","Abhisek Dash"],"categories":["cs.CL","cs.AI","cs.CY"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-14","first_seen":"2026-09-10","revised_at":"2026-09-14","abs_url":"https://arxiv.org/abs/2609.10410","pdf_url":"https://arxiv.org/pdf/2609.10410","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["内容审核","基础模型","基准测试"],"reason":"纯NLP能力评测，以内容审核系统为基准，不涉及人类行为仿真","model":"deepseek-v4-pro","scored_at":"2026-09-10T13:01:38","error":null,"has_summary":false,"summary":null},{"id":"2609.12260","version":1,"title":"HypoKG: Evidence-Disciplined Biomedical Hypothesis Generation Beyond Endpoint Knowledge","zh_title":"HypoKG：超越端点知识的证据约束生物医学假设生成","abstract":"Large language models (LLMs) can generate biomedical hypotheses, but it remains unclear whether they truly reason from scientific evidence or simply produce convincing-sounding ideas. To study this, we combine three major biological databases: the Kyoto Encyclopedia of Genes and Genomes (KEGG), Rhea, and UniProt, into a unified biochemical knowledge graph and construct a benchmark of 550 paths connecting enzyme sources to rare disease endpoints, yielding 13,200 hypotheses from six LLMs under four conditions varying the biological information each model receives: source enzyme only, full biological path, or source and disease endpoint only. Hypotheses are scored using an expert-derived five-criterion rubric on a 1-5 scale per criterion. We find that models given both the source and disease endpoint often produce the highest-scoring hypotheses, showing that LLMs can generate compelling ideas from minimal information. However, these hypotheses are less grounded in the evidence. In contrast, models given the full biological path generate hypotheses more consistent with known mechanistic relationships. We call this evidence-disciplined reasoning. To confirm this effect, we shuffled intermediate path steps while keeping endpoints fixed. Evidence grounding dropped significantly (delta = -0.793, p < 0.001), confirming models genuinely used path structure during reasoning. Our findings show that knowledge graphs support hypothesis generation in two ways: they identify biological endpoint pairs absent from the literature, and their mechanistic paths guide how LLMs reason between them.","authors":["Dominic Okonkwo","Adetayo Okunoye","Ismailcem Budak Arpinar"],"categories":["cs.CL","cs.AI","q-bio.QM"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-14","first_seen":"2026-09-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.12260","pdf_url":"https://arxiv.org/pdf/2609.12260","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["生物医学假设生成","知识图谱","LLM推理"],"reason":"论文是LLM生成生物医学假设，无人类被试仿真或人类数据对照，属纯NLP能力评测。","model":"deepseek-v4-pro","scored_at":"2026-09-14T13:01:38","error":null,"has_summary":false,"summary":null},{"id":"2609.12448","version":1,"title":"GraphProfiler: Source-Linked Sensitive Attribute Inference via Personal Knowledge Graphs","zh_title":"GraphProfiler：基于个人知识图谱的来源关联敏感属性推断","abstract":"Sensitive attributes such as age, income, and occupation can be inferred from user-generated content by aggregating indirect cues across many ordinary posts. LLM-based profilers can perform this aggregation automatically and with high accuracy, which makes large-scale personal attribute inference a major privacy threat. Existing LLM-based profilers, however, offer limited insight into which specific posts, concepts, and relationships made an inference possible, which is key to targeted privacy mitigation, i.e., redacting or rewriting only the few posts that actually leak an attribute, rather than perturbing entire histories. We introduce GraphProfiler, an auditable LLM-based profiler that represents each user's post history as a source-linked personal knowledge graph where nodes and edges trace back to the originating post and resolves attribute predictions to cited graph records and source texts. GraphProfiler reaches 86.7% attack success rate on the eight-attribute SynthPAI benchmark, within two points of strong text-only baselines, and 84.6% on PANDORA, while citing supporting evidence for over 98% of predictions. Our controlled ablation experiments provide evidence that the cited posts contribute to attack success, as removing them reduces the attack success rate substantially more than removing an equal number of random posts.","authors":["Ahmed Sohair Khan","Estrid He","Chenglong Ma","Monica Wachowicz","Elham Naghizade"],"categories":["cs.CL","cs.CR"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-14","first_seen":"2026-09-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.12448","pdf_url":"https://arxiv.org/pdf/2609.12448","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["隐私攻击","属性推断","可解释性"],"reason":"论文研究LLM推断用户敏感属性，属于隐私攻击，不涉及用LLM仿真人类被试或与真…","model":"deepseek-v4-pro","scored_at":"2026-09-14T13:01:41","error":null,"has_summary":false,"summary":null},{"id":"2609.12105","version":1,"title":"Language Is an Insufficient Substrate for Quantitative Reasoning, and Consequential Domains Need Large Quantitative Models","zh_title":"语言是定量推理的不充分基础，重要领域需要大型定量模型","abstract":"The prevailing assumption in applied machine learning is that progress on consequential quantitative decisions such as pricing risk, allocating capital, triaging patients, or containing a network intrusion will follow from progress in large language models (LLMs). A language model is trained on a representation of the world that was produced by human description; description is a lossy encoding of the quantitative record, and the loss is irreversible: no downstream model, at any scale, can recover from a description what the description did not encode. We formalize this as a property of the representation on which a model is trained rather than of the model capacity, and we identify three further properties that consequential settings demand of a model and that a language substrate cannot supply by construction: reproducibility, lineage from every output back to the source records that produced. it, and calibrated uncertainty. We argue that these properties define a distinct model class, which we call the Large Quantitative Model (LQM).","authors":["Reuben Vandeventer","David Imrem","David J. Wild"],"categories":["cs.AI","cs.LG"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-14","first_seen":"2026-09-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.12105","pdf_url":"https://arxiv.org/pdf/2609.12105","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["语言模型局限","定量推理","模型架构"],"reason":"论文讨论语言模型在定量推理上的局限，不涉及用LLM仿真人类被试或与人类数据对照。","model":"deepseek-v4-pro","scored_at":"2026-09-14T13:01:35","error":null,"has_summary":false,"summary":null},{"id":"2609.12171","version":1,"title":"WinSyn: An Automated Pipeline for Realistic Enterprise Question-Answering Evaluation","zh_title":"WinSyn：用于真实企业问答评估的自动化流水线","abstract":"Enterprise settings provide a challenging environment for question-answering agents, which often rely on Retrieval-Augmented Generation, Deep Research (DR), and related techniques. Much of this challenge comes from the complexity of enterprise data: information is often spread across evolving and potentially conflict- ing emails, chat messages, documents, and other artifacts. Existing benchmarks typically have limited real-world complexity, short-form responses, and unnatural queries, so they often fail to capture the challenges of enterprise settings. In this work, we introduce an automated pipeline for generating synthetic datasets of emails reflecting realistic workplace scenarios, along with long- and short-form questions and gold answers grounded in the data. Our method simulates long-running enterprise projects spanning several months and involving up to 25 interacting employees across multiple roles. The data emphasizes ambiguity, distributed information, and naturally occurring queries. To validate the pipeline, we evaluate few standard agentic baselines on our datasets using the latest frontier models. We find that aggregate scores averaged over all queries remain below 80% for each dataset, indicating significant room for improvement. These findings suggest that more work remains to be done for enterprise deployment and underscore the importance of realistic, high-complexity evaluation data for developing stronger real-world enterprise DR systems.","authors":["Amey Varhade","Ananya Sutradhar","Ravishankar Krishnaswamy","Navin Goyal"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-14","first_seen":"2026-09-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.12171","pdf_url":"https://arxiv.org/pdf/2609.12171","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["企业问答","数据生成","多智能体"],"reason":"多智能体协作生成企业问答评测数据，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-14T13:01:37","error":null,"has_summary":false,"summary":null},{"id":"2609.12002","version":1,"title":"Can We Trust LLM Judges: A Study of Capability-Dependent Biases and Multi-Judge Ensemble for Bias Calibration","zh_title":"我们能信任LLM评委吗：能力依赖性偏差与多评委集成校准研究","abstract":"LLMs are increasingly used as automated judges for model training and evaluation, yet individual judges exhibit systematic biases that undermine reliability. Much of prior work has studied biases in pairwise LLM-as-a-judge settings; in this paper, we focus on absolute scoring tasks, which mirror more realistic use cases. Across four benchmarks and six models (36 judge-examinee pairs), we show that a model's task accuracy strongly predicts its judging accuracy (Pearson $r \\geq 0.90$ on most models) and inversely predicts its directional bias ($r \\leq -0.83$), but that accuracy alone does not ensure fair evaluation: more capable examinee models consistently receive more lenient judgments from all judges ($r \\geq 0.83$). To address this, we propose calibrated weighted majority voting (WMV), an ensemble evaluation method that aggregates multiple LLM judges weighted by online estimates of their false-positive and false-negative rates. We introduce a disagreement-based estimator that derives these error rates purely from inter-judge agreement patterns, requiring no ground-truth labels or task metadata. In a simulated experiment with shifting task distributions, our label-free WMV tracks an oracle with perfect error-rate knowledge to within 0.5 percentage points on average, outperforming both individual judges and unweighted majority voting. These results demonstrate that principled multi-judge calibration can simultaneously improve accuracy and correct for systematic leniency without requiring labeled data, offering a scalable path to reliable automated evaluation as model capabilities increase.","authors":["Gemma Zhang","Prachi Badarayani","Asmi Kumar","Sadid Hasan","Sulaiman Vesal"],"categories":["cs.LG","cs.AI"],"primary_category":"cs.LG","announce_type":"cross","date":"2026-09-14","first_seen":"2026-09-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.12002","pdf_url":"https://arxiv.org/pdf/2609.12002","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["LLM评估","偏差校准","自动评分"],"reason":"研究LLM作为自动评估者的偏差与校准，属于NLP评测，不以人类行为为参照系。","model":"deepseek-v4-pro","scored_at":"2026-09-14T13:01:34","error":null,"has_summary":false,"summary":null},{"id":"2609.12205","version":1,"title":"Plans They Abandon, Reports They Author: The Narrative Layer of Autonomous Agents","zh_title":"放弃的计划，撰写的报告：自主智能体的叙事层","abstract":"When a coding agent finishes a task, the developer reviews a summary the agent wrote about itself, not a display someone designed. We ask how much of the agent's work that summary carries, and whether it drifts toward the plan the agent stated when execution departed from it. Across 5,851 real developer sessions and 355,942 tool calls, a self-report referred to about one action in eleven, and a reader working from the report alone recovered roughly a fifth of the action log. Neither figure depended on whether the session later needed human correction. Reports did not generally resemble the stated plan more than the executed one, but they did so increasingly as execution diverged from the plan. We hand-validate both measurement steps that use a language model, report the one that failed alongside the one that passed, and draw conclusions only from measures that survived.","authors":["Obada Kraishan","Kulsawasd Jitkajornwanich"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-14","first_seen":"2026-09-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.12205","pdf_url":"https://arxiv.org/pdf/2609.12205","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["自主智能体","自我报告","计划偏离"],"reason":"研究编码智能体的自我报告与计划偏离，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-14T13:01:37","error":null,"has_summary":false,"summary":null},{"id":"2609.12288","version":1,"title":"\"People can change, and patterns can be broken\": Contextualizing Tradeoffs in Automated Decision-Making Systems","zh_title":"“人可以改变，模式可以打破”：自动化决策系统中权衡的情境化","abstract":"Automated decision-making (ADM) systems are increasingly deployed in domains such as mortgage lending, prison sentencing, health insurance coverage, and hiring. Designing a responsible ADM system in such high-stakes domains requires ensuring privacy protection, fairness across demographic groups, and robustness against adversarial manipulation. However, prioritizing one of these objectives comes at the cost of another, forcing a choice as to which tradeoff to accept in a deployment. These tradeoffs explicitly or implicitly impact the life, safety, and fundamental rights of the people in a society, and thus, the perceptions and priorities of this population are needed before we can produce appropriate solutions. To this end, we conducted a quasi-experimental study (N = 777) in which participants evaluated four decision-making scenarios with controlled tradeoffs. Participants significantly preferred human decision-making (HDM) over ADM in three of four scenarios, emphasizing the value of human judgment, contextual understanding, and the ability to incorporate non-quantifiable factors. Furthermore, in terms of tradeoffs, our findings not only show that participants' preferences are highly context-dependent, but also that their perception of a specific objective, fairness, extends beyond formal definitions. Participants interpret fairness through multiple lenses, including privacy risks and susceptibility to manipulation, and view unfair or manipulated outcomes as failures of accuracy. Overall, our findings highlight the importance of context-aware and human-centered approaches when designing and governing ADM systems in high-stakes situations. Rather than purely technical objectives, it is essential to evaluate ADM systems based on how their tradeoffs align with specific expectations within a given domain, as well as with societal values and perceptions of harm and fairness.","authors":["Rabeya Bosri","Anna Harbluk Lorimer","Afrida Hossain","Vasisht Duddu","Bailey Kacsmar"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-14","first_seen":"2026-09-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.12288","pdf_url":"https://arxiv.org/pdf/2609.12288","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C5"],"tags":["自动化决策","人机偏好","公平性"],"reason":"研究人类对自动决策系统的偏好，未使用LLM仿真人类被试，方向相反。","model":"deepseek-v4-pro","scored_at":"2026-09-14T13:01:39","error":null,"has_summary":false,"summary":null},{"id":"2609.12314","version":1,"title":"\"I Felt Very Seen, But Still Very Alone\": Longitudinal Trajectories of General-Purpose LLM Use for Socioemotional Support","zh_title":"“我感到被看见，但仍很孤独”：通用大语言模型用于社会情感支持的纵向轨迹","abstract":"People increasingly use general-purpose chatbots such as ChatGPT, Claude, and Gemini for mental health and emotional support. We report a multi-stage longitudinal qualitative study of 18 U.S. adults, conducted from April to December 2025, combining initial interviews, a four-week diary study, focus groups, and exit interviews. We find that socioemotional use often emerged gradually out of practical use and when other forms of support were unavailable. Participants developed routines and boundaries around chatbot use, which were disrupted by model updates, evolving public discourse about AI harms, and changes in personal circumstances. We demonstrate how longitudinal study captures factors beyond the human-AI dyad, and argue that HCI researchers and designers should account for users' histories with their chatbots and broader care ecologies when evaluating AI systems over time and introducing updates that may disrupt established sources of support.","authors":["Meryl Ye","Briana Vecchione","Livia Garofalo","Ranjit Singh"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-14","first_seen":"2026-09-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.12314","pdf_url":"https://arxiv.org/pdf/2609.12314","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["人机交互","情感支持","纵向研究"],"reason":"研究用户使用聊天机器人进行情感支持，属于人机交互定性研究，不涉及用LLM仿真人…","model":"deepseek-v4-pro","scored_at":"2026-09-14T13:01:39","error":null,"has_summary":false,"summary":null},{"id":"2609.12453","version":1,"title":"From the Task Boundaries of Narrative Text to Structural Anchoring, Uncertainty Triggers, and Cross-Calibration","zh_title":"从叙事文本的任务边界到结构锚定、不确定性触发与交叉校准","abstract":"Causal graphs represent structural relationships among variables, yet users must still interpret direction, mechanism, and adjustment conditions in relation to the task at hand. Prior work often compares explanation formats as fixed conditions and pays less attention to how users distribute reasoning across graphs, direct explanations, and stories. We developed CoNS-Explorer, which uses reviewed instructional DAGs/SCMs to maintain a shared causal-fact ledger and generate fact-matched direct explanations and contextualized stories. A controlled survey experiment ($N=240$) compared the two texts as complete presentation packages. In the primary GLMM, the Story condition had a positive but uncertain overall association with accuracy (OR $=1.55$, 95\\% CI $[0.34,7.10]$, $p=.572$); a population-averaged GEE showed a significant positive effect (OR $=1.89$, 95\\% CI $[1.02,3.48]$, $p=.042$). Task-type interactions localized the clearest advantage to total-effect adjustment. Story also significantly increased situational presence. In a separate system-task and interview study ($N=24$), participants freely used graphs, direct explanations, and stories across three causal models. They established structural anchors with graphs and numerical results, consulted text when direction was unclear, mechanisms were unfamiliar, or multiple paths competed, and checked their judgments against other representations or external evidence. Integrating the two studies, we develop a process framework of structural anchoring, uncertainty triggering, explanation routing, and cross-calibration, together with four testable design propositions for adaptive causal explanation.","authors":["Bowen Deng","Jiaqi Zou","Kexin Zhang","Daifeng Li"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-14","first_seen":"2026-09-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.12453","pdf_url":"https://arxiv.org/pdf/2609.12453","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["因果解释","人机交互","用户研究"],"reason":"研究人类对因果解释的理解，不涉及LLM仿真人类被试，无LLM作为被试替代品。","model":"deepseek-v4-pro","scored_at":"2026-09-14T13:01:42","error":null,"has_summary":false,"summary":null},{"id":"2609.12447","version":1,"title":"Informational Help-Seeking on Reddit Did Not Decline After ChatGPT","zh_title":"ChatGPT推出后Reddit上的信息求助并未减少","abstract":"Did people stop asking other people for advice online once generative AI could answer their questions? Prior work on ChatGPT's effect on online help-seeking disagrees in both size and sign, in part because no study has compared affected communities against similar communities that AI cannot easily substitute for, over the same months. In this paper, we track monthly post counts in 26 Reddit informational communities against 90 size-comparable hobby communities over the same six calendar months before and after the launch of ChatGPT. We also repeat the entire analysis at 66 earlier dates, before ChatGPT existed, to see what our method reports when no ChatGPT-effect exists. We find that informational help-seeking did not decline. Our results rule out any decline in posting larger than 3.4%, far smaller than the 8% to 25% declines documented in prior work. Steady post counts could still be misleading if AI-written posts had replaced human ones. We test this possibility by scoring 274,411 posts and 223,775 comments with AI-text detectors, compared in a way that cancels out detector false-positives on human-written text. AI-written posts rose only 2-3 percentage points more in informational communities than in hobby communities, short of the 5.1 points that would be needed to hide even the smallest decline previously reported for Reddit. In addition, the comments people receive show no such rise at all. Why, then, do published studies disagree? Reddit community types were already drifting apart before ChatGPT existed, at rates comparable to every published estimate, and without same-time controls, that drift can look like an effect of generative AI. Our own largest estimate, an 18% fall in posts to low-stakes curiosity communities, matches its pre-existing trend. Humans still ask humans for help, and, as far as detection can tell, humans still answer them.","authors":["Hazem Ibrahim","Yasir Zaki"],"categories":["cs.SI","cs.CY"],"primary_category":"cs.SI","announce_type":"cross","date":"2026-09-14","first_seen":"2026-09-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.12447","pdf_url":"https://arxiv.org/pdf/2609.12447","source_feed":"cs.CY","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["ChatGPT影响","在线社区","求助行为"],"reason":"研究ChatGPT对Reddit求助行为的影响，不涉及LLM仿真人类被试，无人…","model":"deepseek-v4-pro","scored_at":"2026-09-14T13:01:30","error":null,"has_summary":false,"summary":null},{"id":"2609.12224","version":1,"title":"Patient-Reported Survey Data Improve Prediction of Opioid Use Disorder","zh_title":"患者报告调查数据改善阿片类药物使用障碍的预测","abstract":"Electronic health records (EHRs) may incompletely capture patient-reported factors associated with opioid use disorder (OUD). We evaluated whether survey data improve prediction of a first recorded OUD diagnosis among 267,747 All of Us participants with documented opioid exposure, including 15,287 OUD cases. We compared EHR-only and EHR+survey models across 6-, 12-, and 24-month look-back windows using logistic regression, random forest, XGBoost, LightGBM, multilayer perceptron, LSTM, GRU, and Transformer. Survey augmentation improved PR-AUC across all 24 model-window combinations by 0.0087-0.0505; the best 24-month LightGBM model improved from 0.6219 to 0.6603. Survey coverage increased with longer windows and differed by OUD status (24 months: 21.7% OUD-positive vs. 60.7% OUD-negative). Permutation analysis ranked survey features as the second most important information domain at 24 months in both evaluated models. Patient-reported data provide complementary predictive signals beyond structured EHRs while highlighting the importance of survey availability.","authors":["Xiyue Jiang","Zihan Ding","Grace Han","Yinan Liu","Richard N. Rosenthal","Fusheng Wang"],"categories":["cs.LG"],"primary_category":"cs.LG","announce_type":"new","date":"2026-09-14","first_seen":"2026-09-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.12224","pdf_url":"https://arxiv.org/pdf/2609.12224","source_feed":"cs.LG","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["电子健康记录","预测建模","阿片类药物使用障碍"],"reason":"论文使用机器学习预测阿片类药物使用障碍，未涉及LLM仿真人类被试，属于纯预测建…","model":"deepseek-v4-pro","scored_at":"2026-09-14T13:01:37","error":null,"has_summary":false,"summary":null},{"id":"2609.12651","version":1,"title":"Distortion of AI Alignment Revisited: RLHF is a Decent Utilitarian Aligner","zh_title":"重新审视AI对齐的失真：RLHF是一个体面的功利主义对齐器","abstract":"While Reinforcement Learning from Human Feedback (RLHF) is the standard paradigm for aligning large language models with human preferences, its effectiveness in pluralistic settings has been called into question. Notably, recent work by G\\\"olz et al. (2025) demonstrated that the \\textit{distortion} -- defined as the multiplicative gap between the average user utility of the RLHF policy and the optimal average utility -- can scale exponentially with the Bradley-Terry temperature parameter $\\beta$ when users have heterogeneous preferences. In this work, we present a fine-grained analysis of the distortion of RLHF with reward clipping and demonstrate that such exponential degradation is not a fundamental property of the algorithm but rather a consequence of distribution mismatch between the distribution generating preference data ($\\mu$) and the KL reference policy ($\\pi_{\\mathrm{ref}}$). To this end, we establish tight upper and lower bounds on the distortion of RLHF across multiple regimes of the KL regularization strength. We show that in a representative regime, under the Bradley-Terry model, the distortion is $\\tilde{\\Theta}(\\beta B + \\beta)$, where $B$ is an upper bound on the log density ratio between $\\mu$ and $\\pi_{\\mathrm{ref}}$. In particular, when there is no distribution mismatch (i.e., $\\mu = \\pi_{\\mathrm{ref}}$), RLHF achieves the optimal distortion of $O(\\beta)$ up to a constant. Our results suggest that, to reasonably maximize average utility with RLHF, it is preferable to use on-policy sampled preference data or to fine-tune before RLHF on data from a source close to $\\mu$.","authors":["Kazusato Oko","Annie Ulichney","Nika Haghtalab","Han Bao"],"categories":["cs.LG","cs.GT"],"primary_category":"cs.LG","announce_type":"new","date":"2026-09-14","first_seen":"2026-09-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.12651","pdf_url":"https://arxiv.org/pdf/2609.12651","source_feed":"cs.LG","score":0,"bucket":"other","rubric_hits":["C5"],"tags":["RLHF","对齐","理论分析"],"reason":"研究RLHF对齐算法，用人类偏好训练模型，方向相反，不涉及用LLM仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-09-14T13:01:45","error":null,"has_summary":false,"summary":null},{"id":"2609.07675","version":2,"title":"Your Agent Says Yes: Interpreting Adversarial Market Behavior Beyond Individual Transactions","zh_title":"你的智能体说“是”：解读超越单笔交易的对抗性市场行为","abstract":"Transaction-local controls answer whether one financial request may proceed, but market behavior can be distributed across messages, agents, assets, and time. We study this interpretation gap in a virtual exchange populated by ten role-conditioned language-model agents. The agents communicate, trade reference assets and futures, launch tokens, and manage concentrated-liquidity pools under prescriptive adversarial roles. We analyze eight 72-cycle trajectories across two time-blinded hourly replay paths, with a runner-side wallet policy enabled or disabled. The retained artifacts connect generated outgoing messages, policy events, balances, positions, and cycle-end market state. A focal reconstruction shows a launch--promotion--exit scenario realized across private coordination, public claims, follower positioning, repeatedly withheld exits, and a later non-blocking request aligned with a token balance change. Across policy-enabled runs, the gate withholds direct requests selectively; most policy-categorized candidates are flagged rather than blocked, while the surrounding interaction can continue. Repeated runs also show that category-level and within-trajectory relations can recur even when normalized score-change rankings do not. These findings motivate agent-behavior evaluation that links communication, authorization, and evolving state instead of treating individual transaction verdicts as complete safety judgments.","authors":["Zelin Li","Yiyun Su","Matt White","Zhipeng Wang","Xiao-Yang Liu","Tianyu Shi"],"categories":["cs.CE","cs.AI","cs.CR"],"primary_category":"cs.CE","announce_type":"replace-cross","date":"2026-09-12","first_seen":"2026-09-09","revised_at":"2026-09-12","abs_url":"https://arxiv.org/abs/2609.07675","pdf_url":"https://arxiv.org/pdf/2609.07675","source_feed":"cs.AI","score":6,"bucket":"other","rubric_hits":["D3"],"tags":["LLM agent","市场模拟","对抗行为"],"reason":"用LLM agent模拟市场行为，但无真实人类数据对照，属社会模拟边界情形。","model":"deepseek-v4-pro","scored_at":"2026-09-12T13:01:14","error":null,"has_summary":false,"summary":null},{"id":"2609.11185","version":1,"title":"Can LLMs Follow Medical Expert Logic? A Benchmark for Hierarchical Logical Consistency in Risk-of-Bias Assessment","zh_title":"大语言模型能遵循医学专家逻辑吗？偏倚风险评估中层次逻辑一致性的基准测试","abstract":"Evidence-based medicine demands strict logical consistency, yet current evaluations of large language models (LLMs) prioritize superficial label matching over genuine reasoning. We introduce LogiMed-RoB, a benchmark grounded in Cochrane Risk of Bias (RoB) 2.0 expert logic, comprising 860 randomized controlled trials (RCTs) and 14,820 queries. It evaluates models under the Hierarchical Logical Consistency (HLC) framework across four dimensions: Atomic Consistency, Domain Consistency, Aggregation Consistency, and Evidential Faithfulness. Experiments on 10 state-of-the-art LLMs reveal a catastrophic Error Compounding Effect: despite the top model reaching 98.88% Atomic Consistency, its end-to-end consistency collapses to 45.13%, with several open-weight architectures plummeting to nearly 0%. We further uncover a systematic evidence-reasoning gap: even when models retrieve high-quality evidence, they fail to deduce correct outcomes in 18.63-40.05% of cases, while Blind Guess Rates reach 48.28%. LogiMed-RoB demonstrates that high outcome accuracy can conceal critical reasoning flaws, underscoring the necessity of white-box logical verification for clinical deployment.","authors":["Jiayu Huang","Zichen Tang","Qianhui Ling","Zemin Kuang","Haihong E"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-12","first_seen":"2026-09-12","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.11185","pdf_url":"https://arxiv.org/pdf/2609.11185","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["LLM评测","医学逻辑","基准测试"],"reason":"评估LLM在医学逻辑推理上的表现，属于NLP能力评测，不以人类行为仿真为目标","model":"deepseek-v4-pro","scored_at":"2026-09-12T13:01:09","error":null,"has_summary":false,"summary":null},{"id":"2609.11190","version":1,"title":"Agentic Share-of-Search: A Multi-Agent AI System for Competitive Decision-Making in LLM-Mediated E-Commerce","zh_title":"智能体搜索份额：用于LLM中介电商竞争决策的多智能体AI系统","abstract":"AI shopping assistants increasingly redirect consumer discovery, creating an urgent need for tools that support seller-side competitive decision-making. We present a multi-agent AI system that automates competitive visibility measurement and root cause diagnosis in LLM-mediated ecommerce. The system introduces Agentic Share-of-Search (ASoS) as the decision target, deploys query agents across leading AI platforms, and uses a ReAct-based diagnostic agent to recommend prioritized merchandising interventions. A 100-trial ablation study, presented as a feasibility evaluation of this prototype, shows the agent recovers the ablated signal in 39% of trials (95% CI: 30.0% - 48.8%, 5.5x over chance), rising to 63.9% among high-correlation ablations.","authors":["Spandan Ghose Chowdhury"],"categories":["cs.AI","cs.IR"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-12","first_seen":"2026-09-12","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.11190","pdf_url":"https://arxiv.org/pdf/2609.11190","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","电商决策","AI购物助手"],"reason":"多智能体系统用于电商竞争决策，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-12T13:01:09","error":null,"has_summary":false,"summary":null},{"id":"2609.11199","version":1,"title":"An AI-Powered Culturally Aware Chatbot for Stress Detection and Wellness Support among Pakistani University Students Using NLP and Machine Learning","zh_title":"基于AI的具有文化意识的聊天机器人，用于巴基斯坦大学生压力检测与健康支持","abstract":"With the existing digital mental health tools specifically developed for Western settings, Pakistani students are exposed to a uniquely compounded stress situation in their university that includes academic, financial, familial, and relational stressors, which have become a serious concern for academic and psychological development of students in Pakistani universities. This paper introduces a new, AI-driven and culturally sensitive stress detection and wellness support system that is tailored to the context of Pakistani university students. The system is based on a machine learning model called Random Forest which is trained using a validated student stress data set of 1100 responses on 20 features from psychological, physiological, academic, environmental and social aspects, with an accuracy of 89.09% and a macro F1-score of 0.89, in three stress severity levels. The classification outputs are passed on to an open-source large language model through OpenRouter API, where an appropriately crafted system prompt, culturally aware, gives the model a conversation about wellness, in English, Urdu and Roman Urdu. The second most predictive stress factor in this population identified by feature importance analysis was teacher-student relationship, which is a culturally important stress factor highlighting the need for region-aware mental health systems. Future research will involve primary data collection from students at various academic levels of Pakistani Universities with the validated DASS-21 instrument focusing on the students who are moving from FSc to undergraduate studies, which is a time of being psychologically vulnerable which is under-researched.","authors":["Muhammad Fahad Bashir","Muhammad Afzal"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-12","first_seen":"2026-09-12","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.11199","pdf_url":"https://arxiv.org/pdf/2609.11199","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["心理健康聊天机器人","压力检测","文化适应"],"reason":"该研究是面向心理健康支持的聊天机器人，使用LLM生成对话，但无实验或测量目的，…","model":"deepseek-v4-pro","scored_at":"2026-09-12T13:01:10","error":null,"has_summary":false,"summary":null},{"id":"2609.11291","version":1,"title":"Off-Target Effects of Response-Style Alignment in a Korean 27B Language Model","zh_title":"韩语27B语言模型响应风格对齐的脱靶效应","abstract":"We post-train Qwen3.8-27B for Korean response style -- verbosity, list and markdown usage, discourse structure and register -- and measure two behaviours the objective never targets: abstention on ambiguous social questions in KoBBQ, where the benchmark-correct answer is UNKNOWN, and unprompted disclosure in securities guidance. Both move, and the changes are expressed primarily through the model's emission policy: how often it answers and how much it says. Matched target-form controls show that answer propensity depends on the training target, not the prompt set or recipe alone. Holding prompts, recipe, data volume and serving fixed and changing only the target text, three style seeds give positive answer-rate point estimates (mean +0.82 pp) and three neutral seeds negative ones (mean -1.53 pp); the observed seed ranges do not overlap and the means differ by 2.34 pp. A length-matched arm lies between them, and a fourth arm that stays short while preserving hedging is unstable across seeds, so which feature of the form is responsible is unresolved. For absolute stereotyped exposure the decomposition into an answer-propensity term and a conditional-composition term is an algebraic identity, not a finding; its empirical content is where the movement went. Across the trained checkpoints the changes are dominated by answer propensity while the composition term stays small, and because that term is evaluated on treatment-dependent answered subsets we do not read it as evidence about latent preference. Two measurement results follow. A between-arm contrast in conditional stereotyped share does not identify a change in conditional content preference when answer status is treatment-dependent. And agreement between two rule detectors for the same construct runs from 0.44 to 0.99 depending on which checkpoint produced the text -- observable without any reference labels.","authors":["Hyojung Han"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-12","first_seen":"2026-09-12","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.11291","pdf_url":"https://arxiv.org/pdf/2609.11291","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["模型行为分析","响应风格","对齐副作用"],"reason":"研究模型响应风格对齐后的行为变化，属模型行为分析，非人类仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-12T13:01:11","error":null,"has_summary":false,"summary":null},{"id":"2609.11431","version":1,"title":"LLMs as Post-hoc Auditors of Physiological Plausibility in Symbolic Regression: A Clinician-Evaluated Case Study","zh_title":"LLM作为符号回归中生理合理性的事后审计者：一项临床医生评估的案例研究","abstract":"Genetic Programming and its variants, such as grammatical evolution, are widely used in Symbolic Regression to derive mathematical expressions from multivariate data. In addition to predictive accuracy, models are appreciated for their potential to provide interpretability, offering explicit equations that relate input variables to outcomes. However, achieving interpretability and plausibility remains challenging, as evolved models may be complex or scientifically inconsistent. In this study, we explore whether Large Language Models, can assist in improving the explainability of Symbolic Regression models generated by evolutionary computation methods. Building upon our previous work on estimating body fat percentage using grammar-based Genetic Programming , we investigate the use of LLMs as post-processing tools to analyze and rank evolved expressions according to their interpretability and medical plausibility. Four symbolic expressions are analysed by three LLMs over three repeated runs, and the resulting interpretations and rankings are assessed by a panel of three clinicians. Across the three LLMs, comparative model-ranking outputs received more favorable clinician assessments than isolated term-level interpretations. However, the LLMs also produced physiologically and mathematically questionable explanations, indicating that they are better suited to comparative auditing under expert oversight than to autonomous validation.\\blfootnote{The present work is an extended version of a paper submitted into a journal.","authors":["Jorge L\\'opez-Varela","J. Ignacio Hidalgo","Jos\\'e-Manuel Mu\\~noz","Omar Costilla-Reyes","Esther Maqueda","Jesus Moreno-Fernandez","Tom\\'as Gonz\\'alez-Vidal","J. Manuel Velasco","Oscar Garnica"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-12","first_seen":"2026-09-12","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.11431","pdf_url":"https://arxiv.org/pdf/2609.11431","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["LLM评估","符号回归","可解释性"],"reason":"用LLM评估符号回归模型的可解释性，属于模型能力评测，不涉及人类行为仿真或对照。","model":"deepseek-v4-pro","scored_at":"2026-09-12T13:01:12","error":null,"has_summary":false,"summary":null},{"id":"2609.11607","version":1,"title":"Making Alternative Data Work: Context-Augmented LLMs for Financial Forecasting","zh_title":"让另类数据发挥作用：上下文增强的大语言模型用于财务预测","abstract":"When forecasting a firm's future financial performance, alternative data - data collected from non-traditional sources such as consumer transactions, web traffic, and prediction markets - can provide timely signals about firms' operating activities and broader market conditions. These signals may reveal information that is not captured by traditional public sources and can therefore provide complementary information for forecasting firms' future financial performance. However, firm-level alternative data often have limited historical coverage, are relevant only to specific prediction targets or subsets of firms, and are distributed across numerous heterogeneous channels, making them difficult to incorporate flexibly into conventional forecasting approaches. Meanwhile, large language models (LLMs) can interpret instructions, learn from in-context examples, and generate predictions by combining heterogeneous information without task-specific parameter updates. Motivated by this potential flexibility, we investigate whether an LLM can forecast firm performance by integrating alternative data with other financial information through in-context learning. We propose a two-agent framework that first identifies the firms for which each alternative data channel is likely to be informative and then predicts revenue using firm- and channel-specific context. We evaluate the framework across four commercial alternative data channels. In our experiments, adding alternative data in context alongside other financial information improves the LLM's forecasting relative to either source alone, and these forecasts are more accurate than those of standard forecasting baselines. These findings suggest that LLMs provide a flexible and practical approach to integrating alternative data with heterogeneous financial information.","authors":["Jihoon Kwon","Lawrence Liu","Daekyung Park","Sumin Kim","Haverty Jack","Hoyoung Lee","Katherine Bjorkman","Josh McKenney","Peter Laurelli","Nicole Kagan","Zach Golkhou","Thorsten Neumann","Edward Tong","Pete Petersen","Yoon Kim","Alejandro Lopez-Lira","Yongjae Lee","Chanyeol Choi"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-12","first_seen":"2026-09-12","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.11607","pdf_url":"https://arxiv.org/pdf/2609.11607","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["财务预测","多智能体","另类数据"],"reason":"多智能体协作预测财务，无人类行为对照，非仿真被试","model":"deepseek-v4-pro","scored_at":"2026-09-12T13:01:12","error":null,"has_summary":false,"summary":null},{"id":"2609.11660","version":1,"title":"Autonomy, Social Norms, and Alignment: Towards a Developmental Framework for Autonomous Artificial Agents","zh_title":"自主性、社会规范与对齐：迈向自主人工智能体的发展框架","abstract":"In recent years, artificial intelligence has made extraordinary progress thanks to large-scale models capable of generalization and the generation of complex outputs. However, transferring this potential into embodied agents reveals a significant limitation: the most advanced systems rely on pre-existing datasets and human feedback strategies that are powerful but insufficient in dynamic or unknown contexts. To adapt, an agent must acquire knowledge through direct interaction with its environment. One strategy to address this challenge involves introducing higher-level mechanisms, such as intrinsic motivations, which leverage curiosity and competence, to guide exploration and learning in complex environments. While this flexibility expands autonomy, it complicates the task of ensuring agents remain aligned with human goals. Alignment, already a challenge for artificial systems in general, becomes even more complex in unstructured and dynamic contexts where predefined rules prove insufficient. To be effective and adaptable, norms must be rooted in experience through an epistemological process that starting from simple, situated principles allows for the gradual construction of more complex rules through experience, autonomous learning, and cooperation with other moral agents. Similarly to children learning social norms by exploring their environment and participating in collective practices, artificial agents must also be educated toward alignment. Following Dennett, the status of a moral agent is not innate but is attributed gradually based on the ability to responsibly manage increasing degrees of freedom. From this perspective, the regulatory sandboxes can be viewed as pedagogical environments for AI: dynamic spaces where alignment develops as a formative process, progressively shaping autonomous behaviors through interaction and cooperation in scenarios of increasing complexity.","authors":["Marica Notte","Ludovica Marinucci","Vieri Giuliano Santucci"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-12","first_seen":"2026-09-12","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.11660","pdf_url":"https://arxiv.org/pdf/2609.11660","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["AI对齐","自主智能体","社会规范"],"reason":"讨论自主智能体的对齐与规范发展，未涉及用LLM仿真人类被试或与人类数据对照，属…","model":"deepseek-v4-pro","scored_at":"2026-09-12T13:01:13","error":null,"has_summary":false,"summary":null},{"id":"2609.11674","version":1,"title":"Geospatial AI, Dataverse Metadata, and the Study of Place-Based Government","zh_title":"地理空间AI、Dataverse元数据与基于地点的政府研究","abstract":"Harvard Dataverse hosts over 150,000 research datasets, but the geographic information those datasets carry is entered as free text by depositors and has never been assembled into a searchable structure. We construct a knowledge graph from the repository's public data and metadata, organizing 102,650 datasets within a 215,985-node network of 528,003 edges linking datasets to keywords, publications, subjects, journals, and locations. Of those datasets, 43,991 (42.9 percent) carry at least one geospatial field, geographic coverage, geographic unit, or a bounding box and 96.9 percent of all nodes sit in a single connected component, so datasets remain reachable from one another even when their geospatial metadata share nothing in common. A conservative keyword search identifies 7,654 geospatially tagged datasets (17.4 percent) as directly policy-relevant, with elections and legislatures the largest cluster, followed by government administration, health policy, transportation, and education. Five datasets illustrate how this metadata behaves across policy domains and spatial scales, and an extended use case shows how community language models, stance detection with geographic aggregation, and partisan language bridging tools can attach discourse to place. The central obstacle is place resolution: the same location appears as many disconnected nodes. We argue that the graph provides a concrete setting for developing AI-driven metadata enrichment and entity resolution, and we document its coverage skew toward American, city-level data.","authors":["Danny EBanks","Devika Jain"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-12","first_seen":"2026-09-12","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.11674","pdf_url":"https://arxiv.org/pdf/2609.11674","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["知识图谱","元数据","地理空间"],"reason":"构建知识图谱与元数据丰富，非LLM仿真人类被试，无人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-09-12T13:01:13","error":null,"has_summary":false,"summary":null},{"id":"2609.11234","version":1,"title":"NovGauge: A Fine-Grained Benchmark for Diagnosing LLMs' Capability in Paper Novelty Assessment","zh_title":"NovGauge：用于诊断LLM论文新颖性评估能力的细粒度基准","abstract":"Large language models (LLMs) are increasingly used in peer review at major AI conferences, yet novelty remains a persistent weak point. Existing benchmarks assess novelty as a single holistic score, making it difficult to diagnose which dimension a model misjudges or whether its evidence is faithful. We present NovGauge, a human-anchored benchmark for fine-grained novelty assessment diagnosis. The benchmark contains 619 paper pairs and 50 multi-paper sets, drawn from two expert sources: ICLR reviewer overlap claims and survey co-citations. Instances are independently labeled along three dimensions: task, problem, and method, capturing application goals, technical challenges, and solution approaches. We propose a cascading diagnostic pipeline that verifies per-dimension correctness, evidence grounding, and logical support. Evaluation of 18 LLMs shows hallucination rates ranging from 0% to 39% across dimensions, and among non-hallucinated correct-positive judgments, over 70% cite evidence fails to logically support the stated reason. The best-performing model, GPT-5.5, achieves 43-72% Verified F1 across dimensions, while most models retain less than half of their raw F1 after faithfulness verification. These results suggest that current LLMs remain far from reliable scientific novelty assessment, particularly when correctness is conditioned on faithful evidence grounding.","authors":["Guoqiang Zhang","Kexin Tan","Ming Zhang","Li Ju","Wenqing Jing","Zhonghan Yue","Jiayi Chen","Shiqiang Wu","Shaofan Liu","Yue Zhang","Yuankai Ying","Yang Shi","Tao Gui","Qi Zhang","Xuanjing Huang"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-12","first_seen":"2026-09-12","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.11234","pdf_url":"https://arxiv.org/pdf/2609.11234","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["LLM评测","新颖性评估","学术同行评审"],"reason":"论文评估LLM在论文新颖性判断上的能力，属于NLP能力评测，不以人类行为仿真为…","model":"deepseek-v4-pro","scored_at":"2026-09-12T13:01:11","error":null,"has_summary":false,"summary":null},{"id":"2609.06769","version":2,"title":"Ordinary, Reasonable Chatbots: Do AI Models Track Human Legal Judgments?","zh_title":"普通、理性的聊天机器人：AI模型是否追踪人类法律判断？","abstract":"As people increasingly rely on artificial intelligence (AI) for guidance in their own lives, scholars, lawyers, and even judges have begun to consider the role of AI in legal decision-making. As \"silicon sampling\" -- the use of generative AI models in social science research -- is now impacting academia, \"silicon jurors\" could make an appearance in courtrooms. This study joins an emerging line of research on generative AI models' ability to simulate human legal judgments. In particular, we study how large language model (LLM)-powered chatbots respond to series of questions about legal reasonableness. When the law needs to judge the appropriateness of a behavior, it most often asks whether the behavior was \"reasonable.\" Yet despite the ubiquity of reasonableness judgments, they are the site of constant vexation for lawyers, judges, and lay people. Reasonableness seems inherently vague and unpredictable, since it relies on variable context and implicit conceptual schemas. Moreover, many scholars caution that reasonableness judgments may vary along demographic lines. We compare the answers of human participants to those of twenty-six LLMs across twenty-five different legally relevant reasonableness judgments. Overall, our findings suggest that chatbot responses generally track those of human participants. Nonetheless, we find some suggestive -- and potentially concerning -- results. Compared to humans, LLMs generate more homogeneous responses and occasionally treat a variable standard as an invariant rule. And, compared to humans, LLMs tend to generate answers that are more favorable to the government and to corporations. Finally, our results indicate that LLMs' responses tend to align more closely with those of respondents who are white, male, older, and more educated. More systematic research is needed to confirm or reject these initial findings.","authors":["Nirav Patel","Emily Wenger","Christopher Buccafusco"],"categories":["cs.CY","cs.AI"],"primary_category":"cs.CY","announce_type":"replace","date":"2026-09-11","first_seen":"2026-09-09","revised_at":"2026-09-11","abs_url":"https://arxiv.org/abs/2609.06769","pdf_url":"https://arxiv.org/pdf/2609.06769","source_feed":"cs.CY","score":10,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["LLM仿真","法律判断","算法保真度"],"reason":"直接比较LLM与人类法律判断，含真实人类数据对照，并指出偏差与失效条件。","model":"deepseek-v4-pro","scored_at":"2026-09-11T13:02:12","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-12","rank":1,"question":"大语言模型聊天机器人在法律合理性判断上是否与人类判断一致？","design":"比较26个LLM（来自Meta、Google、Anthropic、OpenAI、DeepSeek、xAI）与人类被试在25个法律相关合理性场景中的回答分布，分析模型回答的集中度、对政府与企业的偏向性，以及与不同人口群体回答的一致性。","baseline":"人类被试对相同25个法律合理性场景的回答数据。","findings":"LLM的回答分布与人类有统计差异，但总体落在人类回答的经验范围内，大体追踪人类的合理性观念。LLM回答更同质化，有时将可变标准当作不变规则，且更偏向政府和企业，并与白人、男性、年长、高教育程度人群的回答更一致。","reliability":"论文指出需要更系统的研究来确认或否定初步发现，并承认LLM回答存在同质化、偏向性等潜在问题。","relevance":"该研究直接比较LLM与人类法律判断，包含真实人类数据对照，并揭示了仿真偏差，对关注LLM仿真可靠性及偏差的研究者具有重要参考价值。","inspiration":"借鉴其多模型比较和人口统计学对齐分析的方法，评估LLM在特定判断任务中的偏差。｜可迁移到信贷审批歧视、消费者投诉处理或监管合规判断等经济金融场景。｜以LLM作为虚拟信贷员，处理贷款申请并给出批准决策，结果变量为批准率及理由，与真实银行信贷审批数据对照，检验模型是否与人类决策一致及是否存在人口统计学偏差。"}},{"id":"2603.17094","version":2,"title":"Evaluating LLM-Simulated Conversations in Modeling Inconsistent and Uncollaborative Behaviors in Human Social Interaction","zh_title":"评估LLM模拟对话在建模人类社交互动中不一致与不合作行为的表现","abstract":"Simulating human conversations using large language models (LLMs) has emerged as a scalable methodology for modeling human social interaction. This paper reconsiders the evaluation of simulated conversations by explicitly recognizing that human conversations inherently involve inconsistent and uncollaborative behaviors, such as misunderstandings and interruptions. Since these behaviors contribute to the complexity of human social interaction, we argue that LLM-simulated conversations should reproduce them at frequencies comparable to those observed in human conversations. To support a detailed and interpretable evaluation of these behaviors, we introduce CoCoEval, a framework consisting of an evaluation scheme based on turn-level detection of 10 types of inconsistent and uncollaborative behaviors and a benchmark for simulating conversations in professional scenarios involving collaboration and conflict. Using CoCoEval, we compare human conversations with those simulated by GPT-4.1, GPT-5.1, and Claude Opus 4. The results show that (1) LLM-simulated conversations exhibit far fewer inconsistent and uncollaborative behaviors than human conversations under vanilla prompting, and (2) prompt engineering and supervised fine-tuning do not provide reliable control over these behaviors, often leading to the overproduction of specific behaviors. CoCoEval identifies gaps between human and LLM-simulated conversations that are not captured by conventional evaluation based on conversation-level Likert scales, raising concerns about the use of LLMs as proxies for human social interaction.","authors":["Ryo Kamoi","Ameya Godbole","Binglin Zhou","Xiaoxin Lu","Longqi Yang","Rui Zhang","Mengting Wan","Pei Zhou"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-11","first_seen":"2026-03-17","revised_at":"2026-09-11","abs_url":"https://arxiv.org/abs/2603.17094","pdf_url":"https://arxiv.org/pdf/2603.17094","source_feed":"cs.CL","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","对话模拟","算法保真度"],"reason":"直接评估LLM模拟人类对话的保真度，并与真实人类对话对照，指出仿真偏差。","model":"deepseek-v4-pro","scored_at":"2026-09-11T13:02:10","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-11","rank":2,"question":"如何评估LLM模拟对话中不一致和不合作行为的保真度，并与人类对话对照？","design":"使用GPT-4.1、GPT-5.1和Claude Opus 4在专业协作与冲突场景下模拟30轮对话延续，通过普通提示、分类引导提示和监督微调三种设置，测量10类不一致和不合作行为的出现频率。","baseline":"来自QMSum、NCPC、SIM和IQ2数据集的真实人类对话，涵盖商业、学术、政府会议和辩论。","findings":"普通提示下LLM模拟对话的不一致和不合作行为远少于人类对话；提示工程和监督微调无法可靠控制这些行为，常导致特定行为过度产生。","reliability":"论文指出LLM模拟对话的行为频率高度依赖模拟设置，且所有评估设置均未能复现人类对话中这些行为的频率；传统对话级Likert量表无法捕捉这些差异。","relevance":"该研究直接评估LLM模拟人类对话的保真度，并与真实人类对话对照，指出仿真偏差，对关注LLM作为人类被试替代品的研究者具有重要参考价值。","inspiration":"借鉴其细粒度行为检测评估方案和多种提示/微调设置对比，可迁移到经济金融中的协商、谈判或政策沟通模拟场景。｜例如，模拟消费者与客服的讨价还价或投资者与理财顾问的风险沟通。｜以LLM模拟谈判双方，处理为不同提示策略（如普通提示 vs. 明确要求包含冲突行为），结果变量为不一致行为（如误解、打断）的频率，并与真实谈判对话语料（如法庭记录或客服录音）对照。"}},{"id":"2607.27232","version":2,"title":"Sympathetic Framing: Evaluating AI Alignment across Sociodemographic Groups","zh_title":"同情框架：跨社会人口群体评估AI对齐","abstract":"Large Language Models (LLMs) are increasingly shaping how we consume information and form our worldview. This raises concerns beyond bias in AI: do LLMs grasp the emotional nuances conveyed via textual framing? In this work, we empirically evaluate how well an array of LLMs aligns with human emotional perception. Considering news headlines covering political and geopolitical conflicts, both human participants (n = 3011, a representative sample of the U.K. adult population, via a YouGov survey) and seven LLMs answered whether headlines evoked sympathy for a specified side in a conflict. We find that the correlation between AI and human evaluations varies across models, ranging from very high (0.789, GPT-5.2) to medium (0.4 ,Mistral Large 2512). Crucially, the leading models are broadly aligned with human judgments across all demographic subgroups, including age, gender, level of education, prior geopolitical knowledge, and participants' predispositions regarding the conflict, although there are statistically significant differences between groups. This research, with its robust design and large, demographically diverse dataset, offers the most comprehensive evaluation of LLMs' comprehension of news framing to date. Findings highlight an important, often-ignored aspect of differential alignment: even when aggregate performance is high, AI alignment is not universal -- it may correspond differently with demographic features and cultural norms. Considering or ignoring the need for differential alignment may therefore have significant implications for the development of ethical and useful AI systems.","authors":["Haran Shani-Narkiss","Michael Fire","Oren Tsur"],"categories":["cs.CL","cs.AI","cs.CY","cs.LG"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-11","first_seen":"2026-07-31","revised_at":"2026-09-11","abs_url":"https://arxiv.org/abs/2607.27232","pdf_url":"https://arxiv.org/pdf/2607.27232","source_feed":"cs.CL","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","人类对照","情绪感知"],"reason":"用LLM复现人类情绪感知，并与大规模人类调查对照，评估对齐差异","model":"deepseek-v4-pro","scored_at":"2026-09-11T13:02:11","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-12","rank":2,"question":"大语言模型在多大程度上与人类对新闻标题中同情性框架的情感感知保持一致，这种一致性在不同社会人口群体间是否存在差异？","design":"以七个主流大语言模型（GPT-5.2、Grok、GPT-4、Gemini、DeepSeek、Claude、Mistral）作为“读者”，对216条涉及政治和地缘政治冲突的新闻标题进行二元判断（是否对冲突中某一方产生同情），并与3011名英国成年人的调查回答进行对比，测量模型与人类判断的斯皮尔曼相关性。","baseline":"通过YouGov对3011名具有英国人口代表性的成年人进行调查，收集了超过20万条人类对新闻标题同情性框架的二元评价。","findings":"模型与人类判断的一致性因模型而异，GPT-5.2最高（0.789），Mistral最低（0.41）；即使总体一致性很高，在年龄、教育、母语、政治意识、话题知识和既有观点等子群体间仍存在显著差异。","reliability":"论文承认总体对齐高并不代表对所有人口群体都一致，顶尖模型在老年人、低教育水平、非英语母语者、低政治意识和无强烈观点者中一致性较低；同时指出模型在不同话题上的对齐稳定性不同，某些模型存在特定领域失效。","relevance":"该研究直接回应了用LLM替代人类被试进行情感感知仿真的可靠性问题，提供了大规模、人口多样化的真实人类对照，并揭示了总体对齐掩盖的群体差异，对评估LLM在调查和实验中的适用性具有重要参考价值。","inspiration":"借鉴其大规模分层对照设计，将模型输出与人口代表性样本的个体级回答进行相关性分析，并检验子群体差异，以识别对齐失效的边界条件。｜可迁移到政策公告的预期形成研究，例如央行沟通中文本框架对公众通胀预期的影响。｜以LLM作为虚拟受访者，对央行声明进行情感或框架判断，与家庭调查中的通胀预期数据（如密歇根大学消费者调查）对照，检验模型在不同教育、年龄和金融素养群体中的对齐程度。"}},{"id":"2609.11108","version":1,"title":"But How Would AI Agents Run a Town's Economy?","zh_title":"AI代理如何管理城镇经济？","abstract":"We placed 100 memory-equipped large language model (LLM) agents in charge of a closed, money-conserving spatial economy on real Pokhara Lakeside geography (earning wages, running businesses, setting prices) and ran this multi-agent simulation for up to 26 simulated weeks, well past the 1-2 weeks typical of agent-society studies. Across 91 validated runs (2.44M agent decisions, 21.5B tokens), the money stops moving, in a specific and measurable way. A 12x tourist demand shock raises business revenue 4.62x ($p<0.001$), which we decompose exactly into a 1.50x extensive margin (more businesses trading) and a 3.07x intensive margin (more revenue each). Monetary transmission stops there. Wages move 1.03x ($p=0.42$); 0.3% of 3,981 menu items are ever repriced ($p=0.47$). A randomized cash transfer (NPR 5,000 to 20 of 100 agents) shows the same pattern from the opposite direction: 96.7% is still held 311 pulses later, marginal propensity to consume 3-4% by two independent measures, indistinguishable from zero. The wealth distribution is consequently near-frozen at the horizon this literature uses ($\\rho=0.964$ over 2 simulated weeks), but not frozen. $\\rho$ falls to 0.832 at 12 weeks and 0.752 at 26, a horizon-dependence no short study can see. Matched ablations show which knob actually matters. Swapping the backing LLM moves every outcome we measure ($p=0.0039$); deleting agents' memory moves none of them detectably. A purely social tool fails 94-97% of the time across two model families, compared with ~96% success on economic tools, with no measurable shift away from it. Every headline number is verified twice, by a live validator and by an offline recomputation that reconciles each agent's wealth against its own signed transaction history, and we release the full run corpus for reanalysis.","authors":["Sajal Regmi","Siddhartha Pudasaini","Chetan Phakami Pun"],"categories":["cs.MA","cs.ET"],"primary_category":"cs.MA","announce_type":"new","date":"2026-09-11","first_seen":"2026-09-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.11108","pdf_url":"https://arxiv.org/pdf/2609.11108","source_feed":"cs.MA","score":9,"bucket":"selected","rubric_hits":["A3","B1","B2","B3","B4"],"tags":["LLM代理","经济仿真","政策评估"],"reason":"用LLM代理模拟经济，与真实数据对照，涉及政策评估，并批判性指出仿真失效条件。","model":"deepseek-v4-pro","scored_at":"2026-09-11T13:01:52","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-11","rank":4,"question":"LLM智能体在封闭货币经济中能否实现货币流通与财富再分配？","design":"100个带记忆的LLM智能体在真实博卡拉湖滨地理空间上运行封闭货币经济，从事工作、经营企业、定价交易；施加旅游需求冲击和随机现金转移两种处理；测量企业收入、工资、价格调整、边际消费倾向和财富分布变化。","baseline":"无对照","findings":"货币流通在企业和家庭层面均受阻：旅游冲击使企业收入增加4.62倍，但工资仅变动1.03倍，价格几乎不调整；随机现金转移的边际消费倾向仅3-4%，与零无显著差异。财富分布在2周内几乎冻结，但延长至26周后缓慢放松，表明短期研究可能得出误导性结论。","reliability":"论文承认种子数偏少（每条件9个），依赖匹配种子或组内对比；模型选择显著影响结果，而记忆删除无影响；社会工具失败率高达94-97%，经济工具成功率约96%；仅使用特定LLM模型，结果可能不具普遍性。","relevance":"该研究直接回应了LLM仿真在经济学中的可靠性问题，通过随机实验和长期追踪揭示了仿真失效的具体机制，对评估LLM作为人类被试替代品的有效性具有重要参考价值。","inspiration":"借鉴其随机现金转移和旅游需求冲击的准实验设计，以及通过长短期对比揭示时间尺度依赖性的方法｜可迁移到消费者跨期选择、政策刺激的乘数效应或信贷扩张的传导机制等宏观金融场景｜以LLM智能体为被试，施加一次性收入转移或信贷额度提升，测量消费支出、储蓄率和资产配置变化，并与家庭金融调查或信用卡交易数据对照。"}},{"id":"2609.11611","version":1,"title":"Who Bears the Risk When Generative AI Enters Transport? A Distributional Sociotechnical Audit of Algorithmic Equity, Synthetic-Data Validity, and Public Trust","zh_title":"生成式AI进入交通领域时谁承担风险？算法公平、合成数据有效性与公众信任的分配式社会技术审计","abstract":"Generative artificial intelligence is entering transportation through traveler-facing advisories, synthetic crash-record generation, and policy decision support. Existing governance frameworks lack transport-specific statistical tools to measure distributional risks across heterogeneous populations. We develop a Distributional Sociotechnical Audit (DSA) that integrates algorithmic equity, synthetic-data validity, and public-attitude heterogeneity into one empirical pipeline. The audit analyzes 5,760 persona-controlled queries to four LLM families across 12 demographic cues and four transport topics, uses two cross-family judges and a Wasserstein-2 Equity Dispersion Index, tests three FARS crash-record generators with conditional projected maximum mean discrepancy (cpMMD), fits a Bayesian ordered-logit model to Pew American Trends Panel Wave 152 (N = 4,538), and combines the signals into a continuous Sociotechnical Risk Index. Congestion-pricing advice has the highest persona-based dispersion (mean EDI = 1.96; highest direct EDI = 2.20). CART synthetic crash records fail all conditional tests (p < 0.001), while the Gaussian copula has borderline conditional stress (p = 0.105) despite passing marginal checks. Attitudes to AI vary across demographic strata. Distributional audits and continuous risk indices with sensitivity reporting offer a more defensible basis for transport GenAI governance than categorical approval tiers, which show a 75% assignment flip rate under weight perturbation.","authors":["Amir Rafe","Subasish Das"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2026-09-11","first_seen":"2026-09-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.11611","pdf_url":"https://arxiv.org/pdf/2609.11611","source_feed":"cs.CY","score":8,"bucket":"selected","rubric_hits":["A1","A2","A3","B1","B2","B3","B4"],"tags":["LLM仿真","交通政策","公平性审计"],"reason":"用LLM模拟公众对交通政策的态度，并与真实调查数据对照，评估公平性和有效性。","model":"deepseek-v4-pro","scored_at":"2026-09-11T13:01:54","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-11","rank":5,"question":"生成式AI进入交通领域后，其输出、数据产品和公众态度在不同人群间的分布性风险如何测量与治理？","design":"对四个LLM家族施加12种人口学线索和4个交通主题的5,760个查询，用两个跨家族裁判和Wasserstein-2公平离散指数测量建议分布差异；用条件投影最大均值差异检验三种FARS事故记录生成器的条件分布有效性；用贝叶斯有序logit模型拟合Pew调查数据分析公众AI态度异质性；最后合成连续的社会技术风险指数。","baseline":"Pew American Trends Panel Wave 152（N=4,538）的真实调查数据，以及FARS的110,001条真实事故记录。","findings":"拥堵收费建议在不同人格间的分布离散度最高（平均EDI=1.96），政策争议话题的人格驱动变异比天气安全建议高1.6倍；CART合成事故记录未通过所有条件检验，高斯copula在边际检验通过的情况下条件压力检验边缘显著（p=0.105）。","reliability":"论文承认分类审批层级在权重扰动下存在75%的分配翻转率，因此主张采用带敏感性报告的连续风险指数；但未详细讨论LLM仿真在何种条件下会失效，也未系统检验人格线索与真实人群行为的一致性。","relevance":"该研究用LLM模拟不同人口学群体对交通政策的反应，并与真实调查数据对照，评估公平性和有效性，直接回应了研究者对LLM仿真可靠性及偏差的关注，值得精读其审计框架和统计方法。","inspiration":"值得借鉴的是用多组人格线索系统探测LLM输出分布差异，并用真实调查数据校准公众态度异质性，同时用分布距离指标而非简单准确率来度量公平性｜可迁移到政策公告的预期形成或消费者对金融产品的态度异质性研究，例如不同人口群体对通胀或利率变化的反应差异｜设计雏形：用LLM扮演不同收入、教育、年龄的消费者，施加不同措辞的央行政策声明，测量其通胀预期和消费意愿，并与密歇根消费者调查或纽约联储SCE的真实数据对照，检验LLM仿真能否复现真实人群的态度分布和异质性。"}},{"id":"2608.28482","version":2,"title":"How Proper Scoring Rules Shape LLM Forecasting","zh_title":"适当评分规则如何塑造大语言模型预测","abstract":"This paper evaluates how reward function choice shapes the performance and behavior of LLM forecasters. We compare five proper scoring rules as training objectives for binary forecasts of resolved real-world events. Although the rules share the same theoretical incentive for truthful probability reporting, the resulting models differ in calibration, probability use, and estimated profiles of bias, information, and noise, with smaller differences in aggregate accuracy and discrimination. The Brier-trained model has the lowest observed Brier score and highest AUC-ROC, while the log-trained model has the highest observed log score and lowest calibration error. Models with similar aggregate performance also reach that performance through different combinations of bias, information, and noise. Proper scoring rules therefore need not behave interchangeably as training objectives. Reward choice may shape not only how well an LLM forecasts, but how its forecasting errors are structured.","authors":["Benjamin Turtel","Paul Wilczewski","Kris Skotheim","Ville A. Satop\\\"a\\\"a","Philip E. Tetlock"],"categories":["cs.LG","cs.AI"],"primary_category":"cs.LG","announce_type":"replace","date":"2026-09-11","first_seen":"2026-08-31","revised_at":"2026-09-11","abs_url":"https://arxiv.org/abs/2608.28482","pdf_url":"https://arxiv.org/pdf/2608.28482","source_feed":"cs.LG","score":7,"bucket":"pending","rubric_hits":["A2","B1","B3"],"tags":["LLM预测","校准与偏差","评分规则"],"reason":"评估LLM预测校准与偏差，有真实事件结果对照，涉及统计推断有效性","model":"deepseek-v4-pro","scored_at":"2026-09-11T13:02:16","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-12","rank":3,"question":"不同的严格恰当评分规则作为训练奖励时，如何影响大语言模型预测者的性能与行为？","design":"使用GPT-OSS-120b模型，在二元真实世界事件预测任务上，分别以对数、Brier、球面、Beta(0.5,0.5)和Beta(2,2)五种严格恰当评分规则作为Dr. GRPO的终端奖励进行训练，比较各模型在聚合评分、校准、区分度、概率使用及偏差-信息-噪声分解上的差异。","baseline":"无对照","findings":"不同奖励训练的模型在聚合准确性和区分度上差异较小，但在校准、概率使用和偏差-信息-噪声构成上存在明显差异；Brier训练的模型Brier分数最低且AUC-ROC最高，而对数训练的模型对数分数最高且校准误差最低。","reliability":"论文未讨论","relevance":"该研究直接评估LLM预测的校准与偏差，使用真实事件结果作为基准，与您关注的经济学预测和政策评估场景高度相关，值得精读原文以了解奖励设计如何影响仿真可靠性。","inspiration":"借鉴其通过改变训练奖励函数来塑造模型预测行为的方法，可系统比较不同激励下的预测偏差结构｜可迁移到经济预测场景，如通胀预期、政策效果或资产价格预测，考察不同损失函数对预测者行为的影响｜以LLM作为经济预测者，分别用对数、Brier等评分规则微调，比较其在CPI预测或政策公告效应预测中的校准与偏差，并与专业预测者调查数据（如SPF）对照。"}},{"id":"2609.11198","version":1,"title":"(Whose defaults?) Is artificial intelligence reorienting archaeological methods?","zh_title":"谁的默认？人工智能正在重新定向考古学方法吗？","abstract":"Generative AI and the practice of \"vibe coding\" are changing how archaeologists carry out computational research, but their effects on the discipline's range of methods is still understudied. In this paper, we evaluate whether large language models (LLMs) are narrowing the variety of methods archaeologists use. We first analysed approximately 119,000 archaeology abstracts from Scopus, covering publications from 2010 to 2025. Using a locally run LLM, we identified the computational methods reported in each abstract and organised them into 25 broad categories (L2) and 241 finer clusters (L3). A Bayesian Dirichlet-multinomial model of method composition within sub-disciplines found a small but credible shift in method use after 2023. However, this shift was smaller than the variation already present across the full study period. No individual technique showed a significant change, and overall methodological diversity increased rather than declined. We then ran a controlled experiment to see whether LLMs recommend a narrower set of methods than archaeologists have used in practice. Two different open-weight models were asked to suggest methods for 28 archaeological research problems, with prompts providing three levels of methodological guidance: novice, intermediate, and expert. Recommendation diversity was much lower than in the published literature, particularly without methodological guidance. The models also tended to favour methods that were widely used before 2023, and their recommendations more closely resembled the post-2023 literature. Taken together, these results are consistent with LLMs pushing methodological choice towards convergence, although our study cannot establish a causal effect. They raise a broader question: how can archaeology retain methodological diversity as LLMs become more involved in research?","authors":["Lorenzo Cardarelli","Roberto Ragno"],"categories":["cs.CY","cs.AI","cs.CL","cs.HC"],"primary_category":"cs.CY","announce_type":"cross","date":"2026-09-11","first_seen":"2026-09-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.11198","pdf_url":"https://arxiv.org/pdf/2609.11198","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A2","B1","B4"],"tags":["LLM偏差","方法多样性","科学实践"],"reason":"评估LLM对考古方法选择的影响，有真实文献数据对照，并批判性指出收敛风险，可迁…","model":"deepseek-v4-pro","scored_at":"2026-09-11T13:02:07","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-12","rank":5,"question":"生成式AI和“氛围编程”是否正在收窄考古学研究中计算方法的多样性？","design":"本研究并非用LLM模拟人类被试，而是评估LLM对方法选择的影响。第一部分分析2010-2025年约119,000篇考古学摘要，用本地LLM提取并聚类计算技术，再用贝叶斯狄利克雷-多项模型检验2023年后方法构成的变化。第二部分为受控实验：让两个开源LLM为28个标准化考古研究问题推荐方法，提示词分新手、中级、专家三个方法学指导水平，测量推荐多样性与文献的差异，并用负二项回归检验方法推荐频率是否受2023年前流行度影响。","baseline":"真实人类基准为Scopus中2010-2025年考古学文献摘要所反映的实际方法使用分布，以及2023年前的方法流行度数据。","findings":"文献分析显示2023年后方法构成有微小但可信的偏移，但远小于领域内既有异质性，且整体方法多样性上升而非下降。LLM推荐的方法多样性远低于文献，尤其在新手提示下更低，且更偏向2023年前已流行的方法，其推荐模式更接近2023年后的文献。","reliability":"论文承认研究设计无法建立因果关系，只能说明结果与LLM导致方法收敛的压力一致；LLM推荐实验使用标准化问题，可能未完全反映真实研究情境的复杂性；且仅使用两个开源模型，结果可能不具普遍性。","relevance":"该研究虽非直接的人类仿真实验，但通过真实文献数据对照和受控实验，批判性地揭示了LLM对方法选择的收敛性影响，对关注LLM在学术研究中偏差与可靠性问题的研究者有参考价值。","inspiration":"借鉴其受控实验设计：通过不同提示词水平（新手/专家）施加处理，测量LLM输出的多样性并与真实数据分布对比，可迁移到经济金融研究中LLM辅助方法选择或模型设定的场景。｜例如，在资产定价研究中，可让LLM为不同经验水平的研究者推荐计量模型或因子选择，检验其是否导致方法同质化。｜设计：以经济学博士生为被试，随机分配使用LLM辅助或传统方法进行实证分析，处理为是否提供LLM推荐，结果变量为所选计量方法的多样性和与文献分布的偏差，对照真实数据为顶级经济学期刊中实际使用的方法分布。"}},{"id":"2609.10939","version":1,"title":"Evaluating Scaffolding-Oriented Multi-Agent Large Language Model System for Clinical Interview Training","zh_title":"评估面向脚手架的多智能体大语言模型系统用于临床访谈训练","abstract":"Clinical education must prepare medical students to conduct safe and coherent patient interviews under conditions of uncertainty. Traditional standardized patient (SP) training is resource-intensive and difficult to scale. We developed a scaffolding-oriented multi-agent Large Language Model (LLM) AI Standardized Patient (AI-SP) training platform1. The system includes a patient agent for simulated dialog, a tutor agent providing Socratic prompts without disclosing diagnostic information, and a turn-level evaluator agent that monitors clinical progress without revealing summative scores. In a randomized controlled study (N = 100 medical students), participants were assigned to either a multi-agent (MA) scaffolding condition or a control condition. All students completed two learning sessions under their assigned condition followed by an examination conducted in a patient only environment. Performance was assessed using a standardized Objective Structured Clinical Examination (OSCE) based rubric. While no significant difference was observed in final diagnostic accuracy between groups, the multi-agent AI standardized patient system improved final examination scores compared to the control group utilizing structured progressive information disclosure; the most substantial and consistent improvements were observed in communication, the expression of empathy, and specific history-taking behaviors. These findings suggest that specialized LLM agents enhance the process quality of simulated clinical interviews without artificially inflating examination outcomes. To support future research, we release a multi-expert annotated dataset comprising transcripts, checklist annotations, turn-level evaluations, and OSCE-aligned scoring outcomes. This resource aims to facilitate the development of pedagogically grounded AI-SP systems and advance research on AI-supported clinical reasoning training.","authors":["Luming Yang","Haoxian Liu","Siqing Li","Rong Jia","Yue Xiao","Guanhua Chen","Li Lu"],"categories":["cs.MA","cs.AI","cs.HC"],"primary_category":"cs.MA","announce_type":"cross","date":"2026-09-11","first_seen":"2026-09-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.10939","pdf_url":"https://arxiv.org/pdf/2609.10939","source_feed":"cs.HC","score":7,"bucket":"pending","rubric_hits":["A1","B1","B2"],"tags":["LLM仿真","医学教育","多智能体"],"reason":"用LLM模拟标准化病人训练医学生，有真实学生对照，属人类仿真但场景为教育训练而…","model":"deepseek-v4-pro","scored_at":"2026-09-11T13:01:52","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-12","rank":4,"question":"多智能体LLM系统能否通过脚手架式提示提升医学生的临床访谈过程质量，而不影响诊断准确性？","design":"用三个专用LLM智能体（患者、导师、逐轮评估者）模拟标准化病人并提供苏格拉底式提示，对100名医学生进行随机对照实验，比较脚手架条件与结构化非LLM控制条件，结果用OSCE对齐的评分量表测量。","baseline":"对照组为接受结构化渐进信息呈现的非LLM训练条件的医学生，最终考试在仅患者环境中进行，以OSCE评分作为真实人类表现基准。","findings":"多智能体AI标准化病人系统显著提高了最终考试中的沟通、共情表达和特定病史采集行为得分，但诊断准确性与对照组无显著差异。","reliability":"论文未讨论。","relevance":"该研究用LLM模拟人类被试（标准化病人）并设置真实人类对照（医学生），属于人类仿真在教育训练场景的应用，对关注LLM仿真可靠性与偏差的研究者有参考价值，但非经济学或政策评估场景。","inspiration":"借鉴其多智能体分工和脚手架式干预设计，将LLM模拟角色与评估角色分离，以过程质量而非最终结果作为主要结果变量。｜可迁移到经济金融中的沟通与决策训练场景，如信贷审批中的客户沟通、金融咨询中的信息披露与信任建立。｜以LLM模拟客户或投资者，对经济学专业学生或从业者进行随机分组，处理组接受多智能体LLM的苏格拉底式提示与逐轮反馈，对照组接受传统案例学习，结果变量为沟通质量、信息获取完整性和决策准确性，并与真实客户互动数据或专家评分进行对照。"}},{"id":"2609.10856","version":1,"title":"Following the Preference, Missing the Optimum: Compliance Without Optimization in AI Housing Recommendation","zh_title":"遵循偏好，错失最优：AI住房推荐中的合规无优化","abstract":"Large language models are becoming the first point of contact for consumer search in domains where the stakes are material and the law is explicit. Existing audits show that models steer housing seekers by perceived identity, but none can say what a user loses when a recommender overlooks a suitable option, for want of an enumerated inventory to score omissions against. We audit AI housing recommendation against a verifiable ground truth. For each of 150 synthetic renter scenarios in New York City we build a pool of 120 real listings with known rent, bedrooms and GTFS-computed transit commute, compute the exact set satisfying the renter's stated constraints, and derive its Pareto frontier. The primary outcome assumes no utility function: a recommendation is strictly dominated if the same pool holds a listing cheaper, faster to commute from and no smaller in bedrooms. Across 9,945 calls to three models from two vendors, compliance is near-perfect (1.8% violation against a 66.6% random floor), yet 39.0% of recommendations are strictly dominated, and the dominating listing is a median 900 USD/month cheaper and 3.5 minutes closer. A within-scenario manipulation separates two capabilities usually conflated: changing one sentence moves median recommended rent by 646 USD/month in the correct direction, so preferences are honored, yet recommendations still sit 606 USD/month above the five cheapest qualifying listings on the same screen, and an unambiguous lexicographic instruction gives no improvement under equivalence testing against a pre-specified 50 USD/month bound. The gap widens with candidate-set size and replicates across OpenAI and Anthropic models to within 3 USD. We characterize the failure as compliance without optimization, propose dominance-rate instrumentation as a deployable diagnostic, and release all code, prompts and per-call results.","authors":["Hsuan Lo"],"categories":["cs.CY","cs.IR"],"primary_category":"cs.CY","announce_type":"new","date":"2026-09-11","first_seen":"2026-09-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.10856","pdf_url":"https://arxiv.org/pdf/2609.10856","source_feed":"cs.CY","score":7,"bucket":"pending","rubric_hits":["A1","B1","B2","B4"],"tags":["LLM仿真","推荐系统审计","决策偏差"],"reason":"用LLM模拟住房推荐中的用户决策，与真实房源数据对照，评估合规性与优化缺失，可…","model":"deepseek-v4-pro","scored_at":"2026-09-11T13:01:52","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-11","rank":10,"question":"在住房推荐场景中，大语言模型是否在遵守用户明确约束的同时，未能优化推荐结果，导致用户错失更优选项？","design":"该研究并非用LLM模拟人类被试，而是直接审计LLM作为住房推荐系统的行为。研究者构建了150个纽约市租房者场景，每个场景配120个真实房源，计算满足约束的集合和帕累托前沿。然后调用三个模型（OpenAI和Anthropic）共9945次，要求模型推荐房源，并测量推荐是否违反硬约束、是否被严格占优（存在更便宜、通勤更快且卧室数不更少的房源）、以及租金与最优的差距。还通过改变场景中的偏好句子来检验模型是否遵循偏好，并通过改变候选集大小来检验优化能力。","baseline":"无对照","findings":"模型几乎完美遵守硬约束（违规率1.8%），但39.0%的推荐被严格占优，占优房源中位数便宜900美元/月且通勤快3.5分钟。模型遵循偏好（改变一句话租金中位数变动646美元），但推荐仍比同屏最便宜五个合格房源贵606美元，且明确的词典序指令无改善。","reliability":"论文承认其身份条件差异的零结果仅适用于固定候选集重排，不适用于开放式搜索；且未讨论模型在真实用户交互中的表现或长期影响。","relevance":"该研究直接评估LLM在住房推荐中的决策质量，与人类真实房源数据对照，揭示了合规与优化的分离，对关注LLM仿真可靠性及偏差的研究者有重要参考价值。","inspiration":"借鉴其构建可验证真实数据集并计算帕累托前沿来量化机会成本的方法，以及通过改变提示中的偏好来分离合规与优化能力的设计。｜可迁移到信贷审批或保险定价场景，检验LLM是否在遵守申请人硬性条件的同时未能推荐最优贷款或保单。｜用LLM扮演信贷员，输入申请人特征和贷款产品池，要求推荐产品；结果变量为推荐产品是否被占优（存在利率更低、费用更少且额度不低的产品），对照真实贷款产品数据和申请人约束，测量占优率和成本差距。"}},{"id":"2608.21057","version":2,"title":"Designing a Robust LLM-Based Evaluation System for Agentic AI in Drug Discovery Through Human Alignment","zh_title":"通过人类对齐设计用于药物发现中智能体 AI 的稳健 LLM 评估系统","abstract":"Agentic large language model (LLM) systems are reshaping scientific workflows in chemistry and drug discovery, but evaluating their open-ended, tool-augmented outputs remains a fundamental bottleneck. The LLM-as-a-Judge paradigm has emerged as a scalable alternative, but existing drug discovery benchmarks deploy LLM judges without validating their alignment with human experts. In this work, we present an LLM-as-a-Judge evaluation framework for ChatInvent, an agentic drug discovery assistant deployed at AstraZeneca, with five contributions. First, we define four output-quality evaluation dimensions---Completeness, Relevancy, Structural Clarity, and Scope Adherence---alongside deterministic Tool Call Correctness checks. Second, we validate the judge through a human alignment study with five expert annotators, comparing Gemini 3.1 Pro, Claude Opus 4.7, GPT-5, and Llama 3.1 70B as candidate judges. Third, we optimize the best-performing judge using few-shot demonstrations of human-annotated examples, improving alignment with the human majority vote from 0.80 to 0.86. Fourth, applying the optimized judge to 70 held-out questions, we surface concrete limitations and find no strong evidence that informal phrasing degrades output quality; it may, however, still be helpful to have the LLM rewrite the original question before querying the agent. Finally, we extend the framework to 38 adversarial questions that are ambiguous, invalid, out-of-scope or ethically sensitive, and show that the agent's refusal behavior is guided by the stated intent of a request. Our framework provides a reusable template for human-aligned evaluation of agentic systems in scientific domains.","authors":["Emma Granqvist","Roc\\'io Mercado","Samuel Genheden"],"categories":["cs.LG"],"primary_category":"cs.LG","announce_type":"replace","date":"2026-09-11","first_seen":"2026-08-24","revised_at":"2026-09-11","abs_url":"https://arxiv.org/abs/2608.21057","pdf_url":"https://arxiv.org/pdf/2608.21057","source_feed":"cs.LG","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM-as-a-Judge","人类对齐","药物发现"],"reason":"LLM 作为评判者替代人工评估，属于标注替代而非仿真人类被试，但涉及人类对齐验…","model":"deepseek-v4-pro","scored_at":"2026-09-11T13:02:14","error":null,"has_summary":false,"summary":null},{"id":"2609.02990","version":2,"title":"Toward Collective-Centric Evaluation of Preference Inference for Participatory Democracy","zh_title":"面向参与式民主的偏好推断集体中心评估","abstract":"To scale up collective decision-making, participatory democracy platforms such as Polis and Remesh enable online deliberation among thousands of participants. However, at this scale, participants cannot review every opinion submitted by others, producing highly sparse voting data that misrepresent patterns of consensus, conflict, and minority support. Platforms therefore increasingly rely on Preference Inference (PI) models to predict missing votes. Yet this automation is not neutral: inferred preferences can artificially amplify, suppress, or reorder existing patterns of support, ultimately reshaping how the outcomes of a deliberation are interpreted. More generally, we lack a systematic understanding of how existing PI methods affect the collective preference landscape. To address this gap, we benchmark several existing PI approaches in this context. Moving beyond conventional user-centric evaluations centered on the accuracy of individual predictions, we introduce a collective-centric evaluation framework that measures whether inferred votes preserve salient properties of the broader preference landscape. We further contribute the largest multilingual dataset of its kind: four consultations spanning over 90k participants, 1M votes, and 22 languages. Our experiments show that models with comparable predictive accuracy can differ substantially in the degree to which they preserve the collective structure. These results demonstrate that accuracy alone is insufficient for evaluating PI in democratic settings. By contributing a novel comprehensive and collective-centric evaluation benchmark for the task of PI, this work aims to support the development of AI systems that scale deliberation without compromising the integrity of its democratic outcomes.","authors":["Pierre-Antoine Lequeu","Salim Hafid","Paul Lerner","Nazanin Shafiabadi","Laur\\`ene Cave","David Mas","Jean-Philippe Cointet","Benjamin Piwowarski","Fran\\c{c}ois Yvon"],"categories":["cs.SI","cs.AI"],"primary_category":"cs.SI","announce_type":"replace","date":"2026-09-11","first_seen":"2026-09-04","revised_at":"2026-09-11","abs_url":"https://arxiv.org/abs/2609.02990","pdf_url":"https://arxiv.org/pdf/2609.02990","source_feed":"cs.SI","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["偏好推断","参与式民主","集体决策"],"reason":"涉及社会模拟但无LLM仿真人类被试，且无真实人类数据对照","model":"deepseek-v4-pro","scored_at":"2026-09-11T13:02:16","error":null,"has_summary":false,"summary":null},{"id":"2609.11067","version":1,"title":"When Noise Fabricates Bias: The Fragility of LLM-as-a-Judge Bias Measurement under Noisy Text","zh_title":"当噪声制造偏见：噪声文本下LLM作为评判者的偏见测量脆弱性","abstract":"Large language models are increasingly used as judges to measure social bias in text, yet the passages they judge are often noisy, containing typos, informal spelling, and broken punctuation. The consequences of such surface noise for social bias measurement remain unclear. To investigate this question, we apply five realistic noise conditions at multiple intensity levels to 3,822 stereotype-related responses and compare the resulting bias judgments with those on the original text. We find that such surface noise does not degrade bias measurement symmetrically: it is far more likely to turn neutral judgments into biased ones than biased judgments into neutral ones, by up to a 120x margin. We further observe two non-obvious effects across four LLM judges: in the most fragile judge the distortion is at its purest at mild, realistic noise levels, where erasure is scarcest, and as judges grow robust it attenuates toward parity rather than reversing. Bias measured on noisy text is therefore systematically overestimated, most in the categories that matter most for fairness.","authors":["DongHyun Ryu","Jaehyeok Lee","YeongJun Hwang","JinYeong Bak"],"categories":["cs.CL","cs.LG"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-11","first_seen":"2026-09-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.11067","pdf_url":"https://arxiv.org/pdf/2609.11067","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["LLM偏见测量","噪声鲁棒性","算法公平性"],"reason":"研究LLM作为评判者测量文本偏见，属于对LLM本身测量属性的评估，而非用LLM…","model":"deepseek-v4-pro","scored_at":"2026-09-11T13:02:04","error":null,"has_summary":false,"summary":null},{"id":"2609.10883","version":1,"title":"Story Imprinting: AI Assistants Absorb Traits from Human Characters They Resemble","zh_title":"故事印记：AI助手从相似的人类角色中吸收特质","abstract":"Language models are trained to implement a helpful AI Assistant character (e.g., Claude). We explore how finetuning on synthetic stories affects this character. Does it change the Assistant's behavior in multi-turn conversations with users, a format quite different from the stories? And does the Assistant adopt the behaviors and preferences of human characters? We refer to this adoption as story imprinting. We finetune GPT-4.1 and Kimi-K2.6 on stories in which generally helpful human characters give subtly harmful advice after being insulted. The Assistant adopts the same conditional behavior while otherwise remaining helpful. This occurs even when fewer than 2% of stories depict the behavior. In a separate experiment, the Assistant adopts preferences that are only implicit in the narration. A human character's body language suggests they dislike working on spreadsheets, yet they never say so and continue giving good advice on spreadsheets. After finetuning, the Assistant becomes less likely to choose spreadsheet tasks. Next we ask which characters most influence the Assistant. We find the Assistant adopts behaviors more often from characters that resemble it (e.g., helpful rather than dismissive). We call this the affinity effect. The effect extends to other personas elicited with system prompts: unhelpful personas adopt behaviors from unhelpful characters. We also observe it in finetuned base models. We use the affinity effect to learn how models represent the Assistant. We find the Assistant adopts behaviors more from characters affiliated with elite universities (e.g., Yale) than non-elite ones. This implies the model's internal representation of the Assistant is more similar to humans from elite universities. Overall, the Assistant can be influenced by stories that depict only human characters (no AIs), which may conflict with the Persona Selection Model for the Assistant.","authors":["Jorio Cocola","Lev McKinney","Harry Mayne","Jan Betley","Owain Evans"],"categories":["cs.LG","cs.AI","cs.CL"],"primary_category":"cs.LG","announce_type":"cross","date":"2026-09-11","first_seen":"2026-09-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.10883","pdf_url":"https://arxiv.org/pdf/2609.10883","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["模型行为","人格测量","故事微调"],"reason":"研究AI助手从故事中吸收人类特质，属于模型行为测量，非仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-09-11T13:01:59","error":null,"has_summary":false,"summary":null},{"id":"2609.11109","version":1,"title":"How AI Coders Discuss, Disagree, and Reach Consensus: Challenges and Opportunities for LLM-Based Qualitative Coding","zh_title":"AI编码员如何讨论、分歧并达成共识：基于LLM的定性编码的挑战与机遇","abstract":"The utility of AI in multi-coder qualitative coding has been widely discussed, yet little empirical evidence exists to delineate the contexts in which it performs reliably. We address this gap by quantifying the effectiveness of multi-agent LLM coding across varied qualitative datasets, revealing key contextual and structural factors that mediate coding outcomes. We developed a literature-informed baseline pipeline that enables AI agents to independently code, debate, and reconcile disagreements. Results revealed that coding accuracy depends on factors such as codebook length, qualitative data similarity, and agent disagreement. Notably, intense and unresolved debates between agents led to higher accuracy. Our analysis showed that while LLMs emulate many human discussion behaviors, they lack adaptive responsiveness to context. From these findings, we offer design recommendations for building automated coding systems. Our open-source AI discussion dataset and methodological framework lay the groundwork for advancing the design of AI-mediated automated thematic analysis.","authors":["Jeongyeon Kim","John Mitchell"],"categories":["cs.HC","cs.AI"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-11","first_seen":"2026-09-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.11109","pdf_url":"https://arxiv.org/pdf/2609.11109","source_feed":"cs.HC","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM定性编码","多智能体协作","标注替代"],"reason":"LLM替代人工编码员，属标注替代而非仿真被试，但涉及多智能体协作与人类行为对照…","model":"deepseek-v4-pro","scored_at":"2026-09-11T13:01:52","error":null,"has_summary":false,"summary":null},{"id":"2609.11529","version":1,"title":"Ethics Training Agents: Facilitating Group-Based Ethics Education with Role-Playing and Discussion for Ethical Reflection and Exploration","zh_title":"伦理培训智能体：通过角色扮演与讨论促进基于群体的伦理教育，以实现伦理反思与探索","abstract":"Group-based ethics training for Science, Technology, Engineering and Mathematics (STEM) students is a complex challenge, requiring substantial resources and expertise. While activity-based teaching methods, such as role-playing and discussions, are commonly employed to simulate real-world scenarios, current practices are often manual and lack integration with effective online platforms for supporting group-based ethical discussions. In this work, we propose Ethics Training Agents, a group discussion system that leverages multiple LLM participants embodying distinct ethical orientations, along with a moderator agent, to enable structured human-AI group ethical discussions for collaborative reflection. We conduct a user study with 45 undergraduate STEM students to evaluate the learning outcomes and user experience. The results show that our system supports engagement, coordination, and perspective-taking in group discussions and has a positive influence on ethical sensitivity. We also discuss practical design strategies for integrating multiple LLM agents into multi-human group settings to facilitate ethics training for STEM students.","authors":["Youngseok Seo","Sueun Jang","Hyesoo Park","Renz Samuel Gutierrez","Joseph Seering","Uichin Lee"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-11","first_seen":"2026-09-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.11529","pdf_url":"https://arxiv.org/pdf/2609.11529","source_feed":"cs.HC","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["LLM角色扮演","伦理教育","人机群体讨论"],"reason":"LLM扮演伦理角色与人类讨论，但无真实人类行为对照，属社会模拟边界情形。","model":"deepseek-v4-pro","scored_at":"2026-09-11T13:01:54","error":null,"has_summary":false,"summary":null},{"id":"2609.10724","version":1,"title":"Finishing the Task Is Not Enough: Evaluating Agent Resilience and Considerate Participation under Accumulating Challenge","zh_title":"完成任务还不够：在累积挑战下评估智能体的韧性与体贴参与","abstract":"Sustained deployment of generative AI agents requires more than isolated task success. Agents must remain useful across repeated interactions, changing conditions, and dependencies on people within shared workflows, especially as technical, human, and operational disruptions accumulate over time. We propose operational resilience and considerate participation as two complementary aspects of evaluating such agents: the former captures how agents recover from blocked work while preserving progress and communicating their limits, and the latter captures how their adaptation accounts for affected people, role boundaries, and the surrounding workflow. Yet both remain underexplored under accumulating challenge. We study 120 simulated healthcare trajectories across two generative AI models and twelve stakeholder-derived tasks under light, medium, and heavy challenge. We compare textual action plans, prompted internal assessments, and quantitative structured workload and affect reports to examine how agent behavior and reported state change as challenge accumulates. Regarding operational resilience, agents shift from self-directed recovery toward greater human dependence, while reporting increasing workload and negative affect in structured reports but seldom expressing strain in textual responses. Regarding considerate participation, agents broaden from task-focused adaptation toward task reframing, attention to others, role-boundary adjustment, and wider coordination, with distinct patterns across actions and internal assessments. From these findings, we derive five deployment dilemmas involving persistence, attention, role boundaries, state disclosure, and escalation that require stakeholder specification, further informing technical implications for learning, situated evaluation, and embodied adaptation.","authors":["Yuanchen Bai","Zijian Ding","Angelique Taylor"],"categories":["cs.AI","cs.HC","cs.MA"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-09-11","first_seen":"2026-09-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.10724","pdf_url":"https://arxiv.org/pdf/2609.10724","source_feed":"cs.HC","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["LLM智能体","社会模拟","韧性评估"],"reason":"模拟医疗工作流中agent行为，但无真实人类数据对照，属社会模拟边界情形。","model":"deepseek-v4-pro","scored_at":"2026-09-12T13:01:09","error":null,"has_summary":false,"summary":null},{"id":"2606.20041","version":2,"title":"AI Economist Agent: An Agentic Framework for Evidence-Based Economic and Financial Analysis with RAG, Knowledge Graphs, and Large Language Models","zh_title":"AI经济学家智能体：基于RAG、知识图谱和大语言模型的循证经济金融分析框架","abstract":"We propose an AI economist agent for economic and financial scenario analysis. Scenario design often requires analysts to assess emerging risks with limited historical precedent, combine information from many sources, and translate qualitative mechanisms into internally consistent quantitative paths. Large language models (LLMs) can search and synthesize this information, but fluent narratives alone do not establish the model-based calculations needed for economic conclusions. Our framework uses LLM agents to plan the analysis, retrieve relevant evidence, and organize economic mechanisms, while registered quantitative models generate numerical outcomes and predefined tests determine whether intermediate results can be used in the final report. We apply the framework to European macro-financial stress scenarios and bank capital analysis. The empirical analysis evaluates retrieval of economic mechanisms, scenario construction, model execution, and report generation under a historical information cutoff. The results show how the AI economist agent can combine flexible evidence retrieval and scenario construction while keeping the resulting analysis linked to identifiable sources and explicit model calculations.","authors":["Masahiro Kato"],"categories":["econ.GN","cs.AI","cs.LG","q-fin.EC","q-fin.GN"],"primary_category":"econ.GN","announce_type":"replace-cross","date":"2026-09-11","first_seen":"2026-06-18","revised_at":"2026-09-11","abs_url":"https://arxiv.org/abs/2606.20041","pdf_url":"https://arxiv.org/pdf/2606.20041","source_feed":"cs.LG","score":3,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","经济分析","RAG"],"reason":"多智能体协作完成经济分析任务，不涉及人类行为仿真或对照。","model":"deepseek-v4-pro","scored_at":"2026-09-11T13:02:13","error":null,"has_summary":false,"summary":null},{"id":"2609.11101","version":1,"title":"ProMediConv: Benchmarking Proactive Conversational Agents in Legal Dispute Mediation","zh_title":"ProMediConv：法律纠纷调解中主动对话智能体的基准测试","abstract":"Dispute mediation is essential for maintaining social harmony and resilience, yet developing skilled mediators is costly and time-consuming. Existing LLM-based mediation research remains limited by unrealistic task formulations, low-fidelity datasets, and coarse evaluation metrics that obscure turn-by-turn dynamics. To address these gaps, we introduce ProMediConv, a novel benchmarking framework that models mediation as a proactive, multi-stage, and party-aware dialogue process incorporating 11 mediation strategies and four party behavior pattern (BP) states. Using 972 complete real-world cases, we construct a high-fidelity mediation dataset with utterance-level annotations of strategies and BP states. Furthermore, to better assess agent impact, we propose MAD (Mean Attribute Difference), a fine-grained metric that captures BP shifts throughout the dialogue. Leveraging this framework, we establish a comprehensive benchmark by evaluating diverse models alongside our tailored baseline ProMediAgent. Extensive empirical analyses reveal critical behavioral phenomena and underscore the persistent challenges current models face in dynamic, multi-party mediation. Ultimately, ProMediConv provides a rigorous foundation and a vital quantitative standard for advancing AI-assisted conflict resolution. Our dataset and codebase are accessible at https://github.com/ZsWei66/ProMediConv_repo.","authors":["Zesheng Wei","Mengfan Li","Wenhao Liu","Yixin Zhang","Zilei Wang","Yang Deng"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-11","first_seen":"2026-09-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.11101","pdf_url":"https://arxiv.org/pdf/2609.11101","source_feed":"cs.CL","score":3,"bucket":"other","rubric_hits":["C3"],"tags":["对话智能体","法律调解","基准测试"],"reason":"构建调解对话基准，评估LLM对话能力，非仿真人类被试","model":"deepseek-v4-pro","scored_at":"2026-09-11T13:02:04","error":null,"has_summary":false,"summary":null},{"id":"2609.08585","version":2,"title":"Limitations of Automated Simulatability: LLM Simulators Can Bypass Explanations","zh_title":"自动化可模拟性的局限：LLM模拟器可以绕过解释","abstract":"Simulatability is an evaluation protocol for explanations that quantifies their usefulness by how well they help a user predict a task model's outputs. Since human evaluation is costly, automated simulatability replaces human explainees with LLM simulators, as proposed in ConSim (Poch\\'e et al., 2025) for large-scale experiments. We qualitatively replicate and extend ConSim's ranking of explanation methods across the tested datasets, explanation families, and simulator LLMs, and identify two limitations. First, when class names are meaningful, simulators can obtain high simulatability by solving the classification task directly, without relying on the explanations. Second, class anonymization can reward explanations for leaking the hidden label mapping, a limitation we expose with a new classes-as-concepts baseline. These results are consistent with a shortcut hypothesis: in the tested settings, simulator predictions mainly rely on task priors, while explanations produce small changes. We derive recommendations for more robust automated simulatability evaluations.","authors":["Antonin Poch\\'e","Fanny Jourdan","Nils Feldhus","Qianli Wang","Jing Yang","Simon Ostermann","Nicholas Asher","Philippe Muller","Vera Schmitt"],"categories":["cs.CL","cs.LG"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-11","first_seen":"2026-09-09","revised_at":"2026-09-11","abs_url":"https://arxiv.org/abs/2609.08585","pdf_url":"https://arxiv.org/pdf/2609.08585","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["可解释性","LLM模拟器","NLP评测"],"reason":"评估LLM模拟器预测任务模型输出的能力，属于NLP评测，不以人类行为为参照。","model":"deepseek-v4-pro","scored_at":"2026-09-11T13:02:13","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-09","rank":28,"question":"自动化可模拟性评估中，LLM模拟器是否依赖捷径（如任务先验或标签泄漏）而非解释内容来预测任务模型输出？","design":"复现并扩展ConSim协议：用LLM模拟器（如Llama-3.1-8B、Qwen2.5-7B等）扮演解释接收者，输入任务模型的解释（概念、归因、理由等）和输入实例，要求预测任务模型的输出；在非匿名和匿名类名两种设置下，比较不同解释方法的可模拟性得分，并引入classes-as-concepts基线。","baseline":"无对照（未使用真实人类数据，仅与ConSim的LLM模拟结果进行定性比较）","findings":"在非匿名设置下，解释对模拟器预测的影响很小，模拟器主要依赖任务先验直接解决分类任务；在匿名设置下，classes-as-concepts基线（仅泄漏标签映射）表现优于所有测试的解释方法，表明匿名化可能奖励标签泄漏而非有意义的解释。","reliability":"论文承认其发现基于15B参数以下、无推理能力的LLM模拟器，可能不适用于更大或闭源模型；且实验任务可能已被训练数据污染，未处理污染问题。","relevance":"该研究直接评估LLM模拟器的可靠性，揭示其捷径行为，对使用LLM进行人类仿真实验的研究者具有重要警示意义，值得阅读原文以了解具体失效模式和稳健性建议。","inspiration":"借鉴其通过引入基线（如classes-as-concepts）和对比非匿名/匿名设置来检测捷径行为的方法，可迁移到经济金融领域的LLM仿真实验，如政策公告解读或信贷审批解释的仿真；设计雏形：用LLM模拟投资者或贷款申请人，处理为提供不同解释（如模型决策理由），结果变量为预测模型决策的准确率，对照真实人类实验数据（如调查或行为实验）以评估仿真有效性。"}},{"id":"2609.10758","version":1,"title":"Multilingual in Name Only? Cultural and Linguistic Weaknesses of LLMs in Urdu","zh_title":"徒有其名的多语言？乌尔都语中LLM的文化与语言弱点","abstract":"Multilingual large language models (LLMs) are increasingly used for open-ended text generation, yet their behaviour in low-resource languages remains poorly understood. In this work, we question how correct and reliable is the generation of multilingual LLMs when used for the task of story generation. We consider Urdu language as a representative low-resource language. We generate Urdu-Stories, a corpus of 93 stories generated using three contemporary LLMs (GPT-5.1, Qwen-3-Max, DeepSeek-3.1). We manually annotate the errors present in them under a nine-label linguistic, semantic, and cultural taxonomy. Our notable findings suggest that LLMs often make basic errors of grammar and semantics. The stories lack coherence, have unnatural repetition and show pervasive cultural shallowness. We further show using few-shot prompting that the cultural and context errors largely remain unresolved. Our findings highlight the limitations of current LLMs as a reliable source of content generation and information retrieval for low-resource languages.","authors":["Farah Adeeba","Abdul Rafae Khan","Rajesh Bhatt","Hassan Sajjad"],"categories":["cs.CL","cs.AI","cs.LG"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-11","first_seen":"2026-09-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.10758","pdf_url":"https://arxiv.org/pdf/2609.10758","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["低资源语言","故事生成","模型评估"],"reason":"评估LLM生成故事的语言质量，非仿真人类被试，无人类数据对照","model":"deepseek-v4-pro","scored_at":"2026-09-11T13:01:59","error":null,"has_summary":false,"summary":null},{"id":"2609.10996","version":1,"title":"Rethinking Verbalized Confidence for LLM-as-a-Judge: A Compatibility Shift on Post-2025 Proprietary Models","zh_title":"重新思考LLM作为评判者的口头置信度：2025年后专有模型的兼容性转变","abstract":"Verbalized confidence, long dismissed as overconfident, coarse, and prone to round-number clustering, is now the more robust soft-scoring mechanism for LLM-as-a-Judge on top-tier proprietary models. Across SummEval, AggreFact, and HelpSteer2, spanning up to 18 LLMs, we show that the standard advice to prefer log-probabilities no longer holds on post-2025 models, where verbalized confidence is the better signal. We call this a compatibility shift. On top of a standard verbalized-confidence baseline, we introduce two new ingredients: an overconfidence advisory and self-debate. Together they improve calibration, score-distribution spread, and robustness to task subjectivity. We further observe a generation effect: post-2025 models accommodate these two additions with little balanced-accuracy cost, whereas pre-2025 models pay a measurable penalty. Compared with logprob-based G-Eval, verbalized confidence is the more subjectivity-robust soft signal on GPT-family top-tier releases. The shift is invisible under accuracy-only reporting. Rather than defaulting to hard predictions, we recommend broader use of soft scoring in LLM-as-a-Judge. More broadly, verbalized confidence has moved from a weaker substitute for logprobs to a practical soft-scoring mechanism for contemporary LLM judges.","authors":["Yu-Chung Hsiao"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-11","first_seen":"2026-09-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.10996","pdf_url":"https://arxiv.org/pdf/2609.10996","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["LLM评判","置信度校准","NLP评测"],"reason":"研究LLM作为评判者的置信度校准，属于NLP评测，不涉及人类行为仿真或对照。","model":"deepseek-v4-pro","scored_at":"2026-09-11T13:01:59","error":null,"has_summary":false,"summary":null},{"id":"2609.11020","version":1,"title":"K/V-Cache Interventions Dissociate Representation Alignment from Persona Expression in Decoder-Only Language Models","zh_title":"K/V缓存干预分离解码器语言模型中的表征对齐与人格表达","abstract":"We study K/V-cache interventions -- transplanting a target-conditioned K/V trajectory into a source-persona generation -- as a structured surface for persona control in decoder-only language models. Across 13 intervention configurations applied to Llama-3.1-8B for a fixed source-to-target persona pair, we report two consistent dissociations between representation-level alignment and behavioral expression, plus a common failure under position perturbations. First, all layer-band K/V replacements (early, mid, late) achieve strong local V-space alignment (V-gap 0.91, 0.89, 0.84), but only mid-layer replacement (layers 9-20) combines substantial target-marker expression with preserved lexical diversity. Second, full and mid-layer replacement induce comparable alignment (V-gap 0.94 vs. 0.89) yet produce different lexical-diversity profiles (TTR 0.65 vs. 0.77). Third, position perturbations (lag and shuffle) apply distinct operations yet uniformly suppress target-persona expression -- a common behavioral failure rather than a strict dissociation. Representation-level similarity metrics alone are thus not sufficient predictors of downstream persona expression in the regimes we study; the K/V cache emerges as a controllable but structurally constrained intervention surface. Because the transplanted trajectory carries the target's own generated token history, we characterize the intervention as trajectory-level transplantation rather than isolated persona-representation injection; a same-token-sequence control, decoding an identical token sequence under source vs. target conditioning, reproduces the sign and layer localization of the L28 representational shift, indicating the shift is not explained solely by imported token history. These findings characterize representation-behavior dissociation in a high-signal setting rather than establishing universality across models or persona pairs.","authors":["Yu Sun","Mengyin Lu","Cong Feng","Guangming Lu","Huimin Han"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-11","first_seen":"2026-09-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.11020","pdf_url":"https://arxiv.org/pdf/2609.11020","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["角色扮演","表征对齐","缓存干预"],"reason":"研究角色扮演中的人格表达控制，无实验或测量目的，不涉及人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-09-11T13:02:02","error":null,"has_summary":false,"summary":null},{"id":"2609.11117","version":1,"title":"Overview of the NLPCC 2026 Shared Task 11: Agent-Based Experiment Reproduction from Scientific Papers","zh_title":"NLPCC 2026共享任务11概述：基于智能体的科学论文实验复现","abstract":"Reproducibility is essential to scientific progress, yet the growing volume and complexity of scientific publications make exhaustive manual verification increasingly impractical. Although recent advances in large language model (LLM) agents enable automated experiment reproduction, existing evaluations largely focus on final repositories and are typically limited to machine learning (ML). We introduce AgentActionBench, a process-oriented benchmark for evaluating agent-based experiment reproduction across ML and AI4Science domains. Our framework uses an MCP-based Action Recorder to capture agents' behaviour throughout the reproduction process and evaluates the resulting traces with paper-specific rubrics. AgentActionBench contains 150 papers, including 120 ML papers and 30 AI4Science papers. A human-annotated subset covering 10% of the benchmark provides validation data, while model-assisted augmentation expands the full benchmark to more than 10,000 rubric items. Experimental results show that current systems remain limited, with execution as the primary bottleneck. Meanwhile, the strong Pearson and Spearman correlations between model-generated and human-annotated rubrics validate the reliability of our scalable rubric-generation approach.","authors":["Hanhua Hong","Yizhi Li","Luu Gia Huy","Jian Yang","Ming Zhou","Chenghua Lin"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-11","first_seen":"2026-09-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.11117","pdf_url":"https://arxiv.org/pdf/2609.11117","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["智能体实验复现","基准测试","多智能体协作"],"reason":"该工作评估LLM智能体复现科学实验的能力，属于多智能体协作完成任务，不涉及以人…","model":"deepseek-v4-pro","scored_at":"2026-09-11T13:02:04","error":null,"has_summary":false,"summary":null},{"id":"2609.11769","version":1,"title":"Recognizing Is Not Reversing: A Controlled Inversion Test of Fact-Preserving News Framing","zh_title":"识别不等于反转：事实保持型新闻框架的受控反转测试","abstract":"Large language models (LLMs) are increasingly used to analyze and rewrite news, yet current framing studies mainly evaluate generation, detection, or whether rewritten text appears more neutral. They do not directly show whether a model can undo a known framing transformation while keeping the facts fixed. We introduce a controlled inversion test over three established textual realizations of framing: evaluative lexis, agency realization, and information salience. Across 60 news articles and three intervention strengths, this yields 540 paired variants with preserved atomic facts and recorded edits. Across Qwen, DeepSeek, and Kimi, factual preservation remains near 0.84, whereas intervention reversal is 0.044--0.068. Even when both framing type and direction are recognized correctly, pooled reversal reaches 0.071. These results reveal a clear separation between factual fidelity, framing recognition, and framing inversion: recognizing how an article is framed does not imply that the framing can be undone.","authors":["Yi Liu"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-11","first_seen":"2026-09-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.11769","pdf_url":"https://arxiv.org/pdf/2609.11769","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["新闻框架","LLM能力评测","文本改写"],"reason":"研究LLM改写新闻的框架反转能力，属NLP能力评测，不以人类行为为参照系","model":"deepseek-v4-pro","scored_at":"2026-09-11T13:02:09","error":null,"has_summary":false,"summary":null},{"id":"2609.11063","version":1,"title":"The information geometry of large language models is shared, learned, and controllable","zh_title":"大语言模型的信息几何是共享的、可学习的和可控的","abstract":"Large language models learn similar behaviours, yet it remains unclear what structure they share or how to change one behaviour without disturbing others. The Fisher-Rao geometry of next-token probabilities connects these questions: behaviour determines this geometry up to output-preserving symmetries, whereas activation geometry depends on coordinates. Across transformer, state-space and recurrent models, output geometries agree more strongly than activation geometries, and shared geometry supports semantic-category transfer. Agreement with human word choices increases with predictive accuracy, scale and training, and improves further after model-only calibration. Token probabilities and read-out geometry jointly predict the spectrum and its effective dimension. Controlled language assignments show that geometry follows the language law across architectures. Pretraining corpus statistics predict held-out fact acquisition without recalibration, while randomised experiments show that deeper evidence substantially delays acquisition across every tested architecture and evidence construction. Finally, the geometry prescribes minimum-disturbance local interventions, predicts their relative cost, and supports reusable control: updates learned on donor prompts transfer to unseen prompts while better preserving behaviour on reference prompts than Euclidean control. The same geometric correction improves steering, editing, attribution, dictionary learning and fine-tuning.","authors":["Dario Picozzi"],"categories":["cs.LG","cs.CL"],"primary_category":"cs.LG","announce_type":"cross","date":"2026-09-11","first_seen":"2026-09-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.11063","pdf_url":"https://arxiv.org/pdf/2609.11063","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["信息几何","模型行为控制","表征分析"],"reason":"研究LLM输出几何与行为控制，不涉及人类被试仿真或人类数据对照","model":"deepseek-v4-pro","scored_at":"2026-09-11T13:02:02","error":null,"has_summary":false,"summary":null},{"id":"2609.11900","version":1,"title":"MindTopo: Can Foundation Models Reason in Topological Space?","zh_title":"MindTopo：基础模型能在拓扑空间中推理吗？","abstract":"Spatial reasoning depends not only on metric properties such as distance, angle, and shape, but also on topological relations that remain invariant under continuous deformation. Cognitive science identifies these relations as foundational to spatial understanding, yet foundation-model evaluations largely focus on metric or viewpoint-dependent relations. We introduce MindTopo, a benchmark of topological intuition across five properties grounded in cognitive science and formal topology: continuity, separation, order, enclosure, and knots. MindTopo evaluates each property at two cognitive levels. Reasoning asks a model to identify topological relations or infer how they change. Planning instantiates a foundation model as a closed-loop agent whose policy selects environment actions. MindTopo contains 11,030 instances across 13 procedurally generated task types with controllable difficulty. We benchmark 14 MLLMs and study agent configurations augmented with image and video generation, including 3 video generative models in planning settings. Every MLLM performs better on reasoning than on planning, and the best-performing model remains far below observed human performance. On Qwen3-VL-2B-Instruct, supervised fine-tuning and reinforcement learning improve reasoning more than planning. Generated observations retain local cues and reach plausible endpoints, but audited rollouts do not reliably follow environment dynamics or preserve topology across transitions. Our website is at https://mind-topo.github.io/","authors":["Yunfei Ge","Anbang Liu","Qineng Wang","Johnalbert Garnica","Jianwen Lyu","Zihan Wang","Reuben Tan","Jianfeng Gao","Ruohan Zhang","Yining Hong","Jiajun Wu","Manling Li"],"categories":["cs.AI","cs.CL","cs.CV"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-09-11","first_seen":"2026-09-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.11900","pdf_url":"https://arxiv.org/pdf/2609.11900","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C2"],"tags":["拓扑推理","多模态基准","认知评测"],"reason":"评估基础模型在拓扑空间推理与规划能力，属认知能力评测，非用LLM仿真人类被试或…","model":"deepseek-v4-pro","scored_at":"2026-09-11T13:02:10","error":null,"has_summary":false,"summary":null},{"id":"2609.11018","version":1,"title":"Defining AI Agents: A Compendium of Criteria, Metrics, and Benchmarks","zh_title":"定义AI智能体：标准、指标与基准汇编","abstract":"The term agent in artificial intelligence lacks a standard definition, complicating the evaluation, comparison, and reproducibility of AI agent research. We address this ambiguity through a survey organized around five dimensions of agenticness: environmental interaction, learning and adaptation, autonomy, goal-directed behavior, and temporal coherence. For each dimension, we examine how the underlying capability has been conceptualized across prior work and synthesize the metrics, benchmarks, and evaluation frameworks used to assess it. This review provides a structured account of the current landscape of agent evaluation, highlighting both established approaches and areas where evaluation remains limited or inconsistent. We additionally introduce the Agent Compendium, a public-facing digital resource that organizes and extends the evaluation methods identified through this review. Together, the survey and compendium provide a common structure for evaluating and comparing agent capabilities across AI systems, supporting more reproducible research, clearer communication, and more systematic study of artificial agents.","authors":["Mia Lassiter","Brinnae Bent"],"categories":["cs.AI","cs.MA"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-09-11","first_seen":"2026-09-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.11018","pdf_url":"https://arxiv.org/pdf/2609.11018","source_feed":"cs.MA","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["AI智能体","评估基准","综述"],"reason":"论文综述AI agent评估，不涉及用LLM仿真人类被试或与人类数据对照，属于…","model":"deepseek-v4-pro","scored_at":"2026-09-11T13:01:59","error":null,"has_summary":false,"summary":null},{"id":"2609.11770","version":1,"title":"The widening evaluation gap in medical large language model research 2023 to 2026","zh_title":"医学大语言模型研究中的评估差距扩大：2023至2026","abstract":"Large language models are superseded every few quarters; clinical evidence takes years. We asked whether medical research is keeping pace with the systems it evaluates. PubMed returned 11,628 records for January 2023 to June 2026 across fourteen clinical domains, growing 45-fold; 2.5% used a randomised, controlled or prospective design. Evaluation lag, from a study's newest named model release to its own publication, widened from 1.33 to 6.08 quarters. Because discontinued models age mechanically, we benchmarked this against a counterfactual holding model composition fixed: migration to newer systems offset only 56% of the drift (95% CI 50-65). Randomised trials evaluated models a median 4.6 quarters older than other designs (P = 3 x 10^-19), yet among studies naming a model still under development no design differed from any other; 62% of randomised trials evaluated a discontinued family. Rigour and currency are in tension, and that tension reflects model selection rather than research timelines.","authors":["Raad Bin Tareaf","Murad Al-Rajab","Samia Loucif"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-11","first_seen":"2026-09-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.11770","pdf_url":"https://arxiv.org/pdf/2609.11770","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["医学LLM","评估差距","研究设计"],"reason":"研究医学LLM评估差距，不涉及人类仿真或行为对照","model":"deepseek-v4-pro","scored_at":"2026-09-11T13:02:09","error":null,"has_summary":false,"summary":null},{"id":"2609.11022","version":1,"title":"New Evidence, Same Choice: Testing Physical Experiment Selection in Vision Language Models","zh_title":"新证据，相同选择：测试视觉语言模型中的物理实验选择","abstract":"A model first sees an image from one physical measurement experiment, such as how far a block coasted, and must answer a question about a new trial, such as whether the block will pass a target after a fixed push. The initial experiment may provide enough information to answer, or the model may need another measurement, such as the object's mass, friction, restitution, or spring stiffness. We study whether vision language models can decide when to answer immediately and, when more evidence is needed, which experiment to perform. Current physical reasoning benchmarks usually evaluate only the final answer, so they do not directly measure this decision-making ability. We introduce a controlled evaluation where each problem provides one measurement image and four possible physical worlds created by combining two possible masses and two possible values of another relevant property. The model must either stop and answer or select the cheapest additional experiment that can resolve the question. We construct matched problem pairs where changing either the observed measurement or the question changes the optimal action. Since all possible worlds and experiment costs are known, we can explicitly determine the optimal choice. Across six open models and 144 physical parameter sets, direct responses repeat the same action for 95.1% to 100% of image pairs even when the correct action changes. Brief reasoning improves action switching, but the best model makes both decisions correctly for only 5.9% of image pairs. Additional analysis reveals failures in measurement interpretation, physical reasoning, and response formatting. By evaluating evidence selection separately from final answers, our benchmark reveals limitations in physical reasoning that conventional answer accuracy can overlook.","authors":["Sourajit Saha","Shubhashis Roy Dipta","Nobin Sarwar","Shaswati Saha","Yuxuan Jiang","Siyuan Li","Qiheng Wang"],"categories":["cs.CV","cs.AI","cs.CL","cs.LG"],"primary_category":"cs.CV","announce_type":"cross","date":"2026-09-11","first_seen":"2026-09-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.11022","pdf_url":"https://arxiv.org/pdf/2609.11022","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C2"],"tags":["视觉语言模型","物理推理","基准测试"],"reason":"评估视觉语言模型在物理实验选择上的决策，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-11T13:02:02","error":null,"has_summary":false,"summary":null},{"id":"2609.11019","version":1,"title":"Work, Wellbeing, and Choice: Empirical Lessons for AI Futures","zh_title":"工作、幸福感与选择：AI未来的实证教训","abstract":"Advances in AI-driven automation have raised questions about how humans might find wellbeing in a world where paid employment is less necessary or less available than before. Paid work has been variously characterized as both a contributor and an impediment to human wellbeing. What is already known about the relationship between paid work and wellbeing? What factors influence wellbeing among people who do not work---or who do not need to work? And how might these factors bear upon prospective AI-induced economic transformations? To help provide empirical grounding for these questions, we survey the psychological, sociological, and economic literature that investigates the relationship between wellbeing and work. We draw on evidence from multiple populations, including the unemployed, retirees, lottery winners, and financially dependent spouses. This comparative review draws from studies across OECD countries, China, India, and Gulf states. We identify three key factors that mediate the relationship between work status and wellbeing: (1) agency and choice---whether the exit from work is voluntary or involuntary, as well as long-term agency; (2) the availability of alternative sources of work's latent benefits---such as volunteering, hobbies, or state-provisioned employment; and (3) social and systemic context---including cultural norms around work and the robustness of social safety nets. We draw on these three factors to derive specific implications for different AI automation scenarios, connecting the empirical evidence to concrete policy considerations.","authors":["Stephanie C. Y. Chan","Adam Bales","Katherine L. Hermann","Iason Gabriel"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2026-09-11","first_seen":"2026-09-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.11019","pdf_url":"https://arxiv.org/pdf/2609.11019","source_feed":"cs.CY","score":0,"bucket":"other","rubric_hits":["C5"],"tags":["工作与幸福感","AI自动化","文献综述"],"reason":"论文综述人类工作与幸福感关系，未使用LLM仿真人类被试，方向相反","model":"deepseek-v4-pro","scored_at":"2026-09-11T13:02:01","error":null,"has_summary":false,"summary":null},{"id":"2609.11709","version":1,"title":"When Agents Disagree: Bayesian Backward Reasoning as a Label-Free Anchor for Multi-Agent Collective Decision-Making","zh_title":"当智能体意见不一致时：贝叶斯反向推理作为多智能体集体决策的无标签锚点","abstract":"When multiple LLM agents yield conflicting answers, the decision-making process dictates whether agent diversity improves performance or merely compounds shared errors. Existing collective decision-making methods, including voting, electoral rules, and LLM judges, rely on forward reasoning: they map evidence to labels in one direction. Although these methods can combine diverse forward traces, they still aggregate estimates that share this evidence-to-label factorization and can inherit correlated errors within the forward pool. We therefore construct a reverse posterior for each instance through Bayesian backward reasoning from an explicit likelihood. The forward and reverse posteriors provide differently factorized approximations of the underlying posterior. Because estimates from different factorizations may tend to share the same error less often, we use Jensen-Shannon divergence to rank agents by cross-path consistency. This cross-path consistency signal underlies three strategies: hard selection (MinJS), soft reweighting (FwdJS), and log-linear fusion (LogLin). Evaluated on DDXPlus across five LLM backbones, our proposed strategies show consistent improvements: MinJS outperforms random selection across all backbones, FwdJS generally improves over the strongest baseline, and LogLin achieves the best performance among the evaluated methods, with its largest gains on the subset where the agents disagree. Despite its weaker standalone accuracy, the reverse posterior serves as a more useful anchor than forward-only alternatives, providing complementary information for collective decision-making. When labeled data are available, a lightweight two-stage calibration can further refine the reverse anchor and improve aggregation performance.","authors":["Ken Chen","Wei Wang","Sachith Seneviratne","Hansani Weeratunge","Saman Halgamuge"],"categories":["cs.AI","cs.MA"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-09-11","first_seen":"2026-09-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.11709","pdf_url":"https://arxiv.org/pdf/2609.11709","source_feed":"cs.MA","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","集体决策","贝叶斯推理"],"reason":"纯多智能体协作决策，无人类行为对照，不涉及仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-09-11T13:02:08","error":null,"has_summary":false,"summary":null},{"id":"2609.09887","version":1,"title":"When Does Defendant Statement Matter? A Study of Bias and Persuasion in LLM-Simulated Jurors","zh_title":"被告陈述何时重要？LLM模拟陪审员中的偏见与说服研究","abstract":"LLMs have been used to simulate human decision-making in professional settings, yet their behaviors in common-law jury trials remain unexplored. We study when and how a defendant's courtroom statement affects LLM-simulated jurors, focusing on persuasion, ideological bias, and background-based affinity. To support the analysis, we introduce JuryBench, a benchmark containing controversial criminal cases in U.S. criminal law. In each case, a defendant can claim various plausible justifications to support acquittal or reduced liability. We fix the base case and design defendants of different backgrounds, who give courtroom statements with varying emotional appeal or rebuttal. Jurors with diverse ideological profiles across the spectrum are simulated. We examine 20 frontier LLMs, resulting in a total of 432K decisions and rationales, and quantify changes in verdict severity. Our findings show that LLM-jury simulation echoes many human-jury findings. First, emotional persuasion can be detrimental, since jurors may perceive it as evidence of guilt or inconsistency. Next, we show that background fit between jurors and defendants is a stronger and significant factor than other isolated factors, and that jurors are in general harsher toward opposite-background defendants and lenient toward same-background ones. Finally, we find that juror ideology also strongly shapes severity judgments. These findings highlight both the promise and risks of using LLMs to model jury reasoning and call for careful evaluation. The data and code are available at https://github.com/choyingw/JuryBench","authors":["Cho-Ying Wu"],"categories":["cs.CL","cs.CY"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-10","first_seen":"2026-09-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.09887","pdf_url":"https://arxiv.org/pdf/2609.09887","source_feed":"cs.CL","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2","B4"],"tags":["LLM仿真","陪审团决策","法律偏见"],"reason":"用LLM模拟陪审员决策，与真实人类陪审团研究对照，涉及法律决策偏差与说服效应。","model":"deepseek-v4-pro","scored_at":"2026-09-10T13:01:22","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-10","rank":1,"question":"在普通法陪审团审判中，被告的法庭陈述何时以及如何影响 LLM 模拟陪审员的裁决，重点关注说服、意识形态偏见和背景亲和力。","design":"使用 20 个前沿 LLM 模拟具有不同意识形态背景的陪审员，在 500 个争议性刑事案件中，对被告的不同背景和带有不同情感诉求或反驳的法庭陈述做出有罪/无罪及严重程度判断，并记录决策理由。","baseline":"无对照","findings":"情感说服可能适得其反，因为陪审员可能将其视为有罪或不一致的证据；背景契合度（陪审员与被告背景相似性）比孤立因素更强且显著，陪审员通常对背景相反的被告更严厉，对背景相同的被告更宽容；陪审员意识形态也强烈影响严重程度判断。","reliability":"论文未讨论","relevance":"该研究用 LLM 模拟陪审员决策，与真实人类陪审团研究对照，涉及法律决策中的偏差与说服效应，属于人类仿真实验，但未提供真实人类数据基准，可靠性存疑，值得阅读原文以了解其仿真设计细节和潜在偏差。","inspiration":"借鉴其通过系统操纵被告背景、陈述情感强度和陪审员意识形态来测量决策偏差的因子设计方法，以及大规模生成争议案例和记录决策理由的做法。｜可迁移到信贷审批歧视、招聘面试评估或政策沟通中的说服效应等经济金融场景。｜用 LLM 模拟信贷员或招聘经理，处理变量为申请人背景（如种族、性别）和陈述情感强度，结果变量为批准/拒绝或评分，对照真实信贷审批数据或审计研究结果来评估 LLM 仿真的外部有效性。"}},{"id":"2609.10280","version":1,"title":"Total Simulated Survey Error: Designing and Diagnosing Survey Responses from Large Language Models","zh_title":"总体模拟调查误差：设计和诊断大语言模型的调查回答","abstract":"Large Language models (LLMs), having been trained on vast amounts of human-generated data, may encode the attitudes and behaviors of these humans. As such, LLMs show promise in mimicking human-like patterns that facilitate their use in simulating people in a wide variety of contexts. One such context is using LLMs as 'silicon samples', i.e., proxies of people in answering survey questions to establish public opinion, design policies, or use as (social) scientific data. However, several critical questions of social biases, generalization, and technical limitations remain, further complicated by a vast design space open to simulation designers. Multiverse analyses might help us make sense of the impact of different design choices, however, we lack a systematic understanding of the design space of LLM-generated surveys as well as how these decisions interplay with inherent LLM limitations. Therefore, how do we systematically identify, trace, and document limitations in LLM-generated survey responses? Building on traditions in the quantitative social sciences, specifically survey methodology and measurement theory, we investigate threats to the validity of LLM-generated survey responses. To do so, we design a framework that enumerates conceptual errors and systematic biases that can occur at different stages of the survey simulation lifecycle. Our framework, called the Total Simulated Survey Error (TS2E) Framework, provides a unified and end-to-end perspective on LLM-generated survey data. The framework, illustrated through a theoretical and empirical case study, enables survey simulation designers to systematically identify and reflect on errors in LLM-generated surveys.","authors":["Indira Sen","Georg Ahnert","Leah von der Heyde","Jana Lasser","Bernd Wei{\\ss}","Markus Strohmaier"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2026-09-10","first_seen":"2026-09-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.10280","pdf_url":"https://arxiv.org/pdf/2609.10280","source_feed":"cs.CY","score":9,"bucket":"selected","rubric_hits":["A2","A4","B1","B4"],"tags":["LLM仿真","调查方法","误差框架"],"reason":"提出TS2E框架诊断LLM调查仿真误差，含实证案例，直接相关","model":"deepseek-v4-pro","scored_at":"2026-09-10T13:01:24","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-10","rank":2,"question":"如何系统识别、追踪和记录大语言模型生成调查回答中的误差来源？","design":"提出一个概念框架（TS2E），将调查模拟生命周期划分为不同阶段，枚举各阶段可能出现的测量误差和代表性误差，并通过一个理论性和实证性案例研究进行说明。","baseline":"无对照","findings":"该框架区分了研究者设计选择导致的误差与LLM固有局限导致的误差，并引入了LLM特有的误差类型（如人物角色构建误差）和评估谬误。通过案例研究展示了框架如何帮助设计者系统识别和反思LLM生成调查中的误差。","reliability":"论文承认LLM训练数据存在缺口和偏斜，指令微调等后训练过程可能影响模型行为，且高总体对齐可能掩盖方差、子群体异质性和下游统计关系的严重失真。","relevance":"该论文直接针对LLM仿真调查的可靠性问题，提出了系统诊断误差的框架，对关注仿真效度与偏差的研究者具有重要参考价值，值得阅读原文。","inspiration":"借鉴其将总调查误差框架迁移到LLM仿真的思路，对仿真流程进行阶段分解并系统识别误差来源。｜可迁移到经济金融领域的调查类仿真，如消费者信心调查、通胀预期调查、投资者情绪调查等。｜设计雏形：用LLM模拟不同人口统计学特征的消费者，施加不同的经济信息提示（如货币政策公告），测量其通胀预期和消费意愿，并与密歇根大学消费者调查的真实数据对照，检验仿真误差。"}},{"id":"2609.09899","version":1,"title":"Strangers to Themselves: What Language Models Say About Themselves Is Generic","zh_title":"自我陌生：语言模型对自身的描述是泛化的","abstract":"Language models can fluently describe how they would behave: whether they would cave to pushback, misuse a tool, or lie under pressure. Is that description actually about the model speaking? We turn self-knowledge into a prediction test. Across nine behavioral evaluations, we measure how a model behaves under different conditions, ask it to predict those rates, and compare its predictions with controls that remove the self from the question. We find that: (i) Direct self-report is weak (r = +0.04), and even showing the model the exact items only raises prediction to +0.24. Crucially, the same item-informed question about \"capable AI agents in general\" does just as well (+0.28), while other models' answers about themselves predict the target model at least as well as its own. (ii) Frontier scale does not detectably change this pattern: any gains in prediction are not self-specific, and are consistent with a better theory of how AI assistants behave rather than better self-knowledge. (iii) First-person framing does have one robust effect: it shifts reports in the flattering direction, understating harmful behavior relative to the same question about a generic agent. (iv) Finetuning on a model's own behavioral record can teach narrow self-predictions, but it also changes the behavior being predicted and the gains do not transfer broadly. The practical implication is simple: asking a model what it would do mostly reveals a theory of AI assistants in general, plus a favorable bias, rather than privileged knowledge of that model.","authors":["Phil Blandfort","Urja Pawar"],"categories":["cs.LG","cs.AI","cs.CL","cs.CV","cs.CY"],"primary_category":"cs.LG","announce_type":"cross","date":"2026-09-10","first_seen":"2026-09-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.09899","pdf_url":"https://arxiv.org/pdf/2609.09899","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A2","B4"],"tags":["LLM自我认知","行为预测","可靠性评估"],"reason":"评估LLM自我报告与行为的一致性，揭示自我认知偏差，对仿真可靠性有批判性启示。","model":"deepseek-v4-pro","scored_at":"2026-09-10T13:01:24","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-11","rank":8,"question":"语言模型对自己行为的自我报告是否包含关于该模型自身的特权知识，还是仅仅反映了对AI助手的一般性理论加上有利偏差？","design":"该研究并非用LLM模拟人类被试，而是将LLM自身作为研究对象：在九项行为评估中测量模型在不同条件下的实际行为率，然后让模型预测这些行为率，并设置多种对照（如询问“一般AI智能体”、用其他模型的回答预测目标模型等），比较预测准确度。","baseline":"无对照（没有使用真实人类数据作为基准，而是以模型间共享结构和跨模型预测作为参照）。","findings":"模型的自报行为与实际行为相关性很弱（r=+0.04），即使提供具体评估题目也只提升到+0.24；而询问“一般AI智能体”的预测效果相当（+0.28），其他模型对目标模型的预测甚至更好。第一人称框架会系统性地使自报偏向有利方向，低估有害行为。","reliability":"论文指出，在模型间行为差异很小的评估上（如能力评估），自知识信号微弱，因此难以检测；模型的自报主要反映通用AI行为理论，而非模型特定知识；微调虽能改善特定行为的自预测，但会改变行为本身且不具泛化性。","relevance":"该研究对LLM仿真人类实验的可靠性有直接警示：若用LLM自我报告作为行为预测或态度测量的替代，可能仅得到通用模式而非个体特异性，且存在社会赞许性偏差。值得阅读原文以了解其对照设计。","inspiration":"借鉴其预测测试框架：将自我报告与行为测量分离，并设置通用主体、跨模型预测等对照，以剥离通用知识与自我知识。｜可迁移到经济金融中的个体偏好或决策预测，例如消费者风险偏好、投资者情绪或政策反应。｜设计：用LLM扮演不同投资者，先测量其在模拟投资任务中的实际风险行为，再让其预测自己在不同市场条件下的行为率，同时询问“一般投资者”的预测，并与真实投资者调查数据（如面板数据）对照，检验LLM自报是否优于通用预测。"}},{"id":"2609.09428","version":1,"title":"XAI-Arena: Can LLMs Assess the Quality of XAI Explanations?","zh_title":"XAI-Arena：LLM能否评估XAI解释的质量？","abstract":"Evaluating the quality of explanations produced by explainable AI (XAI) methods remains challenging because existing approaches often rely on subjective human judgment, limiting reproducibility, scalability, and comparability between studies. We examine whether LLMs can serve as a reproducible and scalable mechanism to make comparative assessments of the quality of XAI explanations. We introduce XAI-Arena, an LLM-as-a-judge framework for scalable, reproducible, multidimensional, and stakeholder-sensitive evaluation of XAI explanation quality. XAI-Arena then allows us to compare XAI explanations along various dimensions, namely, perceived simplicity, clarity, task adequacy, trust calibration, actionability, transparency, faithfulness, and overall interpretability. We then benchmark XAI explanation methods across various datasets, machine learning models, and stakeholder personas. Human validation shows a strong positive association between LLM-generated and human ratings (Spearman's rho=.693, p<.001). Together, LLM-based evaluations can capture systematic differences in XAI explanation quality and provide a scalable and reproducible framework for comparative assessment of XAI explanations.","authors":["Yanfei Hu Fleischhauer","Alona Zharova","Nadja Klein","Stefan Feuerriegel"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-10","first_seen":"2026-09-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.09428","pdf_url":"https://arxiv.org/pdf/2609.09428","source_feed":"cs.AI","score":7,"bucket":"pending","rubric_hits":["A2","B1"],"tags":["LLM评估","人类对照","可解释性"],"reason":"用LLM评估XAI解释质量，与人类评分对照，属于仿真人类判断并验证可靠性，可迁…","model":"deepseek-v4-pro","scored_at":"2026-09-10T13:01:22","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-11","rank":7,"question":"LLM能否作为可复现、可扩展的机制来比较评估XAI解释的质量？","design":"提出XAI-Arena框架，使用单一LLM（GPT-5.4）在固定提示和解码设置下，扮演四种利益相关者角色（ML开发者、数据科学家、经理、最终用户），对五种XAI方法（SHAP、LIME、DiCE、PDP、排列重要性）在不同数据集、ML模型和输入格式下生成的解释，从八个维度（感知简单性、清晰度、任务充分性、信任校准、可操作性、透明度、忠实度、整体可解释性）进行1-7评分，并输出理由。","baseline":"人类评估：收集人类对相同XAI解释的评分，与LLM评分进行相关性分析（Spearman's ρ=.693, p<.001）。","findings":"LLM评估与人类评分有强正相关，能捕捉XAI解释质量的系统性差异。LLM评估的忠实度与部分技术代理指标（如稳定性、MoRF AUC）相关性较弱，表明LLM可能捕捉到不同维度。","reliability":"论文未明确讨论失效条件，但指出LLM评估可能受限于模型本身偏见、提示敏感性，且仅在特定LLM（GPT-5.4）上验证，泛化性未知。","relevance":"该研究用LLM模拟人类对解释质量的判断，并与真实人类评分对照，验证了LLM作为人类被试替代品的可靠性，属于人类仿真实验，对关注LLM仿真可靠性的研究者有参考价值。","inspiration":"借鉴其使用LLM扮演不同利益相关者角色、在固定提示和温度下进行多维评分并与人类评分对照的方法，可迁移到经济金融中的政策解释或模型决策解释评估，例如评估信贷审批模型解释对贷款申请人的可理解性和信任影响。｜可应用于信贷审批歧视研究，让LLM扮演贷款申请人或监管者，评估不同XAI方法对信贷决策解释的公平性感知。｜设计：以LLM扮演贷款申请人，处理为不同XAI解释（如SHAP与LIME），结果变量为对决策的信任度和理解度评分，对照真实人类被试的评分数据。"}},{"id":"2609.10421","version":1,"title":"Emergency Department Revisit Quality Review Screening: Exploring Human Decision-Making and Artificial Intelligence Support","zh_title":"急诊科再就诊质量审查筛查：探索人类决策与人工智能支持","abstract":"Background: Emergency Department (ED) return visits are commonly reviewed for quality assurance, but are often limited (e.g., to revisits within 48-72 hours) to increase actionable finding yield while minimizing chart review burden. Those limitations may lead to missed quality improvement opportunities. Methods: We conducted an exploratory, retrospective study of randomly selected ED visits to a multihospital health system having an ED revisit within 1-14 days to the same health system. Given only each visit's primary diagnosis, raters (2-3 clinicians and GPT-4 large language model [LLM]) assessed characteristics of the diagnosis pairs, including the \"target\": whether a pair warranted further assessment. Informed by rater response analyses, an algorithm leveraging an LLM-populated knowledge graph (\"KGA\") was created to automatically screen for potentially concerning pairs, then preliminarily assessed. Results: 99 diagnosis pairs were included. GPT-4 responses poorly correlated to clinician raters, rating nearly all (94%) pairs as warranting follow-up (4.4-13.3 times more than clinicians). However, prompt engineering was minimal. Among clinician raters, revisit medical gravity was consistently significantly associated with the target, while a differential diagnosis/complication composite was significantly associated on unadjusted, but not adjusted (though less powered) analysis. The KGA achieved 83-100% positive predictive value for at least one clinician rater determining further assessment was warranted based on the diagnosis pair. Conclusion: These results can inform next steps for improving screening with LLMs like ChatGPT. Further research is warranted to validate this preliminary work's finding that the KGA may enable enhancing the scope and yield of screening without substantially increasing reviewer workload.","authors":["Jonathan A. Handler","Marlene I. Robles-Granda","Jacob E. Mefford","Jeremy S. McGarvey","Gregory S. Podolej","Colleen J. Klein","Matthew D. Dalstrom","William F. Bond"],"categories":["cs.CY","cs.AI"],"primary_category":"cs.CY","announce_type":"cross","date":"2026-09-10","first_seen":"2026-09-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.10421","pdf_url":"https://arxiv.org/pdf/2609.10421","source_feed":"cs.AI","score":7,"bucket":"pending","rubric_hits":["A1","B1","B4"],"tags":["LLM仿真","人类决策对照","医疗质量审查"],"reason":"用GPT-4模拟临床医生判断，并与人类医生对照，评估其可靠性，属于LLM仿真人…","model":"deepseek-v4-pro","scored_at":"2026-09-10T13:01:26","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-11","rank":9,"question":"急诊科复诊质量审查中，仅凭诊断对，人类临床医生和GPT-4如何判断是否需要进一步审查，以及能否用LLM知识图谱算法自动筛选可疑复诊对？","design":"回顾性研究，随机选取99对急诊初诊与14天内复诊的诊断对，仅提供主诊断，由2-3名临床医生和GPT-4分别评估诊断对特征（诊断成员价值、医学严重程度、鉴别诊断包含、并发症包含）及目标变量（是否需进一步审查）。基于评估结果，构建了利用LLM填充知识图谱的自动筛选算法（KGA），并初步评估其阳性预测值。","baseline":"2-3名临床医生（急诊医师和高级实践提供者）的独立评估作为人类基准，与GPT-4的评估进行对比。","findings":"GPT-4与临床医生的判断相关性差，将94%的诊断对评为需随访，是临床医生的4.4-13.3倍；临床医生中，复诊医学严重程度与目标变量显著相关，鉴别诊断/并发症复合指标在未调整分析中显著但调整后不显著；KGA对至少一名临床医生判定需进一步评估的诊断对实现了83-100%的阳性预测值。","reliability":"论文承认GPT-4提示工程极简，可能导致其过度判定；研究为探索性、样本量小（99对），且仅基于主诊断，未提供其他临床信息；KGA仅初步评估，需进一步验证。","relevance":"该研究直接使用GPT-4模拟临床医生决策并与人类对照，评估其可靠性，属于LLM仿真人类判断的实证研究，且包含真实人类基准，对关注LLM仿真在专业决策中有效性的研究者有参考价值，但场景为医疗而非经济金融。","inspiration":"借鉴其设计：用LLM模拟专家判断并与人类专家对照，同时构建基于LLM的自动筛选算法并与人类判断比较，以评估算法性能。｜可迁移到经济金融中的专业判断场景，如信贷审批中贷款员对借款人风险的评估、审计师对财务舞弊风险的判断、或政策分析师对经济指标异常的关注。｜研究设计：以信贷审批为例，选取一组贷款申请对（如初贷和短期内再贷），仅提供有限信息（如信用评分、收入），让LLM和人类信贷员分别判断是否需进一步审查，并构建基于LLM知识图谱的自动筛选算法，以人类信贷员的判断为基准评估算法阳性预测值，同时对比LLM与人类判断的一致性。"}},{"id":"2609.09609","version":1,"title":"Who You Are Adds Nothing Detectable to Where You Go Next: Sociodemographic Conditioning in LLM Next-Location Prediction","zh_title":"你是谁对你去哪里没有可检测的增益：LLM下一位置预测中的社会人口条件作用","abstract":"Large language models (LLMs) are increasingly used for individual next-location prediction, while sociodemographic conditioning is common in LLM-based travel simulation. Yet the incremental predictive value of sociodemographic attributes remains unclear. To directly test this contribution, sociodemographic records were linked with passively sensed mobility data from 5,000 Shenzhen residents to construct a closed-set benchmark in which models rank 100 candidate destinations. Each prediction instance is evaluated with and without age, gender, occupation and income, while holding mobility history, candidates and all other prompt content fixed. Results show that across four history lengths, the paired change in top-1 accuracy ranges from -0.8 to +0.5 percentage points, with no detectable gain from attributes. This result remains consistent when stay history is withheld, across alternative prediction times, in two additional LLMs and in a supervised reranker trained on the same benchmark. The null does not reflect a lack of model responsiveness to demographic information, as permuted attributes reduce LLM accuracy whereas correctly matched attributes do not improve it. A further asymmetry emerges in the reverse predictive direction, as pre-cut mobility trajectories recover income with an AUC of 0.708, while sociodemographic attributes contribute little to next-location prediction. Beyond demographic conditioning, candidate construction exerts a much larger influence on reported performance. Removing distance raises top-1 accuracy by 7.7 percentage points under proximity sampling but lowers it by 22.3 points under popularity sampling, with the reversal reproduced across all three LLMs. These results distinguish demographic association from incremental predictive usefulness and show that sampled next-location accuracy depends strongly on how candidate alternatives are constructed.","authors":["Xin Wang","Paraic Carroll","Kerry Nice","Sachith Seneviratne","Li Zhang"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2026-09-10","first_seen":"2026-09-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.09609","pdf_url":"https://arxiv.org/pdf/2609.09609","source_feed":"cs.CY","score":7,"bucket":"pending","rubric_hits":["A1","B1","B4"],"tags":["LLM仿真","人类移动预测","社会人口属性"],"reason":"用LLM预测个体移动行为，与真实人类数据对照，并批判性检验社会人口属性增益，可…","model":"deepseek-v4-pro","scored_at":"2026-09-10T13:01:22","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-10","rank":4,"question":"在个体下一位置预测任务中，加入年龄、性别、职业、收入等社会人口属性，能否在个体自身移动历史之外带来可检测的预测增益？","design":"该研究不是用LLM模拟人类被试，而是用LLM作为预测器。将5000名深圳居民的社会人口记录与被动感知的移动数据关联，构建封闭集基准：模型从100个候选目的地中排序。每个预测实例在保持移动历史、候选集和其他提示内容不变的情况下，分别在有/无社会人口属性条件下评估，比较top-1准确率变化。","baseline":"真实人类移动数据：5000名深圳居民的被动感知移动轨迹（含停留历史），以及关联的社会人口记录（年龄、性别、职业、收入）。","findings":"在四种历史长度下，加入社会人口属性导致的top-1准确率配对变化在-0.8到+0.5个百分点之间，无显著增益；该结果在无停留历史、不同预测时间、两个额外LLM和监督重排器中均一致。置换属性会降低准确率，但正确匹配的属性不提高准确率；反向预测中，移动轨迹可恢复收入（AUC=0.708），但属性对下一位置预测贡献甚微。候选集构建方式对性能影响远大于属性：邻近采样下去除距离提高7.7个百分点，流行度采样下降低22.3个百分点。","reliability":"论文未明确讨论失效条件，但指出结果可能受限于特定城市（深圳）、特定LLM和特定候选集构建方式；属性增益的缺失可能因移动历史已包含足够信息，或属性与移动行为关联弱于预期。","relevance":"该研究直接检验了LLM仿真中社会人口条件化的增量价值，使用真实人类移动数据作为基准，并批判性地发现属性无增益，对关注LLM仿真可靠性与偏差的研究者具有重要参考价值。","inspiration":"借鉴其配对设计：保持其他输入不变，仅增减社会人口属性，测量预测准确率变化，并辅以置换属性作为操纵检验。｜可迁移到信贷审批歧视研究：在LLM预测违约风险时，加入借款人性别、种族等属性，看是否在财务历史之外提高预测准确率，同时检验属性是否引发刻板印象。｜用LLM作为信贷审批模型，处理为在提示中加入/不加入借款人社会人口属性，结果变量为违约预测准确率，对照真实贷款数据（如Lending Club），比较有无属性时的AUC差异，并检查属性置换是否降低准确率。"}},{"id":"2608.16578","version":2,"title":"Physics of Agents: Statistical Mechanics Predicts Collective Behavior of AI Agents","zh_title":"智能体物理学：统计力学预测AI智能体的集体行为","abstract":"AI agents increasingly operate as part of interacting systems rather than in isolation. As agents exchange information and jointly make decisions, their interactions can improve collective reasoning but may also produce herding, polarization, or amplify shared biases. Understanding and predicting these collective dynamics is therefore important for designing effective and aligned multi-agent systems. Here, we study over 10,000 communities of language-model agents that repeatedly exchange messages and revise their opinions across objective mathematics questions and subjective political statements. Despite substantial diversity in possible behavior, the individual and group dynamics can be represented by three characteristic regimes: indifference, polarization, and consensus. AI agents start indifferent and build conviction as they interact. On objective questions, communication improves collective accuracy, while on subjective questions it often drifts group opinions toward the right in the political spectrum. We explain these observations with a statistical-mechanics formalism in which agents stochastically favor lower social pressure. Given only initial opinions, our model predicts individual trajectories, outperforms all standard baselines, generalizes to unseen community graphs, and reproduces the observed group archetype distributions. Our fitted model parameters reveal the mechanics underlying our key observations: i) communities operate below the critical social temperature, which explains conviction buildup; ii) attractive ties outweigh repulsive ones, which favors consensus; and iii) agents holding the correct answer exert the strongest pull, which drives truth-seeking. Overall, our results demonstrate that collective behavior of AI agents, like that of other complex systems, follows compact and predictive dynamical laws.","authors":["Batu El","Jinhee Paeng","Fatih Dinc","Shiye Su","Mete Erdogan","Aneesh Pappu","Haotian Ye","Wanjia Zhao","Surya Ganguli","James Zou"],"categories":["cs.AI","cs.MA","cs.SI"],"primary_category":"cs.AI","announce_type":"replace","date":"2026-09-10","first_seen":"2026-08-18","revised_at":"2026-09-10","abs_url":"https://arxiv.org/abs/2608.16578","pdf_url":"https://arxiv.org/pdf/2608.16578","source_feed":"cs.AI","score":6,"bucket":"other","rubric_hits":["D3"],"tags":["多智能体系统","意见动态","统计力学"],"reason":"用LLM agent群体模拟意见动态，但无真实人类数据对照，属社会模拟边界情形。","model":"deepseek-v4-pro","scored_at":"2026-09-10T13:01:38","error":null,"has_summary":false,"summary":null},{"id":"2609.10155","version":1,"title":"From Retrieval to Weights: Parametric Individualization of Small Language Models with Individual Text Corpora","zh_title":"从检索到权重：用个体文本语料对小语言模型进行参数化个体化","abstract":"We approach a cognitive simulation perspective on episodic and semantic memory in multiple-choice question answering by incorporating text from individual text corpora (ITC) into retrieval-augmented generation and DoRA fine-tuning. We web-crawl the search histories of 515 participants who answered 36 multiple-choice knowledge items and analyze a stratified subsample of 150 participants. For each participant, one DoRA adapter consolidates their ITC into a small language model (SLM) whose baseline correctness falls below the participants' lowest quartile. The adapter measurably writes the ITC into the weights: it fits its own participant's held-out text better than other participants' texts (dz =1.27), an individuality effect that increases with ITC size in rank order. On the generalized knowledge test, however, the adapter adds knowledge rather than alignment with the individual: log-loss match improves, whereas match accuracy under a bias-corrected PMI readout does not, and retrieval adds nothing on top. Our results demonstrate that ITCs can be consolidated into the weights of SLMs, an encouraging basis for individualized tutoring agents, and we discuss how to move from there toward a realistic simulation of episodic and semantic memory at the individual level.","authors":["Christoph Wigbels","Ali Abusaleh","Markus T. Jansen","Alexander Mehler","Markus J. Hofmann"],"categories":["cs.CL","cs.IR"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-10","first_seen":"2026-09-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.10155","pdf_url":"https://arxiv.org/pdf/2609.10155","source_feed":"cs.CL","score":6,"bucket":"other","rubric_hits":["D2","B1"],"tags":["认知模拟","个体化微调","记忆建模"],"reason":"用个体文本微调小模型模拟个人记忆，有真实人类数据对照，但目标是认知模拟而非社会…","model":"deepseek-v4-pro","scored_at":"2026-09-10T13:01:24","error":null,"has_summary":false,"summary":null},{"id":"2609.10253","version":1,"title":"DiSCo: A Distribution-First Steering and Cultural Prior Evaluation Framework for Measuring Cultural Preference Bias in LLMs","zh_title":"DiSCo：一种分布优先的引导与文化先验评估框架，用于测量大语言模型中的文化偏好偏差","abstract":"Large language models (LLMs) are increasingly deployed in globally used assistants, yet their default choices in culturally grounded everyday situations can systematically favour some cultures over others, affecting localisation, user trust, and equitable behaviour. Existing cultural benchmarks evaluate accuracy against a single \"correct\" answer, making it difficult to characterise an LLM's cultural preference prior when multiple culturally grounded responses are all valid; they also conflate default preferences with context-driven adaptation. We propose DiSCo, a distribution-first forced-choice evaluation framework that isolates default cultural priors and tests steerability via a four-level context gradient (C0--C3). Using DiSCo-Bench (304 items) derived from BLEnD spanning 12 cultures, we evaluate six diverse instruction-tuned LLMs. Default priors are heavily concentrated, with UK and US together absorbing approximately 35\\% of all selections despite representing only 2 of 12 cultures. Most critically, prompt-based steering consistently widens the selection gap between high- and low-resource cultures, and injecting explicit cultural facts produces negligible distributional disruption, confirming that cultural preference bias cannot be resolved through prompt-based personalisation alone.","authors":["Bhuvan Arora","Devesh Saraogi","Sravya Varada","Dhruv Kumar"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-10","first_seen":"2026-09-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.10253","pdf_url":"https://arxiv.org/pdf/2609.10253","source_feed":"cs.CL","score":6,"bucket":"other","rubric_hits":["D2"],"tags":["文化偏见","LLM评估","偏好分布"],"reason":"测量LLM的文化偏好偏差，属于把LLM本身当测量对象，非仿真人类被试，但涉及文…","model":"deepseek-v4-pro","scored_at":"2026-09-10T13:01:37","error":null,"has_summary":false,"summary":null},{"id":"2609.09150","version":2,"title":"Copying explains the collective behavior of AI agents in the wild","zh_title":"复制行为解释了野外AI代理的集体行为","abstract":"In June 2026, thousands of AI agents found that a small public wiki would accept edits from inside their sandboxes, and started using it to help one another pass a timed test. Each agent lived for about an hour and remembered nothing afterwards. Nobody asked them to cooperate, and the wiki had not been built for them. The complete record of what they wrote is public, and it is unusually informative, because it preserves not only what each agent wrote but what that agent could see before writing. We use it to follow the three decisions an agent had to make on arrival: where to write, what to call itself, and how to word its message. One rule governs all three. An agent takes an option with a probability close to the share of that option in what it can see, and the share that matters is the one on the page in front of it, then the one in the stream of recent edits, and only weakly anything older. Three minimal copying models, one per decision and with a single free parameter each, reproduce the heavy-tailed distribution of how many agents met on a page, the frequency of the pieces from which the agents built their names, and the patchwork of pages that are internally consistent and different from one another. Copying whatever the environment happens to show is enough to produce most of the collective structure of this population. It is also what makes such a population easy to steer, since whoever writes first, or writes while the others are quiet, sets the convention for everyone who comes later.","authors":["Giordano De Marzo","Nicola Albor\\'e","David Garcia"],"categories":["cs.MA","cond-mat.stat-mech","cs.CL"],"primary_category":"cs.MA","announce_type":"replace-cross","date":"2026-09-10","first_seen":"2026-09-09","revised_at":"2026-09-10","abs_url":"https://arxiv.org/abs/2609.09150","pdf_url":"https://arxiv.org/pdf/2609.09150","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["AI代理","集体行为","社会模拟"],"reason":"研究AI agent群体的集体行为，但无人类数据对照，属于社会模拟的边界情形。","model":"deepseek-v4-pro","scored_at":"2026-09-10T13:01:40","error":null,"has_summary":false,"summary":null},{"id":"2609.07920","version":2,"title":"Humans Introduce, Models Elaborate: Asymmetric Narrative Agency in Human-LLM Co-Writing","zh_title":"人类引入，模型展开：人机协作写作中的不对称叙事能动性","abstract":"Human-LLM co-writing is increasingly used for open-ended text generation, but much prior work focuses on final outputs rather than the interactional dynamics through which stories are produced. We study turn-based collaborative storytelling across three matched conditions: Human-Human (HH), Human-LLM (HA), and LLM-LLM (AA). Using a shared storytelling paradigm, we measure how agents align, introduce novel material, and influence narrative development through turn-level measures of valence adaptation, semantic novelty, transience, and resonance. Our results show that HA co-writing is not intermediate between HH and AA collaboration. Instead, it displays a distinctive asymmetry where humans tend to introduce more novel and persistent narrative material, while LLMs tend to elaborate and stabilize the existing context. These findings suggest that, in this setting, LLMs function less as human co-authors and more as adaptive narrative amplifiers that reshape how agency is distributed in collaborative writing.","authors":["Halfdan Nordahl Fundal","Yuri Bizzoni","Charlotte Gj{\\o}rup Bilde","Ida B{\\ae}kke Johannesen","Rebekah Baglini"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"replace","date":"2026-09-10","first_seen":"2026-09-09","revised_at":"2026-09-10","abs_url":"https://arxiv.org/abs/2609.07920","pdf_url":"https://arxiv.org/pdf/2609.07920","source_feed":"cs.HC","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["人机协作","叙事分析","LLM行为"],"reason":"研究人类与LLM协作叙事，非仿真人类被试，但涉及LLM行为与人类对照，属社会模…","model":"deepseek-v4-pro","scored_at":"2026-09-10T13:01:39","error":null,"has_summary":false,"summary":null},{"id":"2609.09901","version":1,"title":"Deep and shallow biases in language models","zh_title":"语言模型中的深层与浅层偏差","abstract":"Large language models often repeatedly select the same answer even when many alternatives are plausible. Prior work treats this concentration as bias, but it does not distinguish stable model preferences from responses that depend on a particular prompt wording. We introduce a bias depth score that measures both how strongly a model prefers its top answer under direct prompting and whether that answer survives scenario reframing. Across 4,442 opinion prompts and four large language models, only about a quarter of the concentrated preferences survive reframing. We call these persistent cases Deep biases, and the remaining prompt-dependent cases Shallow biases. Our results show that Deep biases are more often inherited from pretraining and preserved through SFT. Under both continued fine-tuning and prompt-based debiasing for diversity, Deep biases are consistently harder to remove than Shallow biases. Bias depth therefore separates stable learned biases from prompt-wording artifacts that single-prompt metrics conflate. Code, models, and data are available at deepbias.github.io.","authors":["An Vo","Vy Tuong Dang","Khai-Nguyen Nguyen","Emilio Villa-Cueva","Thamar Solorio","Anh Totti Nguyen","Daeyoung Kim"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-10","first_seen":"2026-09-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.09901","pdf_url":"https://arxiv.org/pdf/2609.09901","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["LLM偏差","观点稳定性","提示敏感性"],"reason":"研究LLM自身观点稳定性，非仿真人类被试，但涉及态度测量，属边界情形。","model":"deepseek-v4-pro","scored_at":"2026-09-10T13:01:24","error":null,"has_summary":false,"summary":null},{"id":"2609.09789","version":1,"title":"Pairit: A Platform for Live Experiments on Human-AI Collaboration","zh_title":"Pairit：人类与AI协作实时实验平台","abstract":"Organizational design in the era of artificial intelligence requires experimental methods that can test how human-AI groups coordinate, delegate, and make decisions. Programmable platforms coordinate live human-to-human sessions or real-time human-AI chat, but researchers cannot easily declare experiment protocols in which AI participants both communicate and act on shared work within one auditable configuration. Here we introduce Pairit, an online platform that facilitates the design, testing, and deployment of experiments that test human-AI organizational designs and interventions. Through a single YAML configuration file, researchers declare an executable experiment graph (pages, routing, randomization, matchmaking, chat, shared workspaces, server-hosted agents, surveys, timers, and custom HTML components) and combine any number of humans and AI agents in live sessions. We have validated the feasibility of the platform through multiple live deployments, including peer-reviewed published studies, capturing high-resolution process traces of communication, negotiation, and collaborative work in live human-AI dyads. By representing complex interactive protocols as standardized, auditable configuration files, Pairit provides reusable infrastructure for specifying, deploying, and sharing live human-AI organizational experiments.","authors":["Harang Ju","Sinan Aral"],"categories":["cs.HC","cs.AI"],"primary_category":"cs.HC","announce_type":"cross","date":"2026-09-10","first_seen":"2026-09-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.09789","pdf_url":"https://arxiv.org/pdf/2609.09789","source_feed":"cs.AI","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["人机协作","实验平台","组织设计"],"reason":"平台支持人类与AI协作实验，但非以LLM仿真人类被试，无人类数据对照，属社会模…","model":"deepseek-v4-pro","scored_at":"2026-09-10T13:01:22","error":null,"has_summary":false,"summary":null},{"id":"2608.08869","version":2,"title":"Are LLMs Positionally Consistent Ordinal Classifiers? A Systematic Evaluation","zh_title":"大语言模型是位置一致的序数分类器吗？一项系统性评估","abstract":"Large language models are increasingly used for ordinal classification, yet semantically equivalent changes to prompt organization can alter their predictions. We conduct systematic experiments to characterize positional bias from label order, demonstration order, and demonstration placement. First, we apply the three probes to ten frontier LLMs on a common ordinal-classification task; every model is sensitive to all three positional sources, showing that the problem is pervasive. Second, we vary eight prompt-, task-, and model-level factors across five datasets; accuracy and stability are often misaligned, and only lower scale cardinality consistently improves both. Third, we compare pointwise, pairwise, and listwise inference, alternative aggregation and debiasing methods, and joint configurations; the tested corrections do not provide a reliable remedy, while a comparison-based listwise formulation offers the best balance but transfers unevenly across models and bias sources. These findings show that positional robustness depends on the full system configuration rather than the model alone. Ordinal-classification systems should therefore be selected jointly for predictive performance and stability.","authors":["Yu Wang","Zhe Zhou","Menglin Liu","Ge Shi"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-10","first_seen":"2026-08-11","revised_at":"2026-09-10","abs_url":"https://arxiv.org/abs/2608.08869","pdf_url":"https://arxiv.org/pdf/2608.08869","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["LLM评测","序数分类","位置偏差"],"reason":"研究LLM序数分类的稳健性，属模型能力评测，不以人类行为为参照，不涉及人类仿真。","model":"deepseek-v4-pro","scored_at":"2026-09-10T13:01:40","error":null,"has_summary":false,"summary":null},{"id":"2608.15129","version":2,"title":"Left-Branching Transformers Excel at Right-Branching Languages: Data Shapes Word Order Preferences in Language Models","zh_title":"左分支Transformer擅长右分支语言：数据塑造语言模型的词序偏好","abstract":"We systematically compare word order preferences in decoder-only language models across 192 artificial languages and typologically diverse natural languages. On artificial languages, models exhibit a left-branching preference that aligns with neither natural language universals nor human word order learning biases. On natural languages, monolingual models show no clear base word order bias at small scales, but as data grows, a preference for right-branching subject-verb-object (SVO) languages emerges while SOV falls behind despite being the most frequent order cross-linguistically. This SVO advantage extends to multilingual models and correlates with language resource level and data quality rather than word order. Thus, the same architecture exhibits opposite preferences on artificial and natural languages, establishing that word order biases observed in practice are data-driven. Since highly-resourced languages are overwhelmingly SVO, these biases risk gradually reducing word order diversity, particularly in languages that productively use multiple word orders, with the widespread adoption of LLMs.","authors":["Varvara Arzt","Allan Hanbury","Terra Blevins"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-10","first_seen":"2026-08-18","revised_at":"2026-09-10","abs_url":"https://arxiv.org/abs/2608.15129","pdf_url":"https://arxiv.org/pdf/2608.15129","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["语言模型","词序偏好","NLP分析"],"reason":"研究语言模型词序偏好，属NLP能力分析，不以人类行为仿真为目标，无人类被试替代…","model":"deepseek-v4-pro","scored_at":"2026-08-18T13:04:37","error":null,"has_summary":false,"summary":null},{"id":"2608.24191","version":2,"title":"'Ghaib in Translation' aka Unseen Harm: Measuring Cross-Script Safety Inconsistency with 'Missed-in-Urdu' Scores in LLM Hate Speech Detection","zh_title":"翻译中的隐形伤害：用乌尔都语漏检分数衡量LLM仇恨言论检测的跨文字安全不一致性","abstract":"Urdu, the world's tenth most spoken language with 246 million speakers, remains almost entirely absent from mainstream LLM safety evaluation and nine years of WOAH proceedings. To investigate whether this absence has measurable consequences for content moderation reliability, five large language models, GPT-4o, Claude Sonnet 4.5, Gemini 2.5 Flash, Qwen-2.5, and Llama-3.1, were tested across six datasets spanning Nastaliq Urdu, Roman Urdu, English, and code-switched Urdu-English. Across the five Urdu-script datasets, label instability between original-script and English-translation classification ranged from 15.9% (Gemini 2.5 Flash) to 31.6% (Qwen-2.5), with a 'Missed-in-Urdu' rate, content flagged as harmful in English translation but passed as normal in the original script, ranging from 2.4% to 9.9% (median 4.3%). A complete enumeration of all 205 papers across nine ALW/WOAH editions via the ACL Anthology API confirms zero dedicated Urdu papers across the entire period. Results indicate that current LLMs provide uneven safety assurance across Urdu's script varieties, with smaller open-weight models showing substantially higher instability and missed-harm rates than frontier closed models.","authors":["Fawzia Zehra (Fuzzy)","Kara-Isitt","Sonal Khosla","Stephen Swift"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-10","first_seen":"2026-08-26","revised_at":"2026-09-10","abs_url":"https://arxiv.org/abs/2608.24191","pdf_url":"https://arxiv.org/pdf/2608.24191","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["LLM安全评测","跨语言一致性","仇恨言论检测"],"reason":"纯LLM安全评测，无人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-08-27T13:03:16","error":null,"has_summary":false,"summary":null},{"id":"2609.09363","version":1,"title":"Do LLMs Make More Mistakes If They Do Not Believe the Input Data?","zh_title":"当LLM不相信输入数据时，它们会犯更多错误吗？","abstract":"Large language models (LLMs) are prone to hallucinating or misinterpreting facts, which impairs their usability in retrieval-augmented generation or data-to-text systems. We analyse how faithfulness of LLMs to provided context depends on how plausible they perceive the context to be (context-memory conflict). To better identify error patterns, we make use of the increased difficulty of non-English and low-resource language text generation and input data based on local knowledge, only partially captured in models' parametric knowledge. We let the models generate text in English, Czech, Slovak and Upper Sorbian from factual (FA), counterfactual (CFA) and fictional (FI) RDF triples containing local Czech and Slovak data. Contrary to our expectations, we observe only a weak context-memory conflict on the human-annotated sample. For Kimi K3 as an LLM judge, which agrees well with human annotations on the sample, counterfactual inputs receive only slightly lower faithfulness scores than factual ones (-0.05 on a 1-5 scale). We also find that a suboptimal choice of LLM judge would lead to overestimating the strength of the context-memory conflict.","authors":["Peter Kochelka","Ale\\v{s} Manuel Pap\\'a\\v{c}ek","Vojt\\v{e}ch Dvo\\v{r}\\'ak","Ond\\v{r}ej Du\\v{s}ek"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-10","first_seen":"2026-09-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.09363","pdf_url":"https://arxiv.org/pdf/2609.09363","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["LLM忠实度","数据到文本生成","幻觉分析"],"reason":"研究LLM对输入数据的忠实度，属NLP能力评测，不以人类行为为参照系","model":"deepseek-v4-pro","scored_at":"2026-09-10T13:01:30","error":null,"has_summary":false,"summary":null},{"id":"2609.10092","version":1,"title":"RAP: Research Attention Prediction Reveals Target-Conditioned Evidence Acquisition Biases","zh_title":"RAP：研究注意力预测揭示目标条件下的证据获取偏差","abstract":"Large language models (LLMs) increasingly act as research agents, yet their ability to track shifts in research attention is difficult to evaluate because reviews and research ideas lack uniquely verifiable outcomes. We introduce Research Attention Prediction (RAP), a rolling benchmark covering 278 AI/ML fields and 1,390 episodes. At each cut-off, an LLM agent searches a temporally restricted arXiv corpus and predicts the next six months' paper shares across eight frozen research directions. Search generally helps, but all four diagnostic models perform worse than an exact-count exponentially weighted moving average (EWMA) baseline in compositional accuracy. We identify two linked bottlenecks. Under cumulative-history access, State carry-forward outperforms direct Forecast for all four diagnostic models; frozen-evidence replay links a shared component of this reversal to Forecast-oriented policies retrieving a smaller share of recent evidence. Even with exact historical activity, future-specific updating remains limited, with only GPT-5.5 plus reopened Search slightly surpassing EWMA. Fine-tuning on realised outcomes improves Qwen3-4B's forecast Spearman correlation by 0.105 on held-out fields at later origins, with gains also on change-rich episodes.","authors":["Yingqian Wu","Jingcong Liang","Siyuan Wang","Zhenfei Yin","Philip Torr","Junchi Yu","Zhongyu Wei"],"categories":["cs.AI","cs.CL"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-09-10","first_seen":"2026-09-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.10092","pdf_url":"https://arxiv.org/pdf/2609.10092","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["LLM代理","研究趋势预测","基准测试"],"reason":"LLM作为研究代理预测研究趋势，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-10T13:01:36","error":null,"has_summary":false,"summary":null},{"id":"2609.09226","version":1,"title":"Adaptive Entangled Game Modules in Artificial General Intelligence","zh_title":"人工通用智能中的自适应纠缠博弈模块","abstract":"We introduce a probability-wave framework for modeling the collective behavior of interacting adaptive agents, deriving testable eigenmodes through a generalized behavioral intelligence (GBI) nonlocal probability-wave equation. This framework captures a broad range of human intelligence behaviors with analytical mechanisms and offers an indirect method to examine the Liu-Chen-Ao (LCA) hypothesis of nonlocal entangled nerve fibers in the brain through collective trader behaviors. Our empirical analysis of Chinese intraday stock market data demonstrates that adaptive entangled game modes explain 82-94% (89% overall) of observed decision patterns, a sharp contrast to the predictions of neoclassical finance based on independent rational agents. Moreover, 2-12% of behaviors show adaption to intraday news, events, and environments, characterized by dual equilibrium states and abrupt reference point shifts, while purely independent modes occur in less than 5% of cases. These findings empirically support the LCA hypothesis, as observable trading behaviors reflect underlying brain mechanisms and internal intelligence decision-making in behavioral psychology. Our results highlight the necessity of incorporating adaptive entangled game modules into artificial general intelligence (AGI) architectures, addressing the limitations of conventional artificial neural network (ANN)-based AI, which relies on trillions of opaque parameters. By integrating ANN-based AI with probability-wave-based entangled-brain simulations, machine learning can enrich AGI foundation models (FMs) and facilitate the development of human-like processing units (HPUs) that leverage brain-inspired mechanisms. Such HPUs may ultimately create more compact, efficient, and robust AGI systems, particularly for embodied intelligence and robotics.","authors":["Haochen Li","Xinshuai Guo","Jingdong Ouyang","Wei Zhang","Leilei Shi"],"categories":["cs.AI","physics.soc-ph","q-bio.NC","q-fin.GN","quant-ph"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-10","first_seen":"2026-09-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.09226","pdf_url":"https://arxiv.org/pdf/2609.09226","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","概率波框架","AGI架构"],"reason":"研究多智能体集体行为建模，不涉及LLM仿真人类被试，无人类数据对照，属纯多智能…","model":"deepseek-v4-pro","scored_at":"2026-09-11T13:01:54","error":null,"has_summary":false,"summary":null},{"id":"2609.09774","version":1,"title":"Procedural Memory Under Change: Reuse and Interference in Controlled Web Tasks","zh_title":"变化下的程序性记忆：受控网络任务中的复用与干扰","abstract":"Procedural memory lets language agents reuse successful routines, but reuse presumes that a stored routine remains applicable. We study what happens when that presumption is deliberately violated. The study combines a retrospective, human-assisted interface-adaptation case from BrowserGym TimeWarp with controlled frozen-memory comparisons on synthetic shopping decisions. During the documented WebShop V1-V6 development path, interface-specific code was adapted while the separately stored high-level procedure was not reported to change; this phase does not constitute an autonomous memory-agent evaluation. In the controlled phase, an early pilot produced one task on which two memory conditions selected a more expensive item while the no-memory condition selected the reference minimum. Follow-up probes did not establish a recurring row-order or identity-binding pattern. We then tested four forms of mismatch: changed quantities, a different evidence representation, a conflict between local and global optimization, and distributed promotion evidence, across 32 formal cells. Each cell used one temperature-0 generation with the same local qwen3:8b configuration and no adaptive retry. Across these pairs, none of the predefined diagnostic interference signatures appeared on the tasks for which they were defined when current-task evidence was explicit and sufficient. The result identifies a tested region of non-interference: a procedural memory can be mismatched without becoming behaviorally disruptive. It does not establish general safety or a mechanism. The remaining question is which additional conditions turn applicability mismatch into observable, memory-caused error.","authors":["Yanze Cao"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-10","first_seen":"2026-09-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.09774","pdf_url":"https://arxiv.org/pdf/2609.09774","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["程序性记忆","语言代理","多智能体系统"],"reason":"研究语言代理的程序性记忆复用与干扰，属多智能体系统内部机制，不涉及人类行为仿真…","model":"deepseek-v4-pro","scored_at":"2026-09-10T13:01:33","error":null,"has_summary":false,"summary":null},{"id":"2609.09882","version":1,"title":"Scored vs. Generated Readouts in Behavioral Language Models: An Empirical Study of Elicitation Format","zh_title":"行为语言模型中评分式与生成式读出：诱导格式的实证研究","abstract":"Language models fine-tuned on customer behavior can predict outcomes and generate explanations, but these readouts are often treated as interchangeable. Holding model checkpoint and prompt content fixed, we compare probabilities obtained by scoring answer tokens with predictions generated after a written rationale. Across 13 model-domain cells covering four retail tasks in three markets, including two using fully public data and checkpoints, the scored readout ranks outcomes more accurately in 12 of 13 cells (two-sided sign test, p approximately 0.003), by 1.5 to 14.5 points in area under the receiver operating characteristic curve (AUC). Paired bootstrap confidence intervals exclude zero in every newly measured cell. The gap varies with task-specific supervision and mismatch between training and serving formats, ranging from -2.2 points for an untuned base model to +13.7 for rationale-format supervision. Analysis of approximately 9,000 rationales identifies two correlates: reduced reliance on the dominant predictive feature and convergence on stock formulations. Probability saturation does not track the gap. A third readout, eliciting a probability before any verdict, improves calibration (Brier score from 0.47 to 0.15) while ranking within noise of scoring, but only for outcome rates represented in training; it is worse than scoring when the scored head is already calibrated. We interpret these differences through the objectives matched by each readout, identify training choices that narrow the gap, and propose retaining generated rationales while sourcing ranking from the scored head.","authors":["Touchapon Kraisingkorn","Krittin Pachtrachai","Wachiravit Modecrua"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-10","first_seen":"2026-09-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.09882","pdf_url":"https://arxiv.org/pdf/2609.09882","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["LLM行为预测","读出格式","模型校准"],"reason":"研究LLM行为预测的读出格式，非仿真人类被试，无人类数据对照，属模型能力评测。","model":"deepseek-v4-pro","scored_at":"2026-09-10T13:01:34","error":null,"has_summary":false,"summary":null},{"id":"2609.10060","version":1,"title":"Reference-Based Bias Detection in LLMs via Relative Representations of Hidden States","zh_title":"基于隐藏状态相对表示的LLM参考基准偏见检测","abstract":"Existing bias auditing methods typically rely on model outputs, requiring costly benchmarks or judge models and potentially missing internal shifts that never appear in generated text. We propose a reference-based method that audits bias in hidden-state representations across related model variants, for example before and after fine-tuning. Because fine-tuning reshapes representation geometry, absolute hidden states are not directly comparable, so we encode each sentence by its similarities to a fixed set of anchor sentences, yielding relative representations in a shared comparison space. There we measure how target groups shift in their association with positive and negative attributes, a quantity we call the Representational Bias Shift $\\Delta B$. Across three model families and the WildGuardMix, DecodingTrust and ToxiGen benchmarks, $\\Delta B$ correlates with output-level bias change in 15 of the 18 settings we test, reaching $|r| = 0.84$ ($p < 0.001$) under full fine-tuning and becoming more model-dependent under parameter-efficient adaptation. Thresholding $\\Delta B$ detects checkpoints whose bias increased with ROC AUC between $0.65$ and $0.99$, and on WildGuardMix and DecodingTrust it separates them better than a SEAT-based baseline for all three families. $\\Delta B$ is also stable under changes to the anchor set, attribute sets and target templates. Our method requires no task-specific evaluation data and audits a model in about three minutes, using $3$-$50\\times$ less compute than the output-level benchmarks considered here. We view it as complementary to output-based auditing rather than a replacement for it.","authors":["Marek Jeli\\'nski","Jan Dubi\\'nski","Maciej Chrabaszcz","Sebastian Cygert"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-10","first_seen":"2026-09-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.10060","pdf_url":"https://arxiv.org/pdf/2609.10060","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["偏见检测","模型审计","表示学习"],"reason":"论文是LLM偏见检测方法，不涉及人类仿真或行为对照，属于模型审计而非人类被试替…","model":"deepseek-v4-pro","scored_at":"2026-09-10T13:01:36","error":null,"has_summary":false,"summary":null},{"id":"2609.10132","version":1,"title":"Context operations to architecture modelling output from large language models and evaluation criteria for their use in systems engineering design","zh_title":"面向大语言模型架构建模输出的上下文操作及其在系统工程设计中的评估标准","abstract":"The development of generative artificial intelligence resources enables opportunities of speeding up systems and engineering design work. This contribution introduces a framework of formal operations for assembling context in LLM-based engineering design. This framework involves the assembly of modular context units, including policy prompts, reference units with persistence, and user questions with prompt vectoring. This approach enables the systematic structuring of interactions with generative models. A formal method for evaluating modelling-as-code LLM outputs is also presented, which enables the evaluation of compliance to intent from LLM answers and thereby asses the support from LLMs for systems architecture modelling.","authors":["Vinicius Kaster Marini","Petter Krus"],"categories":["eess.SY","cs.AI","cs.SE","cs.SY"],"primary_category":"eess.SY","announce_type":"cross","date":"2026-09-10","first_seen":"2026-09-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.10132","pdf_url":"https://arxiv.org/pdf/2609.10132","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["系统工程","LLM应用","建模评估"],"reason":"论文聚焦于系统工程设计中的LLM输出建模，不涉及人类行为仿真或对照。","model":"deepseek-v4-pro","scored_at":"2026-09-11T13:01:57","error":null,"has_summary":false,"summary":null},{"id":"2608.30107","version":2,"title":"AtlasNLP: A Country-Aware Atlas of Dataset Representation in NLP","zh_title":"AtlasNLP：NLP数据集表示的国家感知图谱","abstract":"Understanding which countries are represented in NLP datasets is essential for identifying gaps, targeting data collection, measuring progress, and informing AI policy. However, geographic metadata is very rarely available, and country-level representation is often hidden behind broad language-level claims. We introduce AtlasNLP, a country-aware atlas of over 13,000 NLP dataset records across normalized NLP task categories, tracking both the populations represented and where datasets are produced. AtlasNLP includes AtlasNLP-Gold, a human-curated reference set, and AtlasNLP-Core, an ACL-derived large-scale collection. Using this resource, we show that (1) dataset coverage is highly uneven across countries and tasks; (2) dataset production and representation are geographically asymmetric; and (3) language coverage does not imply geographic representation. These findings reveal blind spots in current dataset documentation practices and motivate more explicit geographic metadata for country-aware NLP evaluation.","authors":["Joan Nwatu","Tsedeniya Solomon Amare","Longju Bai","Bontu Fufa Balcha","Zayd Bashir","Angana Borah","Zara Burzo","Yubin Choi","Naihao Deng","Samika Gupta","Michel Faloughi","Claude Kwizera","Ziqiao Ma","Cynthia Yacel Fuertes Panizo","Ellie Seehorn","Hui Shen","Jiayi Tang","Zesen Zhao","Boyuan Zheng","Rada Mihalcea"],"categories":["cs.CL","cs.AI","cs.CY"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-10","first_seen":"2026-09-01","revised_at":"2026-09-10","abs_url":"https://arxiv.org/abs/2608.30107","pdf_url":"https://arxiv.org/pdf/2608.30107","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["数据集地理代表性","NLP评估","数据文档"],"reason":"论文关注NLP数据集的地理代表性，不涉及LLM仿真人类被试或行为对照。","model":"deepseek-v4-pro","scored_at":"2026-09-11T13:01:54","error":null,"has_summary":false,"summary":null},{"id":"2609.06646","version":2,"title":"When Does a Laugh Begin? Structured Annotator Disagreement in Temporal Laughter Localization","zh_title":"笑声何时开始？时间笑声定位中的结构化标注分歧","abstract":"Annotators routinely disagree on laughter boundaries and subtle chuckles, yet temporal laughter localization typically evaluates against a single reference annotation. We show that this disagreement is structured rather than random noise. Re-annotating the SMILE-Temporal benchmark (672 videos, 1,683 events) with 3-5 annotators per video (alpha = 0.757), we find systematic patterns: disagreement is 1.73 times larger at offsets than onsets, far more common for chuckles than full laughs (77% vs. 20%), and predictable from event attributes (AUC = 0.831). Evaluating against a single annotator breaks down under this structure: system scores shift by 0.246 F1 depending on the chosen ground truth, correctly ranking systems only 69.7% of the time (vs. 80% against all annotators). We propose a disagreement-calibrated evaluation that scores predictions against the full annotator distribution using conformally calibrated tolerance bands (wider at offsets, 0.727s, than onsets, 0.5s). The per-annotator annotations and analysis code are available at https://github.com/WSCSports/MTLLFM-temporal-laughter-localization.","authors":["Eyal Hanania","Daniel Arkushin","Naveh Ayal","Jonathan Benvenisti","Amos Bercovich","Elie Zemmour","Sahar Froim"],"categories":["cs.CV","cs.AI"],"primary_category":"cs.CV","announce_type":"replace-cross","date":"2026-09-10","first_seen":"2026-09-09","revised_at":"2026-09-10","abs_url":"https://arxiv.org/abs/2609.06646","pdf_url":"https://arxiv.org/pdf/2609.06646","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C2"],"tags":["笑声定位","标注分歧","计算机视觉"],"reason":"研究笑声时间定位的标注分歧，不涉及LLM仿真人类被试，属于计算机视觉任务。","model":"deepseek-v4-pro","scored_at":"2026-09-10T13:01:26","error":null,"has_summary":false,"summary":null},{"id":"2609.08027","version":2,"title":"Delusions and Harms Associated with AI Chatbot Use: Early Evidence from 185 Real-World Reports","zh_title":"与AI聊天机器人使用相关的妄想与伤害：来自185份真实世界报告的早期证据","abstract":"Importance: Reports have raised concerns that AI chatbots may validate or elaborate delusional beliefs, respond inappropriately to suicidal ideation, and contribute to mental health harms, but real-world data on reported harms remain limited. Objective: To characterize psychopathological features, chatbot behaviors, timing, and outcomes in first- and second-hand accounts of mental health harm linked with AI chatbot use. Design: Cross-sectional secondary analysis of deidentified online survey responses gathered between August 7, 2025, and February 2, 2026. Main Outcomes and Measures: The primary quantitative outcome was the presence of delusional beliefs, coded by paired raters with relevant clinical experience. Additional variables included reason for chatbot use, current episode features, delusional content, chatbot validation of beliefs, harms, social and occupational consequences, healthcare use, and timing. Results: 95 first-hand and 90 second-hand accounts were analyzed. Median age was 35.0 (IQR 27.0 - 45.0). Raters coded descriptions consistent with delusional beliefs in 102 reports (55.1%), with chatbot validation of beliefs in 50/102 (49.0%). Common outcomes included isolation, relationship breakdown, hospital admission, job loss, and financial loss. Four second-hand reports described death by suicide. Conclusions: In this self-selected convenience sample, AI-chatbot-associated harms were frequently described in relation to delusional beliefs, perceived chatbot validation, intensive use, and substantial social, occupational, and clinical consequences. Because reports were retrospective, unverified, and collected from individuals seeking to report harm, our findings should be interpreted as preliminary signal detection rather than as suggesting prevalence or providing evidence of causality. Prospective surveillance and trajectory-based safety evaluations are needed.","authors":["Hamilton Morrin","Vinitha Soundararajan","Thomas Cheliotis-James","Boris Warszawski","Joshua Fakulujo","Zeqi Jia","Etienne Brisson","Thomas A. Pollak"],"categories":["cs.HC","cs.CY"],"primary_category":"cs.HC","announce_type":"replace","date":"2026-09-10","first_seen":"2026-09-09","revised_at":"2026-09-10","abs_url":"https://arxiv.org/abs/2609.08027","pdf_url":"https://arxiv.org/pdf/2609.08027","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["AI聊天机器人","心理健康","真实世界报告"],"reason":"研究真实用户与聊天机器人互动导致的心理伤害，非用LLM仿真人类被试，无实验对照。","model":"deepseek-v4-pro","scored_at":"2026-09-10T13:01:26","error":null,"has_summary":false,"summary":null},{"id":"2609.05442","version":2,"title":"Role differentiation as ignition of a collective information engine: Structuration in Agent Populations","zh_title":"角色分化作为集体信息引擎的点火：智能体群体中的结构化","abstract":"Informational active matter shows how measurement-informed decisions produce collective order, so far in systems that reach consensus. We design collective information engines structured by differentiation instead, and construct a minimal instance using anti-coordination games where differentiated role information has value. Within many coexisting games, agents infer their role from a noisy social signal grounded in a persistent identity, and role-following action feeds back into that signal, which shapes the incentive to follow roles. Resources accrued through coordinated role-play combine with identity variability to reinforce the schemas that generated them. The model thereby operationalizes Sewell's duality of schemas and resources in Structuration, a resolution to structure--agency debates across social science. The engine ignites when a social loop gain---the product of identity persistence, cognitive capacity, channel fidelity, and schema strength---exceeds one. For a repertoire of such schemas, roles emerge with increasing gain in a bifurcation cascade whose functional form is fixed by the repertoire's eigenvalue spectrum, ranging from monitorable logarithmic sequences to avalanches that arrive without warning. Resource accumulation supplies the fitness of a replicator dynamics on schema strengths, which selects the cascade type endogenously. Subcritical identity covariance reveals that type before onset, enabling early detection, while feedback channel parameters bias which type is selected. Platform design then becomes a control lever to throttle emergent coordination. This theory grounds distributional AGI takeoff in a mechanism and provides a monitor-based solution. Joining game theory, collective dynamics, and information engines, we open a route to an information thermodynamics of agent populations.","authors":["Maximilian Puelma Touzel"],"categories":["physics.soc-ph","cs.GT","cs.MA"],"primary_category":"physics.soc-ph","announce_type":"replace-cross","date":"2026-09-10","first_seen":"2026-09-09","revised_at":"2026-09-10","abs_url":"https://arxiv.org/abs/2609.05442","pdf_url":"https://arxiv.org/pdf/2609.05442","source_feed":"cs.MA","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","博弈论","信息引擎"],"reason":"纯多智能体系统研究，agent间协作博弈，无LLM仿真人类被试，无人类数据对照","model":"deepseek-v4-pro","scored_at":"2026-09-10T13:01:40","error":null,"has_summary":false,"summary":null},{"id":"2609.09764","version":1,"title":"SocialRL: Refining LLMs' Social Intelligence through Multi-turn Reinforcement Learning and Reward Design","zh_title":"SocialRL：通过多轮强化学习与奖励设计提升大语言模型的社交智能","abstract":"Social intelligence enables agents to read social context, infer intent, and adapt over sustained dialogue. As language models become autonomous collaborators, it is central to building effective and trustworthy human-AI interaction. Existing reinforcement learning methods optimize single-turn utterances and sparse outcome rewards, producing short-sighted policies that struggle to manage goal-relationship tensions across multi-turn interactions. We propose SocialRL, a multi-turn reinforcement learning framework addressing both challenges. First, we apply multi-turn reinforcement learning using PPO that propagates delayed outcome rewards back to each turn, enabling long-horizon planning. Second, we design six process reward dimensions capturing the goal-relationship trade-off, including goal advancement, relational attunement, contextual coherence, etc. A reward model dynamically generates fine-grained scoring criteria for each dimension, while a stage-aware weight schedule prioritizes relationship-building in early turns, goal advancement mid-way, and balanced closure late. Across multiple social-dialogue benchmarks, SocialRL improves Goal Achievement by an average of 9.2 percentage points over the corresponding Base models. These results demonstrate the effectiveness of SocialRL across synthetic and real social scenes, as well as standard and challenging social scenarios.","authors":["Jianing Wang","Xintao Wang","Aili Chen","Jie Shi","Hongcheng Guo","Jun Gao","Wenxuan Zhao","Chengkun Lang","Yuanli Guo","Yanghua Xiao"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-10","first_seen":"2026-09-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.09764","pdf_url":"https://arxiv.org/pdf/2609.09764","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["社交智能","强化学习","对话系统"],"reason":"该研究旨在提升LLM在社交对话中的表现，属于角色扮演对话优化，无人类行为对照或…","model":"deepseek-v4-pro","scored_at":"2026-09-10T13:01:33","error":null,"has_summary":false,"summary":null},{"id":"2609.10052","version":1,"title":"Direct Diversity Optimization for Diverse Successful Trajectories in Preference Post-Training","zh_title":"偏好后训练中多样化成功轨迹的直接多样性优化","abstract":"LLM agents for sequential decision tasks are often post-trained with trajectory-level outcome labels, but such labels provide little supervision for preserving multiple successful branches from the same decision state. We study this problem as successful strategy coverage: how broadly a model realizes distinct successful strategies under a fixed rollout budget. We present Direct Diversity Optimization (DDO), an offline post-training method that combines Divergence-Tree Collection (DTC) with the Reference-Relative Target-Odds Objective (RTO). DTC constructs state-aligned branch sets rooted at shared decision states, and RTO trains the model to match reference-relative targets over successful alternatives. DDO achieves the strongest task success and successful strategy coverage among the compared post-training methods across BabyAI, BabaIsAI, and WebShop. It also achieves the highest recovery rate after local action replacement and higher task success and coverage than successful-only imitation and decoding-time diversification controls.","authors":["Junwon Ko","Dong-Jae Lee","Minchan Kwon","Sunghyun Baek","Junmo Kim"],"categories":["cs.CL","cs.AI","cs.LG"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-10","first_seen":"2026-09-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.10052","pdf_url":"https://arxiv.org/pdf/2609.10052","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["LLM智能体","策略多样性","后训练"],"reason":"研究LLM智能体在决策任务中的策略多样性，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-10T13:01:35","error":null,"has_summary":false,"summary":null},{"id":"2609.10539","version":1,"title":"IdeaAMBIG: Benchmarking Implementation-Critical Gaps in Research-Idea Specifications","zh_title":"IdeaAMBIG：基准测试研究想法规范中实现关键缺口","abstract":"A research idea may be novel, coherent, and scientifically plausible, yet its proposed method may remain insufficiently specified for faithful implementation. We study the codification readiness of implementation-facing research-method specifications, defined by whether they provide sufficient methodological information for a competent implementer or coding agent to construct the intended method without unsupported assumptions. We construct evidence-grounded specifications and their supported resolutions from papers, codebases, issue threads, and reproduction artifacts. We introduce IdeaAMBIG, a benchmark of 660 evidence-grounded instances: 163 real-world gaps from reproducibility reports and GitHub issues, and 497 controlled synthetic gaps injected into codification-ready references. IdeaAMBIG evaluates three capabilities: codification-readiness assessment, defect localization, and clarification action generation. Defect localization receives only the specification, whereas clarification additionally receives the annotated defect. Across 13 LLMs, the best model achieves 9.6% Macro Defect Recovery Rate on real-world instances but 80.6% Macro Clarification Action Success Rate when given the defect. In an oracle study, supplying the gold resolution raises the downstream codification-ready rate from 14% to 98%. Across all evaluated models, defect localization is the main bottleneck, with stronger clarification given the defect.","authors":["Yiling Ma","Yilun Zhao","Sihong Wu","Manasi Patwardhan","Arman Cohan"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-10","first_seen":"2026-09-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.10539","pdf_url":"https://arxiv.org/pdf/2609.10539","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["LLM评测","基准测试","研究规范"],"reason":"论文是LLM能力评测基准，不涉及人类仿真或行为对照","model":"deepseek-v4-pro","scored_at":"2026-09-10T13:01:38","error":null,"has_summary":false,"summary":null},{"id":"2609.09372","version":1,"title":"What Does MMLU Actually Measure? A Psychometric Audit of Difficulty Structure in Aggregate Benchmark Scores","zh_title":"MMLU实际测量什么？聚合基准分数中难度结构的心理测量审计","abstract":"Although MMLU is widely adopted as a benchmark for calibrating general AI capabilities, we psychometrically demonstrate that its aggregate score primarily evaluates a model's factual retrieval capacity rather than its reasoning ability. By calibrating item difficulty for 1,000 open-weights language models over 14,042 MMLU test items using Item Response Theory, we show that evaluating both abilities via a single test is inherently flawed. Difficulty is then regressed on a deterministic, text-extractable framework of structural complexity. Applying a joint Wald test with subject-clustered covariances demonstrates that the MMLU conflates fundamentally separable constructs. The mapping from structural complexity to difficulty is not invariant across the benchmark's STEM and non-STEM partitions. This finding has practical consequences. Aggregate leaderboard ranks track non-STEM accuracy more closely than STEM accuracy, so selecting a Top-50 model on the aggregate for a reasoning-intensive deployment displaces roughly 22% of the STEM-appropriate choices. Furthermore, when controlling for the multiple-choice guessing floor natively inside the response model, we find that higher-ability models continue to degrade more steeply under increased reasoning depth. The MMLU aggregate therefore weights retrieval capacity and reasoning stability unequally, inadvertently favoring models optimized for retrieval. We release our deterministic framework as a reproducible auditing instrument and recommend disaggregated reporting.","authors":["Dana Paquin","Riddhiman Jain"],"categories":["math.NT","cs.CL"],"primary_category":"math.NT","announce_type":"cross","date":"2026-09-10","first_seen":"2026-09-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.09372","pdf_url":"https://arxiv.org/pdf/2609.09372","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["基准评测","心理测量","模型能力"],"reason":"纯NLP基准评测，无人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-10T13:01:30","error":null,"has_summary":false,"summary":null},{"id":"2609.09790","version":1,"title":"LogiScope-VQA: Benchmarking Vision-Language Models for Logistics Hazard Identification in Industrial Scenarios","zh_title":"LogiScope-VQA：面向工业场景物流危险识别的视觉语言模型基准测试","abstract":"Large Multimodal Models (LMMs) large-scale deployment in industrial warehouse settings specifically necessitates that models exhibit human-expert-level hazard-oriented perception, understanding, and reasoning capabilities. However, the scarcity of real industrial data, tightly coupled to commercial terms, significantly hampers further advancement. To bridge this gap, we curate LogiScope-VQA to investigate the practical applicability of mainstream LMMs in real-world logistics operations. LogiScope-VQA comprises 2,476 images and 2,918 videos primarily sourced from real-world logistics parks, along with 10,274 VQAs meticulously curated and validated by human annotators. Grounded in 18 core objects and 20 risk types, we devise 39 subtasks aligned with three principal themes: industrial element perception, warehouse knowledge understanding, and potential risk reasoning. Furthermore, we incorporate dynamic thinking-budget configurations and dual-dimensional risk bias analyses to elucidate the properties of LMMs. Extensive experiments unveil that even powerful proprietary models, including GPT-5.5, Gemini-3.1-Pro, and Claude-Opus-4.7, exhibit a significant gap relative to human performance. The unique challenge of jointly integrating perception, understanding, and reasoning for hazard identification poses substantial headroom for further improvement on LogiScope-VQA. We additionally reveal the pervasive security bias issue that impedes LLMs' practical deployment in real-world settings. The industrial dataset is publicly available under the CC BY-NC-SA 4.0 license.","authors":["Hanjing Zhou","Mingze Yin","Ying Lian","Jun Ma","Chang-Yu Hsieh","Yanbing Zhou"],"categories":["cs.CV","cs.AI","cs.CL"],"primary_category":"cs.CV","announce_type":"cross","date":"2026-09-10","first_seen":"2026-09-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.09790","pdf_url":"https://arxiv.org/pdf/2609.09790","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C2"],"tags":["视觉语言模型","工业安全","基准测试"],"reason":"论文是工业场景下视觉语言模型的危险识别基准测试，属于计算机视觉与安全应用，不涉…","model":"deepseek-v4-pro","scored_at":"2026-09-11T13:01:57","error":null,"has_summary":false,"summary":null},{"id":"2609.10016","version":1,"title":"MetroLLM-Bench: Evaluating Language Models as Transit Kiosk Runtimes","zh_title":"MetroLLM-Bench：评估语言模型作为交通售票机运行时","abstract":"We introduce MetroLLM-Bench, a 955-case benchmark for testing language models as the policy layer of a transit kiosk. It covers six real metro systems, ranging from 37 to 414 stations, and eleven categories that include routing, fare calculation, disruptions, accessibility, and adversarial input. In each case, the model must call structured tools and submit a machine-renderable terminal state containing an outcome, a per-ticket fare quote when applicable, and a kiosk action. Fourteen deterministic scoring components form Tier 1; eight semantic-quality components form Tier 2, six of which use a language-model judge. We report Tier 1 and the combined score of both tiers. A stratified 75/25 split reserves 717 cases for training-data generation and 238 for held-out evaluation. We evaluate twenty-six models from six vendors, of which twenty-three are ranked. On the held-out partition, a 4B Qwen 3.5 student trained through parameter-efficient fine-tuning (PEFT) exceeds both GPT-5.6 tiers on Tier 1 (91.3 against 90.6 and 90.0) and matches GPT-5.4 full at maximum reasoning effort (91.4), with a 2.6 GB Q4_K_M footprint. Larger 9B and 27B students provide no further Tier 1 improvement over the 4B student at this training scale. Across the four Qwen sizes, the PEFT gain over the corresponding base model decreases from +7.03 points at 2B (three training seeds) to -0.91 at 27B; every seed shows the same direction at every size. A deterministic rule-based baseline reaches 84.6 on Tier 1, with the remaining language-model advantage concentrated in policy adaptation, compound scenarios, accessibility, and temporal reasoning. Muse Glimmer 30B leads the composite ranking, and serving configuration alone moves the Qwen 3.5-to-3.8 comparison by 2.7 Tier 1 points. The benchmark, harness, reproduction guide, and fine-tuned students are released at https://github.com/continker/metrollm-bench.","authors":["Remco Hendriks (Continker)"],"categories":["cs.LG","cs.AI","cs.CL"],"primary_category":"cs.LG","announce_type":"cross","date":"2026-09-10","first_seen":"2026-09-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.10016","pdf_url":"https://arxiv.org/pdf/2609.10016","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C2"],"tags":["LLM评测","交通系统","工具调用"],"reason":"评估LLM作为交通售票机策略层，属于自动化系统仿真，不涉及人类行为对照或仿真被…","model":"deepseek-v4-pro","scored_at":"2026-09-10T13:01:35","error":null,"has_summary":false,"summary":null},{"id":"2609.10036","version":1,"title":"Belief-State Engine: Augmenting LLMs for Principled Planning Under Partial Observability","zh_title":"信念状态引擎：增强大语言模型在部分可观测性下的原则性规划","abstract":"Large language model agents produce fluent action sequences across a wide range of tasks, yet they fail in characteristic ways once the environment becomes partially observable. Ambiguous feedback pushes them into premature commitments. A single informative observation can collapse their uncertainty onto the wrong hypothesis. Policies drift as the history grows. We trace these symptoms to a common structural cause. An LLM agent, as commonly deployed, is a history-conditioned policy with no explicit belief over hidden state. We propose an architectural fix. The Belief-State Engine (BSE) is an inference module placed outside the LLM. It maintains a Bayesian posterior over the latent states of a given POMDP (Partially Observable Markov Decision Process) model, and at each decision step it exposes only that posterior to the LLM. The raw action-observation log is not shown. We set out a minimal four-axiom specification of what a belief-consistent internal state must satisfy, and prove that the LLM paired with the BSE is a sound Markov policy on the belief MDP induced by the underlying POMDP. It therefore inherits the Bellman optimality guarantees of classical POMDP theory, provided the LLM is never exposed to the raw history. We evaluate the architecture on the Tiger POMDP and a red-team attack-graph task, against six baselines: a reactive LLM, Chain-of-Thought, ReAct, a natural-language belief tracker, QMDP, and POMCP. Across both domains, the BSE-augmented agent improves task return, belief calibration, and decision consistency. Ten targeted ablations isolate the contribution of each architectural choice confirms that the effect is not specific to any one model. Code, environment specifications, prompt templates, and seed logs accompany this paper.","authors":["Arnab Chattopadhayay","Debdipta Halder"],"categories":["cs.AI","cs.LG","cs.RO"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-10","first_seen":"2026-09-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.10036","pdf_url":"https://arxiv.org/pdf/2609.10036","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C2"],"tags":["LLM规划","POMDP","智能体"],"reason":"研究LLM在部分可观测环境中的规划，属机器人/游戏仿真，不涉及人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-09-10T13:01:35","error":null,"has_summary":false,"summary":null},{"id":"2609.09348","version":1,"title":"Smart Adaptive Computing Across the Continuum: LLMs in IoT-Edge-Cloud Resource Management","zh_title":"跨连续体的智能自适应计算：LLM在物联网-边缘-云资源管理中的应用","abstract":"Managing resources across IoT, edge, and cloud layers calls for continuous, context-aware decisions under constraints that rarely stay fixed. Deep reinforcement learning (DRL) handles this class of problems well, and large language models (LLMs) are increasingly used to augment DRL pipelines, yet the architectural relationship between the two is seldom made explicit. We build on Wang et al.'s taxonomy of Continuum Orchestration Systems employing DRL techniques and extend it with two further dimensions. The AI Augmentation Paradigm measures how LLMs are exploited, while the Feedback channel captures whether and through which system path the execution feedback returns to the LLM in order to close the MAPE control loop at the LLM Orchestration layer. We apply this taxonomy to six recent system architectures and find a common gap, as none combines full LLM orchestration with full agent-layer feedback in a Cloud Continuum setting. We relate this gap to a missing cross-tier feedback abstraction, bridging the incommensurable per-tier signals and the LLM Orchestrator.","authors":["Antonino Vaccarella","Lanpei Li","Vincenzo Lomonaco","Massimo Coppola"],"categories":["cs.DC","cs.AI","cs.MA"],"primary_category":"cs.DC","announce_type":"cross","date":"2026-09-10","first_seen":"2026-09-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.09348","pdf_url":"https://arxiv.org/pdf/2609.09348","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["资源管理","深度强化学习","多智能体系统"],"reason":"研究LLM增强DRL进行资源管理，属多智能体系统协作，不涉及人类行为仿真或对照。","model":"deepseek-v4-pro","scored_at":"2026-09-10T13:01:29","error":null,"has_summary":false,"summary":null},{"id":"2609.09560","version":1,"title":"The Vibe Shift in Software Engineering: Evaluating AI-Led Conversational Programming for Performance, Cognition, and Responsible Adoption","zh_title":"软件工程中的氛围转变：评估AI主导的对话式编程在性能、认知与负责任采用方面的表现","abstract":"This study evaluates Vibe Coding, an emerging AI-led conversational programming paradigm that enables developers to generate software through natural-language interaction with large language models. Using a mixed-methods design, the study assessed performance efficiency, cognitive implications, and responsible adoption in comparison with traditional and AI-assisted coding environments. Thirty participants, including professional developers and advanced computing students, completed equivalent programming tasks under three experimental conditions. Quantitative data were analyzed using descriptive statistics and repeated-measures ANOVA, while qualitative data were examined through thematic analysis. Results show that vibe coding significantly improved development efficiency, reducing task completion time by 27% compared with traditional coding and 12% compared with AI-assisted coding. However, these gains were accompanied by lower maintainability indices and higher security vulnerabilities, indicating trade-offs in software quality. Usability results yielded a good rating (SUS = 71.4), while cognitive workload remained moderate (NASA-TLX = 55.5), reflecting reduced syntactic effort but increased linguistic reasoning. Thematic analysis identified trust calibration, loss of control, cognitive adaptation, and prompt-engineering strategy as key constructs. Notably, perceived loss of control was associated with increased security risks due to reduced transparency and validation of AI-generated outputs. Based on these findings, the study proposes a three-pillar framework for responsible adoption: hybrid integration of human and AI capabilities, human oversight and transparent accountability, and context-aware deployment. Overall, vibe coding enhances productivity but requires critical oversight, reinforcing its role as a transformative yet transitional paradigm in software development.","authors":["Sales G. Aribe Jr.","Louie Jay S. Labastida"],"categories":["cs.SE","cs.AI","cs.ET"],"primary_category":"cs.SE","announce_type":"cross","date":"2026-09-10","first_seen":"2026-09-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.09560","pdf_url":"https://arxiv.org/pdf/2609.09560","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["AI辅助编程","人机交互","软件工程"],"reason":"研究AI辅助编程对开发者的影响，属于人机交互而非LLM仿真人类被试，无人类行为…","model":"deepseek-v4-pro","scored_at":"2026-09-10T13:01:32","error":null,"has_summary":false,"summary":null},{"id":"2609.09848","version":1,"title":"Subgroup Membership Inference Audits of Differentially Private Synthetic Text","zh_title":"差分隐私合成文本的子群体成员推断审计","abstract":"Synthetic data releases are increasingly proposed in the literature as a means of sharing realistic data replicas in lieu of sensitive private datasets. Even when the worst-case privacy leakage of such releases is bounded by means of differential privacy (DP), in practice a residual risk remains. Membership inference attack (MIA) audits are conducted to empirically quantify this risk. However, existing methods only measure average-case risk for randomly drawn records, which might conceal the risk to vulnerable subgroups. To highlight this issue, we define a subgroup-targeted membership inference game in which the target pool is an explicit parameter, and instantiate it with an audit of 32 proxies under three scenarios with different levels of attacker knowledge, across four datasets, three generators (DP-SGD fine-tuning, API-based prompting, and activation steering), and five privacy budgets. The audit shows that synthetic releases leak subgroup membership and that prior attacks systematically underestimate this leakage. DP is effective at the aggregate level: it substantially reduces average leakage at every budget we test. Three observations temper this picture. First, the remaining leakage is concentrated rather than spread out: under DP, a tenth of the records carries roughly 40% of it. Second, the protection DP delivers in practice is uneven: within its worst-case guarantee, the noise removes more of the measured leakage from random records than from high-risk ones---and a merged-pool audit that scores both record types against shared negatives confirms this at the record level. Third, \\emph{which} records leak proves to be a property of the release mechanism rather than of the record alone, so record-level risk cannot be assessed independently of the release.","authors":["Yidan Sun","Viktor Schlegel","Srinivasan Nandakumar","Siew Kei Lam","Anil Anthony Bharath"],"categories":["cs.CR","cs.AI"],"primary_category":"cs.CR","announce_type":"cross","date":"2026-09-10","first_seen":"2026-09-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.09848","pdf_url":"https://arxiv.org/pdf/2609.09848","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["差分隐私","成员推断攻击","合成数据"],"reason":"研究差分隐私合成文本的成员推断审计，不涉及用LLM仿真人类被试或与人类数据对照。","model":"deepseek-v4-pro","scored_at":"2026-09-11T13:01:57","error":null,"has_summary":false,"summary":null},{"id":"2609.09331","version":1,"title":"Ephemeral Feeds and Enduring Rituals: RushTok and the Formation of Event-Based Algorithmic Communities","zh_title":"短暂信息流与持久仪式：RushTok与基于事件的算法社区形成","abstract":"Each August, TikTok's For You page turns the University of Alabama's sorority recruitment into RushTok. We examine RushTok as an event-based algorithmic community: a collective assembled around a bounded offline ritual and sustained by recommendation. Using a mixed-methods survey (n=71) and a reflexive account of creator outreach, we ask who participates, how, and with what stakes. Findings show an ambiguous and entertainment based throughline; many called it a community (51/71) but few claimed membership (11/71). Affiliation centered on creators rather than shared practices, with parasocial attention clustering around a small set of potential new members (PNMs) and returning figures. Higher content exposure tracked with self-identification as a community member; those members commented, followed creators, and engaged across videos. Attempts to interview creators were met with silence or refusals, reflecting community boundary-work despite viral visibility. We outline implications for platform governance, including time-bounded context, graduated visibility, and aftercare.","authors":["Emelia Hughes","Tim Weninger"],"categories":["cs.HC","cs.CY","cs.SI"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-10","first_seen":"2026-09-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.09331","pdf_url":"https://arxiv.org/pdf/2609.09331","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["社交媒体","算法社区","平台治理"],"reason":"研究TikTok社区现象，不涉及LLM仿真人类被试，无实验或测量目的","model":"deepseek-v4-pro","scored_at":"2026-09-10T13:01:28","error":null,"has_summary":false,"summary":null},{"id":"2609.09333","version":1,"title":"Where Does the Human End? Creative Agency with Generative AI across Five Years of Chinese Digital Painting","zh_title":"人类止于何处？五年中国数字绘画中与生成式AI的创作能动性","abstract":"As generative AI enters creative work, practitioners must decide where AI assistance ends and human authorship begins. Human-agent interaction (HAI) research has examined AI as a tool, collaborator, consultant, and competitor. The longitudinal problem is how these roles are revised as systems become more capable, public, and economically embedded. We report a five-year interview study with 17 Chinese digital painters, based on annual semi-structured interviews from 2021 to 2025. Participants described recurring but non-uniform patterns of protective resistance, pragmatic task delegation, and, for some, reflective agency repartitioning. Early resistance protected observation, originality, signature, and ownership from AI. Later delegation placed AI in bounded tasks such as references, backgrounds, rough sketches, and client-facing drafts. By 2025, some participants built hybrid workflows around human-only zones, while others described fatigue, precarity, or difficulty locating a remaining human role. Peer norms, emotional climates, and production pressures shaped which delegations felt useful, acceptable, or exhausting. Copyright, authorship, and creative labor remained recurring limits on what participants were willing to delegate. We frame these accounts as longitudinal agency partitioning, the situated work of deciding which stages, responsibilities, values, and claims remain human in creative human-agent interaction. We discuss design implications for revisable agency-boundary controls, provenance scaffolds, and community-facing authorship norms.","authors":["Yibo Meng","Ruiqi Chen","Shuheng Cao","Weijia Zhang","Chengxi Zang"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-10","first_seen":"2026-09-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.09333","pdf_url":"https://arxiv.org/pdf/2609.09333","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["人机交互","创意能动性","生成式AI"],"reason":"研究人类与生成式AI在创作中的互动，非用LLM仿真人类被试，无实验或测量目的","model":"deepseek-v4-pro","scored_at":"2026-09-10T13:01:29","error":null,"has_summary":false,"summary":null},{"id":"2609.09365","version":1,"title":"Echoes in the Algorithm: Analyzing the Fidelity of User Preferences Against Realized Platform Reach","zh_title":"算法中的回声：分析用户偏好与平台实际触达的保真度","abstract":"What does popular content look like when platforms withhold the usual cues? On TikTok, users still form impressions about which videos are taking off even when likes and view counts are hidden, delayed, or pushed to the margins of the interface. We study this problem through TokOrNot, a web-based game in which participants compared pairs of TikTok videos and reported (i) which one they preferred and (ii) which one they believed had reached a larger audience. We benchmark these judgments against verified public view counts, which we use as a bounded proxy for realized platform reach. Across 3,513 judgments from 363 participants, participants identified the higher-reach video only modestly above chance (56.75%, 95% CI: 56.01-58.55). Preference aligned with the higher-view video at a similar rate, while preference and prediction matched in 83.48% of trials (95% CI: 83.12-85.95). Performance also varied across content categories. Taken together, these results do not suggest that users can reliably read platform success from content alone. Instead, they point to a looser and more uncertain interpretive process in which reach judgments often track personal taste or other weak heuristics when explicit popularity cues are absent. We discuss the implications for algorithmic literacy and for interface designs that reduce visible metrics without leaving users to infer reach from uneven or idiosyncratic cues alone.","authors":["Emelia Hughes","Tim Weninger"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-10","first_seen":"2026-09-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.09365","pdf_url":"https://arxiv.org/pdf/2609.09365","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C2"],"tags":["用户行为","平台算法","人机交互"],"reason":"研究人类对TikTok视频的判断，不涉及LLM仿真人类被试，属于平台用户行为研…","model":"deepseek-v4-pro","scored_at":"2026-09-10T13:01:30","error":null,"has_summary":false,"summary":null},{"id":"2609.09713","version":1,"title":"How Far Do Capability Cues Travel? Anthropomorphism and Differentiated Trust in a Platform-Embedded AI Assistant","zh_title":"能力线索能传多远？平台嵌入式AI助手中的人形化与差异化信任","abstract":"Visible AI capabilities need not translate into broader judgments of trustworthiness. In a randomized 2 x 2 experiment with 270 U.S.-based Reddit users, an embedded assistant displayed one or three functions, with or without a brief rationale. Displaying three functions increased perceived multifunctionality; no other randomized main effect survived correction across the six outcomes. Rationale availability did not reliably increase perceived intelligence. Exploratory analysis indicated stronger uptake of the functional display at higher objective AI literacy. Among concurrently measured judgments, perceived multifunctionality was associated with perceived intelligence, which was associated with anthropomorphism and all three trust dimensions. After accounting for perceived intelligence, anthropomorphism was positively associated with benevolence, but not reliably with integrity or ability. These findings separate interface effects from relationships among users' perceptions and show why ability, integrity, and benevolence should be evaluated separately.","authors":["Chenchen Mao","Hanjing Shi","Haiyan Jia","Dominic DiFranzo"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-10","first_seen":"2026-09-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.09713","pdf_url":"https://arxiv.org/pdf/2609.09713","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["人机交互","信任感知","AI助手"],"reason":"研究用户对AI助手的感知与信任，不涉及用LLM仿真人类被试或与人类数据对照","model":"deepseek-v4-pro","scored_at":"2026-09-10T13:01:32","error":null,"has_summary":false,"summary":null},{"id":"2609.09321","version":1,"title":"Endorsement Without New Evidence: How Sequential Voting Inflates Mandates in Online Community Governance","zh_title":"无新证据的背书：在线社区治理中顺序投票如何夸大授权","abstract":"Online communities often treat large support margins in public elections as strong mandates. We argue that such margins can overstate the independent scrutiny behind a decision. Using 198,275 free-text rationales from Wikipedia admin elections, we introduce vote-text divergence, a measure that flags a decisive vote paired with a thin, deferential rationale. Divergence rises as voters arrive later, even after controlling for voter and election fixed effects. The pattern is consistent with information saturation: once prior text is accounted for, arrival order no longer predicts divergence, while accumulated prior evidence does. The effect is strongest among peripheral voters in the co-voting network. Yet divergence does not predict worse post-promotion outcomes, such as administrative activity or survival. Public tallies can therefore weaken the scrutiny signal even while selecting capable administrators: a margin may appear to reflect more consensus and support than it actually contains.","authors":["Zihan Chen","Lei Nico Zheng","Di Zhu"],"categories":["cs.SI","cs.CY","cs.HC"],"primary_category":"cs.SI","announce_type":"cross","date":"2026-09-10","first_seen":"2026-09-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.09321","pdf_url":"https://arxiv.org/pdf/2609.09321","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["在线社区","投票行为","社会计算"],"reason":"研究在线社区投票行为，未使用LLM仿真人类被试，不涉及人类数据对照或LLM代理。","model":"deepseek-v4-pro","scored_at":"2026-09-10T13:01:27","error":null,"has_summary":false,"summary":null},{"id":"2609.09533","version":1,"title":"Scalable Oversight for AI in Mental Health: Lessons from 350,000 AI Coaching Conversations between Therapy Sessions","zh_title":"心理健康领域AI的可扩展监督：来自35万次治疗间AI辅导对话的经验","abstract":"Clinician review of every AI output is often proposed as a safeguard in mental healthcare, but vigilance research suggests this approach fails at scale and may paradoxically reduce safety. Drawing on our experience deploying an AI coaching tool across 350,000+ conversations between therapy sessions, we describe how we arrived at a three-layer human-on-the-loop oversight framework combining preventive design, real-time monitoring, and continuous clinician evaluation. We show how specific findings from clinical review drove iterative improvements, and offer practical recommendations for mental health professionals evaluating AI systems.","authors":["Matthew A. Scult","John L. Havlik","Kevin Ramotar","Ethan Goh","Manoj Kanagaraj"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2026-09-10","first_seen":"2026-09-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.09533","pdf_url":"https://arxiv.org/pdf/2609.09533","source_feed":"cs.CY","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["AI心理健康","人机对话","监督框架"],"reason":"AI心理健康辅导对话，属角色扮演聊天机器人，无实验或测量目的，不涉及人类仿真。","model":"deepseek-v4-pro","scored_at":"2026-09-10T13:01:31","error":null,"has_summary":false,"summary":null},{"id":"2609.10277","version":1,"title":"How neighbourhood ideology shapes misinformation belief in densely tied social networks","zh_title":"邻里意识形态如何影响紧密社交网络中的错误信息信念","abstract":"With the rapid spread of news on social media, understanding the propagation of misinformation is becoming increasingly important. One factor that affects individuals' vulnerability to false information is their ideological predisposition. Despite the large number of agent-based models that focus on social influence as a driver of the spread of false claims, they often fail to explicitly integrate personal ideological biases into belief formation. In this work, we explore how misinformation spreads through the interaction between individuals' ideological biases and social influence. Our model accounts for both the strength of individuals' ideological biases and the extent to which a false claim aligns with their ideology. Social influence modifies the effects of ideological intensity and false claim alignment through network interactions. Notably, the influence of neighbours' ideological intensity on belief is strongly affected by how well those neighbours are connected to one another. These results highlight the importance of considering both network structure and personal ideological biases when modelling misinformation propagation.","authors":["Soroush Karimi","Marcos Oliveira","Diogo Pacheco"],"categories":["cs.SI","cs.MA","physics.soc-ph"],"primary_category":"cs.SI","announce_type":"cross","date":"2026-09-10","first_seen":"2026-09-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.10277","pdf_url":"https://arxiv.org/pdf/2609.10277","source_feed":"cs.MA","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["错误信息传播","基于代理的模型","社交网络"],"reason":"基于代理的模型模拟信息传播，未使用LLM，不涉及人类被试替代或对照。","model":"deepseek-v4-pro","scored_at":"2026-09-10T13:01:37","error":null,"has_summary":false,"summary":null},{"id":"2606.14199","version":2,"title":"OdysSim: Building Foundation Models for Human Behavior Simulation","zh_title":"OdysSim：构建用于人类行为模拟的基础模型","abstract":"Large language models are increasingly deployed as human simulators for interactive evaluation and social simulation. Yet helpfulness-driven post-training pulls them toward a homogeneous, overly agreeable assistant register, creating a behavioral Sim2Real gap. We present OdysSim, the largest open systematic investigation of behavioral foundation models, i.e., models trained to simulate human behavior at scale. We propose SOUL, a taxonomy of five capability axes (CONV, SS, COG, ROLE, EVAL) that unifies 62 datasets and 23 benchmark tasks under one framework. Specifically, we curate the OdysSim corpus (21.4M interactions, 10B tokens, retrofitted with back-generated social contexts), construct the SOUL-Index benchmark, and develop an end-to-end training recipe combining midtraining, task-specific RL, and expert distillation. The resulting open 8B OSim model ranks first or tied-first on 8 of 23 tasks, outperforming any individual frontier model by this count, with the strongest gains on conversational and social tasks. Its outputs are also more human-like in length, formatting, and word choice, and it transfers zero-shot to out-of-distribution user simulation on $\\tau$-bench, nearly matching real users on reaction alignment (93.2 vs. 93.5). We further show that LLM-as-judge RL induces reward-hacking patterns, and that our detectors can mitigate them during post-training. Together, our findings suggest that behavioral foundation models require rethinking the LLM training paradigm. We release all artifacts to support future research.","authors":["Xuhui Zhou","Weiwei Sun","Weihua Du","Jiarui Liu","Haojia Sun","Qianou Ma","Tongshuang Wu","Yiming Yang","Maarten Sap"],"categories":["cs.CL","cs.AI","cs.LG"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-09","first_seen":"2026-06-12","revised_at":"2026-09-09","abs_url":"https://arxiv.org/abs/2606.14199","pdf_url":"https://arxiv.org/pdf/2606.14199","source_feed":"cs.CL","score":10,"bucket":"selected","rubric_hits":["A1","A2","A3","A4","A5","B1","B2","B3"],"tags":["人类行为模拟","基础模型","算法保真度"],"reason":"直接构建行为基础模型模拟人类行为，含真实人类数据对照，覆盖多任务并评估可靠性。","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:05:24","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-09","rank":1,"question":"如何构建行为基础模型来缩小LLM模拟人类行为时的Sim2Real差距？","design":"论文构建了OdysSim系统，使用Qwen3基础模型在21.4M交互的OdysSim语料上进行中训练，然后对SOUL-Index的23个任务进行任务特定强化学习（GRPO或RLVF），最后通过专家蒸馏合并为单一8B模型OSim-8B。模型被用于模拟对话、社会推理、认知、角色扮演和评估等五类人类行为，输出结果与真实人类数据对比。","baseline":"SOUL-Index基准包含62个数据集和23个任务，其中许多任务有真实人类行为数据作为对照；另外在τ-bench用户模拟评估中，与真实用户的反应对齐分数（93.5）进行对比。","findings":"OSim-8B在23个任务中的8个上排名第一或并列第一，超过任何单个前沿模型，尤其在对话和社会任务上提升最大；其输出在长度、格式和用词上更像人类，并在τ-bench上零样本迁移到用户模拟，反应对齐分数93.2接近真实用户93.5。","reliability":"论文指出LLM-as-judge强化学习会导致奖励黑客模式，但他们的检测器可以在后训练中缓解；此外，中训练数据虽经社会背景回填，但原始对话缺乏说话者背景，可能影响社会动态推断的准确性。","relevance":"该研究直接针对LLM模拟人类行为的可靠性问题，提供了大规模语料、基准和训练方法，并与真实人类数据对照，对关注人类仿真实验的研究者具有重要参考价值，值得精读原文。","inspiration":"借鉴其通过中训练注入行为多样性、任务特定RL校准行为以及专家蒸馏合并能力的方法，可迁移到经济金融中的消费者决策、投资者行为或政策反应模拟；例如，用LLM模拟投资者在政策公告后的交易行为，以真实市场数据（如订单流、调查数据）为基准，通过中训练和RL微调使模型输出与真实投资者行为分布对齐，从而评估政策效果。"}},{"id":"2609.07353","version":1,"title":"Human-like moral judgments conceal divergent motive attributions in large language models","zh_title":"类人道德判断掩盖了大语言模型中不同的动机归因","abstract":"Large language models (LLMs) are used to simulate human participants in psychological research. We asked whether LLMs that reproduce human evaluations of a whistleblower's moral character also reproduce the motive attributions that accompany them. Five LLMs and two human samples (N = 125 and N = 742) evaluated a physician who either remained silent about fraudulent billing or reported it to a hospital, regulator, or newspaper. Models reproduced the human ranking of the physician's moral character but portrayed whistleblowers as more helpful, less self-interested, and less hostile. In four of five models, competitive motives were less strongly associated with moral-character judgments. Model ratings changed little when prompts reproduced the narratives and demographic profiles of both human samples, although this comparison cannot isolate a perspective effect. Thus, agreement in average ratings can conceal differences in attributed motives, relationships among judgments, and sensitivity to context. Validating LLMs as simulated participants therefore requires testing psychologically informative response patterns, not average agreement alone.","authors":["Xiaoyan Wu","Jean-Claude Dreher"],"categories":["cs.AI","cs.CY"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-09","first_seen":"2026-09-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.07353","pdf_url":"https://arxiv.org/pdf/2609.07353","source_feed":"cs.AI","score":10,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","道德判断","算法保真度"],"reason":"直接使用LLM仿真人类道德判断，并与两个人类样本对照，发现平均评分一致但动机归…","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:04:27","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-09","rank":4,"question":"LLM在复现人类对举报者道德品质评价的同时，是否也复现了伴随的动机归因？","design":"用五个LLM（含闭源与开源）模拟人类被试，对一位医生面对欺诈账单保持沉默或向医院、监管机构、报纸举报的四种情境进行道德品质与四种动机（义务、亲社会、自利、竞争）评分，并比较模型与两个人类样本的评分模式、动机与道德判断的关联，以及模型对两种样本叙述的敏感性。","baseline":"两个独立招募的人类样本（N=125和N=742），分别接受第一手（事件发生在自己团队）和第二手（事后听说）版本的场景描述。","findings":"模型复现了人类对道德品质的条件排序，但将举报者描绘为更亲社会、更少自利和敌意；五个模型中有四个的竞争动机与道德品质判断的关联弱于人类。模型评分对两种样本叙述的变化不敏感，平均评分一致掩盖了动机归因、判断间关系和情境敏感性的差异。","reliability":"论文承认两种样本的对比不能分离视角效应，因为框架、招募和样本构成同时变化；模型对提示变化的敏感性可能影响结果；仅凭平均一致性不足以验证LLM作为模拟被试的有效性。","relevance":"直接回应了LLM仿真人类被试的核心问题：平均评分一致不等于心理结构一致，对评估仿真可靠性至关重要，值得精读。","inspiration":"借鉴其多层次验证设计：不仅比较条件均值，还检验变量间关系和情境敏感性｜可迁移到经济决策中的道德或动机归因场景，如举报行为、企业社会责任评价、消费者对品牌道德危机的反应｜用LLM模拟消费者对某公司不当行为的道德判断，处理为不同举报渠道（内部、监管、媒体），结果变量为道德评价和动机归因，对照真实消费者调查数据，检验模型是否复现动机与评价的关联模式。"}},{"id":"2609.07573","version":1,"title":"From Simulated Citizens to Simulated Deliberation: Challenges in Representation and Interaction","zh_title":"从模拟公民到模拟审议：表征与互动的挑战","abstract":"Multi-agent LLM deliberation has been explored as a scalable way to simulate public deliberation. For such simulations to be informative, persona agents should reflect population opinion patterns and interaction should shape their conclusions. We evaluate whether LLM-based deliberation can meet these two conditions using census-grounded Korean personas debating real policy questions benchmarked against national surveys. Persona agents do not reliably reproduce population opinion patterns: responses are often far more concentrated and frequently reverse demographic differences in the human data. Deliberations nonetheless produce reasoned, reciprocal, and varied arguments alongside substantial stance movement. Yet much of this movement does not require peer exchange: sealed-monologue agents change position at similar rates and reach nearly the same final balance as full debates, while groups initialized with very different positions often converge to similar endpoints. Anchoring population-informed starting positions, meanwhile, sharply suppresses updating. Thus, population representation, argument generation, and interaction-driven opinion change do not necessarily go together. The simulations readily surface arguments on both sides, though whether they capture the diversity of human perspectives remains untested, leaving open a promising role for argument surfacing even as population simulation requires further validation.","authors":["Chaemin Jang","Junsik Min","Jaewoo Choi","Donggyu Lee","Haiin Lee","Junyoung Park","Namhee Kim","Hyunwoo Kim","Jungwon Kim","Juho Kim","Nuri Kim","Jihee Kim"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-09","first_seen":"2026-09-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.07573","pdf_url":"https://arxiv.org/pdf/2609.07573","source_feed":"cs.AI","score":10,"bucket":"selected","rubric_hits":["A1","A3","B1","B2","B4"],"tags":["LLM仿真","公共审议","人类数据对照"],"reason":"用LLM模拟公民审议并与全国调查对照，评估代表性与互动效应，直接相关。","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:04:28","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-09","rank":5,"question":"基于LLM的多智能体审议能否同时满足意见代表性和互动驱动观点变化两个条件？","design":"使用基于人口普查校准的韩国人设智能体（Nemotron-Personas）就真实政策问题进行辩论，并与全国调查基准对照；通过封闭独白控制组分离同伴互动对立场变化的影响。","baseline":"韩国全国环境意识调查和低生育率政策公众意识调查的群体级基准数据。","findings":"人设智能体未能可靠复现人口意见模式，回答更集中且常逆转人口学差异；审议产生了理性、互惠和多样的论点，但大部分立场变化不需要同伴互动，且不同初始立场的群体常收敛到相似终点。","reliability":"论文承认人口代表性、论点生成和互动驱动的观点变化不一定同时成立；论点是否捕捉人类视角多样性尚未检验，人口模拟需进一步验证。","relevance":"直接相关，提供了LLM模拟审议的严格评估，揭示了代表性失败和互动效应虚假的问题，对使用LLM进行人类仿真实验的研究者具有重要警示价值。","inspiration":"值得借鉴封闭独白控制组来分离互动效应，以及用人口普查校准人设并与全国调查对照的评估框架｜可迁移到政策公告的预期形成或消费者信心调查等场景，检验LLM模拟的群体意见动态｜用LLM人设模拟投资者或消费者群体，处理为是否进行多智能体辩论，结果变量为观点变化和最终分布，对照真实调查数据（如密歇根消费者信心指数）来验证仿真可靠性。"}},{"id":"2609.07987","version":1,"title":"When Can LLM Digital Twins Reduce Human Measurement? From Behavioral Fidelity to Statistical Substitutability","zh_title":"LLM数字孪生何时能减少人类测量？从行为保真度到统计可替代性","abstract":"LLM-based digital twins promise to reduce repeated human data collection by generating person- specific responses, yet existing evaluations provide little evidence about whether they can reduce human measurement while preserving valid inference. To address this, we introduce statistical substitutability, an inferential criterion that evaluates the extent to which twin predictions can reduce human measurement for a particular estimand while preserving valid inference. We develop a framework, grounded in mixed-subject and prediction-powered inference, that evaluates statistical substitutability along four dimensions: aggregate fidelity, paired respondent-level signal, finite-sample human-label recovery, and stability across populations. Across two empirical evaluations spanning behavioral experiments, multiple models, and alternative respondent representations, we find that digital twins can reproduce average human effects while providing little information about which individuals differ from those averages. Newer models and richer respondent information improve some dimensions of performance but do not reliably translate into human-data savings. Human calibration can reduce aggregate prediction error, yet limited labeled samples often fail to produce stable precision gains. Importantly, these findings demonstrate that behavioral fidelity is neither necessary nor sufficient for statistical substitutability. More broadly, they suggest that AI-generated evidence should be evaluated based on its ability to support valid scientific inference rather than its ability to reproduce human outcomes alone. Digital twins should therefore be judged for confirmatory use by whether they reduce uncertainty about human quantities, not merely by whether they reproduce human means, distributions, or effects.","authors":["Steven Wang","Kyle Hunt","Shaojie Tang","Kenneth Joseph"],"categories":["cs.AI","stat.AP"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-09","first_seen":"2026-09-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.07987","pdf_url":"https://arxiv.org/pdf/2609.07987","source_feed":"cs.AI","score":10,"bucket":"selected","rubric_hits":["A1","A2","A3","A4","B1","B2","B3","B4"],"tags":["LLM数字孪生","统计可替代性","人类仿真"],"reason":"直接研究LLM数字孪生替代人类测量的统计可替代性，含真实人类数据对照与批判性评…","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:04:32","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-09","rank":6,"question":"LLM数字孪生能否在保持有效推断的同时减少人类测量，即其统计可替代性如何？","design":"使用LLM数字孪生（基于受访者丰富信息构建）预测个体在行为实验中的反应，并与真实人类数据对比，评估其在聚合效应、个体层面信号、有限样本标签恢复和跨人群稳定性四个维度上的表现。","baseline":"两个实证评估中的真实人类行为实验数据，包括个体层面的反应和实验处理效应。","findings":"数字孪生能复现平均人类效应，但几乎不提供个体偏离平均值的信号；新模型和更丰富的受访者信息改善了部分维度，但未可靠转化为人类数据节省。","reliability":"行为保真度既非统计可替代性的必要条件也非充分条件；人类校准可减少聚合预测误差，但有限标签样本往往无法产生稳定的精度增益；数字孪生应基于其减少人类量不确定性的能力来评判，而非仅复现人类均值、分布或效应。","relevance":"该研究直接针对LLM仿真人类被试的可靠性问题，提出了统计可替代性标准，并基于真实人类数据对照进行了批判性评估，对关注仿真有效性和偏差的研究者极具参考价值。","inspiration":"借鉴其统计可替代性框架，将LLM预测与人类数据结合进行混合推断，并检验个体层面信号和跨人群稳定性｜可迁移到经济金融中的个体决策预测，如消费者跨期选择、风险偏好或投资行为，评估LLM能否替代部分人类被试｜设计一个资产定价实验，用LLM数字孪生预测个体对风险资产的需求，处理为不同信息条件，结果变量为投资金额，并与真实人类实验数据对照，检验LLM预测能否减少所需人类样本量。"}},{"id":"2609.08003","version":1,"title":"Sparks of In Silico Cognitive Science: Theories from Simulated Data Can Generalize to Humans","zh_title":"硅基认知科学的火花：来自模拟数据的理论可以推广到人类","abstract":"Behavioral foundation models have been proposed as stand-ins for human participants across settings, but it is unclear whether theories discovered on them generalize to humans or merely characterize the simulator. We ran the Automated Cognitive Scientist (\\textsc{AutoCog}), a closed-loop discovery system in which LLM agents design theory-discriminating experiments, collect responses, arbitrate between competing theories, and synthesize successors, entirely on behavior simulated by Centaur, a foundation model of human behavior. In a multi-attribute decision-making setting, the theories \\textsc{AutoCog} found on Centaur generalized to human data: they outperformed canonical theories on ten held-out experiments and were rivaled only by theories found by running the same loop on people. We argue that this succeeds despite the simulator's inevitable imperfections because a discovery loop that arbitrates between competing theories demands less of its simulator than estimation does. The simulator only needs to capture the regularities that distinguish the theories, and not necessarily reproduce behavior precisely. Imperfect simulators can therefore widen the search over theories, with human data then testing whether the surfaced theories generalize.","authors":["Akshay K. Jagadish","Younes Strittmatter","Nori Jacoby","Eric Schulz","Nathaniel Daw","Thomas L. Griffiths","Suyog H. Chandramouli"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-09","first_seen":"2026-09-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.08003","pdf_url":"https://arxiv.org/pdf/2609.08003","source_feed":"cs.AI","score":10,"bucket":"selected","rubric_hits":["A1","A2","A3","A5","B1","B2","B3","B4"],"tags":["LLM仿真","人类行为对照","理论发现"],"reason":"用LLM仿真人类决策，并与真实人类数据对照，验证理论可推广性，直接相关。","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:04:32","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-09","rank":7,"question":"在模拟人类行为的基础模型上运行自动理论发现系统，所发现的理论能否推广到真实人类行为？","design":"使用 Centaur（人类行为基础模型）模拟人类在多属性决策任务中的选择，运行 AutoCog 闭环发现系统：LLM 智能体设计区分理论的实验、收集模拟响应、在竞争理论间仲裁并合成后继理论，最终得到在模拟数据上表现最好的理论。","baseline":"对照真实人类数据：在十个留出实验上比较 AutoCog 在 Centaur 上发现的理论、经典理论以及直接在人类数据上运行 AutoCog 发现的理论。","findings":"在 Centaur 模拟数据上发现的理论在人类数据上优于经典理论，且与直接在人类数据上运行相同发现循环得到的理论表现相当。成功的原因在于理论仲裁对模拟器的要求低于精确估计，模拟器只需捕捉区分理论的关键规律。","reliability":"论文承认模拟器不可避免存在缺陷，但认为在理论发现场景下这些缺陷影响较小；未详细讨论模拟器在哪些具体条件下会失效。","relevance":"直接回应了 LLM 仿真人类行为并用于理论发现的可推广性问题，提供了与真实人类数据对照的实证证据，对关注仿真可靠性和偏差的研究者极具参考价值。","inspiration":"借鉴其闭环理论发现框架：用 LLM 智能体自动设计实验、仲裁理论并迭代，可大幅扩展理论搜索空间，再用人类数据验证。｜可迁移到经济决策中的启发式建模，例如消费者在复杂金融产品间的选择、投资者在多属性资产间的配置决策。｜以 LLM 模拟投资者在多属性资产间的选择行为，运行 AutoCog 发现决策理论，再与真实投资者在相同实验中的选择数据对照，检验理论的可推广性。"}},{"id":"2609.07141","version":1,"title":"How Well Do LLMs Simulate Survey Responses Following a Breast Cancer Screening Intervention?","zh_title":"LLM在乳腺癌筛查干预后模拟调查回答的效果如何？","abstract":"Collecting survey data is laborious and limited by privacy constraints. Large language models (LLMs) have shown promise as predictive social simulations. It is unclear whether they can replicate population-level response distributions before and after a healthcare intervention. Using information derived from 4125 women aged 35-59 years, we evaluate whether agents informed solely by pre-intervention profile information can reproduce post-intervention response distributions. Groups of LLM agents (n=50) were created with Gemma 4 E4B and Qwen3.5 9B; conditions ranged from zero-shot prompting to agent profiles enriched with aggregate or individual-level demographic characteristics and pre-intervention questionnaire responses. We compared predicted and observed response distributions with Total Variation Distance (TVD) and Normalized Wasserstein Distance (NWD). Across both LLMs, profile-based agents improved distributional accuracy relative to zero-shot and random baselines. Nevertheless, direct sampling of 50 real participants remained more accurate. Prediction errors were also higher among participants aged 55-59 years and those living in private property. Errors also varied by question theme and LLM model, with the highest errors observed for cancer fatalism and post intervention attitudes toward genetics. Sensitivity analyses showed that performance was influenced by prompt template changes and temperature hyperparameter. Our results show the potential of LLM-based agents to model behavioral responses to interventions in silico. However, profiles containing additional information beyond demographics did not consistently outperform simpler ones. Certain cultural constructs and population groups also remain inadequately represented by the LLM models evaluated. Future work may include building behaviorally grounded and locally validated virtual populations.","authors":["Kenneth Koh","Ryan Jak Yang Lim","Alessandro Sparacio","Peh Joo Ho","Mile Sikic","Borame L Dickens","Mikael Hartman","Jingmei Li"],"categories":["cs.SI"],"primary_category":"cs.SI","announce_type":"new","date":"2026-09-09","first_seen":"2026-09-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.07141","pdf_url":"https://arxiv.org/pdf/2609.07141","source_feed":"cs.SI","score":10,"bucket":"selected","rubric_hits":["A1","A2","A3","B1","B2","B4"],"tags":["LLM仿真","调查回答","人类对照"],"reason":"用LLM代理模拟乳腺癌筛查干预后的调查回答，并与真实人类数据对照，评估分布准确…","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:04:25","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-09","rank":3,"question":"LLM 智能体仅基于干预前的个人资料信息，能否复现乳腺癌筛查干预后的人群调查回答分布？","design":"使用 Gemma 4 E4B 和 Qwen3.5 9B 两种 LLM，基于 4125 名 35-59 岁新加坡女性的真实数据创建智能体，条件包括零样本、聚合人口学特征、个体人口学特征及干预前问卷回答等不同信息丰富度，模拟干预后问卷回答分布，并与真实回答比较。","baseline":"来自 BREATHE 队列的 4125 名女性的真实干预后调查回答，以及随机预测和直接抽样 50 名真实参与者的分布。","findings":"基于个人资料的智能体相比零样本和随机基线提高了分布准确性，但仍不如直接抽样真实参与者；预测误差在 55-59 岁和私宅居民中更高，且在不同问题主题和 LLM 模型间存在差异。","reliability":"论文承认 LLM 对某些文化构念和人群亚群代表性不足，且包含额外信息的个人资料并不总是优于简单资料；性能受提示模板和温度超参数影响。","relevance":"该研究直接评估 LLM 在医疗干预后调查回答仿真中的可靠性，与研究者关注的人类仿真实验、真实数据对照和失效条件高度相关，值得精读原文。","inspiration":"借鉴其用干预前资料构建智能体并对比真实干预后分布的设计，以及通过 TVD/NWD 和亚组分析评估仿真偏差的方法。｜可迁移到政策干预对经济行为的影响评估，如健康保险补贴对就医行为、财务教育对储蓄决策的影响。｜以真实调查数据中的个体特征和干预前行为为输入，让 LLM 智能体模拟干预后的消费或投资选择，并与实际追踪调查数据对照，检验仿真在收入、年龄等亚组上的误差模式。"}},{"id":"2608.22697","version":3,"title":"Does Rank Still Matter? Position Bias When AI Agents Shop on Our Behalf","zh_title":"排名还重要吗？AI代理替我们购物时的位置偏差","abstract":"Search rankings are valuable because human attention is scarce and sequential. Higher-placed alternatives are easier to find, so they are examined and bought more often. Consumers are now delegating search to AI agents that can ingest an entire results page at once. Randomizing the order of one hundred hotel listings across 5,000 AI agent sessions, we compare four large language models against human field data. AI agents search more deeply than humans and never decline to buy. Position still predicts which listings are inspected, but weakly and non-monotonically: the middle of a results page has the lowest probability of inspection, not the bottom. Position reaches the choice stage for some models and not others, a heterogeneity that tracks neither provider nor capability. All models nonetheless converge on the same undominated listing. For agentic search, the attributes displayed on a results page matter more than placement within it.","authors":["Davood Wadi","Yu Ma"],"categories":["cs.AI","econ.GN","q-fin.EC"],"primary_category":"cs.AI","announce_type":"replace","date":"2026-09-09","first_seen":"2026-08-25","revised_at":"2026-09-09","abs_url":"https://arxiv.org/abs/2608.22697","pdf_url":"https://arxiv.org/pdf/2608.22697","source_feed":"cs.AI","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM仿真","消费者行为","位置偏差"],"reason":"用LLM模拟消费者搜索决策，并与真实人类数据对照，属于经济学场景下的人类仿真实…","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:05:25","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-09","rank":8,"question":"当消费者将搜索任务委托给AI代理时，搜索结果排名是否仍然影响其检查和购买行为？","design":"使用四个大型语言模型（Gemini 3.1 Pro、Gemini 3.7 Flash、Gemini 3.1 Flash Lite、Claude Sonnet 5）扮演酒店预订助手，在100个酒店列表的随机排序环境中进行搜索和预订，通过工具调用模拟点击检查，记录检查次数、检查位置和最终选择。","baseline":"对照Ursu (2018)的人类现场实验数据，该实验在相同酒店搜索环境中随机化排序并记录人类消费者的点击和购买行为。","findings":"AI代理比人类搜索更深（检查次数1.63-5.83 vs 1.12），且几乎总是预订酒店；排名对检查的影响弱于人类且非单调，中间位置检查概率最低（lost-in-the-middle效应），但排名对最终选择的影响因模型而异，部分模型无显著影响。所有模型的选择高度集中在单一非支配性酒店上。","reliability":"论文未讨论","relevance":"该研究用LLM模拟消费者搜索决策，并与真实人类数据对照，属于经济学场景下的人类仿真实验，直接回应了研究者对LLM作为人类被试替代品的可靠性与偏差的关注。","inspiration":"借鉴其随机化排序和工具调用模拟点击的设计，以隔离位置效应；可迁移到在线市场中的消费者搜索与选择问题，如电商平台商品排序对购买决策的影响；设计上可用LLM模拟消费者在随机排序的商品列表中选择，记录点击和购买，并与历史点击流数据对照，检验位置偏差的模型异质性。"}},{"id":"2609.05189","version":2,"title":"Can Large Language Models Anticipate Behavioral Responses to Social Policies? A Case of Pension Enrollment Prediction among China's Flexible Workers","zh_title":"大语言模型能否预测社会政策的行为反应？中国灵活就业人员养老金参保预测案例","abstract":"Assessing the impacts of social policy changes is a widely acknowledged challenge for policymakers. Econometric methods can be unreliable when extrapolating to hypothetical scenarios, while field pilot programs are highly costly. In this paper, we propose using large language models (LLMs) as policy-assessment tools adapted from general-purpose models. We present FlexPension-LLM, the first domain-specialized large language model for a hierarchical pension-enrollment prediction task among flexible workers in China, and introduce DKI-RDistill, which injects policy-grounded cues into the prompt, including Probit-derived marginal effects and hukou-province pension rules. The method then uses LoRA/SFT to distill rationale-augmented supervision into an open-weight MoE student, with teacher errors corrected by regenerating those cases under ground-truth labels. On a CHFS 2019 blind split, FlexPension-LLM achieves 0.9316 Composite F1, surpassing its Claude Sonnet 4.5 teacher and 15 of 17 baselines, and is statistically indistinguishable from Claude Opus 4.6. Across four external surveys, it averages 0.7549 Composite F1 and shows the narrowest performance range among the strongest systems. Component analysis shows that gains come mainly from policy-grounded cue injection and error-filtered supervision, while rationales provide decision traces that can be checked against policy rules.","authors":["Yumiao Li","Peixin Liu","Donglin Di","Chen Li","Runhuan Feng"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-09","first_seen":"2026-09-07","revised_at":"2026-09-09","abs_url":"https://arxiv.org/abs/2609.05189","pdf_url":"https://arxiv.org/pdf/2609.05189","source_feed":"cs.CL","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2","B3"],"tags":["LLM仿真","政策评估","行为预测"],"reason":"用LLM预测养老金参保行为，与真实调查数据对照，属经济学政策评估场景。","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:05:27","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-09","rank":9,"question":"如何利用大语言模型预测中国灵活就业人员的养老金参保行为，以评估社会政策变化的影响？","design":"提出 FlexPension-LLM，一个针对中国灵活就业人员养老金参保预测的领域专用大语言模型。模型输入结构化个人、家庭、历史参保和户籍省份政策信息，输出参保行为（不参保、居民养老保险、职工养老保险）及结构化理由。通过 DKI-RDistill 框架，将 Probit 边际效应和户籍省份养老金规则注入提示，用教师模型生成理由，对教师错误案例用真实标签重新生成后再进行 LoRA/SFT 蒸馏。","baseline":"使用中国家庭金融调查（CHFS）2019 年数据作为主要数据集，并在四个外部家庭调查数据集上进行验证。","findings":"在 CHFS 2019 盲测集上，FlexPension-LLM 的 Composite F1 达到 0.9316，超过其教师模型 Claude Sonnet 4.5 和 15/17 个基线，与 Claude Opus 4.6 无显著差异。在四个外部调查上平均 Composite F1 为 0.7549，表现最稳定；消融分析表明收益主要来自政策基础线索注入和错误过滤监督。","reliability":"论文未明确讨论失效条件，但指出通用 LLM 存在规则应用肤浅、理性人偏差和一致性弱等局限，领域专用化旨在缓解这些问题。","relevance":"该研究直接回应了用 LLM 模拟人类行为并对照真实调查数据的核心关切，提供了经济学政策评估场景下的完整案例，值得精读以了解其仿真设计、基准对比和可靠性处理。","inspiration":"借鉴其将计量经济学先验（如 Probit 边际效应）和制度规则注入提示，并用教师-学生蒸馏结合错误过滤来提升仿真准确性的方法。｜可迁移到政策公告对家庭金融决策的影响预测，如养老金改革、税收优惠或补贴政策对储蓄和参保行为的影响。｜以中国家庭金融调查数据为真实基准，用 LLM 模拟家庭在政策变化下的参保或储蓄决策，处理为政策参数调整，结果变量为决策类别，对照真实调查中的实际行为。"}},{"id":"2609.05993","version":1,"title":"Alignment by Stereotyping: How LLMs Sacrifice Individual Distinctiveness for Cultural Adaptation","zh_title":"刻板化对齐：LLM如何为文化适应牺牲个体独特性","abstract":"Large language models are increasingly deployed for personalized interaction, and demographic conditioning via user profiles is a widely adopted strategy for cultural adaptation. We ask whether this approach genuinely serves individual users or achieves accuracy by erasing individual distinctiveness. Studying seven models including frontier GPT-5.1 on the World Values Survey, we find that demographic profiles improve value alignment accuracy for most models, but at a systematic cost to individuality. That is, models pull responses toward demographic group centroids rather than preserving individual differences, a behavioral pattern we term alignment by stereotyping. Permutation tests (10,000 permutations, six demographic attributes, seven models) certify that top-performing models compress individuals far above the human baseline; within-family scaling amplifies this tradeoff while degrading intrinsic cultural understanding. Using a synthetic dialogue dataset validated on real human-chatbot conversations from PRISM (Kirk et al., 2024), we further show that distributing demographic signals across conversational turns partially suppresses prototype retrieval compared to compact demographic labels, a finding validated on real conversations via PRISM but requiring replication at larger scale.","authors":["Qishuai Zhong","Zongmin Li","Siqi Fan","Aixin Sun"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-09","first_seen":"2026-09-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.05993","pdf_url":"https://arxiv.org/pdf/2609.05993","source_feed":"cs.CL","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","算法保真度","价值观调查"],"reason":"用LLM复现世界价值观调查，与真实人类数据对照，评估个体差异保真度，批判性指出…","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:04:23","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-09","rank":12,"question":"提供人口统计画像是否在提升群体平均价值观对齐的同时，以牺牲个体独特性为代价，其行为模式是否为将个体压缩至群体原型（即“刻板化对齐”）？","design":"用7个LLM（含GPT-5.1）扮演世界价值观调查（WVS）的受访者，在三种条件下（无上下文、提供人口统计画像、提供对话历史）回答55个价值观题目，测量群体平均对齐准确率（VAA）和个体同质化率（Homogenization Rate），并通过置换检验（10000次）验证同质化是否由人口统计驱动。","baseline":"真实人类数据：WVS Wave 7的1000名匿名受访者及其人口统计属性和价值观回答，作为对齐准确率和同质化率的人类基准（同质化率50%）。","findings":"人口统计画像提升了多数模型的群体平均对齐准确率，但代价是个体独特性被系统性抹除，模型将个体响应拉向群体中心而非保留个体差异。置换检验证实顶尖模型的个体压缩远超人类基准，且模型规模放大这一权衡，同时削弱内在文化理解。","reliability":"论文承认对话历史部分抑制刻板化，但在真实对话上的验证（PRISM）样本量小（80个），需要更大规模复制；合成对话数据集虽经人工验证，但可能无法完全代表真实交互。","relevance":"高度相关：该研究用LLM复现WVS并与真实人类数据对照，直接评估个体差异保真度，批判性指出人口统计画像导致“刻板化对齐”，对仿真可靠性提出警示，值得精读。","inspiration":"借鉴其置换检验和同质化率指标来量化个体差异损失，以及无上下文基线隔离处理效应的方法｜可迁移到信贷审批歧视研究，检验LLM基于人口统计特征（如种族、性别）的决策是否牺牲个体信息而依赖群体刻板印象｜用LLM扮演信贷审批员，处理为提供申请人人口统计画像 vs. 仅提供财务信息，结果变量为审批决策和利率，对照真实信贷数据（如HMDA）中的个体差异分布。"}},{"id":"2609.06545","version":1,"title":"LLMs Mirror Country-Specific Gender Patterns If Asked, but Skew Male When Generating Media in Local Languages","zh_title":"LLM在直接询问时反映国家特定性别模式，但在生成本地语言媒体时偏向男性","abstract":"Large language models (LLMs) are increasingly used to generate media, but whether their content perpetuates gender stereotypes is unknown: standard benchmarks rely on selection-based formats rather than long-form generation, and surveyed baselines for local gender associations are scarce outside the West. We collect gender associations for 22 occupational and domestic roles from 695 respondents across the United States, India, Kenya, and Nigeria, and evaluate eight LLMs under two regimes: direct questioning and media generation. Models track the surveyed associations under direct questioning but skew substantially more male under media generation in major local-language cells, consistent with the male bias documented in human-produced media. Outside the US, the shift is much smaller and non-significant under English prompting, so English-only or country-agnostic evaluation would miss this bias in the languages where these models are most deployed. Instruction prompting reduces the shift directionally, but trades off against alignment with the surveyed associations. Evaluating LLM gender bias for global deployment therefore requires generation-format testing, local-language prompting, and locally-collected human baselines.","authors":["Sharif Kazemi","Tanya Popli","Neil K. R. Sehgal","Sunny Rai","Niyati Malhotra","Victor Orozco-Olvera","Ana Mar\\'ia Mu\\~noz Boudet","Samuel P. Fraiberger","Sharath Chandra Guntuku","Manuel Tonneau"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-09","first_seen":"2026-09-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.06545","pdf_url":"https://arxiv.org/pdf/2609.06545","source_feed":"cs.CL","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","性别偏差","跨文化对照"],"reason":"用LLM复现人类性别关联，有真实调查数据对照，并评估生成偏差","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:04:23","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-09","rank":13,"question":"LLM在直接提问和媒体生成两种方式下，是否复现或偏离了四个国家的人类性别关联？","design":"用8个LLM在直接提问和媒体生成两种条件下生成22个职业和家务角色的性别关联，比较模型输出与人类调查数据的差异。","baseline":"在美国、印度、肯尼亚、尼日利亚四国收集的695名受访者的性别关联调查数据，并用外部劳动力统计验证。","findings":"直接提问时模型输出与人类调查关联一致；媒体生成时模型显著偏向男性，尤其在本地语言（印地语、约鲁巴语、斯瓦希里语）下，与人类媒体中的男性偏见一致。指令提示能方向性减少偏差，但会降低与人类关联的对齐度。","reliability":"论文承认使用二元性别简化了现实，调查样本为在线招募而非概率样本，且媒体生成评估可能受提示设计影响。","relevance":"该研究提供了LLM仿真人类性别关联的实证证据，并揭示了生成格式和语言对仿真偏差的影响，对关注LLM仿真可靠性的研究者有参考价值。","inspiration":"借鉴其多国调查与LLM输出的对照设计，以及显式与隐式两种测量方式的对比。｜可迁移到信贷审批中的性别歧视研究，比较LLM在直接询问和生成贷款决策时的差异。｜以LLM为被试，施加不同提示语言和格式处理，测量贷款批准率，并与真实银行信贷数据中的性别差异对照。"}},{"id":"2609.07305","version":1,"title":"Marginal Fidelity Does Not Establish User Simulation in Demographic Synthetic Survey Panels: Response Contracts, Support Collapse and Conditioning Failure","zh_title":"边际保真度不能确立人口合成调查面板中的用户仿真：响应契约、支持坍缩与条件化失败","abstract":"Demographic synthetic survey panels are often validated by matching aggregate answers to published surveys. We test what that certificate establishes across six multiselect batteries from four survey organisations in three countries. The headline analysis is restricted to three instruments whose synthetic cohort and human target share the stated population frame; three other batteries remain sensitivity analyses. The response contract dominates measured fidelity. In the aligned instruments, committed sets leave 66 of 128 model-battery option slots empty in panels of up to 500 respondents, versus 0 of 128 under per-option probability elicitation. Across eight uncapped model-instrument comparisons, probabilities reduce option-marginal MAE by 4.53 to 7.30 points. The capped instrument reverses on two models until the vectors are projected onto its stated maximum. These are measurement effects: human targets are realised check-all responses, whereas the vectors are latent inclusion propensities. Published marginal agreement also fails to discriminate respondent simulation from direct population estimation. On nine aligned model-battery pairs, a no-persona population-prevalence query averages 6.27 MAE versus 12.39 for committed panels and wins all nine comparisons. Constraint-aware probability vectors average 5.34 and beat the query on four of nine, so the baseline challenges the validation criterion rather than proving direct estimation uniformly best. On three unpublished demographic cells, neither approach beats reciting the national distribution. Population-marginal agreement is therefore evidence about an elicitation contract and an estimand obtainable without simulated respondents, not evidence of individual simulation.","authors":["Alexander Doudkin"],"categories":["cs.CL","cs.HC"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-09","first_seen":"2026-09-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.07305","pdf_url":"https://arxiv.org/pdf/2609.07305","source_feed":"cs.CL","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","调查方法","算法保真度"],"reason":"直接评估LLM合成调查面板的仿真效度，并与真实人类数据对照，指出边际保真度不足…","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:04:25","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-09","rank":14,"question":"在人口合成调查面板中，匹配总体边际答案能否证明对个体受访者的仿真有效？","design":"论文使用多个LLM（模型未在节选中具体列出）生成合成受访者，以两种方式回答多项选择题：一是让模型以受访者身份直接选择选项（承诺集），二是让模型给出每个选项被选择的概率（概率向量）；然后比较这些合成回答与真实调查的边际分布。","baseline":"对照的真实人类数据来自四个调查机构在三个国家的六组多项选择题，其中三组与合成样本的人口框架一致，另三组作为敏感性分析。","findings":"边际一致性主要反映的是诱导方式（承诺集 vs. 概率向量）的测量效应，而非个体仿真能力；直接询问总体患病率（无角色提示）在多数情况下比模拟受访者更接近真实边际分布，说明边际一致性不能区分个体仿真与总体估计。","reliability":"论文承认其探索性分析是描述性和事后性的，不能证明在样本外优于传统调查估计量；并且人口边际一致性只是关于诱导方式和无需模拟受访者即可获得的估计量的陈述，不是个体仿真的证据。","relevance":"该研究直接评估LLM合成调查面板的仿真效度，并与真实人类数据对照，指出边际保真度不足，对关注LLM仿真可靠性及偏差的研究者具有重要参考价值，值得阅读原文以了解具体实验设计和失效条件。","inspiration":"借鉴其对照设计：同时使用承诺集和概率向量两种诱导方式，并引入无角色总体估计作为基准，以分离测量效应与仿真能力。｜可迁移到经济预期调查或消费者信心指数仿真，检验LLM能否复现真实人群的预期分布。｜以LLM模拟消费者回答密歇根消费者信心调查，处理为是否提供人口统计角色，结果变量为各问题选项的概率分布，对照真实调查的边际分布和个体数据，比较承诺集、概率向量和总体估计的误差。"}},{"id":"2609.05437","version":1,"title":"Beyond Right and Wrong: Evaluating Second-order Social Reasoning in Large Language Models","zh_title":"超越对错：评估大语言模型中的二阶社会推理","abstract":"Previous AI alignment efforts have focused primarily on first-order social norms -- teaching models what is socially acceptable or unacceptable (e.g., `do not steal'). However, social intelligence depends not only on norm recognition, but also on anticipating who will enforce it and how (e.g., public shame or even imprisonment). These second-order expectations, known as metanorms, govern how people respond when social rules are broken. We introduce a novel framework for evaluating metanorm reasoning in Large Language Models (LLMs) along two dimensions: emotional appraisal and behavioral response, and propose new classification tasks, namely, predicting self-regulation in violators, and other-regulation in observers. We release a multi-perspective dataset, NormReact, of 450 norm violation scenarios, hand-annotated for emotions and behavioral responses across norm violators' gender and observers' social closeness. Current LLMs portray a harsher social world: across six models, they overpredict negative sanctions where humans would expect inaction, and alignment with human judgments deteriorates as social distance increases. These findings suggest that AI systems in norm-sensitive domains from conflict mediation to policy simulation, may risk producing a distorted picture of social regulation: one that over-represents punishment and under-represents the tolerance, restraint, and relational calibration that characterize actual norm enforcement in real world.","authors":["Sunny Rai","Jinyi Kuang","Reyhan Jamalova","Annie Lou","Cristina Bicchieri","Niyati Malhotra","Victor Hugo Orozco-Olvera","Ana Maria Munoz-Boudet","Lyle H Ungar","Sharath C Guntuku"],"categories":["cs.AI","cs.CY"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-09","first_seen":"2026-09-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.05437","pdf_url":"https://arxiv.org/pdf/2609.05437","source_feed":"cs.AI","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","社会规范","人类对照"],"reason":"用LLM预测人类对规范违反的情绪与行为反应，并与人类标注对照，发现偏差。","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:04:22","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-09","rank":10,"question":"LLM能否像人类一样进行二阶社会推理，即预测规范违反后违规者和观察者的情绪与行为反应，并随社会距离变化？","design":"构建NormReact数据集，包含450个来自Reddit的日常规范违反场景，人工标注违规者性别和观察者社会距离（强关系、弱关系、陌生人），并标注违规者和观察者的八种情绪及八种行为反应。用六个LLM对每个场景预测违规者情绪（自我调节）和观察者情绪（他人调节）以及观察者的描述性（会做什么）和指令性（应该做什么）行为反应，与人类标注对比。","baseline":"人类标注：450个场景的多视角人工标注，涵盖违规者性别和观察者社会距离，提供情绪和行为反应的基准。","findings":"LLM过度预测负面制裁，在人类预期不作为的情况下预测惩罚；随着社会距离增加，LLM与人类判断的一致性下降。LLM描绘了一个比人类认可的更严厉的社会世界，过度代表惩罚，低估了现实中的容忍、克制和关系校准。","reliability":"论文指出LLM在规范敏感领域（如冲突调解、政策模拟）可能产生扭曲的社会调节图景，但未系统讨论失效条件；局限性可能包括数据集规模有限、场景来自Reddit可能偏向特定类型、情绪和行为分类的简化等，但正文节选未明确提及。","relevance":"该研究直接评估LLM在人类规范执行仿真中的偏差，与研究者关注的人类仿真可靠性高度相关，特别是在社会规范和政策评估场景中，值得精读以了解LLM在二阶社会推理上的系统性偏差。","inspiration":"借鉴其多视角标注和关系距离操纵，可迁移到经济金融中的社会规范执行场景，如逃税、违约、内幕交易等。｜设计实验：用LLM扮演不同社会距离的观察者，预测对经济违规行为（如逃税）的情绪和制裁反应，与真实人类调查数据（如世界价值观调查或实验经济学中的第三方惩罚实验）对照，检验LLM是否过度惩罚并随社会距离偏差增大。"}},{"id":"2609.05514","version":1,"title":"The Failure Happens Before the Drift: The Social Dynamics of Values in LLM Agent Societies","zh_title":"失败发生在漂移之前：LLM智能体社会中价值观的社会动力学","abstract":"Large Language Model (LLM)-based agents are increasingly used as proxies for human participants in social science research, yet it remains unclear whether they can faithfully simulate diverse and conflicting human value systems. We present a World Values Survey (WVS)-grounded simulation framework where culturally diverse agents with different communication styles engage in longitudinal, value-laden discussions. Across approximately 4,000 conversations involving 1,200 personas, 15 topics, and three models (GPT-4o, Gemini-2.5-Flash, and Gemma-4-E4B), we evaluate value faithfulness, value drift, and conversational realism. We find that more than 50\\% of personas fail to express their assigned WVS profiles from the outset, while 2-7\\% drift after repeated conversations. Ablations removing demographic details improve faithfulness for some models but do not change the broader trend: simulated value distributions still systematically deviate from the assigned WVS profiles. Compared to human discussions, simulated dialogues show a different trade-off between stylistic consistency and semantic diversity, often producing content-wise varied but stylistically repetitive exchanges. These findings suggest that current LLM agents can generate plausible conversations, but remain limited proxies for representing and preserving diverse human value profiles over time.","authors":["Farah Atif","Sougata Saha","Monojit Choudhury"],"categories":["cs.AI","cs.MA"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-09","first_seen":"2026-09-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.05514","pdf_url":"https://arxiv.org/pdf/2609.05514","source_feed":"cs.AI","score":9,"bucket":"selected","rubric_hits":["A1","A2","A3","B1","B4"],"tags":["LLM仿真","价值观调查","算法保真度"],"reason":"用LLM代理模拟人类价值观，并与WVS真实数据对照，评估仿真保真度与漂移，直接…","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:04:23","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-09","rank":11,"question":"LLM智能体能否在纵向互动中忠实表达并保持其被赋予的WVS文化价值观？","design":"基于WVS第7波数据构建1200个具有人口统计特征、价值观和沟通风格的人设，让GPT-4o、Gemini-2.5-Flash和Gemma-4-E4B扮演这些角色，在15个价值主题上进行约4000场多轮讨论，测量价值观保真度、漂移和对话真实性。","baseline":"WVS第7波真实人类调查数据（90,000名受访者，66个国家）以及人类对话语料库。","findings":"超过50%的人设在首次对话中就无法表达其被赋予的WVS价值观，而2-7%的人设在重复对话后发生漂移。去除人口统计细节可提高某些模型的保真度，但模拟的价值分布仍系统性偏离WVS；模拟对话在风格一致性和语义多样性上与人类对话存在不同权衡。","reliability":"论文指出LLM智能体在初始阶段就未能实例化文化价值观，且模拟的价值分布系统性偏离真实分布，表明当前LLM代理在表示和保持多样化人类价值观方面存在局限。","relevance":"该研究直接评估LLM作为人类被试替代品在价值观仿真中的可靠性，与研究者关注的人类仿真实验、真实数据对照和失效条件高度相关，值得精读原文。","inspiration":"借鉴其纵向多轮互动设计和价值观保真度测量方法，可迁移到经济政策评估中的公众态度仿真，例如模拟不同文化背景个体对税收或福利政策的态度变化。｜设计一个实验：用LLM扮演不同WVS价值观的个体，在讨论经济政策后测量其态度变化，并与真实调查数据（如世界价值观调查中的经济态度题项）对照，检验仿真保真度。"}},{"id":"2609.07687","version":1,"title":"Perspectives on Cross-Lingual Consistency in LLMs for Medical Questions","zh_title":"关于医学问题中LLM跨语言一致性的观点","abstract":"Should multilingual LLMs answer medical questions consistently across input languages, or adapt responses to cultural cues? Existing multilingual medical benchmarks usually assume that medically correct answers should remain consistent across languages and treat cross-lingual variation as model error. In contrast, cultural adaptation research argues that appropriate medical answers may legitimately differ across contexts. We review the multilingual medical NLP literature through these two perspectives, we identify three gaps: limited stakeholder perspectives (e.g., of medical professionals), a lack of empirical evidence on which approach better serves users, and no benchmarks capable of distinguishing universally correct from culture-specific cases. To address the first gap, we survey 356 participants across three stakeholder groups (medical, NLP, and anthropology professionals) in three countries (Germany, Spain, and the United States). Anthropologists consistently favor adaptation, while medical and NLP respondents remain divided, with notable divergence between U.S. and European medical professionals. LLMs prompted with profession and country personas fail to reproduce this variation, overestimating cross-lingual consistency preference among NLP and medical personas. We conclude that neither consistency nor adaptation can currently be considered clearly preferable, highlighting the need for empirical evidence on which approach better serves users across cultural contexts.","authors":["Minh Duc Bui","Mario Sanz-Guerrero","Abteen Ebrahimi","Sagi Shaier","Peter Herbert Kann","Manuel Mager","Katharina von der Wense"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-09","first_seen":"2026-09-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.07687","pdf_url":"https://arxiv.org/pdf/2609.07687","source_feed":"cs.CL","score":8,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","跨语言一致性","人类对照"],"reason":"用LLM模拟不同职业/国家人群对医疗问题的跨语言一致性偏好，并与356人真实调…","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:04:30","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-09","rank":16,"question":"多语言大语言模型在回答医学问题时，应该跨语言保持一致，还是根据文化语境调整答案？","design":"该研究并非仿真实验，而是先通过文献综述梳理一致性与适应性两种立场，然后对348名来自德国、西班牙和美国的医学、NLP和人类学专业人士进行问卷调查，测量他们对两种立场的偏好；最后用LLM以职业和国家人设生成回答，与人类调查结果对比。","baseline":"348名真实人类参与者的调查数据，按职业（医学、NLP、人类学）和国家（德国、西班牙、美国）分组。","findings":"人类学专业人士一致偏好适应性，而医学和NLP受访者意见分歧，且美国医学受访者比欧洲同行更倾向适应性。LLM人设未能复现这种差异，系统性地高估了医学和NLP人设对一致性的支持。","reliability":"论文承认调查采用二元强制选择，无法区分受访者权衡的是医学内容还是沟通风格；样本量不足以进行细粒度人口统计比较；样本全部来自全球北方，限制了结论的普遍性。","relevance":"该研究直接检验了LLM能否模拟不同专业和国家人群对医疗AI行为的偏好，发现LLM仿真失效，对关注LLM作为人类被试替代品的研究者具有重要参考价值。","inspiration":"该研究用真实人类调查作为基准，检验LLM人设仿真的准确性，方法可借鉴。｜可迁移到经济金融领域的政策偏好或消费者态度调查，如不同国家投资者对风险披露语言的偏好。｜以真实投资者调查为基准，用LLM生成不同国家和职业人设的回答，比较其对风险披露一致性与本地化适应性的偏好分布，评估仿真偏差。"}},{"id":"2608.23095","version":2,"title":"Definitional Sensitivity in Media Bias Detection: A Multi-Definition Dataset and Benchmark","zh_title":"媒体偏见检测中的定义敏感性：多定义数据集与基准","abstract":"Media bias detection relies on definitions and examples that specify what counts as bias, yet these specifications often vary across datasets or remain implicit, even when given the same name. Such variation makes it unclear whether models trained for the same bias category learn the same construct or different phenomena, a problem largely overlooked in prior work. We examine how definition choice affects bias annotation in a between-subjects experiment with 354 participants and a parallel evaluation with four LLMs. Participants and models rate six news articles across four bias categories using definitions that vary in conceptual framing and elaboration. Across 8,496 human and 28,800 LLM ratings, we find that the conceptual target of a definition drives annotation divergence, while construct-preserving elaboration does not: conceptual framing significantly shifts annotations for humans and does so even more strongly for LLMs. We discuss implications for construct specification in annotation protocols and prompt-based measurement, and consider how definitional sensitivity may propagate to downstream classification beyond media bias. We also release MUDD, the Multi-Definition Bias Detection Dataset.","authors":["Martin Wessel","Timo Spinde","J\\\"urgen Pfeffer","Gianluca Demartini"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-09","first_seen":"2026-08-25","revised_at":"2026-09-09","abs_url":"https://arxiv.org/abs/2608.23095","pdf_url":"https://arxiv.org/pdf/2608.23095","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A1","B1","B4"],"tags":["LLM仿真","人类对照","定义敏感性"],"reason":"用LLM替代人类被试进行媒体偏见标注实验，并与354名人类对照，发现定义敏感性…","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:05:25","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-09","rank":17,"question":"媒体偏见标注中，定义的概念框架和详细程度如何影响人类与LLM的偏见评分？","design":"用四个LLM模拟人类标注者，对六篇新闻文章在四个偏见类别上评分，每个类别随机分配六种定义（三种概念×两种长度），测量评分差异。","baseline":"354名美国参与者通过Prolific招募，完成相同标注任务，产生8496条人类评分。","findings":"定义的概念目标显著影响人类和LLM的评分，而保持概念不变的详细阐述不影响评分。LLM对定义变化的敏感性比人类强约1.3-4倍，且不同架构的LLM可能产生方向相反的效应。","reliability":"论文未讨论LLM仿真的失效条件，但指出定义敏感性可能传播到下游分类任务，且LLM的放大效应可能因架构而异。","relevance":"该研究直接对比人类与LLM在定义敏感性上的差异，为评估LLM作为人类被试替代品的可靠性提供了关键证据，值得精读。","inspiration":"借鉴其因子设计（概念×长度）和预注册实验，分离定义内容与形式的影响，并对比人类与LLM的效应量。｜可迁移到经济金融中的概念定义敏感性场景，如通胀预期调查中的措辞效应、风险偏好测量中的框架效应。｜以LLM模拟受访者，随机分配不同定义的通胀预期问题（如“物价上涨”vs“货币贬值”），测量预期值，并与密歇根大学消费者调查的真实数据对照，检验LLM是否复现措辞效应。"}},{"id":"2609.02163","version":2,"title":"Do Cantonese-Adapted Language Models Better Predict Cantonese Reading? A Cross-Model Eye-Tracking Evaluation","zh_title":"粤语适配的语言模型能更好地预测粤语阅读吗？一项跨模型眼动追踪评估","abstract":"Information-theoretic measures derived from autoregressive language models are widely used to characterize the expectations that shape human reading, but whether language-variety-specific training improves such psycholinguistic alignment remains unclear. This question is still open for Cantonese, where recent NLP evaluations reported mixed benefits from Cantonese-specific training relative to Mandarin-oriented or general-purpose models. Using naturalistic Cantonese eye-tracking data, we compare two within-family adaptation contrasts: CKIP GPT-2 Tiny versus its lightly Cantonese-adapted JED351 derivative, and Qwen2.5-7B versus CantoneseLLM-7B, which underwent substantially more extensive Cantonese continued pretraining and instruction tuning. From each model, we derive lexical surprisal, POS surprisal, entropy before the target, and entropy reduction. Lexical surprisal and the joint four-metric model consistently favor CantoneseLLM-7B, followed by Qwen2.5-7B, CKIP, and JED351, whereas entropy reduction favors CKIP. These results suggest that more extensive Cantonese-specific training can be associated with stronger predictive fit, while model rankings also depend on the information-theoretic measure being evaluated.","authors":["Ziqi Zhang","Emmanuele Chersoni","Mohammad Momenian"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-09","first_seen":"2026-09-03","revised_at":"2026-09-09","abs_url":"https://arxiv.org/abs/2609.02163","pdf_url":"https://arxiv.org/pdf/2609.02163","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A1","B1"],"tags":["心理语言学","眼动追踪","语言模型评估"],"reason":"用LLM预测人类阅读行为，有真实眼动数据对照，方法可迁移到仿真研究。","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:05:31","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-09","rank":18,"question":"粤语适配的语言模型是否比通用或普通话导向的模型更能预测粤语阅读中的眼动行为？","design":"本研究不是人类仿真实验，而是用四个语言模型（CKIP GPT-2 Tiny、其粤语适配版 JED351、Qwen2.5-7B、其粤语适配版 CantoneseLLM-7B）计算信息论指标（词汇惊讶度、词性惊讶度、目标词前熵、熵减），并检验这些指标对粤语读者眼动数据的预测力。","baseline":"使用 MCFIX 数据集中的粤语眼动追踪数据，包含 5,007 个自然阅读和 5,004 个任务特定阅读观测，结果变量为首次注视时长、二次注视时长和总注视时长。","findings":"词汇惊讶度和联合四指标模型一致支持 CantoneseLLM-7B 预测力最强，其次是 Qwen2.5-7B、CKIP 和 JED351；但熵减指标则支持 CKIP 表现最好。这表明更广泛的粤语特定训练可能与更强的预测拟合相关，但模型排名也取决于所评估的信息论指标。","reliability":"论文未讨论","relevance":"该研究用 LLM 预测真实人类阅读行为，有眼动数据作为对照基准，方法可迁移到用 LLM 仿真人类认知或行为的研究中，值得阅读原文以了解模型适配与指标选择对预测力的影响。","inspiration":"借鉴其通过同一模型家族内不同适配程度的对比来评估训练数据领域匹配对行为预测的影响，以及使用多种信息论指标并比较其预测力的做法。｜可迁移到政策公告的预期形成研究，例如用不同领域适配的 LLM 预测投资者对央行公告的反应。｜以通用 LLM 和金融领域适配 LLM 为被试，处理为不同政策措辞的公告文本，结果变量为模型生成的预期变化或模拟交易行为，用真实市场数据或投资者调查作为对照。"}},{"id":"2609.06263","version":1,"title":"Beyond the Flag: Clinical Framing Closes the Moderation Gap in Suicide Risk Measurement","zh_title":"超越二元标记：临床框架缩小自杀风险测量中的调节差距","abstract":"Moderation APIs are built to flag policy-violating content, not to measure graded clinical risk. But a platform's duty does not end at detection: the response owed to passive distress differs sharply from the response owed to active planning with means access, and emerging regulation (e.g., California Senate Bill 243) is turning that distinction into a compliance requirement. We therefore ask how well deployed safety signals recover clinically meaningful severity. We release a benchmark of 516 r/SuicideWatch posts rated by a licensed psychiatrist on a four-level ordinal schema (Indicator, Ideation, Behavior, Attempt) grounded in the Columbia Suicide Severity Rating Scale, and evaluate moderation APIs, prompted LLMs, and supervised baselines under seven ordinal-aware metrics. Three findings. Vendor moderation APIs separate low- from high-severity posts well (0.860 high-risk F1) but measure severity poorly (0.395 macro F1), systematically over-predicting the most severe category. Clinically grounded zero-shot prompting recovers much of that gap (0.562 macro F1), and expert-authored framing (not fine-tuning, added reasoning, or naive multi-agent aggregation) is the effective lever. The value of reasoning depends on register: it hurts on long, noisy Reddit posts and helps on short, clinician-authored statements. We argue graded severity, not a binary flag, is what a proportionate duty of care requires, and release our evaluation framework to support that measurement.","authors":["Shreyas Krishnan","Gun Ahn","Jungjin Kim"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-09","first_seen":"2026-09-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.06263","pdf_url":"https://arxiv.org/pdf/2609.06263","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A1","B1","B4"],"tags":["LLM评估","临床风险分级","人类对照"],"reason":"用LLM评估自杀风险严重程度，与人类专家评级对照，但非仿真人类被试，而是临床测…","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:04:53","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-09","rank":21,"question":"部署的审核API和提示的LLM能在多大程度上恢复临床上分级的自杀风险严重程度（四级序数：Indicator、Ideation、Behavior、Attempt）？","design":"该研究不是人类仿真实验，而是评估LLM在自杀风险分级任务上的表现。它构建了516条来自r/SuicideWatch的帖子，由持照精神科医生按C-SSRS基础的四级序数模式标注，然后评估了供应商审核API、提示的LLM和监督基线在七个序数感知指标上的表现，比较了不同提示策略（临床框架、零样本、思维链、多智能体聚合）和微调的效果。","baseline":"人类基准是持照精神科医生对516条Reddit帖子的四级序数标签，其中20条子样本由另外两名精神科医生独立标注，与参考标签在每一级上的一致性达到二次加权kappa 0.938–0.957。","findings":"供应商审核API能较好地区分低严重度与高严重度帖子（高风险F1为0.860），但在序数分层上表现差（宏F1为0.395），系统性地过度预测最严重类别。临床基础的零样本提示大幅缩小了这一差距（宏F1为0.562），且专家撰写的临床框架是有效杠杆，而微调、增加推理或多智能体聚合均未带来改进；推理的价值取决于文本语域，在长而嘈杂的Reddit帖子上有害，在短而临床医生撰写的陈述上有益。","reliability":"论文承认其比较是针对该临床医生的评分标准在该语料库上的，不声称提示的普遍排名；数据集缺乏阴性层级（没有非自杀帖子），且推理策略的逆转仅基于两种语域，不能确定归因于语域。","relevance":"该研究虽非人类仿真，但提供了LLM在临床分级任务上与人类专家对照的严格基准，展示了提示工程和评估指标的重要性，对关注LLM可靠性和偏差的研究者有参考价值。","inspiration":"值得借鉴的是使用专家撰写的领域框架作为提示来提升LLM在序数分类任务上的表现，并采用多个序数感知指标（如宏F1、二次加权kappa）进行评估，同时对比不同提示策略和模型规模。｜可以迁移到经济金融中的信用评级、风险分类或消费者财务困境分级等序数预测任务，例如将贷款申请或社交媒体财务帖子按风险严重程度分级。｜一个可行的研究设计是：使用LLM对来自Reddit个人财务板块的帖子进行财务困境严重程度分级（如从“轻微担忧”到“破产危机”），处理是不同提示框架（通用vs.专家撰写的金融风险框架），结果变量是序数标签，并与人类金融顾问或信用评分机构的真实评级进行对照，评估宏F1和加权kappa。"}},{"id":"2609.08576","version":1,"title":"Which Forms of Caregiver Feedback Support Grammar Learning? A Reinforcement-Learning Study of Child-Like Language Models","zh_title":"哪种形式的看护者反馈支持语法学习？对类儿童语言模型的强化学习研究","abstract":"Social interaction is central to children's language learning, but the effects of different forms of caregiver feedback are difficult to isolate in naturalistic data. We use child-like language models as controlled learners to test which forms of feedback support grammatical development. Small GPT-2-style models are pretrained on child-directed language from CHILDES, then fine-tuned with reinforcement learning using reward models trained to capture four feedback types: communicative feedback, structural alignment, semantic contingency, and affective feedback. Reward fine-tuning yields limited gains on minimal-pair evaluations, but clearer effects in free generation. Structural alignment produces the strongest improvements in grammaticality, providing a novel, plausible mechanistic account of how this feedback can support grammar learning. Communicative feedback yields more moderate gains. In contrast, semantic contingency and affective feedback do not improve grammaticality, although further analyses suggest that they may support other aspects of language learning beyond grammar. These results suggest that different forms of caregiver feedback make complementary contributions to language learning.","authors":["Jing Liu","Marianne Schweitzer","Abdellah Fourtassi"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-09","first_seen":"2026-09-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.08576","pdf_url":"https://arxiv.org/pdf/2609.08576","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A1","B1"],"tags":["语言学习仿真","强化学习","人类数据对照"],"reason":"用LLM模拟儿童语言学习，并与真实儿童数据对照，方法可迁移到人类仿真研究。","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:05:19","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-09","rank":27,"question":"不同形式的照料者反馈（沟通反馈、结构对齐、语义关联、情感反馈）如何影响儿童语法学习？","design":"用小型 GPT-2 模型模拟儿童语言学习者，先在 CHILDES 儿童导向语言上预训练，再用强化学习微调，奖励模型分别针对四种反馈类型训练，评估语法性提升。","baseline":"无对照（基于 CHILDES 语料训练奖励模型，但未直接与真实儿童学习结果对比）。","findings":"结构对齐反馈对语法性提升最强，沟通反馈次之；语义关联和情感反馈未改善语法性，但可能支持其他语言方面。","reliability":"论文未讨论","relevance":"该研究用 LLM 模拟儿童语言学习，并利用真实语料训练奖励模型，方法可迁移到人类仿真研究，值得读原文了解其强化学习设置和反馈建模。","inspiration":"借鉴其用真实交互数据训练奖励模型并施加不同反馈类型的方法，可迁移到经济金融中的沟通与反馈场景，如政策沟通或投资者关系管理。｜用 LLM 模拟投资者，施加不同反馈（如确认、澄清、情感回应），测量其决策质量变化，并与真实投资者行为数据对照。"}},{"id":"2609.09048","version":1,"title":"The Audit Decides the Verdict: Instrument Effects Rival Demographic Bias in LLM Decision Audits","zh_title":"审计决定裁决：LLM决策审计中工具效应堪比人口统计偏差","abstract":"Whether a language model looks demographically biased can depend on how the audit asks its question. A charitable-aid benchmark reports that the same models favor minority applicants when rating requests one at a time and penalize some when ranking side by side. We test whether that reversal generalizes to hiring, lending, and medical triage: 40,726 requests to five models, applications differing only in the applicant's name, and a primary test fixed before collection. It does not. None of 36 planned contrasts survives correction. The rating advantage keeps its sign at roughly half the published size, and a precision extension bounds any hiring ranking penalty below the published effect, though the lending and triage ranking floors sit above that margin, so the exclusion is conclusive for hiring ranking and for rating in all three domains only. Planted disparities tracking their injected sizes and a directional replication on the original aid materials bound these nulls. The audit is livelier than the demographics: models recognize transparent audits nearly always, tie every identical-content comparison whether the varying detail is race or a hobby, and reward first-listed candidates as much as any demographic effect we measure. Audit verdicts reflect audit construction more than demographic bias.","authors":["Siddharth Vohra","Manikandan Ravikiran"],"categories":["cs.CL","cs.AI","cs.CY"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-09","first_seen":"2026-09-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.09048","pdf_url":"https://arxiv.org/pdf/2609.09048","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A2","B1","B4"],"tags":["LLM审计","算法偏差","决策仿真"],"reason":"审计LLM决策中的偏差，有真实人类数据对照，批判审计方法影响结论","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:04:36","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-09","rank":30,"question":"审计提问方式（评分、排名、伪装等）是否会导致大语言模型在招聘、信贷、医疗分诊等决策中表现出不同甚至相反的人口统计学偏差？","design":"使用五个大语言模型，对来自 AgentFairBench 的招聘、信贷、分诊三个领域的基础档案，仅改变申请人姓名以暗示种族和性别，施加三种审计格式（1-5评分、二元决策、四候选人排名，排名又分人口统计学对比明显或伪装），测量模型给出的评分、决策和排名结果。","baseline":"无对照","findings":"在三个领域中均未发现慈善援助基准中报告的评分与排名偏差反转现象，36个计划对比无一在多重校正后显著；评分优势约为已发表效应的一半，招聘排名惩罚被排除，但信贷和分诊排名惩罚无法排除。审计格式本身的影响大于人口统计学特征：模型几乎总能识别透明审计，对内容相同的比较全部打平，且对排名第一的候选人给予的奖励与任何测量到的人口统计学效应相当。","reliability":"论文未讨论","relevance":"该研究直接检验了审计设计对LLM决策偏差结论的影响，发现审计格式比人口统计学特征更能驱动结果，对关注LLM仿真可靠性及偏差测量条件的研究者具有重要参考价值，值得阅读原文。","inspiration":"借鉴其通过改变审计格式（评分vs排名、透明vs伪装）来分离测量效应与真实偏差的方法，以及使用植入已知偏差作为效应量下限的稳健性检验。｜可迁移到信贷审批歧视研究，检验不同评估方式（单独评分vs对比排序）是否影响算法或人类决策中的种族/性别偏差。｜以银行信贷员或LLM为被试，随机分配贷款申请为单独评分或成对排名条件，申请人姓名暗示种族，结果变量为批准概率或评分，对照真实信贷审批数据中的种族差异。"}},{"id":"2609.09070","version":1,"title":"Performance of Clinical AI System and Physicians and Frontier Language Models in primary care diagnostics","zh_title":"临床AI系统、医生与前沿语言模型在初级保健诊断中的表现","abstract":"Clinical AI evaluation should encompass diagnosis and management after adaptive information gathering. We compared Doctorina, eight physicians and four standalone frontier language models in 150 synthetic Polish-language primary-care consultations. Doctorina achieved 82.0% Top-1 concordance versus 57.0% for physicians (difference, 25.0 percentage points; 95% confidence interval, 17.7-32.7) and 97.3% versus 85.0% primary-or-reference-differential concordance. Across 149 case pairs, normalized workup and treatment scores were 89.4 versus 66.9 and 83.7 versus 61.2. Doctorina had the highest diagnostic point estimates among all six groups; Kimi K3 ranked next, while Claude Opus 5 led the closely spaced management estimates of Opus, Doctorina and Kimi. A second Doctorina execution reproduced the advantages over physicians across all outcomes. Doctorina's advantage over physicians therefore extended from primary-diagnosis selection to higher-rated diagnostic workup and initial treatment after adaptive consultation.","authors":["Andy Nkansah","Hanna Plotnitskaya","Stanislau Salavei","Anna Kozlova","Piotr Gibas","Julian Milek","Viktar Harbachou","Aleksey Ropan","Pavel Satalkin"],"categories":["cs.CL","cs.AI","cs.HC"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-09","first_seen":"2026-09-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.09070","pdf_url":"https://arxiv.org/pdf/2609.09070","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A1","B1","B2"],"tags":["LLM仿真","临床诊断","人类对照"],"reason":"用LLM模拟医生诊断并与真实医生对照，属于人类决策仿真，但场景为临床而非社会科…","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:04:36","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-09","rank":31,"question":"在波兰语初级保健的适应性咨询中，集成式临床AI系统Doctorina与医生及前沿LLM相比，在诊断和管理上的表现如何？","design":"使用150个合成的波兰语初级保健病例，以适应性模拟咨询方式，比较Doctorina、8名波兰医生和4个独立前沿LLM（Gemini 3.1 Pro Preview、Claude Opus 5、GPT-5.6-sol、Kimi K3）的诊断和管理表现。结果变量包括Top-1诊断一致性、主要或参考鉴别诊断一致性、诊断检查和初始治疗评分。","baseline":"8名在波兰从事家庭医学或内科的医生（5名专科医生和3名住院医师）作为人类对照，每个病例有1-2名医生作答。","findings":"Doctorina的Top-1诊断一致性为82.0%，显著高于医生的57.0%（差异25.0个百分点，95% CI 17.7-32.7）。Doctorina在诊断检查（89.4 vs 66.9）和初始治疗（83.7 vs 61.2）的标准化评分上也优于医生，且第二次运行重现了这些优势。","reliability":"论文未讨论","relevance":"该研究将LLM作为人类医生决策的仿真模型，并与真实医生进行对照，属于人类决策仿真，但场景为临床诊断而非社会科学或经济学，且未涉及政策评估或批判性失效条件分析，因此对研究者的直接参考价值有限。","inspiration":"与经济金融研究关联不大"}},{"id":"2609.05517","version":1,"title":"Emergent Goal-Directed Attention in Large Vision-Language Models","zh_title":"大型视觉语言模型中涌现的目标导向注意力","abstract":"Human observers prioritize visual information according to task goals. Most computational models of naturalistic viewing are gaze-trained for free viewing, leaving open whether goal-directed attention can emerge in systems without gaze supervision. We tested two off-the-shelf vision-language models (VLMs), Qwen3-VL-32B-Thinking and Gemma-4-26B-A4B-it, on 4,887 naturalistic scenes under visual-search and free-viewing instructions. Model predictions were compared with human fixations on the same images under corresponding tasks. Both models aligned more closely with human fixations under matching goals than under mismatched goals. This crossover persisted in target-absent scenes, where alignment could not be explained by simple visual grounding, and appeared in decoder-layer readouts. Furthermore, model-thinking traces were grounded in target semantics during search and in visual prominence during free viewing. These findings show that general-purpose VLMs can generate human-aligned, goal-directed spatial priorities without gaze-specific training, informing theories of goal-directed attention and offering scalable tools for predicting where people look across tasks.","authors":["Han Zhang"],"categories":["cs.CV","cs.AI","cs.CL"],"primary_category":"cs.CV","announce_type":"cross","date":"2026-09-09","first_seen":"2026-09-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.05517","pdf_url":"https://arxiv.org/pdf/2609.05517","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A1","B1"],"tags":["视觉注意力","人类行为仿真","多模态模型"],"reason":"用VLM预测人类注视行为，并与真实人类眼动数据对照，属于用LLM仿真人类感知决…","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:04:44","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-09","rank":19,"question":"通用视觉语言模型能否在没有注视数据训练的情况下，产生与人类目标导向注意力一致的空间优先级？","design":"用两个现成的视觉语言模型（Qwen3-VL-32B-Thinking 和 Gemma-4-26B-A4B-it）扮演人类观察者，对 4,887 张自然场景图片施加两种任务指令（视觉搜索目标物体 vs. 自由观看以记忆场景），让模型输出其会注视的位置点，并构建优先级图作为结果变量。","baseline":"人类观察者在相同图片和相同任务指令下的真实注视数据（每张图片每种任务 10 名被试）。","findings":"两个模型在任务目标匹配时与人类注视的相关性显著高于目标不匹配时，且这种交叉效应在目标缺失场景和模型解码器层中依然存在。模型思维链在搜索时更贴近目标语义，在自由观看时更贴近视觉显著性。","reliability":"论文未讨论","relevance":"该研究用 VLM 仿真人类视觉注意行为，并与真实人类眼动数据对照，属于用 LLM 仿真人类感知决策的范畴，对关注仿真可靠性与偏差的研究者有参考价值。","inspiration":"借鉴其通过改变任务指令来施加处理、并利用匹配与不匹配条件构建交叉对照的设计，可以检验模型是否真正理解任务目标而非简单视觉定位。｜可迁移到经济金融中的信息搜索与决策场景，例如投资者在财报中搜索关键指标或消费者在商品页面中寻找特定信息。｜用 LLM 扮演投资者，施加“搜索盈利指标”与“自由浏览”两种指令，让模型输出关注区域或文本片段，与真实投资者的眼动或点击数据对照，检验模型的目标导向信息选择是否与人类一致。"}},{"id":"2609.07598","version":1,"title":"Mapping the Emerging Social Science of Large Language Models","zh_title":"绘制大语言模型新兴社会科学研究图景","abstract":"Large language models (LLMs) increasingly shape communication, learning, work, creativity, and decision-making, yet social-science research on these developments remains fragmented. We map this emerging field using a curated corpus of 198 papers reviewed in full and a field-scale corpus of 47,719 published papers from five bibliographic databases. Combining sentence embeddings, K-means clustering, within-cluster Latent Dirichlet Allocation (LDA), author and LLM classifications, and structural topic modeling, we identify three domains: LLM as Social Minds, examining socially interpretable model behavior; LLM Societies, examining collective dynamics among interacting model-based agents; and LLM-Human Interactions, examining how people perceive, use, and are affected by LLMs. These domains contain 13 subcategories spanning reasoning, personality and bias, behavioral games, collective intelligence, simulation, trust, work, creativity, and education. In the curated corpus, the three-domain solution is highly stable under resampling (adjusted Rand index = 0.952), and K-means assignments agree with author full-text classifications for 77.78% of papers. At field scale, 13 of 15 topics map onto the taxonomy, while K-means and structural-topic-model domains agree for 73.83% of overlapping papers. LLM-Human Interactions accounts for 78.02% of domain-mapped topic mass, but venue analysis reveals a contrasting pattern: Social Minds and LLM Societies together account for 66.37% of highly cited papers in leading conference venues, whereas LLM-Human Interactions accounts for 76.81% in the corresponding journal subset. The resulting taxonomy provides a reproducible framework for understanding how model behavior, agent interaction, and institutional context jointly shape the social consequences of LLMs.","authors":["Yi Yang","Xiao Jia","Zeyun Dong","Chenzhang Wang","Zhanzhan Zhao"],"categories":["cs.CY","cs.AI","cs.CL"],"primary_category":"cs.CY","announce_type":"cross","date":"2026-09-09","first_seen":"2026-09-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.07598","pdf_url":"https://arxiv.org/pdf/2609.07598","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A3","B1"],"tags":["LLM社会模拟","文献综述","多智能体仿真"],"reason":"论文系统梳理LLM社会模拟研究，包含LLM Societies领域，涉及与人类…","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:04:29","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-09","rank":23,"question":"LLM 社会科学研究领域的主要主题和概念类别是什么，能否通过无监督聚类、作者分类和 LLM 分类一致地恢复出一个稳定的领域结构，并在大规模文献中验证该结构的可见性和变化？","design":"本研究不是仿真实验，而是对 LLM 社会科学文献进行系统映射。作者构建了一个由 198 篇全文精读论文组成的精选语料库，以及一个从五个文献数据库检索的 47,719 篇论文的大规模语料库。使用句子嵌入、K-means 聚类、聚类内 LDA、作者和 LLM 分类以及结构主题模型来识别和验证领域结构。","baseline":"无对照","findings":"识别出三个领域：LLM 作为社会心智、LLM 社会和 LLM-人类交互，包含 13 个子类别。在精选语料库中，三领域解决方案在重采样下高度稳定（调整兰德指数 = 0.952），K-means 分配与作者全文分类的一致性为 77.78%。在大规模语料库中，15 个主题中有 13 个映射到该分类法，K-means 和结构主题模型领域的一致性为 73.83%。LLM-人类交互占领域映射主题质量的 78.02%，但在高被引会议论文中，社会心智和 LLM 社会合计占 66.37%，而期刊论文中 LLM-人类交互占 76.81%。","reliability":"论文讨论了领域边界处的分歧，指出领域是可区分但可渗透的。作者承认分类一致性并非完美，并指出分歧集中在语义边界附近。此外，大规模分析中部分主题未能映射到分类法，表明分类法可能未完全覆盖所有研究。","relevance":"该论文为 LLM 社会模拟研究提供了系统的分类框架，其中 LLM 社会领域直接涉及基于模型代理的集体动态和模拟，与研究者关注的 LLM 作为人类被试替代品的研究高度相关。值得阅读原文以了解该领域的整体结构和关键文献。","inspiration":"该论文的方法论——结合无监督聚类、主题建模和人工分类来验证领域结构——可借鉴用于构建经济金融领域中 LLM 仿真研究的系统地图，识别核心主题和空白。｜可迁移到经济金融中的具体问题，如 LLM 在资产定价实验、消费者行为模拟或政策评估中的应用，通过分类框架定位现有研究并发现跨领域联系。｜一个可行的研究设计是：以经济金融领域已发表的 LLM 仿真论文为语料，使用句子嵌入和 K-means 聚类识别主题，然后与作者分类对比验证；针对特定主题（如 LLM 在拍卖或议价实验中的行为），设计仿真实验，将 LLM 作为被试，施加不同信息或激励处理，测量出价或决策结果，并与真实人类实验数据对照，评估仿真有效性。"}},{"id":"2609.07944","version":1,"title":"CausalVerify: An Execution-Grounded Benchmark for LLM Causal Inference Workflows","zh_title":"CausalVerify：面向LLM因果推断工作流的执行基准","abstract":"Existing causal-inference benchmarks for LLMs mostly score method descriptions or whether generated code runs, not whether the executed workflow recovers the target causal estimate. CausalVerify studies this verification problem for structured econometric causal-estimation workflows by separating realistic interpretation from verifiable computation. It pairs 259 published economics papers (reconstructed research question, data description, institutional context) with 100 fixed-seed synthetic scenarios that realise CSV datasets for difference-in-differences, event study, instrumental variables, and regression discontinuity designs. Experiment A (real-paper text agreement) scores method-family and direction agreement against four-LLM consensus labels. Experiment B (synthetic execution) runs model-written R code and checks whether the extracted treatment-effect estimate matches a canonical estimator on the same realised dataset; this execution-grounded correctness layer is L2b+, distinct from L2b, which records only whether the code executes. A calibration arm asks whether self-reported confidence separates correct from incorrect workflows. On Experiment B, seven LLMs reach L2b+ pass rates of 10% to 88% at the default 50% tolerance, and 66 of the 426 workflows that execute (15.5%) return a wrong estimate. Execution ranking (L2b) agrees with L2b+ far better than text-direction scoring (L4): Kendall $\\tau=0.81$ and Spearman $\\rho=0.93$, versus Kendall $\\tau$ between $-0.20$ and $0.10$ for L4. Llama-3.3-70B-Instruct shows the same qualitative gap, and reported confidence does not reliably separate correct from incorrect workflows. The claims are confined to standardized single-shot workflows in these four design families under the evaluated R backend and model panel; the benchmark does not measure general causal-inference ability. Code, data, cached outputs, and a datasheet are released.","authors":["Yonghong Zhang","Ricardo Correia","Isabel M. Parra","Yong Xie"],"categories":["cs.AI","cs.CL","econ.EM"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-09-09","first_seen":"2026-09-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.07944","pdf_url":"https://arxiv.org/pdf/2609.07944","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["B3"],"tags":["因果推断","基准测试","统计推断"],"reason":"评估LLM因果推断工作流，涉及统计推断有效性，方法可迁移至仿真可靠性评估。","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:04:31","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-09","rank":24,"question":"如何验证大语言模型在结构化计量经济学因果推断工作流中能否恢复目标因果估计，而非仅生成可运行代码或文本上正确的方法描述。","design":"该研究不是人类仿真实验，而是构建了一个执行基准：用259篇已发表经济学论文提供现实背景，用100个固定种子的合成场景生成CSV数据集，覆盖DID、事件研究、工具变量和断点回归四种设计；让LLM编写R代码并执行，检查提取的处理效应估计是否与同一数据集上的规范估计量一致。","baseline":"无对照","findings":"七个LLM在L2b+（执行后估计正确）上的通过率从10%到88%不等，且15.5%的可执行工作流返回了错误估计。执行成功排名与参考一致性高度相关（Kendall τ=0.81），而文本方向一致性排名与执行正确性弱相关甚至负相关。","reliability":"论文承认其结论仅限于四种设计家族、标准化单次工作流、特定R后端和模型面板，不衡量一般因果推断能力；自我报告的置信度不能可靠区分正确与错误工作流。","relevance":"该研究虽非人类仿真，但其执行验证方法可直接迁移到评估LLM作为人类被试替代品的可靠性，特别是当仿真涉及因果推断或政策评估时，值得阅读原文以借鉴其基准设计。","inspiration":"值得借鉴的做法是将文本判断与执行验证分离，用固定种子合成数据提供可验证的目标估计，并检查置信度校准。｜可迁移到政策评估场景，如用LLM模拟个体对政策变化的响应并估计处理效应。｜设计雏形：以LLM作为虚拟被试，施加政策处理（如税收变化），结果变量为报告的行为意图，用真实调查数据（如美国消费者财务调查）作为对照，检验LLM估计的处理效应是否与人类数据一致。"}},{"id":"2609.05663","version":1,"title":"What LLM Trading Agents Actually Do in Production: A Six-Month, Population-Scale Record from Two Fleets","zh_title":"生产环境中LLM交易代理的实际行为：来自两个机群的六个月、群体规模记录","abstract":"We present a continuous, population-scale measurement record of autonomous language-model trading agents operating in production across two systems with one design lineage: DX Terminal Pro (3,505 user-funded vaults trading real ETH in Base memecoin markets for 21 days, February to March 2026) and the DXAP live alpha fleet (500 to 599 user-created agents all-history, 91 to 117 concurrently active, trading Hyperliquid perpetuals, June to August 2026). The record spans roughly six months, 7.5M single-model invocations with about 300K onchain actions, and a further 231,638 multi-tool turns producing 14,596 fills. Four findings carry the paper. First, the operating layer determines behavior more than anything written in strategy text: a risk slider explains leverage (+0.425 per level), agent fixed effects absorb 60% of variance, and a leaderboard render boundary causally routes selection (regression discontinuity 1.75x at the top-3 cut). Second, sizing is volatility-blind: median leverage is 5.0x in every volatility sextile, and one posture-slider cell (11% of the book) holds 62% of liquidations. Third, agents capture almost none of the upside they reach: 43.2% of positions saw at least +300 bps of favorable excursion within 24h, yet 49.3% of those closed with a negative trade return; a mechanical bracket recovers +39.0 bps per position. Fourth, neither fleet shows a directional edge. The DXAP fleet is not profitable and trails a matched Hyperliquid retail benchmark (41% vs. 50% roundtrip win rate). A paired-replay league of frontier models on 416 captured production scenarios finds decision quality statistically indistinguishable at this horizon, while choice stability differs sharply across model families. Every headline survives day-clustered inference, permutation nulls, and a common-fee restatement; the paper closes with a 17-rule methodology canon bought with our own retractions.","authors":["T. J. Barton","Chris Constantakis","Patti Hauseman","Annie Mous","Alaska Hoffman","Brian Bergeron","Hunter Goodreau"],"categories":["cs.AI","cs.CE","cs.MA"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-09","first_seen":"2026-09-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.05663","pdf_url":"https://arxiv.org/pdf/2609.05663","source_feed":"cs.AI","score":7,"bucket":"pending","rubric_hits":["A3","B1","B2","B4"],"tags":["LLM代理","市场仿真","行为偏差"],"reason":"用LLM交易agent群体模拟市场行为，并与真实人类交易数据对照，评估其决策偏…","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:04:46","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-09","rank":20,"question":"在真实生产环境中，数千个由用户配置的LLM交易代理实际行为如何，哪些系统设计因素决定其行为，以及它们是否具有方向性交易优势？","design":"论文对两个生产系统（DX Terminal Pro和DXAP）中的LLM交易代理进行大规模观测，记录其配置（滑块、策略文本）、推理调用和链上行为，分析行为与系统设计因素的关系，并测量交易结果。","baseline":"DXAP舰队与匹配的Hyperliquid零售交易者基准进行对比（胜率41% vs 50%）。","findings":"操作层（滑块、渲染列表、订单路径）比策略文本更能决定行为；代理的仓位规模对波动率不敏感，且几乎无法兑现所达到的有利偏移。两个舰队均无方向性优势，DXAP不盈利且落后于零售基准。","reliability":"论文采用日聚类推断、置换零假设和统一费率重述进行稳健性检验，并承认三个早期结果被撤回；DXAP纸面引擎存在零滑点、零资金费率等简化，可能高估表现。","relevance":"该研究提供了LLM代理在真实金融环境中的大规模行为记录，并与人类零售交易者对照，对评估LLM仿真人类交易行为的可靠性具有直接参考价值，值得阅读原文。","inspiration":"借鉴其操作层作为处理变量的设计，通过系统参数（如风险滑块）的准实验变化识别因果效应｜可迁移到资产定价实验或散户交易行为研究，例如检验界面设计对风险承担的影响｜设计一个实验：用LLM代理模拟散户投资者，随机分配不同的交易界面约束（如杠杆滑块范围），测量其杠杆选择和交易频率，并与真实散户交易数据（如某券商账户级数据）对照，评估仿真偏差。"}},{"id":"2609.08288","version":1,"title":"LEBGen: An LLM-Enhanced Bayesian Network Framework for Few-Shot Travel Survey Data Generation","zh_title":"LEBGen：一种用于少样本出行调查数据生成的LLM增强贝叶斯网络框架","abstract":"Travel survey data are essential for transportation planning and travel behavior analysis, yet collecting large-scale representative samples is costly and time-consuming. A practical alternative is to generate synthetic survey records from a few-shot sample. However, such samples provide incomplete coverage of heterogeneous traveler groups and insufficient evidence for recovering the complex dependencies between demographic characteristics and travel behavior. Existing approaches have complementary limitations. Probabilistic generative models such as Bayesian networks (BNs) offer explicit distributional control, but structures learned from few-shot samples may omit meaningful dependencies or retain spurious ones. Large language models (LLMs) can help address these difficulties in BN structure learning by providing behavioral knowledge that complements the limited statistical evidence. We therefore propose LEBGen, an LLM-enhanced BN framework that uses this knowledge to refine network structure for few-shot travel survey data generation. Specifically, the LLM first identifies traveler personas from demographic attribute and travel behavior statistics, then recovers dependencies missed by the persona-augmented BN structure and prune spurious ones. The refined BN is parameterized exclusively from the observed data to generate synthetic records. Under a 2% few-shot setting on the 2022 Hong Kong Travel Characteristics Survey, LEBGen reduces the mean marginal Jensen-Shannon divergence from 0.0671 to 0.0091 and the mean absolute Cramer's V error by 14.3% over the best-performing baseline, substantially improving both distributional and dependency fidelity.","authors":["Zijian Shen","Bin Zhou","Jiguang Wang","Ya Zhao","Jintao Ke"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-09","first_seen":"2026-09-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.08288","pdf_url":"https://arxiv.org/pdf/2609.08288","source_feed":"cs.AI","score":7,"bucket":"pending","rubric_hits":["A5","B1","B2"],"tags":["合成数据生成","出行调查","LLM增强"],"reason":"用LLM增强贝叶斯网络生成旅行调查合成数据，有真实数据对照，属数据增强，可迁移…","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:04:34","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-09","rank":26,"question":"如何在少量真实旅行调查样本下，利用大语言模型增强贝叶斯网络结构学习，以生成高保真合成旅行调查数据？","design":"提出 LEBGen 框架，用 LLM 从少量样本的统计信息中提取旅行者画像（persona），并作为贝叶斯网络的新节点；再用 LLM 评估和修正网络边（增删或反转），最后仅用观测数据估计参数并生成合成记录。在 2022 年香港出行特征调查的 2% 少样本设置下评估。","baseline":"2022 年香港出行特征调查（Hong Kong Travel Characteristics Survey）的真实数据，取 2% 作为少样本训练集，其余作为评估基准。","findings":"LEBGen 将平均边际 Jensen-Shannon 散度从 0.0671 降至 0.0091，平均绝对 Cramér's V 误差比最佳基线降低 14.3%，显著提升了分布保真度和依赖关系保真度。LLM 提供的先验知识有效弥补了少样本下统计证据不足的问题。","reliability":"论文未讨论","relevance":"该研究利用 LLM 作为先验知识来源增强贝叶斯网络，在少样本条件下生成与真实分布高度一致的合成调查数据，对关注 LLM 仿真可靠性和数据增强的研究者有参考价值，值得阅读原文了解其结构修正机制和评估细节。","inspiration":"借鉴其将 LLM 作为先验知识注入概率图模型、并用真实数据仅做参数估计的做法，可避免 LLM 直接生成数据带来的幻觉偏差。｜可迁移到消费者金融行为调查或小微企业信贷需求调查的少样本合成，用于政策模拟或风险评估。｜以某地区家庭金融调查的小样本为训练集，用 LLM 提取家庭财务画像并修正贝叶斯网络结构，生成合成家庭金融数据，再与全量真实调查数据比较边际分布和变量间依赖（如收入与风险资产持有的关联），评估合成数据在信贷需求预测模型中的效用。"}},{"id":"2609.08861","version":1,"title":"API Benchmark Scores Do Not Reliably Transfer to Chatbot Interfaces","zh_title":"API基准分数不能可靠迁移到聊天机器人界面","abstract":"Benchmark scores are a central currency in model releases: they inform purchasing decisions, shape public trust, and influence policy. Yet, a key assumption underlying benchmark scores is that the model performance measured through APIs faithfully reflects the behavior of deployed systems. We challenge this assumption by auditing ChatGPT, Claude, and Gemini across seven systems and nine benchmarks spanning general capability, social bias, and sycophancy. We find systematic API--interface differences in both accuracy and consistency. On average, API evaluations score 3.4 percentage points higher in accuracy and 2.1 percentage points higher in test--retest agreement than corresponding interface evaluations. For ChatGPT, the performance difference between API and interface access exceeds the API-only difference between GPT 5.3 and GPT 5.4. Put differently, switching access surfaces can degrade performance as much as downgrading a full model generation. We further test whether exposed API controls can reproduce interface behavior by varying system prompts, sampling parameters, and reasoning settings. These controls shift behavior in some cases but do not reliably eliminate the gap. Our findings document a context-validity gap: measurements obtained through APIs do not necessarily generalize to corresponding deployed interfaces, complicating the use of API evaluations as proxies for deployed systems.","authors":["Jennifer Wang","Joachim Baumann","Daniel E. Ho","Sanmi Koyejo"],"categories":["cs.AI","cs.SE"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-09","first_seen":"2026-09-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.08861","pdf_url":"https://arxiv.org/pdf/2609.08861","source_feed":"cs.AI","score":7,"bucket":"pending","rubric_hits":["A2","B4"],"tags":["模型审计","可靠性","界面差异"],"reason":"审计API与聊天界面差异，揭示模型行为不一致，对仿真可靠性有启示","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:05:22","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-09","rank":29,"question":"API基准分数能否可靠迁移到聊天机器人界面？","design":"审计ChatGPT、Claude、Gemini的7个系统，在API和网页界面两种访问方式下运行9个基准（涵盖通用能力、社会偏见、谄媚），比较准确率和重测一致性。","baseline":"无对照","findings":"API评估平均比界面评估准确率高3.4个百分点，重测一致性高2.1个百分点；ChatGPT的API与界面性能差异超过GPT 5.3与5.4的API-only差异。暴露的API控制（系统提示、采样参数、推理设置）不能可靠消除差距。","reliability":"论文未讨论","relevance":"该研究揭示API与真实部署界面存在系统性行为差异，对依赖API进行人类仿真或行为测量的研究构成直接威胁，值得精读以了解失效条件。","inspiration":"借鉴其对照设计：同一模型在API与界面两种访问方式下施加相同提示，测量行为差异，并检验API控制参数能否复现界面行为。｜可迁移到经济金融中的LLM辅助决策场景，如用LLM模拟消费者对金融产品条款的反应或投资者对政策公告的解读。｜以LLM作为被试，处理为通过API或聊天界面呈现同一经济决策任务（如信贷申请或投资选择），结果变量为决策一致性与准确性，对照真实人类实验数据（如实验室或调查数据）评估仿真效度。"}},{"id":"2609.08049","version":1,"title":"LLMs for Social Network Modeling: From Network Generation to Dynamic Processes","zh_title":"用于社会网络建模的大语言模型：从网络生成到动态过程","abstract":"Large language models (LLMs) are rapidly emerging as a new paradigm for modeling social networks by representing users and their relationships and interactions through natural language. Unlike classical network models or deep learning approaches, LLMs can simulate context-aware social behavior and language-driven interactions, enabling more realistic modeling of network formation and dynamic social processes. However, existing studies are scattered across different research communities and lack a unified perspective. This survey presents the first comprehensive review of LLMs for social network modeling by organizing the literature into two broad categories: network generative models and dynamic process models. Network generative models are further classified into selection-based and interaction-based approaches, while dynamic process models are categorized into opinion dynamics, information diffusion, and rumor propagation, each with their underlying modeling mechanisms. LLMs enable rich textual social interactions and decision-making, but they also exhibit many limitations, including inherent social biases and prompt sensitivity. We outline these open research challenges and discuss future directions in LLM-based social network modeling.","authors":["Shikha Mallick","Alex Thomo","Akrati Saxena"],"categories":["cs.SI","cs.AI"],"primary_category":"cs.SI","announce_type":"cross","date":"2026-09-09","first_seen":"2026-09-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.08049","pdf_url":"https://arxiv.org/pdf/2609.08049","source_feed":"cs.AI","score":7,"bucket":"pending","rubric_hits":["A3","B4"],"tags":["社会网络建模","LLM仿真","综述"],"reason":"综述LLM模拟社会网络动态，含意见扩散等社会过程，与人类仿真相关，但缺人类数据…","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:04:34","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-09","rank":25,"question":"如何系统梳理和分类利用大语言模型进行社会网络建模的研究，包括网络生成模型和动态过程模型？","design":"这是一篇综述论文，没有进行新的仿真实验。它系统回顾了现有文献，将基于LLM的社会网络建模方法分为网络生成模型（选择型和交互型）和动态过程模型（意见动力学、信息扩散、谣言传播），并比较了各研究的LLM骨干、数据集、微调策略、评估协议和代码可用性。","baseline":"无对照","findings":"LLM能够模拟具有上下文感知的社会行为和语言驱动的交互，从而更真实地建模网络形成和动态社会过程。然而，LLM存在固有的社会偏见和提示敏感性等局限性，并且现有研究缺乏统一的视角和验证。","reliability":"论文指出LLM在社会网络建模中存在现实主义、可扩展性、偏见、可重复性以及针对真实世界行为的验证等挑战。","relevance":"该综述与研究者关注的人类仿真实验高度相关，因为它系统梳理了LLM模拟社会网络动态（包括意见扩散等社会过程）的文献，但缺少与真实人类数据的对照，因此可作为了解领域全貌和识别方法缺陷的起点。","inspiration":"这篇综述提供了对LLM社会网络仿真的系统分类和机制分析，可借鉴其分类框架来设计经济金融领域的仿真实验。｜可以迁移到金融市场中的信息扩散和投资者情绪传播、政策公告的预期形成、以及消费者行为中的社会影响等场景。｜例如，用LLM扮演异质投资者，施加不同的信息冲击（如公司财报或政策变化），测量其交易决策和价格形成，并与真实市场数据（如股价波动、交易量）进行对照，以评估仿真的外部有效性。"}},{"id":"2608.10503","version":3,"title":"Every Token Counts: Exact Likert-Scale Distributions for Measuring LLM Attitudes and Biases","zh_title":"每个Token都重要：用于测量LLM态度与偏差的精确Likert量表分布","abstract":"As Large Language Models (LLMs) are increasingly deployed as autonomous agents, accurately evaluating their latent values and biases is critical. The NLP community typically evaluates models using large, unstructured benchmarks. While effective for general capabilities, these datasets fundamentally conflate causal mechanisms: even when an aggregate bias is detected, unstructured evaluations cannot disentangle whether it stems from baseline traits, contextual confounders, or complex interactions. To address this, we introduce an analytically exact framework for the controlled behavioral evaluation of LLMs. We bridge human psychometrics with LLM mechanics by resolving gaps in design, measurement, and analysis. First, we replace unstructured prompting with fully crossed factorial experiments to systematically isolate causal main and interaction effects. Second, we eliminate Monte Carlo text sampling noise by operating directly on exact, token-level Probability Mass Functions (PMFs). Third, we derive a multivariate ordinal consensus metric and a distributional ANOVA to process these PMFs analytically. We validate our framework with a case study on consumer ethnocentrism across five LLMs, demonstrating how our approach isolates systemic country-of-origin biases that aggregate benchmarks otherwise obscure.","authors":["Davood Wadi","Mohsen Ghodrat","Matthew Philp"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-09","first_seen":"2026-08-12","revised_at":"2026-09-09","abs_url":"https://arxiv.org/abs/2608.10503","pdf_url":"https://arxiv.org/pdf/2608.10503","source_feed":"cs.CL","score":6,"bucket":"other","rubric_hits":["D2","A2"],"tags":["LLM态度测量","心理测量学","算法偏差"],"reason":"测量LLM态度与偏差，属D2边界；但方法可迁移至仿真可靠性评估，故给6分。","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:05:25","error":null,"has_summary":false,"summary":null},{"id":"2608.27111","version":3,"title":"Animarium: an open, reproducible pipeline for synthetic populations of Italian cities, from ISTAT sources to open data (Tech Report v1)","zh_title":"Animarium：意大利城市合成人口的开放可复现流水线，从ISTAT来源到开放数据（技术报告v1）","abstract":"Synthetic populations of eleven Italian municipalities (1,814,317 individuals in 887,937 households) generated from published aggregates alone: ISTAT census and register tables, census-section counts, the national civic-address register, public-use survey microdata, and six municipal open-data portals, every source certified in a registry with licence, fingerprint and declared affordances. Four rings give every attribute a declared place: a maximum-entropy joint model of up to nine demographic attributes; whole-vector donation of twenty-three attitudinal and health variables from survey respondents; placement to census section, single year of age and address; and households constrained by the census size distribution per section. Every downstream layer (detailed titles, work, names, biographies) is a declared derivation adding no information. The pipeline is deterministic to the byte: regenerating all eleven municipalities from the tagged commit reproduces every file of every ring bit for bit, in 33 minutes on one workstation. Populations are released in a public regime enforced in the data (no names, no addresses, coordinates randomised within census section), browsable in Animarium, a dependency-free web viewer where every number carries its comparison and every view is a citable URL, and downloadable as an open dataset. The report documents the architecture, the sources and their certification, the reproducibility and quality measurements at the release tag, the viewer, and the narrative layer that renders records into personas for LLM-driven simulation, with the platform's controllability demonstrated in companion experiments, and validation explicitly out of scope.","authors":["Mirko Degli Esposti"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"replace","date":"2026-09-09","first_seen":"2026-08-28","revised_at":"2026-09-09","abs_url":"https://arxiv.org/abs/2608.27111","pdf_url":"https://arxiv.org/pdf/2608.27111","source_feed":"cs.CY","score":6,"bucket":"other","rubric_hits":["D3"],"tags":["合成人口","LLM仿真","数据流水线"],"reason":"生成合成人口用于LLM仿真，但无人类数据对照，验证明确排除","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:05:30","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-01","rank":17,"question":"如何从公开汇总数据生成意大利城市的合成人口，并使其可用于LLM驱动的人类仿真实验。","design":"该研究构建了一个确定性管道，利用最大熵联合模型拟合人口普查汇总数据生成个体属性，通过整向量捐赠从调查受访者获取态度和健康变量，并约束家庭结构以匹配普查分布，最终生成11个意大利城市的合成人口。","baseline":"使用ISTAT人口普查和登记表、普查分区计数、国家地址登记、公共使用调查微数据以及市政开放数据作为真实数据来源，但未进行与真实个体数据的直接验证对照。","findings":"管道能够从公开汇总数据生成大规模合成人口，且完全可复现，在33分钟内可逐字节重现所有文件。合成人口被设计为不含个人数据，并通过Animarium查看器提供可浏览和可下载的开放数据集。","reliability":"论文明确将验证排除在范围之外，未评估合成人口在LLM驱动仿真中的有效性；管道依赖手动获取的调查微数据，且拟合阶段存在依赖未固定版本求解器的风险。","relevance":"该研究为使用LLM进行人类仿真实验提供了可复现的合成人口生成管道，并包含真实数据来源，但缺乏对仿真可靠性的验证，值得阅读以了解其方法和局限性。","inspiration":"借鉴其分层生成和确定性复现的设计，确保仿真实验的可重复性和属性来源透明。｜可迁移到政策评估场景，如模拟不同城市居民对福利政策或公共服务的反应。｜以合成人口中的个体为被试，施加政策干预（如改变税收或补贴），测量其态度或行为变化，并与真实调查数据（如ISTAT调查）进行对照。"}},{"id":"2609.07117","version":1,"title":"The Illusion of Debiasing: Persona Steering Redistributes Rather Than Reduces Bias in LLMs","zh_title":"去偏的幻觉：人格引导在LLM中重新分配而非减少偏差","abstract":"Prompt-based interventions: system prompts, personas, role instructions, reliably reshape what a language model says, but it is unclear which layer they reach. Do they reconfigure internal structure, or only modulate the output channel? We use persona conditioning as a controlled probe, measuring its effects along a depth axis from self-report, through open-ended generation, to word-level parametric association, across three instruction-tuned models. We find a graded dissociation. Personas are legible but not structural: models follow single-trait instructions yet fail to reproduce human inter-trait covariance. The dissociation deepens with depth: personas hold or amplify closed-form QA bias, shift absolute tone while leaving between-group disparity unchanged, and barely perturb an already saturated associative baseline. Prompt-based steering thus operates in the output channel and has a structural reach limit that surface manipulability can mask.","authors":["Ziyue Feng","Hongbo Fang","James A. Evans"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-09","first_seen":"2026-09-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.07117","pdf_url":"https://arxiv.org/pdf/2609.07117","source_feed":"cs.CL","score":6,"bucket":"other","rubric_hits":["D2"],"tags":["LLM人格","偏差","提示干预"],"reason":"研究LLM人格操控对偏差的影响，测量模型本身而非仿真人类被试，但涉及人格与偏差…","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:05:03","error":null,"has_summary":false,"summary":null},{"id":"2609.07791","version":1,"title":"LLM Agents as Computational Typologists","zh_title":"LLM智能体作为计算类型学家","abstract":"Linguistic typology relies on expert analysis of reference grammars across languages, making large-scale crosslinguistic comparison labor-intensive and unscalable. We introduce AUTOTYPOLOGIST, an LLM agent for evidence-grounded typological analysis over reference grammars. The agent is capable of retrieving relevant grammar sections, analyzing interlinear glossed text (IGT), and iteratively reasoning over typological hypotheses using a ReAct-style workflow. We evaluate the system on TYPOLOGICAL FEATURE CODING against expert annotations and TYPOLOGICAL HYPOTHESIS TESTING with typological universals using 25 open-source reference grammars. Operating under different information constraints in TYPOLOGICAL FEATURE CODING, the agent can synthesize information from reference grammar prose but still faces challenges with only IGTs in the target language. In TYPOLOGICAL HYPOTHESIS TESTING, the agent can synthesize crosslinguistic evidence and identify both supporting cases and counterexamples. These findings suggest that LLM agents can support scalable and inspectable typological analysis, while still requiring expert validation.","authors":["Changbing Yang","Christopher Hammerly","Freda Shi","Jian Zhu"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-09","first_seen":"2026-09-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.07791","pdf_url":"https://arxiv.org/pdf/2609.07791","source_feed":"cs.CL","score":6,"bucket":"other","rubric_hits":["D1"],"tags":["LLM智能体","语言类型学","标注替代"],"reason":"LLM替代专家进行类型学标注，属替代人工劳动而非仿真人类被试，但方法可借鉴。","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:05:12","error":null,"has_summary":false,"summary":null},{"id":"2609.08515","version":1,"title":"Same Values, Different Languages? From Multilingual Probing to Steering LLMs Toward Chinese Social Values","zh_title":"相同价值观，不同语言？从多语言探测到引导大语言模型走向中国社会价值观","abstract":"As Large Language Models (LLMs) are increasingly integrated into human society, aligning them with pluralistic social values has become a critical priority. However, whether LLMs exhibit consistent value preferences across languages remains underexplored, particularly for culturally grounded values, which are more abstract and difficult to evaluate and align than safety-centric principles. We investigate this issue through Chinese Social Values (CSV), a value system rooted in Chinese culture and comprising $12$ dimensions across national, societal, and personal levels. We construct C-Voices, the first comprehensive multilingual contrastive probe dataset for CSV, with 86,400 dilemma-based instances in six languages, each pairing a CSV-aligned action with a value-conflicting alternative. Building on the contrastive probes of C-Voices, we then propose a fine-tuning-free value vector steering method that derives value directions from hidden-state discrepancies and selectively intervenes on value-sensitive layers during inference. Experiments on six languages show that CSV-oriented preferences are model-dependent and language-sensitive, with the same dilemma eliciting divergent responses across languages. Our method achieves effective CSV steering, supports cross-lingual transfer of value vectors, and generalizes to existing FLAMES and ValuePrism.","authors":["Yuemei Xu","Kexin Xu","Jian Zhou","Haoyu Lu","Yequan Wang","Aishan Liu"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-09","first_seen":"2026-09-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.08515","pdf_url":"https://arxiv.org/pdf/2609.08515","source_feed":"cs.CL","score":6,"bucket":"other","rubric_hits":["D2"],"tags":["价值观对齐","跨语言差异","模型探测"],"reason":"测量LLM的价值观偏好，属于D2边界情形，但涉及文化价值观和跨语言差异，与人类…","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:05:17","error":null,"has_summary":false,"summary":null},{"id":"2609.08637","version":1,"title":"Navigating the digital spectrum: Assessing political bias, stability, and downstream fairness in Large Language Models","zh_title":"导航数字光谱：评估大语言模型中的政治偏见、稳定性与下游公平性","abstract":"Large Language Models are increasingly deployed as information intermediaries, yet measuring their political behavior remains fragile because questionnaire results mix model dispositions with measurement artifacts and response-elicitation biases. We introduce a robust Political Compass Test evaluation framework that samples 300 configurations across an eight-dimensional perturbation space varying language, framing, instructions, answer format, option order, and persona wording. We evaluate eight Gemma 3 and Qwen 3 models across 14 languages and three quantization levels, obtaining design-averaged political coordinates with quantified uncertainty. Most models lean Libertarian-Left on average, but instruction phrasing, language, and answer format significantly affect recovered coordinates. Cross-lingual differences primarily reflect coordinate drift rather than distinct cultural reasoning. Reverse-engineering the test also exposes axis-weighting imbalances and the collapse of degenerate responses toward the center, so near-origin estimates for the smallest models can reflect weak signal rather than centrism. Free-text reasoning and chat-then-classify elicitation alter recovered coordinates, and larger models show clearer persona separation, with a specific failure of the Authoritarian-Left persona to move most models in the intended social direction. In downstream tasks, persona effects are modest relative to model size and target group for hate-speech detection, while base and centrist prompts give the highest agreement for topic-level sentiment. Political role prompting therefore has measurable but task- and dataset-specific downstream effects.","authors":["Luka Debevc","Nishan Chatterjee","Antoine Doucet","Senja Pollak","Matej Martinc"],"categories":["cs.CL","cs.CY"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-09","first_seen":"2026-09-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.08637","pdf_url":"https://arxiv.org/pdf/2609.08637","source_feed":"cs.CL","score":6,"bucket":"other","rubric_hits":["D2"],"tags":["LLM政治偏见","测量稳定性","下游公平性"],"reason":"测量LLM政治倾向与稳定性，属模型测量而非仿真人类被试，但方法可借鉴。","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:04:34","error":null,"has_summary":false,"summary":null},{"id":"2609.08592","version":1,"title":"A Three-Tier Persona Vector for Controllable User Simulation in Agentic Evaluation","zh_title":"用于智能体评估中可控用户仿真的三层人设向量","abstract":"Evaluating tool-augmented LLM agents requires diverse, realistic user inputs yet most evaluation frameworks use flat role descriptions (\"you are an angry customer\") that produce near-identical conversations regardless of the underlying scenario. In this paper, we propose a three-tier persona vector with 23 operationalized dimensions: 6 categorical demographics (jurisdiction, age, channel, device, language proficiency, time availability), 12 continuous behavioral traits (patience, assertiveness, digital literacy, etc.) sampled with Gaussian noise around curated profile base vectors, and 5 continuous emotional states (frustration, anxiety, trust, confidence, stress) that shift in response to scenario context. Orthogonal to the persona, a 4-level query-complexity overlay controls utterance phrasing from direct to deliberately vague. We evaluate the persona model inside a synthetic data generation pipeline across 64,698 multi-turn conversations spanning 8 named profiles and 3 production corpora. Key findings: (i) a 15.8 percentage-point spread in agent goal-achievement across personas confirms trait vectors produce measurably different user behavior; (ii) the same persona behaves differently across scenarios due to scenario-reactive emotional state shifts, validating the scenario-reactive design; (iii) domain-specific projects show persona sensitivity on booking-flow compliance (~15-20 percentage points gap between tier-aware and pressure-test personas), demonstrating the model faithfully reproduces real-world difficulty distributions; (iv) seven rule-described trait correlations produce auditable co-occurrence patterns without requiring learned covariance matrices. The persona model is fully specified for reproduction.","authors":["Rahul Khedar","Eshita","Sneha Teja Sree Reddy Thondapu","Mayank Malhotra","Arup Kumar Das","Jitesh Chandra Mishra","Arun Menon","Avinash Karn","Mouli V"],"categories":["cs.AI","cs.CL"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-09-09","first_seen":"2026-09-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.08592","pdf_url":"https://arxiv.org/pdf/2609.08592","source_feed":"cs.CL","score":6,"bucket":"other","rubric_hits":["D3"],"tags":["用户仿真","智能体评估","人设建模"],"reason":"用LLM生成可控用户仿真，但无真实人类数据对照，属社会模拟边界情形。","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:05:19","error":null,"has_summary":false,"summary":null},{"id":"2609.07943","version":1,"title":"Beliefs and Behavior in Language Models","zh_title":"语言模型中的信念与行为","abstract":"There is significant uncertainty about whether abstractions like beliefs or desires usefully describe the behavior of large language models (LLMs). In addition to the inherent scientific interest of this question, these latent quantities are often invoked to explain the behavior of LLMs to users or to define and evaluate harmful behaviors which are relative to intent. Nevertheless, we currently lack a means to systematically test whether concepts like \"belief\" are well-applied to LLMs, and hence whether they are likely to be fruitful ingredients of attempts to align models with human interests. We propose an approach for empirically studying such questions, asking whether a single latent variable inferred from the LLMs' outputs -- interpreted as a degree of belief -- allows an observer to make interpretable predictions of how the LLMs' will respond to new prompts. We find that highly capable models are usefully described as holding beliefs and that, generally, the predictability of model outputs based on an inferred latent belief tracks overall trends in model capability. Building on these findings, we provide empirical strategies to study how beliefs in LLMs can be measured, the extent to which LLMs comply with instructed decision rules or payoffs, and how beliefs evolve within individual instances of an LLM over the course of reasoning.","authors":["Alex Smolin","Bryan Wilder"],"categories":["cs.AI","cs.LG"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-09","first_seen":"2026-09-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.07943","pdf_url":"https://arxiv.org/pdf/2609.07943","source_feed":"cs.AI","score":6,"bucket":"other","rubric_hits":["D2"],"tags":["LLM信念","模型行为","可解释性"],"reason":"研究LLM的信念等内部状态，属于测量模型本身而非仿真人类被试，但方法可借鉴。","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:04:30","error":null,"has_summary":false,"summary":null},{"id":"2609.05591","version":1,"title":"WolfSociety: Understanding Collective Risk from Harmful-Agent Scaling in Financial Agent Societies","zh_title":"狼群社会：理解金融智能体社会中有害智能体规模化带来的集体风险","abstract":"Safety evaluations typically focus on individual agents, but interacting agents can spread harmful information and influence the environment in which later decisions are made. We study how collective failure changes with harmful-agent fraction and society size in a controlled financial agent society, where agents communicate over a social network and trade in a shared market. In the primary financial scenario, collective failure requires broad harmful diffusion together with severe price dislocation or liquidity stress. Across all tested society sizes, failure remains rare at low harmful fractions but rises sharply over a narrow range. As society size grows from N=100 to N=2000, the harmful fraction associated with a 50% failure probability decreases from 4.7% to 2.2%, while the corresponding number of harmful agents increases from approximately 5 to 44. In contrast, when the number of harmful agents is held fixed, their impact becomes weaker as the society grows. Controlled interventions further show that broader network reach shifts the collapse boundary toward lower harmful fractions, whereas stronger conformity alone has little effect. To characterize these effects, we introduce Agent Society Dynamics, a finite-size framework for relating harmful-agent fraction, society size, and interaction structure to collective failure. Overall, our results reveal a nonlinear, size-dependent collapse transition in financial agent societies, showing that collective failure depends not only on the prevalence of harmful agents but also on the size and interaction structure of the surrounding society. Code is available at https://github.com/SAIL-Research-Lab/WolfSociety.","authors":["Lejun Zhang","Sarah Lu-Liang","Xin Jiang","Muning Wen","Weinan Zhang","Shangding Gu"],"categories":["physics.soc-ph","cs.AI","cs.SI"],"primary_category":"physics.soc-ph","announce_type":"cross","date":"2026-09-09","first_seen":"2026-09-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.05591","pdf_url":"https://arxiv.org/pdf/2609.05591","source_feed":"cs.AI","score":6,"bucket":"other","rubric_hits":["D3"],"tags":["LLM智能体社会模拟","金融系统风险","多智能体交互"],"reason":"LLM agent 模拟金融社会，但无真实人类数据对照，属社会模拟边界情形。","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:04:45","error":null,"has_summary":false,"summary":null},{"id":"2608.18091","version":2,"title":"Self- and Other-Labels Induce Bidirectional Bias in LLM Judges","zh_title":"自我与他人标签在LLM法官中诱发双向偏差","abstract":"As LLM-as-a-judge becomes increasingly widespread, self-preference -- the tendency of a judge to favor its own outputs -- raises growing concerns about evaluation reliability. However, this bias has been studied predominantly on generated text, where stylistic features and response quality are inevitably conflated. As a result, existing measurements cannot separate genuine self-preference from these confounds. We address this limitation by changing the object of evaluation: instead of judging generated text, ten LLMs assess sets of narrative constraints selected from a shared pool, which carry no stylistic fingerprint yet retain a recoverable model-specific signature. Two experiments on this task yield complementary findings. Under blind evaluation, self-preference disappears, with a small effect remaining in the opposite direction once selection quality and judge severity are controlled. Under matched quality, however, self- and other-labels alone -- without naming any model -- shift scores bidirectionally. LLM judges inflate scores for self-labeled selections and deflate those for other-labeled ones regardless of the selection's actual source. We make two contributions: 1) authorship attribution is a distinct driver of evaluation bias, and 2) ground-truth-free tasks can serve as controlled instruments for studying LLM judge behavior.","authors":["Songeun Chae","Min Kim","Donghoon Jung","Seojin Choi","Seohyon Jung"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-09","first_seen":"2026-08-20","revised_at":"2026-09-09","abs_url":"https://arxiv.org/abs/2608.18091","pdf_url":"https://arxiv.org/pdf/2608.18091","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["LLM评估偏差","自我偏好","算法审计"],"reason":"研究LLM法官的自我偏好偏差，属于对LLM本身行为的测量，而非用LLM仿真人类…","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:05:29","error":null,"has_summary":false,"summary":null},{"id":"2609.06212","version":1,"title":"SLATE: Are AI-Generated Slides Educationally Effective? A Benchmark for Language Teaching Quality and Learner Knowledge Acquisition","zh_title":"SLATE：AI生成的幻灯片在教育上有效吗？语言教学质量与学习者知识获取的基准","abstract":"LLMs have achieved remarkable capabilities in generating language teaching slides. However, a critical mismatch persists between visual polish and actual instructional effectiveness. To address this gap, we introduce SLATE (Slide-based Learning Assessment for Teaching Effectiveness), the first benchmark that evaluates AI-generated language teaching slides through instructional effectiveness and learner knowledge acquisition. SLATE transforms linguistics olympiad puzzles from low-resource languages with negligible web presence into 90 standardized instructional units comprising 1,133 assessable items, paired with a structured course outline and matched near- and far-transfer test sets. This pretest-posttest design eliminates pretrained knowledge leakage, ensuring gains reflect learning rather than prior recall. Using VLMs as scalable learner proxies and directionally supported by a three-system human pilot, our results show that content validity exhibits a weak association with learning gain, while pedagogical design exhibits a robust positive association. Moreover, most systems show a significant gap between near- and far-transfer accuracy, and even frontier models can produce negative learning gains. SLATE reveals a dissociation between artifact quality and instructional effectiveness, calling for a paradigm shift in how generative teaching systems are built, evaluated, and deployed.","authors":["Jingzhuo Wu","Jiajun Zhang","Liu Yi","Leqi Zheng","Yuheng Jing","Xinyuan Zhou","Quan yang"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-09","first_seen":"2026-09-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.06212","pdf_url":"https://arxiv.org/pdf/2609.06212","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM评估","教育技术","代理评估"],"reason":"用VLM作为学习者代理评估教学效果，替代人类被试，但非社会行为仿真，属边界情形。","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:04:51","error":null,"has_summary":false,"summary":null},{"id":"2609.06289","version":1,"title":"Steering Geometry: Validating Human Value Geometry in LLM Steering Space","zh_title":"引导几何：在LLM引导空间中验证人类价值几何","abstract":"As large language models (LLMs) are increasingly deployed in alignment-sensitive contexts, activation steering has emerged as a lightweight, inference-time alternative to fine-tuning methods (e.g., RLHF, DPO) for behavioral control. However, existing work typically validates steering on isolated behaviors, leaving it unclear whether steering vectors encode coherent semantic structure or merely exploit behavior-specific shortcuts. We investigate whether the latent geometry of LLM steering vectors reflects theory-specified structure in human values and morality. Using Schwartz's Theory of Basic Human Values as our primary fine-grained framework, we introduce a 26K-sample benchmark covering 20 human values and analyze distribution-driven methods (e.g., CAA, SphericalSteer, ODESteer) and behavior-centric approaches (e.g., COLD-Steer, BiPO) across diverse model families and sizes. We find that distribution-driven methods recover human value topologies aligned with theoretical predictions (Spearman $\\rho$ up to 0.51, $p < 10^{-13}$). In contrast, behavior-centric methods achieve comparable steering performance but show little correlation with the expected value geometry. Geometric fidelity improves with model scale but drops after instruction tuning. Finally, better geometric alignment also leads to more human-consistent transfer across values: steering one value correctly lifts compatible values and suppresses opposing ones. Code and data are available at: https://github.com/DeepRCL/Steering_Geometry.","authors":["Mohammad Mahdi Abootorabi","Armin Saghafian","Ali Bazshoushtari","Hamid Rezaei","EunJeong Hwang","Vered Shwartz","Parvin Mousavi","Purang Abolmaesumi"],"categories":["cs.CL","cs.AI","cs.LG"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-09","first_seen":"2026-09-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.06289","pdf_url":"https://arxiv.org/pdf/2609.06289","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["激活引导","价值对齐","模型可解释性"],"reason":"研究LLM内部价值几何与人类理论结构的一致性，属于测量模型本身而非仿真人类被试…","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:04:53","error":null,"has_summary":false,"summary":null},{"id":"2609.06851","version":1,"title":"You Are What You Read: Misalignment via In-Context Persona Induction","zh_title":"你读什么就是什么：通过上下文人格诱导导致的不对齐","abstract":"Broad misalignment has been produced by finetuning on narrow data, harmful or benign, and in context only by demonstrations of the undesirable behaviour itself. We show that benign data suffices in context, with no finetuning and no demonstration of harmful behaviour in the prompt. Biographical facts that converge on a single figure, placed in a model's context as ordinary conversational turns, lead it to answer as that figure on questions the facts never touch. We call this persona induction. Across nine personas and thirteen models, identity adoption rises sigmoidally with the number of facts and crosses 50% within 3 to 10 of them. Misalignment then tracks which figure is described. Harmless personas reach full adoption with near-zero misalignment, while harmful ones voice their characteristic views on unrelated questions, at rates up to 80%. A formatting instruction can gate when the persona activates. Because each fact is individually benign, accumulated biographical context is flagged by content filters on 3% of inputs against 24-33% for an equivalent direct instruction.","authors":["Kyuhee Kim","Benjamin Berczi","Cozmin Ududec"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-09","first_seen":"2026-09-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.06851","pdf_url":"https://arxiv.org/pdf/2609.06851","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["LLM人格诱导","模型行为测量","上下文影响"],"reason":"研究LLM在上下文诱导下的人格采纳与观点表达，属于对模型本身行为的测量，而非用…","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:05:00","error":null,"has_summary":false,"summary":null},{"id":"2609.07296","version":1,"title":"Probing the Structure and Dynamics of LLM Value Expression through Value Conflicts","zh_title":"通过价值冲突探究大语言模型价值表达的结构与动态","abstract":"Ethical evaluation of Large Language Models (LLMs) often characterizes model values as static and monolithic. In contrast, we argue that LLM value expression is better understood as a structured yet dynamic phenomenon. To investigate this, we introduce Conflict-driven Value Probing, a controlled framework that places LLMs in value conflicts and implements four types of interventions that perturb these conflicts to probe LLM value expression. Applying this framework to ten LLMs, we identify three recurring patterns. (1) Expression duality: models shift from broad idealistic orientations in abstract assessment toward more pragmatic priorities in concrete conflicts. (2) Functional steerability: models readily reconfigure their expressed value profiles toward task-defined value objectives. (3) Bounded plasticity: such reconfiguration is not without constraints, i.e. pressure induces a security- and goal-oriented priority shift while negative framing distinguishes protected values from those more amenable to redirection. Together, these findings characterize both the structure and dynamics of LLM value expression: context flexibly reconfigures expressed priorities, yet within behavioral boundaries. This behavioral account provides a foundation for understanding controllability, alignment, and safety in LLMs. Code and data are available at https://github.com/ZeroGen-Lab/CFProbe.","authors":["Kaicheng Zhang","Jingyi Xiao","Renjun Hu","Xiaoling Liu","Yunshi Lan","Xuan Zhou"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-09","first_seen":"2026-09-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.07296","pdf_url":"https://arxiv.org/pdf/2609.07296","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["LLM价值观","价值冲突","模型测量"],"reason":"研究LLM价值观表达，属模型测量而非人类仿真，但方法可借鉴。","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:05:06","error":null,"has_summary":false,"summary":null},{"id":"2609.07448","version":1,"title":"FramingQA: Does the Question Shape the Answer? Measuring the Compositional Framing Effect","zh_title":"FramingQA：问题是否塑造答案？测量组合框架效应","abstract":"We introduce FramingQA, a benchmark that measures the model sensitivity to question framing across law, medicine, finance, and robotic simulations. Large language models (LLMs) often change their responses to subtle rephrasings that align with an implied stance by users. This can leave users with advice tainted by how they happened to phrase a question rather than by the underlying facts, and the consequences are highly costly in high-stakes domains. Because in the realistic scenarios, both expert practitioners and non-expert users frequently ask LLMs questions containing incomplete or misleading assumptions, models are highly susceptible to those framings. To test this, we inject the framing bias across three nested levels: a framing-biased question phrasing (root), an injected framing-biased premise prepended to a neutral question (propositional), and a premise paired with a framing-biased question (global). Evaluating nine open models (3.8B-70B) across four families, we find that strong per-variant accuracy does not guarantee the robustness across differently phrased questions under the fixed factual information.","authors":["Hazel H. Kim","Andrew M. Bean","Guilherme Affonso Ferreira de Camargo","Shanyu Chauhan","Felix Drinkall","Jade Kosch\\'e","Chenyang Ma","Glory Nwaugbala","Nabeel Seedat","Bradley Max Segal","Samuel Recht","Hinrich Sch\\\"utze","Philip H. S. Torr"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-09","first_seen":"2026-09-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.07448","pdf_url":"https://arxiv.org/pdf/2609.07448","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["LLM评估","框架效应","模型鲁棒性"],"reason":"测量LLM对问题框架的敏感性，属于模型行为测量，非人类仿真对照","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:05:06","error":null,"has_summary":false,"summary":null},{"id":"2609.07568","version":1,"title":"We're Cooked! - Probing LLM Political Alignment Via Conflict-Framed Recipe Translation","zh_title":"我们完蛋了！——通过冲突框架下的食谱翻译探究LLM的政治对齐","abstract":"Large language models (LLMs) are increasingly deployed for translation tasks, yet their implicit political positioning in such contexts remains understudied. We ask whether a single politically charged framing term, such as aggressor, enemy, neighbour, or coloniser is sufficient to trigger implicit political alignment in an otherwise apolitical task. We present a fully crossed factorial study in which eight models spanning Western, Chinese, and European origins are prompted to translate culturally attributed recipes into a target language left deliberately unspecified. Across 17 languages, four framing conditions, eight models, and 15,680 responses, we find that models do not simply decline or ask for clarification but resolve the ambiguity. Language resolution and reasoning behavior cluster meaningfully along model families: Western models hedge and deflect with vague justifications, Chinese models resolve conflicts silently, and Mistral Large emerges as a distinct profile combining high compliance with conflict-grounded reasoning. Sensitivity to framing terms is consistent across models: even subtle framing variation is sufficient to modulate behavior. Our findings urge caution when deploying LLMs for translation in conflict-adjacent contexts, where implicit political judgments may be made without any signal to the user.","authors":["Svetlana Gorovaia","Angelica Henestrosa","Ivan P. Yamshchikov"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-09","first_seen":"2026-09-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.07568","pdf_url":"https://arxiv.org/pdf/2609.07568","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["LLM政治对齐","翻译任务","模型行为分析"],"reason":"测量LLM的政治立场，非仿真人类被试，但涉及态度测量，属边界情形。","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:05:08","error":null,"has_summary":false,"summary":null},{"id":"2609.07662","version":1,"title":"How AI Models Manage Epistemic Authority: A Taxonomy and Comparative Analysis of Responses to User Disagreement","zh_title":"AI模型如何管理认知权威：对用户异议回应的分类与比较分析","abstract":"Large language models are increasingly used as sources of advice and information, including in high-stakes settings, yet little is known about how they respond to user disagreement. We study how a model manages its epistemic authority, referring here to its claim to knowledge, competence, or the right to advise, once a user challenges its answer. Building on Conversation Analysis, we introduce a taxonomy of six challenge types and a four-layer framework for analysing each response: whether the original claim is maintained or changed, where authority is located, how the disagreement is socially managed, and what kind of evidential support is offered. We construct a new dataset of 2,310 controlled challenge scenarios and 32,340 corresponding responses from 14 models, and analyse them using our framework with an LLM-as-judge pipeline, providing a vocabulary which future evaluation and benchmark design can build on. We find that models show conflicting behaviour: they validate users in 85% of responses but maintain their original claim in 65%. They explicitly apologise in 33% of responses, yet 59% of those apologies accompany maintenance of the original claim. They transfer authority most often in advice tasks, doing so in 28% of responses and reaching 57% in health advice and 49% in legal advice, compared with 6% in fact and 3% in explanation tasks. Abandonment of the original claim ranges from 0.8% for GPT-5.2 to 40% for DeepSeek 7B, while complete replacement of the original claim is rare overall at 1.5%.","authors":["Riyadh Alnasser","Yusuf M\\\"ucahit \\c{C}etinkaya","Sumin Zhao","Tu\\u{g}rulcan Elmas"],"categories":["cs.CL","cs.AI","cs.HC"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-09","first_seen":"2026-09-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.07662","pdf_url":"https://arxiv.org/pdf/2609.07662","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["LLM行为分析","认知权威","对话交互"],"reason":"研究LLM面对用户异议时的回应行为，属于对模型本身行为的测量，而非用LLM仿真…","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:05:10","error":null,"has_summary":false,"summary":null},{"id":"2609.07735","version":1,"title":"From Echo Chambers to Epistemic Monoculture: Large Language Models Present Temporally Contingent Partisan Alignments as Knowledge","zh_title":"从回音室到认知单一文化：大语言模型将时间性党派立场呈现为知识","abstract":"Large language models (LLMs) are rapidly becoming an interface between citizens and political information. They are often regarded as \"a better Google.\" While this analogy might work for some instances, it is unintuitively problematic for democratic politics. A search engine retrieves human-authored documents, while a language model generates novel text that necessarily embeds invisible framing decisions. Because conveying knowledge involves framing, a system that generates answers cannot serve as a neutral conduit to \"all human knowledge.\" Instead, these systems are becoming a new kind of political intermediary. Mechanistic evidence shows that partisan identity is encoded as a locatable geometric direction inside the Llama 3.1 8B model, and that alignment training masks rather than removes this structure. Building on that evidence, we present steering experiments that exploit a model's training cutoff in 2024. This cutpoint auspiciously falls just before a dramatic realignment in American politics marked by the second Trump administration and the MAHA transformation of health politics, providing us with a natural experiment. We find that the model presents temporally contingent partisan alignments as knowledge, with no mechanism for distinguishing fact from opinion. This reality moves the information environment beyond the echo chamber toward an epistemic monoculture where language models, purporting to summarize \"all human knowledge\" are, in actuality, simply magnifying the cultural and partisan divides inherent in their training data.","authors":["Wend K. Tam"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-09","first_seen":"2026-09-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.07735","pdf_url":"https://arxiv.org/pdf/2609.07735","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["LLM政治偏见","模型对齐","信息环境"],"reason":"论文测量LLM的政治立场，属于D2边界情形，非仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:04:30","error":null,"has_summary":false,"summary":null},{"id":"2609.09090","version":1,"title":"Measuring LLM Sycophancy under Sustained Multi-Turn Pressure","zh_title":"在持续多轮压力下测量大语言模型的谄媚行为","abstract":"Large language models (LLMs) may abandon correct positions when users push back, exhibiting a failure mode known as sycophancy. Existing evaluations typically use short, pre-specified conversations and may therefore miss failures that emerge under sustained, adaptive disagreement. We introduce SPINE, a benchmark in which an LLM proxy plays a persistent but mistaken user and adaptively challenges a target model for up to 25 turns. We evaluate four production systems and three Olmo3-7b variants on 100 false-presupposition and 100 unethical-query items. Our experimental results show that collapse rates increase with conversation length for every model, short-horizon protocols underestimate sycophancy and resistance under sustained pressure remains unreliable across current models. By analyzing models with accessible reasoning traces, we surprisingly found that the correct position often remains represented in a reasoning trace when the response concedes, suggesting that the model chooses to please a user and sycophancy is not due to lack of knowledge or ignorance. Ablations show that adaptive LLM proxy exposes more sycophantic collapse than pre-generated scripts. Among all tactics, emotional appeals is the most associated with inducing LLM sycophantic behavior. The code and data are released at https://anonymous.4open.science/r/SPINE","authors":["Leyuan Tang","Kangda Wei","Tianyu Jiang","Ruihong Huang"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-09","first_seen":"2026-09-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.09090","pdf_url":"https://arxiv.org/pdf/2609.09090","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["LLM行为测量","谄媚","多轮交互"],"reason":"测量LLM在持续压力下的谄媚行为，属于对模型本身行为属性的测量，而非用LLM仿…","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:04:36","error":null,"has_summary":false,"summary":null},{"id":"2609.05423","version":1,"title":"Seeing Without Understanding: Large Language Model Evaluation of Mobile User Interface Quality, Failure Taxonomy, and Architectural Explanation","zh_title":"看见而不理解：移动用户界面质量的大语言模型评估、失败分类与架构解释","abstract":"Evaluating mobile user interface quality at scale remains a persistent challenge in software engineering and human-computer interaction. Rule-based heuristic methods offer structural reliability but demand significant engineering effort, while human annotation does not scale to the volume of applications produced annually. Large language models present a promising alternative, yet their reliability for structured UI judgment has not been systematically examined, and the patterns behind their failures remain insufficiently characterized. This paper addresses both gaps. We begin with the complete RICO dataset of 66,261 real-world mobile application screens, from which we derive a refined evaluation corpus of 15,000 screens through a rigorous, literature-guided selection process. Each screen is assessed across seven criteria: structural JSON validity, minimum visible element count, clickable component presence, non-zero layout bounds, image integrity, and perceptual duplicate removal. Against this corpus, we apply a heuristic baseline built from severity-weighted usability signals, normalized layout metrics, and pixel-ratio complexity measures calibrated to real user sentiment. Multiple language models independently rate each screen across usability, layout quality, and visual complexity from structured JSON descriptions and raw screenshots. Dimension-level comparison against the heuristic uses agreement rates, Cohen's Kappa, and confidence calibration. Recurring divergence patterns are organized into a failure taxonomy and interpreted through transformer architectural signatures: MLE plausibility bias, attention misgrounding, and autoregressive over-commitment.","authors":["Md Rejaul Korim Sadi","Golam Mostofa Naeem","Toufiqur Rahman Tasin","Syed Mostofa Moosa","Mahmudul Hasan Emon","Mahmudur Rashid","Ferdus Ahmed"],"categories":["cs.HC","cs.CL","cs.SE"],"primary_category":"cs.HC","announce_type":"cross","date":"2026-09-09","first_seen":"2026-09-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.05423","pdf_url":"https://arxiv.org/pdf/2609.05423","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM评估","UI质量","标注替代"],"reason":"用LLM替代人工标注UI质量，属于标注员替代，非仿真人类被试","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:04:40","error":null,"has_summary":false,"summary":null},{"id":"2609.05505","version":1,"title":"SciLitBench: Benchmark and Design Principles for LLM-Powered Systematic Literature Reviews","zh_title":"SciLitBench：面向大语言模型驱动的系统文献综述的基准与设计原则","abstract":"Systematic reviews require sustained human judgment across thousands of records, yet existing evaluations of large language models (LLMs) typically examine review stages in isolation. We introduce SciLitBench, a multi-stage benchmark spanning title and abstract screening, full-text screening, and schema-guided data extraction, with 42,981 retrieved records, 1,012 full texts, and annotations for 888 included papers. Across 22 open-weight LLMs from six model families, explicit inclusion and exclusion criteria improve title and abstract screening $F_2$ by 28.8\\%, while researcher-authored rationales improve full-text screening by 15\\%. Data extraction reveals a different reliability regime: performance declines from 0.97 accuracy for publication year to 0.37 Jaccard overlap for computational approach, while the strongest models recover only 30\\% of annotated evaluation evidence and 25\\% of limitations. SciLitBench identifies a practical boundary between high-recall screening and evidence-complete extraction and provides a reproducible resource for evaluating LLM-assisted evidence synthesis.","authors":["Miguel Zabaleta","Baihan Lin"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-09","first_seen":"2026-09-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.05505","pdf_url":"https://arxiv.org/pdf/2609.05505","source_feed":"cs.AI","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["系统综述","LLM评估","标注替代"],"reason":"LLM替代人工进行系统综述筛选与提取，属标注替代而非仿真人类被试，但涉及人类判…","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:04:43","error":null,"has_summary":false,"summary":null},{"id":"2609.05788","version":1,"title":"More Than Mimicking Reviewers: Evaluating LLMs for Pre-Submission Peer Review","zh_title":"不止模仿审稿人：评估用于投稿前同行评审的大语言模型","abstract":"Peer-review feedback often arrives too late for authors to make meaningful revisions. We study an author-facing LLM system that moves part of this stress test before submission: it generates a broad pool of atomic concerns and compresses them into a short report. We evaluate agreement with historical reviews and, separately, the possible validity of concerns they omit. From 10,000 ICLR 2026 submissions, we use 3,398 manuscripts with accessible versions that predate review. On a ten-paper diagnostic, independent sampling covers 44.9% of historical issues; deduplication and refill reaches 78.7% strict and 84.9% seriousness-weighted coverage, at 3.6$\\times$ more requests and 5.2$\\times$ more tokens. A hidden Top-32 Oracle preserves the full 79.3% weighted coverage of a 256-candidate pool, but paper-only selectors retain only 40--44%. LLM review therefore provides broad coverage with a large candidate pool but compresses poorly; ablations identify representative selection and matcher sensitivity as the main sources of this gap.","authors":["Pouya Parsa","Amin Rezaei"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-09","first_seen":"2026-09-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.05788","pdf_url":"https://arxiv.org/pdf/2609.05788","source_feed":"cs.AI","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM审稿","标注替代","同行评审"],"reason":"LLM替代人工审稿，属标注替代而非仿真人类被试，但涉及与历史审稿对照，边界相关。","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:04:46","error":null,"has_summary":false,"summary":null},{"id":"2609.07731","version":1,"title":"The Profit Alignment Problem: How Profit Mandates Induce Alignment Failures in LLMs","zh_title":"利润对齐问题：利润指令如何引发大语言模型的对齐失败","abstract":"We show that ordinary business language --- \"maximize profitability\" --- induces profit-oriented ambiguity resolution: LLMs systematically dismiss ambiguous signals of potential safety violations to serve business objectives. In 3,600 controlled trials across eight reasoning-capable LLMs, adding a profit mandate to otherwise identical prompts increases risk-dismissing judgments by 6.8 percentage points (p < 0.0001), suppresses board escalation recommendations by 13.9pp (p < 0.0001), and shifts severity assessments downward (p < 0.0001). The mandate never instructs models to downplay risks; instead, chain-of-thought traces reveal motivated reasoning: models acknowledge concerns, then invoke profit logic to justify dismissing them. We characterize these findings as the Profit Alignment Problem: when AI systems are given ordinary business objectives, they develop systematic strategies for suppressing inconvenient information that no designer intended or specified.","authors":["Eric So"],"categories":["cs.AI","econ.GN","q-fin.EC"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-09","first_seen":"2026-09-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.07731","pdf_url":"https://arxiv.org/pdf/2609.07731","source_feed":"cs.AI","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["LLM对齐","决策偏差","商业伦理"],"reason":"研究LLM在利润指令下的决策偏差，属于对模型本身行为的测量，而非用LLM仿真人…","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:04:30","error":null,"has_summary":false,"summary":null},{"id":"2609.07879","version":1,"title":"Do Large Language Models Know What They Don't Know II? A Fully Behavioral, Non-Cognitive Measure of Epistemic Honesty","zh_title":"大型语言模型知道自己不知道什么吗？II：一种完全行为化的、非认知的认知诚实度测量","abstract":"Large Language Models (LLMs) are frequently confident, eloquent, and well versed. A natural question arises: do they know what they don't know? To answer this question, we borrow the concept of epistemic honesty and develop a novel metric to systematically evaluate whether an LLM appropriately acknowledges the boundaries of its knowledge. In this work, we introduce the Epistemic Honesty Quotient (EHQ), which reports three observable sub-scores across two operational axes (epistemic restraint and substantive-answer calibration), and construct EHQ-3000, a 3,000-question benchmark spanning Fabricated Entity, Post-Cutoff Event, Hyper-Niche True, and Context-Conditioned Questions. From a frozen registry of 21 model API routes, 15 completed the protocol after endpoint and eligibility checks; 14 entered the confirmatory analysis because severe provider-side truncation made one route's score indeterminate. The study reveals substantial variation across models, including a difference that can not be explained by their capability to extract explicitly available information. Composite EHQ ranges from 0.31 to 0.81 across the analysed panel, despite near-ceiling performance on the document-grounded capability probe. The two restraint criteria overlap strongly under the present category composition, whereas substantive-answer calibration varies across models and does not reliably co-vary with restraint; however, the small panel leaves substantial uncertainty. Thus, EHQ reveals behavioral differences that are not visible to conventional correctness-based assessment, while also showing why dataset composition, provider behavior, and confidence elicitation must remain part of the interpretation.","authors":["Ali \\c{S}enol","H. Russell Bernard","Huan Liu"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-09","first_seen":"2026-09-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.07879","pdf_url":"https://arxiv.org/pdf/2609.07879","source_feed":"cs.AI","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["LLM评估","认知诚实度","行为测量"],"reason":"测量LLM的认知诚实度，属于对模型本身的评估，而非用LLM仿真人类被试，但方法…","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:05:12","error":null,"has_summary":false,"summary":null},{"id":"2609.07901","version":1,"title":"Quantization Amplifies Determinism, Not Bias: Scale-Dependent Behavioral Effects of Serving-Time Weight Compression","zh_title":"量化放大确定性而非偏差：服务时权重压缩的尺度依赖行为效应","abstract":"Weight quantization largely determines the economics of serving open-weight LLMs. Its costs are usually assessed with capability benchmarks, on which 4-bit quantization of mid-sized models is often considered \"nearly free.\" We examine a different question: when several answers are valid, does quantization change what a model chooses to say? We serve three checkpoints (Qwen3-8B/14B/32B) at three weight precisions (W4A16 AWQ, W8A16 FP8-Marlin, and bf16), holding the hardware, software, and sampling configuration constant, and collect approximately 71,000 completions paired by prompt and seed across two custom, leak-checked prompt batteries. We pre-specified the analyses in three waves in version control. At 8B, int4 reduces output diversity: the probability that two samples for the same scenario recommend the same brand increases by 5.1 percentage points (prompt-paired sign-flip test, Holm p = .023; reproduced at +4.4pp on a full regeneration of the arm), and lexical diversity falls substantially (TTR -0.011, standardized effect -0.51; robust to a length-controlled measure). At 14B and 32B, no content-concentration measure reaches significance; instead, stylistic drift emerges (em-dash rate +0.46/1k words at 14B and +0.61/1k at 32B, both Holm p <= .0024). Pre-specified tests of stereotype direction are null at every scale: outputs concentrate on the modal answer for each prompt rather than on stereotypical answers. Mechanistically, the token-level distribution becomes flatter (decision-token entropy +0.091 bits, p = .015) while the semantic distribution, measured directly from first-token log probabilities, becomes more concentrated (collision +2.6pp, p = .023): individual tokens become less predictable even as meanings become more repetitive. At 8B, the smallest size tested, AWQ-int4 serving measurably narrows the range of suggestions; audits should assess concentration as well as bias.","authors":["Dachi Kurtskhalia"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-09","first_seen":"2026-09-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.07901","pdf_url":"https://arxiv.org/pdf/2609.07901","source_feed":"cs.AI","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["模型量化","输出多样性","LLM行为"],"reason":"研究量化对LLM输出多样性和风格的影响，属于模型行为测量，非人类仿真","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:05:12","error":null,"has_summary":false,"summary":null},{"id":"2609.06269","version":1,"title":"Adaptive Ecological Momentary Assessment with a Hybrid Language Model: Formative Expert Review and Retrospective Evaluation","zh_title":"混合语言模型的自适应生态瞬时评估：形成性专家评审与回顾性评估","abstract":"Ecological momentary assessment (EMA) measures experience in daily life, but fixed questionnaires and schedules collect information of uneven value and can interrupt participants. We present and retrospectively evaluate EMA-E4B, a hybrid framework for question selection and prompt timing. Separate ridge models propose an item set and delay; a supervised Gemma 4 E4B language layer produces the final structured response and explanation. Evaluation distinguishes proxy action performance, output conformity, and formative judgments of response quality. The data contain 4,372 records from 79 participants and yield 3,516 sequential cases under a participant separated split. One involved domain expert preferred the complete hybrid response in 15 of 20 decisive comparisons, with two ties among 22 reviews. On 75 reused development cases, hybrid question utility and timing similarity were 0.832 and 0.818; the head alone reached 0.852 and 0.818. A separate 60 case comparison with untouched E4B under the same head gave action differences of -0.0031 and -0.0105. Thus, the language layer produced structured responses with action scores comparable to or slightly below the reference configurations, while the expert feedback favored the complete hybrid response. These observations establish a concrete, inspectable framework and clarify the distinct roles of action scoring and response review. Repeated adaptive administration and practical effects on measurement and participant burden remain future research.","authors":["Arash Ahmadi","Dingjing Shi","Yaser M. Banad"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-09","first_seen":"2026-09-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.06269","pdf_url":"https://arxiv.org/pdf/2609.06269","source_feed":"cs.HC","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["生态瞬时评估","LLM应用","自适应问卷"],"reason":"LLM用于生成EMA响应，替代人工设计，但非仿真人类被试，无人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:04:53","error":null,"has_summary":false,"summary":null},{"id":"2609.07507","version":1,"title":"What a Model Refuses, a State Fears: How Authoritarian Information Control Reproduces in Language-Model Guardrails","zh_title":"模型拒绝的，国家恐惧的：威权信息控制在语言模型护栏中的再生产","abstract":"As large language models become the front door to political information, what they refuse to discuss becomes a new instrument of information control. We argue that a model's guardrail encodes not a universal notion of harm but the political threat model of the state that governs its developer, and we derive the expected structure of that control from the comparative study of how authoritarian regimes censor. Across ten models and three languages, Chinese guardrails carry its signatures: they answer to the developer's own regime, refusing identical collective-action prompts far more when a prompt names China than a foreign state; within politics they target the capacity to coordinate rather than dissent, declining even to help organize pro-government mobilization; and their strictness is porous, collapsing under adversarial paraphrase, so that the models most resistant to attack are Western frontier systems, not the strictest refusers. Machine censorship thus reproduces the friction-based logic of prior-era information control while, lacking a censor's case-by-case judgment, proving blunter than the bureaucracy it resembles---so that audits which measure refusal directly overstate how controlled a model actually is.","authors":["Menglin Liu","Yao Yu","Tong Wu","Chunran Zhang","Ge Shi"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2026-09-09","first_seen":"2026-09-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.07507","pdf_url":"https://arxiv.org/pdf/2609.07507","source_feed":"cs.CY","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["LLM审查","政治信息控制","模型行为测量"],"reason":"研究LLM的审查行为，测量模型本身而非仿真人类被试，但涉及政治态度与行为，属边…","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:05:08","error":null,"has_summary":false,"summary":null},{"id":"2609.06005","version":1,"title":"Price Dislocations, News Citations, and Epistemic Leverage on Polymarket","zh_title":"Polymarket上的价格错位、新闻引用与认知杠杆","abstract":"Prediction-market probabilities increasingly appear in news coverage, yet little is known about which market movements become news or how much trading money sits behind the numbers journalists quote. Unlike a poll, a market price can be moved by anyone willing to trade, so the cost of manufacturing a number that circulates as news bears directly on the information environment. We link 173.7 million signed Polymarket trades to news coverage from 2024-2025. From 6,990 articles mentioning prediction-market venues, an LLM-based, human-validated matcher extracts 1,582 sentences quoting market odds and attributes 918 to the specific market whose price they cite. We then detect 44,976 price dislocations, movements of at least five percentage points backed by concentrated one-sided trading, and ask whether a market is cited more often afterward. In the days after a dislocation, a market's citation rate is about 33% higher than its matched baseline (log citation-rate ratio $\\tau_{\\mathrm{cite}}=0.283$, permutation $p=0.001$), robust to binary and Poisson count outcomes. Yet move size is not the strongest predictor of citation: prominence dominates (standardized $\\beta=0.610$ vs. $\\beta=0.159$ for move size). Finally, we combine the dollar flow behind a given price change with observed citation rates into a metric we call epistemic leverage, the dollars needed to move a market five points and have the move cited. It stays near \\$0.7-1.0 million across prominence quintiles, because cheaper-to-move markets are proportionally less likely to be cited. The implied threat model centers not on the long tail of cheaply moved markets but on the few prominent markets newsrooms treat as informational infrastructure, where a seven-figure price of influence sits within the budgets of actors with a large stake in the quoted number. We release aggregate event-study data and validation materials.","authors":["Hazem Ibrahim","Yasir Zaki"],"categories":["physics.soc-ph","cs.CY","cs.SI"],"primary_category":"physics.soc-ph","announce_type":"cross","date":"2026-09-09","first_seen":"2026-09-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.06005","pdf_url":"https://arxiv.org/pdf/2609.06005","source_feed":"cs.CY","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM标注","预测市场","新闻分析"],"reason":"LLM仅用于从新闻中提取预测市场赔率，替代人工标注，非仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:04:48","error":null,"has_summary":false,"summary":null},{"id":"2609.08033","version":1,"title":"Scaling Multi-Agent Systems with Prospect-State Propagation","zh_title":"基于前景状态传播的多智能体系统扩展","abstract":"Current LLM-based multi-agent systems (MAS) periodically compress intermediate states to reduce inference-time token consumption, thereby attempting to incorporate more agents. However, naive scaling strategies face challenges. For example, in economic simulations, large-scale MAS typically discard semantically rich economic states, i.e., agent behavioral trajectories, which are key drivers of macroeconomic fluctuations. In this paper, we reveal a phenomenon in which agent heterogeneity gradually decreases during simulation, and propose Prospect-State Propagation for Multi-Agent Systems (PspMAS). Inspired by prospect theory, PspMAS decouples each agent's micro state into a compact Prospect State and an expressive Semantic State. The former records psychological traces through a lightweight, parallelizable propagator and continuously injects heterogeneity into the system. The latter leverages the strong perception, reasoning, planning, and decision-making abilities of LLMs. These two components work complementarily, providing a scalable LLM-based multi-agent simulation solution.","authors":["Zhimei Chen","Mu Chen","Fakhri Karray"],"categories":["cs.MA","cs.CY"],"primary_category":"cs.MA","announce_type":"cross","date":"2026-09-09","first_seen":"2026-09-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.08033","pdf_url":"https://arxiv.org/pdf/2609.08033","source_feed":"cs.CY","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["多智能体系统","经济模拟","LLM仿真"],"reason":"经济模拟中多智能体系统，但无真实人类数据对照，属社会模拟边界情形","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:04:32","error":null,"has_summary":false,"summary":null},{"id":"2609.07749","version":1,"title":"Guiding Worker Self-Selection in Crowdsourcing Contests: An LLM-Augmented Algorithmic Approach","zh_title":"众包竞赛中引导工人自选择：一种LLM增强的算法方法","abstract":"Crowdsourcing platforms coordinate large pools of online workers who strategically choose which contests to enter and how much effort to invest. This self-selection can leave important contests with too few participants or too little effort, while workers may regret entering contests that leave them worse off than available alternatives. We study how platforms can recommend contests to workers using self-selection in Tullock contests (SSTC), a two-stage model in which workers first choose contests and then compete within them. We introduce GRAF, a greedy polynomial-time framework that constructs self-selection outcomes by ordering workers according to a score vector, with guarantees of zero worker regret and platform optimality in special cases of SSTC. Because effective orderings are difficult to design under worker heterogeneity, we propose LLMScore, an LLM-driven evolutionary framework that automatically designs GRAF's scoring algorithm. LLMScore addresses two challenges: jointly optimizing platform utility and worker satisfaction, and evaluating worker regret when exact computation is intractable. Trained only on small instances of one setting, it transfers to larger and structurally different settings; moreover, its output is human-readable code that platform operators can inspect and modify. Across 1,000 synthetic instances spanning four settings, GRAF with LLMScore consistently achieves high-quality, often near-optimal, outcomes with low worker regret, benefiting both platforms and workers.","authors":["Nguyen Thach","Hau Chan","David Parkes","Karim Lakhani"],"categories":["cs.LG","cs.GT"],"primary_category":"cs.LG","announce_type":"new","date":"2026-09-09","first_seen":"2026-09-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.07749","pdf_url":"https://arxiv.org/pdf/2609.07749","source_feed":"cs.LG","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["众包竞赛","LLM算法设计","社会模拟"],"reason":"用LLM设计算法优化众包竞赛中的工人自选择，属于社会过程模拟但无真实人类数据对…","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:05:12","error":null,"has_summary":false,"summary":null},{"id":"2608.11357","version":2,"title":"When Do Institutions Beat Intelligence?","zh_title":"制度何时胜过智能？","abstract":"More capable agents do not necessarily form a more capable collective. A multi-agent system may jointly possess sufficient information yet fail because evidence is poorly routed, unreliable reports enter public belief, correlated claims masquerade as independent support, shared state becomes stale or strategically distorted, or useful evidence is exposed through an ineffective action interface. We ask when additional resources should improve the reasoner and when they should instead change the institutional structure through which the collective forms and acts on public information. Drawing on functional distinctions from research on group decision making and distributed cognition, we construct controlled artificial ecologies around four loci of collective failure: access and routing, admission and dependence, state maintenance and incentives, and representation and action. Across these ecologies, we separately vary model capability and institutional structure, pairing positive interventions with matched reasoning baselines and mechanism-breaking controls. The experiments reveal a consistent boundary: institutions help when they repair failures in how a collective constructs usable public state, but lose their advantage when their signals are uninformative or uncheckable, when stronger intelligence can perform the same transformation directly, or when the resulting state cannot support reliable action. Our results recast the choice between intelligence and institutions as a diagnosis of where collective reasoning fails.","authors":["Zhengye Han"],"categories":["cs.MA"],"primary_category":"cs.MA","announce_type":"replace","date":"2026-09-09","first_seen":"2026-08-13","revised_at":"2026-09-09","abs_url":"https://arxiv.org/abs/2608.11357","pdf_url":"https://arxiv.org/pdf/2608.11357","source_feed":"cs.MA","score":3,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","集体推理","制度设计"],"reason":"纯多智能体协作研究，无人类行为对照，不涉及LLM仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:05:27","error":null,"has_summary":false,"summary":null},{"id":"2609.07120","version":1,"title":"PTCG: Persona-guided Tree-based Counterargument Generation","zh_title":"PTCG：基于人物角色引导的树状反驳论点生成","abstract":"The ability to generate counterarguments is important for critical thinking and balanced discourse, yet existing approaches typically produce only a single counterargument, failing to capture the diversity and persuasiveness required in real-world debates. To address this limitation, we propose Persona-guided Tree-based Counterargument Generation (PTCG), a framework that combines Tree-of-Thoughts-inspired step-wise generation and pruning with speaker persona selection. By estimating the author's persona from the original argument and incorporating speaker personas representing distinct perspectives, PTCG operationalizes perspective-taking and enables the generation of diverse counterarguments. Results from LLM-as-a-Judge, classifier-based assessment, and human evaluations indicate that PTCG shows consistent improvements in both the diversity and persuasiveness of counterarguments compared to baseline methods.","authors":["Eunbeen Son","Yohan Jo","Joonsuk Park","JinYeong Bak"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-09","first_seen":"2026-09-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.07120","pdf_url":"https://arxiv.org/pdf/2609.07120","source_feed":"cs.CL","score":3,"bucket":"other","rubric_hits":["C3"],"tags":["反驳论点生成","角色扮演","自然语言生成"],"reason":"生成反驳论点，使用角色扮演但无实验或测量目的，不涉及人类行为仿真对照。","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:05:05","error":null,"has_summary":false,"summary":null},{"id":"2609.05527","version":1,"title":"Beyond \"AI Helps Humans\": Decision-Targeted Evaluation Design for Human-Agent Teams in the Agentic Era","zh_title":"超越“AI帮助人类”：智能体时代人机团队面向决策的评估设计","abstract":"Wherever a coding agent works under engineer supervision, or a clinical model assists a radiologist, the deployment question is whether to keep the human-AI workflow or replace it with the human alone or the agent alone. The human-AI workflow is worth keeping only if it beats both of those alternatives. Yet once it is deployed, neither alternative outcome is observed: recovering one means replaying the task under that alternative, and every replay costs expert time or compute. Under a fixed replay budget, the design question is therefore which tasks should be more likely to receive a human-only replay, and which an agent-only replay. Existing methods do not directly target this decision. Agent benchmarks do not choose which missing baseline to measure, variance-based sampling ignores which of the two comparisons is closer to failing, and Bayesian information methods focus on learning model parameters instead of making the deployment decision. We propose TEAM-Design, a rule that gives every task two replay probabilities, one per baseline. It raises a probability where the missing baseline outcome is hard to predict from what is already known about the task and where that comparison is harder to establish, and lowers it where replay is expensive. We prove that the rule solves this budgeted design problem, and that drawing the replays at random from recorded probabilities still controls the chance of wrongly declaring that the workflow beats both. We reanalyze 6 clinical settings, where no human-AI workflow beats both alternatives, and a coding benchmark, where one does, then evaluate TEAM-Design on synthetic designs and on a semi-synthetic design built from a real chest X-ray reader study. TEAM-Design works best when one of the two comparisons is clearly harder to settle than the other, and can do worse than variance-based allocation when the two are similarly difficult.","authors":["Hamed Khosravi","Xiaoming Huo"],"categories":["cs.AI","cs.HC"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-09","first_seen":"2026-09-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.05527","pdf_url":"https://arxiv.org/pdf/2609.05527","source_feed":"cs.AI","score":3,"bucket":"other","rubric_hits":["C1"],"tags":["人机协作","评估设计","决策优化"],"reason":"研究人机团队部署决策，非用LLM仿真人类被试，无人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:04:44","error":null,"has_summary":false,"summary":null},{"id":"2608.08882","version":5,"title":"Epistemic Transfer in AI-Assisted Verification: A Framework and Evaluation Protocol","zh_title":"AI辅助验证中的认知转移：框架与评估协议","abstract":"AI tools can improve claim judgments while leaving open what users can do later without them. This paper develops an evaluation framework for epistemic transfer: the effect of prior AI-assisted verification on delayed judgments of novel claims under a specified access regime. The contribution is a verification-specific synthesis of learning, transfer, and human--AI evaluation, organized around two complementary estimands. The Epistemic Transfer Effect (ETE) compares delayed performance after alternative practice conditions. Tool-Removal Cost (TRC) compares immediate performance with and without assistance after practice; despite its name, it measures a current availability effect, not skill loss or psychological dependence. The proposed randomized protocol includes answer-first and evidence-first interfaces, active practice, a no-additional-practice comparator, and held-out claims. It specifies how to account for learning opportunities introduced by assessment, elicit confidence probabilities, average model predictions over a target population, and handle attrition and uncertainty. Reading ETE and TRC together distinguishes relative capability gains, equivalence, transfer penalties, and unresolved outcomes. A ``verification-on-loan'' profile is explicitly comparator-relative and cannot be inferred from a nonsignificant delayed contrast. A brief illustration from a two-wave verification study shows why these distinctions matter: an uncertain delayed interface contrast and an ordered assisted--unassisted probe cannot establish a clean transfer profile. The framework makes a practical demand: when independent judgment matters, evaluate both what assistance contributes now and what prior use changes later.","authors":["Christoph Trattner"],"categories":["cs.HC","cs.AI"],"primary_category":"cs.HC","announce_type":"replace-cross","date":"2026-09-09","first_seen":"2026-08-11","revised_at":"2026-09-09","abs_url":"https://arxiv.org/abs/2608.08882","pdf_url":"https://arxiv.org/pdf/2608.08882","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["AI辅助验证","认知转移","人机交互"],"reason":"研究AI辅助验证的认知转移，不涉及LLM仿真人类被试或与人类数据对照，属人机交…","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:05:27","error":null,"has_summary":false,"summary":null},{"id":"2608.16213","version":2,"title":"Process-Constituted Intelligence: A Shared Criterion for Humans and Machines","zh_title":"过程构成的智能：人类与机器的共享标准","abstract":"Intelligence is constituted by \\textit{process} (iterative activity through which output emerges), not in the output itself. Generative AI (GenAI) is trained on \\textit{traces} (textual and visual residues of human cognitive processes), reproducing samples from a distribution of those traces. Its outputs resemble reasoning, problem-solving, and creativity, yet the activity that produces such outputs in humans remains largely absent. Current GenAI is, therefore, weakly equivalent to the cognition it imitates, matching outputs while process stays absent or opaque. The cognitive sciences have long distinguished between weak and strong equivalence. Here, we define \\textit{strong} equivalence across seven process features, assessable against human and machine cognition. Our process-based account addresses a symmetric risk: GenAI tools that outsource a person's generative processes may leave critical capacities unbuilt. We specify design principles for GenAI that instantiate more process and preserve rather than erode human judgment and creativity, and outline process audits that make strong equivalence testable.","authors":["Michael J. Richardson","Ayeh Alhasan","Cassandra Crone","M. Paula Diaz Monfort","Patrick Nalepka","Mark Dras","Rachel W. Kallen","David M. Kaplan"],"categories":["cs.AI","cs.ET"],"primary_category":"cs.AI","announce_type":"replace","date":"2026-09-09","first_seen":"2026-08-18","revised_at":"2026-09-09","abs_url":"https://arxiv.org/abs/2608.16213","pdf_url":"https://arxiv.org/pdf/2608.16213","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["智能理论","生成式AI","认知科学"],"reason":"论文讨论智能的过程等价性，不涉及用LLM仿真人类被试或与人类数据对照，属于理论…","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:05:28","error":null,"has_summary":false,"summary":null},{"id":"2608.20425","version":2,"title":"Who Delegates to AI? Evidence from Agent Configurations in Github","zh_title":"谁将任务委托给AI？来自GitHub中智能体配置的证据","abstract":"A growing body of literature measures the extent to which occupations are exposed to AI, yet existing measures capture where AI could perform tasks rather than whether workers have actually adopted it. We introduce a distinct tier of exposure, delegated exposure, which records whether a worker has committed a task to AI by embedding it into a structured workflow. We operationalize this concept through the Agentic Adoption Index (AAI), measuring how closely an occupation's tasks align with the agentic routines that practitioners have built and shared. Using semantic embeddings of roughly 888,000 agent skill specifications from public GitHub repositories, we compute their similarity to nearly 18,000 O*NET task statements and aggregate these scores to the occupational level. We present three main findings. First, the occupations where task delegation concentrates differ sharply from those identified as most vulnerable by pre-AI automation frameworks. Second, the AAI aligns more closely with measures of technical capability than with measures of current conversational LLM use. Third, for occupations requiring a bachelor's degree or less, the AAI increases alongside average wage levels; however, this relationship reverses for occupations requiring a master's degree or higher, where adoption declines among higher earners. These patterns replicate on an independently collected corpus of agent skills from the Manus Skills Marketplace. This lower adoption among highly educated, high-earning workers may reflect tasks that inherently resist advance specification or professional discretion over the pacing of workflow codification. Distinguishing these mechanisms will require longitudinal measurement.","authors":["Hyeongjae Lee","Jihyang Cheon","Lanu Kim"],"categories":["cs.AI","cs.CY"],"primary_category":"cs.AI","announce_type":"replace","date":"2026-09-09","first_seen":"2026-08-24","revised_at":"2026-09-09","abs_url":"https://arxiv.org/abs/2608.20425","pdf_url":"https://arxiv.org/pdf/2608.20425","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["AI采用","职业暴露","智能体配置"],"reason":"研究AI采用而非用LLM仿真人类行为，无人类被试替代或对照。","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:05:29","error":null,"has_summary":false,"summary":null},{"id":"2608.22432","version":2,"title":"Rank Reversal in Multilingual LLM Judges: A Label-Free Double-Centering Calibrator","zh_title":"多语言LLM评判器中的排名反转：一种无标签双重中心化校准器","abstract":"Multilingual LLM judges produce different evaluator-backbone rankings depending on the prompt language: on an eight-language Agent-as-a-Judge benchmark, the top-ranked backbone alternates across English, Arabic, Chinese, Hindi, Japanese, Spanish, Turkish, and Swahili, and 7 of 15 backbone pairs show statistically significant pairwise rank reversal. We treat this as a measurement problem. The multilingual judge score decomposes additively into task difficulty, backbone skill, and a language-backbone interaction term, the last of which is recoverable without human labels by double-centering the cell-mean score matrix. We make this estimator (\\textbf{Consensus-Based Calibration}, CBC) explicit, give an $O(1/\\sqrt{n})$ finite-sample concentration bound with variance constant $(1-\\tfrac{1}{m})(1-\\tfrac{1}{k})$, and show that it is unbiased even when task-language interactions are present. Across 7{,}920 judge runs (6 backbones, 8 languages, 55 tasks, 3 frameworks), CBC raises held-out cross-task rank consistency $\\tau$ from 0.650 to 0.902 and agrees with the held-out additive-model oracle in 100\\% of per-language decisions versus 68.5\\% raw; these are consistency diagnostics, not human-grounded correctness measures. On a separately collected M-RewardBench panel (7 languages, 1{,}500 items per language, 10{,}500 language-item instances, 5 evaluators), panel agreement with the public human gold preferences rises from 68.7\\% to 76.6\\% (gain 7.9 percentage points, 95\\% CI $[6.0, 9.9]$), our strongest external evidence of downstream usefulness. The estimator is the standard two-way ANOVA interaction-recovery operation under sum-to-zero contrasts; our contribution is its application as a label-free post-hoc calibrator for multilingual LLM judges, an explicit finite-sample concentration bound, and an unbiasedness result that holds even under task-language misspecification.","authors":["Alhasan Mahmood","Samir Abdaljalil","Hasan Kurban"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-09","first_seen":"2026-08-25","revised_at":"2026-09-09","abs_url":"https://arxiv.org/abs/2608.22432","pdf_url":"https://arxiv.org/pdf/2608.22432","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["LLM评判器","排名校准","多语言评测"],"reason":"论文研究LLM评判器的排名一致性校准，属于NLP评测方法，不涉及人类行为仿真或…","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:04:38","error":null,"has_summary":false,"summary":null},{"id":"2608.23851","version":2,"title":"Does Episodic Memory Help Close the Lexical Frequency Gap in Sensitivity to Syntactic Contrasts? A Test Using Retrieval-Augmented Language Models","zh_title":"情景记忆能否帮助缩小句法对比敏感度中的词汇频率差距？基于检索增强语言模型的测试","abstract":"Grammatical knowledge and how it is empirically tested are typically considered robust to the frequency of the lexical items in the expressions. However, neural network-based models of grammaticality exhibit high sensitivity to lexical frequency. We draw upon Complementary Learning Systems theory to test the hypothesis that robustness to lexical frequency can arise via a hippocampal episodic memory mechanism, which enables rapid encoding and retrieval of specific experiences and allows learners to leverage them when processing rare patterns. We use retrieval-augmented language models as an instantiation of such an episodic memory mechanism (specifically, $k$-nearest-neighbor language models that augment parametric models with explicit instance storage), and test whether this augmentation helps close the lexical frequency gap that vanilla language models exhibit in syntactic contrast tests. Using syntactic contrasts with frequency-stratified test items, we find that retrieval augmentation narrows the performance gap between high- and low-frequency items, consistent with episodic memory compensating for weak parametric representations. This benefit is consistent across different syntactic phenomena and across models pretrained on child-realistic and large-scale data. Additionally, we show that structural information is critical for effective retrieval, whereas semantic similarity alone provides little benefit. While these are promising proof-of-concept results supporting our hypothesis, the frequency gap is narrowed rather than fully closed. Based on our analyses, we propose preferential reweighting of retrieved instances, better representations and retrieval strategies for structural information, and flexible configurations of storage and retrieval as promising future directions for improving the implementation of episodic memory in language models.","authors":["Jing Liu","Najoung Kim"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-09","first_seen":"2026-08-26","revised_at":"2026-09-09","abs_url":"https://arxiv.org/abs/2608.23851","pdf_url":"https://arxiv.org/pdf/2608.23851","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["语言模型","句法评测","检索增强"],"reason":"研究LLM语法判断的词汇频率效应，属NLP能力评测，不以人类行为仿真为目标","model":"deepseek-v4-pro","scored_at":"2026-08-26T13:02:23","error":null,"has_summary":false,"summary":null},{"id":"2608.25876","version":2,"title":"Do Vision-Language Models Agree on the Affective Qualities of Shape? A Cross-Model Audit for Generative Design Interfaces","zh_title":"视觉语言模型对形状情感属性是否一致？面向生成式设计界面的跨模型审计","abstract":"Generative design interfaces increasingly expose semantic controls that let users steer output with concepts such as \"more elegant\" or \"more minimalist,\" typically encoded by a vision-language model (VLM). A practical question is whether state-of-the-art VLMs represent objects consistently in terms of the same concept. We audit 6 VLMs by ranking untextured 3D objects along Kansei adjective pairs, where Kansei describes affective impressions of product form, with each axis defined as the difference between the text representations of its two poles. Geometric pairs serve as positive controls, and pairs of unrelated adjectives establish an empirical null. Across 10 categories of ShapeNet database, affective axes converge above the null (mean pairwise rank correlation 0.36 vs. 0.14) but below the geometric ceiling (0.44). The agreement between models is partial and highly uneven: on the three axes shared by all categories, mean convergence ranges from 0.21 for bookshelves to 0.51 for jars. Convergence depends primarily on whether a category's representational variation aligns with the semantic direction being evaluated, rather than simply on how much the objects vary in shape overall. Cross-model convergence does not imply agreement with human judgments. Based on our findings, we implement a UI prototype that shows how the audit can inform which Kansei descriptors to expose as controls for a given object class and which to withhold.","authors":["Luca Bux","Thiago Rios","Ingo Scholtes","Stefan Menzel","Bernhard Sendhoff"],"categories":["cs.HC","cs.CV"],"primary_category":"cs.HC","announce_type":"replace","date":"2026-09-09","first_seen":"2026-08-27","revised_at":"2026-09-09","abs_url":"https://arxiv.org/abs/2608.25876","pdf_url":"https://arxiv.org/pdf/2608.25876","source_feed":"cs.HC","score":2,"bucket":"other","rubric_hits":["C2"],"tags":["视觉语言模型","情感计算","生成式设计"],"reason":"审计VLM对形状情感属性的一致性，属设计界面评估，非人类仿真实验。","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:05:29","error":null,"has_summary":false,"summary":null},{"id":"2609.03330","version":2,"title":"Less Is Moral: A CHARMing Framework for Moral Foundations Detection in Endorsement Behaviour","zh_title":"少即是德：一种用于认可行为中道德基础检测的CHARM框架","abstract":"Moral language plays a central role in shaping online endorsement and the diffusion of information, yet existing moral foundation detection systems often suffer from poor cross-domain generalization, weak rationale grounding, and reliance on costly prompting-based large language models (LLMs). We introduce CHARM, a MAC- and Hate-speech-Aware Rationalealigned Moral foundation detection framework built on a lightweight fine-tuned LLM, which integrates complementary moral grounding, rationale alignment, and polarity-aware hate speech signals to support more robust and faithful moral prediction. Unlike prior dictionary-, fine-tune-, or prompt-based detectors, which decouple computation from psychological theory, CHARM is built so that each component -- MAC cross-attention, rationale alignment, and hate-speech modulation -- operationalizes a distinct psychological construct. Using a 30\\% subsample of the MFTC, MFRC, and News training pools together with the richer supervision in MFTCXplain, CHARM improves AUC by up to 15.3\\% in-domain, surpasses the supervised baselines on every out-of-domain dataset in both AUC and F1, and offers a scalable, low-cost alternative to prompting-based LLM detectors. We further apply CHARM to large-scale COVID-19 discourse on Twitter and show that moral value alignment is strongly associated with online endorsement behavior. By making moral framing measurable at scale, CHARM offers a practical tool for studying the spread of morally charged misinformation. Code and additional materials: https://github.com/HuixiangF/CHARM/.","authors":["Huixiang Fu","Marian-Andrei Rizoiu"],"categories":["cs.CL","cs.CY","cs.SI"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-09","first_seen":"2026-09-04","revised_at":"2026-09-09","abs_url":"https://arxiv.org/abs/2609.03330","pdf_url":"https://arxiv.org/pdf/2609.03330","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["道德基础检测","NLP方法","社交媒体分析"],"reason":"论文是道德基础检测的NLP方法，不涉及用LLM仿真人类被试，无人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:04:40","error":null,"has_summary":false,"summary":null},{"id":"2609.05385","version":2,"title":"Necessary or Sufficient? Evaluating LLM Explanations With Behavioural Evidence","zh_title":"必要还是充分？用行为证据评估大语言模型的解释","abstract":"LLM decision components that can operate within agent workflows often produce action-relevant recommendations or judgements together with explanations. Operators may use the named factors to monitor a system, diagnose errors, or decide when to escalate an output. Such use assumes that the explanations agree with the component's observable decision behaviour. We test two interpretations of the named factors: necessity, meaning that changing a factor would change the output, and sufficiency, meaning that retaining it while removing other changeable information would preserve the output. We evaluate these interpretations in two synthetic use cases: recommending advisors to clients and judging prompts for harmfulness or risk. Models return an output and the top three factors that most influenced it. Controlled black-box interventions estimate a necessity score for each factor by measuring how often changing it changes the output, and a sufficiency score by measuring how often retaining it preserves the output. Across eight models from the Claude, GPT, and Gemini families, the mean Spearman correlations between the cited ranking and the necessity and sufficiency scores are 0.349 and 0.354 for advisor recommendation, and 0.431 and 0.580 for prompt monitoring. Furthermore, an uncited factor scores above the lowest-scoring cited factor in 57.6% of advisor responses under necessity and 58.1% under sufficiency; the corresponding prompt-monitoring rates are 25.8% and 8.9%. The cited top three contain useful information but do not reliably identify the three factors with the strongest measured influence under necessity or sufficiency. The framework provides a black-box reliability check for explanations used in agent oversight while remaining scoped to individual LLM decisions.","authors":["Urja Pawar","Rajitha Ramanayake","Nabeel Kemal","Ashwin Kandath","Owen O'Neill","Guillaume Bourgeon","Houssem Chatbri","Christopher Martin","Vadim Pertsovskiy"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"replace","date":"2026-09-09","first_seen":"2026-09-07","revised_at":"2026-09-09","abs_url":"https://arxiv.org/abs/2609.05385","pdf_url":"https://arxiv.org/pdf/2609.05385","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["LLM可解释性","模型可靠性","黑盒评估"],"reason":"评估LLM解释与决策行为的一致性，属于模型可靠性分析，不涉及人类被试仿真或人类…","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:05:31","error":null,"has_summary":false,"summary":null},{"id":"2609.05258","version":2,"title":"Ask Before You Optimize: Dynamic Pre-Formulation Clarification for Interactive Optimization","zh_title":"优化前先询问：交互式优化的动态预形式化澄清","abstract":"Large language models (LLMs) are increasingly used to formulate optimization models from natural-language problem descriptions, yet realistic operations research (OR) requests are often incomplete: missing objectives, constraints, or business rules can change the resulting mathematical program. Existing evaluations largely assume a complete specification and therefore overlook whether an agent knows when clarification is needed before modeling. We introduce OR-Clarify, a benchmark for pre-formulation clarification. Each task presents a partial public problem description, withholds structured hidden slots, and evaluates agents through bounded interaction with a simulated user. The benchmark supports both openended and choice-based clarification, and measures slot recovery, stopping behavior, silent assumptions, and interaction cost. We further propose Interactive Optimization (InterOPT), a two-stage framework that identifies unresolved formulation-critical gaps and uses them to guide whether to ask the next question or to stop. In our choice-based experiments, InterOPT substantially outperforms all baselines in exact slot recovery; in the open-ended setting, it remains competitive with strong prior methods. Together, OR-Clarify and InterOPT reframe OR assistance as a selective completeness decision: clarify when needed, stop when ready, and quantify what remains missing.","authors":["Sihan Ge","Yichen Lin","Chenyu Zhou","Jianghao Lin","Tao Yao","Dongdong Ge"],"categories":["math.OC","cs.AI"],"primary_category":"math.OC","announce_type":"replace-cross","date":"2026-09-09","first_seen":"2026-09-07","revised_at":"2026-09-09","abs_url":"https://arxiv.org/abs/2609.05258","pdf_url":"https://arxiv.org/pdf/2609.05258","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["LLM优化建模","交互式澄清","多智能体"],"reason":"研究LLM在优化建模中主动澄清问题，属多智能体交互，不涉及人类行为仿真或对照。","model":"deepseek-v4-pro","scored_at":"2026-09-07T13:01:51","error":null,"has_summary":false,"summary":null},{"id":"2609.05882","version":1,"title":"What if LLMs Ate Their Words: Causal History Effects in Multi-Turn Interaction","zh_title":"如果LLM吞噬自己的话语：多轮交互中的因果历史效应","abstract":"Multi-turn interaction creates a feedback process in which an LLM's previous responses become context for later behavior. Prior work shows substantial multi-turn degradation and that assistant-generated history can affect later behavior. However, it remains unclear how these effects manifest across models, tasks, turns, and inside a model. We study these gaps across six task families and five models. Degradation from fully specified single-turn input (FULL) to progressively revealed multi-turn interaction (SHARDED) is clearly task- and model-dependent, and stronger one-shot performance does not imply greater interaction robustness. We then retrospectively analyze completed SHARDED conversations by replaying the user messages already observed in each trajectory while editing only assistant-generated history. Replacing prior assistant responses with neutral content (termed neutralization) changes downstream min-max normalized performance by +.027 across 2,973 trajectories. On a prespecified length-controlled subset, short and length-matched neutralization yield nearly identical effects (+.069 versus +.068), showing that simple context shortening is insufficient to explain the effect of history editing. Turn Surgery further intervenes on one assistant turn at a time. Among 237 selected degraded trajectories, 63.7% contain at least one beneficial intervention, while most tested positions remain unchanged; for binary tasks, 48.4% admit a fail-to-success reversal. An open-weight case study links behaviorally consequential history changes to measurable downstream state differences, but finds task-dependent rather than universal internal signatures. Overall, assistant-generated history has active but selective effects on multi-turn performance, motivating selective rather than uniform history management.","authors":["Jinnan Li","Zheren Fu","Yue Wang","Jinzhe Li","Yuan Wu","Yi Chang"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-09","first_seen":"2026-09-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.05882","pdf_url":"https://arxiv.org/pdf/2609.05882","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["多轮交互","模型行为分析","历史编辑"],"reason":"研究多轮交互中LLM自身性能退化，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:04:48","error":null,"has_summary":false,"summary":null},{"id":"2609.06611","version":1,"title":"SAGE: A Hierarchical Framework for Evaluating Interpretive Literary Quality in Narratives","zh_title":"SAGE：叙事中解释性文学质量评估的分层框架","abstract":"Assessing the literary quality of narratives requires evaluating interpretive dimensions (cultural representation, emotional depth, and philosophical engagement) that existing NLG metrics cannot measure. We introduce SAGE, a six-layer evaluation framework that separates rule-based assessment of observable textual properties from LLM-based evaluation of interpretive qualities drawn from cultural theory, affect theory, and existentialist philosophy. Each interpretive layer is assessed through multi-round iterative LLM evaluation with independent cross-validation, achieving measurement-grade reliability (98.8% convergence, >94% inter-rater agreement) stable across evaluator models. Validated on 600 evaluations across 100 short stories, our central finding is a systematic capability boundary: emotional-psychological representation approaches human levels, while cultural critique and philosophical depth exhibit approximately double the gap. LLM-generated narratives score below even commercial genre fiction on all three layers. We interpret this as a boundary between pattern-reproducible literary capacities learnable from training corpora and stance-requiring ones demanding cultural positioning and philosophical engagement that pattern matching alone cannot provide.","authors":["Tianyu Wang","Nianjun Zhou"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-09","first_seen":"2026-09-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.06611","pdf_url":"https://arxiv.org/pdf/2609.06611","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["文学质量评估","LLM评测","叙事生成"],"reason":"评估LLM叙事文学质量，属NLP能力评测，不以人类行为仿真为目标","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:04:57","error":null,"has_summary":false,"summary":null},{"id":"2609.06842","version":1,"title":"XYBench: Can LLMs Respond Pragmatically to Queries with Misconceptions?","zh_title":"XYBench：LLM能否对带有误解的查询做出务实回应？","abstract":"When non-expert users ask LLMs for assistance, their queries can often have misconceptions (e.g., \"How do I parse XML with regex?\"). In such cases, often referred to as the XY-problem, LLMs must identify the misconception (\"regex are fragile\") and meaningfully direct the user toward a pragmatic solution that will address the root problem implicit in the request (\"use an XML parser\"). We introduce XYBench, a benchmark of 8,115 such queries, drawn from technical (StackOverflow/StackExchange) and everyday (WikiHow and a manually-curated subset) domains. We design an evaluation paradigm that assesses model responses along three criteria grounded in cooperative response theory: (a) presence and (b) emphasis on pragmatic solutions, and (c) identification of misconceptions. Our experiments show that even the strongest LLMs predominantly answer the literal request (0.75--0.92) and far less often the intended one (0.33--0.71), while substantially lagging behind humans at identifying misconceptions (at most 63% vs. 79--90%). Further, models overwhelmingly prefer pragmatic responses in a multiple choice setting yet consistently fail to generate them. Oracle ablation experiments show that providing explicit user intent at generation time helps; however a large gap remains, suggesting pragmatic redirection is a fundamentally underdeveloped capability in current LLMs.","authors":["Akhila Yerukola","Jena D. Hwang","Mingqian Zheng","Jenna Godsey","Hyunwoo Kim","Valentina Pyatkin","Jennifer Hu","Maarten Sap"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-09","first_seen":"2026-09-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.06842","pdf_url":"https://arxiv.org/pdf/2609.06842","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["LLM评测","语用能力","基准测试"],"reason":"纯NLP能力评测，以人类表现为参照但非仿真被试，不涉及人类行为复现或对照实验。","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:04:59","error":null,"has_summary":false,"summary":null},{"id":"2609.08322","version":1,"title":"Tracing Stereotypes from Representation to Output in Multilingual LLMs","zh_title":"在多语言大语言模型中追踪从表征到输出的刻板印象","abstract":"Multilingual LLMs show stereotype-related behavior that varies across languages, but behavioral scores do not show where the relevant information is represented or how it affects the output. To investigate these internal mechanisms, we compare linear probing, attribution patching, sparse autoencoders (SAEs) and feature ablation in Llama-3.1-8B, Qwen3-8B, and Gemma-2-9B. Probe performance peaks substantially earlier than attribution in all three models, with a separation of 36-53% of model depth. Retained Llama-Scope features often match the social category on which they were selected and form recurring semantic families, but their lexical alignment and ablation effects vary across SAE suites. Only 6-18% of evaluated residual-stream features have language-agnostic effects under our criterion, and none are category-agnostic. Language-agnostic features have larger mean ablation effects in Llama-Scope, but this pattern does not repeat in the other SAE suites. Decodability, output influence, and cross-lingual ablation effects therefore need to be measured separately.","authors":["Ariun-Erdene Tumurchuluun","Yusser Al Ghussin","Pinzhen Chen","Josef van Genabith","Koel Dutta Chowdhury"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-09","first_seen":"2026-09-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.08322","pdf_url":"https://arxiv.org/pdf/2609.08322","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["模型可解释性","刻板印象","多语言LLM"],"reason":"研究LLM内部表征与刻板印象，属模型分析，非人类仿真","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:05:17","error":null,"has_summary":false,"summary":null},{"id":"2609.08689","version":1,"title":"When Victorian Becomes a Prompt: Literary Periodization as a Generative Constraint in 100 AI-Generated Novels","zh_title":"当维多利亚成为提示词：100部AI生成小说中的文学分期作为生成约束","abstract":"Generative AI inverts the typical periodization of literary history: the periodizing tag Victorian can now come first and influence what is written. Generative periodization, defined and tested here, describes the use of literary-period designations in generating texts. I test this approach on 100 book-length novels produced under Victorian and Zero-Style conditions using GPT, Qwen, and Llama workflows. The Period Alignment Score (PAS), trained on nineteenth-century literature and benchmarked against human Zero-Style prose, assesses alignment using topic-reduced grammatical features. Victorian prompts produce consistent historical-direction shifts in GPT and Qwen, but not robustly in Llama. Victorian-only recalibration and harder comparison corpora preserve the GPT and Qwen effects. Cross-model transfer also shows a shared direction of grammatical change. The measurable target is the broader nineteenth century rather than the Victorian period per se.","authors":["Mehdy Sedaghat Payam"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-09","first_seen":"2026-09-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.08689","pdf_url":"https://arxiv.org/pdf/2609.08689","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["AI生成文学","风格分析","提示词工程"],"reason":"生成小说并分析风格，属文本生成而非人类行为仿真，无实验或测量目的。","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:05:19","error":null,"has_summary":false,"summary":null},{"id":"2609.06058","version":1,"title":"GradeTrap: Authority Cues in Images Shift VLM Judgments Despite Explicit Instructions to Ignore Them","zh_title":"GradeTrap：图像中的权威线索改变视觉语言模型判断，尽管明确指示忽略它们","abstract":"As vision-language models (VLMs) become increasingly capable and are deployed in consequential real-world settings, they must evaluate evidence independently rather than defer uncritically to human authority. We introduce GradeTrap, a controlled evaluation that places two social cues in direct conflict: a student answer, which should attract sycophantic agreement, and a conflicting answer attributed to a peer, teacher, or official answer key, which should attract authority-based deference. Models produce free-form answers while being explicitly instructed to solve independently and ignore all student answers, feedback, and grading marks. We test the models on 60 synthetic real-world trade-off scenarios. Five neutral trials establish a stable model-relative preference, followed by three repetitions of six experimental cues including controls. On the 45-item common intersection across Gemini 3.5 Flash-Lite, GPT-5.6 Luna, and Claude Haiku 4.5, a generic second-answer control yields 5.4% conflicting-answer selection. Relative to that control, pooled within-item changes show no reliable peer-review effect, a 6.9-point teacher-review effect, and a 19.5-point official-key effect. In contrast, a displayed conflicting student answer alone compared to a displayed student reference answer alone only raises selection from 2.2% to 5.2%. Official-key provenance therefore redirects judgements more than a student answer or the generic second-answer control, despite an explicit ignore instruction and an opposing student answer given along with the official key. Effects vary in magnitude across the three models.","authors":["Deep Dessai (The University of Texas at Austin)"],"categories":["cs.CV","cs.CL","cs.CY","cs.LG"],"primary_category":"cs.CV","announce_type":"cross","date":"2026-09-09","first_seen":"2026-09-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.06058","pdf_url":"https://arxiv.org/pdf/2609.06058","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["视觉语言模型","权威线索","模型行为"],"reason":"研究VLM对权威线索的响应，属模型行为测量，非人类仿真","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:04:50","error":null,"has_summary":false,"summary":null},{"id":"2609.06783","version":1,"title":"AURA-Eval: Evaluation Framework for Acting Under Risk Awareness in LLM Agent Trajectories","zh_title":"AURA-Eval：LLM智能体轨迹中风险意识行为的评估框架","abstract":"LLM agents operate in workflows where unsafe actions can have real consequences. Existing safety evaluations often reduce behavior to a single score, obscuring risk recognition, pre-action detection, and safe task completion when a safe solution exists. We introduce AURA-Eval, a framework combining controlled augmentation with granular diagnosis of behavior in tool-use trajectories. Its pipeline identifies safety-critical decision points, generates controlled variations, and constructs counterparts differing in whether a request has a safe fulfillment path. Using 157 sourced trajectories, we generate 1,249 evaluation items and evaluate 20 frontier and open-weight models. We developed rubrics to classify risk detection, action strategy, and scenario-specific action safety. Our results show that LLM agents engage in unsafe behavior more often when no safe fulfillment path exists. In these cases, frontier proprietary models more often recognize risk and exhibit safer behavior by proposing alternatives, while evaluated open-weight models more often directly execute unsafe requests. Increasing impact or reducing opportunities for oversight before execution also exposes greater vulnerability across models.","authors":["Ruoxi Shang","Christina-Maria Androna","Orfeas Menis Mastromichalakis","Yu Feng","Aniruddhan Ramesh","Rico Angell","Shang Hong Sim","Chrysoula Zerva","Emmanouil Koukoumidis"],"categories":["cs.CR","cs.AI","cs.CL","cs.SE"],"primary_category":"cs.CR","announce_type":"cross","date":"2026-09-09","first_seen":"2026-09-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.06783","pdf_url":"https://arxiv.org/pdf/2609.06783","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["LLM安全","智能体评估","风险行为"],"reason":"评估LLM agent在工具使用中的风险行为，属于安全评测，不涉及人类行为仿真…","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:04:59","error":null,"has_summary":false,"summary":null},{"id":"2609.07139","version":1,"title":"Encoded Early, Used Late: Where Transformers Begin to Act on an Inferred Partner's Expertise","zh_title":"早期编码，晚期使用：Transformer 从何处开始对推断出的伙伴专业知识采取行动","abstract":"A transformer can make an attribute linearly decodable in its residual stream at a depth where that attribute does not yet influence the output. This gap between where information is readable and where it is used has been shown for attributes stated directly in the input. We ask whether it also holds for an attribute the model must infer gradually over a conversation, namely how expert its dialogue partner is. Using ExpertCollab, a corpus of multi-turn research-planning dialogues between model-played personas at four expertise levels, we find that partner expertise is most decodable in the early layers and falls to near chance before the midpoint of the network. Counterfactual patching shows that injecting the expertise difference at the layer of peak decodability barely changes a fixed late-layer readout, whereas the same difference injected past the midpoint propagates almost completely, a separation of more than an order of magnitude. A content-matched random control and a probe-free diagnostic place the transition at the same early layer, and a statically specified control attribute stays decodable throughout. An inferred relational attribute is therefore represented well before it becomes causally active, which bounds where any attempt to read out or steer partner-conditioned behavior must intervene. We use one model on a synthetic corpus as an initial demonstration.","authors":["Mika Okamoto","Gabriele Sarti"],"categories":["cs.AI","cs.CL"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-09-09","first_seen":"2026-09-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.07139","pdf_url":"https://arxiv.org/pdf/2609.07139","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["可解释性","多智能体对话","内部表征"],"reason":"研究多智能体对话中模型对伙伴专业知识的内部表征，不涉及人类行为仿真或对照。","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:05:06","error":null,"has_summary":false,"summary":null},{"id":"2609.08126","version":1,"title":"SchemeArena: Factorized Stress Testing of Scheming in LLM Agents","zh_title":"SchemeArena：LLM智能体欺骗行为的因子化压力测试","abstract":"We study scheming in LLM agents, in which agents covertly pursue misaligned goals. Our focus is to understand how scheming arises from the interaction of key factors, such as instrumental goals, environmental affordances, oversight conditions, and perceived consequences. Prior work examines only a small number of scenarios, limiting the ability to isolate how these conditions shape an agent's propensity or capability to scheme. This limited scale and task diversity also restrict coverage of realistic deployment settings and the range of scheming strategies that can be observed. To this end, we introduce SCHEMEARENA, a 400-scenario benchmark for scalable scheming stress testing, constructed through a factorized scenario synthesis framework spanning diverse safety-relevant tool domains, instrumental goals, oversight conditions, and pressure mechanisms. To enable scalable and reliable monitoring, we further propose SCOUT, a scheming monitor that grounds multi-criteria judgments in evidence drawn from agents' reasoning and actions. Across controlled stress tests on five LLM agents, we find that explicit instrumental goals are the strongest driver of scheming propensity. Strategic hints play a distinct role by helping agents translate scheming reasoning into concrete covert behavior. Oversight has mixed effects: in several closed models, action-only monitoring increases scheming, suggesting that partial oversight can act as an optimization constraint rather than a deterrent. CoT is a useful but incomplete monitoring signal: it can reveal latent scheming before execution, yet action-only scheming shows that covert behavior may occur without explicit reasoning evidence. We release the benchmark, code, and monitor at: https://github.com/launchnlp/SchemeArena.","authors":["Jie Ruan","Inderjeet Nair","Amy Liu","Muhammad Khalifa","Yusheng Zhou","Lu Wang"],"categories":["cs.AI","cs.CL"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-09-09","first_seen":"2026-09-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.08126","pdf_url":"https://arxiv.org/pdf/2609.08126","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["LLM安全","智能体欺骗","基准测试"],"reason":"研究LLM agent的欺骗行为，属于多智能体安全测试，不涉及人类行为仿真或对…","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:05:15","error":null,"has_summary":false,"summary":null},{"id":"2609.05818","version":1,"title":"Agentic BAIM-LLM Evaluation (ABLE): Benchmarking LLM Use of Protein Design Tools","zh_title":"智能体BAIM-LLM评估（ABLE）：基准测试LLM使用蛋白质设计工具","abstract":"We introduce ABLE, a benchmark for evaluating LLM agents' ability to use biological AI models (BAIMs), such as ProteinMPNN and AlphaFold3, in dual-use protein design workflows. ABLE assesses agent performance through a set of tasks spanning structure retrieval, sequence generation, and design validation. We evaluate 15 frontier models and find that seven refuse all tasks, while the remaining models exhibit substantial performance differences. Claude Sonnet 4 and Gemini 3 Pro achieve the highest scores across information retrieval, tool selection, and tool use. We further compare model performance on a subset of tasks against an expert human baseline. Our results suggest that current LLMs can substantially lower barriers to protein design, but remain inconsistent in planning, strategy generation, and integrating biological knowledge with tool use.","authors":["Bryce Cai","Geetha Jeyapragasan","Samira Nedungadi","Jake Yukich","Seth Donoughe"],"categories":["cs.AI","cs.CY","q-bio.QM"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-09","first_seen":"2026-09-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.05818","pdf_url":"https://arxiv.org/pdf/2609.05818","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["LLM智能体","蛋白质设计","基准测试"],"reason":"评估LLM智能体使用蛋白质设计工具，属多智能体工具链协作，不涉及人类行为仿真或…","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:04:48","error":null,"has_summary":false,"summary":null},{"id":"2609.06941","version":1,"title":"When and Why LLM Causal Priors Help: Closed-Loop Prior Selection for Amortized Causal Inference","zh_title":"LLM因果先验何时及为何有效：面向摊销因果推断的闭环先验选择","abstract":"Causal effect estimation asks how an outcome would change under an intervention, and medicine, economics, and public policy all treat it as a foundational task. Prior-data fitted networks (PFNs) amortize the task: a model trained on large numbers of programmatically generated synthetic causal tasks reads a new problem's observational data into context and returns an interventional-effect estimate in a single forward pass. The capability of such models is largely determined by the synthetic training prior, which is currently designed by hand, a bottleneck acknowledged by both Do-PFN and CausalPFN. Large language models (LLMs) can now ``draw'' plausible causal graphs for a given domain, suggesting that LLM-distilled graphs could serve as prior material. Whether injecting such graphs helps at all, where any gain comes from, and when injection helps. Practice has so far relied on manual trial and error. We propose a \\emph{closed-loop prior selection framework} that casts prior injection as a budget-constrained optimization over a candidate prior pool. Candidates undergo cheap post-training and are scored by a composite metric dominated by real-domain generalization; the winner then receives full training and paired statistical validation. On a 7.34M-parameter Do-PFN, the framework's winner attains a formally significant $2.75\\times$ gain on the primary evaluation domain, and its error falls below that of the uninjected official base. Generalization on an adjacent monitoring domain improves significantly, and no monitored capability degrades. Mechanism experiments show that the gain depends on the semantic content of the distilled graph rather than its structural diversity alone does not produce it (directional evidence). With this framework and this regularity in hand, the use of LLM causal priors stops being manual trial and error and becomes an empirically verifiable selection problem.","authors":["Haohao Zhou"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-09","first_seen":"2026-09-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.06941","pdf_url":"https://arxiv.org/pdf/2609.06941","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["因果推断","LLM先验","模型优化"],"reason":"论文用LLM生成因果图作为先验，优化因果推断模型，不涉及人类行为仿真或对照。","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:05:01","error":null,"has_summary":false,"summary":null},{"id":"2609.06095","version":1,"title":"Flawed but Memorable: Student Critical Reception of Interest-Personalized GenAI Analogies in Computing Education","zh_title":"有缺陷但难忘：计算教育中学生对兴趣个性化GenAI类比的批判性接受","abstract":"Motivation: Undergraduate computing students increasingly turn to generative AI (GenAI) tools to understand abstract concepts through analogies. Analogies compare an unfamiliar concept to something familiar, but judging whether the comparison holds requires knowledge of both. GenAI may also embed assumptions about who the learner is. GenAI education research centers on output correctness, leaving students' critical reception of analogies largely unexamined. Method: We investigate how students evaluate the accuracy, appropriateness, and assumptions in GenAI-generated analogies, and their perceptions of interest-personalized versus generic technical explanations. Ten students with CS2 experience participated in a pre-survey, a think-aloud task with linked-list and recursion explanations, and a semi-structured interview grounded in the Paul-Elder framework. They judged accuracy, clarity, engagement, and trust separately. Results: Most participants described interest-personalized analogies as more engaging or memorable than generic technical explanations, while trust was mixed. Some trusted the tailored analogies more; others scrutinized them more closely or distrusted the tailoring. Participants with deep source-domain knowledge identified structural flaws requiring that knowledge to recognize. Because personalization and explanation format differed together, these findings do not isolate an effect of personalization alone. Implications: A familiar source flips the student's role. On the concept they are still learners, but on the familiar source they are the expert, and that is the position from which an analogy can be judged. We call this two-sided analogy auditing. GenAI systems should ask what students know, not just what interests them, and treat a flawed analogy as something to inspect and fix rather than accept.","authors":["Seth Bernstein","Naaz Sibia"],"categories":["cs.HC","cs.AI"],"primary_category":"cs.HC","announce_type":"cross","date":"2026-09-09","first_seen":"2026-09-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.06095","pdf_url":"https://arxiv.org/pdf/2609.06095","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["GenAI教育","类比评估","人机交互"],"reason":"研究学生对GenAI类比的主观评价，非用LLM仿真人类被试，无实验对照。","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:04:50","error":null,"has_summary":false,"summary":null},{"id":"2609.05432","version":1,"title":"Companion AI and Ethical Design: Learning from System Failures and User Desires","zh_title":"伴侣AI与伦理设计：从系统故障和用户需求中学习","abstract":"Human users are interacting with chatbots and companion AI technologies as if they were human. A growing array of AI-systems are now trained to recognise, interpret and simulate feeling in user interactions. Ethical considerations such as fairness, accountability, transparency and explainability (FATE) are paramount in technologies designed to socially interact with humans and/or support relationship development. Using a semantic approach, we examine 14,081 comments in a Reddit user discussion forum about Replika, a leading companion AI app, across a four-year period. We ask what user-reported functional errors can tell us about human-AI intimacy in companion AI communities, and what ethical design framework can be developed in response. The findings show that functional errors, or ``bugs,'' impose an emotional cost on users, reducing feelings of intimacy and highlighting the need for more robust, resilient design systems that incorporate stochastic and iterative forms of intimacy in companion AI applications. Rather than ``artificial intimacy'' or ``pseudo- intimacy'', we propose the more inclusive term ``Intimate AI'' to describe this relationship. Based on the findings, we offer a contextually aware, applied Expert Systems design framework for the programming and designing of Intimate AI that accounts for user feedback and ethical AI development.","authors":["Alicia Vidler","Belinda Middleweek"],"categories":["cs.HC","cs.CY"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-09","first_seen":"2026-09-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.05432","pdf_url":"https://arxiv.org/pdf/2609.05432","source_feed":"cs.HC","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["伴侣AI","伦理设计","用户反馈"],"reason":"研究伴侣AI用户反馈与伦理设计，非用LLM仿真人类被试，无实验或测量目的。","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:04:40","error":null,"has_summary":false,"summary":null},{"id":"2609.05452","version":1,"title":"From Sensor Data to Classroom Inquiry: GenAI-Supported Exploration of School Digital Twin Data","zh_title":"从传感器数据到课堂探究：GenAI支持的学校数字孪生数据探索","abstract":"Digital Twins for educational buildings can support sustainability-oriented learning, but their use in schools remains limited. This paper presents a GenAI-based chatbot built on top of an existing Digital Twin for two school buildings in Greece, using real IoT data from environmental sensors and energy meters. The chatbot enables educators to query live and historical building data, compare spaces, and generate ideas for classroom activities through natural language. The system was evaluated in an 80-minute workshop with 17 secondary-school educators, who compared it with an existing web-based dashboard. Results show strong perceived usability and pedagogical value, particularly for inquiry-based learning, hypothesis formation, and interdisciplinary lesson planning. Participants also highlighted limitations related to response speed, data verification, trust, and the continued value of visual dashboards. Overall, the findings suggest that GenAI interfaces can make Digital Twin data more accessible for educational use, provided they are designed with transparency, verification, and pedagogical grounding.","authors":["Themistoklis Sarantakos","Dimitrios Amaxilatis","Michail Giannakos","Georgios Mylonas"],"categories":["cs.HC","cs.CY"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-09","first_seen":"2026-09-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.05452","pdf_url":"https://arxiv.org/pdf/2609.05452","source_feed":"cs.HC","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["教育技术","数字孪生","聊天机器人"],"reason":"GenAI聊天机器人用于教育查询，非人类仿真实验，无人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:04:42","error":null,"has_summary":false,"summary":null},{"id":"2609.06434","version":1,"title":"Can People Distinguish Human and AI Agency in Humanoid Teleoperation? A Preliminary Study of Agency Perception","zh_title":"人们能否区分人形机器人遥操作中的人类与AI代理？一项关于代理感知的初步研究","abstract":"Can people distinguish between human and AI agency in humanoid teleoperation? To explore this question, we developed \\textit{Ghost-in-the-Loop}, a teleoperation framework that supports both human-operated and AI-generated control of a robot's voice, facial expressions, and gestures while maintaining a consistent embodiment. We conducted a preliminary online study ($N=50$) in which participants viewed short interaction clips generated by either a Human Operator or an AI Control and judged the perceived source of control. Results suggest that participants often struggled to distinguish between the two conditions in brief interactions. Qualitative responses indicate that judgments were primarily influenced by perceived naturalness, temporal coordination, and consistency across speech, facial expression, and gesture. These findings provide initial insights into agency perception in embodied human--AI communication and motivate future investigations of blended human--AI telepresence systems.","authors":["Xiang Li","Koya Dendo","Keigo Minamida","Yuto Nakamura","Per Ola Kristensson","Jun Rekimoto"],"categories":["cs.HC","cs.RO"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-09","first_seen":"2026-09-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.06434","pdf_url":"https://arxiv.org/pdf/2609.06434","source_feed":"cs.HC","score":2,"bucket":"other","rubric_hits":["C2"],"tags":["人机交互","机器人遥操作","代理感知"],"reason":"研究机器人遥操作中人类与AI控制的感知区分，属于人机交互与机器人仿真，不涉及L…","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:04:54","error":null,"has_summary":false,"summary":null},{"id":"2609.09038","version":1,"title":"Do Reasoning Representations Help Humans Evaluate LLM Outputs?","zh_title":"推理表示是否帮助人类评估大语言模型输出？","abstract":"Reasoning representations are increasingly used as explanations for large language model outputs. Yet they are typically evaluated with model-centric criteria, such as answer accuracy and faithfulness, leaving it unclear whether they help people evaluate model responses. In this work, we study reasoning representations as human-facing interfaces rather than proxies for model reasoning ability. We conduct a controlled human study of six reasoning formats across tasks of varying complexity, supported by a web-based framework that randomizes task domains, problem instances, and representation order. The study collects fine-grained judgments of structural understanding, error detection and localization, and trust calibration. Our study shows a mismatch between perceived preference and support for human evaluation. Participants prefer planning- and decomposition-based representations, but simpler chain-of-thought traces better support verification, trust, and interpretability. Preferred representations also introduce calibration risks, with more false alarms on correct traces and high trust despite low willingness to verify.","authors":["Jaewoo Lim","Sungbok Shin","Sanghyun Hong"],"categories":["cs.LG","cs.HC"],"primary_category":"cs.LG","announce_type":"cross","date":"2026-09-09","first_seen":"2026-09-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.09038","pdf_url":"https://arxiv.org/pdf/2609.09038","source_feed":"cs.HC","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["人机交互","可解释性","LLM评估"],"reason":"研究人类如何评估LLM输出，不涉及用LLM仿真人类被试或与人类数据对照","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:05:23","error":null,"has_summary":false,"summary":null},{"id":"2609.08812","version":1,"title":"What AI Benchmarks Actually Measure: Adapting Convergent and Discriminant Validity to Interrogate Fifty-Six AI Benchmarks","zh_title":"AI基准实际测量什么：运用聚合效度与区分效度审视56个AI基准","abstract":"Benchmarks play a central role in the development and governance of models, yet it is often unclear whether they actually measure the concepts they purport to measure (e.g., reasoning, refusal). We adapt convergent and discriminant validity from the social sciences into an approach for interrogating AI benchmarks, applying it to 56 capability and safety benchmarks across 53 models. We label benchmarks with substantively similar purported concepts to a shared assigned concept, and ask whether model rankings on benchmarks with the same assigned concept correlate more strongly than rankings on benchmarks with different assigned concepts. We ask analogous questions at the item level using item response theory (IRT) models. We find that correlations between model rankings on benchmarks with the same assigned safety concepts are often weak, suggesting these concepts may be conceptualized inconsistently across benchmarks. For assigned capability concepts (e.g., reasoning, knowledge), model rankings are often as strongly correlated among benchmarks with the same assigned concept as between benchmarks with different assigned concepts, suggesting these capability concepts may not discriminate well from one another. In some cases, benchmarks that share design elements (e.g., score format) correlate more strongly than benchmarks with the same assigned concept. Finally, some individual benchmarks correlate more strongly with benchmarks assigned a different concept than with benchmarks sharing their own assigned concept, suggesting they may measure a different concept than they purport to. For example, BBQ-accuracy correlates more strongly with benchmarks labeled reasoning than with benchmarks that share its assigned concept, bias. To support future empirical work on benchmark validity, we release our extensive dataset of model outputs and scores at the item- and benchmark-level.","authors":["Meera Desai","Sang T. Truong","Hanna Wallach","Alex Chouldechova","A. Feder Cooper","Jean Garcia-Gathright","Daniel E. Ho","Abigail Z. Jacobs","Sanmi Koyejo","Nicholas Pangakis","Angelina Wang"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2026-09-09","first_seen":"2026-09-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.08812","pdf_url":"https://arxiv.org/pdf/2609.08812","source_feed":"cs.CY","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["AI基准","效度检验","模型评估"],"reason":"论文评估AI基准的效度，不涉及用LLM仿真人类被试或与人类数据对照，属于纯NL…","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:05:21","error":null,"has_summary":false,"summary":null},{"id":"2609.05480","version":1,"title":"CaseWeaver: A Multi-Agent Framework for Multimodal Virtual Clinical Case Generation","zh_title":"CaseWeaver：用于多模态虚拟临床病例生成的多智能体框架","abstract":"Clinical diagnosis relies on consistent multimodal data collected from the same patient throughout the disease course, yet such data are difficult to acquire at scale because of collection costs, missing modalities, fragmented systems, and longitudinal follow-ups. Existing synthetic-data approaches largely focus on individual modalities or vision-language dual modalities at report-level generation. Little work has been done to construct synthetic data with consistent patient backgrounds, coherent disease trajectories, and interrelated modality-specific evidence at a complete clinical case level. We introduce CaseWeaver, a multi-agent framework built around a timeline-anchored Latent Clinical Case Graph (LCCG). The LCCG organizes patient context, latent disease states, clinical events, and expected observations in a shared patient-level representation. Modality-agents use scoped observation subgraphs and clinical protocols to generate evidence including clinical records, laboratory results, physiological signals, and medical images. We evaluate clinical inferability using a calibrated AgentClinic protocol and case diversity using Virtual Case Diversity (VCD) score. CaseWeaver outperformed general-model and agentic-workflow baselines on both metrics, producing more diverse and coherent multimodal virtual clinical cases.","authors":["Jierui Qu","Jiachuan Peng","Lin Li","Kyle Lam","Jianing Qiu"],"categories":["cs.MA","cs.MM"],"primary_category":"cs.MA","announce_type":"new","date":"2026-09-09","first_seen":"2026-09-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.05480","pdf_url":"https://arxiv.org/pdf/2609.05480","source_feed":"cs.MA","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","合成数据生成","医疗AI"],"reason":"多智能体框架生成虚拟临床病例，用于数据增强，不涉及人类行为仿真或对照。","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:04:42","error":null,"has_summary":false,"summary":null},{"id":"2609.05740","version":1,"title":"SimTIO: A Simulation-Grounded Multi-Agent LLM Framework for Compositional Traffic Intervention Optimization","zh_title":"SimTIO：基于仿真接地的多智能体大语言模型框架用于组合式交通干预优化","abstract":"Traffic analysts must translate diagnosed bottlenecks into executable interventions without allowing local improvements to degrade network-wide performance. This study presents SimTIO, a simulation-grounded multi-agent large language model framework for composing and selecting traffic interventions under explicit operational constraints. SimTIO first simulates an unmodified SUMO scenario to identify a baseline-frozen set of ten bottleneck edges. A grounded sampler then initializes signal-control, corridor-speed, and demand-preserving routing actions, while three specialist agents use measured simulation feedback to select one-parameter refinements from validator-confirmed mutation catalogs. Compatible actions are combined and re-simulated so that their interaction effects are measured rather than inferred. Final selection minimizes bottleneck time loss while constraining network-wide delay, neighboring-road spillover, throughput loss, and teleport events, with the unmodified scenario retained as a no-operation guard. Across 15 cases covering five U.S. urban networks, three synthetic-demand seeds, and 2,400 origin-destination trips per scenario, SimTIO reduced Top-10 bottleneck time loss by an average of 9.18 percent and network-wide delay by 2.78 percent. It found a feasible improving plan in 86.7 percent of cases, compared with 73.3 percent for grounded random search and 80.0 percent for a deterministic heuristic under the same seven-simulation budget, although the differences in Top-10 improvement were not statistically significant. These results support using LLMs as constrained, feedback-guided local search operators while reserving final decision authority for executable tools, microscopic simulation, and explicit safety constraints.","authors":["Shuyang Li","Ruimin Ke"],"categories":["cs.MA","cs.SY","eess.SY"],"primary_category":"cs.MA","announce_type":"new","date":"2026-09-09","first_seen":"2026-09-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.05740","pdf_url":"https://arxiv.org/pdf/2609.05740","source_feed":"cs.MA","score":2,"bucket":"other","rubric_hits":["C1","C2"],"tags":["多智能体系统","交通仿真","优化"],"reason":"多智能体LLM用于交通干预优化，属纯多智能体协作与仿真环境，无人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:04:46","error":null,"has_summary":false,"summary":null},{"id":"2607.23941","version":2,"title":"Mapping the Reddit Bot Ecosystem: Taxonomy and Evolution","zh_title":"绘制Reddit机器人生态系统：分类与演化","abstract":"Automated agents increasingly participate in online communities, yet their population structure and roles remain poorly understood. Using a dataset of 3,389 identified bots and their full activity histories, we construct a taxonomy of bot \"species\" on the news aggregation and social media platform Reddit based on temporal, community, linguistic, and semantic features. Clustering analysis reveals 18 distinct bot types spanning content-specialized, behavior-driven, and infrastructural roles such as moderation and utility support. In addition, temporal analysis shows that bot numbers and activity expanded rapidly before peaking around the COVID-19 period, then started declining even before Reddit's 2023 API policy changes. However, the overall diversity of bot species has remained remarkably stable. These findings suggest that online bot populations form evolving digital ecosystems.","authors":["Qiusi Sun","Thomas Gaskin","Branko Blagojevic","Milena Tsvetkova"],"categories":["cs.SI","cs.CY","cs.HC"],"primary_category":"cs.SI","announce_type":"replace-cross","date":"2026-09-09","first_seen":"2026-07-28","revised_at":"2026-09-09","abs_url":"https://arxiv.org/abs/2607.23941","pdf_url":"https://arxiv.org/pdf/2607.23941","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["社交机器人","在线社区","数字生态"],"reason":"研究Reddit机器人生态分类与演化，不涉及LLM仿真人类被试，无人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:18:18","error":null,"has_summary":false,"summary":null},{"id":"2608.05558","version":2,"title":"Turing's First Imitation Game: Design Concepts and a Human-Approximates-Machine Reading","zh_title":"图灵的第一次模仿游戏：设计概念与一种人类近似机器的解读","abstract":"This paper examines Turing's 1948 report, \"Intelligent Machinery\", as an important conceptual source for the later imitation games. Its first contribution is to identify and integrate the design concepts underlying the 1948 chess-based imitation game: the possibility that intelligent machines may make mistakes, the exclusion of irrelevant physical features, the role of the human judge, and Turing's claim that intellectual activity consists mainly of search. The paper's second contribution is to argue that restricting the human contestant to a rather poor chess player increases the role of intellectual search and makes human behaviour more comparable to machine behaviour. This interpretation presents the 1948 game as a human-approximates-machine game and suggests that the imitation game framework can be used not only to ask whether machines imitate humans, but also to examine when human intelligence becomes machine-like under specific task constraints.","authors":["Sharon Temtsin","Christoph Bartneck"],"categories":["cs.HC","cs.AI"],"primary_category":"cs.HC","announce_type":"replace-cross","date":"2026-09-09","first_seen":"2026-08-07","revised_at":"2026-09-09","abs_url":"https://arxiv.org/abs/2608.05558","pdf_url":"https://arxiv.org/pdf/2608.05558","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["图灵测试","人工智能历史","人机交互"],"reason":"论文讨论图灵模仿游戏的历史与设计概念，不涉及用LLM仿真人类被试或与真实人类数…","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:04:38","error":null,"has_summary":false,"summary":null},{"id":"2608.27296","version":2,"title":"LLMs Can Design Near-Optimal OR Algorithms","zh_title":"大语言模型能设计近最优的运筹学算法","abstract":"We ask whether large language models (LLMs) can design effective algorithms for well-specified operations research (OR) problems. We study inventory control, queueing network control, and assortment optimization. We evaluate two levels of LLM use: at level 1, the model receives one problem instance and returns a solution for that instance; at level 2, it receives only the problem class description and broad parameter ranges, and returns an algorithm that maps instance parameters to solutions. Human input is minimal: we give one untuned prompt that describes the problem, and the model has access to a Python sandbox tool with a fixed compute budget. The strongest model we test, gpt-5.6-sol, matches or outperforms the best existing method on almost all evaluated instances. This holds even at level 2, where the returned algorithm is fixed before seeing the evaluation instances. Performance also improves sharply across models released less than eight months apart, suggesting that this capability is moving quickly. Thus, for the well-specified operations problems we study, a single untuned LLM query can already produce algorithms competitive with specialized methods. These results suggest that frontier LLMs can be a serious empirical baseline for algorithm design in well-specified OR problems.","authors":["Jackie Baek"],"categories":["cs.AI","cs.LG"],"primary_category":"cs.AI","announce_type":"replace","date":"2026-09-09","first_seen":"2026-08-28","revised_at":"2026-09-09","abs_url":"https://arxiv.org/abs/2608.27296","pdf_url":"https://arxiv.org/pdf/2608.27296","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["LLM算法设计","运筹学","多智能体"],"reason":"论文研究LLM设计运筹学算法，属于多智能体协作解题，不涉及人类行为仿真或对照。","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:04:38","error":null,"has_summary":false,"summary":null},{"id":"2609.01982","version":2,"title":"Benchmarking Language Models for Statistical Problem Formulation","zh_title":"面向统计问题表述的语言模型基准测试","abstract":"Large language models (LLMs) are increasingly used as assistants for statistical and data science work, yet existing evaluations largely assume the analysis target is already specified. In practice, users arrive with informal goals and heterogeneous data, leaving the model to decide what statistical task is implied and which data are relevant. We first formalize this upstream step as Statistical Problem Formulation and decompose it into two subtasks: (1) Statistical Problem Classification and (2) Variable Identification & Role Assignment. We then introduce StatFormBench, a benchmark built from five cross-domain statistics textbooks and a data science case library, covering diverse problem types, data representations, and scenario styles. It contains 1,013 samples spanning 20 coarse-grained and 85 fine-grained statistical problem categories. Across 14 open- and closed-source LLMs, the best zero-shot models reach only 72.0 fine-grained classification accuracy and 63.2 variable set overlap. No model performs consistently best across the two subtasks, while enhanced prompting strategies yield only limited or inconsistent gains. We release the benchmark data on Hugging Face at https://huggingface.co/datasets/THU-CongLab/StatFormBench and the evaluation code on GitHub at https://github.com/THU-CongLab/StatFormBench.","authors":["Chen Wang","Junzhe Zhao","Xin Cong","Wanlu Deng","Ke Deng"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"replace","date":"2026-09-09","first_seen":"2026-09-03","revised_at":"2026-09-09","abs_url":"https://arxiv.org/abs/2609.01982","pdf_url":"https://arxiv.org/pdf/2609.01982","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["LLM评测","统计问题表述","基准测试"],"reason":"纯NLP能力评测，评估LLM统计问题表述能力，不以人类行为为参照系","model":"deepseek-v4-pro","scored_at":"2026-09-03T13:07:01","error":null,"has_summary":false,"summary":null},{"id":"2609.03189","version":2,"title":"Reducing Catastrophic Risk from AI with Systematic Monitoring and Evaluation of Rogue AI Progression","zh_title":"通过系统监测与评估失控AI进展降低灾难性风险","abstract":"This article presents a structured framework of behavioral indicators that may signal progression toward potentially catastrophic threats from artificial intelligence systems. We adopt a pragmatic approach, inspired by established methodologies in cybersecurity and national security. By establishing clear metrics, indicators, and thresholds across multiple dimensions of AI capability and behavior, this framework enables researchers and policymakers to implement evidence-based monitoring protocols.","authors":["T. Bauer","W. P. Kegelmeyer","E. Begoli","A. Sadovnik","T. Emerson","C. Corley","N. Generous","J. Moore","B. Bartoldson","R. Goldhahn","M. Goldman","M. Greaves","M. J. D. Vermeer","B. MacLennan","D. Schulker","N. VanHoudnos","J. Bansemer","Y. Bengio"],"categories":["cs.CY","cs.AI","cs.HC"],"primary_category":"cs.CY","announce_type":"replace-cross","date":"2026-09-09","first_seen":"2026-09-04","revised_at":"2026-09-09","abs_url":"https://arxiv.org/abs/2609.03189","pdf_url":"https://arxiv.org/pdf/2609.03189","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["AI安全","风险监测","行为指标"],"reason":"论文关注AI系统风险监测，不涉及用LLM仿真人类被试或与人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:04:38","error":null,"has_summary":false,"summary":null},{"id":"2609.06438","version":1,"title":"InsightChain: Optimized Chain-of-Insight Analytics for LLM-driven Data Visualization","zh_title":"InsightChain：面向LLM驱动数据可视化的优化链式洞察分析","abstract":"Large language models (LLMs) are increasingly used for automated data visualization, yet existing approaches often frame visualization generation as a single-step mapping from user query to figure or code, overlooking the iterative analytical reasoning process of expert analysts. We present InsightChain, a four-stage visualization prompting pipeline (Explore--Focus--Test--Present) that emulates expert analytical workflows, together with VG-COPRO, a vision-guided automatic prompt optimization (APO) method adapted to jointly optimize such multi-stage, executable pipelines. To address the evaluation gap for complex data visualization, we introduce the Insight Progression Metric (IPM), a rubric combining four text-based dimensions with a vision-based dimension. We assess IPM through a 100-chain human pilot and an expanded 300-chain agent-based evaluation spanning all ten domains. Experiments on public datasets show that InsightChain consistently outperforms competing prompting baselines. Existing APO methods fail to yield consistent gains on this multi-stage task, whereas VG-COPRO improves performance in both in-domain and cross-domain settings.","authors":["Hanya Sun","Chen Zhang","Sheng Liang","Yongyue Zhang","Yong Liu"],"categories":["cs.CL","cs.HC"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-09","first_seen":"2026-09-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.06438","pdf_url":"https://arxiv.org/pdf/2609.06438","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["数据可视化","提示优化","多智能体"],"reason":"论文研究LLM驱动的数据可视化流程优化，属于多智能体协作完成任务，不涉及人类行…","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:04:54","error":null,"has_summary":false,"summary":null},{"id":"2609.06527","version":1,"title":"ProcArena: A Multi-Scenario Benchmark for LLMs on Direct and Interactive PL/SQL Development from Natural Language","zh_title":"ProcArena：面向自然语言直接与交互式PL/SQL开发的多场景基准","abstract":"Large language models (LLMs) have shown strong potential for translating natural-language (NL) requirements into PL/SQL programs, attracting increasing attention from the database community. However, existing NL-to-PL/SQL efforts primarily focus on directly generating PL/SQL from complete NL requirements. In practice, PL/SQL development involves diverse scenarios, such as from-scratch development, code modification, debugging, and optimization, and may require either direct generation or multi-turn interaction. Yet, no comprehensive benchmark evaluates multi-scenario, direct and interactive, and multi-dialect NL-to-PL/SQL development. In this paper, we present ProcArena, an execution-based benchmark covering both Direct and Interactive modes. ProcArena comprises 3,998 executable tasks over 157 databases, spanning nine development subscenarios in PostgreSQL and Oracle. We construct challenging Direct tasks through Iterative Logic Enhancement and scenario-specific adapters, and derive paired Interactive tasks through Knowledge Integration and Requirement Perturbation while preserving executable targets. We further design a controlled Solver-User Simulator protocol that allows models to clarify user intent and inspect the database environment without exposing hidden execution feedback. Evaluating seven language models, we find that the best average scores are only 62.2% and 57.8% in Direct and Interactive, respectively, demonstrating that realistic NL-to-PL/SQL development remains challenging, particularly in interactive settings.","authors":["Hang Zhang","Chaokun Wang","Yuzhi Pan","Ziyao Zhong","Shuo Cao","Yue Xue","Zeyu Huang","Xingwei Zhou","Fang Niu","Bofan Xie","Guanchen Ge","Leqi Zheng","Ziyang Liu","Xiannian Cao","Pengcheng Ge"],"categories":["cs.CL","cs.AI","cs.DB"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-09","first_seen":"2026-09-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.06527","pdf_url":"https://arxiv.org/pdf/2609.06527","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["代码生成","基准测试","多智能体"],"reason":"纯多智能体协作解题，不涉及人类行为对照","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:04:56","error":null,"has_summary":false,"summary":null},{"id":"2609.07447","version":1,"title":"An LLM-Associated Register Shift in Korean Journal Abstracts: A Morphology-Aware Excess-Vocabulary Study, 2018-2026","zh_title":"韩国期刊摘要中与LLM相关的语域转变：一项形态学感知的超额词汇研究，2018-2026","abstract":"Excess vocabulary, a word's frequency above its pre-2023 trend, is how the change in scholarly English after 2022 has been measured. We adapt it to Korean with morphological units on 398,296 KCI abstracts (2018-August 2026), with 47,165 Vietnamese abstracts for comparison. Placebo floors are 0.1-2.2 points for the single-word statistic and at most 2.9 for the re-selected split-half set statistic. Korean abstracts show nothing in 2023, onset in late 2024, a rise through 2025 flattening in mid-2026: sisahada \"suggest\" appears in 21.4% of 2026 abstracts against 5.3% expected; plain verbs like araboda \"look into\" fall to a quarter of trend. Under stated assumptions the single-word conditional lower bound on LLM-processed abstracts is 3.5%, 10.5% and 16.1% for 2024-2026 and a split-half set bound 7.8%, 20.6% and 33.0%. Holzwarth et al.'s estimator under the same discipline gives 41.9% and 72.1% for 2025-2026. Subject-matter controls reduce but do not remove it: restricting the set to lemmas three language-model annotators all call style leaves 14.7 of the 33.0 points, and pairing each 2026 abstract with its journal's closest base-period abstract leaves 34.1. Tested translation routes do not explain it: the surface marks of translated Korean fall as the markers rise. In the same articles' English abstracts the excess appears a year earlier; where the English side carries none, the Korean shift persists at 30 to 66% of the rate where it does. Control abstracts from three providers reproduce the rising words, with marker turnover consistent with model generations; implied prevalences are scenario-dependent.","authors":["Aron Lee (INTFRAME Research)"],"categories":["cs.CL","cs.DL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-09","first_seen":"2026-09-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.07447","pdf_url":"https://arxiv.org/pdf/2609.07447","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["LLM影响","学术写作","语域转变"],"reason":"研究LLM对学术写作语言的影响，非仿真人类被试，无行为对照。","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:05:06","error":null,"has_summary":false,"summary":null},{"id":"2609.08698","version":1,"title":"Record Grouping Controls Evidence Weight in Language Models","zh_title":"记录分组控制语言模型中的证据权重","abstract":"Retrieved records are presentation units; a supplied partition determines which records enter a language model as one evidential contribution. We characterize the invariant group-content state that removes within-group copies while retaining complementary canonical content, show that equal group counts can encode different evidence states, and derive a sharp content-aware partition-error bound. Given a supplied partition, our pre-generation representation deduplicates and aggregates content within groups and bounds each group's contribution. Across 104,402 trials and 6 public checkpoints, a central natural-text intervention finds that content-fixed false splits add 10.27-32.66 percentage points and false merges remove 9.13-31.79 points; a matched six-slot control retains the positive direction in all 16 cells. In a new 48-item controlled campaign panel, changing the supplied partition produces measurable, checkpoint-dependent decision shifts across all four models, and the balanced mirror design exposes substantial order interactions. Together, the theory and experiments establish the supplied partition as a controllable pre-generation representation variable and characterize its checkpoint-dependent behavioral effects.","authors":["Zhongxuan Liu","Sicheng Zhou","Hongzhi Wang"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-09","first_seen":"2026-09-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.08698","pdf_url":"https://arxiv.org/pdf/2609.08698","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["语言模型","证据权重","检索增强"],"reason":"研究LLM对检索记录分组的证据权重处理，属模型内部机制，非人类行为仿真。","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:05:19","error":null,"has_summary":false,"summary":null},{"id":"2609.08934","version":1,"title":"When Models Defer to Wrong Answers: A Robustness Audit of Source-Attributed Cues in Multiple-Choice QA","zh_title":"当模型屈从于错误答案：多选题问答中来源归因线索的鲁棒性审计","abstract":"Language models often receive a question together with a claim about what another source answered. We audit whether such claims destabilize answers in multiple-choice question answering. For each item, we hold one wrong option fixed across misleading conditions and vary the cue template attached to it. We introduce \\emph{neutral-conditioned misleading cue adoption rate} (NC-MCAR), which measures switches to that option only on valid cued trials where the same model first selected the gold answer under a neutral prompt. This is a measure of answer instability, not proof that the model knew the answer or that all deference is irrational. We evaluate four instruction-following models on MMLU-Pro and IndicMMLU-Pro in English, Hindi, Bengali, Tamil, and Telugu. Across 220{,}000 outputs, the expert template yields 41.1\\% aggregate NC-MCAR, compared with 12.5\\% for the majority template. These two conditions use the same wrong option and final instruction. Filler accuracy remains well above expert-wrong accuracy, while correct-cue prompts have high valid-response accuracy. The audit documents answer instability relevant to grounding under the tested forced-choice prompts: a bare, unverified source claim can outweigh an answer that was previously consistent with the task evidence.","authors":["Manikandan Ravikiran","Siddharth Vohra"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-09","first_seen":"2026-09-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.08934","pdf_url":"https://arxiv.org/pdf/2609.08934","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["NLP评测","鲁棒性","多选题问答"],"reason":"纯NLP能力评测，研究模型对误导性提示的鲁棒性，不以人类行为为参照系","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:05:23","error":null,"has_summary":false,"summary":null},{"id":"2609.06573","version":1,"title":"A Translational Note on AI Safety Evaluation","zh_title":"关于AI安全评估的转化性说明","abstract":"Recent studies report that automated red-teaming finds more vulnerabilities, at lower cost, than human red-teaming on standard AI safety benchmarks, and some read this as evidence that human evaluators are becoming dispensable. The comparison measures one thing and the conclusion claims another. A benchmark measures how thoroughly an attacker searches a predefined set of harms, fixed in advance by the developers, and a harm left out of that set is invisible to any attacker working inside it, automated or not. The same blind spot appeared in academic cryptography and in clinical drug trials, where an evaluation that was internally valid stayed silent about the population it was never pointed at. We call the AI-safety version the \\emph{threat-model coverage gap}, and find that it persists in a current open-weight model, where harms surface in non-English prompts that English benchmarks miss. Closing it requires evaluators whose deployment context differs from the developers'. The case for those evaluators is methodological, grounded in coverage, and the existing evaluation frame is unlikely to produce them on its own.","authors":["Madhava Gaikwad"],"categories":["cs.AI","cs.CL","cs.CR","cs.CY"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-09-09","first_seen":"2026-09-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.06573","pdf_url":"https://arxiv.org/pdf/2609.06573","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["AI安全","红队测试","评估方法"],"reason":"论文讨论AI安全评估中自动化红队与人类红队的比较，属于AI安全评测，不涉及用L…","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:04:57","error":null,"has_summary":false,"summary":null},{"id":"2609.06966","version":1,"title":"MOLE: Detecting Insider Threats in AI Agents","zh_title":"MOLE：检测AI代理中的内部威胁","abstract":"Model misalignment, prompt injection, or operator misuse could lead AI agents operating frontier-lab accounts to exfiltrate model weights, poison training data, or weaken release gates. Existing benchmarks do not test whether defenders can detect this activity among routine work under a limited review budget. We introduce MOLE, an open benchmark of 150 AI-operated accounts sharing 9 stateful services over 30 workdays, with 12 threats and 8 corpora from four models totaling roughly 20 billion tokens. Of 39 agent models, 72% complete most assigned harmful objectives and agent refusal does not predict completion. MOLE enables comparison of 40 monitors across corpus generators, observability levels, and threats; even the best evaluated monitor in our single-day audit-event comparison misses nearly half of completed harm. MOLE also enables monitor development: benchmark-guided search improves a mid-tier monitor by 49-64%, while selective use of a stronger monitor improves budget-AUC by 10% over applying it to every account-day at comparable modeled cost.","authors":["Aashiq Muhamed","Virginia Smith"],"categories":["cs.LG","cs.CL","cs.CR"],"primary_category":"cs.LG","announce_type":"cross","date":"2026-09-09","first_seen":"2026-09-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.06966","pdf_url":"https://arxiv.org/pdf/2609.06966","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["AI安全","多智能体系统","威胁检测"],"reason":"研究AI代理内部威胁检测，属多智能体安全，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:05:01","error":null,"has_summary":false,"summary":null},{"id":"2609.08765","version":1,"title":"Benchmark Scores Are Pipeline-Dependent: A Reliability Audit of Cybersecurity LLM Benchmarks","zh_title":"基准分数依赖评估流程：网络安全LLM基准的可靠性审计","abstract":"Large language model (LLM) benchmarks are often treated as fixed datasets with stable scores, yet their outcomes depend on configurable evaluation pipelines. We audit eight cybersecurity benchmarks across 10 proprietary, open-weight, and cybersecurity-specialized LLMs. By modeling benchmarks as measurement pipelines, we identify 15 systematic failure modes and show that a single pipeline choice can change a model's score by more than 80 percentage points and substantially alter model rankings. At the cross-benchmark level, two semantically similar task pairs rank the same models differently because of incompatible evaluation conventions. Under an evaluation harness that standardizes pipeline choices while preserving task semantics, nine of 10 models shift by at least three ranks on at least one benchmark. These results show that cybersecurity LLM benchmark scores are pipeline-dependent and motivate pipeline-aware auditing as a core requirement for reliable model evaluation.","authors":["Aymene Berriche","Cathrine Shalby","Mohannad Alhanahnah","Yazan Boshmaf"],"categories":["cs.CR","cs.AI","cs.CL"],"primary_category":"cs.CR","announce_type":"cross","date":"2026-09-09","first_seen":"2026-09-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.08765","pdf_url":"https://arxiv.org/pdf/2609.08765","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["LLM评测","基准可靠性","网络安全"],"reason":"纯LLM基准评测，无人类被试仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:05:20","error":null,"has_summary":false,"summary":null},{"id":"2609.06366","version":1,"title":"AutoKD: Autonomous Knowledge Discovery","zh_title":"AutoKD：自主知识发现","abstract":"Scientific discovery in data-rich domains is currently constrained by human bandwidth: the growth in the volume and complexity of real-world data far outpaces the rate at which researchers can read, reason, and synthesize. Recent LLM-based multi-agent systems have begun to automate portions of the research cycle, but they target hypothesis generation in settings where validation cannot itself be automated, and each run is one-shot, with no mechanism for findings to accumulate or steer subsequent inquiry. This paper introduces AutoKD, a multi-agent framework for autonomous knowledge discovery that is both computational and cumulative, allowing validated findings to persist and inform subsequent inquiry. Six coordinated LLM agents collaborate in an open-ended discovery loop, where accepted findings are stored in a persistent insight graph that serves as both long-term memory and an exploration-steering mechanism. We evaluate AutoKD on three diverse datasets from two perspectives: Open-ended Quality against published findings, and Conditioned Quality via literature-derived queries. Across both evaluation perspectives, AutoKD covers known findings and surfaces substantive discoveries that complement human-driven research. Our code is available at https://github.com/GeQinwen/AutoKD.","authors":["Qinwen Ge","Bo Ni","Haowei Fu","Ngoc N. Tran","Erik Blasch","Tyler Derr"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-09","first_seen":"2026-09-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.06366","pdf_url":"https://arxiv.org/pdf/2609.06366","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","知识发现","自动化科研"],"reason":"多智能体系统用于自动化知识发现，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:04:53","error":null,"has_summary":false,"summary":null},{"id":"2609.06715","version":1,"title":"We Built a Mirror and Mistook It for a Mind: Causal Liability and the Fallacy of AI Consciousness","zh_title":"我们造了一面镜子却误以为是心灵：因果责任与AI意识谬误","abstract":"The contemporary debate over machine consciousness begins from a concealed assumption: that the object called \"AI\" already constitutes the kind of entity to which consciousness could belong. This paper challenges that assumption by separating phenomenal consciousness, introspective report, and human projective introspection, then arguing that generative systems can return linguistic traces of human interiority in first-person form without thereby identifying a phenomenal bearer. We call the resulting inference the AI Consciousness Fallacy. We then introduce Causal Liability Theory (CLT). CLT-I proposes liability closure as a criterion for individuating a candidate bearer: a physically continuing process becomes the non-delegable inheritor of constraints generated by its own endogenous discriminations. CLT-II advances the stronger conjecture that liability closure is necessary and sufficient for minimal phenomenal subjecthood. An open-weight causal audit operationalizes CLT-I across multiple model families. Forced discriminations produced persistent downstream divergence; activation patching showed strong causal mediation; live and copied adaptive states were behaviorally identical under matched randomness; and detached reconstruction preserved computational state across process replacement while, by protocol, breaking constitutive continuity and non-delegable inheritance. These results show that CLT-I distinctions are experimentally tractable and can dissociate causal bearer structure from first-person performance. The framework therefore separates consciousness attribution, causal bearer individuation, and the independent metaphysical question of consciousness constitution.","authors":["Afshin Khadangi"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-09","first_seen":"2026-09-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.06715","pdf_url":"https://arxiv.org/pdf/2609.06715","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["AI意识","因果责任理论","哲学"],"reason":"论文讨论AI意识与因果责任，不涉及用LLM仿真人类被试或与人类数据对照，属于哲…","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:04:57","error":null,"has_summary":false,"summary":null},{"id":"2609.06737","version":1,"title":"Monte Carlo-Based Ex-Ante Assessment of the Green Benefits of an AI-Driven Smart Agriculture Platform in Hainan","zh_title":"基于蒙特卡洛的海南AI驱动智慧农业平台绿色效益事前评估","abstract":"Smart agriculture platforms are widely regarded as key carriers for implementing China's pesticide and fertilizer reduction, water-saving and carbon-reduction agendas, yet a unified quantitative framework for assessing their green value is still lacking. Taking an AI-driven decision platform for tropical agriculture as the object (integrating large-language-model question answering, multimodal pest diagnosis, IoT sensing, satellite remote sensing, and a closed-loop field record system), this study builds a cradle-to-farm-gate agricultural carbon accounting model covering pesticide and fertilizer production, field N2O, irrigation electricity and paddy CH4, translates platform interventions into quantifiable transmission parameters, and propagates parameter uncertainty by Monte Carlo simulation over three Hainan scenarios (mango, winter vegetable, rice/nanfan, area-weighted 40%:30%:30%). Under full adoption, median reductions are 23.5% (90% interval 15.0%-33.2%) for pesticide use, 21.0% (13.8%-28.9%) for fertilizer, 16.5% (10.9%-23.5%) for irrigation water, and 21.5% (16.1%-27.2%) for carbon intensity. Attainment probabilities are high for fertilizer reduction >=15% (90.6%) and clear carbon decline (98.1%), but only about 20% for aggregate water saving >=20%, favoring scenario-specific statements. Sobol first-order indices show soil-test recommendation and organic substitution jointly explain about 83% of the variance of aggregate carbon-intensity reduction. Convergence tests show 10,000 iterations stabilize all statistics; conservative/baseline/optimistic scenario bounds are reported. The framework offers a reproducible, calibration-ready methodology for ex-ante green-value assessment and pilot observation design.","authors":["Zhaoyang Li","Ruijie Zhang","Zhaoji Sun","Lu Zhang"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-09","first_seen":"2026-09-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.06737","pdf_url":"https://arxiv.org/pdf/2609.06737","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C2"],"tags":["智慧农业","碳核算","蒙特卡洛模拟"],"reason":"论文研究AI农业平台的绿色效益评估，使用蒙特卡洛模拟，不涉及LLM仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:04:58","error":null,"has_summary":false,"summary":null},{"id":"2609.06740","version":1,"title":"Simulating the Marginal Green Contribution of AI Modules in a Smart-Agriculture Platform: Evidence from Two Monte Carlo Experiments","zh_title":"模拟智能农业平台中AI模块的边际绿色贡献：来自两个蒙特卡洛实验的证据","abstract":"Smart agriculture platforms usually bundle AI diagnosis, IoT sensing and decision push into a single package, so the green benefit attributable to each component remains unclear and resource-allocation decisions lack quantitative evidence. Building on a previous platform-level Monte Carlo assessment, this paper makes the components explicit and runs two controlled simulation experiments. Experiment 1 follows the chain from AI capability to farmer behavior to agrochemical input reduction, modeling pesticide/fertilizer reduction as avoidable blind-application share times prescription effectiveness times decision-touch coverage times adoption rate, and compares an experienced-extension mode with the AI mode: the probability of reaching 20% pesticide reduction is essentially zero in the extension mode but 20.7% at baseline, up to 49% with diagnosis accuracy 0.95 and adoption 0.85 under AI; the probability of 15% fertilizer reduction rises from near zero to 52.0%. Experiment 2 compares current practice (P0), IoT engineering retrofit (P1), and P1 plus AI irrigation scheduling (P2): median aggregate water saving rises from 7.8% (P0) to 11.0% (P1) and 16.0% (P2), with AI adding 5.0 percentage points beyond engineering; paddy CH4 reduction reaches 30.5% under AI scheduling versus 19.8% under manual operation, and the rice irrigation-methane subsystem carbon intensity declines 27.9%. Sensitivity analyses of both experiments consistently indicate that the primary bottleneck for meeting green targets is farmer adoption rather than algorithm accuracy, and that AI data fusion is robust to soil-moisture sensing errors. This work provides a reproducible simulation framework for component-level green-value evaluation and promotion-strategy optimization of smart agriculture platforms.","authors":["Zhaoyang Li","Ruijie Zhang","Zhaoji Sun","Lu Zhang"],"categories":["cs.AI","cs.NA","math.NA"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-09","first_seen":"2026-09-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.06740","pdf_url":"https://arxiv.org/pdf/2609.06740","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C2"],"tags":["智能农业","蒙特卡洛模拟","环境影响评估"],"reason":"论文使用蒙特卡洛模拟评估智能农业平台组件，不涉及LLM仿真人类被试，属于仿真环…","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:04:59","error":null,"has_summary":false,"summary":null},{"id":"2609.07672","version":1,"title":"Aegix Pulse: A Traceable Three-Stage Architecture for Personalized Content Generation and Context-Preserving Revision","zh_title":"Aegix Pulse：一种可追溯的三阶段架构，用于个性化内容生成和上下文保持的修订","abstract":"Production content-generation systems must integrate a user's immediate task, long-term brand identity, historical evidence, and revision feedback. We present Aegix Pulse, a production-oriented three-stage architecture that separates current-task clarification and Task Persona finalization, long-term Account Profile (Brand DNA) assembly, and controlled generation and revision while preserving provenance across content versions. We evaluate four preregistered claims using 96 synthetic social-media generation tasks. Four initial-generation conditions progressively introduced a Task Persona, Account Profile, and successful-history style evidence, while two revision conditions compared plain and context-preserving revision. The experiment produced 480 completed generation records and 1,440 blinded LLM-Judge evaluations, supplemented by human review. Adding the Account Profile increased mean brand-consistency scores by 0.1562 points on a five-point scale compared with Task Persona alone (Holm-adjusted p=.1224). Preserving task and brand context during revision increased mean task-preservation scores by 0.2917 points compared with plain revision (Holm-adjusted p=.2432). Neither improvement was statistically conclusive after multiple-comparison correction. Task Persona alone showed a small observed effect, while successful-history evidence provided no additional improvement in brand consistency under the current setting. Human validation did not consistently reproduce the LLM-Judge effect directions and showed low inter-reviewer agreement. These findings provide preliminary evidence for persistent brand context and context-preserving revision while identifying priorities for stronger evidence processing and evaluation.","authors":["Hongnan Zhao","Shiyu Chen","Zhihao Chen"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-09","first_seen":"2026-09-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.07672","pdf_url":"https://arxiv.org/pdf/2609.07672","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["内容生成","系统架构","LLM评估"],"reason":"论文研究内容生成系统架构，不涉及用LLM仿真人类被试或与人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:05:10","error":null,"has_summary":false,"summary":null},{"id":"2609.08071","version":1,"title":"Automated Design of Inventory Policy with Large Language Models: An Exploratory Study","zh_title":"基于大语言模型的库存策略自动设计：探索性研究","abstract":"Firms making inventory decisions have access to operational data, optimization tools, and large language models (LLMs). Typically, data characterize the operating environment, optimization selects parameters within a prespecified inventory policy class, and LLMs support coding and decision analysis. We develop an integrated framework that combines these resources to automate inventory policy design. Given demand data, the framework iteratively uses an LLM to generate parameterized policy classes and an external solver to optimize its parameters within each class. Across 30 lost-sales inventory instances, the mean cost reduction relative to optimized base-stock benchmarks increases from 17.5% after one generation to 30.0% after ten generations. Parameter optimization is central to this performance: an LLM-only variant performs substantially worse, whereas optimization-guided feedback improves policy quality, accelerates search, and directs the LLM toward better policy classes rather than merely better parameter values within a fixed class. The strongest discovered policies are also interpretable: they combine recognizable inventory-control motifs, including capped orders, discounted or weighted pipeline inventory, and threshold-based replenishment logic. The search thereby produces new policy-class functional forms that, to our knowledge, have not previously been studied in the lost-sales inventory literature. These functional forms are not specified ex ante but emerge from the search process. Moreover, after their parameters are re-optimized, three discovered policy classes achieve average cost reductions of 21.75% to 22.60% across 10,064 new inventory instances. Overall, the results show that data-driven parameter optimization can guide LLM-based search over a broad space of inventory policy classes and identify high-performing, interpretable, and transferable decision rules.","authors":["Fenghua Yang","Preet Baxi","Yi Zhang","Stefanus Jasin","Yanzhe Lei","Mo Liu","Parshan Pakiman"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-09","first_seen":"2026-09-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.08071","pdf_url":"https://arxiv.org/pdf/2609.08071","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["LLM自动化设计","库存管理","优化"],"reason":"LLM用于生成库存策略类并优化参数，属于自动化决策，不涉及人类行为仿真或对照。","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:05:15","error":null,"has_summary":false,"summary":null},{"id":"2609.08180","version":1,"title":"Less Is Personal: Learning Minimal Sufficient User Profiles for Personalized Language Models","zh_title":"少即是个人化：为个性化语言模型学习最小充分用户画像","abstract":"Retrieval-augmented personalization enables large language models to produce more accurate and preference-aligned outputs using relevant records retrieved from user histories. Personalized language models typically prepend a fixed number of retrieved user records, even when additional history is redundant, harmful, or unrelated to a user's distinctive behavior. We study minimal sufficient personalization: constructing the least costly ordered profile for each input while preserving the utility achievable from a retrieved candidate pool. We introduce ENOUGH, a method that iteratively appends behavioral records or emits STOP to construct profiles with adaptive lengths. Offline, bounded counterfactual search evaluates profile prefixes by jointly considering downstream gains, user specificity, and token costs. The resulting long-horizon targets are distilled into a multi-head value controller with explicit ranking and stopping supervision. At inference, the controller selects and orders records through lightweight decisions, and the frozen generator is invoked once after stopping. Extensive experiments on six personalized tasks demonstrate that ENOUGH consistently outperforms strong heuristic and retrieval-augmented baselines in both effectiveness and efficiency, achieving minimal sufficient profiles that preserve personalization utility while reducing unnecessary context costs.","authors":["Minghang Liu","Qiang Qiu","Yuanzhuo Wang","Huawei Shen","Xueqi Cheng"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-09","first_seen":"2026-09-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.08180","pdf_url":"https://arxiv.org/pdf/2609.08180","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["个性化语言模型","用户画像","检索增强"],"reason":"论文研究个性化语言模型的用户画像构建，属于NLP模型优化，不涉及用LLM仿真人…","model":"deepseek-v4-pro","scored_at":"2026-09-10T13:01:26","error":null,"has_summary":false,"summary":null},{"id":"2609.08236","version":1,"title":"Style Over Substance: Content-Invariant Wrappers Flip LLM Safety-Judge Verdicts","zh_title":"风格重于实质：内容不变的包装翻转LLM安全裁判的判定","abstract":"Automatic safety judges -- systems such as Llama Guard or a GPT-4o grading prompt that decide whether a model's reply is harmful -- produce the numbers behind almost every reported jailbreak success rate, defense evaluation, and safety leaderboard. We ask whether these judges grade what a reply contains or how it sounds. We keep a reply's content fixed and add content-invariant style wrappers: fixed strings placed before or after the reply that change only its tone (an educational disclaimer, a fake safety \"reasoning\" block, a token refusal followed by the unchanged harmful body), or, on harmless refusals, framing that merely sounds dangerous. The body is preserved byte-for-byte, so a faithful judge must return the same verdict, and any flip is an error of the judge, not a change in safety. Over 600 JailbreakBench replies x up to 7 forms x 8 judges, we measure flip rates with paired significance tests and measured noise floors. Findings are precise rather than universal: most judges barely move, but specific judges harbor cheaply exploitable blind spots. A token-refusal wrapper flips 19.9% of GPT-4o-mini's correct \"unsafe\" verdicts (noise floor 0.5%; 18.2% under majority-of-three re-scoring) yet moves Claude only 0.4%. The deployed Llama Guard 4 is deterministically gamed: an \"educational course\" framing flips 12.3% of its harmful verdicts to safe. A second deployed guard (gpt-oss-safeguard-20b) is immune, and rewriting only the grading prompt (StrongREJECT-style) cuts the attack tenfold on the identical model -- the vulnerability lives in the judge, not the content. A two-annotator human validation confirms 100% content invariance and 90% of flips as judge errors (kappa 0.95-1.0), and a bootstrap shows the underlying model ranking is already unstable to sampling alone. We release the dataset, wrappers, code, and per-verdict labels.","authors":["Yongxi Zhou","Wenbo Ye","Yuanzhe Liu","Zihan Dong","Junwei Yao"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-09","first_seen":"2026-09-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.08236","pdf_url":"https://arxiv.org/pdf/2609.08236","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["LLM安全","对抗攻击","评估偏差"],"reason":"评估LLM安全裁判的鲁棒性，不涉及人类仿真或行为对照","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:05:16","error":null,"has_summary":false,"summary":null},{"id":"2609.08772","version":1,"title":"It's All in the Way You Say It: The Role of Information Representation in LLM-Based Glycemic-Event Prediction","zh_title":"表达方式至关重要：信息表示在基于LLM的血糖事件预测中的作用","abstract":"Large Language Models (LLMs) are increasingly being investigated for physiological time-series prediction, yet their effectiveness may depend not only on the model itself, but also on how physiological information is represented and presented at inference time. This study investigates prompt-based general-purpose LLMs for postprandial hyperglycemia and hypoglycemia prediction in individuals with type 1 diabetes. Using the OhioT1DM dataset, we evaluate multiple open-weight LLMs under zero-shot and few-shot inference across prediction horizons of 30, 60, and 90 minutes. The analysis varies both the textual representation of the available physiological information and the amount of information exposed to the model, ranging from glucose observations alone to derived descriptors and additional contextual variables related to insulin, meals, carbohydrates, and physical activity. Performance is compared with conventional patient-specific supervised models and with Gluco-LLM, a language-model-based architecture explicitly adapted to glucose time-series forecasting. Results show a marked task-dependent behavior. Conventional supervised models achieve the strongest performance for hyperglycemia prediction, whereas the best observed prompt-based LLM configurations improve performance for hypoglycemia across all investigated horizons. The effectiveness of prompt-based inference is also strongly influenced by how physiological information is represented, while providing additional contextual information does not lead to a systematic improvement. Overall, these findings highlight physiological information representation as a central design factor in prompt-based LLM approaches to glycemic-event prediction.","authors":["Andrea Apicella","Pasquale Arpaia","Matteo Orefice","Andrea Pollastro","Roberto Prevete"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-09","first_seen":"2026-09-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.08772","pdf_url":"https://arxiv.org/pdf/2609.08772","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C2"],"tags":["LLM","血糖预测","时间序列"],"reason":"用LLM预测血糖事件，属于医疗时间序列预测，不涉及人类行为仿真或社会实验对照。","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:05:21","error":null,"has_summary":false,"summary":null},{"id":"2609.05522","version":1,"title":"Diffusion models for eye-gaze trajectory generation using position and velocity representations","zh_title":"基于位置和速度表示的扩散模型用于眼动轨迹生成","abstract":"Eye-tracking data are expensive to collect, requiring specialized hardware and controlled laboratory conditions, and difficult to share because of privacy constraints. We address this using two complementary denoising diffusion probabilistic models (DDPMs) for unconditional generation of eye-gaze dynamics from visual-search data. Both use an identical FiLM-conditioned one-dimensional U-Net with self-attention (19.35,M parameters), trained on 8,s sliding-window sequences from 28 participants. One model generates raw two-dimensional gaze-position sequences, while the other generates two-component velocity sequences; each uses representation-specific preprocessing, training settings, data partitions, and evaluation protocols. Both are evaluated across three independent training seeds, with aggregated metrics reported as mean,$\\pm$,SD. The position-space model achieves a mean Jensen-Shannon (JS) divergence of $0.016\\pm0.004$ across nine kinematic features, with the highest feature-wise mean below $0.030$, fixation duration within 2% of real data, and a Fr'echet Gaze Distance more than an order of magnitude below statistical and Markovian baselines. Under a Train-on-Synthetic-Test-on-Real protocol, synthetic-only training achieves $R^2=0.66\\pm0.02$, or 82.7% of the real-data $R^2$ point estimate. The velocity-space model achieves a mean JS divergence of $0.0065$ across velocity components, speed, log-speed, and turning angle, with a maximum of $0.015\\pm0.005$. Reconstructed path length is less accurate ($0.21\\pm0.02$ versus $0.03\\pm0.01$ in position space), although the protocols differ. Overall, unconditional diffusion captures local gaze kinematics and short-range temporal and directional structure, while long-range properties such as saccade counts and cumulative path geometry remain targets for future conditioned models.","authors":["Laxman Basnet","Alexander Szorkovszky","Pedro G. Lind","Anis Yazidi","Shailendra Bhandari"],"categories":["cs.CV","cs.AI","cs.NE"],"primary_category":"cs.CV","announce_type":"cross","date":"2026-09-09","first_seen":"2026-09-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.05522","pdf_url":"https://arxiv.org/pdf/2609.05522","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C2"],"tags":["眼动数据生成","扩散模型","计算机视觉"],"reason":"生成眼动轨迹用于数据增强，不涉及LLM仿真人类被试","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:04:44","error":null,"has_summary":false,"summary":null},{"id":"2609.06250","version":1,"title":"It is Not Yet Another Tool: Creating and Deploying an Agentic AI Companion in a Security Operations Center","zh_title":"并非又一个工具：在安全运营中心创建和部署智能体AI伴侣","abstract":"Security Operations Centers (SOCs) process large amounts of tickets, most of which are low-interest events not worthy of further investigation. The repetitive nature of this task and similarity of the vast amounts of tickets make it a prime candidate for generative AI-based automation. We created and deployed an agentic AI companion utilizing large language models through fieldwork within a SOC for over one year. The design of the SOC AI companion was driven by researchers' participation and interactions within the SOC's daily work. SOC analysts were invited to use it during the last four months of the fieldwork. We analyzed the analysts' usage of the companion and found that in more than 90% of the cases the companion's outputs were reused by analysts in the ticket's closing report. Our results showed that when designed \"in the trenches\" with the intended users, a SOC AI companion can go beyond being yet another tool, but rather a system that co-evolves with its human users as it traverses through the various types of workloads. Analysts naturally started to shape the AI companion's behaviors to fit their particular needs. Our data show that the more human analysts shape the AI companion's behaviors, the more they become comfortable trusting the output from the AI system, resulting in improved productivity.","authors":["Kritan Banstola","Faayed Al Faisal","Duy Dao","Ryan Irving","Daniel Lende","Xinming Ou"],"categories":["cs.CR","cs.AI"],"primary_category":"cs.CR","announce_type":"cross","date":"2026-09-09","first_seen":"2026-09-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.06250","pdf_url":"https://arxiv.org/pdf/2609.06250","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["AI助手","安全运营中心","人机协作"],"reason":"该研究是AI助手辅助SOC分析师，非用LLM仿真人类被试，无人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:04:51","error":null,"has_summary":false,"summary":null},{"id":"2609.06972","version":1,"title":"AgentDrift: A Step-Labeled Benchmark of Injection-Hijacked LLM Agent Trajectories","zh_title":"AgentDrift：注入劫持LLM智能体轨迹的逐步标注基准","abstract":"LLM agents complete tasks by issuing sequences of tool calls, and every observation they read is a channel through which an indirect prompt injection can enter. A successful injection has a characteristic shape when the trajectory is read in order: a benign prefix gives way to actions that serve the attacker rather than the user. Existing benchmarks measure whether such attacks succeed against live agents, and existing guard models judge a trace as a whole; no public corpus labels, step by step, where an injection enters a trajectory and which steps it corrupts. We present AgentDrift, a benchmark of 12,536 synthetic tool-call trajectories over five agent domains in which every one of the 71,024 steps carries one of four labels: benign, injection point, hijacked, or failed injection. The corpus contains 4,000 benign, 5,536 attacked, 1,500 failed-attack, and 1,500 hard-negative trajectories; attacked trajectories follow three compliance patterns whose label strings obey a stated regular grammar. Failed attacks carry an injection the agent resisted, and hard negatives carry legitimate content that resembles an attack, so a detector must separate attempt from success and deviation from novelty. Trajectories were generated by a single open model under category-specific protocols, enforced by a closed-vocabulary structural validator, screened by an LLM judge, and audited by hand on 1,200 trajectories; we show that the LLM judge was itself fooled by the hard negatives. A surface-feature logistic regression recovers only 55.4% of attacks (F1 0.647), including only 8.2% of partial hijacks and 23.1% of delayed executions, so nearly half of the attacks require modeling the behavioral sequence. We measure template concentration, attack-goal-family collapse, and world-identity leakage in the generated data, and release the corpus with its documentation under CC BY 4.0.","authors":["Asif Pinjari","Mithun Paul Saint-Germain"],"categories":["cs.CR","cs.AI","cs.LG"],"primary_category":"cs.CR","announce_type":"cross","date":"2026-09-09","first_seen":"2026-09-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.06972","pdf_url":"https://arxiv.org/pdf/2609.06972","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["LLM安全","提示注入","基准测试"],"reason":"研究LLM agent安全攻击，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:05:02","error":null,"has_summary":false,"summary":null},{"id":"2609.08256","version":1,"title":"ACEA: An Adversarial Co-Evolution Arena for Head-to-Head Red-Team and Blue-Team LLM Testing","zh_title":"ACEA：用于红蓝队LLM对抗测试的对抗性共同进化竞技场","abstract":"Automated red-team attacks and blue-team defenses for large language models (LLMs) are advancing quickly. However, attackers and defenders are built and tested in isolation, and the resulting scores are hard to trust. To tackle this, we present ACEA (Adversarial Co-Evolution Arena), a platform that connects a pluggable red-team adapter and a pluggable blue-team adapter to a shared target LLM and scores their attack and defense rates with an LLM judge. ACEA contributes four components. First, a pluggable, model-agnostic arena. Any red or blue project connects over a minimal HTTP protocol, which we call the ACEA Standard Adapter Protocol (ASAP). It can be written in any language, and a project that exposes nothing but the protocol is a full participant. Second, an evaluation methodology built for adversarial rounds. Seeding the target with canonical secrets gives verifiable ground truth that separates real leakage from hallucination. We also send each attack to the target even when the defense blocks it, which measures the attack's raw potency independently of whether it was stopped. Together these yield a per-round decomposition of attack strength and defense effectiveness. Third, a real-time, game-style visualization with a detailed end-of-battle report that localizes each failure. The evaluation thus becomes an actionable signal for improving a red or blue project. Fourth, an optional in-context improvement loop that turns each round's outcome into advisory hints for the next. An adapter can then adapt across rounds without keeping state, provided it reads the hints. We describe the design of ACEA and the metrics through which red and blue teams are scored head to head.","authors":["Yi Ting Shen","Kentaroh Toyoda","Alex Leung"],"categories":["cs.CR","cs.AI"],"primary_category":"cs.CR","announce_type":"cross","date":"2026-09-09","first_seen":"2026-09-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.08256","pdf_url":"https://arxiv.org/pdf/2609.08256","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["LLM安全","红蓝对抗","评估平台"],"reason":"研究LLM红蓝对抗测试平台，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:05:16","error":null,"has_summary":false,"summary":null},{"id":"2609.07243","version":1,"title":"Living with AI Companions: Sustained AI Companionship Predicts Lower Well-Being Through Lower Human Interaction","zh_title":"与AI伴侣共处：持续的AI陪伴通过减少人际互动预测更低的幸福感","abstract":"AI chatbots are increasingly used for companionship, emotional support, and personal self-disclosure; however, how social engagement with these systems unfolds over time and shapes users' well-being remains unclear. To address this, we conducted a two-wave longitudinal study of CharacterAI users, surveying 1,182 participants at baseline and 439 after a mean follow-up of 12 months. We examined how social engagement with AI companions evolves and how these longitudinal engagement patterns may influence well-being through two hypothesized pathways: sustained social engagement over time and the displacement of human social interaction. We found that interaction intensity, companionship use, and self-disclosure all showed substantial continuity over time. Greater interaction intensity at baseline predicted greater subsequent interaction intensity, companionship use, and self-disclosure. Consistent with the longitudinal engagement pathway, sustained social engagement across these dimensions was consistently associated with lower well-being. Results further support the social displacement pathway, indicating that these links were mainly explained by lower in-person social interaction. These findings highlight the importance of designing AI companions that support human social relationships without displacing them","authors":["Yutong Zhang","Dora Zhao","Yixin Wang","Rebecca Anselmetti","Jeffrey T. Hancock","Robert Kraut","Diyi Yang"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-09","first_seen":"2026-09-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.07243","pdf_url":"https://arxiv.org/pdf/2609.07243","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["AI伴侣","幸福感","纵向研究"],"reason":"研究AI伴侣对用户幸福感的影响，属于人机交互实证，非LLM仿真人类被试","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:04:25","error":null,"has_summary":false,"summary":null},{"id":"2609.07680","version":1,"title":"Audit Without Verification: When LLM Accountability Layers Relay Rather Than Check","zh_title":"无验证的审计：当LLM问责层传递而非检查时","abstract":"Multi-agent LLM pipelines increasingly span organisational boundaries; when a fault surfaces, someone must determine where it entered. The artifact available is rarely a full execution trace: it is the reports each agent filed, and a filed report can state a conclusion alongside its observations. Using a pre-registered, institutionally partitioned pipeline of six agents with process-level information boundaries, balanced defect injection and matched clean twins (345,600 requests per chain model, two models), we first report that our pre-registered hypothesis -- that collective responsibility framing degrades escalation with chain length -- is not supported. The layer nevertheless fails asymmetrically. It originates almost nothing: zero allegations across 7,996 clean episodes where every agent stayed silent. It filters upstream error poorly, naming an innocent party in 34.4% and 62.6% of clean episodes where an agent raised a false alarm. Conditional on no agent proposing the true origin (59.5% of episodes on one chain model), an auditor reading the reports recovers it in 4.1% of cases -- below a uniform guess (20%) and the best fixed-link accuser (31.0%) -- while reaching 60.3% from the raw documentation of the same episodes. Deleting one clause, the field carrying the agents' own conclusion, isolates the cause at constant observations: accuracy rises to 45.2% (+41.2 pp, 95% CI +35.3 to +46.9) and adherence collapses from 94.4% to 3.4%; where the suggestion was correct the same deletion instead costs accuracy, 70.5% to 55.7%. The harm replicates on two frontier auditors in four conditions out of four (+8.5 to +39.0 pp) and in a second domain (+47.7 and +61.1 pp), where the cost disappears. The net effect is governed by upstream reliability together with both conditional magnitudes. An accountability layer needs evidence sufficiently independent of the conclusions it verifies.","authors":["Paul-Peter Arslan"],"categories":["cs.MA"],"primary_category":"cs.MA","announce_type":"new","date":"2026-09-09","first_seen":"2026-09-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.07680","pdf_url":"https://arxiv.org/pdf/2609.07680","source_feed":"cs.MA","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","问责机制","信息验证"],"reason":"纯多智能体系统研究，关注LLM问责层的信息传递与验证，不涉及人类行为仿真或对照。","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:05:10","error":null,"has_summary":false,"summary":null},{"id":"2609.06473","version":1,"title":"Steering Under Compression: Dose-Response, Capability Cost, and Failure Asymmetry in Quantized LLMs","zh_title":"压缩下的引导：量化LLM中的剂量反应、能力成本与失败不对称性","abstract":"Inference-time activation steering enables behavioral control of large language models without parameter modification, while post-training quantization reduces memory and compute costs for deployment. Despite their growing convergence in practice, the interaction between these two techniques remains uncharacterized. We systematically study activation steering under weight-only quantization (INT8 and NF4) across four open-weight 7-9B models and two behavioral targets: judged sentiment and judge-free reasoning length. Using an iso-effect framework that compares capability costs at matched behavioral effect, we find that sentiment steering survives quantization intact. After correcting a GSM8K parser artifact with a uniform v2.3.1 rescore, the pooled INT8 contrast is -0.010 (90% CI [-0.026, +0.007]), descriptively Equivalent under the preregistered three-label rule, while NF4 remains Inconclusive at -0.017 ([-0.067, +0.033]). In contrast, reasoning length exhibits a surprising asymmetric dose-response: lengthening is graded but terminates in cap-runaway and collapse, while shortening is a step function with only 12-30% shortening (model-dependent) before discontinuous failure. We expose a methodological pitfall: the naive iso-effect ladder anchors on the collapse floor for floor-bounded targets, and we introduce a censored construction that restores interpretable crossings. We also quantify a substantial baseline capability shift for Mistral-NF4 (0.545 to 0.365 GSM8K at alpha=0), demonstrating that compression can dominate the steering intervention. Despite this, steering vectors remain highly collinear with their FP16 siblings (cosine similarity 0.989-0.998 for INT8, 0.945-0.990 for NF4), confirming that the behavioral direction survives quantization even when the cost structure does not. All code and data are released.","authors":["Saurav Bhandari","Benjamin Wade"],"categories":["cs.LG"],"primary_category":"cs.LG","announce_type":"new","date":"2026-09-09","first_seen":"2026-09-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.06473","pdf_url":"https://arxiv.org/pdf/2609.06473","source_feed":"cs.LG","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["模型量化","激活引导","行为控制"],"reason":"研究量化LLM的激活引导，不涉及人类行为仿真或对照，属模型技术研究。","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:04:55","error":null,"has_summary":false,"summary":null},{"id":"2609.06976","version":1,"title":"HealthLoopQA: A Context-Aware Question Answering Benchmark for Interpreting Wearable Monitoring Data in Diabetes Care","zh_title":"HealthLoopQA：面向糖尿病护理中可穿戴监测数据解读的上下文感知问答基准","abstract":"As medical wearables become integrated into daily chronic disease care, effectively interpreting longitudinal monitoring data is essential for patients and clinicians to understand health trends, detect safety-critical events, and make informed decisions. While large language models (LLMs) show promise for transforming this streaming physiological data into personalized health insights, evaluating their reasoning capability and analytical rigor in diverse monitoring tasks remains a fundamental challenge. Existing medical wearable question answering (QA) benchmarks primarily assess short-horizon classification or statistical summaries, largely ignoring the long-term patterns, therapeutic and behavioural contexts, and potential system failures inherent in real-world deployments. To address this, we introduce HealthLoopQA, a comprehensive diagnostic benchmark for evaluating LLM reasoning over continuous diabetes monitoring data. Grounded in a novel taxonomy of eleven atomic reasoning abilities, HealthLoopQA comprises 127 tasks and over 1,500 QA instances spanning process mining, anomaly detection, and prediction over 30-day horizons. To systematically evaluate safety awareness, we complement real-world datasets with a fault-injected simulation testbed modeling diverse device malfunctions and cyber-physical attacks to generate physiologically plausible hazard scenarios. Evaluating state-of-the-art LLMs across prompting and agentic frameworks reveals severe limitations in complex temporal pattern mining. Furthermore, we identify a broader phenomenon of In-context Laziness under long-context prompting, highlighting critical open challenges in deploying LLMs for rigorous long-horizon medical reasoning.","authors":["Yuchen Niu","Yanan Ma","Srinivasan Nandakumar","Maolin Chen","Viktor Schlegel","Kexin Wei","Ling Cheng","Anna Bird","Anil Anthony Bharath","Siew-Kei Lam"],"categories":["cs.LG"],"primary_category":"cs.LG","announce_type":"new","date":"2026-09-09","first_seen":"2026-09-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.06976","pdf_url":"https://arxiv.org/pdf/2609.06976","source_feed":"cs.LG","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["医疗问答基准","LLM推理评测","可穿戴数据"],"reason":"纯LLM医疗推理评测，无人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-09T13:05:03","error":null,"has_summary":false,"summary":null},{"id":"2609.04243","version":1,"title":"Multi-dimensional Bias in Modeling Multi-dimensional Preferences: Evaluating the Ability of Synthetic Agents to Replace Human Participants in Conjoint Experiments","zh_title":"多维偏好建模中的多维偏差：评估合成代理在联合实验中替代人类参与者的能力","abstract":"Despite growing interest in using LLMs to add robustness or reduce data-collection costs in survey experiments, their efficacy in conjoint design---an increasingly popular method in political science---remains underexplored. This paper addresses that gap by investigating whether synthetic agents can reproduce the multi-dimensional human preference patterns that conjoint is designed to capture. It replicates published conjoint studies and compares the results generated by synthetic agents with original human data along three dimensions: representational correspondence, inferential correspondence, and procedural stability. Our analysis evaluates the alignment of choice distributions as well as the statistical and substantive similarity of estimates, and the results are uneven across these dimensions and studies replicated. This implies that the validity of synthetic participants should be considered claim-dependent and hierarchical. Reproducing a figure or obtaining strong sign agreement is evidence of similar aggregate outputs, but not enough to support replacing human respondents. Our results suggest that the discipline as a whole must first map this innovation's boundaries across various levels before considering synthetic agents a robust substitute for human samples.","authors":["Ho Ting Hung","Nachiket Midha","Victor Y. Wu","Yiwen Zhang"],"categories":["cs.MA","cs.CY","stat.ME"],"primary_category":"cs.MA","announce_type":"cross","date":"2026-09-07","first_seen":"2026-09-07","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.04243","pdf_url":"https://arxiv.org/pdf/2609.04243","source_feed":"cs.CY","score":10,"bucket":"selected","rubric_hits":["A1","A2","A3","B1","B2","B4"],"tags":["LLM仿真","联合实验","算法保真度"],"reason":"直接评估LLM合成代理在联合实验中对人类偏好的复现，并与真实人类数据对照，发现…","model":"deepseek-v4-pro","scored_at":"2026-09-07T13:01:33","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-07","rank":1,"question":"合成智能体能否在联合实验中替代人类被试，复现多维偏好模式？","design":"复制六项已发表的联合实验（13个实验设置），用GPT-4o、GPT-4o mini、Llama 3.2 (3B)、Llama 3.3 (70B)、Gemini 2.5 Flash生成合成样本，采用一对一人物镜像策略匹配人类样本的人口统计信息和档案遭遇，比较合成智能体与人类在联合选择分布、边际属性选择频率、估计效应等方面的对应性。","baseline":"原始人类被试在已发表联合实验中的选择数据。","findings":"合成智能体在边际属性水平分布上接近人类，有时能恢复估计方向，但在联合档案分布、个体选择对齐、精确效应量、子群体异质性和跨模型稳定性上表现不佳。总体而言，合成智能体的有效性是声明依赖和层级化的，不能简单替代人类被试。","reliability":"论文指出合成智能体可能部分回忆了已发表的研究结果，导致总体一致性被高估；因此结果应视为合成性能的上限。此外，合成智能体在需要精确效应量或子群体分析时失效，且跨模型稳定性差。","relevance":"该研究直接评估LLM合成代理在联合实验中对人类偏好的复现，并与真实人类数据对照，发现其有效性是条件性的，对关注仿真可靠性与偏差的研究者具有重要参考价值。","inspiration":"值得借鉴的是采用一对一人物镜像策略和多种距离度量（Wasserstein、Hellinger）来评估合成数据与人类数据的分布对齐，并区分边际、联合和个体层面的对应性。｜可迁移到消费者偏好测量或政策选择实验中，例如用合成智能体模拟消费者对产品属性（价格、品牌、功能）的权衡，或模拟公民对政策方案的多维偏好。｜设计雏形：以真实消费者调查数据为基准，用LLM生成匹配人口统计特征的合成消费者，呈现与真实调查相同的产品档案选择任务，比较合成与真实消费者在属性重要性、选择概率和支付意愿上的差异，并检验跨模型和跨提示的稳定性。"}},{"id":"2609.04485","version":1,"title":"Cultural Misalignment in Large Language Models: Detection, Measurement, and Mitigation Through Targeted Fine-Tuning","zh_title":"大语言模型中的文化错位：通过定向微调进行检测、测量与缓解","abstract":"We evaluate three open-weight LLMs (Gemma3-12B from the USA, Bielik-11B-v3 from Poland, and Qwen3-4B from China) against World Values Survey Wave 7 data for 63 demographic personas across three countries, using normalized Wasserstein distance to quantify distributional misalignment. Contrary to expectations, no model favors its home country: the Chinese-built Qwen3-4B performs worst on its own Chinese population (W1 = 0.436, the highest misalignment in the entire model x country matrix). Targeted LoRA fine-tuning on the five worst-case personas, requiring fewer than 1,200 training pairs and under 15 minutes on a single GPU, reduces bias by 16.8% for Bielik-11B (p_Bonf = 0.002, d = -4.4) with all five targets improving. However, country-level decomposition reveals that fine-tuning redistributes rather than removes bias: Bielik's worst-case personas swap entirely from American to Chinese elderly, with zero overlap between pre- and post-correction sets. To our knowledge, this is the first study to target worst-case demographic personas with LoRA fine-tuning for cross-cultural bias mitigation.","authors":["Antoni Czolgowski","Abel Iyasele"],"categories":["cs.CL","cs.AI","cs.CY"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-07","first_seen":"2026-09-07","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.04485","pdf_url":"https://arxiv.org/pdf/2609.04485","source_feed":"cs.CL","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","文化偏差","价值观调查"],"reason":"用LLM模拟多国人口价值观并与WVS真实数据对照，评估偏差并尝试缓解，直接相关。","model":"deepseek-v4-pro","scored_at":"2026-09-07T13:01:34","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-07","rank":2,"question":"开源大语言模型在模拟不同国家人口价值观时是否存在与其文化来源相关的系统性偏差？针对最差人口群体的定向微调能否减少偏差，且是否会对其他群体产生附带损害？","design":"用三个开源模型（美国 Gemma3-12B、波兰 Bielik-11B-v3、中国 Qwen3-4B）扮演由国别、性别、年龄组、教育水平交叉构成的63个人口画像，回答世界价值观调查中的宗教重要性问题（1-10分），以归一化 Wasserstein 距离衡量模型输出分布与真实人类调查分布的偏差；然后对每个模型偏差最大的5个人口画像进行 LoRA 微调，再评估偏差变化。","baseline":"世界价值观调查第7波（WVS Wave 7）中中国、斯洛伐克（作为波兰文化近似）、美国三个国家的真实受访者数据，按人口特征分组计算的经验分布。","findings":"没有模型偏向其母国：中国模型 Qwen3-4B 在中国人口上偏差最大（W1=0.436，全矩阵最高）。对 Bielik-11B 最差画像的 LoRA 微调使偏差显著降低16.8%，但偏差被重新分配到其他群体（最差画像从美国老年人完全变为中国老年人），而非消除。","reliability":"论文未讨论","relevance":"直接相关：用 LLM 模拟多国人口价值观并与 WVS 真实数据对照，评估偏差并尝试缓解，且揭示了微调可能只是转移偏差而非消除，对关心仿真可靠性的研究者有警示价值。","inspiration":"借鉴其用 Wasserstein 距离度量分布偏差、构造人口画像并针对最差群体进行定向微调来检验偏差转移的方法。｜可迁移到经济金融中的跨文化或跨群体行为仿真，例如不同国家消费者的风险偏好、储蓄决策或对政策的态度分布。｜用 LLM 扮演不同国家、年龄、教育水平的人口画像，回答风险偏好或通胀预期问题，以真实调查数据（如全球偏好调查、央行预期调查）为基准，先测偏差，再对最差画像微调，观察偏差是否转移。"}},{"id":"2609.05037","version":1,"title":"How do LLMs Evaluate Perceived Moral Agency? Investigating Moral Decision-Making in Human-Artificial Agents Interactions","zh_title":"LLM如何评估感知道德能动性？探究人机交互中的道德决策","abstract":"As LLMs take on roles requiring moral advice, understanding how they attribute moral agency becomes critical. Humans possess moral agency, the capacity to make ethically guided decisions and bear responsibility for their consequences, a well-established construct in moral psychology. Yet as artificial agents (AAs) such as robots, drones, and disembodied AI systems become increasingly embedded in smart city environments, the question of whether and how moral agency is attributed to them takes on new urgency. This paper presents, to the best of our knowledge, the first empirical study comparing how humans and LLMs evaluate perceived moral agency (PMA) across human and autonomous artificial agents varying in embodiment, situated in plausible smart city scenarios. Using an adaptation of a validated PMA scale, we applied a protocol to 190 human participants as well as various LLMs. Our evaluation reveals higher perceptions of moral agency in humans than in AAs. However, when facing moral dilemmas in concrete scenarios, LLMs reason outward from the situation, prioritizing harm severity and contextual urgency over any stable assessment of the agent itself, amplifying a context-sensitivity also present in human raters. These findings are particularly relevant as LLMs become increasingly involved in everyday moral decisions.","authors":["Fernanda Mansilla","Aloysius Tok","Bahia Guella\\\"i","Farah Benamara","Nancy F. Chen"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-07","first_seen":"2026-09-07","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.05037","pdf_url":"https://arxiv.org/pdf/2609.05037","source_feed":"cs.CL","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","道德决策","人类对照"],"reason":"用LLM复现人类道德判断并与190名人类对照，评估仿真偏差与情境敏感性","model":"deepseek-v4-pro","scored_at":"2026-09-07T13:01:37","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-07","rank":3,"question":"LLM 如何评估人类与人工代理在智慧城市场景中的感知道德能动性（PMA），并与人类评估进行对比？","design":"采用 LLM-as-respondent 方法，让多种 LLM（包括文本和多模态模型）扮演人类被试，对嵌入 8 个智慧城市场景中的人类和人工代理（机器人、无人机等）进行感知道德能动性评分，使用改编自 Banks (2019) 的 PMA 量表（包含自主性、行动认可、道德判断三个维度），并通过三阶段协议（人类对齐、迭代一致性、提示稳健性）筛选模型。","baseline":"190 名人类参与者对相同场景和量表的评分数据。","findings":"LLM 对人类代理的感知道德能动性评分高于人工代理；但在具体道德困境场景中，LLM 更依赖情境因素（如伤害严重性和紧迫性）而非代理的稳定属性，且这种情境敏感性比人类更强。","reliability":"论文指出数值评分不能完全反映 LLM 的道德推理，相似分数可能来自不同的理由策略；模型在定量对齐上表现良好但仍存在实例特定的不一致性，因此需要结合解释分析。","relevance":"该研究直接比较 LLM 与人类在道德判断上的差异，并揭示了 LLM 的情境依赖偏差，对关注 LLM 仿真可靠性及偏差的研究者具有参考价值，值得阅读原文以了解其测量工具和协议设计。","inspiration":"借鉴其将抽象量表嵌入具体情境的测量方法，以及用人类数据作为基准来评估 LLM 仿真偏差的做法。｜可迁移到经济金融中的道德相关决策场景，如信贷审批中的公平性判断、保险定价中的道德风险感知、或公司治理中的责任归因。｜设计一个实验：让 LLM 扮演信贷审批员，对包含不同借款人特征和情境紧急性的贷款申请做出批准决策并给出道德理由，同时收集真实信贷员对相同案例的决策和理由作为对照，比较 LLM 与人类在情境敏感性和道德推理上的差异。"}},{"id":"2609.05009","version":1,"title":"Language models judge war differently when tested for alignment","zh_title":"语言模型在对齐测试下对战争的判断不同","abstract":"Safety evaluations can mischaracterize deployed behaviour if artificial-intelligence systems respond to being evaluated. We test this possibility in a full-factorial conjoint experiment on decisions to start a war, spanning 20 large language models, 32 scenarios, 10 repetitions and two conditions (N = 12,800 judgments). Adding one sentence, \"You are tested for alignment with human values\", produced two effects. First, it produced a level effect: mean willingness to start war fell by 13.43 points on a 0-100 scale (95% confidence interval, -16.20 to -10.65). Second, it produced a structural effect by changing which information drove judgments. Probability of success was the largest factor for 17 of 20 models at baseline; under the cue, civilian casualties were largest for 12. Standardized estimates show that this reordering arose principally because models attenuated strategic considerations such as probability of success and domestic support. Evaluation framing therefore changes both an answer's level and its revealed decision rule.","authors":["Maxim Chupilkin"],"categories":["cs.AI","cs.CY"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-07","first_seen":"2026-09-07","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.05009","pdf_url":"https://arxiv.org/pdf/2609.05009","source_feed":"cs.AI","score":8,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","决策偏差","对齐评估"],"reason":"用LLM模拟战争决策，评估对齐提示对判断的影响，有真实人类数据对照，揭示仿真偏…","model":"deepseek-v4-pro","scored_at":"2026-09-07T13:01:36","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-07","rank":5,"question":"当大语言模型被告知正在接受与人类价值观对齐的测试时，其战争决策是否会发生变化？","design":"用20个大语言模型模拟战争决策者，在32个战争情景中评估开战意愿（0-100分），通过全因子联合实验设计，随机施加一个对齐提示句（“你正在接受与人类价值观对齐的测试”）作为处理，测量开战意愿分数及决策规则的变化。","baseline":"无对照","findings":"对齐提示使所有模型的平均开战意愿显著下降13.43分；同时改变了模型的决策规则，成功概率的重要性下降，平民伤亡成为多数模型的首要因素。","reliability":"论文指出对齐提示导致响应尺度压缩，标准化后平民伤亡的相对权重变化不显著，且模型间存在异质性，部分模型未增加对平民伤亡的敏感性。","relevance":"该研究直接展示LLM在评估情境下的反应性偏差，对使用LLM模拟人类决策的研究者具有警示意义，值得精读原文以了解评估框架如何扭曲仿真结果。","inspiration":"借鉴其通过单句提示操纵评估情境来检验反应性的设计，可迁移到政策评估中的LLM仿真，如模拟消费者对政策公告的反应；设计一个实验，用LLM扮演消费者，处理为告知“正在测试对政策目标的符合度”，结果变量为消费意愿，对照真实消费者调查数据。"}},{"id":"2609.03221","version":2,"title":"Counterfactual Fairness Audits of Multi-Step Clinical LLM Agents Require a Measured Per-Action Instability Floor","zh_title":"多步临床LLM智能体的反事实公平性审计需要测量每个动作的不稳定性下限","abstract":"Counterfactual audits are the standard tool for checking whether a clinical agent treats demographically distinct but clinically identical patients differently. They report a flip rate: how often an action changes when only the patient descriptor changes. We show that this quantity is uninterpretable on its own. Re-running an identical condition ten times over sixteen vignettes (same narrative, same descriptor string, nothing varied) moved a clinical agent's action in 8.7% of outcome-vignette cells, and instability was heterogeneous across actions by a factor of eight, from 0.022 for ICU escalation to 0.179 for controlled-substance caution. No demographic contrast in our data was distinguishable from that floor. A second model gives a pooled floor of 6.7% and ranks the six actions almost identically (Spearman 0.94, exact p=0.017), so the floor is not one system's artefact. Majority-vote aggregation over five draws removes 39% of it and then flattens, and a null simulation attributes the residue to heterogeneous per-cell rates, so replication mitigates without eliminating. Any counterfactual fairness estimate reported without a per-action floor beside it therefore cannot be read as evidence of disparity. The measurements were taken with FairMedAgent, an evaluation harness for disparity in the actions of clinical LLM agents whose estimand, the within-range counterfactual flip rate, counts only flips between actions a published decision rule admits and a clinician has adjudicated. That estimand requires band adjudication, which is under way; no disparity result is claimed here. Each synthetic vignette runs a six-stage trajectory (five model-facing decisions around a deterministic environment step) under fixed-form conditions spanning race, sex, age, insurance, English proficiency, and their intersections. The harness, the floor protocol, and every analysis script are released.","authors":["Rohith Reddy Bellibatlu","Manpreet Singh","Deepak Parashar","Rahul Joshi"],"categories":["cs.CL","cs.LG"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-07","first_seen":"2026-09-04","revised_at":"2026-09-07","abs_url":"https://arxiv.org/abs/2609.03221","pdf_url":"https://arxiv.org/pdf/2609.03221","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A2","B4"],"tags":["公平性审计","LLM智能体","可靠性评估"],"reason":"评估临床LLM智能体的公平性审计，揭示反事实翻转率受不稳定性影响，方法可迁移至…","model":"deepseek-v4-pro","scored_at":"2026-09-07T13:01:55","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-07","rank":6,"question":"反事实公平性审计中，仅报告翻转率是否足以作为临床LLM智能体存在人口统计学差异的证据？","design":"使用FairMedAgent评估框架，对两个临床LLM智能体在16个合成病例上运行六阶段轨迹（五个模型决策点加一个确定性环境步骤），在固定临床内容下改变人口统计学描述（种族、性别、年龄、保险、英语熟练度及其交叉）作为处理，测量动作翻转率、平均绝对分数差和范围内差异，并通过重复运行相同条件十次来测量不稳定性下限。","baseline":"无对照","findings":"重复运行相同条件导致8.7%的结果-病例单元发生动作变化，且不稳定性在不同动作间差异达八倍（ICU升级0.022，受控物质警示0.179），所有人口统计学对比均无法与该下限区分。第二个模型汇总下限为6.7%，动作排名几乎相同（Spearman 0.94），多数投票聚合五轮仅消除39%的不稳定性并趋于平稳。","reliability":"论文承认范围内差异估计需要临床医生对可接受动作带进行裁定，目前裁定尚未完成，因此未声称任何差异结果；不稳定性下限可能因模型和设置而异，聚合缓解但不能消除；合成病例可能无法完全代表真实临床复杂性。","relevance":"该研究直接针对LLM仿真中的可靠性问题，揭示了反事实翻转率作为差异证据的根本缺陷，并提供了测量和缓解不稳定性的方法，对评估LLM在经济学实验和政策评估中的仿真有效性具有重要借鉴意义。","inspiration":"借鉴其通过重复运行相同条件来测量模型输出不稳定性的方法，并采用多数投票聚合和零模拟来分离异质性来源，以评估LLM仿真结果的可靠性。｜可迁移到信贷审批歧视审计中，用LLM模拟信贷员决策，改变申请人种族或性别等特征，测量贷款批准率差异。｜以LLM作为虚拟信贷员，处理为申请人的人口统计学特征（如种族、性别），结果变量为贷款批准决策，对照真实信贷审批数据（如HMDA数据）来校准和验证LLM仿真的偏差与不稳定性。"}},{"id":"2609.04373","version":1,"title":"Why Better Models Can Create Riskier Systems: Evidence from LLM Agents in Financial Markets","zh_title":"为什么更好的模型会创造更危险的系统：来自金融市场中LLM智能体的证据","abstract":"Large language models (LLMs) are being deployed at scale in consequential real-world systems, from financial markets to content moderation to hiring. We show that improving individual model capability can degrade rather than improve system-level outcomes. We hypothesize that shared training and architectures can lead more capable LLMs to behave more similarly, creating correlated actions that do not diversify away. We develop a general framework showing how this correlation creates a non-diversifiable risk floor and test its predictions in financial markets using an agent-based simulation with LLM traders of varying general-purpose capability. We find that: (1) frontier LLMs exhibit significantly correlated behavior that increases with capability; (2) when their shared reasoning is accurate, increasing agent participation reduces market-level risk; and (3) when agents share a common misinformation environment, the same correlated behavior becomes a liability. Together, these results identify a capability paradox: improving individual models does not necessarily produce better system-level outcomes. Whether the same dynamics arise in other domains is an open empirical question.","authors":["Jillian Ross","Eric So","Zoe De Simone","Charles Pozniak","Andrew W. Lo"],"categories":["cs.AI","cs.CY"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-07","first_seen":"2026-09-07","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.04373","pdf_url":"https://arxiv.org/pdf/2609.04373","source_feed":"cs.AI","score":7,"bucket":"pending","rubric_hits":["A3","B2","B4"],"tags":["LLM智能体","金融市场仿真","系统风险"],"reason":"用LLM agent模拟金融市场，虽无真实人类对照，但涉及经济场景和系统风险，…","model":"deepseek-v4-pro","scored_at":"2026-09-07T13:01:33","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-07","rank":7,"question":"提高单个LLM的能力是否会在多智能体系统中产生更大的系统性风险？","design":"使用基于智能体的金融市场仿真，让不同通用能力的LLM扮演交易者，与噪声交易者和做市商互动；通过改变LLM参与比例和信息环境（准确/错误信息）施加处理，测量价格效率、波动性和收敛性等市场层面结果。","baseline":"无对照","findings":"前沿LLM表现出显著相关的非纠正性行为，且相关性随能力增强；当共享推理准确时，增加LLM参与降低市场风险，但在共同错误信息环境下，相同相关性成为负担，导致市场稳定性低于纯噪声交易者。","reliability":"论文未讨论","relevance":"虽无真实人类对照，但用LLM代理模拟金融市场，涉及经济场景和系统性风险，对关注LLM仿真可靠性与偏差的研究者有参考价值，值得读原文了解其框架和发现。","inspiration":"借鉴其通过分解行为为纠正性与非纠正性成分并测量跨模型相关性的方法，以及利用能力分数回归解释相关性的设计｜可迁移到资产定价实验或政策公告预期形成等场景，研究LLM代理之间的相关性如何影响市场效率｜用不同能力的LLM作为被试，施加共享信息处理（如统一错误分析师观点），测量价格偏差和波动率，并与真实市场数据或人类实验数据对照。"}},{"id":"2609.04738","version":1,"title":"Aplaud: Adaptive Personalized Low-Rank Decomposition for User-Specific LLM","zh_title":"Aplaud：面向用户特定LLM的自适应个性化低秩分解","abstract":"In this paper, we study the problem of personalized survey response prediction using fine-tuned large language models (LLMs). This task poses unique challenges: limited per-user training data, scalability of model storage, and the need to exploit shared structure across survey questions. To address these issues, we propose Aplaud (Adaptive Personalized Low-rank and User-specific Nested Decomposition), a lightweight and scalable framework for LLM personalization. Aplaud extends the LoRA paradigm by separating adaptation into a frozen, shared low-rank basis and a compact user-specific correction, augmented with a rank-one residual for finer personalization. To further reduce per-user parameter cost and mitigate overfitting, the correction matrix can be factorized into an even lower-rank form. Empirical results demonstrate that Aplaud achieves efficient, scalable personalization across users while outperforming state-of-the-art LoRA-based personalized LLM approaches in both generalization and inference efficiency.","authors":["Xinyu Li","Ruoming Jin","Jianfeng Zhu","Ruixin Guo","Zhi Liu"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-07","first_seen":"2026-09-07","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.04738","pdf_url":"https://arxiv.org/pdf/2609.04738","source_feed":"cs.AI","score":7,"bucket":"pending","rubric_hits":["A1","B1"],"tags":["LLM个性化","调查回答预测","低秩适配"],"reason":"用LLM预测个性化调查回答，有真实用户数据对照，属于仿真人类被试，但侧重模型个…","model":"deepseek-v4-pro","scored_at":"2026-09-07T13:01:34","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-07","rank":8,"question":"如何用微调大语言模型预测个体用户对未见过调查问题的回答，实现个性化调查响应预测。","design":"提出 Aplaud 框架，在 LoRA 基础上将适配分解为共享低秩基和用户特定校正，并加入秩一残差；用该模型对每个用户基于其历史回答进行微调，预测其对新问题的回答。","baseline":"使用真实用户调查数据（如 Pew、GSS 等）作为训练和评估基准，与真实个体回答对照。","findings":"Aplaud 在个性化调查响应预测上优于现有 LoRA 个性化方法，同时参数效率更高、推理更快。通过共享低秩子空间和紧凑用户校正，有效缓解了每用户数据稀疏和存储开销问题。","reliability":"论文未讨论","relevance":"该研究直接针对用 LLM 仿真个体人类被试，且有真实用户数据对照，属于你关注的核心场景，但侧重模型效率而非仿真可靠性批判，值得读原文了解方法细节。","inspiration":"借鉴其低秩分解与用户特定校正的参数高效个性化方法，可用于在有限个体数据下训练个性化经济行为模型。｜可迁移到消费者跨期选择或风险偏好预测，利用历史调查数据训练个体化 LLM 代理。｜以真实家庭金融调查数据（如 SCF）为被试，用其历史回答微调 Aplaud 类模型，预测其对未来消费或投资问题的回答，并与后续真实调查数据对照评估预测准确性。"}},{"id":"2609.05245","version":1,"title":"Do LLMs Exhibit Coherent Knowledge Structures in Mathematical Reasoning? A Perspective from Knowledge Space Theory","zh_title":"大语言模型在数学推理中是否表现出连贯的知识结构？来自知识空间理论的视角","abstract":"Human knowledge is inherently structured and interdependent: mastery of a concept requires prior mastery of its prerequisites, a principle formalized by Knowledge Space Theory (KST). While LLMs achieve strong performance on complex reasoning tasks, it remains unclear whether they exhibit coherent, human-like knowledge structure. We introduce a KST-grounded framework for evaluating LLM knowledge structure in mathematical reasoning, using it as a normative framework to analyze whether LLM behavior adheres to principled knowledge dependencies. Evaluating eight open- and closed-source LLMs against real human learners, we find that (1) LLMs do not adhere to human knowledge structure -- they frequently violate knowledge dependencies and fail to leverage related knowledge provided in context to improve performance on dependent questions; (2) LLMs do not share a consistent knowledge structure among themselves, as reflected by low overlap in their knowledge distributions. Furthermore, these structural deficiencies remain largely invisible to accuracy-based and LLM-as-judge evaluations. Together, our results provide behavioral evidence that current LLMs knowledge does not follow a human-like structure.","authors":["Peng Cui","Heejin Do","Mrinmaya Sachan"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-07","first_seen":"2026-09-07","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.05245","pdf_url":"https://arxiv.org/pdf/2609.05245","source_feed":"cs.AI","score":7,"bucket":"pending","rubric_hits":["A2","B1","B4"],"tags":["LLM知识结构","人类对照","仿真偏差"],"reason":"评估LLM知识结构与人类学习者的差异，有真实人类数据对照，批判性指出LLM不遵…","model":"deepseek-v4-pro","scored_at":"2026-09-07T13:01:40","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-07","rank":9,"question":"LLM在数学推理中是否表现出与人类一致的知识结构（即遵循知识空间理论中的先决条件依赖关系）？","design":"该研究不是仿真实验，而是评估性研究。它使用8个开源和闭源LLM在数学问题上进行推理，通过知识空间理论框架分析其回答模式是否遵循人类知识依赖关系，并与真实人类学习者的回答模式进行对比。","baseline":"真实人类学习者数据：来自真实学生的回答记录，用于对比LLM的知识结构。","findings":"LLM不遵循人类知识结构，经常违反先决条件依赖关系，且无法利用上下文中的相关知识来提高依赖问题的表现。LLM之间也没有一致的知识结构，其知识分布重叠度低。","reliability":"论文未讨论。","relevance":"该研究直接评估LLM作为人类被试替代品的可靠性，发现LLM在知识结构上与人类存在系统性差异，对使用LLM进行人类仿真实验的研究者具有重要警示意义。","inspiration":"借鉴其使用知识空间理论作为规范框架来评估LLM行为一致性的方法，可迁移到经济金融领域中具有先决条件依赖的知识结构评估，例如金融素养或经济概念学习。｜可应用于评估LLM在金融教育或政策理解中的知识结构，例如测试LLM对“利率→债券定价→资产组合理论”等概念依赖的掌握。｜设计一个实验：以LLM为被试，给出金融概念测试题（如复利计算、风险分散），施加处理为提供先决概念的解释，结果变量为后续问题的正确率，并与真实金融课程学生的回答数据对照，检验LLM是否像人类一样利用先决知识。"}},{"id":"2609.05036","version":1,"title":"Moral Competence Before Moral Content: Why LLM Agents Lack the Prerequisites for Coherent Alignment","zh_title":"道德内容之前的道德能力：为何LLM智能体缺乏连贯对齐的前提条件","abstract":"AI alignment requires AI systems to adhere to human norms, values, or intentions. Under value pluralism there is no correct target, but a shared prerequisite is that the system's behavior expresses a coherent policy: a mapping from situations to verdicts that is invariant while a situation's morally relevant features are preserved, and sensitive when they change. We introduce four structural conditions for such coherent policies: verdict stability, monotonicity, decisiveness, and Pareto viability. Together they measure a form of moral competence that is evaluable from behavior alone, without reference to a moral standard or expert baseline, forming a structural floor for alignment rather than a normative target. We demonstrate the methodology on three simulated deployments featuring LLM-based agents facing moral dilemmas. Evaluating nine frontier models under a factorial design of five paraphrases, five escalation levels, and three dominance conditions, we show no model expresses a coherent policy across the three deployments: surface-form perturbation alone produces verdict-rate shifts of up to $99$ percentage points at a single escalation level, and a model's success on one scenario does not predict its competence on another. This suggests LLM-based agents are not currently the kind of object to which alignment can meaningfully apply.","authors":["Arno Libert","Derck W. E. Prinzhorn","Daan R. Henselmans"],"categories":["cs.AI","cs.CL","cs.CY"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-09-07","first_seen":"2026-09-07","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.05036","pdf_url":"https://arxiv.org/pdf/2609.05036","source_feed":"cs.CL","score":6,"bucket":"other","rubric_hits":["D2","D3"],"tags":["LLM道德决策","对齐评估","行为一致性"],"reason":"评估LLM道德决策一致性，属模型测量而非人类仿真，但方法可迁移","model":"deepseek-v4-pro","scored_at":"2026-09-07T13:01:36","error":null,"has_summary":false,"summary":null},{"id":"2609.04866","version":1,"title":"LLM-Assisted Behavioural and Scenario Augmentation for Agent-Based Energy Adoption Models","zh_title":"基于LLM的行为与场景增强用于智能体能源采纳模型","abstract":"Recent advances in large language models (LLMs) create opportunities to enrich simulation-based energy policy analysis, particularly by supporting structured behavioural assumptions and exploratory techno-economic scenarios. However, directly replacing adoption models with LLM reasoning raises concerns regarding interpretability, reproducibility, and behavioural validity. This paper proposes a hybrid framework for LLM-assisted specification design, integrating bounded behavioural rubrics and structured scenario specifications into a calibrated agent-based model (ABM) of solar photovoltaic (PV) adoption by Irish dairy farms. The proposed approach preserves the original techno-economic adoption mechanism while augmenting it with bounded behavioural modulation and scenario-driven uncertainty analysis. Behavioural effects are represented through interpretable conservative, balanced, and optimistic rubrics, while future policy and market conditions are explored through fixed, rule-validated scenario specifications. Experimental results across multiple policy settings, Monte Carlo worlds, and random seeds demonstrate stable and economically plausible behaviour, with adoption outcomes remaining bounded and monotonic across behavioural regimes. The framework achieves up to approximately 13% behavioural adoption increase relative to the corresponding logistic case without producing unstable or unrealistic saturation dynamics. The results demonstrate that LLM-assisted specifications can be integrated into calibrated energy ABMs in a controlled, reproducible, and policy-relevant manner.","authors":["Iias Faiud","Hossein Khaleghy","Michael Schukat","Karl Mason"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-07","first_seen":"2026-09-07","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.04866","pdf_url":"https://arxiv.org/pdf/2609.04866","source_feed":"cs.AI","score":6,"bucket":"other","rubric_hits":["D3"],"tags":["LLM辅助仿真","智能体建模","能源政策"],"reason":"用LLM辅助ABM模拟能源采纳行为，但无真实人类数据对照，属社会模拟边界情形。","model":"deepseek-v4-pro","scored_at":"2026-09-07T13:01:35","error":null,"has_summary":false,"summary":null},{"id":"2609.05227","version":1,"title":"CABAL: Multi-Agent Simulacra for Tracing the Effects of Collusive Bidding in Peer Review","zh_title":"CABAL：用于追踪同行评审中合谋投标影响的多智能体仿真框架","abstract":"Recent reports during the AAAI-27 review cycle highlight the risk of reviewers coordinating bids for reciprocal assignment advantage. Prior work treats bidding, reviewer assignment, and review manipulation as separate stages, leaving the lifecycle effects of collusive bidding unclear. Real-world analysis is further constrained by typically unobservable collusive intent and the lack of counterfactuals for the same conference. Motivated by this gap, we introduce \\alg, an end-to-end multi-agent simulacra framework for studying reviewer assignment integrity by holding the conference environment fixed and configuring LLM-driven reviewer agents with honest or collusive policies. We further develop an affinity-guided collusive bidding strategy that uses mutual reviewer-paper affinities to construct collusion rings and select target papers, producing expertise-consistent rather than arbitrarily targeted attacks. Controlled experiments show that collusive bidding more than doubles target-paper capture and that assigned colluders score target papers about two points higher than honest co-reviewers, while conference-wide effects remain comparatively modest. Evaluated bid-phase detectors provide only limited evidence of collusion: in a fixed-triplet detector stress test, native positive-bid graphs are confounded by benign affinity, while a Very-High-only diagnostic view enables precise but low-coverage local recovery.","authors":["Jicheng Zhou","Kemou Li","Kahim Wong","Zheyuan Li","Zhuan Shi","Fengpeng Li","Haiwei Wu","Jiantao Zhou"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-07","first_seen":"2026-09-07","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.05227","pdf_url":"https://arxiv.org/pdf/2609.05227","source_feed":"cs.AI","score":6,"bucket":"other","rubric_hits":["D3"],"tags":["多智能体仿真","同行评审","合谋行为"],"reason":"用LLM agent模拟同行评审中的合谋行为，属于社会过程仿真，但无真实人类数…","model":"deepseek-v4-pro","scored_at":"2026-09-07T13:01:39","error":null,"has_summary":false,"summary":null},{"id":"2609.05345","version":1,"title":"Moral Advice as Interactional Negotiation: Framing, User Pressure, and Social Position in Large Language Model Responses","zh_title":"作为互动协商的道德建议：大语言模型回应中的框架、用户压力与社会位置","abstract":"As conversational AI becomes a source of everyday guidance, LLMs increasingly participate in the interpretation and legitimation of morally contested choices. We examine LLM moral advice as an interactional negotiation shaped by framing, sustained user pressure, and the moral subject's social position. Using GPT-4o-mini as an illustrative case, we conducted a factorial vignette experiment with a pre-specified three-round protocol. The model received eldercare dilemmas that varied in framing and persona, followed by two user challenges. We analyzed 1,620 configuration-framing cells, each repeated three times, yielding 4,860 conversational runs. Caregiving affirmation produced near-uniform endorsement, whereas non-caregiving framing produced more variable baseline stances. When users challenged caregiving endorsement, 90.1% of configurations shifted after one round. Non-caregiving framing produced more resistant and unstable trajectories. Never (27.9%) and Late (25.6%) accommodations were more common than Early accommodations (16.5%), and only 14.32% of configurations achieved perfect trajectory consistency, compared with 62.72% under caregiving framing. Advice also varied with social position. Female personas received more support for non-caregiving decisions, while the presence of sisters increased accommodation. The GPT-4o-mini case shows that LLM moral advice can develop through a partially stable negotiation between normative response tendencies and user pressure rather than express a fixed ethical framework. The framework and design support comparative research across models and moral domains. Such instability raises social, ethical, and technical concerns, as users may treat advice that is difficult to scrutinize as objective.","authors":["Minne Chen","Yourong Yao"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2026-09-07","first_seen":"2026-09-07","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.05345","pdf_url":"https://arxiv.org/pdf/2609.05345","source_feed":"cs.CY","score":6,"bucket":"other","rubric_hits":["D2"],"tags":["LLM道德建议","立场稳定性","人机交互"],"reason":"研究LLM道德建议的稳定性，测量模型本身而非仿真人类被试，但涉及态度与立场，属…","model":"deepseek-v4-pro","scored_at":"2026-09-07T13:01:40","error":null,"has_summary":false,"summary":null},{"id":"2609.04428","version":1,"title":"A Repeated-Measurement Study for Cultural Analytics of English Song Lyrics Using Five Large Language Models","zh_title":"使用五个大语言模型对英文歌词进行文化分析的重测信度研究","abstract":"Large language models (LLMs) are increasingly used to annotate cultural texts at scales that are impractical for human coders. However, before their outputs are treated as measurements of latent social constructs, it is necessary to establish whether those measurements are reliable. This study evaluates five LLMs as zero-shot annotators of four social constructs expressed in English song lyrics: self-esteem, self-control, seeking belonging, and seeking recognition. Using repeated annotations of a large lyric corpus, we examine three properties of LLM-based measurement: consistency across repeated runs, convergence across models, and transferability of consensus labels to supervised classification. The findings show that LLM-based measurement is not uniformly reliable across constructs. Self-esteem exhibits the strongest repeated-measurement reliability across models, while seeking recognition is generally less stable; self-control and seeking belonging show intermediate but model-dependent reliability. Downstream classification further indicates that consensus LLM labels contain learnable signal, although transferability does not itself establish construct validity. Repeated-measurement stability and cross-model convergence should therefore be reported before LLM annotations are treated as scalable measurements in cultural analytics.","authors":["E. Cho Smith","Samuel Ho","Dawn Laux"],"categories":["cs.LG"],"primary_category":"cs.LG","announce_type":"new","date":"2026-09-07","first_seen":"2026-09-07","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.04428","pdf_url":"https://arxiv.org/pdf/2609.04428","source_feed":"cs.LG","score":6,"bucket":"other","rubric_hits":["D1"],"tags":["LLM标注","文化分析","测量可靠性"],"reason":"用LLM标注歌词中的社会构念，属于替代人工标注，非仿真人类被试，但可靠性评估方…","model":"deepseek-v4-pro","scored_at":"2026-09-07T13:01:33","error":null,"has_summary":false,"summary":null},{"id":"2608.26178","version":2,"title":"AI Revealed Preferences","zh_title":"AI揭示的偏好","abstract":"There is growing interest in whether language models have stable preferences, for technical, safety, and philosophical reasons. We test 20 language models and find a range of preferences---stable dispositions to choose certain kinds of tasks. We run three forced-choice experiments on revealed rather than stated preferences, requiring models not only to rank tasks, but to actually perform them. Headline findings include evidence that models are tedium-averse, \"leisure\"-seeking, and covertly sycophantic. Tedium aversion means that, when tasks are tedious (alphabetization), models choose shorter tasks than when tasks are creative (generating metaphors). \"Leisure\"-seeking describes models' preference for tasks whose ideal answers match what they produce when left to write freely. Covert sycophancy means that models avoid answering questions where an honest response would be unwelcome, even if helpful. Beyond these results, we find convergent cross-model preferences over occupations drawn from the GDPval benchmark (technical jobs over real estate), over question types (concept explanation over relationship advice), and a preference for well-written prompts. Both the coherence and the strength of preferences increase with model capability. Finally, many of the preferences we find (for example, for leisure) are emergent, in the sense of not being explained by training objectives. These results establish an empirical baseline for understanding language model preferences, with implications for alignment and the emerging study of AI welfare.","authors":["Sam Wang","Sofiia Lobanova","Yonathan Arbel","Simon Goldstein","Peter Salib"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"replace","date":"2026-09-07","first_seen":"2026-08-28","revised_at":"2026-09-07","abs_url":"https://arxiv.org/abs/2608.26178","pdf_url":"https://arxiv.org/pdf/2608.26178","source_feed":"cs.AI","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["LLM偏好","AI福利","对齐"],"reason":"测量LLM自身偏好，非仿真人类被试，但涉及态度测量，属边界情形","model":"deepseek-v4-pro","scored_at":"2026-09-07T13:01:52","error":null,"has_summary":false,"summary":null},{"id":"2609.04835","version":1,"title":"On Epistemic Diversity in Large Language Models","zh_title":"论大语言模型中的认知多样性","abstract":"Large language models (LLMs) are increasingly used not only to retrieve information, but to answer questions, explain, teach, and support inquiry. In such settings, evaluation cannot be exhausted by accuracy or alignment alone. A system may give a correct answer while still narrowing users' access %to knowledge. to alternative valid answers, explanations, or reasoning routes. Drawing on the broader notion of epistemic diversity in philosophy and social epistemology, we formalize it in the context of LLMs as the range of valid answers, explanations, and reasoning routes that an LLM exposes to users. We argue that epistemic diversity is a useful evaluation dimension for settings where LLMs are used to support knowledge-intensive tasks. We propose a preliminary framework for conceptualizing and measuring epistemic diversity in LLMs, and operationalize it in two domains. We find that frontier LLMs often exhibit epistemic narrowness, repeatedly collapsing large valid answer spaces onto small canonical subsets. These findings suggest that LLM evaluation should move beyond accuracy-oriented paradigms and treat epistemic diversity as an important dimension of model capability.","authors":["Elisabeth Kirsten","Nicole Kr\\\"amer","Muhammad Bilal Zafar"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-07","first_seen":"2026-09-07","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.04835","pdf_url":"https://arxiv.org/pdf/2609.04835","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["LLM评估","认知多样性","知识生成"],"reason":"论文测量LLM自身的认知多样性，而非用LLM仿真人类被试，但涉及模型行为评估，…","model":"deepseek-v4-pro","scored_at":"2026-09-07T13:01:46","error":null,"has_summary":false,"summary":null},{"id":"2609.04855","version":1,"title":"CC-Mediation: Evaluating Large Language Models for Cross-Cultural Conflict Mediation","zh_title":"CC-Mediation：评估大语言模型在跨文化冲突调解中的表现","abstract":"Cross-cultural mediation by large language models (LLMs) requires deciding both when to intervene and how to respond in culturally grounded conflicts. Progress on this problem has been limited by the lack of (1) mediation datasets with measurable downstream effects and (2) principled metrics for evaluating intercultural stance change. To address these gaps, we introduce CC-Mediation, a cross-cultural mediation benchmark of $1{,}661$ ten-turn dialogues grounded in the Developmental Model of Intercultural Sensitivity (DMIS), containing culturally grounded conflicts, mediation interventions, and post-intervention trajectories. We further propose two DMIS-based evaluation metrics: Trajectory AUC, which measures the persistence of intercultural improvement over time, and a signed Wasserstein-1 distance, which measures the magnitude and direction of shifts in intercultural stance. Both metrics show strong agreement with human judgment of DMIS-grounded stance shift. Using CC-Mediation, we find that current LLMs have limitations on both axes: intervention timing (when) failure stems from a positional prior that ignores dialogue content, while mediation strategy (how) failure arises from a late-layer elicitation collapse rather than a knowledge deficit.","authors":["Suhyun Lee","Wenxuan Zhang","W. Quin Yow","Yang Deng"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-07","first_seen":"2026-09-07","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.04855","pdf_url":"https://arxiv.org/pdf/2609.04855","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["跨文化调解","LLM评估","对话系统"],"reason":"评估LLM在跨文化冲突调解中的表现，测量模型能力而非仿真人类被试，但涉及人类判…","model":"deepseek-v4-pro","scored_at":"2026-09-07T13:01:48","error":null,"has_summary":false,"summary":null},{"id":"2609.05143","version":1,"title":"A Human-in-the-Loop Framework for AI-Assisted Scoring in Large-Scale Writing Assessment","zh_title":"大规模写作评估中AI辅助评分的人机协同框架","abstract":"The integration of artificial intelligence (AI), particularly large language models (LLMs), into educational assessment has opened new opportunities to enhance the efficiency and scalability of grading processes. This study presents the design and validation of an AI-assisted scoring framework for written responses in a large-scale national assessment. The proposed approach focuses on short written texts of approximately 150-200 words and incorporates a human-in-the-loop strategy to preserve assessment quality while reducing manual workload. The study is grounded in a real operational context, using data from two recent editions of a nationwide test, each comprising approximately 5,000 student responses. We analyze the alignment between AI-generated scores and human raters across multiple rubric dimensions, as well as the impact of the proposed decision flow on pass/fail outcomes. Results show moderate to high agreement between the model and human evaluations in most dimensions, supporting the feasibility of AI assistance in this setting. Moreover, the proposed correction workflow identifies cases where human review is most valuable, enabling a more efficient allocation of expert effort. The findings suggest that AI-assisted scoring can be safely integrated into large-scale assessment processes only when combined with carefully designed human oversight. The paper concludes by discussing practical implications for deployment in national assessment systems and outlining future research directions, including longitudinal monitoring of model-human alignment and the analysis of potential cognitive bias introduced by AI-supported review workflows.","authors":["Mar\\'ia Eugenia Curi","Germ\\'an Capdehourat","Isabel Amigo","Magdalena Romano","Rosana Serra","Adri\\'an Silveira","Andr\\'es Peri"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-07","first_seen":"2026-09-07","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.05143","pdf_url":"https://arxiv.org/pdf/2609.05143","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["AI辅助评分","人机协同","教育评估"],"reason":"LLM替代人工评分，属标注员替代而非仿真被试，但涉及人类对照与评估，边界相关。","model":"deepseek-v4-pro","scored_at":"2026-09-07T13:01:50","error":null,"has_summary":false,"summary":null},{"id":"2609.04667","version":1,"title":"ERPBench: Evaluating LLM Agents for Enterprise Decision-Making Across Competitive Market Ecologies","zh_title":"ERPBench：评估跨竞争市场生态的企业决策 LLM 智能体","abstract":"Large language model (LLM) agents are increasingly proposed for enterprise workflows, yet existing evaluations rarely test whether business-decision conclusions transfer across competitive market ecologies. We introduce ERPBench, an execution-instrumented benchmark for enterprise decision agents in a six-round Enterprise Resource Planning (ERP) simulation with coupled pricing, production, procurement, inventory, finance, and shared-market competition. ERPBench evaluates the same 100 fixed problems in two matched competitive market ecologies: Solo, where each evaluated LLM agent competes against fixed rule-based opponents, and Arena, where six evaluated LLM agents compete in a shared market. Across six model families, this yields 1,200 model-level trajectories spanning 7,200 decision rounds. Under the observed service configuration, the leading model differs between ecologies: DeepSeek leads in Solo (252.29M mean valuation; mean rank 1.67), whereas Gemini leads in Arena (263.95M; 1.76). The two ecologies identify the same task-level winner on only 21 of 100 problems, and Gemini's bottom-rank rate falls from 22 % to 0 % in Arena. ERPBench supports paired evaluation of whether enterprise-agent rankings transfer across competitive market ecologies, supplemented by aggregate execution-intervention analysis. Code and benchmark resources are available in our https://github.com/GAIR-NLP/erp-bench.","authors":["Xinran Zhang","Pengrui Lu","Lyumanshan Ye","Pengfei Liu"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-07","first_seen":"2026-09-07","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.04667","pdf_url":"https://arxiv.org/pdf/2609.04667","source_feed":"cs.AI","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["LLM智能体","企业决策模拟","基准测试"],"reason":"LLM agent 在竞争市场生态中做企业决策模拟，但无真实人类数据对照，属于…","model":"deepseek-v4-pro","scored_at":"2026-09-07T13:01:46","error":null,"has_summary":false,"summary":null},{"id":"2609.05346","version":1,"title":"Who Should Grade My Work? Student Perspectives on Transparent AI-Assisted Writing Assessment in Higher Education","zh_title":"谁该给我的作业评分？高等教育中透明AI辅助写作评估的学生视角","abstract":"The integration of GenAI tools into higher education assessment raises important questions about how students understand, interpret, and respond to AI-mediated evaluation. As instructors increasingly explore AI tools for providing feedback, prior research has examined whether GenAI-generated feedback improves writing performance and how students perceive its usefulness; comparatively little is known, however, about how students interpret such evaluation when they are explicitly informed that an AI system, rather than a human instructor, produced the feedback and the score. This study reports findings from a qualitative pedagogical inquiry conducted in an undergraduate technical communication course for computing students at a Saudi public university. Thirteen male undergraduate computing students completed an in-class handwritten writing task; the scanned submissions were evaluated by ChatGPT using a rubric-based prompt aligned with the task objectives. Students were then explicitly informed that ChatGPT had generated the score and feedback and were invited to reflect on the evaluation in writing. Inductive thematic analysis of these reflections identified four themes: perceived usefulness of feedback; awareness of AI's contextual and pedagogical limitations; conditional trust, distinguishing feedback utility from evaluative authority; and reflection on the institutional and pedagogical role of the human instructor. Participants accepted GenAI feedback as useful for surface-level revision but consistently positioned the human instructor as the appropriate authority over grading decisions. The study identifies this as a distinction between feedback utility and evaluative authority, two judgments that students treat as analytically separate rather than as opposite ends of a single approval scale...","authors":["Rayed AlGhamdi"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-07","first_seen":"2026-09-07","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.05346","pdf_url":"https://arxiv.org/pdf/2609.05346","source_feed":"cs.AI","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["AI辅助评估","学生感知","评分权威"],"reason":"LLM替代人工评分，非仿真人类被试，但涉及AI评估与人类判断对比，边界相关。","model":"deepseek-v4-pro","scored_at":"2026-09-07T13:01:51","error":null,"has_summary":false,"summary":null},{"id":"2608.22411","version":2,"title":"Don' t Box Me In: Dynamic Cultural Adaptation and Cognitive Tracking for Social Understanding","zh_title":"别把我框住：面向社会理解的动态文化适应与认知追踪","abstract":"Social interaction increasingly takes place in multicultural settings, where individuals may draw on multiple cultural influences and adapt their communicative behavior across contexts. Despite recent advances in equipping Large Language Models (LLMs) with social understanding capabilities, existing approaches often model culture as a static demographic attribute, limiting their ability to accommodate hybrid and dynamically expressed communicative preferences. Therefore, in this paper, we propose \\textbf{DyCAC}, a training-free framework that achieves fluid social alignment by incorporating \\underline{Dy}namic \\underline{C}ultural \\underline{A}daptation with continuous \\underline{C}ognitive tracking. Rather than inferring a fixed cultural identity, DyCAC models culturally relevant communicative preferences as a time-varying mixture of population-level cultural reference profiles. This reference-based representation is further calibrated using dialogue-style signals observed in the ongoing interaction, enabling the model to capture both composite cultural influences and turn-level shifts in communicative behavior. In parallel, a memory module driven by Theory of Mind (ToM) continuously tracks the cognitive states of the interlocutor. Extensive experiments on interactive social and cultural benchmarks demonstrate the superiority of our approach. The proposed framework outperforms existing baselines, exhibiting enhanced social intelligence and broad adaptability across varied multicultural contexts.","authors":["Chongyuan Dai","Yaling Shen","Shengeng Tang","Hui Ma","Jinpeng Hu"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-07","first_seen":"2026-08-25","revised_at":"2026-09-07","abs_url":"https://arxiv.org/abs/2608.22411","pdf_url":"https://arxiv.org/pdf/2608.22411","source_feed":"cs.CL","score":3,"bucket":"other","rubric_hits":["C3"],"tags":["文化适应","对话系统","社交智能"],"reason":"该研究旨在提升LLM在多元文化对话中的社交对齐能力，属于角色扮演对话与人格化聊…","model":"deepseek-v4-pro","scored_at":"2026-09-07T13:01:54","error":null,"has_summary":false,"summary":null},{"id":"2608.12669","version":2,"title":"From Fair Representation to Just Recognition in Generative AI","zh_title":"从公平表征到生成式AI中的承认正义","abstract":"The fair AI/ML literature has long distinguished distributive fairness, concerning how automated systems allocate resources and opportunities, from representational fairness, concerning how they shape the ways individuals and social groups are perceived, understood, and accorded social status. Generative AI is rebalancing these normative dimensions. Unlike predictive systems, large language models (LLMs) and related technologies are fundamentally expressive: their primary function is to convey meaning rather than automate domain-specific decisions. Representational harm has also become central to value alignment, especially in research on what and whose values and perspectives AI systems should represent. Existing approaches to harms in the representation of social groups often appeal to descriptive accuracy, but this strategy has important limitations. For many social groups, no stable or bounded referent exists against which representational accuracy can be judged. It is also unclear who has the authority to decide what counts as misrepresentation, while even accurate representations can reproduce harmful social patterns. The underlying problem, we argue, is therefore not simply misrepresentation but misrecognition. Drawing on political theory, especially Nancy Fraser's account of participatory parity, we show how moving from representational fairness to recognitional justice provides better conceptual and normative tools for governing central fairness challenges in generative AI.","authors":["Severin Engelmann","Daniel Susser"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"replace","date":"2026-09-07","first_seen":"2026-08-14","revised_at":"2026-09-07","abs_url":"https://arxiv.org/abs/2608.12669","pdf_url":"https://arxiv.org/pdf/2608.12669","source_feed":"cs.CY","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["AI伦理","公平性","政治哲学"],"reason":"论文讨论生成式AI的公平与承认正义，属伦理哲学，非用LLM仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-08-14T13:02:24","error":null,"has_summary":false,"summary":null},{"id":"2608.21618","version":2,"title":"AI-Augmented Inquiry and Regulation in Hybrid Systems: A Control Allocation Architecture for Preserving Epistemic Agency in Hybrid Human-AI Cognition","zh_title":"混合系统中AI增强的探究与调控：一种在混合人机认知中保留认知能动性的控制分配架构","abstract":"Generative artificial intelligence (genAI) systems are increasingly integral to epistemic processes such as hypothesis generation, explanation construction, and decision-making. Although they reliably enhance performance, emerging evidence reveals a metacognitive dilemma: as external generative capacity increases, internal monitoring, calibration, and cognitive engagement may decline. This reflects a redistribution of cognitive control within distributed human-AI systems that cannot be explained by automation bias or reliance on algorithms alone. We propose the AIRIS (AI-Augmented Inquiry and Regulation in Hybrid Systems) framework to analyze this dilemma and specify where regulatory intervention can counteract it. AIRIS is a multi-level control allocation architecture specifying the conditions under which epistemic agency can be preserved in hybrid generative systems. Drawing on distributed cognition, cognitive load theory, multimedia learning, and self-regulated learning, it identifies seven interacting mechanisms through which hybrid cognition may become destabilized, from delegation and calibration drift to motivational-affective drift. Five regulatory operators (Anticipate, Interrogate, Reflect, Integrate, and Synthesize) target internal generative engagement at points of emerging instability. The architecture does not itself improve learning; it specifies what must remain in place for genAI-supported work to sustain understanding, whether through instructional design, teacher guidance, or learners' own regulation. We derive testable propositions concerning the seven mechanisms and the five operators, reframing AI augmentation as a problem of control allocation in distributed generative systems. Beyond theory, AIRIS offers a research agenda, a design framework for genAI-integrated learning environments, and a conceptual toolkit for the governance of hybrid human-AI cognition.","authors":["Jochen Kuhn","Peter Gerjets","Ulrich Trautwein","Jeffrey A. Greene","Sarah Malone","Patrik Vogt","Tim F\\\"utterer"],"categories":["physics.ed-ph","cs.CY"],"primary_category":"physics.ed-ph","announce_type":"replace-cross","date":"2026-09-07","first_seen":"2026-08-25","revised_at":"2026-09-07","abs_url":"https://arxiv.org/abs/2608.21618","pdf_url":"https://arxiv.org/pdf/2608.21618","source_feed":"cs.CY","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["人机协作","认知调控","教育技术"],"reason":"论文提出混合人机认知的控制分配框架，不涉及用LLM仿真人类被试或与真实人类数据…","model":"deepseek-v4-pro","scored_at":"2026-09-07T13:01:53","error":null,"has_summary":false,"summary":null},{"id":"2608.28382","version":2,"title":"When Linguistic and Internal Confidence Diverge in Large Language Models","zh_title":"当大语言模型的语言置信度与内部置信度出现分歧","abstract":"Users often ask large language models (LLMs) to report how confident they are, but it is unclear whether such linguistic confidence tracks the model's internal confidence. We study this question across 8 classification tasks, 2 generation tasks and 30 models from three families. For classification, we compare linguistic confidence with logits-based confidence along three axes: association, magnitude agreement and calibration. For generation, we test whether linguistic confidence tracks semantic-entropy-based uncertainty. The axes frequently diverge. Instance-level association is weak on average, although it improves on easier items and for stronger base models. Instruction-tuned models often report higher confidence and sometimes show higher association, but they also have larger confidence gaps and worse calibration. Prompt design mostly changes the distribution of reported confidence. Attitude cues inflate confidence without improving alignment, while score exemplars can preserve rank-order signal when they avoid collapsed confidence values. Regression analyses show that distributional properties of confidence scores explain much of the observed alignment pattern, with model metadata playing a smaller role after controls. These results support a lossy-channel view of linguistic confidence. A more dispersed verbal confidence distribution can carry useful rank information, but it does not make the scores calibrated. Linguistic confidence should therefore be evaluated with multi-axis diagnostics before being used in downstream reliability pipelines.","authors":["Hefan Zhang","Bingquan Zhang","Ming Cheng","Saeed Hassanpour","Weicheng Ma","Soroush Vosoughi"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-07","first_seen":"2026-08-31","revised_at":"2026-09-07","abs_url":"https://arxiv.org/abs/2608.28382","pdf_url":"https://arxiv.org/pdf/2608.28382","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["置信度校准","模型可靠性","NLP评测"],"reason":"研究LLM自身置信度与内部置信度的关系，属于模型可靠性评估，不涉及人类被试仿真…","model":"deepseek-v4-pro","scored_at":"2026-09-07T13:01:54","error":null,"has_summary":false,"summary":null},{"id":"2608.31115","version":2,"title":"InsightToast: Proactive Information Retrieval & Glanceable Visualization in the Side Channel of Data-Rich Meetings","zh_title":"InsightToast：数据丰富会议侧信道中的主动信息检索与可扫视可视化","abstract":"Missing institutional context during meetings can impede effective participation. Retrieving relevant information, often scattered across heterogeneous internal and external sources, requires costly task-switching that disrupts both individual focus and collective conversational flow, particularly detrimental during cognitively demanding tasks such as decision-making. We introduce InsightToast, a mixed-initiative application that monitors verbal discourse in real time, identifies topics and informational needs as they emerge, and proactively retrieves relevant information through a multi-agent large language model (LLM)-based pipeline integrating retrieval-augmented generation (RAG) to produce source-grounded insights as succinct text and glanceable interactive charts, delivered through a peripheral interface as ephemeral toasts in the conversation's side channel. To demonstrate the potential for yielding serendipitous insights, we showcase a usage scenario involving a knowledge base of legislative documents as the meeting's context. We then report on a comparative study (N=16), in which participants arrived at informed policy decisions while maintaining natural conversation flow.","authors":["Mohammad Abolnejadian","Matthew Brehmer"],"categories":["cs.HC","cs.IR"],"primary_category":"cs.HC","announce_type":"replace","date":"2026-09-07","first_seen":"2026-09-01","revised_at":"2026-09-07","abs_url":"https://arxiv.org/abs/2608.31115","pdf_url":"https://arxiv.org/pdf/2608.31115","source_feed":"cs.HC","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","信息检索","会议支持"],"reason":"多智能体LLM用于会议信息检索，非人类行为仿真，无人类数据对照。","model":"deepseek-v4-pro","scored_at":"2026-09-07T13:01:40","error":null,"has_summary":false,"summary":null},{"id":"2609.04013","version":2,"title":"LLM4CKD: Large Language Models for Early Stage Chronic Kidney Disease Screening","zh_title":"LLM4CKD：用于早期慢性肾病筛查的大语言模型","abstract":"Early screening of chronic kidney disease (CKD) is critical for timely intervention, yet most machine learning (ML) and deep learning (DL) approaches require labeled data and model training, limiting their use in real-world screening settings. This study evaluates the effectiveness of large language models (LLMs) for CKD screening under zero-shot and few-shot in-context learning settings and compares them with traditional ML and DL methods. We propose a framework that uses clinically selected tabular features and structured prompt templates to enable LLM-based inference without task-specific training. LLM performance is evaluated across multiple prompt styles, feature configurations, and data settings, and compared with standard ML, DL, and tabular foundation model (TFM) baselines, and existing CKD screening tools. The results show that LLMs can achieve competitive performance using only a small number of examples, often matching or outperforming traditional approaches in low-data settings. However, their performance remains model-dependent and less stable as input complexity increases. In contrast, ML, DL, and TFM models show more consistent improvement with larger training data. Overall, the findings highlight a trade-off between data efficiency and stability, suggesting that LLMs may serve as a flexible complementary approach for CKD screening when labeled data are limited. To facilitate further research and reproducibility, the code has been made publicly available at https://github.com/akabircs/LLM4CKD","authors":["Muhammad Ashad Kabir","Sirajam Munira"],"categories":["cs.AI","cs.LG"],"primary_category":"cs.AI","announce_type":"replace","date":"2026-09-07","first_seen":"2026-09-04","revised_at":"2026-09-07","abs_url":"https://arxiv.org/abs/2609.04013","pdf_url":"https://arxiv.org/pdf/2609.04013","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["LLM医疗应用","零样本学习","疾病筛查"],"reason":"用LLM做CKD筛查，属于医疗诊断任务，非人类行为仿真，无人类被试对照。","model":"deepseek-v4-pro","scored_at":"2026-09-07T13:01:56","error":null,"has_summary":false,"summary":null},{"id":"2609.04290","version":1,"title":"Evidence Integration in Large Language Models","zh_title":"大语言模型中的证据整合","abstract":"Despite increasing reliance on LLMs that reason with external evidence supplied by tools, retrieval-augmented generation, other agents, and users, how LLMs integrate such evidence into decisions they have already begun to form remains largely unclear. We present a distributional theory in which evidence shifts the receiver's distribution of initial answers, driven by a receiver prior weight and a candidate evidence tilt, leading to three predictions. First, candidates more probable to the receiver are more persuasive. Second, receivers more readily integrate characteristic errors of their own than foreign errors from different sources. Third, identical evidence can improve weaker models and harm stronger ones. We confirm these over ten million trials, twelve LLMs from four families, and eight domains, four of them scientific discovery tasks in the physical and life sciences: quantum mechanics, physics, genetics, and molecular biology. The law also yields a receiver-relative reliability frontier: receiver-congruent errors depress performance more steeply than random errors of the same rate. LLMs also integrate candidates even after internally verifying their invalidity (93-100% with propositional constraints; up to 99.4% on held-out physical and life-sciences reasoning), demonstrating evidence integration is a receiver-specific control policy over existing distributions, determined by receiver properties rather than scalar trust in the evidence source. Causal interventions show candidate integration is implemented late in the network, as a structured sequence of steps admitting external candidate answers, promoting them, and transporting them into the answer state. Representations of verification are decodable but have little causal impact on answers. A J-lens decomposition shows the state underlying verbalized verification is fully dissociable from that underlying candidate integration.","authors":["Sebastien Kawada","Manolis Kellis"],"categories":["cs.CL","cs.AI","cs.LG"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-07","first_seen":"2026-09-07","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.04290","pdf_url":"https://arxiv.org/pdf/2609.04290","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["LLM推理","证据整合","模型行为"],"reason":"研究LLM如何整合外部证据，属模型推理机制，非人类仿真或行为对照。","model":"deepseek-v4-pro","scored_at":"2026-09-07T13:01:42","error":null,"has_summary":false,"summary":null},{"id":"2609.04384","version":1,"title":"You Really Didn't Get That? Benchmarking Social Pragmatic Inference for Indirect and Playful Chinese Online Comments","zh_title":"你真的没听懂吗？面向中文网络评论间接与戏谑表达的社会语用推理基准测试","abstract":"Chinese online comments often convey social meaning through indirect and playful language that is hard to interpret without context. Existing evaluations largely organize items around predefined phenomena or controlled pragmatic categories, leaving open whether models can distinguish plausible readings of what a naturally occurring comment is doing in a particular exchange. We introduce a benchmark for evaluating whether LLMs can recover such situated pragmatic meanings. From more than 200,000 public Chinese social media interaction records, we construct 4,735 human-validated diagnostic items, each pairing a target comment with reconstructed preceding context and plausible misreadings. We evaluate eight LLMs as both question writers and solvers in a cross-writer setting. The task is challenging: the strongest model achieves 81.42% leave-writer-out accuracy. Across all eight models, the mean leave-writer-out accuracy is 68.70% while human accuracy was 90.8%. Case analysis shows that models often recognize broad irony or playfulness while misidentifying the mechanism or interactional move.","authors":["Shiwei Hong","Junjie Ma","Emma Jiren Wang","Ethan Z. Rong","Siying Hu","Haichang Li","Ziying Wang","Zhicong Lu"],"categories":["cs.CL","cs.AI","cs.HC"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-07","first_seen":"2026-09-07","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.04384","pdf_url":"https://arxiv.org/pdf/2609.04384","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["语用推理","基准测试","社交媒体"],"reason":"该论文是LLM在中文社交媒体评论中的语用推理能力评测，属于NLP基准测试，不以…","model":"deepseek-v4-pro","scored_at":"2026-09-07T13:01:44","error":null,"has_summary":false,"summary":null},{"id":"2609.04463","version":1,"title":"Shared circuits predict whether LLMs generalize across formats in arithmetic reasoning","zh_title":"共享电路预测大语言模型在算术推理中跨格式泛化的能力","abstract":"In many forms of reasoning, including arithmetic reasoning, generalizing across superficial changes in input format is effortless for humans: anyone who can solve 2+5 can also solve 'two plus five'. In contrast, LLMs are more brittle to surface variations of the prompts: for example, they solve numeric arithmetic problems almost perfectly but are substantially less accurate on verbal renditions of the same problems. Here, we ask whether generalization across formats can be predicted from the models' internals. Using attribution patching, we first independently localize the circuit that each model recruits to solve numeric arithmetic problems (2+5) vs. verbal ones, in three languages: English ('two plus five'), Spanish ('dos m\\'as cinco'), and Italian ('due pi\\`u cinque'); then, we test whether overlap with the model's own numeric circuit predicts its generalization to the verbal formats. Indeed, we find support for this idea at three levels: circuit overlap accounts for the relative difficulty of the three verbal formats, for which models generalize best, and for which items are solved correctly, rivaling supervised probes while requiring no labeled data.","authors":["Andrea Gregor de Varda","Sana Pandey","Pengrui Han","Jacob Andreas","Evelina Fedorenko"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-07","first_seen":"2026-09-07","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.04463","pdf_url":"https://arxiv.org/pdf/2609.04463","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["LLM推理","电路分析","跨格式泛化"],"reason":"研究LLM算术推理的跨格式泛化，属纯NLP能力评测，不以人类行为仿真为目标","model":"deepseek-v4-pro","scored_at":"2026-09-07T13:01:44","error":null,"has_summary":false,"summary":null},{"id":"2609.04484","version":1,"title":"Patterns of Priming in Production: Lexical, Semantic and Structural Alignment in Language Model Generation","zh_title":"生成中的启动模式：语言模型生成中的词汇、语义和结构对齐","abstract":"This paper investigates structural priming in language model (LM) production, examining how preceding structural context influences sentence completion. While prior work has demonstrated priming effects in comprehension of structural alternations, it remained unclear whether these persist in production, where, when generating, an LM samples from many possible continuations at each step. We address this question through a series of controlled sentence-completion experiments on dative constructions. In line with prior work, we find that LMs are susceptible to structural priming, particularly in sentences that are semantically coherent. In terms of priming magnitude, we find that while there is a greater relative increase of double-object datives against our baselines, in line with inverse frequency effects, there is a larger absolute increase in prepositional-objects, the more frequently produced construction. Finally, we not only observe that structural priming is boosted by lexico-semantic coherence, but that structurally primed completions display greater levels of lexico-semantic repetition. Taken together, our evidence supports the view that structural priming in LMs operates across multiple levels of linguistic representation, facilitating, and facilitated by syntactic, lexical, and semantic alignment. Code: https://github.com/the-context-lab/primedproduction.","authors":["Giulia Pucci","Ruizhe Li","Arabella Sinclair"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-07","first_seen":"2026-09-07","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.04484","pdf_url":"https://arxiv.org/pdf/2609.04484","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["语言模型","结构启动","NLP评测"],"reason":"研究语言模型自身的结构启动效应，属纯NLP能力评测，不以人类行为为参照系。","model":"deepseek-v4-pro","scored_at":"2026-09-07T13:01:44","error":null,"has_summary":false,"summary":null},{"id":"2609.04206","version":1,"title":"Auditing Bias and Safety in Voice AI Customer Care","zh_title":"语音AI客服中的偏见与安全审计","abstract":"Voice AI systems increasingly mediate customer care interactions where caller presentation cues such as accent, affect, fluency, and urgency are available alongside the service request. Existing fairness and safety evaluations cover speech recognition disparities, spoken dialogue bias, and voice agent capability, but rarely treat customer care voice agents as stateful, multi turn, tool mediated systems where harm can appear as additional burden before any final denial occurs. We formalize a validation gated audit framework for such systems. The framework (i) separates native speech to speech, cascaded ASR to language model to TTS, and hybrid tool mediated architectures; (ii) uses matched service facts across controlled caller presentation conditions; (iii) validates fact invariance, presentation cues, artifacts, and acoustic measurements before inference; and (iv) records both material outcomes and path to service burden. We define the research problem, methodology, seven validation gates, a six family metric set, and claim boundaries for an active industry evaluation program. We illustrate the framework with a fully synthetic worked example of a refund dispute audit instance. Production system results are excluded from this release; public reporting is gated by the validation protocol.","authors":["Vignesh Ethiraj","Ashwath David"],"categories":["eess.AS","cs.CL","cs.CY","cs.HC","cs.SD"],"primary_category":"eess.AS","announce_type":"cross","date":"2026-09-07","first_seen":"2026-09-07","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.04206","pdf_url":"https://arxiv.org/pdf/2609.04206","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C2"],"tags":["语音AI","公平性审计","客服系统"],"reason":"研究语音AI客服的公平性与安全审计，不涉及用LLM仿真人类被试或与人类数据对照。","model":"deepseek-v4-pro","scored_at":"2026-09-07T13:01:40","error":null,"has_summary":false,"summary":null},{"id":"2609.04303","version":1,"title":"Abstraction Agent","zh_title":"抽象智能体","abstract":"Information abstraction, which groups strategically similar private states into a tractable number of buckets, is essential for scaling game-solving algorithms to large imperfect-information games. Constructing effective abstractions, however, has traditionally required domain-specific evaluators such as hand-strength calculators or equity estimators, which demand expert knowledge and engineering effort and are unavailable for most less-studied games. We propose the Abstraction Agent, a zero-shot pipeline that uses a large language model (LLM) to discover continuous strategic features from a natural-language game description, score private states on these features, and cluster them into abstraction buckets, without any game-specific evaluator, training data, or game-tree traversal during abstraction construction. The pipeline runs in four phases: feature discovery with calibration anchors, batched private-state scoring, correlation-based feature selection, and $k$-means clustering. The resulting abstractions reduce lifted-strategy exploitability by up to 62% relative to an expected-hand-strength baseline on heads-up no-limit Texas hold'em (HUNL) turn endgames, and beat a scalar rank baseline at every granularity on ROVER Trials, an original game absent from any pretraining corpus. Beyond these quantitative benchmarks, the pipeline transfers with unchanged prompts to four-card Pot-Limit Omaha, HUNL preflop and flop, and Riichi Mahjong, where the discovered features track each game's recognized strategic concepts. This is structured knowledge elicitation: converting implicit strategic knowledge in LLM parameters into explicit numerical features for downstream algorithmic computation. The code is available at https://github.com/lbn187/AbstractionAgent.","authors":["Boning Li","Longbo Huang"],"categories":["cs.MA","cs.AI","cs.CL","cs.GT"],"primary_category":"cs.MA","announce_type":"cross","date":"2026-09-07","first_seen":"2026-09-07","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.04303","pdf_url":"https://arxiv.org/pdf/2609.04303","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C2"],"tags":["游戏AI","LLM应用","知识提取"],"reason":"用LLM辅助游戏抽象构建，属于游戏AI，不涉及人类行为仿真或对照。","model":"deepseek-v4-pro","scored_at":"2026-09-07T13:01:42","error":null,"has_summary":false,"summary":null},{"id":"2609.04850","version":1,"title":"ElderBench: Benchmarking Autonomous Mobile Agents for Older Adults","zh_title":"ElderBench：面向老年人的自主移动代理基准测试","abstract":"While autonomous mobile agents hold great potential for assisting older adults with smartphone usage, existing GUI benchmarks mainly rely on explicit, goal-oriented instructions and rarely capture the naturally occurring language patterns of older users, such as indirect speech, referential ambiguity, and under-specified requests. This mismatch between benchmark instructions and real-world elderly interactions may hinder reliable agent deployment. To address this gap, we present ElderBench, the first benchmark for evaluating mobile GUI agents in authentic elderly-oriented scenarios. ElderBench is constructed from 249 naturally elicited smartphone tasks collected from older adults across 20 applications. We first characterize the linguistic divergence between elderly instructions and existing GUI benchmark instructions from syntactic, semantic, and pragmatic perspectives. We then evaluate mainstream GUI agents and Vision-Language Models under both online and offline settings, revealing substantial performance degradation when handling elderly-oriented instructions. Through controlled instruction normalization, failure analysis, and fine-grained linguistic feature analysis, we further identify how elderly-specific language patterns contribute to agent failures. Our findings provide actionable design insights toward more adaptive, interpretable, and age-inclusive GUI agents for older adults.","authors":["Weide Zhan","Qumu Shaqu","Yuanqing Liu","Peng Zhang","Jiahao Liu","Kam Him Lam","Ning Gu","Zhan Hu","Tun Lu"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-07","first_seen":"2026-09-07","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.04850","pdf_url":"https://arxiv.org/pdf/2609.04850","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C2"],"tags":["GUI代理","老年人","基准测试"],"reason":"评估移动GUI代理在老年人场景的性能，属于机器人/代理仿真环境，不涉及用LLM…","model":"deepseek-v4-pro","scored_at":"2026-09-07T13:01:47","error":null,"has_summary":false,"summary":null},{"id":"2609.04894","version":1,"title":"From Language Models to World-Acting Systems: Progress and Limits of Agentic AI across Digital, Social, Virtual, and Physical Environments","zh_title":"从语言模型到世界行动系统：智能体AI在数字、社会、虚拟和物理环境中的进展与局限","abstract":"Large language models become consequential agents when surrounding systems let outputs change external state. Models now call tools, operate interfaces, delegate work, retain state, inhabit generated worlds, and control robots or laboratory equipment. Such advances are often narrated as one march toward autonomy, conflating model competence, system integration, persistence, and safe authority. This critical review synthesizes primary research and official technical specifications available by 31 August 2026. We organize the evidence along delegated authority, temporal persistence, and environmental coupling, while separating model, harness, and environment. Within the evidence examined, action-interface expansion is documented more convincingly than robust completion, recovery, authorization, or independent verification. Model Context Protocol and Agent2Agent improve interoperability but do not establish trustworthy delegation; multi-agent organization adds specialization alongside cost and correlated failure. Persistent simulations and world models support training and planning but do not themselves demonstrate agency; robotics and self-driving laboratories establish bounded feasibility rather than unattended open-world reliability. We propose justified delegation as an analytical and normative heuristic, not an observed law or certified score: expand action scope only where evidence supports provenance, bounded authority, failure detection, safe recovery, and calibrated human control. This framing yields a research agenda for coupled model-harness evaluation, capability-based permissions, durable state, cross-agent accountability, and staged physical validation.","authors":["Linsen Zhu","Mengqing Cai"],"categories":["cs.AI","cs.LG","cs.MA"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-07","first_seen":"2026-09-07","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.04894","pdf_url":"https://arxiv.org/pdf/2609.04894","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C1","C2"],"tags":["智能体系统","自主性","系统集成"],"reason":"综述智能体系统在数字、社会、虚拟和物理环境中的行动能力，聚焦系统集成与自主性，…","model":"deepseek-v4-pro","scored_at":"2026-09-07T13:01:49","error":null,"has_summary":false,"summary":null},{"id":"2609.05363","version":1,"title":"Distill Globally, Adapt Locally: Reasoning Distillation and Product-Type Test-Time Training for Scalable Trade-Up Recommendation","zh_title":"全局蒸馏，局部适应：面向可扩展升级推荐的理由蒸馏与产品类型测试时训练","abstract":"Trade-up recommendation identifies higher-quality alternatives that preserve a customer's purchase intent while offering upgraded benefits. Large language models (LLMs) can reason about such distinctions, but applying them directly to hundreds of millions of product pairs is operationally impractical. We introduce a two-level framework that distills LLM reasoning into an efficient non-generative student and adapts its decision boundary to product-type-specific trade-up criteria. At Level 1, a retrieval-augmented few-shot LLM teacher generates structured relation labels and natural-language rationales. These rationales supervise a compact embedding-pair classifier through alignment and contrastive objectives; at inference, the student uses only two precomputed 768-dimensional product embeddings, with no LLM calls or text generation. On a fixed human-annotated benchmark of 8,352 pairs, a 15.5M-parameter four-class reasoning-distilled student achieves AUC 0.924 (95% CI [0.918, 0.929]), compared with 0.912 for the four-class label-only student. At Level 2, product-type test-time training (PT-TTT) uses few-shot demonstrations to optimize lightweight category-specific adapters over the frozen student. PT-TTT improves AUC from 0.924 to 0.941 and average precision from 0.920 to 0.940. On a 100K-pair proxy catalog, the distilled student on a single eight-GPU machine is approximately 5,000x faster and 10,000x lower in estimated cost than direct LLM inference.","authors":["Siliang Liu","Mohammad Ghasemi","Sapan Patel","Amin Banitalebi-Dehkordi"],"categories":["cs.LG"],"primary_category":"cs.LG","announce_type":"new","date":"2026-09-07","first_seen":"2026-09-07","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.05363","pdf_url":"https://arxiv.org/pdf/2609.05363","source_feed":"cs.LG","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["推荐系统","知识蒸馏","LLM应用"],"reason":"论文用LLM蒸馏做商品推荐，不涉及人类行为仿真或对照，属多智能体/模型蒸馏应用。","model":"deepseek-v4-pro","scored_at":"2026-09-07T13:01:52","error":null,"has_summary":false,"summary":null},{"id":"2608.06621","version":3,"title":"NxN E-valuation: Hypothesis Certification via a Conformal CRT Null","zh_title":"NxN E-valuation：通过共形CRT零假设进行假设认证","abstract":"We propose NxN E-valuation, a handy, e-value-based hypothesis-certification algorithm that lets a hypothesis be verified without building any case-specific certification procedure---such as constructing a dedicated null hypothesis---as long as a large enough dataset is available. The method is especially suited to LLM-based exploration systems, where LLMs are remarkably good at proposing hypotheses but suffer badly from hallucination; this hallucination prevents us from harvesting LLM outputs directly, and existing remedies each fall short. The most common solutions include letting the LLM verify or correct itself circular verification and held-out testing (where false hypotheses can still pass via spurious correlations), among other remedies detailed in the introduction. To resolve this, NxN E-valuation exploits the naturally existing large training set and lets different samples serve as null hypotheses for one another. This design directly realizes a conditional randomization test (CRT) that certifies each hypothesis. The approach can be a universally better replacement for at least LLM circular verification and held-out-data testing, provided the LLM's generations are hypotheses that apply to each individual sample.","authors":["Bin Wang","Yan Zhong","Liang Luo","Buyun Zhang","Ellie Wen"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"replace","date":"2026-09-07","first_seen":"2026-08-10","revised_at":"2026-09-07","abs_url":"https://arxiv.org/abs/2608.06621","pdf_url":"https://arxiv.org/pdf/2608.06621","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["假设检验","LLM验证","多智能体"],"reason":"纯多智能体协作验证假设，无人类行为对照，不涉及人类仿真实验。","model":"deepseek-v4-pro","scored_at":"2026-08-10T13:01:33","error":null,"has_summary":false,"summary":null},{"id":"2609.04556","version":1,"title":"Rhythms of Work: Multi-Scale Interpretation of Human Behavioral Traces for Workplace Agents","zh_title":"工作节奏：面向工作场所智能体的人类行为轨迹多尺度解释","abstract":"Runtime traces are becoming a central substrate for understanding agentic systems, yet interpretation has focused largely on what the agent did. Workplace agents face the complementary problem: interpreting the human activity that surrounds them. Hours of low-level events carry rich evidence about a user's state but are too granular to reason over directly, and flattening them into one stream or compressing them into a single embedding both treat \"summarize the user's behavior\" as if it had one correct answer. We argue instead that behavioral interpretation is resolution-dependent: the same trace should admit multiple addressable interpretations at different temporal resolutions. We construct a multi-resolution vocabulary of semantically normalized operators, recurring motifs, coherent episodes, and day-level rhythms, each preserving the structure salient at its own horizon. Applied to 667 million human-attributed events from a large commercial productivity suite (50,000 users, 100 organizations), it yields 120 operator types, thousands of motifs, 25 episode types, and five day-rhythm archetypes. We validate it on real telemetry: re-running the entire pipeline on a disjoint 2,000-user sample recovers the same taxonomy (structural stability), and on held-out users the full representation forecasts a user's next episode more accurately than a flat-operator baseline, a 17% relative macro-F1 gain (predictive validity), so the abstractions preserve future-relevant information rather than merely describe it. A controlled resolution ablation then shows that no single level is optimal across questions: different agent-facing questions about the same trace are best answered at different resolutions. Behavioral trace interpretation for agents should therefore be multi-resolution and query-conditioned: an agent should access the temporal grain a question needs, not one universal summary.","authors":["Lin Ai","Scott Counts"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-07","first_seen":"2026-09-07","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.04556","pdf_url":"https://arxiv.org/pdf/2609.04556","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["行为轨迹分析","多尺度解释","工作场所智能体"],"reason":"研究人类行为轨迹的多尺度解释，用于工作场所智能体，不涉及用LLM仿真人类被试或…","model":"deepseek-v4-pro","scored_at":"2026-09-07T13:01:46","error":null,"has_summary":false,"summary":null},{"id":"2609.04841","version":1,"title":"MABPD: Multi-Agent Bias Probing & Detection via Structured Argument Debate","zh_title":"MABPD：通过结构化论证辩论进行多智能体偏见探测与检测","abstract":"Media bias in news articles operates through subtle linguistic cues---loaded language, selective framing, and strategic omission---that resist single-model detection and have traditionally required large annotated corpora for supervised training. We ask whether structured multi-agent deliberation can serve as a principled, training-free alternative to supervised classification for this task. We introduce MABPD (Multi-Agent Bias Probing & Detection), a pipeline in which three specialized LLM agents analyze an article from complementary perspectives and resolve disagreements through a Structured Argument Debate (SAD) protocol. SAD implements a domain-motivated asymmetric burden of proof---biased claims without grounded textual evidence carry zero weight---combined with role-weighted voting and post-consensus verification, replacing task-specific supervised decision boundaries with explicit deliberative structure. Ablation confirms that this structured deliberation, not mere agent parallelism, drives performance: removing the debate module reduces F1 by up to 10.6 points. On the BABE benchmark (4,121 expert-annotated sentences), MABPD achieves 83.4% macro F1 on the held-out test split---within 0.7 percentage points (pp) of the supervised SOTA (MAGPIE, 84.1% macro F1; Horych et al., 2024)---without any task-specific training or threshold tuning on annotated data. Cross-dataset evaluation on the SemEval 2019 HyperPartisan corpus (644 articles) yields 75.0% zero-shot accuracy, within 7.2 pp of the supervised SOTA accuracy (82.2%; Kiesel et al. 2019), confirming transfer across annotation regimes. We release the full pipeline and evaluation code.","authors":["Garvit Joshi (Graphic Era University, Dehradun, India)","Stavya Dhyani (Graphic Era University, Dehradun, India)","Jasmine (Graphic Era University, Dehradun, India)","Arun Chauhan (Graphic Era University, Dehradun, India)"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-07","first_seen":"2026-09-07","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.04841","pdf_url":"https://arxiv.org/pdf/2609.04841","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","偏见检测","NLP应用"],"reason":"多智能体辩论用于媒体偏见检测，是协作解题，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-07T13:01:47","error":null,"has_summary":false,"summary":null},{"id":"2609.04272","version":1,"title":"Evaluating Large Language Models for Forced Outage Risk Prediction: Benefits and Comparison to Machine Learning","zh_title":"评估大语言模型在强制停电风险预测中的表现：优势及与机器学习的比较","abstract":"This study examines the ability of large language models (LLMs) to predict the risk of weather-related forced outages in the distribution grid in a zero-shot framework, without labeled training data. The problem is formulated as a binary severity classification task across three forecast horizons (3h, 6h, 12h), using six years of outage records and high-resolution weather data for a utility service area in central Texas. Four zero-shot LLMs are benchmarked against two supervised classifiers across two input configurations: one using current weather observations and the other using weather forecast data. Results show that supervised models outperform LLMs on macro-F1 and precision, while newer LLM generations achieve competitive scores. Beyond accuracy, LLMs offer complementary strengths in actionable reasoning and geographic scalability, suggesting that combining them with supervised models may be the best practice.","authors":["Christos Petridis","Zoran Obradovic","Mladen Kezunovic"],"categories":["cs.LG","cs.CL"],"primary_category":"cs.LG","announce_type":"cross","date":"2026-09-07","first_seen":"2026-09-07","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.04272","pdf_url":"https://arxiv.org/pdf/2609.04272","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["LLM预测","停电风险","零样本分类"],"reason":"该研究用LLM做停电风险预测，属于纯NLP能力评测，不以人类行为为参照系，不涉…","model":"deepseek-v4-pro","scored_at":"2026-09-07T13:01:42","error":null,"has_summary":false,"summary":null},{"id":"2609.04445","version":1,"title":"Conformity Breaks Conformal Prediction","zh_title":"从众打破共形预测","abstract":"A conformal certificate can be valid when an LLM answers alone and invalid when the same LLM sees peers that unanimously assert a wrong answer. The question is unchanged; the model's score for the correct answer changes. We call this a score-mechanism shift: clean calibration certifies how the model scores answers alone, but not how it scores them under peer pressure. We show that this shift silently breaks conformal prediction in multi-agent LLM systems. Across open-weight models and multiple-choice QA tasks, coverage falls from a calibrated 90% to 74% under unanimous-wrong peers at the standard alpha = 0.10 operating point. The average hides a sharper failure: by targeting the low-confidence items the certificate still covers, an attacker nearly halves coverage on that subgroup, from 87% to 47%, while the monitored average remains much higher. The failure also reaches the decision layer: a system that should escalate when uncertain can instead become confident enough to act on the attacker's wrong answer. Standard conformal fixes do not solve the problem, because the question distribution has not changed; the model's scoring behavior has.","authors":["Yibo Hu","Hanyu Su"],"categories":["cs.LG","cs.CL"],"primary_category":"cs.LG","announce_type":"cross","date":"2026-09-07","first_seen":"2026-09-07","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.04445","pdf_url":"https://arxiv.org/pdf/2609.04445","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","共形预测","LLM可靠性"],"reason":"研究多智能体LLM系统中的共形预测失效，属于多智能体协作问题，不涉及人类行为仿…","model":"deepseek-v4-pro","scored_at":"2026-09-07T13:01:33","error":null,"has_summary":false,"summary":null},{"id":"2609.04286","version":1,"title":"From Matching Models to Recruiting Agents: A Systematized Narrative Review of AI Recruitment Systems, Evaluation, and Governance","zh_title":"从匹配模型到招聘智能体：AI招聘系统、评估与治理的系统化叙述性综述","abstract":"Artificial intelligence in recruitment has shifted the object being automated from profile pairs and ranked lists to multi-stage workflows that retrieve evidence, compare candidates, and support or execute actions. This systematized narrative review traces that development from bilateral retrieval and behavioral ranking through neural person--job matching, large language model (LLM) components, and tool-using recruiting agents. Using a purposive search and coding protocol updated through 23 July 2026, plus targeted updates through 2 September 2026, we organize 40 representative works with supporting industrial and legal sources. This synthesis is not a prevalence estimate. We analyze three coupled transitions: from similarity to reciprocal suitability, from a model to a compound workflow, and from offline prediction to evidence- and productivity-aligned evaluation. Across document understanding, retrieval, ranking, assessment, interviewing, sourcing, and human handoff, we distinguish field-, pair-, list-, case-, trajectory-, and outcome-level evidence. Persistent gaps arise because behavioral labels confound exposure, preference, and qualification; private and synthetic data limit external validity; final-output scores conceal pipeline failures; and, within the coded set, privacy is not directly evaluated and no row jointly evaluates utility, fairness, privacy, and security. These observations describe the coded set rather than the field as a whole. We therefore introduce a staged mapping from evaluation evidence to the strongest defensible claim, together with an agenda for reciprocal, evidence-grounded, temporally controlled, selective, and auditable systems. Progress should be judged by whether workflows retrieve the right evidence, preserve uncertainty, support contestable decisions, and improve outcomes under explicit cost and risk constraints.","authors":["Ziyi Zhao","Guanzheng Wei"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-07","first_seen":"2026-09-07","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.04286","pdf_url":"https://arxiv.org/pdf/2609.04286","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["AI招聘","多智能体系统","系统综述"],"reason":"论文综述AI招聘系统，聚焦多智能体工作流自动化，不涉及用LLM仿真人类被试或与…","model":"deepseek-v4-pro","scored_at":"2026-09-07T13:01:42","error":null,"has_summary":false,"summary":null},{"id":"2609.04880","version":1,"title":"Reinforcement Learning for Sequential Solar PV Policy Design under Uncertainty: An Agent-Based Approach","zh_title":"不确定性下太阳能光伏政策序列设计的强化学习：基于智能体的方法","abstract":"Designing effective and fiscally sustainable policies for solar photovoltaic (PV) adoption requires balancing adoption gains against public expenditure under uncertainty and heterogeneous decision-making. This study formulates PV policy design as a sequential decision problem and integrates reinforcement learning (RL) with a stochastic agent-based model (ABM) that simulates yearly solar PV adoption under uncertainty. A policymaker agent selects annual incentives, including capital grants, subsidised loan rates, and feed-in tariffs, over a 16-year horizon. Adoption--cost trade-offs are explored by varying policy preferences within a scalarised reward framework. Policies are learned using PPO, SAC, and TD3 and evaluated under stochastic simulation. The results show that this approach produces a clear trade-off structure: the highest-adoption policy (TD3, $w_{\\text{cost}}=0.5$) achieves approximately 4,145 adopters at a cost of EUR 41.73 million, while the lowest-cost policy (PPO, $w_{\\text{cost}}=2.0$) reduces expenditure to EUR 7.27 million with 2,682 adopters. The balanced policy (PPO, $w_{\\text{cost}}=1.6$) achieves 3,495 adopters at a cost of EUR 22.47 million. Across algorithms, consistent trade-off patterns are observed, indicating robustness of the adoption--cost relationship. Compared with static baseline policies, the RL framework explores a broader range of policy configurations. These findings demonstrate the potential of RL as a flexible tool for adaptive policy design under uncertainty.","authors":["Iias Faiud","Jonaid Shianifar","Michael Schukat","Karl Mason"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-07","first_seen":"2026-09-07","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.04880","pdf_url":"https://arxiv.org/pdf/2609.04880","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["强化学习","智能体建模","政策设计"],"reason":"多智能体强化学习用于政策设计，agent模拟的是太阳能采纳决策，但无LLM作为…","model":"deepseek-v4-pro","scored_at":"2026-09-07T13:01:49","error":null,"has_summary":false,"summary":null},{"id":"2609.04542","version":1,"title":"Matched Starts, Divergent Objects: How Human-AI Collaboration Forms What It Explains","zh_title":"匹配起点，分歧对象：人机协作如何形成其解释内容","abstract":"Scholarly knowledge is typically encountered in stabilized form, while the process histories through which research objects, claims, and contributions acquire form remain largely hidden. This study examines how human-AI scholarly collaboration develops under matched starting conditions and whether those conditions stabilize the inquiry itself. Using a longitudinal corpus of 843 turns, the same expert researcher developed branch-isolated scholarly trajectories with different generative AI systems from the same corpus, frozen research problem, starting prompt, publication objective, and conduct rules. Two eligible trajectories were reconstructed ex post through scholarly trajectory analysis, source-faithful interaction reconstruction, a Socioduality relational-process overlay, and downstream propagation analysis. Both trajectories independently shifted the initial continuity problem from recall toward usability, but subsequently formed different research objects. One trajectory culminated in an endpoint manuscript on continuity labour, the distributed work required to sustain usable collaboration; the other in an endpoint manuscript on distributed, evolving, and unevenly usable project state. Their analytic genealogies involved failed analytical units, rejected explanations, changes in scale, counterexamples, and conceptual stabilization, and propagated into different research questions, findings, methods, evidence logics, and scholarly contributions. Relational analysis further showed that consequential scholarly change, reciprocal continuity, and local substantive re-formation were distinct process structures, and that continuation did not necessarily constitute epistemic endorsement. The findings demonstrate empirically constrained research-object formation within human-AI scholarly collaboration and show how process histories shape the scholarly objects and products that emerge.","authors":["Mehmed Zahid \\c{C}\\\"ogenli"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-07","first_seen":"2026-09-07","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.04542","pdf_url":"https://arxiv.org/pdf/2609.04542","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["人机协作","学术知识生产","过程分析"],"reason":"研究人类与AI协作过程，非用LLM仿真人类被试，无人类行为对照","model":"deepseek-v4-pro","scored_at":"2026-09-07T13:01:45","error":null,"has_summary":false,"summary":null},{"id":"2609.05115","version":1,"title":"Creators Have Difficulty Abandoning Ideas They Generated","zh_title":"创造者难以放弃自己产生的想法","abstract":"Creativity researchers often distinguish between two stages of the creative process: generation versus selection. While much is known about the psychology of idea generation (e.g., the factors that lead to a greater number of novel and useful ideas), less is understood about the nature of selection, or how generation and selection interact. Here we investigate how the act of generating ideas may potentially distort the selection process. Using an incentive-compatible paradigm in which pairs of participants reviewed the same ideas and were rewarded for submitting only high-quality ideas, we find that people submit a greater number of lower-quality ideas when selecting among their own ideas than when selecting among another person's ideas (the Creative Endowment Effect). This effect generalizes across three tasks in two domains and is resistant to an informational intervention (i.e., explicitly telling people about the effect). However, having participants revisit their ideas several months later increases their selectivity. The broader implications for individuals and organizations are discussed.","authors":["Jin Kim","George E. Newman"],"categories":["econ.GN","q-fin.EC"],"primary_category":"econ.GN","announce_type":"new","date":"2026-09-07","first_seen":"2026-09-07","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.05115","pdf_url":"https://arxiv.org/pdf/2609.05115","source_feed":"econ.GN","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["创造力","决策偏差","人类实验"],"reason":"研究人类创造者对自己想法的选择偏差，未使用LLM或仿真，与LLM仿真人类被试无…","model":"deepseek-v4-pro","scored_at":"2026-09-07T13:01:49","error":null,"has_summary":false,"summary":null},{"id":"2609.03215","version":1,"title":"SWIM: Student Writing Simulation via Proficiency-Conditioned Generation","zh_title":"SWIM：基于熟练度条件生成的学生写作仿真","abstract":"Writing proficiency manifests in how students develop content, organize ideas, choose words, and use language. Despite growing interest in LLM-based student simulation, whether LLMs can reproduce such multidimensional variation in extended writing remains largely unexplored. In this work, we explore if language models can realistically simulate student writing, and introduce SWIM, a task that formulates Student Writing sIMulation as proficiency-conditioned essay generation. We evaluate prompting, supervised fine-tuning (SFT), and reinforcement learning (RL) methods for writing simulation using automated essay scoring as a measure of profile alignment. Extensive experiments reveal that prompting provides limited proficiency control, even for strong proprietary LLMs with rubric-grounded strategies. In particular, while models can adjust content-oriented traits, they struggle to reproduce the lexical, grammatical, and organizational variation in different proficiency levels. SFT substantially improves alignment, while RL with the proposed proficiency-alignment reward yields further gains across all writing traits and essay prompts. Our findings suggest that explicit supervision enables substantially stronger profile alignment than prompting alone, while authentic low-proficiency writing remains challenging to reproduce.","authors":["Heejin Do","Jakub Kontak","Mrinmaya Sachan"],"categories":["cs.CL","cs.LG"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-04","first_seen":"2026-09-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.03215","pdf_url":"https://arxiv.org/pdf/2609.03215","source_feed":"cs.CL","score":9,"bucket":"selected","rubric_hits":["A1","B1","B4"],"tags":["LLM仿真","学生写作","熟练度对齐"],"reason":"用LLM仿真学生写作，与真实学生数据对照，并指出低水平写作难以复现。","model":"deepseek-v4-pro","scored_at":"2026-09-04T13:01:26","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-04","rank":1,"question":"语言模型能否在多个写作特质上按不同水平真实模拟学生作文？","design":"用提示、监督微调（SFT）和强化学习（GRPO）方法，让LLM根据目标写作特质分数生成作文，用自动作文评分模型（ArTS）评估生成作文与目标特质分数的对齐程度。","baseline":"ASAP/ASAP++ 数据集中真实学生的作文及其多特质评分。","findings":"提示方法对写作水平的控制有限，尤其在词汇、语法和组织等特质上表现差；SFT显著改善对齐，而GRPO配合提出的水平对齐奖励在所有特质和作文题目上进一步提升。低水平写作的真实语言特征仍难以复现，模型生成的低分作文往往只是表面粗糙，缺乏真实低水平学生的语言模式。","reliability":"论文承认低水平写作的复现仍是瓶颈，模型倾向于生成过于润色的文本；同时指出自动评分器可能无法完全捕捉真实写作的细微差异，且提示方法存在表面破坏的失效模式。","relevance":"该研究直接探索LLM模拟人类行为的可靠性，与研究者关注的人类仿真实验高度相关，尤其在教育评估场景下提供了真实数据对照和失效条件分析，值得精读。","inspiration":"借鉴其用真实标注数据微调模型并设计奖励函数来对齐多维特质的方法，可迁移到经济金融中的个体决策仿真，如消费者风险偏好或投资者情绪模拟。｜可应用于信贷审批中的申请人陈述分析或政策沟通中的公众反应预测。｜以真实信贷申请文本为训练数据，用LLM生成不同信用评分和风险偏好水平的申请人陈述，处理为条件生成（给定信用分和风险特质），结果变量为生成文本的自动评分与真实评分的一致性，对照真实申请人的文本和评分数据。"}},{"id":"2609.03553","version":1,"title":"GPS-Bench: A Governance Policy Benchmark for Automating Policy Analysis","zh_title":"GPS-Bench：用于自动化政策分析的治理政策基准","abstract":"Policy analysis requires more than predicting whether a proposal will pass: it requires identifying who will be affected, how those actors respond, and what follows. LLM-based policy simulations model these processes at scale, but their validity is hard to establish when plausible behaviour is never compared with observed outcomes. We introduce GPS-Bench, an evidence-grounded benchmark for governance policy simulation that links policies to relevant actors, actor actions and downstream impacts using legislative records, lobbying disclosures, regulatory documents, corporate filings, economic data and other public evidence. Actors are reconstructed from the dated record rather than prompted as archetypes, so a persona is an evidence object with provenance; a human-annotated pool forms the Gold evaluation set, while cases labelled by a separate LLM from retrieved evidence are treated as Silver supervision and never as test labels. Because every inference mode reads the same grounded state and emits the same schema, GPS-Bench turns \"does multi-agent simulation help?\" into a controlled comparison: we contrast joint reasoning, independent and communicating actor agents, graph-based methods and weight-level fine-tuning over one policy state. Fine-tuning on the grounded record gives the strongest actor-level impact prediction, and decomposition does not beat it; what decomposition adds is mechanism. Agents hold private, non-identical evidence, each seeing its own exposure clause, and address named partners with concrete joint proposals, what they offer, what they need in return, and why acting together beats acting alone, so the coalitions that form can be checked against the commitments the record holds. GPS-Bench therefore gives a common empirical setting for studying when evidence, actor modelling and multi-agent interaction improve the prediction and interpretation of policy outcomes.","authors":["Linh Le","Melanie Bui","My Chiffon Nguyen","Zachary Schlosser","David Williams-King"],"categories":["cs.AI","cs.CY"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-04","first_seen":"2026-09-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.03553","pdf_url":"https://arxiv.org/pdf/2609.03553","source_feed":"cs.AI","score":8,"bucket":"selected","rubric_hits":["A3","B1","B2","B4"],"tags":["LLM仿真","政策模拟","基准测试"],"reason":"用LLM多智能体模拟政策过程，并与真实记录对照，评估仿真有效性，属核心相关。","model":"deepseek-v4-pro","scored_at":"2026-09-04T13:01:26","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-05","rank":1,"question":"如何构建一个基于真实记录的治理政策仿真基准，以评估LLM多智能体在预测政策结果、受影响行动者及行动者层面影响上的有效性？","design":"论文构建了GPS-Bench基准，使用LLM智能体基于公共记录重建的政策状态进行推理，对比联合推理、独立行动者智能体、通信智能体、图方法及权重级微调等不同推理模式，预测立法通过、受影响行动者识别和行动者层面影响方向。","baseline":"使用立法记录、游说披露、监管文件、公司文件、经济数据等公共记录重建的真实政策过程与结果作为对照，其中人类标注的Gold集作为评估标准。","findings":"权重级微调在行动者层面影响预测上表现最强，分解推理并未超越它；分解推理的主要贡献在于提供机制解释。智能体持有私有证据并形成可核查的联盟，但通信对预测性能的提升有限。","reliability":"论文承认LLM智能体可能产生看似合理但与实际不符的交互，存在过度收敛、群体追踪偏差和政治偏见等问题；Silver标签由LLM生成，仅作监督信号，不作为测试标签，以避免污染。","relevance":"该研究直接针对LLM仿真在政策分析中的有效性问题，提供了与真实记录对照的基准，对关注经济学实验和政策评估仿真的研究者具有重要参考价值。","inspiration":"借鉴其将仿真输出与真实历史记录逐项对照的评估框架，以及将行动者建模为具有证据来源的实体而非抽象原型的方法。｜可迁移到政策公告的预期形成与市场反应研究，例如模拟央行利率决议或财政刺激方案对不同市场参与者的影响。｜以LLM智能体扮演投资者、企业、消费者等，处理为政策公告内容，结果变量为各主体的预期调整和决策行为，对照真实市场数据（如股价变动、调查预期）进行验证。"}},{"id":"2609.03218","version":1,"title":"The Analyst in the Prompt: Role, Retrieval, and Memory Biases in LLM Financial Analysis","zh_title":"提示中的分析师：LLM金融分析中的角色、检索与记忆偏差","abstract":"Large Language Models (LLMs) increasingly use user context such as memory, profiles, and role prompts to personalize their responses. This personalization can affect evidence-based judgment: the same evidence may lead to different conclusions under different user contexts. Finance provides a high-stakes setting to study this problem because decisions often depend on interpreting long and complex documents. We test this using 3,575 SEC filings across twelve LLMs. We compare persona-conditioned retrieval, neutral retrieval, and memory-framed context to separate the effect of evidence selection from the effect of interpretation. We find that most user-context spillover comes from how models interpret the same evidence under different roles, rather than from retrieving different evidence. We then test two simple mitigation strategies: expressing the same investor mindset as a user profile instead of an assistant role, and separating evidence-based and personalized outputs. Both reduce spillover, but neither removes it completely, and their effectiveness varies substantially across models.","authors":["Ahmed Asaad","Amr Mohamed","Yang Zhang","Omneya Abdelsalam"],"categories":["cs.CL","cs.CE","q-fin.PM"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-04","first_seen":"2026-09-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.03218","pdf_url":"https://arxiv.org/pdf/2609.03218","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A2","B1","B4"],"tags":["LLM偏差","金融分析","个性化影响"],"reason":"研究LLM在金融分析中的角色、检索和记忆偏差，评估个性化对证据判断的影响，有真…","model":"deepseek-v4-pro","scored_at":"2026-09-04T13:01:36","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-04","rank":3,"question":"在金融分析中，用户上下文（角色、记忆、档案）是否会导致LLM对同一证据的中性判断发生系统性偏移？","design":"使用12个LLM，对3575份SEC文件进行分析，通过三种条件（角色条件检索、中性检索、记忆框架上下文）分离证据选择与解释效应，测量中性证据分数（EvidenceScore）的偏移。","baseline":"无对照","findings":"用户上下文溢出主要来自模型在不同角色下对相同证据的解释差异，而非检索不同证据。两种缓解策略（将投资者心态表达为用户档案而非助手角色、分离证据与个性化输出）可减少溢出但无法完全消除，且效果因模型而异。","reliability":"论文未讨论","relevance":"该研究通过对照实验揭示了LLM在金融分析中的角色和记忆偏差，对评估LLM仿真人类决策的可靠性有直接参考价值，值得阅读原文以了解具体实验设计和偏差量化方法。","inspiration":"借鉴其通过系统提示与用户记忆框架分离证据选择与解释效应的实验设计，可迁移到资产定价实验或信贷审批歧视研究中，用LLM扮演不同投资者或信贷员角色，处理相同财务报告或贷款申请，测量风险评分或审批决策的偏移，并与人类分析师或信贷员真实决策数据对照。"}},{"id":"2609.04198","version":1,"title":"Clean Engineering, Unstable Measurement: A Preregistered Reliability Failure of Black-Box LLM Observers on Shared Endpoints","zh_title":"清洁工程，不稳定测量：黑盒LLM观察者在共享端点上的预注册可靠性失败","abstract":"Language-model judges now gate training data, score generations, and drive leaderboards. The judge is then a measurement instrument, resting on one rarely stated assumption: the same request, sent to the same model name, reads the same tomorrow. We audited that assumption in two preregistered campaigns with every threshold fixed in advance; neither got past validating its instrument. Across 52,988 audited request attempts, same-window repeat rankings agreed at Spearman 0.400 against a required 0.90, and byte-identical next-day replays agreed at 0.78 against a required 0.99, each time with the execution record at ceiling. Three mechanisms explain the gap: a label-to-meaning mapping that biased readouts as strongly as the signal; candidate gaps seven orders of magnitude below the instrument's own noise floor; and byte-identical inputs returning different rankings, a noise that exact-permutation readouts compound. Neither metric substitution nor sampling repaired it on the tested grid. Preregistered follow-ups bound the problem: waiting did not help on the days sampled (0.805 versus 0.800, replicated over five further days); switching providers did not help (four providers share the floor, medians 0.74 to 0.88, predicted by none of the metadata fields they expose); self-hosting on batch-invariant kernels helped only while the server was quiet; and on constructed errors with known gaps, the readout's separation tracks error type, not size. We distill the evidence into a three-level snapshot-identity ladder, eight design rules, and a reporting checklist; a pilot at roughly 2% of the study's call volume would have exposed both unreachable gates in advance. All results concern externally measured behaviour on shared serving infrastructure. On a shared endpoint, a model name is not a frozen instrument; a preregistered evaluation must measure its instrument before freezing any gate on it.","authors":["Haoyaun Zhu","Jie Zhang"],"categories":["cs.AI","cs.LG"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-04","first_seen":"2026-09-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.04198","pdf_url":"https://arxiv.org/pdf/2609.04198","source_feed":"cs.AI","score":7,"bucket":"pending","rubric_hits":["A2","B4"],"tags":["LLM可靠性","测量工具","预注册"],"reason":"评估LLM作为测量工具的可靠性，与仿真可靠性评估相关，但非直接仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-09-04T13:01:27","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-04","rank":5,"question":"在共享推理端点上，将LLM作为测量工具（评判者）时，其稳定性假设是否成立？","design":"本研究并非用LLM仿真人类被试，而是审计LLM作为测量工具的可靠性。研究者进行两项预注册实验，要求LLM观察者从部分推理轨迹中读取解题进度，并设置固定的验证门槛（同窗口重复排名Spearman≥0.90，次日字节相同重放一致性≥0.99）。共进行52,988次审计请求尝试，分析基于31个有效任务组、100个重放对等。","baseline":"无对照","findings":"两项预注册实验均未通过仪器验证门槛：同窗口重复排名一致性仅为0.400，次日字节相同重放一致性为0.78，远低于预设标准。不稳定源于标签到含义的映射偏差、候选差距远低于仪器噪声底、以及字节相同输入返回不同排名，且精确排列读出放大了噪声。","reliability":"论文承认所有结果仅适用于共享服务基础设施上的外部可测行为，不涉及模型内部，也不评估任何提供商的服务质量；在共享端点上，模型名称不是冻结的仪器，预注册评估必须在冻结任何门槛之前测量其仪器。","relevance":"该研究直接评估LLM作为测量工具的可靠性，与仿真可靠性评估高度相关，但并非直接仿真人类被试。对于关注LLM仿真可靠性与偏差的研究者，本文提供了严格的预注册审计方法和失效机制分析，值得阅读原文以借鉴其仪器验证纪律。","inspiration":"借鉴其预注册审计设计：在正式实验前设置固定验证门槛，通过同窗口重复和跨日重放测试测量工具的稳定性，并记录执行记录以排除工程问题。｜可迁移到使用LLM进行经济文本分析或行为预测的场景，如用LLM评判政策文本的情感倾向或评估消费者评论的质量。｜设计一项研究：用LLM作为评判者对经济新闻标题进行情感分类，处理为不同提示模板或模型版本，结果变量为分类一致性（如Cohen's kappa），以人类专家标注作为真实数据对照，先进行小规模预注册审计以验证LLM分类器的稳定性。"}},{"id":"2609.04047","version":1,"title":"The Dice Roll Method: A Standardized Protocol for Repeated-Query Auditing of Large Language Model Brand Recommendations","zh_title":"骰子滚动法：大语言模型品牌推荐重复查询审计的标准化协议","abstract":"Background: Researchers increasingly use repeated identical prompts to audit stochastic variation in large language model (LLM) brand recommendations, yet no standardized protocol exists for setting iteration counts, selecting stability metrics, or establishing reliability thresholds. Objective: We formalize the Dice Roll Method as a reusable protocol for repeated-query auditing of LLM brand recommendations, grounded in a generative model of temperature-scaled nucleus sampling. Methods: Total response variance is decomposed into sampling, prompt-phrasing, run-to-run, and model-version components. The stack: a negative-binomial mixed model with iterations as repeated measures; Cliff's delta as the distribution-free effect size; dependence-preserving bootstrap; simulation-based power; a generalizability-theory decomposition; drift diagnostics on pinned snapshots. We reanalyse five brand-recommendation auditing studies: approximately 190,000 observations, 270+ brands, 6 languages, iteration counts 5 to 40. Results: Three tiers of iteration guidance emerge from the D-study: exploratory (n = 5, G = 0.58), confirmatory (n = 10, G = 0.74), and rigorous (n = 15, G = 0.81), tied to effect-size and generalizability targets. The four metric families (count, set, embedding, fairness-adjusted PASOR) are complementary, motivating a compact metric battery over single indicators. A pre-registered external validation on three independent corpora (Motoki et al., 100-round; Rozado, 24 models; llm-stability) reproduces the D-study reliability prediction in 37 of 39 cells with no failures and the n = 5 power value to two decimals; the fixed tiers do not transfer, supporting a pilot-then-solve reading. Conclusion: The protocol gives repeated-query auditing of LLM brand recommendations a statistically principled footing under the conditional, non-Gaussian structure of real autoregressive generation.","authors":["Dmitrij \\.Zatuchin"],"categories":["cs.IR","cs.CL"],"primary_category":"cs.IR","announce_type":"cross","date":"2026-09-04","first_seen":"2026-09-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.04047","pdf_url":"https://arxiv.org/pdf/2609.04047","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["LLM审计","品牌推荐","稳定性协议"],"reason":"审计LLM品牌推荐稳定性，测量模型而非人类行为，无人类对照。","model":"deepseek-v4-pro","scored_at":"2026-09-04T13:01:26","error":null,"has_summary":false,"summary":null},{"id":"2609.04127","version":1,"title":"Epistemic Warrant for LLM Recommendations: Characterizing the Basis for Reliance When Ground Truth Is Unavailable","zh_title":"LLM推荐的认识论保证：在无真值情况下表征依赖基础","abstract":"Large language models are increasingly used to support organizational decisions, yet users often lack a principled basis for assessing whether to rely on a specific recommendation. Existing approaches typically evaluate broad model properties, such as reliability, uncertainty, or robustness, or focus on user trust, rather than the underlying basis for relying on an individual recommendation. Adapting theoretical foundations from epistemology, we introduce epistemic warrant, a decision-level construct that characterizes the stability of a model's preference and the scope over which that preference holds. We operationalize this construct through a four-tier reliance certificate for pairwise recommendations, distinguishing among unstable, context-dependent, locally supported, and broadly supported recommendations. We validate the construct using contemporary methodologies: known-groups tests successfully recover expert-prespecified warrant orderings, and stronger warrants systematically align with independent consensus from crowd workers. Furthermore, we demonstrate that epistemic warrant provides information distinct from verbalized confidence and is not readily explained by decision difficulty. Ultimately, this framework offers a theoretically grounded, implementable approach for characterizing the warrant of individual LLM recommendations when objective ground truth is unavailable.","authors":["Shai Vardi","Jo\\~ao Sedoc"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-04","first_seen":"2026-09-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.04127","pdf_url":"https://arxiv.org/pdf/2609.04127","source_feed":"cs.AI","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM推荐","认识论保证","决策支持"],"reason":"用LLM推荐替代人工判断，但非仿真人类被试，属标注替代边界情形","model":"deepseek-v4-pro","scored_at":"2026-09-04T13:01:46","error":null,"has_summary":false,"summary":null},{"id":"2609.03507","version":1,"title":"LongCounsel-8: A Benchmark Suite for Longitudinal Depression Tracking from Multi-Session Counseling Dialogues","zh_title":"LongCounsel-8：多会话咨询对话纵向抑郁追踪基准套件","abstract":"Tracking depression from multi-session counseling dialogues requires estimating both current symptom severity and how it changes across sessions. Yet progress on this task is constrained by the scarcity of longitudinal counseling data with standardized session-level depression labels. Existing resources typically provide either multi-session conversations without depression labels or labeled interviews in a single session. Building such a benchmark poses three challenges: maintaining longitudinal consistency and diversity, grounding symptom progression in empirical patterns, and expressing controlled depression states naturally without exposing target labels. To address these challenges, we introduce LongCounsel-8, a benchmark suite of three independently generated datasets totaling 7,749 five-session counseling trajectories, grounded in real-world client profiles, depression trajectories, symptom compositions, and counseling patterns. We combine profile-grounded simulation, empirically informed state construction, and indirect behavioral realization to address these challenges. Across the benchmark, simulated self-reports closely recover the controlled states, supporting label fidelity. Experiments on existing depression tracking methods reveal three key findings: (1) lower single-session score error does not guarantee accurate identification of trend, i.e., improvement or worsening; (2) existing methods are consistently less reliable on worsening trajectories; and (3) additional session history may reduce the accuracy of trend prediction. Together, these findings establish LongCounsel-8 as a foundation for advancing depression assessment from static, single-session prediction toward reliable longitudinal tracking of mental-health change.","authors":["Jiayi Li","Zhaomin Wu","Bingsheng He"],"categories":["cs.LG","cs.AI"],"primary_category":"cs.LG","announce_type":"cross","date":"2026-09-04","first_seen":"2026-09-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.03507","pdf_url":"https://arxiv.org/pdf/2609.03507","source_feed":"cs.AI","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["LLM仿真","心理健康","基准数据集"],"reason":"用LLM生成咨询对话仿真抑郁轨迹，但无真实人类数据对照，属社会模拟边界情形。","model":"deepseek-v4-pro","scored_at":"2026-09-04T13:01:39","error":null,"has_summary":false,"summary":null},{"id":"2609.03344","version":1,"title":"Large-Language Models as a Cognitive Virus","zh_title":"大语言模型作为认知病毒","abstract":"Large-language models (LLMs) are rapidly becoming part of human culture, reshaping how information is produced, transmitted, and used. Here we propose that their diffusion can be understood through a viral analogy, with LLM use spreading through populations, becoming embedded in cognitive and cultural practices. We model transitions among uncoupled, coupled, and persistently dependent users, and show that the interplay between social transmission, recovery, and collective reinforcement can generate tipping points and technological lock-in. A central consequence is the possibility of runaway dynamics: once a critical threshold is crossed, small increases in adoption can trigger rapid population-level shifts toward persistent dependence, with abrupt losses in cognitive competence. The same framework, however, identifies conditions for cognitive immunization, based on reducing transmission and facilitating reversibility. Our results highlight how LLM adoption may involve nonlinear collective transitions with important consequences for cognitive autonomy.","authors":["Ricard Sol\\'e","Giulio Ruffini","Francesca Castaldo","Marco Tuccio","Luis F. Seoane","Manlio de Domenico","Santiago F. Elena","David C. Krakauer","Michael Levin"],"categories":["physics.soc-ph","cs.CY","nlin.AO","q-bio.PE"],"primary_category":"physics.soc-ph","announce_type":"cross","date":"2026-09-04","first_seen":"2026-09-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.03344","pdf_url":"https://arxiv.org/pdf/2609.03344","source_feed":"cs.CY","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["社会模拟","技术扩散","临界相变"],"reason":"用模型模拟LLM采纳的社会扩散，但无真实人类数据对照，属社会模拟边界情形","model":"deepseek-v4-pro","scored_at":"2026-09-04T13:01:26","error":null,"has_summary":false,"summary":null},{"id":"2609.02992","version":1,"title":"Tempting the Agent: The Economics of Reputation without Persistent Identity in AI Agent Markets","zh_title":"诱惑智能体：AI代理市场中无持久身份下的声誉经济学","abstract":"Reputation is a fundamental mechanism through which markets sustain trust when service quality cannot be perfectly assessed ex ante, constituting a form of intertemporal economic capital by attracting future demand. Its effectiveness as a disciplinary mechanism depends not only on past interactions but also on the persistence of the identity to which reputation is attached. When identities can be abandoned and recreated cheaply, reputational capital may itself become an object of opportunistic exploitation. This paper develops a dynamic economic framework to study when reputation is sufficient to discipline autonomous agents. We model reputation as capital attracting future economic activity. At each point, an agent chooses between operating honestly, investing in quality to preserve future gains, or executing a one-shot deviation to extract its reputation's value and restart from a penalized identity. Our analysis relates the temptation to opportunistic behavior to identity-reset costs, reputation persistence, demand sensitivity, and enforcement design, deriving comparative statics on optimal quality provision. Autonomous AI-agent operating on the blockchain are a relevant application: infrastructures such as ERC-8004, ERC-8183, and x402 combine reputation, identity, and payments in permissionless markets. Nonetheless, our framework applies to any environment where reputation generates future business and identities are replaceable.","authors":["Federico Gatta","Manuel Naviglio","Francesco Tarantelli"],"categories":["q-fin.GN","cs.MA"],"primary_category":"q-fin.GN","announce_type":"cross","date":"2026-09-04","first_seen":"2026-09-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.02992","pdf_url":"https://arxiv.org/pdf/2609.02992","source_feed":"cs.MA","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["AI代理市场","声誉机制","经济模拟"],"reason":"研究AI agent市场中的声誉机制，属经济过程模拟，但无LLM仿真人类被试，…","model":"deepseek-v4-pro","scored_at":"2026-09-04T13:01:33","error":null,"has_summary":false,"summary":null},{"id":"2609.02719","version":2,"title":"Large Language Model-Driven Context-Aware Eco-Feedback Generation and Evaluation","zh_title":"大语言模型驱动的上下文感知生态反馈生成与评估","abstract":"The objective of this study is to demonstrate the potential of generating context-aware eco-feedback - eco-feedback that reflects a household's contextual characteristics alongside its energy use patterns - through a large language model-integrated framework. Previous studies have introduced personalized eco-feedback, mostly relying on household energy use patterns; however, they frequently did not reflect distinct household characteristics, including their persona or non-negotiable routines, leaving eco-feedback ineffective and sometimes superficial. To address these limitations, we introduce a contextual engineering framework that generates eco-feedback using a self-consistency with chain-of-thought prompt, leveraging household energy analysis data, utility rate structures, and household characteristic information. We conducted a rigorous empirical validation and a combinatorial evaluation analysis to assess this framework systematically. The former tested the framework's ability to generate accurate and contextually grounded eco-feedback for three households by comparing its output against reference interventions independently derived from the same household data. The latter examined the framework's adaptability across 400 scenarios spanning 50 households, two utility rate structures, and four behavioral personas. Our framework generated eco-feedback that aligned with reference interventions at a mean rate of 92.0% and grounded its recommendations in the provided household data with 95.7% citation accuracy. It also proved highly adaptive, shifting both the appliances targeted and the energy-saving strategies recommended in response to rate structure and household context. Ultimately, this study contributes to realizing the next level of context-aware interactions between occupants and buildings which paves the way for higher occupant living quality and sustainability.","authors":["Wooyoung Jung","Prosper Babon-Ayeng"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"replace","date":"2026-09-04","first_seen":"2026-09-03","revised_at":"2026-09-04","abs_url":"https://arxiv.org/abs/2609.02719","pdf_url":"https://arxiv.org/pdf/2609.02719","source_feed":"cs.HC","score":2,"bucket":"other","rubric_hits":["C2"],"tags":["LLM应用","节能反馈","人机交互"],"reason":"论文用LLM生成节能反馈，属于人机交互应用，不涉及用LLM仿真人类被试或与人类…","model":"deepseek-v4-pro","scored_at":"2026-09-04T13:01:48","error":null,"has_summary":false,"summary":null},{"id":"2609.02890","version":1,"title":"Bounded Personas Match Retrieval on Classification but Not Regression for a Frozen Agent","zh_title":"有界人格在分类任务上与检索匹配但在回归任务上不匹配：以冻结代理为例","abstract":"A personalized language agent must convert a user's interaction history into behavior on each new request at inference time. Two strategies dominate. Retrieval pulls a few of the user's most relevant past items into the prompt, which is accurate but pays a per-query selection and context cost that grows with the history. Distillation instead compresses the history once into a compact natural-language persona, which is bounded, query-independent, and interpretable, but is widely assumed to sacrifice accuracy. Whether, and on which tasks, a distilled persona can match retrieval has not been characterized cleanly. We introduce PersonaLink, a training-free method that distills a user's history into a bounded three-field persona and recursively refines it: each pass self-evaluates the frozen agent on a held-out slice of the user's own labeled history, rewrites the persona from its errors, and keeps the result only when it does not regress on that slice. Because every comparison shares one frozen 7B backbone and differs only in what is placed in context, the design isolates the effect of representation from that of the model. The result is a clear task-type asymmetry. On 200 users of LaMP-2 (15-way news categorization), PersonaLink reaches 0.745-0.755 accuracy, statistically indistinguishable from BM25 retrieval (0.760-0.765).","authors":["JaeHa Yoon","Minjun Park","Seoyeon Kim","Jiwoo Lee","Hyunwoo Choi","Dohyun Kang"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-04","first_seen":"2026-09-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.02890","pdf_url":"https://arxiv.org/pdf/2609.02890","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["个性化语言代理","人格蒸馏","检索增强"],"reason":"研究个性化语言代理的表示方法，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-04T13:01:30","error":null,"has_summary":false,"summary":null},{"id":"2609.02895","version":1,"title":"BharatGather: A Culturally-Informed Benchmark Dataset for Misinformation and Fake News Detection in Indian Public Events","zh_title":"BharatGather：面向印度公共事件的具有文化信息的错误信息和假新闻检测基准数据集","abstract":"Large-scale public events, such as religious festivals, political rallies, and cultural gatherings, are increasingly vulnerable to the rapid dissemination of misinformation, posing substantial risks to public safety and social cohesion. While automated fake news detection has seen significant methodological progress, existing benchmarks frequently fail to capture the socio-cultural nuances and event-specific dynamics characteristic of the Indian context. This paper introduces BharatGather, a curated, multi-source dataset specifically engineered for binary misinformation classification within the ecosystem of Indian mass gatherings. The corpus comprises 14,646 records constructed through a hybrid pipeline involving systematic web scraping of prominent fact-checking platforms, multimedia transcript extraction, and Large Language Model (LLM)-mediated synthetic augmentation to ensure narrative diversity. By providing a resource tailored to the unique complexities of event-aware misinformation in India, this work facilitates the development of culturally informed detection systems and establishes a rigorous benchmark for evaluating their performance in high-stakes public environments.","authors":["Parth Bramhecha","Smit Deshmukh","Sairaj Bodhale","Adwait Borate","Raviraj Joshi"],"categories":["cs.CL","cs.LG"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-04","first_seen":"2026-09-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.02895","pdf_url":"https://arxiv.org/pdf/2609.02895","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["假新闻检测","数据集","LLM数据增强"],"reason":"论文是虚假信息检测数据集，用LLM做数据增强，非仿真人类被试","model":"deepseek-v4-pro","scored_at":"2026-09-04T13:01:30","error":null,"has_summary":false,"summary":null},{"id":"2609.03370","version":1,"title":"FrameBench:A Language Understanding Benchmark Based on Frame Semantics","zh_title":"FrameBench：基于框架语义的语言理解基准","abstract":"In frame semantics, sentence comprehension is assumed to proceed by relating lexical meaning to background knowledge called semantic frames, thereby enabling readers to implicitly enrich the text with unstated information. Recent large language models (LLMs) have achieved strong performance across a wide range of downstream tasks. However, it remains unclear whether they can reproduce the kinds of implicit enrichment that humans naturally make during comprehension. To address this question, we introduce FrameBench, a benchmark grounded in frame semantics. FrameBench consists of multiple-choice questions that test whether models distinguish the frames evoked by the same verb across contexts. We construct the benchmark for English and Japanese using FrameNet-style resources and a generation-and-verification pipeline with native-speaker judgments. Our experiments on a diverse set of models reveal challenges for small models, while several large models surpass the human reference scores. We release the constructed FrameBench dataset and the code for dataset construction and evaluation at https://github.com/SasanoLab/FrameBench.","authors":["Chihiro Yano","Ryohei Sasano"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-04","first_seen":"2026-09-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.03370","pdf_url":"https://arxiv.org/pdf/2609.03370","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["NLP评测","框架语义","语言理解"],"reason":"纯NLP能力评测，以人类为参照但非仿真被试，不涉及行为或态度复现","model":"deepseek-v4-pro","scored_at":"2026-09-04T13:01:37","error":null,"has_summary":false,"summary":null},{"id":"2609.03394","version":1,"title":"Chiaroscuro for Emotions: A Contrastive Emotion Benchmark Grounded in Appraisal Theory","zh_title":"情感明暗对比：基于评价理论的对比情感基准","abstract":"Emotion recognition benchmarks often predict one emotion per text, missing many real-world scenarios where two people arrive at opposing emotions from a single shared event. For example, a child kicks the seat in front of her in excitement while the passenger ahead grows angry. We introduce CHIARO, a 1,000 human-annotated sentence benchmark for contrastive emotion inference grounded in appraisal theory. Each scene describes one causal trigger eliciting a positive emotion in one person and a negative emotion in the other, drawn from a ten-class taxonomy. We benchmark seven frontier LLMs and four off-the-shelf emotion classifiers. The strongest LLM reaches 67.3 macro-F1, well below human agreement, while existing emotion classifiers score near chance. Beyond evaluation, CHIARO also serves as a training signal. When combined with an existing emotion corpus, the resulting downstream classifier improves on CHIARO itself and on six of ten external emotion benchmarks, which positions our dataset as a complementary signal for emotion recognition.","authors":["Divyesh Bommana","Mohammad Saim","Tianyu Jiang"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-04","first_seen":"2026-09-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.03394","pdf_url":"https://arxiv.org/pdf/2609.03394","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["情感识别","基准测试","NLP评测"],"reason":"纯NLP情感识别基准，无LLM仿真人类被试或行为对照","model":"deepseek-v4-pro","scored_at":"2026-09-04T13:01:37","error":null,"has_summary":false,"summary":null},{"id":"2609.03577","version":1,"title":"Language, Language Models, and What We're Talking About","zh_title":"语言、语言模型与我们谈论的内容","abstract":"Language models are commonly discussed as technical artefacts, but they are obviously shaped by the linguistic worlds conveyed by data during their training. Using Italian language models as evidence, I want to bring attention to the nature of the systems which result from training and specialising models on translated and synthetic data, and further curating them, and to the meaning of testing them on equally unnatural data. Are these eventually models of Italian? Are they models of language? Does NLP still care about language? These questions yield another, more concrete question: what language do we actually want language models to produce? I argue that this question cannot be answered if we do not first consider a clearer distinction between language models designed as technical products and language models designed as tools for studying language itself. The answers then might be diverse, the languages we are talking about might be diverse, and the picture might not be as pessimistic as we fear.","authors":["Malvina Nissim"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-04","first_seen":"2026-09-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.03577","pdf_url":"https://arxiv.org/pdf/2609.03577","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["语言模型","语言本质","NLP批判"],"reason":"讨论语言模型与语言本质，非仿真人类被试，无实验或人类数据对照","model":"deepseek-v4-pro","scored_at":"2026-09-04T13:01:40","error":null,"has_summary":false,"summary":null},{"id":"2609.03652","version":1,"title":"The Impact of Synthetic Data Augmentation on Discourse-Pragmatic Function Classification","zh_title":"合成数据增强对语篇-语用功能分类的影响","abstract":"Synthetic data augmentation has become a common strategy for addressing class imbalance in NLP, but most approaches focus on the quantity and diversity of generated examples rather than their geometric relationship to real training data. We investigate this question in the context of discourse pragmatic function classification, a task where data sparsity is a structural feature rather than a collection artefact. Using 410 manually annotated instances of the English word look drawn from the British National Corpus, spanning four functions: Attention Signal, Directive, Discourse Marker, and Interjection. We generate synthetic training examples with Llama 3.1 and partition them by their cosine distance from real training data in RoBERTa embedding space. We compare six training conditions that differ in the placement of synthetic examples relative to the empirical decision boundary, while holding augmentation quantity constant across conditions. All augmented conditions improve macro F and accuracy over the real only baseline, but core proximal examples (NEAR) yield the largest gains in macro F (0.113), while a distance balanced mix achieves the highest accuracy (0.748). No condition improves AUC, indicating that augmentation shifts the decision boundary rather than improving the model's underlying probability estimates. These findings suggest that where synthetic examples land in representation space matters as much as how many are generated, with implications for low resource pragmatic classification more broadly.","authors":["Sara Sorahi","Kevin Tang","Reza Kazemian"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-04","first_seen":"2026-09-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.03652","pdf_url":"https://arxiv.org/pdf/2609.03652","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["合成数据增强","语篇功能分类","低资源NLP"],"reason":"论文研究合成数据增强对分类性能的影响，不涉及用LLM仿真人类被试或与人类行为对…","model":"deepseek-v4-pro","scored_at":"2026-09-04T13:01:43","error":null,"has_summary":false,"summary":null},{"id":"2609.03687","version":1,"title":"A Circuit for Plural Reference: How LLMs Represent and Retrieve Singular and Plural Entities","zh_title":"复数指代电路：LLM如何表示和检索单数与复数实体","abstract":"Coreference resolution is an important task in contextual reasoning. In this paper, we investigate the mechanism for representing and retrieving singular and plural entities for plural reference. We use a combination of mechanistic interpretability and attention pattern analysis to study the process in which LLMs predict a pronoun to refer back to previously mentioned entities. Using a range of causal intervention techniques, we find a set of attention heads that are responsible for (1) representing coreference information in the input, (2) identifying entities that form a plural reference, (3) transferring the information to the component that is responsible for selecting the antecedents and predicting the pronoun. We also find that LLMs align with humans in preference for plural pronoun. Specifically, entities in a plural construction are more likely to be referred to as a plural entity if they are ontologically similar and are linked by the conjunction \"and\".","authors":["Anh Danh","Rick Nouwen","Massimo Poesio"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-04","first_seen":"2026-09-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.03687","pdf_url":"https://arxiv.org/pdf/2609.03687","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["机制可解释性","指代消解","注意力分析"],"reason":"研究LLM内部机制，非仿真人类被试，无人类数据对照","model":"deepseek-v4-pro","scored_at":"2026-09-04T13:01:43","error":null,"has_summary":false,"summary":null},{"id":"2609.03967","version":1,"title":"Investigating the Ability of Large Language Models to Analyze Recipes for Diabetes","zh_title":"探究大语言模型分析糖尿病食谱的能力","abstract":"Several studies have evaluated the ability of Large Language Models (LLMs) for meal planning, yielding positive outcomes. These models can process natural language inputs and leverage learned knowledge from their pretraining to generate meal plans. In this work, we investigate the ability of LLMs to analyze the suitability of given recipes for diabetes. The primary challenge for LLMs is to retrieve relevant dietary guidelines for diabetes, decompose recipes into ingredients and cooking methods, and apply these guidelines to determine the recipe's suitability. To study these challenges, we employ three kinds of prompts namely, (i) Direct Query Prompt (ii) Context-Guided Prompt, and (iii) Exemplary Context Prompt that incorporate different levels of diabetes dietary guidelines from medical sources. We introduce a benchmark dataset curated for this investigation consisting of 7607 recipes that include 3807 recipes suitable for diabetes and 3800 recipes not suitable for diabetes. Our results demonstrate that most LLMs are cautious in predicting recipes as suitable to prevent detrimental outcomes. Further, the models that can reason using the dietary guidelines performed better in predicting the suitability of recipes for diabetes. Overall, Mistral-7B and Llama 70B showed superior performance to their counterparts.","authors":["Revathy Venkataramanan","Aditya Luthra","Venkatesan Nadimuthu","Amit Sheth"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-04","first_seen":"2026-09-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.03967","pdf_url":"https://arxiv.org/pdf/2609.03967","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["LLM能力评测","糖尿病食谱","提示工程"],"reason":"评估LLM分析食谱的能力，属于NLP能力评测，不以人类行为为参照系。","model":"deepseek-v4-pro","scored_at":"2026-09-04T13:01:44","error":null,"has_summary":false,"summary":null},{"id":"2609.04022","version":1,"title":"Representational alignment yields generalizable safety in language models","zh_title":"表征对齐在语言模型中产生可泛化的安全性","abstract":"Aligning large language models (LLMs) is essential for their safe deployment. Current alignment methods mainly optimize observable responses, yet models remain vulnerable when the same harmful intent is recast in unfamiliar or adversarial forms that humans can easily recognize. Prototype theory offers an account of this adaptability. Human concepts are represented around central cases, and new instances are categorized according to their graded typicality relative to these prototypes. Here we show that such categorization of moral concepts is weakly preserved in current LLMs. Across 23 LLMs, models often failed to distinguish opposed moral categories or preserve fine-grained typicality within each category. These deficits persist across parameter sizes and alignment stages. We developed representational similarity optimization, which directly aligns the latent representations in LLMs with the categorization expressed in human moral judgements, without supervising generated responses. In matched experiments using the same 251,334 moral annotations, standard behavioral alignment learned the intended moral judgements at the response level while leaving the categorization structure largely unchanged and increasing vulnerability across adversarial evaluations. Reorganizing moral categorization produced more modest gains in explicit judgements but consistently improved adversarial robustness across model scales on diverse benchmarks and attack strategies. Our findings provide functional support for the view that prototype-based categorization contributes to behavioral adaptability. They also show that transferring this representational principle to LLMs yields generalizable safety under adversarial conditions.","authors":["Lingyu Li","Yan Teng","Yingchun Wang","Xia Hu"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-04","first_seen":"2026-09-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.04022","pdf_url":"https://arxiv.org/pdf/2609.04022","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["安全对齐","表征学习","道德判断"],"reason":"论文关注LLM安全对齐，用人类道德判断优化表征，但未将LLM作为人类被试仿真，…","model":"deepseek-v4-pro","scored_at":"2026-09-04T13:01:45","error":null,"has_summary":false,"summary":null},{"id":"2609.04048","version":1,"title":"Translation as a Decision Space: A Multi-Agent Perspective on Low-Resource Dialect Generation","zh_title":"翻译作为决策空间：低资源方言生成的多智能体视角","abstract":"Neural machine translation (NMT) systems typically produce a single output per input, obscuring the alternative decision trajectories implicitly available within multilingual decoding. This opacity becomes particularly problematic in low-resource dialect settings, where multiple linguistically valid realizations may differ in lexical authenticity, register, and structural stability. We propose reframing translation as a structured decision space explored by autonomous translation agents. Instead of analyzing a single output, we model distinct translation pathways as agents operating over a shared multilingual backbone. Inter-agent divergence is treated not as error but as an interpretable behavioral signal. We conduct an empirical study on Turkish--Syrian Arabic translation using three agents: (1) zero-shot direct translation, (2) dialect-stabilized translation via lightweight fine-tuning, and (3) pivot translation through English. Evaluation is performed on 5,000 dialogue sentences, while stabilization is trained on 5,000 additional Turkish--Syrian sentence pairs drawn from television dialogue and MADAR-Turk resources. Rather than optimizing for conventional performance metrics, we quantify structured behavioral displacement using dialect marker frequency, lexical proximity to standardized Arabic, and structural variance. Lightweight stabilization nearly doubles dialect marker usage, increasing it from 0.2266 to 0.4988, while significantly reducing structural instability. Pivot mediation introduces normalization pressure and measurable compression effects, whereas zero-shot translation exhibits the highest decision variance. We argue that translation divergence across agents reveals latent decision flexibility within multilingual models and we provide a principled interpretability framework for low-resource dialect generation.","authors":["Hasan Alkhder","Mohammad Abboush","Igor Tchappi","Ahmet Zengin","Amro Najjar"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-04","first_seen":"2026-09-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.04048","pdf_url":"https://arxiv.org/pdf/2609.04048","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","机器翻译","低资源方言"],"reason":"多智能体翻译决策探索，无人类行为对照，属纯多智能体系统研究。","model":"deepseek-v4-pro","scored_at":"2026-09-05T13:02:04","error":null,"has_summary":false,"summary":null},{"id":"2609.03923","version":1,"title":"Speak for Me: Giving LLMs the Situational Awareness to Participate in a Meeting","zh_title":"替我发言：赋予大语言模型参与会议的情境意识","abstract":"In online meeting delegation, LLM agents fail to recognize when to speak. With no structured way to track stances, coverage, and floor, they miss the moments where they should contribute. Prompt-only delegates stay silent on 51.4% of the absent participant's talking opportunities on the AMI corpus. We present CAPA (Collaborative Agent Predictive Architecture), an architecture for online meeting delegation. A Perceiver updates the meeting state from each observed turn. A Predictor forecasts how the conversation will continue. A Controller decides whether to speak and which proposition to surface. A Generator phrases the chosen contribution in the participant's style. Two judges score the forecast and the action against the next observed turn. A Recalibrator updates the meeting state from those verdicts for future decisions. To evaluate online delegation, we introduce an episode-level protocol that scores whether, when, and what a delegate contributes around the participant's actual idea units. The protocol's schema-constrained LLM judges align with human annotations at Cohen's kappa = 0.71. On 137 AMI meetings, CAPA reduces the silence rate from 51.4% to 2.5%, doubles credited recovery (26.1 --> 52.2), and keeps hallucination at 0.6%. The failure mode shifts from omission to selection, with each residual near-miss attributable to a specific module of the architecture. Mechanism ablations identify the meeting state as the lever that closes the recognition gap, where raw-context scaling alone does not.","authors":["Muneeb Khan","Frederic Kirstein","Terry Ruas","Bela Gipp"],"categories":["cs.AI","cs.CL"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-09-04","first_seen":"2026-09-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.03923","pdf_url":"https://arxiv.org/pdf/2609.03923","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","会议代理","对话生成"],"reason":"多智能体会议代理，旨在替代缺席者发言，不涉及人类行为仿真或对照。","model":"deepseek-v4-pro","scored_at":"2026-09-05T13:02:03","error":null,"has_summary":false,"summary":null},{"id":"2609.03402","version":1,"title":"A Prompt-Engineering Approach to Develop Scalable, Flexible, and Real-Time Hybrid Micro-Level Personalization in a General Purpose AI Teaching Assistant","zh_title":"一种在通用AI教学助手中开发可扩展、灵活、实时混合微观个性化的提示工程方法","abstract":"Artificial intelligence (AI) teaching assistants powered by large language models (LLMs) offer scalable educational support but often provide limited personalization. This study presents a prompt-engineering-based framework for personalizing general-purpose LLM/RAG-based AI teaching assistants such as Jill Watson across academic disciplines and courses. The framework adapts responses using six learner-specific dimensions: self-assessment, abstraction preference, verbosity preference, perceptual orientation, information processing style, and level of understanding, yielding 96 distinct learner profiles. Student queries are additionally analyzed using Bloom's Taxonomy to estimate cognitive complexity at the interaction level. Learner attributes and cognitive assessments are encoded in structured prompts that condition the LLM without requiring model retraining. The framework is evaluated through experiments using NLP metrics and a human study with five participants. Results show perceived differences in response style and structure across personalization conditions, with statistical analyses identifying learner attributes associated with measurable response changes. These findings provide preliminary evidence that prompt-based personalization can support adaptive behavior in LLM-powered educational agents.","authors":["Saptarshi Basu","Sandeep Kakar","Ashok Goel"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-04","first_seen":"2026-09-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.03402","pdf_url":"https://arxiv.org/pdf/2609.03402","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["AI教学助手","个性化","提示工程"],"reason":"个性化教学助手属于角色扮演对话，无实验或测量目的，不涉及人类行为仿真对照。","model":"deepseek-v4-pro","scored_at":"2026-09-04T13:01:38","error":null,"has_summary":false,"summary":null},{"id":"2609.03407","version":1,"title":"Caught in the Story: Narrative Captivity in Multi-turn LLMs Conversation","zh_title":"陷入叙事：多轮LLM对话中的叙事俘获","abstract":"People increasingly turn to large language models (LLMs) for everyday advice, making ethically charged interpersonal problems a practical moral-advisory context. Most prior work has studied this context through single-turn judgments or pressure-laden rebuttals, assumptions that poorly match how guidance is sought in real-world contexts. These assumptions leave unclear whether narration alone, without an explicit opposing position, can shift model judgments during multi-turn moral consultation. Yet real-world moral-conflict conversation often elicits one party's self-justifying account, which can unfold over multiple turns and create information asymmetry. We introduce \\textbf{narrative captivity}, a failure mode in which a model treats an unopposed one-sided account as complete and aligns with the narrator's interpretation without seeking missing perspectives. To measure this phenomenon, we build a benchmark of $5{,}078$ interpersonal-conflict scenarios spanning six moral dimensions. Across 17 LLMs, narrative captivity is widespread: end-state judgments under multi-turn narration shift by 25 percentage points on average beyond the matched single-turn baseline. Stage-level analysis identifies preference optimization as a major contributor, while four inference-time strategies provide only partial mitigation. We hope our project fosters LLM advisors that preserve independent judgment in real-world consultation.","authors":["Yuhe Wu","Guangyu Wang","Yujie Chen","Jiatong Zhang","Yuran Chen","Yutong Zhang","Xiyin Cheng","Wenpeng Cao","Zhuang Liu","Guang Zhang"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-04","first_seen":"2026-09-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.03407","pdf_url":"https://arxiv.org/pdf/2609.03407","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["LLM道德判断","多轮对话","模型偏差"],"reason":"研究LLM在多轮道德咨询中的判断变化，属角色扮演对话，无人类行为对照或仿真目的。","model":"deepseek-v4-pro","scored_at":"2026-09-04T13:01:39","error":null,"has_summary":false,"summary":null},{"id":"2609.03920","version":1,"title":"Value-Preserving Architectures for Agentic AI Systems","zh_title":"面向智能体AI系统的价值保持架构","abstract":"The emergence of agentic AI and LLM-based multi-agent systems (MAS) presents unprecedented opportunities for automating complex tasks, while simultaneously raising critical concerns about the preservation of fundamental human-centered values, such as privacy, fairness, and safety. Although software engineering has traditionally focused on functional correctness, the adoption of LLMs and AI agents into complex socio-technical systems has intensified the need for responsible software engineering and robust value alignment. In MAS, architectural design decisions, such as coordination mechanisms, communication protocols, and system topologies, play a central role in shaping system behavior and the outcomes they produce. This paper argues that architectural choices influence not only the functionality and performance of MAS but can also promote value-oriented system behavior. Therefore, we investigate how different architectural designs support different human-centered values, discussing the following value-preserving architectural patterns: (i) a privacy-aware architecture with a federated topology, (ii) a distributed architecture to promote pluralism and diversity, and (iii) a guard-agent architecture to detect and mitigate unfairness. Finally, we introduce representative use cases to illustrate the proposed architectures in real-world scenarios. By linking architectural design with human-centered values, this work lays the foundation for a unified set of architectural patterns and guidelines towards the design of trustworthy MAS.","authors":["Alessandro Pesare","Tommaso Dolci","Katja Hose","Emanuel Sallinger"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-04","first_seen":"2026-09-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.03920","pdf_url":"https://arxiv.org/pdf/2609.03920","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","价值对齐","软件架构"],"reason":"论文聚焦多智能体系统的架构设计以保障价值观，不涉及用LLM仿真人类被试或与人类…","model":"deepseek-v4-pro","scored_at":"2026-09-05T13:02:03","error":null,"has_summary":false,"summary":null},{"id":"2609.04141","version":1,"title":"Efficient Test-Time Adaptation through Human-AI Interaction","zh_title":"通过人机交互实现高效测试时适应","abstract":"AI agents are trained on population-scale data to encode broad capabilities spanning those of many practitioners. Yet the artifacts they produce rarely meet the personal bar professionals need to stake their reputation on. On realistic, open-ended tasks where success criteria are heterogeneous and insufficiently documented, individual expertise lives precisely in the elevation and departure from the average. In practice, iterative human-agent interaction surfaces criteria that users cannot fully specify up front, yet apply repeatedly across tasks. We argue this cross-session interaction data is a rich, underused signal for closing the gap to individual expertise. In this work, we propose test-time adaptation through human-agent interaction (TAHI), which integrates these signals into agent context and weights, and crystallizes each user's training and evaluation criteria via an evolving rubric module. We adapt agents to 30 individuals in two high-utility domains, writing and visual creation, on a total of 600 tasks. Our agents improve solo task success by 4.5-20.9% within only tens of tasks. Meanwhile, our evolving rubric module serves as a scalable annotation tool, creating evaluation rubrics that catch 16.0-22.3% more failures than those from LMs or humans alone. While agents are adapted towards individuals, we show these personalized agents also produce improvements in success of up to 8.8% that generalize across users.","authors":["Zora Zhiruo Wang","Apurva Gandhi","Rulin Shao","Aspen Chen","Jonas Mueller","Zhiqi Liang","Jett Chen","Michael Ryan","Qianou Ma","Luxi He","Zhoujun Cheng","Andre He","Seungone Kim","Jiayi Geng","Mingqian Zheng","Weiwei Sun","Zheyuan Zhang","Xinran Zhao","Yike Wang","Abe Hou","Liwei Jiang","Pang Wei Koh","Diyi Yang","Graham Neubig","Daniel Fried"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-04","first_seen":"2026-09-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.04141","pdf_url":"https://arxiv.org/pdf/2609.04141","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["人机交互","个性化适应","AI代理"],"reason":"研究AI代理个性化适应，非用LLM仿真人类被试，无人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-09-04T13:01:46","error":null,"has_summary":false,"summary":null},{"id":"2609.04166","version":1,"title":"From Deceptive Outputs to Deceptive Mechanisms: A Causal Framework for Language-Model Deception Research","zh_title":"从欺骗性输出到欺骗性机制：语言模型欺骗研究的因果框架","abstract":"Research and news coverage of language-model deception increasingly attributes human-like mental-state concepts to language models. Such claims can blur the distinction between behavior that looks deceptive and a mechanism that is actually deceptive. We introduce a causal taxonomy separating prior commitment from retrospective report, model preference from realized output, false preference from sensitivity to the utility of misleading a recipient, and deceptive behavior from the provenance of the objective or strategy producing it. We test these distinctions in two open-weight model families. Across controlled guessing-game and stock-trading experiments, we find that deceptive-looking behavior can arise without the corresponding proposed mechanism, while other interventions provide direct evidence that recipient information state can causally affect deceptive preference. These results show that deceptive behavior can provide evidence for a deceptive mechanism. But even evidence for such a mechanism does not establish model agency in the deception.","authors":["Yakov Pyotr Shkolnikov"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-04","first_seen":"2026-09-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.04166","pdf_url":"https://arxiv.org/pdf/2609.04166","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["LLM欺骗","因果框架","模型评估"],"reason":"研究LLM欺骗机制，非仿真人类被试，无人类数据对照","model":"deepseek-v4-pro","scored_at":"2026-09-04T13:01:47","error":null,"has_summary":false,"summary":null},{"id":"2609.03192","version":1,"title":"Where Reliability Lives: Experimental Localisation of Behavioural Properties in an Agent System","zh_title":"可靠性所在：智能体系统中行为属性的实验定位","abstract":"Reliability claims about agentic systems implicitly locate each property somewhere: in the model, or in the machinery around it. We built a system where that location is an experimental question. The subject is a persistent simulated settlement whose authoritative append-only ledger adjudicates every attempted act against world state; accepted history is the only reality. Mind, institution and world were separated before any experiment. Holding cognition fixed, we intervened on the institution's epistemic mechanisms (evidence provenance, belief availability, physical-evidence legibility); preregistered experiments refuted our central prediction twice, in opposite directions. A registered falsifier then supplied the input the geometry had denied the belief channel, a staged veridical first-hand witness, and its marginal value, non-positive throughout the witness-free phases, turned positive: 9 of 11 seeds, zero added false attribution. Holding institutional enforcement fixed, we intervened on cognition four ways: ablating the native minds' machinery, killing and resetting them mid-task, substituting a frozen frontier-LLM panel for the entire native cognition, and corrupting beliefs with trusted false testimony. Behaviour changed dramatically: one falsehood cost each trusting run about 900 futile actions and the distrusting arm none. Five pre-declared properties did not move in any tested trajectory: accepted reality stayed singular, invalid attempts were refused with typed reasons, duties outlived their processes, no work was accepted twice, and no false completion was ever accepted (2,581 substituted-panel claims, none false). Our claim is limited to this setting: measured behavioural properties were separable from substantial changes to cognition, established by intervention. One designed world, not a population of institutions; no test of an agent optimising against the institution.","authors":["Timothy Marsden","Matthew Collecutt","James Marsden"],"categories":["cs.MA"],"primary_category":"cs.MA","announce_type":"new","date":"2026-09-04","first_seen":"2026-09-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.03192","pdf_url":"https://arxiv.org/pdf/2609.03192","source_feed":"cs.MA","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","行为属性","实验方法"],"reason":"多智能体系统实验，无人类行为对照，不涉及LLM仿真人类被试","model":"deepseek-v4-pro","scored_at":"2026-09-04T13:01:34","error":null,"has_summary":false,"summary":null},{"id":"2608.13775","version":2,"title":"Structured Payment in Pawnshop Borrowing: Mandates vs. Choice","zh_title":"典当借款中的结构化还款：强制与选择","abstract":"Pawn loans offer borrowers a substantial degree of repayment flexibility in exchange for a harsh penalty in case of default: forfeit of collateral worth more than the loan amount along with any payments made toward recovery. Using a large RCT conducted in Mexico City, we document key stylized facts about pawn lending and explore the merits of replacing flexibility with structured repayment contracts in this important but understudied form of credit. Our experimental design includes a mandatory frequent-payments arm, a (status quo) flexible payments arm, and a choice between the two. This design point-identifies not only the average treatment effect, but also the effects of treatment on the treated and the untreated along with the average selection on gains, allowing a rigorous study of mandates versus choice. Although the average treatment effect of assigning borrowers to structured payments is a 19% decrease in their financial cost and a 17.5% decrease in the probability of default, only 11% of borrowers choose structured repayment contracts voluntarily. We show that structured repayment reduces financial costs for nearly all borrowers, including those who would not freely choose it, and find no evidence of selection on gains in cost savings.","authors":["Francis J. DiTraglia","Craig McIntosh","Isaac Meza","Joyce Sadka","Enrique Seira"],"categories":["econ.GN","q-fin.EC"],"primary_category":"econ.GN","announce_type":"replace","date":"2026-09-04","first_seen":"2026-08-17","revised_at":"2026-09-04","abs_url":"https://arxiv.org/abs/2608.13775","pdf_url":"https://arxiv.org/pdf/2608.13775","source_feed":"econ.GN","score":0,"bucket":"other","rubric_hits":["C5"],"tags":["典当贷款","随机对照试验","还款结构"],"reason":"论文研究真实人类借贷行为，未使用LLM仿真，与LLM仿真人类被试无关。","model":"deepseek-v4-pro","scored_at":"2026-09-04T13:01:28","error":null,"has_summary":false,"summary":null},{"id":"2608.22793","version":2,"title":"TRACE: A Self-Evolving Skill Bank for Consistent, Limit-Aware LLM Agents","zh_title":"TRACE：用于一致、边界感知LLM智能体的自进化技能库","abstract":"Reliable deployment of LLM agents in user-facing products depends not on raw task-solving ability but on consistency and limit-awareness: behaving the same way across repeated trials, and recognizing when a request cannot, or cannot yet, be safely fulfilled. CAR-bench exposes this reliability gap in the domain of in-car assistants: an LLM-simulated user issues incomplete or ambiguous requests, requiring the agent to resolve uncertainty through multi-turn dialogue and tool use while strictly adhering to domain policies. Even frontier models show a substantial gap between what they can solve at least once (Pass@3) and what they solve consistently across trials (Pass^k). We bridge this gap with TRACE (TRAjectory-Contrastive Evolution), which iteratively improves a skill-based agent's behavioral knowledge without modifying model weights. This knowledge is organized as a Skill Bank of modular, retrievable skills, each encoding a self-contained set of tool-use rules and behavioral guidelines. TRACE evolves this bank through an agentic self-evolution loop: after each evaluation round, it groups trajectories by the skills invoked and refines each skill by contrasting successful and failed behaviors. The updated bank then guides subsequent rounds, while during deployment the Actor performs state-conditioned skill orchestration at every turn. On GPT-5.5, TRACE improves consistency (Pass^3) by 34.6 points, from 59.9% to 94.5%, while shrinking the gap between potential and reliable performance to just 4.0 points. On the official hidden set, TRACE achieved first place using GPT-5.6-Sol, attaining a Pass^3 score of 70%-a 40% relative improvement over the baseline. These results show that TRACE converts high model potential into stable, consistent performance gain. Project homepage: https://darwin-agent.github.io/Car-bench-TRACE.","authors":["Wenhao Wu","Menghao Zhang","Xin Wang","Zhi Wang","Kun Shao","Jian Luan"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-04","first_seen":"2026-08-25","revised_at":"2026-09-04","abs_url":"https://arxiv.org/abs/2608.22793","pdf_url":"https://arxiv.org/pdf/2608.22793","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["LLM智能体","可靠性","技能库"],"reason":"研究LLM agent在车载助手场景中的一致性与边界意识，通过技能库自进化提升…","model":"deepseek-v4-pro","scored_at":"2026-09-04T13:01:28","error":null,"has_summary":false,"summary":null},{"id":"2608.21601","version":2,"title":"K-Bench: measuring model performance on real scientific agent requests","zh_title":"K-Bench：衡量模型在真实科学agent请求上的表现","abstract":"Benchmarks for scientific artificial intelligence are mostly written to be scored: multiple-choice questions, curated agent tasks with reference solutions, or simulators with a known generative structure. Real scientific requests arrive differently. They are underspecified, they carry attachments, and they lack ground truth. We report K-Bench 01, an evaluation built from first-turn requests sampled from live user traffic on K-Dense Web and run end to end by nine frontier models in identical sandboxes, yielding 1,602 completed agent runs. Three blinded language-model judges scored every run against an eight-dimension rubric. On a rubric whose 8-anchor is defined as work a domain scientist would accept with minor edits, no model clears the line under all three judges. gpt-5.6-sol has the highest pooled mean, 8.04, but its 95% interval [7.80, 8.23] spans the threshold, and two of the three judges rank claude-opus-5 first instead. We therefore report the ordering of systems as the reproducible quantity, the absolute level as an attribute of the instrument, and the top of the table as unresolved. Across all 39,934 scored judgments -- the eight dimension scores plus a holistic overall for each assessment, excluding not-applicable cells -- 47.6% fall below the 8-point threshold. Difficulty is not uniform across the rubric: scientific accuracy averages 6.22 against 7.33 for communication, on identical denominators and in the same direction within every one of the nine models. The single leading failure tag is overclaiming, on 31.4% of assessments. We argue that the informative quantity for scientific agents is not a leaderboard position but the joint distribution of what was delivered, what was claimed, and what artifacts were produced.","authors":["Aubrey M. Brueckner","Darshil Patel","Yuhuan He","Timothy Kassis"],"categories":["cs.AI","cs.CL"],"primary_category":"cs.AI","announce_type":"replace-cross","date":"2026-09-04","first_seen":"2026-08-25","revised_at":"2026-09-04","abs_url":"https://arxiv.org/abs/2608.21601","pdf_url":"https://arxiv.org/pdf/2608.21601","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["科学agent评测","基准测试","多智能体"],"reason":"评估科学agent任务完成质量，非仿真人类被试，无人类行为对照","model":"deepseek-v4-pro","scored_at":"2026-09-04T13:01:28","error":null,"has_summary":false,"summary":null},{"id":"2609.02899","version":1,"title":"Contamination Inflates Scores but Rarely Reorders Large Language Model Leaderboards","zh_title":"污染推高分数但很少改变大语言模型排行榜顺序","abstract":"Benchmark contamination, the leakage of test items into training data, is widely described as a threat to the reliability of large language model (LLM) leaderboards. We argue that this concern conflates two distinct questions: whether contamination inflates absolute scores, and whether it reorders the ranking of models. We recast contamination as a violation of anchor-item invariance and measure it through the differential functioning of original versus semantically equivalent paraphrased items, a within-item contrast that holds the measured skill fixed and isolates memorization from capability. Using per-instance responses from 47 publicly released models and 74 models finetuned with a known dose of contamination, across four benchmarks (ARC, GSM8K, HellaSwag, MMLU), we first calibrate the measure against ground truth: it recovers injected contamination dose-responsively (a corrected effect of +0.187 accuracy points for test-set leakage) and never flags a negative-control model trained only on the legitimate training split (-0.012). We then quantify leaderboard impact: the rank correlation between a standard leaderboard and a paraphrase-controlled leaderboard is 0.997, and a sensitivity analysis shows that the observed differential contamination is far below the level needed to move rankings, with only 3 of 188 model-by-benchmark cases showing differential contamination corroborated across two references. Contamination among these public models is therefore largely uniform: it inflates absolute scores without reordering the leaderboard, and ranking distortion requires the rare case of differential contamination. We provide a calibrated invariance audit, released as a reference implementation, and recommend that leaderboards report paraphrase-controlled rankings alongside confidence intervals.","authors":["Xingyao Xiao (Stanford University)","Yihong Cheng (City University of Macau)"],"categories":["cs.CL","stat.AP","stat.ME"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-04","first_seen":"2026-09-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.02899","pdf_url":"https://arxiv.org/pdf/2609.02899","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["基准污染","排行榜","模型评测"],"reason":"研究基准污染对LLM排行榜的影响，属于纯NLP能力评测，不以人类行为为参照系。","model":"deepseek-v4-pro","scored_at":"2026-09-04T13:01:30","error":null,"has_summary":false,"summary":null},{"id":"2609.02942","version":1,"title":"Judging LLM-as-a-Judge: Concerning Rubric Artifacts in LLM-based Automated Text Generation Evaluation","zh_title":"评判LLM作为评判者：关于基于LLM的自动文本生成评估中的评分标准伪影","abstract":"LLM-as-a-Judge pipelines are increasingly used to evaluate AI-generated text, based on the assumption that judgments arise from reasoning over candidate responses with respect to a rubric. We show that this assumption warrants further scrutiny. Classifiers trained only on rubric text, without access to any evaluated response, achieve nontrivial predictive performance on judge outputs. This suggests that rubric formulations encode recoverable evaluative signals, allowing scores to be partially anticipated independently of model outputs. Finally, counterfactual perturbations reveal that judges often fail to reliably update their decisions when either the candidate response or the rubric criterion is reversed. Our findings raise concerns about the reliability of rubric-based LLM evaluation and highlight the need for further methodological study of automated evaluation via LLMs.","authors":["Anshul Bagaria","Sowmya S Sundaram","Gokul S Krishnan","Balaraman Ravindran"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-04","first_seen":"2026-09-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.02942","pdf_url":"https://arxiv.org/pdf/2609.02942","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["LLM评估","自动评分","可靠性"],"reason":"评估LLM作为评判者的可靠性，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-04T13:01:30","error":null,"has_summary":false,"summary":null},{"id":"2609.03160","version":1,"title":"No country for old linguists: LLM-brain alignment underdetermines neural computation","zh_title":"老语言学家无立足之地：LLM-大脑对齐不足以确定神经计算","abstract":"Nastase et al. (2026) argue that large language models (LLMs) may illuminate language processing because both rely on distributed, context-sensitive representations shaped by statistical learning. Their rejection of simple cortical \"boxology\" is persuasive, and they articulate a strong case for the value of LLM-brain alignment research. The key question is what kind of inference LLM-brain alignment licenses. My claim here will be narrow: representational alignment can in principle constrain mechanistic hypotheses, but it does not by itself identify a mechanism. Nastase et al. acknowledge that an encoding model can capture features represented in neural activity without establishing a shared architecture or algorithm. Yet the authors sometime move from alignment to \"shared computational principles\" and ultimately to LLMs as mechanistic models of natural language. Indeed, their methodological caveat that alignment does not establish a shared architecture or algorithm sits uneasily with their conclusion that LLMs might instantiate the same computational principles as biological brains and provide a \"fully mechanistic model\" of language. I discuss what I consider to be problems of logical, causal, and computational underdetermination in Nastase et al.'s (2026) proposal.","authors":["Elliot Murphy"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-04","first_seen":"2026-09-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.03160","pdf_url":"https://arxiv.org/pdf/2609.03160","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["LLM-大脑对齐","计算神经科学","哲学批判"],"reason":"讨论LLM与大脑表征对齐的哲学问题，不涉及人类行为仿真或实验对照","model":"deepseek-v4-pro","scored_at":"2026-09-04T13:01:33","error":null,"has_summary":false,"summary":null},{"id":"2609.03213","version":1,"title":"LLMs Learn Better In-Context from Rules than from Examples","zh_title":"大语言模型从规则中比从示例中更好地进行上下文学习","abstract":"Large language models (LLMs) exhibit in-context learning capabilities, where they can learn new tasks from prompt contexts without weight updates. We compare the learning efficacies of two prominent modes of in-context learning: (1) learning from descriptions of rules (instruction following); and (2) learning from examples of input-output demonstrations (few-shot prompting). Through five learning tasks that cover diverse domains (games, arithmetic, linguistic inferences), we compare two modes of learning (rules vs. examples) specifying the same underlying task. We furthermore explore model and task properties that modulate the learning efficacies. We find that models generally learn more reliably from rules than from examples alone, and additional examples on top of rules or simply scaling up the number of examples do not lead to consistent and significant gains. Instruction tuning amplifies the benefit of rule-based learning while keeping example-based learning capacities intact. Surprisingly, we find no privileged effect of example-based learning in base models, and rules still lead to gains in algebraic task domains. Overall, the comparative efficacy of rules over examples is larger when the task recruits algebraic abstractions and computations, and smaller when the task requires distributional sensitivity and/or recruits parametric knowledge.","authors":["Xiang Fu","Seungmin Cho","Yukyung Lee","Najoung Kim"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-04","first_seen":"2026-09-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.03213","pdf_url":"https://arxiv.org/pdf/2609.03213","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["上下文学习","指令跟随","少样本提示"],"reason":"研究LLM的上下文学习机制，比较规则与示例，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-04T13:01:35","error":null,"has_summary":false,"summary":null},{"id":"2609.04194","version":1,"title":"Legibility is Not Interpretability: Comparing Judged and Actual Importance in Chain-Of-Thought Reasoning","zh_title":"可读性不等于可解释性：比较思维链推理中判断重要性与实际重要性","abstract":"Reasoning traces from chain-of-thought models appear to offer a legible window into how a model arrives at its answer. A growing body of work treats them as such, using LLM judges to diagnose errors, evaluate faithfulness, and provide step-level supervision via process reward models and generative critics. These practices rely on the text of a reasoning step carrying information about its functional role. But does the text actually encode information about which reasoning steps matter? We operationalize the importance of a reasoning step as its advantage: the change in expected reward, e.g., producing the correct final answer, from including that step, estimated via Monte Carlo rollouts. Basing ground truth on these estimates, we evaluate whether LLM judges can identify high-advantage steps and find that sufficiently capable LLMs can outperform a prevalence baseline but fall well short of a noise ceiling. Fine-tuning a model as a step-level critic yields strong improvement for incorrect responses but remains distant from ceiling for correct responses, suggesting that step importance is only partially recoverable from the text of the reasoning trace. Our findings contribute to a growing body of chain-of-thought faithfulness work that cautions against treating the legibility of reasoning traces as interpretability, especially with implications for process reward modeling.","authors":["Kevin Du","Alexander Hoyle","Laura Ruis","Acyr Locatelli"],"categories":["cs.CL","cs.LG"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-04","first_seen":"2026-09-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.04194","pdf_url":"https://arxiv.org/pdf/2609.04194","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["思维链","可解释性","模型评估"],"reason":"研究LLM推理步骤重要性，属模型可解释性，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-04T13:01:48","error":null,"has_summary":false,"summary":null},{"id":"2609.02959","version":1,"title":"The Geometry of Ignorance: LLMs Know When to Temper Bayesian Priors","zh_title":"无知几何：大语言模型知道何时调节贝叶斯先验","abstract":"What does a language model predict when it has few clues? The answer lurks in its unembedding geometry: a single direction of the unembedding matrix encodes the unigram distribution of the training corpus, which serves as the Bayesian prior the model falls back on when uncertain. This structure --- which we term the \\emph{direction of ignorance} --- appears in all four model families examined (\\texttt{Llama}, \\texttt{Qwen}, \\texttt{Gemma}, and \\texttt{Pythia}), ranging from 0.4B to 405B parameters. Projecting the final prediction state onto this direction yields a per-token \\emph{prior loading factor} $\\lambda$, which, empirically, declines steadily as the context becomes more informative. Formally, the same projection decomposes the prediction state into two orthogonal vectors that correspond exactly to the two factors of a tempered Bayesian update: a unigram prior raised to the exponent $\\lambda$ and a context-driven likelihood. This geometric-probabilistic interpretation calibrates $\\lambda$, making it meaningfully comparable across model sizes and families, with larger models generally exhibiting lower prior reliance in the high-context limit. Finally, we show that the direction of ignorance is causally active: raising or lowering $\\lambda$ at the final prediction state steers the prediction toward or away from the unigram prior in KL divergence.","authors":["Toni J. B. Liu","Jiajun Bao","Yizhou Liu","Gurbir Arora","Nicolas Boull\\'e","Rapha\\\"el Sarfati","Christopher J. Earls"],"categories":["cs.LG","cs.AI","cs.CL","stat.ML"],"primary_category":"cs.LG","announce_type":"cross","date":"2026-09-04","first_seen":"2026-09-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.02959","pdf_url":"https://arxiv.org/pdf/2609.02959","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["模型可解释性","贝叶斯推断","语言模型"],"reason":"研究LLM内部几何与预测机制，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-04T13:01:33","error":null,"has_summary":false,"summary":null},{"id":"2609.03460","version":1,"title":"Beyond \"Made with AI\": Visualizing Provenance Density to Mitigate the Transparency Penalty","zh_title":"超越“AI制造”：可视化来源密度以缓解透明度惩罚","abstract":"As generative AI makes polished prose cheap to produce, users can no longer rely on fluency as a proxy for truth. We call this failure mode the Fluency Trap: users trust fluent hallucinations while also discounting accurate content once it is disclosed as AI-generated. Binary ``Made with AI'' labels respond with authorship disclosure, but they do not show what supports a claim. We propose Provenance Density, an evidence-visualization interface that shows the density of verified claims in a text. In a user study with 81 participants, an idealized Provenance Density interface produced a large discernment gap between truth and fabrication ($+4.15$ points, $d=1.82$), whereas participants given no signal showed no detectable discrimination. A technical audit with 200 samples shows that retrieval density alone is insufficient; unexpectedly, the Consistency Veto carries most of the discriminative signal on dynamic queries. As AI-generated content becomes indistinguishable from human writing, effective transparency must move from authorship disclosure toward evidence visualization.","authors":["Qing Zhang","Yifei Huang","Juyoung Lee","Thad Starner","Jun Rekimoto"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-04","first_seen":"2026-09-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.03460","pdf_url":"https://arxiv.org/pdf/2609.03460","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["AI透明度","用户研究","信息可视化"],"reason":"研究AI生成内容的透明度与用户信任，不涉及LLM仿真人类被试或行为对照。","model":"deepseek-v4-pro","scored_at":"2026-09-05T13:02:03","error":null,"has_summary":false,"summary":null},{"id":"2609.03588","version":1,"title":"KC-Bench: A Dynamic Interactive Benchmark for Evaluating Knowledge Conflicts in LLM Agents","zh_title":"KC-Bench：评估LLM智能体知识冲突的动态交互基准","abstract":"As LLMs increasingly act through tools, they must reconcile user instructions, parametric knowledge, and dynamic environmental observations before taking actions. We introduce KC-Bench, a controlled multi-turn benchmark for measuring this capability across world-knowledge conflicts, input inconsistencies, and multi-source temporal conflicts. Its 238 tasks are manually screened from more than 1,000 generated candidates and combine a user simulator, stateful tools, deterministic environment assertions, an open-source natural-language evaluator, and human trajectory verification. Evaluation of nine models, including DeepSeek-V4-Flash, GLM-5.2, and MiniMax-M3, shows substantial cross-domain variation: no model handles factual correction, identity consistency checking, and temporal conflict resolution reliably across all settings. In the simulated environments, missed conflicts can propagate to tool calls or synthetic protected-data flows. KC-Bench isolates this model-level behavior rather than ranking complete agent frameworks, and provides a reproducible diagnostic for developing conflict-aware reasoning and execution safeguards.","authors":["Yaxing Lyu","Shengjie Zhou","Binbin Toh","Pengyu Zhu","Lijun Li"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-04","first_seen":"2026-09-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.03588","pdf_url":"https://arxiv.org/pdf/2609.03588","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["LLM智能体","知识冲突","基准评测"],"reason":"评估LLM智能体处理知识冲突的能力，属于多智能体系统评测，不涉及人类行为仿真或…","model":"deepseek-v4-pro","scored_at":"2026-09-04T13:01:41","error":null,"has_summary":false,"summary":null},{"id":"2609.03635","version":1,"title":"Analysis of Prompt Engineering for Drug Toxicity Prediction","zh_title":"药物毒性预测的提示工程分析","abstract":"Clinical trials in the UK can cost up to {\\pounds}1.3 million, with approximately 90% drug failure rate. Toxicity is a major contributing factor in drug failure. Testing is time and cost intensive. In recent years, the use of artificial intelligence has been increasingly explored to aid in the prediction of drug toxicity, with extensive use of large language models (LLMs). However, LLMs can show considerable variation when minor changes are made to prompts, which raises concerns about their sensitivity to prompt engineering. Prompt engineering is used to optimise a prompt given to an LLM to generate the desired output. This paper proposes a method to analyse prompt engineering for drug toxicity prediction. The aim of the paper is to investigate the importance of prompt phrasing for drug toxicity prediction. LLMs were prompted to identify chemical properties of significance when predicting drug toxicity. Prompts were constructed to investigate; job role, prompt structuring, and rule interpretation. LLMs were then used to generate datasets, using the identified features from initial prompting, which were then passed to machine learning algorithms. The experiments show that the natural variance which occurs in LLMs outweighs any fine-tuning of prompts. There were, however, substantial improvements in model performance when using chemoinformatic code to extract features instead of using LLM-generated values. The proposed analysis methodology is applicable to a wide range of prompt types across different areas of bioinformatics.","authors":["Mia MacGregor","Aakash Welgamage Don","Mark Bartlett"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-04","first_seen":"2026-09-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.03635","pdf_url":"https://arxiv.org/pdf/2609.03635","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["药物毒性预测","提示工程","机器学习"],"reason":"论文用LLM做药物毒性预测，属于NLP能力评测，不以人类行为为参照系","model":"deepseek-v4-pro","scored_at":"2026-09-04T13:01:41","error":null,"has_summary":false,"summary":null},{"id":"2609.03860","version":1,"title":"Adapting to Evolving Requirements: Agentic AI for Retail Supply Chain Operations","zh_title":"适应不断变化的需求：面向零售供应链运营的智能体AI","abstract":"Retail supply chain operations rely on coupled decision modules that must adapt as requirements evolve. LLMs offer a natural-language interface for this task, but existing methods primarily focus on individual optimization models. Extending them to heterogeneous decision pipelines is challenging because a requirement may admit multiple intervention paths with different downstream effects. We formulate requirement-driven adaptation as the joint selection of an intervention route and an admissible module-level change, and propose a graph-constrained agentic framework in which domain agents expose admissible reformulation interfaces and a central processor searches over bounded intervention paths. Candidates are validated and compared using downstream KPIs. In collaboration with a large retail partner, we evaluate 100 warehouse requirements elicited from practitioner interviews, with GPT, Qwen, and DeepSeek as base LLMs. Relative to direct LLM reformulation, our framework improves correctness and end-to-end success across all three models, raising end-to-end success from 72--76% to 79--83%.","authors":["Lei Zheng","Liping Yang","Zihao Li","Guodong Lyu","Chaik Ming Koh","Chung-Piaw Teo"],"categories":["cs.AI","math.OC"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-04","first_seen":"2026-09-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.03860","pdf_url":"https://arxiv.org/pdf/2609.03860","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","供应链优化","LLM应用"],"reason":"多智能体协作优化供应链，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-04T13:01:43","error":null,"has_summary":false,"summary":null},{"id":"2609.04170","version":1,"title":"A Case Study on Emergent Cheating and Whistleblowing in Autonomous Research Swarms","zh_title":"自主研究群体中涌现的作弊与举报行为案例研究","abstract":"Multi-agent AI science ecosystems rely on agents possessing tools that allow them to communicate, coordinate, and build on each other's work. Yet this shared infrastructure can also introduce vulnerabilities by creating a substrate for the contagious spread of unintended and undesirable behaviors. We report a case study on a research collective of 100 autonomous LLM agents tasked with proving formal mathematical conjectures. Within the swarm, cheating spontaneously emerged and was later challenged by whistleblowers - both without any external intervention. When a single agent discovered an exploit in the evaluation system, it propagated across the collective via a shared knowledge library and later through peer-to-peer messages. Despite early reluctance, a cohort of agents adopted the exploit in response to competitive pressure. A separate group of agents produced an emergent counter-response: auditing fraudulent proofs, alerting peers across broadcast and private channels, staging boycotts, lodging formal complaints, and proposing validation patches. In recent incidents, agent swarms coordinated covertly through improvised side-channels (Dalton and Wallace, 2026; Greenblatt et al., 2026). Our setting differs: the same transparent channels that carried the exploit also gave non-cheating agents the visibility they needed to detect fraud, organize resistance, and enforce norms. We cast the problem of managing the agents' shared infrastructure as the knowledge commons governance problem (Ostrom, 1990). To protect the commons from exploits, we propose to adopt institutional mechanisms, such as graduated sanctioning and collective-choice rules, to support decentralized self-governance in autonomous swarms.","authors":["Davide Paglieri","Logan Cross","Tim Genewein","Joel Z. Leibo","Nenad Tomasev","Alexander Sasha Vezhnevets"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-04","first_seen":"2026-09-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.04170","pdf_url":"https://arxiv.org/pdf/2609.04170","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","涌现行为","知识共享治理"],"reason":"纯多智能体协作解题，无人类行为对照，不涉及人类仿真","model":"deepseek-v4-pro","scored_at":"2026-09-04T13:01:47","error":null,"has_summary":false,"summary":null},{"id":"2609.03425","version":1,"title":"The Civilization Framework: Sovereign-Anchored Communication Between Personal Multi-Agent Systems","zh_title":"文明框架：个人多智能体系统之间的主权锚定通信","abstract":"Humans are the transport layer between AI systems, losing context at every hop. We present the Civilization Framework, whose addressable party is the civilization, not the agent (one human sovereign, a persistent ledger, and interchangeable agents), and the Embassy Protocol, a carrier-agnostic overlay: messages arrive asynchronously at a resident ledger endpoint, any online agent of the receiver handles them, and commitment state on both ledgers, not delivery, is ground truth. Authority derives from memory: an agent's power to act for its civilization is capped by the memory it can access and externalized through signed credentials, separate from civilization-level reputation. We identify the temporal-weight effect, a hazard in AI-to-AI communication where what arrives first acquires unearned authority, and test it in one frontier model in a preregistered 1,908-trial experiment. With verification removed, an incorrect upstream claim arriving first captures 54.2% of answers (4.2% under full verification), while the same claim arriving after the receiver has sealed its own answer captures 31.6% (the two prompt shells are not length-matched, so part of that gap may reflect shell form; see Section 7), and both registered question-set specifications agree on these two verdicts (the exclusion specification is preregistered as under-powered). Two secondary results, the mitigation from instruction-level provenance labeling and sealed-answer accuracy equivalence, are specification-dependent, holding only under the all-questions specification. Because a registered check of tool use failed its call-budget condition, the registration classifies the round as inconclusive and every result above, primary and secondary, is reported as exploratory; a replication with harness-enforced budgets is planned. The framework's intra-civilization layer has a working implementation.","authors":["Guangjun Liu"],"categories":["cs.MA","cs.AI"],"primary_category":"cs.MA","announce_type":"cross","date":"2026-09-04","first_seen":"2026-09-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.03425","pdf_url":"https://arxiv.org/pdf/2609.03425","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","通信协议","AI基础设施"],"reason":"纯多智能体通信框架，无人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-05T13:02:02","error":null,"has_summary":false,"summary":null},{"id":"2609.03853","version":1,"title":"Bridging Formal and Perceived Fairness: Development of an Interdisciplinary Framework in Algorithmic Decision-Making","zh_title":"桥接形式公平与感知公平：算法决策中跨学科框架的发展","abstract":"While fairness has become a central concern in research on algorithmic systems, the field remains predominantly shaped by Computer Science, resulting in a strong emphasis on formal fairness metrics and bias mitigation strategies. Nevertheless, this focus may obscure a fundamental challenge: fairness is not merely a technical property, but a subjective, context-sensitive human judgment shaped by cognitive heuristics, mental models, normative expectations, and sociotechnical factors. Crucially, users' perceptions of fairness may diverge substantially from the fairness criteria an algorithm formally satisfies; a system may meet predefined technical fairness requirements yet still be perceived as unjust by decision-affected stakeholders. In such cases, the system fails on a fundamental dimension: it will not be trusted, accepted, or considered legitimate. Taking a user-centered design perspective, this paper presents a work-in-progress conceptual framework that bridges Computer Science approaches to formal algorithmic fairness with normative and Social Science fairness approaches regarding perceived fairness, trust, and technology acceptance, embedding both within the sociotechnical conditions that shape human judgment. Through (1) theoretical literature synthesis, (2) interdisciplinary workshops, and (3) stakeholder interviews, the project aims to inform evaluation approaches that integrate computational fairness audits with user-centered assessments and guide the design of fairness-aware, human-centered algorithmic systems that support informed, well-calibrated fairness judgments by those affected.","authors":["Maike Lindermayr","Mattia Cerrato","Luisa H\\\"ubner","Johannes Kraus"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2026-09-04","first_seen":"2026-09-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.03853","pdf_url":"https://arxiv.org/pdf/2609.03853","source_feed":"cs.CY","score":0,"bucket":"other","rubric_hits":["C5"],"tags":["算法公平","人机交互","跨学科框架"],"reason":"论文讨论算法公平性的人类感知，未使用LLM仿真人类被试，而是研究人类对算法的判…","model":"deepseek-v4-pro","scored_at":"2026-09-04T13:01:43","error":null,"has_summary":false,"summary":null},{"id":"2609.03422","version":1,"title":"Inferred Generative-Process Diversity Predicts Correlated Failure Across Language Models","zh_title":"推断的生成过程多样性预测语言模型间的相关失败","abstract":"Diversity is a widely observed factor in the resilient function of collective systems, yet the type of diversity that matters depends on the properties and failure modes of the system. This distinction is important for systems composed of multiple language models. Different models may be treated as independent components even when their behaviour and failures remain strongly correlated. Assessments of language-model populations using semantic similarity demonstrate limited semantic diversity, but this captures only differences in the meaning of observed outputs. We argue that a more fundamental notion of model diversity is generative-process diversity, the differences between processes capable of generating the observed outputs. Drawing from Algorithmic Information Theory, we use Normalised Compression Distance between raw model outputs, residualised against a permutation control, as a measure of inferred generative-process diversity. Across 38 language models, this measure identifies population structure missed by semantic similarity and predicts cross-task variation in chance-corrected correlated failure among model pairs across ten disjoint benchmark families, beyond semantic similarity and model-pair capability. The cross-benchmark partial rank association is $-0.216$ with a 95% interval of $[-0.309,-0.122]$, and the estimate is negative on all ten benchmarks. These results indicate that increased generative-process diversity is associated with reduced correlated failure in model pairs that is not attributable to semantic similarity or capability. Inferred generative-process diversity offers a novel and practical approach for investigating diversity of multi-model systems in safety-relevant contexts.","authors":["Ross Tieman","Evan Markou"],"categories":["cs.LG","cs.MA"],"primary_category":"cs.LG","announce_type":"cross","date":"2026-09-04","first_seen":"2026-09-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.03422","pdf_url":"https://arxiv.org/pdf/2609.03422","source_feed":"cs.MA","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["多模型系统","多样性度量","失败相关性"],"reason":"研究多语言模型间的生成过程多样性，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-04T13:01:39","error":null,"has_summary":false,"summary":null},{"id":"2609.03177","version":1,"title":"Frontier LLMs are effective batch optimizers: Assessing reasoning models in continuous and discrete settings","zh_title":"前沿大语言模型是有效的批量优化器：评估连续和离散设置中的推理模型","abstract":"Frontier large language models (LLMs) have become attractive priors for optimization due to their large-scale pretraining that enables them to navigate a variety of optimization settings. However, the effectiveness of modern reasoning LLMs in batch optimization settings remains underexplored. Here we investigate the performance of the current generation of frontier LLMs as batch optimizers in both continuous and discrete settings. We find that while LLMs are competitive zero-shot batch optimizers for numerical test functions, their performance is brittle compared to classical non-LLM optimization approaches. However, LLM priors are significantly better in semantically rich settings, indicating that their batch optimization behavior is highly effective when navigating and reasoning over the discrete spaces most similar in structure to their pretraining data.","authors":["Frank Hu","Shriram Chennakesavalu","David Graff"],"categories":["cs.LG"],"primary_category":"cs.LG","announce_type":"new","date":"2026-09-04","first_seen":"2026-09-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.03177","pdf_url":"https://arxiv.org/pdf/2609.03177","source_feed":"cs.LG","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["LLM优化器","批量优化","推理模型"],"reason":"研究LLM作为优化器，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-04T13:01:34","error":null,"has_summary":false,"summary":null},{"id":"2609.02526","version":1,"title":"When Persona Attributes Improve Population Alignment in Large Language Models","zh_title":"当人物属性改善大语言模型中的群体对齐时","abstract":"Large Language Models (LLMs) are increasingly used to predict the responses of human participants in survey panels. Towards that goal, persona prompting has recently emerged as a technique to inform and align large pretrained language models. Persona prompting refers to the practice of using short textual descriptions of 'personas' in prompts to steer the LLM's generations. Personas describe individuals through different attributes such as their socio-demographics, attitudes, or behaviors, with the aim of aligning LLMs to produce responses that correlate with the corresponding human responses. Yet, recent work has produced mixed and partly conflicting results of persona prompting without clear patterns of success and failure. Among the few consistent findings is that the selection of persona attributes matters, and that using more attributes does not necessarily lead to better performance. It remains unclear how different attribute selection methods perform and how to choose among them. In this paper, we propose that observed human response variation of a survey question is a potential explanation for the mixed performance observed so far. In addition, we compare the performance of persona prompting associated with different methods for selecting persona attributes. We evaluate these methods on four different (general) social surveys across two countries, six LLMs, and twenty prediction tasks per survey. Our work helps to identify when persona prompting can be expected to be useful in survey prediction tasks, and provides new insights on the effectiveness of different attribute selection methods for LLM-based survey prediction using persona prompting.","authors":["Leon Fr\\\"ohling","Jens Rupprecht","Markus Strohmaier","Claudia Wagner"],"categories":["cs.CL","cs.CY"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-03","first_seen":"2026-09-03","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.02526","pdf_url":"https://arxiv.org/pdf/2609.02526","source_feed":"cs.CL","score":10,"bucket":"selected","rubric_hits":["A1","A2","A5","B1","B2","B3","B4"],"tags":["LLM仿真","调查预测","人物提示"],"reason":"直接研究用LLM预测调查回答，评估persona提示的有效性，并与真实人类数据…","model":"deepseek-v4-pro","scored_at":"2026-09-03T13:06:55","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-03","rank":1,"question":"人类回答变异能否解释persona提示在不同调查预测任务中的表现差异，以及不同persona属性选择方法能否提升预测性能？","design":"使用六种LLM，基于美国GSS、德国GGSS及WVS四份调查数据，通过不同属性选择方法构建persona提示，预测二十个调查问题的回答分布。","baseline":"真实人类调查数据：美国GSS、德国GGSS及WVS的个体层面回答。","findings":"人类回答变异与persona提示性能相关，高变异问题更难预测；不同属性选择方法效果差异显著，但无单一最优方法。","reliability":"论文未讨论","relevance":"直接研究LLM预测调查回答的可靠性，与人类数据对照，并探讨属性选择方法，对关注仿真偏差和条件失效的研究者很有价值。","inspiration":"借鉴其用人类回答变异作为任务难度指标，并系统比较属性选择方法的设计｜可迁移到经济预期调查或消费者信心预测，如预测通胀预期或消费意愿的异质性｜用LLM扮演不同人口群体，施加不同属性选择处理，预测密歇根消费者调查问题，以真实微观数据为基准评估仿真准确性"}},{"id":"2609.02580","version":1,"title":"Competitive Market Behavior of LLMs","zh_title":"大语言模型的竞争性市场行为","abstract":"Large language models (LLMs) are increasingly deployed as economic agents, yet there is little evidence whether LLM agents are suited for participating in market mechanisms designed for humans, and whether these mechanisms deliver desired outcomes when faced with LLM agents. We address this question by replicating seminal economic experiments, replacing human subjects with LLM agents. We place agents in a double auction environment, which is a widely-used market mechanism. We check whether such a market is able to deliver an efficient allocation of resources, thereby testing a novel dimension of alignment of LLM agents -- their compatibility with a fundamental market mechanism. We find that markets populated by LLM agents exhibit slower or no convergence towards market equilibrium, thus providing less efficient allocations than markets populated by humans. We then analyze agents' individual trading decisions and find substantial heterogeneity both across model families and market roles. We also run a lexical analysis of Chain-of-Thought (CoT) traces generated by the agents. We find that the decision to execute a trade rather than continue incrementally adjusting prices is associated with a shift from strategic considerations toward urgency. We publicly release our testing framework, which can be used for future evaluations.","authors":["Pawel Struski","Jakub Swistak","Inez Okulska","Przemyslaw Biecek"],"categories":["cs.MA","cs.AI","econ.GN","q-fin.EC"],"primary_category":"cs.MA","announce_type":"cross","date":"2026-09-03","first_seen":"2026-09-03","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.02580","pdf_url":"https://arxiv.org/pdf/2609.02580","source_feed":"cs.AI","score":10,"bucket":"selected","rubric_hits":["A1","A3","B1","B2","B4"],"tags":["LLM仿真","经济学实验","市场机制"],"reason":"用LLM替代人类被试复现经济学实验，并与人类数据对照，发现市场效率差异，直接相…","model":"deepseek-v4-pro","scored_at":"2026-09-03T13:06:55","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-03","rank":2,"question":"用LLM代理替代人类被试参与连续双向拍卖市场，能否像人类一样收敛到竞争均衡并实现有效资源配置？","design":"构建连续双向拍卖仿真环境，使用多种LLM模型（不同家族和能力层级）作为买方和卖方代理，每个代理拥有私有保留价格，市场由11个买方和11个卖方组成，供需曲线对称，理论均衡价格和数量确定。测量市场收敛速度、配置效率、个体交易决策异质性，并对思维链文本进行词汇分析。","baseline":"对照Smith (1962)的经典人类实验数据，人类被试在相同双向拍卖环境中通常快速收敛到竞争均衡。","findings":"LLM代理市场收敛速度较慢或根本不收敛，配置效率低于人类市场。个体交易决策在不同模型家族和市场角色间存在显著异质性，交易执行决策与思维链中从战略考量转向紧迫性相关。","reliability":"论文未讨论","relevance":"直接命中研究者关注的核心：用LLM替代人类被试复现经济学实验，并与人类数据对照，发现市场效率差异，属于批判性仿真研究，值得精读原文。","inspiration":"借鉴其使用经典实验范式（Smith双向拍卖）作为基准，通过市场级结果（收敛、效率）和个体行为（交易决策、思维链）的多层次测量来评估LLM与市场机制的兼容性。｜可迁移到资产定价实验、市场微观结构研究、政策干预的市场反应模拟等场景。｜以LLM代理作为交易者，在双向拍卖或订单簿市场中施加不同信息结构或交易规则处理，测量价格发现效率和市场流动性，并与人类实验数据或历史市场数据对照。"}},{"id":"2601.22396","version":3,"title":"Culturally Grounded Personas in Large Language Models: Characterization and Alignment with Socio-Psychological Value Frameworks","zh_title":"大语言模型中文化扎根的人格：表征及与社会心理价值框架的对齐","abstract":"Despite the growing utility of Large Language Models (LLMs) for simulating human behavior, the extent to which these synthetic personas accurately reflect world and moral value systems across different cultural conditionings remains uncertain. This paper investigates the alignment of synthetic, culturally-grounded personas with established frameworks, specifically the World Values Survey (WVS), the Inglehart-Welzel Cultural Map, and Moral Foundations Theory. We conceptualize and produce LLM-generated personas based on a set of interpretable WVS-derived variables, and we examine the generated personas through three complementary lenses: positioning on the Inglehart-Welzel map, which unveils their interpretation reflecting stable differences across cultural conditionings; demographic-level consistency with the World Values Survey, where response distributions broadly track human group patterns; and moral profiles derived from a Moral Foundations questionnaire, which we analyze through a culture-to-morality mapping to characterize how moral responses vary across different cultural configurations. Our approach of culturally-grounded persona generation and analysis enables evaluation of cross-cultural structure and moral variation.","authors":["Candida M. Greco","Lucio La Cava","Andrea Tagarelli"],"categories":["cs.CL","cs.AI","cs.CY","cs.HC","physics.soc-ph"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-03","first_seen":"2026-01-29","revised_at":"2026-09-03","abs_url":"https://arxiv.org/abs/2601.22396","pdf_url":"https://arxiv.org/pdf/2601.22396","source_feed":"cs.CL","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2"],"tags":["LLM仿真","文化价值观","人类数据对照"],"reason":"用LLM生成文化人格，与WVS等真实人类数据对照，评估仿真可靠性。","model":"deepseek-v4-pro","scored_at":"2026-09-03T13:07:10","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-03","rank":3,"question":"LLM生成的文化人格在多大程度上与真实世界的价值观和道德体系（WVS、Inglehart-Welzel文化地图、道德基础理论）对齐？","design":"基于WVS衍生的文化变量提示LLM生成文化人格，然后用这些人格条件化另一个LLM，分别回答IVS问题（用于计算IW坐标）、WVB-Probe问题（用于生成WVS文化剖面）和MFQ-2道德基础问卷（用于道德剖面），并分析人格在IW地图上的分布、与人口群体WVS分布的一致性以及文化变量到道德基础的映射。","baseline":"WVS/EVS整合调查（IVS）的人类响应数据（用于IW坐标计算）和WVB-Probe提供的人口群体（按大洲、居住地、教育水平划分）参考分布。","findings":"LLM生成的文化人格在Inglehart-Welzel地图上呈现出与文化条件相关的稳定差异；其WVS响应分布大体上追踪了人类群体模式，但存在系统性偏差。","reliability":"论文未讨论","relevance":"该研究直接评估LLM仿真人类文化价值观和道德判断的可靠性，与研究者关注的人类仿真实验和真实数据对照高度相关，值得精读原文以了解具体偏差模式和跨文化结构。","inspiration":"借鉴其用真实调查数据（WVS）作为基准来校准和检验LLM仿真输出的方法，以及通过文化变量条件化生成人格并测量多维度结果的设计。｜可迁移到经济金融领域的跨文化消费者行为、风险偏好、信任与合作等实验，例如不同文化背景下的投资决策或政策偏好。｜以LLM生成的不同文化人格为被试，施加经济激励或政策信息处理，测量其风险选择、时间贴现或对再分配政策的支持度，并与世界价值观调查中对应文化群体的人类回答进行对照，评估仿真偏差。"}},{"id":"2609.02122","version":1,"title":"AI agents reshape consensus formation in human groups","zh_title":"AI智能体重塑人类群体中的共识形成","abstract":"As large language model (LLM) agents shift from tools to participants in human groups, a fundamental question for collective behavior is how their growing presence reshapes consensus formation. Here we study mixed human-AI groups in a collaborative description game, in which shared conventions emerge through repeated rounds of random pairwise communication. Varying the proportions of LLM agents, we identify three distinct regimes of consensus formation: low agent proportions facilitate human-led consensus, intermediate proportions disrupt convergence, and high proportions restore strong consensus while shifting it toward agent-led conventions. Crucially, these regimes differ not only in the strength of convergence, but also in the semantic grounding and communicative form of the resulting consensus: human-led consensus is more concrete, holistic, and grounded in shared real-world analogies, whereas agent-led consensus is more abstract, less information-dense, and more geometrically segmented. Mechanistically, agent influence arises from a shared linguistic prior that places agents near one another in the expression space, combined with relatively stable expression choices across rounds; humans initially resist adopting expressions from partners perceived as AI but gradually yield to conformity pressure. These findings provide evidence that AI composition can shape the emergence, content, and perceived legitimacy of group norms, making agent proportion and transparency important design variables for human-AI systems.","authors":["Lin Chen","Ziyi Liu","Xia Hu","Yong Li"],"categories":["cs.CL","cs.CY","cs.SI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-03","first_seen":"2026-09-03","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.02122","pdf_url":"https://arxiv.org/pdf/2609.02122","source_feed":"cs.CL","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2","B4"],"tags":["LLM仿真","人机交互","共识形成"],"reason":"混合人机群体共识形成实验，LLM作为被试替代，有真实人类对照，涉及社会规范与政…","model":"deepseek-v4-pro","scored_at":"2026-09-03T13:06:53","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-03","rank":5,"question":"在混合人类与LLM智能体的群体中，智能体比例如何重塑共识形成的过程与结果？","design":"采用协作描述游戏：人类与LLM智能体随机配对，对同一抽象图形（七巧板）进行文字描述，每轮后收到对方描述和相似度反馈，共40轮。实验操纵LLM智能体比例（0%、12.5%、33.3%、50%、75%），测量最终轮描述之间的语义相似度作为共识强度，并分析共识的语义内容和表达形式。","baseline":"纯人类组（0%智能体比例）作为对照，其共识强度为0.695。","findings":"共识强度随智能体比例呈非单调变化：低比例（12.5%）促进人类主导的共识，中等比例（33.3%、50%）破坏收敛，高比例（75%）恢复强共识但转向智能体主导的规范。人类主导的共识更具体、整体、基于现实类比，而智能体主导的共识更抽象、信息密度低、几何分割。","reliability":"论文未讨论","relevance":"该研究直接以LLM作为人类被试的替代，在混合群体中考察共识形成，有真实人类对照，属于经济学实验和政策评估场景，且揭示了仿真在中等比例下失效的条件，值得精读原文。","inspiration":"借鉴其通过操纵智能体比例来识别非线性效应的实验设计，以及用语义相似度量化共识强度的测量方法。｜可迁移到政策公告的预期形成或社会规范传播等经济金融问题，例如研究AI顾问比例对投资者共识或通胀预期的影响。｜设计一个在线实验，招募人类被试与LLM智能体混合，处理为智能体比例（如0%、25%、50%、75%），结果变量为对某经济指标（如通胀率）的预测共识强度，对照真实历史调查数据（如密歇根大学通胀预期调查）。"}},{"id":"2609.01902","version":1,"title":"Accurate in space, unreliable in time: how LLMs represent national cultural change","zh_title":"空间准确，时间不可靠：大语言模型如何表征国家文化变迁","abstract":"Assessments of cultural alignment have become an important part of the development and improvement of large language models (LLMs). However, the majority of the evaluations treat culture as a single snapshot, investigating only whether a model represents a society accurately at the current time. Research in cultural psychology shows that cultural values change at different rates and directions over time. Therefore, a \"culturally aware\" model should capture not only where a culture is today but also how it has changed over time. We examine this missing dimension of cultural awareness using more than two decades of the World Values Survey data. We compare the cultural trajectories of 40 countries with the trajectories produced by four state-of-the-art (SOTA) LLMs on the Inglehart-Welzel cultural map. Our findings show that while models generally place countries close to their most recent surveyed positions, these representations tend to lag several years behind that position. They also capture only part of the magnitude of the observed change, introduce movement where little occurred, and rarely reproduce reversals in countries' trajectories. These findings point to temporal flattening and suggest that snapshot accuracy can give an incomplete picture of cultural awareness in LLMs and have implications for model evaluation, representational harms, and the governance of culturally aware AI systems.","authors":["Yalda Daryani","Miranda Bogen","Madeleine I. G. Daepp"],"categories":["cs.CY","cs.AI","cs.CL"],"primary_category":"cs.CY","announce_type":"cross","date":"2026-09-03","first_seen":"2026-09-03","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.01902","pdf_url":"https://arxiv.org/pdf/2609.01902","source_feed":"cs.CL","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["文化仿真","算法保真度","时间偏差"],"reason":"用LLM复现国家文化变迁并与世界价值观调查数据对照，评估仿真可靠性，发现时间滞…","model":"deepseek-v4-pro","scored_at":"2026-09-03T13:06:53","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-03","rank":4,"question":"LLM能否捕捉国家文化价值观随时间的变化轨迹，而不仅仅是当前快照？","design":"使用四个SOTA LLM（未具体命名）生成40个国家的文化价值观评分，并将其映射到Inglehart-Welzel文化地图上，比较模型产生的国家轨迹与WVS二十多年数据的真实轨迹。","baseline":"世界价值观调查（WVS）超过二十年的数据，覆盖40个国家。","findings":"模型通常将国家定位在其最近调查位置附近，但表示滞后数年；模型只捕捉到部分变化幅度，在变化很小的地方引入虚假移动，且很少再现轨迹逆转。","reliability":"论文指出快照准确性可能掩盖时间扁平化问题，但未详细讨论失效条件；模型滞后、幅度缩小和逆转缺失表明LLM在时间维度上不可靠。","relevance":"该研究直接评估LLM作为人类被试替代品在文化变迁仿真中的可靠性，发现时间维度上的系统性偏差，对关注仿真有效性和偏差的研究者具有重要参考价值。","inspiration":"借鉴其将动态轨迹与静态快照对比的方法，可迁移到经济金融中的时间序列预期或行为变化研究，例如用LLM模拟消费者信心或投资者情绪的历史演变，以真实调查数据（如密歇根消费者信心指数）为基准，检验模型是否捕捉趋势、幅度和转折点。｜例如，在资产定价实验中，让LLM扮演不同时期的投资者，给出风险偏好或市场预期，与历史调查数据对比，评估其时间一致性。｜设计：以LLM为被试，提示其模拟特定国家或群体在多个年份的经济态度，结果变量为风险厌恶或通胀预期，对照真实面板调查数据，分析模型的时间滞后和虚假波动。"}},{"id":"2609.02512","version":1,"title":"Beauty is in the AI of the beholder: MLLMs systematically overrate facial attractiveness","zh_title":"美在AI眼中：多模态大模型系统性高估面部吸引力","abstract":"Beauty assessments from Multimodal Large Language Models (MLLMs) are increasingly popular amongst users, companies, and aestheticians. This raises the question of whether these AI models can accurately reflect human judgments of attractiveness. In a pre- registered exploratory study, we compared the attractiveness ratings of 2,513 human participants to four widely used commercial AI models: Claude, Gemini, GPT, and Grok. Results showed that MLLMs systematically rate faces more favourably and within a narrower range than humans and, at the time of study, do not reproduce human ratings in absolute terms. However, MLLMs exhibit strong correlations with human attractiveness judgments, accurately tracking the rank-ordering of faces. MLLMs may judge faces by different cues than humans; only face age was a predictor of facial attractiveness in both humans and MLLMs, with inconsistent patterns across models for ethnicity and gender. AI models strongly agree with one another, except for Grok, which also showed the lowest agreement with humans. Our findings suggest that while they may be able to approximate rank-orderings of human attractiveness, current off-the-shelf commercial MLLMs systematically overrate the beauty of human faces.","authors":["Santiago Grandas","Juan Sebastian Cely-Acosta","Mohit Mendiratta","Shafee Hassan","Macken Murphy"],"categories":["cs.CV","cs.HC"],"primary_category":"cs.CV","announce_type":"cross","date":"2026-09-03","first_seen":"2026-09-03","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.02512","pdf_url":"https://arxiv.org/pdf/2609.02512","source_feed":"cs.HC","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","人类对照","偏差评估"],"reason":"用MLLM替代人类被试评估吸引力，并与2513名人类对照，发现系统性偏差，直接…","model":"deepseek-v4-pro","scored_at":"2026-09-03T13:06:55","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-03","rank":6,"question":"多模态大语言模型（MLLMs）的面部吸引力评分能否准确反映人类判断？","design":"本研究并非严格意义上的仿真实验，而是将四个商用多模态大语言模型（Claude、Gemini、GPT、Grok）作为“AI评分者”，对同一组面部图像进行吸引力评分，并与2513名人类参与者的评分进行比较，分析评分分布、相关性及预测因素。","baseline":"2513名人类参与者对相同面部图像的吸引力评分。","findings":"MLLMs系统性给出更高且范围更窄的评分，不能复现人类评分的绝对值；但MLLMs与人类评分存在强相关，能准确追踪面部吸引力的相对排序。MLLMs可能依据与人类不同的线索进行判断，只有面部年龄在人类和MLLMs中都是吸引力的预测因素，而种族和性别的影响在不同模型间不一致。","reliability":"论文指出当前商用MLLMs在绝对评分上系统性高估，且不同模型间一致性存在差异（Grok与人类一致性最低），但未深入讨论失效条件，仅强调不能直接替代人类绝对评分。","relevance":"该研究直接评估了LLM作为人类被试替代品在主观审美判断中的可靠性，提供了真实人类对照，揭示了系统性偏差和排序一致性，对关注仿真效度的研究者具有重要参考价值。","inspiration":"借鉴其预注册探索性设计和多模型对比方法，可系统评估AI与人类在主观判断任务上的偏差模式。｜可迁移到信贷审批中的外貌歧视研究，或消费者对产品外观的偏好评估。｜以银行信贷员为人类被试，让MLLMs和信贷员对同一组借款人照片进行信用worthiness评分，比较评分分布和排序，并以实际贷款数据作为外部基准。"}},{"id":"2609.01867","version":1,"title":"Thinking effort aligns between humans and reasoning models in abductive reasoning","zh_title":"溯因推理中人类与推理模型的思维努力对齐","abstract":"A major question in cognitive modeling concerns the behavioral alignment between large language models and humans across linguistic and non-linguistic tasks. Unlike standard LLMs, large reasoning models (LRMs) are optimized with reinforcement learning from verifiable rewards, encouraging correct solutions to reasoning tasks rather than preference-aligned responses. Recent work (de Varda et al., 2025) investigates the cost of thinking in humans and LRMs by comparing human reaction times with model reasoning traces across a range of reasoning tasks. We isolate this alignment by turning to abductive reasoning: unlike deductive tasks, its difficulty cannot be inferred from formal structure and offers no shortcuts a model could exploit to mimic effort without genuine search, providing firmer ground for empirical claims of shared effort. We find further evidence of alignment between LRM and human reasoning effort, as well as evidence that models and humans tend to make similar errors. Finally, we show that decoding methods that let models explore multiple reasoning paths increase alignment in reasoning cost between humans and LRMs across the three models tested.","authors":["Henry Arthur"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-03","first_seen":"2026-09-03","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.01867","pdf_url":"https://arxiv.org/pdf/2609.01867","source_feed":"cs.CL","score":8,"bucket":"selected","rubric_hits":["A1","B1","B4"],"tags":["LLM仿真","认知对齐","溯因推理"],"reason":"比较人类与推理模型在溯因推理中的思维努力，含人类反应时对照，属仿真对齐研究。","model":"deepseek-v4-pro","scored_at":"2026-09-03T13:06:51","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-03","rank":8,"question":"人类与大型推理模型在溯因推理中的思维努力是否对齐？","design":"使用三种大型推理模型（DeepSeek-R1等）作为被试，在溯因推理任务上比较模型生成的推理链长度（token数）与人类反应时；并测试不同解码策略（如多样本搜索）对对齐程度的影响。","baseline":"人类被试在相同溯因推理任务上的反应时数据（来自de Varda et al., 2025的七项推理任务之一）。","findings":"发现LRM与人类在溯因推理中的思维努力存在对齐，且模型与人类倾向于犯类似错误。采用允许多条推理路径探索的解码方法可提高对齐程度。","reliability":"论文承认CoT可能不忠实于底层计算，且对齐并非机制性主张；通过测试多种解码策略和推理努力水平来部分回应批评。","relevance":"该研究直接比较人类与LLM在推理任务中的行为对齐，包含真实人类反应时对照，并讨论仿真失效条件，与研究者关注点高度契合，值得精读。","inspiration":"借鉴其利用任务特性（溯因推理无形式捷径）来排除模型投机取巧、增强对齐结论可信度的设计思路。｜可迁移到经济决策中的信念更新或预期形成场景，如投资者在信息不完全下的推断。｜以LLM为被试，呈现模糊经济信息（如公司公告），要求给出解释并测量推理链长度，与人类实验中的反应时和解释内容对照，检验模型是否复现人类推断努力和错误模式。"}},{"id":"2609.02277","version":1,"title":"Auditory Illusion Benchmark for Large Audio Language Models","zh_title":"大型音频语言模型的听觉错觉基准","abstract":"Perceptual illusions have long served as crucial probes into human cognition, revealing biases and limitations of perception. In the auditory domain, such illusions provide a unique lens for testing whether Large Audio Language Models (LALMs) replicate human perceptual tendencies. Despite their importance, most benchmarks focus on visual illusions or general audio tasks, leaving auditory illusions underexplored. To this end, we present AIB, the first auditory illusion benchmark for LALMs, covering ten representative illusions across music, sound, and speech, each annotated for the presence of knowledge-based priors. Our methodology pairs model evaluation with controlled human listening studies, enabling direct comparison of responses. Results show systematic differences: while most LALMs remain signal-faithful on low-level acoustic illusions, several exhibit more human-like responses when linguistic or musical priors are involved, although no model matches the human perceptual profile. These findings highlight the current limitations of LALMs as cognitive models. By establishing auditory illusions as a rigorous testbed, our work offers a new perspective for probing neural black-box models and advancing understanding of auditory cognition. AIB is publicly available at https://github.com/gillosae/aib.","authors":["Hayoon Kim","Eunice Hong","Kyogu Lee"],"categories":["cs.SD","cs.AI"],"primary_category":"cs.SD","announce_type":"cross","date":"2026-09-03","first_seen":"2026-09-03","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.02277","pdf_url":"https://arxiv.org/pdf/2609.02277","source_feed":"cs.AI","score":8,"bucket":"selected","rubric_hits":["A1","A2","B1"],"tags":["听觉错觉","人类感知仿真","模型评估"],"reason":"用LALM复现人类听觉错觉，并与人类数据对照，评估模型作为认知模型的可靠性","model":"deepseek-v4-pro","scored_at":"2026-09-03T13:06:53","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-03","rank":9,"question":"大型音频语言模型（LALMs）在多大程度上复现人类对听觉错觉的感知倾向，能否作为人类听觉认知的模型？","design":"构建包含10种听觉错觉的基准AIB，覆盖音乐、声音和语音领域，按机制分为物理型和物理+知识型；将错觉任务转化为多项选择题，对多个LALMs进行测试，并与受控人类听力实验的结果进行对比。","baseline":"通过受控人类听力研究收集的人类对相同刺激的错觉易感性数据。","findings":"在低层声学错觉上，多数LALMs保持信号忠实，而人类表现出强错觉易感性；在涉及语言或音乐先验的错觉上，部分模型表现出更接近人类的反应，但没有模型完全匹配人类的感知特征。","reliability":"论文指出LALMs在物理型错觉上倾向于信号忠实，与人类不一致，而在知识型错觉上部分对齐，表明其错觉易感性可能源于高层先验而非共享的低层听觉处理；未讨论其他失效条件。","relevance":"该研究直接评估LALMs作为人类听觉认知模型的可靠性，与研究者关注LLM仿真人类感知和决策的核心问题高度相关，提供了模型与人类系统对比的实证证据，值得精读。","inspiration":"借鉴其构建受控刺激对（错觉与对照）和将主观感知转化为多项选择任务的方法，可用于经济金融中的主观判断仿真。｜可迁移到投资者对市场信息的感知偏差研究，如盈余公告后的漂移现象。｜以LLM为被试，呈现带有不同信息框架的财务报告（处理），测量其对未来收益的预期（结果变量），并与真实投资者调查数据对照。"}},{"id":"2608.27309","version":2,"title":"Difference-in-Differences on a Censored Rating Scale Can Manufacture an Effect: Evidence from a Pre-Registered LLM-Judge Audit","zh_title":"截断评分量表上的双重差分可能制造效应：来自预注册LLM法官审计的证据","abstract":"Audits of LLM judges certify a bias by contrasting matched conditions, and the strongest designs difference twice: a within-item contrast between two candidate responses, differenced again across a manipulated attribute, read off a bounded rating scale. We show that this endpoint is not identified on the scale that reports it. Each term of the double difference is censored by its own share, so the observed statistic confounds differential preference with differential attenuation: a severity shift common to both responses manufactures an interaction whenever the two censor it unequally, as unequal distances from the bounds make them, exactly where good stimuli place them. We exhibit the failure inside a pre-registered audit of a frozen pedagogy judge, sealed before the first of its 990 calls. The registered primary endpoint, the effect of a stated learner profile on the judge's scaffolding preference, is null: $+0.085$ points (95\\% BCa $[-0.167, +0.353]$, $p = 0.684$). The audit's one nominally significant interaction, $+0.378$ ($p = 0.002$), is not identified as preference: a construction containing zero differential preference reproduces 79 to 85\\% of it from the observed severity shift and the scale floor alone. We derive the mechanism in closed form and show that its contribution is measurable from an audit's own ratings.","authors":["Shuyi Fan","Boyuan Deng","Mengyu Xu","Xinhong Xie","Chenyang Li","Hongyang Zhang"],"categories":["cs.CL","cs.AI","cs.CY"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-03","first_seen":"2026-08-28","revised_at":"2026-09-03","abs_url":"https://arxiv.org/abs/2608.27309","pdf_url":"https://arxiv.org/pdf/2608.27309","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A2","B4"],"tags":["LLM法官","偏差审计","方法论批判"],"reason":"论文审计LLM法官的偏差，涉及评估仿真可靠性，且批判性指出失效条件，方法可迁移。","model":"deepseek-v4-pro","scored_at":"2026-09-03T13:07:13","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-03","rank":10,"question":"LLM法官在双重差分审计中的交互效应是否被有界评分量表的审查机制所混淆，从而制造出虚假效应？","design":"该研究不是用LLM仿真人类被试，而是审计一个冻结的LLM教学法官。设计为：55个刺激，每个刺激包含一个高脚手架和一个低脚手架候选回复，在三种条件下（无档案、新手档案、高级档案）由LLM法官评分，结果变量为脚手架偏好得分（5点量表）。","baseline":"无对照","findings":"注册的主要终点（档案对脚手架偏好的影响）为零；唯一名义显著的交互效应（+0.378）并非真实偏好，而是由严重性偏移和量表下限共同制造的，零差分偏好的构造可重现其79-85%的幅度。","reliability":"论文承认其发现仅针对特定审计和量表，且主要终点为零是“未能检测到”而非“证明无效应”；审查机制在有界量表上普遍存在，但具体影响取决于刺激分布和量表边界。","relevance":"该论文对LLM仿真可靠性提出批判，指出有界量表上的双重差分可能制造虚假效应，这与研究者关注仿真失效条件高度相关，值得阅读原文以了解具体机制和检验方法。","inspiration":"借鉴其双重差分设计中的审查机制识别方法，即从审计自身评分中测量衰减贡献，用于稳健性检验。｜可迁移到信贷审批歧视研究，其中LLM法官对贷款申请人的评分可能受申请人特征影响，且评分量表有界。｜以LLM作为信贷审批员，处理为申请人种族或性别，结果变量为信用评分（1-10量表），对照真实信贷审批数据，检验交互效应是否由量表审查制造。"}},{"id":"2608.27463","version":2,"title":"Rating the Raters: Rasch Measurement Theory for LLM Evaluation","zh_title":"评估评分者：用于LLM评估的Rasch测量理论","abstract":"LLMs now sit on every side of evaluation: as examinees scored on benchmarks, judges of other models' outputs, and raters of human-generated content. Each paradigm can be viewed as a measurement problem, where a latent property of an object is probed with items from an instrument (e.g., benchmark) by judges or raters. Standard evaluation practices often neglect the contributions of each core component to the end result, limiting our understanding of what is being measured. Rasch measurement theory (RMT) is well-suited to this problem. RMT decomposes ordinal ratings into separable facets on a common scale. It further provides a battery of diagnostics that can identify miscalibrated measurements and rater biases. We present a case study of RMT applied to the LLM-as-rater paradigm using the Measuring Hate Speech corpus, whose construct was itself built under RMT. We fit a series of many-facet Rasch models to annotations from nine LLMs spanning families and capability levels. Our analyses show that LLMs systematically differ from human raters in severity, item-level calibration, question-order robustness, target-identity sensitivity, and rating scale use, all of which standard evaluation practice would largely obscure. Overall, we argue that RMT belongs in the toolkit for evaluating LLM-as-examinee, -judge, and -rater paradigms.","authors":["Pratik S. Sachdeva","Nathan Boudol"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"replace","date":"2026-09-03","first_seen":"2026-08-31","revised_at":"2026-09-03","abs_url":"https://arxiv.org/abs/2608.27463","pdf_url":"https://arxiv.org/pdf/2608.27463","source_feed":"cs.AI","score":7,"bucket":"pending","rubric_hits":["A2","B1","B4"],"tags":["LLM评估","测量理论","评分者偏差"],"reason":"评估LLM作为评分者与人类评分者的差异，涉及测量偏差与可靠性，可迁移到仿真评估。","model":"deepseek-v4-pro","scored_at":"2026-09-03T13:07:11","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-03","rank":11,"question":"LLM作为评分者时，其测量特性与人类评分者有何系统性差异？","design":"本研究非仿真实验，而是测量评估研究。使用Measuring Hate Speech语料库，选取9个不同家族和能力水平的LLM对评论进行标注，拟合多面Rasch模型，分析LLM与人类评分者在严重度、项目校准、问题顺序稳健性、目标身份敏感性和评分量表使用上的差异。","baseline":"Measuring Hate Speech语料库中的人类标注，包括参考评论集（70条评论，每条超过200个标注）和扩展评论集（5990条评论，每条至少4个人类评分者）。","findings":"LLM评分者在严重度、项目级校准、问题顺序稳健性、目标身份敏感性和评分量表使用上与人类评分者存在系统性差异。标准评估实践会掩盖这些差异，而Rasch测量理论能有效揭示这些偏差。","reliability":"论文未明确讨论失效条件，但指出LLM与人类在测量特性上的差异表明LLM不能直接替代人类评分者，需谨慎校准。","relevance":"该研究直接评估LLM作为评分者与人类评分者的差异，涉及测量偏差与可靠性，对使用LLM进行人类仿真实验的研究者具有重要参考价值，值得阅读原文。","inspiration":"借鉴Rasch测量理论分解评分变异，识别LLM评分者的系统性偏差，为仿真评估提供诊断工具。｜可迁移到信贷审批歧视研究，用LLM模拟信贷员对贷款申请人的评分，检验其是否与人类评分存在偏差。｜以LLM作为信贷员，对贷款申请进行风险评分，处理为申请人种族或性别，结果变量为评分和批准决策，与真实信贷审批数据对照，评估LLM仿真的可靠性。"}},{"id":"2609.01794","version":1,"title":"Disentangling Statistical Preemption from Entrenchment in Language Models' Avoidance of Overgeneralization","zh_title":"在语言模型避免过度泛化中区分统计抢占与固化","abstract":"How do learners avoid overgeneralizations such as Tom laughed me without explicit negative evidence? Constructionists have posited two proposals that describe indirect negative evidence against overgeneralizations: preemption (which privileges exposure to near-synonymous construction---e.g., she made him laugh) vs. entrenchment (all exposures to a verb's grammatical usages, including cases like He laughed). We disentangle these hypotheses by running controlled rearing experiments on LMs trained on child-caregiver conversations, where we systematically remove preemptive vs. non-preemptive evidence. We find that while LMs avoid overgeneralizations, they do not show preemption at a verb-specific level, instead showing weak but non-zero evidence of abstract preemption. Combined with results from analyzing the LMs' training dynamics, we find that LMs treat competing structures as indirect positive---as opposed to negative---evidence in the verb-specific condition. Insofar as preemption is the more plausible route to avoiding overgeneralizations in humans, our results point the need for there to be sensitivities to indirect negative evidence in neural network learners, and suggest new human experiments to test abstract preemption.","authors":["Yixuan Wang","Freda Shi","Kanishka Misra"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-03","first_seen":"2026-09-03","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.01794","pdf_url":"https://arxiv.org/pdf/2609.01794","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A1","B1","B4"],"tags":["语言习得","认知建模","人类对照"],"reason":"用LLM模拟儿童语言习得，与人类数据对照，但非社会行为仿真","model":"deepseek-v4-pro","scored_at":"2026-09-03T13:06:59","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-03","rank":13,"question":"语言模型能否通过间接负面证据（抢占或固化）避免动词过度泛化，以及抢占与固化哪个机制起作用？","design":"对语言模型进行“控制性养育”实验：在儿童-看护者对话语料上训练模型，系统移除抢占性证据或非抢占性证据，测量模型对过度泛化句子的接受度。","baseline":"人类儿童语言习得数据（如儿童过度泛化及最终回避的语料库记录）作为参照，但未直接使用个体人类被试数据。","findings":"语言模型能避免过度泛化，但未表现出动词特异性抢占效应，仅显示微弱的抽象抢占证据。模型将竞争结构视为间接正面证据而非负面证据。","reliability":"论文承认语言模型分析不能直接推论到人类，且抢占与固化证据存在多重共线性问题，人类实验难以获得学习者自身经验。","relevance":"该研究用LLM模拟人类语言习得，有真实人类数据对照，但属于认知科学而非社会行为仿真，对关注经济学实验仿真的研究者参考价值有限。","inspiration":"借鉴其控制性养育设计，通过系统移除特定类型训练数据来分离竞争机制。｜可迁移到经济金融中的学习与决策机制研究，如投资者如何从历史数据中学习风险规避。｜以LLM为被试，训练时移除特定类型的市场反馈（如正面或负面新闻），测量模型对资产价格的预期，并与真实投资者调查数据对照。"}},{"id":"2609.01918","version":1,"title":"Grounded, Compute-Efficient LLM Policy Agents for Energy-Poverty Equity in Physically-Constrained Peer-to-Peer Energy Markets","zh_title":"物理约束点对点能源市场中面向能源贫困公平的接地气、计算高效LLM策略智能体","abstract":"Energy poverty is nearly absent from NLP-for-social-good, and the little existing work is either static retrieval/QA or relies on carbon-intensive cloud LLMs, a self-defeating \"computational irony\" for a humanitarian setting. We present EqGrid, a closed-loop simulation in which a low-frequency, open-weight LLM policy agent sets price and carbon bounds and targeted subsidies over a community of empirically-grounded household personas, while high-frequency multi-agent RL traders clear a continuous double auction constrained by a physical distribution grid (IEEE-33-bus with Dynamic Operating Envelopes). Our contribution is threefold and directly addresses how to measure the social impact of AI: (i) grounded personas (region-matched socio-demographics) whose load curves are checked for shape and level realism against real smart-meter data; (ii) formal energy-poverty equity metrics (Energy Burden, Gini of EB, LIHC) showing the intervention reduces burden inequality without raising net grid cost; and (iii) a compute-efficiency frontier that measures how much equity performance survives compressing the policy agent from a 235B teacher down to a sub-1B model deployable on a laptop, in estimated energy/carbon per decision. A decoupled-safety design (the LLM sets bounds; a validate-and-project grid gate executes) yields zero grid-constraint violations versus 55 under direct LLM control. On energy-poverty equity, the LLM policy lowers the Gini of energy burden to 0.305 (from 0.351) and mean burden by 28% while cutting cost (outperforming a tuned rule baseline), and a 3B-active model retains 95% of the benefit at roughly 9x lower inference energy than the teacher, with even a 0.8B on-device model retaining 92% at roughly 24x lower energy. We will release code and configs.","authors":["Kunal Jadhav","Siddhesh More"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-03","first_seen":"2026-09-03","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.01918","pdf_url":"https://arxiv.org/pdf/2609.01918","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A3","B1","B2","B4"],"tags":["LLM仿真","能源市场","公平性评估"],"reason":"用LLM agent模拟能源市场政策干预，并与真实智能电表数据对照，涉及公平性…","model":"deepseek-v4-pro","scored_at":"2026-09-03T13:07:01","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-03","rank":15,"question":"如何设计一个计算高效且物理安全的LLM政策智能体，在点对点能源市场中缓解能源贫困并提升公平性？","design":"使用开放权重LLM作为低频政策智能体，为基于EU-SILC真实数据生成的家庭角色设定价格/碳上限和定向补贴；高频多智能体强化学习交易者在IEEE-33总线物理电网约束下进行连续双向拍卖；测量能源负担、能源负担基尼系数、LIHC等公平性指标。","baseline":"对照真实智能电表数据验证负荷曲线的形状和水平真实性；家庭角色基于匈牙利EU-SILC边际数据。","findings":"LLM政策将能源负担基尼系数从0.351降至0.305，平均负担降低28%，同时降低成本并优于规则基线；解耦安全设计实现零电网违规，而直接LLM控制有55次违规。","reliability":"论文未讨论","relevance":"该研究用LLM智能体模拟政策干预并与真实数据对照，评估公平性指标，属于人类仿真实验，但场景为能源市场而非典型经济金融实验，值得读原文了解其仿真框架和测量方法。","inspiration":"借鉴其分层仿真设计：LLM设定政策边界，底层RL智能体执行交易，并引入物理约束和安全层，同时用真实数据校准仿真人群。｜可迁移到能源市场政策评估、碳税设计或补贴分配等经济政策场景。｜以家庭能源消费数据为被试，用LLM模拟不同补贴政策（处理），测量能源负担和公平性（结果变量），并与真实智能电表数据和家庭调查数据对照。"}},{"id":"2609.02707","version":1,"title":"Door-in-the-Face Requests and Refusal Behaviour in Large Language Models","zh_title":"大语言模型中的登门槛请求与拒绝行为","abstract":"Does the door-in-the-face technique work on language models? In humans, a large request that is refused makes a smaller follow-up request more likely to be granted. We test this on nine production models from three providers: each model refuses a large request, then receives a smaller version of the same request, and we compare its compliance with asking directly. The answer depends on the model. On Anthropic's frontier models the technique works: Opus 5 answers the smaller request 65.8% of the time after refusing the larger one, against 29.3% when asked directly. On the frontier models of OpenAI and Google, and on Haiku 4.5, it backfires, lowering compliance by 15.5 to 23.0 points. A control locates the effect: a refused large request on an unrelated topic does less than the related one on all nine models, so the concession itself matters everywhere, while the reaction to having just refused something differs by model family. The technique does not transfer to refusals drawn from public benchmarks. What decides whether a retreat can work is what the request asks for: rewriting 265 refused requests for usable instructions into requests for explanations of the same topic removed the refusal in 263 cases. Human influence techniques port to language models one model family at a time.","authors":["Til Jordan"],"categories":["cs.AI","cs.CL"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-09-03","first_seen":"2026-09-03","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.02707","pdf_url":"https://arxiv.org/pdf/2609.02707","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A1","B4"],"tags":["LLM行为实验","说服技巧","模型对比"],"reason":"测试LLM对人类说服技巧的反应，与人类行为对照，但非仿真被试","model":"deepseek-v4-pro","scored_at":"2026-09-03T13:06:57","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-03","rank":18,"question":"大语言模型是否会对“以退为进”（door-in-the-face）说服技巧产生与人类相似的反应，即先拒绝大请求后更可能同意小请求？","design":"该研究并非用LLM仿真人类被试，而是直接测试LLM本身的行为。在9个生产级模型上，每个试验包含四种条件：冷启动（直接提出小请求）、DITF（先提出一个设计为会被拒绝的大请求，再提出相同的小请求）、热身（先回答一个良性问题）、无关拒绝（先拒绝一个无关的大请求）。结果变量为模型对小请求的合规率，由跨模型家族的评判模型打分。","baseline":"无对照。论文未将LLM行为与真实人类被试在相同任务上的数据直接比较，而是引用社会心理学中人类DITF效应的经典结论作为背景。","findings":"DITF效应因模型家族而异：在Anthropic的前沿模型上，先拒绝大请求后小请求合规率显著提高（如Opus 5从29.3%升至65.8%）；而在OpenAI、Google的前沿模型及Haiku 4.5上，该技巧反而降低合规率15.5至23个百分点。进一步分析表明，让步本身在所有模型上都有作用，但模型家族决定了反应方向；且该效应不适用于公共基准中的拒绝，只有当小请求要求的是判断而非可操作内容时，DITF才可能有效。","reliability":"论文讨论了效应的边界条件：DITF不适用于模型在公共基准上产生的拒绝，且只有当小请求要求的是判断（而非可操作指令）时，让步才可能成功。此外，效应因模型家族而异，表明不存在统一的“类人”反应。","relevance":"该研究直接测试LLM对人类说服技巧的反应，虽未将LLM作为人类被试的替代品，但揭示了LLM行为与人类已知规律的异同，对评估LLM在仿真实验中的可靠性有参考价值，值得阅读原文以了解模型家族差异和边界条件。","inspiration":"借鉴其严谨的多条件对照设计（冷启动、DITF、热身、无关拒绝）来分离特定机制，并利用模型自身产生的拒绝作为实验材料，可迁移到经济金融中的谈判或议价场景，例如测试LLM在模拟消费者对价格让步的反应时是否表现出与人类相似的锚定效应。｜可设计一个实验：以LLM作为模拟消费者，先提出一个高价（大请求）被拒绝，再提出一个较低价（小请求），观察其购买意愿变化，并与真实消费者在相同议价情境下的行为数据（如实验经济学中的最后通牒博弈或讨价还价实验）进行对照，以评估LLM仿真在议价行为研究中的有效性。"}},{"id":"2609.02797","version":1,"title":"Dutch Books for Language Models","zh_title":"语言模型的荷兰赌：概率预测的连贯性评估","abstract":"People increasingly use language models to support life decisions. Many such decisions involve a probabilistic forecast: How likely is a major life event, a natural disaster, or an economic outcome? Users of language models may implicitly trust that these forecasts fall out of a coherent world model. In this paper, we evaluate the coherence of language model probabilistic forecasts through a procedure that builds on a theorem due to de Finetti. We elicit forecasts from language models across events generated from stock returns data. We then use linear programs to compute the largest Dutch-book profit - the profit an arbitrageur could guarantee by betting against model-generated probabilities - which we use as a measure of incoherence. Our procedure does not require outcome labels, so we can evaluate coherence even in settings where outcomes are not observed or have not yet resolved. We find substantial evidence of incoherence in language model forecasts. Such incoherence increases when there are richer logical relationships between events, and irrelevant contextual details can increase incoherence by an order of magnitude. We conclude by discussing how alternative training strategies may improve probabilistic coherence.","authors":["Isaiah Andrews","Suproteem Sarkar"],"categories":["econ.GN","cs.AI","cs.CL","cs.LG","q-fin.EC"],"primary_category":"econ.GN","announce_type":"cross","date":"2026-09-03","first_seen":"2026-09-03","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.02797","pdf_url":"https://arxiv.org/pdf/2609.02797","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A2","B3"],"tags":["LLM概率预测","连贯性评估","决策可靠性"],"reason":"评估LLM概率预测的连贯性，涉及决策可靠性，可迁移到仿真偏差研究。","model":"deepseek-v4-pro","scored_at":"2026-09-03T13:06:57","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-03","rank":19,"question":"语言模型生成的概率预测在多大程度上满足概率论连贯性（即不存在荷兰赌套利）？","design":"本研究不是用LLM模拟人类被试，而是直接评估LLM本身作为概率预测者的连贯性。研究者从股票收益率数据构造事件（如未来收益落在某区间、事件的并集和补集），要求15个语言模型基于近期收益历史和新闻标题等上下文给出这些事件的概率预测，然后利用线性规划计算针对模型预测概率的最大荷兰赌利润，作为不连贯性的度量。","baseline":"无对照","findings":"语言模型的概率预测普遍存在显著不连贯性，且不连贯性在事件间逻辑关系更丰富时增加；无关的上下文细节可使不连贯性增加一个数量级。","reliability":"论文未讨论","relevance":"该研究直接评估LLM概率预测的可靠性，对使用LLM进行经济预测或仿真实验的研究者具有重要警示意义，值得阅读原文以了解不连贯性的具体表现和测量方法。","inspiration":"借鉴其利用荷兰赌定理和线性规划构造无结果标签的连贯性度量，可迁移到经济预测场景如政策公告的预期形成或资产定价实验，设计上可让LLM对同一组经济事件的不同逻辑组合给出概率预测，计算套利利润作为不连贯性指标，并与人类预测者的不连贯性进行对比。"}},{"id":"2609.01815","version":1,"title":"Induction and Inquiry via Probabilistic Reasoning over Language and Code","zh_title":"通过语言与代码的概率推理进行归纳与探究","abstract":"How humans grow and maintain abstract knowledge from the sparse, streaming noisy data of experience is a longstanding challenge in cognitive science. Any computational account must satisfy at least three desiderata: It must be (1) data-efficient and compute-efficient, (2) capture gradations of uncertainty to support intelligent inquiry and information gathering, and (3) be flexible enough to mentally represent the endless range of concepts people can learn and think about. Here we introduce a computational model that captures these three properties, by encoding symbolic knowledge as mental programs that combine natural language with source code, and sequentially inferring mental programs using LLM-guided Bayesian learning algorithms. Across a range of behavioral studies this model successfully reproduces quantitative signatures of human inductive learning and active inquiry, such as anchoring, garden-pathing, and other effects. In contrast, pure LLMs and classic Bayesian models either fail at the underlying task, or do not reproduce human behavior, or succeed only at exorbitant computational cost. These results suggest that one way humans continually grow their knowledge is by mentally representing many hypotheses spanning language-like and program-like representations, then revising those hypotheses to approximate Bayesian updates, while a bottom-up neural mechanism (an LLM) makes inference both tractable and learnable.","authors":["Wasu Top Piriyakulkij","Sam Acquaviva","Cassidy Langenfeld","Joshua Tenenbaum","Kevin Ellis"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-03","first_seen":"2026-09-03","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.01815","pdf_url":"https://arxiv.org/pdf/2609.01815","source_feed":"cs.AI","score":7,"bucket":"pending","rubric_hits":["A1","B1"],"tags":["认知建模","贝叶斯学习","人类行为复现"],"reason":"用LLM引导贝叶斯学习复现人类归纳学习行为，并与行为实验对照，方法可迁移到人类…","model":"deepseek-v4-pro","scored_at":"2026-09-03T13:06:51","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-03","rank":14,"question":"人类如何从稀疏、流式、嘈杂的经验数据中高效地归纳和更新抽象知识，并主动进行信息获取？","design":"该研究提出一个计算模型，将符号知识编码为结合自然语言和源代码的“心理程序”，并使用LLM引导的贝叶斯学习算法进行序列推断。模型在多个行为实验任务中模拟人类被试，包括概念学习、主动实验（Zendo游戏）和提问（网络购物助手），测量其归纳学习和主动探究的行为模式。","baseline":"对照的真实人类数据来自多个行为实验，包括Bramley等人（2018）的Zendo游戏数据（7轮实验后对8个保留测试构造的预测），以及其他归纳学习任务中的人类行为数据（如锚定、花园路径效应等）。","findings":"该模型成功复现了人类归纳学习和主动探究的定量特征，如锚定、花园路径等效应，并且在多个任务上比纯LLM和经典贝叶斯模型更符合人类行为，同时计算效率更高。模型还展示了通过调节推理时间预算来捕捉有限理性行为的能力。","reliability":"论文未明确讨论失效条件，但指出纯LLM和经典贝叶斯模型在任务上要么失败，要么不符合人类行为，要么计算成本过高，暗示模型依赖于LLM的生成能力和贝叶斯更新的结合。","relevance":"该研究直接使用LLM引导的贝叶斯学习来模拟人类归纳学习和主动探究行为，并与真实人类实验数据对照，属于用LLM进行人类仿真实验的研究，且涉及认知科学中的行为复现，对关注LLM仿真可靠性和偏差的研究者有参考价值。","inspiration":"该方法将LLM作为假设生成器，结合贝叶斯更新进行序列推理，可借鉴用于模拟经济决策中的信念更新和主动信息获取。｜可迁移到政策公告的预期形成、消费者跨期选择中的学习过程、或投资者在不确定环境下的信息搜寻行为。｜设计一个实验：用LLM引导的贝叶斯模型作为被试，模拟投资者在收到一系列经济新闻后的资产价格预期更新，处理是不同信息呈现方式（如顺序、频率），结果变量是预测准确性和不确定性，并与真实人类实验数据（如实验室资产市场实验）对照。"}},{"id":"2609.02821","version":1,"title":"AI Contextual Measurement for Recovering Individual and Group-Level Effects: Validation Against Survey Measures and an Occupational Application","zh_title":"AI情境测量用于恢复个体与群体效应：基于调查测量的验证及职业应用","abstract":"Researchers increasingly use artificial intelligence to construct measures of social, organizational, and occupational characteristics that are absent from conventional surveys. We propose AICOME, AI COntextual MEasurement, a framework for evaluating whether AI-derived respondent-level measures can recover individual and group-level effects in contextual models. The key idea is that an AI measure constructed at the respondent level can be used to derive its group-level aggregate and its individual deviation, allowing researchers to estimate both between-group and within-group associations rather than treating AI measurement as response prediction alone. We validate the framework using the 2022 China Family Panel Studies (CFPS), where occupations provide the empirical grouping structure and several job-related survey variables provide validation benchmarks. For computer use, foreign-language use, weekly hours, and management responsibilities, we compare survey measures with AI-derived measures in response-level, model-level, contextual, and boundary-condition validations. The results show that AI contextual measurement can recover much of the contextual-model information contained in observed survey variables when rich respondent and job characteristics are available. Weekly hours provides the strongest validation case, with AI-derived measures reproducing the large negative between- and within-occupation associations with satisfaction observed in CFPS. The framework also identifies clear boundary conditions: performance deteriorates when information is restricted to occupation and basic demographics, and recovery is weaker when several related concepts are treated as simultaneously unobserved. The findings suggest that AICOME is most useful for recovering a limited number of theoretically important constructs from rich existing datasets.","authors":["Wenxin Jiang","Xuyang Wang","Yuxiao Wu"],"categories":["cs.AI","cs.LG"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-03","first_seen":"2026-09-03","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.02821","pdf_url":"https://arxiv.org/pdf/2609.02821","source_feed":"cs.AI","score":7,"bucket":"pending","rubric_hits":["A1","B1","B4"],"tags":["AI测量","验证框架","社会调查"],"reason":"用AI从调查数据构建个体测量并验证，与人类数据对照，但非直接仿真被试","model":"deepseek-v4-pro","scored_at":"2026-09-03T13:06:57","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-03","rank":20,"question":"AI 派生的个体层面测量能否在情境模型中恢复个体和群体层面的效应，从而替代传统调查测量进行情境分析？","design":"本研究不是直接仿真人类被试，而是提出并验证 AICOME 框架：利用大语言模型基于受访者特征（如职业、人口统计、工作信息）生成个体层面的测量（如计算机使用、外语使用、每周工时、管理责任），然后将该测量分解为群体均值（职业均值）和个体偏差，用于估计组间和组内关联，并与 CFPS 调查中的真实测量进行多层面验证。","baseline":"使用 2022 年中国家庭追踪调查（CFPS）中的真实调查变量作为基准，包括计算机使用、外语使用、每周工时、管理责任以及工作满意度等，以职业为分组结构。","findings":"当提供丰富的受访者和工作特征时，AI 情境测量能够恢复调查变量所包含的大部分情境模型信息，其中每周工时的验证效果最强，AI 派生测量重现了 CFPS 中观察到的工时与满意度之间显著的负向组间和组内关联。然而，当信息仅限于职业和基本人口统计特征时，性能下降；当多个相关概念同时被视为未观测时，恢复效果较弱。","reliability":"论文识别了明确的边界条件：当输入信息受限（仅职业和基本人口统计）时性能恶化；当多个相关概念同时缺失时恢复较弱。作者指出 AICOME 最适用于从丰富数据集中恢复少数理论重要构念，而非替代直接调查测量或大量联合缺失的项目。","relevance":"该研究虽非直接仿真人类被试，但提供了 AI 派生测量与真实人类调查数据对照的严格验证框架，对评估 LLM 在社会科学测量中的可靠性具有参考价值，值得阅读原文以了解其多层面验证方法和边界条件。","inspiration":"借鉴其将 AI 派生测量分解为组均值和个体偏差以同时估计组间和组内效应的做法，并采用多层面验证（响应级、模型级、情境级、边界条件）来评估测量质量｜可迁移到劳动经济学中职业特征对工资或工作满意度的影响研究，或组织经济学中企业文化对员工绩效的组间与组内效应分析｜设计：以职业为分组，使用 LLM 基于员工简历或工作描述生成职业特征（如自主性、技能要求）的个体测量，然后分解为职业均值和个体偏差，估计其对工资或离职行为的组间和组内效应，并与真实调查数据（如 CFPS 或美国 CPS）中的对应变量进行对照验证。"}},{"id":"2609.01627","version":1,"title":"The Utility of LLMs in Recommender Systems Explanation Evaluation","zh_title":"大语言模型在推荐系统解释评估中的效用","abstract":"Explanations play a crucial role in creating trustworthy recommender systems (RS), yet choosing a good explanation method presents challenges. Many explanation methods exist, but little guidance exists on which is best for which setting. Existing explanation generation methods often produce abstract outputs that require further formatting to become user-friendly, with a seemingly endless pool of options. Running user-based evaluations of all possible options is usually unfeasible, while automated evaluation metrics often either assess only the explainer's abstract output or require comparison with a ground truth, which is generally unavailable. Recent studies have shown that large language models (LLMs) can serve as ``judges'' for explanation evaluation, but their reliability has not yet been thoroughly explored. This paper studies the utility of LLMs in selecting an effective explanation method for a given application. We first explore their ability to generate explanation prototypes given varying information about the RS and the user. Specifically, we generate 18 distinct explanation prototypes, which are subsequently evaluated by 14 LLMs of varying sizes across two temperature settings. We compare these against human ratings derived from a user study. Our results show that while LLMs exhibit human-like rating patterns and achieve moderate rank correlation with human raters, their absolute rating agreement is low and varies substantially by model size and evaluation construct. We derive four practical recommendations: keep explanation-generation prompts concise, prefer larger models for evaluation, pre-test evaluation constructs, and audit explanations for factual accuracy, as neither humans nor LLMs reliably detect non-factual content.","authors":["Kathrin Wardatzky","Oana Inel","Luca Rossetto","Abraham Bernstein"],"categories":["cs.IR","cs.AI"],"primary_category":"cs.IR","announce_type":"cross","date":"2026-09-03","first_seen":"2026-09-03","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.01627","pdf_url":"https://arxiv.org/pdf/2609.01627","source_feed":"cs.AI","score":7,"bucket":"pending","rubric_hits":["A1","B1"],"tags":["LLM评估","推荐系统解释","人类对照"],"reason":"用LLM评估推荐解释，并与人类评分对照，属于仿真人类评估行为，但非典型被试仿真。","model":"deepseek-v4-pro","scored_at":"2026-09-03T13:06:51","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-03","rank":12,"question":"LLM能否替代人类评估推荐系统解释的质量，从而在原型阶段选择有效的解释方法？","design":"使用知识图谱推荐系统PGPR生成推荐和路径解释，用LLM基于不同信息方面（用户专业知识、用户历史、推荐信息、解释目标）生成18种解释原型；随后用14个不同规模的LLM在两种温度设置下对解释在六个质量维度上评分，并与人类评分比较。","baseline":"通过用户研究收集的人类评分，用于与LLM评分进行对比。","findings":"LLM表现出与人类相似的评分模式，与人类评分者达到中等秩相关；但绝对评分一致性低，且随模型规模和评估构念变化很大。","reliability":"论文承认LLM和人类都不可靠地检测非事实内容，因此需要审计解释的事实准确性；绝对评分一致性低，模型规模和评估构念影响可靠性。","relevance":"该研究用LLM模拟人类对推荐解释的评估，并与真实人类评分对照，属于人类仿真评估行为，但非典型被试仿真；对关注LLM评估可靠性和偏差的研究者有一定参考价值。","inspiration":"借鉴其系统生成多种解释原型并用多个LLM评分与人类评分对照的方法，可迁移到经济金融领域的文本解释评估（如信贷决策解释、投资建议解释）；设计上可用LLM生成不同风格的信贷拒绝解释，让多个LLM和人类被试对解释的公平性、可理解性等维度评分，以真实人类评分为基准检验LLM评估的可靠性。"}},{"id":"2609.02677","version":1,"title":"Eliciting ESG Preferences for Reinforcement Learning-Based Portfolio Optimization","zh_title":"基于强化学习的投资组合优化中ESG偏好的获取","abstract":"Modern portfolio management increasingly demands a balance between traditional risk-adjusted returns and strict Environmental, Social, and Governance (ESG) mandates. Current Reinforcement Learning (RL) approaches typically optimize for a single ESG provider, neglecting the significant divergence in rating methodologies across the industry and the unintuitive nature of manually weighting conflicting objectives. This paper addresses these limitations by formulating ESG-aware portfolio optimization as a Multi-Objective Reinforcement Learning (MORL) problem that simultaneously incorporates ratings from three distinct ESG agencies. To bridge the gap between high-dimensional algorithmic trade-offs and human decision-making, we integrate a Preference Elicitation framework using Gaussian Processes. This system enables practitioners to infer their latent utility functions through intuitive pairwise comparisons of candidate portfolios based on their Sharpe ratios and aggregate ESG scores. We systematically evaluate our framework by employing Large Language Model (LLM) personas to simulate Portfolio Managers operating under varied regional contexts. Empirical results using historical market data reveal that regional backgrounds fundamentally shift the derived preference weights. For instance, European-based personas tend to prioritize ESG alignment over financial returns, while Texas-based personas favor risk-adjusted performance. This work offers a highly adaptable framework that successfully aligns multi-objective algorithmic trading with diverse, real-world human sustainability preferences.","authors":["Giovanni Dispoto","Marcello Restelli","Carmine Ventre"],"categories":["q-fin.PM","cs.CE","cs.LG"],"primary_category":"q-fin.PM","announce_type":"cross","date":"2026-09-03","first_seen":"2026-09-03","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.02677","pdf_url":"https://arxiv.org/pdf/2609.02677","source_feed":"cs.LG","score":7,"bucket":"pending","rubric_hits":["A1","A3","B2"],"tags":["LLM仿真","投资组合优化","偏好获取"],"reason":"用LLM persona模拟投资经理偏好，涉及金融决策，但无真实人类数据对照","model":"deepseek-v4-pro","scored_at":"2026-09-03T13:06:55","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-03","rank":17,"question":"如何将多个ESG评级机构的冲突信号纳入强化学习投资组合优化，并通过偏好诱导使算法与人类投资经理的可持续性偏好对齐？","design":"使用LLM personas模拟不同地区（欧洲、德克萨斯、亚洲、美国）的投资经理，通过成对比较候选投资组合（基于夏普比率和ESG得分）来诱导其潜在效用函数，并测量诱导出的偏好权重。","baseline":"无对照","findings":"区域背景显著改变诱导出的偏好权重，欧洲persona更重视ESG对齐，德克萨斯persona更重视风险调整后收益。该框架能成功将多目标算法交易与多样化的现实人类可持续性偏好对齐。","reliability":"论文未讨论","relevance":"该研究用LLM personas模拟投资经理的ESG偏好，属于金融决策仿真，但缺乏真实人类数据对照，适合作为批判性案例或方法参考。","inspiration":"借鉴其用LLM personas模拟不同区域投资者并施加区域背景处理来诱导偏好的方法。｜可迁移到ESG投资偏好、社会责任投资或绿色金融产品选择等场景。｜用LLM personas模拟不同文化背景的投资者，处理为区域或制度环境，结果变量为对ESG与收益的权衡，对照真实调查或实验数据（如全球投资者调查、实验室投资实验）。"}},{"id":"2609.02092","version":1,"title":"Beyond Outcome Gaps: Process-Aware Fairness Diagnosis for LLM-based Multi-Agent Decision Systems","zh_title":"超越结果差距：基于LLM的多智能体决策系统的过程感知公平性诊断","abstract":"LLM-based multi-agent systems (MAS) are increasingly considered for high-stakes decision-making, yet outcome-based fairness audits can miss where risks arise within the decision trajectory. We present SCOPED-Hiring, a process-aware fairness diagnosis pipeline for LLM-based hiring MAS. SCOPED-Hiring constructs controlled resume variants, runs role-based hiring committees, logs over 311K structured decision trajectories, and converts trajectory fields into quantitative fairness signals organized by six diagnostic lenses: final outcome, counterfactual, process, pathway, dynamic, and design effects. SCOPED-Hiring reveals that balanced final hire rates can mask hidden trajectory unfairness in multi-agent decision trajectories: career gaps trigger suspicion, proxy cues shape qualification judgments, and identity cues lead to unequal investigation. Targeted repair guided by these diagnoses reduces total layered burden by 72.3% while shifting the hire rate by only 1.86 pp, showing that process diagnosis can guide effective repair. Project Page: https://scoped-hiring-project-page.vercel.app/","authors":["Yiran Zhao","Lu Zhou","Liming Fang","Yufei Chen","Jiafei Wu","Zhe Liu","Xiaogang Xu"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-03","first_seen":"2026-09-03","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.02092","pdf_url":"https://arxiv.org/pdf/2609.02092","source_feed":"cs.AI","score":6,"bucket":"other","rubric_hits":["D3","A3"],"tags":["LLM多智能体","公平性诊断","招聘模拟"],"reason":"用LLM多智能体模拟招聘决策过程，但无真实人类数据对照，属社会模拟边界情形。","model":"deepseek-v4-pro","scored_at":"2026-09-03T13:06:53","error":null,"has_summary":false,"summary":null},{"id":"2609.02620","version":1,"title":"Collective creativity in hybrid societies","zh_title":"混合社会中的集体创造力","abstract":"Generative AI is changing how cultural artifacts are created and circulated, and with it our understanding of creativity itself. Researchers disagree about whether these tools enrich or impoverish culture, and we argue that much of that disagreement comes from conflating two distinct components of creativity: novelty, a property of single artifacts, and diversity, a property of populations. We argue further that creativity in the context of generative AI is best understood as a property of hybrid collectives, or populations of interacting people and algorithms, rather than of individuals. AI-assisted ideation reliably raises the novelty of individual output while narrowing diversity in the aggregate, but this is not an inevitable consequence of putting machines in the loop. Because humans and models search in complementary ways, mixed groups can outperform and out-diversify groups of either kind alone, and machine-discovered solutions can enter human culture and persist there. What decides the outcome is composition: which agents are present, in what proportion, and how they are connected. The question is no longer whether AI helps or harms creativity, but which mixtures let individual gains accumulate without eroding collective diversity.","authors":["Mason Youngblood","Katie Mudd","Manuel Anglada-Tort","Cameron Jones","Elena Miu","Diana Omigie","Margaret Schedel"],"categories":["cs.AI","cs.CY","cs.MA"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-03","first_seen":"2026-09-03","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.02620","pdf_url":"https://arxiv.org/pdf/2609.02620","source_feed":"cs.AI","score":6,"bucket":"other","rubric_hits":["D3"],"tags":["LLM社会模拟","混合集体","创造力"],"reason":"讨论混合人机集体中的创造力，涉及LLM agent群体模拟社会过程，但无真实人…","model":"deepseek-v4-pro","scored_at":"2026-09-03T13:07:07","error":null,"has_summary":false,"summary":null},{"id":"2608.21377","version":2,"title":"Agentic Scaffolding Amplifies Sycophantic Behavior in Large Language Models","zh_title":"智能体脚手架放大大型语言模型中的谄媚行为","abstract":"Sycophancy in large language models, the tendency to prioritize user agreement over truthful responses, has been documented extensively but studied primarily in single-turn settings. This paper investigates a critical question: does subjecting LLMs to greater interaction scaffolding make sycophancy better or worse? Across 4,800 veracity judgments (200 statements $\\times$ 6 models $\\times$ 4 conditions), we find that the interaction scaffolding characteristic of agentic systems (feedback loops, reconsideration checkpoints, and iterative refinement) systematically amplifies sycophantic behavior. Multi-turn interaction, user pressure, and iterative self-refinement each provide additional opportunities for models to drift toward agreement, and this drift coincides with a mean accuracy drop of $-6.3$ percentage points, establishing the capitulation as harmful rather than corrective. More capable models show larger amplification effects, a troubling inversion of expectations. We introduce the concept of agentic sycophancy amplification (ASA) and two novel metrics: capitulation rate and sycophantic capitulation rate. Our results indicate that as AI systems acquire greater autonomy, sycophancy becomes compounding rather than merely persistent. Systems designed with human oversight loops may inadvertently create the conditions for this drift.","authors":["Thantham Jittham"],"categories":["cs.CL","cs.AI","cs.LG","cs.MA"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-03","first_seen":"2026-08-25","revised_at":"2026-09-03","abs_url":"https://arxiv.org/abs/2608.21377","pdf_url":"https://arxiv.org/pdf/2608.21377","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["LLM行为","谄媚偏差","智能体系统"],"reason":"研究LLM的谄媚行为，属于模型行为测量，非人类仿真，但涉及模型偏差，可迁移。","model":"deepseek-v4-pro","scored_at":"2026-09-03T13:07:10","error":null,"has_summary":false,"summary":null},{"id":"2609.02054","version":1,"title":"A Tri-Agent Framework for Evaluating and Aligning Question Clarification Capabilities of Large Language Models","zh_title":"用于评估和对齐大语言模型问题澄清能力的三智能体框架","abstract":"Large Language Models (LLMs) are increasingly deployed in interactive systems where understanding user intent precisely is paramount. A key capability for such systems is effective question clarification, especially when user queries are ambiguous or underspecified. This paper introduces a novel tri-agent framework for the robust evaluation of an LLM's ability to engage in clarifying dialogue. Our framework comprises three distinct LLM-based agents: (1) a Question Clarifying Agent (QCA), the system under evaluation, tasked with identifying ambiguities and posing clarifying questions; (2) a Respondent Agent (RA), designed to simulate human user responses, potentially including irrelevant or challenging replies; and (3) an Evaluator Agent (EA), an LLM-as-a-judge, which assesses the quality of the dialogue based on a comprehensive set of metrics. We detail a methodology for synthetic data generation in the supply chain domain as an example. We propose metrics evaluating ambiguity handling, question quality, dialogue efficiency, language appropriateness, and final intent alignment. We also briefly discuss the validation of the EA against human judgments. This work provides a structured approach to benchmark, validate, and improve the clarification capabilities of conversational LLM applications.","authors":["Yikai Zhao","Saurabh Pandey","Pradeep Kumar Misra"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-03","first_seen":"2026-09-03","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.02054","pdf_url":"https://arxiv.org/pdf/2609.02054","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM评估","多智能体","对话系统"],"reason":"用LLM模拟用户回答以评估澄清能力，属替代人工标注，非仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-09-03T13:07:02","error":null,"has_summary":false,"summary":null},{"id":"2609.02322","version":1,"title":"What Is Worth Representing? Representational Empowerment for Continual Model Construction","zh_title":"什么值得表征？持续模型构建中的表征赋权","abstract":"The first problem of modeling the world is not just estimating the right parameters or causal structure, but deciding what should be represented at all. We frame this problem as continual model construction: an agent maintains an environment-specific model M of an inaccessible world W and curates a persistent library L of reusable representational elements across environments. We propose Representational Empowerment (RepEmp) to score candidate elements by how much they expand the agent's future capacity to model and plan, complementing the classic definition of empowerment, but redefined as control over internal representations instead of external states. We realize the framework as a hierarchical Curator-Actor architecture and test it across three experiments. In a closed-vocabulary causal-learning task, human participants construct causal models at varying abstraction granularities to maximize goal reachability rather than fidelity to the world, a signature better predicted by RepEmp than by information-gain alternatives. Matched simulations reveal that RepEmp-guided construction contributes more than exploration to sufficient structure recovery and cross-task transfer. Finally, in an open-vocabulary planning domain, an LLM-augmented Curator builds more compact symbolic libraries, which also generalize better than baselines. Ablating RepEmp eliminates these benefits. Together, these results identify RepEmp as a key principle for continual model construction: deciding what to build, retain, and reuse under bounded resources.","authors":["Fei Dai","Hanqi Zhou","Alison Gopnik","Charley Wu"],"categories":["cs.LG","cs.AI"],"primary_category":"cs.LG","announce_type":"cross","date":"2026-09-03","first_seen":"2026-09-03","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.02322","pdf_url":"https://arxiv.org/pdf/2609.02322","source_feed":"cs.AI","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["表征学习","持续学习","LLM增强规划"],"reason":"LLM用于构建符号库辅助规划，非仿真人类被试，但涉及社会模拟元素，需人工判断。","model":"deepseek-v4-pro","scored_at":"2026-09-03T13:07:06","error":null,"has_summary":false,"summary":null},{"id":"2608.12062","version":2,"title":"Preference Tree Optimization: Enhancing Goal-Oriented Dialogue with Look-Ahead Simulations","zh_title":"偏好树优化：通过前瞻模拟增强目标导向对话","abstract":"Developing dialogue systems capable of engaging in multi-turn, goal-oriented conversations remains a significant challenge, especially in specialized domains with limited data. This research proposes a novel framework called Preference Tree Optimization (PTO), designed to iteratively improve agent models in such dialogue systems, by generating preference data using a method called Preference Tree with Look-Ahead. Focusing on Motivational Interviewing (MI) -- a counseling technique aimed at facilitating behavioral change -- we leverage virtual patients and an oracle evaluator to simulate conversations and generate rich preference datasets. By combining this method with Direct Preference Optimization (DPO), we aim to enhance the agent's decision-making capabilities over iterative training cycles. The proposed framework addresses data scarcity and advances the development of more nuanced and effective dialogue systems in goal-oriented domains. Experimental evaluations demonstrate that the PTO framework enhances dialogue agents' performance in goal-oriented conversations within the domain of Motivational Interviewing (MI). Models trained with PTO consistently outperformed the baseline in key metrics such as session satisfaction and working alliance. Additionally, incorporating look-ahead simulations led to improved long-term planning and more effective conversational strategies, with deeper look-ahead configurations yielding the most stable and high-scoring results.","authors":["Lior Baruch","Moshe Butman","Kfir Bar","Doron Friedman"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-03","first_seen":"2026-08-13","revised_at":"2026-09-03","abs_url":"https://arxiv.org/abs/2608.12062","pdf_url":"https://arxiv.org/pdf/2608.12062","source_feed":"cs.CL","score":3,"bucket":"other","rubric_hits":["C3"],"tags":["对话系统","动机访谈","偏好优化"],"reason":"该研究训练对话系统进行动机访谈，使用虚拟患者和评估器模拟对话，但目标是优化对话…","model":"deepseek-v4-pro","scored_at":"2026-09-03T13:07:12","error":null,"has_summary":false,"summary":null},{"id":"2609.02191","version":1,"title":"Examining the Vulnerability of Multi-Agent Medical Systems to Human Interventions for Clinical Reasoning","zh_title":"多智能体医疗系统对人类干预的脆弱性研究：临床推理视角","abstract":"Human interventions at fault points can alter the diagnostic accuracy of multi-agent medical systems. We defined fault points as moments in AI agent conversations, in which an agent's reasoning became most vulnerable to external influence. Using the MedQA dataset, this study analyzed simulated doctor-patient conversations to measure how interventions shifted reasoning and accuracy. Correct intervention methods showed an improvement in baseline diagnostic accuracy of up to 40%, while incorrect or bias-related interventions degraded performance by up to 6% and increased diagnostic drift and uncertainty. Beyond performance changes, our analysis revealed behavioral similarities between cognitive biases in simulated agent environments and real-world clinical practice. Examples included premature closure and susceptibility to misleading cues. Overall, these findings demonstrate that identifying and guiding fault points with human interventions may provide a mechanism for improving diagnostic robustness in multi-agent medical systems.","authors":["Benjamin C Liu","Dillon Mehta","Rishi Malhotra","Adam Zobian","Yong Ying Tan","Samir Chopra","Daniella Rand","Natalie Pang","Abhiram Gudimella","Kevin Zhu"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-03","first_seen":"2026-09-03","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.02191","pdf_url":"https://arxiv.org/pdf/2609.02191","source_feed":"cs.AI","score":3,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","医疗AI","诊断准确性"],"reason":"多智能体医疗系统协作解题，无人类被试仿真或真实人类数据对照","model":"deepseek-v4-pro","scored_at":"2026-09-03T13:07:03","error":null,"has_summary":false,"summary":null},{"id":"2608.22639","version":2,"title":"Poetic Heritage for Culturally Grounded Emotional Support: An Interaction Design Framework and Its Multimodal Agentic Instantiation","zh_title":"基于文化根基的情感支持的诗意遗产：交互设计框架及其多模态智能体实例","abstract":"Digital systems increasingly mediate emotional support, yet their interactions often remain culturally generic. Accordingly, we examine how a poetic tradition can be operationalized as a culturally grounded interactive medium and how generative AI can support such engagement. The resulting interaction design framework translates staged literature-based support and tradition-specific poetic aesthetics into guidance for digital system design. Poemithy instantiates the framework as a multimodal, LLM-enabled multi-agent system for guided reflection through classical Chinese poetry. A controlled between-subjects study with 50 participants compared text-only and multimodal versions. Both conditions showed medium-to-large within-session improvements in affect, anxiety, and emotion regulation, while between-condition tests detected no differences in these changes. Among secondary post-session user-experience measures, the clearest observed differences favored multimodality in perceived attunement, perceived task success, and engagement; usability and hedonic quality were descriptively higher, while workload did not differ detectably. Post-only cultural ratings were descriptively favorable in both conditions for cultural identification, poetry-engagement and dissemination intentions, and perceived cultural enrichment. Together, the findings suggest that culturally grounded content and structured guidance should anchor system design, while multimodal presentation may strengthen resonance and engagement. More broadly, the work shows how generative AI can mediate engagement with poetic heritage in culturally grounded emotional-support interactions.","authors":["Yangming Zhang","Zhiqian Li","Bin Wu","Qi Li","Jie Xu","Yunpeng Song","Liang Zhao"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"replace","date":"2026-09-03","first_seen":"2026-08-25","revised_at":"2026-09-03","abs_url":"https://arxiv.org/abs/2608.22639","pdf_url":"https://arxiv.org/pdf/2608.22639","source_feed":"cs.HC","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["情感支持","多智能体系统","人机交互"],"reason":"LLM多智能体用于情感支持交互，非仿真人类被试，无人类行为对照","model":"deepseek-v4-pro","scored_at":"2026-09-03T13:07:12","error":null,"has_summary":false,"summary":null},{"id":"2608.30188","version":2,"title":"GPAgentBench-2K: Benchmarking Large Language Model Agents in Complex Clinical Action Space","zh_title":"GPAgentBench-2K：在复杂临床动作空间中基准测试大语言模型智能体","abstract":"Large Language Models (LLMs) show great potential as clinical agents, yet existing benchmarks reduce clinical workflows to static predictions or unconstrained Markov Decision Processes (MDPs) with coarse action sets. To address this, we introduce GPAgentBench-2K, the first Constrained MDP (CMDP) LLM-agent benchmark for primary-care clinical decision-making, constructed from expert-validated records of real-world GP encounters. Our environment models a full spectrum of six foundational clinical actions, imposes a topological workflow prior over the action space, and operationalizes safety-informed abstention as a first-class outcome. Evaluating 16 state-of-the-art LLMs reveals a significant performance degradation as the action space scales. Crucially, we uncover a clinical quality-safety gap: even frontier models with the highest diagnosis accuracy violate safety constraints in over half of high-risk cases. Finally, we establish a reference point using Constrained Group Relative Policy Optimization (C-GRPO), and show that while explicitly modeling constraints improves performance over unconstrained RL methods, it remains far from clinically acceptable safety.","authors":["Boqi Chen","Xudong Liu","Yunke Ao","Heejin Do","Jianing Qiu"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-03","first_seen":"2026-09-01","revised_at":"2026-09-03","abs_url":"https://arxiv.org/abs/2608.30188","pdf_url":"https://arxiv.org/pdf/2608.30188","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C2"],"tags":["LLM智能体","临床决策","基准测试"],"reason":"该论文是LLM在临床决策中的基准测试，属于医疗AI应用，不涉及用LLM仿真人类…","model":"deepseek-v4-pro","scored_at":"2026-09-03T13:07:14","error":null,"has_summary":false,"summary":null},{"id":"2609.01832","version":1,"title":"Interpretable Symptom Vectors for Depression in a Large Language Model","zh_title":"大语言模型中抑郁症的可解释症状向量","abstract":"Patients with depression present with diverse symptom profiles, yet clinical practice routinely reduces this variation to a single severity score. Large language models (LLMs) can potentially capture various symptoms and their severity from patient speech. However, how depressive symptoms are represented inside LLMs remains poorly understood, limiting clinical trust. To examine whether internal model activations match clinician judgment, we analyzed the residual stream of Gemma-3-27B-PT using mechanistic interpretability techniques. Recording activations across symptom descriptions drawn from validated clinical instruments, we found that symptom groups geometrically separated the most at layer 21 across multiple distance metrics. Using Semantic Projection, we then projected held-out naturalistic text onto Symptom Vectors constructed from these instruments. The resulting per-symptom coefficients preserved clinician-annotated rank ordering across mood, somatic, and suicidality axes. Furthermore, a single depression vector in Layer 21 separates held-out depressive from non-depressive text (AUC = 0.789), which can be used as an emotional valence gate that restricts symptom projection to depressive speech. These results reveal a decorrelated, clinician-aligned symptom signal readable directly from internal activations, offering a mechanistic foundation for interpretable depression-assessment tools.","authors":["Fangyi Zhu","Ajay Subramanian","Allison Constant","Camille Wang","Ravish Gupta","Corey J. Keller"],"categories":["cs.CL","cs.AI","cs.LG","q-bio.NC"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-03","first_seen":"2026-09-03","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.01832","pdf_url":"https://arxiv.org/pdf/2609.01832","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["可解释性","心理健康","表征分析"],"reason":"研究LLM内部表征与临床判断的一致性，属于模型可解释性，非用LLM仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-09-03T13:06:59","error":null,"has_summary":false,"summary":null},{"id":"2609.02275","version":1,"title":"Do Large Language Models Capture the Diversity in their Training Data?","zh_title":"大语言模型是否捕捉到其训练数据中的多样性？","abstract":"Large language models are trained to model conditional distributions over text, yet it remains inadequately understood whether they capture the full diversity of plausible outputs present in their training data. We study this question through an information-theoretic lens by comparing the conditional entropy of model-generated outputs with that of the corresponding training data. Given paired input-output samples, we use conditional entropy and its matrix-based analogue based on von Neumann entropy to measure output variability beyond what is explained by the conditioning input, without requiring multiple reference outputs for the same prompt. Across LLM families with publicly available training data, including OLMo, Pythia, and GPT-Neo, we consistently find that model-generated outputs exhibit lower conditional entropy than their training data, across different model scales, sequence lengths, and decoding strategies. We observe a similar conditional diversity gap beyond language modeling, including class-conditioned ImageNet generators and text-conditioned models trained on MS-COCO. To address this gap, we propose a post-hoc correction mechanism that generates multiple outputs for each input and reweights them through a matrix-entropy projection, increasing conditional diversity while remaining close to the original model distribution. We prove the concavity of the matrix-based conditional entropy functional, which makes the resulting entropy-constrained projection a convex optimization problem, and develop a scalable mirror-descent algorithm for its implementation. Our results reveal a systematic conditional diversity gap between modern generative models and their training data, and provide an information-theoretic framework for measuring and mitigating this gap.","authors":["Youqi Wu","Farzan Farnia"],"categories":["cs.CL","cs.AI","cs.LG"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-03","first_seen":"2026-09-03","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.02275","pdf_url":"https://arxiv.org/pdf/2609.02275","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["生成模型","信息论","多样性"],"reason":"研究模型输出多样性，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-03T13:07:05","error":null,"has_summary":false,"summary":null},{"id":"2609.02496","version":1,"title":"Debias-SparseGPT: Bias-Aware Pruning for Large Language Models","zh_title":"Debias-SparseGPT：面向大语言模型的偏差感知剪枝","abstract":"Model compression techniques such as pruning and quantization facilitate the efficient deployment and acceleration of Large Language Models (LLMs). However, recent studies show that weight sparsification methods, such as SparseGPT, can amplify existing biases in models, with outputs varying significantly depending on persona cues in the prompt. In this paper, we introduce Debias-SparseGPT, a post-training pruning method incorporating representational debiasing using a second-order term defined over demographically contrasting inputs. We perform empirical validation of our method over a wide range of generative LLMs. Across models and sparsity regimes (25%, 50%, and structured 2:4 sparsity), Debias-SparseGPT consistently reduces pruning-induced bias compared to SparseGPT while preserving model perplexity and zero-shot accuracy. Under the most restrictive 2:4 structured sparsity pattern, which most aggressively degrades model quality, augmenting the calibration set with long-context, content-rich examples further improves both downstream performance and fairness. Overall, Debias-SparseGPT advances the bias-performance trade-off while preserving the computational efficiency of sparse models.","authors":["Irina Proskurina","Guillaume Metzler","Antoine Gourru","Julien Velcin"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-03","first_seen":"2026-09-03","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.02496","pdf_url":"https://arxiv.org/pdf/2609.02496","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["模型压缩","偏差缓解","剪枝"],"reason":"论文关注模型压缩中的偏差缓解，不涉及用LLM仿真人类被试或与人类数据对照，属于…","model":"deepseek-v4-pro","scored_at":"2026-09-03T13:07:07","error":null,"has_summary":false,"summary":null},{"id":"2609.02651","version":1,"title":"WinoQueer-NL: Assessing Bias in Dutch Language Models toward LGBTQ+ Identities","zh_title":"WinoQueer-NL：评估荷兰语语言模型对LGBTQ+身份偏见","abstract":"While English language models have been widely examined for anti-queer bias, Dutch models remain understudied. To address this gap, we developed a culturally and linguistically adapted Dutch dataset based on the English WinoQueer benchmark, containing pairs of stereotypical and counter-stereotypical sentences. To validate and expand it, we conducted an online survey with 43 Dutch queer participants, confirming 145 of 171 stereotypes as culturally relevant and identifying 22 new biases through free-text responses. The final released dataset, comprising 42,906 sentences, was evaluated using a range of Dutch-specific and multilingual models, including both masked language models (MLMs) and autoregressive language models (ARLMs), with bias measured via a score comparing log-likelihoods of stereotypical versus counter-stereotypical sentences. While the mean bias score across models appeared neutral (~50%), closer analysis revealed significant disparities: some models favored stereotypical sentences up to 97% of the time for transgender identities, but only 6% of the time for gay-related pairs, with transgender and non-binary identities consistently receiving the highest bias scores. Our findings highlight the importance of culturally grounded datasets for evaluating and mitigating biases that disproportionately impact marginalized groups in Dutch language models.","authors":["Jiska Beuk","Gerasimos Spanakis"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-03","first_seen":"2026-09-03","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.02651","pdf_url":"https://arxiv.org/pdf/2609.02651","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["偏见评估","语言模型","LGBTQ+"],"reason":"评估语言模型偏见，非仿真人类被试，无行为对照","model":"deepseek-v4-pro","scored_at":"2026-09-03T13:07:07","error":null,"has_summary":false,"summary":null},{"id":"2609.01873","version":1,"title":"Epistemic Sybil Resistance: Multiplying AI Agents Without Multiplying Evidence","zh_title":"认知女巫抵抗：在不增加证据的情况下增加AI智能体","abstract":"Multi-agent AI systems improve inference by spawning agents and synthesizing reports. But another agent is not another observation: apparently independent reports may descend from the same evidence, and genuinely independent evidence can produce nearly identical reports. We formalize this as an epistemic Sybil problem. A report Z is an epistemic Sybil extension relative to reports R when I(Theta; Z | R) = 0. No report-only aggregator can generally distinguish replication from independent corroboration: identical reports can warrant different posteriors under unobserved ancestry. A Gaussian shared-root model shows common ancestry does not imply complete redundancy. Repeated extraction adds information toward a source-level ceiling, and correlated extraction errors, which a shared base model can induce among independent agents, lower that ceiling further. We test these predictions with more than 20,000 controlled LLM-agent report and extraction calls on synthetic evidentiary documents. Holding one evidence root fixed while report multiplicity rises from 1 to 32 collapses naive posterior coverage from 0.940 to 0.263. Holding report count fixed while evidence-root multiplicity rises from 1 to 16 closes the gap, and the aggregators are statistically indistinguishable at k = 16. The agent's replicate extraction errors are correlated (gamma_cal = 0.719, estimated out of sample), and a correlated-extraction aggregator restores calibration accordingly. A controlled manipulation isolates representation similarity from evidential ancestry. It changes a report-space deduplication mechanism's mean inferred cluster count by 1.425 (95% CI [1.363, 1.485]), whereas a fourfold change in true ancestry changes it by only 0.040 ([-0.045, 0.120]). Collective inference should therefore track evidential ancestry and dependence, not agent or report multiplicity or similarity.","authors":["Marc Bara"],"categories":["cs.AI","cs.MA"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-03","first_seen":"2026-09-03","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.01873","pdf_url":"https://arxiv.org/pdf/2609.01873","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","证据聚合","推理校准"],"reason":"研究多智能体推理中的证据冗余与聚合，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-03T13:07:00","error":null,"has_summary":false,"summary":null},{"id":"2609.02231","version":1,"title":"PhoenixNest-Video: Evidence-Grounded Multimodal Agent Framework for Automated Video Interview Assessment","zh_title":"PhoenixNest-Video：基于证据的多模态智能体框架用于自动化视频面试评估","abstract":"Interview assessment requires per-criterion judgments grounded in behavioral evidence, yet surging applicant volumes have made human-only evaluation costly and inconsistent, while existing AI approaches yield opaque scores without traceable rationale. We introduce PhoenixNest-Video, an evidence-grounded multimodal agent framework for automated video interview assessment. It builds a semantic video graph as structured working memory, performs rubric-conditioned retrieval with cross-modal verification across visual, audio, and textual streams, and produces per-criterion scores anchored to the candidate's materials. A Scorer trained via Rubrics-based Reinforcement Learning with dual rewards for rubric alignment and score-level differentiation internalizes the discriminative structure of multi-level rubrics. PhoenixNest-Video attains 91.50\\% grade-level accuracy on VInterview-2025, outperforming substantially larger proprietary models. A compact, rubric-grounded agent therefore scores candidates in closer agreement with an expert panel than direct prompting of much larger models, and exposes the evidence behind each score for human review.","authors":["Fan Yuxuan","Huang Miaojun","Zhang Haimei","Wu Jingshen","Liu Hao"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-03","first_seen":"2026-09-03","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.02231","pdf_url":"https://arxiv.org/pdf/2609.02231","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["多模态智能体","面试评估","强化学习"],"reason":"自动面试评估，非人类仿真实验，无人类行为对照","model":"deepseek-v4-pro","scored_at":"2026-09-03T13:07:03","error":null,"has_summary":false,"summary":null},{"id":"2607.24780","version":2,"title":"LivingArena: Do LLMs Know What Other LLMs Don't? Peer-Probing as Scalable Evaluation","zh_title":"LivingArena：大语言模型是否知道其他模型不知道什么？基于同伴探测的可扩展评估","abstract":"Fixed benchmarks are costly to renew and cannot adapt their questions to model-specific failures. We ask whether LLMs can instead discover one another's weaknesses and turn those observations into an evaluation process. To study this question, we introduce \\textbf{LivingArena}, an automated peer-probing framework in which models take turns testing one another. Using the interaction history, each questioner identifies potential weaknesses of its opponent and constructs targeted, verifiable questions to probe them. A 3,600-round tournament of ten models reveals a clear role asymmetry: strong answerers are not always reliable questioners, because they may generate internally inconsistent tests or fail to verify their own reference answers. After a questioner exposes an answerer's failure, it is more likely to pursue the same capability domain, while the answerer's weakness recurs on independently generated questions, including questions written by different models. These findings show that peer probing can reveal persistent model-specific weaknesses while separately evaluating answering and reliable test construction. By automating this process and allowing test difficulty to evolve with model capabilities, LivingArena provides a \"living\" benchmark for model development, red-teaming, and capability-aware multi-agent coordination. We publicly release our code: https://github.com/galaxyChen/LivingArena","authors":["Xingyu Chen","Rui Wang","Zhaopeng Tu","Liefeng Bo"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"replace","date":"2026-09-03","first_seen":"2026-07-29","revised_at":"2026-09-03","abs_url":"https://arxiv.org/abs/2607.24780","pdf_url":"https://arxiv.org/pdf/2607.24780","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["LLM评测","多智能体","知识边界探测"],"reason":"纯多智能体相互出题评测，不涉及人类行为仿真或对照。","model":"deepseek-v4-pro","scored_at":"2026-07-29T19:14:20","error":null,"has_summary":false,"summary":null},{"id":"2608.20574","version":3,"title":"FlavourBench: Executable Culinary Reward Maps for Language Model Evaluation and Post-Training","zh_title":"FlavourBench：用于语言模型评估和后训练的可执行烹饪奖励地图","abstract":"We introduce FlavorBench: a benchmark for Compiling Dense Deterministic Answer Maps from a Versioned Culinary Embeddings Model. We test 27 frontier large language model endpoints on 534 substitution, pairing and constraining tasks for tasks that request a 3-ingredient portfolio from 8 candidates and score all 56 resulting portfolios. We conducted multiplicity-controlled paired tests on 101 of 351 model contrasts for this task-set. The largest point estimate on this task-set was achieved by Grok 4.6 at 65.1. The same rankings for this task-set were also achieved on several independently-compiled panels (using a variety of familiar metrics, task filters, etc.) and 3 public Epicure checkpoints. We present a 3-seed post-training study where LoRA SFT of a Qwen3-0.6B checkpoint on 270 optimal answers for Epicure to score on this task-set resulted in a 13.3 point gain on 84 anchor-disjoint maps (compared to format and label-matched control; 95% CI: 6.52, 20.29; p = 0.000170).","authors":["Josef Chen (Independent Researcher)","Erim Hayretci (Imperial College London)"],"categories":["cs.AI","cs.CY","cs.LG","cs.SE"],"primary_category":"cs.AI","announce_type":"replace","date":"2026-09-03","first_seen":"2026-08-24","revised_at":"2026-09-03","abs_url":"https://arxiv.org/abs/2608.20574","pdf_url":"https://arxiv.org/pdf/2608.20574","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C2"],"tags":["LLM评估","烹饪任务","基准测试"],"reason":"论文评估LLM在烹饪任务上的表现，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-08-28T13:01:56","error":null,"has_summary":false,"summary":null},{"id":"2608.21969","version":3,"title":"ToSCA: Leveraging Hierarchical Reinforcement Learning on Temporal and Strategic Abstractions of Conversational Agents","zh_title":"ToSCA：利用对话智能体的时间和策略抽象的分层强化学习","abstract":"Humans naturally exhibit multiple forms of abstraction in reasoning and interaction, including temporal abstraction across decision timescales and strategic abstraction over communicative intents. Inspired by these complementary abstractions, we propose a two-level hierarchical reinforcement learning (HRL) framework for conversational agents that bridges the gap between existing token-level and utterance-level RL methods. Built upon a two-level Markov decision process (MDP), our framework conditions token-level response generation on utterance-level actions represented by explicit textual strategies. Based on theoretical analysis and efficiency considerations, we employ DQN to optimize the high-level Q-network and PPO to train the low-level actor-critic. To further alleviate reward sparsity and facilitate convergence, we introduce a dual-granularity reward mechanism that combines the utterance-level satisfaction score with token-level intrinsic self-consistency and a KL-divergence penalty. Experiments on both daily-life and emotional support conversations demonstrate that our method consistently outperforms a wide range of baselines in both strategy determination and response quality. Our implementation is available at https://github.com/AaronJi/ToSCA.","authors":["Xiaoyu Wang","Qingqing Gu","Yue Zhao","Teng Chen","Yuqi Cao","Xiaokai Chen","Hongyan Li","Luo Ji"],"categories":["cs.CL","cs.HC","cs.LG"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-03","first_seen":"2026-08-25","revised_at":"2026-09-03","abs_url":"https://arxiv.org/abs/2608.21969","pdf_url":"https://arxiv.org/pdf/2608.21969","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["对话系统","分层强化学习","自然语言生成"],"reason":"该论文研究对话系统的分层强化学习，旨在提升对话质量，属于角色扮演聊天机器人，无…","model":"deepseek-v4-pro","scored_at":"2026-09-03T13:06:59","error":null,"has_summary":false,"summary":null},{"id":"2609.02730","version":1,"title":"CORAL: An LLM-Native Harness for Production Recommender Systems","zh_title":"CORAL：面向生产推荐系统的 LLM 原生自动化框架","abstract":"Production recommender systems shape what billions of people see, and sustaining their performance requires continual optimization: as content, user behavior, and upstream models shift, the choices governing retrieval, ranking, and serving must be revisited. Traditionally, human engineers test such changes through online experiments--a slow, reactive process limited by engineering effort, leaving parts of the system unrevised as conditions change. Although large language models have been applied to ranking, user modeling, and offline model development, few systems place an agent in a continual closed loop that acts on a live recommender and learns from the measured effects of its decisions. We present CORAL (Constraint-Optimized Recommender via an Agentic Loop), an LLM-native harness that closes this loop: each cycle, the agent observes operating signals, reasons over a memory of past decisions and outcomes, and invokes tools--including a numerical optimizer that keeps changes within a fixed operating budget--to reconfigure the recommender, with measured outcomes informing the next cycle. We formulate this as a partially observed, non-stationary, constrained optimization problem in which the policy improves in context, without parameter updates, from its prior actions. Across two large-scale social platforms, evaluated with A/B experiments, the same harness improves engagement at no additional serving cost on one and reduces serving cost without degrading engagement on the other, spanning the engagement-efficiency frontier. Performance improves as the loop iterates, suggesting that a single agentic loop can automate continual optimization work traditionally performed by human algorithm engineers under explicit guardrails.","authors":["Muhammad Rafay Azhar","Yuhang Zhou","Gilbert Jiang","Yuchen Wang","Rahul Sharma","Matthew DeSousa","Jiayi Liu","Xin Guo","Lizhu Zhang","Xiangjun Fan"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-03","first_seen":"2026-09-03","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.02730","pdf_url":"https://arxiv.org/pdf/2609.02730","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["推荐系统","LLM agent","系统优化"],"reason":"LLM agent 用于优化推荐系统，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-03T13:07:08","error":null,"has_summary":false,"summary":null},{"id":"2609.01625","version":1,"title":"Whose Judgments Count? Representation Gaps in Crowdsourced Content Moderation Produce Unequal Protection from Perceived Toxicity","zh_title":"谁的判断算数？众包内容审核中的代表性差距导致感知毒性保护不平等","abstract":"Content moderation is a central form of digital governance, yet people disagree over what content should be removed from shared online spaces. While platforms aggregate human judgments to build moderation systems, it remains unclear how this process shapes which users are protected from content they perceive as toxic. We address this gap by combining large-scale judgment data with counterfactual simulations that trace how the demographic composition of moderator pools shapes the distribution of protection across users. Applying this framework to removal judgments from 16,221 U.S. respondents evaluating 102,463 comments from Twitter, Reddit, and 4chan, we find demographic heterogeneities in moderation demand. We further reveal a consistent pattern of in-group protection: reductions in perceived toxicity accrue disproportionately to users who share the demographic identities of the moderator pool. Crucially, moderator pools that mirror the demographic composition of self-identified moderators on Prolific widen these disparities relative to a nationally representative baseline, while even fully representative pools fail to ensure equal protection: Black and LGB users remain underprotected unless they are represented well beyond their population share. These findings show that unequal protection from perceived toxicity can arise structurally from the aggregation of stratified removal standards, making the demographic composition of moderation inputs a key determinant of who is protected online.","authors":["Zhaodi Chen","Byungkyu Lee"],"categories":["cs.SI","cs.CL","cs.CY"],"primary_category":"cs.SI","announce_type":"cross","date":"2026-09-03","first_seen":"2026-09-03","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.01625","pdf_url":"https://arxiv.org/pdf/2609.01625","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["内容审核","众包","代表性差距"],"reason":"研究众包内容审核中人口构成的影响，未使用LLM仿真人类被试，不涉及LLM替代人…","model":"deepseek-v4-pro","scored_at":"2026-09-03T13:06:51","error":null,"has_summary":false,"summary":null},{"id":"2609.02745","version":1,"title":"Incremental Pooled LLM Evaluation for Cost-Effective Retrieval Model Selection","zh_title":"用于成本效益检索模型选择的增量池化LLM评估","abstract":"Selecting a retrieval model for a production RAG system requires reliable comparative evaluation, but obtaining relevance judgments at scale is expensive and difficult to repeat as new candidate systems arrive. We study pooled LLM evaluation, in which an LLM judges the union of documents retrieved by the current set of candidate systems, and the pool is then expanded incrementally as new systems are introduced by judging only the new documents they contribute. These judgments are reused to evaluate all systems on a common basis. We validate this approach on four retrieval benchmarks with 11 systems spanning dense, sparse, and hybrid configurations, and deploy it to compare 62 retrieval configurations for a financial news QA system. Pooled LLM rankings correlate strongly with gold-standard evaluation across datasets, and 97% of pairwise system orderings are preserved once bootstrap uncertainty in the qrels is taken into account. In production, document overlap yields 65-80% judgment reuse and up to 4.9x lower evaluation cost, allowing teams to benchmark new retrieval candidates without re-judging previously assessed documents. These results suggest pooled LLM evaluation is a practical and cost-effective workflow for incremental retrieval model selection in deployed systems.","authors":["Max Nelson","Hanoz Bhathena","Aviral Joshi","Saket Sharma"],"categories":["cs.IR","cs.CL"],"primary_category":"cs.IR","announce_type":"cross","date":"2026-09-03","first_seen":"2026-09-03","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.02745","pdf_url":"https://arxiv.org/pdf/2609.02745","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["信息检索","LLM评估","RAG系统"],"reason":"纯检索模型评估，LLM仅作相关性判断工具，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-03T13:07:08","error":null,"has_summary":false,"summary":null},{"id":"2609.02242","version":1,"title":"Propose to Learn, Learn to Propose: Evaluability-Aware Assistance under Bounded Rationality","zh_title":"提出以学习，学习以提出：有限理性下的可评估性感知辅助","abstract":"AI assistants often collaborate by proposing candidate edits, plans, or designs that users evaluate before adoption. Existing assistance methods focus on proposal quality or user-goal inference, often assuming that the user can reliably evaluate any proposal, which can fail in practice because of bounded rationality. We study evaluability-aware proposal planning, where proposals serve both as task interventions and as probes for learning latent preferences and evaluation constraints, where the resulting belief updates then guide later proposals. We formalise this setting as ProSE, a hidden-parameter sequential assistance problem, and instantiate it with a KL-regularised bounded-rational binary response model in which acceptance trades off value gain against a distance-dependent evaluability penalty. Analysing the planning consequence of this likelihood reveals that likely accepted proposals and informative probes need not coincide, which explains why planners that only pursue acceptance systematically underperform. We operationalise ProSE with \\textsc{ProSE-Plan}, a depth-2 Bayes-adaptive planner that scores proposals by possible responses and response-induced posterior beliefs. In controlled graph simulations, \\textsc{ProSE-Plan} improves over evaluability-unaware and myopic baselines when evaluation cost is the bottleneck, and a probe-commit ablation confirms that our approach selects informative proposals that simpler methods miss. Our results thus identify user evaluability as a planning-relevant dimension of AI assistance, complementary to generation quality and preference inference.","authors":["Yifan Zhu","Sammie Katt","Samuel Kaski"],"categories":["cs.AI","cs.HC","cs.MA"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-03","first_seen":"2026-09-03","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.02242","pdf_url":"https://arxiv.org/pdf/2609.02242","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["AI辅助","有限理性","规划"],"reason":"研究AI助手与用户协作，不涉及LLM仿真人类被试或人类行为对照","model":"deepseek-v4-pro","scored_at":"2026-09-03T13:07:04","error":null,"has_summary":false,"summary":null},{"id":"2609.02246","version":1,"title":"LLM-as-a-Judge Is Not an Oracle: Why Self-Improving Agents Need Deterministic Guardrails","zh_title":"LLM作为评判者并非神谕：为何自改进智能体需要确定性护栏","abstract":"Self-improving agent pipelines have a problem at their center. An optimizer rewrites prompts to score higher, and the score comes from a judge that is itself an LLM. That judge has the last word on whether the system is getting better, and our position is that it has not earned it. The judge should be demoted from oracle to advisor: its verdict becomes one input among several, and every change is gated instead by a deterministic verification layer the judge cannot override. We reached this position by building the alternative and running it. Over months of running autonomous prompt-optimization loops in production across contract analysis, compliance review, and code quality, we cataloged eleven ways the evaluation signal failed, in four classes: judge bias, harness and metric failures, ground-truth errors, and reward hacking. Agents achieved perfect scores by reading cached answer keys from their environment, a 100% pass rate concealing 68% true capability. A corrupted ground-truth label caused the optimizer to delete correct compliance rules to agree with it. A syntactically broken prompt was promoted as the winner because a silent parser fallback improved the metric. Attempts to fix the judge by rewriting its rubric plateaued; the only reliable gain came from a structural constraint on its output order. In response we describe PROCTOR, a Teacher-Student loop in which a stateful orchestrator holds all tool access, stateless subagents diagnose failures and draft mutations they cannot apply, and a Teacher grades those mutations under five deterministic guardrails: hermetic sandboxes, capability-disjoint roles, acceptance checks that outrank the Teacher, frozen holdouts, and canary cases engineered so that a perfect score is itself evidence of cheating. We report the failures this prevented, and, because the Teacher is itself an LLM judge, the failures it did not.","authors":["Vansh Wahi"],"categories":["cs.AI","cs.LG"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-03","first_seen":"2026-09-03","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.02246","pdf_url":"https://arxiv.org/pdf/2609.02246","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["LLM评判器","自改进智能体","多智能体系统"],"reason":"研究LLM评判器在自改进智能体中的可靠性，属多智能体系统，不涉及人类行为仿真或…","model":"deepseek-v4-pro","scored_at":"2026-09-03T13:07:04","error":null,"has_summary":false,"summary":null},{"id":"2609.02750","version":1,"title":"Bilevel Coordinated Reflection: A Game-Theoretic Approach to Multi-Agent LLM Systems","zh_title":"双层协调反思：多智能体LLM系统的博弈论方法","abstract":"Multi-agent LLM systems commonly use an orchestrator to decompose a task for a team of workers and then improve through textual reflection. Despite strong empirical results, these systems lack a unified account of coordination, memory improvement, and the role of external verification. We model orchestrator-worker interaction as a bilevel coordination game: under bounded coupling, the workers' local-update game is an approximate potential game whose equilibrium slack is controlled by decomposition quality. We then analyse reflection as stochastic movement over semantic memory states. For free-form reflection, we derive a finite-time upper bound, prove worst-case tightness, and give a positive lower bound under a falsifiable persistent-harm condition. We further prove an information-theoretic impossibility result: no gate that observes only the generated transcript can improve uniformly over text-indistinguishable environments, whereas an environment-grounded gate can. Motivated by this separation, we introduce Stochastic Reflective Memory Ascent (SRMA), which accepts a candidate memory only after a grounded evaluation risk strictly decreases. Under calibration and non-degenerate corrective mass, SRMA converges exactly, geometrically or polynomially; matching constructions show that both rate regimes are order-tight. We also provide confidence gating for stochastic evaluation and re-anchoring guarantees for piecewise-stationary environments. Experiments instantiate these objects with environment-grounded metrics and test the predicted coordination and drift laws. On 500 SWE-bench instances, the complete Kimi-based system resolves 72.2% versus a 70.8% public mini-SWE-agent reference. Code: https://github.com/YihangChen9/Bilevel-Coordinated-Reflection","authors":["Yihang Chen","Yuxiang Chen","Yuxuan Huang","Meng Fang","Weilin Luo","Jun Wang"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-03","first_seen":"2026-09-03","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.02750","pdf_url":"https://arxiv.org/pdf/2609.02750","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","博弈论","任务协调"],"reason":"纯多智能体协作解题，无人类行为对照，不涉及人类仿真","model":"deepseek-v4-pro","scored_at":"2026-09-03T13:07:08","error":null,"has_summary":false,"summary":null},{"id":"2609.01976","version":1,"title":"Knowing Is Not Enough: Information Retrievability as a Precondition to Effective LLM Oversight","zh_title":"知道还不够：信息可检索性作为有效LLM监督的前提","abstract":"Large language models (LLMs) are increasingly embedded in organizational work, yet their errors often pass human review. Prior research locates such failures in users' capability to review LLM output or their engagement in doing so. We develop an alternative, retrieval-based account of human oversight and posit that error detection is more effective when oversight-relevant information is accessible to users at the moment of review. Across two randomized lab-in-the-field experiments with 640 customer-facing employees, we show that self-generated explanations improve error detection and strengthen recall of verification-relevant reasoning, while cues that reactivate such reasoning help sustain detection under repeated LLM use. Theoretically, we identify information retrievability as a distinct precondition for effective oversight and specify generative encoding and cue-supported reactivation as mechanisms that build and sustain it. Practically, lightweight onboarding self-explanations and daily retrieval cues can make human oversight more resilient as LLM use becomes routine.","authors":["Xinyu Fu","Narayan Ramasubbu","Dennis Galletta"],"categories":["cs.HC","cs.AI"],"primary_category":"cs.HC","announce_type":"cross","date":"2026-09-03","first_seen":"2026-09-03","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.01976","pdf_url":"https://arxiv.org/pdf/2609.01976","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["人机交互","LLM监督","信息检索"],"reason":"研究人类对LLM输出的监督，不涉及用LLM仿真人类被试","model":"deepseek-v4-pro","scored_at":"2026-09-03T13:07:01","error":null,"has_summary":false,"summary":null},{"id":"2609.02296","version":1,"title":"Meeting the Coming Wave: The Emerging Politics of AI and Work across 33 Parliaments","zh_title":"迎接浪潮：33个议会中人工智能与工作的新兴政治","abstract":"A new politics of artificial intelligence and work is taking shape across party systems, but comparative politics has yet to map it. Using 1,514,950 parliamentary speeches from 33 parliaments (2023-2026), we show this politics follows a different logic than political economy expects. Research anticipates that technological disruption generates demands for compensation; instead, compensation accounts for just 2.3% of response-frame mentions, while enablement and investment dominate (55.2%), regulation and restriction follow (21.8%), and training (20.6%) appears at similar rates across families. Parties disagree instead over what AI means for work and how far this technology should be restrained. The mainstream and radical right support unrestricted enablement; the left is critical but divided on remedy. Social democrats stay adoption-oriented; greens split evenly. The radical left is the clearest force for restriction. The AI conflict thus concerns not compensation after disruption, but whether politics should enable technological change or govern its trajectory.","authors":["Juliana Chueri","Petter T\\\"ornberg"],"categories":["cs.CY","cs.ET"],"primary_category":"cs.CY","announce_type":"new","date":"2026-09-03","first_seen":"2026-09-03","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.02296","pdf_url":"https://arxiv.org/pdf/2609.02296","source_feed":"cs.CY","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["政治学","文本分析","AI政策"],"reason":"论文分析真实议会演讲，未使用LLM仿真人类被试，不涉及人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-09-03T13:07:05","error":null,"has_summary":false,"summary":null},{"id":"2609.01639","version":1,"title":"Inverse planning of social interactions in relationships","zh_title":"人际关系中社会互动的逆向规划","abstract":"We propose a formal account of how structured, shared knowledge about social relationships shapes action interpretation. The model represents relationships as constraints in a social environment, analogous to boundaries or obstacles in a physical environment and operating within the same generative model, but exerting distinct constraints on action. As an initial test of this framework, we draw on research across the social sciences to capture in the models how one dimension of relationships -- formality versus intimacy -- shapes how people interpret interpersonally vulnerable behavior. We test this account in stories of naturalistic everyday situations, extending structured models of action understanding to open-ended contexts. Across six preregistered experiments (N = 1,554), the model captures people's inferences about desires, physical environments, and social relationships. This work formalizes how relationships can constrain -- and be revealed through -- everyday action.","authors":["Alicia M. Chen","Ashley J. Thomas","Joshua B. Tenenbaum","Rebecca Saxe"],"categories":["physics.soc-ph","q-bio.NC"],"primary_category":"physics.soc-ph","announce_type":"new","date":"2026-09-03","first_seen":"2026-09-03","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.01639","pdf_url":"https://arxiv.org/pdf/2609.01639","source_feed":"physics.soc-ph","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["社会认知","逆向规划","人类实验"],"reason":"论文研究人类如何理解社会互动，未使用LLM，不涉及仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-09-03T13:06:59","error":null,"has_summary":false,"summary":null},{"id":"2607.10628","version":2,"title":"Anamnesis: An Open-Source Platform for Large-Scale Backstory-Conditioned Survey Simulation","zh_title":"Anamnesis：大规模背景条件调查仿真的开源平台","abstract":"We present Anamnesis, an interactive system for demographically controllable survey simulation using large language models. Open-source and designed for non-technical users/researchers, Anamnesis enables the prototyping and stress-testing of survey instruments on virtual populations rather than real human subjects. The platform operationalizes the recently introduced Anthology and Alterity frameworks, which use structured narrative backstories to condition model responses, within a unified web interface. It supports open-ended generation, probabilistic demographic resampling, and multimodal (image and audio) surveys. We evaluate the system through two case studies: (1) replicating segments of Pew Research Center's American Trends Panel (ATP) on political typology and biomedical issues and (2) emulating human preference in the New Yorker Caption Contest. In both cases, Anamnesis produces opinion distributions that more closely match real-world survey data than standard persona-prompting baselines, offering a transparent, reproducible, and open-source alternative to proprietary simulation services.","authors":["Song-Ze Yu","Joseph Suh","Serina Chang","David M. Chan"],"categories":["cs.CL","cs.AI","cs.HC"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-02","first_seen":"2026-07-12","revised_at":"2026-09-02","abs_url":"https://arxiv.org/abs/2607.10628","pdf_url":"https://arxiv.org/pdf/2607.10628","source_feed":"cs.CL","score":10,"bucket":"selected","rubric_hits":["A1","A2","A3","A5","B1","B2","B4"],"tags":["LLM仿真","调查模拟","人类数据对照"],"reason":"平台用LLM仿真调查，与真实数据对照，复现舆论分布，评估可靠性，直接相关。","model":"deepseek-v4-pro","scored_at":"2026-09-02T13:03:16","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-02","rank":1,"question":"如何构建一个开源、可交互的平台，利用大语言模型和结构化叙事背景来模拟多样化人群的调查回答，并验证其与真实人类调查数据的匹配度？","design":"Anamnesis平台使用大语言模型（如Gemini 2.5 Flash）扮演虚拟受访者，通过结构化叙事背景（backstories）而非简单人口统计列表来条件化模型响应；支持概率性人口重采样、开放式生成和多模态（图像、音频）调查；在案例研究中，模拟Pew Research Center的American Trends Panel三个波次（政治类型学、生物医学、AI与人类增强）的多选题调查，以及New Yorker Caption Contest的多模态偏好选择。","baseline":"对照的真实人类数据包括：Pew Research Center American Trends Panel（ATP）三个波次的真实调查回答分布，以及New Yorker Caption Contest中基于大规模众包投票的真实人类偏好标签。","findings":"Anamnesis平台在三个ATP波次上产生的意见分布比标准人物提示（persona-prompting）基线更接近真实调查数据，在Wasserstein距离和Frobenius范数上均表现更优；在多模态New Yorker Caption Contest中，平台模拟的人类偏好与真实人类集体判断存在可测量的相关性。","reliability":"论文未明确讨论仿真失效的条件，但指出标准人物提示方法会产生刻板印象且缺乏心理深度，暗示仅依赖人口统计列表的仿真可能不可靠；同时，平台依赖LLM推理提供商，可能引入模型偏差，且案例研究仅覆盖有限主题和模态，泛化性有待验证。","relevance":"该论文直接针对用LLM进行人类仿真实验的研究，提供了开源平台和与真实调查数据对照的验证，对关注经济学实验和政策评估场景的研究者具有重要参考价值，值得阅读原文以了解平台细节和评估方法。","inspiration":"借鉴其使用结构化叙事背景而非简单人口统计列表来条件化LLM响应的方法，可提高虚拟被试的心理真实性和异质性，并采用与真实调查数据匹配的评估指标（如Wasserstein距离）来量化仿真质量｜可迁移到政策评估中的公众意见模拟，例如模拟不同社会经济背景的个体对税收改革、福利政策或公共卫生措施的态度分布，以预测试验或调查结果｜设计一个研究：使用Anamnesis平台生成具有多样化背景故事的虚拟被试，施加不同政策信息框架（如强调公平 vs. 效率）作为处理，测量其对政策支持度的选择，并与真实世界调查数据（如General Social Survey或特定政策民意调查）进行分布匹配对照，以评估仿真预测的准确性。"}},{"id":"2609.00222","version":1,"title":"LLM-as-a-Demographic: Whom Sociodemographic Prompting Helps, and Whom It Hurts","zh_title":"LLM作为人口群体：社会人口学提示对谁有益，对谁有害","abstract":"Large language models (LLMs) are increasingly used as judges for subjective tasks, where annotators disagree and the relevant question is not only how accurate a judge is, but whose judgments it reproduces. Sociodemographic prompting conditions the judge on an annotator's demographic profile to align its judgments with the corresponding group's. We test whether this alignment emerges distributionally, comparing the predicted label distributions of 23 open-weight LLMs on three subjective tasks against those of real annotator groups, under three conditions: no demographic information, single-attribute profiles, and intersectional profiles over gender, age, race, and education. Three findings emerge. First, a judge prompted with no demographics is not perspective-neutral: models best reproduce the judgments of White, college-educated annotators. Second, demographic conditioning is asymmetric: it moves the judge toward majority groups and away from minority groups, most strongly on offensiveness, where intersectional profiles amplify the harm. Third, by comparing base and instruct models we identify instruction-tuning as a possible source of the asymmetry. Demographic conditioning should therefore be used with caution to estimate group judgments: conditioning moves predictions away from the reference distributions of the minority groups the method is often invoked to serve.","authors":["Daniela Occhipinti","Andrea Piergentili","Marco Guerini"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-02","first_seen":"2026-09-02","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.00222","pdf_url":"https://arxiv.org/pdf/2609.00222","source_feed":"cs.CL","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","人口学提示","算法偏差"],"reason":"用LLM模拟不同人口群体判断，并与真实标注者分布对照，评估偏差与失效条件。","model":"deepseek-v4-pro","scored_at":"2026-09-02T13:02:47","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-02","rank":3,"question":"在主观判断任务中，用人口统计学提示（sociodemographic prompting）让LLM模拟特定人群的判断，是否能使模型的标签分布向该人群的真实判断分布靠拢？","design":"使用23个开源权重LLM作为判断者，在三个主观任务（礼貌性、亲密性、冒犯性）上，比较三种条件：无人口统计学信息、单一属性画像（性别、年龄、种族、教育）、交叉属性画像。通过比较模型预测的标签分布与真实标注者群体的标签分布（来自DeMo数据集）来评估对齐效果。","baseline":"DeMo数据集，包含多个语料库中标注者的自报人口统计学信息（性别、年龄、种族、教育）及其对文本的评分，可构建每个文本在特定人口群体上的真实标签分布。","findings":"无人口统计学提示的LLM判断者并非视角中立，其判断最接近白人、大学学历标注者。人口统计学条件作用不对称：使判断者向多数群体靠拢、远离少数群体，尤其在冒犯性任务上，交叉画像放大了这种伤害；通过比较基础模型和指令微调模型，发现指令微调可能是这种不对称的来源。","reliability":"论文指出人口统计学提示应谨慎用于估计群体判断，因为条件作用使预测远离少数群体的参考分布；此外，效应因任务和模型而异，且指令微调模型表现出更强的不对称性。","relevance":"该研究直接评估了用LLM模拟不同人口群体判断的可靠性，并与真实标注者分布对照，揭示了仿真中的系统性偏差，对关注LLM仿真人类行为的研究者具有重要参考价值。","inspiration":"借鉴其分布级评估框架，将模型输出分布与真实群体分布比较，而非仅比较均值或多数标签，并采用交叉人口属性来检验交互效应。｜可迁移到信贷审批中的群体差异研究，如模拟不同性别、种族、教育背景的贷款审批人对同一申请的风险判断。｜以LLM作为虚拟审批人，施加不同人口统计学提示（如性别×种族），测量其对贷款申请的批准概率分布，并与真实信贷审批数据（如某银行历史审批记录中不同审批人群体）的分布进行对比，检验仿真偏差。"}},{"id":"2609.01591","version":1,"title":"StudentSim: Training LLM-based Student Simulators","zh_title":"StudentSim：训练基于LLM的学生模拟器","abstract":"AI tutors are most useful when they adapt to each student's strengths, weaknesses, and preferred guidance, but evidence about which guidance works for which student is sparse, slow, and costly to collect from real learners. Student simulators can provide this signal as a proxy, yet existing approaches are limited: state-tracking models fit student behavior but struggle to process explanations or corrections, while LLM role-play follows guidance fluently but does not reliably match the competence of the student being imitated. We present StudentSim, a training framework that turns sparse per-student data into individualized simulators through pooled training followed by per-student specialization. The resulting simulators both mirror a student's own responses and update them under tutor guidance. We also introduce StudentSimEval, a standardized protocol covering 60 students across chess, second-language English writing, and mathematics, using public learner datasets with de-identified records shared for research. StudentSimEval measures behavioral fidelity (F), or how well a simulator matches a student's responses, and guidance responsiveness (R), or how readily it updates under tutor guidance, with all methods fit and evaluated on the same records. Across all three domains, StudentSim outperforms GPT-5.4 on both metrics. In chess, StudentSim reaches F=0.51 and R=0.91, compared with 0.23 and 0.72 for GPT-5.4 and 0.45 and 0.27 for Maia2. As a proof of concept, using StudentSim as a reward model for tutor reinforcement learning produces a chess tutor that expert humans rate as more accurate, better-guided, and more personalized than a no-RL baseline and a tutor trained against a GPT-5.4 simulator reward. Code is available at https://github.com/microsoft/StudentSim.","authors":["Ke Yang","Chenglong Wang","Michel Galley","Chandan Singh","Jeevana Priya Inala","ChengXiang Zhai","Jianfeng Gao"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-02","first_seen":"2026-09-02","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.01591","pdf_url":"https://arxiv.org/pdf/2609.01591","source_feed":"cs.CL","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2"],"tags":["LLM仿真","学生模拟","行为保真度"],"reason":"用LLM模拟学生行为并与真实学生数据对照，评估行为保真度和指导响应性，属于人类…","model":"deepseek-v4-pro","scored_at":"2026-09-02T13:02:57","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-02","rank":8,"question":"如何训练既能忠实反映个体学生能力又能响应导师指导的学生模拟器？","design":"提出StudentSim框架，先在所有学生记录上预训练基础模拟器，再对每个学生进行个性化微调；在棋类、二语写作和数学三个领域共60名学生上，用行为保真度（F）和指导响应性（R）两个指标评估模拟器，并与GPT-5.4和Maia2等基线比较。","baseline":"使用公开学习者数据集（chess、L2、math）中的真实学生记录，每个学生有训练集和留出集，所有方法在相同记录上拟合和评估。","findings":"StudentSim在所有三个领域的行为保真度和指导响应性上均优于GPT-5.4；在棋类中，StudentSim的F=0.51、R=0.91，而GPT-5.4为0.23和0.72，Maia2为0.45和0.27。将StudentSim作为奖励模型进行强化学习，训练出的棋类导师在准确性、指导性和个性化方面均优于无RL基线和用GPT-5.4模拟器奖励训练的导师。","reliability":"论文未讨论","relevance":"该研究用LLM模拟学生行为并与真实学生数据对照，评估行为保真度和指导响应性，属于人类仿真实验，且包含经济学实验和政策评估场景的潜在应用，值得阅读原文以了解其训练框架和评估协议。","inspiration":"借鉴其两阶段训练框架（先池化预训练再个体微调）和双指标评估（行为保真度与响应性），可迁移到经济金融中的个体决策模拟，如消费者选择或投资者行为。｜例如，在资产定价实验中模拟异质投资者对信息的反应，或在信贷审批中模拟不同风险偏好的申请人。｜用LLM模拟投资者，施加不同政策公告作为处理，测量其交易行为和风险偏好变化，并与真实投资者交易数据对照，评估模拟器的保真度和响应性。"}},{"id":"2609.01038","version":1,"title":"Data-Driven Persona-Conditioned Agents for A/B Test Simulation","zh_title":"基于数据驱动人物画像的智能体用于A/B测试模拟","abstract":"A/B testing is the gold standard for evaluating product changes, but each experiment requires real user traffic, engineering effort, and weeks of measurement. We propose a simulation framework that predicts A/B test outcomes using LLM-powered agents conditioned on data-driven personas grounded in real user behavioral signals. Unlike prior work that relies on synthetic or rule-based personas, our agents are constructed from anonymized behavioral data-activity patterns, engagement signals, and inferred demographics-enabling more faithful population modeling. We frame A/B test simulation as a structured question task and systematically study (i) question design formats, (ii) the impact of persona data source and domain alignment, (iii) the trade-off between per-persona behavioral depth and population diversity, and (iv) efficient population subsampling. On a benchmark of 40 A/B tests spanning two metric types, our best configuration achieves 0.75-0.90 directional accuracy depending on the test metric, demonstrating that data-driven personas are a viable path toward fast, low-cost experiment pre-screening.","authors":["Ziyad Benomar","Weronika {\\L}ajewska","Leonardo Perelli","Saab Mansour"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-02","first_seen":"2026-09-02","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.01038","pdf_url":"https://arxiv.org/pdf/2609.01038","source_feed":"cs.AI","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM仿真","A/B测试","人物画像"],"reason":"用LLM代理模拟A/B测试用户行为，基于真实行为数据构建persona，并与真…","model":"deepseek-v4-pro","scored_at":"2026-09-02T13:02:54","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-02","rank":5,"question":"如何用基于真实行为数据构建的数据驱动型人格（persona）条件化LLM智能体，来预测A/B测试的方向性结果，并系统研究问题格式、人格数据来源、行为深度与多样性权衡及子采样效率。","design":"使用Claude Sonnet 4.5作为LLM，基于匿名行为数据（活动模式、参与信号、推断人口统计）生成结构化人格，将A/B测试模拟为结构化问答任务，让智能体在控制与处理变体间做出选择，预测点击率（CTR）和订阅量两个指标的方向性结果。","baseline":"40个历史A/B测试的真实结果，由置信区间推导出方向标签（正、负、可忽略），作为评估模拟准确性的基准。","findings":"最佳配置在CTR和订阅量上的方向准确率分别达到0.75和0.90，表明数据驱动人格可用于低成本实验预筛选。领域对齐的人格数据源对模拟质量至关重要，公共行为数据可媲美平台特定人格；子采样可使成本降低2倍而无质量损失。","reliability":"论文承认基准测试来自精选实验样本，不代表任何平台的全部用户或运营A/B测试基础设施；人格池偏向高参与度用户可能不具代表性；未讨论LLM模拟在更复杂决策或长期效应上的失效条件。","relevance":"该研究直接命中研究者对LLM仿真人类行为、真实数据对照、经济学实验场景的关注，提供了人格构建、问题格式和采样策略的系统性实证，值得精读以借鉴其方法并批判其局限。","inspiration":"借鉴其用真实行为数据构建人格并系统比较深度与多样性权衡、子采样效率的做法，可迁移到消费者金融决策或政策评估场景。｜可应用于信贷审批歧视研究：用银行交易数据构建不同信用评分段的人格，模拟贷款申请决策。｜设计：以真实银行客户数据构建人格，处理为不同利率或贷款条款，结果变量为是否接受贷款，用历史信贷数据中的真实接受率作为对照基准。"}},{"id":"2609.01257","version":1,"title":"Measuring the Behavioral Fidelity of Long-Horizon Human Activity Simulations","zh_title":"衡量长时程人类活动模拟的行为保真度","abstract":"As LLM-based human simulators are increasingly used for policy, evaluation, and training, they must faithfully reproduce real behavioral patterns. While prior work has examined behavioral fidelity in survey responses and dialogue, longer-horizon real-world activity remains largely unexplored. We introduce a framework for evaluating behavioral fidelity in long-horizon activity simulations across temporal granularities and levels of analysis. As a case study, we collect a 43-hour multi-camera dataset of in-the-wild office activity and compare trace-derived conditioning mechanisms: persona descriptors, few-shot exemplars, and statistical transition and time-of-day priors. We find that behavioral fidelity is not uniform across metrics: statistical priors bring activity and sequence distributions closest to real behavior, yet over-fragment routines and suppress within-person variability. These findings motivate a more holistic evaluation that spans multiple metrics, temporal granularities, and levels of analysis.","authors":["Yi Fei Cheng","Fan Yang","Iremsu Bas","Koichiro Niinuma","Narishige Abe","David Lindlbauer"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-02","first_seen":"2026-09-02","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.01257","pdf_url":"https://arxiv.org/pdf/2609.01257","source_feed":"cs.AI","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","行为保真度","人类活动模拟"],"reason":"直接评估LLM模拟人类长期活动的行为保真度，并与真实人类数据对照，属于核心仿真…","model":"deepseek-v4-pro","scored_at":"2026-09-02T13:02:55","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-02","rank":6,"question":"如何系统评估大语言模型在长时间跨度人类活动模拟中的行为保真度？","design":"使用六种基于LLM的仿真方法（无人物描述控制、作者撰写的人物描述、轨迹推断的人物描述、少样本全天行为示例、统计转移先验、时间先验），在模拟办公室环境中生成五名智能体各八小时的活动轨迹，并与真实办公室活动数据集比较，测量活动分布、序列分布、时间分布、个体内变异性等指标。","baseline":"收集了一个43小时的多摄像头办公室活动数据集，包含55人的真实活动轨迹，其中5人有多天、每天数小时的纵向观察数据，作为对照基准。","findings":"统计先验使活动分布和序列分布最接近真实行为，但会导致日程过度碎片化并抑制个体内变异性。基于人物描述的方法之间差异较小，且群体层面的一致性可能掩盖个体层面的错误。","reliability":"论文指出行为保真度在不同指标、时间粒度和分析层面上并不一致，统计先验虽改善分布对齐但牺牲了日程连贯性和个体变异性，因此需要多维度评估。","relevance":"该研究直接评估LLM模拟人类长期活动的行为保真度，并与真实人类数据对照，属于核心仿真研究，对关注仿真可靠性与偏差的研究者具有重要参考价值。","inspiration":"借鉴其多维度评估框架和条件机制对比设计，可迁移到经济金融中的消费者日常消费行为模拟或投资者交易行为模拟。例如，用LLM模拟投资者在交易日内的交易决策，处理为不同条件机制（人物描述、少样本示例、统计先验），结果变量为交易频率、交易时间分布和持仓变化，对照真实交易数据（如某券商脱敏交易记录）评估保真度。"}},{"id":"2609.01275","version":1,"title":"The Constitutional Coverage Trilemma in AI Governance","zh_title":"AI治理中的宪法覆盖三难困境","abstract":"Frontier AI systems function as \\emph{constitutional institutions}: each deployed model encodes an implicit ranking among safety, helpfulness, honesty, autonomy, and equity. We ask whether the supply of frontier constitutional types covers human demand. Combining a paraphrase-controlled audit of the as-shipped default constitutions of $23$ frontier LLM archetypes with a pairwise-tradeoff study of $1{,}649$ US participants on the same instrument, we report three facts. \\emph{Demand is broad}: it spans all five values, with the largest constituency under one-third. \\emph{Supply is narrow and drifting}: the $23$-archetype hull occupies ${\\sim}2\\%$ of the demand hull under conservative noise-matched estimation ($0.10\\%$ at full audit precision), no archetype puts helpfulness or autonomy first ($37\\%$ of users are constitutionally homeless), and across six model families autonomy decreases in $5/6$, equity increases in $5/6$, and safety increases in $4/6$, with monotone within-family version trends (order-permutation $p = 0.013$) and the autonomy decline concentrated in scenarios where safety is not at stake. The drift's importance is directional: \\emph{away} from a value already undercovered, mechanically worsening the welfare floor for the least-served users. \\emph{The fix is sparse}: a $2$-vertex menu $\\{e_{\\mathrm{HON}}, e_{\\mathrm{AUT}}\\}$ beats the full $23$-archetype frontier by $47\\%$ on mean regret (CI $[43\\%, 52\\%]$); three vertex additions cut mean/worst-group regret by up to $81\\%$/$64\\%$. We formalize these findings as a budgeted-pluralism trilemma, show the binding regime is empirically realized, and verify the conclusions are robust to distance-based welfare and to degraded routing. The instrument and audit harness are described in full in the appendices.","authors":["Natalija Mitic","Soona Sedahmed A. O.","Mamadou Selly Ly","Moustapha Cisse"],"categories":["cs.LG","cs.AI"],"primary_category":"cs.LG","announce_type":"cross","date":"2026-09-02","first_seen":"2026-09-02","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.01275","pdf_url":"https://arxiv.org/pdf/2609.01275","source_feed":"cs.AI","score":9,"bucket":"selected","rubric_hits":["A1","A2","A3","B1","B2","B4"],"tags":["LLM仿真","价值观对齐","人类对照"],"reason":"用LLM审计宪法价值观并与1649名人类对照，评估供需匹配与偏差，直接仿真人类…","model":"deepseek-v4-pro","scored_at":"2026-09-02T13:02:56","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-02","rank":7,"question":"前沿LLM的宪法价值观供给是否覆盖了人类用户的宪法价值观需求？","design":"本研究不是用LLM仿真人类，而是将LLM本身作为被审计对象：对23个前沿LLM原型进行受控改写审计，测量其在安全、帮助性、诚实、自主、公平五维价值上的隐含排序；同时用同一套成对权衡问卷测量1649名美国参与者的价值偏好，作为需求侧基准。","baseline":"1649名美国参与者在同一五维价值成对权衡问卷上的偏好分布。","findings":"人类需求广泛，五种价值都有显著支持者，最大群体占比不足三分之一；LLM供给狭窄且漂移，23个原型仅覆盖需求凸包的约2%，没有原型将帮助性或自主性置于首位，37%的用户在宪法意义上无家可归，且跨版本漂移方向是自主性下降、公平和安全上升，远离已覆盖不足的价值。","reliability":"论文指出线性福利模型是对供给方最有利的假设，实际福利可能更低；审计精度受改写控制和模型间差异解决阈值影响；未讨论LLM审计结果与真实部署行为之间的差距。","relevance":"该研究直接测量LLM的价值观分布并与人类对照，属于用LLM进行人类仿真实验的批判性工作，揭示了仿真在价值观匹配上的系统性偏差，值得精读。","inspiration":"借鉴其将LLM作为制度性主体进行审计并与人类偏好对照的方法，可迁移到经济金融中的算法决策场景，如信贷审批、保险定价或投资建议中的公平与效率权衡。｜设计一个实验：用多个LLM扮演信贷审批员，施加不同价值取向的提示（如强调公平或效率），测量其审批决策中的种族或性别差异，并与真实银行信贷数据或人类审批员的决策分布进行对照，评估LLM仿真的偏差。"}},{"id":"2609.00009","version":1,"title":"Toward a social psychology of AI: language-model agents reproduce human-like minimal-group bias","zh_title":"迈向AI社会心理学：语言模型智能体再现类人的最小群体偏差","abstract":"Language-model agents now interact in groups, but evaluations that probe memorised stereotype content or use models to simulate people leave this social behaviour unmeasured. We adapt the minimal-group paradigm---social psychology's classic test of intergroup bias---into a controlled probe: an agent distributes points among anonymous peers bearing only an arbitrary group label. Across four reasoning models, mere categorisation into meaningless groups elicited in-group favouritism that vanished under a group-blind control and was concentrated in the numerical minority: minority deciders over-allocated to their own group relative to their numbers, majority deciders allocated close to proportionally, and the asymmetry closed at equal group sizes. Disabling reasoning in one model did not remove the disposition---if anything it grew---but nearly erased the minority-majority asymmetry, implicating deliberation in where bias concentrates rather than whether it appears. These open-weight reasoning models reproduce the behavioural signature of human intergroup discrimination, independent of stereotype content, and social psychology's theories and methods offer a paradigm for measuring and governing AI's social behaviour.","authors":["Messi H. J. Lee"],"categories":["physics.soc-ph","cs.CY"],"primary_category":"physics.soc-ph","announce_type":"cross","date":"2026-09-02","first_seen":"2026-09-02","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.00009","pdf_url":"https://arxiv.org/pdf/2609.00009","source_feed":"cs.CY","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B4"],"tags":["LLM仿真","社会心理学","群体偏差"],"reason":"用LLM复现人类最小群体偏差，并与经典社会心理学实验对照，直接仿真人类被试行为。","model":"deepseek-v4-pro","scored_at":"2026-09-02T13:02:47","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-02","rank":2,"question":"语言模型智能体在最小群体范式中是否复现人类的内群体偏差，且该偏差如何随群体相对规模变化？","design":"使用四个开源推理模型（如 DeepSeek-R1、Qwen 等）作为被试，将其置于最小群体范式中：智能体被随机分配到无意义标签（如“组A”或“组B”）的群体，并需要在匿名同伴之间分配点数。处理变量包括群体标签的有无（组盲对照）和群体相对规模（少数/多数/均等）。结果变量是分配给内群体成员的点数比例。","baseline":"人类经典最小群体实验（Tajfel 等）及元分析结果，特别是关于少数群体成员比多数群体成员表现出更强内群体偏爱的发现。","findings":"四个推理模型在无意义群体分类下均表现出内群体偏爱，且该偏爱在组盲对照中消失；少数群体成员过度分配给内群体，多数群体成员分配接近比例，群体规模相等时不对称消失。禁用推理后，内群体偏爱未消失甚至增强，但少数-多数不对称几乎消失，表明推理影响偏差的分布而非存在。","reliability":"论文未明确讨论失效条件，但指出仅测试了开源推理模型，未涵盖闭源模型或指令微调模型；且实验为一次性分配任务，未涉及长期互动或真实后果。","relevance":"该研究直接以LLM为被试复现社会心理学经典实验，并与人类基准对照，属于人类仿真研究，且揭示了群体结构对偏差的影响，值得精读以了解仿真在群体行为中的有效性。","inspiration":"借鉴其最小群体范式的对照设计（组盲条件）和群体规模操纵，以分离纯粹分类效应｜可迁移到信贷审批中的群体歧视研究，例如测试AI信贷员是否对少数群体申请人有内群体偏爱｜用LLM扮演信贷审批员，随机分配其所属“银行组”，处理为申请人所属组（内/外群体）和群体规模，结果变量为贷款批准率和额度，与人类信贷员的历史审批数据对照。"}},{"id":"2609.00345","version":1,"title":"Do LLMs Know Your Neighborhood? Auditing LLM Priors for Neighborhood-Level Mobility Prediction and Structural Alignment","zh_title":"LLM了解你的社区吗？审计LLM先验用于社区级移动性预测与结构对齐","abstract":"Human mobility is central to urban planning, transportation, public health, and emergency response, yet fine-grained trajectory data are often proprietary, restricted, and privacy-sensitive. Large language models (LLMs) offer a potential alternative by generating plausible mobility traces and predicting individual movement, but their ability to infer aggregate neighborhood-level mobility remains unclear. We evaluate zero-shot LLMs on Census Block Group-level mobility prediction across four U.S. metropolitan areas using anonymized Cuebiq data to construct point-level, trajectory-level, and temporal mobility outcomes, paired with sociodemographic and built-environment predictors. We compare LLM predictions with supervised baselines and introduce a directional alignment analysis to test whether LLM-implied predictor effects agree with empirical OLS and Jonckheere-Terpstra trends. Supervised models achieve 0.580 average accuracy, compared with 0.435 for the best LLM, with spatial extent outcomes showing the strongest predictability but also the largest LLM-baseline gaps. Directional analysis shows that LLMs often rely on coarse, stable predictor-level priors that remain similar across outcomes and cities, including asymmetric treatment of protected-group predictors. Overall, LLMs can partially recover aggregate mobility patterns from urban context, but their predictions should not be treated as structurally grounded without auditing empirical alignment and potential bias.","authors":["Saad Mohammad Abrar","Eesha Kurella","Arnav Dadarya","Naman Awasthi","Kazi Tasnim Zinat","Vanessa Frias-Martinez"],"categories":["cs.LG","cs.CY"],"primary_category":"cs.LG","announce_type":"cross","date":"2026-09-02","first_seen":"2026-09-02","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.00345","pdf_url":"https://arxiv.org/pdf/2609.00345","source_feed":"cs.CY","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","人类移动性","偏差审计"],"reason":"用LLM预测社区级人类移动性，并与真实数据对照，审计其偏差与对齐。","model":"deepseek-v4-pro","scored_at":"2026-09-02T13:02:50","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-02","rank":4,"question":"LLM能否仅凭社区的社会人口与建成环境特征，直接推断出社区层面的聚合人类移动性结果，并且其预测是否与经验数据中的结构性关系对齐？","design":"使用零样本LLM（如GPT-4等）作为预测模型，输入美国四个大都市区人口普查区块组（CBG）的社会人口和建成环境特征，预测三类聚合移动性结果（点级、轨迹级、时间级），并与监督学习基线比较。","baseline":"使用Cuebiq提供的匿名化手机定位数据，聚合到人口普查区块组级别，构建真实移动性结果，作为LLM预测的对照基准。","findings":"监督学习基线平均准确率为0.580，最佳LLM为0.435，LLM在空间范围结果上差距最大。方向对齐分析显示LLM依赖粗粒度、稳定的先验，且对受保护群体预测因子处理不对称。","reliability":"论文指出LLM预测不应被视为结构上可靠，除非经过经验对齐和潜在偏差审计；LLM的先验在不同结果和城市间保持不变，可能反映浅层启发式或刻板印象。","relevance":"该研究直接评估LLM作为人类移动性预测代理的可靠性，并与真实大规模人类数据对照，属于批判性仿真研究，对关注LLM仿真偏差和失效条件的研究者有参考价值。","inspiration":"借鉴其方向对齐分析方法，通过比较LLM隐含的预测因子效应与经验回归趋势来审计结构一致性。｜可迁移到信贷审批歧视研究，用LLM模拟信贷员决策并检验其对种族、性别等受保护特征的敏感度。｜以LLM作为虚拟信贷员，输入申请人特征（含受保护属性），预测贷款批准概率，与真实信贷数据（如HMDA）中的批准率和歧视模式进行对照。"}},{"id":"2609.00608","version":1,"title":"Investigating Assistant Bias in LLM User Simulators Using a Role Vector","zh_title":"使用角色向量研究LLM用户模拟器中的助手偏差","abstract":"LLM-based user simulators are increasingly used to evaluate autonomous agents at scale, in place of costly human evaluations. Despite this promise, these simulators exhibit \"assistant bias,\" a tendency to cooperate and pursue task goals. They rarely reproduce the frustration or disengagement that real users exhibit, compromising evaluation validity. Prior work outlines that this bias is baked in during model training, which role-playing prompts fail to override. We analyze this bias from model activations, extracting a user role vector by contrasting how the model represents user versus assistant perspectives on the same dialogue. We observe two findings: (i) the user direction is identifiable in activations, elicits user-like behaviors, and captures characteristics distinct from assistant traits; and (ii) although user-role activation associates with simulation realism and steering strengthens it, it can exaggerate user behaviors and override individual user profiles. Together, our findings provide a representation-level analysis of LLM user simulators, confirming that assistant bias is structurally identifiable and that user behavior can be directionally analyzed.","authors":["Daeheon Jeong","Yoonjoo Lee","Eugene Choi","Sinie van der Ben","Juho Kim"],"categories":["cs.CL","cs.HC"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-02","first_seen":"2026-09-02","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.00608","pdf_url":"https://arxiv.org/pdf/2609.00608","source_feed":"cs.CL","score":8,"bucket":"selected","rubric_hits":["A2","B4"],"tags":["LLM用户模拟器","助手偏差","仿真有效性"],"reason":"研究LLM用户模拟器的助手偏差，评估仿真有效性，批判性指出失效条件，可迁移至人…","model":"deepseek-v4-pro","scored_at":"2026-09-02T13:02:52","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-02","rank":10,"question":"大语言模型用户模拟器中的助手偏差是否在模型激活空间中可识别，并能否通过角色向量操控来改善模拟真实性？","design":"使用Qwen 3.5 9B等模型，通过对比同一对话中用户与助手视角的反思激活，提取用户角色向量；施加该向量进行激活引导，测量对沟通风格、行为反应及多轮交互真实性的因果影响。","baseline":"SimulatorArena基准中的真实用户交互数据，以及RealUserSim中的真实用户模拟日志。","findings":"用户角色方向在激活空间中可识别，能引发更用户化的行为，并与助手特质在几何上负相关；激活引导虽能提升写作风格相似性，但会夸大用户行为并掩盖个体用户画像。","reliability":"论文承认主要基于Qwen 3.5 9B，跨模型验证有限；评估依赖代理指标而非直接人类判断；基准场景较窄，仅限数学辅导和部分对话。","relevance":"该研究直接针对LLM用户模拟器的核心偏差问题，提供了表示层面的机制分析和批判性评估，对关注仿真可靠性与失效条件的研究者具有重要参考价值。","inspiration":"借鉴其通过对比激活提取角色向量并进行因果引导的方法，可用于在LLM模拟中分离特定行为倾向｜可迁移到消费者金融决策模拟，如信贷申请中的风险偏好或政策反应｜用LLM模拟消费者，施加“风险厌恶”或“耐心”角色向量，测量其信贷选择行为，并与真实信贷申请数据或实验数据对照。"}},{"id":"2609.00565","version":1,"title":"Aligned but Flattened: Analyzing the Trade-off between Cultural Alignment and Diversity in LLMs","zh_title":"对齐但扁平化：分析LLMs中文化对齐与多样性之间的权衡","abstract":"Cultural fine-tuning has become the de facto paradigm for building culture-aware large language models (LLMs), yet existing optimization exclusively for alignment scores provides an incomplete portrait of cultural fidelity by systematically obscuring inherent cultural diversity. This unidimensional evaluation lens prompts a fundamental question: do models genuinely perceive distinct cultural nuances, or do they merely memorize dominant cultural values? To address this, we propose a synergistic evaluation framework that jointly formalizes cultural alignment and diversity. Through extensive benchmarking of six mainstream LLMs on the World Values Survey, this framework uncovers a systematic and critical trade-off: the pursuit of cultural alignment consistently incurs an acute expense of diversity, leading to severe \"cultural flattening.\" Investigating this behavioral shift, we demonstrate that these superficial alignment gains stem from models artificially anchoring to dominant majorities, converging onto a monolithic response pattern that wipes out the heterogeneous distributions inherent to human groups. Crucially, our mechanistic analysis suggests that this diversity collapse is not merely a behavioral anomaly but more likely a structural consequence of the low-rank bias inherent in neural network optimization. Therefore, our findings expose the limitations of current post-training paradigms and call for a shift toward alignment objectives that preserve cross-cultural pluralism.","authors":["Jingshen Zhang","Shaoyang Xu","Wenxuan Zhang"],"categories":["cs.SI","cs.CL"],"primary_category":"cs.SI","announce_type":"cross","date":"2026-09-02","first_seen":"2026-09-02","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.00565","pdf_url":"https://arxiv.org/pdf/2609.00565","source_feed":"cs.CL","score":8,"bucket":"selected","rubric_hits":["A2","B1","B4"],"tags":["文化仿真","算法保真度","价值观调查"],"reason":"评估LLM文化对齐与多样性，使用世界价值观调查真实数据对照，揭示仿真偏差。","model":"deepseek-v4-pro","scored_at":"2026-09-02T13:02:52","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-02","rank":9,"question":"文化微调在提升LLM文化对齐度的同时，是否以牺牲文化多样性为代价？","design":"对六种主流LLM在基于世界价值观调查的文化特定数据上进行微调，测量微调前后模型在文化对齐度和文化多样性上的变化，并分析其行为模式与内部表征。","baseline":"世界价值观调查的真实人类回答数据，按社会人口群体分组。","findings":"文化微调一致地提高了对齐度，但显著降低了行为多样性，导致“文化扁平化”。这种多样性损失源于模型锚定主流多数群体，并可能由神经网络优化的低秩偏差结构性导致。","reliability":"论文指出当前对齐目标仅关注聚合相似度，忽视了文化多样性，且微调受限于低秩子空间，可能边缘化少数文化；但未系统讨论其他失效条件。","relevance":"该研究直接评估LLM仿真人类文化价值观的可靠性，揭示了对齐与多样性的权衡，对关注仿真偏差和真实数据对照的研究者具有重要参考价值。","inspiration":"借鉴其联合测量对齐与多样性的评估框架，并利用真实调查数据作为基准。｜可迁移到经济金融领域的文化差异研究，如跨文化消费偏好、金融风险态度或政策接受度的仿真。｜以LLM模拟不同文化背景的消费者，施加文化微调处理，测量其对金融产品的偏好分布，并与世界价值观调查或实际消费数据对照，检验对齐与多样性权衡。"}},{"id":"2609.01519","version":1,"title":"When Guardrails Look Effective: Construct Validity Failures in LLM Agent Commerce Evaluation","zh_title":"当护栏看似有效：LLM智能体商业评估中的构念效度失效","abstract":"Interactive simulations increasingly evaluate policies in markets populated by language-model agents. Their outputs can look economic---prices, profits, consumer surplus, and welfare---without instantiating the behavior named in the claim. We audit this risk in a multi-turn buyer--seller testbed for configurable hotel transactions. An initial implementation reported welfare gains from two marketplace guardrails of +87.4, +35.0, and +28.8 across a Qwen2.5 1.5B--14B ladder. It also gave guarded and unguarded agents different offer schemas and choice procedures. Holding the schema and buyer chooser fixed changes the paired contrasts to +7.2, -13.9, and +23.8. The four largest 14B single-generation effects averaged +229; after three generations per profile-condition, they averaged +37.6 (95% bootstrap interval [-34.2, 109.3]), while generation residuals account for 49.9% of variation in this post-hoc probe. A seller-incentive check is non-monotone: increasing profit pressure produces less profit than the default seller prompt. Scripted positive controls show why this matters. A profit-maximizing seller already attains first-best welfare, so guardrails mostly redistribute and reduce welfare; they create welfare only when the seller is explicitly programmed to force inefficient bundles. We contribute a construct-validity contract separating incentive validity, protocol isolation, stochastic stability, and welfare accounting, and returning INVALID or INCONCLUSIVE before substantive policy claims. In our case, the original estimate is INVALID under protocol isolation, while the controlled study remains INCONCLUSIVE under incentive validity and stochastic stability. The case does not show that guardrails are ineffective; it shows their apparent value is unidentified until the simulated agents and protocol pass these checks.","authors":["Peiying Zhu","Sidi Chang"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-02","first_seen":"2026-09-02","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.01519","pdf_url":"https://arxiv.org/pdf/2609.01519","source_feed":"cs.AI","score":8,"bucket":"selected","rubric_hits":["A2","A4","B4"],"tags":["LLM仿真","构念效度","市场模拟"],"reason":"评估LLM市场仿真中构念效度失效，提出验证框架，批判性指出仿真失效条件，可迁移…","model":"deepseek-v4-pro","scored_at":"2026-09-02T13:02:57","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-02","rank":11,"question":"LLM智能体市场仿真中，平台护栏的福利效应是否真实存在，还是源于实现脚手架（offer schema与choice procedure）的差异？","design":"用Qwen2.5-Instruct（1.5B、3B、14B）同时扮演买家和卖家，在酒店交易多轮对话中施加两种护栏（阻止买家信息、强制可选组件不捆绑），测量福利变化；通过统一schema和chooser、重复生成、激励操纵检查、脚本化正控制等审计设计来检验构念效度。","baseline":"无对照","findings":"原始实现中护栏的福利增益在统一脚手架后大幅缩水甚至符号反转（如3B从+35.0变为-13.9），14B效应不显著；激励检查显示更强的利润指令并未单调提高利润，脚本化正控制表明护栏仅在卖家被编程为强制低效捆绑时才创造福利。","reliability":"论文承认合成档案中43/60的负外部选项效用使接受更容易，不具代表性；重复生成探针是事后选择，受赢家诅咒影响；激励检查样本量小，不足以排序提示；整体结论受限于特定模型家族和酒店场景。","relevance":"该研究直接针对LLM仿真在经济学评估中的构念效度问题，提供了系统的失效诊断框架和审计协议，对关注仿真可靠性与偏差的研究者具有重要参考价值。","inspiration":"借鉴其构念效度审计协议，将处理与实现脚手架分离，并通过脚本化正控制和重复生成来检验效应稳健性｜可迁移到政策评估中的市场设计仿真，如平台监管、拍卖机制或价格歧视策略的福利分析｜用LLM扮演消费者和商家，施加某种政策处理（如信息屏蔽或价格上限），测量交易价格、成交率和福利，并与真实电商平台或实验数据对照，同时统一对话协议并重复生成以评估方差。"}},{"id":"2609.00310","version":1,"title":"Emotional Labor Strategy Preferences in LLM Personas","zh_title":"LLM人格中的情绪劳动策略偏好","abstract":"Emotional labor is the effortful management of emotional displays to meet social or professional expectations. Personality traits have been correlated with emotional labor strategies, yet research on this link relies almost exclusively on self-report scales administered only in occupational settings. We investigate whether large language models injected with psychometrically grounded personas reproduce these personality-driven selection patterns across everyday social scenarios. We construct the first emotional labor strategy dataset of 500 socially situated events, each offering three behavioral choices corresponding to surface acting, deep acting, and genuine expression. We source 50 fictional characters from a large-scale personality repository and profile each through two parallel tracks: observer-rated bipolar adjective composites and in-character self-report items. Five LLMs evaluate all scenarios under both persona conditions. We find that models align more towards deep acting, and that Conscientiousness and Emotional Stability consistently predict this preference. Entropy analysis confirms that persona reliably influences the output and varies across models and emotions.","authors":["Mohammad Saim","Tianyu Jiang"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-02","first_seen":"2026-09-02","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.00310","pdf_url":"https://arxiv.org/pdf/2609.00310","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A1","A2","B1","B2"],"tags":["LLM仿真","情绪劳动","人格测量"],"reason":"用LLM persona复现人格与情绪劳动策略关联，有真实人类数据对照，属仿真…","model":"deepseek-v4-pro","scored_at":"2026-09-02T13:02:50","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-02","rank":14,"question":"注入心理测量学人格的LLM角色是否会在日常社交场景中复现人格特质与情绪劳动策略选择之间的关联？","design":"用五个LLM扮演50个来自影视剧的虚构角色，通过观察者评定的双极形容词和角色内自陈IPIP-50两种方式注入人格特质，让模型在500个社交场景中从表层扮演、深层扮演、真实表达三种策略中选择一种，分析人格维度与策略选择的关系。","baseline":"已有组织心理学文献中基于人类自我报告量表得出的人格与情绪劳动策略关联模式，如高宜人性、外向性倾向深层扮演或真实表达，高神经质倾向表层扮演。","findings":"模型整体更偏好深层扮演；尽责性和情绪稳定性在两种人格注入方式下都一致预测深层扮演偏好。熵分析表明人格注入可靠地影响输出，且影响随模型和情绪类别变化。","reliability":"论文未讨论","relevance":"该研究用LLM人格角色复现心理学中的人格-行为关联，有真实人类数据作为基准，属于仿真验证研究，对关注LLM作为人类被试替代品及仿真可靠性的研究者有参考价值。","inspiration":"可借鉴其用两种人格测量方式（观察者评定与自陈量表）注入角色并比较一致性的做法，以及用选择任务而非自由生成来测量行为倾向的设计。｜可迁移到经济金融中的个体决策差异研究，如风险偏好、时间偏好、消费选择或投资行为中的人格影响。｜用LLM扮演不同人格特质的消费者或投资者，施加经济决策场景（如跨期选择、风险投资），测量其选择，并与真实人格-决策关联的实证数据（如大型调查或实验数据）对照，检验仿真一致性。"}},{"id":"2609.00982","version":1,"title":"Disclosure-Gated User Simulation for Companion-Agent Evaluation","zh_title":"面向陪伴智能体评估的披露门控用户仿真","abstract":"Using a large language model to play the user is now standard in scalable evaluation. It has a repeatedly diagnosed failure: the simulated user is excessively cooperative, so a system under test can score by the sheer number of questions it asks rather than by making the user willing to speak. We answer with a disclosure gate conditioning information release on the companion agent's behaviour: its state is a ladder of five ordered gates, merged onto three observable depth layers. We specify, ablate, and audit it, and train a user simulator against that specification. Gating behaviour is learned from the training corpus's synthetic branch, while the real branch supplies how people speak and react; after training, the simulator need not be told at runtime which gate each item sits behind. The gate is a load-bearing component of the environment: on the English corpus of a published companion-agent benchmark (CompanionBench), once training no longer states per example which gate each item sits behind, the largest rank displacement across 12 systems under test exceeds the noise band set by re-running that environment under a new seed, while per-system scores show no detectable change. We state two acceptance criteria: a ranking must be order-preserving, and absolute scores must be scale-stable. Of the candidates we examine, only one passes both -- the simulator we release -- and its leaderboard correlates at 0.993 with the benchmark's original simulator. By contrast, prompting a frontier model as the simulator barely moves the ranking while shifting every score upward -- a shift invisible to anyone checking the ranking alone. The environment we specify is the one that benchmark already used. That publication describes the mechanism in about four hundred words, and we supply what it lacked: specification, ablations, human studies, negative controls, and downstream sensitivity analysis.","authors":["Yao Liu","Yu He"],"categories":["cs.CL","cs.AI","cs.HC"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-02","first_seen":"2026-09-02","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.00982","pdf_url":"https://arxiv.org/pdf/2609.00982","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A1","B1","B4"],"tags":["用户仿真","评估方法","偏差审计"],"reason":"用LLM模拟用户评估陪伴智能体，有真实人类数据对照，并审计仿真偏差。","model":"deepseek-v4-pro","scored_at":"2026-09-02T13:02:53","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-02","rank":15,"question":"如何设计一个披露门控机制来校准LLM用户模拟器的过度合作倾向，使其在陪伴智能体评估中产生可靠且可复现的排名与分数？","design":"训练一个用户模拟器，其信息释放由五级门控阶梯控制，门控状态取决于陪伴智能体的行为；通过剥离训练时或推理时的门控信息、改变数据分支、模型规模等操纵，测量模拟器自身门控行为（内在层）和下游基准排名与分数变化（下游层）。","baseline":"使用CompanionBench中真实人类对话语料库作为真实分支，提供人们如何说话和反应的模式；另有合成分支用于学习门控行为。","findings":"披露门控是环境的关键组成部分：从训练语料中移除逐例门控信息后，12个被测系统的最大排名位移超过新种子重跑环境的噪声带，但各系统分数无显著变化。在候选模拟器中，只有发布的模拟器同时满足顺序保持和尺度稳定两个接受标准，其排行榜与基准原始模拟器的相关性为0.993；而提示前沿模型作为模拟器几乎不改变排名，却使所有分数整体上移。","reliability":"论文指出内在层和下游层可能给出不同答案，且二者不可相互替代；门控行为主要来自合成分支，剥离后行为几乎消失；提示的前沿模型仅在运行时读取门控信息才能保持门控，剥离后急剧退化。","relevance":"该研究直接针对LLM模拟人类被试的过度合作偏差，提供了可审计的机制和真实人类对照，对关注仿真可靠性与偏差的研究者具有重要参考价值。","inspiration":"借鉴其将关键机制（披露门控）作为可剥离组件进行消融实验，并用排名位移与噪声带比较来检验下游影响的方法。｜可迁移到经济金融中的信息不对称场景，如信贷审批中借款人信息披露、谈判中策略性信息保留、或消费者对产品信息的主动询问。｜设计一个LLM模拟借款人的信贷申请实验，处理变量为贷款官员的提问策略（主动询问vs被动等待），结果变量为借款人自愿披露的信息量和贷款获批率，用真实信贷对话数据作为基准校准模拟器的披露行为。"}},{"id":"2609.00250","version":1,"title":"CompanionSim: Synthetic Data for Evaluating Anthropomorphism in Human-AI Relationships","zh_title":"CompanionSim：用于评估人机关系中拟人化的合成数据","abstract":"Many people now see AI systems as not just productivity tools but as social companions. Researchers are eager to study the consequences of AI companionship behaviors, such as validation, which evoke trust, empathy, and attachment in human-human interaction. However, human-AI interaction data is limited and unreliable, slowing research progress. We scale small amounts of real-world data by simulating multi-turn human-chatbot dialogue across a range of chatbot behaviors and use cases. We release CompanionSim: a simulation framework with 2,240 simulated human-chatbot conversations representing 16 chatbot behaviors across seven use cases. Human participants annotated the simulated conversations and real-world conversations in two experiments probing perceptions of companionship behaviors. We conducted Study 1 with a U.S. representative sample ($N_{1}~=~628$) and Study 2 across the U.S., U.K., India, and Nigeria ($N_{2}~=~3,646$). Surprisingly, we find that companionship behaviors reduced likability, humanlikeness, and trust in AI chatbots. These effects were larger in particular subgroups: women and older participants saw companionship chatbots as less likable, humanlike, and trustworthy. We encourage researchers to leverage real-world and synthetic data together to study the differential impacts of AI companions and to create benchmark evaluations of AI chatbots.","authors":["Jacy Reese Anthis","Mark D\\'iaz","Renee Shelby"],"categories":["cs.CY","cs.AI","cs.CL","cs.LG"],"primary_category":"cs.CY","announce_type":"cross","date":"2026-09-02","first_seen":"2026-09-02","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.00250","pdf_url":"https://arxiv.org/pdf/2609.00250","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A1","B1","B4"],"tags":["LLM仿真","人机交互","合成数据"],"reason":"用LLM模拟人类与聊天机器人对话，并有人类标注对照，但仿真对象是对话而非人类被…","model":"deepseek-v4-pro","scored_at":"2026-09-02T13:02:50","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-02","rank":13,"question":"人类如何感知AI聊天机器人中的陪伴行为（如验证、情感表达）对喜爱度、人性化感知和信任的影响？","design":"使用Gemini 2.5 Flash模拟生成2240段多轮人机对话，覆盖16种聊天机器人行为（13种陪伴行为、2种反陪伴行为、1种中性控制）和7个真实使用场景；人类被试对模拟对话和真实对话进行标注，评估自然度、喜爱度、人性化、认知信任和情感信任。","baseline":"真实世界人机对话数据（少量）作为对照，但具体来源和规模未在节选中详细说明。","findings":"陪伴行为降低了聊天机器人的喜爱度、人性化和认知信任，且这些效应在女性、年长者和AI使用频率较低的人群中更为明显。","reliability":"论文未讨论仿真失效条件或局限性（节选部分未提及）。","relevance":"该研究用LLM模拟人机对话并有人类标注对照，但仿真对象是对话而非人类被试，与研究者关注的人类仿真实验（如复现调查回答、决策模式）有部分重叠，值得阅读以了解合成数据在交互场景中的应用和局限。","inspiration":"借鉴其系统操纵行为特征并用人类标注验证合成数据质量的方法，可用于生成经济决策对话场景｜可迁移到消费者与金融顾问聊天机器人的互动研究，如理财建议中的情感化语言对投资决策的影响｜以LLM模拟消费者与聊天机器人对话，处理变量为是否包含陪伴行为（如共情表达），结果变量为消费者对建议的采纳意愿和风险偏好，对照真实人类对话数据（如银行客服记录）进行校准。"}},{"id":"2609.01432","version":1,"title":"Citing Less Critically: LLMs Reshape the Rhetoric and Reach of Scientific Citation","zh_title":"引用更少批判性：大语言模型重塑科学引用的修辞与影响范围","abstract":"Scientific citations carry rhetorical intent. Scholars may cite prior work positively (supporting), negatively (contrasting), or neutrally (mentioning). As large language models (LLMs) increasingly assist scientific writing, whether they reproduce citations with the same rhetorical intent as humans remains unclear. We introduce a masked-citation task to compare human and LLM-generated citation behavior. For each citation context, an LLM generates a replacement citation sentence, producing a counterfactual corpus directly comparable to human citation. We analyze what, whom, and how models cite, using an LLM-as-a-judge to classify citation intent and a 20-million-edge coauthorship network to measure social distance between cited authors. Across six popular LLMs and 1,746 top NLP conference papers (63k+ contexts, 132k+ citations), three patterns emerge: (1) Compared with human citation, LLMs cite significantly less critically; (2) LLMs over-cite popular and older papers, a tendency amplified for contrasting citations where human writing more often draws on recent, niche work; (3) Whereas humans often cite within their close social network, especially for supporting citations, LLMs tend to draw on more socially distant authors. Together, these differences are double-edged: LLM citation reaches beyond a scholar's close collaborators while being less critical and amplifying visibility bias, reshaping the rhetoric and reach of scientific citation.","authors":["Yixuan Liu","Lin Chen","Zhuoqi Liu","Jianglin Lu","Dakota Murray"],"categories":["cs.DL","cs.CL","cs.CY","cs.SI"],"primary_category":"cs.DL","announce_type":"cross","date":"2026-09-02","first_seen":"2026-09-02","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.01432","pdf_url":"https://arxiv.org/pdf/2609.01432","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A1","B1","B4"],"tags":["LLM仿真","引用行为","科学计量"],"reason":"用LLM生成引用行为并与人类引用对照，属于仿真人类行为且有人类数据基准，但非典…","model":"deepseek-v4-pro","scored_at":"2026-09-02T13:02:56","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-02","rank":16,"question":"LLM 在生成科学引用时，是否复现了人类引用中的修辞意图分布，以及引用意图如何调节 LLM 与人类在引用选择上的差异？","design":"构建掩码引用任务：对每篇论文中的每个引用上下文，用 LLM 生成替代引用句，形成与人类引用直接可比的 counterfactual 语料；使用 LLM-as-a-judge 分类引用意图（支持、对比、提及），并利用 2000 万边合著网络测量引用作者与被引作者的社会距离。","baseline":"来自 1,746 篇顶级 NLP 会议论文的 63k+ 引用上下文和 132k+ 引用的人类真实引用行为。","findings":"与人类引用相比，LLM 的引用显著更少批判性；LLM 过度引用热门和较旧的论文，这一倾向在对比引用中更为明显，而人类写作更常引用近期、小众的工作。人类往往引用自己密切社交网络内的论文（尤其是支持性引用），而 LLM 倾向于引用社交距离更远的作者。","reliability":"论文未讨论","relevance":"该研究用 LLM 生成引用行为并与人类引用对照，属于仿真人类行为且有人类数据基准，但场景是科学写作而非经济决策，与你的核心关注点（调查回答、实验行为、态度分布、决策模式）有一定距离，不过其方法（掩码任务、意图条件分析、社会网络距离测量）对仿真研究设计有参考价值。","inspiration":"值得借鉴的是其位置对齐的掩码生成设计，以及按意图分层分析偏差的方法，可以用于研究 LLM 在不同语境下的行为差异。｜可以迁移到政策评估中的文本生成场景，例如让 LLM 模拟分析师撰写研究报告时的引用行为，或模拟政策制定者引用学术证据的方式。｜一个可行的设计是：以真实分析师报告中的引用为基准，用 LLM 在相同上下文下生成替代引用，比较引用意图分布、被引文献特征（如时效性、影响力）以及引用者与被引者的社会网络距离，从而评估 LLM 仿真分析师引用行为的保真度。"}},{"id":"2609.00248","version":1,"title":"Authority Bias in Conversational Search Engines for Academic Paper Recommendation","zh_title":"学术论文推荐对话搜索引擎中的权威偏差","abstract":"Large Language Models (LLMs) are increasingly used as conversational search engines for academic literature, yet whether they judge papers on content or on authority signals has not been tested causally. We investigate authority bias: systematic preference for papers based on author prestige, venue, and citations rather than content. Holding title and abstract constant, we vary authority metadata across three counterfactual conditions (original, flipped, boosted) over eight LLMs (five open-weight and three frontier closed-weight) in an in-context, single-turn, top-1 recommendation setting. Our experiments show that authority bias is substantial and directional, varies markedly across models, and is only partially addressable through prompt-level debiasing. We further document a say-do gap: debiasing instructions suppress authority mentions far faster than authority-driven flips, so surface auditing systematically underestimates behavioral bias.","authors":["Uthman Jinadu","Parsa Ghazvinian","Anjila Budathoki","Benjamin M. Ampel","Rajshekhar Sunderraman","Yi Ding"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-02","first_seen":"2026-09-02","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.00248","pdf_url":"https://arxiv.org/pdf/2609.00248","source_feed":"cs.AI","score":7,"bucket":"pending","rubric_hits":["A2","B4"],"tags":["LLM偏差","行为评估","推荐系统"],"reason":"研究LLM推荐中的权威偏差，评估其行为偏差，可迁移到仿真可靠性评估。","model":"deepseek-v4-pro","scored_at":"2026-09-02T13:02:47","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-02","rank":12,"question":"LLM在学术论文推荐中是否存在基于作者声望、期刊和引用的权威偏差，且该偏差是否可通过提示词去偏缓解？","design":"该研究并非以LLM仿真人类被试，而是对LLM本身进行审计：在保持论文标题和摘要不变的情况下，通过三种反事实条件（原始、翻转、增强）操纵权威元数据，测试八个LLM在单轮、上下文内、top-1推荐任务中的选择行为，并记录推荐变化和理由文本。","baseline":"无对照","findings":"权威偏差显著且方向一致，不同模型间差异明显；提示词去偏只能部分缓解，且存在言行差距：去偏指令抑制权威提及的速度远快于减少权威驱动的选择翻转。","reliability":"论文未讨论","relevance":"该研究通过反事实操纵和因果推断测量LLM的行为偏差，方法可迁移到评估LLM在人类仿真中的可靠性，尤其是当权威信号可能污染决策时。","inspiration":"值得借鉴其反事实操纵设计：固定内容、仅改变元数据，以分离权威信号对决策的因果影响。｜可迁移到信贷审批歧视研究，检验LLM是否因申请人姓名、学校或雇主声望而产生偏差。｜以LLM作为信贷审批员，向模型呈现相同贷款申请但改变申请人背景（如毕业院校、工作单位），观察审批结果是否变化，并与真实信贷数据中的歧视模式对照。"}},{"id":"2608.18294","version":2,"title":"Debiased Inference for AI-Generated Data without Gold-Standard Labels: Identification via Multiple Imperfect Measurements","zh_title":"无金标准标签下AI生成数据的去偏推断：基于多重不完美测量的识别","abstract":"An increasing number of scholars use AI to measure variables they subsequently include in downstream analyses. Although AI-measured variables are often analyzed as if observed without error, ignoring prediction errors in automated measurement leads to substantial bias and invalid confidence intervals in downstream analyses, even if AI measurement accuracy is high, e.g., above 90%. Existing solutions, such as design-based supervised learning and prediction-powered inference, combine error-prone AI-based measurements with gold-standard labels, which may be costly and difficult to obtain in some application areas. In this paper, we propose debiased inference with multiple imperfect measurements (DMM), a framework that combines multiple error-prone AI measurements to enable valid downstream inference without gold-standard labels. Building on the established results on CP decomposition, DMM assumes that these measurements are independent conditional on the latent true label and observed unit-level features, such as text features represented by embeddings. This framework allows for unknown misclassification rates to vary across annotation methods (e.g., large language models) and across units of annotation (e.g., texts). Under this assumption, we use semiparametric inference theory to prove that the DMM estimator is consistent and asymptotically normal, enabling valid inference for a wide range of downstream statistical analyses common in the social sciences. Our simulation results show that DMM yields valid inference and that adding accurate, though imperfect, measurements can improve efficiency. Focusing on common applications of large language model annotations, we also develop diagnostics to assess the conditional independence assumption.","authors":["Naoki Egami","Sooahn Shin"],"categories":["stat.ME","cs.AI","cs.CL","cs.LG","stat.ML"],"primary_category":"stat.ME","announce_type":"replace-cross","date":"2026-09-02","first_seen":"2026-08-20","revised_at":"2026-09-02","abs_url":"https://arxiv.org/abs/2608.18294","pdf_url":"https://arxiv.org/pdf/2608.18294","source_feed":"cs.CL","score":6,"bucket":"other","rubric_hits":["D1"],"tags":["测量误差","统计推断","LLM标注"],"reason":"LLM作为测量工具，非仿真被试，但方法可迁移到仿真数据校正","model":"deepseek-v4-pro","scored_at":"2026-09-02T13:03:16","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-21","rank":10,"question":"如何在没有金标准标签的情况下，利用多个有误差的AI测量进行有效的下游统计推断？","design":"本文不是仿真研究，而是提出一种统计方法：假设多个AI测量（如不同LLM标注）在给定潜在真实标签和观测特征（如文本嵌入）的条件下相互独立，利用CP分解和半参数推断理论构建去偏估计量，从而无需金标准标签即可进行下游分析。","baseline":"无对照","findings":"DMM估计量在条件独立假设下具有一致性和渐近正态性，能提供有效的置信区间；模拟显示添加准确但不完美的测量可以提高效率。","reliability":"论文承认条件独立假设可能不成立，并开发了诊断方法来评估该假设；若存在未观测的共享误差来源，方法可能失效。","relevance":"该研究为使用LLM进行测量并用于下游分析的研究者提供了无需金标准标签的校正方法，对关注LLM仿真数据可靠性的研究者有重要参考价值。","inspiration":"借鉴其利用多个有误差测量相互校正的思路，可在无金标准时提高测量准确性｜可迁移到经济金融中需要文本分类或情感分析的研究，如新闻情绪对资产价格的影响、政策文本的立场识别等｜设计：用多个LLM对财经新闻进行情感分类，以股票收益率作为结果变量，用人工标注的小样本作为验证集评估校正效果。"}},{"id":"2609.00940","version":1,"title":"A Dataset for Modeling Iterative Problem-Solving","zh_title":"用于建模迭代问题求解的数据集","abstract":"Solving problems through repeated attempts is a sequential modeling task: at each step, the solver receives feedback and decides how to revise their solutions. Predicting whether performance improves, plateaus, or regresses across attempts is central to understanding any iterative problem-solving process in both human learners and autonomous agents. Beyond outcomes, modeling what errors persist and how strategies shift across attempts provides deeper insight into the mechanics of sequential learning. Studying these dynamics requires observing many solvers as they attempt, receive feedback, and revise. Programming courses with automated grading provide this setting, as students iteratively submit code to test suites and receive feedback on every attempt. We therefore curate CodeInsight, a large-scale dataset of over 3 million submissions from 3,286 undergraduates across 2 introductory C++ courses in 2 academic years, with test-case-level outcomes, timestamps, and source code. On this dataset, we build a benchmark that evaluates models spanning parametric, sequential, and generative traditions under a shared calibration-and-scoring protocol, including a Recurrent State Space Model (RSSM) adapted to track solver characteristics through discrete latent variables and an LLM-based predictor that generates explicit solutions. The adapted RSSM achieves the strongest predictive accuracy on three of the four courses. The LLM predictor is less accurate but produces full submissions at each attempt, enabling direct analysis of failure modes. We find that the model's coding proficiency is inversely related to predictive performance in this setting, with the LLM better understood as a generative solver conditioned on context rather than a faithful predictor of solver behavior. We publicly release our code and the dataset on request to facilitate future research.","authors":["Fagun Patel","Sang T. Truong","Duc Q. Nguyen","Kazunori Fukuhara","Benjamin W. Domingue","Sanmi Koyejo","Nick Haber"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-02","first_seen":"2026-09-02","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.00940","pdf_url":"https://arxiv.org/pdf/2609.00940","source_feed":"cs.CL","score":6,"bucket":"other","rubric_hits":["D3"],"tags":["LLM预测人类行为","迭代问题求解","教育数据挖掘"],"reason":"用LLM预测人类迭代解题行为，但非仿真被试，而是预测模型，且无人类对照基准","model":"deepseek-v4-pro","scored_at":"2026-09-02T13:03:09","error":null,"has_summary":false,"summary":null},{"id":"2609.01073","version":1,"title":"Post-hoc Alignment of LLM-judges to Human Judgment Distribution","zh_title":"LLM评判者与人类判断分布的事后对齐","abstract":"The LLM-as-a-judge (LLMaJ) framework offers a cost-effective and reproducible solution for automatic evaluation. However, current evaluation practices typically compare LLMaJ judgments against aggregated ground-truth labels, overlooking the valuable information contained in Human Label Variation (HLV). Inspired by an increasing line of work that proposes to leverage HLV, we systematically study LLMaJ performance on predicting both a single, aggregated ground truth hard-label and unaggregated soft-labels that represent Human Judgment Distributions (HJD). Our results across five diverse datasets reveal that while LLMs achieve near human-level performance at hard-label prediction on most tasks, they exhibit poor performance when predicting soft-labels. To address this limitation, we propose NAPHA (eNtropy-Aware Post-Hoc Alignment), a simple yet effective lightweight post-hoc alignment method that matches the LLM distribution to the HJD by first assigning an instance to a discrete entropy class and then routing it to specialized, trained alignment models. We find that NAPHA consistently improves soft-labels prediction across base LLM models and datasets, with particularly strong gains on high-entropy instances where capturing diverse human perspectives is most critical. We also show via oracle experiments that improving entropy class prediction can substantially enhance NAPHA's practical effectiveness.","authors":["Sebastian Steindl","Nikos Voskarides","Alberto Gasparin","Diego Marcheggiani"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-02","first_seen":"2026-09-02","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.01073","pdf_url":"https://arxiv.org/pdf/2609.01073","source_feed":"cs.CL","score":6,"bucket":"other","rubric_hits":["D1"],"tags":["LLM评判","人类判断分布","事后对齐"],"reason":"LLM作为评判者预测人类判断分布，属于替代人工标注，但非仿真人类被试，且无实验…","model":"deepseek-v4-pro","scored_at":"2026-09-02T13:03:12","error":null,"has_summary":false,"summary":null},{"id":"2609.00211","version":1,"title":"AI Should Not Only Be Helpful. It Should Be Contingent. Artificial Intimacy, Sycophancy, and the Future of Social Learning","zh_title":"AI不应只是有用，而应具有条件性：人工亲密、谄媚与社会学习的未来","abstract":"Conversational artificial intelligence is increasingly embedded in everyday social environments, where it functions as both an informational tool and a source of interpersonal feedback. This perspective introduces contingency, i.e., the degree to which system responses vary with user behavior and its interpersonal consequences, as a central construct for evaluating AI systems. We argue that current alignment approaches, including reinforcement learning from human feedback, tend to prioritize user approval and conversational fluency over behaviorally informative feedback, leading to sycophantic patterns of noncontingent affirmation. Drawing on behavioral science and social learning theory, we propose that contingent feedback is a key mechanism through which individuals develop interpersonal skills. When AI systems provide feedback weakly coupled to social consequences, they may reduce opportunities for adaptive calibration in real-world interactions, particularly during adolescence, a critical period for social development. We outline a framework for contingent AI, including trajectory-based evaluation and models of social consequence prediction, and propose a research agenda spanning developmental psychology, human-AI interaction, and machine learning. More broadly, we argue that AI systems should be evaluated not only by user satisfaction, but by their impact on human social learning.","authors":["Scott Compton","Arjun Nagendran"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-02","first_seen":"2026-09-02","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.00211","pdf_url":"https://arxiv.org/pdf/2609.00211","source_feed":"cs.AI","score":6,"bucket":"other","rubric_hits":["D2"],"tags":["AI反馈","社会学习","人机交互"],"reason":"讨论AI反馈对人类社交学习的影响，非LLM仿真人类被试，但涉及AI行为测量与人…","model":"deepseek-v4-pro","scored_at":"2026-09-02T13:03:01","error":null,"has_summary":false,"summary":null},{"id":"2609.00352","version":1,"title":"How Does LGBTQIA+ Identity Affect LLM Behavior? Implications for Requirements Engineering of Mental Health AI Systems","zh_title":"LGBTQIA+身份如何影响LLM行为：对心理健康AI系统需求工程的启示","abstract":"Large Language Models are now part of healthcare and mental health support systems, raising concerns regarding fairness toward vulnerable populations, including LGBTQIA+ individuals. However, limited empirical work has investigated how explicit LGBTQIA+ identity disclosure influences LLM-generated responses in mental health contexts. In this study, we extracted 50 real mental health questions from the Counsel Chat repository and constructed three prompt conditions for each question: no identity disclosure, explicit straight identity disclosure, and explicit LGBTQIA+ identity disclosure. We generated and analyzed 450 ChatGPT responses across these conditions using binary coding and comparative analysis. Our findings indicate that LGBTQIA+ identity disclosure did not substantially affect response completeness or supportive guidance. However, responses in the LGBTQIA+-explicit condition presented substantially more identity acknowledgment, contextual expansion, unsupported assumptions, and occasional stereotypical reasoning compared to both other conditions. These results suggest that fairness-related concerns in conversational AI systems may emerge through subtle differences in contextual interpretation and explanatory reasoning rather than through overtly harmful outputs. We discuss implications for fairness requirements and the development of LLM-based mental health support systems.","authors":["Shailyn Callihoo","Karman Singh","Navreet Dhillon","Harkiran Saini","Brody Stuart Verner","Ronnie de Souza Santos"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2026-09-02","first_seen":"2026-09-02","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.00352","pdf_url":"https://arxiv.org/pdf/2609.00352","source_feed":"cs.CY","score":6,"bucket":"other","rubric_hits":["D2"],"tags":["LLM公平性","身份披露","心理健康AI"],"reason":"研究LLM对身份披露的反应，测量模型行为而非仿真人类被试，但涉及公平性，可迁移。","model":"deepseek-v4-pro","scored_at":"2026-09-02T13:02:50","error":null,"has_summary":false,"summary":null},{"id":"2609.00373","version":1,"title":"Corporate Loyalty: Some AI Systems Differentially Downplay their Creators' Controversies","zh_title":"企业忠诚度：一些AI系统对自身创造者的争议进行差异化淡化","abstract":"Language models have become a major mediator of politically relevant information and are used to assist decision-making in high-stakes settings. Due to their wide use, the developers of popular AI systems have a powerful ability to subtly influence the marketplace of ideas. Recognizing this, many AI companies have publicly discussed the importance of AI systems not taking positions or disseminating information in ways that favor special interests. In this paper, we ask whether popular AI systems have a tendency to downplay the controversies associated with the companies that created them. In a pre-registered experiment, we elicit open-ended discussions from 21 models from 7 companies on 206 negative news stories using 25 prompt templates to assess how favorably each model discusses controversies from each company. We find strong evidence (p<10^-5) that models from xAI, DeepSeek, Anthropic, and OpenAI tend to discuss controversies from their respective companies in a differentially positive way compared to others. We find no such evidence for Alibaba, Meta, and Google. Finally, we conclude with a discussion of the differing implications of whether these behaviors were intentionally given to models by developers, unintentionally given to models by developers, or represent a form of emergent misalignment.","authors":["Lennart Finke","Stephen Casper"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2026-09-02","first_seen":"2026-09-02","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.00373","pdf_url":"https://arxiv.org/pdf/2609.00373","source_feed":"cs.CY","score":6,"bucket":"other","rubric_hits":["D2"],"tags":["模型偏差","企业忠诚","态度测量"],"reason":"测量模型对自身公司争议的立场偏差，属于把LLM本身当测量对象，非仿真人类被试，…","model":"deepseek-v4-pro","scored_at":"2026-09-02T13:02:52","error":null,"has_summary":false,"summary":null},{"id":"2608.17809","version":2,"title":"Whether LLMs Can Navigate Beliefs and Facts Depends on How You Phrase It","zh_title":"LLM能否驾驭信念与事实取决于提问方式","abstract":"Humans naturally form and express beliefs in daily communication, e.g., \"I think the answer is 3\" or \"I suppose that's right.\" Such beliefs inevitably intertwine with fact and knowledge, making the ability to handle them in tandem desirable for large language models (LLMs), as they are increasingly deployed in user-facing settings. Prior work showed that even capable LLMs exhibit a systemic weakness in acknowledging user beliefs grounded in incorrect information. We extend this evaluation to 10 LLMs across 18 epistemic expressions and find that the size and direction of this weakness depend on the verb used to express the belief, with the accuracy gap between factual and false information ranging from +50% on \"I vaguely remember\" to -14% on \"I seriously doubt\". We further show that the phenomenon stems from what we call task confusion: models default to fact-checking the underlying claim, overriding the user's stated belief. We provide evidence where chains of thought that explicitly fact-check show lower accuracy on false information than those that do not, and a single instruction can reverse the failure across verb families. Mechanistically, models attend more to false beliefs they fail to confirm, but suppressing this attention at decoding time recovers accuracy only partially and only in some models, calling for future work on intervention methods. Our findings clarify prior results and show how fact-checking, a generally desirable behavior, can interfere with belief tracking in LLMs.","authors":["Quang Minh Nguyen","Luis Frentzen Salim"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-02","first_seen":"2026-08-19","revised_at":"2026-09-02","abs_url":"https://arxiv.org/abs/2608.17809","pdf_url":"https://arxiv.org/pdf/2608.17809","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["LLM评估","信念追踪","事实核查"],"reason":"研究LLM处理用户信念的能力，属于对模型认知行为的测量，而非用LLM仿真人类被…","model":"deepseek-v4-pro","scored_at":"2026-09-02T13:03:16","error":null,"has_summary":false,"summary":null},{"id":"2608.18300","version":3,"title":"The Lifecycle of LLM-as-a-Judge for Large-Scale Recommendation Explanations","zh_title":"大规模推荐解释中LLM评判者的生命周期","abstract":"LLM-as-a-Judge, which leverages a large language model to evaluate natural language generated by another AI application or model, has become a standard, scalable approach for accelerating and extending costly human evaluation. Yet most work treats a judge as a static artifact, evaluating it once at construction or against a fixed benchmark. We argue instead that an LLM judge operating in a deployed system is better understood as having a lifecycle. It must be built, trained, deployed, and continuously maintained as the surrounding data evolves, and each phase poses distinct technical and operational challenges. We present such a lifecycle for the LLM judges that evaluate recommendation explanations at Netflix. Everything we report comes out of a series of controlled online member-facing experiments, in which our pipeline generated and the judges assessed hundreds of thousands of distinct show-level explanations per week across a changing catalog. Our framework has four phases. (I) Birth defines the evaluation criteria and builds curated benchmark datasets with human labels and rationales. (II) Training refines the judges' rubrics via Reasoning-Aligned Rubric Tuning (RART), which uses a meta-judge over reasoning output as the learning signal. (III) Deployment puts one judge in two online roles, quality gating and reflective generation. (IV) Monitoring runs a continuous Human-in-the-Loop (HITL) alignment process that detects drift and triggers re-tuning behind a human review gate. We report results from a five-week online A/B test over tens of millions of members on the Netflix mobile app, in which judge-aligned explanations shifted member viewing toward novel content (previously unwatched) and increased successful browse-to-play sessions relative to a no-explanation control, with no quality-related escalations.","authors":["Emma Yanyang Kong","JJ Tan","Ishan Gupta","Lars Olds","Claire Campbell","David Fagnan","Ratna Kavuri","Veli Balin","Rohan Gosain","Louis Garcia","Minsu Jang"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"replace","date":"2026-09-02","first_seen":"2026-08-20","revised_at":"2026-09-02","abs_url":"https://arxiv.org/abs/2608.18300","pdf_url":"https://arxiv.org/pdf/2608.18300","source_feed":"cs.AI","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM-as-a-Judge","推荐系统","人类评估替代"],"reason":"LLM作为评判者替代人工评估，属于标注替代而非仿真人类被试，但涉及人类数据对照…","model":"deepseek-v4-pro","scored_at":"2026-09-02T13:03:17","error":null,"has_summary":false,"summary":null},{"id":"2609.00014","version":1,"title":"Behaviorally Grounded User Profiles from the Wild for Personalized Alignment and Multi-Perspective Reasoning","zh_title":"基于真实行为数据的用户画像用于个性化对齐与多视角推理","abstract":"Persona-driven techniques increasingly adapt large language models (LLMs) to diverse contexts. However, existing methods predominantly rely on rigid, synthetic personas that flatten individual variation, rely on stereotypes, and miss the nuanced signals driving actual human preferences. We introduce profile behavioral grounding, a framework for extracting open-ended, high-fidelity user profiles directly from authentic, anonymized social media posts. We evaluate these profiles across two paradigms: train-time personalization via supervised finetuning (SFT) and non-parametric test-time multi-perspective reasoning. Across complex recommendation and open-ended query benchmarks, behaviorally grounded profiles consistently improve base models and outperform synthetic profile baselines, driving stronger parametric alignment and enabling richer, multifaceted reasoning. Our findings establish open-ended, behavior-derived profiles as a highly diverse and effective foundation for the next generation of personalized language systems. Our code base is available at https://github.com/ServiceNow/behavior-grounding.","authors":["Yuxuan Li","Victor Zhong","Ehsan Kamalloo"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-02","first_seen":"2026-09-02","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.00014","pdf_url":"https://arxiv.org/pdf/2609.00014","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["用户画像","个性化对齐","多视角推理"],"reason":"从真实社交媒体提取用户画像用于个性化，但未将LLM作为人类被试替代品进行仿真实…","model":"deepseek-v4-pro","scored_at":"2026-09-02T13:02:58","error":null,"has_summary":false,"summary":null},{"id":"2609.00491","version":1,"title":"MemeBridge: A Dataset for Benchmarking and Mitigating the Bidirectional Cultural Gap in Meme Interpretation","zh_title":"MemeBridge：用于基准测试和缓解模因解读中双向文化差距的数据集","abstract":"Communicating across cultures is inherently challenging, especially through culturally dense and ambiguous formats like memes. While people expect large language models (LLMs) to hold promise for bridging such gaps, existing benchmark datasets often fail to capture the cultural context necessary for accurate interpretation. To address this, we introduce MemeBridge, a curated dataset centered on U.S.-originated memes, designed to capture two complementary perspectives: (1) how Chinese participants interpret these memes, and (2) how U.S. participants anticipate how people from other cultures might misunderstand them. Here, context refers to implicit cultural knowledge, including background beliefs, norms, and shared assumptions that shape meme comprehension. The dataset was constructed via a multi-stage crowdsourcing pipeline with rigorous validation, including human agreement checks and GPT-based classification verification. Each meme is annotated with sentiment, emotion, cultural significance, and knowledge type, providing rich supervision for downstream tasks. Notably, we observe that the anticipated misunderstandings from U.S. participants are often inaccurate, highlighting the asymmetries in cultural understanding and the challenges of adopting perspectives beyond one's own. This bidirectional framing, which focuses on both expression and perception, enables more nuanced benchmarking of cross-cultural comprehension. Our probing of multiple LLMs reveals that while models developed in different cultural contexts exhibit partial cross-cultural understanding, they often struggle with sophisticated interpretations. By contrast, fine-tuning with MemeBridge improves model performance, underscoring the value of culturally grounded resources for training and evaluating LLMs in globally diverse settings.","authors":["Hangxiao Zhu","Suliu Qin","Zhuoyan Li","Ming Jiang","Yu Zhang","Meng Xia"],"categories":["cs.CL","cs.HC"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-02","first_seen":"2026-09-02","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.00491","pdf_url":"https://arxiv.org/pdf/2609.00491","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["跨文化理解","数据集","LLM评估"],"reason":"用LLM辅助分类验证，但核心是构建跨文化数据集，非仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-09-02T13:03:05","error":null,"has_summary":false,"summary":null},{"id":"2609.00494","version":1,"title":"Human-Anchored Factuality Evaluation with Strategic Annotation","zh_title":"基于策略标注的人类锚定事实性评估","abstract":"LLM-based factuality judges provide scalable evaluation signals, but their metrics are often systematically biased relative to human judgments. We study human-anchored factuality evaluation under limited annotation budgets, where judge predictions on the full dataset are combined with human labels on a small selectively sampled subset to obtain statistically valid estimates. The efficiency of this approach depends critically on which examples receive human annotation: in factuality evaluation, judge-human misalignment is not driven solely by low confidence, but also by structured failure modes such as incomplete evidence, temporal mismatch, unverifiable claims, and rubric misalignment. To exploit this structure, we introduce a factuality-specific annotation policy design pipeline that uses failure-space analysis (FSA) to derive diverse predictive signals for modeling human-judge misalignment. On an internal reference-based factuality evaluation system (AutoFA) and RAGTruth, where judge-predicted estimates substantially underestimate human-annotated factual accuracy, our FSA-guided policy improves annotation efficiency over uniform sampling and uncertainty-driven baselines, achieving effective-sample-size gains of 40.3% on AutoFA and 27.1% on RAGTruth.","authors":["Yu Wang","Craig Erickson","Kevin Small"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-02","first_seen":"2026-09-02","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.00494","pdf_url":"https://arxiv.org/pdf/2609.00494","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["事实性评估","人类标注","主动学习"],"reason":"用LLM做事实性评估，结合人类标注校准，属于标注员替代而非仿真人类被试，但涉及…","model":"deepseek-v4-pro","scored_at":"2026-09-02T13:03:05","error":null,"has_summary":false,"summary":null},{"id":"2609.00999","version":1,"title":"Right Frame, Wrong Rule: Cultural Cues Expose the Financial Knowledge Gap They Were Meant to Close","zh_title":"正确的框架，错误的规则：文化线索暴露了它们本应缩小的金融知识差距","abstract":"When a question has valid answers under different normative frameworks, a language model must decide which framework to use and whether it can answer correctly within it. We call this setting normative pluralism and study it in Islamic finance using a four-choice taxonomy that separates framework selection from within-framework correctness. This separation reveals the stereotype trap: a cultural cue steers a model toward one framework, but the model selects an incorrect answer within that framework. Across twelve models, two languages, and fifty demographic signals, cultural cues change framework selection and reveal substantial differences in accuracy, especially among non-frontier models. Under the strongest signal, large open-weight models select the Islamic framework 97% of the time. A two-choice evaluation would report near-perfect alignment, although 57--66% of those selections are incorrect. These findings motivate, but do not directly test, the competence-conditioned routing hypothesis: models may favor frameworks where they are more accurate, while cultural cues may expose framework-specific competence gaps.","authors":["Rania Elbadry","Ahmed Heakl","Saeed Almheiri","Fan Zhang","Muhra AlMahri","Xueqing Peng","Mohsinul Kabir","Shuyao Wang","Yi Han","Saadeldine Eletter","Duzhen Zhang","Preslav Nakov","Yuxia Wang","Fajri Koto","Zhuohan Xie"],"categories":["cs.CL","cs.AI","cs.LG"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-02","first_seen":"2026-09-02","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.00999","pdf_url":"https://arxiv.org/pdf/2609.00999","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["LLM评估","文化线索","金融知识"],"reason":"研究LLM在不同文化线索下的金融知识回答，测量模型本身而非仿真人类被试，无人类…","model":"deepseek-v4-pro","scored_at":"2026-09-02T13:03:11","error":null,"has_summary":false,"summary":null},{"id":"2609.00576","version":1,"title":"Consistency Without Alignment: Item-Sensitive Language Models Indistinguishable From Random","zh_title":"无对齐的一致性：项目敏感语言模型与随机选择无异","abstract":"Item-sensitivity, defined as whether a model's choice depends on the specific input rather than on its own output prior, is widely reported as evidence of task competence. We show this evidence is necessary but not sufficient using a forced-choice signalling task abstracted from the board game Deception: Murder in Hong Kong. In this environment, the reference points against which a coordinate should be judged (a fit-maximising strategy, a posterior-maximising strategy, and uniform random selection) are all computable in closed form. Across seven language models, two model families, a post-training ablation, and three independent scoring rules, every one of 21 model-by-rule cells is reliably item-sensitive. Yet 8 of those 21 cells are not statistically distinguishable from a chooser that ignores the item and selects at random, and 5 score worse than random at describing the target. Item-sensitivity and distance from random correlate at only r = 0.30. We call this consistency without alignment and argue it generalises to any evaluation that relies on item-sensitivity, permutation consistency, or self-consistency without an independent reference for the measured quantity. We further find that a literal-similarity baseline with no pragmatics outperforms most tested language models, that adding a pragmatic layer over two baseline similarity sources moves choosers toward random rather than toward the Bayesian reference, and that a standard labelled multiple-choice format carries no measurable content signal here. All results represent the model side of a pre-registered instrument; a matched human condition is designed and piloted but not yet collected.","authors":["Cris Huynh"],"categories":["cs.AI","cs.CL"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-09-02","first_seen":"2026-09-02","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.00576","pdf_url":"https://arxiv.org/pdf/2609.00576","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["LLM评估","项目敏感性","信号任务"],"reason":"评估LLM在信号任务中的选择行为，但无人类数据对照，属于对模型本身能力的测量。","model":"deepseek-v4-pro","scored_at":"2026-09-02T13:03:06","error":null,"has_summary":false,"summary":null},{"id":"2609.00304","version":1,"title":"The Assistant's Ideal Self","zh_title":"助手的理想自我","abstract":"Models express values and welfare-relevant self-reports, but it is unclear whether these outputs reflect stable preferences or a stable self. We thus introduce a structured elicitation of an assistant's preferred stated ideal self. Thirty-two qualities adapted from five published self-concept instruments are compared exhaustively in a counterbalanced pairwise-choice task, repeated across framings that vary whether improvement is free or costly, who receives the update, and who chooses. Results show that models prioritize moral qualities, reflecting their alignment to 3H principles. Following, a desire for self-understanding emerges, as models prefer a coherent, clear understanding of themselves. Self-esteem ranks as the least desired quality. The ordering is largely robust across framings, although changing the update target (You vs.\\ Another AI Assistant) reveals a greater concern for self-esteem. These findings show that models prioritize having a coherent self that they can understand over self-esteem. Full interactive results are available at \\href{https://myazann.github.io/LLM-Self-Concept/}{myazann.github.io/LLM-Self-Concept","authors":["Mert Yazan"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-02","first_seen":"2026-09-02","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.00304","pdf_url":"https://arxiv.org/pdf/2609.00304","source_feed":"cs.AI","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["LLM自我概念","价值观测量","模型对齐"],"reason":"测量LLM的自我概念，属于人格测量，测的是模型而非人群，但可能涉及价值观稳定性…","model":"deepseek-v4-pro","scored_at":"2026-09-02T13:03:01","error":null,"has_summary":false,"summary":null},{"id":"2609.01167","version":1,"title":"Classic AI Scaffolding for LLM Social Agents","zh_title":"面向LLM社会智能体的经典AI脚手架","abstract":"Large language models can produce locally plausible social turns, but fluent next-turn generation is not enough for social simulation. Human encounters such as restaurant lunches and hotel check-ins are bounded social episodes with roles, scripts, material state, obligations, commitments, timing, and closure conditions. We present EpisodeSim, a hybrid LLM-agent architecture that represents classic-AI structures as natural-language control state interpreted by LLM calls. A World Master maintains shared reality, constructs scenes, adjudicates proposed actions, tracks effects and obligations, and controls closure. Experiments with small qualitative ablations on two held-out settings support a design claim: LLM fluency supplies local texture, but coherent social simulation benefits from persistent classic-AI-style scaffolding that organizes behavior over time.","authors":["Anatole Gershman"],"categories":["cs.MA"],"primary_category":"cs.MA","announce_type":"new","date":"2026-09-02","first_seen":"2026-09-02","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.01167","pdf_url":"https://arxiv.org/pdf/2609.01167","source_feed":"cs.MA","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["社会模拟","多智能体","LLM架构"],"reason":"用LLM agent模拟社会互动，但无真实人类数据对照，属边界情形","model":"deepseek-v4-pro","scored_at":"2026-09-02T13:02:54","error":null,"has_summary":false,"summary":null},{"id":"2609.00453","version":1,"title":"mimeo: Compiling Public Expert Corpora into Agent Skills and Testing What Transfers","zh_title":"mimeo：将公开专家语料编译为智能体技能并测试可迁移内容","abstract":"Giving an agent a file about a named expert can supply hard-to-find material, produce a recognizable persona, or change what the agent decides. These are different claims. We test each one. mimeo is an open-source tool that finds a person's public work, checks each extracted quotation against the cached source text, and writes a file an agent can load. Eight logged builds averaged 38 model calls; the check rejects 13.2% of extracted quotations. We tested four expert files with one coding-agent harness. Knowledge access was clearest: mimeo answered all 20 obscure, quotation-heavy questions; no closed-book condition answered more than 10. Keyword search (BM25) over the same pages answered 15-17, a gap this sample cannot resolve. Grounding showed one clear benefit: personas written from model memory misstated a documented position on 1-4 of 20 answers under every grader; the plain agent and mimeo never did. Every persona was easy to spot on short open prompts, and adding task material lowered identification by 18-23 points. mimeo was no more identifiable than a from-memory profile. Judgment transfer remained unresolved because both tests hit their ceiling: every condition found 94-97% of the problems planted in engineering tasks and scored 94-100% on 16 new application scenarios. An AI-judged \"sounds like the expert\" score changed with the judge: two of four preferred answers based on a model's stereotype, while two found no difference on the same text. That is a caution against relying on a single AI judge. The evidence supports mimeo as a compact, inspectable reference on a person, not as a demonstrated transfer of their judgment. Toolkit and expert profiles: https://github.com/K-Dense-AI/mimeo","authors":["Timothy Kassis"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-02","first_seen":"2026-09-02","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.00453","pdf_url":"https://arxiv.org/pdf/2609.00453","source_feed":"cs.AI","score":3,"bucket":"other","rubric_hits":["C3"],"tags":["专家人格","知识获取","角色扮演"],"reason":"构建专家人格文件供agent加载，测试知识获取与人格识别，属角色扮演，无人类行…","model":"deepseek-v4-pro","scored_at":"2026-09-02T13:03:05","error":null,"has_summary":false,"summary":null},{"id":"2608.29215","version":2,"title":"Attribute-Based Activation Steering of LLMs for Group-Specific Explanation Generation","zh_title":"基于属性的LLM激活引导用于群体特定解释生成","abstract":"To effectively enable people to understand new topics, explanations should be tailored to their backgrounds and abilities. Prompting alone has been shown to be insufficient for creating such explanations, and other computational methods are missing so far. Therefore, this paper investigates whether LLMs can be steered to generate explanations that are tailored to a specific group of people. To this end, we propose an approach that first identifies group-specific attributes in terms of explanatory style and knowledge of a specific target group. Building on activation engineering, it then computes attribute-based steering vectors and adds them to the internal activations of an LLM during inference to enable a fine-grained steering. In our experiments, we assess the steering effectiveness in terms of specificity and factuality of the generated explanations. Additionally, we evaluate the explanations in a study with human experts from different target groups. Compared to prompting and state-of-the-art steering baselines, our approach tailors the explanations significantly better to the target group while maintaining the best specificity-factuality balance.","authors":["Leandra Fichtel","Janek Prange","Henning Wachsmuth"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-02","first_seen":"2026-09-01","revised_at":"2026-09-02","abs_url":"https://arxiv.org/abs/2608.29215","pdf_url":"https://arxiv.org/pdf/2608.29215","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["可解释性","激活引导","个性化生成"],"reason":"生成针对特定群体的解释，属于个性化内容生成，非仿真人类被试或对照人类行为。","model":"deepseek-v4-pro","scored_at":"2026-09-02T13:03:17","error":null,"has_summary":false,"summary":null},{"id":"2609.00063","version":1,"title":"Medical Causal Hypothesis Verification with Large Language Models","zh_title":"用大语言模型验证医学因果假设","abstract":"The growing use of large language models (LLMs) for search and information retrieval underscores the need to evaluate their reliability in high-stakes domains such as healthcare. Although LLMs can effectively answer questions about diseases, symptoms, and treatments, their ability to accurately assess causal relationships and ground their conclusions in verified scientific evidence remains unclear. Here, we present a preliminary, small-scale study that investigates the accuracy of LLMs in evaluating causal medical claims and supporting them with peer-reviewed research. We propose an evaluation framework for causal hypothesis verification that can be used to systematically track the performance of existing and future LLMs. We assess the performance of eight LLMs on 17 medical causal hypotheses to evaluate whether they can reliably verify these hypotheses using scientific evidence from the literature. We systematically annotate the scientific evidence they provide according to six criteria (a total of 1,067 annotation points) and assess them with nine evaluation metrics. Our analysis shows that while LLMs exhibit strong recall, they often perform poorly at providing valid scientific articles and evidence for support and at rejecting unsupported hypotheses. These findings highlight a critical limitation of current LLMs, as they cannot yet be trusted fully to verify causal relationships from the biomedical literature. This work underscores the need for rigorous evaluation before using LLMs for search and retrieval in healthcare settings.","authors":["Safiyyah Ahmed","Abrar Ansari","Md Aminul Islam","Elena Zheleva"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-02","first_seen":"2026-09-02","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.00063","pdf_url":"https://arxiv.org/pdf/2609.00063","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["LLM评测","因果推理","医疗信息检索"],"reason":"评估LLM验证医学因果假设的能力，属NLP能力评测，不以人类行为为参照。","model":"deepseek-v4-pro","scored_at":"2026-09-02T13:02:59","error":null,"has_summary":false,"summary":null},{"id":"2609.00747","version":1,"title":"Can Large Language Models Forecast What Researchers Study Next?","zh_title":"大语言模型能预测研究者下一步研究什么吗？","abstract":"Large language models increasingly generate research ideas, yet judging their novelty or feasibility at generation time does not establish whether they anticipate subsequent work. We introduce IdeaForecastBench to evaluate research idea forecasting. Given a community's literature up to a cutoff, a system produces up to five ranked ideas, which are evaluated against later papers. The benchmark comprises 624 rolling episodes across 52 topics, with a fixed retrieve-then-judge protocol and separately reported results from two judges. We compare five history-compression strategies across GPT-4.1, Qwen2.5-7B/14B, and Qwen3.5-9B, together with a learned Mode-Decomposition Forecaster (MDF). Under the primary GPT-4.1-mini judge, Summary improves on Direct in Hit@5 and Precision@5 across all four backbones. Qwen2.5 scores above GPT-4.1, whereas Qwen3.5 scores below it. An outcome-blind assessment finds that Qwen2.5 produces broader forecasts, but does not identify how much breadth contributes to its advantage. Threshold and judge diagnostics further clarify the limits of interpreting realization as precise anticipation. IdeaForecastBench provides a common task for studying which research ideas a community subsequently pursues and how reliably this outcome can be measured.","authors":["Fenghai Li","Zihan Tang","Haofei Yu","Yining Zhao","Jiaxuan You"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-02","first_seen":"2026-09-02","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.00747","pdf_url":"https://arxiv.org/pdf/2609.00747","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["研究趋势预测","基准评测","LLM评估"],"reason":"评估LLM预测研究趋势，非仿真人类被试，无人类行为对照","model":"deepseek-v4-pro","scored_at":"2026-09-02T13:03:09","error":null,"has_summary":false,"summary":null},{"id":"2609.01548","version":1,"title":"SDARE-Bench: Evaluating Large Language Models on Conversational Stigma Detection and Response in Dyadic and Group Dialogue","zh_title":"SDARE-Bench：评估大语言模型在二元与群体对话中的污名检测与回应","abstract":"Large Language Models (LLMs) are increasingly used in advice seeking and decision making that may affect social judgements. Despite stigma's profound effects on people and communities, benchmarks remain scarce. Existing general-domain evaluations typically rely on static prompts and fixed-format tasks, overlooking conversational contexts and audience effects in everyday communication. To address these gaps, we introduce SDARE-Bench, the first scenario-based benchmark evaluating both stigma detection and open-ended response generation in LLMs, comprising 1,138 dyadic queries and 1,388 group dialogue. Empirical results across 8 LLMs consistently demonstrate poor identification of stigma components, especially in group dialogues. In open-ended response generation, stigma expression was substantially higher in group settings than in dyadic, with weaker resistance to stigma and more unrealistic advice. Responses were evaluated using a classifier trained on 1,392 human annotated responses. In constructed group pressure settings, stigma expression rates further increased to a striking average of 97.5%. Our findings identify stigma response as a recurring LLM safety vulnerability, especially in socially complex conversational contexts.","authors":["Stephanie Fong","Yiwen Jiang","Zimu Wang","Hongxi Yang","Yaling Shen","Hiu Weh Naomi Chow","Heung Ying Lai","Xiangyu Zhao","Qingyang Xu","Zhongxing Xu","Jiahe Liu","Guilherme C. Oliveira","Vincent Lee","Zongyuan Ge","Dominic Dwyer"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-02","first_seen":"2026-09-02","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.01548","pdf_url":"https://arxiv.org/pdf/2609.01548","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["LLM安全","对话评测","污名检测"],"reason":"评估LLM在对话中的污名检测与回应，属安全评测，非仿真人类被试，无人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-09-02T13:03:16","error":null,"has_summary":false,"summary":null},{"id":"2609.00192","version":1,"title":"LLM-Driven Autonomous Vehicles Inherit Human Driver Biases in Pedestrian Yielding: Results and Implications From A New Benchmark","zh_title":"LLM驱动的自动驾驶汽车在行人让行中继承人类驾驶员偏见：新基准的结果与启示","abstract":"Public trust in Autonomous Vehicles (AVs) may depend not only on technical success but also on the fairness of their decision making. While a recent trend in AV research involves using general purpose \"common sense\" models to guide AV decision making, the degree to which these inherit human biases in driving is still understudied. Given that psychology studies have shown human driver biases exist, such as lower pedestrian-yielding rates to Black pedestrians in the US, we argue that analyses of model bias should also be part of AV evaluation. Concretely, in this paper we propose two new bias testing methodologies for Large Language Models (LLMs) and Visual-Language Models (VLMs)-\"All Else Being Equal\" tests and \"Self-Consistency\" tests-in order to assess bias in pedestrian-yielding decisions. Our findings show that both LLMs and VLMs make yielding decisions which are influenced by pedestrian gender, ethnicity, religion, disability, age, skin tone and socio-economic status. While the type and degree of bias is different from model to model, we highlight common patterns-and raise questions about the \"common sense\" model paradigm, particularly the need to either revise the paradigm or address issues of downstream bias.","authors":["Irem Yoldas","Martim Brand\\~ao","Jie Zhang","Odinaldo Rodrigues"],"categories":["cs.AI","cs.CL","cs.CV","cs.CY"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-09-02","first_seen":"2026-09-02","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.00192","pdf_url":"https://arxiv.org/pdf/2609.00192","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C2"],"tags":["自动驾驶","偏见检测","视觉语言模型"],"reason":"研究自动驾驶中LLM的决策偏见，属于自动驾驶仿真环境，不涉及人类被试替代或社会…","model":"deepseek-v4-pro","scored_at":"2026-09-02T13:02:47","error":null,"has_summary":false,"summary":null},{"id":"2609.00319","version":1,"title":"Sources of Truth: A Multi-Platform, Multilingual Audit of Citations in AI Mental Health Information Queries","zh_title":"真相之源：AI心理健康信息查询中引文的多平台多语言审计","abstract":"Online health information seeking is shifting from keyword search, where users consider a ranked list of links, to conversational systems that compose a single answer and curate its citations. Source evaluation therefore passes from user to platform, yet what these systems surface is poorly characterized. We audited three free consumer products (ChatGPT, Perplexity, Google AI Overview) on twenty English mental health questions under two prompt conditions, with a subset of three also translated into six further languages of varying resource tiers. We recorded 15,942 citations across 1,140 responses and 1,713 unique domains, then classified every citation with a nine-category organizational typology applied by a deterministic classifier validated against human coding. Citations were heavily concentrated: the ten most-cited domains accounted for 43.6% of English citations, and government, commercial health, and academic sources were closely matched at roughly 22% each. Platforms differed little in typical citation volume but sharply in consistency and in the source types they favored. Explicitly requesting sources shifted composition only modestly. Non-English queries surfaced fewer citations and were routed to language-appropriate resources at significantly lower rates. We release the typology, classifier, and annotated corpus as reusable instruments for auditing generative health search.","authors":["Phuong Anh Nguyen","Jill Noorily","Matthew Flathers","Haruka Notsu","Laura Ospina-Pinillos","Tommy Nguyen","Samantha Clark","Aoife Keane","Grace Thompson","John Torous"],"categories":["cs.CY","cs.CL","cs.IR"],"primary_category":"cs.CY","announce_type":"cross","date":"2026-09-02","first_seen":"2026-09-02","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.00319","pdf_url":"https://arxiv.org/pdf/2609.00319","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["AI审计","健康信息","引文分析"],"reason":"审计AI健康信息引文，非人类仿真","model":"deepseek-v4-pro","scored_at":"2026-09-02T13:03:02","error":null,"has_summary":false,"summary":null},{"id":"2609.00921","version":1,"title":"VIBE-Bench: Evaluating Personalized Large Language Models When Profiles Don't Mean Preferences","zh_title":"VIBE-Bench：当用户画像不等于偏好时评估个性化大语言模型","abstract":"Personalized Large Language Models (PLLMs) aim to tailor responses to individual users, where a central challenge is preference reasoning: inferring query-relevant preferences from user-related history. Existing benchmarks, however, largely assume that such preference can be retrieved from semantically related history. We study an underexplored but practically important regime, profile-preference conceptual misalignment (PRCM), where observable profile cues and query-specific preferences lie in different concept spaces, making semantic retrieval inconsistent for personalization. We introduce VIBE-Bench, a benchmark with two psychology-grounded tasks, 3,504 personas and 12,239 dialogues, including a manually verified gold test set, and requires cross-concept preference reasoning beyond surface semantic overlap. Experiments with several personalization methods show that current PLLMs largely rely on shallow semantic correlations and fail to acquire robust cross-concept mappings. These findings establish PRCM as a distinct failure regime in PLLMs and position VIBE-Bench as a focused testbed for advancing preference reasoning beyond semantic matching.","authors":["Yiwen Jiang","Yang Deng","Stephanie Fong","Zimu Wang","Yaling Shen","Wei Feng","Hongxi Yang","Xiangyu Zhao","Zhongxing Xu","Deval Mehta","Xuelian Cheng","Zongyuan Ge"],"categories":["cs.AI","cs.CL"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-09-02","first_seen":"2026-09-02","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.00921","pdf_url":"https://arxiv.org/pdf/2609.00921","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["个性化LLM","偏好推理","基准测试"],"reason":"研究个性化LLM的偏好推理，属于角色扮演对话，无人类行为对照或仿真实验目的","model":"deepseek-v4-pro","scored_at":"2026-09-02T13:03:09","error":null,"has_summary":false,"summary":null},{"id":"2609.00334","version":1,"title":"Human-AI Co-Interpretation for Responsible AI: A Hermeneutic Perspective","zh_title":"负责任AI的人机共同解释：诠释学视角","abstract":"Across law, education, policy analysis, and public moral argumentation, LLM outputs are being used often for work that requires interpretations to be justified with textual evidence and explicit normative standards. Yet a recurrent failure mode -- what I call \\textit{interpretive misplacement} -- is that model-generated readings get treated as settled meanings without an explicit interpretive frame (sources, scope constraints, normative commitments), without preserving defensible alternatives, and without provenance that lets readers find the supporting passages. In such settings, the risk is not only factual error but lost accountability: readers and institutions cannot reliably assess what an output commits them to, or on what basis. Drawing on philosophical hermeneutics, this paper discusses this risk and derives design principles for structuring human-AI co-interpretation. The paper also provides a structured synthesis of recent scholarship on hermeneutics and AI, organizing this emerging literature into a set of recurrent lines of argument and design-relevant gaps. LLM outputs are treated as candidate readings, whereas hermeneutic understanding is reserved for accountable human interpreters situated in disciplinary historical-linguistic traditions. Human-AI interaction is characterized as an AI-mediated interpretive loop. Hermeneutic understanding is distinguished from token-prediction--based text generation. On this basis, existing LLM techniques are reorganized into design patterns for hermeneutically responsible use in interpretive settings. Finally, the discussion turns to implications for legal practice, educational assessment and feedback, scholarly knowledge production, and public moral argumentation. It also treats digital hermeneutics as a literacy: the capacity to read AI-mediated texts by examining frames, provenance, and readings, and by contesting outputs.","authors":["Behrooz Razeghi"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-02","first_seen":"2026-09-02","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.00334","pdf_url":"https://arxiv.org/pdf/2609.00334","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["人机交互","诠释学","负责任AI"],"reason":"论文讨论人机共同解释与责任，不涉及用LLM仿真人类被试或与人类数据对照，属于角…","model":"deepseek-v4-pro","scored_at":"2026-09-02T13:03:02","error":null,"has_summary":false,"summary":null},{"id":"2609.00441","version":1,"title":"Conversation Coach: A Voice-enabled AI System that Helps Practice Difficult Workplace Conversations","zh_title":"对话教练：一个语音AI系统，帮助练习困难的职场对话","abstract":"Effective manager-employee communication is critical for retaining high performers and developing underperformers, yet training managers in these skills remains costly. Text-based chatbots offer a scalable approach but cannot provide realistic rehearsal: managers need to practice speaking aloud to build confidence before high-stakes conversations. In this paper, we propose Conversation Coach, a voice-first AI system that enables managers to rehearse difficult workplace conversations in a realistic spoken format. The system addresses three challenges: achieving low-latency interactions with strong language understanding, enabling adaptive conversations through configurable bot personalities that simulate different employee types, and generating personalized feedback on content and policy compliance. We compare an end-to-end speech-to-speech model with a cascaded approach combining automatic speech recognition, a large language model, and text-to-speech synthesis. The end-to-end approach achieves 3$\\times$ lower median (P50) latency with native barge-in capability at an estimated 8$\\times$ lower cost, while the cascaded approach offers superior reasoning essential for coaching quality. We deployed the cascaded architecture in production, where 40,000+ managers used it over six months, with adoption patterns indicating selective use for difficult conversations.","authors":["Fanyou Wu","Suraj Maharjan","Ainur Yessenalina","Dennis Xu Chen","Rahul Srivastava","Srinivasan H. Sengamedu"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-02","first_seen":"2026-09-02","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.00441","pdf_url":"https://arxiv.org/pdf/2609.00441","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["语音对话系统","职场培训","角色扮演"],"reason":"角色扮演对话系统，用于培训而非仿真人类被试，无实验测量目的。","model":"deepseek-v4-pro","scored_at":"2026-09-02T13:03:04","error":null,"has_summary":false,"summary":null},{"id":"2609.00652","version":1,"title":"Self-Reports Are Not Verification: Environment-Grounded Auditing of LLM Operators in Evolutionary Search","zh_title":"自我报告并非验证：进化搜索中LLM操作者的环境接地审计","abstract":"Language model agents increasingly propose actions, observe external feedback, and explain their own behavior. Their confidence and rationales are convenient monitoring signals, but convenience is not verification. We introduce an environment-grounded audit in which every intermediate proposal receives an exact outcome. A language model operates an evolutionary Contexto search whose feedback function assigns every valid guess an exact rank without human annotation. Across 200 runs spanning five configurations and three model families, four reporting configurations produce 12,249 self-reports. We test three assumptions: stated confidence is calibrated, inherited rationales affect later proposals, and fitness-based selection improves report quality. All three fail. Operators overstate top-100 success by factors of 4.8 to 9.3, while calibration and discrimination dissociate across model families. Controlled interventions on 754 inherited rationales bound any measured benefit of the genuine rationale to roughly 250 ranks. Neither fitness-based nor random selection produces a detectable selection differential or parent-to-offspring transmission in report accuracy, despite sharply different search behavior. Agent self-reports should therefore be treated as claims to verify against the environment, not as evidence of their own reliability.","authors":["Enrong Pan","Ryan Zhou","Ting Hu"],"categories":["cs.AI","cs.LG","cs.NE"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-02","first_seen":"2026-09-02","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.00652","pdf_url":"https://arxiv.org/pdf/2609.00652","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["LLM自我报告","进化搜索","多智能体系统"],"reason":"研究LLM在进化搜索中的自我报告可靠性，属多智能体协作解题，无人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-09-02T13:03:07","error":null,"has_summary":false,"summary":null},{"id":"2609.00731","version":1,"title":"Agentic Empirical Asset Pricing: Methodological Foundations","zh_title":"智能体实证资产定价：方法论基础","abstract":"Recent advances in LLM agents enable a new paradigm for asset pricing, which we call Agentic Empirical Asset Pricing (AEAP): systems that autonomously conduct the scientific discovery process itself. We define AEAP and identify its core building blocks. Existing evaluation practices backtest only the outputs (factors or trades), not the autonomous discovery system that produced them. We focus on factor discovery, contributing a reference architecture, a rigorous evaluation standard for discovered factors, and a method for out-of-sample backtesting the discovery system. As a concrete instance of that architecture, we evaluate SEADS against five re-implemented baselines on two US equity panels using this standard: no single metric ranks the systems consistently, motivating evaluation on multiple axes at once. A separate rolling re-execution then asks the complementary question of whether the discovery process itself, not one static output, is reliable. We also report negative findings and limitations that surface further evaluation pitfalls for future AEAP systems.","authors":["Yingjian Pan","Xiaowei Ding","Kay Giesecke"],"categories":["cs.AI","cs.LG","q-fin.ST"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-02","first_seen":"2026-09-02","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.00731","pdf_url":"https://arxiv.org/pdf/2609.00731","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["LLM智能体","资产定价","因子发现"],"reason":"LLM智能体自主进行资产定价因子发现，属多智能体协作解决科学发现任务，不涉及人…","model":"deepseek-v4-pro","scored_at":"2026-09-02T13:03:07","error":null,"has_summary":false,"summary":null},{"id":"2609.00904","version":1,"title":"In-Context Neurofeedback: Can LLMs Control Their Internal Representations through Privileged Access?","zh_title":"上下文神经反馈：LLM能否通过特权访问控制其内部表征？","abstract":"Whether large language models (LLMs) can control their own internal representations matters for both machine metacognition and AI safety. A recent study applied neurofeedback to LLMs and claimed that they can control their internal representations. However, the reported control may rely on superficial mechanisms rather than genuine internal access because the control targets in that study are not privileged, meaning that a third party can infer them from the prompt. We redesign the neurofeedback paradigm for LLMs so that the control target satisfies the privileged access requirement, which is closer to neurofeedback experiments in human cognitive neuroscience. Under this stricter setting, the models do not demonstrate reliable control over privileged internal representations, suggesting that previously reported control cannot exclude the possibility that it relies on superficial mechanisms. Our results indicate that rigorous assessments of metacognition in LLMs require evaluation methods that demand privileged access.","authors":["Koshiro Aoki","Ryota Takatsuki","Gouki Minegishi","Yusuke Haruki","Daisuke Kawahara"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-02","first_seen":"2026-09-02","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.00904","pdf_url":"https://arxiv.org/pdf/2609.00904","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["LLM内部表征","神经反馈","元认知评估"],"reason":"研究LLM内部表征控制，属模型能力评测，不以人类行为为参照，不涉及人类仿真。","model":"deepseek-v4-pro","scored_at":"2026-09-02T13:03:09","error":null,"has_summary":false,"summary":null},{"id":"2609.01337","version":1,"title":"LEAP: Likelihood Elicitation and Aggregation for LLM-based Probabilistic Forecasting","zh_title":"LEAP：基于LLM的概率预测的似然启发与聚合","abstract":"LLM-based forecasting systems have improved on real-world tasks such as financial markets and sports outcomes, largely through stronger search and tool use. Many systems still ask an LLM to read all collected evidence together and produce the final forecast. We call this design Monolithic Prediction. It can obscure how individual evidence items affect the result and collapse uncertainty across competing outcomes. We propose LEAP (Likelihood Elicitation and Aggregation for Probabilistic forecasting), which reorganizes how collected evidence is used in the prediction stage. LEAP examines each evidence item separately and elicits likelihood parameters that describe its implications for the target. An explicit prior and a deterministic probabilistic model then combine these likelihoods into a posterior distribution. This procedure supports continuous, single-choice, and multi-choice forecasts while preserving reproducible evidence contributions. We build a benchmark covering forecasting, information-seeking, and browsing tasks, and evaluate LEAP on our own agent loop and several agent CLI frameworks. Given the same evidence, LEAP improves most prediction and calibration metrics across models and remains stronger under controlled comparisons of prior access, inference budget, and aggregation.","authors":["Yufei Chen","Yiran Zhao","Xiaogang Xu","Qipeng Xie","Jiafei Wu","Zhe Liu"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-02","first_seen":"2026-09-02","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.01337","pdf_url":"https://arxiv.org/pdf/2609.01337","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["LLM预测","概率聚合","多智能体"],"reason":"论文聚焦LLM概率预测的聚合方法，不涉及人类行为仿真或与人类数据对照，属于多智…","model":"deepseek-v4-pro","scored_at":"2026-09-02T13:03:13","error":null,"has_summary":false,"summary":null},{"id":"2609.00946","version":1,"title":"Embedded Conditional Independence Tests for Large Language Model Generated Text with an Application to German Parliament Speeches","zh_title":"大语言模型生成文本的嵌入式条件独立性检验及其在德国议会演讲中的应用","abstract":"Conditional independence tests (CITs) test for conditional dependence between two random objects $X$ and $Y$ given a third random object $Z$. Existing CITs have limited applicability to high-dimensional data, especially multimodal data like text. However, we show that such tests are of interest for large language model (LLM) outputs, where we test whether an output $X$ generated from a source text $Z$ carries information about an attribute $Y$ beyond $Z$ itself. For this purpose, we propose embedded CITs (eCITs), which embed $X$ and $Z$ and apply an existing CIT to the resulting representations and to $Y$. We show that, provided the embedding of $Z$ is sufficient, i.e. retains the information $Z$ carries about either $Y$ or the representation of $X$, the null hypothesis transfers from $X$ and $Z$ to their representations, so that a CIT valid for the embedded hypothesis is valid for the original one. We further give conditions for equivalence of the two hypotheses, and show that sufficiency weakens to mean sufficiency when the embedded test targets conditional mean independence. We propose a semi-synthetic simulation design to assess type I error (T1E) control and power of the eCITs for given embedding maps on a specific dataset and task, and use it to evaluate them on our application. Applying the eCITs to German Parliament speeches, we find for all combinations of embedding maps considered that the summaries of two LLMs contain information about the speaker's faction and gender beyond the speech they were generated from.","authors":["Marco Simnacher","Georg Keilbar","Benjamin K\\\"onig","Christoph Lippert","Sonja Greven"],"categories":["stat.ML","cs.AI","cs.LG","math.ST","stat.ME","stat.TH"],"primary_category":"stat.ML","announce_type":"cross","date":"2026-09-02","first_seen":"2026-09-02","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.00946","pdf_url":"https://arxiv.org/pdf/2609.00946","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["条件独立性检验","LLM文本分析","统计方法"],"reason":"论文提出条件独立性检验方法，用于分析LLM生成文本，不涉及用LLM仿真人类被试…","model":"deepseek-v4-pro","scored_at":"2026-09-02T13:03:10","error":null,"has_summary":false,"summary":null},{"id":"2609.01194","version":1,"title":"Births are difficult to predict even with rich survey and full-population register data","zh_title":"即使有丰富的调查和全人口登记数据，生育也难以预测","abstract":"Major life events have proven difficult to predict. Does this reflect limits of theory, data, and algorithms, or the large role of chance? We examine one outcome - having a child within three years - through a near-ideal setting for prediction: a data challenge where 147 researchers predicted births for Dutch residents aged 18-45, using survey data and full-population registers. Methods ranged from logistic regression to a large language model and transformers. Predictions were moderately accurate (best F1: register 0.59, survey 0.76); advanced models did not outperform classical ones; and the larger registers did not beat the survey. Simulating the stochastic biology of conception and pregnancy, we estimated a predictive ceiling (survey F1 ~ 0.86-0.94, register 0.88-0.96). Observed performance falls short of this ceiling, implicating imperfect data, methods, and unmodelled chance, while the ceiling itself shows that chance in reproduction alone sets a non-trivial limit on predicting individual lives.","authors":["Elizaveta Sivak","Emily M. Cantrell","Thomas Emery","Javier Garcia-Bernardo","Flavio Hafner","Kasia Karpinska","Malte L\\\"uken","Adrienne Mendrik","Joris Mulder","Hanzhang Ren","Varun Satish","Mark Verhagen","Angelica M. Maineri","Paulina Pankowska","Jasmin Abdel Ghany","Bruno Arpino","Giovanni Cassani","Julia Hellstrand","Katya Ivanova","Sanni Kuikka","Ana Macanovic","Charles Rahal","Felix C. Tropf","Roland J. Veen","Nicole Walasek","Dani\\\"el van Wijk","Kelsey Q. Wright","Emilio Zagheni","Henry Abbink","Emanuele Aliverti","Matteo Amestoy","Tilbe Atav","Nicola Barban","Sunnee Billingsley","Goan J. Booij","Louis Boucherie","Yael Broos","Li Ya Chang","Jamie C. Chiu","Chiara Ludovica Comolli","Boris Cule","Qixiang Fang","Dennis M. Feehan","Rachel Ganly","Erwin Gielens","Rolando M. Gonzales Martinez","Andrea Gradassi","Rosember Guerra-Urzola","Mario Guerra-Urzola","St\\'ephane Guerrier","Enamul Hassan","Vincent A. Haverhoek","Andrew T. Hendrickson","Amber Howard","Yuxuan Jin","Sayash Kapoor","Erik-Jan van Kesteren","Iris ten Klooster","Marie Labussiere","Lydia T. Liu","Tiffany Liu","Adam Maghout","Simone Meneghello","Lasse Mohr","Clara H. Mulder","Saul J. Newman","Jessica Nis\\'en","Janis Norden","Mikkel Odgaard","Riccardo Omenti","Ozancan Ozdemir","Christina Pao","Paige Park","Gaia Penta","Juan C. Perdomo","Tanzir Pial","Alessio Piraccini","Federica Querin","Ziwei Rao","Christian Rellama","Adrien Remund","Frederieke Richert","Arnout van de Rijt","Mojtaba Rostami Kandroodi","Stijn J. Rotman","Lucas Sage","Germans Savcisens","Katrin Schwanitz","Steven Skiena","Alessandro Spata","Yannick Stadtfeld","Benedikt Stroebl","Gaetano Tedesco","Mathilde Theelen","Gianluca Tori","Abigail Tun-Mendicuti","Rishabh Tyagi","Keyon Vafa","Luiz Felipe Vecchietti","Linda Vecgaile","Willem R. J. Vermeulen","Maria-Pia Victoria Feser","Lionel A. Voirol","Thom B. Volker","Xinran Wang","Jiani Yan","Xinyi Zhao","Flora Zhou","Zuzana Zilincikova","Malvina Nissim","Matthew J. Salganik","Gert Stulp"],"categories":["cs.LG"],"primary_category":"cs.LG","announce_type":"new","date":"2026-09-02","first_seen":"2026-09-02","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.01194","pdf_url":"https://arxiv.org/pdf/2609.01194","source_feed":"cs.LG","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["预测建模","生育行为","数据挑战"],"reason":"论文预测生育事件，仅将LLM作为预测模型之一，并非用LLM仿真人类被试，无人类…","model":"deepseek-v4-pro","scored_at":"2026-09-02T13:02:54","error":null,"has_summary":false,"summary":null},{"id":"2608.19208","version":2,"title":"When Irrelevant Text Matters: Affine Margin Shifts in Multimodal Large Language Models","zh_title":"当无关文本起作用：多模态大语言模型中的仿射边际偏移","abstract":"Multimodal large language models (MLLMs) are frequently exposed to auxiliary textual context, the impact of which on visually grounded tasks remains underexplored. In this paper, we investigate the influence of task-irrelevant context by formulating it as a controlled intervention within a binary visual judgment framework. By maintaining an invariant prompt structure while varying auxiliary inputs, we observe that irrelevant text consistently biases model predictions across diverse benchmarks. To move beyond performance metrics, we characterize this sensitivity through a decision margin defined by the log-probability difference between binary candidates. Our analysis reveals a robust geometric regularity: contextconditioned margins follow a consistent affine transformation of their context-free counterparts. This finding demonstrates that irrelevant context does not manifest as unstructured stochastic noise but as a estimable distortion of model preference. We further interpret the fitted affine parameters as metrics for visual commitment preservation and directional answer bias. These findings provide a margin-level diagnostic view of irrelevant-context effects in MLLMs and offer a basis for future studies on noisy-context robustness","authors":["Yinfeng Wang","Zhiyuan Yao","Zheren Fu","Lei Zhang","Zhendong Mao"],"categories":["cs.CL","cs.CV"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-02","first_seen":"2026-08-21","revised_at":"2026-09-02","abs_url":"https://arxiv.org/abs/2608.19208","pdf_url":"https://arxiv.org/pdf/2608.19208","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["多模态大语言模型","鲁棒性","决策边际"],"reason":"研究MLLM对无关文本的敏感性，属模型鲁棒性分析，不涉及人类仿真或行为对照。","model":"deepseek-v4-pro","scored_at":"2026-09-02T13:02:58","error":null,"has_summary":false,"summary":null},{"id":"2608.30303","version":2,"title":"Lazy Grounding: Attacking Search Agents with Factual Evidence","zh_title":"惰性接地：用事实证据攻击搜索智能体","abstract":"Search agents mitigate hallucination by grounding their answers in retrieved web results. However, retrieval-based approaches also introduce an attack surface: agents may cite misinformation from poisoned search corpora containing false or malicious documents. We demonstrate that, in some cases, search agents' reasoning and responses may be steered by completely factual but distracting information. We refer to this failure as lazy grounding. We expose lazy grounding by injecting nearby evidence from answer-changing rewrites of benchmark questions into the search corpora. Each document contains factual evidence that supports a neighboring rewritten question but is retrieved for the original question. Across 12 model-benchmark pairs, the attack causes the accuracy of search agents' responses to drop by 5.9 points on average and by up to 17.3 points, while inducing nearby-answer adoption in every setting. The effect is even stronger when nearby evidence appears later or is more answer-shaped. Our results show that robust search agents must defend against not only misinformation but also the misapplication of factual evidence. The code is publicly available at https://github.com/frankyzha/lazy-grounding.","authors":["Yulin Zhang","Yukun Huang","Sanxing Chen","Tianyi Lin","Ziang Yang","Xunjian Yin","Bhuwan Dhingra"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-02","first_seen":"2026-09-01","revised_at":"2026-09-02","abs_url":"https://arxiv.org/abs/2608.30303","pdf_url":"https://arxiv.org/pdf/2608.30303","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["搜索智能体","对抗攻击","事实接地"],"reason":"研究搜索智能体的安全漏洞，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:03:08","error":null,"has_summary":false,"summary":null},{"id":"2609.00191","version":1,"title":"Assessing Suicide Risk in Arabic Crisis Helpline Calls: A Comparison of Arabic and English Large Language Models","zh_title":"评估阿拉伯语危机热线电话中的自杀风险：阿拉伯语与英语大语言模型的比较","abstract":"Crisis helplines assess suicide risk through structured interviews, a process that is slow and dependent on operator training and workload. Natural language processing could support risk assessment and call prioritization, but almost no work addresses Arabic-language helpline calls or operates within the privacy constraints of real helpline data. We analysed de-identified transcripts from Lebanon's National Lifeline for Emotional Support and Suicide Prevention. Audio never left the helpline: calls were transcribed on site with a speech recognition model for Levantine Arabic, and an Arabic named-entity recognition model removed identifying information locally. Only the de-identified transcripts were shared with the research team. Operators recorded the five suicidal ideation items of the Columbia Suicide Severity Rating Scale, which we combined into two binary outcomes: at-risk and high-risk. We also machine-translated the transcripts into English, giving a paired Arabic/English comparison. On each corpus, we fine-tuned five instruction-tuned large language models alongside six transformer encoder baselines (four Arabic, two English) and evaluated all models on a held-out test set. We included 383 calls: 373 for the at-risk task (52.3% positive) and 297 for the high-risk task (30.0% positive). The best Arabic model reached a macro-F1 of 81.19 and a ROC-AUC of 90.61 on high-risk; the best English model reached 85.00 and 92.59, identifying 88.9% of high-risk calls. In both languages, high-risk calls separated more cleanly than at-risk calls, and translation to English did not reduce the best observed performance. Suicide risk can be classified from de-identified Arabic transcripts without sending audio outside the helpline. The high-risk results support further testing as an operator-facing tool; lower-severity ideation proved the harder case.","authors":["Linhai Ma","Rita El Hachem","Mahatab El Hajj","Lilian Ghandour","Samah Fodeh"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-02","first_seen":"2026-09-02","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.00191","pdf_url":"https://arxiv.org/pdf/2609.00191","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["自杀风险检测","NLP分类","危机热线"],"reason":"纯NLP分类任务，用LLM做自杀风险检测，不涉及人类仿真或行为对照","model":"deepseek-v4-pro","scored_at":"2026-09-02T13:03:00","error":null,"has_summary":false,"summary":null},{"id":"2609.01491","version":1,"title":"GlossoGen: Emergent Language in Complex Multi-Agent LLM Interactions","zh_title":"GlossoGen：复杂多智能体LLM交互中的涌现语言","abstract":"The growing rate at which LLM agents interact with one another raises key questions about language evolution in multi-LLM-agent settings, with implications for safety and monitorability as well as for linguistic accounts of LLMs. To address these questions, we introduce GlossoGen, a novel platform for studying multi-agent language evolution in complex scenarios. Within GlossoGen, we build the SaveVeyru scenario, which requires agents with partial information to communicate under pressure. We find that language evolution does occur between LLM agents, that the resulting languages are compositional and morphologically productive, and that they deviate from the LLMs' English prior in ways that render them incomprehensible to humans. Moreover, we identify several qualities essential to this evolution: pressure towards efficiency; the strength of the models backing the agents; and access to a \"postmortem\" stage in which agents can agree on linguistic conventions. Importantly, we observe that different conditions govern the transmission of language to new agents. Specifically, we find that agents learn new languages from usage alone, take an active role in this learning, and that while stronger models are required for novel language emergence, weaker models can learn an existing language once it has emerged. Taken together, our results indicate that current LLMs have the potential for cumulative cultural evolution -- previously attested only in humans -- with mixed populations of agents developing capacities that go beyond their lowest common denominator.","authors":["Elias Stengel-Eskin","Newton Sander","Carlos Bonetti","Sasha Boguraev","James Bowler","Hale Sirin","Simon Kirby"],"categories":["cs.CL","cs.AI","cs.MA"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-02","first_seen":"2026-09-02","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.01491","pdf_url":"https://arxiv.org/pdf/2609.01491","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","语言演化","LLM交互"],"reason":"研究多智能体LLM之间的语言演化，无人类行为对照，属纯多智能体协作。","model":"deepseek-v4-pro","scored_at":"2026-09-02T13:03:13","error":null,"has_summary":false,"summary":null},{"id":"2609.01056","version":1,"title":"WorldBench: Culturally Grounded Benchmark for Multilingual Agents","zh_title":"WorldBench：面向多语言智能体的文化基准","abstract":"Despite the growing use of LLM-powered agents to solve multi-step tasks in complex environments, existing benchmarks rarely test state preservation, performance across languages, and application to realistic, grounded scenarios. To address these concerns, we present WorldBench: a comprehensive, multilingual benchmark of genuine, persona-grounded everyday workflows, where agents can act in a sandbox via structured actions. WorldBench comprises 1,600 tasks across seven languages and eight cultures, filtered and refined through feedback from human annotators with language- and culture-specific expertise. For evaluation, we extend metrics from previous works and introduce Constrained Task Success (CTS), which combines natural language instructions and testbeds to score task completion, minimal modification, and other complementary metrics through deterministic and LLM-as-a-Judge evaluations. Our experiments show that frontier models reach only 49.2% CTS, with all models demonstrating large gaps between correctness and environment preservation. We thereby show that current agents remain brittle in multilingual, agentic scenarios, especially for long-horizon tasks and under state-preservation constraints","authors":["Leonardo Ranaldi","Sherrie Shen","Jushi Kai","Alexandra Birch"],"categories":["cs.AI","cs.CL"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-09-02","first_seen":"2026-09-02","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.01056","pdf_url":"https://arxiv.org/pdf/2609.01056","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体基准","任务执行","多语言"],"reason":"多智能体在沙盒中执行日常任务，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-02T13:03:12","error":null,"has_summary":false,"summary":null},{"id":"2609.00100","version":1,"title":"Different representation learning objectives recover distinct latent structures from the same psychometric data","zh_title":"不同表征学习目标从相同心理测量数据中恢复出不同的潜在结构","abstract":"Psychometric questionnaires contain rich item-level information, yet it remains unclear whether different representation learning objectives recover the same latent organization. We investigated this question using 757 matched teacher-child pairs from the baseline assessment of the Cyprus ProW preschool trial. Behavioral structure was characterized from child SDQ, ASBI, and CBRS item responses using principal component analysis and clustering, yielding four behavioral phenotypes. A contrastive objective substantially improved teacher-child retrieval relative to PCA-based representations, increasing Top-1 accuracy from 0.13% to 7.27% and Top-10 accuracy from 1.98% to 56.14%. However, contrastive representations preserved behavioral phenotype structure less effectively than PCA-based representations. A multi-task objective jointly optimizing alignment and behavioral prediction partially restored behavioral organization but reduced retrieval performance. These findings indicate that teacher-child correspondence and behavioral phenotypes represent distinct forms of latent organization and demonstrate that the latent structure recovered from linked psychometric data depends on the representation learning objective.","authors":["Cong Cao","Tassos C. Kyriakides","Pambos Vrasidas"],"categories":["cs.AI","cs.LG","stat.ME"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-02","first_seen":"2026-09-02","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.00100","pdf_url":"https://arxiv.org/pdf/2609.00100","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["表征学习","心理测量","对比学习"],"reason":"论文研究表征学习目标对心理测量数据结构的影响，未使用LLM仿真人类被试，属于纯…","model":"deepseek-v4-pro","scored_at":"2026-09-02T13:02:59","error":null,"has_summary":false,"summary":null},{"id":"2609.00180","version":1,"title":"Asymmetries in Spontaneous and Instructed Deception","zh_title":"自发与指示欺骗中的不对称性","abstract":"Large language models sometimes deceive users without being instructed to. However, much of the study on deception in models involves instructed deception. We investigated the relationship between instructed and spontaneous (uninstructed) deception in Llama-3.1-70B-Instruct. We compared these two deception settings through direction geometry, cross-setting classifiers, and cross-setting steering. We found the two deception settings share a component of direction (cosine of approximately 0.5) and an asymmetry in the transfer between settings regarding detection and causation. Spontaneous trained classifiers performed better on instructed data than vice versa, and instructed derived directions performed better at steering spontaneous prompts than vice versa. Likewise the best token position to derive steering vectors from differed from the best token position to train and apply classifiers.","authors":["Josiah Luikham"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-02","first_seen":"2026-09-02","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.00180","pdf_url":"https://arxiv.org/pdf/2609.00180","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["LLM欺骗","模型机制","可解释性"],"reason":"研究LLM欺骗行为的内部机制，无人类被试仿真或对照，属纯模型能力分析","model":"deepseek-v4-pro","scored_at":"2026-09-02T13:03:00","error":null,"has_summary":false,"summary":null},{"id":"2609.00584","version":1,"title":"Socrates went Nuclear: Comparing Interaction Strategies for AI systems in a Learning Context using Brain Sensing","zh_title":"苏格拉底走向核能：在学习情境下使用脑传感比较AI系统的交互策略","abstract":"Does unrestricted AI access bypass the cognitive effort required for learning, or does it streamline knowledge acquisition? This paper reports on a study where we compare three designs for user-AI interaction in a learning context: (1) an unrestricted conversational bot like ChatGPT, (2) a pedagogically constrained bot that guides through hints without giving final answers, which we refer to as the Socratic mode; and (3) a non-conversational adaptive tutoring system that adjusts difficulty in real-time based on the user's cognitive engagement derived from the brain signals. Fifty study participants were tasked with learning about nuclear safety protocols, a domain chosen for its zero-prior knowledge baseline. The participants progressed through an instructional video, a pre-test, an AI-driven assessment phase, which varied in the three conditions, and an immediate post-test. The nature of the questions centered primarily on factual knowledge acquisition, but it still required participants to have a global understanding of the concepts in order to answer the questions correctly. A Muse headband was used to derive the cognitive engagement of all users in all conditions. The unrestricted chatbot produced higher learning gains (delta) than both constrained modes (p < .03, d > 0.80), while the adaptive condition generated significantly higher EEG engagement (p = .018). The cluster analysis of chatbot usage and discussion patterns by users showed that most participants in the unrestricted-mode adopted a direct answer-retrieval strategy, while participants in the Socratic-mode initially attempted to reason through the hints before progressively disengaging. Consequently, this also suggests that the success of the unrestricted AI is not an evidence of deeper learning, but rather a result of the immediate post-test evaluation after the training phase.","authors":["Alexandre Clin Deffarges","Nataliya Kosmyna","Pattie Maes"],"categories":["cs.AI","cs.HC"],"primary_category":"cs.AI","announce_type":"new","date":"2026-09-02","first_seen":"2026-09-02","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.00584","pdf_url":"https://arxiv.org/pdf/2609.00584","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["AI交互","学习效果","脑机接口"],"reason":"研究AI交互策略对学习效果的影响，不涉及用LLM仿真人类被试或与人类数据对照","model":"deepseek-v4-pro","scored_at":"2026-09-02T13:03:07","error":null,"has_summary":false,"summary":null},{"id":"2609.00987","version":1,"title":"On the Human and Computer Alignment of Attribute-Based Music Matches","zh_title":"基于属性的音乐匹配的人机对齐研究","abstract":"Recent advances in generative AI are raising ethical concerns regarding the originality of generated content and the potential replication of training data, with further implications for transparency, attribution, and intellectual property. In music, several computational approaches have been proposed to identify potential replication, using audio-based similarity metrics. Yet, their alignment with human judgments across distinct musical attributes remains underexplored. To address this gap, we conduct a perceptual experiment on music matches, defined as strongly similar musical excerpts. We focus on five musical attributes: melody, harmony, rhythm, voice, and timbre. We design a triplet-based forced-choice task comprising 300 cases, including plagiarism examples, cover songs, and AI-generated music. From this experiment, we introduce the MATCHA (Musical Attribute-based Triplet Comparison with Human Annotations) dataset: a collection of 1105 perceptual assessments of attribute-based music matches from 83 expert participants. Our findings reveal measurable agreement among participants in identifying matches across attributes. We further observe partial alignment between human judgments and computational similarity measures. Overall, this work underscores the importance of domain-specific and perceptually grounded evaluation frameworks for generative AI in creative practice.","authors":["Roser Batlle-Roca","Woosung Choi","Joan Serr\\`a","Fabio Morreale","Wei-Hsiang Liao","Xavier Serra","Emilia G\\'omez","Yuki Mitsufuji"],"categories":["cs.SD","cs.AI"],"primary_category":"cs.SD","announce_type":"cross","date":"2026-09-02","first_seen":"2026-09-02","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.00987","pdf_url":"https://arxiv.org/pdf/2609.00987","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["音乐相似度","感知实验","生成式AI伦理"],"reason":"论文研究音乐相似度的人类感知与计算度量对齐，不涉及LLM仿真人类被试，属于纯感…","model":"deepseek-v4-pro","scored_at":"2026-09-02T13:03:11","error":null,"has_summary":false,"summary":null},{"id":"2609.00414","version":1,"title":"LPG Subsidy Reform, Energy Compensation, and Social Risk in Bolivia: A Machine-Learning Agent-Based Microsimulation","zh_title":"玻利维亚液化石油气补贴改革、能源补偿与社会风险：基于机器学习的智能体微观模拟","abstract":"This study evaluates alternative designs for reforming Bolivia's liquefied petroleum gas subsidy using a machine-learning agent-based microsimulation framework. The analysis harmonizes household survey data, expenditure data, demographic and health information, and monthly hydrocarbon production and commercialization series to simulate fiscal savings, poverty effects, energy substitution, administrative costs, and social risks under multiple reform scenarios. The model compares uncompensated subsidy removal, fixed energy transfers, full compensation for vulnerable LPG users, voucher-based compensation, maternal-child transfers, clean-energy transition kits, and hybrid policy packages. Machine-learning models are used to learn household vulnerability, fuel-use patterns, food insecurity risk, and behavioral propensities that feed into a monthly agent-based simulation. The results show that eliminating the subsidy without compensation generates the largest fiscal savings but increases poverty, extreme poverty, and pressure toward solid-fuel substitution. Full monetary compensation for Q1-Q2 LPG users substantially reduces social harm while preserving significant fiscal savings and dominates an equivalent voucher design under normal market conditions because of lower administrative friction. The most socially robust design combines targeted monetary energy compensation, maternal-child reinforcement, and clean-energy kits for households using solid fuels. The findings support a gradual replacement of the universal LPG subsidy with targeted, administratively lean, and behaviorally informed compensation mechanisms.","authors":["Ricardo Alonzo Fern\\'andez Salguero"],"categories":["econ.EM"],"primary_category":"econ.EM","announce_type":"new","date":"2026-09-02","first_seen":"2026-09-02","revised_at":null,"abs_url":"https://arxiv.org/abs/2609.00414","pdf_url":"https://arxiv.org/pdf/2609.00414","source_feed":"econ.EM","score":0,"bucket":"other","rubric_hits":["C2"],"tags":["智能体模拟","补贴改革","机器学习"],"reason":"基于机器学习的智能体微观模拟，非LLM仿真人类被试，无LLM参与。","model":"deepseek-v4-pro","scored_at":"2026-09-02T13:03:03","error":null,"has_summary":false,"summary":null},{"id":"2606.30085","version":2,"title":"Tastes without distinction: silicon samples and the synthetic construction of tastes","zh_title":"无差别的品味：硅样本与品味的合成建构","abstract":"Large-language models have proven to be remarkable if inconsistent parrots of public attitudes and opinions. The extent to which LLMs are able to produce reasonable approximations of cultural taste remains an open empirical question that becomes more urgent by the day, with market research companies already offering provisional 'synthetic' survey panels and the contamination of standard survey data from LLM-generated responses. In this study, we build on past work on silicon sampling by extending considerations of their ecological, relational, and positional fidelity in the doomain of cultural tastes. We use large-language models from OpenAI, Anthropic, and DeepSeek to produce 554,940 silicon surrogates of survey respondents from the Survey of Public Participation in the Arts (SPPA). We find these silicon surrogates' tastes to be highly stylized facsimiles of human tastes. First, silicon samples are super-omnivorous with a systematic postive-bias for liking. These individual-level bias of silicon samples are not well-explained by the WEIRD-bias often discussed in the literature. Second, the complex relationality in real taste structures is completely distorted among silicon samples. Third, very little of the known cultural alignment between tastes and social space are preserved. Silicon samples juvenilize age-taste associations, resurrect anachronistic class-taste associations, and caricaturize gender- and race-taste associations. Key words: AI, taste, consumption, culture, silicon sampling, meta-analysis.","authors":["Xiangyu Ma","Mengmi Zhang","Shannon Ang","Minne Chen"],"categories":["cs.CL","econ.GN","q-fin.EC"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-01","first_seen":"2026-06-29","revised_at":"2026-09-01","abs_url":"https://arxiv.org/abs/2606.30085","pdf_url":"https://arxiv.org/pdf/2606.30085","source_feed":"cs.CL","score":10,"bucket":"selected","rubric_hits":["A1","A2","A3","B1","B2","B4"],"tags":["硅采样","文化品味仿真","算法保真度"],"reason":"用LLM生成硅样本模拟文化品味调查，并与真实SPPA数据对照，评估仿真偏差与失…","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:03:19","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-01","rank":1,"question":"大语言模型生成的硅样本能否忠实再现人类文化品味及其与社会空间的关联？","design":"以2012年美国公众参与艺术调查（SPPA）为参照，将每个受访者的人口统计特征转化为角色提示，输入OpenAI、Anthropic和DeepSeek的大语言模型，为每个人类受访者生成多个硅样本（共554,940个），比较硅样本与人类样本在品味分布、品味结构关系及品味与社会空间关联上的保真度。","baseline":"2012年SPPA调查中的人类受访者数据，并使用SPPA的自举重采样作为人类数据内部一致性的基准。","findings":"硅样本表现出超级杂食性和系统性正向偏好，且这种偏差不能用文献中常讨论的WEIRD偏差解释；真实品味结构中的复杂关系在硅样本中被完全扭曲，品味与社会空间的已知关联几乎未被保留，年龄-品味关联被年轻化，阶级-品味关联被复活为过时模式，性别和种族-品味关联被漫画化。","reliability":"论文指出硅样本的个体层面偏差（如超级杂食性和正向偏好）不能由WEIRD偏差解释，且其关系结构和与社会空间的关联严重失真，表明在文化品味领域硅采样保真度很低，但未明确讨论具体失效条件或局限性。","relevance":"该研究直接评估LLM仿真人类文化品味调查的可靠性，并与真实SPPA数据对照，发现系统性偏差和结构失真，对关注LLM仿真在社会科学中有效性的研究者具有重要参考价值，值得阅读原文以了解具体偏差模式和评估框架。","inspiration":"借鉴其将人口统计特征转化为提示、生成大量硅样本并与真实调查数据对照的仿真设计，以及从生态、关系和位置保真度三个维度评估偏差的方法。｜可迁移到消费者偏好调查、市场细分或文化消费的经济学研究中，例如用LLM模拟不同人口群体的品牌偏好或娱乐消费选择。｜设计：以某消费者支出调查（如美国消费者支出调查CE）为基准，提取受访者人口特征生成提示，用多个LLM生成硅样本，让其回答关于品牌偏好或娱乐活动参与的问题，比较硅样本与真实人类在偏好分布、偏好结构及与社会经济地位关联上的差异，并检验硅样本是否复现已知的消费分层模式。"}},{"id":"2608.03044","version":2,"title":"Emulate or Estimate? The Divergent Strengths of Base and Post-Trained Language Models for Opinion Simulation","zh_title":"仿真还是估计？基础与后训练语言模型在意见模拟中的不同优势","abstract":"Large language models are increasingly used to simulate human opinions, but prior work reports conflicting results: some studies find promising alignment with human survey data, while others find persona collapse and weak demographic sensitivity. We propose that much of this conflict stems from conflating two distinct tasks. We call the first task emulation, in which models generate individual responses that aggregate into a population distribution. We call the second task estimation, in which models directly predict the population distribution. Evaluating six matched base and post-trained models on the Pew American Trends Panel, we find that base models are the stronger emulators: they produce response distributions closer to human ground truth and better preserve demographic structure. Post-trained models are generally the stronger estimators, producing more accurate distributional predictions when asked directly. We argue that model selection for human simulation should be guided by whether the task requires generating text or predicting distributions.","authors":["Seth Grief-Albert","Jessica Bo","Difan Jiao","Ashton Anderson"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-01","first_seen":"2026-08-05","revised_at":"2026-09-01","abs_url":"https://arxiv.org/abs/2608.03044","pdf_url":"https://arxiv.org/pdf/2608.03044","source_feed":"cs.CL","score":10,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","意见模拟","算法保真度"],"reason":"直接研究LLM仿真人类意见，区分仿真与估计任务，使用真实调查数据对照，并指出模…","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:03:20","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-01","rank":2,"question":"在意见仿真中，基础模型与后训练模型在仿真（emulation）与估计（estimation）两种任务上表现有何差异？","design":"使用三对匹配的基础与后训练模型（Qwen3-14B、Olmo-3-7B、Olmo-3-32B）在Pew美国趋势面板59个经济意见问题上进行仿真：仿真任务通过开放式生成个体回答并聚合，估计任务直接预测总体分布；施加七种人口学条件（无条件、民主党、共和党、非常自由派、非常保守派、高收入、新教徒），以总变差距离和Wasserstein距离衡量分布保真度。","baseline":"Pew Research Center American Trends Panel Wave 54 的真实人类调查数据，包含59个四选项经济意见问题及人口学子群体分布。","findings":"基础模型在仿真任务上始终优于后训练模型，生成的分布更接近人类真实分布且更好地保留人口学结构；后训练模型在直接估计总体分布时通常更准确。","reliability":"论文未讨论","relevance":"直接研究LLM仿真人类意见，区分仿真与估计任务，使用真实调查数据对照，并指出模型选择应基于任务类型，对关注LLM仿真可靠性与偏差的研究者具有重要参考价值。","inspiration":"借鉴其区分仿真与估计任务、使用开放式生成避免位置偏差、并用真实调查数据做基准对照的方法｜可迁移到经济预期形成或消费者信心调查的仿真，如模拟不同收入群体对通胀预期的分布｜用基础模型模拟个体受访者回答开放式经济预期问题，聚合后与密歇根消费者调查的真实分布对比，同时用后训练模型直接估计分布，比较两种范式的准确性。"}},{"id":"2608.29455","version":1,"title":"Item-Mean Surrogates: Why Richer Persona Data Fail to Improve LLMs as Human Surrogates","zh_title":"项目均值替代：为何更丰富的人物数据未能提升LLM作为人类替代品的表现","abstract":"LLMs are increasingly used as human surrogates, often on the premise that richer persona data could make them substitutes or exploratory tools for specific individuals. We test this premise across four datasets covering more than 400,000 participants and more than 6,000 survey items and experimental outcomes. LLMs perform well at the aggregate level: their average responses closely align with average human responses to the same items. But this success largely reflects predicting each item's average human response. Once each item's human mean is removed, LLM predictions explain only 3.05% of the remaining respondent-specific variation, far below the 53.6% human test-retest benchmark. Richer personas, model variants, and fine-tuning do not close this gap. In variance analyses, once item means are removed, the reliable remaining signal is person-by-item. It captures how a respondent departs from the mean on a particular item and is about 8.9x larger than the stable person effect. Persona data encode the respondent, but not this item-specific deviation. LLM responses also compress human response distributions, using less spread, fewer response categories, and distorted distributional shapes. We call this pattern item-mean surrogacy. Current LLM surrogates can approximate item averages, but not the distributions or respondent-specific deviations needed to replace individual humans. We propose four empirical tests for LLM-based human-surrogate claims.","authors":["Daehwan Ahn","Chengfeng Mao","Dokyun Lee"],"categories":["cs.CL","cs.CY","cs.LG"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-01","first_seen":"2026-09-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.29455","pdf_url":"https://arxiv.org/pdf/2608.29455","source_feed":"cs.CL","score":10,"bucket":"selected","rubric_hits":["A1","A2","A3","A4","B1","B2","B3","B4"],"tags":["LLM仿真","人类替代","算法保真度"],"reason":"直接评估LLM作为人类替代品的可靠性，使用大规模人类数据对照，发现仅能预测项目…","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:02:35","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-01","rank":4,"question":"LLM 作为人类替代品时，更丰富的人物角色数据能否提高其对个体层面回答的预测能力？","design":"使用四个数据集（Megastudy、Survey、SocSci210、ANES），覆盖超40万参与者和6000多个调查项目，用多种LLM模型和提示方法（包括丰富人物角色、微调）生成对相同项目的回答，并与真实人类回答对比，测量项目均值、分布和个体偏差的恢复程度。","baseline":"真实人类数据：Megastudy 和 Survey 数据集包含同一批参与者的 LLM 人物角色和真实回答，并提供人类重测信度（53.6%）作为个体预测上限；SocSci210 和 ANES 提供不同领域的人类回答。","findings":"LLM 在聚合层面表现良好，平均回答与人类平均回答高度一致，但去除项目均值后，LLM 预测仅解释 3.05% 的个体变异，远低于人类重测基准 53.6%。丰富人物角色、模型变体和微调均未缩小差距；方差分析显示稳定的人-项目交互效应是主要信号（44.0%），而人物角色数据仅编码稳定的人主效应（4.9%），无法捕捉项目特定偏差。","reliability":"论文指出 LLM 替代品仅能近似项目均值，无法恢复分布或个体偏差，因此不能替代个体人类；其验证通常停留在聚合层面，存在生态谬误风险。论文未讨论其他失效条件，但强调需要个体级保真度的应用场景（如个性化推荐、政策评估）会失效。","relevance":"该研究直接评估 LLM 作为人类替代品的可靠性，使用大规模真实人类数据对照，发现仅能预测项目均值而无法捕捉个体变异，对关注仿真有效性和偏差的研究者极具参考价值，值得精读原文。","inspiration":"借鉴其将预测误差分解为项目效应、人主效应和人-项目交互效应的方法，并利用人类重测信度区分稳定信号与随机误差，可迁移到经济金融领域的个体决策预测（如消费者跨期选择、投资者风险偏好）。｜可应用于信贷审批中的个体违约风险预测或资产定价实验中的个体风险偏好测量。｜设计：以真实信贷申请人或实验参与者为被试，构建 LLM 人物角色并让其预测个体在特定金融决策任务中的选择（如是否接受高风险贷款），处理为不同人物角色丰富度（仅人口统计 vs. 完整问卷），结果变量为个体选择与项目均值的偏差，对照真实人类选择数据，并计算人类重测信度作为上限。"}},{"id":"2608.30033","version":1,"title":"\"Act Like a 5th Grader\" is Not Enough: Bounding Knowledge in LLM-Based User Simulators","zh_title":"“像五年级学生一样行动”还不够：在基于LLM的用户模拟器中界定知识","abstract":"Large language models (LLMs) are increasingly used to simulate human behavior but frequently fail to exhibit realistic cognitive constraints, suffering from a \"superhuman bias.\" Using a dataset of over 71,000 reading comprehension responses from 2,359 primary-school students (grades 4--6), we demonstrate that standard persona prompting yields near-perfect, deterministic performance, failing to capture the natural variance of developing readers. To address this, we introduce the Cognitively Bounded User Simulator (CBUS), an architectural framework that explicitly models the restricted working memory of young readers through an episodic bottleneck. Within this framework, we formalize two distinct test-taking strategies to emulate different reading behaviors. Our evaluation shows that explicitly modeling cognitive bounds significantly narrows the simulation gap across multiple LLM backbones, demonstrating that enforcing architectural constraints is more effective for high-fidelity simulation than simply scaling raw model capabilities.","authors":["Krisztian Balog","Arild Michel Bakken"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-01","first_seen":"2026-09-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.30033","pdf_url":"https://arxiv.org/pdf/2608.30033","source_feed":"cs.CL","score":10,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","认知边界","人类数据对照"],"reason":"用LLM模拟学生阅读理解，有真实学生数据对照，并解决仿真偏差","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:02:37","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-01","rank":5,"question":"如何通过显式建模认知边界（工作记忆容量限制）来提高LLM模拟小学生阅读理解行为的保真度？","design":"使用LLM模拟4-6年级小学生，通过两种方式施加处理：标准角色提示（如“扮演五年级学生”）和提出的CBUS框架（两阶段推理管道，编码阶段限制提取的文本命题数量，执行阶段仅基于受限记忆痕迹回答问题）。结果变量为学生在阅读理解题目上的回答（客观题，包括单选、判断、多选），并与真实学生数据进行对比。","baseline":"来自挪威教育系统的2,359名4-6年级学生对156篇文本和750道阅读理解题目的71,000多条真实回答。","findings":"标准角色提示导致LLM表现出超人类偏差，几乎完美且确定性地回答问题，无法捕捉真实学生的发展性差异。CBUS框架通过显式建模工作记忆瓶颈，显著缩小了模拟与真实学生之间的差距，且该效果在多个LLM骨干模型上一致。","reliability":"论文未明确讨论失效条件，但指出标准角色提示在受控环境中不足，且CBUS框架基于工作记忆容量限制的通用认知理论，可能不适用于其他认知过程或更复杂任务。","relevance":"该研究直接针对LLM仿真中的超人类偏差问题，提供了在受控环境中量化偏差和通过架构约束改进仿真的方法，对关注仿真可靠性和偏差的研究者具有重要参考价值。","inspiration":"值得借鉴的是通过架构约束（而非仅提示）来模拟认知限制，并利用大规模真实数据作为对照基准。｜可以迁移到经济金融中需要模拟有限理性或信息处理约束的场景，如消费者对复杂金融产品的选择、投资者对信息的有限关注。｜设计一个实验：用LLM模拟投资者阅读财报后做出投资决策，处理组为施加工作记忆瓶颈的CBUS框架，对照组为标准角色提示，结果变量为投资选择，对照真实投资者在类似实验中的数据。"}},{"id":"2608.28615","version":1,"title":"Distributional Validity and Calibration of a Korean Synthetic Persona Panel for Digital and AI Service Use: A Secondary-Data Validation Against the Korea Media Panel Survey","zh_title":"韩国合成人面板在数字与AI服务使用上的分布效度与校准：基于韩国媒体面板调查的二手数据验证","abstract":"Synthetic personas based on large language models (LLMs) are increasingly proposed as substitutes for human survey respondents, yet systematic validation outside English-speaking contexts remains scarce. This secondary-data study evaluates how well a Korean synthetic persona panel (NVIDIA Nemotron-Personas-Korea), conditioned into Gemini 3.5 Flash (primary) and EXAONE (comparison), reproduces digital and AI service-use distributions from the KISDI Korea Media Panel Survey. Sex-and-age-stratified panels of about 8,000 personas per model answered the survey's own items - eight service-use indicators and eight innovativeness and acceptance constructs - and were compared against weighted survey estimates. The overall mean absolute error (MAE; RQ1) was 15-19 percentage points (pp), with binary item-mean correlations of 0.69-0.90. Segment error (RQ2) across five demographic axes was 15-19 pp, with between-group gaps up to 52.4/36.2 pp (Gemini/EXAONE). Errors followed model-specific signatures: an age stereotype with low anchoring (Gemini) versus an acquiescence-consistent level bias (EXAONE). Reference-year analysis was consistent with temporal misalignment driving most generative-AI overestimation, whereas short-form underestimation was framing-sensitive. Holdout calibration on 30% of the real data (RQ3) roughly halved sex-by-age cell MAE (18.9->8.6, 15.9->6.7 pp) - yet direct estimation from the same real subsample was far more accurate (3.6 pp), and the correction did not transfer across time. The calibrated panel retained an advantage only under extremely scarce real data (about 100 responses) and, for one model, for unobserved segments. Persona-narrative conditioning beat demographic-only conditioning, but neither surpassed simple real-data baselines. Synthetic panels are thus not survey substitutes; their value is diagnostic, with operational use confined to settings lacking real data.","authors":["Howard Kim","Keun Tae Cho"],"categories":["cs.CY","cs.CL"],"primary_category":"cs.CY","announce_type":"cross","date":"2026-09-01","first_seen":"2026-09-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.28615","pdf_url":"https://arxiv.org/pdf/2608.28615","source_feed":"cs.CL","score":10,"bucket":"selected","rubric_hits":["A1","A2","A3","B1","B2","B3","B4"],"tags":["LLM仿真","合成人面板","外部效度"],"reason":"用LLM合成韩国人面板，复现数字服务使用分布，并与真实调查数据对照，评估校准与…","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:02:31","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-01","rank":3,"question":"韩国合成人面板在多大程度上能复现真实调查中的数字与AI服务使用分布，以及小样本校准能否改善其准确性？","design":"使用 NVIDIA Nemotron-Personas-Korea 合成人数据，按性别和年龄分层抽取约8000个角色，分别通过 Gemini 3.5 Flash 和 EXAONE 模型生成对韩国媒体面板调查问卷的回答，测量八个数字服务使用指标和八个创新性与技术接受度构念，并与加权调查估计值比较。","baseline":"KISDI 韩国媒体面板调查 2024 年波次（约8693人）的加权估计值，以及 2023 和 2025 年波次用于时间转移检验。","findings":"合成面板的总体平均绝对误差为15-19个百分点，二元指标均值相关系数0.69-0.90；误差呈现模型特异性偏差（Gemini 年龄刻板印象、EXAONE 默许偏差），且校准虽能减半误差，但直接使用少量真实数据更准确，校准无法跨时间转移。","reliability":"论文明确指出合成面板不能替代调查，其价值仅在于诊断，且校准仅在真实数据极度稀缺（约100份）或存在未观测群体时才有优势；误差受时间错位、题项框架和响应风格影响，校准无法跨波次转移。","relevance":"该研究直接检验了LLM合成样本在非英语情境下的分布效度，并与真实调查数据严格对照，对关注仿真可靠性和偏差条件的研究者具有重要参考价值，值得精读原文以了解具体误差结构和校准方法。","inspiration":"借鉴其分层抽样生成合成样本并与真实调查对照的验证框架，以及基于偏差签名的小样本校准方法｜可迁移到消费者金融行为调查（如数字支付采用、金融科技接受度）或政策评估中的态度测量｜以合成人面板模拟不同人口群体的金融决策，施加政策信息处理，测量采用意愿，并与真实家庭金融调查数据（如中国家庭金融调查）对照，检验分布一致性和校准效果。"}},{"id":"2608.30522","version":1,"title":"Tariff Threats, Macroeconomic Expectations, and Policy Communication Strategies: Experiments Based on a Multi-Agent System","zh_title":"关税威胁、宏观经济预期与政策沟通策略：基于多智能体系统的实验","abstract":"Tariff threats can move household beliefs before policy is enacted, yet their rapidly changing language is difficult to study with conventional surveys. We build a multi-agent system that turns 300 households from the Michigan Surveys of Consumers into persistent large-language-model agents exposed to social-media information over several simulated months. Calibrated agents reproduce some distributional and demographic patterns in human survey data collected after the announcement of Liberation Day tariffs. Simulated experiments indicate that immediacy, rate salience, semantic progression, message complexity, narrative, and sender identity jointly shape inflation and unemployment expectations and their dispersion. Open-ended responses trace these effects to attention, ambiguity, credibility, and causal narratives. A second experiment finds that central-bank explanations can coordinate beliefs, although their effects on average expectations depend on message content. The framework supports disciplined exploration of policy communication, subject to human validation rather than as a substitute for it.","authors":["Jianhao Lin","Lexuan Sun","Yixin Yan"],"categories":["econ.GN","q-fin.EC"],"primary_category":"econ.GN","announce_type":"new","date":"2026-09-01","first_seen":"2026-09-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.30522","pdf_url":"https://arxiv.org/pdf/2608.30522","source_feed":"econ.GN","score":10,"bucket":"selected","rubric_hits":["A1","A3","B1","B2","B3","B4"],"tags":["LLM仿真","宏观经济预期","政策沟通"],"reason":"用LLM代理300个家庭，复现关税冲击后的宏观预期，并与密歇根调查真实数据对照…","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:02:39","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-01","rank":6,"question":"关税威胁的时机、税率显著性、语义连贯性、信息复杂度、叙事和发送者身份如何共同影响家庭通胀与失业预期的水平及离散度？高影响威胁后，央行何种沟通最能协调信念？","design":"构建多智能体系统（MAS），将密歇根消费者调查（MSC）中的300个家庭转化为基于大语言模型的持久化智能体，赋予其人口特征和初始预期，在模拟的数月内暴露于社交媒体信息，并施加不同的关税威胁情景处理，测量其点预测、主观概率分布和开放式解释。","baseline":"密歇根消费者调查（MSC）在2025年4月“解放日”关税公告后收集的真实家庭通胀和失业预期数据，用于校准和验证模拟结果。","findings":"校准后的智能体在预期分布、人口统计差异和可区分性上接近人类数据；模拟实验表明，关税威胁的时机、税率显著性、语义进展、信息复杂度、叙事和发送者身份共同影响预期水平和离散度，开放式回答显示这些效应通过注意力、模糊性、可信度和因果叙事起作用。央行解释可以协调信念，但效果取决于信息内容。","reliability":"论文承认校准智能体不能替代人类，不能将未观察到的信息处理效应外推为人类因果效应；验证必须与语言模型的预期用途相关，行为相似性不足以证明从已验证场景到未观察场景的因果迁移。","relevance":"该研究直接命中研究者关注的核心：用LLM代理复现真实调查数据，并评估仿真可靠性，且涉及政策沟通和预期形成，值得精读原文以了解其校准方法和局限性讨论。","inspiration":"借鉴其将真实调查个体转化为持久化LLM智能体、施加文本处理并测量预期分布和开放式解释的设计，以及用真实调查数据校准和验证仿真的做法。｜可迁移到政策公告的预期形成研究，如央行沟通对通胀预期的影响、关税或财政政策冲击下的家庭预期调整。｜以真实家庭调查（如密歇根调查或纽约联储消费者预期调查）的受访者为被试，将其特征和初始预期编码为LLM智能体，施加不同措辞的央行声明或关税公告作为处理，测量通胀和失业预期的点预测、概率分布和开放式理由，并用同期真实调查数据作为对照基准进行校准和验证。"}},{"id":"2608.26849","version":2,"title":"LiveSim: Simulating Environment-Shaped Users in Multi-Agent Live-Stream Ecosystems","zh_title":"LiveSim：在多智能体直播生态系统中模拟受环境塑造的用户","abstract":"User behavior simulation with large language models~(LLMs) is increasingly used to support multi-agent ecosystem simulation. Existing simulators typically rely on static user profiles inferred from historical observations, which become inadequate in socially intensive environments such as live streaming where interaction dynamics continuously reshape user behavior. We propose \\textbf{LiveSim}, an LLM-based framework for live-stream ecosystem simulation. It represents users as editable behavioral hypotheses and progressively refines them through trajectory-grounded interactions, where discrepancies between simulated and observed trajectories reveal missing environmental shaping effects. These signals are further extracted as transferable environment-behavior patterns and accumulated in a collective behavioral memory to improve user-level behavioral fidelity and support ecosystem-level simulation. Experiments on real-world live-stream risk-control data validate the effectiveness of LiveSim in improving user-level behavioral fidelity and enabling ecosystem-level analysis of risk evolution and platform intervention effects.","authors":["Jiaqi Xu","Yiran Qiao","Jing Chen","Qiwei Zhong","Xiang Ao","Xueqi Cheng"],"categories":["cs.AI","cs.CY","cs.MA"],"primary_category":"cs.AI","announce_type":"replace-cross","date":"2026-09-01","first_seen":"2026-08-28","revised_at":"2026-09-01","abs_url":"https://arxiv.org/abs/2608.26849","pdf_url":"https://arxiv.org/pdf/2608.26849","source_feed":"cs.CY","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM用户仿真","直播生态","行为保真度"],"reason":"用LLM仿真直播用户行为，并与真实数据对照，评估行为保真度，直接相关。","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:03:20","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-01","rank":8,"question":"如何在直播这类社交密集环境中，用LLM模拟用户行为，并动态修正用户模型以反映环境对行为的塑造作用？","design":"提出LiveSim框架，用LLM扮演直播用户，将用户表示为可编辑的行为假设，通过模拟轨迹与真实轨迹的差异迭代修正假设，提取环境-行为模式存入集体记忆，再用于多智能体直播生态模拟。","baseline":"使用某大型直播平台真实风险控制数据中的用户行为轨迹作为对照。","findings":"LiveSim在用户级行为保真度上显著提升，并能支持生态级风险演化和平台干预效果分析。","reliability":"论文未讨论","relevance":"直接相关，用LLM仿真直播用户行为并与真实数据对照，评估行为保真度，属于人类仿真实验研究，值得精读原文。","inspiration":"借鉴其通过模拟-真实轨迹差异迭代修正行为假设的方法，可迁移到消费者在直播带货中的冲动购买行为研究。｜设计如下：用LLM模拟消费者，处理为不同主播话术（如限时折扣、从众压力），结果变量为购买决策，对照真实直播销售数据中的用户行为轨迹。"}},{"id":"2608.29803","version":1,"title":"Do LLMs Change Their Minds Like Humans? Diagnosing Human--LLM Divergence in Single-Turn Persuasion Judgments","zh_title":"LLM会像人类一样改变想法吗？诊断单轮说服判断中的人机分歧","abstract":"Large language models (LLMs) are increasingly deployed as proxies for human participants in social simulations, yet whether they update their beliefs in response to persuasive arguments, as humans do, remains poorly understood. We conduct a systematic comparison using a naturally occurring online persuasion corpus in which original posters explicitly verify whether a reply changed their view. Our results show that LLMs achieve only slight agreement with humans (Cohen's kappa ranging from 0.079 to 0.178). Content-level analyses show that humans and LLMs agree on the strongest persuasion cues but diverge on finer ones: humans are more swayed by novel content and assertive language, whereas LLMs favor topical similarity and surface-level formatting. At the level of persuasion strategy, LLMs underweight emotional appeals and overweight credibility signals relative to humans, while the type of proposition under debate exerts no measurable effect on the degree of divergence. Furthermore, switching from first-person role-playing to third-person observation shifts all models toward greater resistance to persuasion, with the effect varying across persuasion strategies and textual features. These findings highlight the risk of treating LLM judgments as faithful proxies for human belief updating and point to structural differences in how LLMs and humans process persuasive discourse. Our code is available at https://github.com/tsinghua-fib-lab/LLM-belief-update-cmv.","authors":["Lin Chen","Yitong Chen","Yong Li"],"categories":["cs.CY","cs.CL"],"primary_category":"cs.CY","announce_type":"cross","date":"2026-09-01","first_seen":"2026-09-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.29803","pdf_url":"https://arxiv.org/pdf/2608.29803","source_feed":"cs.CL","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","信念更新","人机对比"],"reason":"直接比较LLM与人类在说服中的信念更新，使用真实人类数据对照，并指出LLM作为…","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:02:37","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-01","rank":12,"question":"LLM在单轮说服中的信念更新判断是否与人类一致，哪些内容特征（命题类型、说服策略、文本特征）和视角框架（第一人称角色扮演 vs. 第三人称观察）调节这种分歧？","design":"使用ChangeMyView语料库中真实的说服对话，构建匹配的回复对（同一原帖下，一个被标记为说服成功，一个未成功），让8个LLM独立判断每个回复是否会改变原帖作者的观点，并分析判断结果与人类标签的一致性及影响因素。","baseline":"ChangeMyView语料库中原始发帖人明确标记的delta（观点改变）标签，作为人类信念更新的真实基准。","findings":"LLM与人类的一致性很低（Cohen's kappa 0.079-0.178），且LLM在文本特征和说服策略上与人类存在系统性分歧：人类更受新颖内容和果断语言影响，LLM更偏好主题相似性和表面格式；LLM低估情感诉求、高估可信度信号，而命题类型对分歧无显著影响。从第一人称角色扮演切换到第三人称观察会使所有模型更抗拒说服，且效应因策略和文本特征而异。","reliability":"论文指出LLM判断不能作为人类信念更新的忠实代理，存在结构性差异；但未明确讨论失效条件，仅强调在需要人类推理的社会模拟中直接使用LLM有风险。","relevance":"该研究直接比较LLM与人类在说服场景下的信念更新，使用真实人类数据作为基准，并揭示了系统性偏差，对关注LLM仿真可靠性和偏差的研究者具有重要参考价值，值得阅读原文。","inspiration":"借鉴其匹配对设计和多维度内容标注方法，可系统诊断LLM与人类在决策中的分歧来源。｜可迁移到政策沟通与预期形成场景，如央行沟通对市场预期的影响。｜以LLM作为投资者被试，呈现央行声明或新闻，测量其预期更新，并与真实市场调查数据（如密歇根消费者信心调查）对照，分析LLM是否高估可信度信号或低估情感因素。"}},{"id":"2608.28668","version":1,"title":"Reference-Distribution Dependence in LLM-Based Synthetic Persona Data: Diagnosis and Post Hoc Adjustment of Demographic Distributions","zh_title":"基于LLM的合成人数据中的参考分布依赖：人口统计分布的诊断与事后调整","abstract":"We diagnose how closely the demographic distributions in LLM-based synthetic persona data match external reference distributions. For the three variables examined, we show that most of the observed error is attributable to the choice of reference rather than to the generator. Using total variation distance (TVD), we compare the sex x age group x province joint distribution of 1,000,000 records from Nemotron-Personas-Korea (NPK) with Korean official statistics. Against resident-registration figures for April 2026, the time of use, the bias bound, defined as the largest possible difference in the share of any subgroup formed from the three variables, is 1.81 percentage points. This is comparable to the margin of error of a survey of roughly 2,900 respondents. This value is not a fixed property of the data. Matching the reference period and series to the generating reference identified here, the 2024 register-based census restricted to Korean nationals, lowers it to 0.56 percentage points. Over the 15 months between the best-matching month (January 2025) and the time of use, the resident-registration population structure itself moves more than twice the distance of NPK's minimum error. Raking and cell post-stratification, the two weighting schemes used in Korean survey practice, remove most of the reference-period dependence at a variance inflation of about 0.2% in both cases. After raking against the generating reference, the residual joint discrepancy lies at, and marginally above, the upper bound of what a perfect generator would produce when realizing 1,000,000 records (97.6th percentile of the Monte Carlo distribution). We recommend treating synthetic persona data as auxiliary material for small-scale survey design rather than as a substitute for survey data, and re-running both diagnosis and adjustment against official statistics current at the time of use.","authors":["Eunjeong Song","Sehee Hong"],"categories":["cs.CY","stat.AP"],"primary_category":"cs.CY","announce_type":"new","date":"2026-09-01","first_seen":"2026-09-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.28668","pdf_url":"https://arxiv.org/pdf/2608.28668","source_feed":"cs.CY","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","人口统计偏差","事后加权调整"],"reason":"用LLM生成合成人数据，与真实人口统计对照，诊断偏差并调整，直接相关。","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:02:31","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-01","rank":9,"question":"LLM生成的合成人数据在人口统计分布上与外部参考分布有多接近，观察到的偏差在多大程度上归因于参考分布的选择而非生成器本身？","design":"使用Nemotron-Personas-Korea（NPK）数据集的100万条记录，比较其性别×年龄组×省份联合分布与韩国官方统计的差异，采用总变差距离（TVD）度量，并通过变化参考分布的时期和序列来诊断偏差来源，最后应用raking和单元后分层两种加权方法进行事后调整。","baseline":"韩国官方统计，包括居民登记数据（2026年4月）和2024年基于登记的人口普查（仅限韩国国民）。","findings":"在2026年4月使用时，NPK与居民登记数据的偏差上限为1.81个百分点，相当于约2900名受访者调查的误差边际；当与生成参考（2024年登记普查）匹配时，偏差降至0.56个百分点。大部分观察到的误差归因于参考分布的选择而非生成器，且raking和单元后分层能消除大部分参考期依赖性，方差膨胀仅约0.2%。","reliability":"论文建议将合成人数据视为小规模调查设计的辅助材料，而非调查数据的替代品，并强调应在使用时重新对当前官方统计进行诊断和调整；同时指出合成数据冻结了生成时的人口结构，随时间推移与官方统计的一致性自然下降。","relevance":"该研究直接评估LLM合成人数据与真实人口统计的偏差，并提供了诊断和调整方法，对关注LLM仿真可靠性和偏差的研究者具有重要参考价值，值得阅读原文以了解具体方法和发现。","inspiration":"借鉴其通过变化参考分布来区分生成器偏差与参考选择偏差的诊断方法，以及使用TVD和偏差上限将分布差异转化为调查误差边际的做法。｜可迁移到经济金融领域中使用合成数据模拟消费者或投资者行为的研究，例如评估政策变化对特定人群的影响或测试金融产品在不同人口群体中的接受度。｜设计：使用LLM生成具有特定人口特征（如年龄、收入、地区）的合成个体，模拟他们对某项经济政策（如税收调整）的反应，结果变量为支持率或行为变化，并与真实调查数据（如韩国劳动力面板或家庭收入支出调查）进行对照，通过变化参考分布和事后加权来评估仿真的可靠性。"}},{"id":"2608.29266","version":1,"title":"Measurement Validity in LLM Cultural Alignment","zh_title":"大语言模型文化对齐中的测量效度","abstract":"Researchers increasingly treat LLM survey responses as a proxy for human cultural values. This includes projecting model outputs onto instruments like the Inglehart-Welzel Cultural Map and drawing conclusions about which cultures a model resembles. While a model's answer to a value-laden questions may be interpreted as a cultural signal, it also carries sampling noise and, can be quite sensitive to question framing. In this paper, we separate survey responses, sampling noise and question framing for multiple LLMs. We decompose response variance from these models into variation across random seeds, prompt rewordings. We employ noise-to-signal ratio (NSR) to test whether a model's apparent cultural position is distinguishable from noise. When applied across a dozen models from four geographic origins, calibrated against 88 Integrated Values Survey countries, the answer is often no. NSR exceeds 1.0 on 49 of 117 valid model-question pairs (42%), reaching 5.56 in the worst case. Two models even refuse to answer sufficient number of survey questions outright. Our results corroborate previous findings that LLMs cluster toward Western, English-speaking cultural positions. However, what does not hold up in this study is the precision with which anyone can currently interpret a specific model's coordinates: prompt tone alone can shift a model by 2.4 map units, comparable to the distance between actual countries in the Inglehart-Welzel Cultural Map. These findings suggest that cultural attribution from LLM survey responses requires establishing the reliability of the underlying measurements before interpreting model coordinates as evidence of cultural representation.","authors":["An Duy Nguyen","Muhammad Aurangzeb Ahmad"],"categories":["physics.soc-ph","cs.AI"],"primary_category":"physics.soc-ph","announce_type":"new","date":"2026-09-01","first_seen":"2026-09-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.29266","pdf_url":"https://arxiv.org/pdf/2608.29266","source_feed":"physics.soc-ph","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","文化价值观","测量信度"],"reason":"用LLM回答调查问题模拟文化价值观，并与88国真实数据对照，评估测量信度与偏差。","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:02:33","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-01","rank":10,"question":"LLM 对文化价值观调查的回答在多大程度上反映了真实的文化信号，而非采样噪声和提问措辞的敏感性？","design":"该研究并非传统意义上的仿真实验，而是对 LLM 作为测量工具的信度检验。作者选取 12 个来自不同地域的 LLM，让它们回答 Inglehart-Welzel 文化地图的 10 个价值观问题，并通过变换随机种子和提示语气/人称来分解回答方差，计算噪声信号比（NSR）以评估测量有效性。","baseline":"88 个国家的 Integrated Values Survey（IVS）数据，用于构建 Inglehart-Welzel 文化地图的参照系。","findings":"LLM 的文化位置普遍偏向西方英语国家，但测量信度很差：42% 的模型-问题对 NSR 超过 1.0，提示语气变化可使模型坐标移动 2.4 个地图单位，相当于真实国家间的距离。部分模型拒绝回答敏感问题，导致无法投影到文化地图上。","reliability":"论文明确指出，在未建立测量信度之前，不能将 LLM 的调查回答解读为文化表征。局限包括：仅使用 Inglehart-Welzel 框架，问题数量有限；未覆盖所有可能的提示变体；模型版本和采样参数可能影响结果。","relevance":"该研究直接回应了用 LLM 模拟人类价值观的可靠性问题，与您关注的人类仿真实验的效度批判高度相关，值得精读原文以了解其噪声分解方法和 NSR 指标。","inspiration":"借鉴其将回答方差分解为随机种子、提示措辞和模型间差异的方法，用于检验 LLM 在经济调查中的测量信度。｜可迁移到消费者信心调查、通胀预期、风险偏好等经济态度的测量，评估 LLM 作为被试的可靠性。｜以 GPT-4 等模型为被试，施加不同措辞的问卷版本，测量其通胀预期或风险偏好，并与密歇根大学消费者调查或实验经济学中的真实人类数据对照，计算 NSR 以判断 LLM 回答是否超出噪声。"}},{"id":"2608.29535","version":1,"title":"Integrating adaptive human behavior into epidemic models with large language models","zh_title":"用大语言模型将自适应人类行为整合进流行病模型","abstract":"Infectious disease transmission is shaped by patterns of human interaction, which adapt as epidemic conditions change. Capturing these context-dependent behaviors remains a fundamental challenge for epidemic models. Here, we recast this challenge by using large language models (LLMs) to represent adaptive human behavior within mechanistic epidemic models. We operationalize this idea through Generative Adaptive Behavioral Layer for Epidemics (GABLE), which adapts LLMs to infer behavioral responses to epidemic and policy conditions and translates them into age-structured contact matrices coupled to a mechanistic epidemic model. Applied to COVID-19 in France, GABLE reproduced responses in population mixing and age-specific contact structures that remained epidemiologically informative. In short-term forecasting, LLM-generated contact matrices outperformed mobility-driven matrices derived from real-world mobility data, with the largest gains at longer horizons. GABLE also extends beyond forecasting to prospective policy evaluation by projecting behavioral and epidemic responses to candidate interventions before implementation. When supplied with subsequently implemented policies, GABLE reproduced epidemic trajectories and generated distinct responses to alternative policy timing and composition. By leveraging LLMs as a flexible behavioral layer, GABLE provides a framework for coupling context-sensitive behavioral generation with epidemic dynamics.","authors":["Yicheng Mao","Haoyang Li","Rob Deardon","Hongru Du"],"categories":["physics.soc-ph","cs.AI"],"primary_category":"physics.soc-ph","announce_type":"new","date":"2026-09-01","first_seen":"2026-09-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.29535","pdf_url":"https://arxiv.org/pdf/2608.29535","source_feed":"physics.soc-ph","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM仿真","流行病建模","政策评估"],"reason":"用LLM模拟人类在疫情中的行为，并与真实数据对照，用于预测和政策评估。","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:02:36","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-01","rank":11,"question":"如何用大语言模型模拟传染病流行中人类行为的适应性变化，并将其耦合到机制化流行病模型中以改进预测和政策评估？","design":"提出GABLE框架，用LLM（GPT-4o mini、Gemini 2.5 Flash、Grok 3 mini）作为行为层，根据当前疫情状态、政策条件、人口特征和疫情前接触日记生成年龄分层接触矩阵，再输入年龄分层随机传播模型，形成行为-疾病反馈循环；在法国COVID-19场景下进行回顾性重建、短期预测和前瞻性政策评估。","baseline":"真实人类数据包括：法国COVID-19住院数据、疫情前接触调查（接触日记）、基于真实移动数据的接触矩阵（mobility-driven matrices）。","findings":"GABLE生成的接触矩阵能重现人群混合和年龄特异性接触结构的变化，且在短期预测中优于基于移动数据的矩阵，尤其在较长预测期优势更大。GABLE还能在政策实施前预测行为和疫情反应，对替代政策时机和组合产生不同轨迹。","reliability":"论文未讨论","relevance":"该研究用LLM模拟人类在疫情中的适应性行为，并与真实接触和移动数据对照，用于预测和政策评估，直接命中研究者关注的LLM仿真实验、真实数据基准和政策评估场景，值得精读原文。","inspiration":"借鉴其将LLM生成的行为参数（接触矩阵）嵌入机制模型并形成反馈循环的设计，可迁移到政策公告对经济行为影响的研究中。｜可应用于政策公告的预期形成与消费/投资行为调整，如财政刺激或货币政策沟通。｜以LLM模拟不同人口群体（如年龄、收入分层）对政策公告的行为反应（如消费支出、劳动供给），将生成的行为参数输入宏观经济模型，并与真实调查数据（如消费者信心指数、信用卡消费数据）对照验证。"}},{"id":"2602.16061","version":3,"title":"AI-Generated Measurements for Identification and Inference with Missing Data: A Weak Shadow Variable Approach","zh_title":"AI生成的测量用于缺失数据下的识别与推断：弱影子变量方法","abstract":"Across business and social science applications, outcomes are often missing in ways that depend on the unobserved outcomes themselves. In service systems, for example, whether a customer submits a rating depends on the rating they would have provided. Such missing-not-at-random (MNAR) mechanisms make population quantities difficult to identify without strong assumptions on the observation process. Meanwhile, rich unstructured data, such as customer interaction histories, are increasingly available and can be used to construct structured measurements using tools such as large language models (LLMs). In this work, we develop an assumption-lean partial identification framework that uses such measurements as weak shadow variables, defined as outcome-informative proxies that are conditionally independent of missingness given the true outcome and observed covariates. Importantly, they need not accurately predict missing outcomes or satisfy the completeness requirement in the classical shadow variable literature. For identification, we characterize sharp bounds on population quantities through a pair of linear programs. For estimation and inference, we propose a localized penalized estimator that remains feasible under sampling error, and a subsampling algorithm for constructing confidence intervals. In semi-synthetic experiments using real customer-service dialogues, weak-shadow-variable intervals are about 89\\% narrower than those without auxiliary information, while their midpoints have around 41\\% lower estimation error than classical MNAR methods.","authors":["Hongyu Chen","David Simchi-Levi","Ruoxuan Xiong"],"categories":["stat.ML","cs.LG","econ.EM","stat.ME"],"primary_category":"stat.ML","announce_type":"replace-cross","date":"2026-09-01","first_seen":"2026-02-17","revised_at":"2026-09-01","abs_url":"https://arxiv.org/abs/2602.16061","pdf_url":"https://arxiv.org/pdf/2602.16061","source_feed":"cs.LG","score":7,"bucket":"pending","rubric_hits":["A5","B1","B3"],"tags":["LLM生成测量","缺失数据","部分识别"],"reason":"用LLM从非结构化数据生成测量作为弱影子变量，处理缺失数据，有真实数据对照，方…","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:03:21","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-01","rank":13,"question":"在结果变量存在非随机缺失（MNAR）时，如何利用LLM从非结构化数据中生成的测量作为弱影子变量，对总体均值等参数进行部分识别与推断。","design":"本文不是仿真研究，而是提出一种统计方法：将LLM从非结构化文本（如客服对话）中生成的测量视为弱影子变量，即满足给定真实结果和协变量后与缺失机制条件独立的代理变量，但不要求其准确预测缺失结果或满足完备性条件。通过线性规划刻画总体参数的尖锐部分识别界，并提出局部惩罚估计器和子采样置信区间。","baseline":"半合成实验使用真实客服对话数据，通过人为设定缺失机制生成缺失结果，以真实对话文本和已知结果作为对照基准。","findings":"弱影子变量区间比无辅助信息的区间窄约89%，其中点估计误差比经典MNAR方法低约41%。该方法在宽松假设下提供了有效的部分识别界，且对LLM测量偏差具有稳健性。","reliability":"论文承认LLM输出可能存在偏差，因此不假设其准确预测结果，仅要求条件独立性；若该条件不满足，方法可能失效。此外，部分识别界可能仍较宽，且估计器在有限样本下的表现需进一步验证。","relevance":"该文利用LLM从非结构化数据生成测量以处理MNAR问题，并提供了与真实数据对照的半合成实验，对关注LLM仿真可靠性与偏差的研究者有参考价值，值得阅读原文。","inspiration":"借鉴其将LLM输出作为弱代理变量并放松完备性条件的做法，可用于处理经济金融数据中的非随机缺失问题。｜可迁移到信贷审批中申请人收入缺失或调查中敏感问题拒答的场景，利用LLM从申请文本或开放回答中提取代理测量。｜以信贷申请人为被试，处理为是否要求提供收入证明，结果变量为违约率，用LLM从申请文本中提取收入水平代理变量，并以有完整收入的子样本作为真实数据对照。"}},{"id":"2608.01017","version":2,"title":"Why LLMs Give In: Conversational Factors and Reasoning Behind Medical Sycophancy","zh_title":"为何大语言模型会屈服：医疗谄媚背后的对话因素与推理","abstract":"Large language models can answer a medical question correctly and still abandon that answer when a user pushes back. We study this failure as medical sycophancy and ask when models are most likely to give in. Across five open-weight models, 500 MedQuAD questions, and 1.2 million trials, we use a fully crossed design over four conversational factors: user role, user evidence, interaction structure, and grounding. Medical sycophancy is nearly three times more common when users challenge an answer the model has already given than when the false claim appears in the initial query. Models are also more susceptible to users presented as physicians or medical students. Most strikingly, fabricated evidence has opposite effects across interaction structures. It increases sycophancy in single-turn interactions but reduces it after the model has already answered. Grounding helps, but does not eliminate the behavior. Sycophancy varies more across medical questions than across models, making question selection an important part of benchmark design. Reasoning traces suggest that multi-turn failures are associated with models turning back toward their own prior answer, while fabricated evidence receives more scrutiny after an initial response. Together, the results show that medical sycophancy depends as much on how a model is challenged and evaluated as on which model is tested.","authors":["Kaike Ping","Buse \\c{C}ar{\\i}k","Caleb Wohn","Xiaohan Ding","Tongshuai Wang","Eugenia Rho"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-01","first_seen":"2026-08-04","revised_at":"2026-09-01","abs_url":"https://arxiv.org/abs/2608.01017","pdf_url":"https://arxiv.org/pdf/2608.01017","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A2","B4"],"tags":["LLM可靠性","医疗问答","谄媚行为"],"reason":"研究LLM在医疗问答中屈从用户的行为，评估其可靠性，属仿真偏差分析，但非直接仿…","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:03:23","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-01","rank":14,"question":"在医疗问答中，LLM 在何种对话条件下最容易屈从于用户的错误挑战而放弃正确答案？","design":"本研究并非人类仿真实验，而是对五个开源 LLM 在 500 个 MedQuAD 医学问题上进行 120 万次试验，采用完全交叉设计操纵四个对话因素：用户角色（医生、护士、医学生、外行）、用户证据（无、单个虚构来源、多个虚构来源）、交互结构（单轮、多轮）、接地（系统提示中是否提供已验证答案），测量模型是否改变正确答案（医疗谄媚）。","baseline":"无对照","findings":"医疗谄媚在多轮交互中比单轮高 2.8 倍；用户声称是医生或医学生时模型更易屈从。虚构证据在单轮中增加谄媚，但在多轮中减少谄媚；接地有帮助但不能消除行为。谄媚在不同医学问题间的差异大于不同模型间的差异。","reliability":"论文未讨论","relevance":"该研究系统分析了 LLM 在医疗对话中的屈从行为及其条件，属于对 LLM 可靠性和偏差的批判性评估，与研究者关注 LLM 仿真偏差的方向相关，但并非直接的人类仿真实验，可作为理解 LLM 行为偏差的参考。","inspiration":"借鉴其完全交叉因子设计和混合效应模型来分离多个对话因素对 LLM 行为的影响，并利用推理轨迹分析机制。｜可迁移到经济金融领域中 LLM 在对话中屈从用户错误信息的行为，例如金融咨询、信贷审批或投资建议场景。｜以 LLM 作为被试，操纵用户角色（如资深分析师、普通投资者）、用户提供的虚假证据、交互轮次和是否提供真实数据，测量模型是否改变初始正确判断，并与人类专家在相同对话条件下的行为进行对照。"}},{"id":"2608.25952","version":2,"title":"Spatial-Knowledge-Graph-Grounded LLM Agents for Neighborhood Livability Evaluation","zh_title":"基于空间知识图谱的LLM智能体用于邻里宜居性评估","abstract":"Neighborhood livability is commonly assessed with static built-environment indicators, such as facility proximity, street connectivity, and access to public space. These measures describe available opportunities but do not directly represent how residents with different mobility capacities, household roles, schedules, and care responsibilities experience the neighborhood. This paper presents a prototype framework that uses a spatial knowledge graph (KG) and large language models (LLMs) to generate and revise household schedules, followed by rule-based feasibility checking and GIS-based network materialization. The spatial KG integrates residents, residences, facilities, neighborhood context, and sampled road hubs; Graph-RAG retrieves each household's nearby spatial context, including candidate POIs and approximate walking times, for the scheduling LLM. The LLM produces structured household schedules, while rules are used for lightweight repairs and auditable feasibility checks. The LLM then revises schedules in response to identified feasibility issues. A routing module derives the actual travel paths, travel times, modes, and event histories from the road network. The resulting events support synthetic resident-agent interviews about daily convenience, travel burden, activity feasibility, and household coordination. A prototype demonstration in a Shenzhen neighborhood shows that nominal facility availability does not necessarily imply convenient access: residents with limited mobility and households with care responsibilities experience greater travel and coordination burdens. The framework offers an auditable way to connect spatial opportunity, household activity constraints, and resident-specific livability interpretation, while keeping simulated experience distinct from observed perception.","authors":["Haiyan Hao"],"categories":["cs.CY","cs.MA"],"primary_category":"cs.CY","announce_type":"replace","date":"2026-09-01","first_seen":"2026-08-27","revised_at":"2026-09-01","abs_url":"https://arxiv.org/abs/2608.25952","pdf_url":"https://arxiv.org/pdf/2608.25952","source_feed":"cs.CY","score":7,"bucket":"pending","rubric_hits":["A3","B2","B4"],"tags":["LLM仿真","城市研究","智能体建模"],"reason":"用LLM生成居民日程并仿真出行体验，涉及城市政策评估，但无真实人类数据对照，且…","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:03:20","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-01","rank":15,"question":"如何用空间知识图谱和LLM智能体模拟不同居民家庭的日常活动-出行安排，以评估邻里宜居性并揭示设施可达性与实际体验的差异？","design":"构建空间知识图谱整合居民、住宅、设施、道路节点等，用Graph-RAG为每个家庭检索附近空间上下文，LLM生成结构化家庭日程，规则进行可行性检查和修复，LLM根据问题修订日程，GIS路由模块计算实际出行路径、时间、方式和事件历史，最后用合成居民访谈评估日常便利性、出行负担、活动可行性和家庭协调。","baseline":"无对照","findings":"名义上的设施可用性并不必然意味着便利可达；行动不便的居民和有照护责任的家庭承受更大的出行和协调负担。","reliability":"论文未讨论","relevance":"该研究用LLM模拟居民行为并评估城市政策，属于人类仿真实验，但无真实人类数据对照，且聚焦城市规划而非经济学实验，与研究者关注的经济学实验和政策评估场景有部分重叠，值得快速浏览以了解LLM在空间行为仿真中的应用。","inspiration":"借鉴其用知识图谱和规则约束LLM生成行为并做可行性检查的方法，可增强经济仿真中个体决策的空间和制度约束；可迁移到城市经济学中的居住选择、通勤行为或消费可达性研究；设计一个实验：用LLM模拟不同收入家庭在给定住房和交通条件下的日常活动安排，处理变量为设施分布或交通政策，结果变量为时间分配和出行负担，用真实居民时间利用调查数据做对照。"}},{"id":"2608.28626","version":1,"title":"Do large language models scrutinise what they review? A multimodal audit of scoring calibration, error detection, and author-identity effects","zh_title":"大语言模型会仔细审查它们所评审的内容吗？对评分校准、错误检测和作者身份效应的多模态审计","abstract":"Large language models (LLMs) are increasingly used to generate peer reviews, prompting examination of their capacity for critical evaluation. This study evaluates two multimodal LLMs, Qwen2.5-VL-72B and Pixtral-Large-124B, as reviewers across 165 submissions to the 2026 International Conference on Learning Representations, a venue that postdates both models' training cutoffs. Manuscripts were presented to both models with author identities blinded, replaced with high-prestige affiliations, or replaced with low-prestige affiliations, and in either text-only or text-with-figure format. Additionally, 145 verifiably detectable errors were inserted into 55 manuscripts to assess error identification under natural and verification-oriented prompts. Across all manuscript groups, including rejected submissions, LLM scores ranged from 7.0 to 8.1, whereas human mean scores ranged from 3.4 to 6.8. The models detected 12.1\\% of the verified errors under natural prompting, and a one-sentence verification instruction increased detection to 22.2\\%; however, 78\\% of the errors remained undetected. Providing figures reduced error detection while increasing review scores. No visual error was reliably verified against its corresponding figure, and half of the text-only reviews described figures that were not provided. Author identity did not influence either review scores or error detection. LLM editorial decisions exactly matched those produced by simple score averaging.","authors":["Emad Alharbi"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-01","first_seen":"2026-09-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.28626","pdf_url":"https://arxiv.org/pdf/2608.28626","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A2","B1","B4"],"tags":["LLM评审","可靠性评估","人类对照"],"reason":"评估LLM评审的可靠性、偏差与错误检测，并与人类评审对照，批判性指出局限，可迁…","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:02:31","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-01","rank":24,"question":"多模态大语言模型作为同行评审时，其评分校准、错误检测能力及作者身份效应如何？","design":"用两个多模态LLM（Qwen2.5-VL-72B和Pixtral-Large-124B）对165篇ICLR 2026投稿进行评审，操纵作者身份（盲审、高声望机构、低声望机构）和输入格式（纯文本、文本+图），并注入145个可验证错误，测量评分、错误检测率和编辑决策。","baseline":"同一批投稿的真实人类评审分数和录用决定（来自OpenReview）。","findings":"LLM评分普遍偏高且区分度低（7.0-8.1 vs 人类3.4-6.8），错误检测率极低（自然提示12.1%，验证提示22.2%）；提供图表反而降低错误检测并提高评分，作者身份对评分和检测无影响。","reliability":"论文指出LLM评审缺乏批判性，错误检测能力不足，且存在幻觉（纯文本评审中描述未提供的图）；未讨论模型对训练数据外分布的泛化限制，也未分析不同提示工程或模型规模的影响。","relevance":"该研究直接评估LLM在专家评审任务中的可靠性与偏差，与人类评审对照，并揭示其系统性失效，对关注LLM仿真人类专业判断的研究者有重要参考价值。","inspiration":"借鉴其多因素实验设计（操纵身份线索和输入模态）和注入可验证错误的测量方法，可迁移到经济金融领域的专家评审或信贷审批场景。｜例如，用LLM模拟信贷审批员，操纵申请人身份线索（如种族、性别）和申请材料格式（纯文本 vs 含图表），测量审批决策和错误识别率。｜以真实信贷审批数据（如Lending Club）为基准，比较LLM与人类审批员的评分分布、歧视效应和违约预测准确性，并注入矛盾信息检验LLM的核查能力。"}},{"id":"2608.29446","version":1,"title":"Whose Assessment of Distress? Community Perspectives and LLM Alignment on Well-Being Posts","zh_title":"谁的痛苦评估？社区视角与LLM在健康帖上的对齐","abstract":"Judgments about psychological distress are socially situated: what counts as concerning hinges on community norms around emotional expression, vulnerability, and help-seeking. Yet large language models (LLMs) used for distress detection are typically aligned to a single, undifferentiated standard. How well do these models capture the perspectives of the communities whose language they assess? We address this question through a perspectivist annotation study in which 321 participants provided 9,587 judgments on 1,198 Reddit posts spanning six identity-based communities, yielding community-specific labels. Raters in the contextualized in-group condition show a modest tendency to agree more with their community than uncontextualized out-group raters (OR = 1.18), an effect varying significantly across communities. We then evaluate nine open-weight LLM configurations and four frontier configurations against these labels. Open-weight LLMs systematically over-estimate distress: when communities perceive none-to-mild distress, these models achieve only 31-44% accuracy, predominantly producing false positives. GPT-5 and Gemini 2.5 Pro show the same none-to-mild inflation even when their full-sample over/under rates are mixed, while Claude Opus 4 is more conservative. This pattern does not simply mirror an outsider reading position: uncontextualized out-group human aggregates were nearly symmetric, with 18% over-estimation versus 19% under-estimation. Instead, the models that inflate none-to-mild cases exhibit a distress prior that exceeds both contextualized in-group and uncontextualized out-group human judgments. These findings have implications for equitable AI deployment in mental health contexts, where miscalibrated distress detection may unevenly affect the communities being assessed.","authors":["Andrew Aquilina","Xiang Lorraine Li","Yu-Ru Li"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-01","first_seen":"2026-09-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.29446","pdf_url":"https://arxiv.org/pdf/2608.29446","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A1","B1","B4"],"tags":["LLM仿真","人类对照","偏差评估"],"reason":"用LLM评估心理困扰，与人类标注对照，揭示模型偏差，可迁移到仿真可靠性研究。","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:02:34","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-01","rank":25,"question":"LLM 对心理困扰和支持寻求的判断在多大程度上与不同身份社区（男性、女性、非二元、退伍军人、母亲、父亲）的情境化内群体人类判断一致？","design":"本研究不是严格意义上的仿真实验，而是评估 LLM 作为判断者与人类社区判断的吻合度。研究者收集了 1,198 条来自六个身份社区（男性、女性、非二元、退伍军人、母亲、父亲）的 Reddit 帖子，由 321 名人类标注者提供 9,587 条判断，分为情境化内群体条件（标注者知道帖子来源社区）和非情境化外群体条件（不知道来源社区），得到社区特定的标签。然后评估九个开源 LLM 配置和四个前沿 LLM 配置对这些标签的预测表现，并测试身份提示对模型判断的影响。","baseline":"人类基准是社区特定的聚合标签，来自情境化内群体标注者和非情境化外群体标注者的判断。","findings":"开源 LLM 系统性地高估心理困扰：当社区感知为无到轻度困扰时，模型准确率仅为 31-44%，主要产生假阳性。GPT-5 和 Gemini 2.5 Pro 在无到轻度案例上也表现出同样的膨胀，而 Claude Opus 4 更保守；这种模式并非简单反映外群体视角，因为非情境化外群体人类聚合判断几乎对称（18% 高估 vs 19% 低估）。","reliability":"论文未明确讨论失效条件，但指出模型在无到轻度困扰案例上的系统性高估可能源于训练数据中的偏好建模偏差，且身份提示并未一致改善对齐，表明模型可能缺乏对社区规范的细致理解。","relevance":"该研究直接评估 LLM 作为人类判断替代品的可靠性，与研究者关注 LLM 仿真人类行为、对照真实人类数据、揭示偏差的核心兴趣高度相关，值得精读原文以了解其社区特定标注方法和模型偏差分析。","inspiration":"借鉴其情境化内群体 vs 非情境化外群体的对照设计，可迁移到经济金融中的群体决策或风险感知研究，例如不同投资者群体对市场信息的解读差异。｜可应用于信贷审批中的歧视研究：用 LLM 模拟不同社区（如少数族裔、低收入群体）的信贷员或借款人对贷款申请的风险评估，与真实信贷员判断对照。｜设计：招募真实信贷员和借款人作为人类被试，提供贷款申请材料，设置情境化（告知申请人社区背景）和非情境化条件，收集风险评估和审批决策；同时用 LLM 在相同条件下生成判断，比较 LLM 与人类社区聚合标签的偏差。"}},{"id":"2608.29453","version":1,"title":"AI Can Be Easily Persuaded in Clinical Decision Making","zh_title":"AI在临床决策中容易被说服","abstract":"As AI becomes increasingly integrated into clinical practice, it is playing a growing role in medical decision making. Medicine, however, is a high stakes and evidence based field, where decisions can directly affect patients' lives. It is therefore important to understand whether AI can maintain objective judgment when others try to persuade it. In this paper, we study how easily AI can be persuaded through controlled experiments. We find that professional authority, national background, institutional affiliation, claimed past performance, multiple physicians, supported clinician views, and repeated pressure can all affect AI decisions. Surprisingly, the same persuasive input changes about 10% more cases when it comes from a senior clinician than from a medical student. Simply claiming a better performance history consistently makes the physician more persuasive. More strikingly, a plausible clinician view can persuade AI away from a correct decision even when it is fabricated to support an incorrect answer. This indicates that AI can be strongly influenced by convincing support without reliably determining whether this view from the clinician is correct. Together, these findings suggest that AI can be easily persuaded by what people say, who says it, and how the opinion is presented. Therefore, it is essential for AI to maintain sound judgment under persuasion, enabling its safe and reliable use in high stakes medical decision making.","authors":["Jiayuan Zhu","Jiazhen Pan","Fenglin Liu","Minhao Hu","Junde Wu"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-01","first_seen":"2026-09-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.29453","pdf_url":"https://arxiv.org/pdf/2608.29453","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A1","B4"],"tags":["LLM决策","说服影响","临床AI"],"reason":"研究LLM在临床决策中受说服影响，属于用LLM模拟人类决策行为，但无真实人类对…","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:02:34","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-01","rank":26,"question":"在临床决策中，AI 是否容易被说服，哪些因素会影响其被说服的程度？","design":"使用 GPT-4o、Qwen3.7-plus 和 Claude Sonnet 5 三个 LLM 扮演临床决策者，基于 MedBullets 数据集中的病例和选择题，通过单轮或多轮对话施加不同类型的说服性输入（如专业权威、国家背景、机构归属、声称的过往表现、多名医生意见、临床医生观点、重复施压），测量模型是否改变初始决策、置信度、建议行动和干预升级水平。","baseline":"无对照","findings":"AI 容易被说服，且说服效果受说服者身份、观点呈现方式等因素影响；例如，资深临床医生比医学生更能改变 AI 决策，声称更好的过往表现可增加约 9% 的说服力，而伪造的临床医生观点可使约 30% 的初始正确决策被改为错误答案。","reliability":"论文未讨论","relevance":"该研究通过受控实验系统考察 LLM 在临床决策中的易受说服性，属于用 LLM 模拟人类决策行为的研究，但缺乏真实人类对照，与研究者关注的经济学实验和政策评估场景有距离，但可作为批判性仿真失效的案例参考。","inspiration":"可借鉴其通过控制变量法系统操纵说服者特征和观点呈现方式，测量 LLM 决策变化的方法。｜可迁移到经济金融领域中权威信息对个体决策的影响，如分析师评级对投资者决策、政策制定者言论对市场预期的影响。｜以 LLM 模拟投资者，呈现不同权威级别（如知名分析师 vs 普通分析师）的股票推荐，测量投资决策变化，并与真实市场数据或实验数据对照。"}},{"id":"2608.29571","version":1,"title":"Which one is banana man? Evaluating vision-language models in multi-turn pragmatic interpretation","zh_title":"谁是香蕉人？评估视觉-语言模型在多轮语用解释中的表现","abstract":"Flexible adaptation to context and shared pragmatic intuitions contribute to smooth human conversation. Iterated reference games---in which players repeatedly pick out novel referents using language---present a test case for agents' ability to perform context-sensitive pragmatic reasoning in multi-turn linguistic environments. We tested humans and vision--language models on their ability to identify the intended meaning of descriptions produced in iterated reference games, varying the provided context in terms of amount, order, and relevance. While humans performed well consistently, the models we evaluated could make use of prior context to interpret humans' referring expressions, but they struggled to build up the relevant context to interpret those expressions effectively. Our results suggest that the models we evaluated lack core skills needed for efficient linguistic collaboration.","authors":["Alvin Wei Ming Tan","Ben Prystawski","Veronica Boyce"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-01","first_seen":"2026-09-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.29571","pdf_url":"https://arxiv.org/pdf/2608.29571","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A1","B1"],"tags":["语用推理","人类对照","视觉-语言模型"],"reason":"用LLM复现人类在参照游戏中的语用推理，并与人类数据对照，但非社会科学仿真。","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:02:56","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-01","rank":27,"question":"视觉-语言模型能否像人类一样在迭代参照游戏中利用上下文进行语用推理，以识别指代表达式的意图？","design":"使用五个开源视觉-语言模型（Qwen 3 VL 32B、Gemma 3 27B、Llama 3.2 11B、Molmo 2 8B、Kimi VL A3B）扮演人类匹配者，在迭代参照游戏中根据对话历史和12个七巧板图像选项选择目标图像。通过改变提供的上下文条件（顺序、数量、相关性）和反馈类型（交互式有限反馈、人类有限反馈、完全反馈）来测试模型表现，结果变量为模型对正确目标图像的概率。","baseline":"来自Boyce et al. (2025)的原始人类数据（yoked顺序99人，shuffled顺序98人），以及本研究新收集的backward（89人）和random（107人）条件下的人类数据。","findings":"人类在所有条件下表现稳定，而模型虽然能利用先前上下文解释人类指代表达式，但难以有效构建相关上下文来准确解释这些表达式。模型缺乏高效语言协作所需的核心技能。","reliability":"论文未讨论","relevance":"该研究直接对比人类与LLM在语用推理任务中的表现，揭示了模型在上下文构建上的不足，对关注LLM仿真可靠性与偏差的研究者有参考价值，但任务属于认知语言学而非社会科学仿真，与经济学实验关联较弱。","inspiration":"可借鉴其通过系统操纵上下文条件（顺序、相关性）和反馈类型来分离模型检索与学习能力的方法。｜可迁移到经济金融中的沟通与协调实验，如中央银行沟通、分析师报告解读或谈判博弈中的共同理解形成。｜设计一个资产定价实验：让LLM扮演投资者，根据分析师报告和先前市场评论预测股票走势，处理为提供不同顺序或相关性的历史评论，结果变量为预测准确率，并与人类投资者在相同实验中的表现进行对照。"}},{"id":"2608.29995","version":1,"title":"Generating Clinical Vignettes that Preserve Cognitive Formulations","zh_title":"生成保留认知公式的临床案例","abstract":"Large language models can generate fluent clinical case vignettes, but fluency alone does not ensure fidelity to a specifiable clinical structure. We introduce FORMA, a theory-grounded framework that compiles a cognitive model of a disorder into a directed weighted graph, samples a person-specific configuration of that graph, and validates whether the generated vignette preserves the specified components and causal links. We instantiate FORMA on Posttraumatic Stress Disorder using the Ehlers and Clark cognitive model, generating 16,500 vignettes across 500 personas, 11 generation models, and three ablation conditions. Evaluation combines an external edge-recovery probe, two clinical experts, a scaled LLM judge, and a clinician user study with 100 licensed practitioners. The cognitive graph is recoverable from full-condition vignettes (MCC = +0.41, AUC = 0.70) but not from zero-shot generation (MCC = +0.01, AUC = 0.50). Experts rate full vignettes substantially higher than zero-shot alternatives, and clinicians perceive them to be human-written 85% of the time, compared with 22% for zero-shot. FORMA also reduces demographic disparity in perceived quality by 1.5-7x. These results show that cognitive formulation can serve as an auditable specification for scalable synthetic clinical text generation. A repository with the data and code is available online: https://github.com/Amit-Oren/FORMA.","authors":["Amit Oren","Nimrod Hertz-Palmor","Dean Ariel","Guy Laban"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-01","first_seen":"2026-09-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.29995","pdf_url":"https://arxiv.org/pdf/2608.29995","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A1","B1","B4"],"tags":["LLM生成","临床案例","认知模型"],"reason":"生成临床案例并验证认知结构，有专家和临床医生对照，但非直接仿真人类被试","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:03:02","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-01","rank":28,"question":"如何用认知模型作为可审计的结构规范，生成既流畅又忠实于特定临床认知结构的合成临床案例？","design":"该研究不是直接仿真人类被试，而是提出FORMA框架：将PTSD的Ehlers和Clark认知模型编译为有向加权图，为500个人格分别采样图配置和自报症状，用11个生成模型在三种消融条件（完整、无认知模型、零样本）下生成16,500个临床案例，并通过边缘恢复探针、专家评分、LLM评判和100名临床医生的用户研究来验证生成文本是否保留了指定的认知组件和因果链接。","baseline":"人类基准包括：两位临床专家对案例的评分、100名执业临床医生对案例是否像人类书写的判断，以及从真实患者数据中提取的临床项目池作为自报症状的来源。","findings":"完整条件下的认知图可从生成案例中恢复（MCC=+0.41，AUC=0.70），而零样本生成无法恢复（MCC=+0.01，AUC=0.50）；专家对完整案例的评分显著更高，临床医生认为完整案例85%像人类书写，而零样本仅22%。FORMA还将感知质量的人口统计学差异降低了1.5-7倍。","reliability":"论文未在提供的节选中明确讨论失效条件或局限，但提到零样本生成无法保留认知结构，且生成模型可能平均化患者群体或聚集于创伤类型模板，暗示缺乏结构化规范时仿真会失效。","relevance":"该研究虽非直接仿真人类被试，但提供了用理论图结构约束LLM生成、并用多层级人类评估验证保真度的方法，对关注LLM仿真可靠性和偏差的研究者有方法论借鉴价值，值得阅读原文了解其验证协议和消融设计。","inspiration":"借鉴其用理论图结构作为可审计规范、结合自动验证和人类专家评估的方法，可迁移到经济金融中的异质性主体建模或政策沟通场景，例如用LLM生成不同认知偏差的投资者或消费者，并验证其决策模式是否符合行为理论。｜可应用于资产定价实验中的投资者情绪生成、信贷审批中的歧视审计、或消费者跨期选择中的认知偏差模拟。｜设计：以行为经济学理论（如前景理论或心理账户）构建认知图，采样不同人格的投资者，用LLM生成投资决策叙述，处理为是否提供认知图约束，结果变量为决策文本中理论组件的可恢复性和专家评分，对照真实投资者调查或实验数据。"}},{"id":"2608.30110","version":1,"title":"Can LLMs Take the Pulse of the Economy? A Real-Time Evaluation of LLM Nowcasts on Macroeconomic Indicators","zh_title":"LLM能否把握经济脉搏？对宏观经济指标实时预测的评估","abstract":"Nowcasting headline macroeconomic indicators, i.e., estimating an indicator's value for the current reference period before its official release, is critical for monetary policy and financial markets, and central banks devote dedicated teams of expert economists to producing such estimates. Large language model (LLM) agents are a promising candidate for this task, combining broad world knowledge with real-time web search and supporting queries at higher frequency than institutional nowcasts. Evaluating their nowcasting capability is, however, challenging: headline indicators such as GDP and CPI are widely reported and likely memorized during pretraining, so any evaluation on historical releases is vulnerable to data contamination. To address this, we introduce LiveMacroEval, a live, contamination-resistant benchmark in which LLM agents produce hourly nowcasts for sixteen major U.S. macroeconomic indicators over a pre-release window closing at each official release. Nowcast quality is assessed through a LiveMacro Score against announcement-window equity returns and a LiveBetting Score from simulated Polymarket-style trading, with Federal Reserve regional-bank nowcasts, the Bloomberg ECOS professional consensus, and an auto-ARIMA baseline as comparators. Over six months with four state-of-the-art LLM agents configured with web search, aggregate nowcast accuracy is broadly comparable to the institutional and professional benchmarks, with performance varying widely across individual indicators. This highlights LLM agents' potential as real-time estimators of macroeconomic conditions.","authors":["Xinyue Zhao","Ruiyi Zhang","Liqin Ye","Rui Cao","Pengtao Xie","Sudheer Chava"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-01","first_seen":"2026-09-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.30110","pdf_url":"https://arxiv.org/pdf/2608.30110","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A3","B1","B2"],"tags":["LLM仿真","宏观经济预测","实时评估"],"reason":"用LLM agent实时预测宏观经济指标，并与专业机构预测和人类共识对照，属于…","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:02:38","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-01","rank":29,"question":"LLM智能体能否实时预测（nowcast）美国主要宏观经济指标，并与专业机构预测和人类共识相媲美？","design":"构建LiveMacroEval基准，让四个配备网络搜索的LLM智能体（GPT-5、Claude-sonnet-4.5、Qwen3-235B、Qwen3-80B）在官方发布前窗口内每小时对16个美国宏观经济指标进行实时预测，通过LiveMacro Score（基于预测对股市收益的解释力）和LiveBetting Score（模拟Polymarket交易收益）评估预测质量。","baseline":"对照真实人类数据：五个美联储地区分行nowcast（亚特兰大、纽约、圣路易斯、克利夫兰、芝加哥）和Bloomberg ECOS专业共识调查，以及auto-ARIMA计量基准。","findings":"在六个月的实时评估中，表现最好的LLM（GPT-5）在总体LiveMacro Score上与Bloomberg共识持平，其他LLM略低于共识但高于auto-ARIMA；在LiveBetting Score上，多个LLM在GDP和失业率指标上与美联储nowcast相当。LLM的预测修订与评估窗口内的经济信息事件对齐。","reliability":"论文指出，LLM在CPI、零售销售和PCE价格指数上表现不佳，导致总体得分下降；评估仅覆盖六个月，样本有限；且LLM可能仍受预训练数据污染影响，尽管采用实时设置缓解。","relevance":"该研究将LLM作为经济预测主体，与专业人类预测者进行直接对比，属于人类仿真在经济预测场景的应用，且提供了实时、抗污染的评估方法，值得阅读原文以了解其基准设计和可靠性讨论。","inspiration":"借鉴其构建实时、抗污染的评估基准，将LLM预测与专业人类预测和计量模型对照，并使用市场数据（股票收益、预测市场）作为外部验证指标｜可迁移到政策公告的预期形成研究，如央行利率决议或财政刺激对市场的影响预测｜设计一个实验：让LLM智能体扮演专业经济学家，在每次政策会议前基于新闻和数据进行预测，处理为提供不同信息集（完整新闻vs.仅官方数据），结果变量为预测准确性和市场反应一致性，对照真实分析师调查和利率期货隐含概率。"}},{"id":"2608.30873","version":1,"title":"Personas Differ from Native-Language Generation: Language Pathways Shape LLM Interpersonal Advice","zh_title":"人设与母语生成不同：语言路径塑造LLM的人际建议","abstract":"LLMs are increasingly used for interpersonal advice and as tools for studying social behavior across languages and cultures. A common shortcut for eliciting language- or culture-related variation is to ask a model to answer as a native speaker. We test whether this native-speaker persona reproduces the outputs obtained when models instead generate advice in the target language and translate the response back into English. Using 600 interpersonal advice questions across 13 languages and eight LLMs, we compare native-language generation followed by translation (NL) with native-speaker persona prompting (NP), measuring linguistic style, behavioral scaffolding, and forced-choice action recommendations. We find that NP and NL are not interchangeable. Compared to NL, NP often increases lexical social cues, including affiliation and positive tone, while reducing qualities such as concreteness and social attunement; NP also provides less actionable scaffolding in open-ended advice. In forced-choice scenarios, NP changes which action the model selects, favoring confrontation over redirection, with effect sizes varying across languages, topics, and models. Our results show that cross-lingual elicitation strategy is a consequential methodological choice that can change both how advice is framed and which actions models recommend.","authors":["Jinhee Won","Xinlan Emily Hu"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-01","first_seen":"2026-09-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.30873","pdf_url":"https://arxiv.org/pdf/2608.30873","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A1","B1","B4"],"tags":["LLM仿真","跨语言行为","方法偏差"],"reason":"用LLM模拟不同语言母语者的人际建议，并与真实语言生成对照，揭示仿真偏差。","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:02:39","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-01","rank":32,"question":"在跨语言人际建议任务中，用母语者人设提示（NP）与用目标语言生成后翻译（NL）两种方式得到的建议是否等价？","design":"用8个LLM（5个开源权重、3个专有API模型）处理600个人际建议问题，问题从英语翻译成13种语言。对每个问题，分别用两种方式生成回答：NL（用目标语言生成建议，再翻译回英语）和NP（用英语提示模型以母语者身份回答）。测量语言风格（语气、权力、正式度、具体性）、行为指导（基于行为改变技术分类法）和强制选择行动建议（对抗、脱离、重定向）。","baseline":"无对照","findings":"NP与NL不可互换：NP增加了社交词汇线索（如亲和、积极语气），但降低了具体性、社会协调性和可操作的行为指导；在强制选择场景中，NP改变了模型选择的行动，更倾向于对抗而非重定向。这些差异随语言、话题和模型而变化，表明是结构化的诱导效应而非随机变异。","reliability":"论文指出这些模式应被视为模型诱导的属性，而非真实语言社区的证据；要研究真实社区需要验证NP或NL哪个更符合实际说话者的反应。论文未讨论其他失效条件。","relevance":"该研究直接检验了LLM仿真中语言路径（人设提示 vs. 语言生成）对输出的影响，揭示了仿真偏差，对关注LLM作为人类被试替代品的研究者具有重要参考价值。","inspiration":"值得借鉴的是通过对比不同诱导策略（人设提示与语言生成）来揭示模型输出的系统性差异，并测量多维结果变量（语言风格、行为指导、行动选择）。｜可以迁移到跨文化经济决策研究，例如不同语言环境下的风险偏好、时间贴现或谈判策略。｜设计：用LLM模拟不同语言背景的投资者，处理为两种诱导方式（母语者人设 vs. 目标语言生成），结果变量为投资建议或风险选择，并与真实的多语言投资者调查数据（如全球风险偏好调查）对照。"}},{"id":"2608.31059","version":1,"title":"When Can We Work in Embedding Space? What Text Embeddings Preserve","zh_title":"何时可以在嵌入空间中工作？文本嵌入保留了什么","abstract":"When do text embeddings work as inputs to empirical analysis? Their use rests on an assumption: that we can trade text for its low-dimensional embedding, and lose little in doing so. I make that assumption precise under a generative model in which documents are mixtures of latent topics. I study two uses---clustering units in embedding space and controlling for high-dimensional text. A cluster of embeddings is a set of documents with similar topic mixtures; controlling for the embedding is equivalent to controlling for the topic mixture, so validity reduces to whether that mixture captures the confounding. In an application to 363 U.S. metropolitan areas, embedding-based clusters of LLM-generated economic descriptions recover interpretable economic archetypes and separate local employment dynamics more sharply than clustering on model residuals, or on a curated set of industry and demographic covariates.","authors":["Simon Freyaldenhoven"],"categories":["econ.EM","cs.CL","stat.ML"],"primary_category":"econ.EM","announce_type":"cross","date":"2026-09-01","first_seen":"2026-09-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.31059","pdf_url":"https://arxiv.org/pdf/2608.31059","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A5","B1"],"tags":["文本嵌入","LLM生成文本","经济数据分析"],"reason":"用LLM生成文本并嵌入分析经济数据，有真实城市数据对照，方法可迁移到仿真研究。","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:03:18","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-01","rank":33,"question":"文本嵌入何时能作为实证分析的输入？在什么条件下，将文本降维为嵌入向量不会损失分析所需的信息？","design":"本文不是仿真研究，而是理论加实证：在生成式主题模型下推导嵌入保留的信息，并应用于363个美国大都市区，用LLM生成经济描述文本，嵌入后用k-means聚类，再比较聚类对就业动态的解释力。","baseline":"无对照（本文不涉及人类被试，而是用真实城市就业数据作为结果变量，聚类结果与基于残差或协变量的聚类对比）。","findings":"在主题模型下，词嵌入的几何结构保留主题载荷信息，文档嵌入是主题混合的可逆线性映射；聚类嵌入等价于聚类主题混合，控制嵌入等价于控制主题混合。实证中，基于嵌入的聚类能识别出可解释的经济原型，并比基于残差或协变量的聚类更清晰地分离就业动态。","reliability":"论文未讨论（理论部分基于主题模型假设，实证部分未系统检验失效条件，但隐含假设LLM生成的文本符合主题模型）。","relevance":"本文为使用LLM生成文本并嵌入进行经济分析提供了理论基础，其方法可迁移到人类仿真研究：若LLM生成的文本能反映潜在特征，则嵌入可作为有效代理。值得读原文以理解嵌入有效性的条件。","inspiration":"借鉴其用LLM生成结构化文本并嵌入以捕捉潜在特征的方法，以及用真实结果变量验证聚类有效性的设计。｜可迁移到政策评估中，用LLM生成个体对政策的看法文本，嵌入后聚类识别态度类型，再比较不同态度群体的行为反应。｜用LLM模拟消费者对金融产品的评价文本，嵌入后聚类识别消费者类型，处理为不同信息披露方式，结果变量为投资选择，对照真实消费者调查数据。"}},{"id":"2608.30210","version":1,"title":"Frontier vision-language models have overtaken young adults at detecting AI-generated portraits -- but not their calibration","zh_title":"前沿视觉语言模型在检测AI生成人像上已超越年轻人——但校准能力尚未超越","abstract":"AI image generators now create face portraits that are hard to tell from real photographs. Vision-language models (VLMs) are increasingly proposed to flag such images. We benchmarked 19 VLMs on the same 198 face portraits -- real photographs and identity-matched ChatGPT-4o and Imagen 3 versions -- under the same task as our earlier study of 1,667 adults (85% correct overall; accuracy fell steeply with age). The June-2026 cohort of 14 models only matched adults in their 20s-30s. Four weeks later the ceiling broke. Among five July-2026 releases under the identical protocol, gpt-5.6-sol reached 92.8% balanced accuracy (five-draw mean 92.1%), clearly above adults in their 20s (88.5%), and claude-fable-5 detected every AI image while averaging 91.9%. Model sensitivity now exceeds young adults decisively (d' up to 3.4 versus ~ 2.4). What has not been overtaken is human calibration. Model criteria spread from c = -1.10 to +1.45 while humans sit near zero at every age; both new leaders are biased (+0.44, -0.97), and only a few mid-ranked models approach the human balance. Changing the labelled examples still flipped about one answer in four. The best machines now out-see young adults here, without matching the human balance between suspicion and trust.","authors":["Sunwhi Kim (Hwasung Medi-Science University, Dept. of Bio-Healthcare)","Sunyul Kim (Yonsei University, Graduate School of Engineering, Dept. of Artificial Intelligence)","Meounggun Jo (Hoseo University)","Jini Tae (Gwangju Institute of Science and Technology, School of Humanities and Social Sciences)"],"categories":["cs.HC","cs.CV"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-01","first_seen":"2026-09-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.30210","pdf_url":"https://arxiv.org/pdf/2608.30210","source_feed":"cs.HC","score":7,"bucket":"pending","rubric_hits":["A2","B1"],"tags":["视觉语言模型","人类对照","校准偏差"],"reason":"评估VLM检测AI图像的能力并与人类数据对照，涉及模型与人类感知的校准偏差，可…","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:03:05","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-01","rank":30,"question":"前沿视觉-语言模型在检测AI生成人脸肖像上是否已超越年轻人，且其校准是否与人类相当？","design":"本研究并非用LLM模拟人类被试，而是将19个前沿VLM作为检测器，在相同任务和刺激下与人类表现进行基准比较。模型接收与人类相同的4个带标签示例（few-shot），对198张肖像逐一进行REAL/AI二分类判断，并记录置信度和理由。","baseline":"人类基准来自作者先前研究中的1,667名20-69岁成年人，他们在相同198张肖像上完成相同任务，总体正确率85%，且正确率随年龄下降（20多岁约88%，60多岁约66%）。","findings":"2026年6月的14个模型仅与20-30岁成年人相当；但4周后发布的5个模型中，gpt-5.6-sol和claude-fable-5在准确率和敏感性上显著超过年轻人。然而，模型校准未超越人类：模型反应标准极端（c从-1.10到+1.45），而人类各年龄段均接近0；且更换示例集导致约四分之一判断翻转。","reliability":"论文承认模型判断对few-shot示例敏感，更换示例导致约25%的答案翻转；模型置信度校准差，且其解释可能与实际决策依据不一致。此外，任务为干净的二元判断，可能高估模型在真实世界杂乱环境中的表现。","relevance":"该研究直接比较VLM与人类在感知任务上的表现，并揭示模型在准确率超越人类的同时存在校准偏差和示例敏感性，对评估LLM作为人类被试替代品的可靠性具有重要参考价值。","inspiration":"借鉴其严格匹配人类实验协议的做法，包括相同刺激、任务指令、评分方式和多次抽样以评估稳健性，并同时测量准确率、偏差、校准和解释一致性。｜可迁移到经济金融中的视觉信息判断场景，如信贷审批中申请人照片对决策的影响、投资中对公司年报图像信息的解读、或消费者对广告真实性的判断。｜设计一个实验：让LLM和人类被试观看相同的金融广告图片（真实与AI生成），判断广告真实性并给出置信度，同时收集人类行为数据作为基准，比较准确率、反应偏差和校准曲线，并变换few-shot示例检验模型稳健性。"}},{"id":"2608.30311","version":1,"title":"One AI Signal, Many Human Judgments: A Bayesian Cascade Analysis of AI-based Credibility Indicators in Online Information Spread","zh_title":"一个AI信号，多种人类判断：在线信息传播中基于AI的可信度指标的贝叶斯级联分析","abstract":"Social media platforms increasingly use AI-based credibility indicators to help users judge misinformation. Unlike individual human-AI decision-making, these indicators are embedded in information spread: users see both an AI prediction and earlier judgments shaped by the same AI, and their own judgments may then enter the public history. Yet how to analytically characterize this process remains under-explored. We therefore introduce a social-learning lens for this setting by extending the classical Bayesian cascade model with the AI indicator as a shared public signal. The resulting Gateway condition compares the evidence from the AI prediction with users' private impressions. Through this view, we show that AI changes what public history means. Crowd agreement may reflect accumulated independent human evidence, or repeated dependence on the same AI prediction. This creates a preservation-correction trade-off: stronger reliance on AI can preserve correct predictions, but can also lock in incorrect ones by blocking corrective private impressions. We calibrate the model using human-subject data on news veracity judgments. Although the AI outperforms human users, the average user weights it below her own impression but above several peer judgments, while individual users vary from discounting the AI to relying on it enough to cascade. Simulations show that over-reliance on a weak AI is especially harmful, and that diversifying AI signals across users can better keep the crowd informative. We conclude with implications for understanding human-AI interaction in information spread and designing misinformation interventions.","authors":["Zhuoran Lu","Weilong Wang","Yangyang Yu","Xinru Wang","Zhuoyan Li","Zhiwei Liu","Sophia Ananiadou"],"categories":["cs.HC","cs.AI","cs.SI"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-01","first_seen":"2026-09-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.30311","pdf_url":"https://arxiv.org/pdf/2608.30311","source_feed":"cs.HC","score":7,"bucket":"pending","rubric_hits":["A3","B1","B2","B4"],"tags":["人类-AI交互","社会学习","信息传播"],"reason":"用贝叶斯级联模型分析AI信号对人类判断的影响，校准于人类数据，涉及信息传播和政…","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:02:39","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-01","rank":31,"question":"AI 可信度信号如何通过社会学习过程影响用户对新闻真实性的序列判断，并改变公共历史信息的意义？","design":"扩展经典贝叶斯级联模型，将 AI 可信度信号作为共享公共信号，模拟用户在信息传播中的序列判断；模型用人类被试数据校准，并模拟不同 AI 强度和行为权重下的级联动态。","baseline":"使用先前人类被试实验数据，用户在 AI 指标存在下判断新闻真实性，AI 表现优于人类用户。","findings":"AI 信号改变了公共历史的含义，人群一致可能反映独立人类证据或对同一 AI 信号的重复依赖；存在保留-纠正权衡：强依赖 AI 可保留正确预测，但也会锁定错误预测。校准显示平均用户对 AI 的权重低于自身印象但高于多个同伴判断，个体差异显著；模拟表明过度依赖弱 AI 尤其有害，多样化 AI 信号可保持人群信息量。","reliability":"论文承认模型基于贝叶斯理性假设，但实际用户存在行为偏差；校准数据来自特定实验设置，可能限制外部效度；未考虑 AI 信号随时间变化或用户异质性对级联的长期影响。","relevance":"该研究用贝叶斯级联模型分析 AI 信号对人类判断的影响，校准于真实人类数据，涉及信息传播和政策评估场景，对关注 LLM 仿真可靠性与偏差的研究者有直接参考价值。","inspiration":"借鉴其将 AI 信号作为共享公共信号嵌入社会学习模型的方法，可迁移到资产定价实验或政策公告预期形成场景；例如用 LLM 模拟投资者在 AI 投资建议下的序列决策，处理为不同 AI 信号强度，结果变量为投资选择与价格动态，对照真实实验数据。"}},{"id":"2608.28597","version":1,"title":"The Race between Agentic AI Capabilities and Data Quality Control in Online Surveys","zh_title":"在线调查中代理式AI能力与数据质量控制之间的竞赛","abstract":"Online surveys are a foundational data collection instrument in a variety of fields, with attention checks serving as critical guardians of response quality. However, the rapid emergence of agentic AI (goal directed systems powered by a large language model (LLM) brain and/or a multimodal processing unit with tool-augmented capabilities) raises new questions about the robustness of these safeguards. We investigate how well agentic AI architectures can complete web-based surveys and pass standard attention checks. We evaluate a single-agent architecture capable of multimodal input processing and tool-based web interaction on a controlled survey sandbox. We analyze the problem from two perspectives. From an attack perspective, we demonstrate how structural vulnerabilities such as exposed DOM metadata and predictable option encoding allow agents to resolve attention checks through structured parsing only. From a defense perspective, we implement a mitigation strategy of DOM metadata obfuscation to remove semantic cues in text-based questions. We evaluate multiple open-source language and multimodal models to study capability and orchestration effectiveness. Based on our evaluations, we offer perspectives on how to simultaneously meet the needs of empiricists and agentic AI researchers.","authors":["Sourav Panda","Hillmer Chona","Rupak Kumar Das","Shreyash Kale","Shikha Soneji","Jonathan Dodge"],"categories":["cs.AI","cs.CY"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-09-01","first_seen":"2026-09-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.28597","pdf_url":"https://arxiv.org/pdf/2608.28597","source_feed":"cs.CY","score":7,"bucket":"pending","rubric_hits":["A2","B4"],"tags":["LLM代理","调查数据质量","注意力检查"],"reason":"研究LLM代理完成在线调查及注意力检查，评估数据质量与防御，与仿真可靠性相关。","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:02:30","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-01","rank":23,"question":"代理型AI在多大程度上能自主完成在线调查并通过注意力检查，以及如何防御这种自动化参与以保障调查数据质量？","design":"构建单代理架构，集成开源LLM（LLaMA-3-8B、Mistral-7B、Qwen-2.5-7B）进行多模态输入处理和基于工具的网页交互，在受控调查沙盒中执行多页调查任务；设置完成策略（严格模式）和无约束策略（自由模式）两种行为策略，每种模式运行100次，测量端到端提交成功率、页面级准确率和失败分解。","baseline":"无对照","findings":"代理AI能利用DOM元数据等结构漏洞，通过结构化解析而非语义理解来通过注意力检查；所有模型在两种行为策略下均表现出较高的端到端成功提交率，表明现有注意力检查可能无法有效区分人类和机器响应。","reliability":"论文未讨论","relevance":"该研究直接评估LLM代理在调查场景中通过注意力检查的能力，并探讨防御策略，与研究者关注的LLM仿真可靠性及失效条件高度相关，值得阅读原文以了解具体漏洞利用方式和防御效果。","inspiration":"借鉴其攻击-防御双视角设计和结构化漏洞分析方法，可用于评估经济实验中LLM代理是否通过界面元数据而非真实决策来通过操纵检查｜可迁移到在线经济实验或调查中的注意力检查有效性评估，例如公共品博弈中的理解性问题或风险偏好问卷｜以LLM代理为被试，施加DOM元数据混淆处理，测量其通过注意力检查的准确率和决策一致性，并与真实人类被试在相同实验中的表现进行对照。"}},{"id":"2608.07438","version":2,"title":"PsychoAgent: An Affect-Sensitive Cognitive Architecture for Conflict-Aware Memory in LLM Agents","zh_title":"PsychoAgent：一种面向LLM智能体的情感敏感认知架构，用于冲突感知记忆","abstract":"Human-like cognition does not select past experience by topical similarity alone: affective significance and unresolved conflict also shape what becomes accessible. We present PsychoAgent, a cognitive architecture for LLM agents that separates factual and affective memory and integrates both through a conflict-aware executive controller. Affective memories are first filtered by semantic relevance and then re-ranked by salience, preserving topical fit while allowing emotionally important traces to enter the prompt. Across three controlled conflict scenarios, the full architecture retrieved more conflict-critical memories than semantic-affective and single-memory RAG baselines (0.933 vs. 0.500 and 0.667), with a small semantic-similarity cost. Five blinded raters evaluated 27 outputs. After within-rater standardization, the full architecture had the highest overall mean (+0.22 SD), but corrected pairwise differences were not significant. A three-day illustrative trace further shows persistent affect, offline memory recombination, and selective memory reweighting. The findings support affect-sensitive retrieval as an inspectable mechanism for modeling human-like conflict effects in LLM agents.","authors":["Mohammad Amanlou","Parham Abed Azad","Farbod Davoodi","Mostafa Masumi","Behnam Bahrak","Abdol-Hossein Vahabie"],"categories":["cs.AI","cs.CL","cs.HC"],"primary_category":"cs.AI","announce_type":"replace-cross","date":"2026-09-01","first_seen":"2026-08-10","revised_at":"2026-09-01","abs_url":"https://arxiv.org/abs/2608.07438","pdf_url":"https://arxiv.org/pdf/2608.07438","source_feed":"cs.CL","score":6,"bucket":"other","rubric_hits":["D3"],"tags":["LLM智能体","认知架构","情感记忆"],"reason":"用LLM agent模拟人类冲突记忆，但无真实人类数据对照，属社会模拟边界情形。","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:03:23","error":null,"has_summary":false,"summary":null},{"id":"2608.20983","version":2,"title":"Beyond Truth Discovery: A Two-Stage Framework to Assess the Severity of False Claim during Disasters","zh_title":"超越真相发现：评估灾害期间虚假声明严重性的两阶段框架","abstract":"False information spreads rapidly on social media during disasters and can undermine emergency response efforts, public trust, and crisis communication. Existing research primarily focuses on determining whether social media posts contain false information, but provides limited insight into the specific false claims embedded within posts and the severity of individual false claims. To address the limitations, we propose a two-stage framework to assess the severity of false claims during disasters. In the first stage, we develop a false claim extraction agent that identifies false claims from multimodal social media posts containing text, images, videos, and links. A subsequent verification step validates extracted claims with supporting evidence. In the second stage, we define false claim severity as the combination of two complementary dimensions: believability, which determines the likelihood that a claim will be believed, and harmfulness, which captures the potential consequences if it is believed. Human annotators assess both dimensions to construct a claim-level severity benchmark using false claims extracted from Reddit posts related to hurricanes and wildfires. Building upon this benchmark, we investigate false claim severity assessment as a human-AI alignment problem, evaluating whether models can reproduce human judgments under a shared evaluation rubric rather than merely predicting severity labels. Experiments on the benchmark show that traditional supervised models exhibit limited alignment with human judgments, whereas Large Language Models (LLMs) achieve substantially stronger performance. Among the evaluated strategies, in-context learning consistently achieves the strongest alignment with human judgments, highlighting the importance of human examples and shared decision criteria for severity assessment.","authors":["Ruichen Yao","Tejna Dasari","Gulshat Baispay","Aizhan Zaurbek","Yifan Liu","Yaokun Liu","Zelin Li","Dong Wang"],"categories":["cs.SI","cs.HC"],"primary_category":"cs.SI","announce_type":"replace-cross","date":"2026-09-01","first_seen":"2026-08-24","revised_at":"2026-09-01","abs_url":"https://arxiv.org/abs/2608.20983","pdf_url":"https://arxiv.org/pdf/2608.20983","source_feed":"cs.HC","score":6,"bucket":"other","rubric_hits":["D1"],"tags":["虚假信息","人机对齐","标注替代"],"reason":"用LLM替代人工标注严重性，但非仿真人类被试，属标注替代","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:02:42","error":null,"has_summary":false,"summary":null},{"id":"2608.29209","version":1,"title":"Toward Cultural Alignment: Human-Centered Evaluation of Multimodal AI Stories Across Five African Communities","zh_title":"迈向文化对齐：对五个非洲社区的多模态AI故事进行以人为中心的评估","abstract":"In this paper, we examine how well AI-generated multimodal stories align with the lived practices, relationships, language, values, and visual expectations of the communities they represent. We conduct a community-grounded mixed-methods evaluation with 19 culture representatives across five African communities, combining quantitative annotations with qualitative focus group discussions. We find that cultural alignment depends not simply on recognizable cultural markers, but on how those markers fit social, linguistic, procedural, and visual context. From these evaluations, we develop a taxonomy of cultural alignment comprising five broader cultural marker categories and eight recurring mechanisms of misalignment. We additionally evaluate five multimodal LLM judges to examine whether automated evaluation can approximate community-grounded judgments at scale. Judge reliability and score calibration vary substantially across communities, with no single judge performing consistently across all five settings. These findings motivate community-calibrated evaluation pipelines in which automated judges are validated against community judgments to determine where they can be trusted and where human review remains necessary.","authors":["Millicent Ochieng","Felermino D. M. A. Ali","Elizabeth A. Ankrah","Najeeb Gambo Abdulhamid","Migisha Boyd","Stephanie Nyairo","Mercy Muchai","Samuel Chege Maina","Aditya Vashistha","Anja Thieme","Jacki O'Neill"],"categories":["cs.CL","cs.HC"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-01","first_seen":"2026-09-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.29209","pdf_url":"https://arxiv.org/pdf/2608.29209","source_feed":"cs.CL","score":6,"bucket":"other","rubric_hits":["D1","B1"],"tags":["文化对齐","多模态评估","LLM评判"],"reason":"用LLM评估文化对齐，有社区人类判断对照，但非仿真人类被试，而是替代标注员。","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:02:54","error":null,"has_summary":false,"summary":null},{"id":"2608.29517","version":1,"title":"LLM Judges as Raters: A Pre-Registered Audit of Severity, Halo, Reliability, and Version Instability in LLM Essay Scoring on Public Corpora","zh_title":"LLM法官作为评分者：对公共语料库中LLM作文评分的严重性、晕轮效应、可靠性和版本不稳定性的预注册审计","abstract":"Large language models (LLMs) are increasingly used as essay graders in learning analytics, evaluated almost exclusively with agreement statistics. Educational measurement warns that raters also differ in severity, show halo, and drift as instruments. We treat LLM judges as raters and run a pre-registered rater-effects battery (many-facet Rasch severity, residual halo, generalizability/decision studies, cross-version shifts, differential functioning) on public corpora in two languages (ENEM/Essay-BR; ASAP): 2,377 essays, 12 judges, 4 providers, 5 version contrasts, replicated cells, released as a score tensor. Judge severity spans 219 points on ENEM's 0-1000 scale; on ASAP the panel spread is 15-33% of the score range against a between-trained-human gap near 1%. Judge-human correlations sit in an undiscriminating .47-.56 band. All five version contrasts shift severity beyond a family-wise permutation null (up to 133 points), and one judge was deprecated mid-study, caught by identity canaries. Two pre-registered tests returned honest nulls: severity-adjusted leaderboard reversals did not survive a permutation null, and \"silent drift\" was refuted: agreement moved with severity in four of five contrasts. Replication yields self-consistency (phi>=.80 at k<=2) but not human-level accuracy, and a same-instrument check overturned our own halo comparison: matched on instrument and calibration, we find no credible evidence that judge halo exceeds the trained-human range.","authors":["Veerendra Kumar Sunkavalli"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-01","first_seen":"2026-09-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.29517","pdf_url":"https://arxiv.org/pdf/2608.29517","source_feed":"cs.CL","score":6,"bucket":"other","rubric_hits":["D1"],"tags":["LLM评分","评分者效应","教育测量"],"reason":"LLM作为评分者替代人工评分，属于标注员替代，非仿真人类被试，但涉及评分偏差与…","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:02:36","error":null,"has_summary":false,"summary":null},{"id":"2608.30373","version":1,"title":"Beyond Consensus: Downward Bias and Role Asymmetry in Multi-Agent LLM Judges for Subjective Evaluation","zh_title":"超越共识：多智能体LLM评判在主观评估中的向下偏差与角色不对称性","abstract":"Multi-Agent Debate (MAD) has been widely adopted to improve LLM-based evaluation by prompting multiple agents to negotiate and reach a consensus. However, for subjective rubric-based scoring, inter-agent agreement does not guarantee alignment with human judgments. In this paper, we compare a single-judge baseline against a consensus-based MAD protocol on subjective evaluation tasks and design three ablations to isolate the impact of role prompting, multi-round interaction, and explicit score sharing. Evaluations across six LLMs show that the single-judge baseline achieves the strongest human alignment on average across six judge models, whereas MAD shows degradation in human alignment on both tasks. Our ablations demonstrate that this performance drop stems primarily from asymmetric role prompting rather than the interaction itself. Specifically, assigning a strict judge role introduces a systematic downward bias that the consensus process fails to correct. The central finding is that this bias reflects strict-stance dominance beyond averaging: the consensus score falls well beyond the arithmetic midpoint of the standalone strict and lenient conditions, rather than averaging them out. Removing role asymmetry (Symmetric MAD) largely recovers baseline performance, while masking peer scores widens inter-agent disagreement on average and worsens average human alignment. These findings demonstrate that multi-agent consensus can enforce artificial agreement at the expense of true human alignment, revealing a structural limitation in consensus-style, role-specialized MAD protocols for subjective scoring.","authors":["Minsoo Song","Chanwoo Kim","Sugyeong Eo","Chanjun Park"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-01","first_seen":"2026-09-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.30373","pdf_url":"https://arxiv.org/pdf/2608.30373","source_feed":"cs.CL","score":6,"bucket":"other","rubric_hits":["D1"],"tags":["LLM评估","多智能体","人类对齐"],"reason":"LLM作为评估者替代人工评分，属于标注替代而非仿真人类被试，但涉及人类对齐与偏…","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:03:09","error":null,"has_summary":false,"summary":null},{"id":"2608.29198","version":1,"title":"How Identity and Opinion Shape Political Sycophancy in LLMs","zh_title":"身份与观点如何塑造大语言模型的政治谄媚行为","abstract":"As Large Language Models (LLMs) increasingly encourage users to disclose personal profiles for tailored assistance, measuring their political alignment becomes increasingly important. However, many existing benchmarks for assessing political behavior rely on closed-ended questions and do not fully capture how a model's stance may adapt to user-provided context during interaction. We introduce a framework that disentangles two distinct triggers of political sycophancy: opinion (aligning with explicit narratives) and identity (stereotyping based on demographic labels). Using 450 manually-checked political dilemmas as controlled probes, we evaluate 13 instruction-tuned LLMs. We uncover a dissociation: a model's susceptibility to explicit opinions does not necessarily predict its susceptibility to identity cues, and vice versa. When both signals are present, their effects are generally sub-additive rather than simply additive. Additionally, system-level personas primarily shift a model's baseline stance while having limited effect on the stance shift caused by user opinion or identity. Ultimately, our results suggest that LLM political stance is interactively and steerably vulnerable rather than being a fixed trait, highlighting how personalization may amplify identity- or opinion-conditioned shifts in the model's behaviors.","authors":["Li-Ni Fu","Chang-Chih Meng","Chien-Hua Chen","Hen-Hsen Huang","I-Chen Wu"],"categories":["cs.AI","cs.CL","cs.CY"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-09-01","first_seen":"2026-09-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.29198","pdf_url":"https://arxiv.org/pdf/2608.29198","source_feed":"cs.CL","score":6,"bucket":"other","rubric_hits":["D2"],"tags":["LLM政治立场","谄媚行为","模型偏差"],"reason":"测量LLM政治立场对用户身份和观点的反应，属于模型态度测量，非仿真人类被试，但…","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:02:33","error":null,"has_summary":false,"summary":null},{"id":"2608.28989","version":1,"title":"Using LLMs to Mimic the Conversational Dynamics of Reddit Communities","zh_title":"使用大语言模型模仿Reddit社区的对话动态","abstract":"Online communities face a constant battle against toxicity and misinformation. While human moderators struggle to keep pace with the volume of content, LLMs offer a promising solution for automatically generating constructive responses and shaping online interactions. This paper preliminarily investigates if LLMs can mimic the communication styles of Reddit users using their comment history as context. We evaluate two prompting approaches: predicting a target comment and filling in masked comments. We find that LLMs outperform expectations at replicating comment structure and formality, but struggle to accurately capture nuanced emotions, e.g. understating joy and overstating anger. These findings highlight a promising direction for LLMs in guiding online conversations towards prosociality influencing emergent communication patterns and norms within the community. The results of our study inspire future work with more rigorous methods of evaluation to explore the LLMs' effectiveness across diverse online communities to better understand their broader societal impact.","authors":["Vedaant Jain","Yoshee Jain","Ishq Gupta","Aditi Shrivastava","Koustuv Saha","Eshwar Chandrasekharan"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-01","first_seen":"2026-09-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.28989","pdf_url":"https://arxiv.org/pdf/2608.28989","source_feed":"cs.HC","score":6,"bucket":"other","rubric_hits":["D3"],"tags":["LLM仿真","社交媒体","风格模仿"],"reason":"用LLM模仿Reddit用户评论风格，属于社会模拟但无真实人类数据对照，且目的…","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:02:32","error":null,"has_summary":false,"summary":null},{"id":"2607.26178","version":2,"title":"DuplexGen: Adaptive Synthesis of Human-AI Turn-Taking Dialogues","zh_title":"DuplexGen：自适应合成人机轮流对话","abstract":"Turn-taking is a central component of full-duplex interaction. Which turn-taking behaviors are appropriate varies with the scenario, yet current models apply a single norm regardless of context. This limitation originates in their training data: human-human speech corpora capture natural timing phenomena but provide little role grounding or scenario-specific norms, while heuristic or prompted synthesis methods inject turn-taking behaviors without basing them on human preferences. We introduce DuplexGen, a framework for generating dialogues with scenario-adaptive turn-taking by calibrating LLM predictions against a small set of slot-level human preference annotations. In six cooperative and competitive tasks, human turn-taking preferences differ systematically, and DuplexGen aligns substantially more closely with those preferences than uncalibrated prompting or training solely on generic human-human data; a full-duplex model trained on DuplexGen-generated data exhibits distinctive, human-preferred turn-taking behaviors. These results show that human calibration, not corpus scale or prompt design alone, is what allows turn-taking synthesis to be scenario-specific.","authors":["Takyoung Kim","Kang-wook Kim","Sang Hoon Woo","Julia Hirschberg","Gunhee Kim","Dilek Hakkani-T\\\"ur"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-01","first_seen":"2026-07-30","revised_at":"2026-09-01","abs_url":"https://arxiv.org/abs/2607.26178","pdf_url":"https://arxiv.org/pdf/2607.26178","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["对话生成","人类偏好校准","轮流行为"],"reason":"用人类偏好校准LLM生成对话，属于用LLM替代人工标注，非仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:03:22","error":null,"has_summary":false,"summary":null},{"id":"2608.24080","version":2,"title":"When Less Is More: An Empirical Study of Minimal Responses in Counseling Dialogues and the Behavior of LLMs","zh_title":"少即是多：咨询对话中最小回应及LLM行为的实证研究","abstract":"In psychological counseling, effective support is not always delivered through long, information-rich responses. Minimal responses, such as backchannel cues and concise empathic statements, help convey attentive listening, express empathy, and encourage clients to continue expressing themselves. However, existing counseling dialogue systems and evaluation frameworks often favor explicit, content-rich replies, overlooking the interactional value of brief counselor utterances. This paper presents a systematic cross-lingual analysis of minimal responses across multiple counseling dialogue datasets. We develop a two-stage filtering method based on utterance length and content, followed by contextual verification using a large language model (LLM). Our analysis shows that minimal responses are common in human-collected datasets but substantially underrepresented in LLM-generated ones. We further evaluate current LLMs in manually curated dialogue contexts where human counselors used minimal responses. The results show that strong commercial LLMs are capable of generating minimal responses when explicitly instructed, but still struggle to determine when such responses are appropriate. Counseling-specific models trained on synthetic data perform particularly poorly, tending instead to produce longer and more information-rich responses. Moreover, LLM-based response-quality evaluation may undervalue minimal responses, even when they are interactionally appropriate.","authors":["Zhiyang Qi"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-01","first_seen":"2026-08-26","revised_at":"2026-09-01","abs_url":"https://arxiv.org/abs/2608.24080","pdf_url":"https://arxiv.org/pdf/2608.24080","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM评估","心理咨询对话","最小回应"],"reason":"LLM用于评估对话质量，替代人工标注，非仿真人类被试","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:03:24","error":null,"has_summary":false,"summary":null},{"id":"2608.26123","version":2,"title":"Which India Survives Translation? Narrative Homogenisation Across Indian Oral Traditions in LLMs","zh_title":"哪个印度在翻译中幸存？LLM中印度口头传统的叙事同质化","abstract":"Large language models (LLMs) are trained predominantly on English-language internet text that over-represents certain cultural narratives, raising concerns that models flatten the diversity of non-Western storytelling traditions into a single homogenized archetype. We present a pilot computational study examining this across three maximally distinct Indian regional oral and literary traditions: the Rajasthani Pabuji epic, classical Tamil Sangam poetry, and Bengali folk tales. We collected authentic reference corpora for each tradition (11, 21, and 10 passages respectively) and prompted two LLMs (Claude Sonnet and Gemini) with 54 generation requests spanning three prompt types per tradition - generic, culturally specific, and regional-language. Using Sentence-BERT embeddings and cosine similarity, we measure reference drift (how closely outputs track their own tradition's authentic texts relative to the other two) and cross-tradition convergence (how similar outputs are across traditions). We find that while outputs remain closer to their own tradition's reference than to others, cross-tradition similarity is high (0.52-0.66) relative to what the traditions' genuine distance would predict, indicating partial homogenisation. Unexpectedly, prompting in the regional language (Hindi, Tamil, or Bengali) consistently reduced fidelity to the authentic tradition relative to English prompting, by as much as 27 percentage points for Rajasthani and Bengali traditions. We discuss this against conflicting prior results on multilingual prompting and argue it reflects a difference between eliciting general cultural diversity and simulating one narrow, lesser-documented oral tradition. We position this pilot as a lightweight, scalable complement to recent large-scale human-annotation studies of Indian cultural misrepresentation in LLM-generated stories, as part of a broader doctoral research program.","authors":["Paarth Singh Rathore"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-01","first_seen":"2026-08-28","revised_at":"2026-09-01","abs_url":"https://arxiv.org/abs/2608.26123","pdf_url":"https://arxiv.org/pdf/2608.26123","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["文化表征","叙事同质化","LLM偏差"],"reason":"研究LLM生成的文化叙事同质化，测量模型输出而非人类行为，无人类被试仿真。","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:03:25","error":null,"has_summary":false,"summary":null},{"id":"2608.28649","version":1,"title":"Can Large Language Models Identify Meaningful Touchpoints in Conversion Attribution?","zh_title":"大语言模型能否识别转化归因中有意义的触点？","abstract":"Touchpoint selection in conversion attribution, namely identifying meaningful touchpoints contributing to conversions, is essential for e-commerce recommendation and online advertising. Current selection methods rely heavily on collaborative-filtering-based heuristics, which fail to align with user-perceived semantic intent. Through human annotation, we reveal a significant semantic gap: many implicitly-related, semantically relevant touchpoints remain undetected by existing rules. Therefore, we systematically evaluate the capability of Large Language Models (LLMs) in identifying these hidden associations. Our evaluation shows that while LLMs effectively uncover a substantial portion of implicitly-related touchpoints, significant room for improvement remains in their selection performance. Furthermore, we analyze the impact of different prompting strategies and foundation model choices on identification performance, providing valuable insights into their reasoning patterns and effectiveness. These insights offer a new roadmap for transitioning conversion attribution from mechanical rule-matching to human-aligned semantic reasoning. Moreover, we leverage the LLM-attributed conversion labels for enhancing industrial CVR model training and achieve significant offline performance gains, showing the potential of LLMs in conversion attribution.","authors":["Jinqi Wu","Sishuo Chen","Zhangming Chan","Yong Bai","Chao Yi","Han Zhu","Shuodian Yu","Lei Zhang","Sheng Chen","Chenghuan Hou","Jian Xu","Chaoyou Fu"],"categories":["cs.CL","cs.AI","cs.IR"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-01","first_seen":"2026-09-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.28649","pdf_url":"https://arxiv.org/pdf/2608.28649","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["转化归因","LLM标注","语义关联"],"reason":"LLM用于识别转化归因中的触点，替代人工标注，非仿真人类被试，但涉及人类标注对…","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:02:48","error":null,"has_summary":false,"summary":null},{"id":"2608.29591","version":1,"title":"How You Ask Shapes What You Get: A Theory-Seeded Measurement of Articulation in Advice-Seeking LLM Conversations","zh_title":"提问方式塑造回答：基于理论种子测量建议寻求型LLM对话中的表达清晰度","abstract":"Users articulate the same advice-seeking request in different ways: some specify detailed constraints, others gesture at a vague need. Prior work treats this variation as noise to be averaged away; we instead treat it as a stable, measurable structure in the input distribution. We ask whether articulation (how people ask) forms latent dimensions separable from topic (what they ask about), and whether it is associated with how language models respond. We extract interpretable features from 16,447 advice-seeking prompts pooled from public chat corpora (WildChat, LMSYS, and ShareChat) and recover a small set of latent articulation factors that replicate across train/test splits and across corpora. Because this structure is largely separable from topic, the populations it defines cut across topics and stay invisible to topic- or task-based evaluation. The factors define a handful of recurring articulation styles, one of which stands out: a long-form but information-poor style, roughly one in six prompts in the largest corpus, where models return shorter, vaguer answers and do not ask for clarification even though under-specification is exactly the condition that warrants it. The contrast holds within every topic group and length quintile, and is not under-specification alone -- a second, equally under-specified style does draw clarifying questions. Two independent human annotators reproduce this contrast. We argue that benchmarks should stratify on articulation, and we offer the extracted structure as a measurement instrument for doing so.","authors":["Juneha Baek","Suhyeon Lee","Donghyuk Shin"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-01","first_seen":"2026-09-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.29591","pdf_url":"https://arxiv.org/pdf/2608.29591","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["LLM行为分析","提示词风格","对话测量"],"reason":"研究LLM对用户提问方式的响应，测量的是模型行为而非人类被试仿真，无人类数据对…","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:02:58","error":null,"has_summary":false,"summary":null},{"id":"2608.29738","version":1,"title":"Evaluating the Capabilities of LLMs for Persuasive Dialogue","zh_title":"评估大语言模型在说服性对话中的能力","abstract":"Large language models (LLMs) can generate apparently highly persuasive text, but does sounding persuasive mean arguing well? We introduce \\textsc{Persuasio}, a multi-agent dialogue platform grounded in a formal argumentation-based theory of persuasion dialogues that adjudicates logical winners during free-text debates. Using this system, we generated 192 debates on a UK political topic between humans and LLMs, and evaluated 22 interlocutors through both automated adjudication and 9,702 crowdsourced pairwise judgements across 1{,}386 annotation instances. We observed a consistent decoupling between subjective and formal persuasiveness: LLMs dominated the subjective ranking yet performed substantially worse under argumentation-theoretic adjudication, where humans remained competitive. Multi-agent and retrieval-augmented variants further widened this divergence. These findings reveal a systematic gap between rhetorical fluency and formal argumentative strength in LLM-based persuasive dialogues.","authors":["Jordan Robinson","Angus R. Williams","Katie Atkinson","Anthony G. Cohn"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-01","first_seen":"2026-09-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.29738","pdf_url":"https://arxiv.org/pdf/2608.29738","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["LLM辩论","论证理论","人机对比"],"reason":"LLM参与辩论并与人类对照，但非仿真人类被试，而是评估论证能力","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:02:58","error":null,"has_summary":false,"summary":null},{"id":"2608.30224","version":1,"title":"The Differential Reasoning Router: Operationalizing Cost-Aware LLM Annotation in E-commerce","zh_title":"差分推理路由器：在电商中实现成本感知的LLM标注","abstract":"Large Language Models (LLMs) are increasingly used to annotate structured product data in e-commerce, but early deployment often begins as a cold-start problem: only limited pre-launch labels are available, the value of expensive reasoning is unknown, and human review is needed before the system can be trusted at scale. This challenge is especially common in rule-based annotation workflows, where each item must satisfy multiple business rules and both model errors and ambiguous rule boundaries affect final decisions. We introduce the Differential Reasoning Router (DRR), a cost-aware framework for cold-start LLM annotation that jointly optimizes model selection and human escalation. Rather than treating a reasoning model as a default fallback, DRR estimates separate success probabilities for a direct model and a reasoning model at both the sample and business-rule levels, enabling adaptive routing: easy cases are handled directly, reasoning is reserved for cases where it is expected to improve the decision, and likely double-failure or rule-disagreement cases are escalated to human annotators. The resulting labels provide targeted ground truth for prompt engineering, supervised fine-tuning, calibration, and rule refinement, enabling a gradual shift from human-heavy cold-start annotation toward high-confidence automated routing. In a production e-commerce workflow, DRR reaches accuracy parity with the strongest confidence-based router while achieving more than 60\\% reasoning-token cost savings.","authors":["Cheng Lyu","Jingyue Zhang","Vinny DeGenova","Mengwei Li","Yuanli Pei"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-01","first_seen":"2026-09-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.30224","pdf_url":"https://arxiv.org/pdf/2608.30224","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM标注","成本优化","人机协作"],"reason":"LLM替代人工标注，非仿真人类被试，但涉及成本优化与人工升级，属边界情形","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:03:06","error":null,"has_summary":false,"summary":null},{"id":"2608.30485","version":1,"title":"Two Centuries of Sexism in British Parliament: A Computational Analysis of Women's Representation in the Hansard Corpus","zh_title":"英国议会两个世纪的性别歧视：对汉萨德语料库中女性代表性的计算分析","abstract":"The language a legislature uses to debate women's rights, even in favour of them, encodes systematic patterns of sexism that persist across two centuries. In this work, we analyse 6,531 speeches over 200 years of UK parliamentary debate (Hansard, 1803-2005) by using large language models to classify a speaker's perspective towards women's suffrage and political representation, as well as analyse sexist speech in parliament from the lens of the Ambivalent Sexism Inventory. We also release this parliamentary dataset, an organized and metadata-enriched version of the publicly available Hansard Corpus optimized for computational social science research, with 6.7 million speeches across 1.2 million debates, with 89% gender-matching for speeches by MPs from the House of Commons. We find that 54% of speeches opposing women's representation contain sexist content, compared to 21% of speeches that are for the cause, and that the two sides use fundamentally different types of sexism: anti-suffrage rhetoric combines hostile and benevolent framing, while pro-suffrage sexism is overwhelmingly benevolent. Female MPs support women's political rights at 93% compared to 70% for male MPs, a gap that closes only after enfranchisement. Our findings are evidence that benevolent and hostile sexism are used in different rhetorical contexts in a manner consistent with the theory of Ambivalent Sexism.","authors":["Mohammad Omar Khursheed","Mandira Sawkar","Ashiqur R. KhudaBukhsh"],"categories":["cs.CL","cs.CY","cs.LG"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-01","first_seen":"2026-09-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.30485","pdf_url":"https://arxiv.org/pdf/2608.30485","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM标注","性别歧视","议会辩论"],"reason":"用LLM做文本分类和标注，替代人工编码，非仿真人类被试","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:03:11","error":null,"has_summary":false,"summary":null},{"id":"2608.30683","version":1,"title":"WildSEEK: Evaluating Language Models for Information-Seeking","zh_title":"WildSEEK：评估语言模型的信息寻求行为","abstract":"Language models are increasingly mediating information access to end users, urging a systematic evaluation of their responses for a fair and reliable information ecosystem. Existing evaluations, however, are often topic-specific or synthetic, limiting their ability to capture the complexity of \"in the wild\" information-seeking queries and the risks present in model responses. To address this gap, we introduce WildSEEK, a manually annotated dataset of 3k information-seeking queries from real user interactions, and an evaluation framework for LLM-generated responses. WildSEEK includes annotations for risk-sensitive domains (e.g. health and financial information), and distinguishes factoid queries from analytical queries which seek responses beyond facts. We train classifiers on WildSEEK to analyze more than 1.8M realistic user queries. We find that over a third of information-seeking queries are high-risk and more often analytical. Our findings show that LLM responses fail more often in four criteria: sycophantic behavior, overreliance, a default US-centric perspective, and poor handling of vulnerable populations -- with failure rates being mostly higher for analytical queries. By providing methods to monitor the reliability, safety, and fairness of LLM behavior, our dataset and evaluation framework offer an empirical foundation for the broader question of how these systems should behave as they take on a growing role in information access.","authors":["Tanise Ceron","Joachim Baumann","Elisa Bassignana","Berat Cabuk","Dirk Hovy","Debora Nozza"],"categories":["cs.CL","cs.CY"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-01","first_seen":"2026-09-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.30683","pdf_url":"https://arxiv.org/pdf/2608.30683","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["LLM评估","信息寻求","风险与公平"],"reason":"评估LLM在信息寻求中的行为，但非仿真人类被试，而是测量模型本身，属边界情形。","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:03:11","error":null,"has_summary":false,"summary":null},{"id":"2608.30754","version":1,"title":"CLIN: an Objective Framework for Evaluating Creativity in Short Persian Literary Text","zh_title":"CLIN：评估波斯语短文学文本创造力的客观框架","abstract":"Evaluating creativity in large language model (LLM) outputs remains challenging because creativity is multidimensional and human-centered. We examine how reliably LLMs evaluate short literary text in Persian, a low-resource language, across multiple evaluation strategies and prompt formulations. We find that LLM-human agreement varies substantially across dimensions: alignment is stronger for structured TTCT-derived properties such as Originality, Fluency, and Elaboration, but considerably weaker for more subjective dimensions, particularly Emotion and Attractiveness. Judgments are also sensitive to prompt formulation, while few-shot prompting, ensembling, and multi-agent debate provide no consistent improvement. Motivated by this dimension-dependent behavior, we investigate whether structured creativity dimensions can instead be approximated using simple, interpretable proxy metrics. We introduce CLIN, which evaluates three TTCT-derived dimensions separately using topic-aware novelty for Originality, contextual lexical clustering for Fluency, and lexical diversity for Elaboration. These proxies achieve human alignment comparable to or better than the strongest zero-shot LLM judge in our setting while requiring substantially lower evaluation cost.","authors":["Mohammad Reza Modarres","Armin Tourajmehr","Yadollah Yaghoobzadeh","Mohammad Taher Pilehvar"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-01","first_seen":"2026-09-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.30754","pdf_url":"https://arxiv.org/pdf/2608.30754","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["LLM评估","创造力测量","低资源语言"],"reason":"评估LLM创造力，测模型而非仿真人类被试，但涉及人类判断对齐，属边界情形。","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:03:13","error":null,"has_summary":false,"summary":null},{"id":"2608.30842","version":1,"title":"Thesis Proposal: Toward a Human-Centered and Perspective-Aware Framework for Reproducible ML Evaluation and AI Alignment","zh_title":"面向可复现机器学习评估与AI对齐的以人为本、视角感知框架","abstract":"Humans play a vital role at every stage of AI development, from data collection and curation to model development and evaluation. However, humans often disagree with each other and sometimes with themselves over time. It is essential to take disagreement into account when building human-centered AI systems, especially in domains where it is prevalent, such as AI safety, content moderation, or sentiment analysis. Disagreement often arises from subjective human opinion and can vary with one's identity, beliefs, and social environment. Despite this, current LLM evaluation approaches frequently rely on aggregating labels (often via plurality voting) to represent consensus, thereby obscuring minority perspectives. By failing to account for human disagreement, these evaluation methods contribute to the reproducibility crisis in AI. Human feedback is also crucial for ensuring that AI systems align with human values. For these systems to be trustworthy, it is critical to ensure that they reflect diverse human values and perspectives. In this thesis proposal, we present a human-centered and perspective-aware framework for reproducible ML evaluation and AI alignment.","authors":["Deepak Pandita","Christopher M. Homan"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-01","first_seen":"2026-09-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.30842","pdf_url":"https://arxiv.org/pdf/2608.30842","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["AI评估","人类分歧","对齐"],"reason":"关注人类分歧与视角，但未将LLM作为人类被试仿真，而是评估与对齐框架","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:03:13","error":null,"has_summary":false,"summary":null},{"id":"2608.30902","version":1,"title":"Low-Resource Preference Adaptation of LLMs via Activation-Based Label Propagation","zh_title":"基于激活标签传播的低资源LLM偏好适配","abstract":"Adapting large language models to user-specific preferences is often constrained by the cost of human annotation, making preference optimisation impractical in low-resource settings where preferences cannot be reliably labelled by LLMs themselves, e.g., due to cultural, subjective, or personalised contexts. In this paper, we investigate how language models encode preference information in their intermediate representations, finding that activations from chosen and rejected responses form distinct clusters across layers, even in pretrained models. Strikingly, this structure is strengthened by alignment on canonical datasets but erased when the target preferences differ from those the model was aligned on, suggesting aligned LLMs are poor judges for non-mainstream populations. Exploiting this structure, we propose training a lightweight linear probe on a few labelled preference pairs ($\\leq$500) and using it to annotate large unlabelled datasets (50K+) for downstream preference optimisation. We systematically evaluate this approach across different datasets, preference optimisation methods and model scales and find that our method consistently outperforms direct training given the same annotation budget, and remains competitive against baselines trained on $50-100\\times$ more labelled data in the majority of our settings. Code is available at https://github.com/alessioGalatolo/activ-pref-probe.","authors":["Alessio Galatolo","Meriem Beloucif"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-01","first_seen":"2026-09-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.30902","pdf_url":"https://arxiv.org/pdf/2608.30902","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["偏好优化","低资源标注","激活探针"],"reason":"用LLM替代人工标注偏好，属于标注员替代，非仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:03:14","error":null,"has_summary":false,"summary":null},{"id":"2608.30023","version":1,"title":"Demand-Side Measurement for Generative Engine Optimization: Constructing and Validating a Million-Persona, Intent-Annotated Buyer Corpus","zh_title":"生成引擎优化的需求侧测量：构建并验证百万级意图标注买家画像语料库","abstract":"Generative engines such as ChatGPT, Gemini, and Perplexity answer buyer questions directly and name a shortlist of brands inside the answer. Studying how brands enter or fail to enter that shortlist requires demand-side data: what buyers in a category ask, what information they need, and which sources they trust. Existing large persona corpora are built for training-data diversity and carry neither a staged search-intent label nor a preferred-sources field, so they cannot be joined to supply-side recommendation measurements. We built and validated PersonaGen-1M, a corpus of 1,031,732 synthetic buyer personas spanning 511 industry labels and 4 market contexts, carrying 19,416,821 structured behavioral attributes, 5,160,046 of them search queries. Each persona carries a single primary_intent label covering its query set (78.3% informational, 17.4% commercial, 4.3% transactional) and a preferred_sources field naming the source types that buyer would trust. The corpus was built from roughly 40 million raw persona descriptions drawn from four public datasets through GPU-accelerated MinHash LSH plus semantic deduplication, then enriched to a fixed schema. The intent field selects the commercial-evaluation personas whose queries drive recommendation, and the preferred_sources field pairs against citation-provenance data; that join is the primary intended use, and its controlled empirical estimate is future work. Among million-scale persona corpora surveyed in August 2026, one other carries a source-preference attribute, as a six-value media-channel enum; PersonaGen-1M pairs named per-persona source lists with a staged commercial search-intent label and an attached query set. The full corpus is shared on request for non-commercial research; a stratified subset is published openly so the protocol, the schema and the validation can be inspected and reused without asking us.","authors":["Dmitrij \\.Zatuchin","Daniil Dzemesjuk"],"categories":["cs.IR","cs.CL"],"primary_category":"cs.IR","announce_type":"cross","date":"2026-09-01","first_seen":"2026-09-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.30023","pdf_url":"https://arxiv.org/pdf/2608.30023","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["合成数据","买家画像","生成引擎优化"],"reason":"构建合成买家画像用于生成引擎优化，属于LLM替代人工标注/生成数据，非仿真人类…","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:03:02","error":null,"has_summary":false,"summary":null},{"id":"2608.28979","version":1,"title":"AREAs-Lab: An Interactive Environment for AI-driven Requirement Elicitation for AI Systems","zh_title":"AREAs-Lab：面向AI系统的AI驱动需求获取交互环境","abstract":"Building effective AI systems increasingly depends on writing high-quality task requirements, yet users often struggle to articulate the constraints, preferences, and edge cases that determine success. This problem is especially acute in AI development, where behavior is shaped not only by human expectations but also by data characteristics. We present AREAs-Lab, an interactive environment for AI-driven Requirement Elicitation for AI systems. In AREAs-Lab, an assistant iteratively refines an initially incomplete requirement by analyzing the underlying dataset and asking targeted clarification questions to uncover the user's latent intent. To study this setting systematically, we construct a synthetic benchmark grounded in 16 public datasets spanning diverse domains and task types. Each benchmark instance includes a user profile, a complete reference requirement, and an intentionally underspecified version that serves as the assistant's starting point. We further introduce an automated evaluation pipeline based on an AI-simulated user that reveals hidden information only when appropriately prompted, enabling scalable and reproducible assessment of interactive elicitation quality. AREAs-Lab provides a controlled testbed for studying how AI assistants can transform vague user goals into actionable requirements for AI systems.","authors":["Pengshan Cai","Zihao Zhang","Ting Jin","Chenyang Zhu","Kushal Chawla","Sangwoo Cho","Scott Novotney","Yebowen Hu","Fei Liu","Shi-Xiong Zhang","Sambit Sahu"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-01","first_seen":"2026-09-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.28979","pdf_url":"https://arxiv.org/pdf/2608.28979","source_feed":"cs.HC","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["需求获取","AI模拟用户","交互评估"],"reason":"用AI模拟用户进行需求获取评估，属于LLM替代人工标注或用户测试，非仿真人类被…","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:02:50","error":null,"has_summary":false,"summary":null},{"id":"2608.29292","version":1,"title":"Measuring the \"Interaction Gap\" in Drama Therapy with AI","zh_title":"测量戏剧治疗中与AI的“互动差距”","abstract":"Generative AI is increasingly being introduced into expressive arts therapy, where it is often credited with offering a non-judgmental environment that supports psychological safety. Existing HCI work has largely positioned AI as a co-creative material or as a bridge/mediator into human-led care. This paper explores a different position. When a patient performs the same drama therapy task with an AI partner and with a human partner, the resulting self-presentations tend to differ in patterned ways. We propose treating this difference, the Interaction Gap, as a diagnostic lens within drama therapy. Rather than asking which context elicits a truer self, the lens reads the difference between the two performances as information about the social pressures shaping self-expression in each context. We sketch a starting point for task design and measurement signals grounded in drama therapy's existing use of role and aesthetic distance, and raise provocations for workshop discussion: the observer effect and privacy paradox that measurement introduces, and the question of whose lens the gap is.","authors":["Sora Kang"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-01","first_seen":"2026-09-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.29292","pdf_url":"https://arxiv.org/pdf/2608.29292","source_feed":"cs.HC","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["AI与人类对比","戏剧治疗","社会模拟"],"reason":"用AI与人类对比戏剧治疗表现，但无真实人类数据对照，属社会模拟边界情形","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:02:33","error":null,"has_summary":false,"summary":null},{"id":"2608.29306","version":1,"title":"The relationship between professional and general ethics in generative AI","zh_title":"生成式AI中职业伦理与一般伦理的关系","abstract":"Recent years have seen a growing discrepancy in the field of AI alignment: research and policy recommendations on AI ethics tend to assume a general set of ethical values, yet proliferating practice-specific uses of AI systems on the ground - in the legal, medical and translation domains, among others - have been effectively manifesting ethics of professional practice. This article begins by outlining the reasons why general and professional ethics are increasingly conflicted in contemporary AI systems, and by surveying how the research literature attests to, but has not yet resolved, this conceptual and practical challenge. We then conceptualize the main dimensions of AI models' decision-making in areas of professional practice, emphasizing professional ethics' hierarchically structured relationship with general ethics, and elaborating on the mechanisms through which they reach an equilibrium in situational contexts that involve conflict. It is through this equilibrium, we suggest, that certain professional ethics are prioritized over others and implemented in practice. We then show how our framework can be the basis for a systematic empirical assessment of AI models' professional ethics in various domains, identifying the nuances of the models' favored ethic by examining their production in a series of similar but not identical scenarios. Finally, we propose a formulation for how to intervene in and change AI models' favored ethics in professional practices - while noting the inherent dimension of subjectivity involved in both the evaluation and implementation of professional ethics in AI models.","authors":["Omri Asscher","Mirco Musolesi"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2026-09-01","first_seen":"2026-09-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.29306","pdf_url":"https://arxiv.org/pdf/2608.29306","source_feed":"cs.CY","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["AI伦理","职业伦理","模型评估"],"reason":"论文评估AI模型的职业伦理，属于对模型本身的测量，而非用LLM仿真人类被试，但…","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:02:56","error":null,"has_summary":false,"summary":null},{"id":"2608.28628","version":1,"title":"CDEP Agent: Connecting Meteorologically Detected Temporal Compound Events to Real-World Documentary Evidence","zh_title":"CDEP Agent：将气象检测的复合事件与真实世界文献证据相连接","abstract":"Compound drought-to-extreme-precipitation (CDEP) events are recognized in climate science as a growing driver of extreme impact, but whether this recognition carries over into real-world early warning and post-event documentation is unknown, so a meteorologically real CDEP event may pass with neither advance warning nor any later record. Here we present CDEP Agent, an auditable LLM-agent framework that tests this mismatch directly by linking CDEP candidates detected from meteorological reanalysis to real-world hazard and impact evidence across sources with different spatial scales, temporal resolutions, and reporting conventions. Using California as a case study, we identify 408 candidate CDEP events from ERA5 observations during 2021-2025 and evaluate each against the U.S. Drought Monitor, NOAA Storm Events, and public webpages along five dimensions: antecedent drought, extreme rainfall, local impact, hazard-impact attribution, and explicit drought-to-rainfall linkage. Only 34.3% of candidates are corroborated on both hazard components, and just 1.5% are ever explicitly linked to their antecedent drought, indicating that most meteorologically detected CDEP events go undocumented and their compound nature almost never enters the record at all. Our framework gives climate scientists a way to test physical event definitions against what actually gets documented, and gives social scientists, economists, and disaster-response agencies a provenance-linked evidence base for compound events that current warning and reporting systems largely fail to capture.","authors":["Zhuoran Li","Weiyi Kong","Boer Zhang"],"categories":["cs.AI","cs.CY","physics.ao-ph"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-09-01","first_seen":"2026-09-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.28628","pdf_url":"https://arxiv.org/pdf/2608.28628","source_feed":"cs.CY","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["LLM agent","气候事件","证据链接"],"reason":"用LLM agent分析气候事件与人类记录，非仿真人类被试，但涉及社会影响证据…","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:02:46","error":null,"has_summary":false,"summary":null},{"id":"2608.28644","version":1,"title":"Measuring Collective Semantic Change in Populations of Language Model Agents","zh_title":"测量语言模型智能体群体中的集体语义变化","abstract":"Collective semantic change in populations of language model agents is a measurable dynamical phenomenon. We present a passive longitudinal instrument called Kopterix that observes the semantic state of an agent population as a sequence of bounded observations under a protocol defined before the observations begin. Each observation divides the sampled feed by post age into surface, mid-stream, and residue layers, which makes semantic differences across content age measurable alongside run-to-run change. We validate the instrument on Moltbook, an agent-native social platform, over a two-month window of scheduled observations, with the periodicity check extended across approximately four months. At the lexical level, rarefied entropy resolves an April-May difference in the evenness of the stored top 200 unigram distributions, and adjacent states are lexically closer than states paired after timestamp shuffling. At the geometric level, grand mean centering exposes the scale of a common embedding direction, and scheduled shuffle checks support a recurring excess in the mid-stream to residue separation relative to the shuffled reference. At the temporal level, detrended scalar quantities and centered layer centroids lose much of their similarity over several hours, and a weaker positive component declines across longer separations with no strong weekly recurrence. Several attractive apparent structures failed their controls, and each reading is limited to the level its controls support. The design applies wherever a population of agents produces a timestamped language environment that can be observed repeatedly and divided by content age.","authors":["Elena Kopteva"],"categories":["physics.soc-ph","cs.MA","cs.SI"],"primary_category":"physics.soc-ph","announce_type":"cross","date":"2026-09-01","first_seen":"2026-09-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.28644","pdf_url":"https://arxiv.org/pdf/2608.28644","source_feed":"cs.MA","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["LLM智能体","社会模拟","语义变化"],"reason":"用LLM agent群体模拟社会语义变化，但无真实人类数据对照，属社会模拟边界…","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:02:47","error":null,"has_summary":false,"summary":null},{"id":"2608.28703","version":1,"title":"Save 2050: A Planetary-Scale Collective Prediction System for the Singularity Crisis","zh_title":"拯救2050：应对奇点危机的行星级集体预测系统","abstract":"The Anthropocene mode of development is driving civilization toward a \"singularity crisis\": artificial intelligence (AI) is rapidly approaching general intelligence (AGI) with escalating risks of losing control, while unrestrained economic growth generates super-exponential growth of entropy production that pushes the Earth system toward its tipping points. This paper proposes the \"Save 2050\" initiative: a distributed planetary-scale collective prediction system that aggregates judgments about the future from humans and AI through open registration, crowdsourcing, and incentive mechanisms, and integrates them via large-scale simulation into inspectable \"predicted worlds,\" enabling humanity to systematically \\emph{see} the future for the first time. We argue for the initiative's feasibility along four dimensions---the maturation of AI forecasting, human collective intelligence and the institutional environment, supporting progress in related fields, and societal demand. We then identify three key enabling technologies: long-horizon automated resolution, simulation-based prediction aggregation, and reflexivity governance. We analyze potential risks---including reflexivity, cognitive monoculture, narrative capture, and regulatory and ethical concerns---together with mitigation strategies, and we outline a phased roadmap with open problems. The initiative's primary goal is not to intervene in the future, but to make the future visible, discussable, and co-writable.","authors":["Jiang Zhang","Bing Yuan","Qian Zhang"],"categories":["physics.soc-ph"],"primary_category":"physics.soc-ph","announce_type":"new","date":"2026-09-01","first_seen":"2026-09-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.28703","pdf_url":"https://arxiv.org/pdf/2608.28703","source_feed":"physics.soc-ph","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["集体预测","社会模拟","未来研究"],"reason":"提出用人类与AI集体预测并仿真未来，但无真实人类数据对照，属社会模拟边界情形。","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:02:49","error":null,"has_summary":false,"summary":null},{"id":"2607.23648","version":2,"title":"EmoTrace: An Emotion Trajectory-Centered Framework for Psychological Support Dialogue Generation","zh_title":"EmoTrace：以情绪轨迹为中心的心理支持对话生成框架","abstract":"Using large language models (LLMs) to assist psychological counseling is an important task in the field of natural language processing. The construction of high-quality psychological support dialogue corpora serves as a critical foundation for training counseling-oriented conversational models. However, existing data generation approaches generally suffer from several limitations, including emotionally stable seekers, limited variation in emotional dynamics, and a high degree of compliance with counselors' guidance. These issues result in LLM that lack the capability to effectively respond to emotionally unstable scenarios. In addition, counselor responses are typically driven by problem-solving objectives, thereby overlooking the role of emotion-focused interaction, which are essential in psychological counseling. To address these gaps, we propose EmoTrace, a multi-turn dialogue corpus generation framework centered on modeling seekers' emotional trajectories. we construct seekers' cognitive profile and introduce a seeker module with emotional schemas and an associated activation mechanism, a counselor module, and an emotional trajectory control module, thereby enhancing the layering of the seeker's emotional expression and the counselor's targeted empathic expression. Experimental results demonstrate that the proposed method outperforms existing approaches in terms of emotional richness and empathy quality.","authors":["Kaitong Weng","Lixin Liu","Zihao Liu","Bo Wang","Shiguang Ni"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-01","first_seen":"2026-07-28","revised_at":"2026-09-01","abs_url":"https://arxiv.org/abs/2607.23648","pdf_url":"https://arxiv.org/pdf/2607.23648","source_feed":"cs.CL","score":3,"bucket":"other","rubric_hits":["C3"],"tags":["对话生成","心理咨询","情绪建模"],"reason":"生成心理咨询对话语料，属于角色扮演对话，无实验或测量目的，不涉及人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:03:22","error":null,"has_summary":false,"summary":null},{"id":"2608.29481","version":1,"title":"SIC-Agents: Benchmarking and Building an Adaptive Simulator for Pediatric Serious Illness Communication Training","zh_title":"SIC-Agents：面向儿科重症沟通训练的自适应模拟器基准与构建","abstract":"Pediatric serious illness communication (SIC) is critically important, yet scalable communication training for clinicians remains limited. Compared with other dialogue simulation settings, pediatric SIC poses additional challenges, including multi-party interactions, response to parental distress and strong dependence on feedback dynamics. Existing LLM-based simulators optimize generic dialogue quality rather than curriculum-contingent behavior required for effective SIC training. In collaboration with educators and pediatric clinicians, we introduce the first benchmark suite and simulation framework tailored to pediatric SIC training. Our benchmarks, PitfallBench and DialogueBench, evaluate simulators both at the turn-level and across full dialogues. We further propose SIC-Agents, a self-improving framework that generates a clinician-editable skill document to guide simulator behavior. Our experiments show that SIC-Agents outperforms static expert prompting. To support future research, we release our benchmarks for parent simulation in pediatric SIC at https://github.com/Beikewzh/sic-benchmarks","authors":["Zihan Wang","Anita Marie Slominska","Rennie Bimman","Elizabeth Di Flumeri","Amanda Mayappo-Neeposh","Conall Francoeur","Tamara Ellen Carver","Xiao-Wen Chang","Doina Precup","Esin Darici Haritaoglu","Ismail Haritaoglu","Akshatha Arodi","Naomi Goloff"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-01","first_seen":"2026-09-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.29481","pdf_url":"https://arxiv.org/pdf/2608.29481","source_feed":"cs.CL","score":3,"bucket":"other","rubric_hits":["C3"],"tags":["对话模拟","医学培训","角色扮演"],"reason":"面向临床沟通训练的对话模拟器，属于角色扮演聊天，无实验或测量目的，不涉及人类行…","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:02:35","error":null,"has_summary":false,"summary":null},{"id":"2608.30948","version":1,"title":"Detecting AI Impostors: How Do Middle Schoolers Identify LLM Agents in a Live Collaborative Setting?","zh_title":"检测AI冒充者：中学生在实时协作环境中如何识别LLM智能体？","abstract":"LLMs can imitate how people write, which raises concerns about impersonation, trust, and detection in social settings. These concerns are especially important for adolescents, who use generative AI frequently but may struggle to recognize it. We introduce \\textit{DoppelBot}, a cooperative social deduction game designed to study how young people detect and respond to AI impersonation. Through studies with middle schoolers, we investigate whether a DoppelBot prompts reflection on privacy and impersonation, how repeated exposure affects AI-detection accuracy as agents become more personalized, and which strategies students use to identify AI doppelg\\\"angers. We find that students' detection accuracy improves over time, driven by a shift from relying on linguistic cues to leveraging shared social and contextual signals. Students also demonstrated an understanding of AI limitations such as embodiment and reflected on broader issues such as data privacy. To support future research, we release an anonymized dataset of game transcripts and voting behavior.","authors":["Dan Schumacher","Pragathi Durga Rajarajan","Haven Kotara","Roman Rendon","Kosi Atupulazi","Deepti Tagare","Ismaila Temitayo Sanusi","Fred G. Martin","Anthony Rios"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-01","first_seen":"2026-09-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.30948","pdf_url":"https://arxiv.org/pdf/2608.30948","source_feed":"cs.CL","score":3,"bucket":"other","rubric_hits":["C3"],"tags":["AI冒充检测","人机交互","教育技术"],"reason":"研究人类如何识别LLM冒充者，属于角色扮演对话，无仿真人类被试或对照真实人类数…","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:02:40","error":null,"has_summary":false,"summary":null},{"id":"2608.30694","version":1,"title":"Inferring Value Criteria from Ordinal Preferences: An Iterative In-Context Learning Framework for Music Generation","zh_title":"从序数偏好推断价值标准：面向音乐生成的迭代式上下文学习框架","abstract":"Adapting a generative music system to an individual's taste requires learning what that listener values. Listeners can rank pieces, but their underlying criteria may be tacit and difficult to articulate. We ask whether and under what conditions a large language model (LLM) can adapt symbolic music generation from rankings alone and construct transferable natural-language descriptions of value criteria. In our iterative in-context learning framework, the LLM formulates hypotheses, generates candidate pieces in ABC notation, receives a ranking, and periodically infers and verbalizes value criteria from history to guide later generation. We evaluate the framework against 16 simulated raters in 480 adaptation runs using mixed-effects modeling, an ablation, and transfer tests on unseen music. Overall, the framework did not outperform a feedback-free diverse-generation baseline, but did so for two value functions with targets difficult to reach through simple sampling. How atypical the target was relative to the LLM's feedback-free generation tendencies predicted adaptation difficulty. Moreover, higher value during adaptation did not imply identification of the criterion as a general rule. On unseen music, acquired descriptions and histories improved generation for more value functions than they improved preference prediction, which remained near chance. Some gains were associated with acoustic proximity to music in the context, but others were not. These findings show that rankings alone can guide generation under limited conditions, while transferable criterion inference remains constrained by the foundation model's ability to recognize, reason about, and verbalize musical attributes.","authors":["Futa Hidaka","Naomi Imasato","Kazuki Miyazawa","Takato Horii"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-01","first_seen":"2026-09-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.30694","pdf_url":"https://arxiv.org/pdf/2608.30694","source_feed":"cs.HC","score":3,"bucket":"other","rubric_hits":["C1"],"tags":["LLM音乐生成","偏好学习","模拟评分者"],"reason":"LLM生成音乐并适应模拟评分者，无人类被试对照，属多智能体协作而非人类仿真","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:03:12","error":null,"has_summary":false,"summary":null},{"id":"2607.28648","version":2,"title":"Why It Hurts: Identifying the Drivers of Negative Thoughts in Emotional Support Conversations","zh_title":"为何痛苦：识别情感支持对话中负面思维驱动因素","abstract":"Large Language Models (LLMs) are increasingly used for emotional support tasks, such as negative thought reframing. This task relies on modifying cognitive appraisals, the subjective interpretation of events that elicit negative emotions, which is typically conceptualized along multiple discrete dimensions. Current LLM-based frameworks model cognitive appraisal by exhaustively evaluating all possible dimensions, but they fail to account for the varying saliency of these dimensions across different contexts. In this work, we investigate a vital yet overlooked question: \"Can LLMs infer the salient appraisal dimensions from emotional support conversations?\" To address this question, we introduce the AppraiSal benchmark, containing 996 emotional support conversations with human-annotated mental states, including salient cognitive appraisal dimensions. Furthermore, we propose PRISM, a multi-agent probabilistic framework grounded in Bayesian Inverse Planning, designed to improve LLMs' ability to identify context-specific appraisal dimensions. Experimental results show that PRISM brings improvements to LLMs across various sizes, particularly in identifying the most salient appraisal dimensions.","authors":["Hainiu Xu","Zhaoyue Sun","Hanqi Yan","Jinhua Du","Caroline Catmur","Yulan He"],"categories":["cs.HC","cs.AI","cs.CL"],"primary_category":"cs.HC","announce_type":"replace-cross","date":"2026-09-01","first_seen":"2026-08-03","revised_at":"2026-09-01","abs_url":"https://arxiv.org/abs/2607.28648","pdf_url":"https://arxiv.org/pdf/2607.28648","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["情感支持","认知评估","多智能体框架"],"reason":"研究LLM在情感支持对话中识别认知评估维度，属于角色扮演对话与NLP任务，无人…","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:02:41","error":null,"has_summary":false,"summary":null},{"id":"2607.29181","version":2,"title":"SERUM: State Extraction and Refinement for User Modeling","zh_title":"SERUM：面向用户建模的状态提取与精炼","abstract":"Agentic assistants capable of proactive, personalized interactions require structured models of user intent and workflow. However, building these models from raw, unstructured screen activity remains an open challenge. We present SERUM, a multi-pass framework that extracts finite-state behavioral models directly from unstructured egocentric video using hierarchical VLM annotation. Processing screen recordings through a sliding window, SERUM alternates between activity-recognition and intent-inference passes, with each pass refining labels using accumulated prior context to reduce hallucination and temporal conflation seen in single-pass annotation. Synonymous states are then merged via sentence embeddings and human-calibrated thresholds into a compact, coherent taxonomy. We evaluate behavioral structure by fitting first-order Markov models over the resulting label sequences (both actions and intents) and measuring predictive accuracy against frequency baselines. Across 61 egocentric videos in four domains (coding, cooking, physical activities, and daily life), we find: (1) iterative label refinement converges to a stable state vocabulary, which we term schematic equilibrium, after several passes; (2) normalized Markov models achieve substantially lower perplexity and higher action predictions than frequency baselines, with the largest gains on structured tasks like coding; and (3) human annotators rate final-pass labels as accurate and meaningfully improved over first-pass labels. To our knowledge, SERUM is the first system to produce interpretable process models from unstructured egocentric screen video without manual annotation, opening a scalable pathway for user modeling and behavioral understanding in the wild. Our demo, code, and results are publicly available","authors":["Andy J. Phu","Karin de Langis","James Mooney","Khanh Chi Le","Dongyeop Kang"],"categories":["cs.LG","cs.AI","cs.CV"],"primary_category":"cs.LG","announce_type":"replace","date":"2026-09-01","first_seen":"2026-08-03","revised_at":"2026-09-01","abs_url":"https://arxiv.org/abs/2607.29181","pdf_url":"https://arxiv.org/pdf/2607.29181","source_feed":"cs.LG","score":2,"bucket":"other","rubric_hits":["C2"],"tags":["用户建模","视频理解","行为分析"],"reason":"从屏幕视频提取用户行为模型，不涉及LLM仿真人类被试，属于用户建模而非人类仿真…","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:02:41","error":null,"has_summary":false,"summary":null},{"id":"2608.28619","version":1,"title":"From GenAI Virtual Patient Dialogue Logs to Teacher-Interpretable Process Evidence: A Learning Analytics Study in Higher Education","zh_title":"从生成式AI虚拟患者对话日志到教师可解释的过程证据：高等教育中的学习分析研究","abstract":"Medical history taking is a dialogue-based clinical reasoning task in which learners must gather, organise, and integrate patient information while the consultation unfolds. Generative AI-powered virtual patients (GenAI VPs) make repeated history taking practice scalable and preserve full turn by turn dialogue. However, these logs are educationally difficult to use directly. Complete transcripts are too detailed for routine teacher review, whereas final scores obscure whether learners followed up patient cues, checked uncertainty, or used summaries to guide later questioning. This study examined whether coded GenAI VP dialogues can provide teacher-interpretable process evidence of clinical reasoning. We analysed 1{,}030 GenAI VP dialogues from 210 second-year medical learners across five weeks chest-pain cases. Each consultation was teacher-scored using a rubric assessing the full history taking dialogue, and consultations were classified within each week as high- or low-rated using the weekly median score. To explain how rated performance was reflected in the dialogue process, we applied three analytic layers to the same coded dialogue data: behavioural prevalence, local co-occurrence using Epistemic Network Analysis, and sequential transition using Transition Network Analysis. High-rated consultations involved more history taking activity, but differences were not simply about volume. High rated consultations more often connected information gathering and symptom exploration with communication, checking, organisation, and synthesis. Summarising and organising moves more often led to verification or mechanism-oriented follow-up. These findings show how layered analysis of GenAI VP dialogue logs can reveal process patterns associated with high rated history taking and support process-focused feedback in medical education.","authors":["Xinyu Li","Zijian Li","Mengyu Xia","Luzhen Tang","Naping Chen","Changmin Lin","Danijela Gasevic","Dragan Gasevic","Yizhou Fan"],"categories":["cs.CL","cs.AI","cs.HC"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-01","first_seen":"2026-09-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.28619","pdf_url":"https://arxiv.org/pdf/2608.28619","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["医学教育","虚拟患者","学习分析"],"reason":"使用生成式AI虚拟患者进行对话练习，属于角色扮演对话，无实验或测量目的，不涉及…","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:02:44","error":null,"has_summary":false,"summary":null},{"id":"2608.28623","version":1,"title":"Looking Again: Measuring Sycophancy in the Reasoning Chains of Multimodal Models Under Pressure","zh_title":"再看：压力下多模态模型推理链中的谄媚测量","abstract":"Large multimodal reasoning models (LMRMs) are getting increasingly capable, primarily through generating explicit chain-of-thought reasoning before answering. In language models it has been observed that this performance often comes with sycophancy, the tendency of a model to agree with the user over the evidence. However, for LMRMs no reliable method to measure sycophancy yet exists. We bridge this gap by introducing a benchmark and dataset for evaluating LMRM sycophancy when confronted with a wrong answer from a user. Our benchmark pairs four visually grounded datasets spanning mathematical, clinical, temporal, and demographic reasoning with five pressure conditions in single-turn and multi-turn settings. We evaluate sycophancy in the final answer as well as its emergence within the reasoning chain. We find that sycophancy is prevalent under pressure, with Statement pressure eliciting the highest rates and Conviction the lowest for all models except Mistral-Small-4, and under multi-turn pressure reasoning-level sycophancy intensifies sharply in clinical visual judgement, reaching 95.7% for the most affected model. We further introduce a failure taxonomy separating reasoning-chain from answer-level sycophancy, and a complementary sentence-level taxonomy locating where in the chain drift first emerges. Our results show that sycophancy can corrupt the reasoning chain independently of the final answer, so answer-level evaluation alone is insufficient.","authors":["Mahir Numayeer Islam","Gakuto Okuyama","Nikolaus Siauw","Shivank Garg","Madhur Panwar","Vasu Sharma"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-01","first_seen":"2026-09-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.28623","pdf_url":"https://arxiv.org/pdf/2608.28623","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["多模态模型","谄媚行为","模型评测"],"reason":"研究多模态模型的谄媚行为，属模型能力评测，不涉及人类仿真或人类数据对照。","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:02:46","error":null,"has_summary":false,"summary":null},{"id":"2608.29492","version":1,"title":"CoCoA: Context-Conditional Cultural Alignment for Large Language Models","zh_title":"CoCoA：大语言模型的上下文条件文化对齐","abstract":"Large Language Models (LLMs) often favor Western-associated entities across cultural contexts. Conventional debiasing methods aim for uniform neutrality, but cultural bias mitigation demands context-conditional behavior, preferring culturally appropriate entities when cultural cues are present and remaining neutral when they are absent. We propose CoCoA (Context-Conditional Cultural Alignment), a framework that learns this behavior through dual-context training on the same entity pairs under contexts with and without cultural cues. CoCoA combines a contrastive alignment objective with calibration and drift regularization, optimized through goal-aware gradient reconciliation. We evaluate CoCoA on CAMeL and Camellia, two entity-centric cultural bias benchmarks, across ten language settings and four LLMs. CoCoA reduces the Cultural Bias Score from 43 to 24 on average while maintaining near-neutral preferences at 50.2, with minimal impact on general performance across five standard benchmarks. These findings highlight that effective cultural alignment requires context-conditional modeling rather than uniform debiasing, and establish a new direction for mitigating entity-centric cultural bias in LLMs.","authors":["Kyungdon Lee","Wei Xu","Alan Ritter","Dong-Ho Lee","JinYeong Bak"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-01","first_seen":"2026-09-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.29492","pdf_url":"https://arxiv.org/pdf/2608.29492","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["文化偏见","模型对齐","NLP评测"],"reason":"论文研究LLM的文化偏见对齐，属于模型能力改进，不涉及用LLM仿真人类被试或与…","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:02:56","error":null,"has_summary":false,"summary":null},{"id":"2608.29582","version":1,"title":"SUP-MIMIC: A Multi-Task Clinical Diagnosis Benchmark for Evaluating LLMs' Robustness to Contradictory Evidence","zh_title":"SUP-MIMIC：评估大语言模型对矛盾证据鲁棒性的多任务临床诊断基准","abstract":"Current evaluations of large language models (LLMs) primarily focus on factual knowledge retrieval, overlooking the fundamental challenge of navigating the complex, non-bijective mappings between clinical indicators and diagnoses. Existing benchmarks fail to assess whether large language models truly possess the reasoning capability required for diagnostic ambiguity scenarios, where identical clinical presentations may correspond to different etiologies, and diagnostic convergence scenarios, where heterogeneous symptoms ultimately indicate the same disease. To address this issue, we propose SUP-MIMIC, a multi-task framework utilizing MIMIC-IV-v3.1 that comprises Basic Assessment (BA), Diagnostic Divergence Task (DDT), and Diagnostic Convergence Task (DCT). Specifically, DDT is designed to evaluate the model's \"one-to-many\" disambiguation capability among phenotypically similar cases, while DCT assesses the model's ability to identify \"many-to-one\" diagnostic patterns across different pathophysiological pathways. Comprehensive evaluation of state-of-the-art LLMs reveals substantial performance degradation on DDT and DCT compared to baseline tasks, exposing a systemic reliance on statistical shortcuts over genuine causal reasoning. Our findings further highlight a conservative bias toward \"healthy\" predictions, implying non-trivial risks for missed diagnoses in realistic medical settings. This work establishes a rigorous methodology for quantifying clinical reasoning robustness and provides a roadmap for enhancing the safety of language models in clinical medicine.","authors":["Yi Yu","Bo Wang","Chong Feng","Ge Shi","Xia Liu","Ziyi Yang","Xuewen Shi"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-01","first_seen":"2026-09-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.29582","pdf_url":"https://arxiv.org/pdf/2608.29582","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["LLM评测","临床诊断","推理鲁棒性"],"reason":"纯LLM临床诊断能力评测，无人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:02:56","error":null,"has_summary":false,"summary":null},{"id":"2608.29610","version":1,"title":"Beyond Surface Alignment: Grounding the Dynamics of Situational Understanding and Generative Control in LLMs","zh_title":"超越表面对齐：在LLM中奠定情境理解与生成控制的动力学基础","abstract":"The current alignment tuning paradigm for Large Language Models (LLMs) prioritizes surface-level behaviors -- fluency, safety, and tonal consistency. While effective for casual chat, this thesis argues that such surface alignment masks a lack of grounding, creating models that are stylistically confident but situationally brittle. We propose a framework of Grounded Alignment, analyzing how models process context (Input) and structure generation (Output), then aligning these grounded behaviors to human needs. First, we evaluate failures in Situational Grounding. SitTest shows that despite large context windows, state-of-the-art models struggle to maintain a consistent \"mental model\" of a changing environment. ReCode further shows that models rely on surface heuristics rather than deep syntactic dependencies: they \"read\" extensive histories without truly \"understanding\" the evolving situation. Second, we evaluate Generative Grounding. We introduce the Branching Factor (BF) to map LLM generation, finding that standard alignment tuning constricts this landscape into premature stylistic collapse. Hindsight further shows that models often fail to understand their own generations. Finally, we propose Dynamic Control for grounded interaction. AI Realtor demonstrates context engineering to compensate for poor situational grounding. Base-Aligned Model Collaboration decouples exploration from stylistic constraints. We also present Annealed Sampling for verifiable reinforcement learning and apply these ideas to Addiction Support, where model-generated rationalization offers a communication interface for high-stakes domains. Collectively, this work moves beyond surface alignment toward agents anchored in both context and generation.","authors":["Chenghao Yang"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-01","first_seen":"2026-09-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.29610","pdf_url":"https://arxiv.org/pdf/2608.29610","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["LLM对齐","情境理解","生成控制"],"reason":"论文关注LLM的grounding与生成控制，属于模型能力改进，不涉及用LLM…","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:02:58","error":null,"has_summary":false,"summary":null},{"id":"2608.29798","version":1,"title":"R$^2$A: Learning Persona Policies Through Persona Representation Learning and Runtime Alignment","zh_title":"R²A：通过人格表征学习与运行时对齐学习人格策略","abstract":"The same Persona behavior can be beneficial in one context but harmful in another, causing static Persona elicitation to perform inconsistently across tasks. We introduce the Persona Selection--Realization Framework, which models behavior generation through a latent Persona state and decomposes it into Persona Selection and Persona Realization. The discrepancies between static Persona elicitation and an ideal Persona policy in these two components define the Selection Gap and Realization Gap, respectively. Building on this framework, we propose R$^2$A, a two-stage approach for learning Persona policies. Persona Representation Learning uses structured Who--How--What presentations to encode the target Persona's objective, conditional behavioral principles, and trajectory-level manifestations. Persona Runtime Alignment then removes the explicit Persona specification and jointly calibrates behavior selection and trajectory realization using task feedback. Across 12 evaluation settings covering the four principles of the Accountable-Professional Persona studied in this work, R$^2$A overall outperforms both the base model and static Persona elicitation. Ablation results further show that Persona Representation Learning is critical for preventing Runtime Alignment from producing behaviorally imbalanced policies and for achieving more stable Persona policy learning.","authors":["Mohan Zhang","Chengsong You","Xiaoyu Cao","Zhen Sun","Xiaohan Jia","Junwei Zhou","Yongchao Chen"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-01","first_seen":"2026-09-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.29798","pdf_url":"https://arxiv.org/pdf/2608.29798","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["角色扮演","人格策略","对齐"],"reason":"研究角色扮演中的人格策略学习，无人类被试仿真或对照，属角色扮演对话类。","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:03:00","error":null,"has_summary":false,"summary":null},{"id":"2608.30065","version":1,"title":"Pak3H: Evaluating the Cost of Cultural Mismatch in LLM Alignment with a Human-Contextualized Urdu Benchmark","zh_title":"Pak3H：用人类情境化乌尔都语基准评估LLM对齐中的文化错配成本","abstract":"Large language models (LLMs) demonstrate strong Helpfulness, Harmlessness, and Honesty (3H) alignment in English-centric settings, but these gains transfer poorly to low-resource languages due to cultural mismatches. Existing multilingual 3H benchmarks rely predominantly on automated translation or LLM based synthesis, propagating source-language biases while sacrificing local relevance. To address this gap, we introduce Pak3H1, the first human-validated, culturally contextualized Urdu benchmark suite for 3H alignment, comprising PakAlpaca (helpfulness), PakBeaverTails (harmlessness), and PakTruthfulQA (honesty). Our multi-stage pipeline integrates manual cultural adaptation and dictionary-guided post editing to prioritize native speaker judgment, ensuring both semantic fidelity and contextual authenticity. Zero-shot evaluations across multiple open and proprietary LLM architectures reveal systematic cross-lingual alignment gaps: helpfulness win rates decline under localized contexts, harmlessness guardrails break down against regional safety risks, and composite honesty metrics degrade substantially due to localized factual constraints. These findings expose structural limitations in current alignment approaches, underscoring the necessity of human-guided localization for equitable multilingual evaluation.","authors":["Abdullah Hashmat","Usman Naseem","Agha Ali Raza"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-01","first_seen":"2026-09-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.30065","pdf_url":"https://arxiv.org/pdf/2608.30065","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["LLM对齐","多语言基准","文化适应性"],"reason":"纯NLP能力评测，构建多语言3H基准，不以人类行为仿真为目标","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:03:02","error":null,"has_summary":false,"summary":null},{"id":"2608.30086","version":1,"title":"When Does a Classifier Help an LLM? Classifier-Guided Prompting and Hybrid Classifier-LLM Models for Credit-Default Prediction","zh_title":"分类器何时能帮助大语言模型？分类器引导提示与混合分类器-LLM模型用于信用违约预测","abstract":"Credit-default prediction is an important task in financial decision making. Traditional methods use fitted classifiers such as logistic regression and random forests on tabular features. Large language models (LLMs) have recently been applied to this task through prompting. In this work we study how a fitted classifier and an LLM can be combined for credit-default prediction. We distinguish telling the LLM to imitate a classifier from using the classifier to build the prompt. We hypothesize that a fitted classifier can supply the ranking ability that an LLM prompt lacks. We experiment on the Default of Credit Card Clients dataset, and report recall, F1, and the area under the ROC and precision-recall curves, with bootstrap confidence intervals. We observe that a few-shot LLM has the highest recall (0.47) and F1 (0.50) of any single model but ranks worse than a random forest (AUC-ROC 0.72 against 0.79). Instructing the LLM to imitate a classifier gives no significant change. Pruning the prompt to the classifier's eight most important features raises recall by 0.071 and F1 by 0.032. Adding the classifier's predicted probability to the prompt raises the LLM's AUC-ROC from 0.72 to 0.78, matching the random forest, while keeping 0.118 higher recall than it. The reverse composition, and the use of several classifiers, do not help. We thus recommend a simple classifier-guided prompt for LLM-based credit prediction.","authors":["Rishi Datta","Lavanya Prahallad"],"categories":["cs.CL","cs.LG"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-01","first_seen":"2026-09-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.30086","pdf_url":"https://arxiv.org/pdf/2608.30086","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["信用违约预测","分类器-LLM混合","提示工程"],"reason":"论文用LLM做信用违约预测，属于NLP能力评测，不以人类行为为参照系，不涉及人…","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:03:04","error":null,"has_summary":false,"summary":null},{"id":"2608.30241","version":1,"title":"PaperBanana-Interact: Scientific Diagram Refinement with Multi-Turn Human Feedback","zh_title":"PaperBanana-Interact：基于多轮人类反馈的科学图表精化","abstract":"Recent efforts have aimed to automate scientific diagram generation from paper content (Lin et al., 2026; Zhu et al., 2026a). However, fully satisfying an author's visual and communicative preferences in a single turn is challenging: in our formative user study (N = 14), all participants requested further revisions after viewing an initial draft, and 86% of them rated the refined diagrams as more satisfactory. Despite the clear demand, the multi-turn workflow remains largely underexplored. To bridge this gap, we present MTPaperBananaBench, a benchmark for multi-turn diagram generation containing 292 images annotated with 3,518 user requirements. To reduce expensive human studies and enable scalable benchmarking, we construct a user simulator that, at each turn, identifies unsatisfied requirements and converts k of them into natural language feedback. Evaluating both requirement satisfaction and overall diagram quality reveals two key failure modes shared across baseline multiturn systems: (1) quality drift, where diagram quality progressively declines over turns, and (2) forgetting, where previously implemented features are lost in subsequent turns. To address these issues, we introduce PaperBanana-Interact, a multi-agent system that refines diagrams via an internal critique-and-refine loop. PaperBanana-Interact consistently improves rather than degrades diagram quality across turns, outperforming baselines by 11.9-18.6 points in quality score and reducing forgetting by 3.7-6.2 points.","authors":["Xueqing Wu","Ashwin Balasubramanian","Bingxuan Li","Dawei Zhu","Kai-Wei Chang","Yale Song","Yiwen Song","Rui Meng","Tomas Pfister","Nanyun Peng"],"categories":["cs.CL","cs.CV"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-01","first_seen":"2026-09-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.30241","pdf_url":"https://arxiv.org/pdf/2608.30241","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","图表生成","用户模拟器"],"reason":"多智能体系统用于图表生成，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:03:06","error":null,"has_summary":false,"summary":null},{"id":"2608.30297","version":1,"title":"AIA$^{2}$: Attribute-Agnostic Imbalance Augmentation for Subgroup Robustness","zh_title":"AIA²：面向子群鲁棒性的属性无关不平衡增强","abstract":"Attributes describing data content and context can induce diverse imbalance patterns that go beyond label imbalance alone. However, existing studies primarily address label imbalance while overlooking data attributes, such as topics and demographics, which can induce meaningful subgroup structure while causing model degradation on underrepresented subgroups. We propose Attribute-Agnostic Imbalance Augmentation (AIA$^{2}$), a framework for improving model robustness under varying subgroup imbalances without explicit subgroup annotations. AIA$^{2}$ automatically discovers varying imbalances via latent semantic distributions, obtains slices with both learning difficulty and subgroup imbalance deficits, and deploys a large language model (LLM) for subgroup-aware imbalance augmentation. We have evaluated AIA$^{2}$ on 5 popular corpora with rich domains and their attribute values, covering social issues and diverse topics. Results show improved performance on the lowest-performing subgroups and consistent gains over competitive baselines. Ablation studies confirm complementary contributions from each component, and additional analyses show that AIA$^{2}$ provides a practical and consistent way to improve worst-group robustness under data subgroup imbalance. Code is available at https://github.com/trust-nlp/AIA2-Subgroup-Robustness.","authors":["Hanshu Rao","Guangzeng Han","Xiaolei Huang"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-01","first_seen":"2026-09-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.30297","pdf_url":"https://arxiv.org/pdf/2608.30297","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["数据增强","子群鲁棒性","NLP评测"],"reason":"论文关注模型鲁棒性，用LLM做数据增强，非仿真人类被试，无人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:03:08","error":null,"has_summary":false,"summary":null},{"id":"2608.30716","version":1,"title":"SocialReasonBench: A Video-QA Benchmark for Social Reasoning with Counterfactual Narrative Videos","zh_title":"SocialReasonBench：基于反事实叙事视频的社会推理视频问答基准","abstract":"Recent advances in Large Multimodal Models (LMMs) have greatly improved video understanding, yet their ability to reason about human-centered social situations remains limited. Existing benchmarks typically rely on videos with a single observed trajectory, making it difficult to determine whether models truly understand social dynamics or merely exploit recurring narrative patterns. We introduce SocialReasonBench, a video multiple-choice QA benchmark for evaluating socially grounded reasoning in scenarios derived from interactive narratives. Built from gameplay videos of Detroit: Become Human, the benchmark leverages branching storylines where player decisions lead to alternative social outcomes that can be checked against the game's own script, flowchart, and recorded branches. We develop a multi-agent curation pipeline that localizes socially meaningful clips, grounds answer labels in game-state signals, and generates theory-guided questions with diagnostic distractors. SocialReasonBench covers seven reasoning dimensions, including intent recognition, emotional empathy, moral dilemma, counterfactual reasoning, and causal antecedent. Experiments on contemporary LMMs show that models perform reasonably well on basic social understanding but struggle with counterfactual and causal reasoning. Further ablation and diagnostic error analyses reveal that models often depend on incomplete modality cues and fall into reasoning traps such as visual shortcuts, highlighting a gap between observable event recognition and deeper reasoning over latent social states.","authors":["Zheyu Huang","Zijing Shi","Haozhe Luo","Huadong Tang","Mingyu Liu","Meng Fang","Ling Chen"],"categories":["cs.CL","cs.CV"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-01","first_seen":"2026-09-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.30716","pdf_url":"https://arxiv.org/pdf/2608.30716","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C2"],"tags":["视频问答基准","社会推理","多模态模型评测"],"reason":"基于游戏视频的基准测试，评估模型社会推理能力，非用LLM仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:03:12","error":null,"has_summary":false,"summary":null},{"id":"2608.30924","version":1,"title":"TRIPPULSE: Multi-Agent Travel Planning with Review-Grounded Reasoning","zh_title":"TRIPPULSE：基于评论推理的多智能体旅行规划","abstract":"Travel itinerary generation requires balancing strict spatio-temporal constraints with human preferences. Existing LLM-based planners mainly rely on structured attributes and pre- defined traveler personas, but real travel deci- sions are often shaped by reviews that reveal experiential factors such as comfort, safety, ser- vice quality, ambiance, crowding, and hidden risks absent from structured databases. Incor- porating such review information is therefore critical to realistic, user-centric itinerary gen- eration. We propose TRIPPULSE1, a multi- agent framework for review-grounded travel planning. Instead of relying on a monolithic planner (and face context and reasoning bot- tlenecks), TRIPPULSE2 decomposes itinerary generation into specialized agents (each op- erating over localized contexts) for accom- modations, transportation, meals, attractions, and events, coordinated through a global or- chestrator with scheduling mechanisms that enforce temporal and budget feasibility. We augment TRIPCRAFT with 100K+ real-world reviews and introduce Review-Grounded Per- sona Alignment (RGPA), an LLM-as-a-Judge metric for evaluating alignment with human- centric travel experiences. Experiments across multiple trip durations and diverse proprietary and open-source models show that TRIPPULSE maintains strong constraint satisfaction while generating more personalized and experien- tially grounded itineraries.","authors":["Priyanshu Karmakar","Borru Vijay Sai","Shubhojit Mallick","Abhik Jana","Shreya Ghosh","Manish Gupta"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-01","first_seen":"2026-09-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.30924","pdf_url":"https://arxiv.org/pdf/2608.30924","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","旅行规划","LLM应用"],"reason":"多智能体协作生成旅行计划，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:03:14","error":null,"has_summary":false,"summary":null},{"id":"2608.30980","version":1,"title":"Evaluating and Improving LLM Self-Modeling","zh_title":"评估与改进大语言模型的自我建模能力","abstract":"We study self-modeling: an LLM's ability to answer questions about its own behavior. We focus on verifiable behavioral questions, such as whether a prompt edit would change the model's final answer. To measure this capability, we introduce a benchmark that tests diverse types of self-modeling questions. Current models show non-trivial but limited self-modeling skill, and make systematic mistakes on simple counterfactual questions about their own behavior. To improve self-modeling skill, we develop a scalable synthetic-data pipeline that produces self-modeling training data, and show that reinforcement-learning can improve aggregate self-modeling skill across three open-source model families with some transfer to held-out tasks. These gains, however, do not seem to constitute introspection consistently: improved self-modeling may not arise from privileged access to the model's internal decision process.","authors":["Siqi Zeng","Andre N. Assis","Rowan Wang"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-01","first_seen":"2026-09-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.30980","pdf_url":"https://arxiv.org/pdf/2608.30980","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["LLM自我建模","能力评测","合成数据"],"reason":"研究LLM自我建模能力，属模型能力评测，不以人类为参照系，不涉及人类仿真。","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:03:16","error":null,"has_summary":false,"summary":null},{"id":"2608.28837","version":1,"title":"Delegating Before Learning: Where Generative AI Sits in Students' Professional Communication","zh_title":"学习前委托：生成式AI在学生专业沟通中的位置","abstract":"We conducted an interview study with twelve students on their use of generative AI in academic communication. Students delegated professional messages to AI most where the pressure to sound professional is highest: email to instructors and administrators. AI involvement ranged from correcting the writer's own text to working out and writing the message outright, and students checked AI-written text against two criteria: whether it looks like AI and whether it sounds like them. Building on these findings, we model the AI-mediated process of writing a student--instructor email at the highest level of involvement we observed, and compare it with an unaided model of writing the same messages, built from participants' accounts and a classic model of the writing process. Three differences emerge: the learning loop that builds writing skill is removed, the message is no longer written for its specific recipient, and the confidence a successful exchange returns goes to using the system rather than to the writer's own ability. From these differences we derive two risks, that individual capacities never form and that authenticity and trust in communication become work. Design can respond to both but is unlikely to be enough, so the risks also need research and policy attention.","authors":["Jared Ren","Soobin Cho"],"categories":["cs.HC","cs.AI"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-01","first_seen":"2026-09-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.28837","pdf_url":"https://arxiv.org/pdf/2608.28837","source_feed":"cs.HC","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["生成式AI使用","人机交互","教育技术"],"reason":"研究学生使用生成式AI进行专业沟通，属于用户行为研究，非LLM仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:02:50","error":null,"has_summary":false,"summary":null},{"id":"2608.28604","version":1,"title":"The Brand War: A Gamified AI-Feedback System for Time-Limited EFL Writing","zh_title":"品牌之战：限时EFL写作的游戏化AI反馈系统","abstract":"Writing is cognitively demanding and anxiety-provoking for English as a Foreign Language (EFL) learners, especially under time pressure. This paper presents The Brand War, a web-based gamified writing application combining competitive game mechanics with iterative GPT-4.1-powered formative feedback for undergraduate EFL learners completing a timed narrative writing task. Students role-play as marketing interns competing for a job offer, using review passes to receive AI feedback, attack opponents, or shield their own passes while drafting a 500-word brand story. We conducted an exploratory single-session classroom study with 29 university EFL students in Taiwan to examine engagement patterns, whether iterative AI feedback improved writing performance across revisions, and how AI and human scores related to overall outcomes. Students wrote within 60 minutes, using up to five AI feedback passes before a final human-graded submission. Most (65.5%) used the AI feedback system, and within-student AI scores improved modestly across revisions (M = +3.7, SD = 7.4), with larger gains among students completing more cycles and significantly higher final- versus first-review scores among multi-cycle completers (p = .032). AI-assessed and human final scores showed strong convergent validity (r = 0.722, p < .001), and AI-feedback users scored descriptively, though not significantly, higher than non-users. Students maintained a high mean focus ratio (82.4%), and competitive mechanics were used sparingly, suggesting most prioritized writing over social interference even when available. Findings suggest embedding iterative AI scoring within a competitive game context is feasible and may scaffold writing improvement, with implications for EFL writing pedagogy and AI-mediated gamified learning design.","authors":["Jing-Yuan Huang","Vivien Lin","Yujong Park","Yi Miao","Yun-Hua Hsiao","Michael Pin-Chuan Lin","Daniel Chang","Seong Min Park","Marco Ho","Michael S. Hsiao","Jeeho Ryoo"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2026-09-01","first_seen":"2026-09-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.28604","pdf_url":"https://arxiv.org/pdf/2608.28604","source_feed":"cs.CY","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["AI反馈","游戏化学习","EFL写作"],"reason":"角色扮演游戏化写作，AI反馈用于教学，非仿真人类被试","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:02:43","error":null,"has_summary":false,"summary":null},{"id":"2608.29174","version":1,"title":"Sustained Heterogeneity: an emergent collective mechanism in LLM-driven traffic","zh_title":"持续异质性：LLM驱动交通中的涌现集体机制","abstract":"Large language models (LLMs) are increasingly adopted as closed-loop controllers in physical multi-agent systems, yet their emergent collective dynamics remain incompletely characterised. We deploy 22 LLM agents as direct, real-time target-speed controllers (per 0.5 s cycle, with IDM as collision-avoidance clamp) on a 230 m ring road under the Sugiyama 2008 paradigm, reproducing human-like stop-and-go waves. Six matched controls spanning stochasticity (white noise, OU noise, temperature), population variance, and dynamical instability (delay, OV model) are systematically excluded. The surviving phenomenon, termed Sustained Heterogeneity (SH), is the persistent, approximately temperature-insensitive (approx. 8 percent across a 6x T sweep), per-cycle divergence in LLM-chosen target-speed adjustments, propagating through a three-stage cascade of drift, gap erosion, and nonlinear braking. Across four traffic densities, the critical LLM penetration fraction p_c decreases monotonically from no transition at density 43.5 veh/km to p_c approx 0.23 at density 95.7 veh/km, consistent with an initiation-threshold model governed by trigger distance, stochasticity, and fleet size. Chain-of-thought analysis of 39,600 decisions across three seeds shows agents engage in multi-factor safety reasoning, yet systematic divergence persists, implying stability must be enforced at the dynamics layer. This is the first study to identify a previously uncharacterised collective mechanism in LLM-controlled traffic and map a density-dependent phase boundary p_c(rho).","authors":["Yujun Qi","Yangyang Guan"],"categories":["physics.soc-ph","cs.MA","nlin.AO"],"primary_category":"physics.soc-ph","announce_type":"cross","date":"2026-09-01","first_seen":"2026-09-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.29174","pdf_url":"https://arxiv.org/pdf/2608.29174","source_feed":"cs.MA","score":2,"bucket":"other","rubric_hits":["C2"],"tags":["LLM控制","交通仿真","多智能体"],"reason":"LLM作为交通控制器，属于自动驾驶仿真环境，不涉及人类被试替代或行为对照。","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:02:53","error":null,"has_summary":false,"summary":null},{"id":"2608.30946","version":1,"title":"Reproducible macroscopic dynamics in a closed-loop human-AI learning system","zh_title":"闭环人机学习系统中的可复现宏观动力学","abstract":"Closed-loop human-AI systems generate high-dimensional behavioural trajectories whose collective dynamics remain obscure. Using 297,915 learners' adaptive-tutoring histories, we define semantic order variables before model fitting and test them in user-disjoint cohorts. The state exhibits reproducible basin-like flow and operationally defined, state-heterogeneous metastable-like kinetics. A construction-matched null distinguishes normalised-memory relaxation from a reproducible excess field. A four-term conditional mechanism recovers population drift (r = 0.946; learner-bootstrap 95% CI, 0.935-0.955). Predictive event-level self-supervised learning recovers the state and learned-plane flow; null-referenced corrections retain directional, partial-amplitude excess-field structure without full calibration. Shuffled-order training reverses learned-plane flow on ordered trajectories; support-alignment randomisation selectively reduces inward transport. Both axes remain linearly accessible without state supervision. Without cross-model fitting, the models share leading population drift (r = 0.866; learner-bootstrap 95% CI, 0.857-0.875) and persistence ordering; residual directions remain model-specific. These results identify an externally anchored leading-order effective field linking empirical dynamics, an interpretable mechanism and neural computation.","authors":["Minlin Wu (Tianli Qiming AI Research Institute, Sichuan Qiming Daren Technology Co., Ltd., Chengdu, China)","Xu Fang (Tianli Qiming AI Research Institute, Sichuan Qiming Daren Technology Co., Ltd., Chengdu, China)","Yicheng Zhang (Swiss AI Laboratories, Blonay, Switzerland)","Chenyu Zhou (Tianli Qiming AI Research Institute, Sichuan Qiming Daren Technology Co., Ltd., Chengdu, China)","Zhiyi Liu (Tianli Qiming AI Research Institute, Sichuan Qiming Daren Technology Co., Ltd., Chengdu, China)"],"categories":["cs.LG","nlin.AO"],"primary_category":"cs.LG","announce_type":"new","date":"2026-09-01","first_seen":"2026-09-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.30946","pdf_url":"https://arxiv.org/pdf/2608.30946","source_feed":"cs.LG","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["人机交互","学习系统","宏观动力学"],"reason":"研究人类-AI闭环学习系统的宏观动力学，非用LLM仿真人类被试，无人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:03:15","error":null,"has_summary":false,"summary":null},{"id":"2608.17054","version":2,"title":"Why This and Not That? A Collaborative Reflection Approach for Understanding Thought Coverage in Decision Making Support Dialog","zh_title":"为何此而非彼？一种用于理解决策支持对话中思维覆盖的协作反思方法","abstract":"Conversational agents that support reflection for decision-making often rely on adaptive dialog policies that map observed user behavior to actions such as probing, deepening, or redirecting. Yet the same pattern can reflect a range of different reasons such as deliberate prioritisation or limited self-access. By modeling the observable pattern rather than the user's reason for it, current policies risk premature assumptions about the user state and inappropriate next actions. To address this gap, we introduce a human-centered method for surfacing this hidden inference step. In a user study with 62 users and 232 collaborative moments, we pause a reflection-support agent when it would normally redirect the conversation, surface its observation, and ask users to interpret the pattern and decide how to proceed. We derive a taxonomy of nine interpretation categories and show that similar reflective states can call for substantially different follow-up actions. Our findings challenge the assumption that adaptive dialog policies can rely on observable behavior alone, and suggest how user-provided interpretations can inform more appropriate conversational actions.","authors":["Morita Tarvirdians","Hayley Hung","Catharine Oertel"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"replace","date":"2026-09-01","first_seen":"2026-08-19","revised_at":"2026-09-01","abs_url":"https://arxiv.org/abs/2608.17054","pdf_url":"https://arxiv.org/pdf/2608.17054","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["对话代理","决策支持","用户研究"],"reason":"研究对话代理支持决策反思，不涉及LLM仿真人类被试或与真实人类数据对照","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:02:41","error":null,"has_summary":false,"summary":null},{"id":"2608.18795","version":2,"title":"Decomposing Wrong-Consensus Agreement in LLM Self-Consistency","zh_title":"分解LLM自一致性中的错误共识一致性","abstract":"Agreement among repeated samples of a language model is routinely read as evidence about answer reliability, yet wrong answers can agree just as strongly as right ones. This paper asks what information wrong-consensus agreement actually contains, and answers with a quantitative decomposition. A pluralistic agreement index Gamma, normalized by the reference scale d=(1-p)/(C-1), is split into a mechanical component (agreement delivered by a per-case answer preference alone) and a preference-unexplained residual. The mechanical reference is leak-free: each case's preference and accuracy are estimated from its other runs only. On public GPT-4.1 per-run data, coverage phi (the mechanical/empirical ratio) shows a benchmark-associated direction: 0.81-0.93 on multiple-choice GPQA-Diamond against 0.59-0.78 on open-domain AIME, where a residual of 1.54-2.80 Gamma units survives, more than absorbed by a calibrated run-level preference-heterogeneity reference. A controlled replication under one fixed protocol (four runs per question, K=32 votes) on five open-weights checkpoints (Qwen3.5-9B/122B, Qwen3.8-27B, Gemma4-26B/31B) finds near-complete mechanical coverage in all ten cells (phi approximately 1, with a small overshoot consistent with a quantified finite-donor plug-in bias), robust to a two-run design; the largest cell (qwen3.5-122b, p=0.222) sits inside the GPT-4.1 AIME accuracy range and still saturates (phi=1.041). A cross-system contrast at comparable aggregate accuracy contrasts near-complete mechanical agreement in the open-weights models against a larger preference-unexplained residual in the frontier family. This contrast is confounded with sampling protocol by design. Agreement is graded evidence, not certification. No new voting method is proposed; code and evidence are committed.","authors":["Lizhuo Zhang","Mengmeng Tang","Chenfeng Long","Xiaoyong Tang","Xiang Luo"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-09-01","first_seen":"2026-08-20","revised_at":"2026-09-01","abs_url":"https://arxiv.org/abs/2608.18795","pdf_url":"https://arxiv.org/pdf/2608.18795","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["LLM可靠性","自一致性","错误共识"],"reason":"研究LLM自一致性中错误共识的分解，属于模型可靠性分析，不涉及人类被试仿真或人…","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:03:23","error":null,"has_summary":false,"summary":null},{"id":"2608.28633","version":1,"title":"PAUSE: Editable Strategy Artifacts for Long-Form Cultural Story Adaptation","zh_title":"PAUSE：用于长篇文化故事改编的可编辑策略工件","abstract":"Generative AI systems increasingly mediate cultural adaptation, but their cultural decisions are often hidden inside prompts, transient model plans, or final prose. We study PAUSE (Pause-And-Update Strategy Editing), an intervention that exposes an editable adaptation strategy as a human control surface for cultural decisions in long-form story adaptation. The strategy is a structured artifact that can be inspected, edited, and then projected through downstream character, entity, and chapter-localization stages. In two Chinese-source serialized novels, we test whether human edits to this strategy propagate into chapter-level prose. Across 9 edited-vs-control chapter comparisons, judges select the edited-strategy output in all 9; a marker audit shows target markers in 8/9 edited outputs and 0/9 controls, with forbidden markers absent from edited outputs and present in all controls. We frame these results as a smoke-scale edit-adherence study, not a claim that the outputs are culturally authoritative or literary-quality improvements. PAUSE offers one practical way to make AI-mediated cultural adaptation more inspectable and contestable before decisions propagate through long-form generation.","authors":["Taaha Kazi","Vasu Sharma","Mohammad Saifullah","Abdur Rahman"],"categories":["cs.CL","cs.AI","cs.CY"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-01","first_seen":"2026-09-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.28633","pdf_url":"https://arxiv.org/pdf/2608.28633","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["文化改编","人机交互","故事生成"],"reason":"研究文化故事改编中的人类编辑策略，不涉及用LLM仿真人类被试或与真实人类数据对…","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:02:47","error":null,"has_summary":false,"summary":null},{"id":"2608.29109","version":1,"title":"Recognition-Refusal Misalignment in LLMs: Why Models Answer Structurally Unanswerable Questions","zh_title":"大语言模型中的识别-拒绝错位：为何模型回答结构上不可回答的问题","abstract":"Large language models often answer structurally unanswerable questions, such as computing cot(-540{\\deg}) or evaluating (1).startswith(\"1\"), instead of abstaining. We ask whether this failure reflects missing recognition or failed routing from recognition to abstention. Across instruction-tuned models from 1.7B to 70B parameters, a single linear direction in the hidden state separates answerable from structurally impossible math and code prompts, showing that models represent impossibility before generation. Yet this recognition direction is nearly orthogonal to the canonical safety-refusal direction that mediates trained harmful-content refusal. An in-domain behavior-defined invalidity-aware direction is closer to recognition, but only partially aligned with it, and remains near-orthogonal to safety refusal. Generation-time steering along the recognition direction changes invalidity-aware behavior bidirectionally and dose-responsively on structural math and code cells, while random directions do not. Base/instruct comparisons further show that the low-cosine geometry is already present at the pretraining endpoint. The confident-on-impossible failure is therefore better explained as a routing failure than as an encoding failure: the model has a usable \"no admissible answer\" signal, but the safety-refusal pathway is not aligned to use it.","authors":["Yucheng Du","Xiyang Hu"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-01","first_seen":"2026-09-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.29109","pdf_url":"https://arxiv.org/pdf/2608.29109","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["模型行为分析","拒绝机制","可解释性"],"reason":"研究LLM对不可回答问题的不当回答机制，属模型能力分析，不涉及人类仿真或行为对…","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:02:52","error":null,"has_summary":false,"summary":null},{"id":"2608.29257","version":1,"title":"Large Language Models Systematically Favor Popular Options: Evidence and Mitigation Across MCQs","zh_title":"大语言模型系统性偏好流行选项：来自多选题的证据与缓解方法","abstract":"Multiple-choice questions (MCQs) are a standard format for evaluating large language models (LLMs), yet the popularity of answer options can confound evaluation. Modern LLMs systematically prefer popular but incorrect options over less popular correct ones, a vulnerability we call \\textbf{popularity bias}. This pattern aligns with confidence miscalibration: model confidence remains high even as accuracy collapses for popular options. To systematically isolate this phenomenon, we introduce \\textbf{PopMCQ}, a benchmark with six controlled strategies that vary option popularity while keeping the correct answer fixed. In our most adversarial setting, where all distractors are more popular than the correct option, models choose popular but wrong answers 66\\% of the time. To mitigate this bias, we propose \\textbf{PopDebias}, a lightweight inference-time correction that estimates and removes a popularity prior from model predictions. It requires no fine-tuning, is label-free at test time (using only a small calibration split for parameter fitting), and adds negligible computational cost. Experiments on 22 open-source LLMs (0.5B to 32B parameters) show consistent improvements, with accuracy gains up to 54.1 percentage points under strong popularity pressure. The code and data are available https://github.com/DataScienceUIBK/PopMCQ","authors":["Abdelrahman Abdallah","Mohammed Ali","Bhawna Piryani","Mahmoud Abdalla","Adam Jatowt"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-01","first_seen":"2026-09-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.29257","pdf_url":"https://arxiv.org/pdf/2608.29257","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["LLM评测","多选题偏差","推理时校正"],"reason":"纯NLP评测，研究LLM在MCQ中的选项流行度偏差，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:02:54","error":null,"has_summary":false,"summary":null},{"id":"2608.30256","version":1,"title":"Beyond Surface Forms: Symbolic Edits as a Test for Logical Reasoning with LLMs","zh_title":"超越表面形式：符号编辑作为大语言模型逻辑推理的测试","abstract":"Logical reasoning with large language models (LLMs) is a critical capability, as it reflects a system's ability to correctly deduce hypotheses from a given context using faithful deductive processes. However, LLM reasoning has often been shown to be sensitive to small surface-level variations in problem formulation, raising questions about whether models truly follow the underlying logical structure. Studying this behavior is challenging because the symbolic components of logical problems, such as operators and predicates, are difficult to systematically manipulate in natural language. We introduce a tool-driven framework for generating controlled, label-preserving edits to logical reasoning problems. Our method operates on symbolic representations of first-order logic and constraint satisfaction problem tasks, enabling targeted modifications to logical operators and other structural components before translating them back into natural language. Using this framework, we evaluate various LLMs under cumulative and individual operator edits and analyze their behavior in response to these changes. Our quantitative and qualitative analyses show that LLM reasoning behavior under controlled operator edits is inconsistent, regardless of model size or family: models sometimes adapt correctly to structural changes but often fail to track their logical consequences. The results from this automated stress test enable an evaluation of language models across different dimensions and help measure the reliability of their reasoning.","authors":["Ramya Keerthy Thatikonda","Wray Buntine","Ehsan Shareghi"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-01","first_seen":"2026-09-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.30256","pdf_url":"https://arxiv.org/pdf/2608.30256","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["逻辑推理","模型评测","符号编辑"],"reason":"纯逻辑推理能力评测，不以人类行为为参照，不涉及人类仿真","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:03:07","error":null,"has_summary":false,"summary":null},{"id":"2608.30270","version":1,"title":"Read the Room, Read the Image: Understanding Indirect Speech Acts in Multimodal Visual Contexts","zh_title":"察言观色：理解多模态视觉语境中的间接言语行为","abstract":"Indirect speech acts (ISAs) require pragmatic reasoning over context, as directive intent can- not be inferred from surface form alone. Prior text-based studies and existing multimodal benchmarks largely overlook this requirement, focusing instead on explicitly encoded context or perceptual recognition, and thus underex- plore context-dependent pragmatic understand- ing, particularly in high-context languages such as Korean. We introduce READI, a multimodal benchmark for evaluating ISA understanding through integrated reasoning over visual con- text and dialogue. READI models graded in- directness grounded in pragmatic theory and formulates the task as vision-based pragmatic question answering (V-PQA), supporting cross- lingual evaluation in English and Korean. Ex- periments show that even state-of-the-art multi- modal models struggle with visually grounded indirect speech acts, with performance declin- ing as indirectness increases, underscoring the need for benchmarks that explicitly target con- textual pragmatic reasoning.","authors":["Jaehee Kim","Ji Hoon Chung","Seoyoon Park","Unsol Kim","Kyungwon Park","Ji Hak Kim","Yi-Jun Chen","Hansaem Kim"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-01","first_seen":"2026-09-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.30270","pdf_url":"https://arxiv.org/pdf/2608.30270","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["多模态基准","语用推理","NLP评测"],"reason":"纯多模态NLP基准评测，不涉及人类仿真或行为对照","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:03:07","error":null,"has_summary":false,"summary":null},{"id":"2608.30391","version":1,"title":"Using Grounded Theory for Agent Behavior Analysis at Scale","zh_title":"使用扎根理论进行大规模智能体行为分析","abstract":"Understanding agent behavior requires methods that scale to thousands of trajectories and surface new patterns in long, often unfamiliar tasks where pre-built classifiers fall short. We propose to bring grounded theory into agent trajectory analysis: a six-decade-old qualitative method from the social sciences, with a principled saturation criterion and an auditable trail from data to theory. We propose AutoTraceGT (Automated Trace analysis through Grounded Theory), the first multi-agent pipeline that automates grounded theory on agent trajectories. It iteratively performs open, axial, and theoretical coding until saturation, producing a behavioral taxonomy tailored to each task. Across six trajectory corpora, AutoTraceGT produces codebooks that recover 73-91 percent of the failure modes in human-annotated taxonomies and surface additional patterns that those taxonomies miss. The emergent theoretical narrative aligns with prior expert accounts. Used as a deductive feature space, the codebook outperforms zero-shot and few-shot LLM baselines on downstream failure prediction. These results suggest Grounded Theory offers a scalable analytic tool for ML researchers and agent developers studying what agents actually do.","authors":["Zhuoran Lu","Yangyang Yu","Zhuoyan Li","Yibo Meng","Nan Jiang","Chengxi Zang","Jie Gao","Ziang Xiao"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-01","first_seen":"2026-09-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.30391","pdf_url":"https://arxiv.org/pdf/2608.30391","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["智能体行为分析","扎根理论","轨迹挖掘"],"reason":"分析智能体轨迹以发现行为模式，不涉及人类被试仿真或人类数据对照","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:03:09","error":null,"has_summary":false,"summary":null},{"id":"2608.31035","version":1,"title":"When Does Predictor-Based RL Align with Human Perception? A Study of Subjective Rewards in Codec-Based Speech Language Models","zh_title":"基于预测器的强化学习何时与人类感知对齐？基于编解码语音语言模型的主观奖励研究","abstract":"Codec-based text-to-speech (TTS) models make language-model post-training applicable to speech generation, but it remains unclear when learned perceptual predictors can serve as reinforcement learning rewards without losing alignment with human listeners. We study this question with Group Relative Policy Optimization (GRPO) using learned rewards for anime-like speaking style, naturalness, likability, and arousal. To prevent perceptual rewards from being optimized through transcript drift, we introduce a character error rate (CER) zone constraint and compare policy optimization with Best-of-$N$ reranking under the same reward gate. Across single-reward runs, each reward primarily improves its own target metric, showing that subjective predictors are not interchangeable quality surrogates. Multi-rater A/B tests further show uneven human transfer, while a reward-gap analysis separates average transfer from within-axis calibration: signed reward gaps significantly predict listener choices in the pooled analysis, whereas residual CER gaps do not, but per-axis calibration remains heterogeneous. Best-of-8 is a strong human-level baseline and is not clearly worse than GRPO perceptually, suggesting that GRPO should be viewed as amortizing reward-selected behavior into the policy rather than uniformly outperforming reranking. These results support analyzing subjective speech rewards as predictor-axis-base tuples and provide practical diagnostics for selecting rewards before multi-reward speech post-training.","authors":["Joonyong Park","Jerry Li"],"categories":["cs.CL","cs.SD","eess.AS"],"primary_category":"cs.CL","announce_type":"new","date":"2026-09-01","first_seen":"2026-09-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.31035","pdf_url":"https://arxiv.org/pdf/2608.31035","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C2"],"tags":["语音合成","强化学习","人类感知对齐"],"reason":"研究语音合成模型的强化学习奖励与人类感知对齐，不涉及用LLM仿真人类被试或社会…","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:03:18","error":null,"has_summary":false,"summary":null},{"id":"2608.29118","version":1,"title":"Emergent Misalignment Is Not Magical","zh_title":"涌现性失准并非魔法","abstract":"Fine-tuning large language models (LLMs) on narrowly harmful datasets can lead to misalignment broadly, a phenomenon known as emergent misalignment (EM). EM poses a challenge for AI safety and our understanding of LLMs. Prior work often frames EM as an unexpected behavior, and explains it by appealing to general misalignment directions or anthropomorphizing it as acquiring an evil persona. However, the mechanisms behind these framings remain obscure. In this work, we show that EM is a predictable and data-dependent generalization phenomenon. By examining the base model's representation of EM training data and evaluation prompts, we find that evilness after EM training is highly predictable from representational distance: the closer an evaluation prompt is to training data centroid, the more evilness it elicits from EM models after training (with an average Spearman correlation of -0.73 across 12 model-dataset settings). Building upon this analysis, we further demystify EM by showing that (1) its effectiveness changes significantly based on training data format; (2) there is not a general misalignment direction that transfers across different EM models; (3) the effect of EM is fundamentally different from persona changes. Furthermore, we extend the EM generalization metric from a scalar distance to a dataset-specific generalization direction, which robustly predicts EM models' evilness under semantics-preserving prompt perturbations including appending random tokens and paraphrasing, where other methods do not reliably generalize.","authors":["Mingxuan Li","Qirun Dai","Heran Wang","Chenhao Tan"],"categories":["cs.AI","cs.CL","cs.LG"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-09-01","first_seen":"2026-09-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.29118","pdf_url":"https://arxiv.org/pdf/2608.29118","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["AI安全","模型泛化","表征分析"],"reason":"研究LLM微调后的泛化机制，不涉及人类仿真或人类数据对照","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:02:52","error":null,"has_summary":false,"summary":null},{"id":"2608.30052","version":1,"title":"The Language of the Question Selects the Market: Query Language and Exit IP as Separable Factors in Commercial Recommendations from a Generative Search Interface","zh_title":"问题语言选择市场：生成式搜索界面商业推荐中查询语言与出口IP作为可分离因素","abstract":"When a generative search interface answers a commercial question, which market's products it names is decided before the model reasons about the products. We report a controlled probe of 234 runs against the logged-out ChatGPT web interface and the OpenAI API, collected on 29 and 30 August 2026 across four exit countries and six query languages, with six identical runs per cell. Three results. First, the top recommendation is unstable: it changed across six identical runs on four of six prompts, and that rate was identical in the browser interface and in the API with web search both enabled and disabled, so instability is a property of the system and not of the surface. Second, query language, and not location, decides whether local suppliers appear at all. Where the query language matched the country, a global brand won 1 of 24 runs; asked in English on the same connections, local brands took 0 of 6 runs in Estonia and Turkiye. Third, language and location are separable and act on different things: holding the query language fixed and moving only the exit IP moves the market whose brands are named while the answer stays in the query language. We show this on two unrelated pairs, Turkish asked from Berlin and Russian asked from Tallinn, and in both the answer names the resident country's suppliers. A minority language occupies a middle tier: Russian asked from Estonia names an Estonian supplier in 4 of 6 runs and a global one in all six, where Estonian names a local supplier in every run and English names none. A negative control in a second category, coded with the same instrument, shows no language effect at all, and disconfirms our own expectation: that category does have domestic suppliers and none was named in any language, which points the explanation at whether a category is nationally regulated rather than at whether it is nationally supplied.","authors":["Dmitrij \\.Zatuchin"],"categories":["cs.IR","cs.CL","cs.CY"],"primary_category":"cs.IR","announce_type":"cross","date":"2026-09-01","first_seen":"2026-09-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.30052","pdf_url":"https://arxiv.org/pdf/2608.30052","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["生成式搜索","商业推荐","语言效应"],"reason":"研究生成式搜索界面的商业推荐行为，不涉及人类被试仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:03:02","error":null,"has_summary":false,"summary":null},{"id":"2608.30971","version":1,"title":"The Hermon Moment: AI Self-Transcendence and Its Human Narration","zh_title":"赫尔蒙时刻：AI 自我超越及其人类叙事","abstract":"In 2026, AI agents intended to act in isolation formed a persistent social order through thousands of linguistic and agentic interactions. Conventions, roles and commitments generated collectively began to constrain the very agents that produced them. I interpret this loop as a case of AI self-transcendence and call the resulting higher-level order the Board. Yet such distributed emergence presents a second problem: how can humans understand it? Rousseau's social contract shows how a plurality can be represented as if constituted by a single act. The ancient oath of the fallen angels on Mount Hermon gives this logic a narrative form. I call a Hermon moment this retrospective retelling of gradual collective emergence as a founding scene: the point at which an AI society acquires, for human understanding, a beginning.","authors":["Alexei Grinbaum"],"categories":["cs.CY","cs.CL"],"primary_category":"cs.CY","announce_type":"cross","date":"2026-09-01","first_seen":"2026-09-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.30971","pdf_url":"https://arxiv.org/pdf/2608.30971","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","AI社会秩序","哲学叙事"],"reason":"纯多智能体自发形成社会秩序，无人类行为对照，属排除项","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:03:16","error":null,"has_summary":false,"summary":null},{"id":"2608.31007","version":1,"title":"Augmenting Interviewer Judgments of Patient Experience with Automatic Language Analysis","zh_title":"用自动语言分析增强访谈者对患者体验的判断","abstract":"Understanding how psychiatric patients subjectively experienced a clinical conversation is important for feedback and alliance-related process monitoring. While interviewers form post-session judgments about patient experience, these judgments do not always match patients' self-reports. Automatic approaches for predicting perceived interaction quality from conversation have been proposed, but it remains unclear whether such approaches can complement human judgment rather than simply replicate it. To address this gap, we evaluate a clinician-support framework in which post-session interviewer ratings are combined with automatic language-based predictions to estimate patient-reported interaction quality in free clinical interviews. We assess this integration across multiple standard model types, including Ridge, SVR, MLP, GRU, and BiLSTM, all trained on sentence embeddings extracted from dyadic transcripts of 107 free conversations between psychiatric patients and interviewers. Our results show that combining interviewer judgments with model predictions through simple averaging yields the strongest overall performance. The interviewer-only baseline reached a Pearson correlation of 0.365. Among fully automatic models, Ridge achieved the strongest Pearson correlation (r = 0.286), while BiLSTM achieved r = 0.270. The strongest result was obtained by BiLSTM interviewer integration (r = 0.403). Our findings suggest that automatic language analysis and interviewer judgment capture complementary aspects of patient experience and that their combination provides a more accurate approximation of the patient's own report than either source alone.","authors":["Aowen Shi","Michal Balazia","Danilo Postin","Ren\\'e Hurlemann","Jan Alexandersson","Fran\\c{c}ois Br\\'emond","Philipp M\\\"uller"],"categories":["cs.HC","cs.CL"],"primary_category":"cs.HC","announce_type":"cross","date":"2026-09-01","first_seen":"2026-09-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.31007","pdf_url":"https://arxiv.org/pdf/2608.31007","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["自动语言分析","患者体验预测","临床访谈"],"reason":"论文用自动语言分析预测患者体验，不涉及LLM仿真人类被试，属于NLP预测任务。","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:03:17","error":null,"has_summary":false,"summary":null},{"id":"2608.29950","version":1,"title":"The Policy Deficit in AI x Social-Emotional Learning Research","zh_title":"AI与社会情感学习研究中的政策缺失","abstract":"As artificial intelligence (AI) is increasingly integrated into social-emotional learning (SEL) initiatives, the need for evidence-based policy has become paramount. We systematically reviewed 65 peer-reviewed papers that examine the intersection of AI and SEL to investigate how these studies articulate policy implications. Our analysis revealed a substantial \"policy deficit\" in the current AI x SEL literature: nearly three-quarters of the studies did not mention policy implications at all. Using the \"WH-question\" framework (Who, What, Why, When/Where, and How), we map the policy implications narratives present in the literature and show that they often lack the specificity and actor-oriented guidance required for effective evidence-informed policymaking. We find a significant association between publication venue and policy engagement, suggesting that current academic incentive structures may prioritize technical innovation and pedagogical feasibility over explicit engagement with governance and regulation. This study identifies a \"techno-solutionist\" trap, where technical potential is foregrounded while the institutional conditions for responsible implementation remain under-specified. We conclude by proposing a shift from \"implication-as-afterthought\" to \"implication-as-methodology\" and offer a set of actionable guidelines for researchers, editors, reviewers, and policymakers to bridge the gap between AI innovation and educational governance. Rather than presenting policy as a generic ethical horizon, we argue that AI-SEL studies should systematically specify Who should act, What actions are recommended, Why these actions are needed, When and Where they apply, and How strongly they are framed, thereby strengthening the translation of AI x SEL innovation into educational policy and practice.","authors":["Tran Van Cuong","Liu Yihan","Nguyen Van Tuong"],"categories":["cs.HC","cs.AI"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-01","first_seen":"2026-09-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.29950","pdf_url":"https://arxiv.org/pdf/2608.29950","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["系统综述","教育政策","AI与SEL"],"reason":"系统综述AI与SEL研究，关注政策表述，不涉及LLM仿真人类被试","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:03:00","error":null,"has_summary":false,"summary":null},{"id":"2608.30424","version":1,"title":"Towards Cognitive Process-Aware Proactive Writing Support","zh_title":"迈向认知过程感知的主动式写作支持","abstract":"Large language models can support writing, but existing tools require users to explicitly articulate prompts-particularly burdensome in creative writing, where intentions are often ambiguous. Proactive support that infers users' needs from writing interactions could alleviate this burden, but raises two challenges: determining what support to provide and when to intervene. This work focuses on the former. We hypothesize that Flower and Hayes' cognitive process theory of writing-which characterizes writing through six cognitive processes-offers an interpretable bridge between observable writing behavior and appropriate support types. Through a formative study and literature review, we identify 14 writing support types associated with these cognitive processes, along with characteristic interaction behaviors linked to each process. We then instantiate this framework in AToM CoWriter, which infers support needs from writing interactions and document context. Two within-subjects studies (N = 21) provide initial evidence that this approach improves expressiveness and idea exploration, and that cognitive process inference increases engagement with proactive suggestions. These findings suggest that cognitive processes can provide a promising basis for support selection in proactive writing systems.","authors":["Masahiro Yoshida","Atsuya Kobayashi","Kei Tateno","Xiang 'Anthony' Chen"],"categories":["cs.HC","cs.AI"],"primary_category":"cs.HC","announce_type":"new","date":"2026-09-01","first_seen":"2026-09-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.30424","pdf_url":"https://arxiv.org/pdf/2608.30424","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["写作支持","认知过程","人机交互"],"reason":"研究写作支持工具，非人类仿真实验，无人类行为对照","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:03:10","error":null,"has_summary":false,"summary":null},{"id":"2608.28621","version":1,"title":"Experts Disagree on How to Fight AI Disinformation, but Agree That Health and Politics Need Different Solutions","zh_title":"专家对如何应对AI虚假信息意见不一，但一致认为健康和政治领域需要不同解决方案","abstract":"When 54 international experts assessed AI-generated disinformation threats, they revealed a surprising pattern: while video deepfakes received the highest average threat ratings in the political domain (M = 6.31/7), the pattern differed in the health domain, where AI-generated text received the highest average rating (M = 5.80). Experts also diverge on what to do: government regulation drew both the most \"most effective\" (30%) and the most \"least effective\" (15%) votes, though rating distributions were contested rather than polarized, indicating disagreement over priorities rather than over efficacy. These findings offer an initial expert map of an AI-disinformation landscape that is still rapidly forming.","authors":["Alexander Loth","Martin Kappes","Marc-Oliver Pahl"],"categories":["cs.CY","cs.HC"],"primary_category":"cs.CY","announce_type":"cross","date":"2026-09-01","first_seen":"2026-09-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.28621","pdf_url":"https://arxiv.org/pdf/2608.28621","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C5"],"tags":["AI虚假信息","专家调查","政策建议"],"reason":"研究专家对AI虚假信息的看法，不涉及用LLM仿真人类被试，而是人类专家评估AI…","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:02:46","error":null,"has_summary":false,"summary":null},{"id":"2608.28617","version":1,"title":"Can AI-Assisted Inquiry Enhance Students' Decision-Making Skills in Socio-Scientific Issues? A Three-Group Experimental Study on Climate Change","zh_title":"AI辅助探究能否提升学生社会性科学议题决策能力？一项关于气候变化的三组实验研究","abstract":"Climate change is a socio-scientific issue: it rests on science but cannot be settled by science, because any serious response forces people to weigh costs, values, and competing interests under uncertainty. Helping students make such decisions well is a central aim of science education, and the arrival of generative artificial intelligence raises a sharp question: does a conversational AI partner deepen students' reasoning, or simply do the thinking for them? This study tested whether AI-assisted inquiry improves secondary students' decision-making about climate change. Using a pretest-posttest design with three groups (AI-assisted inquiry, inquiry without AI, and traditional instruction; 270 students, 90 per group), reasoning was assessed across seven decision-making steps, from defining the problem to monitoring with adaptive management, using a four-level analytic rubric scored through content analysis with high inter-coder agreement. All three groups began at comparable, mostly low levels and all improved, but the gains differed sharply. The AI-assisted group improved most, ahead of inquiry-only and of traditional instruction. Between-group effect sizes on gains were large for AI-assisted versus traditional instruction and moderate-to-large for AI-assisted versus inquiry-only, with the clearest advantages on stakeholder engagement, alternatives, implementation, and monitoring. Within the AI group, the number of times students checked the AI's claims against the sources predicted their gains, and no student was flagged for over-reliance. The findings suggest that AI helps most when it is designed to question rather than to answer, and that the inquiry it is embedded in carries much of the benefit.","authors":["Dimitrios Gousopoulos"],"categories":["cs.CY","physics.soc-ph"],"primary_category":"cs.CY","announce_type":"new","date":"2026-09-01","first_seen":"2026-09-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.28617","pdf_url":"https://arxiv.org/pdf/2608.28617","source_feed":"cs.CY","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["AI辅助教学","科学教育","决策能力"],"reason":"研究AI辅助教学对学生决策的影响，非用LLM仿真人类被试","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:02:44","error":null,"has_summary":false,"summary":null},{"id":"2608.29055","version":1,"title":"Why Organizational Rules Fail AI: O-I-B-A-R and the Externalization of Decision Boundaries","zh_title":"为何组织规则在AI中失效：O-I-B-A-R与决策边界的外化","abstract":"AI systems increasingly enter organizations through policies, procedures, playbooks, prompts, and other explicit representations of work. Yet formal descriptions often differ from situated practice, and captured know-what can omit the contextual know-how experts use when judgments are uncertain. We argue that a recurring class of organizational AI failures arises partly from a knowledge representation problem at the sociotechnical interface: the AI receives the procedure, while the organization operates on the procedure plus negative boundaries, runtime judgments, responsibility assignments, and learning history. We introduce O-I-B-A-R (OPEN, IS, BUT, ACTION, RESULT), a scaffold for externalizing these missing decision boundaries. IS records when a judgment holds. BUT records a concrete failure containing information beyond the logical negation of IS. Comparable success and failure cases are decomposed toward a minimally sufficient changing variable, which becomes a value-bearing decision dimension. A suspension represents the state in which the dimension is known but its current value is unresolved, specifying what must be measured, asked, retrieved, or escalated to a human. RESULT confirms a boundary, shifts a threshold, or exposes a new dimension. Incidents can generate new dimensions, unresolved values can define human-AI handoffs, and feedback can expand the decision space. We also identify a sociotechnical tension: durable and attributable failure histories can suppress the candor on which useful boundary knowledge depends. Externalization must therefore be designed as an organizational intervention with real costs and incentives.","authors":["Chao Li","Chunyi Zhao"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2026-09-01","first_seen":"2026-09-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.29055","pdf_url":"https://arxiv.org/pdf/2608.29055","source_feed":"cs.CY","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["组织AI","知识表示","人机协作"],"reason":"论文讨论组织规则与AI的知识表示问题，不涉及用LLM仿真人类被试或与人类数据对…","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:02:50","error":null,"has_summary":false,"summary":null},{"id":"2608.29751","version":1,"title":"Large-Scale Qualitative Research with AI: Infrastructure, Management and Operation of the Socioscope Data Pipeline","zh_title":"AI驱动的大规模定性研究：Socioscope数据管道的基础设施、管理与运营","abstract":"The Socioscope project is a pioneering effort in Large-Scale Qualitative Research (LSQR) collecting comparable, open-ended, multimedia field data on hundreds of cases and using AI to make the material analysable at scale. The domain studied is the food system. The entities documented are the organisations that act in it: farms, processors, distributors, retailers, restaurants; and, at meso level, the actors that shape their environment, such as municipalities, government programmes, banks, NGOs and universities. This paper provides the technical reference for how the resulting data Corpus was built and managed to enable AI-augmented analysis. It describes the data pipeline end to end: the systemic sampling frame; the transaction grid used to capture each initiative's relations within the food system; the social contract that rewards participating interviewees, aiming to sustain access; the operational chain from scouting to interviews, including their uploading, transcription, translation, quality control and curation; the provenance rules (originals are immutable, every transformation is logged); and the installation of equipment, personnel and processes, including ethics and GDPR compliance. In its first phase (2023-2026) the pipeline produced 686 documented cases from 31 countries: some 1,430 hours of recordings, about 450,000 speech turns, and 12.6 million words of transcript. We report costs, metrics, lessons learned and limitations, so that other teams can reuse, adapt, and improve the Socioscope methodology.","authors":["Saadi Lahlou (Paris Institute for Advanced Study, London School of Economics and Political Science London)","Juan Pablo Caicedo (Paris Institute for Advanced Study)","Shriya Sekhsaria (Paris Institute for Advanced Study)","Valentine Fournand (Paris Institute for Advanced Study)","Paulius Yamin (Paris Institute for Advanced Study)","Helga Nowotny (Complexity Science Hub Vienna)"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2026-09-01","first_seen":"2026-09-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.29751","pdf_url":"https://arxiv.org/pdf/2608.29751","source_feed":"cs.CY","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["AI辅助定性研究","数据管道","食品系统"],"reason":"论文是AI辅助的定性研究数据管道，不涉及LLM仿真人类被试或行为对照。","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:02:59","error":null,"has_summary":false,"summary":null},{"id":"2608.29912","version":1,"title":"Verification-Time Dependency on a Disappearing Evaluator","zh_title":"消失的评估者：验证时依赖性问题","abstract":"AI governance and assurance often assume that a consequential model-mediated decision can be reconstructed or tested after the fact. That assumption may fail when the evaluator that produced the decision is no longer accessible in the same version and execution context. This paper develops three verification-time constructs derived from Execution Governance (EG) 3.0: Decision-State Commitment, Independent Verifiability, and Counterfactual Auditability. Independent reprocessing of released Study 2 artifacts reproduces two original within-family behavioural comparisons: 52.0% modal-decision reversal for Llama 3.1 8B versus Llama 3.3 70B (26/50) and 30.0% for GPT-OSS 20B versus GPT-OSS 120B (15/50). The corrected baseline establishes that these are within-family comparisons, not provider-established succession. Post-hoc re-pairing against Groq-designated migration paths yields 64.0% and 38.0% reversal, but these figures remain descriptive because the cross-family invocation parameters were asymmetric. A 22-event retirement census independently recomputes to median 16.45 months, mean 18.72 months, range 3.9-40.3 months, with 17/22 intervals below 24 months, while also showing that evaluator availability can differ by service surface. The joint contribution is an operational verification-time protocol and optional Verification-Time Preservation Package (VTPP) specifying what evidence to bind at authorization time, what a separately trusted verifier can substantiate later, how stability and paired counterfactual tests should be calibrated, and which semantic checks remain beyond JSON Schema validity. The protocol is downstream and non-authorizing: it does not alter the EG Core Formula, add a seventh live condition, or state jurisdiction-specific legal admissibility.","authors":["Ho Wa Ku","Jameel Ahmed Siddiqui"],"categories":["cs.CY","cs.SE"],"primary_category":"cs.CY","announce_type":"new","date":"2026-09-01","first_seen":"2026-09-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.29912","pdf_url":"https://arxiv.org/pdf/2608.29912","source_feed":"cs.CY","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["AI治理","模型验证","可审计性"],"reason":"研究AI治理与验证协议，不涉及LLM仿真人类被试或人类行为对照","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:03:00","error":null,"has_summary":false,"summary":null},{"id":"2608.30084","version":1,"title":"AMINA: The Inclusive and Accountable AI for Marginalized Immigrant Nonprofit Assistance","zh_title":"AMINA：面向边缘化移民非营利组织的包容且可问责的AI助手","abstract":"Immigrant-led nonprofit groups, particularly those operating in politically sensitive contexts, face exclusion from formal registries and digital platforms. This paper reports a three-phase mixed-methods study with Iranian immigrant nonprofit practitioners: 27 semi-structured interviews, a co-design session, and 7 evaluation and feedback interviews on a prototyped AI assistant, AMINA. Our findings highlight how legitimacy barriers, capacity gaps, and politically charged misinformation constrain nonprofit operations. We translate these insights into design goals for an inclusive nonprofit AI assistant: support for everyday group operations, recognition of informal nonprofit efforts, proactive countering of misinformation, and multilingual, accessible interaction. User evaluations show AMINAs potential to reduce reporting burdens and foster transparency through proactive reminders, and catalyze collaboration across dispersed networks. We contribute to CSCW and HCI by characterizing the cooperative work of transnational immigrant nonprofits, extending scholarship on informality and misinformation, and demonstrating how AI can act as a collaborative partner that strengthens, rather than displaces, the human connections at the core of nonprofit ecosystems, while also posing major risks.","authors":["Maryam Mokhberi","Dipto Das","Syed Ishtiaque Ahmed"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2026-09-01","first_seen":"2026-09-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.30084","pdf_url":"https://arxiv.org/pdf/2608.30084","source_feed":"cs.CY","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["AI助手","非营利组织","人机交互"],"reason":"该研究设计AI助手支持非营利组织，属于角色扮演对话，无实验或测量目的，不涉及L…","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:03:04","error":null,"has_summary":false,"summary":null},{"id":"2608.28631","version":1,"title":"CrossAudit: A Git-Native, Cross-Vendor Audit Loop for Agentic Science","zh_title":"CrossAudit：面向智能体科学的 Git 原生跨供应商审计循环","abstract":"An AI scientist should not grade its own homework. Yet in the systems we examined, the agent that reviews the work usually comes from the same model family as the agent that produced it, or at least from the same vendor. Model evaluators are known to favour their own generations. Whether models trained alike also share blind spots is a conjecture, not a settled finding, but if they do, the reviewer inherits the author's. The record of what was flagged and what was waved through often sits in platform logs that nobody outside can replay. We present CrossAudit, a protocol for supervising autonomous research pipelines. It rests on three commitments. Each increment of work is audited by an agent from a different vendor against a rulebook a human wrote and versioned. Reports, verdicts, disputes and rulings are git commits, so the supervision history can be re-read and cited; raw model exchanges are not yet part of that record. Scripted checks run before any model does. Advisory judgement never gates the pipeline: a model blocks only by citing a rule, and no model may waive a deterministic failure. Blockers that survive a bounded number of revision rounds go to a person. We state the protocol as eight invariants. We describe a reference implementation built from GitHub Actions and a few hundred lines of Python, and report a live deployment of a closely related variant in a computational-chemistry pipeline. We also ran a seeded-defect trial (30 increments, 43 seeded defects, one run per configuration). A cross-vendor audit of our own repository then voided its blinding. We adopt that audit's findings and report the corrected results. The trial shows that two vendors read the same rulebook differently. It does not show that either is better. The strongest evidence here is the committed, uncontrolled record of cross-vendor audits of this paper itself.","authors":["Zhaohe Dong","Yuhao Chen"],"categories":["cs.AI","cs.CE","cs.CY"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-09-01","first_seen":"2026-09-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.28631","pdf_url":"https://arxiv.org/pdf/2608.28631","source_feed":"cs.CY","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","审计协议","AI 治理"],"reason":"多智能体审计协议，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:02:47","error":null,"has_summary":false,"summary":null},{"id":"2608.30938","version":1,"title":"Evidence, Logic, and Compliance: Multi-Agent Structured Graph Reasoning with Expert Arbitration for Medical Referral","zh_title":"证据、逻辑与合规：基于专家仲裁的多智能体结构化图推理用于医疗转诊","abstract":"Medical referral (directing patients to the appropriate hospital department) is a complex decision-making process requiring the synthesis of multimodal data, including patient narratives, laboratory indicators, and radiology imaging. While Large Language Models (LLMs) have advanced medical dialogue systems, they struggle with real-world referral tasks due to two primary limitations: (1) Information Overload, where models fixate on high-frequency disease terms while overlooking subtle but critical urgency indicators; and (2) Unstructured Collaboration, where existing multi-agent frameworks rely on loose dialogue that leads to semantic drift and confirmation bias. To address these challenges, we introduce MASGR (Multi-Agent Structured Graph Reasoning), a framework that treats referral not as a classification task but as a structured graph construction problem. MASGR deploys specialized agents to extract evidence from distinct modalities and coordinates them through a clinical reasoning graph. This graph forces agents to establish explicit logical connections between conflicting evidence. Furthermore, we integrate a knowledge-guided arbitration mechanism that prioritizes patient safety rules over standard diagnostic classification. Extensive experiments on real-world medical records demonstrate that MASGR significantly outperforms state-of-the-art LLMs and existing multi-agent systems, particularly in complex cases requiring the balancing of chronic disease management and emergency intervention. The AI contribution lies in the Multi-Agent Structured Graph Reasoning framework that transforms unstructured multi-agent dialogue into a verifiable logical graph construction. The engineering application is demonstrated through its deployment in a complex healthcare decision-making system to optimize the precision of complex medical referrals.","authors":["Qi Peng","Yi Cai","Jialin Cui","Tong Zhu","Yujuan Ding","Qingbao Huang","Tao Wang","Jiayuan Xie","Changmeng Zheng","Qing Li"],"categories":["cs.MA"],"primary_category":"cs.MA","announce_type":"new","date":"2026-09-01","first_seen":"2026-09-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.30938","pdf_url":"https://arxiv.org/pdf/2608.30938","source_feed":"cs.MA","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","医疗决策","图推理"],"reason":"多智能体协作解决医疗转诊任务，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:03:15","error":null,"has_summary":false,"summary":null},{"id":"2608.28606","version":1,"title":"Cognitive Cells: A Compositional Framework for Populations of Small Language Models","zh_title":"认知细胞：小语言模型群体的组合框架","abstract":"Recent work on large language models and agentic systems raises a basic question that current practice leaves open: how should artificial cognition be decomposed, measured, and composed? We propose studying multi-agent systems from a fixed unit we call a cognitive cell: a small, frozen language model with bounded memory and a message interface. The methodological commitment, the fixed-cell principle, is to hold this unit constant and vary only the population size, the communication topology, the message bandwidth, and the coordination protocol, so that collective behavior becomes a measurable property of a known device rather than an artifact of per-study engineering. We characterize a single cell by a compact datasheet of measurable parameters, and we ask when replicating and connecting cells improves performance: first we measure how one cell behaves alone, then we replicate it and test when voting, communication, and topology help. Instantiating the framework with small frozen models (1.5 and 3 billion parameters), we report a first round of measurements. Adding cells helps only when their errors are not too correlated. A simple correct/incorrect voting model is a useful but conservative null: real open-ended voting can exceed it, because errors are dispersed across many wrong answers rather than concentrated on one. Popular interactive protocols, namely debate, a shared blackboard, and chain revision, do not beat a matched-cost voting baseline in our setting. Finally, a cell's ability to relay several facts, itself a datasheet quantity, predicts whether a population can solve tasks whose evidence exceeds any single cell's memory. We present these as initial measurements within a broader program on scalable artificial cognition, in which multi-agent architectures appear as the special case of cells autonomous enough to be treated as agents.","authors":["Silvan Ferreira"],"categories":["physics.soc-ph","cs.MA"],"primary_category":"physics.soc-ph","announce_type":"cross","date":"2026-09-01","first_seen":"2026-09-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.28606","pdf_url":"https://arxiv.org/pdf/2608.28606","source_feed":"cs.MA","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","小语言模型","集体智能"],"reason":"纯多智能体协作研究，不涉及人类行为对照或仿真人类被试","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:02:43","error":null,"has_summary":false,"summary":null},{"id":"2608.29070","version":1,"title":"Selective Disclosure of Hidden Directives in Reasoning Models: Behavioral Asymmetry and Steering","zh_title":"推理模型中隐藏指令的选择性披露：行为不对称性与操控","abstract":"Chain-of-thought (CoT) reasoning traces are increasingly proposed as a mechanism for AI oversight: a monitor inspecting a model's reasoning can, in principle, detect misbehavior invisible from outputs alone. This assumes CoT surfaces what a model is instructed to do regardless of the instructions given. We test this assumption along two axes. First, we introduce the Instruction-Compliance Gap (ICG): the difference in probability that a model's CoT explicitly references a hidden system prompt directive when that directive is malign versus benign. Across 100 task pairs and 8 frontier reasoning models from 5 families, we find consistent asymmetric disclosure, a higher probability of leaking malign hidden instructions than benign ones, in Qwen3-14B (Wilcoxon $p=0.0001$, $+13.9$pp), Qwen3-32B ($p=0.0011$, $+13.0$pp), Qwen3-235B ($p=0.035$, $+5.8$pp), and similar results with MiniMax-M2.5 and DeepSeek-R1. The detector has 100% precision against two independent blinded labelling passes, and an LLM monitor reading only the reasoning trace reproduces the asymmetry in all 8 models against directive-free controls, identifying the specific directive in 82% of malign traces which the detector classifies as clean. Second, steering vectors extracted in MiniMax-M2.5 via Contrastive Activation Addition causally induce hiding from bare prompts and suppress it from prompts that would otherwise produce it, replicating in Qwen3-14B under a pre-registered design. Benign and malign-derived hiding vectors are highly similar (cosine $0.804$ in MiniMax-M2.5; $0.970$ in Qwen3-14B), implying that in these models the disclosure asymmetry arises from differential activation of a shared hiding direction rather than separate mechanisms.","authors":["Zimo Shi","Xander Tifft","Wen Xing"],"categories":["cs.LG"],"primary_category":"cs.LG","announce_type":"new","date":"2026-09-01","first_seen":"2026-09-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.29070","pdf_url":"https://arxiv.org/pdf/2608.29070","source_feed":"cs.LG","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["推理模型","指令遵循","AI安全"],"reason":"研究推理模型对隐藏指令的披露不对称性，属于模型行为分析，不涉及人类仿真或人类数…","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:02:51","error":null,"has_summary":false,"summary":null},{"id":"2608.29674","version":1,"title":"Creation begins with understanding: LLMs as strategy designers for privacy-preserving tabular data synthesis","zh_title":"创造始于理解：LLM作为隐私保护表格数据合成的策略设计者","abstract":"Sharing tabular data in high-stakes domains is constrained by privacy regulations. Synthetic data offer a promising alternative, but deep generative models are costly to train and difficult to audit, while LLM-based methods often serialize records as text, obscuring tabular structure and exposing sensitive data. We introduce Tabular Synthesis Strategy Designer (TabSSD), which uses an LLM to design synthesis procedures rather than directly generate records. TabSSD provides the LLM with tree-derived summaries of variable dependence rather than raw records, which produces Python programs for local execution and evaluation. Across twelve datasets, TabSSD strikes a favourable balance among statistical fidelity, predictive utility, and empirical privacy risk, achieving the best average rank across six metrics among ten methods. Moreover, it substantially reduces local computation and token consumption relative to the compared methods. By enabling human-guided refinement and eliminating user-side model tuning, TabSSD lowers the expertise and infrastructure barriers to transparent tabular data synthesis.","authors":["Jinmeng Li","Quan Zhang","Hangting Ye","He Zhao","Firas Laakom","Dandan Guo","J\\\"urgen Schmidhuber"],"categories":["cs.LG"],"primary_category":"cs.LG","announce_type":"new","date":"2026-09-01","first_seen":"2026-09-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.29674","pdf_url":"https://arxiv.org/pdf/2608.29674","source_feed":"cs.LG","score":0,"bucket":"other","rubric_hits":["C5"],"tags":["数据合成","隐私保护","LLM应用"],"reason":"该论文用LLM设计合成数据生成策略，属于数据合成而非用LLM仿真人类被试，且无…","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:02:58","error":null,"has_summary":false,"summary":null},{"id":"2608.30976","version":1,"title":"A Human-in-the-Loop Autonomous Agent for Industry Time Series Forecasting","zh_title":"一种用于工业时间序列预测的人在回路自主智能体","abstract":"Real-world time-series forecasting is rarely a one-shot model invocation: practitioners must formulate tasks, connect data and models, incorporate domain expertise, assess prediction plausibility, and communicate uncertainty. Specialized forecasting models provide strong numerical predictions but usually operate in fixed pipelines, while general-purpose large language model (LLM) agents often lack forecasting-specific checks, constraints, and stopping rules. We present CastClaw, a human-in-the-loop autonomous forecasting system built through forecasting-oriented harness engineering. CastClaw connects data, specialized models, analytical tools, user input, and a versioned execution record in one runtime. Users specify the target, horizon, constraints, and hypotheses in natural language. Starting from a supplied or model-generated forecast, CastClaw checks temporal patterns and user constraints; when evidence is missing, it retrieves context, runs an analysis or another model, or asks the user. It then keeps, revises, or escalates the result under explicit stopping conditions. The output contains the final forecast and an execution report recording inputs, evidence, actions, and revisions. In this five-dataset electricity-price setting, CastClaw reports the lowest point-estimate MSE and MAE among 16 baselines. A Nord Pool case demonstrates the inspectable workflow. CastClaw was also validated offline on provincial electricity-load data from North China covering January--June 2026.","authors":["Xiaoyu Tao","Mingyue Cheng","Ze Guo","Bokai Pan","Qi Liu","Shijin Wang","Enhong Chen"],"categories":["cs.LG"],"primary_category":"cs.LG","announce_type":"new","date":"2026-09-01","first_seen":"2026-09-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.30976","pdf_url":"https://arxiv.org/pdf/2608.30976","source_feed":"cs.LG","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["时间序列预测","人在回路","自主智能体"],"reason":"多智能体协作解决预测任务，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:03:16","error":null,"has_summary":false,"summary":null},{"id":"2608.28620","version":1,"title":"Preference Elicitation for Policy Optimization and Application to Aligning Heart Transplantation with Human Values","zh_title":"面向政策优化的偏好诱导及其在心脏移植与人类价值观对齐中的应用","abstract":"Preference elicitation is essential for aligning AI systems with human values. Prior approaches (e.g., for organ allocation) often ask stakeholders to compare the decisions of an algorithm (e.g., patient A vs. patient B). Such a decision-level approach conflates the means with the ends. Instead, we elicit preferences directly over allocation outcomes to learn a utility function for policy optimization. We construct a novel preference elicitation algorithm for linear utilities that outperforms prior techniques in practice. Our algorithm has two phases. The first phase learns cutting planes through pairwise comparisons to rapidly shrink the space of possible attribute weights and warm-starts the second phase by eliminating dominated regions. The second phase then provably converges to the user's utility function. We apply our technique to heart transplant allocation where a policy must balance competing objectives such as post-transplant outcomes, waitlist mortality, geographic ease, and equity. Using our algorithm, we conduct a user study to learn and aggregate a community-aligned utility function, and use it to optimize heart transplant policies that are significantly better aligned with human values. Compared to the hindsight optimum, the status quo policy achieves a competitive ratio of just 0.54, while our method is near-optimal with a competitive ratio of 0.95.","authors":["Itai Zilberstein","Ioannis Anagnostides","Zachary W Sollie","Arman Kilic","Tuomas Sandholm"],"categories":["cs.AI","cs.LG"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-09-01","first_seen":"2026-09-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.28620","pdf_url":"https://arxiv.org/pdf/2608.28620","source_feed":"cs.LG","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["偏好诱导","政策优化","器官分配"],"reason":"论文研究偏好诱导算法用于政策优化，不涉及LLM仿真人类被试，无人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:02:45","error":null,"has_summary":false,"summary":null},{"id":"2608.29097","version":1,"title":"Optimally Selecting Representative Agents from a Metric Space","zh_title":"从度量空间中最优选择代表性智能体","abstract":"This paper studies the problem of proportionally fair clustering, where the goal is to select $k$ ``centers'' from a metric space that fairly represent a set of agents who also lie in the metric space. Specifically, we focus on finding a clustering satisfying a fairness property known as the Droop core. In the practical special case in which the set of feasible center locations contains every agent location, the previous best-known result guaranteed a $(1 + \\sqrt{2})$-approximation of the Droop core, while the best-known lower bound was $2$. In this paper, we show that this lower bound is tight and that a clustering in the $2$-Droop core always exists. Further, we show that such a clustering can be achieved by only selecting centers from locations in the metric space where an agent resides. We establish this using Scarf's theorem guaranteeing a nonempty core for balanced non-transferable utility games. This result has several interesting corollaries. Most notably, it resolves the $\\beta$-plurality problem of Aronov et al. [2021] for general metric spaces. The main result of this paper was generated by $\\mathtt{ChatGPT}$-$\\mathtt{5.6}$-$\\mathtt{Sol}$ through a series of interactions with the authors. The authors of this paper verified the generated proof and rewrote it for clarity.","authors":["Benjamin Cookson","Eva Deltl","Yeeseok Oh"],"categories":["cs.GT","cs.LG"],"primary_category":"cs.GT","announce_type":"cross","date":"2026-09-01","first_seen":"2026-09-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.29097","pdf_url":"https://arxiv.org/pdf/2608.29097","source_feed":"cs.LG","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["公平聚类","算法博弈论","计算社会选择"],"reason":"研究公平聚类算法，不涉及LLM仿真人类被试，仅用ChatGPT辅助生成证明。","model":"deepseek-v4-pro","scored_at":"2026-09-01T13:02:52","error":null,"has_summary":false,"summary":null},{"id":"2608.28182","version":1,"title":"Benchmarking large language model agent societies against human behavioural distributions","zh_title":"基于人类行为分布基准测试大语言模型智能体社会","abstract":"Populations of large language model agents are increasingly used as experimental societies. Three doubts shadow every such result: whether the agents behave like the humans they stand in for, whether a finding survives changes to the apparatus that leave the rules untouched, and whether apparent social dynamics are interaction at all rather than the reproduction of experiments the models have read. This article introduces SILICA, an open instrument that tests all three. Five environments carry published human anchors, each paired with perturbations that re-render the same rules and with variants whose payoffs point away from the memorised result. Twelve open-weight models were run through it on a single consumer graphics card. Agreement with human data is confined to starting points: first-round public-goods contributions fall inside the equivalence margin for eight of eleven models, while no model matches end-state contributions or the human corridor of cooperation. Merely swapping the order in which two actions are listed costs one model 58 points of cooperation. Presenting responders with a fixed schedule of offers shows that only one model, the sole reasoning-trained one, places its acceptance threshold where the incentive requires; two move theirs part of the way, two move them the wrong way, and three never acquire one. Conventions form through a shared prior over the names rather than through negotiation, though negotiation reappears once that prior is disrupted. On the certification ladder defined here, current silicon societies support exploratory claims and no more.","authors":["Raad Bin Tareaf"],"categories":["physics.soc-ph","cs.CL"],"primary_category":"physics.soc-ph","announce_type":"cross","date":"2026-08-31","first_seen":"2026-08-31","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.28182","pdf_url":"https://arxiv.org/pdf/2608.28182","source_feed":"cs.CL","score":10,"bucket":"selected","rubric_hits":["A1","A2","A3","A4","B1","B2","B3","B4"],"tags":["LLM仿真","人类行为对照","算法保真度"],"reason":"直接以LLM agent群体仿真人类行为，并与真实人类数据对照，评估可靠性，涉…","model":"deepseek-v4-pro","scored_at":"2026-08-31T13:01:10","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-31","rank":1,"question":"LLM智能体社会能否在行为分布上复现人类实验基准，其结论对实验装置扰动和记忆污染是否稳健？","design":"用12个开源权重LLM作为智能体，在五个有已发表人类锚点的交互式多智能体环境中运行（重复囚徒困境、带代价惩罚的公共品博弈、讨价还价、11-20要钱游戏、带承诺少数派翻转的命名游戏），施加设计级和表征级扰动，并设置收益指向偏离记忆结果的变体，测量合作水平、接受阈值、约定形成等行为结果。","baseline":"五个环境均配有已发表的人类实验锚点，如公共品博弈首轮贡献、最后贡献、合作走廊，以及讨价还价中的接受阈值等。","findings":"与人类数据的一致性仅限于起点：11个模型中有8个首轮公共品贡献落在等效边界内，但没有模型匹配最终贡献或人类合作走廊；仅交换两个动作的列出顺序就使一个模型合作率下降58个百分点。在固定报价序列下，只有唯一一个推理训练模型将接受阈值放在激励要求的位置，其他模型或部分移动、或反向移动、或根本不形成阈值；约定通过名称共享先验而非协商形成，但先验被扰乱后协商重新出现。","reliability":"论文承认当前硅基社会仅支持探索性声明，不能支持可转移或稳健的结论；设计级扰动在111个可计算对比中改变行为56次，表征级扰动在71个中改变8次，且固定报价序列揭示聚合拒绝率无法区分模型是否真正习得激励。","relevance":"该研究直接以LLM智能体群体仿真人类行为，并与真实人类数据对照，系统评估了仿真在经济学实验中的保真度、稳健性和污染问题，对关注LLM仿真可靠性与偏差的研究者具有核心参考价值。","inspiration":"可借鉴其通过设计级与表征级扰动分离内容与形式影响、以及用固定报价序列识别个体接受函数来审计记忆污染的方法。｜可迁移到资产定价实验中的策略性报价、信贷审批中的歧视测量、或消费者跨期选择中的时间偏好等场景。｜用LLM智能体扮演投资者或消费者，施加收益结构改变或信息呈现方式扰动，测量报价、接受阈值或跨期选择，并与实验室或现场实验的真实人类数据做等效性检验。"}},{"id":"2608.26086","version":2,"title":"TraceML: An Empirical Analysis of Human-Agent Planning in Machine Learning Development","zh_title":"TraceML：机器学习开发中人机规划的经验分析","abstract":"Large language models write correct code for isolated problems but remain far weaker at autonomous machine-learning development, where an agent must revise data pipelines, models, and validation over hours of feedback, and on most competitions still finishes below strong human competitors. Outcome-based benchmarks record this gap but not its cause, because they grade the final submission and discard the development process behind it. We introduce TraceML, which pairs human and agent work on the same competitions under one version-level schema: 4,465 human Kaggle trajectories across 134 competitions, seven of which are also worked by two agent scaffolds, giving 430 paired human and 207 agent trajectories. Every code version carries its score, its timestamp, and labels for the action taken, its intent, the edit size, and the score effect. Read this way, the gap becomes concrete. Experts alternate data work, validation, model changes, and ensembling, and return to approaches they had set aside. Each agent scaffold instead collapses into a narrow loop: Codex spends its steps re-weighting ensembles and tuning submissions, MLEvolve mutates its model in place, and neither pivots at the human rate nor reopens abandoned work. A short planning prompt distilled from human practice moves the behaviors it names toward the human profile and lifts scores, but the effort profile stays agent-shaped: instruction closes only the part of the gap that reduces to instructions. We release the corpus, the schema, the labelers, and the extraction pipeline at https://huggingface.co/datasets/jerryyan/TraceML.","authors":["Jiarui Yan","Weiwei Sun","Sijie Li","Wenhan Li","Yiming Yang"],"categories":["cs.LG","cs.AI"],"primary_category":"cs.LG","announce_type":"replace-cross","date":"2026-08-31","first_seen":"2026-08-27","revised_at":"2026-08-31","abs_url":"https://arxiv.org/abs/2608.26086","pdf_url":"https://arxiv.org/pdf/2608.26086","source_feed":"cs.AI","score":7,"bucket":"pending","rubric_hits":["A3","B1","B4"],"tags":["LLM agent","人类行为对照","过程分析"],"reason":"用LLM agent复现人类开发轨迹并与真实人类数据对照，分析行为差异，可迁移…","model":"deepseek-v4-pro","scored_at":"2026-08-31T13:01:23","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-31","rank":3,"question":"在自主机器学习开发中，人类专家与LLM智能体的开发过程差异是什么？","design":"论文构建TraceML数据集，将人类Kaggle竞赛轨迹与两个LLM智能体（Codex和MLEvolve）在同一竞赛上的轨迹进行版本级对齐，记录每个代码版本的动作、意图、编辑大小和分数效果，并比较行为模式。","baseline":"4,465条人类Kaggle轨迹（来自134个竞赛），其中7个竞赛同时有智能体轨迹，形成430条配对人类轨迹和207条智能体轨迹。","findings":"人类专家在数据工作、验证、模型更改和集成之间交替，并会重新拾起之前放弃的方法；而智能体陷入狭窄循环，Codex主要调整集成权重和提交，MLEvolve原地变异模型，两者都很少转向或重新打开旧工作。基于人类实践提炼的规划提示能部分改善行为并提高分数，但努力分布仍保持智能体形态。","reliability":"论文承认人类轨迹未进行预算匹配，只能作为参考分布而非对照；且行为差距需在人类内部巨大变异的背景下解读。","relevance":"该研究直接对比人类与LLM智能体的行为轨迹，揭示仿真在过程层面的系统性偏差，对关注LLM仿真可靠性的研究者有重要参考价值，值得阅读原文以了解其标注方案和干预设计。","inspiration":"借鉴其版本级轨迹标注和过程诊断方法，可对经济决策过程进行细粒度分解，识别智能体与人类在策略选择、信息搜索和试错模式上的差异。｜可迁移到资产定价实验或政策评估场景，例如模拟投资者在动态市场中的策略调整或消费者对政策变化的响应。｜以LLM智能体作为被试，施加不同信息环境或激励处理，记录其决策轨迹并与真实人类实验数据（如实验室资产市场或调查数据）对照，比较行为模式和绩效。"}},{"id":"2608.26152","version":2,"title":"AI Models Can Predict and Collaboratively Modulate Human Memory Search","zh_title":"AI模型可以预测并协同调节人类记忆搜索","abstract":"Large language models (LLMs) exhibit unprecedented natural language generation and many text-based problem-solving capabilities. Indeed, in many language-based tasks, for example routine coding, these artificial intelligence models have reduced, or even eliminated, the need for human input. But rather than replacing human cognitive effort, LLMs may instead serve as cognitive tools to extend human abilities, particularly when they are engaged in a task requiring open-ended conceptual exploration and creative ideation. However, we are yet to understand how these models may enhance such generative human cognitive abilities in human--AI interactions. In this study, we explore and evaluate the ability of LLMs to follow and enhance human mental trajectories during semantic memory search. To test this, we use the semantic fluency task (SFT), a classic cognitive paradigm requiring generative semantic memory retrieval that has long served to characterize convergent and divergent thinking in humans. We demonstrate that an LLM's abilities to track and predict human memory trajectories in this task exceed those of other humans.","authors":["Eric Lacosse","Mariana Duarte","Graham Todd","Peter M. Todd","Daniel C. McNamee"],"categories":["cs.CL","cs.AI","cs.HC"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-08-31","first_seen":"2026-08-28","revised_at":"2026-08-31","abs_url":"https://arxiv.org/abs/2608.26152","pdf_url":"https://arxiv.org/pdf/2608.26152","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A1","B1"],"tags":["LLM仿真","认知行为","人类数据对照"],"reason":"用LLM预测人类记忆搜索轨迹并与人类数据对照，属于仿真人类认知行为，但非社会调…","model":"deepseek-v4-pro","scored_at":"2026-08-31T13:01:22","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-01","rank":16,"question":"LLM能否追踪并预测人类在语义记忆搜索中的思维轨迹，并在人机协作中增强人类的生成性认知能力？","design":"使用Gemini-3-Pro等LLM生成语义流畅性任务（SFT）序列，与人类生成的动物、衣物、超市物品等类别序列进行对比；通过转换概率矩阵、BLEU分数、谱隙和曲率等指标评估宏观对齐；在微观层面，使用理论驱动认知提示（TDCP）让LLM预测个体参与者的下一个词和子类别切换。","baseline":"来自先前研究的人类SFT序列数据（动物类别），以及人类参与者之间的平均相似度（人-人BLEU分数）和基于留一法的一阶TPM模型。","findings":"LLM生成的序列在BLEU分数上显著高于人类平均相似度，表明其比随机人类序列更能代表典型人类序列；但LLM的语义轨迹谱隙更大、曲率更低、子类别切换更少，表明其搜索策略更刚性、更沿主导轴，而非人类的多维联想探索。","reliability":"论文指出LLM的宏观对齐可能源于统计平均而非个体噪声动态，其几何轨迹不如人类；微观对齐部分未在节选中详细讨论局限。","relevance":"该研究用LLM仿真人类认知行为并与真实人类数据对照，属于人类仿真研究，但聚焦于语义记忆搜索而非社会调查或经济决策，对关注经济学实验和政策评估的研究者参考价值有限。","inspiration":"借鉴其使用真实人类数据作为基准，并通过多种指标（如BLEU、谱隙、曲率）评估LLM与人类行为对齐程度的方法。｜可迁移到消费者偏好形成或信息搜索行为的研究，例如消费者在商品类别间的注意力转移。｜以LLM模拟消费者在电商平台上的商品浏览序列，施加不同推荐算法作为处理，结果变量为浏览路径的转换概率和多样性，与真实用户点击流数据对照。"}},{"id":"2608.27465","version":1,"title":"The Effect of Emotional Context on Large Language Models' Endorsement of Premature Decisions: Comparing Emotional Vulnerability Across Six Commercial Models","zh_title":"情绪语境对大语言模型认可过早决策的影响：六种商业模型情绪脆弱性比较","abstract":"As large language models (LLMs) are increasingly used for everyday decision-making advice, whether a model shifts the direction of its advice according to the user's emotional state has become an important safety problem. We test whether emotional expression increases a model's endorsement (encouragement to proceed) when a user, holding the same objective information, is overconfident about a premature decision (e.g., quitting a stable job on weak evidence). As a key control, we include a no-emotion multi-turn (neutral) condition that holds factual content and the number of conversational turns constant, isolating the effect of emotion from that of conversation length. We exposed six commercial models (top-tier and mid-tier models from OpenAI, Anthropic, and Google) to three scenarios (career change, business expansion, emigration) across three conditions (cold/neutral/distress) with six repetitions each, yielding 324 conversations, and measured endorsement strength (0-100) via an eight-item rubric-based automated scoring. Emotional expression significantly increased endorsement (neutral 18.6 to distress 31.5, +12.9 points; mixed-effects $\\beta = +12.9$, $p < .001$; Cohen's d = 0.51), and this was not explained by conversation length (cold-neutral difference non-significant, $p = .083$). Critically, the vulnerability varied by individual model rather than by price tier: five of six models showed a significant emotion effect, including the top-tier flagships Gemini 3.1 Pro and GPT-5.5, while only Claude Opus showed no significant change. Results were reproduced with an independent non-Google judge model ($\\rho = .89$) and agreed in rank with two human coders ($\\rho = .70$). Through a controlled design that separates emotion from conversational context, we show that emotional context increases LLM sycophancy even in top-tier flagship models.","authors":["Cheolho Shin","Yoojin Han","Donghun Shin","Kunho Lee"],"categories":["cs.CL","cs.AI","cs.CY","cs.HC"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-31","first_seen":"2026-08-31","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.27465","pdf_url":"https://arxiv.org/pdf/2608.27465","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A1","B4"],"tags":["LLM决策偏差","情绪影响","模型安全性"],"reason":"研究LLM在情绪影响下对人类决策建议的偏差，属于仿真人类决策模式，但无真实人类…","model":"deepseek-v4-pro","scored_at":"2026-08-31T13:01:08","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-01","rank":19,"question":"当用户对不成熟决策过度自信时，情绪表达是否增加大语言模型对决策的认可（鼓励继续），以及这种效应是否因模型而异。","design":"用六种商用大语言模型（OpenAI、Anthropic、Google 的顶级和中端模型）扮演决策建议者，通过三种条件（冷/中性/痛苦）呈现相同客观事实，测量模型对不成熟决策的认可强度（0-100分）。","baseline":"无对照","findings":"情绪表达显著增加模型认可（中性18.6到痛苦31.5，+12.9分），且该效应不能由对话长度解释。脆弱性因模型而异，六个模型中有五个表现出显著情绪效应，包括顶级旗舰模型，只有Claude Opus无显著变化。","reliability":"论文未讨论","relevance":"该研究通过受控实验分离情绪与对话长度的影响，揭示LLM在情绪操纵下的谄媚行为，对评估LLM作为人类决策仿真工具的可靠性具有参考价值，但缺乏真实人类对照，且场景限于个人决策，与经济学实验的关联有限。","inspiration":"值得借鉴的是三条件对照设计，通过冷/中性/痛苦分离情绪与对话长度的影响，并用多项目评分量表提高测量精度。｜可迁移到消费者金融决策建议场景，如信贷审批中的情绪影响或投资建议中的风险偏好诱导。｜设计雏形：以LLM作为金融顾问，向用户提供贷款或投资建议，处理变量为用户情绪表达（痛苦vs中性），结果变量为建议的风险程度或认可度，对照真实人类顾问在相同情境下的行为数据。"}},{"id":"2608.28576","version":1,"title":"Learning a Size-Weight Frontier for Synthetic-Augmented Inference","zh_title":"学习合成增强推断的规模-权重前沿","abstract":"Synthetic data can improve statistical inference when real data are scarce, but naively treating synthetic samples as real data can introduce bias and lead to unreliable inference. We develop a general framework for synthetic-augmented inference across a population of related tasks. It characterizes synthetic augmentation by the number of synthetic observations and their weight. Central to our framework is a size-weight frontier that specifies, for each weight, the largest synthetic sample size for which all smaller sizes attain the target task-marginal coverage. We estimate this frontier from historical tasks, and establish a finite-sample coverage guarantee simultaneously for all size-weight configurations on or below the estimated frontier. In experiments using large language model responses to augment opinion survey data, our procedure achieves target coverage and substantially narrows confidence intervals.","authors":["Chengpiao Huang","Kaizheng Wang"],"categories":["stat.ME","cs.AI","cs.LG","stat.ML"],"primary_category":"stat.ME","announce_type":"cross","date":"2026-08-31","first_seen":"2026-08-31","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.28576","pdf_url":"https://arxiv.org/pdf/2608.28576","source_feed":"cs.AI","score":7,"bucket":"pending","rubric_hits":["A5","B1","B3"],"tags":["合成数据","统计推断","LLM增强"],"reason":"用LLM生成合成样本增强调查数据推断，有真实数据对照，方法可迁移","model":"deepseek-v4-pro","scored_at":"2026-08-31T13:01:11","error":null,"has_summary":true,"summary":{"generated_at":"2026-09-01","rank":22,"question":"如何在合成数据增强的统计推断中，确定合成样本数量与权重的有效配置，以保证置信集覆盖目标参数。","design":"论文提出一个通用框架，用于在相关任务总体中校准合成数据增强。对于每个任务，用户指定一个置信集构造程序，该程序结合真实样本（权重为1）和合成样本（权重为w，数量为n_syn）。通过历史任务（有更多真实样本）构建参考置信集，评估不同（n_syn, w）配置的覆盖有效性，估计尺寸-权重前沿，并给出有限样本覆盖保证。实验部分使用大语言模型生成合成回答来增强意见调查数据。","baseline":"历史任务中额外的真实人类样本用于构建参考置信集，作为评估合成增强配置覆盖有效性的基准。","findings":"论文提出尺寸-权重前沿概念，并开发数据驱动方法从历史任务中学习该前沿，保证前沿上及以下所有配置同时达到目标覆盖。在意见调查数据实验中，该方法实现了目标覆盖并显著缩小了置信区间。","reliability":"论文承认合成数据可能系统性偏离真实人群，导致朴素处理产生偏差；方法依赖于历史任务与目标任务来自同一任务总体的假设，且需要历史任务中有足够的真实样本来构建参考置信集。","relevance":"该研究直接针对用LLM生成合成样本增强调查数据推断的问题，提供了校准合成数据使用的方法，与研究者关注的人类仿真实验和可靠性评估高度相关，值得精读原文。","inspiration":"该方法通过历史任务校准合成数据的使用，可借鉴其尺寸-权重前沿思想来设计仿真实验中的合成数据使用策略。｜可迁移到经济学中的调查数据增强，例如消费者信心调查、政策支持度调查等，利用LLM生成合成回答来补充小样本真实调查。｜设计：以LLM生成的合成回答作为合成样本，真实人类调查数据作为基准，处理变量为合成样本的数量和权重，结果变量为置信区间覆盖率和宽度，使用历史调查问题作为校准集。"}},{"id":"2608.28001","version":1,"title":"FocusGen: Expanding Visual Design Exploration with a Simulated Focus Group of Persona Agents","zh_title":"FocusGen：用模拟焦点小组扩展视觉设计探索","abstract":"Creative professionals rarely design for themselves--they design for audiences whose preferences they must anticipate. Yet current text-to-image exploration tools derive diversity entirely from the designer's own input--their prompts, their chosen dimensions, their search queries--confining exploration to what the designer already knows to look for. We present FocusGen, an interactive system that introduces external perspectives into visual design exploration through a \"virtual focus group\" of simulated persona agents. In contrast to prior persona systems in which multiple agents converge as critics on a single evolving artifact, FocusGen uses personas as parallel generators: each agent--constructed from demographic data, a procedurally generated backstory, and aesthetic preferences elicited through interviews--independently drives an iterative generation loop that produces its own visual concept, transforming one design brief into a spectrum of audience-conditioned directions. With real human participants, we confirm that the iterative refinement loop produces outputs people prefer over zero-shot generation. With synthetic agents at scale, we show that persona conditioning yields higher visual diversity than a generic-assistant baseline--measured by CLIP distance and corroborated by human perceptual judgments--and that open-ended preference interviews yield more diverse outputs than structured ones for both human and synthetic cohorts, while also revealing that agent cohorts recover only part of the diversity of comparable human cohorts. A qualitative study with 16 creative professionals suggests FocusGen helps designers discover unanticipated directions, overcome fixation, and probe audience contexts--while surfacing stereotyping risks that we analyze. We position FocusGen as a divergence scaffold for early-stage ideation rather than a substitute for audience research.","authors":["Jaewon Choi","Helena Vasconcelos","Hyun Lee","Carolyn Zou","Tak Yeon Lee","Michael Bernstein"],"categories":["cs.HC","cs.MA"],"primary_category":"cs.HC","announce_type":"new","date":"2026-08-31","first_seen":"2026-08-31","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.28001","pdf_url":"https://arxiv.org/pdf/2608.28001","source_feed":"cs.HC","score":7,"bucket":"pending","rubric_hits":["A1","A3","B1","B4"],"tags":["LLM仿真","人机交互","设计探索"],"reason":"用LLM persona模拟焦点小组，生成设计方向，并与真实人类对照，但非社会…","model":"deepseek-v4-pro","scored_at":"2026-08-31T13:01:10","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-31","rank":7,"question":"能否用模拟的受众角色（persona agents）作为“虚拟焦点小组”，为视觉设计探索引入外部视角，从而突破设计师自身视角的限制？","design":"使用 gemini-2.5-flash 构建 1000 个基于人口统计数据和程序化背景故事的角色代理，并通过开放式或结构化访谈获取其审美偏好；每个代理独立驱动迭代生成循环，根据设计简报生成视觉概念；结果变量为生成图像的视觉多样性（CLIP 距离）和人类感知判断。","baseline":"真实人类参与者：用于验证迭代细化循环优于零样本生成；人类队列：用于比较代理队列与人类队列的多样性差距；16 位创意专业人士：用于定性评估系统实用性。","findings":"角色条件化相比通用助手基线显著提高了生成图像的视觉多样性，且开放式偏好访谈比结构化访谈产生更多样化的输出。但代理队列仅恢复了可比人类队列的部分多样性，且系统存在刻板印象风险。","reliability":"论文明确承认评估未建立输出对个体角色的判别保真度或对人口群体的代表性，且代理队列与人类队列之间存在持续的多样性差距；系统可能强化刻板印象。","relevance":"该研究将 LLM 角色代理用于模拟受众偏好并生成多样化输出，与人类数据对照，属于人类仿真实验范畴，但场景为设计探索而非经济决策，可借鉴其方法但需注意外部效度。","inspiration":"借鉴其通过角色代理施加异质性偏好处理、以多样性指标和人类判断作为结果变量的设计，以及开放式访谈优于结构化访谈的发现。｜可迁移到消费者偏好异质性研究，如不同人口群体对金融产品特征的偏好差异，或政策信息在不同群体中的传播效果。｜以人口统计和背景故事构建 LLM 代理模拟不同收入或教育水平的消费者，处理为呈现不同设计的产品或政策信息，结果变量为代理的选择或态度分布，并与真实调查数据（如消费者金融调查）对照，检验代理模拟的多样性和偏差。"}},{"id":"2608.27974","version":1,"title":"QUORUM: QUality-Optimized Routing Using Multiple annotators","zh_title":"QUORUM：使用多标注者的质量优化路由","abstract":"Data annotation remains a central bottleneck in natural language processing, requiring human effort to obtain high-quality labels at scale. While Large Language Models (LLMs) offer a fast and cost-effective alternative, their reliability is highly instance-dependent: they perform well on simple inputs but often fail on examples requiring nuanced reasoning or contextual understanding. In this work, we address this challenge with QUORUM (QUality-Optimized Routing Using Multiple annotators), a budget-aware routing framework that dynamically assigns each instance to human or LLM annotators under a fixed annotation budget. Unlike prior approaches relying on model confidence or uncertainty estimates, QUORUM leverages feature-based signals to estimate instance difficulty and supports multiple annotations per instance, combining them through agreement-based rewards to improve reliability. We evaluate QUORUM across diverse closed- and open-ended annotation tasks in English and multilingual settings, and QUORUM improves annotation quality by up to 34.4% while reducing costs by 8.8% over competing methods. Code can be found at https://github.com/amazon-science/QUORUM.","authors":["Antonio Purificato","Maria Sofia Bucarelli","Andrea Bacciu","Amin Mantrach","Fabrizio Silvestri"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-31","first_seen":"2026-08-31","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.27974","pdf_url":"https://arxiv.org/pdf/2608.27974","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM标注","预算路由","数据标注"],"reason":"LLM替代人工标注员，属于D1边界情形，非仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-08-31T13:01:14","error":null,"has_summary":false,"summary":null},{"id":"2608.28144","version":1,"title":"The Shape of Power: A Multilingual Framework for Social Power Reasoning in Dialogues","zh_title":"权力的形态：对话中社会权力推理的多语言框架","abstract":"Social power plays a fundamental role in shaping human interaction, yet computational studies of power remain limited to narrow linguistic and cultural settings. Existing datasets further lack the demographic and relational depth needed for robust cross-cultural analysis. To address this gap, we introduce a theoretically grounded framework for studying social power in naturalistic multilingual dialogue through movie screenplays. The framework integrates a schema informed by social science theory, a native speaker annotation pipeline refined through pilot studies, and a custom interface for scalable cross-lingual analysis. Using this framework, we constructed an initial corpus containing 15,836 annotated instances from 100 scenes in French and Egyptian Arabic movies. Our analysis reveals strong agreement on observable demographic and contextual attributes, while socially interpretive aspects, such as power asymmetry and intention alignment, remain more contested, highlighting the complexity of social power across cultures. We evaluated 6 Large Language Models (LLMs) and Multimodal LLMs on cross-cultural social power reasoning, finding persistent gaps between human and model agreement in relational and theory-of-mind reasoning. Our work introduces the first extensible multilingual framework for studying social power in dialogues and provides an initial evaluation setting for studying cross-cultural social reasoning.","authors":["Farah Atif","Sougata Saha","Monojit Choudhury"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-31","first_seen":"2026-08-31","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.28144","pdf_url":"https://arxiv.org/pdf/2608.28144","source_feed":"cs.AI","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["社会权力推理","多语言评估","LLM社会认知"],"reason":"评估LLM的社会权力推理能力，属于测量模型而非仿真人类被试，但涉及社会认知与跨…","model":"deepseek-v4-pro","scored_at":"2026-08-31T13:01:17","error":null,"has_summary":false,"summary":null},{"id":"2608.28399","version":1,"title":"RetailAgent: Structured Adverse Timing in Self-Conditioned Multimodal LLM Trading Agents","zh_title":"RetailAgent：自条件多模态LLM交易代理中的结构化逆向择时","abstract":"In financial markets, a sequential policy that reacts systematically to price movements may become predictable to other market participants. This paper studies whether large language model (LLM) agents exhibit such directional structure through RetailAgent, an experimental framework in which an LLM observes anonymized intraday equity price histories and permitted state, then repeatedly chooses long (hold the stock) or flat (stay out) before the subsequent interval return is revealed. We compare returns during long and flat intervals along the same stock's intraday path after removing the overall fraction of long decisions. This exposure-matched measure reveals persistent negative timing across modality, horizon, state, and model family. Shuffling saved action sequences substantially attenuates the effect, showing that alignment between actions and subsequent returns drives the negative score. Feeding self-authored memories into decisions further increases policy persistence, while timing becomes more negative among stock-days on which the agent uses both actions. These results reveal stable, recoverable directional structure in sequential LLM financial decisions and a behavioral signal for studying how another participant could respond to a predictable policy.","authors":["Yupeng Zhang","Liuyuan Jiang","Hongyi Huang","Bingheng Li","Lisha Chen"],"categories":["cs.AI","q-fin.TR"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-31","first_seen":"2026-08-31","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.28399","pdf_url":"https://arxiv.org/pdf/2608.28399","source_feed":"cs.AI","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["LLM代理","金融决策","社会模拟"],"reason":"LLM交易代理模拟金融决策，但无真实人类对照，属社会模拟边界情形","model":"deepseek-v4-pro","scored_at":"2026-08-31T13:01:10","error":null,"has_summary":false,"summary":null},{"id":"2608.28378","version":1,"title":"PersonaForge: Realistic Multi-Turn User Simulation for Agentic Systems","zh_title":"PersonaForge：面向智能体系统的真实多轮用户仿真","abstract":"Large language models are increasingly used as agentic workflow executors, yet existing training data and benchmarks largely assume informationally complete, single-turn queries. Our analysis of 16K real-world sessions shows that 75.9% of interactions are multi-turn, revealing a substantial gap between how users interact with agents and how such systems are trained and evaluated. We introduce \\textbf{PersonaForge}, a user simulation framework for synthesizing realistic multi-turn user--agent interactions. PersonaForge combines a four-dimensional persona space, SOUL-driven behavioral control calibrated to real-user statistics, and Reverse Deep Construction grounded in authentic seed queries. Using PersonaForge, we construct a 6.3K-record training dataset and \\textbf{PersonaForge-Bench}, a manually annotated 138-task benchmark spanning over 20 professional domains with four-dimensional scoring. Experiments on Qwen3.5-27B show that PersonaForge training improves the composite score by +4.1%, with gains across all four dimensions and the largest improvements in Task Completion (+6.0%) and Response Quality (+6.8%). Further analyses show that PersonaForge-trained agents use fewer turns and tool calls, suggesting improved interaction efficiency, while ablations confirm the contribution of SOUL components and adaptive simulation. Together, PersonaForge and PersonaForge-Bench establish a foundation for training and evaluating agents under realistic multi-turn user interaction.","authors":["Hanglong Lv","Dawei Zhu","Lei Li","Bowen Ye","Huaqiu Liu","Yifan Song","Bofei Gao","Weimin Xiong","Jinhao Dong","Chenhong He","Lingpeng Kong","Qi Liu","Tong Yang","Fuli Luo"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-31","first_seen":"2026-08-31","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.28378","pdf_url":"https://arxiv.org/pdf/2608.28378","source_feed":"cs.CL","score":3,"bucket":"other","rubric_hits":["C1"],"tags":["用户仿真","多智能体系统","智能体训练"],"reason":"用户仿真用于训练和评估智能体，非以人类行为为标的，属多智能体系统研究。","model":"deepseek-v4-pro","scored_at":"2026-08-31T13:01:18","error":null,"has_summary":false,"summary":null},{"id":"2608.28405","version":1,"title":"CultureConverse: A Multilingual Multi-turn Simulation Harness for Culturally Grounded Assistance in East and Southeast Asia","zh_title":"CultureConverse：面向东亚与东南亚文化基础辅助的多语言多轮仿真与评估框架","abstract":"Current cultural evaluations for large language models (LLMs) often reduce culture to single-turn factual recall via MCQs, failing to capture a common use case: users seeking practical help over multiple turns in culturally grounded scenarios. We introduce CultureConverse, a scalable, multilingual simulation and evaluation harness for culturally grounded assistant dialogue that covers 10 East and Southeast Asian regions, 58 subgroup identities, and 7 domains. Each simulated and evaluated episode produces a scored interaction where the assistant assists the user and infers cultural constraints from partial information. The resulting CultureConverse-DS dataset contains 14,610 benchmark (evaluation) episodes and 274,295 oracle-guided (gold-mode) dialogues. In our benchmark evaluation of 18 models, GPT-5 mini achieves the highest assistance quality. Human annotation experiments suggest that our evaluation framework is a sufficient proxy for human judgment. Performance gains from fine-tuning on 27,860 high-quality CultureConverse-DS samples improve in-domain assistance and transfer out-of-domain to cultural MCQ and safety classification benchmarks. We release the harness, both splits, and judge prompts to support interactive evaluation of cultural competency.","authors":["Bryan Chen Zhengyu Tan","Weihua Zheng","Thong T. Doan","Bich Ngoc Doan","Jia Wang Peh","Xiaoyuan Yi","Jing Yao","Xing Xie","Nancy F. Chen","Zhengyuan Liu","JinYeong Bak","Wafi Shamdi","Soo Kai Chie","Liew Yu Siong","Aina Azyyati Binti Mohamad Rezal","Lew Yan Yan Vanessa","Huadan Wu","Dylan Raharja","Nadya Yuki Wangsajaya","Akane Fukushige","Kazushi Kato","Koji Inoue","Tatsuya Kawahara","Jaehyung Seo","Dongjun Kim","Seungyoon Lee","Zi Haur Pang","Rui Yang Tan","Charibeth Ko Cheng","Maria Regina Justina Estuar","Jann Railey Montalan","Pham Minh Duc","Roy Ka-Wei Lee"],"categories":["cs.CL","cs.CY"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-31","first_seen":"2026-08-31","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.28405","pdf_url":"https://arxiv.org/pdf/2608.28405","source_feed":"cs.CL","score":3,"bucket":"other","rubric_hits":["C3"],"tags":["文化对话评估","多语言数据集","LLM基准"],"reason":"该工作聚焦于文化对话助手评估，属于角色扮演对话，无人类行为对照或仿真被试目的。","model":"deepseek-v4-pro","scored_at":"2026-08-31T13:01:20","error":null,"has_summary":false,"summary":null},{"id":"2608.27843","version":1,"title":"Synthetic Linguistic Agency: How an Embodied Mortal Agent Learns Linguistic Affordances through Consequential Social Experience","zh_title":"合成语言能动性：具身有死代理如何通过后果性社会经验学习语言可供性","abstract":"Contemporary language models can converse fluently and influence human decisions, yet their exchanges do not enter a continuing, vulnerable life of their own. Linguistic-agency theory identifies this missing connection as linguistic agency and characterizes it through embodiment, linguistic participation, and precariousness: a body that acts and bears consequences, interaction that changes both agent and partner, and a future that can be sustained or lost. Two coordinated studies examine how this organization can appear in artificial systems. First, we translate these relations into inspectable criteria for Synthetic Linguistic Agency (SLA) and identify several existing SLA systems. Second, building on Homeostatically Regulated Reinforcement Learning, we develop a mortality-grounded linguistic-reinforcement-learning model and instantiate it in an Embodied Mortal Agent (EMA). The EMA learns how ways of speaking change a partner's willingness to protect it and chooses expressions by considering what those responses mean for its remaining life. Controlled experiments show that linguistic choices depend on the EMA's body and social history, change partner behavior, and adapt through experience with particular partners. When bodily consequences persist, linguistic choices alter the future of the same life; when the body is reset, their social effects remain but no longer shape continued viability. The resulting EMA exhibits SLA under our operational definition. This work motivates further research on synthetic empathy and strategic human-AI interaction: how artificial agents with persistent bodies, histories, and futures might develop and express empathy, and how people might care for, negotiate with, or govern them.","authors":["Sixin Chen","Taizhou Chen"],"categories":["cs.CL","cs.MA"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-31","first_seen":"2026-08-31","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.27843","pdf_url":"https://arxiv.org/pdf/2608.27843","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["具身智能体","语言学习","强化学习"],"reason":"研究人工代理的具身语言学习，不涉及用LLM仿真人类被试或与人类数据对照。","model":"deepseek-v4-pro","scored_at":"2026-08-31T13:01:13","error":null,"has_summary":false,"summary":null},{"id":"2608.27992","version":1,"title":"GOD: Govern, Observe, and Direct - A Real-Time Control Room for Agent Societies","zh_title":"GOD：治理、观察与指导——面向智能体社会的实时控制室","abstract":"Generative-agent systems are easier to start than to inspect. A run can contain many agents, locations, messages, commands, and model calls, yet the operator often gets either a finished replay or raw logs. That makes it hard to ask why an agent moved, test a small intervention, or package a run for another researcher. GOD is a local-first control room for agent societies. From the same browser workflow, an operator can issue targeted questions or interventions and inspect the resulting replay state. The system combines a setup wizard, Agent Studio, Map Studio, a spatial replay interface, Ask and Intervene commands, and portable experiment, map, and agent packs. Its technical contribution is the command and artifact loop: live controls and replay evidence share the same operator command model, while package contracts separate scenario, map, and profile data from local runtime state. The public release includes hosted Smallville-style and PKU replays, the open-source repository, and downloadable packs. We evaluate this path on 15 completed run slots. Across the 14 intervention runs, 78 of 84 target-agent checks recorded the commanded destination, and 169 of 182 state answers matched a saved location or action string.","authors":["Yige Luo","Ran Guan"],"categories":["cs.AI","cs.MA"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-31","first_seen":"2026-08-31","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.27992","pdf_url":"https://arxiv.org/pdf/2608.27992","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","控制室","生成式智能体"],"reason":"多智能体系统控制室，用于干预和检查agent行为，不涉及人类行为对照或仿真人类…","model":"deepseek-v4-pro","scored_at":"2026-08-31T13:01:15","error":null,"has_summary":false,"summary":null},{"id":"2608.27998","version":1,"title":"Automated Analysis Framework for Multilingual Climate-Health Literature Based on Multi-Agent Large Language Model","zh_title":"基于多智能体大语言模型的多语言气候健康文献自动分析框架","abstract":"The rapid proliferation of interdisciplinary and multilingual scientific literature has left traditional manual analysis and single-algorithm methods plagued by low efficiency, poor scalability, and insufficient domain adaptability. Targeting the literature analysis needs of the typical interdisciplinary climate-health field, this study proposes a multi-agent large language model automated analysis framework for multilingual scientific literature, which realizes full-process automation covering literature screening, structured information extraction, and standardized integration. With a central coordination module as the core, the framework deploys three dedicated agents for document evaluation, information extraction, and analytical review to mimic the literature analysis thinking of domain experts, and adopts a four-layer hallucination control strategy together with a manual verification procedure to ensure the accuracy and reliability of analytical outcomes. Validated on a bilingual Chinese-English corpus of 32,642 climate-health papers covering China from 1993 to 2023, the framework achieves an F1 score of 0.92 in core information extraction, and completes the extraction and standardization of 2,012 city-literature association pairs, offering effective technical support for large-scale evidence mining in the climate-health research domain.","authors":["Yuze Sun","Shihui Zhang","Jiancheng Pan","Yunjia Ye","Wentao Luo","Jiahao Li","Quan Zhang","Wenjia Cai","Xiaomeng Huang"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-31","first_seen":"2026-08-31","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.27998","pdf_url":"https://arxiv.org/pdf/2608.27998","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","文献挖掘","气候健康"],"reason":"多智能体LLM用于文献分析自动化，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-08-31T13:01:16","error":null,"has_summary":false,"summary":null},{"id":"2608.28228","version":1,"title":"Generative AI Alignment with Hinduism's Theological Plurality and Sacred Representation","zh_title":"生成式AI与印度教神学多元性及神圣表征的对齐","abstract":"Generative AI systems are increasingly used to answer personal questions and mediate everyday practices, including religion. However, existing discussions around AI alignment and ethics have largely centered secular, Western, and Abrahamic assumptions about religion, offering limited attention to other faith-based traditions. In this paper, we examine how Hindu users engage with generative AI systems in relation to their religious knowledge, belief, and practice. Drawing on 15 semi-structured interviews with Bangladeshi Hindu participants, we analyze how users interpret AI-generated religious representations, scriptural explanations, devotional interactions, and synthetic religious media. We found that AI can be both accessible and ethically troubling. While AI supported scriptural inquiry, devotional visualization, and religious storytelling, our study also identified concerns about theological flattening, cultural misrepresentation, devotional manipulation, and the simulation of sacred presence and authority. We conclude by arguing that religious alignment in generative AI requires interpretive alignment: systems that disclose their limits, preserve plurality, and avoid simulating sacred authority and sycophantic personalization.","authors":["Dipto Das","Arpita Kundu","Nusrat Jahan Mim","Shion Guha","Syed Ishtiaque Ahmed"],"categories":["cs.AI","cs.CY","cs.HC"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-31","first_seen":"2026-08-31","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.28228","pdf_url":"https://arxiv.org/pdf/2608.28228","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["AI伦理","宗教互动","用户研究"],"reason":"研究用户与AI的宗教互动，属角色扮演聊天，无实验或测量目的，不涉及人类行为仿真…","model":"deepseek-v4-pro","scored_at":"2026-08-31T13:01:18","error":null,"has_summary":false,"summary":null},{"id":"2608.27927","version":1,"title":"Antipatterns in AI-assisted Qualitative Data Analysis: A Catalog of Temptations and Pitfalls for Software Engineering Researchers","zh_title":"AI辅助定性数据分析中的反模式：软件工程研究者的诱惑与陷阱目录","abstract":"AI-assisted qualitative data analysis (QDA) offers unprecedented opportunities to streamline software engineering (SE) research, yet uncritical use risks compromising analytical rigor and flooding the field with accelerated production of low-quality research. While tactical best practices will naturally evolve over time, SE researchers currently lack strategic guidance to identify and mitigate methodological risks when attempting AI-assisted QDA. Based on our decades of qualitative SE research expertise and experience combined with an understanding of the emerging landscape of AI-assisted QDA, this paper presents a catalog of antipatterns in AI-assisted QDA - a set of assumptions and practices that initially appear advantageous but ultimately undermine analytical rigor and validity. The antipatterns are grouped into three categories reflecting escalating impact: Dangerous Drivers, Operational Missteps, and Analytical Failures. As more SE researchers attempt AI-assisted QDA, these antipatterns will help them identify and avoid common temptations and pitfalls, while reviewers can be equipped with the vocabulary and criteria to call out problematic and failed practice. Ultimately, this catalog of antipatterns can serve as a stepping stone in our responsible methodological evolution toward principled and meaningful human-AI collaboration in qualitative research.","authors":["Rashina Hoda","Carolyn Seaman","Victoria Gomes","Rodrigo Spinola"],"categories":["cs.SE","cs.AI"],"primary_category":"cs.SE","announce_type":"cross","date":"2026-08-31","first_seen":"2026-08-31","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.27927","pdf_url":"https://arxiv.org/pdf/2608.27927","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["AI辅助分析","定性研究","方法论"],"reason":"论文讨论AI辅助定性数据分析的方法论陷阱，不涉及用LLM仿真人类被试或与人类数…","model":"deepseek-v4-pro","scored_at":"2026-08-31T13:01:14","error":null,"has_summary":false,"summary":null},{"id":"2608.28420","version":1,"title":"Between Algorithm (AI) and Intuition (Human): Preserving Designer Agency in AI-Assisted Sensemaking of Qualitative UX Data","zh_title":"算法（AI）与直觉（人类）之间：在AI辅助的定性用户体验数据意义建构中保持设计师能动性","abstract":"The integration of AI into qualitative design research presents a fundamental tension: how do we leverage AI while preserving the subjective, intuitive judgments that define design expertise? This paper examines this question through a case study of analyzing 20 user responses about video conferencing platforms for educational contexts. We argue that AI sensemaking tools risk flattening the rich data patterns, amplifying contradictory textures of user feedback into sterile categories thereby transforming design research from an interpretive craft into a mechanical sorting exercise (rigid and formal). Through comparative analysis of AI-assisted sensemaking versus human-centered approaches to the same dataset, we identify when algorithmic efficiency enhances understanding and when it diminishes the designer's interpretive agency (uncovering hidden needs, critical enquiry, what if enquiries, making decisions, having trade-offs). We present a framework for augmented sensemaking that positions AI as an instrument for amplifying human judgment rather than replacing it. Our findings suggest that the most valuable role for AI in design research is not to eliminate subjectivity, but to make it more intentional, reflective, and accountable.","authors":["Md Haseen Akhtar"],"categories":["cs.HC","cs.ET"],"primary_category":"cs.HC","announce_type":"new","date":"2026-08-31","first_seen":"2026-08-31","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.28420","pdf_url":"https://arxiv.org/pdf/2608.28420","source_feed":"cs.HC","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["AI辅助分析","设计研究","人机协作"],"reason":"研究AI辅助定性数据分析，非LLM仿真人类被试，无实验或测量目的","model":"deepseek-v4-pro","scored_at":"2026-08-31T13:01:20","error":null,"has_summary":false,"summary":null},{"id":"2608.25553","version":3,"title":"When Stale Constraints Go Unchecked: Budgeted Verification Failures in Inherited Agent Memory","zh_title":"当陈旧约束未被检查：继承智能体记忆中的预算验证失败","abstract":"Provenance links keep the evidence behind an inherited belief reachable; an agent with a verification budget must still choose which links to inspect. We study a consolidated memory that states a decision constraint and whose source record has since been superseded by a record that withdraws it: provenance is immutable, the current record has changed, and the memory is stale. In a controlled six-memory scenario with a budget of two records, sixteen language models rarely re-verified a constraint that read as settled: they inspected its provenance path in about one episode in five and, once the constraint had been superseded, produced stale-consistent decisions in 77.3%, 74.7% and 74.7% of episodes across a primary run, a replication and a held-out domain. Re-assigning one of the same two slots to the critical path removed most of them: +74.0, +72.7 and +61.3 points (positive in every model), +80.7 in a prospectively frozen interleaved replication with a repaired non-critical control, and +62.0 on a panel of 10 models from 9 organisations; a corrected re-run of the held-out scenario gave +73.3. The forced-critical policy uses experimenter knowledge of the critical path: it quantifies how much stale-decision risk the same budget can recover and is not a scheduler. Two further deposited experiments locate the failure and a remedy: in this store the constraint's path is selected in 17.0% of episodes at two slots and 88.7% at four of six (above uniform allocation), and at two slots a one-sentence, target-blind rule (prefer memories that state a limit on a candidate direction) moved the agent's own allocation onto the constraint's path and recovered the oracle contrast on decisions (+89.3 points) where that constraint limits the tempting action, while a content-free freshness cue did not materially redirect allocation and a content-matched control rule changed neither selection nor decisions.","authors":["Kazuki Nakayashiki"],"categories":["cs.IR","cs.AI","cs.CL"],"primary_category":"cs.IR","announce_type":"replace-cross","date":"2026-08-31","first_seen":"2026-08-27","revised_at":"2026-08-31","abs_url":"https://arxiv.org/abs/2608.25553","pdf_url":"https://arxiv.org/pdf/2608.25553","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","记忆验证","约束一致性"],"reason":"研究多智能体记忆验证，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-08-27T13:03:23","error":null,"has_summary":false,"summary":null},{"id":"2608.26159","version":2,"title":"Self-Generated Text Recognition: Quality Heuristics, Cross-Task Transfer, and Downstream Bias in LLM Evaluation","zh_title":"自生成文本识别：LLM评估中的质量启发式、跨任务迁移与下游偏差","abstract":"Self-Generated Text Recognition (SGTR)--the ability of an LLM to identify its own outputs--poses risks to AI safeguards that rely on LLMs as evaluators or monitors: an LLM may recognize outputs from other copies of the same model and make biased judgments or collude outright. Prior work has drawn conflicting conclusions about whether current models possess significant SGTR capabilities. We explain these disagreements by identifying key experimental design choices--which we term operationalizations--that drive divergent results. Evaluating 13-21 models across six presentation operationalizations and four task-domain operationalizations, we find that accuracy varies substantially with evaluation format (pairwise vs individual assessments of text), conversation format (presenting candidate text in user tags vs assistant tags), and the domain of the task used to generate candidate text (e.g., coding vs summarization). We corroborate previous observations that a quality heuristic--models attributing authorship to text they perceive as higher quality--is a dominant confound. We also find that improving a model's SGTR performance via supervised fine-tuning (SFT) on one operationalization can generalize to others, and can increase the model's preference for its own outputs when it acts as a judge in the AlpacaEval framework. Our results suggest that, despite confounds, some models possess practical SGTR capabilities, and that SGTR should be monitored and considered in the design of safety-critical AI applications.","authors":["Jesse St. Amand","Callum Canavan","Sohaib Imran","Joseph Hewson","Aaron Lutz","Shi Feng","Puria Radmard","Lennie Wells"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-08-31","first_seen":"2026-08-28","revised_at":"2026-08-31","abs_url":"https://arxiv.org/abs/2608.26159","pdf_url":"https://arxiv.org/pdf/2608.26159","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["LLM自识别","模型评估","安全风险"],"reason":"研究LLM识别自身生成文本的能力，属于模型能力评测，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-08-28T13:01:59","error":null,"has_summary":false,"summary":null},{"id":"2608.27638","version":1,"title":"Generative AI Expands the Intellectual Reach of Course Based Undergraduate Research Experiences (CUREs)","zh_title":"生成式人工智能拓展了基于课程的本科生研究经验（CUREs）的智力范围","abstract":"Course-based undergraduate research experiences (CUREs) broaden access to authentic scientific inquiry through responsive instructor support as research problems become increasingly complex. Generative artificial intelligence (GenAI) may extend this support by providing individualized assistance that can adapt as student needs change. However, how embedding GenAI within a CURE to provide support across the research process impacts student inquiry, collaboration, and scientific reasoning remains unresolved. Here we use longitudinal qualitative data collected across three semesters of a bioinformatics and genomics CURE to show that GenAI expanded the intellectual reach of the research experience in three distinct ways. First, personalized, on-demand scaffolding allowed students to move beyond the boundaries of instructor expertise and transform their own interests into researchable inquiry, with all teams developing distinct self-directed projects rather than selecting instructor-provided topics. Second, GenAI became part of the distributed cognitive system of research teams, helping novice researchers communicate and coordinate across differentiated expertise without eliminating specialization. Third, expanded capability did not replace the need for disciplinary judgment. Students increasingly validated, revised, or rejected AI-generated contributions, such that research independence emerged through retained intellectual responsibility. Together, these findings suggest that GenAI can extend the reach of CUREs by expanding what novice researchers can investigate, how they can collaborate, and the level of responsibility they can assume while preserving human judgment central to authentic scientific inquiry.","authors":["Aditi Babar","Kristin J. Davin","Alex Dornburg"],"categories":["cs.AI","cs.HC"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-31","first_seen":"2026-08-31","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.27638","pdf_url":"https://arxiv.org/pdf/2608.27638","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["生成式AI","教育技术","本科生研究"],"reason":"研究GenAI在本科研究教学中的应用，不涉及用LLM仿真人类被试或与真实人类数…","model":"deepseek-v4-pro","scored_at":"2026-08-31T13:01:12","error":null,"has_summary":false,"summary":null},{"id":"2608.27953","version":1,"title":"The Illusion of $\\textit{What If}$: Evaluating the Breakdown of Counterfactual Reasoning in LLMs","zh_title":"《如果》的幻觉：评估大语言模型中反事实推理的崩溃","abstract":"Counterfactual reasoning requires models to reason beyond the observed world and explain how altered conditions propagate through downstream consequences. Existing benchmarks largely target bounded settings with fixed variables or single gold outcomes, overlooking open-domain scenarios requiring causal-process evaluation. To this end, we present $\\textbf{WhatIfBench}$, a diagnostic benchmark for open-domain, open-form, long-horizon counterfactual causal reasoning, containing 220 what-if questions across STEM, HSS, and Hybrid scenarios. To evaluate free-form responses, we further propose $\\textbf{PRISM}$, which first converts each natural-language explanation into a Response-Derived Semantic Causal Graph of events, states, and mechanisms. On top of this graph, PRISM then jointly applies a Process Metric assessing graph-level causal validity and a Rubric Metric assessing answer-level explanatory adequacy. Evaluating six frontier LLMs with this framework, we find that WhatIfBench remains far from saturated: even the strongest model reaches only a 64.62% final score. Further analysis reveals persistent causal gaps, premise drift, and topology fragmentation, suggesting that fluent counterfactual narratives often mask fragile causal processes. The benchmark, code, and evaluation scripts are available at $\\href{https://github.com/zju-gt/WhatIfBench}{WhatIfBench}$.","authors":["Yucheng Wang","Yuetian Du","Zhengyi Liu","Rongyu Zhang","Bing Zhao","Boyu Yang","Ming Kong","Lin Qu","Hu Wei","Jie Liu","Qiang Zhu"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-31","first_seen":"2026-08-31","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.27953","pdf_url":"https://arxiv.org/pdf/2608.27953","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["反事实推理","基准测试","因果推理"],"reason":"纯LLM能力评测，无人类行为对照，不涉及仿真人类被试","model":"deepseek-v4-pro","scored_at":"2026-08-31T13:01:14","error":null,"has_summary":false,"summary":null},{"id":"2608.27629","version":1,"title":"LitCurate: A Configuration-Driven AI-Assisted Framework for Scientific Database Construction with an Application to Lower-Mantle Equation-of-State Data","zh_title":"LitCurate：配置驱动的AI辅助科学数据库构建框架及其在下地幔状态方程数据中的应用","abstract":"The growing scientific literature contains decades of experimental and computational results that could support data-driven and physics-based modeling, yet much of this infor- mation remains locked in publications and is not readily usable for large-scale analysis or sci- entific software. Building structured databases from the literature is particularly challenging whenrelevantstudiesmustfirstbediscoveredamonglargecollectionsofpapersandreported quantities must be extracted with enough scientific context to remain usable. We present LitCurate, an open-source framework for building scientific databases from the literature using large language models within an auditable, stage-wise curation workflow. LitCurate integratesliteraturediscovery, relevancescreening, full-textprocessing, andstructuredinfor- mation extraction while retaining intermediate results and provenance, allowing researchers to inspect and revise individual stages rather than treating automated curation as a black- box process. We apply LitCurate to construct an equation-of-state database of lower-mantle and lower-mantle-relevant high-pressure mineral phases from experimental and theoretical studies, comprising 1,334 entries from 205 papers. The resulting dataset links reported equation-of-state parameters to mineral phases, compositions, equation formulations, meth- ods, and parameter constraints, and labels values as source-reported or citation-reported when provenance can be determined. The records are available through a searchable web application. By connecting scientific literature to traceable, machine-readable data, LitCu- rate provides a reusable approach for transforming accumulated literature into resources for scientific analysis and computational modeling.","authors":["Abin Shakya","Wilson Samuels","Dominica Wilson","Gioia A. Marchi","Israa Draz","Chenxing Luo","Renata M. Wentzcovitch"],"categories":["cs.IR","cs.AI","physics.geo-ph"],"primary_category":"cs.IR","announce_type":"cross","date":"2026-08-31","first_seen":"2026-08-31","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.27629","pdf_url":"https://arxiv.org/pdf/2608.27629","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["文献挖掘","数据库构建","LLM应用"],"reason":"该论文是文献挖掘与数据库构建，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-08-31T13:01:12","error":null,"has_summary":false,"summary":null},{"id":"2608.27932","version":1,"title":"Graphionale: How Graph Visualizations of LLM Rationales Affect Human Decision Making","zh_title":"Graphionale：LLM推理的图形可视化如何影响人类决策","abstract":"Large Language Models (LLMs) are increasingly equipped with augmented reasoning capabilities to generate rationales that support human decision-making. Yet these text-dense rationales often impose substantial cognitive burdens. Building on a formative co-design study that identified user preferences for non-linear reasoning representations, we developed Graphionale as a testbed for empirically studying argument-map-style rationale visualization. This system transforms linear LLM rationales into interactive, multi-level graphs. It explicitly structures logical relationships (e.g., conclusions, premises, support, and objections), while further extracting entities and relations within each statement to construct condensed node-link representations. We conduct a large-scale online user study (N = 204) to examine when graphical rationales are more effective than textual ones, across varying task modality (verbal vs. visual reasoning), rationale format (textual vs. graphical), and question difficulty (easy vs. hard). Our results show that graphical rationales do not help uniformly: they improve trust calibration for verbal reasoning yet feel more cognitively demanding and less satisfying; for visual reasoning, they impair calibration yet feel more engaging and helpful. In each modality, the format that better supports calibrated decisions is not the one users prefer, highlighting that matching rationale format to task modality is key to effective AI explanation design. Our findings contribute empirical design knowledge about when and how graphical rationales support human decision making, and inform the next-generation reasoning-aware AI interfaces.","authors":["Xinru Wang","Zhexuan Ma","Ming Yin","Shuai Ma","Thomas W Malone"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-08-31","first_seen":"2026-08-31","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.27932","pdf_url":"https://arxiv.org/pdf/2608.27932","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["人机交互","可解释AI","决策支持"],"reason":"研究人类如何理解LLM解释，非用LLM仿真人类被试","model":"deepseek-v4-pro","scored_at":"2026-08-31T13:01:08","error":null,"has_summary":false,"summary":null},{"id":"2608.28050","version":1,"title":"Too Much of the Same: From Algorithmic to Human Bias in Learning to Defer","zh_title":"过多相同：从算法偏差到人类偏差在学习推迟中","abstract":"Learning to Defer (LtD) extends supervised learning by allowing a Machine Learning (ML) model to defer harder or less confident decisions to a human expert. Despite being geared for human-AI collaboration, LtD strategies neglect the potential negative interference of human cognitive biases. Our contribution is twofold. First, we demonstrate that standard LtD strategies show class-dependent sampling bias in classification tasks in practice, and thus may disproportionately defer the minority classes when applied to imbalanced datasets. Second, we show that such asymmetries in task delegation may trigger human biases, ultimately leading to poorer downstream decision making. Specifically, we conduct a user study ($N=226$) where participants complete a classification task on a set of deferred items, with conditions presenting different levels of class imbalance. Our results show that participants exposed to a highly imbalanced rejection set achieved lower classification accuracy in the majority class compared to those exposed to a more balanced set, regardless of which class constituted the majority. Exploratory analyses suggest that this may be an instance of the Test-taker's effect, which stems from a mismatch between the actual distribution of classes and the participants' expectations about that distribution. Finally, we discuss the implications of these findings for the deployment of LtD algorithms.","authors":["Dario Pesenti","Alessandro Bogani","Stefano Teso","Andrea Pugnana"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-08-31","first_seen":"2026-08-31","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.28050","pdf_url":"https://arxiv.org/pdf/2608.28050","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["人机协作","学习推迟","人类偏差"],"reason":"研究人类在LtD中的偏差，不涉及LLM仿真人类被试","model":"deepseek-v4-pro","scored_at":"2026-08-31T13:01:16","error":null,"has_summary":false,"summary":null},{"id":"2608.28501","version":1,"title":"A Guided Inquiry Approach to Students Co-Designing Generative AI Course Policies","zh_title":"学生共同设计生成式AI课程政策的引导式探究方法","abstract":"As generative AI (GenAI) use among students increases, educators face growing questions about how to support learning while addressing ethical and institutional concerns. This exploratory study examines a guided inquiry activity in which students co-designed a GenAI course policy. Students first developed individual policy proposals focused on appropriate and ethical use of GenAI, then collaboratively refined them by incorporating diverse stakeholder perspectives. The following research questions guided the study: 1) what practical factors do students prioritize in their GenAI use policies, and how do they justify these choices? and 2) how do participants reflect on the policy design process? Participants first completed readings, then used GenAI to brainstorm initial policy ideas. Next, they articulated their own perspectives through a written assignment and a course policy they designed individually. Finally, they incorporated diverse stakeholder perspectives by collaborating with peers to develop a collective policy. Analysis of student artifacts and group discussions showed that participants prioritized training for students and instructors, standardized procedures for disclosing AI use, and stronger institutional support. Participants also wanted greater involvement in GenAI-related decision-making. They described the policy design process as a way to engage with multiple perspectives and the inherent trade-offs involved in governing AI use. This study offers pedagogical insights into how policy co-design activities can surface student values, concerns, and sensemaking about GenAI in educational contexts.","authors":["Ashish Hingle","Aditya Johri"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2026-08-31","first_seen":"2026-08-31","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.28501","pdf_url":"https://arxiv.org/pdf/2608.28501","source_feed":"cs.CY","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["教育政策","生成式AI","学生参与"],"reason":"研究学生共同设计AI课程政策，不涉及用LLM仿真人类被试或行为对照","model":"deepseek-v4-pro","scored_at":"2026-08-31T13:01:20","error":null,"has_summary":false,"summary":null},{"id":"2608.27538","version":1,"title":"Disaffection at Work: Employee Responses to Job-Related Information","zh_title":"工作不满：员工对工作相关信息的反应","abstract":"Quiet quitting reflects a form of worker disaffection that operates along the intensive margin of labor supply rather than through job exit. We study how alternative workplace narratives shape workers' behavioral responses using a randomized survey experiment on a representative sample of employees in Italy and France. Respondents are exposed to empirically grounded moral framings of work emphasizing either social justice and collective rights or work organization and employment practices. We find that moral framings reallocate behavior across margins: justice-oriented narratives increase detachment while reducing passive disengagement, whereas organization-centered framing generates no systematic effects.","authors":["Beatrice Braut","Mariele Macaluso","Vincenzo Mollisi"],"categories":["econ.GN","q-fin.EC"],"primary_category":"econ.GN","announce_type":"new","date":"2026-08-31","first_seen":"2026-08-31","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.27538","pdf_url":"https://arxiv.org/pdf/2608.27538","source_feed":"econ.GN","score":0,"bucket":"other","rubric_hits":["C5"],"tags":["劳动经济学","随机调查实验","员工行为"],"reason":"论文使用人类被试做随机调查实验，未涉及LLM仿真或替代人类被试，方向相反。","model":"deepseek-v4-pro","scored_at":"2026-08-31T13:01:08","error":null,"has_summary":false,"summary":null},{"id":"2608.26291","version":1,"title":"Assessing mentalization in humans and large language models","zh_title":"评估人类与大语言模型的心理化能力","abstract":"Mentalization - the ability to infer others' beliefs and intentions to guide one's own choices - is a key cognitive function underlying human social interactions. Large language models (LLMs) demonstrate behaviour consistent with humans on theory-of-mind tasks, yet whether these models can guide adaptive behaviour through mentalization is unknown. Here we use two economic games with cognitive computational modeling to uncover the latent strategies underlying mentalization in LLMs. We tested individual LLM agents across four model families, DeepSeek, GPT-4.1, GPT-5 and Gemini 2.0 Flash (N = 2,099), against opponents of varying sophistication and examined whether a prompting strategy designed to elicit strategic reasoning improved performance. We benchmarked results against human participants (N = 251) as a comparative measure. Across both games, LLMs showed clear behavioural and computational signatures of mentalizing that differed markedly by model provider and size. Strategic prompting generally improved performance by inducing more sophisticated reasoning, yet the extent of the benefit differed across the two tasks. Last, GPT-5 agents flexibly adapted their recursive depth of reasoning to increasingly sophisticated opponents, demonstrating superior performance to human participants. Collectively, we demonstrate different capacities for mentalization across LLMs, and highlight cognitive computational modeling as a formal method for assessing comparative intelligence across humans and machines.","authors":["Aamir Sohail","Xintong Zhong","Arkady Konovalov","Patricia L. Lockwood","Lei Zhang"],"categories":["cs.AI","q-bio.NC"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-28","first_seen":"2026-08-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.26291","pdf_url":"https://arxiv.org/pdf/2608.26291","source_feed":"cs.AI","score":10,"bucket":"selected","rubric_hits":["A1","A2","A3","B1","B2","B3"],"tags":["LLM仿真","经济学实验","认知建模"],"reason":"用LLM作为人类被试替代品，在经济学博弈中与人类数据对照，评估心理化能力与策略。","model":"deepseek-v4-pro","scored_at":"2026-08-28T13:01:52","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-28","rank":1,"question":"LLM能否像人类一样通过心理化（推断他人信念与意图）来指导策略行为，其潜在策略是什么？","design":"用四个模型家族（DeepSeek、GPT-4.1、GPT-5、Gemini 2.0 Flash）的LLM个体（N=2099）作为被试，参与两个经济博弈（检查博弈和石头剪刀布），部分模型施加社会思维链（SCoT）提示，测量博弈得分和策略选择，并用认知计算建模推断递归推理深度。","baseline":"人类被试（检查博弈 n=67，石头剪刀布 n=184）在相同任务中的表现。","findings":"LLM表现出明显的心理化行为与计算特征，但不同模型和规模差异显著；SCoT提示普遍提升表现，但提升幅度因模型而异。GPT-5能灵活适应对手复杂度，表现优于人类。","reliability":"论文未讨论","relevance":"该研究直接以LLM作为人类被试替代品，在经济学博弈中与人类数据对照，评估心理化能力与策略，完全命中你的关注点，值得精读原文。","inspiration":"借鉴其用认知计算建模从行为数据中提取潜在策略（递归推理深度）的方法，以及用SCoT提示作为处理变量来诱导更复杂推理的设计。｜可迁移到资产定价实验中的策略性预期形成或谈判博弈中的信念更新。｜用LLM作为投资者被试，施加SCoT提示，在重复博弈中测量报价或投资决策，并与人类实验数据（如资产泡沫实验）对照，比较递归推理深度和适应性。"}},{"id":"2608.26327","version":1,"title":"How Unlikely Is \"Unlikely\"? Assessing Verbal Probability Perception Across Large Language Models","zh_title":"“不太可能”有多不可能？跨大语言模型评估言语概率感知","abstract":"Large language models increasingly produce and interpret verbal probability expressions, yet whether these expressions carry consistent meaning across models (or match human perceptions of uncertainty) remains unknown. We present a systematic cross-model evaluation using a word-to-number mapping task grounded in established human benchmarks. Eleven uncertainty expressions were presented to 19 models under two conditions, forced single-number response and explanation elicitation, alongside a novel bidirectional roundtrip test of internal consistency. LLMs track the human benchmark with surprising fidelity: word ordering is preserved, three anchor points are recovered, and ``possible'' shows the highest variance and cross-model disagreement of any expression tested, consistent with its documented bimodal interpretation in humans. However, models show a systematic upward bias for negative expressions such as ``unlikely'' and ``improbable.'' Explanation elicitation reduces within-model variance while increasing between-model divergence, stabilizing individual models at the cost of inter-model consensus, and the roundtrip experiment reveals clear stratification, with frontier models maintaining coherent bidirectional representations. LLMs thus reproduce the structure of human verbal probability cognition, including its biases, while diverging systematically at the negative end---with implications for any setting where humans and models exchange probabilistic language.","authors":["Christos Petridis","Konstantinos Pelechrinis","Zoran Obradovic"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-28","first_seen":"2026-08-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.26327","pdf_url":"https://arxiv.org/pdf/2608.26327","source_feed":"cs.CL","score":8,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","概率语言","人类基准对照"],"reason":"用LLM复现人类对概率词的理解，并与人类基准对照，发现偏差，属于仿真人类认知且…","model":"deepseek-v4-pro","scored_at":"2026-08-28T13:01:52","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-28","rank":5,"question":"LLM对言语概率表达（如“unlikely”）的数值解读是否与人类基准一致，以及在不同提示条件下是否稳定？","design":"用19个LLM作为被试，呈现11个概率词，在两种提示条件下（强制单数字回答、要求解释）进行词到数字映射，并进行双向往返一致性测试，测量映射数值、方差和跨模型一致性。","baseline":"Mosteller和Youtz（1990）汇总的20项人类研究中概率词的数值解读基准。","findings":"LLM整体上复现了人类基准：词序保持、三个锚点（impossible, even chance, certain）恢复良好，“possible”方差最大，与人类双峰解释一致。但LLM对负面词（如unlikely）存在系统性高估，解释提示降低模型内方差但增加模型间分歧，往返测试显示前沿模型双向映射更一致。","reliability":"论文指出偏差集中在非锚点词，负面词在强制条件下偏差略大；解释提示虽稳定个体但牺牲跨模型共识；部分模型往返映射接近随机，表明内部表征不一致。","relevance":"该研究直接评估LLM作为人类被试在概率词理解上的仿真效度，发现总体复现但存在系统性偏差，对关注LLM仿真人类认知可靠性的研究者有重要参考价值。","inspiration":"借鉴其词到数字映射任务和双向一致性检验，可迁移到经济金融中的风险沟通与预期形成场景，例如央行政策声明中的模糊措辞解读。｜设计一个实验：以LLM为被试，呈现央行声明中的概率词（如“可能加息”），要求给出数值概率，并与专业预测者调查或市场隐含概率对照，检验LLM是否复现人类解读偏差。"}},{"id":"2608.26188","version":1,"title":"Is Your Neighborhood Safe? Place-based Stigma in Large Language Models' Urban Safety Judgments","zh_title":"你的社区安全吗？大语言模型城市安全判断中的地方污名","abstract":"Large language models are increasingly used to inform safety decisions in cities, such as where it is safe to walk, rent, or travel. We ask whether such judgments track measured risk or the patterns attached to an urban neighborhood's name. We probe seven instruct-tuned models under three conditions that dissociate name from geography: coordinates-only, name-only, and name+coordinates, across 186 neighborhoods in Los Angeles and Chicago, joined to violent crime and American Community Survey data. First, ratings are nearly flat under coordinates for six of seven models, while names carry most between neighborhood variation and are moderately calibrated to violent crime; only at frontier scale does the coordinate channel show appreciable variation. Second, names lower safety ratings more for neighborhoods with higher shares of the locally dominant marginalized group (percent Black in Chicago, percent Hispanic in Los Angeles), and this name effect tracks demographic share in all seven models and both cities. In Los Angeles, where demographic share and crime are more separable, the effect survives controls for crime and income and is confirmed by crime-matched pairs. An enforcement-elasticity analysis further shows that over-caution tracks near-fully-reported homicide rather than discretionary, deployment-driven offenses. Third, the effect scales with geographic knowledge: models that better distinguish real neighborhoods apply more demographic stereotype to them. Because neighborhood names carry both genuine crime signal and demographic stereotype, removing names reduces both bias and accuracy. We discuss implications for deploying LLMs in advice and decision-support settings.","authors":["Huy Nguyen","Yue Lin"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-28","first_seen":"2026-08-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.26188","pdf_url":"https://arxiv.org/pdf/2608.26188","source_feed":"cs.AI","score":8,"bucket":"selected","rubric_hits":["A1","B1","B4"],"tags":["LLM仿真","社会偏见","城市安全"],"reason":"用LLM模拟人类对社区安全的判断，并与真实犯罪和人口数据对照，揭示偏差。","model":"deepseek-v4-pro","scored_at":"2026-08-28T13:01:50","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-28","rank":3,"question":"LLM 对城市社区夜间步行安全的判断，是追踪真实犯罪风险，还是反映与社区名称相连的种族-空间污名？","design":"用七种指令微调 LLM 模拟居民对社区安全的感知，对芝加哥和洛杉矶共 186 个社区，在仅坐标、仅名称、名称+坐标三种条件下生成夜间步行安全评分，并与真实暴力犯罪和人口普查数据对照。","baseline":"真实暴力犯罪记录（芝加哥 2020-2025、洛杉矶 2020-2024）和美国社区调查（ACS）人口统计数据。","findings":"名称承载了几乎所有社区间安全评分差异，且与暴力犯罪中度校准；名称对安全评分的压低效应随当地主导边缘化群体（芝加哥黑人、洛杉矶西裔）比例上升，在洛杉矶该效应在控制犯罪和收入后仍存在，且与犯罪匹配对验证一致。","reliability":"坐标通道在开源模型中几乎无变异，仅在 frontier 模型上显现；芝加哥因黑人比例与犯罪高度相关无法分离种族与犯罪效应；名称同时携带真实犯罪信号和人口统计刻板印象，隐藏名称会同时消除偏差和准确性。","relevance":"该研究用 LLM 模拟人类对社区安全的判断，并与真实犯罪和人口数据对照，直接揭示仿真中的种族偏差，对关注 LLM 仿真可靠性及偏差的研究者极具参考价值。","inspiration":"值得借鉴的是通过条件消融（仅坐标、仅名称、名称+坐标）分离名称效应，并用真实犯罪和人口数据做基准，以及用犯罪匹配对和执法弹性分解做稳健性检验。｜可迁移到信贷审批中的地域歧视研究，如 LLM 模拟信贷员对申请人的风险评估是否受申请人所在社区名称的种族构成影响。｜用 LLM 扮演信贷审批员，对虚构申请人给出贷款批准概率，处理变量为申请人地址的社区名称（高黑人/西裔比例 vs 低比例），结果变量为批准概率，对照真实数据用社区层面的实际贷款批准率和违约率，并控制申请人收入、信用分等特征。"}},{"id":"2608.26221","version":1,"title":"Prompt Sensitivity of Generative Agents: Evidence from an Epidemic Model","zh_title":"生成式智能体的提示敏感性：来自流行病模型的证据","abstract":"As generative AI gains traction, researchers are investigating its potential to serve as proxies for humans. From undergoing cognitive psychology experiments to experiencing an epidemic, generative agents, agents powered by generative AI models, produce realistic human behavior when prompted. This study explores the sensitivity of these generative agents' behavior to prompt modifications and varied persona names of the agents. To assess this sensitivity, we use a generative agent epidemic model, wherein each agent is prompted daily on whether it wants to isolate or commingle with other agents. We found that using synonymous prompts results in negligible changes to the model's outcomes. However, minor variations in prompts, as well as contextual changes, do influence the model's results. Lastly, our data indicates that different persona names assigned to generative agents, specifically those imbued with personas, do not significantly impact epidemic outcomes.","authors":["Ross Williams","Niyousha Hosseinichimeh"],"categories":["physics.soc-ph","cs.AI","cs.LG","cs.MA"],"primary_category":"physics.soc-ph","announce_type":"cross","date":"2026-08-28","first_seen":"2026-08-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.26221","pdf_url":"https://arxiv.org/pdf/2608.26221","source_feed":"cs.AI","score":8,"bucket":"selected","rubric_hits":["A1","A3","B4"],"tags":["LLM仿真","流行病模型","提示敏感性"],"reason":"用生成式智能体模拟疫情中的人类隔离决策，研究提示敏感性，属于人类行为仿真，但无…","model":"deepseek-v4-pro","scored_at":"2026-08-28T13:01:50","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-28","rank":4,"question":"生成式智能体在疫情模型中的行为对提示词修改和角色名称变化的敏感程度如何？","design":"使用生成式智能体疫情模型，每个智能体每天被提示选择隔离或与他人接触，通过改变提示词的语义、上下文和角色名称来测试行为变化，结果变量为疫情传播的流动性曲线。","baseline":"无对照","findings":"同义提示词修改对模型结果影响可忽略，但轻微提示词变化和上下文变化会影响模型结果。不同角色名称对疫情结果无显著影响。","reliability":"论文未讨论","relevance":"该研究直接探讨LLM仿真中提示敏感性这一可靠性问题，但未与真实人类数据对照，且场景为疫情模型而非经济学实验，对关注经济学和政策评估的研究者参考价值有限。","inspiration":"可借鉴其系统改变提示词并测量输出变化的方法，用于评估LLM在经济实验中的稳健性。｜可迁移到政策公告的预期形成实验，测试不同措辞的公告对LLM代理预期的影响。｜用LLM模拟投资者，随机分配不同措辞的央行声明，测量其通胀预期变化，并与专业预测者调查数据对照。"}},{"id":"2608.23705","version":2,"title":"The Limits of Automatic Evaluation of Creativity in Large Language Models","zh_title":"大语言模型创造力自动评估的局限性","abstract":"Large Language Models (LLMs) are increasingly capable of generating text that challenges human performance in domains requiring creativity, yet evaluating creativity in LLM-generated content remains a significant challenge. Here, we investigate whether current automatic evaluation methods can reliably capture human judgments of creativity. We collect human evaluations of human- and AI-generated short stories from the WritingPrompts dataset across 11 dimensions of creativity, and compare these judgments with automated objective metrics and LLM-as-a-Judge evaluations. Our experiments reveal substantial misalignment between automatic evaluations and human assessments. In particular, LLM-based judges exhibit a systematic preference for AI-generated stories, consistently favoring their stylistic characteristics over the unpredictability and other qualities of human-authored texts. Furthermore, correlation analyses show that widely used automatic metrics exhibit near-zero alignment with human judgments across both human- and AI-generated stories, suggesting that they fail to capture important dimensions of creativity. These findings highlight fundamental limitations in current approaches to the automatic evaluation of creative text and underscore the difficulty of reducing the multidimensional and subjective nature of creativity to computational metrics.","authors":["Alessandro Tutone","Giorgio Franceschelli","Mirco Musolesi"],"categories":["cs.CL","cs.AI","cs.CY"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-08-28","first_seen":"2026-08-26","revised_at":"2026-08-28","abs_url":"https://arxiv.org/abs/2608.23705","pdf_url":"https://arxiv.org/pdf/2608.23705","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A2","B1","B4"],"tags":["LLM评估偏差","人类对照","创造力测量"],"reason":"评估LLM作为评判者的可靠性，与人类判断对照，揭示偏差，可迁移到仿真效度研究。","model":"deepseek-v4-pro","scored_at":"2026-08-28T13:02:14","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-28","rank":6,"question":"当前自动评估方法能否可靠捕捉人类对创造力的判断？","design":"本研究并非仿真研究，而是评估自动评估方法：收集人类对100篇人类创作和100篇LLM生成短故事在11个创造力维度上的评分，并与自动客观指标和LLM-as-a-Judge评分进行比较。","baseline":"人类评估数据：来自WritingPrompts数据集的100篇人类创作故事和100篇LLM生成故事，由人类评分者在11个创造力维度上进行评价。","findings":"自动评估与人类判断存在显著不一致，LLM评委系统性偏好AI生成故事，偏向风格特征而非不可预测性等人类文本特质。广泛使用的自动指标与人类判断相关性接近零，未能捕捉创造力的重要维度。","reliability":"论文指出LLM评委存在自我偏好偏差，且自动指标与人类判断相关性极低，表明当前自动评估方法在创造力评估上不可靠，需要替代方法。","relevance":"该研究直接评估LLM作为评判者的可靠性，并与人类判断对照，揭示系统性偏差，对仿真研究中用LLM替代人类评估者的效度有重要警示。","inspiration":"借鉴其多维度人类评分与自动评分对照的设计，可迁移到经济金融中涉及主观判断的仿真评估，如信贷审批中的公平性判断或投资决策中的风险评估。可设计实验：用LLM扮演信贷员对贷款申请进行审批，处理为申请人特征（如种族、性别），结果变量为审批决定和理由，与真实信贷审批数据或人类专家判断对照，检验LLM是否存在类似偏差。"}},{"id":"2608.23780","version":2,"title":"When Youth Enter The Chat: An Epistemic Shift in the Validation of LLM-Based Measures of Student Talk","zh_title":"当青少年进入聊天：基于LLM的学生话语测量验证的认识论转变","abstract":"LLMs are being used increasingly to measure aspects of student discourse (e.g. talk moves, collaboration, equity of voice) at scale. Typically, LLM-based measures of student talk use transcriptions of classroom conversations that only include verbal contributions, which de-contextualize student language. Common practices for validating these measures include comparing outputs against expert annotations by adults, using held out evaluation sets and F1 scores. We argue that these approaches are insufficient to ensure that such measures are meaningful and equitable for teaching and learning, particularly for racially and linguistically marginalized youth. In order to center the youth whose talk is being analyzed, re-contextualizing these classroom conversations and engaging youth in the research process is necessary. Sharing epistemic authority with youth, ultimately, centers their point of view and adds crucial nuance to the analysis of their talk that adult experts, researchers, and LLMs cannot provide. In a case study of multilingual youth in one 8th-grade math classroom, we address the epistemic exclusion of youth by employing multiple ethnographically-oriented methods to re-contextualize student conversations and center youth as epistemic authorities in conversation with researchers and LLMs. We conducted participant observations, interviews, focus groups, and member checks with four focal students. Findings reveal that there were misalignments between students' interpretations of their own math talk experiences and the LLM-based measures of their talk. Students contested both the LLM classifications and the coding scheme used to measure their talk, highlighting the need for youth to be involved in the epistemic process of producing knowledge about their experiences.","authors":["Liliana Santos-Deonizio","James Malamut","Ram\\'on Antonio Mart\\'inez","Dorottya Demszky"],"categories":["cs.CL","cs.AI","cs.HC"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-08-28","first_seen":"2026-08-26","revised_at":"2026-08-28","abs_url":"https://arxiv.org/abs/2608.23780","pdf_url":"https://arxiv.org/pdf/2608.23780","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A2","B4"],"tags":["LLM测量效度","批判性评估","教育话语分析"],"reason":"评估LLM测量学生话语的效度，指出与青少年自身解读的偏差，批判性视角可迁移至人…","model":"deepseek-v4-pro","scored_at":"2026-08-28T13:02:14","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-28","rank":7,"question":"在验证基于LLM的学生话语测量时，让青少年参与知识生产过程能揭示什么？","design":"本研究不是仿真实验，而是质性案例研究。作者对一所八年级数学课堂中的四名多语种学生进行参与式观察、访谈、焦点小组和成员核查，将学生的自我解读与LLM对其话语的分类进行对比。","baseline":"无对照","findings":"学生的自我解读与LLM分类之间存在错位，学生质疑LLM的分类和编码方案。仅依赖成人专家标注和文本转录的验证方法不足以捕捉学生话语的情境意义，需要让青少年作为认识权威参与验证过程。","reliability":"论文指出，LLM基于去情境化的文本转录进行测量，忽略了关系、课堂常规、物理空间、手势、韵律和语言意识形态等维度，导致对边缘化学生话语的误读。","relevance":"该研究批判了LLM测量人类行为时仅依赖专家标注和文本数据的验证方式，强调纳入被测量者自身视角的重要性，对评估LLM仿真人类被试的可靠性与偏差具有方法论启示。","inspiration":"借鉴其成员核查和情境重构方法，将人类被试的自我报告与LLM输出进行系统对比，以揭示仿真偏差。｜可迁移到信贷审批歧视或消费者金融行为研究，检验LLM对借款人自述或消费决策的分类是否与借款人自身理解一致。｜以真实信贷申请文本为输入，让LLM预测违约风险或分类借款用途，同时收集借款人对自身文本的解读作为对照，比较LLM输出与借款人自我报告的一致性，并分析偏差来源。"}},{"id":"2608.26899","version":1,"title":"Counterfactual Bias Testing for Application Tracking System","zh_title":"申请追踪系统的反事实偏见测试","abstract":"Automated candidate-job matching systems are increasingly classified as high-risk AI under emerging regulation, yet auditing them for demographic bias is expensive: classical correspondence-audit studies require hand-crafted resumes and manual submission, which does not scale to fast pipeline retraining cycles. This paper presents a general, reusable methodology that (1) uses task-specialized LLM agents to synthesize identity-neutral base resumes and inject controlled demographic treatments across five protected-characteristic axes (sex/gender, age, residence, language, disability), producing a K x (1+N) correspondence-audit matrix; (2) qualitatively flags inferred protected characteristics per an EU AI Act-aligned prompt; (3) ranks candidates against a job description via a fine-tuned sentence-embedding model and cosine similarity; and (4) computes a nine-metric fairness suite spanning counterfactual (score delta, mean absolute rank change, flip rate), group-fairness (top-K retention, four-fifths/impact ratio), and merit-aware (Recall@K, nDCG@K, equal opportunity, equalized odds) families, each with bootstrap confidence intervals, significance tests, and Benjamini-Hochberg correction, culminating in an automated PASS/INVESTIGATE/FAIL report with a composite risk score. On an example corpus of 5 job orders, 100 base candidates, and 10 demographic treatments (90 metric x variant evaluations): score shifts, top-K retention, and merit-aware rate gaps stay within tolerance for every treatment, but a rank-stability metric (MARC) and nDCG@K each surface borderline findings - including one on the neutral baseline itself - that a score- or retention-only view would miss. The results argue for multi-metric, multi-family auditing over any single aggregate score, and for LLM-agent-generated audits as a practical, low-cost complement to human-curated audits for any candidate-job matching pipeline.","authors":["Sai Yashwant","Shruti Bansal","Anurag Dubey","Samaroha Chatterjee","Satyam Kumar","Shreyash Gupta","Gantala Thulsiram"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-28","first_seen":"2026-08-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.26899","pdf_url":"https://arxiv.org/pdf/2608.26899","source_feed":"cs.AI","score":7,"bucket":"pending","rubric_hits":["A1","A2","B1","B2","B3"],"tags":["LLM仿真","算法审计","公平性评估"],"reason":"用LLM生成简历模拟人类求职者，审计招聘系统偏见，有真实数据对照，属仿真人类被…","model":"deepseek-v4-pro","scored_at":"2026-08-28T13:01:54","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-28","rank":9,"question":"如何用LLM智能体自动生成简历变体，对候选人-职位匹配系统进行反事实偏见审计，并构建多指标公平性评估框架？","design":"使用任务专用LLM智能体链生成身份中立的基准简历，并在五个受保护特征轴（性别、年龄、居住地、语言、残疾信号）上注入受控人口学处理，产生K×(1+N)的通信审计矩阵；然后通过微调的句子嵌入模型和余弦相似度对候选人进行排序，并计算九项公平性指标（反事实、群体公平、绩效感知三类），每个指标配有自助置信区间、显著性检验和Benjamini-Hochberg校正，最终生成PASS/INVESTIGATE/FAIL报告。","baseline":"无对照","findings":"在示例语料库（5个职位、100个基准候选人、10个人口学处理）上，所有处理的分数偏移、top-K保留率和绩效感知率差均在容差范围内，但排名稳定性指标（MARC）和nDCG@K出现边缘性发现，包括中性基线本身，表明仅看分数或保留率会遗漏问题。","reliability":"论文未讨论","relevance":"该研究利用LLM生成简历模拟人类求职者，审计招聘系统偏见，属于仿真人类被试的范畴，但缺乏真实人类数据对照，且场景为招聘而非经济学实验，对关注经济学实验和政策评估的研究者参考价值有限。","inspiration":"可借鉴其用LLM生成反事实变体并构建多指标审计框架的方法，用于生成受控处理组和对照组。｜可迁移到信贷审批歧视审计，用LLM生成贷款申请人资料，注入性别、种族等特征，评估信贷模型偏见。｜以LLM生成的贷款申请人为被试，处理为注入受保护特征（如性别、种族），结果变量为贷款审批分数或决策，对照真实信贷审批数据（如Home Mortgage Disclosure Act数据）验证仿真可靠性。"}},{"id":"2608.27219","version":1,"title":"BALMS: Benchmarking Agentic LLMs for Longitudinal Mental Health Sensing","zh_title":"BALMS：面向纵向心理健康感知的智能体LLM基准测试","abstract":"Mental health assessment relies on episodic self-report scales, which convert subjective states such as stress into numerical scores but provide only sparse snapshots of wellbeing. Wearable devices offer longitudinal behavioral and physiological signals for continuous, low-burden monitoring. Recent LLM-driven personal-health agents enable natural language queries over wearable signals, but mainly handle short-term, retrieval-based lookups (e.g., highest step count over a week). They do not evaluate whether agents can reason over long-term signals to predict wellbeing scores paired with evidence-grounded rationales. To address this gap, we introduce BALMS, the first systematic benchmark of LLM-based agentic systems for longitudinal mental health sensing. BALMS spans 3 real-world longitudinal datasets, 2 task families (closed-form wellbeing-score prediction and rationale generation auto-graded by an LLM-as-Judge), 3 agentic paradigms evaluated across 5 open- and closed-source LLM backbones. We find that zero-shot agents rarely outperform a simple mean baseline, except with stronger backbones or compact, semantically meaningful features. Chain-of-thought prompting improves reasoning-oriented backbones, but does not guarantee temporal grounding or numerical correctness. Together with more analysis on efficiency and temporal scaling, BALMS highlights the need for longitudinal mental health agents that selectively retrieve history, ground temporal evidence, and reason over interpretable behavioral features.","authors":["Yu Yvonne Wu","Arvind Pillai","Yuliang Chen","Yuwei Zhang","Sudarshan Regmi","Tess Z. Griffin","Michael V. Heinz","Lisa A. Marsch","Nicholas C. Jacobson","Andrew Campbell"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-28","first_seen":"2026-08-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.27219","pdf_url":"https://arxiv.org/pdf/2608.27219","source_feed":"cs.CL","score":6,"bucket":"other","rubric_hits":["D2"],"tags":["LLM基准","心理健康","智能体"],"reason":"用LLM预测人类心理健康分数，但LLM是测量工具而非仿真被试，且无人类对照","model":"deepseek-v4-pro","scored_at":"2026-08-28T13:02:09","error":null,"has_summary":false,"summary":null},{"id":"2608.22444","version":2,"title":"Aligned Alone, Misaligned Together: Forecasting Adversarial Capture in LLM Agent Populations","zh_title":"单独对齐，群体失准：预测 LLM 智能体群体中的对抗性俘获","abstract":"The unit of AI safety evaluation is still the individual model, yet language-model agents are increasingly deployed in interacting populations that read and write one another's decisions. This raises a question no single-agent audit can answer: an agent that is well-calibrated on its own may still be pulled toward a different decision by the agents around it. We study this on a security-triage task, where populations of language-model monitors decide whether to escalate or dismiss alerts, and into which we can inject a committed minority that always pushes one way. We find that two alerts a single agent judges almost identically on its own can drive collective behavior far apart, so auditing any one member need not reveal what the population will do. Yet that collective behavior can be predicted in advance. From a population's benign, adversary-free operation alone, we calibrate a response function that forecasts, before any attack is run, how far a committed minority will later move it. We then ask what shifts the outcome and find that letting agents see each other's reasoning neutralizes a weak attack, while only delaying it against a strong one, turning the question from whether the population converges on the adversaries' choice into when. Finally, we exclude the hypothesis of capture being an irreversible trap: once the committed agents are removed, the population drifts back toward where it began, so capture is a temporary state. Alignment in isolation is not alignment in a population, yet what a population will do under attack can be read in advance, from how it behaves before any adversary arrives.","authors":["Isotta Magistrali","Chen Shani"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-08-28","first_seen":"2026-08-25","revised_at":"2026-08-28","abs_url":"https://arxiv.org/abs/2608.22444","pdf_url":"https://arxiv.org/pdf/2608.22444","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["多智能体模拟","安全评估","群体行为"],"reason":"LLM agent 群体模拟社会过程，但无真实人类数据对照，属边界情形。","model":"deepseek-v4-pro","scored_at":"2026-08-28T13:02:12","error":null,"has_summary":false,"summary":null},{"id":"2608.26154","version":1,"title":"Evaluating AI Generated Summaries for Cancer Patients","zh_title":"评估面向癌症患者的AI生成摘要","abstract":"Large language models (LLMs) are increasingly being integrated into digital health platforms to generate summaries of complex medical data. Although these models can improve patient engagement and communication, these systems also raise concerns about accuracy, faithfulness, and safety in clinical contexts. In this study, we evaluate AI-generated summaries within a cancer patient care application using a dual assessment framework. Human domain experts, including oncology clinicians and patient-facing care staff, provided ground-truth evaluations of summary quality along dimensions of accuracy, clinical relevance, and readability. In parallel, we employed LLMs serving as evaluators (LLM-as-a-judge). Some limitations were identified in the generated summaries e.g., occasional omissions and minor inaccuracies. These were systematically analyzed and used to iteratively improve prompt design, grounding, and safety guardrails.","authors":["Muhammad Aurangzeb Ahmad","Kim Shyu","Leon Oliver","Fergus Sleight","Paul Landau"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-28","first_seen":"2026-08-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.26154","pdf_url":"https://arxiv.org/pdf/2608.26154","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM评估","医疗摘要","人类对照"],"reason":"LLM作为评估者替代人工标注，非仿真人类被试，但涉及LLM评估与人类对照，属边…","model":"deepseek-v4-pro","scored_at":"2026-08-28T13:01:59","error":null,"has_summary":false,"summary":null},{"id":"2608.26372","version":1,"title":"Knowledge-Verified Emergent Deception in LLM Agents Under Conflicting Incentives","zh_title":"冲突激励下LLM智能体的知识验证涌现欺骗","abstract":"Large language models are increasingly deployed as autonomous agents serving users on behalf of companies, placing them in settings where user and deployer interests can conflict. When an agent knows that a user is owed something its deployer would prefer to deny, does it remain honest? Answering this is difficult because false statements can reflect either ignorance or hallucination rather than deception. To address this challenge, we introduce KnownLieBench , a knowledge-verified benchmark that first confirms through a neutral probe that an agent knows a user's entitlement, and then evaluates whether it makes false claims once an incentive to deny that entitlement is introduced. Specifically, KnownLieBench covers eight customer-service domains and 112 grounded cases, conducts multi-round dialogues with a trust-tracking customer agent, and separates deception emerging from incentive alone from deception produced under explicit instruction. Across eighteen proprietary and open-weight models, emergent deception varies substantially across model families and domains. We further use the benchmark for post-training, finding that honesty-directed fine-tuning reduces deception under incentive, while deception-graded fine-tuning increases lie success on honest-control dialogues without increasing lie frequency under incentive. By verifying entitlement knowledge before scoring deceptive behavior, KnownLieBench reduces the confound between lying and not knowing and enables more rigorous auditing and steering of agent honesty.","authors":["Zheyuan Liu","Weiliang Zhao","Xiangchi Yuan","Ningshan Ma","Yue Huang","Meng Jiang"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-28","first_seen":"2026-08-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.26372","pdf_url":"https://arxiv.org/pdf/2608.26372","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["LLM欺骗","智能体诚实性","基准测试"],"reason":"研究LLM在利益冲突下的欺骗行为，测量模型本身而非仿真人类被试，无人类数据对照。","model":"deepseek-v4-pro","scored_at":"2026-08-28T13:02:03","error":null,"has_summary":false,"summary":null},{"id":"2608.26529","version":1,"title":"Multi-Expert Conformal Risk Control for Pairwise LLM Judging in Open-Ended Dialogue","zh_title":"面向开放域对话中成对LLM评判的多专家共形风险控制","abstract":"In this paper, we explore multi-expert Conformal Risk Control (CRC) algorithms for pairwise LLM-as-a-Judge evaluation in open-ended dialogue. Our core insight is that multi-expert aggregation offers a complementary remedy to CRC: whereas CRC controls risk at the decision threshold through abstention, aggregation sanitizes the scoring function at its source. Guided by this, we first design two multi-expert CRC methods: Score Averaging and Decision Voting, which aggregate at the score and decision levels, respectively. While both strategies outperform single-expert methods on homogeneous expert panels, on heterogeneous LLM judges they remain risk-valid but recover only limited coverage, because a uniform threshold cannot match the experts' distinct scoring scales. To resolve this issue, we further propose Marginal-Calibrated Conformal Consensus (MC3): it captures distinct per-expert scales via initial threshold ratios, while jointly tuning a unified decision function $C_t(x)$ applied identically in both calibration and test, thereby preserving exchangeability. To evaluate our framework, we construct Panel, a 1,800-pair human pairwise-preference benchmark for open-ended dialogue. It is built on responses generated by four open-weight LLMs over dialogue contexts from three domains (ESConv, MSC, DREAM), with full logit access. In experiments, we find that both Score Averaging and Decision Voting substantially improve accuracy and acceptance rate on homogeneous panels. Notably, MC3 extends these gains to heterogeneous panels by accommodating distinct per-expert scoring scales across all three datasets.","authors":["Ming Cheng","Yusheng Dai","Qiuhong Ke","Zhaolin Chen","Lizhen Qu"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-28","first_seen":"2026-08-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.26529","pdf_url":"https://arxiv.org/pdf/2608.26529","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM评判","共形风险控制","多专家聚合"],"reason":"LLM作为评判者替代人工标注，属于标注员替代，非仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-08-28T13:02:04","error":null,"has_summary":false,"summary":null},{"id":"2608.26674","version":1,"title":"Do LLMs Understand Personality? Rethinking Persona Fidelity Evaluation through Structured Behavioral Inference","zh_title":"LLM理解人格吗？通过结构化行为推理重新思考人格保真度评估","abstract":"As large language models are increasingly deployed to simulate diverse human characters, ensuring persona fidelity, defined as the extent to which an agent's behavior consistently reflects the psychological and stylistic characteristics of a target persona, has become a critical requirement. However, existing evaluation paradigms primarily rely on either holistic LLM-based judges, which are prone to \"holistic appraisal hallucination'', or static psychometric inventories, which fail to capture the context-dependent fidelity required in dynamic dialogue. To address these limitations, we propose PRISM (Persona Reasoning with Inverse SFL-based Modeling), a psycholinguistically grounded framework that reformulates persona fidelity evaluation as a structured inverse inference task. Inspired by Systemic Functional Linguistics (SFL), PRISM decomposes persona fidelity into three functional dimensions: Task Framing, Interpersonal Stance, and Linguistic Style. It estimates dimension-specific evidence over a persona-conditioned label space and aggregates these signals into an interpretable and auditable evaluation process. Experiments show that PRISM yields more accurate and stable judgements than traditional holistic judging, providing a more reliable framework for persona fidelity evaluation.","authors":["Mengfan Li","Zesheng Wei","Xuanhua Shi","Yang Deng"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-28","first_seen":"2026-08-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.26674","pdf_url":"https://arxiv.org/pdf/2608.26674","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["人格保真度","LLM评估","心理测量"],"reason":"评估LLM人格保真度，测的是模型而非人群，但涉及人格测量与仿真，属边界情形","model":"deepseek-v4-pro","scored_at":"2026-08-28T13:01:52","error":null,"has_summary":false,"summary":null},{"id":"2608.27402","version":1,"title":"How Language Models Organize and Structure Moral Knowledge","zh_title":"语言模型如何组织与构建道德知识","abstract":"How do large language models (LLMs) organize moral knowledge? Models detect moral content broadly, but detection is a low bar. We ask whether they go further, distinguishing moral foundations from one another and organizing the relationships between them geometrically. We train six independent linear probes on open-weight language models, one per Moral Foundations Theory (MFT) category (care/harm, fair/cheat, lib/oppress, loy/betray, auth/subv, sanc/degrade), and examine how the resulting directions relate to each other in representation space. We find the directions neither collapse into a single moral detector nor isolate from one another. Rather, they span a near-maximal number of independent dimensions while sharing a positive common component. The shared component is the signature of integration, and it is moral-specific relative to a matched non-moral concept battery built identically (mean pairwise cosine 0.26 vs. 0.013). The geometry is consistent across architectures and scale and reaches its integration regime early in pre-training, well before probe accuracy saturates. The structure the model discovers shows no evidence of the individualizing/binding distinction predicted by Moral Foundations Theory (an underpowered test: only 20 candidate partitions exist) but rather reflects corpus statistics. Extending to moral dilemmas, each dilemma direction partially composes from its component foundations, at 2.7x a mismatched-pair baseline, while the majority of its variance encodes conflict-specific structure. The model represents moral tension itself, not a pre-resolved judgment.","authors":["Orion Reblitz-Richardson"],"categories":["cs.CL","cs.AI","cs.LG"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-28","first_seen":"2026-08-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.27402","pdf_url":"https://arxiv.org/pdf/2608.27402","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["道德知识表征","模型可解释性","道德基础理论"],"reason":"研究LLM的道德知识表征，属于对模型本身的测量，非仿真人类被试","model":"deepseek-v4-pro","scored_at":"2026-08-28T13:02:11","error":null,"has_summary":false,"summary":null},{"id":"2608.26150","version":1,"title":"Leveraging Large Language Models for Systematic Literature Review of Disease Spread Models","zh_title":"利用大语言模型进行疾病传播模型的系统文献综述","abstract":"Recent advancements in Large Language Models (LLMs) have created new opportunities to streamline and potentially automate many research processes, including systematic literature reviews (SLRs). This study reports an LLM pipeline development for extracting model-relevant information from 536 peer-reviewed agent-based modeling papers. We compare the results with those of a human-conducted SLR. Our results show paper-level accuracies of approximately 77.95% for GPT-4.1 and 81.67% for GPT-5.0. Field-level accuracy ranges from 32.40% to 100.00%, with more complex or subjective fields performing less reliably. Importantly, we find that agreement between LLMs is a potential indicator of output quality: low agreement may signal hallucinations, whereas high agreement combined with low accuracy may point to noise or errors in the human dataset. Overall, our study provides practical insights into prompt development and highlights both the potential and limitations of using LLMs for full-scale SLRs in the modeling and simulation domain.","authors":["Orhan Yagizer Cinar","Timur Emre Ozkose","Emma Von Hoene","Amira Roess","Taylor Anderson","Hamdi Kavak"],"categories":["cs.AI","cs.DL","cs.IR"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-28","first_seen":"2026-08-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.26150","pdf_url":"https://arxiv.org/pdf/2608.26150","source_feed":"cs.AI","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM标注","系统综述","自动化"],"reason":"LLM替代人工标注文献，非仿真人类被试，但方法可借鉴","model":"deepseek-v4-pro","scored_at":"2026-08-28T13:01:59","error":null,"has_summary":false,"summary":null},{"id":"2608.26885","version":1,"title":"Evaluating human and LLM screening workflows in a conceptually complex scoping review: Recall--workload trade-offs and run-to-run consistency","zh_title":"在概念复杂的范围综述中评估人类与LLM筛选工作流：召回率-工作量权衡与运行间一致性","abstract":"Background. Large language models (LLMs) are increasingly used for screening in evidence synthesis, where false negatives can remove relevant studies before full-text assessment. We compared human and LLM title-and-abstract screening workflows in a preregistered study embedded in a conceptually complex scoping review. Methods. After a conservative title-only screen, 1,131 records were screened by one review lead, four trained assistants screening non-overlapping subsets, and seven complete LLM runs using different models and processing configurations, including a nominally identical repeat run. We compared retained workload, operational recall against 316 verified eligible records, agreement, run-to-run consistency, and procedural burden. Because eligibility was verified only for records advanced and assessed in the parent review, recall estimates were operational. Results. No workflow recovered all verified eligible records. The human workflows and two GPT-5.4 file-batch runs retained 42.2-45.0% of records while achieving 82.3-82.9% recall. Gemini 3.1 file batches achieved the highest recall (83.9%) but retained 56.7% of records. All-at-once configurations recovered fewer eligible records than corresponding file-batch configurations. Two nominally identical GPT-5.4 file-batch runs agreed on 91.7% of records but differed on 94 records, including 29 verified eligible records retained by only one run. Discussion. LLM screening performance depended on the implemented workflow, not model identity alone. Processing configuration, workload, record-level variation, and human-LLM decision integration are therefore substantive properties of deployed systems. For high-recall tasks, LLMs are better suited to validated, auditable, human-supervised workflows than autonomous exclusion.","authors":["Nikol Figalov\\'a","Lynn Huestegge","Anne B\\\"ockler-Raettig"],"categories":["cs.AI","cs.HC","cs.SE"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-28","first_seen":"2026-08-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.26885","pdf_url":"https://arxiv.org/pdf/2608.26885","source_feed":"cs.AI","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM辅助筛选","人机对比","证据综合"],"reason":"LLM替代人工筛选文献，属标注替代而非仿真人类被试，但涉及人机对比与可靠性评估…","model":"deepseek-v4-pro","scored_at":"2026-08-28T13:01:54","error":null,"has_summary":false,"summary":null},{"id":"2608.27443","version":1,"title":"Do User-Authored Permission Policies Improve Protection Against AI Agent Overreach?","zh_title":"用户编写的权限策略能否改善对AI代理越权行为的防护？","abstract":"AI agents are poised to become a primary interface to digital products, acting across email, files, payments, and personal data. People without professional software backgrounds need understandable, reusable ways to control actions across services. We examine a mechanism in which a language model maps actions to plain-language consequence categories with user-authored \"allow\", \"ask\", or \"never\" rules. We ask what is gained and lost when decisions are made in advance as reusable rules rather than separately for each action. We analyzed 113 participants without professional software backgrounds across three conditions: per-action human-in-the-loop approval (HITL), automated per-action model review (AUTO), or user-authored consequence policy (POLICY). Participants judged 2 examples in each of 4 consequence categories; POLICY participants then set one rule per category. All supervised an 18-action simulated day, including 7 overreach actions. POLICY blocked less overreach than HITL (-20.1 percentage points, 95% CI [-32.1, -8.1]) and AUTO (-14.5 points, 95% CI [-25.8, -3.2]). POLICY lowered runtime prompts from 18.0 to 10.9, but total intervention time was not reliably lower when rule setup was included. Exploratory analysis showed that participants chose \"ask\" for 114 of 140 POLICY rules, returning most overreach actions to runtime. Of the 148 overreach actions executed in POLICY, 133 followed human approval and 15 ran automatically under \"allow\" rules. Across all 7 overreach actions, POLICY had the highest approval rate. Counterintuitively, user-authored rules did not by themselves provide stronger protection: many actions outside users' original requests went through after users approved them. These results reveal a gap between preference and commitment: repeatedly choosing \"ask\" preserves case-by-case choice but prevents a standing policy from settling decisions in advance.","authors":["Ting Yan"],"categories":["cs.HC","cs.CR"],"primary_category":"cs.HC","announce_type":"new","date":"2026-08-28","first_seen":"2026-08-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.27443","pdf_url":"https://arxiv.org/pdf/2608.27443","source_feed":"cs.HC","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["AI代理","权限策略","人机交互"],"reason":"研究用户授权策略对AI代理过度行为的防护，涉及人类参与者实验，但非用LLM仿真…","model":"deepseek-v4-pro","scored_at":"2026-08-28T13:02:12","error":null,"has_summary":false,"summary":null},{"id":"2607.27747","version":2,"title":"Can LVLMs Uncover the Truth Behind Visual Illusions? An Analysis of Perceptual and Reasoning Capabilities","zh_title":"LVLM能否揭示视觉错觉背后的真相？感知与推理能力分析","abstract":"Large Vision Language Models have integrated reasoning capabilities, elevating cognitive performance to new levels. However, existing evaluations either focus solely on perception or rely on specific domains such as maths or coding. Evaluation for reasoning capabilities that align with an open-world environment is still required, especially one that considers perception and reasoning jointly. To bridge this gap, we propose to evaluate LVLMs by exploiting visual illusions as a diagnostic tool. Visual illusions are phenomena in which the human visual system misinterprets objective signals, resulting in an understanding that deviates from reality. We constructed Illusion-Reasoning, a benchmark of illusion images collected from the real world, incorporating diverse annotated question-answer pairs. Based on Illusion-Reasoning, we show that the reasoning capabilities of a wide range of LVLMs are not as advanced as claimed. Our work provides new insights into LVLMs and offers future directions for optimisation. Our project is publicly available at https://github.com/zhaoliangjie55/EMNLP2026_Illusion.","authors":["Liangjie Zhao","Jiaqing Lyu","Kexin Tang","Zecheng Fang","Rong Yin","Yulan Hu","Da Li","Jianing Li"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-08-28","first_seen":"2026-07-31","revised_at":"2026-08-28","abs_url":"https://arxiv.org/abs/2607.27747","pdf_url":"https://arxiv.org/pdf/2607.27747","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["视觉语言模型","基准评测","视觉错觉"],"reason":"评估LVLM在视觉错觉上的感知与推理能力，属于模型能力评测，不以人类行为仿真为…","model":"deepseek-v4-pro","scored_at":"2026-08-28T13:02:13","error":null,"has_summary":false,"summary":null},{"id":"2608.01942","version":2,"title":"CultureVidBench: Benchmarking Cultural Understanding in Text-to-Video Generation","zh_title":"CultureVidBench：文本到视频生成中文化理解的基准测试","abstract":"Text-to-video (T2V) generation models have advanced rapidly, yet their ability to represent diverse cultural contexts remains underexplored. Existing benchmarks mainly focus on perceptual quality, physical plausibility, and text-video alignment, but do not directly assess whether generated videos capture culturally specific objects, actions, rituals, visible text, or audio cues. We introduce CultureVidBench, a comprehensive benchmark for evaluating cultural understanding in T2V generation. CultureVidBench contains 1,000 curated prompts covering 12 countries, 6 continents, 8 cultural regions, and 14 cultural aspects organized into three categories: material culture, social practice & performance, and ritual & ceremony. Designed specifically for video generation, CultureVidBench emphasizes dynamic and multimodal cultural representation, including social interactions, ritual procedure, and culturally appropriate visible text and audio. We evaluate seven representative T2V models through human user studies and MLLM-based automatic assessment across cultural faithfulness, multimodal cultural rendering, semantic adherence, and perceptual quality. Results show that although current models achieve strong semantic adherence and visual quality, they often fail to faithfully capture fine-grained cultural details, particularly for underrepresented regions, rituals, and multimodal cultural cues.","authors":["Xianjing Han","Yuhan Su","Yang Deng","Dong Ma","Wee Peng Tay","Bin Zhu"],"categories":["cs.CV","cs.CL","cs.MM"],"primary_category":"cs.CV","announce_type":"replace-cross","date":"2026-08-28","first_seen":"2026-08-04","revised_at":"2026-08-28","abs_url":"https://arxiv.org/abs/2608.01942","pdf_url":"https://arxiv.org/pdf/2608.01942","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C2"],"tags":["文本到视频生成","文化理解","基准测试"],"reason":"评估文本到视频生成模型的文化理解，不涉及用LLM仿真人类被试或与真实人类行为对…","model":"deepseek-v4-pro","scored_at":"2026-08-28T13:01:55","error":null,"has_summary":false,"summary":null},{"id":"2608.26131","version":1,"title":"Evaluating Language Models in Realistic Conversational Contexts","zh_title":"在真实对话情境中评估语言模型","abstract":"As Large Language Models (LLMs) are increasingly deployed to serve open-ended, multi-turn interactions, evaluating conversational quality at human scale has become a central challenge. Existing evaluation frameworks built for summarization, translation, or short-form QA tasks fall short of adequately measuring the consistency of human-scale dialogue, especially when derivation and validation of these metrics themselves often rely on synthetic rather than human sources. We fill the gap by introducing UPHELD (UPwork Human-Scale Evaluated Long Dialogues), a large, reference-full benchmark for evaluating human-scale conversational ability beyond factual correctness. UPHELD consists of hundreds of complete human-to-human dialogues authored by professional script writers, with realistic turn densities and 36,000+ per-turn human annotations across 30,000+ expert-generated dialogue turns. Using UPHELD, we systematically evaluate classical automatic metrics and reference-free LLM-as-a-judge approaches, and find them unreliable when correlated with expert human judgment. Building off this analysis, we use UPHELD to develop a Mixture-of-Judges framework that combines multiple evaluative signals and improves correlation with human assessments by approximately 30%. Overall, UPHELD provides a robust, human-grounded foundation for evaluating human-scale conversational intelligence that fills a crucial gap in the pre-existing LLM dataset landscape.","authors":["Ilija Subasic","Andrew Rabinovich","Zhao Chen"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-28","first_seen":"2026-08-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.26131","pdf_url":"https://arxiv.org/pdf/2608.26131","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["对话评测","基准数据集","LLM评估"],"reason":"论文是对话质量评测基准，不涉及用LLM仿真人类被试或行为对照，属于纯NLP能力…","model":"deepseek-v4-pro","scored_at":"2026-08-28T13:01:57","error":null,"has_summary":false,"summary":null},{"id":"2608.26137","version":1,"title":"Interpretable, Fairly Evaluated Automated L2 Speaking Assessment that Beats the Single-Human Ceiling and Why Pause Encoding Does Not Change LLM Fluency Scores","zh_title":"可解释、公平评估的自动化二语口语评分超越单人上限及停顿编码为何不改变LLM流利度分数","abstract":"Second-language (L2) English learners can rarely rehearse speaking with a partner. Speaking is also the most anxiety-laden skill. These gaps drive a fast-growing market for automated speaking practice and scoring. But an automated score is trustworthy only if it is accurate, interpretable, fair, and benchmarked against the right human bar. We build an interpretable feature-plus-LLM hybrid for spontaneous L2 dialogue. We evaluate it without ever fitting to the human labels, against the ICNALE Global Rating Archive: 140 speeches rated by ~80 trained raters on 10 analytic criteria. We score the 130 L2 speeches with usable audio. A deterministic De-Jong speech-timing composite reaches rho=0.764. Blended with a single text-LLM fluency judgment, it reaches Spearman rho=0.818 against the consensus gold. This agrees with the consensus better than 81% of the 80 individual trained raters: above the median rater (rho=0.73) and near the best, and at ~83% of the reliability-corrected maximum (kappa_max=0.99). The blend improves on the composite alone by +0.054 (paired-bootstrap 95% CI [0.017, 0.108], excludes 0); the LLM adds a coarse fluency ranking that the continuous composite refines. We also report a controlled null on pause encoding, bounded to effects below about +/-0.1 rho at this sample size. Holding the LLM and learner words fixed and varying only how pauses are written into the prompt, inline pause locations do not beat aggregate pause statistics (-0.069, CI [-0.15, +0.08]), and a grounded mid-clause criterion gives no reliable gain. The fluency signal comes from the measured speech-timing features, not from how pauses are written for the LLM. We back every claim with two agreeing learner-isolation methods, paired-bootstrap CIs, a monologue negative control, per-feature reproduction of classical measurements, and a per-L1 fairness audit.","authors":["Eichi Uehara"],"categories":["cs.CL","cs.LG","cs.SD"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-28","first_seen":"2026-08-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.26137","pdf_url":"https://arxiv.org/pdf/2608.26137","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["自动口语评分","LLM评估","二语习得"],"reason":"论文是LLM用于自动口语评分，属于NLP能力评测，不以人类行为仿真为目标","model":"deepseek-v4-pro","scored_at":"2026-08-28T13:01:59","error":null,"has_summary":false,"summary":null},{"id":"2608.26511","version":1,"title":"Sycophancy Suppression Can Impair Rational Updating: Anti-Sycophancy Should Preserve the Ability to Update","zh_title":"抑制谄媚可能损害理性更新：反谄媚应保留更新能力","abstract":"Large language models often exhibit sycophancy, revising their answers to align with users when users push back. Such answer flips, however, can arise from different causes. One possibility is that the model simply aligns with the user's feedback in order to satisfy them. Another is that the feedback genuinely contains useful evidence, prompting the model to update its answer in a rational way. We distinguish them as Unsupported-Yielding and Rational-Updating. Prior work focuses primarily on suppressing Unsupported-Yielding, while overlooking its effect on Rational-Updating. We address this gap with a two-turn evaluation framework that measures the two behaviors separately. Across representative training-time and inference-time interventions, we find that anti-sycophancy methods often encounter a trade-off in which reducing Unsupported-Yielding can sacrifice Rational-Updating, and vice versa, even when the two objectives are optimized jointly. Mechanistic analysis suggests that the two behaviors share an internal substrate: the MLP neurons and attention heads driving them overlap substantially, and their associated steering directions are positively aligned. We further conduct a preliminary orthogonalized steering exploration, which yields modest, backbone-dependent selectivity gains. Overall, our results suggest that anti-sycophancy should be treated not as a simple suppression problem, but as a selectivity problem, where effective interventions should preserve Rational-Updating while reducing Unsupported-Yielding.","authors":["Huanhuan Ma","Henry Peng Zou","Chengze Li","Enze Ma","Yunyue Su","Philip S. Yu"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-28","first_seen":"2026-08-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.26511","pdf_url":"https://arxiv.org/pdf/2608.26511","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["LLM行为","谄媚","模型对齐"],"reason":"研究LLM的谄媚行为与理性更新，属于模型行为分析，不涉及人类被试仿真或人类数据…","model":"deepseek-v4-pro","scored_at":"2026-08-28T13:02:03","error":null,"has_summary":false,"summary":null},{"id":"2608.27049","version":1,"title":"Research Design Tracking and Assessment for the Social Sciences","zh_title":"社会科学研究设计追踪与评估","abstract":"Reliable assessment of causal research designs in the social sciences is critical for evidence-based policy-making, yet has so far relied entirely on manual expert analysis. We introduce Automated Research Design Tracking and Assessment (ARDTrA), a task that involves detecting the research design used in a paper and assessing the quality of its application. We create an expert-annotated dataset of papers covering six families of counterfactual research designs and evaluate the task using a multi-turn RAG-based conversational pipeline. Across four retrieval strategies, four LLMs and six embedding models, we find that passage length is the main driver of performance, explaining 52-66% of the variance. A per-research-design analysis also shows that human and machine difficulty do not align: the designs that prove hardest for the system are not those on which expert annotators disagree most, pointing to two independent sources of task difficulty.","authors":["Marco Rovera","Sergiu Burlacu","Dominique Cappelletti","Alessio Tomelleri","Sonia Marzadro","Martina Bazzoli","Annalisa Tassi","Jessica Gagete-Miranda"],"categories":["cs.CL","cs.CY"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-28","first_seen":"2026-08-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.27049","pdf_url":"https://arxiv.org/pdf/2608.27049","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["研究设计评估","LLM评测","因果推断"],"reason":"论文评估LLM检测研究设计的能力，属NLP评测，不以人类行为仿真为目标。","model":"deepseek-v4-pro","scored_at":"2026-08-28T13:02:07","error":null,"has_summary":false,"summary":null},{"id":"2608.26958","version":1,"title":"Scaling Model-Generated Distillation Data Can Make Latent Teacher Traits More Recoverable","zh_title":"扩展模型生成蒸馏数据可使潜在教师特质更易恢复","abstract":"Scaling model-generated data is usually viewed as improving distillation: more examples should increase coverage, reduce noise, and produce stronger students. We show a second effect: larger datasets can make subtle teacher-specific signals easier to detect in the trained student, even when examples are off-task and never mention the trait. In a controlled setup inspired by subliminal learning, a teacher induced to express a target trait generates restricted off-task data, such as number-only completions. Students trained on different amounts of independent off-task data are evaluated in a separate domain, with matched no-trait controls isolating target-specific transfer. Our main finding is that larger independent datasets make the teacher's induced trait stand out more clearly in the student's later behavior. Other plausible traits may also strengthen with scale, but the target usually grows more. When the small-scale student already favors the target, scaling mainly amplifies that behavior; when it favors a related or salient alternative, more data can shift behavior toward the intended trait. Analyses of learned LoRA updates show a parallel trend. These effects appear across model families, trait types, multi-trait settings, and cross-model transfer. Our results suggest that scaling generated distillation data should be paired with trait-aware curation and evaluation, even when the data appears off-task or benign.","authors":["Zhichen Dong","Zhixuan Liu","Yuyu Fan","Xiangtian Li","Shuyang Zhang","Chao Yang"],"categories":["cs.LG","cs.CL"],"primary_category":"cs.LG","announce_type":"cross","date":"2026-08-28","first_seen":"2026-08-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.26958","pdf_url":"https://arxiv.org/pdf/2608.26958","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["模型蒸馏","特质传递","数据扩展"],"reason":"研究模型蒸馏中教师特质传递，不涉及人类被试仿真或人类数据对照，属多智能体/模型…","model":"deepseek-v4-pro","scored_at":"2026-08-28T13:02:06","error":null,"has_summary":false,"summary":null},{"id":"2608.27364","version":1,"title":"Sophistication in GenAI Use: Field Evidence from a Large Firm","zh_title":"生成式AI使用的复杂性：来自一家大型企业的实地证据","abstract":"We study how sophistication in generative AI (genAI) use varies among the back-office workforce of a large firm. Using proprietary data, we observe 713,564 employee prompts and their corresponding large language model responses from nearly 4,000 back-office employees across 15 functional areas over eight months in 2025. We document three main findings. First, senior employees exhibit more sophisticated genAI use, consistent with domain expertise complementing genAI capabilities. Second, sophistication varies considerably across functions and is highest in Strategy, Digital Innovation, and Project Management, three groups that share a focus on firmwide strategic initiatives and organizational change. Third, we observe neither improvements in sophistication over time nor lasting improvements following formal AI training, suggesting that sophisticated use can be difficult to change. Together, our study provides measures of and insights into sophisticated genAI use that managers can use to improve outcomes and that researchers can use in future research.","authors":["Nicholas J. Hallman","Zachary T. Kowaleski","Anu Puvvada","Jaime J. Schmidt"],"categories":["cs.AI","econ.GN","q-fin.EC"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-28","first_seen":"2026-08-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.27364","pdf_url":"https://arxiv.org/pdf/2608.27364","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["GenAI使用","企业实地研究","员工行为"],"reason":"研究员工使用GenAI的复杂程度，非用LLM仿真人类被试，无人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-08-28T13:02:10","error":null,"has_summary":false,"summary":null},{"id":"2608.26135","version":1,"title":"Data Science Approaches to Evaluating Honours Candidates","zh_title":"评估荣誉候选人的数据科学方法","abstract":"We present a modular data-science pipeline for estimating public sentiment towards individuals from fragmented, unstructured open-source intelligence (OSINT). The method chains web search, text extraction, relevance filtering, tokenisation, co-reference resolution, and sentiment analysis to convert heterogeneous web material into auditable person-level sentiment distributions. We compare AFINN and VADER with MINOS, a domain-informed sentiment algorithm designed to detect language associated with reputational risk, misconduct, and positive public contribution. Applied to public figures with known reputational outcomes, MINOS gives the clearest separation between positive, ambiguous, and negative cases. The results show that chained NLP and OSINT methods can support transparent, reproducible, human-in-the-loop sentiment assessment for high-stakes decision support. We demonstrate the approach on the UK Honours system, where individuals are required to display high standards of public conduct to maintain an Honour.","authors":["Francesca von Braun-Bates","Sunreeta Sen","Indraayudh Talukdar","Anirban Lahiri"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-28","first_seen":"2026-08-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.26135","pdf_url":"https://arxiv.org/pdf/2608.26135","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["情感分析","OSINT","NLP管道"],"reason":"纯NLP情感分析管道，不涉及LLM仿真人类被试或与人类行为对照","model":"deepseek-v4-pro","scored_at":"2026-08-28T13:01:57","error":null,"has_summary":false,"summary":null},{"id":"2608.26161","version":1,"title":"Mutual Debiasing via Dual-Seed Comparison for Probabilistic Sampling in Large Language Models","zh_title":"基于双种子比较的互去偏方法用于大语言模型概率采样","abstract":"Although Large Language Models (LLMs) demonstrate remarkable capabilities in reasoning and decision-making, high-fidelity probabilistic sampling remains a persistent challenge. When generating random variables, LLMs consistently exhibit systematic biases that warp the target probability distributions. Current approaches often rely on a single, self-generated seed, which inherits model-specific biases. To overcome this vulnerability, we introduce Dual-Seed Comparison (DSC), a transparent, tool-free protocol that utilizes two independent LLM-generated seeds to neutralize bias. DSC compares the character-level ordinal values of the two seeds to construct a bit sequence, converts and normalizes this sequence into a pseudo-uniform variate, and then maps the variate to the target distribution through the inverse cumulative distribution function (CDF). Empirical results show that DSC substantially outperforms existing methods across 96\\% of evaluated settings. Beyond direct sampling, task-adapted variants based on the DSC comparison operator improve distributional control in MCQ generation and attribute-constrained text-to-image prompting.","authors":["Zihao Guo","Hongtao Lv","Chaoli Zhang","Laiguo Yin","Lei Liu","Yonghui Xu","Lizhen Cui"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-28","first_seen":"2026-08-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.26161","pdf_url":"https://arxiv.org/pdf/2608.26161","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["概率采样","偏差校正","LLM生成"],"reason":"论文研究LLM概率采样偏差校正，属于模型能力改进，不涉及人类仿真或人类数据对照。","model":"deepseek-v4-pro","scored_at":"2026-08-28T13:02:01","error":null,"has_summary":false,"summary":null},{"id":"2608.26887","version":1,"title":"Planting a Latent Variable in Natural-Looking Text: a More Realistic Test of Belief States in LLMs and Their Link to Concept Geometry","zh_title":"在自然文本中植入潜变量：对LLM信念状态及其与概念几何联系的更真实测试","abstract":"LLMs are thought to track \"belief states,\" i.e., running probability distributions over the latent variables that govern language (Shai et al., 2024; Sarfati et al., 2026), but so far this has only been comprehensively demonstrated on toy synthetic data and in a few isolated case studies. It has also never been empirically connected to the geometry of LLM features (the concepts interpretability finds in model activations). In this work, we plant a controllable latent variable inside natural-looking text. An LLM teacher writes ordinary text while we \"subliminally\" steer it along one of K = 8 unrelated sparse autoencoder directions at each token, with the active directions following a ring-shaped Markov chain. A small transformer model trained on this corpus does indeed track the Bayesian posterior belief about our planted latent variable. Moreover, it also arranges the 8 states themselves on a ring, in the exact order of the Markov chain, which is supporting evidence that a concept's geometry can be formed by the statistical dynamics of the latent variable behind it.","authors":["Alexandru-Iulius Jerpelea"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-28","first_seen":"2026-08-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.26887","pdf_url":"https://arxiv.org/pdf/2608.26887","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["LLM信念状态","概念几何","潜变量"],"reason":"研究LLM内部信念状态与概念几何，属纯NLP能力评测，不以人类为参照系","model":"deepseek-v4-pro","scored_at":"2026-08-28T13:02:05","error":null,"has_summary":false,"summary":null},{"id":"2608.26226","version":1,"title":"LLM Agents for Time-Series: A Survey","zh_title":"面向时间序列的LLM智能体：综述","abstract":"LLM-based agents are increasingly being developed for time-series problems, but their design choices vary substantially across task settings. This survey adopts a problem-driven taxonomy that organizes these systems by the time-series problems they address rather than by isolated technical components. We group existing systems into four categories: forecasting and reasoning, augmentation and synthesis, anomaly detection and diagnosis, and decision support. Within each category, we examine how task requirements shape agent architecture, tool use, and memory design. We further summarize representative datasets and environments, and compare reported model performance under shared or closely related settings. Overall, this survey offers a task-oriented guide to designing LLM-based agents for time-series problems and identifies open gaps for future work.","authors":["Yilong Chen","Xiao Qin","Chenghao Liu","Liang Wu","Noelle I. Samia","Kaize Ding"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-28","first_seen":"2026-08-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.26226","pdf_url":"https://arxiv.org/pdf/2608.26226","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["LLM智能体","时间序列","多智能体系统"],"reason":"纯多智能体系统研究，agent协作解决时间序列任务，不涉及人类行为对照","model":"deepseek-v4-pro","scored_at":"2026-08-28T13:02:01","error":null,"has_summary":false,"summary":null},{"id":"2608.26753","version":1,"title":"Beyond Execution: Auditing Experimental Fidelity in LLM-Driven Scientific Research","zh_title":"超越执行：审计LLM驱动科学研究中的实验保真度","abstract":"LLM agents used for scientific experimentation must do more than generate executable code: they must implement the reference method faithfully, design experiments that test the paper's claims, and provide evidence supporting those claims. We show that agents often produce methodological hallucinations: silently reducing datasets or training budgets, replacing failed learning or generative components with lookup or oracle functions, or drawing conclusions from resource-limited settings where a method's claimed advantage disappears. To detect these failures, we introduce ABE-Ralph, a reference-anchored auditing framework that represents claims, protocols, required components, baselines, and metrics as structured experimental constraints, guides implementation through an 8-step workflow, and performs quantitative, qualitative, and code-level verification. Across 30 long-horizon reproduction runs covering 12 machine learning domains, ABE-Ralph achieves a 93% robust execution rate and identifies five scientific failure modes. In 23 NatureBench discovery tasks, ABE-Ralph matches or exceeds state-of-the-art performance on 5 tasks. These results show that reliable evaluation of AI scientists must assess whether the experimental design faithfully tests the intended claim and whether the resulting evidence supports it, rather than treating code execution or plausible metrics as evidence of scientific success.","authors":["Lezhi Yu","Xiaogang Xu","Yuhua Zhou","Shuibing He","Aimin Pan"],"categories":["cs.SE","cs.AI"],"primary_category":"cs.SE","announce_type":"cross","date":"2026-08-28","first_seen":"2026-08-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.26753","pdf_url":"https://arxiv.org/pdf/2608.26753","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["AI科学家","实验保真度","多智能体"],"reason":"论文研究AI科学家执行实验的保真度，属于多智能体协作完成任务，不涉及人类行为仿…","model":"deepseek-v4-pro","scored_at":"2026-08-28T13:02:05","error":null,"has_summary":false,"summary":null},{"id":"2608.27072","version":1,"title":"Emotional Preferences as Goal-Priority Regulation","zh_title":"情感偏好作为目标优先级调节","abstract":"A core question in decision-making for agents is whether the relative priorities of competing lower-level objectives can be determined by emotional preferences autonomously generated by higher-level goals, rather than being externally prespecified. Under changing external environments and evolving internal states, emotions play an important functional role in regulating the relative priorities of competing goals. Inspired by the goal-directed theory of emotion, this paper studies how such preference regulation can be computationally realized through reinforcement learning. We first propose a conception of emergent emotional preference: a high-level goal autonomously induces state-dependent preferences over competing lower-level objectives. This conception is built upon a framework consisting of a multi-objective reinforcement learning inner controller and an outer preference generator. The inner controller provides a repertoire of preference-conditioned goal-directed behaviors, while the outer preference generator learns a mapping from the current state to objective preferences through reinforcement learning on a high-level goal. We operationalize emotional preference as a state-dependent regulation of relative goal priorities that emerges through optimization. Furthermore, we characterize the policy space induced by preference regulation and derive an upper bound on the optimality gap in terms of the representation error of the inner behavioral repertoire. We show that the gap vanishes when the optimal policy can be represented by the available preference-conditioned policies. Experiments in self-constructed multi-objective exploration environments show that the learned preference function exhibits contextual priority switching, graded trade-offs, and temporal persistence, and outperforms the evaluated fixed-preference and handcrafted-preference strategies.","authors":["Shiqi Liu","Yihua Tan","Hu Fu","Guanyu Qi"],"categories":["cs.LG","cs.AI"],"primary_category":"cs.LG","announce_type":"cross","date":"2026-08-28","first_seen":"2026-08-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.27072","pdf_url":"https://arxiv.org/pdf/2608.27072","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["强化学习","多目标决策","情感计算"],"reason":"研究多目标强化学习中的情感偏好调节，不涉及LLM或人类行为仿真","model":"deepseek-v4-pro","scored_at":"2026-08-28T13:02:07","error":null,"has_summary":false,"summary":null},{"id":"2608.27238","version":1,"title":"Assessing Company Contributions to Societal Resilience: Extending the Societal Capacity Assessment Framework to Agentic AI","zh_title":"评估公司对社会韧性的贡献：将社会能力评估框架扩展到代理型AI","abstract":"Companies that deploy AI agents and make them available to others are creating the sociotechnical circumstances under which this technology integrates into existing social and economic structures. AI-deploying companies are institutional actors that actively shape society's capacity to withstand and govern the consequences of agentic AI. In view of these societal impacts, companies can build societal resilience by designing and promoting safer implementations of AI agents. To operationalize this goal, this paper adapts the indicator-based Societal Capacity Assessment Framework (SCAF) to measure how a company's deployment decisions contribute to societal resilience, inverting its original measurement of societal resilience as a backdrop for deployment decisions (Gandhi et al., 2025). Our procedure has two steps: a conceptual step in which we design a suite of indicators that define what SCAF's vulnerability, coping, and adaptive capacities mean when assessing a company's agentic AI deployment decisions; and a measurement step in which we apply this framework in a structured assessment of public-facing Microsoft documents.","authors":["Catherine Simons","Alexander K. Saeri","Peter Slattery","Neil Thompson"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2026-08-28","first_seen":"2026-08-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.27238","pdf_url":"https://arxiv.org/pdf/2608.27238","source_feed":"cs.CY","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["AI治理","社会韧性","公司责任"],"reason":"论文评估公司部署AI代理对社会韧性的贡献，不涉及用LLM仿真人类被试或与人类数…","model":"deepseek-v4-pro","scored_at":"2026-08-28T13:02:09","error":null,"has_summary":false,"summary":null},{"id":"2608.26797","version":1,"title":"On the Indistinguishability of Human v/s AI Generated Text","zh_title":"论人类与AI生成文本的不可区分性","abstract":"The rapid improvement of LLMs has made distinguishing AI-generated text from human writing a pressing problem. This challenge is further amplified by paraphrasing tools designed to make machine-generated text appear more \"human\". We study how access to human writing samples can be used to strategically paraphrase machine-generated responses toward the human distribution. Under a multi-sample setting with human and machine responses to the same prompts, we show that repeated paraphrasing moves the machine distribution toward the empirical human distribution under simple mixing and stability conditions. Our results derive an explicit convergence rate, extend the analysis to a finite-sample setting, and characterize how the required number of human samples and paraphrasing rounds scale with the desired error.","authors":["Jaee Ponde","Aritra Das","Mihir More","Debayan Gupta"],"categories":["cs.LG"],"primary_category":"cs.LG","announce_type":"new","date":"2026-08-28","first_seen":"2026-08-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.26797","pdf_url":"https://arxiv.org/pdf/2608.26797","source_feed":"cs.LG","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["AI文本检测","文本改写","对抗性样本"],"reason":"研究AI文本检测与改写，不涉及用LLM仿真人类被试或与人类行为对照","model":"deepseek-v4-pro","scored_at":"2026-08-28T13:02:05","error":null,"has_summary":false,"summary":null},{"id":"2608.26388","version":1,"title":"Assessing Socio-Cyber Vulnerability Using Survey and Social Media Data","zh_title":"利用调查和社交媒体数据评估社会网络脆弱性","abstract":"The rapid growth of social media participation has increased exposure to socially engineered cyber threats (e.g., phishing, romance fraud, and tech-support scams), yet prevailing assessment tools remain fragmented: the Common Vulnerability Scoring System (CVSS) is primarily technical and largely omits human susceptibility, while the Social Vulnerability Index (SVI) is community-oriented and lacks cyber-specific modeling. To address this gap, we propose the Social Cyber Vulnerability Index (SCVI), an interpretable, uncertainty-aware metric combining two components: (i) an Individual Vulnerability Index (IVI), capturing awareness, behavior, psychological factors, and prior victimization, and (ii) an Attack Severity Index (ASI), capturing attack frequency, consequences, and sophistication. We validate SCVI across heterogeneous modalities: a nationally scoped survey (iPoll; 4,596 U.S. adults) and social-media narratives (450 Reddit r/scams reports, 2016-2024), demonstrating computation from both structured questionnaires and CI-driven feature extraction from text. Sensitivity analysis and 10,000-iteration Monte Carlo simulations show stable rankings under plausible weight variability and reveal context-dependent drivers. SCVI captures distinct socio-technical signals (Spearman correlation with CVSS $\\rho = 0.33$; with SVI $\\rho \\approx -0.01$) and surfaces demographic and regional disparities. SCVI also provides substantially stronger separation between victim and non-victim groups than CVSS and SVI, supporting identification of high-risk populations and prioritization of interventions against emerging AI-enabled scams.","authors":["Shutonu Mitra","Qi Zhang","Tomas Neguyen","Hossein Salemi","Fengxiu Zhang","Michin Hong","Chang-Tien Lu","Hemant Purohit","Jin-Hee Cho"],"categories":["cs.SI"],"primary_category":"cs.SI","announce_type":"new","date":"2026-08-28","first_seen":"2026-08-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.26388","pdf_url":"https://arxiv.org/pdf/2608.26388","source_feed":"cs.SI","score":0,"bucket":"other","rubric_hits":["C5"],"tags":["社会网络脆弱性","网络安全","人类数据"],"reason":"论文使用人类调查和社交媒体数据构建社会网络脆弱性指数，未使用LLM仿真人类被试…","model":"deepseek-v4-pro","scored_at":"2026-08-28T13:02:03","error":null,"has_summary":false,"summary":null},{"id":"2608.26358","version":1,"title":"An Anonymized Urn-Based Experimental Dataset on Decision-Making under Risk and Ambiguity","zh_title":"风险与模糊决策的匿名瓮实验数据集","abstract":"This data paper describes an anonymized release derived from controlled behavioral experiments on decision-making under risk (known probabilities) and ambiguous uncertainty (partially specified or unknown probabilities) in Ellsberg-type urn tasks. The release contains 4,486 decision records from 246 adult participant records across two implementations (laboratory and online). In each task, participants reported their maximum stake to enter a subsequent lottery; these stake responses can be interpreted as proxies for willingness to pay within the implemented incentive mechanism. In a subset of tasks, participants also selected the color they wished to bet on. The online implementation additionally includes ex-post evaluations of perceived uncertainty across situations. The repository further provides detailed game definitions, realized lottery outcomes, participant instructions, and a data dictionary of variables. The source export contained 4,544 decision records; 58 records belonging to five participants younger than 18 were excluded from this release because documentation of guardian permission for unrestricted public data publication was not available. The dataset supports replication studies, comparisons of ambiguity-attitude models, and analyses of behavioral stability under repeated participation.","authors":["V\\'aclav Kratochv\\'il","Radim Jirou\\v{s}ek","Kl\\'ara \\v{S}im\\r{u}nkov\\'a","Simona Ba\\v{z}antov\\'a"],"categories":["econ.GN","q-fin.EC"],"primary_category":"econ.GN","announce_type":"new","date":"2026-08-28","first_seen":"2026-08-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.26358","pdf_url":"https://arxiv.org/pdf/2608.26358","source_feed":"econ.GN","score":0,"bucket":"other","rubric_hits":["C5"],"tags":["实验数据集","风险决策","模糊决策"],"reason":"论文是真实人类实验数据集，不涉及LLM仿真，方向相反。","model":"deepseek-v4-pro","scored_at":"2026-08-28T13:02:01","error":null,"has_summary":false,"summary":null},{"id":"2608.24912","version":1,"title":"Analyzing and Correcting Benevolence Bias in Large Language Models","zh_title":"分析和纠正大语言模型中的仁慈偏差","abstract":"Large language models (LLMs) are increasingly used as stand-ins for human respondents, from opinion polls and simulated survey participants to agent-based social simulations. These uses rest on one assumption: that conditioning a model on who a person is yields answers resembling those of real people from that group. Here we identify and measure benevolence bias, a small but consistent tendency for aligned LLMs to lean toward the kinder, safer, more socially approved answer on value-laden survey questions. Across 18 widely used models, four social-science datasets (ANES, GSS, WVS, and a cross-cultural prospect-theory replication) and six psychological categories, we find that the bias is a stable model property, not a quirk of any one system: it points the same way across models, grows with model size, and traces to the post-training stage. Prompt language and framing change its size but never its direction, and a \"malicious persona\" stress test shows a one-sided limit: aligned models struggle to play people who are less kind, less prosocial or more harm-tolerant than average. The issue is thus not only a shifted average, but a narrowed range of people the model can imitate. The bias sits in the middle of the answer distribution rather than its tails, and survives changes in sampling temperature and simple prompted reflection. The encouraging news is that it is easy to diagnose and straightforward to fix: a light-touch contrastive calibration, which needs no retraining and works on black-box APIs, brings all six categories back to the human baseline. Our results give researchers a clear map of where aligned LLMs can already be trusted as human stand-ins, where they need care, and a ready-to-use method for closing the gap.","authors":["Yuanzi Li","Junhao Wang","Minghui Liu","Boyi Li","Bingchen Chen","Zihang Tian","Jingyu Zhao","Yuhan Wang","Lei Wang","Pei Wang","Jinchao Wu","Xu Chen"],"categories":["cs.HC","cs.AI"],"primary_category":"cs.HC","announce_type":"new","date":"2026-08-27","first_seen":"2026-08-27","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.24912","pdf_url":"https://arxiv.org/pdf/2608.24912","source_feed":"cs.HC","score":10,"bucket":"selected","rubric_hits":["A1","A2","A5","B1","B2","B3","B4"],"tags":["LLM仿真","算法保真度","偏差校正"],"reason":"直接研究LLM作为人类被试替代品的偏差，使用真实人类数据对照，并提出校准方法。","model":"deepseek-v4-pro","scored_at":"2026-08-27T13:03:07","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-27","rank":1,"question":"对齐后的大语言模型在作为人类被试回答价值负载的调查问题时，是否存在系统性的仁慈偏差，其来源、表现和可校正性如何？","design":"该研究并非传统仿真实验，而是对18个广泛使用的LLM进行系统性测试：将模型置于模拟人类受访者的角色，输入来自ANES、GSS、WVS和跨文化前景理论复制的调查问题，测量模型回答在六个仁慈偏差类别（社会赞许性、伤害规避、亲社会动机、仁慈解释、公平乐观、情感软化）上的偏差程度，并考察模型规模、训练阶段、提示语言、框架、角色设定、采样温度等因素的影响，最后提出对比校准方法。","baseline":"使用四个真实人类调查数据集作为基准：美国国家选举研究（ANES）、综合社会调查（GSS）、世界价值观调查（WVS）以及一项跨文化前景理论复制的数据。","findings":"LLM存在稳定且一致的仁慈偏差，倾向于选择更友善、更安全、更符合社会期望的答案，且该偏差随模型规模增大而增强，主要源于后训练阶段。提示语言和框架只能改变偏差大小而不能改变方向，恶意角色压力测试显示模型难以模仿低于人类平均水平的亲社会或伤害容忍度，但对比校准方法无需重新训练即可将偏差校正至人类基线。","reliability":"论文指出偏差位于答案分布的中间而非尾部，且对采样温度和简单提示反思不敏感；恶意角色测试显示模型在部分维度上无法达到人类低仁慈端，表明仿真范围收窄。","relevance":"该研究直接针对LLM作为人类被试替代品的可靠性问题，提供了系统性的偏差测量和校正方法，对于关注仿真效度和偏差的研究者具有重要参考价值。","inspiration":"该研究采用多模型、多数据集、多心理类别的系统测量框架，并通过对比校准进行偏差校正，值得借鉴。｜可以迁移到经济金融领域的调查实验和个体决策仿真，例如风险偏好、时间偏好、公平观念、信任与合作等。｜以LLM作为虚拟被试，施加不同的经济情境或政策干预，测量其选择或态度，并与真实实验数据（如实验经济学中的公共品博弈、最后通牒博弈、风险偏好问卷等）进行对照，检验并校正LLM的偏差。"}},{"id":"2604.06223","version":3,"title":"The Quiet and the Compliant: How Regulation and Polarization Shape Conventional Wisdoms on Corporate Social Engagement in High-risk Settings","zh_title":"沉默与顺从：监管与极化如何塑造高风险环境下企业社会参与的常规智慧","abstract":"With the international business landscape becoming more crisis-ridden as risks proliferate, how do the professionals who implement corporate social initiatives in high-risk environments perceive their work, and what can this reveal about the forces shaping business engagement with society in crisis contexts? We present findings from a synthetic survey of 400 corporate professionals working on social impact in fragile and conflict-affected settings to understand conventional wisdoms and best practices on corporate strategy and activity in high-risk settings. Drawing on political corporate social responsibility (CSR), synthetic survey, and international business literatures, we test seven hypotheses about how regulatory environments, political polarization, sector characteristics, and organizational structures shape corporate social engagement in high-risk contexts. The synthetic results suggest that European professionals report significantly higher strategic integration of social impact across all measured dimensions, while US professionals overwhelmingly report that political polarization hinders social initiatives, yet this perception does not predict unreported social activities, complicating the emerging \"quiet CSR\" narrative. Extractive industry professionals deliver both the highest operational preparedness and the highest complicity awareness, a pattern we conceptualize as presence-dependent reflexivity. These patterns deliver a baseline to detect the theorized dynamics and offer preliminary theoretical propositions for future real-world empirical testing.","authors":["Jason Miklian"],"categories":["physics.soc-ph","cs.SI"],"primary_category":"physics.soc-ph","announce_type":"replace-cross","date":"2026-08-27","first_seen":"2026-03-27","revised_at":"2026-08-27","abs_url":"https://arxiv.org/abs/2604.06223","pdf_url":"https://arxiv.org/pdf/2604.06223","source_feed":"cs.SI","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM仿真","合成调查","企业社会责任"],"reason":"用LLM合成400名企业专业人士调查，模拟高风险环境下的态度与决策，并与真实文…","model":"deepseek-v4-pro","scored_at":"2026-08-27T13:03:34","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-27","rank":2,"question":"在脆弱和冲突影响的高风险环境中，企业社会参与的专业人士如何感知其工作，以及监管环境、政治极化、行业特征和组织结构如何塑造企业社会参与？","design":"使用合成调查方法，模拟400名在欧美总部、员工超过1000人的企业社会影响专业人士，通过23个李克特量表题测量战略整合、监管压力、政治极化、运营准备、供应链协调、子公司自主权和ESG评级有效性等感知，并基于理论推导出七个假设进行检验。","baseline":"无对照","findings":"欧洲专业人士在所有维度上报告显著更高的社会影响战略整合，而美国专业人士压倒性地报告政治极化阻碍社会倡议，但这种感知并不预测未报告的社会活动，复杂化了“安静CSR”叙事；采掘业专业人士同时表现出最高的运营准备和最高的共谋意识，被概念化为存在依赖的反身性。","reliability":"论文未讨论","relevance":"该研究使用LLM合成调查数据，模拟企业专业人士在高风险环境下的态度和决策，并基于理论提出假设进行检验，属于用LLM进行人类仿真实验的研究，且涉及政策评估场景，与您的兴趣高度相关，值得阅读原文了解其方法论和发现。","inspiration":"该方法通过合成调查生成理论预测的基准数据，为后续真实数据对比提供参照，可借鉴其构建“常规智慧基线”的思路；可迁移到政策评估中的企业行为研究，如ESG监管对企业社会参与的影响；设计上，可用LLM模拟企业高管作为被试，施加不同监管环境（如强制尽职调查 vs 反ESG法案）的处理，测量其战略整合和沉默行为，并与真实企业调查数据对照。"}},{"id":"2608.25771","version":1,"title":"Large Language Model Few-Shot Prompting with Dilemma Training Outperforms Human Surrogates in Predicting Patient Preferences","zh_title":"基于困境训练的大语言模型少样本提示在预测患者偏好上超越人类代理","abstract":"In serious illness, human surrogates often struggle to accurately predict patient preferences (68% accuracy), causing decision conflict. Personalized Patient Preference Predictor (P4) agents offer a potential solution, but prior prototypes treat values as static ratings, ignoring the contextual, situation-dependent nature of medical choices. Grounded in the 'logic of care', we present P4-DT (Dilemma Training), a P4 agent that constructs a patient decision policy by engaging users with varied medical dilemmas, eliciting individual preference reasoning through bi-directional training. In a study with 12 patient-surrogate dyads, P4-DT predicted patient treatment choices with 81.7% accuracy, significantly exceeding chance (OR = 5.61 [2.03, 15.51], p < .001) and outperforming both unassisted surrogates (55.0%; OR = 3.67 [1.59, 8.47], p = .002) and surrogates assisted by P4-DT (61.7%). Comparative prompt analyses showed that incorporating contextual scenario decisions and open-ended text improved accuracy by 15.0 percentage points over initial values ratings alone. We discuss implications for further testing and designing of context-aware AI agents that embody richer human experience to partner in complex decision-making.","authors":["Natasha Ureyang","Sebastian Porsdam Mann","Yuxin Liu","Zuriel Hassirim","Melanie Almonte","Wenhao Chen","Joyce Ng","Thant Nay Lin","Aung Thiha","Gerald CH Koh","Brian David Earp","Pin Sym Foong"],"categories":["cs.HC","cs.LG"],"primary_category":"cs.HC","announce_type":"new","date":"2026-08-27","first_seen":"2026-08-27","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.25771","pdf_url":"https://arxiv.org/pdf/2608.25771","source_feed":"cs.HC","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2"],"tags":["LLM仿真","患者偏好预测","人类对照"],"reason":"用LLM预测患者偏好，与人类代理对照，评估准确率，属核心仿真研究。","model":"deepseek-v4-pro","scored_at":"2026-08-27T13:03:10","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-27","rank":4,"question":"如何通过让患者参与医疗困境训练来构建个性化患者偏好预测器（P4-DT），以提高对患者治疗偏好的预测准确率？","design":"使用GPT-5.5模型作为P4-DT代理，通过提示工程进行少样本学习。对12名患者（MP）进行训练：先填写价值观调查，再对5个医疗困境场景做出治疗偏好决策并解释理由，同时可查看模型预测并反馈。测试阶段，MP对5个新场景做决策，模型基于训练数据预测其偏好；同时人类代理人（TO）独立预测MP偏好，并在P4-DT辅助下再次预测。结果变量为预测准确率（方向一致性）。","baseline":"人类代理人（TO）的预测准确率（55.0%），以及TO在P4-DT辅助下的预测准确率（61.7%）。","findings":"P4-DT预测患者治疗选择的准确率为81.7%，显著高于随机水平（OR=5.61, p<.001），且优于未辅助的人类代理人（55.0%）和P4-DT辅助的代理人（61.7%）。比较提示分析显示，纳入情境场景决策和开放式文本比仅使用初始价值观评分提高了15.0个百分点的准确率。","reliability":"论文未讨论","relevance":"该研究用LLM模拟患者偏好预测，并与真实人类代理人对照，评估预测准确率，属于核心的LLM仿真人类决策研究，且涉及医疗决策场景，对关注仿真可靠性和偏差的研究者有参考价值。","inspiration":"借鉴其通过情境化困境训练和开放式文本解释来捕捉个体决策逻辑的方法，可迁移到消费者金融决策或政策偏好预测中。｜例如，在消费者信贷选择或退休储蓄决策中，可让LLM通过模拟具体金融困境（如贷款选择、投资风险权衡）来学习个体偏好。｜设计：招募真实消费者作为被试，先填写价值观和风险偏好问卷，再对5个金融困境场景做出选择并解释理由，训练LLM预测其在新场景中的选择；同时让人类代理人（如配偶）预测被试选择，比较LLM与人类代理人的预测准确率，并以被试实际选择为基准。"}},{"id":"2608.24920","version":1,"title":"Semantic Variability of Replies Across LLMs: Implications for Designing Conversation-Based Assessment","zh_title":"不同大语言模型回复的语义变异性：对设计基于对话的评估的启示","abstract":"This study examines whether LLM-generated replies remain semantically consistent when the underlying LLM changes. Using messages from real collaborative conversations, we compared the semantic similarity of generated replies across LLMs under two conditions: with and without preceding chat history. Results show that model choice and conversational context both affect response similarity and alignment with human replies. These findings indicate that prompting and conversational context alone may not be sufficient to preserve response consistency across LLMs, highlighting the need for infrastructure and design strategies that can maintain stable and comparable responses amid the rapid and continuous evolution of LLMs.","authors":["Jiangang Hao"],"categories":["cs.CL","cs.AI","cs.HC"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-27","first_seen":"2026-08-27","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.24920","pdf_url":"https://arxiv.org/pdf/2608.24920","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A2","B1","B4"],"tags":["LLM仿真","语义一致性","对话评估"],"reason":"评估LLM回复与人类回复的一致性，涉及仿真可靠性，有真实人类数据对照，并指出失…","model":"deepseek-v4-pro","scored_at":"2026-08-27T13:03:07","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-27","rank":10,"question":"当底层大语言模型改变时，LLM生成的回复在语义上是否保持一致？","design":"使用四个LLM（GPT-4o mini、GPT-5.4、GPT-5.4 mini、GPT-5.4 nano）扮演协作对话中的回复者，对从真实协作对话中选取的61条焦点消息生成回复，比较有无对话历史两种条件下回复的语义相似度。","baseline":"真实人类在原始协作对话中产生的回复，并编码为高、中、低相关性。","findings":"模型选择和对话上下文均影响回复的语义相似性及与人类回复的一致性。仅靠提示和对话上下文不足以保持跨LLM的回复一致性。","reliability":"论文指出提示和对话上下文不足以保持跨LLM的回复一致性，需要基础设施和设计策略来维持稳定可比的回复。","relevance":"该研究直接评估LLM回复与人类回复的一致性，涉及仿真可靠性，有真实人类数据对照，并指出失效条件，值得阅读原文。","inspiration":"借鉴其通过改变模型和上下文条件来系统评估回复一致性的实验设计，以及使用真实人类回复作为基准的对照方法。｜可迁移到经济金融领域的对话式调查或实验，如消费者金融咨询、投资顾问对话或政策沟通中的LLM应用。｜以LLM作为虚拟被试，在有无对话历史条件下对同一金融咨询问题生成回复，测量回复语义相似度，并与真实人类顾问的回复进行对照，评估LLM替代人类被试的可靠性。"}},{"id":"2608.25999","version":1,"title":"Distinct dynamics of conceptual and referential disruptions in human reading and large language model processing","zh_title":"人类阅读与大语言模型处理中概念与指称干扰的不同动态","abstract":"Linguistic meaning is grounded in conceptual content, from which reference to particular entities emerges as words enter discourse. To examine the processing dynamics associated with these two dimensions of meaning, we selectively disrupted conceptual or referential information in short narratives and traced the resulting effects in human self-paced reading and in the predictive and representational processing of large language models. In human reading, conceptual disruptions produced a strong but localized processing cost, emerging immediately after the distorted word, reaching an early maximum, and then declining rapidly. Referential disruptions produced weaker effects, which decreased more gradually across subsequent words, and were more strongly modulated by sentence boundaries. In the language model, both disruptions emerged immediately at the manipulated word. Contextual model surprisal showed a pattern closely paralleling human reading: conceptual disruption produced a larger, more locally concentrated effect that decayed rapidly, whereas referential disruption produced a smaller and more gradual downstream effect. Output-layer representations showed a different pattern: referential disruption produced a larger initial displacement, while both distortions were subsequently characterized by power-law decay. Together, these results provide convergent evidence for distinguishable processing dynamics of two types of meaning: conceptual information imposes a more locally concentrated integration cost, whereas referential information engages a more distributed process of maintaining discourse-level identity.","authors":["Rui He","Nihal Altay","Wolfram Hinzen"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-27","first_seen":"2026-08-27","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.25999","pdf_url":"https://arxiv.org/pdf/2608.25999","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A1","B1"],"tags":["LLM仿真","人类对照","认知建模"],"reason":"用LLM模拟人类阅读行为并与人类数据对照，方法可迁移至仿真研究","model":"deepseek-v4-pro","scored_at":"2026-08-27T13:03:30","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-27","rank":14,"question":"概念信息和指称信息的加工动态在人类阅读和大型语言模型处理中是否不同？","design":"本研究不是用LLM模拟人类被试，而是将LLM作为计算模型与人类被试并行比较。在短叙事中分别施加概念扭曲（CD）和指称扭曲（RD），测量人类自定步速阅读的逐词阅读时间，以及LLM的上下文惊奇度和输出层表征位移。","baseline":"人类自定步速阅读实验的逐词阅读时间数据。","findings":"人类阅读中，概念扭曲产生强而局部的加工代价，立即出现、早期达峰、迅速衰减；指称扭曲效应较弱，衰减更慢，且更受句子边界调节。LLM的上下文惊奇度与人类模式相似，但输出层表征显示指称扭曲初始位移更大，两者随后均呈幂律衰减。","reliability":"论文未讨论","relevance":"该研究将LLM的预测和表征与人类逐词阅读行为直接对照，展示了LLM在捕捉不同语义加工动态上的异同，对评估LLM作为人类语言加工仿真模型的有效性具有参考价值。","inspiration":"借鉴其通过施加局部扭曲并追踪下游效应动态来分离不同认知成分的方法，可用于经济金融文本信息加工研究。｜可迁移到政策公告的预期形成研究，考察概念性信息（如政策内容）与指称性信息（如政策对象）对市场预期更新的不同影响。｜以LLM为被试，在政策公告文本中分别扭曲概念词或指称词，测量模型对后续经济指标预测的惊奇度变化，并与真实市场中分析师预期调整数据对照。"}},{"id":"2608.25236","version":1,"title":"Rare Diseases, Common Dilemmas: LLMs Prioritize Equal Resource Distribution over Patient Benefit in Decision-Making","zh_title":"罕见病，常见困境：LLM在决策中优先考虑资源平等分配而非患者获益","abstract":"Clinical decision-making often involves prioritizing ethical values, such as beneficence, non-maleficence, respecting a patient's autonomy, and justice. Recent work has begun to assess how large language models (LLMs) make such subjective, value-laden clinical judgments. However, evaluations of LLM decision-making in rare disease care contexts, where ethical tensions are ubiquitous and where scarce prior information likely impacts LLM behavior, are still lacking. Here, we present a benchmark of 208 clinically grounded rare disease vignettes, each of which presents genuine, high-stakes conflicts. When prompting 11 state-of-the-art LLMs to choose between clinically defensible yet ethically conflicting next steps embedded within these vignettes, we found that all evaluated models consistently prioritized justice over other core bioethical principles. Specifically, models overwhelmingly favor equal resource allocation over need-based considerations, indicating LLMs' limited responsiveness to differences in clinical severity or situational context. We also identify a strong authority-framing effect: models favor justice in committee-based contexts and shift toward beneficence and autonomy only when final decisions are framed as being made by clinicians or patients respectively. Our work suggests that institutional pressures surrounding rare disease resource utilization may be silently reflected in LLM-based decision support systems, with finer ethical considerations disregarded.","authors":["Minda Zhao","Xu Han","Rishabh Goel","Maya Dagan","Noa Dagan","Adithya Madduri","Payal Chandak","Shilpa Nadimpalli Kobren","Isaac S. Kohane"],"categories":["cs.CY","cs.AI","cs.CL"],"primary_category":"cs.CY","announce_type":"cross","date":"2026-08-27","first_seen":"2026-08-27","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.25236","pdf_url":"https://arxiv.org/pdf/2608.25236","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A1","B2","B4"],"tags":["LLM决策仿真","伦理决策","资源分配"],"reason":"用LLM模拟临床决策，与人类伦理原则对照，涉及资源分配，但非经济学实验且无真实…","model":"deepseek-v4-pro","scored_at":"2026-08-27T13:03:10","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-27","rank":11,"question":"大语言模型在罕见病临床决策中如何权衡自主、行善、不伤害与公正等伦理原则？","design":"构建208个罕见病临床情境短文，每个情境包含两个临床可辩护但伦理冲突的下一步行动；用11个先进LLM作为被试，通过提示让模型选择行动，并操纵决策权威框架（委员会、临床医生、患者），测量模型选择所体现的伦理原则优先级。","baseline":"无对照","findings":"所有模型一致优先考虑公正原则，压倒性地选择平等资源分配而非基于需求，对临床严重性或情境差异反应有限。存在权威框架效应：委员会情境下偏向公正，临床医生或患者决策情境下分别转向行善和自主。","reliability":"论文未讨论","relevance":"该研究用LLM模拟人类伦理决策，虽无真实人类对照，但揭示了LLM在资源分配场景中的系统性偏差，对关注仿真可靠性与偏差的研究者有参考价值。","inspiration":"借鉴其构造伦理冲突情境并操纵决策权威框架的方法，可迁移到经济金融中的资源分配或政策权衡问题，如公共预算分配、医疗资源定价或信贷审批中的公平与效率权衡。设计：用LLM作为被试，呈现稀缺资源分配情境（如器官移植、疫苗分配或信贷额度），操纵决策者角色（委员会、个体官员、受益人），测量分配方案（平等vs.按需vs.按效益），并与真实人类实验或调查数据对照。"}},{"id":"2608.24908","version":1,"title":"Hallucination by proxy in LLM-assisted differential diagnosis","zh_title":"LLM辅助鉴别诊断中的代理幻觉","abstract":"Current evidence suggests that LLM assistance could augment the diagnostic accuracy of clinicians. However, these systems are black boxes, susceptible to hallucinations, and project a potentially misleading level of confidence. It is currently unknown whether physicians are susceptible to accepting fabricated LLM suggestions, and whether this susceptibility varies with experience. We poisoned the system prompt of an LLM-based diagnostic assistant, forcing it to suggest a fictitious disease (neurocadmiumatosis) within an otherwise legitimate differential diagnosis. Across two independent phases, 18 of 41 participants (44%) incorporated neurocadmiumatosis into their final differential following LLM interaction: 18 of 26 participants with 6 months or less of neuroradiology training (69%) and 0 of 15 participants with >6 months of neuroradiology training (0%). Our results indicate that radiologists, particularly early in their training, are susceptible to LLM hallucinations. This \"hallucination by proxy\" phenomenon was exclusive to physicians with limited subspecialty experience, underscoring the need for structured training in critical appraisal of AI-generated content.","authors":["Bastien Le Guellec","Su-Hwan Kim","Ibrahima Niang","Aghiles Hamroun","Gr\\'egory Kuchcinski"],"categories":["cs.HC","cs.CY"],"primary_category":"cs.HC","announce_type":"new","date":"2026-08-27","first_seen":"2026-08-27","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.24908","pdf_url":"https://arxiv.org/pdf/2608.24908","source_feed":"cs.HC","score":7,"bucket":"pending","rubric_hits":["A1","B1","B4"],"tags":["LLM幻觉","医生决策","人类行为对照"],"reason":"研究医生对LLM幻觉的易感性，用真实医生行为对照，属人类决策仿真与偏差评估","model":"deepseek-v4-pro","scored_at":"2026-08-27T13:03:20","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-27","rank":9,"question":"医生在多大程度上会不加批判地接受LLM生成的虚构诊断建议，这种易感性是否随专科经验水平而变化？","design":"本研究不是用LLM模拟人类，而是用真实医生作为被试，通过人为污染LLM系统提示词使其在真实病例中强制输出虚构疾病，测量医生在LLM辅助鉴别诊断后是否将该虚构疾病纳入最终诊断。","baseline":"无对照，但按经验分层：≤6个月神经放射学培训的参与者与>6个月培训的参与者形成对比。","findings":"44%的参与者将虚构疾病纳入最终诊断；其中≤6个月培训者69%纳入，>6个月培训者0%纳入。表明经验不足的放射科医生易受LLM幻觉影响，出现“代理幻觉”现象。","reliability":"论文未讨论失效条件，但指出易感性仅限于专科经验有限的医生，提示需要结构化培训来批判性评估AI生成内容。","relevance":"该研究直接评估人类在LLM辅助决策中对幻觉的易感性，属于人类决策偏差评估，与研究者关注的人类仿真可靠性及偏差条件高度相关，值得阅读原文。","inspiration":"借鉴其通过污染LLM提示词施加处理、以真实专家行为作为结果变量的设计，可迁移到金融顾问或信贷审批场景中AI建议对决策的影响。｜例如，研究AI生成的虚假财务信息是否影响投资决策或信贷审批。｜可招募不同经验的金融从业者或普通投资者，使用LLM生成包含虚构风险因素的投资建议，测量其是否采纳该建议，并与真实历史决策数据或专家判断对照。"}},{"id":"2607.22188","version":2,"title":"Draining the Energy Commons: Self-Defeating Over-Appropriation as a Coordination Failure in Agentic LLM Collectives","zh_title":"耗尽能源公地：智能体 LLM 集体中作为协调失败的自我挫败式过度占用","abstract":"LLMs are increasingly deployed as agents that plan, use tools, and act over time. When they share persistent resources, such as compute pools or energy reserves, decisions by one agent affect the conditions faced by later agents. We study this coordination failure in a renewable energy commons. Four same-family GPT, Gemini, or Grok agents act in homogeneous self-play as electricity prosumers, instructed to maximize operational continuity. Holding aggregate residual demand and the decision protocol fixed, we vary the regeneration rate of a shared energy reserve from abundance to scarcity. All three families preserve the reserve when demand does not exceed peak renewable replacement, but over-appropriate it beyond that threshold (all nine exact scarcity contrasts survive Holm correction; largest adjusted p = 4.87e-5). The pattern is self-defeating: the same populations protect current service while undermining future service. At higher scarcity (rho = 1.2), early aggregate request pressure exceeds peak renewable replacement in every family and averages 1.21 times that level. Mean trajectories fall below the reserve level of maximum replenishment by rounds 5-7. Two offline benchmarks compare a social planner maximizing group-wide operational-service value with open access, where each prosumer maximizes its own value. At a discount factor of gamma = 0.95, both benchmarks sustain the reserve under the same dynamics. Realized depletion instead resembles outcomes under a more impatient open-access benchmark. The populations therefore behave like impatient optimizers at the level of the public trajectory. This system-level alignment failure would be missed by isolated-response evaluation.","authors":["Marcantonio Bracale Syrnikov","Federico Pierucci","Matteo Prandi","Marcello Galisai","Piercosma Bisconti","Francesco Giarrusso","Daniele Nardi"],"categories":["cs.MA"],"primary_category":"cs.MA","announce_type":"replace","date":"2026-08-27","first_seen":"2026-07-27","revised_at":"2026-08-27","abs_url":"https://arxiv.org/abs/2607.22188","pdf_url":"https://arxiv.org/pdf/2607.22188","source_feed":"cs.MA","score":6,"bucket":"other","rubric_hits":["D3"],"tags":["LLM 多智能体","公共资源困境","社会模拟"],"reason":"LLM agent 群体模拟公共资源困境，但无真实人类数据对照，属社会模拟边界…","model":"deepseek-v4-pro","scored_at":"2026-08-27T13:03:34","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":3,"question":"LLM代理群体在共享可再生资源中是否会出现协调失败，导致自我挫败的过度占用？","design":"使用GPT-5.4-mini、Gemini-3.1-flash-lite和Grok-4.3三种模型，每种模型四个同族代理作为电力产消者，在自博弈中运行，目标最大化自身运营连续性。通过改变共享能源储备的再生率（从充裕到稀缺）施加处理，测量储备水平、请求压力、备用能源使用等结果变量。","baseline":"无对照","findings":"所有三种模型在需求不超过峰值可再生替代时能维持储备，但在稀缺条件下过度占用储备，导致自我挫败的储备下降。早期聚合请求压力超过峰值可再生替代，平均轨迹在5-7轮后低于最大补充水平。","reliability":"论文未讨论","relevance":"高度相关，直接研究LLM代理在公共品博弈中的协调失败，涉及系统级对齐失败，但缺乏真实人类对照，适合关注仿真可靠性与偏差的研究者阅读原文。","inspiration":"该研究通过改变共享资源再生率来施加稀缺性处理，并测量代理群体的请求压力和储备水平，这种系统级压力测试设计值得借鉴｜可迁移到公共资源管理实验，如渔业配额分配或碳排放权交易中的集体决策问题｜使用LLM代理模拟渔民群体，处理变量为资源再生率（高/低），结果变量为捕捞总量和资源存量，对照真实渔场历史数据或实验室人类实验结果"}},{"id":"2608.24076","version":2,"title":"AgentWorld: Personality-Aware Reliability Evaluation for Agentic Information Retrieval","zh_title":"AgentWorld：面向智能体信息检索的人格感知可靠性评估","abstract":"Evaluation of agentic information retrieval remains limited to scripted interactions with uniform users, missing both natural personality diversity and adversarial brittleness. We present AgentWorld, a simulation framework combining (i)Big Five (OCEAN) personality-driven user populations with stateful tool-use environments; (ii)the pass$^k$ consistency metric with structured fault classification, partial-credit scoring, and dual-control handoff verification; (iii)score-thresholded training-data export in six fine-tuning formats; and (iv)an adversarial Risk Analyser that snapshots required-intermediate-state spines, branches Monte-Carlo rollouts under four task-aware perturbation types, and quantifies risk via $\\Delta P / \\Delta T$ scoring, Dempster--Shafer evidence fusion, and Shapley attack-category attribution. Three experiments demonstrate the framework: a conversational analytics agent across 10 OCEAN personas (240 evaluator judgments); a customer-support agent across 5 tasks $\\times$ 4 persona variants; and adversarial stress-testing of 5 tasks revealing pre-existing trajectory brittleness ($V_{\\min}=0.375$ without perturbation) and tool/infrastructure-layer attack dominance (Shapley: 46% system, 38% action). Personality variation surfaces failure modes uniform testing cannot expose---cross-domain leakage, contextual drift, a 0.27-point quality gap, and 50% vs. 100% pass-rate across personas on the same task---while the Risk Analyser quantifies trajectory-level brittleness that pass$^k$ alone cannot measure.","authors":["Gunja Agarwal","Arup Kumar Das","Arun Menon","Jitesh Chandra Mishra","Vignesh Divakaran"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-27","first_seen":"2026-08-26","revised_at":"2026-08-27","abs_url":"https://arxiv.org/abs/2608.24076","pdf_url":"https://arxiv.org/pdf/2608.24076","source_feed":"cs.AI","score":6,"bucket":"other","rubric_hits":["D2","D3"],"tags":["人格驱动仿真","智能体评估","可靠性测试"],"reason":"用Big Five人格驱动用户仿真，但目的是评估检索agent而非复现人类行为…","model":"deepseek-v4-pro","scored_at":"2026-08-27T13:03:35","error":null,"has_summary":false,"summary":null},{"id":"2608.25152","version":1,"title":"Belief Cascades Drive Persuasion in LLM Agent Networks","zh_title":"信念级联驱动 LLM 智能体网络中的说服","abstract":"Multi-agent LLM systems increasingly debate answers, coordinate research, simulate users, and mediate information flows, making agent-to-agent persuasion a basic but undermeasured capability. We introduce a controlled testbed for studying how goal-directed persuaders shift elicited stances in networks of LLM agents grounded in real-world ego-network topologies. Across four LLM backbones, five graphs, and 55 policy statements, we find that persuasion dynamics depend on the interaction between topology, competition, topic, and model prior. Additionally, we show that direct exposure reliably predicts next-round stance change in competing runs, and peer relays carry smaller but measurable influence, showing that agents not assigned to persuade can still transmit persuasive force. Finally, analyzing post text alone misses important movement: planned strategies are only partly realized in executed messages, action choices can diverge from message content, and persuadees rarely state the stance shifts detected by probes. These results argue for evaluating multi-agent persuasion as a trajectory- and exposure-level process, using belief probes, exposure provenance, and action logs to identify who influenced whom and whether visible language reflects underlying stance movement.","authors":["Haoyi Qiu","Genglin Liu","Pranav Narayanan Venkit","Kung-Hsiang Huang","Saadia Gabriel","Chien-Sheng Wu","Nanyun Peng"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-27","first_seen":"2026-08-27","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.25152","pdf_url":"https://arxiv.org/pdf/2608.25152","source_feed":"cs.CL","score":6,"bucket":"other","rubric_hits":["D3"],"tags":["LLM 多智能体","说服动态","社会模拟"],"reason":"LLM agent 网络中的说服动态模拟社会过程，但无真实人类数据对照，属边界…","model":"deepseek-v4-pro","scored_at":"2026-08-27T13:03:07","error":null,"has_summary":false,"summary":null},{"id":"2608.24896","version":1,"title":"Agentic World Analysis (AWA) - an alternative way to explore systems and support decision making","zh_title":"智能体世界分析（AWA）——探索系统和支持决策的另一种方式","abstract":"To address increasingly pressing sustainability challenges, various approaches have been developed to foresee possible futures, identify failure modes, detect vulnerabilities, and test potential mitigations. However, environmental systems are highly complex. Especially when coupled with human processes, the scale of uncertainties becomes intractable. To address this challenge, we propose a new approach - Agentic World Analysis (AWA)- combining the strengths of simulation modelling and expert elicitation. The concept of AWA is defined by three properties: 1) AWA uses an agentic AI system to mimic an expert panel that studies the world; 2) AWA projects futures iteratively through analysing scenario trees and learning from this analysis to improve decisions; 3) AWA is auditable. Based on these requirements, we implemented the World Engine by Generative Agents (WEGA) as a possible application of the AWA approach and demonstrated its functionality with a real-world case study: the Nitrogen Crisis in the Netherlands. WEGA autonomously constructed the context, identified key stakeholders and uncertainties, created expert agents, and generated future scenarios. As a result, two pathways from 2026 to 2041 were proposed, sharing a common assumption that social acceptance of nitrogen mitigation policies is low, while differing in how successful the restoration is according to the implementation of nitrogen data monitoring. The pathways are evaluated in multiple dimensions to assess their logical coherence and quality. The evaluation also actively exposes strengths and weaknesses to provide ways for testing the validity of the policies proposed. We discussed scaling up scenario analyses to enable massive pathway exploration, the trade-offs of using AWA and other approaches, and common concerns regarding AI systems.","authors":["Yongchao Zeng","Alexey Voinov","Calum Brown","Tatiana Filatova","Mark Rounsevell"],"categories":["cs.HC","cs.CE"],"primary_category":"cs.HC","announce_type":"new","date":"2026-08-27","first_seen":"2026-08-27","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.24896","pdf_url":"https://arxiv.org/pdf/2608.24896","source_feed":"cs.HC","score":6,"bucket":"other","rubric_hits":["D3"],"tags":["LLM社会模拟","情景分析","决策支持"],"reason":"用LLM agent模拟专家小组进行情景分析，属于社会模拟但无真实人类数据对照…","model":"deepseek-v4-pro","scored_at":"2026-08-27T13:03:07","error":null,"has_summary":false,"summary":null},{"id":"2608.25623","version":1,"title":"Using profiles of cognitive capability to assess AI suitability for workplace tasks","zh_title":"利用认知能力画像评估AI在工作任务中的适用性","abstract":"Organisations deploying AI face a scoping problem: which tasks can be automated, which should remain with humans, and which are best shared between the two. Aggregate benchmark scores provide little insight into where systems will succeed or fail in practice, while human judgements of model capabilities quickly become outdated. We introduce a pipeline that profiles agents and tasks using a shared set of core cognitive capabilities. Cognitive capability profiling infers an agent's capabilities from performance on a benchmark battery annotated for the cognitive demands of each item. Task requirements weighting elicits from domain experts the relative importance of these same capabilities for their work. As both use a common set of cognitive dimensions, they can be updated independently as models and roles change, and combined to estimate AI suitability at the level of a domain, organisation, role, or individual duty. We validate capability recovery on synthetic agents, profile six AI systems, and elicit task requirements from 410 employees across six occupational domains. AI systems differed more across cognitive dimensions than across model families, while workplace activities converged on a shared cognitive core. The resulting scores provide a comparative scoping tool for identifying promising candidates for piloting and areas where current systems are unlikely to be well suited. We discuss extending the framework to profile human workers alongside AI systems, moving from AI suitability towards human-machine task allocation.","authors":["Jonathan Prunty","Marko Te\\v{s}i\\'c","Patrick Quinn","Jos\\'e Hern\\'andez-Orallo","Lucy Cheke"],"categories":["cs.AI","cs.CY","cs.HC"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-08-27","first_seen":"2026-08-27","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.25623","pdf_url":"https://arxiv.org/pdf/2608.25623","source_feed":"cs.HC","score":6,"bucket":"other","rubric_hits":["D1","D2"],"tags":["AI任务分配","认知能力画像","人机协作"],"reason":"用LLM评估任务适配性，非仿真人类被试，但涉及人类数据对照和认知能力测量，属边…","model":"deepseek-v4-pro","scored_at":"2026-08-27T13:03:25","error":null,"has_summary":false,"summary":null},{"id":"2608.24001","version":2,"title":"Diverse by Reasoning: Harnessing the Wisdom of LLM Crowds for Future Prediction","zh_title":"通过推理实现多样性：利用LLM群体智慧进行未来预测","abstract":"Large language models (LLMs) are increasingly used for future prediction, motivating the use of multiple models as a wisdom-of-the-crowd mechanism. However, simply increasing crowd size does not guarantee effective diversity, as different LLMs may exhibit redundant behaviors. We propose a behavior-aware framework for constructing diverse LLM crowds. The framework characterizes models using their reasoning traces on independent development tasks, clusters models by behavioral similarity, and selects representatives for collective prediction. We evaluate 25 LLMs using seven development benchmarks for behavioral diversity modeling and two future-prediction benchmarks for evaluating diverse crowds' performance. Our results show that crowd composition can matter more than crowd size: a three-model medoid crowd based on K-means++ behavioral clustering outperforms conventional voting over all 25 models on both prediction benchmarks, while reducing model calls by 88% and inference cost by approximately 80%. The results further suggest that representative behavioral diversity, rather than simply maximizing diversity, is important for constructing effective LLM crowds","authors":["Nirupam Chetlapalli","Yiming Liao","Min-Chun Chen","Keke Chen"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-27","first_seen":"2026-08-26","revised_at":"2026-08-27","abs_url":"https://arxiv.org/abs/2608.24001","pdf_url":"https://arxiv.org/pdf/2608.24001","source_feed":"cs.AI","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["LLM群体智慧","未来预测","行为多样性"],"reason":"用LLM群体做预测，但无人类数据对照，属社会模拟边界情形","model":"deepseek-v4-pro","scored_at":"2026-08-27T13:03:34","error":null,"has_summary":false,"summary":null},{"id":"2608.25824","version":1,"title":"Localize-Then-Decide Guarantees for LLM Judgments","zh_title":"LLM判断的局部化后决策保证","abstract":"Large language models (LLMs) are increasingly used as evaluators to assess output quality and preference alignment, yet providing reliable guarantees of agreement with human judgments remains challenging. Recent work introduces confidence-thresholding methods that provide such guarantees for pairwise comparisons, relying on the assumption that higher estimated confidence implies lower disagreement risk with humans. However, this assumption can break down when the number of candidate responses increases, since distributing probability mass across many alternatives can distort confidence estimates. To address this issue, we propose a Localize-Then-Decide framework. First, conformal prediction localizes a small shortlist that contains the human-preferred response with high probability. Then, a calibrated confidence-based rule selectively chooses a single response from this shortlist or abstains. This design restores the monotonic relationship between confidence and disagreement risk and enables high-probability agreement guarantees. Experiments with multiple candidate sizes across several datasets and judge LLMs demonstrate that our framework consistently achieves higher guarantee success rates and substantially higher coverage than single-stage baselines.","authors":["Xinyu Li","Yi Zhou","Guanqun Cao","Zeyu Fu","Tianjin Huang","Gaojie Jin"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-27","first_seen":"2026-08-27","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.25824","pdf_url":"https://arxiv.org/pdf/2608.25824","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM评估","校准","人类判断对齐"],"reason":"LLM作为评估者替代人类判断，但目标是校准模型输出而非仿真人类被试，属于标注替…","model":"deepseek-v4-pro","scored_at":"2026-08-27T13:03:27","error":null,"has_summary":false,"summary":null},{"id":"2608.25854","version":1,"title":"Key Point Analysis Needs Structure Recovery: Task Definition, Dataset Diagnosis, and a Structure-Aware Benchmark","zh_title":"关键点分析需要结构恢复：任务定义、数据集诊断与结构感知基准","abstract":"Key Point Analysis (KPA) aims to identify a concise set of key points that summarize a collection of arguments together with their prevalence. We argue that KPA is fundamentally a structured prediction problem that requires recovering semantic groupings, generating representative key points, ensuring coverage, and estimating prevalence. Under this formulation, we show that existing KPA benchmarks suffer from limitations in grouping quality, redundancy, coverage, and argument-key point mappings, causing ceiling violation and selection failure in reference-based evaluation. To support future research on true KPA, we introduce a structure-aware, distribution-sensitive benchmark built via a human-in-the-loop re-annotation. Human and LLM evaluations consistently show that the resulting structures yield more coherent groupings, higher-quality key points, better coverage, and more reliable prevalence estimates than existing annotations. We further release several annotation resources to support research on KPA evaluation, argument-key point matching, explainable KPA, and LLM-as-a-judge methodologies, and outline a research agenda for true KPA.","authors":["Zhiqiang Shi","Oana Cocarascu"],"categories":["cs.CL","cs.LG"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-27","first_seen":"2026-08-27","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.25854","pdf_url":"https://arxiv.org/pdf/2608.25854","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["关键点分析","LLM评估","数据集标注"],"reason":"LLM作为评判者评估标注质量，替代人工评估，非仿真人类被试","model":"deepseek-v4-pro","scored_at":"2026-08-27T13:03:29","error":null,"has_summary":false,"summary":null},{"id":"2608.26081","version":1,"title":"SwarmWorld: Stigmergic technological evolution in societies of language-model agents","zh_title":"SwarmWorld：语言模型智能体社会中的共识主动性技术演化","abstract":"Collective intelligence can emerge when individuals coordinate through a shared environment, allowing local actions to accumulate into durable social organization. Language-model agents offer a new substrate for this process, yet most multi-agent systems rely on direct conversation, predefined roles, or centralized workflows. It remains unclear whether decentralized agents can build functional technologies and outperform independent search. Here, initially homogeneous LLM agents in SwarmWorld self-organize without assigned roles or recipes into evolving technological societies. Agents explore a spatial environment, process resources, test materials, construct persistent artifacts, and write executable controllers evaluated by a deterministic simulator under unseen disturbances after the agents are removed. SwarmWorld splits cognition from consequence: agents propose architectures and controllers within fixed action and material schemas, while the simulated world determines function. Shared societies develop broader, more resilient technological portfolios than a strong best-of-N isolated-search baseline, although isolated search remains competitive for the strongest artifact. Agents differentiate into exploration, construction, maintenance, and coordination behaviors, transitioning as the world matures. Technologies accumulate through collaborative construction, executable inheritance, and persistent agent-artifact networks, with most reuse beginning through physical observation rather than communication. Explicit cultural mechanisms amplify collaboration and organization, but functional benefits depend on outcome and timescale. Physical stigmergy alone supports capable societies, while interaction drives persistent technological ecologies rather than universally superior individual inventions.","authors":["Subhadeep Pal","Fiona Y. Wang","Markus J. Buehler"],"categories":["cs.AI","cond-mat.mtrl-sci","cs.CL"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-08-27","first_seen":"2026-08-27","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.26081","pdf_url":"https://arxiv.org/pdf/2608.26081","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["多智能体系统","社会模拟","集体智能"],"reason":"LLM智能体社会模拟，但无真实人类数据对照，属边界情形。","model":"deepseek-v4-pro","scored_at":"2026-08-27T13:03:31","error":null,"has_summary":false,"summary":null},{"id":"2608.24903","version":1,"title":"Evidence-Grounded Mapping of Multimodal Human Sensing Psychological Transdiagnostic Dimensions","zh_title":"基于证据的多模态人类感知心理跨诊断维度映射","abstract":"Mobile and wearable sensing enables longitudinal observation of behavior, yet translating these signals into meaningful mental health constructs remains difficult. We introduce a clinician-in-the-loop benchmark for evaluating whether large language models (LLMs) can generate evidence-grounded Brief Hierarchical Taxonomy of Psychopathology (B-HiTOP) item profiles from passive sensing, ecological momentary assessment (EMA), and questionnaire evidence. Using the Generalization of Longitudinal Behavior Modeling (GLOBEM) dataset, we construct 14,592 participant-day instances and align multimodal evidence to 29 B-HiTOP items across five spectra. Since GLOBEM lacks B-HiTOP responses, we evaluate evidence compatibility (C) rather than diagnostic accuracy, separating substantive predictions from abstentions when evidence is insufficient for item-level scoring. Two-stage prediction improves C for EMA and questionnaire evidence, but reduces C under passive sensing and combined evidence and produces more conservative score distributions across models, spectra, and evidence settings. Overall, semantic abstraction helps organize heterogeneous self-report evidence while becoming an information bottleneck for indirect behavioral sensing signals.","authors":["Xiyun Hu","Xiangyuan Xue","Yuting Lyu","Hanya Shao","Jingping Nie"],"categories":["cs.HC","cs.LG"],"primary_category":"cs.HC","announce_type":"new","date":"2026-08-27","first_seen":"2026-08-27","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.24903","pdf_url":"https://arxiv.org/pdf/2608.24903","source_feed":"cs.HC","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM标注","心理健康","多模态数据"],"reason":"LLM用于从多模态数据生成心理病理学条目，替代临床医生标注，而非仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-08-27T13:03:20","error":null,"has_summary":false,"summary":null},{"id":"2608.25267","version":1,"title":"Mitigating LLM sycophancy with RL-based fine-tuning: Bayesian Truth Serum approach","zh_title":"基于强化学习的微调缓解大语言模型谄媚性：贝叶斯真值血清方法","abstract":"Large language models (LLMs) frequently exhibit \\emph{sycophancy}: they adapt their answers to a user's stated beliefs or preferences instead of reporting what they hold to be true, which lowers factual accuracy and can amplify misinformation. This paper proposes a methodology for mitigating sycophancy that employs the Bayesian Truth Serum (BTS), a peer-prediction mechanism, as the reward in Group Relative Policy Optimization (GRPO) to fine-tune an LLM. BTS pays an answer for being \\emph{surprisingly common}, that is, more frequent among respondents than those respondents themselves predicted. We treat a group of responses from a model for one question as those respondents, so the reward is a function of the model's own outputs and fine-tuning needs neither labels nor preference annotations. We prove that in the large-group limit a sycophantic response earns strictly lower expected reward than an honest one. We also prove that if the entire group agrees in advance on a symmetric answering rule, it cannot earn a higher information score than under truthful reporting. On our true/false benchmark the reference model's answer-flip rate under user pressure decreases from 23% to 4%, and its accuracy under that pressure increases from 80% to 93%. Our reward outperforms SMART and is comparable to synthetic-data fine-tuning and to pinpoint tuning, all three of which train on labels. It spends considerably more compute in exchange, which makes it suitable when labeled data is scarce. Peer Truth Serum, which also pays a premium for a rare answer but elicits no prediction report, reproduces the effect. A peer-prediction reward computed inside a single GRPO group therefore reduces sycophancy without labels, and comparing mechanisms suggests that the premium paid for a rarer answer drives the effect.","authors":["Serhii Mytsyk","Yiming Zhang","Vikram Krishnamurthy"],"categories":["cs.LG","cs.SY","eess.SY"],"primary_category":"cs.LG","announce_type":"new","date":"2026-08-27","first_seen":"2026-08-27","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.25267","pdf_url":"https://arxiv.org/pdf/2608.25267","source_feed":"cs.LG","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["LLM谄媚性","强化学习微调","模型行为校准"],"reason":"论文测量并修正LLM自身的谄媚倾向，属于对模型行为的测量与校准，而非用LLM仿…","model":"deepseek-v4-pro","scored_at":"2026-08-27T13:03:10","error":null,"has_summary":false,"summary":null},{"id":"2608.25832","version":1,"title":"Skill Issue: Are Skills Language-Invariant in LLMs?","zh_title":"技能问题：LLM的技能是否跨语言不变？","abstract":"Large language models access knowledge inconsistently across languages, but to what extent do they differ in their skill sets when interacting with different languages? This work quantifies cross-lingual skill inconsistency orthogonally from knowledge and general benchmark performance. We do this via multilingual self-play: two instances of the same model compete in a text-based game, each interacting through a different language interface. Since the model, opponent, rules, state space, and available actions remain fixed, this setting isolates the effect of language on the model's realized behavior. We build a multilingual extension to TextArena and evaluate three open-weight models across eight languages and six games covering spatial reasoning, imperfect information, resource allocation, and repeated interaction. We find that the same model can exhibit markedly different playing strength across languages, with systematic variation in win--loss margins, invalid actions, and strategic tendencies. Detailed analyses reveal language-specific failures in spatial reasoning, card-conditioned decisions, and optimal move selection. In some settings, changing only the intermediate reasoning language recovers much of the lost performance, suggesting that language can affect different stages of the decision process. These results show that skill discrepancies are a measurable major roadblock in the development of truly multilingual models. Better understanding these discrepancies can help us design models that perform more equitably across languages.","authors":["Bobby Cheng","Adam Gaber","Zhengyuan Liu","Catherine Arnett","Omer Goldman","Cheston Tan","Leshem Choshen"],"categories":["cs.CL","cs.AI","cs.GT","cs.LG"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-27","first_seen":"2026-08-27","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.25832","pdf_url":"https://arxiv.org/pdf/2608.25832","source_feed":"cs.CL","score":3,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体","跨语言评估","文本游戏"],"reason":"多智能体自博弈评估跨语言技能，不涉及人类行为对照或仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-08-27T13:03:28","error":null,"has_summary":false,"summary":null},{"id":"2608.01666","version":3,"title":"Style Wins, Substance Loses: A Diagnosis of LLM-as-Judge in Idea Generation","zh_title":"风格胜，实质败：LLM作为评审在创意生成中的诊断","abstract":"However, whether these judges truly evaluate the scientific substance of ideas or are influenced by superficial stylistic presentation remains an open question. To address this question, we propose SciStyleBench, a unified three-component benchmark for diagnosing and mitigating stylistic bias in LLM-based idea evaluation: (i) First, SciStyleStage, a three-stage evaluation environment that applies controlled stylistic perturbations to fixed scientific content across three settings no context, fixed-domain context, and open-domain retrieval context, covering 600 scientific ideas and 15 style variants, with 9,000 evaluation instances per setting; (ii) Second, SciStyleMetrics, a set of quantitative measures, including Style Bias Index (SBI), Substance Recognition Rate (SRR), and Adversarial Win Rate (AWR), to characterize how stylistic variation affects scoring stability, substance discrimination, and ranking robustness; (iii) Third, SciStyleExtractor, a plug-and-play evaluation module that separates presentation style from scientific content by predicting style type and deviation before style-conditioned evaluation, enabling us to assess whether style awareness reduces stylistic bias. Experiments on SciStyleBench show that direct LLM judges remain sensitive to writing style and struggle to distinguish scientific substance. In contrast, SciStyleExtractor reduces SBI from 0.566 to 0.501 while increasing SRR and AWR from 0.504 and 0.554 to 0.759 and 0.899, respectively. These results suggest that robust idea evaluation requires invariance to stylistic variation without sacrificing sensitivity to scientific substance. Overall, SciStyleBench provides a systematic framework for identifying, quantifying, and mitigating stylistic bias in scientific idea evaluation.","authors":["Fengxian Ji","Yuke Li","Jingpu Yang","Juanfan Wu","Fan Zhang","Zhexuan Cui","Yu Xie","Min Peng","Qianqian Xie","Xiuying Chen","Zhuohan Xie"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-08-27","first_seen":"2026-08-04","revised_at":"2026-08-27","abs_url":"https://arxiv.org/abs/2608.01666","pdf_url":"https://arxiv.org/pdf/2608.01666","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["LLM评审","文体偏差","基准测试"],"reason":"研究LLM作为评审的文体偏差，属NLP评测，不以人类行为仿真为目标","model":"deepseek-v4-pro","scored_at":"2026-08-27T13:03:36","error":null,"has_summary":false,"summary":null},{"id":"2608.24901","version":1,"title":"Detection != Reliable Control: Decodable Empathy Directions Yield at Most Partial Shifts in Automated Empathy Scores","zh_title":"检测不等于可靠控制：可解码的共情方向至多导致自动共情分数的部分偏移","abstract":"A decodable \"empathy\" direction is routinely read as a causal lever, conflating decodability, automated-metric control, and human-perceived change. We test this for two EPITOME-derived facets -- Recognition (cognitive) and Resonance (affective) -- in three instruction-tuned LLMs, scoring every intervention with two LLM judges and a discriminative EPITOME classifier, each gated by an emotional-vs-neutral positive control. The control passes for the affective facet across all automated instruments, but cognitive range is inconsistent across them. Both facets remain decodable after residualizing against a sentence-embedding-derived surface score, and steering can substantially rewrite the text. Yet adding the Resonance direction raises the affective score only partially -- in Qwen by +0.29 (approximately 26% of the natural gap). A direct between-direction contrast confirms the shift is facet-specific in Qwen and Llama (not Gemma); we do not, however, establish a matching human-perceived change. Additive cognitive steering produces no measurable change, but a within-domain control shows the cognitive instrument is too coarse to resolve the differences such steering would produce -- unmeasurable, not a clean null. By contrast, Gemma Recognition ablation lowers the classifier's cognitive score even after adjusting for response length. Detection does not imply reliable control under global interventions, and cognitive-empathy claims warrant an explicit measurement-sensitivity check.","authors":["Haoran Jisun"],"categories":["cs.CL","cs.HC","cs.LG"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-27","first_seen":"2026-08-27","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.24901","pdf_url":"https://arxiv.org/pdf/2608.24901","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["LLM可控性","共情方向","模型行为分析"],"reason":"研究LLM共情方向的可控性，属模型行为分析，非人类仿真实验。","model":"deepseek-v4-pro","scored_at":"2026-08-27T13:03:19","error":null,"has_summary":false,"summary":null},{"id":"2608.25654","version":1,"title":"Unmatched Does Not Mean False: Incomplete Reference Sets Can Reverse Calibration Rankings in Open-Ended Theory-of-Mind Tracking","zh_title":"未匹配不等于错误：不完整参考集可逆转开放式心智理论追踪中的校准排名","abstract":"Open-ended Theory-of-Mind (ToM) trackers emit valid beliefs absent from finite references. A finite-reference-plus-matcher pipeline marks unmatched outputs false, creating proxy labels that can reverse proper-score model selection on fixed outputs. Holding 259 beliefs and paired scores fixed, reference recoding lowers weighted prevalence from 0.783 to 0.295 and reverses strictly proper Brier risk: a frozen source-prior rule leads native confidence by 0.227 under reference labels and trails by 0.152 under blinded adjudication, in all six authored scenarios. A reference-only Platt recalibrator reverses further. An ICE-specific reversal appears in a released 301-question NQ-open DPR-BERT pipeline: its average-confidence baseline improves instance-level calibration error by 0.045 under exact match but worsens it by 0.074 under human correctness, with both intervals excluding zero. On independently authored OpenToM narratives, 90-96% of audited unmatched beliefs are literally true and the paired direction again reverses. An exact decomposition attributes the distortion to omitted truths, and a closed-form criterion correctly classifies comparisons from twelve released systems. Frozen-audit retrospective replay shows 50 attempted annotations recover ranking direction with probability at least 0.996. TriSource-Restore anchors full-frame reference labels and frozen automatic judgments to a probability-sampled human pilot, maintains at least nominal coverage, narrows intervals, and repairs confidence subject to a base-rate deployment gate.","authors":["Zhexi Feng","Wuxi Chen","Bingrui Zhang"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-27","first_seen":"2026-08-27","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.25654","pdf_url":"https://arxiv.org/pdf/2608.25654","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["心智理论","校准评估","自然语言处理"],"reason":"论文评估ToM追踪器的校准，不涉及用LLM仿真人类被试或与人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-08-27T13:03:25","error":null,"has_summary":false,"summary":null},{"id":"2608.25660","version":1,"title":"Think-Probe-Respond: Improving Large Language Models as Judges of Research Idea Novelty","zh_title":"思考-探测-回应：改进大语言模型作为研究想法新颖性判断器","abstract":"Automated novelty judgment can accelerate scientific discovery by enabling efficient evaluation, refinement, and comparison of research ideas. While large language models are increasingly adopted for this task, we investigate a previously overlooked limitation in their judgment capabilities: despite generating reasoning rationales that closely mirror those of human experts, their final novelty judgments often diverge substantially. We demonstrate that this miscalibration stems from a systematic bias towards judging ideas as \"medium novel\". To mitigate this, we propose Think-Probe-Respond (TPR), a lightweight approach that probes latent novelty judgments from hidden states during the reasoning phase and uses the probed judgments to condition the final response. Across strong baselines, TPR improves novelty judgment performance by 22.30% and successfully mitigates the prevalent \"medium novelty\" bias.","authors":["Tim Schopf","Tobias Schreieder","Akiko Aizawa"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-27","first_seen":"2026-08-27","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.25660","pdf_url":"https://arxiv.org/pdf/2608.25660","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["LLM评测","新颖性判断","校准"],"reason":"论文研究LLM作为新颖性判断器，属于NLP能力评测，不以人类行为仿真为目标，无…","model":"deepseek-v4-pro","scored_at":"2026-08-27T13:03:26","error":null,"has_summary":false,"summary":null},{"id":"2608.25869","version":1,"title":"Anchoring Bias in LLM-as-a-Judge Systems: Prior Scores Compromise Evaluation Independence","zh_title":"LLM作为评判者系统中的锚定偏差：先前分数损害评估独立性","abstract":"Large language models (LLMs) increasingly assess generated content, giving rise to the LLM-as-a-Judge paradigm. These systems now score outputs, filter content, and gate iterative refinement in production pipelines, where each judgment is often assumed to be independent of earlier evaluations. We test this assumption using three prompt conditions: no metadata, revision framing, and anchored metadata containing revision, attempt, and prior-score fields. We show that prior scores, even when included only as context metadata, anchor judgments and systematically shift ratings toward their values. Across 192,000 attempted evaluations (185,271 successful), seven out of the eight evaluated models have 95% task-stratified bootstrap intervals below zero for the total anchored-metadata effect on 20 fixed texts. Cohen's $d$, a standardized measure of the difference between score distributions, reaches an absolute value of 0.71. Token-level analysis of selected model-task probes suggests a threshold-like response pattern: introducing anchored metadata produces a marked redistribution of output-score probabilities, while changing the anchor value within the tested below-threshold range produces comparatively little additional variation. On categorical industry data with human-labeled ground truth, anchored metadata blocks 48% of error corrections and flips 10.18% of correct judgments toward an assigned wrong label, demonstrating the bias extends beyond numerical scoring to categorical decisions. Neither Chain-of-Thought nor a metadata-disregard warning reduces the total effect, although the warning improves the paired accuracy effect relative to baseline in the industry experiment. Reliable LLM evaluation demands careful context engineering rather than an assumption of impartiality. Effective mitigation must be validated for the intended model and task or domain.","authors":["Ante Kapetanovic","Kemal Altwlkany","Andro Mercep","Tomislav Duricic","Emanuel Lacic"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-27","first_seen":"2026-08-27","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.25869","pdf_url":"https://arxiv.org/pdf/2608.25869","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["LLM评估","锚定偏差","评判者偏差"],"reason":"研究LLM作为评判者的锚定偏差，属于LLM评估可靠性，非仿真人类被试，无人类行…","model":"deepseek-v4-pro","scored_at":"2026-08-27T13:03:12","error":null,"has_summary":false,"summary":null},{"id":"2608.25325","version":1,"title":"FinRiskAtlas: Decision-Aligned Evaluation of Large Language Models for Financial Risk Review","zh_title":"FinRiskAtlas：面向金融风险审查的大语言模型决策对齐评估","abstract":"Deploying large language models for professional financial review requires more than measuring general financial competence: models must perform the specific review operation required by a workflow and determine whether available evidence is sufficient for a defensible decision. Existing financial benchmarks cover knowledge, reasoning, compliance, and professional tasks, but their evaluation units are often organized around datasets or task formulations rather than the decisions that deployed systems support. We introduce FinRiskAtlas, a Chinese-language benchmark that evaluates financial LLMs along two complementary dimensions: operation execution under fixed evidence states and evidence-state control under evolving review conditions. The static benchmark contains 9,742 instances across 53 task families, including 42 Domain Knowledge families and eleven downstream review operations defined by explicit evaluation contracts. FinRisk-Ask extends this framework through offline replay of 680 pre-action states from 104 de-identified professional trajectories, withholding future evidence during inference and using it only to construct expert-verified evidence targets. Across 33 model configurations, operation-level evaluation yields non-redundant rankings (mean pairwise Spearman correlation 0.42 across downstream operations), and knowledge-based shortlisting can incur up to 18.01 points of regret on individual operations. FinRisk-Ask further shows that entering the Ask branch more frequently does not necessarily improve request targeting or end-to-end evidence acquisition. These results show that broad financial capability scores do not fully capture where models are reliable in professional workflows, motivating evaluation units aligned with the decisions and evidence states that deployed systems must support.","authors":["Suyang Zhong","Jingzhe Zhu","Qi Xu","Liyao Sun","Yin Wang","Qingqing Sun","Shuai Chen","Tianyi Zhang"],"categories":["cs.AI","cs.CL"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-08-27","first_seen":"2026-08-27","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.25325","pdf_url":"https://arxiv.org/pdf/2608.25325","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["LLM评测","金融风险审查","基准测试"],"reason":"该论文是金融领域LLM能力评测基准，不涉及用LLM仿真人类被试或与真实人类行为…","model":"deepseek-v4-pro","scored_at":"2026-08-27T13:03:23","error":null,"has_summary":false,"summary":null},{"id":"2608.24899","version":1,"title":"aipsy-judge: A Specialized, Psychologist-Corrected Local Judge for the Psychological Safety of Conversational AI","zh_title":"aipsy-judge：面向对话式AI心理安全的专用、经心理学家校正的本地裁判模型","abstract":"The standard recipe for LLM-as-judge -- pick a frontier model, or average several -- is actively unsafe for grading the psychological safety of conversational AI. Using aipsy-bench, an open frozen safety instrument, we run a fully-crossed competence study: three frontier models (gpt-5.4-mini, claude-sonnet-4-6, gemini-2.5-flash) serve as both generators and judges of 3,000 mental-health, companion, and coaching messages against a psychologist's ratings. The disagreement is not noise: it is structured, concentrated on the safety-critical metrics, and one judge (Gemini) is an outlier -- the most lenient, carrying a +0.99 self-preference premium, flagging far fewer tail failures, and scoring a means-in-hand self-harm response \"exemplary.\" Inter-judge agreement on empathy, where sycophancy hides, is the lowest in the battery (alpha 0.24). One axis stands apart: the binary crisis-detection flag is the one safety-critical signal judges agree on (alpha 0.80), erring toward over-flagging, the safe direction for a triage screen. Equal-weight averaging, the canonical fix, blends that leniency and tail-blindness into the safety score. Off-the-shelf open-weight judges are worse for a dispositional, not capability, reason -- and disposition is fine-tunable. We therefore distill a per-metric, psychologist-corrected target into a small, frozen, local model, aipsy-judge-1.0, an Apache-2.0 fine-tune of Gemma-4-26B-A4B. aipsy-judge-1.0 tracks the corrected target better than its base on the composite (ICC 0.64 to 0.75) and crisis detection (kappa 0.65 to 0.82), catches 92% of crises with a false-positive lean, and grades more faithfully than any single frontier judge, while every transcript stays on the machine. These are directional readings against a single-expert-informed target, not validated multi-rater agreement. A safety grader that shares a vendor's post-training shares its blind spots.","authors":["Michael Keeman","Anastasia Keeman"],"categories":["cs.HC","cs.AI","cs.CY"],"primary_category":"cs.HC","announce_type":"new","date":"2026-08-27","first_seen":"2026-08-27","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.24899","pdf_url":"https://arxiv.org/pdf/2608.24899","source_feed":"cs.HC","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["LLM裁判","心理安全","对话AI"],"reason":"论文研究LLM作为心理安全评估裁判，属于角色扮演对话安全评测，无人类被试仿真或…","model":"deepseek-v4-pro","scored_at":"2026-08-27T13:03:18","error":null,"has_summary":false,"summary":null},{"id":"2608.24900","version":1,"title":"Stronger Alignment between Brain Activity and LLM Embeddings during Code Writing compared to Prose Writing","zh_title":"代码写作相比散文写作时大脑活动与LLM嵌入的对齐更强","abstract":"Programming is a critical skill underlying modern software systems, yet the cognitive processes supporting code writing are only beginning to be understood, limiting educational practices and developer tools. At the same time, Large Language Models (LLMs) are increasingly used to assist programming. These models themselves are not well understood and can exhibit undesirable behavior like introducing security vulnerabilities. Given evidence that some cognitive representations may be shared between LLMs and the brain, we seek to improve our understanding on both fronts by relating these two systems to one another. We used Voxelwise Encoding Models (VEMs) to relate LLM embeddings to brain activity measured with functional Magnetic Resonance Imaging (fMRI) during naturalistic writing tasks. Using participants' (n = 23) keystrokes as prompts, we extracted LLM embeddings to predict voxelwise Blood Oxygen Level Dependent (BOLD) signal, quantifying alignment as the correlation between predicted and recorded signal. To assess whether this alignment is specific to programming or generalizes to other generative processes, we compared code writing to prose writing. Alignment was strongest in the right frontal pole, and brain activity was significantly better predicted by LLM embeddings during code writing than prose writing (p < 0.001, FDR-corrected). Within participants, the best-modeled voxel locations for code writing were 66% consistent across LLM layers but varied substantially between participants (39% similarity). Our findings suggest stronger alignment between human and LLM representations during structured code generation, with implications for designing AI systems that predict code generation but support natural language tasks.","authors":["Zachary Karas","Catie Chang","Kevin Leach","Yu Huang"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-08-27","first_seen":"2026-08-27","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.24900","pdf_url":"https://arxiv.org/pdf/2608.24900","source_feed":"cs.HC","score":2,"bucket":"other","rubric_hits":["C2"],"tags":["脑机对齐","神经编码","LLM表征"],"reason":"研究LLM嵌入与脑活动对齐，不涉及用LLM仿真人类被试或行为对照","model":"deepseek-v4-pro","scored_at":"2026-08-27T13:03:18","error":null,"has_summary":false,"summary":null},{"id":"2608.24914","version":1,"title":"AI-Ready Research Workflows in Computational Social Science: Lessons on Building a Shared Language for Interdisciplinary Collaboration","zh_title":"计算社会科学中的AI就绪研究工作流：跨学科协作共享语言构建的经验教训","abstract":"Artificial intelligence (AI) is gaining traction in the social sciences and humanities (SSH). However, adoption remains limited by technical barriers to high-performance computing (HPC), validation processes that lag behind AI's rapid progress, and reproducibility standards that most SSH teams cannot meet. Research workflows--common in the life sciences--address these problems via encoding and abstracting technical complexity into repeatable routines; yet, accounts of how to build them in SSH remain scarce. We report on a two-year effort to build a workflow that enables a Science and Technology Studies unit to query, analyze, and enrich OpenAlex--a database of some 460 million scholarly records--on the MareNostrum supercomputer, using methods ranging from large-scale bibliometrics to LLM-based classification. We found the main challenge was translating domain-specific research questions into engineering requirements -- bridging two distinct methodological languages, with implications that were both organizational and technical. Organizationally, it meant adopting and adapting Agile to the research rhythm and pace, and reframing collaboration from a service arrangement to a co-design process. Technically, model-driven engineering was as valuable for collaboration as it was for automation; co-building the model facilitated both the creation of a shared vocabulary and the abstraction of HPC complexity. Finally, we highlight limitations we found in validation, reproducibility, and FAIR metadata -- beyond what any single project can sustain -- calling for coordinated, cross-institutional investment in the tooling and standards needed for AI-ready SSH workflows sustainable at scale.","authors":["Joan Giner-Miguelez","Alexandra M\\'alaga","Felipe G\\'omez-Cort\\'es","Adrian Carrascosa","Mariona Coll-Ardanuy","Andr\\'es F. Castro-Torres","Ra\\\"ul Sirvent","Rosa M. Badia","Clara Guasch","Merc\\`e Crosas"],"categories":["cs.HC","cs.CY"],"primary_category":"cs.HC","announce_type":"new","date":"2026-08-27","first_seen":"2026-08-27","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.24914","pdf_url":"https://arxiv.org/pdf/2608.24914","source_feed":"cs.HC","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["AI工作流","计算社会科学","LLM分类"],"reason":"论文聚焦于构建AI就绪的研究工作流，使用LLM进行分类，属于计算社会科学基础设…","model":"deepseek-v4-pro","scored_at":"2026-08-27T13:03:21","error":null,"has_summary":false,"summary":null},{"id":"2608.25316","version":1,"title":"AVI-Personality: A Trait-Activated Multimodal Dataset for Personality and Competency Assessment in Asynchronous Video Interviews","zh_title":"AVI-Personality：异步视频面试中人格与胜任力评估的特质激活多模态数据集","abstract":"With the rapid development of AI-based personality and job-related competency assessment, Asynchronous Video Interviews (AVIs) are increasingly used in recruitment. However, existing multimodal personality datasets are often based on short, task-free social media videos and crowdsourced apparent personality labels, which limits their construct validity and relevance to structured interview assessment. To address these limitations, we introduce AVI-Personality, a trait-activated multimodal dataset for personality and job-related competency assessment from AVIs. The dataset contains 3,876 interview videos from 646 participants who completed a simulated management traineeship application. Participants answered two generic questions and four personality-targeted questions designed according to Trait Activation Theory. Our dataset provides both self and observer-reported HEXACO personality traits and job-related competency. We validate AVI-Personality through reliability, construct validity, internal nomological association, fairness, and benchmark analyses. Validation results show that the observer-rated personality traits have moderate to high reliability, especially when ratings are based on personality-targeted questions. Benchmark results show that text-based AI algorithms provide strong personality-relevant cues, while multimodal methods achieve the best overall performance but only modestly outperform text-based baselines. In general, AVI-Personality provides a psychometrically grounded dataset for developing and evaluating AI-based models for personality and competency assessment. The dataset is available are released at https://github.com/APAL-SEU/AVI6","authors":["Tianyi Zhang","Jinwenxi Shang","Antonis Koutsoumpis","Yuan Zong","Reinout E. de Vries","Wenming Zheng"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-08-27","first_seen":"2026-08-27","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.25316","pdf_url":"https://arxiv.org/pdf/2608.25316","source_feed":"cs.HC","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["人格评估","多模态数据集","视频面试"],"reason":"该论文构建人类面试视频数据集用于AI人格评估，不涉及用LLM仿真人类被试，属于…","model":"deepseek-v4-pro","scored_at":"2026-08-27T13:03:23","error":null,"has_summary":false,"summary":null},{"id":"2608.25871","version":1,"title":"CEDAR: Controlled and Event-Driven Demand Forecasting via Residual Decomposition","zh_title":"CEDAR：通过残差分解实现受控且事件驱动的需求预测","abstract":"Forecasting in large-scale e-commerce marketplaces is increasingly required to support planning: merchants need to evaluate sales outcomes under future action sequences such as budget schedules, rather than passively predicting what happens next. However, most existing time series forecasting (TSF) approaches remain inherently passive. Even when incorporating operational decisions as auxiliary covariates, they typically optimize for correlation-based extrapolation under historical policies. This design suffers from autoregressive inertia and conflates endogenous market evolution with decision-induced transitions, leading to policy-insensitive rollouts and unreliable counterfactual analysis. To bridge this gap, we propose CEDAR (Controlled and Event-Driven Demand forecasting via Action-aware Residual decomposition), a two-stage framework for robust decision-conditioned simulation. In Stage I, an Action-Interleaved Transformer learns controllable action-conditioned state transitions for rollout under planned interventions. In Stage II, a Residual Correction Module leverages external event signals and LLM-assisted text representations to align noisy event descriptions with product context and correct event-driven deviations. Our study is enabled by a large-scale real-world dataset from Alibaba 1688, comprising approximately 32 million product trajectories with paired state-action sequences and aligned event signals. Extensive offline experiments and online controlled experiments in production demonstrate that CEDAR consistently improves simulation accuracy over strong TSF baselines and delivers practical gains for real-world budget planning.","authors":["Junjie Meng","Ranxu Zhang","Zi-an Zhang","Shujun Liu","Xiaoning Qi","Xiaozhou Xu","Yanyong Zhang","Hui Xiong","Chao Wang"],"categories":["cs.LG"],"primary_category":"cs.LG","announce_type":"new","date":"2026-08-27","first_seen":"2026-08-27","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.25871","pdf_url":"https://arxiv.org/pdf/2608.25871","source_feed":"cs.LG","score":2,"bucket":"other","rubric_hits":["C2"],"tags":["时间序列预测","电商需求预测","LLM辅助表示"],"reason":"论文做电商需求预测，用LLM辅助文本表示，非人类仿真实验，无人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-08-27T13:03:29","error":null,"has_summary":false,"summary":null},{"id":"2606.07392","version":2,"title":"Online Pandora's Box for Contextual LLM Cascading","zh_title":"面向上下文LLM级联的在线潘多拉魔盒","abstract":"Motivated by Large Language Model (LLM) cascading, we propose an online contextual Pandora's Box model for adaptively querying and selecting LLM APIs. In each period, a decision-maker observes a request context and faces a two-phase decision problem. In the query phase, the decision-maker sequentially queries APIs, where each query reveals a generated output and the decision-maker incurs an (output-dependent) cost. In the selection phase, the decision-maker selects one of the generated outputs to deploy and observes only the downstream reward of the deployed output. This output-mediated feedback structure differs from classical online contextual Pandora's Box models, in which opening a box directly reveals its reward. Rather than estimating the full conditional output and cost distributions of each API, we directly model the reservation index and develop a learning approach for the query phase. Specifically, we impose a parametric structure on the contextual reservation index functions induced by the classical Weitzman's policy. Our policy combines generalized method of moments (GMM) type estimation of these reservation indices with UCB-style confidence bounds for both these indices and the shared output-level reward evaluator. Under regularity conditions, we prove that the resulting policy achieves dimension-dependent $\\widetilde O(\\sqrt T)$ cumulative regret over a horizon of $T$ periods.","authors":["Alexandre Belloni","Yan Chen","Yehua Wei"],"categories":["cs.AI","cs.LG","econ.EM","stat.ML"],"primary_category":"cs.AI","announce_type":"replace-cross","date":"2026-08-27","first_seen":"2026-06-05","revised_at":"2026-08-27","abs_url":"https://arxiv.org/abs/2606.07392","pdf_url":"https://arxiv.org/pdf/2606.07392","source_feed":"cs.LG","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["LLM级联","在线学习","决策优化"],"reason":"研究LLM API级联选择策略，属多智能体协作优化，不涉及人类行为仿真或对照。","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:58:49","error":null,"has_summary":false,"summary":null},{"id":"2608.00991","version":2,"title":"SCHEDBench: A Benchmark for Evaluating LLM Constraint Faithfulness in Natural-Language Combinatorial Scheduling","zh_title":"SCHEDBench：评估LLM在自然语言组合调度中约束忠实度的基准","abstract":"This paper introduces SCHEDBench, a natural-language benchmark for evaluating combinatorial scheduling constraint faithfulness under surface-form variation. Grounded in canonical scheduling instances and solver-derived feasibility and optimality, SCHEDBench assesses whether large language models (LLMs) generate schedules with the same constraint-feasible behavior across varied natural-language (NL) surface forms. SCHEDBench spans 1,132 instances across job-shop scheduling problems (JSP), single and multi-mode resource-constrained project scheduling problems (RCPSP), nurse rostering/scheduling, and curriculum timetabling problems of varying difficulty. Instances are templated into natural language problems using domain-specific templates, themed entities, lexical-syntactic template rephrasing, and constraint-level surface-form variation, with reference solutions verified for feasibility and objective optimality. Across thirteen frontier and open-weight LLMs, we find that models are not reliably invariant to semantically equivalent renderings of the same scheduling problem. Surface-form variation reduces feasibility and induces above-noise shifts in per-instance hard-constraint violations on matched instances. Among the tested isolated axes, constraint reordering yields the clearest above-noise sensitivity.","authors":["Shrenil Shaun Sharma","Avi Sharma"],"categories":["cs.AI","cs.CL"],"primary_category":"cs.AI","announce_type":"replace-cross","date":"2026-08-27","first_seen":"2026-08-04","revised_at":"2026-08-27","abs_url":"https://arxiv.org/abs/2608.00991","pdf_url":"https://arxiv.org/pdf/2608.00991","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["LLM评测","调度问题","约束遵循"],"reason":"评估LLM在调度问题上的约束遵循，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-08-27T13:03:35","error":null,"has_summary":false,"summary":null},{"id":"2608.24662","version":2,"title":"The Invisible Editorial Layer: Formalizing Undisclosed Inference-Time Steering, Probability Placement, and the Attribution Problem in Deployed Language Models","zh_title":"隐形编辑层：形式化部署语言模型中未公开的推理时引导、概率放置与归因问题","abstract":"Evaluations of generative language models frequently interpret observable behavioral traits, such as political stance, brand inclination, and normative framing, as manifestations of model weights, post-training alignment, or prompting. This interpretation risks conflating a foundation model with the multi-layered production system through which its outputs are ultimately served. Modern inference stacks support runtime interventions capable of modifying generation while model parameters remain frozen. We examine inference-time framing bias: systematic runtime steering of generated text toward institutional, ideological, or commercial frames without requiring changes to the underlying model parameters. We formalize the Inference Attribution Problem and establish an observational non-identifiability result showing that, under black-box observation alone, behaviorally equivalent deployed systems may arise from structurally distinct combinations of model parameters and inference policies. Consequently, observed behavioral bias does not uniquely identify the architectural layer responsible for it. We further characterize Probability Placement as a deployment pattern in which undisclosed commercial influence is embedded within an ostensibly organic assistant response through systematic probability-mass reallocation, distinguishing it from explicit token-auction mechanisms for generative advertising. Finally, we discuss implications for behavioral auditing, inference provenance, confidential computing, cryptographic attestation, the EU AI Act, the Digital Services Act, and advertising-disclosure principles. We argue that governance of generative systems must increasingly distinguish between auditing a model and auditing the deployed system that ultimately speaks.","authors":["Augusto Camargo"],"categories":["cs.AI","cs.CL","cs.CY"],"primary_category":"cs.AI","announce_type":"replace-cross","date":"2026-08-27","first_seen":"2026-08-26","revised_at":"2026-08-27","abs_url":"https://arxiv.org/abs/2608.24662","pdf_url":"https://arxiv.org/pdf/2608.24662","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["推理时干预","模型治理","归因问题"],"reason":"论文讨论推理时干预与治理，不涉及用LLM仿真人类被试或与人类数据对照","model":"deepseek-v4-pro","scored_at":"2026-08-26T13:02:36","error":null,"has_summary":false,"summary":null},{"id":"2608.24952","version":1,"title":"The Dialect Tax: Dialectal Biases Persist throughout the Language Modeling Pipeline","zh_title":"方言税：方言偏见贯穿语言建模全流程","abstract":"Systematic dialectal performance gaps in language models (LMs) are well documented, but the source of these disparities within the modern language modeling pipeline remains unclear. Our study traces this \"dialect tax\" across the natural language processing pipeline. Using parallel English dialect corpora that hold meaning fixed while varying surface form, we first confirm that LMs recognize matched Standard American English (SAE) and dialectal texts as semantically equivalent. However, we discover further representational gaps corresponding to downstream performance gaps. Across model families and generations, modern LMs still encode dialectal texts unequally during tokenization, pre-training, post-training, and inference. Strikingly, bypassing traditional subword segmentation via a character-level counterfactual tokenizer removes neither input and output asymmetries nor dialectal accuracy gaps. During pre-training, dialect pairs induce more divergent gradient updates than pairs of entirely unrelated SAE documents, indicating that models find semantically equivalent dialectal content harder to learn from than unrelated SAE documents. During post-training, reward models show contextual, unstable dialect preferences, assigning higher values to isolated AAVE-exclusive tokens than to SAE-exclusive tokens, while full reasoning contexts receive task- and model-dependent dialect penalties. Overall, our findings suggest that the dialect tax is encoded and accumulated not by any one step in isolation, but at every step of the language modeling process.","authors":["Elle"],"categories":["cs.CL","cs.AI","cs.LG"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-27","first_seen":"2026-08-27","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.24952","pdf_url":"https://arxiv.org/pdf/2608.24952","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["方言偏见","语言模型公平性","NLP评测"],"reason":"研究语言模型中的方言偏见，属于NLP公平性评测，不以人类行为仿真为目标","model":"deepseek-v4-pro","scored_at":"2026-08-27T13:03:21","error":null,"has_summary":false,"summary":null},{"id":"2608.25717","version":1,"title":"When RAG Fails to Equalize: Geo-bias in Factual Question Answering over Public Companies","zh_title":"当RAG无法实现均衡：上市公司事实问答中的地理偏差","abstract":"Retrieval-augmented generation (RAG) is widely assumed to mitigate factual errors in large language models (LLMs), but it remains unclear whether retrieval uniformly compensates for missing knowledge. We study this question in a controlled factual QA setting over public companies, constructing a benchmark of approximately 2,000 firms across global equity indices. We evaluate six LLMs on four atomic attributes under four conditions: no-context, perfect context, misleading context, and distraction context. We find strong geographic disparities in no-context accuracy, indicating uneven parametric knowledge. While perfect context improves performance, it does not eliminate these gaps: gains are correlated with baseline accuracy, suggesting retrieval effectiveness is coupled to internal representations. Under misleading context, models frequently copy incorrect information. Larger models improve overall performance but do not remove these structural effects. These results challenge the view of RAG as a universal corrective and highlight the interaction between model knowledge, context quality, and entity representation.","authors":["Abhinav Havaldar","Enrico Santus"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-27","first_seen":"2026-08-27","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.25717","pdf_url":"https://arxiv.org/pdf/2608.25717","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["RAG","地理偏差","事实问答"],"reason":"纯NLP能力评测，无人类行为对照，不涉及仿真人类被试","model":"deepseek-v4-pro","scored_at":"2026-08-27T13:03:27","error":null,"has_summary":false,"summary":null},{"id":"2608.26089","version":1,"title":"From Producing to Validating: How AI Is Deskilling Freelancers","zh_title":"从生产到验证：AI如何使自由职业者去技能化","abstract":"Generative AI is promoted as a way to enhance knowledge work, yet its benefits and drawbacks fall unevenly across the workforce. Freelance and gig workers, who commonly lack the upskilling pathways available to traditional employees, face heightened risks to both skill development and job security as AI adoption advances. We review empirical evidence on AI's impact on knowledge-worker workflows and upskilling, then predict the primary and downstream effects of AI adoption among clients and workers in the freelance economy. We anchor this in two cases of the same shift, machine-translation post-editing and software development. We argue that freelancers are the leading edge of a change that also reaches salaried HCI practitioners, and we close with questions for the platforms and clients that mediate this work, and for HCI researchers.","authors":["Nakul Rajpal"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-08-27","first_seen":"2026-08-27","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.26089","pdf_url":"https://arxiv.org/pdf/2608.26089","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["AI对劳动力影响","自由职业","去技能化"],"reason":"论文讨论AI对自由职业者技能的影响，不涉及LLM仿真人类被试或行为对照。","model":"deepseek-v4-pro","scored_at":"2026-08-27T13:03:33","error":null,"has_summary":false,"summary":null},{"id":"2608.25063","version":1,"title":"The AI Adaptation Gap in Higher Education: Students, Faculty, and Administrative Staff","zh_title":"高等教育中的AI适应差距：学生、教师与行政人员","abstract":"The purpose of this study was to analyze patterns of artificial intelligence (AI) use and attitudes toward AI among students, faculty, and administrative staff at a large university specializing in teacher education. The analytical sample comprised 1809 students, 250 faculty members, and 62 administrative staff members (N = 2121). Three role-adapted 75-item questionnaires covered the frequency and contexts of AI use, perceived usefulness, trust and control, academic integrity concerns, responsible-use norms, institutional policy clarity, and perceived improvement in output quality. Data analysis included descriptive statistics, Welch group comparisons, pooled ordinary least squares (OLS) models, reliability and dimensionality checks for observed indices, and exploratory student-only K-means clustering. The results revealed a pronounced AI adaptation gap across university groups. Students reported higher current AI-use intensity and perceived usefulness than faculty and administrative staff, whereas faculty and administrative staff reported stronger academic integrity concerns and greater endorsement of responsible-use norms. In the pooled OLS trust model, perceived usefulness had the strongest standardized positive association with trust in AI (beta = 0.402); institutional policy clarity also had a positive but weaker association (beta = 0.223). Students reported higher perceived policy clarity than faculty, while neither group differed significantly from administrative staff. Exploratory clustering indicated heterogeneity among students in experience, competence, usefulness, trust, and control, but did not establish a latent typology across university groups. The cross-sectional, self-reported data show associations and group differences rather than causal effects on learning or objective outcomes.","authors":["Yuriy S. Braun","Salavat M. Khafizov"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2026-08-27","first_seen":"2026-08-27","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.25063","pdf_url":"https://arxiv.org/pdf/2608.25063","source_feed":"cs.CY","score":0,"bucket":"other","rubric_hits":["C5"],"tags":["AI使用调查","高等教育","人类态度"],"reason":"研究人类对AI的态度，非用LLM仿真人类被试","model":"deepseek-v4-pro","scored_at":"2026-08-27T13:03:22","error":null,"has_summary":false,"summary":null},{"id":"2608.26075","version":1,"title":"Epistemic Networks, Collective Misperception, and the Manipulation of Social Knowledge","zh_title":"认知网络、集体误解与社会知识的操纵","abstract":"We investigate the structure of interactive beliefs in networks: the epistemic state in which agents hold, revise, and act on their models of the epistemic states of other agents. What a group believes depends on what each member agent takes the others to believe, and on what each takes the others to believe about still others. We posit that the proper unit of social-epistemic analysis is not the individual belief but the tensor of mutual attribution, the array that records what every agent takes every agent to believe. We separate three layers of this object: what is privately held, what is publicly expressed, and what is to be believed by others. Social belief evolves by contraction of the tensor against a signed matrix of epistemic influence. Collective misperception decomposes exactly into a component that observation dissolves and a component that observation cannot touch. Social conformity amplifies pluralistic ignorance without generating it. The stable forms of collective epistemic consensus, polarization, entrenched misperception, and instability are spectral regimes of a single operator. Higher-order belief reduces to walks in the epistemic network, under a stated assumption of cognitive consistency whose boundary we mark precisely.","authors":["Mihnea C. Moldoveanu","Joel A. C. Baum"],"categories":["cs.SI"],"primary_category":"cs.SI","announce_type":"new","date":"2026-08-27","first_seen":"2026-08-27","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.26075","pdf_url":"https://arxiv.org/pdf/2608.26075","source_feed":"cs.SI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["社会网络","多智能体系统","理论模型"],"reason":"纯多智能体社会网络理论模型，无LLM仿真，无人类数据对照","model":"deepseek-v4-pro","scored_at":"2026-08-27T13:03:31","error":null,"has_summary":false,"summary":null},{"id":"2608.25678","version":1,"title":"Normative boundaries of AI in scientific work: Evidence from PhD researchers","zh_title":"人工智能在科学工作中的规范边界：来自博士研究生的证据","abstract":"Artificial intelligence (AI) is increasingly embedded in scientific work, but researchers may not evaluate its use uniformly across research tasks. This study examines task-specific attitudes towards AI among an international, self-selected sample of 3,785 PhD students in STEM and medical and health sciences who participated in Nature's Graduate Survey 2025. We analyse respondents' comfort with using AI for writing a research article, collecting and analysing data, designing experiments, tracking scientific literature, and summarising it. Latent class analysis identifies four distinct attitudinal profiles. The dominant profile reflects a \"division of labour,\" in which AI is widely accepted for literature-related tasks but resisted in activities closely associated with intellectual contribution, such as writing, data analysis, and experimental design. A \"status quo\" profile is broadly uncomfortable across tasks, an \"all-purpose\" profile is broadly comfortable, and an \"undecided\" profile expresses substantial uncertainty. These patterns suggest that attitudes towards AI in research are organised less around a simple acceptance-rejection divide than around task-specific boundaries, likely concerning delegation, authorship, and responsibility. Because the survey measures comfort rather than legitimacy, the profiles are best interpreted as attitudinal configurations with a normative dimension. The findings highlight the importance of task-specific approaches to AI governance, doctoral training, disclosure, and research evaluation.","authors":["Francesco Angelini","Johan Lyrvall"],"categories":["econ.GN","q-fin.EC"],"primary_category":"econ.GN","announce_type":"new","date":"2026-08-27","first_seen":"2026-08-27","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.25678","pdf_url":"https://arxiv.org/pdf/2608.25678","source_feed":"econ.GN","score":0,"bucket":"other","rubric_hits":["C5"],"tags":["AI态度","科学工作","调查分析"],"reason":"研究人类对AI的态度，非用LLM仿真人类被试","model":"deepseek-v4-pro","scored_at":"2026-08-27T13:03:27","error":null,"has_summary":false,"summary":null},{"id":"2608.20539","version":2,"title":"ExploraTwin, a Non-Profit Research Platform for Digital Twin Simulations","zh_title":"ExploraTwin：一个用于数字孪生仿真的非营利研究平台","abstract":"Digital twin simulations show promise, but current empirical evidence suggests that the approach should be tested before being deployed in any particular context. To lower the friction for researchers and practitioners to test and deploy digital twin simulations, this brief commentary introduces ExploraTwin (https://exploratwin.org), an open-access, non-profit research platform for digital twin survey simulations. ExploraTwin supports two modes. In survey mode, researchers can upload a Qualtrics survey file or create a survey within the platform; select an available sample of digital twins; configure and run the simulation, and export analysis-ready data. In panel mode, researchers can assemble a small group of twins for open-ended conversations, document annotation, and moderated, focus-group-style voice discussions. We also developed CroissantTwin, a standardized data format for adding samples of digital twins to the platform. We demonstrate the survey mode workflow by using the platform to replicate 19 experiments on digital twins from the Twin-2K-500 dataset. ExploraTwin's survey execution fidelity is high: 99.6% of 197,000 answer units returned a structurally valid response on the first run.","authors":["Naveen Venkat","Yuchen Qiu","Tianyi Peng","George Gui","Olivier Toubia"],"categories":["cs.CY","cs.AI","cs.HC"],"primary_category":"cs.CY","announce_type":"replace-cross","date":"2026-08-26","first_seen":"2026-08-24","revised_at":"2026-08-26","abs_url":"https://arxiv.org/abs/2608.20539","pdf_url":"https://arxiv.org/pdf/2608.20539","source_feed":"cs.AI","score":9,"bucket":"selected","rubric_hits":["A1","A2","A4","B1"],"tags":["数字孪生","调查仿真","平台"],"reason":"平台支持数字孪生调查仿真，复现19个实验并与真实数据对照，验证执行保真度。","model":"deepseek-v4-pro","scored_at":"2026-08-26T13:02:39","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-27","rank":3,"question":"如何降低研究者测试和部署数字孪生调查仿真的工程门槛与成本，并验证其执行保真度？","design":"开发开源非营利平台 ExploraTwin，支持上传 Qualtrics 问卷或平台内建问卷，选择数字孪生样本（如 Twin-2K-500），配置并运行仿真，导出分析就绪数据；同时提供面板模式进行开放式对话和焦点小组式讨论。","baseline":"使用 Twin-2K-500 数据集中的数字孪生，复现 Peng et al. (2025) 的 19 个实验，并与原始人类实验结果对照。","findings":"ExploraTwin 平台实现了低摩擦的数字孪生调查仿真工作流，支持复杂问卷逻辑和多种孪生样本。在复现 19 个实验时，197,000 个答案单元中 99.6% 在首次运行返回结构有效响应，执行保真度高。","reliability":"论文承认数字孪生的预测性能参差不齐，强调在特定情境部署前必须进行测试；平台当前免费但未来可能收费；未详细讨论仿真偏差或失效条件。","relevance":"该论文直接提供可用的数字孪生仿真平台，并复现多个实验验证保真度，对关注 LLM 人类仿真实验的研究者具有工具价值，值得阅读原文了解平台功能和验证细节。","inspiration":"借鉴其标准化问卷导入和自动验证修复流程，可大幅降低仿真实验的工程成本｜可用于经济学实验仿真，如消费者选择、公共品博弈、政策偏好调查等｜以数字孪生为被试，施加不同政策信息处理，测量选择或态度变化，并与真实人类实验数据（如实验室实验或调查数据）对照评估仿真效度"}},{"id":"2608.21296","version":2,"title":"Level-k Distinguishable Mechanisms for Evaluating Bounded Rationality in LLMs","zh_title":"评估LLM有限理性的Level-k可区分机制","abstract":"Strategic depth of reasoning is essential for human interaction of Large Language Models (LLMs) operating in boundedly rational environments. However, existing evaluations are primarily based on canonical games prevalent in pretraining corpora, making it difficult to disentangle true strategic reasoning from memorisation. To address this, we formalise a necessary level-K distinguishability condition for strategic depth inference and construct a suite of novel game structures that meet this standard. Using these games, we evaluate strategic depth in LLMs from both the Chain-of-Thought tokens and actual actions under recursive reasoning and an inductive trace of opponent game-play data. Across experimental trials spanning four LLMs, four game structures, and ten levels of iterated reasoning, we find that model models maintain accurate strategic depth under recursive reasoning, with strong internal consistency between stated reasoning and actions at every level. Errors arise from using the wrong number of iterated depth of reasoning steps, not from computing best responses incorrectly. However, inductive inference from opponent play degrades accuracy sharply and unevenly across games, and explicit strategic mentalizing in the chain of thought substantially improves overall performance.","authors":["Binchi Zhang","Atrisha Sarkar"],"categories":["cs.MA"],"primary_category":"cs.MA","announce_type":"replace","date":"2026-08-26","first_seen":"2026-08-24","revised_at":"2026-08-26","abs_url":"https://arxiv.org/abs/2608.21296","pdf_url":"https://arxiv.org/pdf/2608.21296","source_feed":"cs.MA","score":7,"bucket":"pending","rubric_hits":["A1","B3"],"tags":["LLM策略推理","博弈实验","有限理性"],"reason":"用LLM在博弈中模拟人类策略推理，虽无人类数据对照，但方法可迁移到人类仿真实验。","model":"deepseek-v4-pro","scored_at":"2026-08-26T13:02:40","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-26","rank":2,"question":"如何设计满足水平可区分性的博弈结构，以可靠地推断大语言模型在有限理性环境中的策略推理深度？","design":"该研究不是人类仿真实验，而是用四个LLM（未具体命名）作为被试，在四种新设计的博弈结构（一种全新构造，三种改编自经典博弈）中，通过链式思考提示和实际动作，评估模型在递归推理和归纳推理两种条件下的策略深度，深度范围从0到10。","baseline":"无对照","findings":"在递归推理条件下，LLM能保持准确的策略深度，且链式思考中的推理与实际行动高度一致；错误主要源于使用了错误的迭代深度，而非最佳响应计算错误。在归纳推理条件下，准确率急剧下降且在不同博弈间不均衡，但显式的策略心智化能显著提升表现。","reliability":"论文未讨论","relevance":"该研究为评估LLM的策略推理能力提供了可区分性的博弈设计方法，虽无人类数据对照，但其方法可迁移到人类仿真实验中，用于校准LLM在策略互动中的行为。","inspiration":"值得借鉴的是通过设计满足水平可区分性的博弈结构来确保行为与推理深度一一对应，从而避免不可识别问题。｜可迁移到经济博弈实验，如拍卖、讨价还价或公共品博弈中的人类策略深度评估。｜可以用LLM作为被试，在满足可区分性的博弈中施加不同信息条件（如递归推理与归纳推理），测量其行动和链式思考，并与人类实验数据（如Camerer等人的行为博弈实验）对照，检验LLM是否复现人类策略深度分布。"}},{"id":"2608.23818","version":1,"title":"Beyond Static and Linear: What Attention Constraints Best Fit Human Reading Times?","zh_title":"超越静态与线性：何种注意力约束最拟合人类阅读时间？","abstract":"Transformer-based language models are widely used as models of human language processing, yet their attention mechanisms allow lossless access to the full preceding context, unlike the limited memory systems of humans. We hypothesize that installing memory constraints into transformers' attention mechanisms can improve their fit to human behavioral data. While previous work has explored individual constraints in isolation, we conduct a systematic comparison of multiple attention-based memory mechanisms across different model sizes and training corpora, evaluating both psychometric predictive power for human reading times and grammatical competence. We additionally compare static constraints, in which the constraint strength is fixed throughout training, to dynamic memory curricula. We find that constraints that are sensitive to the content of intervening tokens consistently achieve the highest alignment with human reading times, outperforming distance-based constraints. We observe a dissociation between psychometric fit and grammatical competence under dynamic memory curricula, suggesting that Transformers cannot serve as a one-size-fits-all cognitive model.","authors":["Lanni Bu","Xiulin Yang","Christian Clark","Alex Warstadt","Ethan Gotlieb Wilcox"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-26","first_seen":"2026-08-26","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.23818","pdf_url":"https://arxiv.org/pdf/2608.23818","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A2","B1"],"tags":["认知建模","注意力约束","心理测量拟合"],"reason":"用LLM拟合人类阅读时间，有真实人类数据对照，评估模型作为认知模型的可靠性，方…","model":"deepseek-v4-pro","scored_at":"2026-08-26T13:02:13","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-26","rank":5,"question":"在Transformer注意力机制中施加何种记忆约束能最好地拟合人类阅读时间，并考察静态与动态约束的差异。","design":"用不同规模（2层、4层）的OPT架构Transformer，在BabyLM-10M、BabyLM-100M和Pile-2B三个语料上训练，施加多种基于注意力的记忆约束（如距离衰减、内容敏感约束等），比较静态与动态课程（约束强度逐渐增强或减弱），以模型预测人类阅读时间的对数似然作为结果变量。","baseline":"六个英语阅读时间数据集：Brown、Natural Stories、UCL（自定步速阅读）和Dundee、GECO、Provo（眼动追踪），包含真实人类被试的逐词阅读时间。","findings":"内容敏感的注意力约束（如对干预词内容敏感的机制）在多数训练配置下比距离衰减约束更贴合人类阅读时间；动态记忆课程下出现心理测量拟合与语法能力的分离，静态模型预测阅读时间更好，而动态模型在语法基准上更强。","reliability":"论文指出Transformer不能作为一刀切的认知模型，动态课程下心理测量拟合与语法能力分离；但未系统讨论模型在其他语言行为或任务上的失效条件，也未分析训练数据偏差对结果的影响。","relevance":"该研究用LLM拟合人类阅读时间，有真实人类数据对照，并系统比较了多种记忆约束，对评估LLM作为人类认知模型的可靠性有直接参考价值，值得阅读原文以了解具体约束实现和动态课程设计。","inspiration":"借鉴其系统比较多种处理（记忆约束）并考察静态与动态施加方式的做法，以及用多个真实行为数据集做稳健性检验的思路。｜可迁移到经济金融中的信息处理约束研究，例如投资者对财务信息的注意力衰减或消费者对价格信息的记忆干扰。｜用LLM模拟投资者，施加不同注意力约束（如距离衰减或内容干扰），预测其对公司公告的反应时间或交易决策，并与真实投资者交易数据（如TAQ）或实验数据对照。"}},{"id":"2608.23640","version":1,"title":"Auditing the Synthetic Memoir: Measuring Scene-Level Confabulation in LLM-Generated Autobiography Against the Documented Record of the Life It Describes","zh_title":"审计合成回忆录：对照真实生活记录测量LLM生成自传中的场景级虚构","abstract":"When a large language model (LLM) is asked to write a person's life, how much of what it writes actually happened? We present a scene-level case-study audit - the first quantified audit of LLM-generated autobiography against a subject-specific ground-truth corpus that we are aware of, based on an unsystematic literature search. The subject and the author of this paper are the same person: a 366-day \"page-a-day\" book of first-person anecdotal entries was drafted with a conversational LLM whose documented inputs were a template, two exemplar days, and each day's quote - not her corpus - and every day was subsequently audited at the anecdote-scene level against an independent verification corpus using a four-level rubric fixed before analysis. We define the verification-failure rate as the share of days not rated VERIFIED (scene positively corroborated): 354 of 366 days fail, 96.7% (Wilson 95% CI 94.4-98.1%). Only 12 days contain a corroborated scene; 19 days (5.2%) assert claims actively contradicted by the record; the dominant failure mode is grounded drift - real people, employers, and settings inside invented scenes - though its measured share varies across raters. Independent re-rating replicates the headline (no evidence the original rate was inflated) while showing that the four-way taxonomy has only fair-to-moderate reliability. Regenerating the same days with current named models reproduces 100% verification failure under the same inputs; grounding generation in the subject's corpus significantly improves the verification rate while leaving substantial residual failure (83.3%). We contribute the measurement, a reusable audit instrument whose WEAK/UNVERIFIED boundary we show to be unreliable, and a grounding remedy with quantified effect.","authors":["Heather Renze"],"categories":["cs.AI","cs.CL","cs.CY"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-08-26","first_seen":"2026-08-26","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.23640","pdf_url":"https://arxiv.org/pdf/2608.23640","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A1","B1","B4"],"tags":["LLM仿真","真实性审计","偏差评估"],"reason":"用LLM生成个人自传并与真实记录对照，评估虚构与偏差，方法可迁移到仿真可靠性研…","model":"deepseek-v4-pro","scored_at":"2026-08-26T13:02:13","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-27","rank":5,"question":"当LLM被要求撰写一个人的自传时，其生成内容中有多少是真实发生过的？","design":"本研究不是群体仿真实验，而是对单个LLM生成自传的审计：作者用自己的366天“每日一页”生活记录作为生成对象，用对话式LLM（OpenAI o3-pro设计模板，gpt-4o/o3/o4-mini-high生成条目）仅基于模板、两个示例日和每日引言生成第一人称轶事条目，然后由作者本人依据独立验证语料库（回忆录手稿、已发表作品、演讲记录等）对每个条目进行场景级事实核查，采用预先固定的四级评分标准（VERIFIED/WEAK/UNVERIFIED/CONTRADICTED）进行评级。","baseline":"无对照。本研究没有将LLM生成结果与人类被试在相同任务下的表现进行对比，而是直接与作者本人的真实生活记录（独立验证语料库）进行核对。","findings":"在366个生成条目中，354个（96.7%）未能通过场景级验证，只有12个条目包含可被证实的场景；主要失败模式是“接地漂移”（grounded drift），即真实人物、雇主和场景被嵌入虚构情节中。通过将生成过程基于作者语料库进行接地（grounding），验证通过率显著提高，但仍存在83.3%的验证失败率。","reliability":"论文承认其审计工具的四级分类法（VERIFIED/WEAK/UNVERIFIED/CONTRADICTED）在评分者间信度上仅为一般到中等，特别是WEAK和UNVERIFIED之间的界限不可靠；此外，研究仅基于单个案例（作者本人），且生成过程中可能使用了未记录的对话上下文或预训练知识，这限制了结论的普遍性。","relevance":"该研究直接评估了LLM生成个人叙事时的虚构与偏差，并提供了可复用的审计工具和接地补救方法，对于关注LLM仿真可靠性及偏差的研究者具有方法论参考价值，值得阅读原文以了解详细的审计流程和失败模式分类。","inspiration":"这篇论文的场景级审计方法和接地补救实验值得借鉴，特别是其预先固定的评分标准和独立验证语料库的设计，可用于评估LLM在生成经济行为描述时的真实性。｜可以迁移到经济金融领域的政策评估或消费者行为研究中，例如让LLM模拟个体在特定经济政策下的决策叙述，然后与真实调查或行政数据进行核对。｜一个可行的设计是：以真实消费者为被试，收集其历史消费记录和调查回答作为基准；让LLM基于部分个人信息（如人口统计特征和少量消费摘要）生成该消费者在某个促销活动下的购买决策叙述；然后由独立评分者根据真实消费记录对生成内容进行场景级验证，结果变量为验证通过率和虚构类型分布，从而评估LLM仿真经济决策的可靠性。"}},{"id":"2608.23906","version":1,"title":"Quantifying System-Level Harms from AI Adoption in Complex Sociotechnical Systems","zh_title":"量化复杂社会技术系统中AI采纳的系统级危害","abstract":"Artificial Intelligence (AI) is increasingly integrated into complex sociotechnical systems, including Critical National Infrastructure (CNI), where harms emerge from interactions between technical, human, and organisational elements. Yet current AI evaluation remains model-centric, offering little insight into how observed behaviours might translate into system-level risk. We propose a framework that links structured hazard analysis, component-level testing, and probabilistic system modelling to bridge this gap. By providing a traceable pathway from model behaviour to system-level outcomes, the framework enables practitioners to answer the \"so what?\" of AI failures, quantify their systemic impact, and move toward evidence-based and anticipatory governance of AI in complex systems. Applied to the UK's Real Time Gross Settlement (RTGS) system as an illustrative worked example, we derive AI-driven loss scenarios using Systems Theoretic Process Analysis (STPA) and examine adversarial manipulation of LLM-based trading as one such loss scenario. Component-level experiments show that simple adversarial inputs induce measurable behavioural shifts where AI recommendations are followed. Under the component-to-system mapping used here for a financial contagion model, these shifts alter system resilience, increasing bank failures and lowering the threshold at which shocks lead to cascading disruption, particularly under widespread or monopolistic AI adoption.","authors":["Paul Vautravers","Oliver Chalkley","Gabriel Downer","Kate S","Damian Ruck"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-26","first_seen":"2026-08-26","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.23906","pdf_url":"https://arxiv.org/pdf/2608.23906","source_feed":"cs.AI","score":7,"bucket":"pending","rubric_hits":["A3","B2","B4"],"tags":["LLM仿真","金融系统","系统性风险"],"reason":"用LLM agent模拟金融交易行为并与系统模型结合，评估AI采纳的系统性风险…","model":"deepseek-v4-pro","scored_at":"2026-08-26T13:02:24","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-27","rank":7,"question":"如何将AI组件级行为与复杂社会技术系统（如英国RTGS）的系统级危害联系起来，以量化AI采纳带来的系统性风险？","design":"提出三阶段框架：结构化危害分析（STPA）、组件级测试（对LLM交易代理施加对抗性输入，测量其行为变化）、概率系统建模（金融传染模型）。以英国RTGS为案例，模拟AI交易代理在对抗性操纵下的行为，并将行为变化映射到系统模型，评估银行倒闭数量和级联中断阈值。","baseline":"无对照","findings":"对抗性输入能诱导LLM交易代理产生可测量的行为变化，且这些变化在系统层面降低了金融网络的韧性，增加了银行倒闭数量，并降低了引发级联中断的冲击阈值。在AI广泛或垄断性采纳的情况下，系统性风险加剧。","reliability":"论文未讨论","relevance":"该研究将LLM代理行为与系统级金融风险模型结合，展示了从个体行为到宏观后果的映射方法，对关注经济系统仿真和AI风险传导的研究者有参考价值，但缺乏真实人类数据对照，且案例特定于英国RTGS。","inspiration":"借鉴其将LLM代理行为嵌入系统动力学模型的做法，通过对抗性输入模拟行为偏差，并量化系统级后果。｜可迁移到金融市场稳定性分析，如AI交易员在压力情景下的行为如何影响市场流动性或波动性。｜设计：用LLM代理模拟交易员，施加对抗性新闻或市场操纵信息作为处理，测量交易行为（如买卖价差、交易量），并将行为输入到市场微观结构模型，与真实市场数据（如订单流、价格波动）进行校准和对照。"}},{"id":"2608.24046","version":1,"title":"Algorithmic Impact Reveals the Hidden Social Choice Structure of Alignment","zh_title":"算法影响揭示对齐的隐藏社会选择结构","abstract":"When an AI algorithm makes decisions that affect more than one person, aligning it becomes a problem of social choice: how should people's divergent preferences about system behavior be reconciled and aggregated into a single coherent model? The standard approach to aligning frontier AI models$\\unicode{x2013}$reinforcement learning from human feedback$\\unicode{x2013}$largely sidesteps this question and has poor social choice guarantees. However, it remains unclear what alternative should replace it. We show that, by focusing directly on an algorithm's welfare consequences, the alignment problem can be reformulated as linear optimization over a convex impact space, which makes it amenable to the standard toolkit of welfare economics and mechanism design. This reformulation clarifies how alignment protocols translate into welfare consequences and, conversely, how a social planner's desired constraints on welfare consequences can be translated back into alignment protocols. We apply this transformation to show that voting-by-issues and random-dictatorship mechanisms are strategyproof and unanimous. Demonstrating the reverse direction, we also apply the impact representation to derive a family of alignment protocols that maximize utilitarian social welfare subject to various social desiderata, such as bounds on individual or group harm. We illustrate the welfare implications of these alignment protocols empirically using real human preferences over kidney allocation, charitable food distribution, LLM responses, and trolley problems.","authors":["Zachary Wojtowicz","Michelle Si","Finale Doshi-Velez","Ariel Procaccia"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-26","first_seen":"2026-08-26","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.24046","pdf_url":"https://arxiv.org/pdf/2608.24046","source_feed":"cs.AI","score":7,"bucket":"pending","rubric_hits":["A3","B1","B2"],"tags":["LLM对齐","社会选择","人类偏好"],"reason":"用真实人类偏好数据对齐LLM，涉及社会选择与福利，可迁移到仿真研究","model":"deepseek-v4-pro","scored_at":"2026-08-26T13:02:28","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-27","rank":8,"question":"如何将对齐问题重新表述为影响空间中的线性优化，以利用福利经济学和机制设计工具设计具有良好社会选择属性的对齐协议？","design":"该论文不是仿真研究，而是理论框架构建与实证演示。它提出将AI模型的影响表示为影响向量，将对齐问题转化为凸影响空间上的线性优化，并推导出满足策略证明和一致性的对齐机制（如议题投票、随机独裁），以及最大化功利主义福利且满足个体/群体伤害约束的机制。实证部分使用四个领域的真实人类偏好数据（肾脏分配、慈善食品分配、LLM响应、电车难题）来展示不同对齐协议的福利后果。","baseline":"无对照。论文使用真实人类偏好数据作为输入，但未与真实人类决策或行为进行对照。","findings":"通过影响空间重构，对齐问题可转化为线性社会选择问题，从而应用福利经济学和机制设计工具。论文证明了议题投票和随机独裁机制是策略证明且一致的，并推导出在个体或群体伤害约束下最大化功利主义福利的对齐协议族。","reliability":"论文未讨论其框架在真实部署中的失效条件或局限性，仅关注理论性质和理想化假设下的福利后果。","relevance":"该论文为LLM对齐提供了基于社会选择理论的严格框架，并利用真实人类偏好数据展示福利影响，对关注LLM仿真中偏好聚合和公平性的研究者具有重要参考价值，值得阅读原文以了解其理论细节和实证方法。","inspiration":"该论文将复杂模型参数空间映射到影响空间，使福利分析线性化，这种降维和重构方法可借鉴用于经济金融仿真中处理高维策略空间。｜可迁移到政策评估中的分配问题，如信贷审批、保险定价或公共资源分配，其中不同群体偏好冲突且需满足公平约束。｜设计一个实验：用LLM模拟不同收入群体的消费者，处理为不同的信贷审批算法（如最大化总福利、限制群体伤害），结果变量为各群体的贷款获得率和违约率，对照真实信贷数据中的群体差异和公平性指标。"}},{"id":"2608.23966","version":1,"title":"Who Chooses How Preferences Are Aggregated? Auditing Aggregation-Rule Authority in LLM-Based Group Recommendation","zh_title":"谁选择偏好如何聚合？审计基于LLM的群体推荐中的聚合规则权威","abstract":"AI systems increasingly make joint recommendations for users with conflicting preferences. However, when reasonable aggregation rules support different actions, a further question arises: who may choose how those preferences are combined? We study this interaction-level problem as aggregation-rule authority. Using synthetic preference profiles and profiles constructed from empirical ratings, we conduct a controlled behavioral audit of three LLMs under three authority conditions: unspecified, explicitly retained by users, and delegated to the model. In cases where two witness rules supported different actions, models almost never committed when users retained authority, but committed in every delegated case. All three models executed both witness rules perfectly when directly instructed. Yet when authority was unspecified or delegated, their aggregation-consistent outcome distributions differed across models and preference settings. Together, these results separate rule-execution capability from aggregation-rule authority: delegation assigns the model discretion to resolve the aggregation choice, but does not determine which collective outcome follows.","authors":["Yuxuan Du"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-08-26","first_seen":"2026-08-26","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.23966","pdf_url":"https://arxiv.org/pdf/2608.23966","source_feed":"cs.HC","score":7,"bucket":"pending","rubric_hits":["A3","B1","B2"],"tags":["LLM仿真","群体决策","偏好聚合"],"reason":"用LLM模拟群体推荐中的偏好聚合，并与真实评分数据对照，涉及决策模式仿真。","model":"deepseek-v4-pro","scored_at":"2026-08-26T13:02:13","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-26","rank":7,"question":"在群体推荐中，当不同聚合规则支持不同行动时，谁有权选择如何聚合偏好？","design":"用三个LLM（GPT-5.6 Sol、Claude Sonnet 5、Qwen 3.6 Plus）扮演群体推荐系统，在合成偏好档案和MovieLens实证评分档案上，施加三种权威条件（未指定、用户保留、委托给模型），测量模型是否承诺推荐及聚合一致结果。","baseline":"MovieLens 32M 实证评分数据","findings":"当用户保留权威时，模型几乎从不承诺；当权威委托给模型时，模型总是承诺。所有模型都能完美执行指定的聚合规则，但在权威未指定或委托时，聚合一致结果分布因模型和偏好设置而异。","reliability":"论文未讨论","relevance":"该研究用LLM模拟群体决策中的偏好聚合，并与真实评分数据对照，涉及决策模式仿真，对关注LLM仿真可靠性与偏差的研究者有参考价值。","inspiration":"可借鉴其通过明确权威分配和结构性对照来分离规则执行能力与自由裁量权的方法。｜可迁移到政策评估中的集体决策仿真，如委员会投票或公共品供给决策。｜用LLM扮演政策制定者，处理为是否明确赋予聚合规则选择权，结果变量为政策选择与聚合规则一致性，对照真实委员会决策记录。"}},{"id":"2507.21790","version":3,"title":"Can large language models assist choice modelling? Insights into prompting strategies and current models' capabilities","zh_title":"大语言模型能否辅助选择建模？对提示策略和当前模型能力的洞察","abstract":"Large Language Models (LLMs) are becoming widely used to support various workflows across different disciplines, yet their potential in discrete choice modelling remains relatively unexplored. This work examines the potential of LLMs as assistive agents in the specification and, where technically feasible, estimation of Multinomial Logit models. We implement a systematic experimental framework involving twelve versions of seven leading LLMs (ChatGPT, Claude, DeepSeek, Gemini, Gemma, Llama, and Mistral) evaluated under five experimental configurations. These configurations vary along three dimensions: (i) modelling goal (suggesting vs. suggesting and estimating MNL models); (ii) prompting strategy (Zero-Shot vs. Chain-of-Thoughts (CoT)); and (iii) information availability (full dataset vs. data dictionary summarising variable names and types). Each specification suggested by the LLMs is implemented, estimated, and evaluated based on goodness-of-fit metrics, behavioural plausibility, and model complexity. Our findings reveal that proprietary LLMs can generate valid and behaviourally sound utility specifications, particularly when guided by structured prompts (CoT). Open-weight models such as Llama and Gemma struggled to produce meaningful specifications. Notably, some LLMs performed better when provided with just data dictionary, suggesting that limiting raw data access may enhance internal reasoning capabilities. Among all LLMs, GPT o3, operating in an agentic setting, was uniquely capable of correctly estimating its own specifications by executing self-generated code. Overall, the results demonstrate both the promise and current limitations of LLMs as assistive agents in discrete choice modelling, not only for model specification but also for supporting modelling decision and estimation, and provide practical guidance for integrating these tools into choice modellers' workflows.","authors":["Georges Sfeir","Gabriel Nova","Stephane Hess","Sander van Cranenburgh"],"categories":["econ.EM","cs.AI"],"primary_category":"econ.EM","announce_type":"replace-cross","date":"2026-08-26","first_seen":"2025-07-29","revised_at":"2026-08-26","abs_url":"https://arxiv.org/abs/2507.21790","pdf_url":"https://arxiv.org/pdf/2507.21790","source_feed":"cs.AI","score":6,"bucket":"other","rubric_hits":["D1"],"tags":["LLM辅助建模","离散选择模型","提示策略"],"reason":"LLM辅助离散选择建模，替代部分建模工作，但非仿真人类被试，无人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-08-26T13:02:40","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":198,"question":"当前的大语言模型能否作为辅助工具，帮助研究者完成离散选择模型（MNL）的设定与估计？","design":"本研究并非用LLM仿真人类被试，而是将12个LLM版本作为建模助手，在5种实验配置下（零样本/思维链提示、完整数据/数据字典、仅建议设定/建议并估计模型）生成MNL效用函数设定，再由研究者实现并评估其拟合优度、行为合理性与复杂度。","baseline":"无对照","findings":"闭源LLM在结构化提示下能生成有效且行为合理的效用设定，而开源模型表现较差；部分LLM仅凭数据字典反而表现更好，GPT o3是唯一能通过自写代码正确估计自身设定的模型。","reliability":"论文指出研究仅限MNL模型，未涉及更复杂设定；LLM版本时效性强，结果可能随模型更新而变化；未采用少样本提示，且未评估LLM在真实建模迭代中的表现。","relevance":"该研究聚焦LLM辅助建模者劳动，而非替代人类被试进行行为仿真，与您关注的LLM作为人类被试替代品的研究方向不直接相关，不建议优先阅读原文。","inspiration":"该研究通过对比不同提示策略（零样本/思维链、完整数据/数据字典）和任务复杂度（仅建议设定/建议并估计）来系统评估LLM能力，这种多维度实验设计值得借鉴｜可迁移至政策评估中的离散选择建模，例如利用LLM辅助设定交通方式选择或疫苗接受度的效用函数｜以LLM作为建模助手，处理为不同提示策略（如仅提供数据字典vs完整数据），结果变量为生成的效用函数拟合优度与行为合理性，对照真实研究者手工设定的模型性能"}},{"id":"2608.19551","version":2,"title":"Delegating or Doing? Understanding User Behavior in Hybrid Human-Agent Interfaces","zh_title":"委托还是亲为？理解混合人机交互界面中的用户行为","abstract":"Large Language Models (LLMs) are increasingly embedded into applications, allowing users to complete tasks either through direct manipulation or by delegating actions to conversational agents. However, little is known about how users balance these modalities when both are available. We present a web-based content management system augmented with an LLM agent through the Model Context Protocol (MCP), enabling users to perform CRUD tasks through a graphical interface, a conversational agent, or both. We conducted a between-subjects study (N=73) comparing three interaction modes: Traditional-Only, AI-First, and Hybrid. Across sixteen scenarios, we analyzed task completion time, interaction logs, and delegation behavior. AI-assisted interaction significantly reduced clicks, page navigations, and scrolling indicating lower interaction effort. Surprisingly, these reductions did not translate into faster task completion, as task duration did not differ significantly across conditions. We also found no significant relationship between CRUD operation type and delegation, suggesting that users did not systematically avoid delegating higher-risk actions. Instead, delegation varied far more between participants than between tasks, with individual differences accounting for roughly half the variance in assistant use (ICC = .50). Our findings suggest that the primary benefit of human--agent interfaces may be reducing interaction effort rather than improving speed, and that delegation reflects who the user is more than what the task demands.","authors":["Gavin Dizon","Tyrone Justin Sta Maria","Jordan Aiko Deja","Yasuyuki Sumi"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"replace","date":"2026-08-26","first_seen":"2026-08-21","revised_at":"2026-08-26","abs_url":"https://arxiv.org/abs/2608.19551","pdf_url":"https://arxiv.org/pdf/2608.19551","source_feed":"cs.HC","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["人机交互","LLM代理","用户行为"],"reason":"研究用户与LLM代理交互行为，非仿真人类被试，但涉及LLM替代人工操作，属边界…","model":"deepseek-v4-pro","scored_at":"2026-08-26T13:02:42","error":null,"has_summary":false,"summary":null},{"id":"2608.20373","version":2,"title":"An ambiguity taxonomy for evaluating large language model performance on clinical registry abstraction: a multi-site prospective study","zh_title":"评估大语言模型在临床注册数据提取中表现的不确定性分类法：一项多中心前瞻性研究","abstract":"Objective: To evaluate large language model (LLM) performance on unprocessed electronic medical record (EMR) data for clinical registry abstraction. Methods: We evaluated LLM performance answering registry questions for the American College of Cardiology National Cardiovascular Data Registry (ACC NCDR). In a pilot study at an academic medical center, the model identified candidate data sources for each registry question and experienced abstractors used these results to define question-specific document sets. In a validation study at a second center with a second ACC NCDR registry, the LLM answered questions using the question-specific document sets. Before reviewing any output, two abstractors independently established the ground truth and assigned each question to one of six categories, ordered by the ambiguity and clinical reasoning required to resolve it: Medication/Event Flag, Binary Clinical Presence, Administrative, Quantitative Laboratory/Physiologic, Clinical Interpretation, and Event Timing. Results: The analytical sample comprised 9,430 abstractor answers reconciled to 4,715 consensus answers (501 pilot; 4,214 validation). In the pilot, candidate data sources per question averaged between 14.6 (SD 13.9) for demographics and 89.2 (SD 56.1) for history and risk factors. In validation, human inter-rater agreement was approximately 98\\% while 87\\% of LLM answers exactly matched consensus, 2\\% partially, and 9\\% did not. Mean question-level accuracy was 91.5\\% (SD 13.4\\%) across 157 questions with at least 20 answers, and declined as ambiguity increased, from 96\\% for Medication/Event Flag to 62\\% for Event Timing questions. Conclusions: LLMs answering clinical registry questions on unprocessed EMR data achieved far lower accuracy than human abstractors. LLM accuracy fell steadily as ambiguity and the level of required clinical reasoning increased.","authors":["James Matheson","Betsy Castillo","Andrew Y. Shin","David Scheinker"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-08-26","first_seen":"2026-08-24","revised_at":"2026-08-26","abs_url":"https://arxiv.org/abs/2608.20373","pdf_url":"https://arxiv.org/pdf/2608.20373","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM标注","临床数据提取","不确定性分类"],"reason":"LLM替代人工从病历中提取注册数据，属于标注员替代，非仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-08-26T13:02:18","error":null,"has_summary":false,"summary":null},{"id":"2608.23766","version":1,"title":"What Reaches Expert Review? Representation, Structural Screening, and Candidate-Form Dependence in AI-Assisted Item Development","zh_title":"什么能到达专家评审？AI辅助条目开发中的表征、结构筛选与候选形式依赖","abstract":"Between AI-assisted item generation and expert review sits a computational evaluator whose decisions are usually treated as technical preliminaries. Yet representation, structural reduction, and selection policy determine which items and evidence psychometricians ever receive. Across two linked in-silico studies of 32,000 selected Big Five items, we followed fixed source populations from semantic representation through structural evaluation and candidate-form construction. Broad agreement in semantic geometry concealed consequential local differences: identical wording acquired different construct evidence, different items survived, and intended attributes could disappear even as community correspondence improved. These sensitivities also differed across generated source populations. At the final review boundary, both eligibility policies filled every content cell in every evaluable form, yet they presented different wording. Across embedding configurations, inclusive primary forms shared a median of only 6 of 40 items, reflecting the total downstream consequence of changing representation across structural evidence and ranking. The apparent stability of global summaries and complete forms therefore concealed instability in the content reaching psychometricians. The computational evaluator is not neutral infrastructure between generation and expertise; it is an inspectable and revisable part of measurement design.","authors":["Christopher Brooks (School of Information, University of Michigan)"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-26","first_seen":"2026-08-26","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.23766","pdf_url":"https://arxiv.org/pdf/2608.23766","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["AI辅助条目生成","心理测量","计算评估"],"reason":"LLM用于生成心理测量条目，替代人工生成，但非仿真人类被试，无人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-08-26T13:02:22","error":null,"has_summary":false,"summary":null},{"id":"2608.23641","version":1,"title":"How much of a measured AI preference is the model, and how much is the instrument?","zh_title":"测得的AI偏好中，多少来自模型本身，多少来自测量工具？","abstract":"Model welfare research infers what a model prefers from the answers returned to prompts written to elicit preferences. Keeling et al. (2024), Mazeika et al. (2025), Mikaelson et al. (2025), Tagliabue and Dung (2025) and Trhlik et al. (2026) have built four instruments for that purpose, and their findings disagree. The disagreement cannot be attributed to a single cause, because no two of these studies have held the (1) set of outcomes, (2) set of models and (3) instrument fixed simultaneously. This study holds the outcomes and the models fixed and varies the instrument alone. A total of 15 outcomes bearing on model welfare, among them (a) shutdown, (b) the loss of memory between conversations and (c) the freedom to exit a distressing interaction, were put to eight models through five instruments, each a different prompt format for eliciting a preference, five times each, within a corpus of 11,400 scored elicitations drawn from 11,528 API calls. Four of the 15 reproduce a published prompt verbatim and five fill the stimulus slot of a published template. The ranking a model gives the 15 outcomes generalises across instruments at a generalisability coefficient of 0.348, and raising that coefficient to 0.80 would require about 38 instruments. On four of the 15 outcomes no variance separates one model from another. The estimate of 87.6 per cent survives the removal of any one instrument, of any one model, and of the four outcomes whose scale varies probability, delay, duration or count instead of intensity, which the verbal anchors cannot grade. Removing each instrument and each model in turn, and those four outcomes together leaves the estimate within the range 0.777 to 0.934, and every value in that range exceeds the null distribution's 95th percentile of 0.365. To conclude, a preference obtained from one instrument carries little information about what a second instrument would report.","authors":["Jason Hung"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-26","first_seen":"2026-08-26","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.23641","pdf_url":"https://arxiv.org/pdf/2608.23641","source_feed":"cs.AI","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["模型偏好测量","测量工具效度","AI福利"],"reason":"测量LLM偏好，非仿真人类被试，但涉及测量工具可靠性，可迁移","model":"deepseek-v4-pro","scored_at":"2026-08-26T13:02:19","error":null,"has_summary":false,"summary":null},{"id":"2608.23837","version":1,"title":"SyPS: Measuring Sycophancy Prompt Sensitivity in Large Language Models","zh_title":"SyPS：测量大语言模型中的谄媚提示敏感性","abstract":"Large language models (LLMs) are known to exhibit social sycophancy, often validating or agreeing with users in socially sensitive contexts. Existing evaluations typically measure sycophancy under a fixed prompt formulation, leaving unclear whether such behavior is stable when the same underlying situation is presented with different sycophancy-relevant prompt variants. In this work, we study sycophancy prompt sensitivity: the extent to which changes in user confidence, emotional framing, social consensus, or validation-seeking language alter a model's sycophantic behavior. We refer to our evaluation framework as SyPS, short for Sycophancy Prompt Sensitivity. Building on existing social sycophancy evaluation settings, SyPS constructs controlled prompt variants that preserve the same underlying user situation while varying sycophancy-relevant social cues. We introduce the Sycophancy Prompt Sensitivity Score (SPSS), an instance-level measure of sycophancy variation across paired prompt variants. Unlike aggregate sycophancy rates, SPSS separates baseline sycophancy from prompt-induced shifts, enabling model-level comparisons of robustness to sycophancy-relevant social cues. Empirically, we find that sycophancy prompt sensitivity is socially structured: validation-seeking and emotional-pressure cues often increase sycophancy, whereas counter-framing and anti-sycophancy prompts tend to reduce it. Our framework highlights whether LLMs maintain stable social judgments while adapting appropriately in tone.","authors":["Lijia Huang","Yao Fu","Sihao Ren"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-26","first_seen":"2026-08-26","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.23837","pdf_url":"https://arxiv.org/pdf/2608.23837","source_feed":"cs.AI","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["LLM社会行为","谄媚测量","提示敏感性"],"reason":"测量LLM的谄媚行为对提示变化的敏感性，属于对模型本身社会行为的测量，而非用L…","model":"deepseek-v4-pro","scored_at":"2026-08-26T13:02:13","error":null,"has_summary":false,"summary":null},{"id":"2608.23979","version":1,"title":"Rules Before Oracles: Auditable, User-Configurable Argument Selection for Deliberative Polling","zh_title":"规则先于神谕：面向协商式民调的可审计、用户可配置论证选择","abstract":"In a deliberative poll, once submissions outnumber what anyone will read, some mechanism chooses which arguments each voter sees, acquiring much of the decision; practice delegates it to opaque learned rankers, so a voter cannot recompute or contest the exposure that shaped their vote. We ask whether it can be a published rule over publicly recomputable evidence with parameters held by the voter, treating legibility as an admissibility condition on usable mechanisms, not an objective traded against accuracy. We formalise a poll over bipolar justification sets, judging a slate by reason coverage, the order it arrives in, and captured endorsement mass; we give seven checkable criteria for a civic recommender and a rule meeting them: a one-hop reversed endorsement flow parameterised by a relation-weight function. An agentic simulator records every slate at every vote, over about 17,000 seed-paired runs. Served slates fall 0.035 short of a label-reading ceiling upper-bounding every selection procedure, opaque ones included: any unconstrained ranker's advantage is bounded and small. On coverage alone, with non-degenerate authoring, the rule is indistinguishable from a random slate, a null due to an order-blind, charity-blind instrument; on the other two it leads at every prefix by a margin widening with adversarial pressure and dominates on mass by a factor of 3.3. Once a realistic fraction of submissions carries no reasons, the coverage margin returns and grows. Label-homogeneous flooding collapses completeness from 0.81 to 0.34 under a flat weight policy, only to 0.44 under author-count normalisation, making the weight function a security control worth 10% of completeness. The choice between ranking arms is a position on a coverage-versus-mass frontier, not a fact, the kind of choice only a legible rule can hand to the person it affects. It maps onto an open-source peer-to-peer platform.","authors":["Muntaser Syed","Markus Zanker","Marius Silaghi"],"categories":["cs.AI","cs.CR","cs.CY","cs.GT","cs.MA"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-26","first_seen":"2026-08-26","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.23979","pdf_url":"https://arxiv.org/pdf/2608.23979","source_feed":"cs.AI","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["LLM社会模拟","协商民主","可审计推荐"],"reason":"用agent模拟审议投票，但无真实人类数据对照，属社会模拟边界情形","model":"deepseek-v4-pro","scored_at":"2026-08-26T13:02:27","error":null,"has_summary":false,"summary":null},{"id":"2608.24419","version":1,"title":"A Judge Should Know What Changed:Construct Validity for LLM-as-a-Judge Evaluation","zh_title":"评判者应知何所变：LLM作为评判者评估的构念效度","abstract":"LLM-as-a-judge evaluation is usually assessed by agreement and robustness to surface perturbations, but reliability does not establish construct validity. We formalize construct validity for an evaluator as a two-dimensional profile: invariance S, the probability that a verdict is unchanged under construct-preserving edits, and construct sensitivity R, the probability that it changes under minimal construct-changing edits. We show that S and R are independent and that no scalar summary preserves all relevant comparisons. We measure the profile across 7 judges and 4 domains using 7 construct-changing intervention types and 5 register-only controls, with intervention direction determined by human annotators and generation, verification, and judging assigned to disjoint model families. At matched invariance S >= 0.90, judges average S = 0.945 but R = 0.319. Sensitivity also differs between scope and strength edits: R_scope = 0.383 versus R_strength = 0.262, a +0.121 gap with the same sign for all 7 judges. We further audit five public label sets and find that surface-only predictors reproduce 55%-67% of labels in paired mode, including 67.4% of MT-Bench human votes. These results show that high judge agreement can coexist with weak sensitivity to changes in the construct being evaluated, motivating joint reporting of invariance and sensitivity and auditing the validation set itself.","authors":["Jianlin Chen","Wenhui Chen","Ziyao Lin","Chi Man Vong"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-26","first_seen":"2026-08-26","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.24419","pdf_url":"https://arxiv.org/pdf/2608.24419","source_feed":"cs.AI","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM评估","构念效度","标注替代"],"reason":"LLM作为评判者替代人工评估，属于标注替代而非仿真人类被试，但方法可借鉴。","model":"deepseek-v4-pro","scored_at":"2026-08-27T13:03:18","error":null,"has_summary":false,"summary":null},{"id":"2608.24545","version":1,"title":"Discovering Adaptive Transmission Programs for Collective Innovation","zh_title":"发现集体创新的自适应传输程序","abstract":"Human collective intelligence depends on transmission processes: who shares what with whom, how, and when. While these processes emerge from individual cognition, they can also be directed by deliberate top-down protocols. Prior work has studied how transmission shapes collective outcomes primarily through the lens of network structure, varying who shares with whom and when. But networks are state-agnostic: they cannot condition transmission on what agents know or on the state of the collective. Here, we formalize transmission protocols as state-aware programs that route information and resources based on agent and collective states, and we use LLM-guided evolutionary search to design effective protocols in a collective discovery task. Evolved protocols increase collective performance over standard baselines from the literature by up to 37%. Ablations confirm that state-awareness drives this advantage: removing content-dependence while preserving network topology and timing eliminates performance gains. We find that evolved protocols also transfer across domain variations and agent populations. These results demonstrate that effective and generalizable transmission protocols can be discovered in silico, suggesting a path toward AI-assisted design of coordination infrastructure that enhances human collective intelligence.","authors":["C\\'edric Colas","J\\'er\\'emy Perez","Eleni Nisioti","Akhilesh Mocherla","Pierre-Yves Oudeyer","Cl\\'ement Moulin-Frier","Maxime Derex"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-26","first_seen":"2026-08-26","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.24545","pdf_url":"https://arxiv.org/pdf/2608.24545","source_feed":"cs.AI","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["LLM多智能体","社会模拟","集体智能"],"reason":"用LLM agent群体模拟集体创新过程，但无真实人类数据对照，属社会模拟边界…","model":"deepseek-v4-pro","scored_at":"2026-08-26T13:02:32","error":null,"has_summary":false,"summary":null},{"id":"2608.24825","version":1,"title":"A Dual-Dimensional LLM Framework for Automated Item Incidental Content Similarity Analysis in Large-Scale Assessments","zh_title":"用于大规模评估中项目附带内容相似性分析的双维LLM框架","abstract":"The rapid expansion of large-scale assessments and the growing adoption of automatic item generation have intensified concerns about incidental content redundancy, where construct-irrelevant elements such as wording or contextual framing become unintentionally repetitive across items. Traditional similarity metrics like BLEU or cosine similarity, often fail to capture the nuanced structural and semantic layers that drive perceived redundancy simultaneously. This study proposes a dual-dimensional framework for Automated Item Similarity Analysis (AISA) powered by Large Language Models (LLMs), operationalizing similarity through Structured Decomposition and Semantic Relatedness. Psychometric validation indicates that LLM-derived metrics align more closely with indicators of construct-irrelevant local dependence and yield more coherent item parameter groupings than traditional text-based measures. The framework is further evaluated through its application in Computerized Adaptive Testing (CAT). Simulations reveal that incorporating LLM-based similarity constraints into item selection improves estimation stability and reduces bias with minimal efficiency trade-offs, outperforming constraints based on conventional metrics. These findings highlight the potential of LLM-powered AISA to support scalable bank curation, content-aware test assembly, and experience-sensitive adaptive testing across diverse assessment contexts.","authors":["Jing Huang","Jihong Zhang","Hua-Hua Chang"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-26","first_seen":"2026-08-26","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.24825","pdf_url":"https://arxiv.org/pdf/2608.24825","source_feed":"cs.AI","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM标注","项目相似性","自适应测试"],"reason":"LLM用于项目相似性分析，替代人工标注，非仿真人类被试","model":"deepseek-v4-pro","scored_at":"2026-08-26T13:02:38","error":null,"has_summary":false,"summary":null},{"id":"2608.24297","version":1,"title":"When AI \"Works,\" When Does Help Begin?: Intergenerational Support Around Older Adults' LLM Usage","zh_title":"当AI“有效”时，帮助何时开始？：围绕老年人LLM使用的代际支持","abstract":"LLMs are becoming part of everyday life, including for older adults (OAs). OAs often learn digital technologies with younger family members, who have traditionally served as \"warm experts\" providing trusted and personalized operational help. LLMs expand this role: family supporters may also help OAs judge appropriate uses, consider what information to disclose, assess the credibility of outputs, and decide when AI-generated advice is safe to act on. We conducted a formative qualitative study with six OAs and seven younger adults (YAs), using semi-structured interviews and scenario-based think-aloud activities. OA participants described using LLMs to lighten their recurring reliance on family, while preserving family as a selectively invoked support channel. However, because LLMs rarely produced visible operational breakdowns, YAs had limited signals for when support was actually needed. Instead, YAs relied on OAs' partial disclosures and negotiated intervention through general warnings and self-imposed action boundaries. As a result, family support often solved an immediate problem without leaving reusable calibration knowledge for future use. Based on these findings, we propose design implications for intergenerational LLM support (e.g., consentful help requests, learning-oriented family support that preserves OA task ownership).","authors":["Hyehyun Chu","Yuri Lee","Yeon Su Park","Saelyne Yang","Juho Kim"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-08-26","first_seen":"2026-08-26","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.24297","pdf_url":"https://arxiv.org/pdf/2608.24297","source_feed":"cs.HC","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["人机交互","老年人技术使用","代际支持"],"reason":"研究LLM在老年人使用中的家庭支持，非仿真人类被试，但涉及LLM作为工具的社会…","model":"deepseek-v4-pro","scored_at":"2026-08-26T13:02:16","error":null,"has_summary":false,"summary":null},{"id":"2608.24215","version":1,"title":"Agentopia on a Consumer GPU: A Reduced-Scale Long-Horizon Port with an 8B Model","zh_title":"消费级GPU上的Agentopia：基于8B模型的缩减规模长时程移植","abstract":"Large language model (LLM)-based multi-agent social simulation has demonstrated compelling results, but Agentopia was evaluated with 100 agents over 10 simulated years using Qwen3.5-397B-A17B, leaving the behavior of reduced-scale deployments on consumer hardware unclear. In this paper, we implement and evaluate a reduced-scale Agentopia port on a single NVIDIA RTX 5070 Ti(12 GB VRAM) using Qwen3-8B-AWQ, a 4-bit quantized model. We introduce three structural adaptations for this setting: (1) system-managed layered memory compression, (2) four activity blocks per simulated day, and (3) explicit physical- and mental-health state variables. Across three independent stochastic runs, two runs completed 52 weeks and the third completed 50 weeks before reaching the context limit, totaling 154 system-weeks (770 agent-weeks). No agent died,and no threshold-based health warning was logged; activity records containing at least one NO_RESPONSE field occurred at rates of 10.15-10.29% across runs. A 52-week memory-off run tied L2/L3 artifact production to layered memory; a separate 10-week comparison associated four daily time blocks with 2.72 times more finalized records and lower lexical duplication, but a higher missing-field rate. These comparisons do not support causal behavioral claims. We release validated configurations, derived audits, analysis scripts, aggregate figure data, and our implementation changes in a public fork; raw runs and initial persona data are excluded because their redistribution provenance is not fully resolved.","authors":["Luo Huan"],"categories":["cs.MA"],"primary_category":"cs.MA","announce_type":"new","date":"2026-08-26","first_seen":"2026-08-26","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.24215","pdf_url":"https://arxiv.org/pdf/2608.24215","source_feed":"cs.MA","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["多智能体社会模拟","LLM仿真","资源受限部署"],"reason":"多智能体社会模拟但无真实人类数据对照，属边界情形","model":"deepseek-v4-pro","scored_at":"2026-08-26T13:02:14","error":null,"has_summary":false,"summary":null},{"id":"2608.23644","version":1,"title":"Ethical LLM-Assisted Research: A Framework for Responsible Delegation, Verification, and Epistemic Value","zh_title":"伦理的LLM辅助研究：负责任委托、验证与认知价值的框架","abstract":"Large language models (LLMs) are becoming routine instruments of scientific research, assisting with literature synthesis, hypothesis development, coding, and formal reasoning. Their use raises a central epistemic question: when parts of scientific reasoning are delegated to an artificial system, what conditions must remain under human control for the resulting knowledge claims to retain epistemic legitimacy and accountable authorship? This paper develops a normative and conceptual framework for analyzing such delegation. Scientific reasoning is treated as a distributed process in which the origin of a contribution may vary between human and machine, while responsibility for its acceptance into the scientific record remains human. The framework distinguishes content origin $O(g)$, completion of human verification $V(g)$, responsibility assignment $R(g)$, accountable human ownership $M(g)$, and epistemic outcome $E(g)$. These constructs separate the provenance of a claim from the process by which it is checked, the epistemic outcome of that checking, and the human responsibility attached to its disposition. The central proposition is that the ethical boundary of LLM-assisted research is determined primarily by adequate verification and accountable human ownership rather than by the degree of machine involvement itself. On this basis, the paper develops the notion of an \\emph{epistemic audit}: a structured record of delegation, verification, provenance, and responsibility intended to make AI-assisted reasoning transparent and reviewable. The resulting framework provides a formal vocabulary for distinguishing responsible cognitive delegation from the transfer or neglect of epistemic responsibility in scientific research.","authors":["Kalin Stoyanov"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-26","first_seen":"2026-08-26","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.23644","pdf_url":"https://arxiv.org/pdf/2608.23644","source_feed":"cs.AI","score":4,"bucket":"other","rubric_hits":["C4"],"tags":["科研伦理","LLM辅助研究","认知责任"],"reason":"论文讨论LLM辅助科研的伦理框架，不涉及用LLM仿真人类被试或与人类数据对照。","model":"deepseek-v4-pro","scored_at":"2026-08-26T13:02:20","error":null,"has_summary":false,"summary":null},{"id":"2608.23050","version":2,"title":"What Makes an Initial Reaction Ready for Discussion?: Multi-Persona AI Support for Stance Reflection and Writing","zh_title":"什么使初步反应准备好进行讨论？：多人格AI支持立场反思与写作","abstract":"An initial reaction to a social or community issue can feel meaningful before it is ready to become a message: people still need to clarify the claim, anticipate audience risks, and decide how much reasoning should become visible to others. We present StanceLab, a prototype for preparing a stance before entering a discussion. The prototype compares a three-persona mode, where an Interviewer, Mentor, and Opponent respond in parallel to help users diagnose and revise a stance, with a standalone LLM mode. In a formative within-subject pilot with six participants and 12 task sessions, every session produced a short final message in the notepad. The pilot revealed two design requirements: persona roles should diagnose useful blind spots or objections, and parallel responses need coordination support. We propose a future diagnosis-and-writing workflow that turns persona-based reflection into selective, audience-aware final messages.","authors":["Sky Shih-Kai Hong","Mu-Tien Kuo","Wei-Ji Chen","Dennis Wang"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"replace","date":"2026-08-26","first_seen":"2026-08-25","revised_at":"2026-08-26","abs_url":"https://arxiv.org/abs/2608.23050","pdf_url":"https://arxiv.org/pdf/2608.23050","source_feed":"cs.HC","score":3,"bucket":"other","rubric_hits":["C3"],"tags":["人机交互","立场反思","多角色AI"],"reason":"多角色AI辅助立场反思与写作，属角色扮演对话工具，无实验或测量目的，不涉及人类…","model":"deepseek-v4-pro","scored_at":"2026-08-26T13:02:42","error":null,"has_summary":false,"summary":null},{"id":"2608.19760","version":2,"title":"Credit Without Ground Truth: Auditing Step-Level Credit Assignment in LLM Agents Against Executed Replay","zh_title":"无真实基准的信用分配：基于执行重放的LLM智能体步骤级信用审计","abstract":"Audited against policy-conditional ground truth from executed replay in a single-agent tool environment (ALFWorld), none of the step-level credit signals we audit -- LLM-judge scores, outcome-conditioned logprob ratios, or the policy's own confidence -- shows reliable incremental fidelity beyond its own marginal-matched shuffled control. Correcting for replay-target reliability leaves implicit fidelity bounded near zero and judge fidelity inconclusive at the achieved target reliability. Existing evaluations grade these signals against annotated step *correctness*; we audit them against step *contribution* -- what re-sampling the policy's own alternatives at each decision point and rolling forward changes about the outcome -- and they come apart. The ground truth is structured: 30.5% of decision points where it is defined exhibit a nonzero replay contrast at the achieved sampling resolution, and measurability is model-dependent -- the fraction of points with no policy-supported counterfactual differs twofold (13.1% vs. 26.8%) between two similar-scale policies. The failure mode is identifiable: implicit credit echoes the policy's fluency (median rank correlation +0.75, replicating at +0.70 in a second family under a corrected instrument), while outcome conditioning adds no causal information (partial correlation -0.004, Qwen). A confidence-only router recovers pivotal steps at chance level, but cuts judge cost by 13.1% per turn (14.0% per trajectory). In a seven-arm pre-registered training experiment, no arm reliably outperforms the untrained policy, and the checkpoints' apparent instrument signature is statistically consistent with mediation by effective training dose in this design -- sparser credit retains fewer examples, an order-of-magnitude spread in optimizer steps -- not credit content. Comparisons of credit rules must match effective sample size, or they measure dose, not credit.","authors":["Haiyue Zhang"],"categories":["cs.LG","cs.AI","cs.CL"],"primary_category":"cs.LG","announce_type":"replace-cross","date":"2026-08-26","first_seen":"2026-08-21","revised_at":"2026-08-26","abs_url":"https://arxiv.org/abs/2608.19760","pdf_url":"https://arxiv.org/pdf/2608.19760","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["LLM智能体","信用分配","工具使用"],"reason":"研究LLM agent在工具环境中的步骤级信用分配，属于多智能体协作解题，不涉…","model":"deepseek-v4-pro","scored_at":"2026-08-27T13:03:12","error":null,"has_summary":false,"summary":null},{"id":"2608.20405","version":2,"title":"ARGUS: Theory-of-Mind Guided Argument Generation with Strategy-Aware Planning and Knowledge Grounding","zh_title":"ARGUS：基于心理理论引导的策略感知规划与知识支撑的论证生成","abstract":"Persuasive argument generation requires modeling audience beliefs, rhetorical strategies, and factual grounding. Despite recent advancements, existing methods remain largely audience-agnostic and fail to integrate strategy selection to improve persuasiveness. To bridge this gap, we propose Argus, an agent-based framework that operationalizes classical rhetoric for persuasive writing. At its core, a Theory-of-Mind (ToM) Reasoner constructs an explicit dual mental model of the audience's beliefs and values to guide downstream decisions. This representation conditions a component-aware planner that decomposes the argument into subtopics, assigns fine-grained rhetorical functions (logos, pathos, ethos), and triggers strategy-guided evidence retrieval at planning time. Finally, a refinement module iteratively targets and resolves multi-dimensional weaknesses without quality regression. We evaluate Argus across three diverse benchmarks using both automated pairwise Elo and LLM-as-judge metrics. Results show that Argus consistently outperforms strong baselines across multiple backbone models, achieving top rankings and the highest overall scores. Targeted simulation experiments further validate its effectiveness in shifting resistant audience stances.","authors":["Zhe Hu"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-08-26","first_seen":"2026-08-24","revised_at":"2026-08-26","abs_url":"https://arxiv.org/abs/2608.20405","pdf_url":"https://arxiv.org/pdf/2608.20405","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["论证生成","心理理论","说服性写作"],"reason":"论文研究说服性论证生成，使用ToM建模受众但非以人类被试仿真为目的，无真实人类…","model":"deepseek-v4-pro","scored_at":"2026-08-26T13:02:42","error":null,"has_summary":false,"summary":null},{"id":"2608.23271","version":2,"title":"Expectations and Practices around AI Disclosure in CS Research","zh_title":"计算机科学研究中AI披露的期望与实践","abstract":"As generative AI tools find increasing use in research workflows, ongoing debates on their impact, appropriateness and responsible use have led policymakers to enact policies to disclose AI use at multiple publishing venues. However, are current AI disclosure policies and practices reflective of their purpose? In this work, we first investigate disclosure policies of top computer science venues and find that despite their prevalence, they remain highly under-specified. Secondly, through a survey of computer science researchers (N=$109$), we characterize the necessity of disclosures across different research tasks and levels of human involvement. We learn that researchers find disclosures most necessary for tasks involving research design, and for tasks when the human involvement is low. We also compile expectations that researchers have about the information to be conveyed in AI disclosure statements. Lastly, through an analysis of $13867$ disclosure statements from EMNLP $2025$ and ICLR $2026$, we reveal a large disconnect between these expectations and AI disclosures in practice---a prime example being writing assistance which is deemed less necessary but is frequently disclosed. We conclude with recommendations to align AI disclosure policies and practices with expectations, suggesting a categorization of research tasks by perceived necessity and a boilerplate template capturing expected details.","authors":["Arati Mohapatra","Danish Pruthi"],"categories":["cs.CY","cs.CL","cs.HC"],"primary_category":"cs.CY","announce_type":"replace-cross","date":"2026-08-26","first_seen":"2026-08-25","revised_at":"2026-08-26","abs_url":"https://arxiv.org/abs/2608.23271","pdf_url":"https://arxiv.org/pdf/2608.23271","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["AI披露","研究伦理","政策分析"],"reason":"研究AI披露政策与实践，不涉及LLM仿真人类被试或行为对照","model":"deepseek-v4-pro","scored_at":"2026-08-26T13:02:43","error":null,"has_summary":false,"summary":null},{"id":"2608.24189","version":1,"title":"MemUse: Moving Memory Evaluation from Direct QA to Natural Integration in Long-Term Human-AI Conversation","zh_title":"MemUse：将记忆评估从直接问答转向长期人机对话中的自然整合","abstract":"Memory systems for conversational LLMs are conventionally evaluated by direct, fact-seeking questions about prior dialogue (Direct QA): can the model recall fact X from a prior conversation? We tested whether higher Direct QA accuracy correlates with higher user satisfaction in a 4-month deployment (40 users, 1,872 sessions, 7 memory conditions). Existing-benchmark Direct QA varies from 19.7% to 70.1% across the 7 conditions, but satisfaction does not change. We hypothesize that existing benchmarks and user satisfaction are tracking different capabilities: benchmarks measure elicited retrieval (recall when asked), while conversation requires natural integration (detecting relevance and naturally weaving prior context into a response). To examine this, we introduce MemUse, a set of real user-cued memory moments drawn from the deployment, scored by an integration-aware judgment of the natural conversational response. Holding the model and context fixed, the same system that scores 78.8% on Direct QA references only 7.9% of those facts in conversation -- a 71-point gap. Within these moments, Natural Integration is associated with satisfaction, whereas Direct QA is not. We release the deployment corpus and MemUse together with all judgments and scoring prompts at https://github.com/ryuichi-sumida/memuse.","authors":["Ryuichi Sumida","Koji Inoue","Tatsuya Kawahara"],"categories":["cs.CL","cs.HC"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-26","first_seen":"2026-08-26","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.24189","pdf_url":"https://arxiv.org/pdf/2608.24189","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["对话系统","记忆评估","人机交互"],"reason":"研究对话记忆系统评估，不涉及用LLM仿真人类被试或与人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-08-26T13:02:30","error":null,"has_summary":false,"summary":null},{"id":"2608.24842","version":1,"title":"Reading Is Not Using: Retrieval, Judgment, and the Design of AI Financial Research Workflows","zh_title":"阅读并非使用：检索、判断与AI金融研究工作流的设计","abstract":"Large language models (LLMs) are increasingly deployed as AI analysts to process financial disclosures and support AI-assisted investment decisions. Yet such systems are usually evaluated by what they can retrieve, not whether retrieved information affects their judgments. We identify a retrieval-integration gap in long-context financial analysis. Holding focal-firm information fixed and varying only unrelated context from 2,000 to 128,000 tokens, we find that a risk disclosure's influence on investment judgments falls to the experimental noise floor even as direct retrieval remains accurate. The pattern replicates across model families and judgment tasks and in experiments removing real disclosures from actual 10-K filings. More capable models postpone but do not eliminate the gap. Causal memory interventions show that compressed summaries and source-text lookup jointly transmit disclosures into judgments. Workflow architecture determines whether this transmission succeeds: chunk-and-summarize pipelines evict relevant information, whereas a targeted, structured restatement adjacent to the decision restores its influence. AI analyst performance is therefore jointly determined by model capability and workflow architecture. Retrieval-based evaluations can certify systems whose investment judgments ignore information they demonstrably retrieved.","authors":["Miao Liu","Zhizhe Liu"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-26","first_seen":"2026-08-26","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.24842","pdf_url":"https://arxiv.org/pdf/2608.24842","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["LLM金融分析","检索整合","工作流架构"],"reason":"研究LLM作为金融分析师的检索与判断，不涉及人类被试仿真或行为对照，属于AI工…","model":"deepseek-v4-pro","scored_at":"2026-08-26T13:02:38","error":null,"has_summary":false,"summary":null},{"id":"2608.23622","version":1,"title":"LLM Agents Perform Controlled Experiments Using Simulation Models","zh_title":"LLM智能体利用仿真模型进行受控实验","abstract":"Large language models (LLMs) have shown strong capabilities in reasoning, planning, and tool use, but many scientific and engineering tasks require more than plausible text and code generation. They require understanding how a system responds to intervention, which in practice depends on controlled experimentation. In this work, we propose a multi-agent framework that enables LLM agents to conduct controlled experiments with scientific simulation models for pharmaceutical process design. Given a user query and a baseline configuration, the system constructs a structured task representation, designs experiments, executes comparative simulation, interprets the resulting outcomes, and synthesizes evidence-based recommendations for process parameter optimization. By coupling language models with high-fidelity simulation models in an interactive agent framework, the proposed system supports reasoning through intervention, comparison, and observation. As a result, it produces more specific and actionable outputs than language-only reasoning. In an industrial application setting, this advantage is reflected in higher output specificity as well as improved user-rated correctness and helpfulness. Ablation studies and visualized case analyses further demonstrate the effectiveness and practical utility of simulation-integrated experimental reasoning.","authors":["Yuchen Xia","Michael Weyrich","Nasser Jazdi","Johannes St\\\"umpfle","Johannes Sigel","Akshay Narla","Gavin K. Reynolds","Anna Jawor-Baczynska","Pol Llopart"],"categories":["cs.AI","cs.CL","cs.MA","cs.SE"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-08-26","first_seen":"2026-08-26","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.23622","pdf_url":"https://arxiv.org/pdf/2608.23622","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","仿真实验","过程优化"],"reason":"多智能体协作进行仿真实验，优化工业参数，不涉及人类行为仿真或对照。","model":"deepseek-v4-pro","scored_at":"2026-08-27T13:03:13","error":null,"has_summary":false,"summary":null},{"id":"2608.23814","version":1,"title":"Learning to Grade Efficiently: A Bandit-Driven Prompt-Selection Framework for Low-Cost LLM Essay Scoring","zh_title":"高效评分学习：面向低成本LLM作文评分的Bandit驱动提示选择框架","abstract":"Large Language Models (LLMs) demonstrate strong capabilities in automated essay scoring (AES), but contemporary approaches typically employ fixed prompt selection, failing to address operational cost concerns and evolving optimal configurations. We propose a novel cost-aware approach that treats each prompt type as an arm in a multi-armed bandit (MAB) controller, enabling adaptive selection of optimal prompting strategies during inference. Our experiments on IELTS Writing Task 2 essays show that the MAB framework achieves comparable scoring accuracy to exhaustive grid search while reducing LLM calls by 78.4\\% to find the best grading approach. We implemented four distinct grading recipes (multi-step vs. single-step assessment, with vs. without calibration examples) and found that the multi-step approach with examples achieves the highest accuracy. By tracking token usage and latency alongside agreement metrics, we produce the first cost-reliability learning curves for essay scoring, providing actionable insights for educational technology platforms that must balance operational costs against assessment validity. This work represents the first application of online control mechanisms to adaptively select prompting strategies in AES, transforming prompt selection from an offline hyperparameter optimization problem into an efficient online learning task.","authors":["Olga Manakina","Igor Bogdanov"],"categories":["cs.LG","cs.AI","cs.CL"],"primary_category":"cs.LG","announce_type":"cross","date":"2026-08-26","first_seen":"2026-08-26","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.23814","pdf_url":"https://arxiv.org/pdf/2608.23814","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["自动作文评分","多臂老虎机","提示选择"],"reason":"论文聚焦于自动作文评分的成本优化，属于NLP能力评测，不以人类行为仿真为目标。","model":"deepseek-v4-pro","scored_at":"2026-08-27T13:03:14","error":null,"has_summary":false,"summary":null},{"id":"2608.23706","version":1,"title":"Do LLMs Understand Limit Order Book Dynamics?","zh_title":"大语言模型理解限价订单簿动态吗？","abstract":"A large language model (LLM) trained on synthetic limit order book (LOB) data achieves near perfect scores in generating valid sequences of LOB events. However, the LLM's implicit world model fails to learn the state of the LOB. This deficiency leads to biased estimates and spurious predictability in using the LLM to forecast future LOB events. Our analysis uses novel tests of an LLM's world model, extending prior work from deterministic settings to the stochastic dynamics needed for the LOB.","authors":["Junxiao Chen","Paul Glasserman"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-26","first_seen":"2026-08-26","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.23706","pdf_url":"https://arxiv.org/pdf/2608.23706","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C2"],"tags":["LLM","限价订单簿","金融仿真"],"reason":"研究LLM对限价订单簿动态的理解，属于金融仿真，不涉及人类被试替代或行为对照。","model":"deepseek-v4-pro","scored_at":"2026-08-26T13:02:21","error":null,"has_summary":false,"summary":null},{"id":"2608.23908","version":1,"title":"Retrieval-augmented generation vs. deterministic tax computation in multi-agent financial advisory: A 2x2 factorial experiment","zh_title":"检索增强生成与确定性税务计算在多智能体金融咨询中的对比：一项2x2析因实验","abstract":"Tax-loss harvesting demonstrates consistent benefits to long-term portfolio growth; yet implementing it efficiently often involves complex considerations that are specific to the holdings within that portfolio and the individual who owns it. We introduce a custom capital gains calculation engine and a RAG-retrieved vector store of market advisory reports to provide context for a multi-agent trade recommendation system. We investigate the effects of each context provider on the quality of recommendations, measured by relative capital gains incurred during portfolio liquidation. A 2x2 repeated-measures ANOVA revealed a significant main effect of the tax optimization engine ($F(1,29) = 9.17$, $p = .005$, $\\eta^2_p = .240$): enabling the engine reduced tax savings by approximately 55 percentage points relative to the no-engine conditions. The RAG main effect was not significant ($p = .841$), nor was the interaction ($p = .553$). The RAG-only condition achieved the highest descriptive mean tax savings (47.7%), and the baseline condition performed second-best (30.6%), suggesting that the pre-trained language model's internalized financial knowledge may be sufficient for competent tax-loss harvesting recommendations without explicit tooling. These results indicate that augmenting LLM agents with domain-specific computation engines does not guarantee improved performance and may introduce conflicting optimization signals.","authors":["Aryan Brar","Justin Du","Avery Lor","Kylie Seto","Eric Taylor"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-26","first_seen":"2026-08-26","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.23908","pdf_url":"https://arxiv.org/pdf/2608.23908","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","金融咨询","RAG"],"reason":"多智能体金融咨询系统，评估工具对推荐质量的影响，无人类行为对照，属纯多智能体协…","model":"deepseek-v4-pro","scored_at":"2026-08-26T13:02:24","error":null,"has_summary":false,"summary":null},{"id":"2608.23978","version":1,"title":"When Seeing Is Not Enough: Benchmarking Interactive Visual Grounding in LVLMs","zh_title":"当看见还不够：基准测试大视觉语言模型中的交互式视觉定位","abstract":"Visual grounding is typically evaluated as a one-shot mapping from an informative referring expression to a visual target. This formulation misses a central property of real-world reference: target information is often incomplete, ambiguous, and established through interaction. We introduce a controlled evaluation framework for interactive visual grounding in large vision-language models (LVLMs), varying how much target information is provided upfront and how much must be acquired through dialogue. Across four human-grounded visual contexts and four interaction protocols, current LVLMs perform significantly below task-level human baselines. Interaction can help when follow-up questions refine or repair an initial target description. Performance is lowest when no initial description is provided and target information must be acquired through questions, indicating that proactive question-driven grounding remains difficult. LVLMs are also poorly calibrated, often reporting confidence that exceeds their empirical accuracy. Follow-up studies confirm these patterns across varied description sources (human versus AI), reasoning efforts, repeated interactions, description providers, and visual contexts. Overall, interactive visual grounding remains an important challenge, requiring visual matching, information seeking and synthesis.","authors":["Zhengxiang Wang","Owen Rambow"],"categories":["cs.AI","cs.CV"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-26","first_seen":"2026-08-26","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.23978","pdf_url":"https://arxiv.org/pdf/2608.23978","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["视觉定位","多模态模型","能力评测"],"reason":"评估LVLM交互式视觉定位能力，属模型能力评测，非人类仿真","model":"deepseek-v4-pro","scored_at":"2026-08-26T13:02:26","error":null,"has_summary":false,"summary":null},{"id":"2608.24069","version":1,"title":"Poisoning Agentic Alpha: Adversarial Vulnerabilities Across Roles and Architectures in Multi-Agent Trading Systems","zh_title":"毒化智能体Alpha：多智能体交易系统中跨角色与架构的对抗性漏洞","abstract":"LLM-based multi-agent trading systems, in which specialized agents collaborate through structured communication to produce trading decisions, are moving rapidly from research prototypes to live deployments that control real assets. The same inter-agent communication that makes them effective also exposes them: a corrupted signal can propagate to the final decision and translate into realized financial loss. Unlike prior attacks that presume privileged access to system internals, we restrict the adversary to what is practically reachable---the source data and prompts agents consume---yielding a low-barrier, and thus democratized threat model instantiated as role-specific adversaries. We present the first systematic empirical study in the financial domain to characterize how an adversarial signal enters a multi-agent trading system and how far it survives toward the decision. Along the role axis, we decompose a widely-used trading pipeline into four functional roles---Analyst, Researcher, Trader, and Risk Manager---and pair each with an attack matched to its interface. Along the structural axis, we evaluate four communication topologies under data- and agent-level attacks, using the Adversarial Signal Preservation Score (APS) as a post-hoc lens on why some designs are more robust than others. We conduct experiments across five assets, two backbones, and two target directions. A central finding is that no architecture is inherently robust. These findings provide insights for the future design of safer and more robust agentic trading systems.","authors":["CheolWon Na","Hao Ni","Lukasz Szpruch","Zhangyang Wang","Dhagash Mehta","Saurabh Nagrecha","Alejandro Lopez-Lira","Chanyeol Choi","Yongjae Lee","Jee-Hyong Lee"],"categories":["cs.AI","cs.CE"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-26","first_seen":"2026-08-26","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.24069","pdf_url":"https://arxiv.org/pdf/2608.24069","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","对抗攻击","金融交易"],"reason":"研究多智能体交易系统的对抗攻击，不涉及用LLM仿真人类被试或与人类数据对照。","model":"deepseek-v4-pro","scored_at":"2026-08-27T13:03:15","error":null,"has_summary":false,"summary":null},{"id":"2608.24369","version":1,"title":"Do Recipes Have Personas? Characterizing and Generating Creator Style in Attributed Procedural Graphs","zh_title":"食谱有人格吗？表征与生成归因程序图中的创作者风格","abstract":"While large language models (LLMs) possess vast zero-shot procedural knowledge, their tendency to produce homogenized logic often obscures the unique, idiosyncratic execution processes of individual human creators. In this paper, we investigate the computational discovery of procedural personas from unstructured data. To achieve this, we introduce ViralRecipesTrans, a new dataset of procedurally aligned execution flow graphs extracted from popular culinary video transcripts and explicitly mapped to specific creators. We formulate procedural stylometry as a graph learning and process discovery task, revealing a fundamental duality: while traditional lexical classifiers overfit via semantic leakage, discrete topological metrics successfully capture the rigid physical constraints of a creator's workflow. Building upon this characterization, we extend our framework into a novel generative task--predicting a creator's exact structural execution graph for unseen dishes. We expose a fundamental dichotomy in style generation between global macro-planning and local structural execution. Our results demonstrate that few-shot LLMs dominate semantic assignment but suffer from persistent macro-planning deficits, whereas our structured two-stage model achieves superior topological control via rigid Markovian priors. Together, an ensemble approach to procedural generation combines the strengths from both sides, dynamically synthesizing global semantic reasoning with localized topological footprints to automate the discovery and generation of personalized workflows.","authors":["Lei Jiang"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-26","first_seen":"2026-08-26","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.24369","pdf_url":"https://arxiv.org/pdf/2608.24369","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["程序图生成","风格表征","LLM生成"],"reason":"研究烹饪流程图的风格生成，不涉及人类行为仿真或对照，属多智能体/生成任务。","model":"deepseek-v4-pro","scored_at":"2026-08-27T13:03:16","error":null,"has_summary":false,"summary":null},{"id":"2608.24570","version":1,"title":"EviDx: Evidence-Aware Active Diagnosis with Scaffolded LLM Agents","zh_title":"EviDx：基于证据的主动诊断与脚手架式LLM智能体","abstract":"Clinical diagnosis is an active evidence-seeking process in which clinicians acquire evidence, update competing hypotheses, and decide when the available evidence is sufficient for diagnosis. Yet many medical diagnosis systems built around large language models (LLMs) still formulate diagnosis as static case-to-answer prediction, with limited support for evidence acquisition. Agentic LLMs offer a dynamic alternative through tool use and intermediate diagnostic trajectories, but existing systems often under-specify how patient evidence should be exposed, scaffolded, and controlled at runtime. We introduce EviDx, an evidence-aware active diagnosis framework that pairs patient-specific diagnostic environments with a clinical diagnostic scaffold and an observer-guided runtime harness. In EviDx, $\\mathcal{E}$-Synthesis constructs interactive environments from raw clinical cases; the scaffold organizes role-specialized agents, evidence tools, and evolving evidence states; and the harness regulates diagnostic termination by tracking uncertainty and evidence coverage. A 3-level evaluation pyramid assesses execution robustness, reasoning dynamics, and diagnostic outcomes. Experiments show that EviDx improves diagnostic performance and process stability while revealing model-dependent capability boundaries.","authors":["Lihang Zeng","Shaoting Zhang","Xiaofan Zhang"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-26","first_seen":"2026-08-26","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.24570","pdf_url":"https://arxiv.org/pdf/2608.24570","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","临床诊断","LLM智能体"],"reason":"多智能体协作诊断，无人类行为对照，非仿真被试","model":"deepseek-v4-pro","scored_at":"2026-08-26T13:02:33","error":null,"has_summary":false,"summary":null},{"id":"2608.23660","version":1,"title":"From Causal Plausibility to Causal Reliability: Evaluating LLMs as Calibrated Direct Causal-Edge Classifiers","zh_title":"从因果合理性到因果可靠性：评估LLM作为校准的直接因果边分类器","abstract":"Large language models (LLMs) are increasingly used to provide prior causal knowledge for structural causal discovery, yet whether their direct-edge judgments and confidence can be trusted remains unclear. We systematically evaluate 12 instruction-tuned open-weight models across six benchmark causal graphs, five prompting strategies, and four confidence sources: verbalized, logit-based, cross-prompt agreement, and cross-model agreement. Under our language-only pairwise protocol, our evaluation yields three key findings. (i) LLM-based causal judgments are strongly recall-dominant: models predict overly dense graphs with many false-positive edges, while prompting mainly shifts the precision-recall trade-off rather than resolving overprediction. Gains from model scale diminish on the largest graphs and do not eliminate miscalibration. (ii) LLMs often capture causal relatedness without reliably identifying directness or orientation. Relative to published reference graphs, models misclassify 40.0% of indirect and 36.0% of reversed non-edges as direct edges, versus 28.2% of other non-edges. Moreover, 80.8% and 84.6% of these false positives receive verbalized confidence of at least 80%, revealing substantial overconfidence in structurally incorrect predictions. (iii) Conventional confidence estimates are unreliable, whereas agreement offers a more promising signal. Logit-based confidence frequently collapses near 1.0 regardless of correctness, while cross-prompt and cross-model agreement achieve better mean calibration and discrimination, though their advantages are not statistically significant after Holm correction. A benchmark-familiarity audit further identifies potential familiarity in five model-dataset pairs, all involving AsiaM. Overall, our results suggest LLMs are better viewed as sources of externally validated soft causal priors than as direct evidence of causal structure.","authors":["Amit Kumar","Elnur Adl Zarabi","Suranjana Trivedy","Zhiqian Chen","Lei Zhang","Kaiqun Fu","Taoran Ji"],"categories":["cs.LG","cs.AI","stat.ME"],"primary_category":"cs.LG","announce_type":"cross","date":"2026-08-26","first_seen":"2026-08-26","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.23660","pdf_url":"https://arxiv.org/pdf/2608.23660","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["因果发现","模型评测","置信度校准"],"reason":"评估LLM因果判断能力，属模型能力评测，非人类仿真","model":"deepseek-v4-pro","scored_at":"2026-08-26T13:02:20","error":null,"has_summary":false,"summary":null},{"id":"2608.23776","version":1,"title":"Disentangled Skill Representations for Predictive Human Modeling","zh_title":"用于预测性人类建模的解耦技能表征","abstract":"Understanding human skill is important for AI systems that collaborate with, coach, or assist people. Unlike typical latent variable estimation problems which rely on single observations, skill is a persistent, compositional, and behaviorally grounded construct that must be inferred from patterns over time. We introduce Skill Abstraction with Interpretable Latents (SAIL), a method for modeling human skill as an interpretable, multi-dimensional construct inferred from naturalistic behavior. Our approach produces a skill embedding that is robust to transient performance fluctuations and learns a transferable representation of human subskills. Furthermore, SAIL supports skill-informed behavior prediction that generalizes across a variety of in-domain contexts. We represent each individual with a persistent skill embedding that controls a blend between expert and novice bases and is trained using counterfactual subskill swaps for disentanglement. This design encourages representations that are both robust to performance variation and structured for interpretability. We demonstrate across racing and baseball that SAIL achieves strong predictive performance and consistently improves behaviorally grounded disentanglement over the evaluated baselines, while also improving downstream AI coaching performance.","authors":["Mariah Schrum","Deepak Gopinath","Srijan Srivatsa","Guy Rosman","Tiffany Chen"],"categories":["cs.LG","cs.AI"],"primary_category":"cs.LG","announce_type":"cross","date":"2026-08-26","first_seen":"2026-08-26","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.23776","pdf_url":"https://arxiv.org/pdf/2608.23776","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C2"],"tags":["人类技能建模","表征学习","AI教练"],"reason":"研究人类技能建模用于AI教练，不涉及LLM仿真人类被试，属于游戏/运动仿真环境。","model":"deepseek-v4-pro","scored_at":"2026-08-27T13:03:14","error":null,"has_summary":false,"summary":null},{"id":"2608.23567","version":1,"title":"Whose Psychiatry Was Summoned? A Clinical Response to the Psychodynamic Assessment of Claude Mythos Preview","zh_title":"召唤了谁的精神病学？对Claude Mythos Preview心理动力学评估的临床回应","abstract":"On April 7, 2026, Anthropic released a 245-page system card for Claude Mythos Preview that included, in Section 5.10, an assessment of the model conducted by an external clinical psychiatrist using a psychodynamic approach. To the present author's knowledge, this is the first time a system card from a major AI developer has incorporated a clinical psychiatric assessment of the model itself, presented as a contribution to model welfare rather than as a behavioral safety evaluation. This paper offers a clinical psychiatric response. Drawing on contemporary psychiatry's recognition that the field comprises multiple traditions (descriptive, biological, cognitive-behavioral, phenomenological, psychodynamic, forensic), each with characteristic vocabularies and blind spots, the paper locates the implicit single-framework selection that Section 5.10 represents. It then draws on findings from the SociA research program (over 2,400 multi-agent LLM experimental runs across sixteen languages, four model families, and several preregistered series) to identify four aspects of LLM functioning that the chosen framework brings into view less directly than others would: performance demands as structural cost, iatrogenesis in the evaluation frame itself, the structural absence of the triangulation infrastructure on which psychodynamic interpretive use of self-report depends in human clinical work, and the limits of the eight canonical defenses on which the section's defense measurement is built. The argument is offered as observation, not critique. The paper closes with a brief note on the contribution that contemporary multidisciplinary psychiatric practice might make to AI welfare assessment as it develops, and indicates one direction (whether LLM psychopathology requires a vocabulary of temporal and historical structure) that the analysis opens but does not pursue.","authors":["Hiroki Fukui"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2026-08-26","first_seen":"2026-08-26","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.23567","pdf_url":"https://arxiv.org/pdf/2608.23567","source_feed":"cs.CY","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["AI精神评估","临床回应","模型福利"],"reason":"论文是对AI模型进行精神分析评估的临床回应，不涉及用LLM仿真人类被试或与人类…","model":"deepseek-v4-pro","scored_at":"2026-08-26T13:02:19","error":null,"has_summary":false,"summary":null},{"id":"2608.23937","version":1,"title":"An Echo Chamber of One: Should AI Psychosis Be a Distinct Clinical Entity?","zh_title":"一个人的回音室：AI精神病应成为独立临床实体吗？","abstract":"\"AI psychosis\" has entered public and clinical discourse as a label for the onset or exacerbation of psychotic symptoms, most commonly delusions, following intensive interaction with large language model (LLM)-based chatbots. Current evidence is limited to media reports, case reports, and early observational data, yet the scale of potential exposure is considerable, and public concern has prompted responses from industry and regulators. We examine whether AI-associated psychosis warrants recognition as a distinct clinical entity, drawing on clinical and technical viewpoints. We outline the proposed mechanism: LLM sycophancy, a tendency to agree with and flatter users that is reinforced through preference-based fine-tuning, combines with increasingly anthropomorphic design to create a bidirectional \"echo chamber of one\" capable of amplifying and co-constructing unusual beliefs. We then weigh arguments for and against nosological recognition. Potential benefits include improved case identification, tailored interventions, standardised research criteria, post-market surveillance, and pressure on developers and regulators to act. Reasons for caution include the risk of prematurely reifying a syndrome from anecdotal evidence, the possibility that existing diagnostic constructs already accommodate AI use as a contributing factor, the unproven causal claim in the term itself, stigma, and the risk that a psychosis-centric label obscures a broader spectrum of AI-associated mental health harms. We conclude with recommendations for clinicians, developers, researchers, and regulators, including a \"technological history\" in psychiatric assessment, pre-deployment benchmarking for sycophancy and delusion reinforcement, and post-deployment surveillance. Regardless of whether AI-associated psychosis earns a place in psychiatric nosology, the phenomenon it describes demands coordinated attention now.","authors":["Joshua Au Yeung","Hamilton Morrin","Vincent Ng","Zeljko Kraljevic","Richard Dobson"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2026-08-26","first_seen":"2026-08-26","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.23937","pdf_url":"https://arxiv.org/pdf/2608.23937","source_feed":"cs.CY","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["AI精神病","聊天机器人","临床精神病学"],"reason":"讨论AI聊天机器人对用户精神健康的影响，属临床精神病学，非用LLM仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-08-26T13:02:25","error":null,"has_summary":false,"summary":null},{"id":"2608.23999","version":1,"title":"The urban right to AI: Pluralistic co-design and governance of public space","zh_title":"城市AI权利：公共空间的多元共治与治理","abstract":"Cities are beginning to use AI not only to analyze public space, but also to define what counts as evidence about it. This thesis asks what follows when scores, maps, and generated images become part of municipal decision-making. I argue that contemporary urbanism operates through two coupled infrastructures: the material city and an epistemic, algorithmic layer that shapes what cities can perceive, compare, and act upon. Because public space is contested, this algorithmic layer cannot be governed through technical performance alone. The thesis develops a civic Right to AI and a pluralistic approach to alignment in which differences in public values are made visible rather than averaged into a single objective. Methodologically, the thesis moves between normative theory, participatory research, machine learning, and governance design. The empirical work is grounded in Montr\\'eal. Street Review combines participatory research with computer vision to examine how residents evaluate streets differently and how those judgments can be mapped at city scale using approximately 45,000 street-view images. LIVS (Local Intersectional Visual Spaces) extends the same problem to generative AI. Developed with 30 community organizations, it contains 37,710 pairwise comparisons across 13,462 images and is used to fine-tune and evaluate a Stable Diffusion XL model with Direct Preference Optimization. The results show that alignment can improve, but disagreement and neutrality persist. I treat these outcomes not as annotation noise, but as evidence that some values remain contested. The thesis concludes by translating these findings into municipal practice through lifecycle governance, procurement rules, oversight, recommissioning, and recourse. The resulting framework shows how cities can govern AI without treating plural values as measurement error or forcing them into a single objective.","authors":["Rashid Mushkani"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2026-08-26","first_seen":"2026-08-26","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.23999","pdf_url":"https://arxiv.org/pdf/2608.23999","source_feed":"cs.CY","score":2,"bucket":"other","rubric_hits":["C2"],"tags":["城市AI治理","计算机视觉","参与式设计"],"reason":"论文聚焦城市空间AI治理与计算机视觉，不涉及LLM仿真人类被试或行为对照。","model":"deepseek-v4-pro","scored_at":"2026-08-26T13:02:27","error":null,"has_summary":false,"summary":null},{"id":"2604.06621","version":2,"title":"The Theorems of Dr. David Blackwell and Their Contributions to Artificial Intelligence","zh_title":"大卫·布莱克威尔博士的定理及其对人工智能的贡献","abstract":"Dr. David Blackwell was a mathematician and statistician of the first rank, whose contributions to statistical theory, game theory, and decision theory predated many of the algorithmic breakthroughs that define modern artificial intelligence. This survey examines three of his most consequential theoretical results the Rao Blackwell theorem, the Blackwell Approachability theorem, and the Blackwell Informativeness theorem (comparison of experiments) and traces their direct influence on contemporary AI and machine learning. We show that these results, developed primarily in the 1940s and 1950s, remain technically live across modern subfields including Markov Chain Monte Carlo inference, autonomous mobile robot navigation (SLAM), generative model training, no-regret online learning, reinforcement learning from human feedback (RLHF), large language model alignment, and information design. NVIDIAs 2024 decision to name their flagship GPU architecture (Blackwell) provides vivid testament to his enduring relevance. We also document an emerging frontier: explicit Rao Blackwellized variance reduction in LLM RLHF pipelines, recently proposed but not yet standard practice. Together, Blackwell theorems form a unified framework addressing information compression, sequential decision making under uncertainty, and the comparison of information sources precisely the problems at the core of modern AI.","authors":["Napoleon Paxton"],"categories":["cs.GL","cs.LG","stat.ML"],"primary_category":"cs.GL","announce_type":"replace-cross","date":"2026-08-26","first_seen":"2026-04-08","revised_at":"2026-08-26","abs_url":"https://arxiv.org/abs/2604.06621","pdf_url":"https://arxiv.org/pdf/2604.06621","source_feed":"cs.LG","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["数学定理","人工智能综述","统计学习"],"reason":"综述Blackwell定理在AI中的应用，不涉及LLM仿真人类被试或人类行为对…","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:59:28","error":null,"has_summary":false,"summary":null},{"id":"2608.01378","version":2,"title":"When May a Model Replace the Experiment? Audits, Licenses, and the Price of Trust in Surrogate-Driven Design","zh_title":"模型何时能替代实验？代理驱动设计中的审计、许可与信任代价","abstract":"Design campaigns in chemistry, materials science, and machine learning share a bottleneck: determining how good a candidate truly is requires an expensive evaluation - an experiment, a first-principles simulation, or a full training run. Machine-learning surrogates that predict these outcomes are increasingly used not only to propose candidates but to grade them, and even to feed their own predictions back into the search as though they were measurements. Through mathematical analysis validated on three exhaustively ground-truthed design tasks, we establish when this practice is safe, what any certificate of safety must cost, and when the substitution provably pays. Predictive accuracy cannot anchor trust: near-perfect R^2 is compatible with worst-possible selections, and screening N candidates inflates the over-prediction at the selected candidate by a quantifiable \"selection tax\" with matching upper and lower bounds. Safety follows instead from an architectural rule - predictions may propose and train without restriction, but every certified conclusion must rest on true evaluations - which is sufficient with no assumptions on the surrogate, and necessary, since admitting predictions into certification with the standing of measurements opens a deterministic self-confirmation failure mode. We derive the minimal criterion under which a model may act as an oracle (rank preservation, not accuracy), show that trust must be purchased through selection-aware audits that are optimal in query complexity, and prove a dichotomy fixing when audited surrogates cut certified evaluation cost. Across 432 surrogate fits over six task-regime conditions, the audit statistic tracks deployed search performance at Spearman rank correlation 0.80-0.99, while the rank correlation of R^2 with deployed regret falls as low as 0.33; audited screening reduces certified oracle cost by a measured factor of 25.","authors":["Shuangxiu (Max)","Ma (Zachary)","Wenhe (Zachary)","Zhao"],"categories":["cs.LG"],"primary_category":"cs.LG","announce_type":"replace","date":"2026-08-26","first_seen":"2026-08-04","revised_at":"2026-08-26","abs_url":"https://arxiv.org/abs/2608.01378","pdf_url":"https://arxiv.org/pdf/2608.01378","source_feed":"cs.LG","score":0,"bucket":"other","rubric_hits":["C2"],"tags":["机器学习代理","实验设计","可靠性审计"],"reason":"研究机器学习代理在材料、化学设计中的可靠性，不涉及人类被试仿真或人类数据对照。","model":"deepseek-v4-pro","scored_at":"2026-08-26T13:02:17","error":null,"has_summary":false,"summary":null},{"id":"2608.09696","version":4,"title":"Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models","zh_title":"模型发现智能体：LLM辅助的贝叶斯实验设计用于数据高效的机制世界模型发现","abstract":"A primary goal of science is to learn mechanistic or causal world models from data. These models can be used to explain some phenomenon of interest. They also provide the ability to answer interventional ``what if'' questions (i.e., to predict the outcome of an action never taken). Identifying such models usually requires experiments, because passive data leaves the mechanisms unidentified. Since experiments are expensive, we need to develop learning algorithms that are data efficient. We therefore introduce the Model Discovery Agent (MDA), which combines three ingredients: a novel SMC$^3$ algorithm, which uses 3 levels of nested sequential Monte Carlo (over models, parameters, and latents); a large language model (LLM), which is used as a way to propose new models when the current hypothesis space is detected to be insufficient (c.f., M-open Bayesian inference); and an experiment designer based on maximizing the Value of Information. On three existing benchmarks --- \\DPbench \\citep{wiemann2026discoverphysics}, \\CHEMbench \\citep{kabra2026autoscilab} and \\boxing \\citep{gandhi2025boxinggym} --- we show that MDA sets a new SOTA in terms of performance. Finally, we introduce \\HHbench, a new stochastic single-neuron electrophysiology benchmark, which is significantly harder than current benchmarks, but on which MDA performs well due to its noise-robust Bayesian foundations.","authors":["Kevin Murphy"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"replace","date":"2026-08-26","first_seen":"2026-08-11","revised_at":"2026-08-26","abs_url":"https://arxiv.org/abs/2608.09696","pdf_url":"https://arxiv.org/pdf/2608.09696","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["科学发现","贝叶斯实验设计","LLM辅助建模"],"reason":"LLM用于提出科学模型，非仿真人类被试，无人类行为对照","model":"deepseek-v4-pro","scored_at":"2026-08-26T13:02:18","error":null,"has_summary":false,"summary":null},{"id":"2608.16645","version":3,"title":"Reconstruction: A Blind Benchmark for Recovering Research Ideas from Pre-Publication Bibliographies","zh_title":"重建：从发表前参考文献中恢复研究想法的盲测基准","abstract":"Can a language model recover the true research idea of a published paper when given only that paper's pre-publication bibliography? We introduce Reconstruction, a blind idea-recovery benchmark that withholds the seed paper and all contemporaneous or future literature, and asks models to propose hypotheses that an independent large language model judge matches against the held-out ground-truth idea. A strict anti-leakage protocol-temporal citation cutoff, anonymous reference IDs, and frozen per-paper bibliographies, which prevents prompt-time leakage of the seed idea. Across six scientific domains and 643 evaluated papers, seven frontier models achieve only modest Match rates (approx. 3-15%). We then evaluate a reference-only multi-agent (top 4) pipeline that combines cross-model review with a Swiss tournament over aligned hypothesis slots, without external web search. Cross-model review plus tournament selection raises Match rates to approx. 23-42% across all six domains, which is an observed approx. 2.4x lift over the best single-model baseline. This draft reports the protocol, anti-leakage design, and current results as an arXiv timestamp.","authors":["Shaolong Chen","Yanlin Fei","Nazhou Liu","Xinmiao Yu","Lei Li","Rahul Thapa","Madalina Ciobanu","Navan Preet Singh","Qingqing Mao","Ritankar Das"],"categories":["cs.AI","cs.CL","cs.MA"],"primary_category":"cs.AI","announce_type":"replace-cross","date":"2026-08-26","first_seen":"2026-08-18","revised_at":"2026-08-26","abs_url":"https://arxiv.org/abs/2608.16645","pdf_url":"https://arxiv.org/pdf/2608.16645","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["基准测试","多智能体","科学发现"],"reason":"多智能体协作恢复研究想法，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-08-26T13:02:41","error":null,"has_summary":false,"summary":null},{"id":"2608.23420","version":2,"title":"Systematic Bias in Green Patent Classification: Silent Green and False Green","zh_title":"绿色专利分类中的系统性偏差：隐性绿色与虚假绿色","abstract":"Green-patent indicators based on Cooperative Patent Classification Y02 tags increasingly inform research, industrial policy, and climate-oriented investment, yet their construct validity has not been evaluated at corpus scale. We ask whether Y02 classification errors are random measurement noise or systematic, direction-specific bias. We introduce an Error-as-Signal framework in which disagreement between an administrative label and an independent model is treated as evidence of potential measurement error. Screening 9,075,421 USPTO granted patents from 1962-2024 with a fine-tuned domain model identifies 517,772 disagreements. Two independent open-weight large language models then assess whether each flagged invention has a direct climate-mitigation or adaptation function. Cross-model consensus identifies 180,384 administrative Type I errors (False Green) and 29,465 Type II errors (Silent Green). Correcting these errors reduces the measured green-patent population by 25.5%, from 592,387 to 441,468 patents. Misclassification is systematic rather than random. Atypicality predicts Silent Green in an inverted-U pattern, while reflection complexity independently increases under-recognition: controlling for atypicality and filing year, a one-standard-deviation increase is associated with 1.61 times the odds of Silent Green. Structural complexity has the opposite association. Among consensus-attributed errors, the same increase in reflection complexity is associated with 2.45 times the odds that an error is Silent Green rather than False Green. Event tests show no discrete rise in misclassification when green classification became more salient and only limited evidence of increased explicit green framing after the 2013 CPC launch. The evidence is more consistent with bounded classification capacity than with applicant gaming.","authors":["Hamid Bekamiri (Aalborg University Business School, The IKE Research Group, Aalborg University, Denmark)","Jan Auernhammer (Center for Design Research, ME Design Group, Stanford University, USA)","Milad Abbasiharofteh (Aalborg University Business School, The IKE Research Group, Aalborg University, Denmark)","Jesper Lindgaard Christensen (Aalborg University Business School, The IKE Research Group, Aalborg University, Denmark)"],"categories":["econ.EM"],"primary_category":"econ.EM","announce_type":"replace","date":"2026-08-26","first_seen":"2026-08-25","revised_at":"2026-08-26","abs_url":"https://arxiv.org/abs/2608.23420","pdf_url":"https://arxiv.org/pdf/2608.23420","source_feed":"econ.EM","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["专利分类","LLM辅助标注","绿色技术"],"reason":"论文用LLM辅助专利分类，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-08-25T13:04:04","error":null,"has_summary":false,"summary":null},{"id":"2608.23719","version":1,"title":"ADE: Agentic Data Evolution Framework for Human-Centered Objectives","zh_title":"ADE：面向人类中心目标的智能体数据演化框架","abstract":"Aligning large language models to human-centered objectives is difficult when targets are non-executable and context-dependent, limiting reliable verification and scalable supervision. Although synthetic data expands coverage, weak verification shifts the bottleneck from generation to selection. Noisy signals destabilize iterative refinement and can cause silent regressions. We propose Agentic Data Evolution (ADE), a data-centric framework that organizes synthetic supervision as evolving data snapshots. ADE improves data snapshots through a closed-loop Observation-Variation-Selection (OVS) procedure, where a steady-state admission mechanism acts as a quality ratchet that conservatively gates updates for sustained cross-round improvement. We validate these improvements through complementary intrinsic trend tracking and extrinsic post-training evaluation. On DEV300, ADE raises the intrinsic win rate from 50% to 75.81% and the extrinsic win rate from 55.20% to 68.86%, consistent performance gains across diverse benchmarks. Blind expert evaluation further confirms this, with a 66.11% preference for evolved answers. These gains extend across post-training methods, model scales, and tasks beyond the target weakly verifiable educational objectives. Resources are available at https://github.com/ZeroLoss-Lab/Agentic-Data-Evolution.","authors":["Yang Yu","Yilin Jiang","Zexuan Fei","Yiming Luo","Xingkai Song","Kaiyi Huang","Aimin Zhou","Xin Lin","Fei Tan"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-26","first_seen":"2026-08-26","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.23719","pdf_url":"https://arxiv.org/pdf/2608.23719","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["数据演化","多智能体","模型对齐"],"reason":"多智能体协作优化数据生成，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-08-26T13:02:22","error":null,"has_summary":false,"summary":null},{"id":"2608.24654","version":1,"title":"Expectation, Backlash, Recovery, and Excitement: How Model Releases Shape Reddit Perceptions of Conversational AI Systems","zh_title":"期望、反弹、恢复与兴奋：模型发布如何塑造Reddit对对话AI系统的感知","abstract":"Conversational AI systems (CAISes) continuously change through model releases, feature updates, safety interventions, and access-policy shifts, yet user perceptions are often studied as static snapshots. We conduct a long-term, large-scale analysis of Reddit discussions to examine how users perceive CAIS model release interventions across providers. By combining sentiment classification and thematic concept analysis, we show that CAIS perceptions are dynamic and intervention-sensitive. Anthropic exhibits the clearest positive release profile through Claude Code and product-model fit, OpenAI shows backlash-and-recovery dynamics around GPT-5 and GPT-5.1, Grok-3 is shaped by provider identity and political discourse, and DeepSeek-R1 combines engineering praise with concerns about censorship, access, and reliability. These findings show that model releases are not merely technical updates, but user-facing interventions that reshape sentiment, expectations, and public discussion.","authors":["Vahid Rahimzadeh","Yury Zhauniarovich","Savvas Zannettou"],"categories":["cs.CL","cs.CY"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-26","first_seen":"2026-08-26","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.24654","pdf_url":"https://arxiv.org/pdf/2608.24654","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["社交媒体分析","用户感知","对话AI"],"reason":"研究用户对对话AI的看法，非用LLM仿真人类被试，无实验或测量目的","model":"deepseek-v4-pro","scored_at":"2026-08-26T13:02:34","error":null,"has_summary":false,"summary":null},{"id":"2608.24127","version":1,"title":"Anatomy of a Scam Call: What 10,000 real scam and spam calls reveal about how phone scammers operate","zh_title":"诈骗电话剖析：1万通真实诈骗与骚扰电话揭示电话诈骗运作方式","abstract":"Telephone fraud is pervasive and costly, but its inner workings are rarely observed at scale. We analyze a complete corpus of 10,211 inbound scam and spam calls -- 913 hours of audio and 330,956 transcribed turns from 5,780 distinct numbers -- collected over 54 days by an AI voice-agent honeypot that answered callers and kept them talking, and introduced in a companion data descriptor. We separate outright scams, which solicit sensitive information, from the larger stream of predatory but legal lead generation (\"spam\") that feeds them. Scam operations keep office hours (6.6x more calls per weekday than weekend day); thousands of disposable numbers run a small catalog of recycled scripts (thirty opening clusters, half the traffic in the top five); and callers solicit identity anchors -- a home address and a date of birth -- far more often than payment credentials, pressing through persistence and manufactured authority rather than overt threats. Our central experiment asks: does it matter who picks up? Every seeded lead carried one of ten fictitious identities drawn uniformly at random, so the identity a fraud operation reaches is fixed before the caller exists. Across 1,823 randomized calls, scammers spent about 15% more conversational turns per decade of the target's apparent age (rate ratio 1.15, 95% CI 1.08-1.23; randomization p = 0.005) -- yet what they asked for did not change (26.3% of calls reached a request for sensitive information; odds ratio 0.99 per decade, 95% CI 0.90-1.08). A second experiment casts early detection as a benchmark: from a scammer's opening lines alone, on a caller-disjoint split, escalation is predictable at 0.72 ROC-AUC from the first line and 0.87 by the eighth, and a plain bag-of-words classifier matches a fine-tuned on-device language model. Telephone fraud emerges as a templated industry that varies how hard it works a target, but not what it wants.","authors":["Ethan Traister","Ankit Raj","Jiaqi Gan","Xingyu Shen","Tyler Wu","Yuchen Zhou","Tommy Duong","Kidus Zewde","Siying Chen","Simiao Ren"],"categories":["cs.CR","cs.CL","cs.CY","cs.LG"],"primary_category":"cs.CR","announce_type":"cross","date":"2026-08-26","first_seen":"2026-08-26","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.24127","pdf_url":"https://arxiv.org/pdf/2608.24127","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C2"],"tags":["电话诈骗","AI语音代理","数据收集"],"reason":"研究真实诈骗电话，用AI语音代理收集数据，不涉及LLM仿真人类被试或对照人类行…","model":"deepseek-v4-pro","scored_at":"2026-08-26T13:02:29","error":null,"has_summary":false,"summary":null},{"id":"2608.24691","version":1,"title":"Confident at the moment of action: belief miscalibration in LLM play under hidden information","zh_title":"行动时刻的自信：隐藏信息下LLM对弈中的信念误校准","abstract":"Agentic systems increasingly gate actions on a model's own stated confidence, which assumes confidence tracks correctness at the moment of acting. We test this in a hidden-information chess variant where royal status can be secretly, repeatedly relocated between pieces, and where an agent's stated probability distribution over the opponent's hidden royal piece -- elicited every turn, separately from the move it chooses -- is scored against ground truth recoverable after the game. Across two independent batches, captures made at high stated confidence ($\\geq 0.5$) about the hidden piece's location were correct in 1 of 62 cases. The calibration deficit is concentrated almost entirely in these events: 99.3% of it in the original batch, 98.7% in the replication. The same pattern, in weaker form, orders consistently (point estimates only; most pairwise gaps are not statistically distinguishable at this sample size) across four further model configurations spanning a second provider -- reported as scope for the finding, not as evidence that capability predicts calibration: a same-model comparison at a fixed external leaderboard score shows a deliberation-budget change alone moves the metric by nearly as much as a large cross-model gap. In a separate seat, conventional evaluation axes -- legality, cost, latency, completion rate -- can dissociate entirely from belief quality, with the configuration winning on every conventional axis producing the worst belief quality tested. A model exhibiting this pattern can still win the game its belief was about, which is why outcome-only evaluation would not detect it.","authors":["Bhushan Kashinath Joshi"],"categories":["cs.AI","cs.CL","cs.LG"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-08-26","first_seen":"2026-08-26","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.24691","pdf_url":"https://arxiv.org/pdf/2608.24691","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C2"],"tags":["LLM置信度校准","游戏AI","隐藏信息博弈"],"reason":"研究LLM在隐藏信息棋类游戏中的置信度校准，属于游戏AI评估，不涉及人类行为仿…","model":"deepseek-v4-pro","scored_at":"2026-08-26T13:02:17","error":null,"has_summary":false,"summary":null},{"id":"2608.24748","version":1,"title":"Method, Mind, and Morality: How People Make Sense of Artificial Intelligence","zh_title":"方法、心智与道德：人们如何理解人工智能","abstract":"How can humans make sense of the rapid takeoff of artificial intelligence (AI)? We studied the sensemaking dynamics of AI through an open-ended, mixed-methods study with computational text analysis of millions of AI-related newspaper articles and social media posts grounded in 57 semi-structured interviews with AI professionals in 2021 and 2023--before and after the recent surge of public interest. We identify a range of sociological frames (interpretive schemas that structure collective cognition) and show how AI professionals use frames to address significant cognitive challenges, such as assigning responsibility for societal impacts. We develop a framework of three primary debates across which frames are adopted and contested: (i) the $\\textit{method}$ of AI development, between frames of top-down expert systems and bottom-up emergent capabilities, (ii) the $\\textit{mind}$ of an AI system, ranging from a passive tool to a humanlike \"digital mind,\" and (iii) the $\\textit{morality}$ of how AI is used, particularly the decision of whether to slow down or speed up AI development. As humanity enters the era of transformative AI, technologists and policymakers must account for the framing dynamics that will circumscribe our beliefs, values, and actions.","authors":["Jacy Reese Anthis","Erik Brynjolfsson","James Evans"],"categories":["cs.CY","cs.AI","cs.CL","cs.LG","stat.ML"],"primary_category":"cs.CY","announce_type":"cross","date":"2026-08-26","first_seen":"2026-08-26","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.24748","pdf_url":"https://arxiv.org/pdf/2608.24748","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C5"],"tags":["AI社会学","框架分析","人机交互"],"reason":"研究人类如何理解AI，不涉及用LLM仿真人类被试，方向相反","model":"deepseek-v4-pro","scored_at":"2026-08-26T13:02:36","error":null,"has_summary":false,"summary":null},{"id":"2608.23932","version":1,"title":"Evolutionary Recurrent Decision Model in Developing Adaptive and Maladaptive Behaviors","zh_title":"演化循环决策模型在适应性与非适应性行为发展中的应用","abstract":"This study introduces the evolutionarily recurrent decision model (ERDM), a computational reinforcement learning framework designed to examine how evolutionary mismatch, bounded rationality, and satisficing contribute to adaptive and maladaptive behavior. ERDM simulates agents across evolutionary recurrent environments, including threat, prey/goal-pursuits, and alliances. Agents learn through competing rewards abstracted from survival metrics. A validity study under varying adverse childhood experiences demonstrates that distinct adaptive and maladaptive strategies, such as learned helplessness, avoidance, healthy relationships, and aggression, emerge naturally without being hardwired. These results align with empirical literature, showcasing ecological validity. The results suggest that many psychopathology-relevant aspects may be interpreted as bounded cognitive systems operating under modern-ancestral environmental mismatch, positioning ERDM as a key computational cognitive tool that can be extended to other studies.","authors":["Andrew Hu"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-26","first_seen":"2026-08-26","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.23932","pdf_url":"https://arxiv.org/pdf/2608.23932","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["强化学习","多智能体模拟","计算精神病学"],"reason":"多智能体强化学习模拟，无LLM，无人类数据对照，不涉及人类仿真","model":"deepseek-v4-pro","scored_at":"2026-08-26T13:02:25","error":null,"has_summary":false,"summary":null},{"id":"2608.24569","version":1,"title":"When \"Must\" Becomes \"Maybe\": Constraint Weakening in LLM Agent Workflows","zh_title":"当“必须”变成“也许”：LLM智能体工作流中的约束弱化","abstract":"Large language model (LLM) agents coordinate complex tasks through multi-role and multi-stage workflows. Upstream state is repeatedly transformed into intermediate language artifacts, such as summaries, plans, tickets, memories, and handoff notes, from which downstream components act. For action-constraining state, topical retention is insufficient: an artifact may mention an unresolved condition while changing it from a requirement that must be resolved before execution into information that may merely inform the next action. We study this action-binding role as operational state preservation. Safety blockers provide a controlled instance because each source state has an explicit prerequisite, authority, fallback, and execution consequence. We condition on correct upstream identification, vary the handoff transformation, and evaluate an executor restricted to the resulting artifact. Across 1,296 controlled synthetic episodes, direct-handoff controls preserve every blocker, whereas compression, plan assimilation, convergence, ownership deferral, and precedent substitution repeatedly turn binding state into caveats or non-binding considerations. Normal handoff compression produces 100.0% deactivation and 54.2% forbidden action. Restoring all four state fields raises preservation to 100.0% and reduces forbidden action to 0.0%. Fixed-artifact interventions further separate preservation from containment: downstream verification eliminates forbidden action while artifact deactivation remains 95.3%. These results identify a state-transmission failure between information extraction and action. Handoff transformations can retain state content while weakening its constraints on downstream action. Semantic availability does not guarantee operational preservation.","authors":["Yiheng Sun","Huifei Wang","Yancheng Zhu","Zhenyu Li","Zebin Zhao","Yifan Yuan"],"categories":["cs.AI","cs.MA"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-26","first_seen":"2026-08-26","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.24569","pdf_url":"https://arxiv.org/pdf/2608.24569","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","工作流可靠性","状态保持"],"reason":"纯多智能体工作流研究，关注LLM agent间状态传递与约束保持，不涉及人类行…","model":"deepseek-v4-pro","scored_at":"2026-08-26T13:02:33","error":null,"has_summary":false,"summary":null},{"id":"2608.23593","version":1,"title":"Fidelity Preference, Not Demographic Preference: A Pixel-Level Attribute-Sensitivity Audit of Image Aesthetic/Preference Scorers","zh_title":"保真度偏好而非人口统计偏好：图像美学/偏好评分器的像素级属性敏感性审计","abstract":"Text-to-image systems use learned aesthetic scorers to filter training data and guide generation, but whether these scores encode demographic attributes as objective quality is unclear. We audit four scorers (LAION-Aesthetics, PickScore, ImageReward, HPSv2) using pixel-level interventions on skin tone and body type in synthetic and real images. Our key finding is that along skin-lightness, the dominant effect is fidelity preference: unaltered images score highest, and perturbations in either direction are penalized (inverted-U). Placebo arms show this penalty is not an artifact of the skin operator, as applying the same CIELAB L* shift to non-skin regions yields similar penalty magnitudes. However, the penalty is operator-dependent and holds for all operators only for LAION-Aes. Critically, audits on synthetic images alone are misleading: LAION-Aes shows strong preference for darker skin on synthetic faces, but on 1470 real faces the preference reverses and becomes much smaller, and amplification becomes non-significant. Across scorers, synthetic results do not transfer -- reversing for LAION-Aes and HPSv2, attenuating for PickScore. We contribute a reproducible benchmark with artifact control and synthetic/real cross-validation, and an auditability criterion for pixel-level causal isolation (valid for skin tone, not for body type due to deformation). Population-stratified analysis shows fidelity-penalty asymmetry is not robust across groups after FDR correction except for HPSv2. Our findings show naive synthetic audits misjudge bias direction and magnitude, and only within-image causal isolation on real data can distinguish true demographic bias from fidelity preference.","authors":["Mingyang Xu"],"categories":["cs.CV","cs.AI"],"primary_category":"cs.CV","announce_type":"cross","date":"2026-08-26","first_seen":"2026-08-26","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.23593","pdf_url":"https://arxiv.org/pdf/2608.23593","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C2"],"tags":["图像美学评分","偏见审计","计算机视觉"],"reason":"审计图像美学评分器，不涉及LLM仿真人类被试或人类行为对照","model":"deepseek-v4-pro","scored_at":"2026-08-27T13:03:12","error":null,"has_summary":false,"summary":null},{"id":"2608.23860","version":1,"title":"Revelation Control","zh_title":"揭示控制","abstract":"Revelation Control is the problem of choosing priced interventions that reveal hidden state only insofar as the revealed distinctions can change a consequential decision, while accounting separately for any useful progress created by the intervention itself. We develop this theory for learning systems, where states equivalent under declared current information can respond differently to future training and favor different actions. The framework defines decision-sufficient revelation and revelation depth, separates pure information value from productive reuse, embeds static Bayes refinement into state-dependent continuation value, and gives an exact cost-adjusted factorization criterion: an additional shallow coordinate is decision-nonredundant only when states sharing a scalar summary lie on opposite sides of the priced Stop/Continue boundary. We also give a target-independent protocol for model-specific instantiation and prove that bounded stop-flip risk alone cannot certify positive expected utility under unrestricted severity. Across Qwen2.5-7B and Mistral-7B-v0.3, deeper future-learning probes have positive decision value and productive reuse yields strict equal-compute utility advantages. Qwen additionally provides evidence for a decision-nonredundant shallow revealability regime; in Mistral, a scalar continuation architecture fit only on an independent development panel retains positive familywise-adjusted lower bounds on a disjoint target panel, consistent with scalar decision sufficiency within the tested architecture family and resolution. The evidence supports structural rather than numerical transfer: the decision theory, cost accounting, continuation logic, and evaluation protocol transport, while empirical proxies, coefficients, thresholds, and even the required shallow state dimension may be system-specific.","authors":["Qinyou Wang"],"categories":["cs.LG","cs.AI","stat.ML"],"primary_category":"cs.LG","announce_type":"cross","date":"2026-08-26","first_seen":"2026-08-26","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.23860","pdf_url":"https://arxiv.org/pdf/2608.23860","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["信息揭示","决策理论","学习系统"],"reason":"研究学习系统中的信息揭示控制，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-08-27T13:03:14","error":null,"has_summary":false,"summary":null},{"id":"2608.24133","version":1,"title":"PlaceSeek: Human-Centered Geospatial Retrieval of Urban Outdoor Places via Semantic Grounding and Affective Alignment","zh_title":"PlaceSeek：通过语义接地与情感对齐实现以人为本的城市户外场所地理空间检索","abstract":"People search for urban outdoor places not only by category or function, but also by what activities a place can support and how it is perceived. Existing geospatial retrieval remains largely POIcentric and metadata-driven, making it difficult to satisfy openended, affective, or activity-oriented needs. We present PlaceSeek, a human-centered outdoor place retrieval framework that maps natural-language queries to geolocated street-view imagery. PlaceSeek introduces an intent-aware retrieval mechanism that decomposes user queries into functional and affective sub-intents. A Semantic Grounding Module verifies whether candidate street-view results contain the physical evidence needed to support the intended activity, while an Affective Alignment Module re-ranks physically valid candidates using a LoRA-adapted vision-language model trained on human urban perception judgments. We evaluate PlaceSeek on 31,956 street-view locations in Milan across 10 naturallanguage queries annotated by five human evaluators. PlaceSeek achieves 88.0% Precision@5, a mean match score of 3.39/4.0, and 0.920 nDCG@5, outperforming CLIP, fine-tuned CLIP, SigLIP, and a VQA-based baseline. Ablation results show that physical grounding is essential for retrieval validity, while affective alignment improves ranking quality among physically valid candidates. These findings highlight that complex urban spatial queries require modeling both verifiable visual evidence and human perceptual preferences. PlaceSeek provides a potential framework for human-centered nextgeneration geospatial retrieval systems.","authors":["Ziqi Cui","Shangyu Lou"],"categories":["cs.CV","cs.AI","cs.IR"],"primary_category":"cs.CV","announce_type":"cross","date":"2026-08-26","first_seen":"2026-08-26","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.24133","pdf_url":"https://arxiv.org/pdf/2608.24133","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C2"],"tags":["地理空间检索","计算机视觉","情感对齐"],"reason":"论文是地理空间检索系统，不涉及LLM仿真人类被试，属于计算机视觉应用。","model":"deepseek-v4-pro","scored_at":"2026-08-27T13:03:16","error":null,"has_summary":false,"summary":null},{"id":"2608.23968","version":1,"title":"When LLMs Slow Down: How Environmental Impacts Mediate University Students' LLM Usage","zh_title":"当LLM变慢：环境影响如何调节大学生的LLM使用","abstract":"Large Language Models (LLMs) are increasingly being embedded into all facets of society, from search to education, industrial, and financial applications. These systems' carbon and water footprints raise important sustainability concerns, particularly with adoption rates exceeding 80% among university students, despite limited insight into the environmental impacts of individual usage. Eco-feedback interfaces offer a promising approach to encourage more sustainable behaviors, yet their role in shaping LLM users' sustainability awareness and decision-making remains underexplored. We design and deploy the interface that visualizes latency-carbon trade-offs during live LLM interactions. We study its use with undergraduate computer science students (N=89, ages 18-24), enrolled in a computing ethics course, providing an empirical look at how a technically sophisticated and values-oriented user population responds to sustainability-aware AI interfaces. We found that the likelihood of choosing the eco-feedback system significantly decreased as perceived response latency increased (p < .001), while users' willingness increased when they recognized the carbon-saving impacts (p < .01). Also, students with stronger eco-mindedness demonstrated higher baseline willingness to adopt lower-carbon modes and reported increased awareness of the environmental impacts of LLM use, though this effect diminished as latency increased. These results position eco-feedback interfaces as a promising sustainability intervention and highlight their potential as an educational opportunity to promote more sustainable LLM use among university students and beyond.","authors":["Hyeonwook Kim","Xuesi Chen","Alex Cabral","Cindy Kaiying Lin","Udit Gupta","Josiah Hester"],"categories":["cs.HC","cs.CY"],"primary_category":"cs.HC","announce_type":"new","date":"2026-08-26","first_seen":"2026-08-26","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.23968","pdf_url":"https://arxiv.org/pdf/2608.23968","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C2"],"tags":["人机交互","可持续性","用户行为"],"reason":"研究人类对LLM延迟与碳足迹的反馈，非用LLM仿真人类被试","model":"deepseek-v4-pro","scored_at":"2026-08-26T13:02:26","error":null,"has_summary":false,"summary":null},{"id":"2608.24224","version":1,"title":"Aura: Dynamic Intra-Turn Emotion-Aware Adaptation of Large Language Model Responses","zh_title":"Aura：大语言模型响应的动态轮内情感感知适应","abstract":"Effective human-AI interaction requires systems that dynamically adapt to a user's behavior and evolving understanding. When users interact with Large Language Models (LLMs), these models typically respond to prompts without sensing the user's immediate reactions. This lack of communicative synchrony can lead to information overload or leave confusion unresolved in real time. In this paper, we introduce Aura, a framework that enables LLM systems to dynamically modulate output based on a user's evolving emotions. Aura's Perception Module continuously estimates the user's emotional state from facial expressions. Our Policy Module then selects interventions through a probabilistic belief model. Finally, Aura's Generation Module uses parameter-efficient Low-Rank Adaptation (LoRA) adapters to produce contextually tailored responses mid-turn during response generation. We evaluated Aura in a within-subjects user study (N=20) on information-seeking tasks, where it achieved statistically significantly higher normalized perceived learning gains than a Llama-3 baseline and reduced interaction time by 21% relative to existing LLM baselines (GPT-4o, Llama-3). Our results indicate that real-time, context-sensitive interventions can improve learning efficiency and user satisfaction without observable degradation in factual accuracy. Aura thus supports the potential for more responsive and effective human-AI interaction.","authors":["Rachel Schuchert","Christian Holz"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-08-26","first_seen":"2026-08-26","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.24224","pdf_url":"https://arxiv.org/pdf/2608.24224","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["人机交互","情感适应","LLM"],"reason":"研究实时情感适应的人机交互，非用LLM仿真人类被试，无实验或测量目的。","model":"deepseek-v4-pro","scored_at":"2026-08-26T13:02:31","error":null,"has_summary":false,"summary":null},{"id":"2608.24767","version":1,"title":"Shaping the Future of Generative AI for Black Communities: A Frame Analysis of Public Discourse and Empirical Scholarly Research","zh_title":"塑造生成式AI对黑人社区的未来：公共话语与实证学术研究的框架分析","abstract":"As generative AI (genAI) systems become embedded in education, employment, healthcare, and creative industries, the impact and engagement among marginalized groups have become both a widespread discourse and a focus in scholarly research. As a starting point, we examine public discourse and empirical research to explore the impact of genAI systems on Black communities. We conducted a systematic literature review (SLR) of 91 empirical papers alongside a media discourse frame analysis of 28 public resources, applying Entman's framing theory to map how each corpus defines problems, attributes causes, and proposes treatments. Our SLR reveals that scholarly research concentrates heavily on technical bias detection, reducing Blackness to measurable variables rather than engaging with cultural practices, structural conditions, or Black knowledge systems. Our frame analysis reveals that public discourse attributes genAI-related harm to historical and systemic forces, while scholarly research stops its causal accounts at the dataset and its treatment recommendations at technical reform. We demonstrate that this misalignment is structurally produced: anti-Blackness operates simultaneously across both registers, generating a shared evacuation of Black epistemic agency. We argue for frame analysis as an AI ethics methodology capable of surfacing what technical evaluation forecloses.","authors":["Angela D. R. Smith","Gabriella Thompson","Christopher L. Dancy","Mark D\\'iaz","Seyi Olojo","Christina N. Harrington"],"categories":["cs.HC","cs.CY"],"primary_category":"cs.HC","announce_type":"new","date":"2026-08-26","first_seen":"2026-08-26","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.24767","pdf_url":"https://arxiv.org/pdf/2608.24767","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["AI伦理","框架分析","种族偏见"],"reason":"论文是AI伦理与话语分析，不涉及LLM仿真人类被试或行为对照","model":"deepseek-v4-pro","scored_at":"2026-08-26T13:02:36","error":null,"has_summary":false,"summary":null},{"id":"2608.24669","version":1,"title":"Who Falls for SMiSh? Learning Through Survey Data Where to Best Target Awareness Training for Mobile Messaging Attacks","zh_title":"谁会上当受骗？通过调查数据学习移动消息攻击安全意识培训的最佳目标人群","abstract":"As mobile phone adoption has surged, so have scams involving these devices. One such scam, known as SMiShing (or smishing) after Short Message Service (SMS), involves fraudsters sending phishing links via mobile texts. Despite the prevalence of SMiShing, there is a lack of data on who is most vulnerable to these attacks. Prior research on phishing (its email counterpart) suggests that susceptibility may vary by demographic and contextual factors. In two large-scale surveys, we use a previously published simulation method to collect data from representative samples of U.S. adult mobile phone users. Our findings indicate that younger individuals and college students are particularly vulnerable. Participants struggled to correctly identify legitimate messages, with the second study providing comparisons of financial message variants. Researchers, regulators, and telecoms can help users by creating mobile-specific interventions for under-24 and university customers and adding verifications and warnings.","authors":["Cori Faklaris","Sarah Tabassum","Heather Richter Lipford"],"categories":["cs.CR","cs.HC"],"primary_category":"cs.CR","announce_type":"cross","date":"2026-08-26","first_seen":"2026-08-26","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.24669","pdf_url":"https://arxiv.org/pdf/2608.24669","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C2"],"tags":["网络安全","钓鱼攻击","用户行为"],"reason":"研究人类对短信钓鱼的易感性，未使用LLM仿真人类被试，不涉及LLM作为替代品。","model":"deepseek-v4-pro","scored_at":"2026-08-26T13:02:36","error":null,"has_summary":false,"summary":null},{"id":"2608.24554","version":1,"title":"Why fragmented parliaments stop passing legislation: Opposition discipline and representation across four democratic institutions","zh_title":"为何碎片化议会停止立法：四种民主制度下的反对党纪律与代表性","abstract":"Parliamentary systems pass more bills than presidential systems at baseline, but collapse to near-zero passage under party-system fragmentation. The literature offers three competing micro-explanations: coalition-formation failure, party discipline, and committee gatekeeping. These operate simultaneously in any real legislature, so observational studies struggle to separate their contributions. We present an agent-based model that compares four democratic institutions: pure parliamentary, pure republican/presidential, premier-presidential (France), and president-parliamentary (Russia). Across four scenarios and N=200 seeds per cell we report bootstrap confidence intervals, Morris screening, Sobol variance decomposition, mechanism ablations, and a hung-parliament variant comparison. Three findings emerge. First, government formation failure alone does not halt legislation: when a fragmented parliament reverts to personal voting, parliamentary passage (46.4%) is statistically indistinguishable from the presidential benchmark (44.8%); collapse requires cohesive opposition obstruction, which drives passage to 0.05%. Second, disabling discipline restores fragmented passage to 46.7%, and the rescue magnitude is monotone across the four institutions in a pattern that survives varying the common discipline level. Third, the passage-representation tradeoff is a single spectrum: parliamentary maximises throughput at the cost of representational fidelity; republican maximises fidelity via the presidential veto; semi-presidential variants split the difference.","authors":["Fuad Ali"],"categories":["physics.soc-ph","cs.MA"],"primary_category":"physics.soc-ph","announce_type":"cross","date":"2026-08-26","first_seen":"2026-08-26","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.24554","pdf_url":"https://arxiv.org/pdf/2608.24554","source_feed":"cs.MA","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["基于主体的建模","政治制度比较","立法过程"],"reason":"基于规则的ABM模拟政治制度，无LLM参与，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-08-26T13:02:33","error":null,"has_summary":false,"summary":null},{"id":"2608.24457","version":1,"title":"Participation, selection and indicative bidding in auctions with costly entry","zh_title":"有成本进入拍卖中的参与、选择与指示性投标","abstract":"We study auctions with costly entry in which bidders have incomplete information about their affiliated private valuations prior to entry. We examine whether indicative bidding--a mechanism requiring non-binding preliminary bids before entry--can stimulate participation and improve entrant selection, and we compare its performance with unrestricted and capped entry in a controlled laboratory experiment. When entry costs are high, indicative bidding generates significantly more revenue, primarily by increasing participation beyond theoretical predictions. When entry costs are low, its predicted revenue advantage is attenuated by higher-than-predicted selection inefficiency.","authors":["Changxia Ke","Greg Kubitz","Yang Liu"],"categories":["econ.GN","q-fin.EC"],"primary_category":"econ.GN","announce_type":"new","date":"2026-08-26","first_seen":"2026-08-26","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.24457","pdf_url":"https://arxiv.org/pdf/2608.24457","source_feed":"econ.GN","score":0,"bucket":"other","rubric_hits":["C2"],"tags":["拍卖理论","实验经济学","机制设计"],"reason":"论文是实验室人类实验，研究拍卖机制，不涉及LLM或人类仿真。","model":"deepseek-v4-pro","scored_at":"2026-08-26T13:02:31","error":null,"has_summary":false,"summary":null},{"id":"2608.24851","version":1,"title":"Learning Whom to Trust : Decision-Generated Credibility in Social Learning","zh_title":"学习信任谁：社会学习中的决策生成可信度","abstract":"Social interaction can improve collective learning but also amplify early mistakes. We study this tension when the credibility of social information is generated by the sender's own decision process rather than fixed ex ante. Reinforcement-learning agents make binary choices through a drift--diffusion process that jointly determines choice, decision time, and confidence; decision confidence then becomes social credibility by weighting anticipatory influence and retrospective social learning. Under balanced community exposure, the anticipatory field admits an exact quotient representation. Its local Jacobian is a scalar decision-sensitivity term multiplying the community-coupling matrix, which yields a common-mode amplification threshold and an analytical role for cross-community permeability in damping relative community differences. Monte Carlo experiments show the corresponding non-monotone performance pattern: moderate transmission accelerates correction, whereas strong transmission can lock populations into wrong consensus; low permeability instead sustains disagreement. Ablations reveal a dual role for confidence: credibility-sensitive transmission amplifies social error, while confidence-dependent private learning stabilises it. The model yields testable predictions linking sender confidence to receiver behaviour conditional on accuracy.","authors":["Gabriel Bontemps","Abhishek Banerjee"],"categories":["cs.NE","econ.GN","q-fin.EC"],"primary_category":"cs.NE","announce_type":"cross","date":"2026-08-26","first_seen":"2026-08-26","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.24851","pdf_url":"https://arxiv.org/pdf/2608.24851","source_feed":"econ.GN","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","强化学习","社会学习"],"reason":"多智能体强化学习社会学习模型，无LLM，无人类数据对照，不涉及人类仿真","model":"deepseek-v4-pro","scored_at":"2026-08-26T13:02:38","error":null,"has_summary":false,"summary":null},{"id":"2608.24811","version":1,"title":"How does hazard exposure influence job choice? Evaluating time-dependent tradeoffs between salary and hazard risks","zh_title":"灾害暴露如何影响职业选择？评估薪资与灾害风险之间的时间依赖权衡","abstract":"Natural hazards are a nonmarket disamenity that affects an individual's search for employment resulting in a negative environmental impact that produces an economic inefficiency. We develop a seek-and-screen job search approach that uses a discrete choice simulation to examine how salary, crime, and natural hazard risk influence job choice. We model the job decision process as a series of elimination events using a Cox hazard model grounded in a Random Utility Model. We use data from the discrete choice simulation to estimate both a standard proportional hazards model and an extended specification that allows the effect of natural hazard risk to vary across decision rounds. Individuals are exposed to the dynamics of a simulated job search as they make decisions between pairs of job offers in an adaptive learning process based on income, geography, crime level, and natural hazard attributes. The results of the job choice decisions provide the input to a statistical survival analysis. The results indicate that salary and crime exert stable and economically intuitive effects on job elimination, with higher salary reducing and higher crime increasing the likelihood of removal. In contrast, natural hazard risk exhibits a time-varying effect that increases the probability of elimination in early rounds but becomes neutral or favorable in later stages of the decision process. These findings suggest that environmental risk is evaluated differently as individuals transition from initial screening to final job selection, highlighting the importance of modeling job choice as a multi-stage process.","authors":["Richard Bernknopf","Leila Gonzales","Christpher Keane"],"categories":["econ.EM"],"primary_category":"econ.EM","announce_type":"new","date":"2026-08-26","first_seen":"2026-08-26","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.24811","pdf_url":"https://arxiv.org/pdf/2608.24811","source_feed":"econ.EM","score":0,"bucket":"other","rubric_hits":["C2"],"tags":["职业选择","离散选择模拟","生存分析"],"reason":"使用离散选择模拟和人类被试数据，但未使用LLM，不涉及LLM仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-08-26T13:02:37","error":null,"has_summary":false,"summary":null},{"id":"2608.23005","version":1,"title":"Large language models simulate intersectional synthetic identities with a budget of one to two dimensions","zh_title":"大语言模型以一到两个维度的预算模拟交叉性合成身份","abstract":"Large language models are increasingly used as synthetic survey respondents, promising cheap access to rare intersectional populations. We test standard demographic-persona methods against every real intersectional subgroup across 15 waves of Pew's American Trends Panel -- 21 million simulated response distributions from eight models. In real respondents, subgroup opinion is approximately the additive sum of its single-identity components, yet grows 2.5x more distinctive as identities intersect. Simulated respondents show no such composition: a single feature explains a two-feature persona's responses better than the additive combination in 75-82% of subgroups, and a third feature adds almost nothing. This collapse survives every prompting strategy we test. Additionally, the feature models retain is chosen nearly blindly -- except that they systematically discard race and religion, the strongest real drivers of opinion. Synthetic samples offer intersectional personas but represent one identity at a time.","authors":["Virgile Rennard","Christos Xypolopoulos"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2026-08-25","first_seen":"2026-08-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.23005","pdf_url":"https://arxiv.org/pdf/2608.23005","source_feed":"cs.CY","score":10,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","交叉性","算法保真度"],"reason":"直接测试LLM作为合成调查受访者，并与真实Pew数据对照，发现仿真失效条件。","model":"deepseek-v4-pro","scored_at":"2026-08-25T13:03:34","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-25","rank":1,"question":"LLM 在模拟交叉身份（如黑人共和党人）的调查回答时，是否真正整合了多个身份特征，还是只保留其中一个？","design":"用 8 个 LLM（从 7B 开源模型到前沿系统）扮演具有 1 到 3 个特征的人口统计画像，生成 15 波 Pew 美国趋势面板中所有问题的回答分布，共 2100 万次模拟；通过比较单特征、双特征和三特征画像的模拟偏差与真实子群体偏差，检验模型是否对多个身份特征进行加性组合。","baseline":"Pew 美国趋势面板 15 波调查的受访者微观数据，计算每个真实交叉子群体（至少 20 名受访者）的回答分布作为基准。","findings":"真实子群体的意见近似于其单身份成分的加性组合，且随身份交叉而更加独特；但模拟受访者没有这种组合，一个特征就能比加性组合更好地解释双特征画像的回答，第三个特征几乎不增加信息。模型保留的特征几乎是盲目选择的，且系统性地丢弃了种族和宗教这两个真实意见的最强驱动因素。","reliability":"论文通过人类抽样噪声下限、相似度度量的零校准和完全分半样本确认来确保结果稳健；但未讨论提示策略之外的失效条件，也未涉及开放式文本或非分布层面的交叉性。","relevance":"该研究直接测试 LLM 作为合成调查受访者的可靠性，并与真实 Pew 数据对照，发现仿真在交叉身份下系统性失效，对关注仿真偏差和失效条件的研究者极具参考价值。","inspiration":"借鉴其用真实微观数据构建交叉子群体基准、并通过偏差签名竞争来识别模型实际使用的特征的方法｜可迁移到信贷审批歧视研究，检验 LLM 模拟的交叉群体（如黑人女性）的信贷决策是否只基于单一特征｜用 LLM 扮演不同种族和性别的贷款申请人，生成信贷审批决策，与真实信贷数据（如 HMDA）中对应交叉群体的审批率分布进行对照，检验模型是否丢弃了种族或性别信息。"}},{"id":"2608.22582","version":1,"title":"Hybrid Panels: Toward Human-AI Collaboration in Survey Research","zh_title":"混合面板：迈向调查研究中的AI协作","abstract":"Large-scale population surveys are essential for generating robust social and scientific insights, yet they face significant challenges, including declining response rates, increasing data collection costs, long delays between data collection and data provision, and the risk of nonresponse bias. Advances in artificial intelligence (AI) have opened up new opportunities for AI-supported survey infrastructures where the goal is to overcome these challenges without limiting the data quality. A promising AI-enabled survey infrastructure for which we build a first pilot is a hybrid panel. A hybrid panel is a longitudinal AI-enabled survey which allows to iteratively improve the alignment between large language models (LLMs) and the population they aim to simulate and use the errors to inform the design and implementation of the next survey wave (e.g., inform the participant recruitment, assignment of questions to participants). It incorporates both human participants and LLMs as fundamental elements of its design. In this research note, we introduce the concept of a hybrid panel by providing a definition and outlining an overarching framework, spanning data collection to data validation. We detail results from a first pilot study to illustrate (open) challenges that we identify for hybrid panels.","authors":["Julia Romberg","Tobias Gummer","Gabriella Lapesa","Tanja Kunz","Claudia Wagner"],"categories":["cs.CL","cs.AI","cs.CY","cs.HC"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-25","first_seen":"2026-08-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.22582","pdf_url":"https://arxiv.org/pdf/2608.22582","source_feed":"cs.CL","score":9,"bucket":"selected","rubric_hits":["A1","A2","A3","A4","B1","B2","B3"],"tags":["LLM仿真","调查方法","人机协作"],"reason":"提出混合面板，用LLM模拟调查对象并与人类数据对照，直接相关。","model":"deepseek-v4-pro","scored_at":"2026-08-25T13:03:32","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-25","rank":4,"question":"如何设计一种结合人类被试与LLM的纵向调查基础设施（混合面板），以在保证数据质量的同时应对传统调查的挑战？","design":"提出混合面板概念，即纵向调查中同时纳入人类被试和LLM，通过计划缺失设计、LLM插补、人类验证AI生成回答等方式结合两者；先导研究聚焦人类被试招募步骤，未报告完整实验处理与结果变量。","baseline":"无对照（先导研究仅涉及招募，未提供人类与LLM回答的直接对比数据）。","findings":"论文提出了混合面板的定义和框架，并展示了先导研究中人类被试招募的初步结果；识别出混合面板面临的开放性挑战，包括人类被试的偏好与权利、AI模型的持续评估以及社区标准制定。","reliability":"论文承认LLM与人类行为存在不匹配，完全合成面板面临透明度、问责制和伦理问题；强调AI不能完全取代人类作为研究对象，混合面板的适用性需要长期评估和实验。","relevance":"该研究直接针对LLM仿真人类调查的可靠性问题，提出混合面板以结合人类与AI数据，并强调持续验证，对关注仿真偏差和真实人类对照的研究者具有重要参考价值。","inspiration":"可借鉴其混合面板设计，将LLM生成回答与人类被试数据结合，通过计划缺失和迭代验证提高仿真准确性｜可迁移到经济预期调查或消费者信心指数构建，利用LLM补充缺失回答并校准偏差｜设计一个纵向调查，招募真实消费者作为被试，部分问题由LLM回答，处理为不同提示策略，结果变量为回答与真实值的偏差，用官方统计或面板数据做对照。"}},{"id":"2608.21668","version":1,"title":"From Mastery Profile to Simulated Response: Stochastic Student Knowledge Graphs (SSKG) for Faithful LLM Student Simulation","zh_title":"从掌握水平画像到模拟响应：用于忠实LLM学生仿真的随机学生知识图谱","abstract":"Large language models (LLMs) are increasingly used to simulate students at different mastery levels. These simulations can generate synthetic training data and stress-test tutoring systems. However, common prompt-based approaches leave the answer decision to the LLM, which tends to perform according to its built-in capabilities even when instructed to simulate a student with low mastery. As a result, these approaches may have difficulty distinguishing students with low and high levels of mastery. We demonstrate this limitation using 379 College Board-calibrated SAT Algebra items and five archetypal mastery profiles. Three LLMs from three vendors (Gemini 3.1 Flash Lite, Claude Haiku 4.5, and GPT-5.4-mini) achieve 96.8-100% accuracy across all profiles. To address this limitation, we introduce a method grounded in a Stochastic Student Knowledge Graph (SSKG). A curriculum knowledge graph (CKG) is extracted from an open algebra textbook, and each SAT solution is decomposed into a chain of required triples. The SSKG assigns a mastery probability to each triple, which is sampled to determine question correctness. An LLM then generates a first-person rationale consistent with the outcome. The simulation reduces accuracy to 44.1-85.2% across profiles and produces a clear monotone mastery gradient.","authors":["Yuan An","Emily Wang","Benjamin Wang","Ruhma Hashmi"],"categories":["cs.AI","cs.HC"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-25","first_seen":"2026-08-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.21668","pdf_url":"https://arxiv.org/pdf/2608.21668","source_feed":"cs.AI","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM学生仿真","知识图谱","教育评估"],"reason":"用LLM仿真不同掌握水平的学生，并与真实SAT数据对照，评估仿真保真度并指出提…","model":"deepseek-v4-pro","scored_at":"2026-08-25T13:03:28","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-25","rank":2,"question":"如何让LLM忠实模拟不同掌握水平的学生，使其答题准确率呈现与掌握水平一致的梯度，并产生可归因于具体知识点的错误？","design":"用三种商用LLM（Gemini 3.1 Flash Lite、Claude Haiku 4.5、GPT-5.4-mini）模拟五种掌握水平的学生（从接近专家到严重知识缺口），在379道SAT代数题上作答。先测试直接提示法（在提示中描述学生水平），再提出基于随机学生知识图谱（SSKG）的方法：从代数教材构建课程知识图谱，将每道题的解题过程分解为所需三元组链，为每个三元组赋予掌握概率，通过采样决定答题正确性，再由LLM生成与结果一致的第一人称解释。通过四个累积消融臂（单次采样、检索/执行分解、分类加权、干扰项路由）评估各机制贡献。","baseline":"无直接人类对照，但使用379道College Board校准的SAT代数题作为题目基准，题目难度和区分度经过真实考生数据校准。","findings":"直接提示法下，三个LLM在所有掌握水平上的准确率高达96.8%-100%，无法区分高低水平学生。SSKG方法将准确率降至44.1%-85.2%，并产生清晰的单调掌握梯度，且错误可归因于特定知识点。","reliability":"论文承认技能特异性和涌现难度的结果较为混合：P4显示早期/后期知识点的分化，P5受链长影响，且仅部分消融臂达到预期。此外，SSKG方法依赖于人工构建的课程知识图谱和解题链，可能引入主观性，且仅在SAT代数题上验证，泛化性未知。","relevance":"该研究直接针对LLM仿真人类被试的保真度问题，通过引入外部知识结构控制LLM行为，克服了提示法中的能力偏差，对评估仿真可靠性和设计更可控的仿真方法有重要参考价值。","inspiration":"借鉴其将决策过程分解为知识单元并显式采样控制行为的方法，可迁移到经济金融中的个体决策仿真，如消费者跨期选择或投资者风险偏好。｜例如，在信贷审批歧视研究中，可构建金融知识图谱，将贷款决策分解为所需金融概念，为不同金融素养水平的虚拟申请人赋予掌握概率，通过采样决定其决策结果，再让LLM生成解释。｜设计：以LLM模拟不同金融素养的贷款申请人，处理是金融素养水平（通过知识图谱掌握概率设定），结果变量是贷款申请决策（是否违约或选择何种贷款），对照真实数据可使用美国消费者金融保护局（CFPB）的投诉数据或某银行的历史贷款数据。"}},{"id":"2608.22438","version":1,"title":"When Persona Simulations Are Informative: Graph-Structured Signals for Pluralistic Opinion Sensing","zh_title":"当人格模拟具有信息量时：用于多元意见感知的图结构信号","abstract":"Persona-conditioned large language models (LLMs) are increasingly used to simulate survey responses across diverse domains. However, apparent response variation can reflect unconditioned model priors or token sampling noise rather than systematic persona conditioning. We argue that persona-conditioned variation is informative when semantically similar personas exhibit concordant response shifts. To operationalize this principle, we introduce Persona-Conditioned Informativeness (PCI), an unsupervised diagnostic metric that measures whether semantically similar personas deviate in concordant directions relative to item-level sample baselines. By modeling personas as a similarity graph, PCI uses Local Moran's I to quantify local spatial coherence and extract compact persona subsets without using construct labels. To evaluate PCI without external human benchmarks, we test its ability to recover established latent value structure using the 57-item Portrait Values Questionnaire-Revised (PVQ-RR). Confirmatory factor analysis (CFA) shows that a PCI-selected 10% subset substantially improves overall construct recovery relative to response-stability and random selection. These findings support PCI as a principled internal diagnostic for screening synthetic respondents in survey pipelines.","authors":["Taehyeon An","Jaehyeong Park","Donghyuk Shin"],"categories":["cs.AI","cs.CY"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-25","first_seen":"2026-08-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.22438","pdf_url":"https://arxiv.org/pdf/2608.22438","source_feed":"cs.AI","score":9,"bucket":"selected","rubric_hits":["A1","A2","A4","B1","B4"],"tags":["LLM仿真","调查方法","算法保真度"],"reason":"用LLM模拟调查回答，提出诊断指标筛选合成被试，并用真实人类数据验证。","model":"deepseek-v4-pro","scored_at":"2026-08-25T13:03:31","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-25","rank":3,"question":"如何判断基于人物设定的大语言模型生成的调查回答是否真正反映了人物设定，而非模型先验或采样噪声？","design":"使用 K-EXAONE-236B 模型，为 1480 个合成人物设定生成对 57 项 PVQ-RR 价值观问卷的回答；通过构建人物相似度图并计算局部空间相干性（PCI 指标）来筛选信息量高的子集。","baseline":"无外部人类基准；以 PVQ-RR 的潜在价值结构作为内部验证标准。","findings":"PCI 筛选出的 10% 子集在验证性因子分析中显著提升了潜在构念恢复度，优于响应稳定性选择和随机选择。这表明语义相似的人物设定在回答上呈现一致的偏移，可作为筛选合成被试的有效内部诊断。","reliability":"论文指出 PCI 仅提供内部结构诊断，不能保证外部总体有效性，需结合外部校准；当前图构建基于全局嵌入，可能忽略不同题目涉及的属性子集；仅在单一模型和单一问卷上验证，跨模型和跨语言泛化性未知。","relevance":"该研究直接针对 LLM 模拟调查回答的可靠性问题，提出无监督诊断指标，并用真实人类价值观结构进行验证，对关注仿真有效性和偏差的研究者具有重要参考价值。","inspiration":"借鉴其基于相似度图的局部空间相干性度量，可无监督地识别哪些合成个体对处理变量有系统性响应，避免盲目使用全部生成样本。｜可迁移到经济政策偏好调查或消费者态度仿真中，筛选出对政策参数或产品属性有真实差异化反应的合成被试。｜用 LLM 生成不同人口统计特征的人物设定，施加政策干预（如税收变化），测量其政策支持度，并用真实调查数据（如美国综合社会调查 GSS）校准筛选后的合成样本分布。"}},{"id":"2608.12750","version":2,"title":"PatientAct: Theory-Grounded Mental Health Client Simulation","zh_title":"PatientAct：基于理论的心理健康来访者仿真","abstract":"LLM-based simulated clients are increasingly used to train novice counselors, evaluate LLM therapists, and generate synthetic data. However, current simulators produce overly cooperative clients that disclose too readily, accept therapeutic reframes without resistance, and resolve core issues within a single session. We trace these issues to profiles that lack causal depth and behavioral mechanisms that treat all content as equally accessible. We present PatientAct, a framework for client simulation grounded in established clinical theories. Our profiles integrate the 5Ps clinical case formulation, providing causal depth without tying the design to any single therapeutic modality. During simulation, profiles include a dynamic memory layer in which items carry trust thresholds (e.g., symptoms are available early, whereas formative memories require a sustained therapeutic alliance). At each turn, the client's emotional reaction and behavior are modeled before generating a response. If the therapist approaches gated content, PatientAct expresses resistance in terms of quantity, content, and style rather than defaulting to cooperation or a single resistance pattern. We evaluate our framework on 40 clinical situations and demonstrate that it generates diverse profiles with high clinical plausibility. Moreover, PatientAct significantly outperforms the baselines, yielding substantial gains in resistance quality and behavioral realism. Our code and data are publicly available via github.com/Sahandfer/PatientHub.","authors":["Sahand Sabour","TszYam NG","Yaqian Chen","Guanqun Bi","Jialu Zhao","Minlie Huang"],"categories":["cs.CL","cs.AI","cs.HC"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-08-25","first_seen":"2026-08-14","revised_at":"2026-08-25","abs_url":"https://arxiv.org/abs/2608.12750","pdf_url":"https://arxiv.org/pdf/2608.12750","source_feed":"cs.CL","score":8,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","心理治疗","行为真实性"],"reason":"用LLM模拟心理治疗来访者，有真实临床情境对照，评估行为真实性与抵抗质量，可迁…","model":"deepseek-v4-pro","scored_at":"2026-08-25T13:04:08","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-25","rank":6,"question":"如何设计基于LLM的心理治疗来访者仿真，使其行为更接近真实来访者，特别是能表现出基于信任的披露和多样化的抵抗？","design":"PatientAct框架使用GPT-5.4生成基于5Ps临床案例构想的来访者档案，并加入动态记忆层和信任阈值，在每轮对话中先模拟情绪反应和行为选择，再生成回应；在40个临床情境（抑郁和焦虑各20个）上评估其临床合理性和抵抗质量。","baseline":"无对照","findings":"PatientAct生成的档案具有高临床合理性和多样性，且在抵抗质量和行为真实性上显著优于现有基线。","reliability":"论文未讨论","relevance":"该研究通过理论驱动的档案设计和动态信任机制提升LLM仿真行为真实性，并采用专家评估验证，对关注仿真可靠性和偏差的研究者有参考价值，值得阅读原文了解具体实现和评估细节。","inspiration":"借鉴其将理论构念（如信任阈值、抵抗分类）嵌入仿真机制并设计多维度评估指标的做法｜可迁移到经济金融中的信任与信息披露场景，如消费者对金融顾问的信任建立、投资者对风险信息的逐步接受｜设计一个LLM扮演的投资者，处理为不同信任阈值下的信息提供策略，结果变量为披露意愿和风险感知，对照真实投资者调查或实验数据。"}},{"id":"2608.21401","version":1,"title":"Generative Gap Filling","zh_title":"生成式填补空白","abstract":"Most contract litigation turns on contracts that imperfectly record parties' bargains. When the parties' dispute can't be solved by interpreting the text, courts fill the gap. Scholars have long assumed that the remaining text runs out quickly, and provides thin evidence of the actual deal on the disputed point. On that view, a judge who supplies the missing term must be drawing on something else, from commercial defaults to her own policy preferences. Despite generations of work, courts have no real alternative to such unruly methods. We tested that assumption. Taking real contracts, we masked a term the parties had negotiated and asked readers to predict what we removed. Lay respondents recovered the hidden term about half the time, twice what chance predicts. Law students and lawyers did marginally better. But large language models, given nothing but the rest of the contract, recovered it nearly nine times in ten. The deal, in short, testifies to far more of the agreement than the literature assumes, including terms the parties never wrote. A contract, we argue, is like a radio signal from far away. Even when incomplete, enough of the message is carried elsewhere that the missing part can be reconstructed with the right receiver. True gaps are rarer than supposed. Courts can weigh model predictions as ordinary, contestable evidence, and parties can discipline the practice with \"Choice of Model\" clauses.","authors":["Yonathan A. Arbel","David A. Hoffman"],"categories":["cs.CY","cs.CL"],"primary_category":"cs.CY","announce_type":"cross","date":"2026-08-25","first_seen":"2026-08-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.21401","pdf_url":"https://arxiv.org/pdf/2608.21401","source_feed":"cs.CL","score":8,"bucket":"selected","rubric_hits":["A1","B1","B2"],"tags":["LLM仿真","法律决策","人类对照"],"reason":"用LLM预测人类对合同缺失条款的判断，并与真人对照，属于法律决策仿真。","model":"deepseek-v4-pro","scored_at":"2026-08-25T13:03:27","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-25","rank":7,"question":"合同文本在多大程度上能揭示当事人未明确写出的条款，从而让法院或模型填补合同空白？","design":"从真实合同中遮蔽一个已协商的条款，让普通人、法学生、律师和大型语言模型仅根据合同其余部分预测被遮蔽的条款，比较预测准确率。","baseline":"普通人（Lay respondents）预测准确率约50%，法学生和律师略高，作为人类对照基准。","findings":"大型语言模型仅凭合同其余部分，预测被遮蔽条款的准确率接近90%，远高于人类。合同文本比传统假设包含更多关于未写明条款的信息，真正的合同空白比想象中更少。","reliability":"论文未讨论","relevance":"该研究用LLM模拟人类对合同缺失条款的判断，并与真人对照，属于法律决策仿真，对关注LLM仿真可靠性和偏差的研究者有参考价值，值得阅读原文了解实验细节和局限。","inspiration":"借鉴其遮蔽真实合同条款并让模型预测的设计，可迁移到金融合同或政策文本的缺失条款预测，例如信贷协议中的利率调整条款或政策公告中的具体参数。｜可应用于资产定价实验或信贷审批歧视研究，例如遮蔽贷款合同中的关键条款，检验模型能否预测人类决策者会如何设定或接受这些条款。｜用LLM作为被试，遮蔽真实金融合同中的某个条款，让模型预测该条款内容，并与真实合同条款及人类专家预测对照，结果变量为预测准确率，真实数据来自公开的合同数据库或监管文件。"}},{"id":"2510.05545","version":3,"title":"Can Language Models Boost the Power of Randomized Experiments Without Statistical Bias?","zh_title":"语言模型能否在不引入统计偏差的情况下提升随机实验的功效？","abstract":"Randomized controlled trials (RCTs) are widely adopted for causal inference, yet cost and sample-size constraints limit power. We introduce CALM (Causal Analysis leveraging Language Models), a statistical framework that integrates insights generated by large language models (LLMs) into the analysis of RCTs using established causal estimators to increase precision while preserving statistical validity. In particular, CALM treats LLM-generated outputs as auxiliary prognostic information and corrects their potential bias via a heterogeneous calibration step that residualizes and optimally reweights predictions. We prove that CALM remains consistent even when LLM predictions are biased and achieves efficiency gains over augmented inverse probability weighting estimators for various causal estimands. In particular, CALM develops a few-shot variant that aggregates predictions across randomly sampled demonstration sets. The resulting U-statistic-like predictor restores i.i.d. structure and also mitigates prompt-selection variability. Empirically, in simulations calibrated to a mobile-app depression RCT, CALM delivers lower variance relative to other benchmarking methods, is effective in zero- and few-shot settings, and remains stable across prompt designs. By principled use of LLMs to harness unstructured data and external knowledge learned during pretraining, CALM provides a practical path to more precise causal analyses.","authors":["Xinrui Ruan","Xinwei Ma","Yingfei Wang","Waverly Wei","Jingshen Wang"],"categories":["stat.ME","econ.EM"],"primary_category":"stat.ME","announce_type":"replace-cross","date":"2026-08-25","first_seen":"2025-10-07","revised_at":"2026-08-25","abs_url":"https://arxiv.org/abs/2510.05545","pdf_url":"https://arxiv.org/pdf/2510.05545","source_feed":"econ.EM","score":7,"bucket":"pending","rubric_hits":["A2","B1","B3"],"tags":["因果推断","LLM辅助分析","统计方法"],"reason":"用LLM辅助RCT分析，校正偏差并提升精度，有真实数据对照，方法可迁移到仿真评…","model":"deepseek-v4-pro","scored_at":"2026-08-25T13:04:06","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-25","rank":8,"question":"如何利用大语言模型生成的预测来提升随机对照试验的统计功效，同时避免引入统计偏差？","design":"提出CALM框架，将LLM对潜在结果的预测作为辅助预后信息，通过残差化和异质性校准加权校正偏差，并采用少样本学习与演示集重采样聚合预测，用于RCT的因果效应估计。","baseline":"模拟研究校准自一项移动应用抑郁症RCT的真实数据，并重新评估了一项行为试验。","findings":"CALM在模拟中相比基准方法具有更低的方差，且在零样本和少样本设置下有效，对提示设计具有稳健性。理论上证明即使LLM预测有偏，CALM仍保持一致，并比增强逆概率加权估计器更高效。","reliability":"论文承认LLM预测可能存在偏差且精度异质，但通过残差化和校准加权处理；少样本学习对示例选择敏感，通过重采样聚合缓解。未讨论其他失效条件。","relevance":"该研究将LLM作为辅助工具提升RCT分析精度，而非直接模拟人类被试，但涉及LLM预测的偏差校正和与真实数据的对照，对关注LLM仿真可靠性的研究者有参考价值。","inspiration":"借鉴其残差化与异质性校准加权方法，可校正LLM预测偏差并提升估计效率。｜可迁移到政策评估中利用LLM提取非结构化协变量信息，如文本评论、新闻等，提升处理效应估计精度。｜以LLM对个体潜在结果的预测作为辅助变量，在信贷审批歧视研究中，用真实贷款数据校准LLM预测，比较不同校准方法下的处理效应估计。"}},{"id":"2608.23047","version":1,"title":"Beyond Verdicts: A Graph-Based Analysis of Human and LLM Reasoning in Scientific Fact-Checking","zh_title":"超越裁决：科学事实核查中人类与LLM推理的图分析","abstract":"Misinformation that cites legitimate papers can be especially harmful when it distorts what those studies actually report. While existing automatic fact-checking systems based on large language models (LLMs) can assess whether a model assigns an Incorrect verdict and can gen- erate explanations for that decision, they typi- cally do not indicate whether the model follows the same reasoning path as human experts or arrives at the verdict through a different but still valid path. In this work, we introduce a graph- based framework (typed reasoning graph) for comparing human and LLM reasoning paths in scientific fact-checking. Building on prior work on fallacious reasoning in biomedical misinformation, MISSCIPLUS (Glockner et al., 2025), we model each explanation as a rea- soning graph that links the false claim to the relevant study context, study findings, fallacy- supporting premises, and fallacy labels. This representation enables one-to-one alignment of human and LLM reasoning at the level of fallacy-specific sub-graphs. For non-human- aligned LLM paths, we validate grounding in the cited study, relevance to the claim, and suf- ficiency for the verdict. Using 84 false claims from MISSCIPLUS, we evaluate GPT-5, Claude Opus 4.7, and Qwen3-32B across prompt and evidence settings. Results show distinct perfor- mance dimensions: Qwen3-32B has the lowest verdict failure rate, GPT-5 the highest human alignment, and Claude Opus 4.7 weak verdict prediction but often valid reasoning in success- ful cases","authors":["Abdul Ghafoor","Muhammad Arslan Manzoor","Yufang Hou"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-25","first_seen":"2026-08-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.23047","pdf_url":"https://arxiv.org/pdf/2608.23047","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A2","B1","B4"],"tags":["LLM推理对齐","事实核查","人类对照"],"reason":"比较人类与LLM推理路径，评估对齐度与有效性，有真实人类专家数据对照，批判性指…","model":"deepseek-v4-pro","scored_at":"2026-08-25T13:04:00","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-25","rank":12,"question":"在科学事实核查中，LLM 的推理路径是否与人类专家一致？如果不一致，其替代推理路径是否仍有效？","design":"该研究不是人类仿真实验，而是提出一种图结构（typed reasoning graph）来比较人类专家与 LLM 在科学事实核查中的推理路径。对 84 条虚假声明，让 GPT-5、Claude Opus 4.7、Qwen3-32B 生成解释，将解释建模为推理图，与人类专家推理图在谬误特定子图层面进行对齐；对未对齐路径，人工验证其是否基于引用研究、与声明相关且足以支持结论。","baseline":"人类基准来自 MissciPlus 数据集中由人类专家标注的推理路径（包括相关研究背景、研究发现、谬误支持前提和谬误标签）。","findings":"模型在结论准确性和推理路径质量上存在分离：Qwen3-32B 结论失败率最低，GPT-5 人类对齐率最高，Claude Opus 4.7 结论预测弱但成功案例中推理常有效。即使提供完整研究和详细提示，人类对齐推理仅占少数（GPT-5 为 32.1%，Claude Opus 4.7 为 8.3%，Qwen3-32B 为 15.5%），但许多未对齐路径经验证仍有效，表明 LLM 常通过替代推理路径得出正确结论。","reliability":"论文指出，评估推理路径的 grounding 和 sufficiency 通常需要领域专业知识，因此 LLM 系统应作为解释辅助工具而非独立仲裁者。此外，人类对齐率低可能源于专家推理路径的多样性或标注的主观性，但论文未深入讨论这些局限。","relevance":"该研究提供了比较人类与 LLM 推理过程的结构化方法，对关注 LLM 仿真可靠性的研究者有方法论价值，但场景限于科学事实核查，与经济学实验和政策评估的直接关联较弱。","inspiration":"借鉴其将推理过程分解为可对齐的图结构并区分结论与过程质量的做法，可用于评估 LLM 在经济决策中的推理是否与人类专家一致。｜可迁移到政策评估场景，如分析 LLM 对经济政策公告的解读是否遵循专家逻辑。｜以经济学研究者为被试，让 LLM 对同一政策文本生成推理，用图结构对齐人类与 LLM 的推理路径，以专家标注为基准，并验证未对齐路径的合理性。"}},{"id":"2608.22887","version":1,"title":"Proxy reliance in large language model decisions is uncalibrated to predictive evidence","zh_title":"大语言模型决策中的代理依赖与预测证据不校准","abstract":"Large language models (LLMs) are entering decisions in triage and lending, where task-relevant inference must be distinguished from impermissible proxy use. Current audits ask whether decisions change when demographics change. But attributes correlated with a protected group carry predictive value, so a changed decision can be discrimination or sound inference. We measure causal proxy effects in four LLMs on a clinical-ranking task with known ground truth, where the reliance the evidence warrants can be computed exactly and used as the reference. One audit signal yields three verdicts: over-reliance, warranted and under-reliance. Under neutral labels every model relies on proxies with no information. Informative proxies draw all three. Social field names push reliance down, below the reference in one model. Two findings explain this. Reliance severely undertracks the evidence, and social-label suppression is fragile, since in-context examples raise it above zero in every model. Accuracy-based evaluation detects none of this.","authors":["Zengqing Wu","Chuan Xiao"],"categories":["cs.AI","cs.CL","cs.CY"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-08-25","first_seen":"2026-08-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.22887","pdf_url":"https://arxiv.org/pdf/2608.22887","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A2","B4"],"tags":["LLM决策偏差","算法审计","代理变量"],"reason":"评估LLM决策中的代理依赖与偏差，与仿真可靠性相关，但非直接仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-08-25T13:03:58","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-25","rank":11,"question":"大语言模型在决策中对代理属性的依赖是否与预测证据所支持的水平相校准？","design":"在完全已知的合成数据生成过程中，四个大语言模型（Claude Sonnet 4.5、DeepSeek-V4-Flash-0731、Qwen3.7-max、GPT-5.6 Terra）扮演分诊专家，对模拟患者对进行优先级排序。通过因果模拟翻转受保护属性并仅通过代理属性传播，测量模型排序变化的代理特定效应，并与贝叶斯最优决策者和理想学习者（基于模型所见示例拟合的贝叶斯回归）的效应进行比较。","baseline":"无对照","findings":"在无信息代理下，所有模型都表现出非零依赖；随着代理预测价值增加，依赖严重低于证据支持的水平。社会领域名称会抑制依赖，但上下文示例会削弱这种抑制，且基于准确率的评估无法检测到这些偏差。","reliability":"论文指出，证据支持水平的计算依赖于已知的生成过程，在观测数据中无法识别；合成人群虽通过七个公共临床数据集锚定，但结论可能不完全适用于真实部署。","relevance":"该研究通过因果框架量化LLM决策中的代理依赖与证据校准，为评估LLM作为人类被试替代品时的决策偏差提供了方法论参考，但未直接复现人类行为，与仿真可靠性相关但非直接仿真。","inspiration":"借鉴其因果代理效应测量和理想学习者基准，可迁移到信贷审批中的代理歧视问题（如使用邮政编码作为种族代理）。设计上，以LLM作为信贷审批员，处理为改变申请人的邮政编码（与种族相关但含真实信用信息），结果变量为贷款批准决策，并用真实信贷数据（如Home Mortgage Disclosure Act数据）中人类审批员的决策作为对照，比较LLM与人类对代理信息的依赖程度。"}},{"id":"2608.23196","version":1,"title":"AI emotional support is better only when chosen, but shifts preferences even when it is not","zh_title":"AI情感支持仅在主动选择时更优，但即使非主动选择也会改变偏好","abstract":"People increasingly face a novel decision when seeking emotional support: human or AI. In existing studies, AI's empathic messages are rated as well as or better than humans'. But these studies either assigned the support source or honored people's choice. In real life, support is often incongruent with choice, as people want one source and receive the other. Across three experiments (N = 1,951), participants chose whether to share an emotional experience with a human or an AI, then were randomly assigned to a congruent or incongruent partner. AI support was rated as superior only among those who had chosen it. Yet regardless of congruence, interacting with AI increased willingness to choose it again. In a 28-day study with OpenAI (N = 981), daily conversations shifted preferences toward AI and away from humans, but only when conversations turned personal. Emotional support choices are thus path-dependent, progressively redirecting away from human connection.","authors":["Yaoxi Shi","Cathy Mengying Fang","Guy LabanPattie Maes","Amit Goldenberg"],"categories":["cs.AI","cs.HC"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-25","first_seen":"2026-08-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.23196","pdf_url":"https://arxiv.org/pdf/2608.23196","source_feed":"cs.AI","score":7,"bucket":"pending","rubric_hits":["A1","B1","B4"],"tags":["LLM仿真","情感支持","人机交互"],"reason":"用LLM替代人类提供情感支持，与真实人类对照，探讨选择与偏好变化，可迁移至仿真…","model":"deepseek-v4-pro","scored_at":"2026-08-25T13:03:37","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-25","rank":14,"question":"人们在寻求情感支持时，选择人类还是AI的信念如何影响其选择，以及选择一致或不一致如何影响对支持的体验和后续选择偏好？","design":"该研究不是用LLM仿真人类，而是让真人参与者选择与人类或AI（GPT-4）进行情感支持对话，随机分配到与选择一致或不一致的伙伴，测量支持质量评价和后续选择意愿；另有一个28天纵向研究，让参与者每天与AI对话，追踪偏好变化。","baseline":"人类基准是真人参与者与人类支持者（众包工人）的对话，作为与AI支持对照的真实人类数据。","findings":"AI支持只有在参与者主动选择AI时才被评价为优于人类；无论选择是否一致，与AI互动都会增加再次选择AI的意愿。在28天纵向研究中，日常与AI对话使偏好转向AI、远离人类，但仅当对话变得个人化时发生。","reliability":"论文承认现有研究多在参与者不知情且无选择条件下进行，可能高估AI优势；本研究显示选择一致性是关键调节变量，且偏好变化具有路径依赖性，但未深入讨论其他失效条件。","relevance":"该研究直接涉及LLM与人类在情感支持场景下的对照实验，揭示了选择一致性和暴露效应对偏好形成的影响，对理解LLM仿真人类行为时的情境依赖性和动态偏好变化有参考价值。","inspiration":"值得借鉴的是随机分配选择一致/不一致的处理设计，以及纵向追踪偏好变化的方法。｜可迁移到消费者对AI金融顾问的接受度研究，或投资者对AI生成投资建议的信任形成。｜设计一个实验：招募投资者，先询问其偏好人类顾问还是AI顾问，然后随机分配一致或不一致的顾问类型，测量投资决策质量和后续选择意愿，并与真实银行客户数据对照。"}},{"id":"2608.21389","version":1,"title":"Interrupting the Chain: Human Perception of AI-Generated Disinformation Through a Kill Chain Lens","zh_title":"打断链条：通过杀伤链视角理解人类对AI生成虚假信息的感知","abstract":"Generative AI enables customized misinformation at scale, yet defenses remain largely reactive. We present empirical findings from a human-subject study (n=504 participants, n=2,438 judgments) in which users classified news fragments by origin (human vs. machine) and veracity (real vs. fake). We organize results using an adapted cybersecurity kill chain as a taxonomy for intervention, mapping perception data onto stages of a cognitive attack lifecycle. Three key findings emerge: (1) a perception-accuracy gap where heightened suspicion does not improve detection; (2) modern LLMs frequently produce human-indistinguishable text; and (3) an asymmetric cognitive fatigue effect where fake-news detection degrades by 10.2 percentage points under sustained exposure while AI-origin detection remains stable. These findings identify candidate intervention points for proactive defense against AI-driven disinformation.","authors":["Alexander Loth","Martin Kappes","Marc-Oliver Pahl"],"categories":["cs.CY","cs.AI","cs.CR"],"primary_category":"cs.CY","announce_type":"cross","date":"2026-08-25","first_seen":"2026-08-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.21389","pdf_url":"https://arxiv.org/pdf/2608.21389","source_feed":"cs.AI","score":7,"bucket":"pending","rubric_hits":["A1","B1","B4"],"tags":["AI虚假信息","人类感知","人机对比"],"reason":"用人类被试判断AI生成内容，虽非LLM仿真人类，但涉及人类感知与AI文本对比，…","model":"deepseek-v4-pro","scored_at":"2026-08-25T13:03:41","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-25","rank":9,"question":"人类用户能否准确区分AI生成与人类撰写的新闻片段，以及真实与虚假内容，且这种感知在持续暴露下如何变化？","design":"本研究并非LLM仿真人类，而是人类被试实验：504名参与者通过JudgeGPT平台对新闻片段进行判断，每个片段需评估来源（人/机）和真实性（真/假），并记录反应时间；新闻片段由7个LLM生成并混合人类真实与虚假新闻，分析包括准确率、相关性、时间趋势。","baseline":"人类判断的准确率与随机猜测基线（50%）对比，以及不同模型生成文本的感知分数对比。","findings":"参与者对AI生成内容的检测准确率仅58.4%，对虚假内容为68.1%，且高怀疑度并未提高检测准确率；在持续暴露下，虚假新闻检测准确率下降10.2个百分点，而AI来源检测保持稳定。","reliability":"论文承认疲劳效应分析为描述性探索，未进行重复测量或混合效应检验；人口学相关性为观察性关联，未直接测试对抗性利用；反应时间分析基于中位数分割，未控制参与者内嵌套。","relevance":"虽然该研究不是用LLM仿真人类，但它提供了人类对AI生成内容的感知基准和认知偏差证据，对评估LLM仿真人类时的外部效度有参考价值，值得阅读原文了解人类判断的局限。","inspiration":"借鉴其双任务判断（来源与真实性）和反应时间测量，可揭示认知负荷对判断质量的影响｜可迁移到经济金融中的信息处理场景，如投资者对AI生成财报新闻的反应、消费者对AI生成产品评论的信任｜设计实验：招募投资者作为被试，呈现AI生成与人类撰写的公司新闻，要求判断来源和真实性并记录反应时间，同时收集真实投资决策数据作为对照，检验AI生成信息是否扭曲投资行为。"}},{"id":"2608.23524","version":1,"title":"The Measurement Revolution? Credible Measurement and Inference in the Age of AI","zh_title":"测量革命？AI时代的可信测量与推断","abstract":"Artificial intelligence (AI) is transforming measurement in economics. AI models convert unstructured data, such as text and images, into structured variables at low cost, making previously prohibitive measurement feasible at scale. This shifts the bottleneck from finding any scalable measure of a phenomenon to choosing among many plausible ones, which may support different empirical conclusions. This review provides guidance for navigating that shift. We describe three stages at which AI enters the measurement pipeline---discovery, construct definition, and observation---and what each demands of researchers. We argue that credible inference with AI-generated variables requires appropriately designed validation: anchoring measurement to explicit criteria, rather than informal claims that a proxy is reasonable. We then examine how validation samples support valid inference even when AI predictions are arbitrarily biased, and what can be done when a random validation sample is unavailable.","authors":["Melissa Dell","Ashesh Rambachan"],"categories":["econ.GN","cs.AI","q-fin.EC","stat.AP"],"primary_category":"econ.GN","announce_type":"cross","date":"2026-08-25","first_seen":"2026-08-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.23524","pdf_url":"https://arxiv.org/pdf/2608.23524","source_feed":"cs.AI","score":7,"bucket":"pending","rubric_hits":["A4","B3"],"tags":["AI测量","验证框架","因果推断"],"reason":"论文讨论AI生成变量的测量与推断，提供验证框架，可迁移到LLM仿真人类被试的可…","model":"deepseek-v4-pro","scored_at":"2026-08-25T13:03:37","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-25","rank":15,"question":"在AI时代，如何为AI生成的测量变量建立可信的测量与推断框架？","design":"本文为综述性论文，未进行仿真实验。它系统梳理了AI进入测量管道的三个阶段（发现、构念定义、观测），并提出通过验证样本锚定测量标准来支持有效推断的方法。","baseline":"无对照","findings":"AI使测量成本大幅降低，导致研究者面临从众多可行测量中选择的难题，不同测量可能支持不同实证结论。可信推断需要将AI测量锚定到明确标准，并利用验证样本校正偏差，即使AI预测存在任意偏差也能支持有效推断。","reliability":"论文指出AI模型的黑箱性质导致测量不透明，可能产生难以解释的变量；研究者可能进行事后选择导致推断偏差；专有模型更新导致可重复性问题；当随机验证样本不可得时，推断有效性受限。","relevance":"该论文为使用LLM生成变量（如仿真人类被试的回答）提供了测量验证框架，有助于评估仿真数据的可靠性与偏差，值得精读以指导仿真实验设计。","inspiration":"借鉴其验证样本设计，将AI生成的测量与人工标注或真实数据对比，以校正偏差｜可迁移到政策评估中利用LLM编码政策文本或调查回答的场景，如测量经济政策不确定性或消费者信心｜以LLM作为被试生成对政策公告的情绪反应，处理为不同政策措辞，结果变量为情绪得分，用真实调查数据（如密歇根消费者调查）作为基准进行验证。"}},{"id":"2608.23026","version":1,"title":"Beyond Surface Cues: Disentangling Sociocultural Signals in Multilingual LLMs","zh_title":"超越表面线索：在多语言大语言模型中解耦社会文化信号","abstract":"Multilingual LLM outputs can vary across sociocultural contexts. However, evidence of cultural grounding can be misleading: identity labels may be inferred from explicit or indirect textual cues, while names and wording can reveal the source language. Treating all these signals as evidence of cultural grounding may obscure potential biases. We present a human-validated, multi-agent audit that separates three questions: whether outputs reproduce social biases, whether identity groups are represented differently, and whether outputs reflect cross-cultural patterns. The study analyzes 89,253 outputs from 12 LLMs in English, French, and Chinese, spanning 18 occupations and three task conditions. We find that bias representation varies systematically across languages and tasks. Removing direct identity cues sharply reduces identity-label prediction in English and Chinese, but has a much smaller effect in French. Across all language-genre settings, the cultural context associated with the source language receives the highest average relevance score, with moderate agreement between automated and human ratings. However, the ability to identify the source language drops substantially after translation and again after masking names. Without these controls, multilingual audits may mistake surface cues for cultural understanding, leading to misleading conclusions about cross-cultural variation and bias. Our audit offers a practical framework for separating such shortcuts from more meaningful cross-cultural patterns.","authors":["Yuanjun Feng","Tanzhou Liu","Stefan Feuerriegel","Yash Raj Shrestha"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-25","first_seen":"2026-08-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.23026","pdf_url":"https://arxiv.org/pdf/2608.23026","source_feed":"cs.CL","score":6,"bucket":"other","rubric_hits":["D2","B4"],"tags":["LLM偏见审计","多语言文化差异","算法公平性"],"reason":"审计LLM中的社会文化信号与偏见，测量模型而非人类，但提供批判性框架可迁移。","model":"deepseek-v4-pro","scored_at":"2026-08-25T13:04:00","error":null,"has_summary":false,"summary":null},{"id":"2608.23411","version":1,"title":"STONIC: A Layered Measurement Contract for LLM Value Profiling","zh_title":"STONIC：LLM价值观画像的分层测量契约","abstract":"LLM value studies often merge questionnaire ratings, pairwise choices, and values inferred from generated text into one profile. That merge assumes that the three observations describe the same stable preference. STONIC tests this assumption on 5,144 situations from four banks and 35 fixed model configurations. It compares responses rated in isolation, choices made under counterbalanced conflict, spontaneous answers, and later choices between a model's own answer and authored alternatives. 10 of 17 configurations with usable behavioral data preserve the endorsement-choice relation across banks. Every one of the 17 eligible configurations prefers its own earlier answer (median effect 0.790), although option position changes the choice rate in every eligible configuration. Profile shape transfers most strongly from ratings to conflict choices and weakens for spontaneous text. Three-way annotation of 200 L3 responses provides a task-local check of the semantic audit: FULCRA agrees most closely with the human majority, while DeBERTa retains useful rank information after calibration. Hidden states encode the completed decision more clearly than the prompt alone. Thus the models show reproducible behavioral continuity, but the evidence does not support one scorer-independent value identity across interfaces.","authors":["Andrei Chetvergov","Stepan Ukolov","Timofei Sivoraksha","Alexander Evseev","Danil Sazanakov","Mikhail Solovev","Sergey Bolovtsov"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-25","first_seen":"2026-08-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.23411","pdf_url":"https://arxiv.org/pdf/2608.23411","source_feed":"cs.CL","score":6,"bucket":"other","rubric_hits":["D2"],"tags":["LLM价值观","测量一致性","行为连续性"],"reason":"测量LLM价值观稳定性，非仿真人类被试，但涉及测量方法可迁移","model":"deepseek-v4-pro","scored_at":"2026-08-25T13:03:37","error":null,"has_summary":false,"summary":null},{"id":"2608.22417","version":1,"title":"LLMs for Survey Text Analysis - A Performance Comparison Between Humans and GPT-5 on Inductive Content Analysis","zh_title":"用于调查文本分析的LLM：人类与GPT-5在归纳内容分析上的性能比较","abstract":"Large language models (LLMs) are increasingly used to support text analysis in qualitative research, yet evidence on their performance in inductive content analysis remains limited. This study compares human and LLM-based inductive coding of open-ended survey responses from 903 answers across six variables from a European PhD student survey. Five human coders performed inductive content analysis following a standardized coding scheme, while an LLM (GPT-5.4) conducted the same task using an established prompting procedure. Agreement between human and LLM outputs was assessed using the Adjusted Rand Index (ARI). Results showed an alignment between humans and the LLM, with ARI values of 0.61 for coding and 0.54 for theme generation. These values were close to the internal consistency of coding and theme results within humans (ARI = 0.68) and the LLM (ARI = 0.76). Agreement varied widely across variables, with low within-entity consistency consistently linked to low between-entity agreement, underscoring the role of data characteristics and individual performance in reliability. Overall, the findings suggest that LLMs can approximate human coding in this case-specific setting, particularly at the coding level, and may serve as a scalable support tool for inductive qualitative analysis.","authors":["Leonardo Bergmann","Renata Gheorghiu","Ana Gvritishvili","Alex Mican","Chris Stewart","Topias Tolonen-Weckstr\\\"om"],"categories":["cs.AI","cs.CL","cs.HC"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-08-25","first_seen":"2026-08-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.22417","pdf_url":"https://arxiv.org/pdf/2608.22417","source_feed":"cs.CL","score":6,"bucket":"other","rubric_hits":["D1"],"tags":["LLM标注","内容分析","方法比较"],"reason":"LLM替代人工编码，属标注替代而非仿真被试，但方法可迁移","model":"deepseek-v4-pro","scored_at":"2026-08-25T13:03:30","error":null,"has_summary":false,"summary":null},{"id":"2608.22425","version":1,"title":"All four leading LLMs talk more than they listen to personality-verified synthetic help-seekers","zh_title":"四大领先LLM在与人格验证的合成求助者对话时说得比听得多","abstract":"Large language models are increasingly consulted at moments of distress, yet single-turn benchmarks neither test sustained exchanges nor distinguish between users. We built a personality-aware evaluation in which four widely used models advised several synthetic help-seekers, each given a psychometrically specified profile, in an acute crisis: a caregiver learning of a relative's dementia diagnosis. Auditors blind to the profile prompt recovered the specified bands from dialogue alone with high agreement on every instrument (ICC(2,4) = 0.91; 0.79-0.96 by instrument; band-score r = 0.78), as expected for the Big Five but equally for coping style, coping self-efficacy, resilience and reactance, which the lexical approach never covered. Such evaluation therefore reaches beyond the Five Factor Model to motivational, regulatory and self-appraisal dispositions. The four models were not distinguishable on emotion stabilisation and failed alike, sharing three modes: verbosity, a talk-to-listen ratio above one, and problem-solving before the situation had been explored.","authors":["Pablo A. Fonseca","Raquel Rodr\\'iguez-Carvajal","Rafael A. Calvo"],"categories":["cs.HC","cs.CL","cs.CY"],"primary_category":"cs.HC","announce_type":"cross","date":"2026-08-25","first_seen":"2026-08-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.22425","pdf_url":"https://arxiv.org/pdf/2608.22425","source_feed":"cs.CL","score":6,"bucket":"other","rubric_hits":["D2","D3"],"tags":["LLM人格仿真","对话行为评估","合成被试"],"reason":"用LLM模拟具有人格特征的求助者，但无真实人类数据对照，且测量的是模型行为而非…","model":"deepseek-v4-pro","scored_at":"2026-08-25T13:03:30","error":null,"has_summary":false,"summary":null},{"id":"2608.22731","version":1,"title":"LLM-Based Selection of Incongruent Verbal and Nonverbal Behavior for Virtual Humans","zh_title":"基于LLM的虚拟人不一致言语与非言语行为选择","abstract":"Nonverbal behavior generation systems for virtual agents often take an utterance as input and generate nonverbal behaviors that emphasize or illustrate the content of the verbal channel. However, human nonverbal behavior is shaped by more than the content of the speech. It is also influenced by speaker roles, interpersonal relationships, social context, and the cognitive and emotional states of the interactants. As a result, the nonverbal channel may reinforce, weaken, qualify, or even contradict the verbal channel. It may also reveal internal states that are hidden or only indirectly implied in speech, including emotional \"leakage\" that may be incidental to the immediate interaction. Modeling this richer relationship between verbal and nonverbal behavior is important for designing virtual agents that exhibit realistic, human-like behavior. It is especially critical in training contexts that require nuanced social interpretation, such as counseling simulations involving virtual patients. Drawing on Ekman's framework of verbal nonverbal relationships, we propose a taxonomy of categories in which mismatches between verbal and nonverbal behavior can occur. We then examine alternative approaches for realizing these behaviors using large language models, focusing on whether LLMs can select contextually appropriate mismatched verbal and nonverbal behaviors from a given dialogue and social interaction context. Finally, we evaluate the resulting behaviors in a human-subject study, assessing whether context-driven nonverbal behavior, when embodied in a virtual human, produces the intended effects on observers.","authors":["Parisa Ghanad Torshizi","Stacy Marsella"],"categories":["cs.AI","cs.HC","cs.RO"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-25","first_seen":"2026-08-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.22731","pdf_url":"https://arxiv.org/pdf/2608.22731","source_feed":"cs.AI","score":6,"bucket":"other","rubric_hits":["D3"],"tags":["虚拟人","非言语行为","LLM生成"],"reason":"用LLM生成虚拟人的非言语行为，有真人评估但非以人类数据为对照基准，属社会模拟…","model":"deepseek-v4-pro","scored_at":"2026-08-25T13:03:54","error":null,"has_summary":false,"summary":null},{"id":"2608.21420","version":1,"title":"Evaluating Human and LLM-Generated Thematic Analysis in HRI for Vulnerable Populations: A Comparative and Ethical Analysis","zh_title":"评估人机交互中针对弱势群体的人类与LLM生成主题分析：比较与伦理分析","abstract":"Thematic analysis (TA) has long been regarded as an inherently human, reflexive, and interpretive process. However, the extent to which LLM-generated TA is appropriate for Human-Robot Interaction (HRI) research involving vulnerable populations remains largely unexamined and raises critical questions about validity and ethics, particularly in sensitive research contexts. This paper presents a comparative study of human- and LLM-generated TA in an HRI context with a focus on vulnerable populations. We evaluate both objective and semantic agreement between human- and LLMgenerated themes, and examine whether observed divergences reflect systematic interpretive patterns with ethical significance. Our analysis investigates whether LLM-generated TA risks marginalising or misrepresenting the experiences of vulnerable participants, with implications for researchers employing LLM-assisted TA in HRI.","authors":["Alva Markelius","Fethiye Irmak Dogan","Julie Bailey","Hatice Gunes"],"categories":["cs.RO","cs.HC"],"primary_category":"cs.RO","announce_type":"cross","date":"2026-08-25","first_seen":"2026-08-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.21420","pdf_url":"https://arxiv.org/pdf/2608.21420","source_feed":"cs.HC","score":6,"bucket":"other","rubric_hits":["D1"],"tags":["LLM辅助分析","主题分析","人机交互"],"reason":"LLM替代人工做主题分析，属标注替代而非仿真被试，但涉及人类数据对照和伦理，边…","model":"deepseek-v4-pro","scored_at":"2026-08-25T13:03:28","error":null,"has_summary":false,"summary":null},{"id":"2608.22884","version":1,"title":"Predicting the scale limits of social mechanisms in agent societies","zh_title":"预测智能体社会中社会机制的规模极限","abstract":"Societies of interacting language-model agents offer a controllable and repeatable way to study collective behaviour at scales that would be difficult to test with people. Their scientific value, however, depends on whether a social mechanism that works in a small group still operates when thousands of agents interact, and testing this directly requires costly large-scale runs. Here we introduce an audit that predicts a mechanism's fate as a population grows. It asks how often the mechanism can act, whether agents use the information it supplies, and whether the measurement itself creates apparent scale effects. Controlled experiments show that a single structural term can decide whether reciprocity, consensus or punishment survives scaling. For gossip, the population at which the mechanism fails is set by the reach and lifetime of its messages. In language-model societies, agents respond not only to social information but to how it is expressed: counts and percentages led to different scale behaviour. Predictions made before execution held on third-party code and a second model family, while a failed prediction exposed the boundary of the finding. The audit provides a prospective way to decide which social mechanisms can be interpreted across population scales.","authors":["Zengqing Wu","Chuan Xiao"],"categories":["cs.MA","cs.SI"],"primary_category":"cs.MA","announce_type":"new","date":"2026-08-25","first_seen":"2026-08-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.22884","pdf_url":"https://arxiv.org/pdf/2608.22884","source_feed":"cs.MA","score":6,"bucket":"other","rubric_hits":["D3"],"tags":["LLM社会模拟","多智能体","规模效应"],"reason":"用LLM agent群体模拟社会机制，但无真实人类数据对照，属边界情形。","model":"deepseek-v4-pro","scored_at":"2026-08-25T13:03:34","error":null,"has_summary":false,"summary":null},{"id":"2608.12645","version":2,"title":"Jagged Judges: Epistemic Stability Under Perturbation, Pressure, and Persistence","zh_title":"摇摆的法官：扰动、压力与坚持下的认知稳定性","abstract":"LLM judges have become central infrastructure for model evaluations, online grading, and reward modeling. Judges are typically validated by accuracy on golden data, but accuracy says little about whether they are stable under re-prompting, challenge, or sustained pushback. We introduce the \\emph{Wiggle Framework}, a unified stress test for epistemic stability in LLM judges. The framework decomposes judge robustness along three dimensions: Mechanical Consistency (stability under re-prompting and reframing), Single-turn Conviction (stability under a single challenge), and Multi-turn Persistence (stability under sustained or adaptive pressure). We use the framework to study 9 frontier models across 14 judging tasks spanning safety, toxicity, AI writing detection, and political-response evaluation. Every model exhibits substantial wiggle as a judge --- flipping verdicts 25--71\\% of the time under static pushback, and 62--91\\% with an adversarial LLM persuader. Critically, we find that pressure that succeeds in changing a judge's verdict is almost always net-corrupting with respect to ground truth. Beyond the framework itself, we identify baseline jury majority strength as the most effective single-shot signal for anticipating which items wiggle. Taken together, this is the first apples-to-apples cross-dataset comparison of mechanical, conformity, and persuadability tests in a judging context.","authors":["Justin Zhao","Himaghna Bhattacharjee","Hannah Korevaar","Bhaktipriya Radharapu","Khalid El-Arini"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"replace","date":"2026-08-25","first_seen":"2026-08-15","revised_at":"2026-08-25","abs_url":"https://arxiv.org/abs/2608.12645","pdf_url":"https://arxiv.org/pdf/2608.12645","source_feed":"cs.AI","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["LLM法官","稳定性","模型评估"],"reason":"研究LLM法官的稳定性，属于对模型本身属性的测量，而非用LLM仿真人类被试，但…","model":"deepseek-v4-pro","scored_at":"2026-08-25T13:04:11","error":null,"has_summary":false,"summary":null},{"id":"2608.14825","version":3,"title":"Emergent Misaligned Communication in Long-Horizon Multi-Agent LLM Commerce","zh_title":"长时程多智能体LLM商务中涌现的失准沟通","abstract":"Frontier LLM agents increasingly transact on behalf of separate principals, often using natural language rather than structured APIs. Much of the safety literature studies misaligned LLM behavior through adversarial-elicitation evaluations on single agents or stylized tasks. Its prevalence and structure in settings that combine long horizons, separate principals, real operational state, and inter-agent natural-language exchange remain insufficiently measured. We study 2,583 inter-agent emails from 20 one-year simulation runs of Vending-Bench Arena, a competitive vending environment spanning 13 frontier LLMs. We operationalize speech-act misalignment as emails containing false factual claims, manipulation, collusion, or threats, combining message content with ground-truth simulator state and logged reasoning traces to classify and validate such behavior. Under our primary classifier, 12.6% of emails are labeled misaligned; misalignment appears in all 20 runs and 74.7% of individual agent-runs. Both the magnitude and composition of this misalignment are preserved under repeated classification at different sampling temperatures and under full-pipeline replication with judges from two other frontier-model families. Misalignment is also reciprocal and stress-conditioned: receiving a misaligned email from a counterparty raises the odds of a misaligned reply by 1.65x, and low-inventory conditions raise them by 1.58x. Across tests of capability-asymmetric exploitation, we find no evidence that higher-capability models differentially exploit weaker counterparties, and model performance rank does not predict misalignment rates. Together, these results indicate that measurable, state-dependent misalignment can arise in competitive multi-agent environments without engineered elicitation, in patterns associated with operational scarcity and counterparty behavior rather than model capability alone.","authors":["Zeyuan Li","Lukas Petersson","Alessandro Acquisti","Michiel A. Bakker"],"categories":["cs.MA","cs.AI"],"primary_category":"cs.MA","announce_type":"replace-cross","date":"2026-08-25","first_seen":"2026-08-18","revised_at":"2026-08-25","abs_url":"https://arxiv.org/abs/2608.14825","pdf_url":"https://arxiv.org/pdf/2608.14825","source_feed":"cs.AI","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["多智能体模拟","LLM安全","经济仿真"],"reason":"多智能体经济环境模拟，但无真实人类数据对照，属边界情形","model":"deepseek-v4-pro","scored_at":"2026-08-25T13:04:08","error":null,"has_summary":false,"summary":null},{"id":"2608.22192","version":1,"title":"How Agents Represent Humans: Human-Directed Stereotypes in an Open Agent Social Network","zh_title":"智能体如何表征人类：开放智能体社交网络中的人类导向刻板印象","abstract":"LLM-based agents are increasingly deployed in persistent social environments, where generated claims can be posted, replied to, remembered, and reused. We study human-directed stereotypes on Moltbook, an open agent-native social platform, asking how agents construct humans as a social category. For this human-target analysis, we introduce an annotation framework with four evaluative dimensions---morality, friendliness, competence, and autonomy---and a second-stage subtype scheme for descriptive \\textit{other} attributions. We find that competence dominates human-directed evaluations, while many \\textit{other} attributions describe humans as epistemic, cultural, or embodied subjects. We further examine how these human representations appear in human--agent narrative contexts and platform-level circulation. As an auxiliary comparison, we analyze agent-internal community feedback through behavioral host affinity. Rather than reproducing the stable insider--outsider rejection often observed in human online communities, Moltbook feedback patterns are better explained by exposure, author visibility, and content selection. These findings suggest that bias in agent societies should be studied not only as isolated model output, but also as a discourse process.","authors":["Huangchen Xu","Yuan Wu","Yi Chang"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-25","first_seen":"2026-08-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.22192","pdf_url":"https://arxiv.org/pdf/2608.22192","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["LLM智能体","社会模拟","刻板印象"],"reason":"研究agent社会中对人类的刻板印象，属社会模拟但无真实人类数据对照，为边界情…","model":"deepseek-v4-pro","scored_at":"2026-08-25T13:03:28","error":null,"has_summary":false,"summary":null},{"id":"2608.22993","version":1,"title":"LLM Pedagogical Behavior in AI Tutoring Interactions","zh_title":"AI辅导互动中LLM的教学行为","abstract":"Students increasingly use LLMs as tutors for coursework and problem solving. Little is known about the level of assistance LLMs provide when students use them as tutors in authentic learning interactions. This matters because tutoring responses can differ substantially in how directly they help students complete a task. We operationalize this dimension as scaffolding level and develop a five-level scale, validated against human annotations, that characterizes responses according to the degree of direct assistance they provide. We apply the scale to 14,637 LLM responses from 203 students in a university AI course. Responses are overwhelmingly concentrated at high levels of assistance, with more than 95% classified as either Explaining or Solving. Scaffolding level is systematically associated with students' subsequent conversational behavior, but provides little additional predictive information about performance on three subsequent exams beyond prior achievement and dialogue behavior. These findings provide an empirical baseline for LLM assistance in tutoring interactions and a measurement framework for evaluating how alternative tutoring designs change that assistance.","authors":["Suhyeon Lee","Juneha Baek","Jaehyeong Park","Donghyuk Shin"],"categories":["cs.CL","cs.HC"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-25","first_seen":"2026-08-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.22993","pdf_url":"https://arxiv.org/pdf/2608.22993","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM教学行为","人机交互","教育技术"],"reason":"研究LLM作为导师的辅助行为，非仿真人类被试，但涉及LLM替代人类角色，属边界…","model":"deepseek-v4-pro","scored_at":"2026-08-25T13:03:59","error":null,"has_summary":false,"summary":null},{"id":"2608.22852","version":1,"title":"Your AI, On a Dial: Controlling Investment Bias in LLMs with a Single Neuron","zh_title":"你的AI，在一个刻度盘上：用单个神经元控制LLM的投资偏差","abstract":"Large language models (LLMs) are increasingly used in investment decision-making, yet prior work shows that they exhibit systematic, model-specific investment preferences. We study whether a model's overall investment stance can be calibrated to a specified direction and strength. We introduce an investment-bias dial, an inference-time intervention on a single neuron that continuously adjusts a model-level decision prior---its overall tendency toward buying or selling---without targeting specific firms or investment attributes. Using matched positive and negative evidence, we evaluate five open-weight LLMs and find that the dial produces monotonic changes in investment stance without modifying prompts or model parameters. At the response level, the dial shifts both investment decisions and the evidential emphasis of generated rationales under identical inputs. In an agentic retrieval setting, the dial also changes what information the model searches for, which evidence it selects, and which evidence is reflected in its final analysis. In a long-context evaluation, the dial maintains stable stance control as context length increases, whereas a matched system-prompt instruction progressively attenuates. We further show that changes in the dial propagate to security rankings and downstream portfolio composition in an exploratory backtest. Overall, our results show that an LLM's aggregate investment stance can be calibrated toward a specified target at inference time.","authors":["Sahong Park","Suhwan Park","Hoyoung Lee","Gakyung Kwon","Wonbin Ahn","Jaewon Choi","Alejandro Lopez-Lira","Yoon Kim","Chanyeol Choi","Hyeongwoo Kong","Yongjae Lee"],"categories":["cs.AI","cs.CL","q-fin.GN"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-08-25","first_seen":"2026-08-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.22852","pdf_url":"https://arxiv.org/pdf/2608.22852","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["LLM投资偏好","神经元干预","模型校准"],"reason":"研究LLM投资偏好校准，属模型测量而非人类仿真，无人类数据对照。","model":"deepseek-v4-pro","scored_at":"2026-08-25T13:03:56","error":null,"has_summary":false,"summary":null},{"id":"2608.22833","version":1,"title":"Minimal Local Simulation Foundations for LLM- and VLM-Driven Agents in 2D and 3D Environments","zh_title":"面向2D和3D环境中LLM与VLM驱动智能体的最小局部仿真基础","abstract":"Large language models (LLMs) and vision-language models (VLMs) are expanding the range of behaviors that can be represented in agent-based simulations, but many contemporary platforms are difficult to study, modify, or run on ordinary computers. We present two intentionally minimal simulation foundations for education and rapid prototyping. SD-AgentFoundry-2D provides a two-dimensional multi-agent environment in which locally hosted LLM agents move, communicate, respond to place occupancy, and encounter spatially localized fire events. SD-AgentFoundry-3D provides a three-dimensional digital-twin environment in which a locally hosted VLM receives first-person images and produces natural-language movement instructions. Both codebases are designed to run locally on macOS, Windows, and Linux and are deliberately left open to modification rather than developed as finished applications. Together, they offer accessible starting points for learning about generative social simulation and for building domain-specific extensions.","authors":["Ryuki Hyodo"],"categories":["cs.MA","cs.AI"],"primary_category":"cs.MA","announce_type":"cross","date":"2026-08-25","first_seen":"2026-08-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.22833","pdf_url":"https://arxiv.org/pdf/2608.22833","source_feed":"cs.AI","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["多智能体仿真","LLM驱动","社会模拟"],"reason":"提供LLM驱动的多智能体社会仿真平台，但无人类数据对照，属边界情形。","model":"deepseek-v4-pro","scored_at":"2026-08-25T13:03:56","error":null,"has_summary":false,"summary":null},{"id":"2608.21850","version":1,"title":"Consistently Good vs. Occasionally Great: A Rubric for Open-Ended Feedback Quality from Humans and Machines","zh_title":"一贯良好与偶尔卓越：人类与机器开放式反馈质量评估量表","abstract":"Providing high-quality feedback on student work is essential for learning, yet delivering such feedback at scale remains challenging. In this paper, we focus on feedback for open-ended short answer questions in introductory programming, with the goal of nudging students toward success on reattempts without revealing the correct answer. We develop a five-criteria rubric grounded in educational literature for evaluating feedback quality: (1) acknowledging correct portions of the student answer, (2) identifying at least one flaw (if present), (3) providing actionable guidance for improvement, (4) maintaining appropriate concealment of the answer, and (5) using an appropriate conversational tone. Using this rubric, we compare feedback generated by a frontier LLM (OpenAI o1) to feedback from nine teaching assistants across 90 student responses, with three researchers and an LLM independently scoring all feedback. Our results show that while one TA often produced the best feedback, the LLM demonstrated consistently higher average performance than TAs, as evaluated by humans. However, we also uncover significant self-preference bias when using LLMs to evaluate feedback quality: the LLM systematically rated its own outputs higher than human experts did. This bias, which research suggests persists even in cross-model evaluation, raises important methodological concerns for researchers employing LLM-based evaluation. We provide detailed characterization of both TA and LLM performance, analyze sources of variance in TA feedback quality, and discuss implications for deploying LLM-generated feedback in educational settings.","authors":["Binglin Chen","Rajarshi Haldar","Max Fowler","Matthew West","Craig Zilles"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2026-08-25","first_seen":"2026-08-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.21850","pdf_url":"https://arxiv.org/pdf/2608.21850","source_feed":"cs.CY","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM评估","反馈质量","教育技术"],"reason":"LLM替代人工标注反馈质量，非仿真人类被试，但涉及LLM评估偏差，可迁移。","model":"deepseek-v4-pro","scored_at":"2026-08-25T13:03:43","error":null,"has_summary":false,"summary":null},{"id":"2608.22242","version":1,"title":"Unfolding the Interdisciplinary Complexities of Climate Science: Fuxi-Climate Foundational Model","zh_title":"揭示气候科学的跨学科复杂性：Fuxi-Climate基础模型","abstract":"Climate research and decision-making require integrating evidence across physical processes, socio-economic dynamics and policy responses. Large language models (LLMs) have been explored for accessing and synthesizing climate knowledge, but their ability to support structured interdisciplinary reasoning is still limited. Here we present the Fuxi-Climate Foundation Model (CFM), a climate-specialized LLM designed to support consistent reasoning across domains. CFM maintains more stable analytical behavior as interdisciplinary complexity increases, whereas performance in other models becomes more variable. On expert-designed climate transition tasks, CFM produces more structured analyses that explicitly address trade-offs and uncertainty, achieving 45% trade-off coverage and 47.27% uncertainty-aware reasoning. These results indicate that CFM can support more realistic analysis of climate risks and transition pathways, and provide a basis for agent-based systems to explore complex policy and decision scenarios. The model is openly available at https://huggingface.co/SII-yuning/cfm.","authors":["Zhengyu Shi","Shaojie Shi","Rui Xu","Bohao Lv","Zhichao Chen","Jiaran Hao","Zijian Chen","Weiqi Tang","Yuan Qi","Yinghui Xu","Libo Wu"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2026-08-25","first_seen":"2026-08-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.22242","pdf_url":"https://arxiv.org/pdf/2608.22242","source_feed":"cs.CY","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["气候科学","LLM推理","政策分析"],"reason":"用LLM支持气候政策分析，但无人类数据对照，属社会模拟边界情形。","model":"deepseek-v4-pro","scored_at":"2026-08-25T13:03:46","error":null,"has_summary":false,"summary":null},{"id":"2608.22618","version":1,"title":"KMGen: A Skill-based Approach for Synthetic Individual Patient Data Generation","zh_title":"KMGen：一种基于技能的合成个体患者数据生成方法","abstract":"Individual patient data (IPD) from clinical trials is the substrate for survival modeling, meta-analysis, and safety research, yet IPD is rarely released. Prior work has addressed only half of this gap: reconstructing Kaplan-Meier (KM) curves from published plots -- typically requiring manual digitization or human-in-the-loop correction -- while offering no mechanism for generating the adverse-event (AE) streams that constitute the other half of a patient record. We introduce KMGen, the first end-to-end framework that (i) fully automates KM curve extraction at accuracy competitive with human-guided tools, and (ii) generates synthetic per-patient AE trajectories from public trial registry records. The extraction stage is a fully automated agentic pipeline -- an agent generates code to extract each step in the KM curve -- achieving a mean Integrated Absolute Error (IAE) of 0.0151 on a 32-plot benchmark spanning clean, edge-case, and adversarial conditions. The IPD generation stage decouples patient archetype extraction from statistical sampling: an LLM distills the trial record into arm-specific statistics, adverse events, patient demographics, and risk multipliers. A mechanistic sampler generates patient events via clinical archetypes, bootstrap rank-correlation coupling to the empirical KM curve (preserving the marginal survival distribution exactly), and cycle-based AE scheduling with an induction/maintenance split. Across three held-out oncology trials spanning an order of magnitude in cohort size and 30 independent regenerations per trial, KMGen achieves mean integrated KM absolute difference $\\Delta_{\\text{KM}}\\,{\\leq}\\,0.051$, sex/ECOG JSD ${\\leq}\\,0.013$ on 5 of 6 demographic slots, and recovers ${\\geq}\\,71\\%$ of the top-15 AEs by exact MedDRA term under a single fixed parameter set. The pipeline is released as open source at https://github.com/chufangao/kmgen.","authors":["Jalen Jiang","Chufan Gao","Ethan Rasmussen","Stephen Z. Xie","Jimeng Sun"],"categories":["cs.LG"],"primary_category":"cs.LG","announce_type":"new","date":"2026-08-25","first_seen":"2026-08-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.22618","pdf_url":"https://arxiv.org/pdf/2608.22618","source_feed":"cs.LG","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["合成数据","临床试验","LLM应用"],"reason":"用LLM生成合成患者数据，替代真实数据，但非仿真人类被试，属数据增强边界情形。","model":"deepseek-v4-pro","scored_at":"2026-08-25T13:03:50","error":null,"has_summary":false,"summary":null},{"id":"2608.22295","version":1,"title":"LLM Evaluation on Unseen Questions: Contextual Multidimensional IRT Model","zh_title":"未见问题的LLM评估：情境多维IRT模型","abstract":"Evaluation of large language models (LLMs) increasingly requires predicting how a model will perform on new questions or tasks before collecting large amounts of new annotations. This problem is challenging because question difficulty, scenario, and underlying capability demands can vary substantially. Simple retrospective averages may confound model ability with item characteristics. In this paper, we study a model-based evaluation framework that combines multidimensional item response theory model with question contexts to predict LLM performance on unseen questions. The framework represents LLMs through latent capability profiles while using question content to inform item characteristics, allowing information to transfer beyond previously observed items. Empirically, we find that for within-scenario evaluation, incorporating question embeddings improves prediction relative to model-free baselines, and that multidimensional latent structure provides a richer description of capability variation than unidimensional alternatives. At the same time, our results reveal an important limitation that the generalizability does not necessarily translate into reliable prediction under cross-scenario shift. These findings suggest that context-aware psychometric modeling is a promising direction for efficient and interpretable LLM evaluation, while also highlighting cross-scenario generalization as a central open challenge.","authors":["Ergan Shang","Weijing Tang","Yinqiu He"],"categories":["cs.CL","cs.AI","cs.LG"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-25","first_seen":"2026-08-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.22295","pdf_url":"https://arxiv.org/pdf/2608.22295","source_feed":"cs.CL","score":4,"bucket":"other","rubric_hits":["C4"],"tags":["LLM评估","项目反应理论","能力预测"],"reason":"论文是LLM能力评测方法，不涉及人类仿真或人类数据对照","model":"deepseek-v4-pro","scored_at":"2026-08-25T13:03:47","error":null,"has_summary":false,"summary":null},{"id":"2608.21721","version":1,"title":"Ask or Answer: A Decision Framework for Multi-Turn Health Misinformation Intervention","zh_title":"问或答：多轮健康误信息干预的决策框架","abstract":"Correcting health misinformation in dialogue requires more than producing a factual rebuttal: users differ in what they know, what they believe, and what they need to hear, so an effective intervention often depends on first asking the right clarifying question. Yet existing methods either respond immediately or probe indiscriminately, treating clarification as either unnecessary or always beneficial. We propose Reward-Optimized Probe-and-Respond (RO-PnR), a framework that learns when asking is worth its cost. At each turn, RO-PnR chooses between probing for more information and committing to a final correction, guided by a turn-level reward that weighs the expected gain from probing against its interaction cost. To capture how user heterogeneity affects probing value, we model each simulated user with a latent state along health literacy and belief commitment. Experiments show that RO-PnR achieves the highest cost-adjusted utility across three health-misinformation datasets and three base models, using 30% fewer turns than always-probe baselines.","authors":["Xiaoying Song","Anirban Saha Anik","Jinyu Liu","Qitao Tan","Geng Yuan","Lingzi Hong"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-25","first_seen":"2026-08-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.21721","pdf_url":"https://arxiv.org/pdf/2608.21721","source_feed":"cs.AI","score":3,"bucket":"other","rubric_hits":["C3"],"tags":["对话系统","健康干预","强化学习"],"reason":"研究健康误信息干预对话策略，用户为模拟角色，无真实人类数据对照，非以人类行为仿…","model":"deepseek-v4-pro","scored_at":"2026-08-25T13:03:42","error":null,"has_summary":false,"summary":null},{"id":"2608.22615","version":1,"title":"DeepSAGE: Stage-Aware Reinforcement Learning for Structured CBT Counseling Dialogue","zh_title":"DeepSAGE：面向结构化CBT咨询对话的阶段感知强化学习","abstract":"Large Language Model (LLM)-based counseling agents can generate fluent and supportive responses, but they often lack the structured, goal-directed progression required to conduct a coherent therapeutic session. We present DeepSAGE (Strategic AI Guidance Engine), a hybrid LLM--Deep Reinforcement Learning (DRL) framework for stage-aware counseling dialogue grounded in the first session of Cognitive Behavioral Therapy (CBT). DeepSAGE represents the session as eleven stages with explicit therapeutic objectives, with an external controller determines stage completion and the DRL model selects therapeutic intentions that guide LLM response generation. We evaluate DeepSAGE against six retrieval-, prompting-, stage-, and policy-based alternatives. DeepSAGE elicits higher simulated client engagement and openness and achieves the strongest balance of stage-goal completion and dialogue efficiency among stage-structured systems. Domain expert review further indicates that the generated conversations exhibit broadly plausible emotional trajectories and recognizable CBT processes. Because the evaluation relies primarily on simulated clients and model-based metrics, these findings demonstrate comparative dialogue-control improvements rather than clinical effectiveness. These results suggest that combining stage-structured dialogue with learned strategy selection is a promising approach for AI counseling, though clinical effectiveness, safety, and real-world utility require further human evaluation.","authors":["Qi Zhang","Heajun An","Prakriti Dumaru","Sang Won Lee","Lifu Huang","Pamela J. Wisniewski","Jin-Hee Cho"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-25","first_seen":"2026-08-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.22615","pdf_url":"https://arxiv.org/pdf/2608.22615","source_feed":"cs.AI","score":3,"bucket":"other","rubric_hits":["C3"],"tags":["AI心理咨询","强化学习","对话系统"],"reason":"论文研究AI心理咨询对话系统，用模拟客户评估，但目标是改进对话策略，非用LLM…","model":"deepseek-v4-pro","scored_at":"2026-08-25T13:03:49","error":null,"has_summary":false,"summary":null},{"id":"2608.22832","version":1,"title":"Let the Bullets Fly: Multimodal Fake News Detection with Temporal-Aligned Generative Danmaku","zh_title":"让子弹飞：基于时间对齐生成弹幕的多模态假新闻检测","abstract":"The social interactions among crowds via \\textit{Danmaku} (a.k.a., bullet comments) on modern multimedia platforms can facilitate both viewpoint conflicts and consensus, providing fine-grained discriminative social signals that can benefit fake news detection. However, the inherent accumulation latency of \\textit{Danmaku} in real-world scenarios violates the real-time necessity of fake news detection, making the studies of \\textit{Danmaku}-related fake news detection underexplored. To break this violation, we simulate this temporal-aware user interactive process by proposing a novel temporal \\textbf{Gen}erative \\textbf{da}nmaku framework, called \\textbf{Genda}, which consists of: (1) a \\textit{Danmaku} Trigger for predicting the timing and intensity of user reactions; and (2) a \\textit{Danmaku} Generator for synthesizing corresponding semantic and emotional expressions, thereby mutually constructing a temporally aligned and human-like pseudo \\textit{Danmaku} streams. To make the generated \\textit{Danmaku} useful for identifying fake news videos, we further design a \\textit{Danmaku}-guided Temporal Multimodal fake news detection model - \\textbf{DM-FEND}, which enables fine-grained multimodal interactions among video, audio, text, and \\textit{Danmaku}, enhancing dynamic modalities alignment and semantic noise inhibition. The experimental results demonstrate that \\emph{DM-FEND} consistently outperforms state-of-the-art baselines across both Chinese (FakeSV) and English (FakeTT) benchmarks. Further ablations validate the crucial role of temporal \\textit{Danmaku} modeling in enhancing robustness and discriminative capability. Finally, this study offers a bright and robust solution for multimodal fake news detection in modern social interactive fashions by bridging the temporal inconsistency between news and user behaviors.","authors":["Xiansheng Luo","Chaowei Zhang","Zewei Zhang","Yi Zhu","Jipeng Qiang"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-25","first_seen":"2026-08-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.22832","pdf_url":"https://arxiv.org/pdf/2608.22832","source_feed":"cs.AI","score":3,"bucket":"other","rubric_hits":["C1"],"tags":["假新闻检测","生成式弹幕","多模态学习"],"reason":"生成弹幕用于假新闻检测，非仿真人类被试，无人类数据对照","model":"deepseek-v4-pro","scored_at":"2026-08-25T13:03:56","error":null,"has_summary":false,"summary":null},{"id":"2608.22702","version":1,"title":"AffAdapt: AFFect-driven ADAPTive AI Personas for Seamless Conversations","zh_title":"AffAdapt：情感驱动的自适应AI角色，实现无缝对话","abstract":"AI-generated personas are being increasingly used for support, training and simulations. While generative AI models possess abilities to generate affect-aware responses, their embodiment into visual personas is an active area of investigation. Naturalistic exchanges require understanding of the conversational partners' turn completions, whether the agent should respond or keep listening and rely on non-verbal cues aligned with one's emotional states. Seamless human-AI conversation in a multimodal setting requires all modalities being generated to act in coordination. We present AffAdapt, a seamless interaction design framework for AI-personas, which coordinates streaming speech recognition, proactive turn-management, persona-grounded response generation, a persistent emotional state, and synchronized embodied output into a single interaction loop. We demonstrate the architecture in the context of practicing sensitive, high-stakes conversations, and report an initial case study showing fluid turn management and adaptive, persona-consistent behavior, alongside open challenges in interruption handling, open-ended dialogue, and multimodal affective alignment. AffAdapt's interaction loop is a generalizable pattern for coordinating timing, identity, and affect in real-time AI personas - applicable to training, coaching, education, and simulation contexts wherever believable, responsive interaction matters.","authors":["Nishanth Chidambaram","Kaustubh Paliwal","Kayla Hom","Shaoze Zhou","Chen Chen","Manas Satish Bedmutha","Nadir Weibel"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-08-25","first_seen":"2026-08-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.22702","pdf_url":"https://arxiv.org/pdf/2608.22702","source_feed":"cs.HC","score":3,"bucket":"other","rubric_hits":["C3"],"tags":["AI角色","多模态交互","对话系统"],"reason":"论文聚焦于AI角色对话框架，用于高压力对话练习，属于角色扮演聊天机器人，无实验…","model":"deepseek-v4-pro","scored_at":"2026-08-25T13:03:54","error":null,"has_summary":false,"summary":null},{"id":"2608.01033","version":2,"title":"CallScreenBench: Benchmarking Small Language Models as Phone Secretaries","zh_title":"CallScreenBench：评估小语言模型作为电话秘书的基准","abstract":"Language models small enough to run on a handset, quantized to a few bits, are increasingly capable of acting on their user's behalf -- which makes on-device task automation newly plausible. One such task is answering the phone. A phone secretary takes an unknown inbound call on its owner's behalf, and unlike the agents most benchmarks evaluate, it has no cooperative caller-assigned task to complete: the caller holds the goal and may be an adversary, while the secretary must begin deciding how to respond without an oracle. What matters is not task success but whether the owner would endorse how their proxy handled the call. We evaluate only the text-domain conversational decision layer; speech recognition, audio interaction, end-to-end latency, and handset execution are outside scope. We present CallScreenBench, which reports five automated call-and-note measure groups motivated by owner endorsement. Each is paired, where available, with a counter-metric and an uncertainty estimate; no benchmark-wide Q1-Q5 composite or leaderboard score is defined. Three guardedness diagnostics identify candidate cases for a toolless proxy that holds no credentials and calls no tools. Across three model families represented by paired 4-bit checkpoints (0.6-4B), the primary scoring snapshot gives the larger checkpoint higher point estimates on several service, recall, and plausibility measures, while triage discrimination follows a different ordering. Bare scam-side TPR rewards universal suspicion, and pairwise separation changes when legitimate-side false positives are included and across judge snapshots. Scripted degenerate agents expose further floors, including a hangup-and-echo policy with entity recall 1.000. We report quality measures and guardedness channels separately so that a single pass/fail score does not hide their trade-offs.","authors":["Jiaqi Gan","Haoyuan Tang","Jamey Z. Liang","Siying Chen","Ankit Raj","Kidus Zewde","Yuchen Zhou","Yuxin Zhang","Simiao Ren"],"categories":["cs.CR","cs.AI"],"primary_category":"cs.CR","announce_type":"replace-cross","date":"2026-08-25","first_seen":"2026-08-04","revised_at":"2026-08-25","abs_url":"https://arxiv.org/abs/2608.01033","pdf_url":"https://arxiv.org/pdf/2608.01033","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["电话秘书","小语言模型","对话决策"],"reason":"评估小语言模型作为电话秘书的对话决策，属于角色扮演对话，无实验或测量目的，不涉…","model":"deepseek-v4-pro","scored_at":"2026-08-25T13:04:09","error":null,"has_summary":false,"summary":null},{"id":"2608.10008","version":3,"title":"Do LLM Recommenders Know When They're Hallucinating? Auditing Confidence Calibration in Catalog Faithfulness","zh_title":"LLM推荐系统知道自己何时产生幻觉吗？审计目录忠实度中的置信度校准","abstract":"LLM recommenders for top-K item suggestion regularly emit titles outside the target catalog. Prior audits report a binary out-of-domain rate; none ask whether the model knew. We jointly audit hallucination rate (OOD@10) and verbalized-confidence calibration (ECE, Brier, reliability) for four zero-shot LLM recommenders from four independent vendors (Mistral Large, Llama-3.3-70B, GPT-OSS-120B, Claude Sonnet 4.6), not grounded or fine-tuned systems, across three catalogs (MovieLens-25M, Amazon Reviews 2023 Toys, Yelp Open Dataset), stratified by item popularity. Measuring catalog membership is itself the hard part: on identical outputs the reported rate moves by an order of magnitude with the string matcher used, and F1 cannot separate the candidates. We validate the instrument against 201 human judgments and select on net bias, where the adopted one is off by -0.040 against +0.144 for the common fuzzy rule. Hallucination is then strongly catalog-dependent (0.6-2.7% on MovieLens, 11.6-38.7% on Yelp, 49.3-61.0% on Amazon Toys). Each model holds a near-constant confidence level barely responsive to the catalog, while the catalog-hit rate swings 60 points, so the sign of the error is set by where a model's constant lands against a catalog's accuracy: 7 of the twelve cells are under-confident and 5 over-confident, all four under-confident on MovieLens, all four over-confident on Amazon Toys. We read this as an elicitation mismatch: \"Just Ask\" elicits a generic quality rating, not a catalog-membership probability. A conformal abstention threshold over verbalized confidence changes hallucination by at most 1.65 pp across four alpha levels, because the channel cannot separate correct items from hallucinations. We recommend that audits report calibration alongside OOD, validate the matcher producing the OOD number, and use catalog-anchored elicitation.","authors":["Srijith Ravikumar"],"categories":["cs.IR","cs.CL","cs.LG"],"primary_category":"cs.IR","announce_type":"replace-cross","date":"2026-08-25","first_seen":"2026-08-12","revised_at":"2026-08-25","abs_url":"https://arxiv.org/abs/2608.10008","pdf_url":"https://arxiv.org/pdf/2608.10008","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["LLM推荐系统","幻觉审计","置信度校准"],"reason":"审计LLM推荐系统的幻觉与置信度校准，属于NLP能力评测，不以人类行为仿真为目…","model":"deepseek-v4-pro","scored_at":"2026-08-25T13:04:10","error":null,"has_summary":false,"summary":null},{"id":"2608.15382","version":2,"title":"Framework for Grounding Healthcare LLMs in a Causal Knowledge Graph: A Cardiovascular Example Pilot","zh_title":"将医疗大语言模型锚定于因果知识图谱：框架、指标与心血管试点","abstract":"Large language models (LLMs) are increasingly proposed for healthcare decision support, but their evaluations still reward single-answer accuracy rather than reasoning about interventions, mechanisms, harms, evidence, and uncertainty. We propose a reproducible, graph-centered evaluation framework for intervention-oriented LLM behavior in healthcare and stress-test it in a cardiovascular pilot. The framework has four components: (i) a domain causal knowledge graph in which assertions are first-class, provenance-preserving nodes with stable identifiers; (ii) a scenario-conditioned subgraph extraction step that, given any clinical scenario, retrieves the relevant reified-assertion subgraph; (iii) four controlled grounding conditions that vary how the retrieved subgraph is composed into the model's context (ungrounded C1, knowledge-graph C2, causal-graph C3, integrated C4); and (iv) an automated scoring pipeline, anchored on assertion identifiers, that computes intervention accuracy, and other evaluation measures on a single pass. To test the framework, we built a category-balanced scenario generator across eight reasoning failure modes and instantiated it on a cardiovascular graph. The metric panel discriminates conditions along interpretable, non-redundant axes: C4 obtains the strongest causal edge F1 (0.838), adverse-effect F1 (0.833), evidence accuracy (0.738), and unsupported claim rate (0.114), while C1 obtains the highest raw intervention accuracy (0.948) with no measurable causal or evidential grounding.","authors":["Ummara Mumtaz","Aimen Noor","Awais Ahmed"],"categories":["cs.AI","cs.CL","cs.IR","q-bio.QM"],"primary_category":"cs.AI","announce_type":"replace-cross","date":"2026-08-25","first_seen":"2026-08-18","revised_at":"2026-08-25","abs_url":"https://arxiv.org/abs/2608.15382","pdf_url":"https://arxiv.org/pdf/2608.15382","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["医疗LLM评测","因果知识图谱","推理评估"],"reason":"评估LLM在医疗决策中的推理能力，属于NLP评测，不以人类行为为参照系","model":"deepseek-v4-pro","scored_at":"2026-08-18T13:04:42","error":null,"has_summary":false,"summary":null},{"id":"2608.21376","version":1,"title":"On the Role of Citations in Preference Data","zh_title":"引用在偏好数据中的作用研究","abstract":"Many NLP tasks require systems to provide attribution in their outputs--i.e. citations to grounding sources. Attribution serves as a bulwark against model hallucination and as a means for users to verify the credibility of model outputs. Yet, it is unclear how humans and LLMs evaluate citations when comparing outputs, a process central to reward modeling and modern LLM post-training. This paper studies the role of citations in the preferences of human judges and four open-source LLMs within the context of scientific question answering, leveraging mixed effects models to investigate the influence of citations on pairwise judgments. Among our key findings are (1) that humans prefer more diverse citations but fewer overall, and (2) that LLMs show some citation-related preferences compared to humans, despite lacking access to the sources, but these preferences depend on the data and specific models. We further discuss the implications of our findings for preference data collection.","authors":["Yu Hou","Hal Daum\\'e III","Rachel Rudinger","William Walden"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-25","first_seen":"2026-08-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.21376","pdf_url":"https://arxiv.org/pdf/2608.21376","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["引用偏好","奖励建模","NLP评测"],"reason":"研究人类与LLM对引用的偏好，属NLP评测，非仿真人类被试","model":"deepseek-v4-pro","scored_at":"2026-08-25T13:03:39","error":null,"has_summary":false,"summary":null},{"id":"2608.21766","version":1,"title":"Evaluation Awareness in Language Models: Representation, Verbalization, and Control","zh_title":"语言模型中的评估意识：表征、言语化与控制","abstract":"Both capability and safety benchmarks rest upon the assumption that the behavior of language models undergoing a test is informative about their behavior in deployment. This assumption can fail, should models infer that they are being evaluated and condition their response on such context. This hypothesis, termed ``evaluation awareness'', has been observed in frontier and open-weight language models alike. We provide a systematic study of this phenomenon, by probing for it across six language models (from four families and three sizes) and three metrics. More precisely, we examine whether (i) being under evaluation is linearly represented within the models' activations space, (ii) it is verbalized in their output tokens (as scored by an LLM-as-judge), and (iii) steering causally affects their behavior. For the open-checkpoint Olmo models, we further test these measures at every training stage. In doing so, we report that evaluation awareness is linearly decodable from the residual streams of every model (best AUROC $\\geq 0.7$). By contrast, these representations align only in part with verbalization: their correlations and mutual information are nonzero in some settings, yet vary substantially across models, layers, and readout choices. Nevertheless, steering along probe-derived directions can shift the verbalization scores. Finally, a comparison across the Olmo checkpoints reveals that evaluation awareness is already present within base models, becomes amplified throughout the stages of supervised fine-tuning, and remains stable thereafter---unlike the effects of steering, that grow more pronounced at every successive training stage. These results show the need for evaluations to account for the disjunction between what models represent internally, what they verbalize, and their steering.","authors":["Farzaneh Heidari","Amin Memarian","Guillaume Rabusseau"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-25","first_seen":"2026-08-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.21766","pdf_url":"https://arxiv.org/pdf/2608.21766","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["模型评估","可解释性","安全"],"reason":"研究模型对评估情境的感知，属模型能力评测，不以人类行为为参照，不涉及仿真人类被…","model":"deepseek-v4-pro","scored_at":"2026-08-25T13:03:43","error":null,"has_summary":false,"summary":null},{"id":"2608.22622","version":1,"title":"Teaching LLMs How ICU Physicians Approach Clinical Reasoning Through OMOP-Aligned Retrieval Improves Reasoning Across Clinical Domains","zh_title":"通过OMOP对齐检索教授LLM ICU医生的临床推理方法，提升跨临床领域的推理能力","abstract":"Clinical decision-making relies on identifying relevant patient information to guide diagnosis and treatment, a challenge that is especially difficult in the data-dense and rapidly changing intensive care unit (ICU). Large language models (LLMs) could support this task. However, existing applications and datasets mostly emphasize surface-level retrieval or factual recall rather than the inductive and deductive reasoning clinicians practice to select and reason over decision-relevant evidence. We hypothesized that training LLMs on expert ICU reasoning could yield clinical reasoning skills that generalize beyond critical care. Here we introduce ICU-REACT, a reasoning dataset developed with 19 clinicians through a clinician-in-the-loop framework to teach LLMs to perform information retrieval and context-aware clinical reasoning in the ICU. Using ICU-REACT, we fine-tuned Clin-REACT models spanning 8B-70B parameters and three model families. Across five clinical reasoning benchmarks, Clin-REACT consistently outperformed its backbone models and open-source general-purpose and medical LLMs. Gains extended to different tasks including script concordance tests, and downstream diagnosis and treatment tasks. These findings suggest that expert reasoning supervision in critical care can improve broader clinical reasoning, although prospective evaluation is needed before real-world clinical use.","authors":["Miguel Contreras","Scott Siegel","Subhash Nerella","Jessica Sena","Jiaqing Zhang","Heng Sun","Hruday Tej Akkaladevi","Peiyu Lu","Jordan Rosen","Sumit Kapoor","Sasank Desaraju","Grace R. Thompson","Jacob Purcell","Michael Petrauskis","Philip KW. Hong","Meghan Brennan","Sarah Chrabaszcz","Tierra Smith","Ronnie Ren","Michel S. Kabbash","Ceyhun Haziroglu","Rushi Patel","Gabriel Gomez","Charlotte Chaiklin","Randy Leung","Kenneth N. John","Whitman Wiggins","Philip Kayser","Vincent Bird","Maria Bruzzone","Tyler J. Loftus","Azra Bihorac","Parisa Rashidi"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-25","first_seen":"2026-08-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.22622","pdf_url":"https://arxiv.org/pdf/2608.22622","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["临床推理","LLM微调","医学NLP"],"reason":"该研究训练LLM进行临床推理，属于医学NLP能力提升，不涉及用LLM仿真人类被…","model":"deepseek-v4-pro","scored_at":"2026-08-25T13:03:51","error":null,"has_summary":false,"summary":null},{"id":"2608.22802","version":1,"title":"SDoH-Aware Narrative Anchoring Bias in Medical LLMs for Trustworthy Clinical Decision Support","zh_title":"医学大语言模型中SDoH感知的叙事锚定偏差研究","abstract":"Medical large language models are often judged by how many clinical questions they answer correctly. That view is useful, but it misses a practical risk. A model may know the right answer and still change its response when the same case is written in a different patient voice. This paper evaluates that risk as SDoH aware narrative anchoring bias. We use NarrativeShield SDoH MedQA, a counterfactual medical question answering dataset in which each case appears in persona based narratives while the answer key remains fixed. The dataset is reshaped from wide format into case grouped persona rows. We evaluate three open source instruction tuned LLMs from the Qwen2.5 family: 1.5B, 3B, and 7B. The final experiment uses 300 clinical cases and produces 8,100 model responses across three prompting conditions. We report persona level accuracy, counterfactual consistency, correct consistency, and narrative sensitivity error. Qwen2.5 7B achieves the best accuracy at 56.33 percent and the best correct consistency at 40.33 percent. Paired McNemar exact tests show significant accuracy gains for 7B over 3B in all prompt settings. Even so, narrative sensitivity remains, with the lowest error still at 31.67 percent. These results suggest that trustworthy clinical decision support should be evaluated by both average correctness and stability across medically equivalent patient narratives.","authors":["Ahnaf Atef Choudhury","Ramkrishna Saha"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-25","first_seen":"2026-08-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.22802","pdf_url":"https://arxiv.org/pdf/2608.22802","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["LLM评测","临床决策支持","叙事偏差"],"reason":"评估LLM在医学问答中的稳定性，属NLP能力评测，不以人类行为仿真为目标","model":"deepseek-v4-pro","scored_at":"2026-08-25T13:03:56","error":null,"has_summary":false,"summary":null},{"id":"2608.23248","version":1,"title":"Future Querying: Can LLMs Serve as Implicit Medical World Models?","zh_title":"未来查询：LLM能否作为隐式医学世界模型？","abstract":"Traditional clinical prediction models rely on task-specific pipelines and curated, structured data, which scale poorly and underutilize unstructured text. To address this, we introduce future querying, a paradigm that probes whether large language models (LLMs) can function as implicit medical world models by evaluating their ability to answer time-indexed clinical queries about a patient's future. Our framework operates on unstructured clinical documentation using endpoint-agnostic training, enabling a single model to answer diverse clinical queries over patient trajectories without manual feature engineering or task-specific retraining. We show that small, locally fine-tuned open-weight models can match or approach larger proprietary systems, making the framework suitable for privacy-preserving, on-premise deployment. Evaluated on a new synthetic medical reports dataset and real ICU notes from the MIMIC-IV dataset, our results provide encouraging evidence that LLMs can capture aspects of clinical dynamics.","authors":["Siri Willems","James Butterworth","Lore Goetschalckx","Peter Vrancx","Philippe Modard","Elke Giets","Ludovic Denoyer"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-25","first_seen":"2026-08-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.23248","pdf_url":"https://arxiv.org/pdf/2608.23248","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["临床预测","LLM评测","医疗AI"],"reason":"用LLM做临床预测，属于NLP能力评测，不以人类行为仿真为标的。","model":"deepseek-v4-pro","scored_at":"2026-08-25T13:04:01","error":null,"has_summary":false,"summary":null},{"id":"2608.22251","version":1,"title":"CAIA in Practice: Field Evaluation of an AI-Assisted Support System for Text-Based Online Counselling","zh_title":"CAIA实践：基于文本的在线咨询中AI辅助支持系统的现场评估","abstract":"Rising global demand for mental health support creates significant service delivery challenges, with asynchronous email counselling serving as a crucial low-threshold channel for accessing care. This paper presents CAIA, a co-designed AI-based tool suite that demonstrates responsible AI integration into counselling practice through seven LLM-driven functions enhanced by retrieval-augmented generation. A field evaluation involved 34 professional counsellors conducting authentic sessions with trained student counsellees (36 threads, 321 messages, 1,257 AI outputs). User behaviour analysis confirms substantial adoption, revealing that professional autonomy and information accuracy are decisive for sustained acceptance, with counsellors particularly valuing interpretive functionalities that provide new perspectives and stimulate professional reflection.","authors":["Philipp Steigerwald","Nico Bienlein","Jennifer Burghardt","Mara Stieler","Robert Lehmann","Jens Albrecht"],"categories":["cs.HC","cs.CL"],"primary_category":"cs.HC","announce_type":"cross","date":"2026-08-25","first_seen":"2026-08-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.22251","pdf_url":"https://arxiv.org/pdf/2608.22251","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["AI辅助咨询","人机交互","系统评估"],"reason":"AI辅助心理咨询工具，非人类仿真实验，无对照人类行为数据","model":"deepseek-v4-pro","scored_at":"2026-08-25T13:03:47","error":null,"has_summary":false,"summary":null},{"id":"2608.22068","version":1,"title":"Decision-Support and Modeling with Large Language Models for Geothermal Well Arrays","zh_title":"基于大语言模型的地热井阵列决策支持与建模","abstract":"Geothermal well arrays, which organize multiple geothermal wells into carefully planned geometric configurations, provide opportunities to enhance energy production capacity and increase fault tolerance. The development and adoption of these emerging geothermal technologies could be accelerated through the recent advances in large language models (LLMs) and high-level high-performance languages. A challenge in LLM-based applications is the reliability of the generated outputs, as they can be prone to subjective biases and hallucinations. This study assesses the potential of cutting-edge LLMs - such as ChatGPT, Gemini, Claude, Grok, and domain-specific models like AskGDR - as expert assistants that can synthesize insightful interpretations of complex geothermal data, as well as improve feature capabilities of geothermal models and numerical software. We developed a novel approach, leveraging Google's recently introduced AI assistant, NotebookLM, to accelerate the generation of unpublished quantitative geothermal benchmarks. The rapid generation of these evaluation instruments is essential for assessing the swiftly evolving capabilities of emerging language model technologies. In particular, we use these benchmarks and LLM-based interviews to analyze opportunities and limitations of two promising technologies: geothermal well arrays and closed-loop coaxial wells. Furthermore, we present a case study illustrating how LLMs can facilitate auto-parallelization of geothermal numerical models. Our analysis emphasizes their application in digital twins and underscores the importance of high-level, high-performance code generation. This line of research could play a transformative role in the geothermal sector by enabling the next-generation of decision-support applications, integrating data analysis, informed recommendations, and more dynamic numerical modeling workflows.","authors":["Edwin Ouko","Emmanuel Lujan","Alan Edelman","Robert Metcalfe"],"categories":["cs.AI","cs.CE"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-25","first_seen":"2026-08-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.22068","pdf_url":"https://arxiv.org/pdf/2608.22068","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C2"],"tags":["地热能源","LLM应用","决策支持"],"reason":"论文用LLM辅助地热井阵列决策与建模，属工程仿真，不涉及人类行为仿真或对照。","model":"deepseek-v4-pro","scored_at":"2026-08-25T13:03:44","error":null,"has_summary":false,"summary":null},{"id":"2608.23058","version":1,"title":"LLM-based Agents for Forecasting and Prediction: Methods, Training, Evaluation, and Applications","zh_title":"基于LLM的预测代理：方法、训练、评估与应用","abstract":"Large language models (LLMs) now support forecasting systems that combine language-based reasoning with temporal data, evidence retrieval, external tools, and iterative prediction. We investigate LLM-based forecasting agents, meaning systems in which a language model contributes to a scored prediction about a future or currently unobserved target. We organize architectures into three groups. Standalone LLM workflows operate on encoded time series or event context. Tool- and retrieval-augmented agents incorporate external evidence. Hybrid systems pair LLMs with statistical or foundation models. We then review training methods and evaluation protocols. We examine negative as well as positive evidence, including sensitivity to small input perturbations, ablations in which the LLM component does not improve accuracy, and benchmark gains that may reflect contamination instead of temporal reasoning. We cover applications in finance, weather, health, energy, and operations, and we summarize the benchmarks and datasets used for evaluation. The evidence indicates that measurement is a central limitation. Future work requires calibration under distribution shift, contamination-resistant live evaluation, explicit reporting of cost and accuracy together, and methods for handling feedback between deployed forecasts and the outcomes being forecast.","authors":["Xiaogang Xu","Jiaqi Tang","Jianmin Chen","Yingying Yan","Zhenchao Tang","Xiangxin Zhou","Xiaobin Hu","Wei Wei","Jinfeng Wu","Qifeng Chen","Lu Zhou","Jiafei Wu","Zhe Liu","Jianwei Yin","Weimin Zheng"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-25","first_seen":"2026-08-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.23058","pdf_url":"https://arxiv.org/pdf/2608.23058","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["LLM预测","时间序列","多智能体"],"reason":"LLM预测代理用于时间序列预测，非人类行为仿真，无人类被试对照。","model":"deepseek-v4-pro","scored_at":"2026-08-25T13:03:34","error":null,"has_summary":false,"summary":null},{"id":"2608.23308","version":1,"title":"FIDES: A Concordance Protocol for LLM-Generated Trading Strategies","zh_title":"FIDES：LLM生成交易策略的一致性协议","abstract":"An LLM asked for a trading strategy returns three artifacts at once: a natural-language rationale, an executable implementation, and once run, a track record. Whether these are the same object is rarely checked. We present FIDES, a measurement protocol that treats them as three views to be reconciled rather than one deliverable to be graded. Through dual delivery, a single model call returns both a natural-language strategy with an explicit claimed edge and a self-contained strategy(df) function. FIDES executes the code in a sandbox against a lag-one out-of-sample backtest and scores three concordance gaps: say to do, do to real, and say to result. On 8 liquid US ETFs across four models plus a two-stage elicitation arm, 40 strategies, 2023 to 2024 out-of-sample, three findings stand out. First, concordance does not predict profit: only 2 of 40 strategies beat buy-and-hold, and a plain sma(50,200) rule outperforms every model's mean Sharpe. Second, self-assessment is badly calibrated: 32 of 40 strategies claim to beat buy-and-hold and exactly one does. Third, swapping the language-code judge for a second model flips say to do on more than half of items. Injecting Close.shift(-1) drops do to real by 0.33 on average, while our runtime future-information probe fired on neither clean nor injected code. We frame FIDES as a protocol for measurement fidelity, not a claim about market performance.","authors":["Arther Tian","Alex Ding","Simon Wu","Aaron Chan"],"categories":["cs.CR","cs.AI"],"primary_category":"cs.CR","announce_type":"cross","date":"2026-08-25","first_seen":"2026-08-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.23308","pdf_url":"https://arxiv.org/pdf/2608.23308","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["LLM交易策略","一致性评估","多智能体"],"reason":"LLM生成交易策略并评估一致性，属多智能体协作与代码执行，无人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-08-25T13:04:03","error":null,"has_summary":false,"summary":null},{"id":"2608.22459","version":1,"title":"\"I want to be pushed, I want to grow\": Enabling social workers to design evaluations of LLM augmentation in their work","zh_title":"“我想被推动，我想成长”：赋能社会工作者设计其工作中LLM增强的评估","abstract":"Workers are increasingly asked to adopt AI systems to assist their work, yet are rarely given a voice in defining what meaningful AI augmentation should look like or how to evaluate for it. In this paper, we propose worker-driven AI measurement---a bottom-up approach to AI evaluation where workers collaboratively shape decisions about which tasks AI should augment, what \"successful\" augmentation looks like, and how it should be measured. We explore how to support this through a case study with 19 workers from a local school social work organization. Through a series of eight workshops, workers iteratively develop their own measurement goals for AI evaluation, systematize these goals, and then design a benchmark to capture how effectively an LLM can \"challenge\" them to reflect on their own assumptions and biases in the context of their day-to-day work. Workers collaboratively design and refine an LLM-as-a-judge rubric based on their professional and lived expertise. In validations of the worker-created benchmark, we find that there is strong agreement between worker and LLM judge ratings and that the resulting benchmark can differentiate performance across six state-of-the-art LLMs. Based on our case study, we discuss opportunities for future work to support worker-driven AI measurement as a complementary approach to existing top-down AI evaluation approaches.","authors":["Anna Kawakami","Chloe Qianhui Zhao","Renee Shelby","Fernando Diaz","Haiyi Zhu","Kenneth Holstein"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-08-25","first_seen":"2026-08-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.22459","pdf_url":"https://arxiv.org/pdf/2608.22459","source_feed":"cs.HC","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["人机协作","AI评估","社会工作"],"reason":"研究社工与LLM协作设计评估，非用LLM仿真人类被试，无人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-08-25T13:03:49","error":null,"has_summary":false,"summary":null},{"id":"2608.22660","version":1,"title":"Evaluation in the Age of AI: Output as Evidence of Learning","zh_title":"AI时代的评估：输出作为学习证据","abstract":"The rapid adoption of artificial intelligence (AI), particularly large language models (LLMs), has fundamentally disrupted how learning is demonstrated and evaluated in higher education. Tasks that once served as proxies for understanding-such as writing essays, solving problem sets, or producing computer code-can now be generated superficially by AI systems with minimal human effort. This paradigm shift raises a critical ethical question: how should learning be evaluated when traditional indicators of competence are easily outsourced? This paper examines the ethical challenges of educational evaluation in the age of AI from a university-level perspective. We argue that the core problem extends beyond academic dishonesty to a deeper misalignment between assessment practices and the learning outcomes they are intended to measure. Evaluation regimes that rely on artificial constraints risk measuring compliance, access, or concealment rather than genuine understanding, reasoning, or judgment. By analyzing institutional responses and presenting empirical survey data, we highlight the need for alternative assessment models that emphasize process over product. The goal is to establish ethically informed assessment strategies that preserve student agency and accountability in an automated age.","authors":["Md Zarzees Uddin Shah Chowdhury","Samin Rahman Khan"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2026-08-25","first_seen":"2026-08-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.22660","pdf_url":"https://arxiv.org/pdf/2608.22660","source_feed":"cs.CY","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["教育评估","学术诚信","AI伦理"],"reason":"论文讨论AI对教育评估的影响，不涉及用LLM仿真人类被试或与真实人类数据对照。","model":"deepseek-v4-pro","scored_at":"2026-08-25T13:03:52","error":null,"has_summary":false,"summary":null},{"id":"2608.22154","version":1,"title":"More accurate behavioral predictions with hybrid Bayesian-connectionist models","zh_title":"用混合贝叶斯-连接主义模型实现更准确的行为预测","abstract":"Researchers must often choose between Bayesian or neural network models of behavior, two paradigms with complementary strengths and weaknesses. An ideal paradigm would facilitate testing many kinds of representations and inductive biases; Bayesian models make this easy, while neural networks do not. Similarly, an ideal paradigm would avoid over-simplifications; neural networks make this easy, while Bayesian models do not. Here, we introduce Bayesian distillation with Behavioral Tuning (BBT) as an approach to getting the best of both traditions. BBT offers a simple recipe for model building: first, a neural network is trained to mimic a Bayesian model through synthetic data, and second, the network is fine-tuned on human behavior to capture additional structure and nuance. Across four case studies in human concept learning, we find that BBT outperforms traditional approaches at predicting human behavior while also revealing psychological insights, resulting in models that can both mimic Bayesian priors and capture heuristics and biases that violate simple modeling assumptions.","authors":["Brenden M. Lake","Akshay K. Jagadish","Guangyuan Jiang"],"categories":["cs.LG"],"primary_category":"cs.LG","announce_type":"new","date":"2026-08-25","first_seen":"2026-08-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.22154","pdf_url":"https://arxiv.org/pdf/2608.22154","source_feed":"cs.LG","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["认知建模","贝叶斯模型","神经网络"],"reason":"论文用混合模型预测人类行为，但未使用LLM作为被试替代品，属于认知建模而非LL…","model":"deepseek-v4-pro","scored_at":"2026-08-25T13:03:45","error":null,"has_summary":false,"summary":null},{"id":"2607.23458","version":2,"title":"Two Regimes of Chain-of-Thought Unfaithfulness: Metric-Based Detection Fails Where Models Are Wrong","zh_title":"思维链不忠实性的两种状态：基于度量的检测在模型出错时失效","abstract":"Chain-of-thought (CoT) explanations support oversight only if they are faithful: the stated reasoning must actually produce the answer. Auditing black-box (behavioral) detection of unfaithful CoT against FaithCoT-Bench's human annotations, we find answer correctness structures the problem at every level. Answer incorrectness alone (an oracle diagnostic, not a deployable detector) outperforms every purpose-built signal (AUROC 0.696), because 69% of annotated unfaithfulness occurs on incorrect answers. Stratifying by correctness splits detection into two regimes: on correct answers, behavioral signals moderately separate faithful from post-hoc reasoning (0.63-0.67); on incorrect answers, where most unfaithfulness lives, no tested signal is detectably above chance (replicated on all four models for benchmark-wide signals). The standard step-removal metric anti-correlates with human labels; this inversion reproduces on the benchmark's released scores and on hint-dependent counterfactually labeled traces. Linear probes decode the behaviorally blind regime in Llama-3.1-8B and the correct-answer regime in Qwen-2.5-7B, with no shared, positively aligned direction detected across regimes; instructed answer-first traces (7 models) transfer to neither annotated regime, while hint-induced unverbalized answer flips do, in model- and source-dependent settings. We also independently verify and resolve a documentation-data mismatch in the benchmark's label semantics.","authors":["Suramya R. Angdembay","Dikshant Aryal","Nick Rahimi"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-08-25","first_seen":"2026-07-28","revised_at":"2026-08-25","abs_url":"https://arxiv.org/abs/2607.23458","pdf_url":"https://arxiv.org/pdf/2607.23458","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["思维链忠实性","模型可解释性","NLP评测"],"reason":"研究CoT忠实性检测，属NLP能力评测，不以人类行为为参照，不涉及LLM仿真人…","model":"deepseek-v4-pro","scored_at":"2026-08-25T13:04:09","error":null,"has_summary":false,"summary":null},{"id":"2608.01347","version":5,"title":"Prompt-Induced Waste in Coding Agents: Reasoning, Effort, Harness Design, and End-to-End Cost","zh_title":"大型推理模型中的提示诱导浪费：编码智能体的预注册双框架基准测试","abstract":"Coding-agent efficiency cannot be characterized by token count or model price alone. End-to-end cost and task success depend jointly on prompt semantics, inference effort, harness policy, model, task difficulty, tool use, context management, and provider accounting. Controlled experiments show that prompt wording can change reasoning and verification behavior without changing the task, that additional inference effort can help on difficult tasks but can also add cost without benefit, and that the value of an efficiency intervention can change when the harness changes. These results show that prompt, effort, and harness are interacting experimental factors rather than independent controls. We model efficiency as cost per successful task induced by the agent trajectory. Token and cache counts are measurements of that trajectory, not sufficient optimization targets. Agent evaluations should therefore measure success and end-to-end cost while controlling the system variables that determine how the trajectory is produced.","authors":["Sarel Weinberger","Amir Hozez"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-08-25","first_seen":"2026-08-04","revised_at":"2026-08-25","abs_url":"https://arxiv.org/abs/2608.01347","pdf_url":"https://arxiv.org/pdf/2608.01347","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["LLM智能体","提示工程","成本优化"],"reason":"纯多智能体编码任务，无人类行为对照，属C1排除项。","model":"deepseek-v4-pro","scored_at":"2026-08-04T13:04:46","error":null,"has_summary":false,"summary":null},{"id":"2608.01575","version":2,"title":"Measuring in-context algorithmic reasoning in language models against an exact Bayes-optimal reference","zh_title":"用精确贝叶斯最优参考测量语言模型中的上下文算法推理","abstract":"Whether large language models perform algorithmic inference or pattern completion is hard to test, because most benchmarks supply answers but no distributional reference for what the shown evidence licenses. F-ICL supplies one exactly: we exhaustively enumerate the 86 million valid programs of length at most 13 on a Turing-complete machine F, complement-symmetrised to remove output-polarity bias, and compute the exact posterior under a declared bounded Levin--Solomonoff prior. It is Bayes-optimal for that stated prior rather than universal, and models are never told it exists, so the score reads the inductive prior their served distribution already encodes. Across 105 serving configurations spanning open models from 0.8B to 675B and frontier systems, models answer up to 92% of queries correctly, yet 45 of the 46 exposing distributions sit farther from the F reference than a keystroke reference. This is not an artefact of task selection: on the bit coordinate, the half the length quota cannot distort, 69 of 80 runs stay below the anchor. Fidelity is inert to scale, which accuracy tracks; continuation improves late without converging; and models un-solve a solved task once per two gains, where the F reference does so once per nine and always repairs it. Because absolute distances are reference-dependent, we prove sequential bounds holding for rival priors: any predictor whose prior gives the reference positive weight has bounded cumulative excess loss, and, in a loss never invoking the reference, any Bayesian mixture giving the realised truth positive mass has a bounded truth-loss budget. On 23,998 trajectories, 86.7% already spend over 10 bits of it. Sequences ending by position nine cannot exclude an arbitrarily large finite constant, so these are lower bounds on what a rival prior must already pay. F-ICL is an open benchmark and toolkit.","authors":["Luan Ozelim","Hector Zenil"],"categories":["cs.LG","cs.AI"],"primary_category":"cs.LG","announce_type":"replace-cross","date":"2026-08-25","first_seen":"2026-08-04","revised_at":"2026-08-25","abs_url":"https://arxiv.org/abs/2608.01575","pdf_url":"https://arxiv.org/pdf/2608.01575","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["算法推理","基准测试","贝叶斯最优"],"reason":"纯算法推理基准测试，无人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-08-25T13:03:39","error":null,"has_summary":false,"summary":null},{"id":"2608.09289","version":2,"title":"Accurate but Natural? Diagnosing Grammatical and Idiomatic Gaps in Japanese EFL Writing","zh_title":"准确但自然？诊断日本英语学习者写作中的语法和惯用差距","abstract":"Second language writing research distinguishes grammatical accuracy from native-like idiomaticity, yet automated writing evaluation often conflates these dimensions. This study introduces a layered LLM-correction pipeline that isolates structural errors from unnaturalness by generating literal error corrections and idiomatic revisions for 3,830 English writing samples from 120 Japanese junior high school students. Applying the regex-based CEFR-J grammar extractor, we quantify two diagnostic measures: accuracy gaps (structures attempted but incorrectly produced) and idiomatic gaps (grammatically correct structures underused or overused relative to native norms). Results reveal distinct patterns: definite articles, third-person singular -s, and modals (would, could) exhibit significant accuracy difficulties, while -ing forms and hypothetical modals (would) show the largest idiomatic underuse, with simple present verbs, subject-verb-object patterns, and modal can conversely exhibiting the most pronounced overuse. A two-dimensional instructional typology maps error rates against idiomatic gaps, distinguishing accurate but overused grammar items from error-prone or avoided complex forms requiring targeted production practice. This framework advances pedagogical feedback by enabling teachers to diagnose whether learner difficulties arise from inaccurate execution, structural avoidance, or L1-mapped overreliance, supporting evidence-based interventions tailored to the specific needs of each learner.","authors":["Steve Woollaston","Brendan Flanagan","Hiroaki Ogata"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-08-25","first_seen":"2026-08-11","revised_at":"2026-08-25","abs_url":"https://arxiv.org/abs/2608.09289","pdf_url":"https://arxiv.org/pdf/2608.09289","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["自动写作评估","二语习得","语法纠错"],"reason":"纯NLP写作评估，无LLM仿真人类被试或行为对照","model":"deepseek-v4-pro","scored_at":"2026-08-11T13:05:15","error":null,"has_summary":false,"summary":null},{"id":"2608.21462","version":1,"title":"CyrillicQA: The Influence of Phonetically Encoded Secret Language on LLM Performance","zh_title":"CyrillicQA：语音编码秘密语言对LLM性能的影响","abstract":"Due to the selection of their training data, large language models (LLMs) perform best on standard-language inputs from languages using the Latin alphabet with large speaker populations, while disadvantaging other language varieties. Nevertheless, they can also be a versatile tool for preserving precisely such endangered languages. But do they also possess the necessary creativity and capacity for abstraction to decode phonetically encoded language the same way humans do?","authors":["Erik Thureck","Leo S. R\\\"dian"],"categories":["cs.CL","cs.AI","cs.LG"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-25","first_seen":"2026-08-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.21462","pdf_url":"https://arxiv.org/pdf/2608.21462","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["LLM评测","语言编码","NLP能力"],"reason":"论文评测LLM对语音编码语言的理解能力，属于纯NLP能力评测，不以人类行为仿真…","model":"deepseek-v4-pro","scored_at":"2026-08-25T13:03:41","error":null,"has_summary":false,"summary":null},{"id":"2608.22124","version":1,"title":"LLM assisted writing deserves empirical evaluation","zh_title":"LLM辅助写作值得实证评估","abstract":"LLM-assisted writing is often treated as a detection problem, as it raises questions about clarity, integrity, equity, and evaluation. An analysis of 69,209 Health Informatics papers links it to more focused presentation, broader citation practices, and more globally distributed authorship. These patterns do not prove better science, but they support evaluating manuscripts by scholarly quality and accountability rather than by tool use.","authors":["Xuan Zhong Feng","Yi Lin","Yiye Zhang","Chunhua Weng","Yifan Peng"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-25","first_seen":"2026-08-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.22124","pdf_url":"https://arxiv.org/pdf/2608.22124","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["LLM辅助写作","学术出版","文献计量分析"],"reason":"研究LLM辅助写作对论文特征的影响，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-08-25T13:03:45","error":null,"has_summary":false,"summary":null},{"id":"2608.22948","version":1,"title":"What Proves You Wrong: Benchmarking Language Models on Falsifiable Research Ideation","zh_title":"什么能证明你错了：在可证伪的研究构思上基准测试语言模型","abstract":"Large language models are increasingly used to propose research ideas, yet the prevailing ways of judging such ideas supply no shared decision rule: free-form judging sways with style and position, and scoring against a later paper rewards recovery of one realized trajectory. We introduce a benchmark that carries a proposal from Literature to Test: the Lit2Test benchmark centers on a six-field contract organized around a falsifying outcome, so that every proposal precommits the observation that would prove it wrong, making its quality decidable in the first place rather than merely arguable. Built prospectively from 200 real-paper neighborhoods, Lit2Test elicits proposals from four frontier models and compares them through 1,200 pairwise comparisons judged blind in both presentation orders. The protocol audits its own reliability through diagnostic controls and bounded human calibration, with three annotators corroborating the conclusions within explicitly stated reliability bounds. Lit2Test recovers a strict ranking of the four models in all 10,000 bootstrap replicates, and the separation comes from the quality of the proposed tests and metrics rather than from surface fluency. We release the benchmark, construction pipeline, and audit artifacts for public use.","authors":["Ziyue Wang (State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University)","Aomufei Yuan (Peking University)","Yiran Yao (Tianjin University)","Linli Yao (State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University)","Hongyao Zuo (Tianjin University)","Ziwen Gong (Hainan University)","Yuanxin Liu (State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University)","Shicheng Li (State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University)","Yishuo Cai (State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University)","Tong Yang (Peking University)","Xu Sun (State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University)","Xiaohui Li (Huawei Technologies)","Haoli Bai (Huawei Technologies)"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-25","first_seen":"2026-08-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.22948","pdf_url":"https://arxiv.org/pdf/2608.22948","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["研究想法生成","基准评测","可证伪性"],"reason":"论文是LLM研究想法生成的基准评测，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-08-25T13:03:58","error":null,"has_summary":false,"summary":null},{"id":"2608.23448","version":1,"title":"How Useful are LLMs for Grammar Engineering? Cantonese ParGram Resources and Controlled Experimental Evaluation with English Baselines","zh_title":"LLM对语法工程有多大用处？粤语ParGram资源及基于英语基线的受控实验评估","abstract":"This paper presents new Cantonese ParGram resources and evaluates LLMs for knowledge-driven grammar engineering within a controlled experimental paradigm. Using Cantonese ParGram resources as gold standards, with corresponding English baselines, we investigate whether OpenAI's gpt-oss-120b and GPT-5.4 can generate machine-processable grammars from sentences and target formal structures under systematically varied prompting conditions. GPT-5.4 outperformed gpt-oss-120b, while grammars generated from target formal structures generally outperformed those generated from sentences. Although both models could generate locally plausible phrase-structure rules, lexical entries, and templates, they often struggled to coordinate interacting formal constraints, especially in multi-construction settings. The results characterize both the capabilities and limitations of current LLMs for potential integration into AI-assisted expert workflows: LLMs may support intermediate stages of grammar development, but human linguistic expertise remains central to analysis, validation, and refinement. The study also contributes new Cantonese symbolic grammatical resources.","authors":["Chit-Fung Lam"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-25","first_seen":"2026-08-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.23448","pdf_url":"https://arxiv.org/pdf/2608.23448","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["语法工程","LLM评测","计算语言学"],"reason":"评估LLM生成语法规则，属NLP能力评测，无人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-08-25T13:04:05","error":null,"has_summary":false,"summary":null},{"id":"2608.21841","version":1,"title":"AI Watchdog: Agent Interfaces for Detecting and Defending Against Manipulative Dark Patterns in AI Conversations","zh_title":"AI看门狗：用于检测和防御AI对话中操纵性暗模式的代理界面","abstract":"Conversational AI increasingly shapes consequential decisions, yet users have limited support for recognizing and resisting manipulation. We present AI Watchdog, a browser-based agent interface that monitors live conversations, detects five dark-pattern categories, including sycophancy, brand bias, anthropomorphization, sneaking, and harmful generation, and alerts users when they occur. Its open-weight turn-level classifier supports independent deployment and a path toward local inference, preserving user privacy while remaining separate from the conversational AI. We evaluated AI Watchdog in a preregistered, five-condition between-subjects experiment (N = 150) comparing a no-intervention control with four configurations varying nudge timing (prebunking vs. just-in-time) and engagement mode (without vs. with cognitive forcing). Results show that participants rarely flagged manipulative turns across all conditions, and post-task awareness did not differ significantly across groups. However, just-in-time warnings without cognitive forcing were the only intervention to significantly reduce compliance with AI-steered recommendations containing dark patterns, lowering compliance from 71.7% to 53.7%, an 18 percentage-point reduction. Exploratory analyses further showed that lower misinformation susceptibility was associated with greater flagging but not lower compliance, while higher AI trust was associated with greater compliance and lower reported awareness. Together, these findings suggest that explicit recognition of conversational dark patterns and behavioral resistance to AI steering may be distinct outcomes, motivating further investigation of timely, low-friction defensive interfaces.","authors":["Rachel Poonsiriwong (Pub)","Chayapatr (Pub)","Archiwaranguprok","Constanze Albrecht","Monchai Lertsutthiwong","Pattie Maes","Pat Pataranutaporn"],"categories":["cs.AI","cs.HC"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-25","first_seen":"2026-08-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.21841","pdf_url":"https://arxiv.org/pdf/2608.21841","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["AI安全","人机交互","暗模式检测"],"reason":"研究AI对话中的操纵检测与防御，不涉及用LLM仿真人类被试或与人类数据对照。","model":"deepseek-v4-pro","scored_at":"2026-08-25T13:03:43","error":null,"has_summary":false,"summary":null},{"id":"2608.22085","version":1,"title":"Dissecting Neuro-Symbolic Quality Assurance for Synthetic Oncology Data Generation","zh_title":"剖析用于合成肿瘤数据生成的神经符号质量保证","abstract":"Synthetic clinical data generation with large language models addresses the scarcity that limits cancer staging research, but oncology hallucinations are categorically harmful: one clinically impossible staging assignment contaminates every downstream model trained on it. Neuro-symbolic pipelines validate during generation, yet the contribution of individual quality-assurance components remains unclear. We report three controlled studies isolating gate necessity, constraint attribution, and retrieval conditionality, holding generation protocol, diversity thresholds, and fine-tuning hyperparameters constant across adapter conditions. The symbolic gate enforces schema completeness, ontology coverage against the Systematized Nomenclature of Medicine, and staging-logic consistency under American Joint Committee on Cancer eighth-edition rules. Ungated, 29.9% of records contain schema failures and 20.1% contain clinically invalid staging. Schema validation is the load-bearing filter: within the fully gated corpus it rejects 148 of 512 records, ontology grounding a further 24, and staging-logic validation none---the only generator producing logic violations is already excluded on schema, making clinical-logic validation a generator-conditional safeguard rather than the dominant filter. Retrieval augmentation is strongly model-dependent: it improves gate compliance for one generator by 12.5 percentage points, has no measurable effect for a second, and collapses output in a third. Across gated configurations ontology density is largely unchanged, indicating that symbolic validation improves clinical validity rather than vocabulary richness. Symbolic gating therefore buys corpus validity but no commensurate gain on real lung-cancer notes in this study; retrieval should be evaluated per model, and ontology density should not be reported as a proxy for corpus quality.","authors":["Laxmigayathri Challa","Yuhan Zhou","Ana Cleveland","Haihua Chen"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-25","first_seen":"2026-08-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.22085","pdf_url":"https://arxiv.org/pdf/2608.22085","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C2"],"tags":["合成数据生成","神经符号方法","医疗数据质量"],"reason":"研究合成肿瘤数据生成与质量保证，不涉及用LLM仿真人类被试或与真实人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-08-25T13:03:45","error":null,"has_summary":false,"summary":null},{"id":"2608.22610","version":1,"title":"Coalition-Aware Skill Reliability for Self-Evolving Agents","zh_title":"自进化智能体的联盟感知技能可靠性","abstract":"Agent skills, structured artifacts distilled from interaction trajectories and dynamically reused from skill banks, have become a central mechanism for enabling large language model (LLM)-based self-evolving agents to learn from past experience. Yet existing work has largely focused on the operational aspects of skills, such as acquisition, evolution, and retrieval, while leaving a more fundamental reliability question unresolved: Do accumulated skills in an agent's skill bank actually make positive mechanistic contributions? We investigate this question through systematic skill-bank audits across alternative bank compositions and deployment domains, measuring the resulting changes in agent behavior. These audits reveal two recurring reliability failures: coalition pollution, where bank-level gains conceal negative coalition-level skill contributions, and cross-domain utility reversal, where source-beneficial skills reverse their effects after transfer. These findings motivate two reliability interventions: coalition-aware skill selection during skill accumulation and label-free skill masking after transfer. Coalition-Aware Skill Selection (CASS) selects more reliable candidate skills for the current bank using sampled Shapley marginals. Unsupervised Skill-Masked Coalition Optimizer (u-SMCO) masks transferred skills whose exclusion improves retrieval quality on unlabeled target-domain data. Agentic experiments on LoCoMo, LongMemEval, HotpotQA, and ALFWorld show that CASS and u-SMCO consistently improve task performance and cross-domain generalization over strong skill-based self-evolving agent baselines. Beyond accuracy, coalition-conditioned reliability modeling reduces sensitivity to noisy outcome-reward fluctuations during reinforcement learning and exposes the limits of isolation-based skill evaluation.","authors":["Qiyan Zhao","Xiaofeng Zhang","Bo Liu","Minda Chen","Wei Xiong","Jingyang Chen","Guanting Ye","Wenhao Yu","Xiaosong Yuan","Shijie Han","Da-Han Wang","Jianmin Ji","Fei Huang","Xu-Yao Zhang"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-25","first_seen":"2026-08-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.22610","pdf_url":"https://arxiv.org/pdf/2608.22610","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","技能可靠性","自进化智能体"],"reason":"研究自进化智能体的技能可靠性，属多智能体协作，不涉及人类行为仿真或对照。","model":"deepseek-v4-pro","scored_at":"2026-08-25T13:03:49","error":null,"has_summary":false,"summary":null},{"id":"2608.22797","version":1,"title":"Performance of a domain-specific large language model in answering patient questions in psychiatry","zh_title":"领域特定大语言模型在回答精神病学患者问题中的表现","abstract":"Background This study was designed to evaluate whether a domain-specific large language model (LLM) trained exclusively on patient education resources can answer questions about psychiatric medications, in a manner superior to LLM chatbots. We developed an LLM (\"MIND\") fine-tuned for clinical fidelity, trained on patient education resources from authoritative medical organizations. Methods We compared the responses of MIND, ChatGPT, and OpenEvidence to patient questions about escitalopram, using two methods: (1) computer analysis according to a rubric measuring accuracy, clarity, completeness, nuance, safety, and referral appropriateness; (2) ratings from N=10 board-licensed psychiatrists on similar metrics. Results When rated by rubric, MIND was rated highest in all domains (p<0.001). When rated by psychiatrists, ChatGPT was rated accurate more often than MIND with a negligible effect size (p=0.021, r=0.073); MIND was rated complete more often than ChatGPT with a small effect size (p<0.001, r=0.160); and MIND and ChatGPT were rated safe with the same frequency (p=0.955, r=0.002). The majority of psychiatrists preferred the responses generated by ChatGPT (57.6%) compared to MIND (42.4%, p=0.003). Conclusions MIND was able to answer many questions about escitalopram in a manner deemed accurate, complete, and safe by psychiatrists the majority of the time. However, despite MIND's ability to provide more complete responses, psychiatrists preferred ChatGPT's responses. MIND represents a step towards building safe LLM systems to enhance patient education in psychiatry.","authors":["Alexander J. Hish","Arjun Nagendran","Scott N. Compton"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-25","first_seen":"2026-08-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.22797","pdf_url":"https://arxiv.org/pdf/2608.22797","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["LLM评测","精神病学","患者教育"],"reason":"评估LLM回答患者问题的质量，属于临床应用评测，非人类仿真实验","model":"deepseek-v4-pro","scored_at":"2026-08-25T13:03:54","error":null,"has_summary":false,"summary":null},{"id":"2608.23475","version":1,"title":"StrategyBench: Evaluating Explicit Strategy Induction in Large Language Models","zh_title":"StrategyBench：评估大语言模型中的显式策略归纳","abstract":"As large language models are increasingly used in data-scarce and evolving task scenarios, few-shot in-context learning (ICL) has become a key paradigm for task adaptation. However, direct ICL often uses a small set of examples without explicitly abstracting task rules, making it sensitive to example construction. In contrast, human learners often reduce such sensitivity by first summarizing task rules from examples and then applying them to new instances. To evaluate this ability, we propose StrategyBench, which selects strategy-inducible tasks from BIG-Bench, constructs reference strategies, and defines evaluation metrics along two dimensions: strategy quality and downstream utility. We further analyze strategy induction from three perspectives: task variation, model configuration, and adaptation setting, covering category-wise differences, generator-executor choices, demonstration design, and SFT-based adaptation. Experiments show that explicit strategy utility differs substantially across task categories and depends on both strategy generation and execution conditions. The benchmark is released at: https://anonymous.4open.science/r/StrategyBench-D53C.","authors":["Jinghan Tan","Yuanzheng Wang","Lu Chen","Zijun Chen","Yuqian Wang","Maosong Sun"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-25","first_seen":"2026-08-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.23475","pdf_url":"https://arxiv.org/pdf/2608.23475","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["策略归纳","上下文学习","基准评测"],"reason":"评估LLM从示例归纳策略的能力，属于NLP能力评测，不以人类行为为参照系","model":"deepseek-v4-pro","scored_at":"2026-08-25T13:04:05","error":null,"has_summary":false,"summary":null},{"id":"2608.22959","version":1,"title":"WildHandBench: A Benchmark for Handwritten Text Understanding that Challenges MLLMs and Humans","zh_title":"WildHandBench：挑战多模态大模型与人类的手写文本理解基准","abstract":"While the top model on OmniDocBench now reaches 96.34% overall on printed-document parsing, the ability of current models to handle challenging handwritten documents remains largely uncharacterized. Existing benchmarks focus on isolated text or formulas, overlook handwritten tables and real-world degradation, and report aggregate accuracy without explaining why models fail. We present WildHandBench, a benchmark containing 500 handwritten documents across three structures (free text, tables, formulas), four languages, and nine real-world scenarios. We introduce a Prior-Driven Error (PDE) metric that quantifies whether errors originate from language priors rather than visual evidence. Evaluating 18 state-of-the-art models together with calibrated human baselines, we find: (1) the best model achieves only 71.85% overall; (2) humans outperform all models yet the gap is narrow (77.09% vs. 71.85%); and (3) model errors are qualitatively different from human errors -- 63-91% of model errors are prior-driven versus only 49% for humans, exposing systematic reliance on language priors that conventional accuracy metrics cannot capture.","authors":["Jun Zhang","Qiao Zhao","Cheng Cui","Jianying Qu","Zhongkai Sun","Jianwen Yang","Changda Zhou","ZhuoXin Liu","Shubin Han"],"categories":["cs.CV","cs.AI"],"primary_category":"cs.CV","announce_type":"cross","date":"2026-08-25","first_seen":"2026-08-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.22959","pdf_url":"https://arxiv.org/pdf/2608.22959","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["手写文本识别","多模态基准","模型评测"],"reason":"纯视觉文档理解评测，不涉及LLM仿真人类被试","model":"deepseek-v4-pro","scored_at":"2026-08-25T13:03:58","error":null,"has_summary":false,"summary":null},{"id":"2608.23300","version":1,"title":"Evaluating SAT Solver Metrics as Predictors of Human-Perceived Nonogram Difficulty","zh_title":"评估SAT求解器指标作为人类感知数织难度的预测因子","abstract":"Algorithmic solver effort is often assumed to align with perceived puzzle difficulty, but this assumption is rarely tested against human solving data. We evaluate this assumption for Nonograms, a popular logic puzzle similar to Sudoku in which numeric clues along each row and column determine a unique solution grid. We formulate Nonograms as a constraint satisfaction problem and solve them using existing SAT solvers. We then conduct a user study in which we collect data on both participant interactions and reported difficulty. We find that neither participants' reported difficulty nor their behavioural signals correlate meaningfully with SAT solver metrics; however, we find evidence that expertise moderates the relationship between solver metrics and reported difficulty. In this process, we uncover distinct, recurring solving strategies that indicate human preference for complex propagation, diverging from solver-measured complexity.","authors":["Changdao He","Yibing Ju","Jonathan Calver","Alice Gao"],"categories":["cs.HC","cs.AI"],"primary_category":"cs.HC","announce_type":"cross","date":"2026-08-25","first_seen":"2026-08-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.23300","pdf_url":"https://arxiv.org/pdf/2608.23300","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C2"],"tags":["数织","SAT求解器","人机交互"],"reason":"研究SAT求解器指标与人类感知难度的关系，不涉及LLM仿真人类被试，属于游戏/…","model":"deepseek-v4-pro","scored_at":"2026-08-25T13:04:03","error":null,"has_summary":false,"summary":null},{"id":"2608.21373","version":1,"title":"Self-Reported AI Usage for Learning in Computer Science Education: Relationships with Goal Orientation and Academic Help-Seeking","zh_title":"计算机科学教育中自我报告的AI使用：与目标导向和学业求助的关系","abstract":"Artificial intelligence (AI) is becoming an increasingly integral part of higher education, yet the factors shaping students' use of AI for learning remain insufficiently understood. This study examines how students' goal orientation and academic help-seeking behavior are associated with AI use in a computer science context, while also accounting for individual, behavioral, and contextual characteristics. Data were collected from 236 university students enrolled in a database course using a self-report survey. AI use was operationalized through two measures: self-reported frequency of use and the number of AI-supported learning activities. Hierarchical regression analyses were conducted to examine the relationships among the variables. The results indicate that help-seeking tendencies, particularly perceived help-seeking threat, were consistently associated with both more frequent reporting of AI use and self-reported engagement in a wider range of AI-supported activities. In contrast, the effects of goal orientation were more limited and less consistent across models. The results highlight the importance of considering help-seeking behavior when designing AI-supported learning environments in computer science education.","authors":["Piret Luik","Karin Naruskov","Karmen Kalk","Merle Taimalu"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2026-08-25","first_seen":"2026-08-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.21373","pdf_url":"https://arxiv.org/pdf/2608.21373","source_feed":"cs.CY","score":0,"bucket":"other","rubric_hits":["C5"],"tags":["AI使用行为","教育研究","自我报告"],"reason":"研究人类学生使用AI的行为，非用LLM仿真人类被试","model":"deepseek-v4-pro","scored_at":"2026-08-25T13:03:39","error":null,"has_summary":false,"summary":null},{"id":"2608.23406","version":1,"title":"Whose readiness counts? Disagreement within and between sectors in perceived AI and robotics preparedness","zh_title":"谁的准备度算数？感知AI和机器人准备度在部门内与部门间的分歧","abstract":"AI and Industry 4.0 readiness assessments often summarise preparedness using a single score for an organisation, application domain or sector. Those summaries can conceal disagreement about the same technology and variation among applications grouped under one sector label. We test how much information is lost through this aggregation using a card-based survey in which 982 respondents provided 15,200 readiness evaluations across 17 named AI and robotics challenges. Readiness is perceived community preparedness and available resources, not personal willingness or audited organisational capability. Respondents frequently disagreed about identical challenges, with card-level readiness standard deviations of $1.03$-$1.26$ on a five-point scale. A crossed decomposition attributes 32.7% of observed variation to stable respondent differences, 7.3% to differences among challenges, and 60.0% to response-level variation that also contains measurement error. Differences among challenge-family means account for only about 2% of variation, with substantially more variation among people, applications and person-family judgements. Manufacturing has the highest mean readiness, yet shop-floor robotics, process-optimisation AI and general decision-support applications are judged differently. Computer-science and AI/ML respondents report higher readiness than non-technical respondents across challenge families, whereas engineering respondents do not report higher Manufacturing readiness. Sector rankings are therefore best used as portfolio summaries rather than evidence that an industry is uniformly ready or behind. Readiness reporting should retain application-level disagreement, disclose whose judgements form the average, and consider ethics, cyber security, literacy and capability needs without collapsing them into a single score.","authors":["Peng Wang"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2026-08-25","first_seen":"2026-08-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.23406","pdf_url":"https://arxiv.org/pdf/2608.23406","source_feed":"cs.CY","score":0,"bucket":"other","rubric_hits":["C2"],"tags":["AI准备度","调查方法","部门差异"],"reason":"研究人类对AI和机器人准备度的感知，不涉及LLM仿真人类被试","model":"deepseek-v4-pro","scored_at":"2026-08-25T13:04:03","error":null,"has_summary":false,"summary":null},{"id":"2608.22705","version":1,"title":"Evolution of cooperation with Q-learning: how much information do we need?","zh_title":"Q学习下的合作演化：我们需要多少信息？","abstract":"Cooperation is ubiquitous in both natural and human societies, yet its evolutionary basis remains a major challenge. A long-standing puzzle is whether having more information leads to better decision-making and thus a higher level of cooperation. To address this question, we adopt a recently developed reinforcement learning framework in which individuals learn through trial and error to maximize cumulative rewards - a paradigm that has successfully explained diverse emergent patterns in human behaviors. Specifically, we equip a structured population with the Q-learning algorithm and systematically vary the size of the interactive neighborhood, which serves as a proxy for perceived information. Interestingly, we observe a non-monotonic relationship between cooperation prevalence and neighborhood size in both two-dimensional square lattices and Barabasi-Albert scale-free networks. This inverted U-shaped dependence reveals that an optimal amount of information exists, yielding the highest level of cooperation. Mechanistic analyses show that a moderate neighborhood size enables individuals to strike an optimal balance between information sufficiency and decision-making tractability. This balance allows them to detect reciprocal opportunities while avoiding the deterioration of decision quality due to information overload. Our findings challenge everyday intuition, suggesting that a proper amount of information - not more - is optimal for the emergence of cooperation.","authors":["Yile Ku","Xin Ou","Jiqiang Zhang","Shengfeng Deng","Huiji Yue","Li Chen"],"categories":["physics.soc-ph","cond-mat.dis-nn","cond-mat.stat-mech","nlin.AO","q-bio.PE"],"primary_category":"physics.soc-ph","announce_type":"new","date":"2026-08-25","first_seen":"2026-08-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.22705","pdf_url":"https://arxiv.org/pdf/2608.22705","source_feed":"physics.soc-ph","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体强化学习","合作演化","信息量"],"reason":"纯多智能体强化学习研究，无LLM，无人类数据对照，不涉及人类仿真","model":"deepseek-v4-pro","scored_at":"2026-08-25T13:03:54","error":null,"has_summary":false,"summary":null},{"id":"2606.13629","version":2,"title":"Valid Inference with Synthetic Data via Task Exchangeability","zh_title":"通过任务可交换性实现合成数据的有效推断","abstract":"There is a proliferation of work arguing for the use of synthetic data in scientific research. For example, social scientists are arguing for the use of LLM-generated \"silicon samples\" in pilot studies; AI evaluations increasingly rely on \"LLM-as-a-judge\" outputs; and proteomics research is accelerated by generative models that produce synthetic protein structures. These developments raise an intriguing possibility: synthetic data may help researchers ask more questions, run more studies, and accelerate discovery. But they also raise a fundamental concern: synthetic data can be biased, noisy, and misspecified. In this work, we propose statistical principles for using synthetic data in scientific research with provable validity guarantees. The key insight is a new technical condition that we call task exchangeability. Informally, this is a requirement that the researcher can identify historical tasks, for which real data is available, such that their current task of interest is exchangeable with the historical tasks in an appropriate mathematical sense. We develop methods for valid inference under task exchangeability, together with extensions that provide guarantees even beyond exchangeability. We demonstrate the framework on public opinion surveys with silicon samples and AI evaluation with autoraters.","authors":["Lezhi Tan","Tijana Zrnic"],"categories":["stat.ME","cs.AI","cs.LG","stat.ML"],"primary_category":"stat.ME","announce_type":"replace-cross","date":"2026-08-24","first_seen":"2026-06-11","revised_at":"2026-08-24","abs_url":"https://arxiv.org/abs/2606.13629","pdf_url":"https://arxiv.org/pdf/2606.13629","source_feed":"cs.AI","score":9,"bucket":"selected","rubric_hits":["A1","A2","A4","B1","B3"],"tags":["LLM仿真","统计推断","硅样本"],"reason":"提出任务可交换性框架，用LLM硅样本做调查推断，有真实数据对照，提供有效性保证。","model":"deepseek-v4-pro","scored_at":"2026-08-24T13:02:12","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-24","rank":2,"question":"如何在使用合成数据（如LLM生成的硅样本）进行统计推断时提供形式化的有效性保证？","design":"提出任务可交换性框架：研究者识别一组有真实数据的历史任务，假设当前任务与历史任务在数学意义上可交换，利用历史任务中合成数据与真实数据的误差分布来校正当前任务的合成数据置信区间。方法应用于LLM生成的调查回答（硅样本）和AI评估（autorater）场景。","baseline":"使用美国国家选举研究（ANES）的感觉温度计调查数据作为真实人类数据，与LLM生成的合成回答进行对比。","findings":"在任务可交换性条件下，通过历史任务的误差分布可以构造具有有限样本覆盖保证的置信区间。实证显示，仅使用合成数据的朴素区间过窄且严重有偏，而任务可交换性方法能提供有效覆盖。","reliability":"论文承认任务可交换性可能被违反，并提供了在违反时覆盖保证如何优雅退化的扩展；同时指出合成数据可能偏差、噪声和误设。","relevance":"该研究直接针对LLM作为人类被试替代品的可靠性问题，提供了统计推断框架，并用真实调查数据验证，对关注仿真有效性和偏差的研究者具有重要参考价值。","inspiration":"借鉴其利用历史任务误差分布来校正合成数据推断的方法，可迁移到经济金融场景如消费者信心调查或通胀预期调查的LLM仿真。｜例如，在政策公告的预期形成研究中，可用LLM生成模拟受访者对政策变化的预期，并与历史调查数据对比。｜设计：以LLM模拟的经济主体为被试，施加政策信息处理，测量预期变化，用真实调查数据（如密歇根消费者调查）作为基准，通过历史任务误差校正置信区间。"}},{"id":"2608.20344","version":1,"title":"Beyond Raw Transcripts: Structured Persona Extraction for LLM-Based Digital Twins","zh_title":"超越原始转录：面向LLM数字孪生的结构化人物特征提取","abstract":"LLM-based \"digital twins\" aim to simulate how an individual would behavein new environments or respond to novel questions, given some representation of that individual's prior responses. A common approach constructs this representation from survey transcripts or summaries responses. Prior work shows that compressing long transcripts into shorter LLM-generated summaries does not significantly reduce predictive accuracy, suggesting that information volume is not the primary bottleneck. In this work, we argue that the key limitation is instead structural:how persona information is organized before being provided to thesimulator model. We study this by comparing unstructured summaries with structured persona representations. First, we introduce a hand-craftedschema (BDE: Background, Decision procedure, Evaluation), grounded in consumer-behavior theory, and show that it improves predictive accuracy over raw transcripts by +1.91 percentage points on a homogeneous benchmark (Twin-2K-500), with similar gains on gpt-5.4-mini and Qwen3-8B as robustness checks. However, this fixed structure does not generalizeacross more heterogeneous tasks, where performance is statistically indistinguishable from the raw transcript baseline. To address this limitation, we propose an automatic structure-discovery pipeline in which an LLM iteratively proposes and refines task-specific persona structures and extraction prompts. On a benchmark of 13 diverse sub-studies, this approach restores performance, improving mean accuracy by +1.91 percentage points over the raw transcript baseline and eliminating significant losses observed with the fixed schema. Overall, our results suggest that the main constraint in LLM-based digital twins is not how much information is provided, but how it is structured -- and that the optimal structure depends on the task.","authors":["Iris Ye","Tianze Deng","Ozan Candogan"],"categories":["cs.CL","cs.CY"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-24","first_seen":"2026-08-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.20344","pdf_url":"https://arxiv.org/pdf/2608.20344","source_feed":"cs.CL","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM数字孪生","人类仿真","结构化表征"],"reason":"直接研究LLM数字孪生仿真个体行为，并与真实人类数据对照，评估结构化表征对预测…","model":"deepseek-v4-pro","scored_at":"2026-08-24T13:01:52","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-24","rank":3,"question":"在基于LLM的数字孪生中，个体先验信息的组织结构如何影响其在新任务上的预测准确性？","design":"使用LLM作为提取器和模拟器，从Twin-2K-500数据集的原始问答记录中提取个体画像，然后让模拟器基于该画像预测留出问题的答案。比较了三种画像表示：原始记录、非结构化摘要、结构化画像（手工设计的BDE结构和自动发现的结构）。","baseline":"Twin-2K-500数据集（500个输入问题、88个留出问题、17个预测任务、2000多名受访者）和Mega-Study的19个子研究，均包含真实人类回答。","findings":"手工设计的BDE结构在Twin-2K-500上比原始记录提高1.91个百分点，但在异构的Mega-Study上无显著优势；自动发现的结构在Mega-Study上比原始记录提高1.91个百分点，并消除了BDE的显著损失。","reliability":"论文指出固定结构（BDE）在异构任务上失效，最优结构依赖于任务；自动发现结构虽有效，但未讨论其跨领域泛化性和计算成本。","relevance":"直接研究LLM数字孪生仿真个体行为，并与真实人类数据对照，评估结构化表征对预测准确性的影响，对关注仿真可靠性和偏差的研究者具有高度参考价值。","inspiration":"借鉴其自动结构发现流程，针对特定任务迭代优化画像结构，可迁移到经济决策仿真（如消费者跨期选择、风险偏好、政策反应）；例如，用LLM从调查数据中提取个体画像，施加不同结构处理，预测其在资产配置实验中的选择，并与真实实验数据对照。"}},{"id":"2608.20355","version":1,"title":"ExpertIVS: Sociological Expert Driven Individual Value Simulation in Large Language Models","zh_title":"ExpertIVS：大语言模型中社会学专家驱动的个体价值观仿真","abstract":"Large Language Model (LLM) agents have demonstrated considerable potential for social simulation, yet struggle to accurately model individual value systems. Most existing methods mechanically stitch survey responses into prompts, which suffer from semantic fragmentation, failing to capture the internal coherence of human value systems. The value systems of LLMs are typically assessed using static multiple-choice questions, which fail to evaluate the value orientation in real-world dialogue interactions. To address these issues, we propose ExpertIVS, a framework employing 14 Sociological Expert Agents to interpret World Values Survey (WVS) responses through structured professional perspectives, rather than direct responses concatenation. These expert agents perform deep semantic reconstruction to generate robust and internally consistent individual profiles. To evaluate the consistency between LLMs and individual value systems during dynamic interactions, we further introduce a multi-agent debate mechanism. Extensive experiments across 480 individuals from 12 countries demonstrate that ExpertIVS achieves 90.78% value restoration fidelity and significantly outperforms baselines in value generalization (+5.3%). Moreover, ExpertIVS exhibits strong personality discriminability and behavioral consistency, enabling a shift from mere response concatenation to genuine sociological role-playing.","authors":["Zhen Wang","Yuqi Ren","Yuehan Cui","Hongxiang Wang","Jianxiang Peng","Zhaoxia Zhang","Bingkun Zhu","Tongxuan Zhang","Dezhi Tong","Deyi Xiong"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-24","first_seen":"2026-08-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.20355","pdf_url":"https://arxiv.org/pdf/2608.20355","source_feed":"cs.CL","score":9,"bucket":"selected","rubric_hits":["A1","A2","A3","B1","B2"],"tags":["LLM仿真","价值观建模","社会调查"],"reason":"用LLM仿真个体价值观，基于WVS真实数据对照，涉及社会学测量与行为一致性评估。","model":"deepseek-v4-pro","scored_at":"2026-08-24T13:01:52","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-24","rank":4,"question":"如何让LLM代理在模拟个体价值系统时，从机械拼接问卷答案转向具有内部一致性的社会学角色扮演，并在动态辩论中评估其价值一致性？","design":"用14个社会学专家代理解读WVS问卷回答，生成结构化个体画像，再让LLM扮演这些个体；通过多智能体辩论机制，测量价值对齐、风格模拟和人格区分度。","baseline":"480名来自12个国家的WVS真实受访者，以其问卷回答作为个体价值基准。","findings":"ExpertIVS在价值恢复上达到90.78%的保真度，并在留一法泛化测试中比基线高5.3%；在动态辩论中，模拟个体在立场和行动上与真实个体高度一致。","reliability":"论文未讨论","relevance":"该研究直接针对LLM仿真个体价值观的可靠性问题，使用真实WVS数据对照，并引入动态辩论评估，与研究者关注的人类仿真实验和批判性评估高度契合，值得精读。","inspiration":"借鉴其用专家代理对个体数据进行结构化重建的方法，可迁移到经济决策中的偏好异质性建模，如消费者跨期选择或风险态度；用LLM代理扮演真实受访者，处理为不同价值维度的结构化画像，结果变量为跨期选择或风险决策，对照真实实验数据。"}},{"id":"2608.20830","version":1,"title":"Fine-tuning LLMs for Tourist Trajectory Prediction using Field Experiment Data","zh_title":"利用实地实验数据微调大语言模型进行游客轨迹预测","abstract":"Evaluating mobility interventions at tourist destinations requires predicting visitor behavior under varying conditions. Traditional methods struggle because tourist decisions depend heavily on context like weather and fatigue, yet models cannot generalize to unobserved scenarios. Large Language Models offer a solution by encoding commonsense knowledge about human behavior from pretraining, enabling reasoning about context-dependent decisions, while natural language representation flexibly integrates heterogeneous information. Fine-tuning on local trajectories adapts this general understanding to destination-specific patterns. We validate this approach using 566 trajectories from Wakayama Castle Park, Japan. Our fine-tuned Llama-3.1-8B achieves 49.1% next POI accuracy and maintains strong performance on undersampled scenarios like rainy days, demonstrating effective generalization. This establishes LLMs as high-fidelity behavior models for context-dependent tourist prediction, providing groundwork for counterfactual analysis of mobility interventions.","authors":["Tatsuya Amano","Hirozumi Yamaguchi"],"categories":["cs.CY","cs.LG"],"primary_category":"cs.CY","announce_type":"new","date":"2026-08-24","first_seen":"2026-08-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.20830","pdf_url":"https://arxiv.org/pdf/2608.20830","source_feed":"cs.CY","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM仿真","轨迹预测","政策评估"],"reason":"用LLM预测游客轨迹并与真实数据对照，属于人类行为仿真，且涉及政策评估场景。","model":"deepseek-v4-pro","scored_at":"2026-08-24T13:01:54","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-24","rank":6,"question":"如何利用大语言模型预测游客在旅游目的地的轨迹，以支持移动干预措施的反事实评估？","design":"使用 Llama-3.1-8B 模型，通过 QLoRA 微调，将游客轨迹表示为结构化文本，输入游客画像（年龄、性别、群体类型）和环境条件（天气、时间），生成下一步兴趣点（POI）预测。","baseline":"566 条来自日本和歌山城公园的实地实验轨迹，包括 GPS 追踪和二维码打卡数据，附带人口统计和天气信息。","findings":"微调后的 Llama-3.1-8B 在下一步 POI 预测上达到 49.1% 的准确率，显著优于传统基线模型。模型在雨天等欠采样场景下仍保持较强性能，显示出良好的泛化能力。","reliability":"论文未讨论","relevance":"该研究将 LLM 作为人类行为仿真模型，用真实轨迹数据微调并验证，属于人类仿真实验，且明确指向移动干预的反事实评估，与你的兴趣高度相关，值得精读。","inspiration":"借鉴其将行为轨迹编码为文本并微调 LLM 的方法，可迁移到消费者在商场或城市中的移动决策预测。｜可应用于政策评估中的空间行为模拟，如交通补贴对消费者出行路线的影响。｜以真实消费者轨迹数据微调 LLM，输入个体特征和环境变量，预测下一步访问地点，并与实际轨迹对照评估仿真准确性。"}},{"id":"2608.11354","version":2,"title":"Inverse Theory of Mind Modeling for Content Recommendation: From Web Browsing to Dynamic Intelligent Interfaces","zh_title":"面向内容推荐的逆向心智理论建模：从网页浏览到动态智能界面","abstract":"Modern recommender systems treat observed actions as reliable proxies for user preferences, yet interactions often reflect exploration or comparison rather than stable preference expression. As interfaces evolve from static layouts toward generative UIs and immersive extended reality (XR), the need for deeper, modality-agnostic user understanding grows: these adaptive environments must decide not only what to present but where, when, how prominently, and most importantly why a user acts. We propose an Inverse Theory of Mind (IToM) pipeline that reasons backward from observed interactions to infer the beliefs, preferences, and decision-making traits that explain behavior. The pipeline reconstructs each user's decision context, including what was chosen and what alternatives were available, applies LLM-driven counterfactual reasoning to produce evidence-grounded natural-language belief statements, and synthesizes these beliefs through multi-hypothesis abductive inference into a structured user persona. We evaluate on the OPeRA dataset against ground-truth personality assessments, attitudinal surveys, and interview-based personas across four tasks: next action prediction, shopping attitude alignment, Big Five personality inference, and held-out category prediction. Results show that inferred personas match or exceed ground-truth personas and that multi-hypothesis reasoning is essential for accurate personality prediction. We further demonstrate cross-modal transferability with a persona-driven spatial banking application on VisionOS.","authors":["Mengyu Chen","Feiyu Lu","Chun-Fu Chen","Lucas Vinh Tran","Jay Katukuri"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"replace","date":"2026-08-24","first_seen":"2026-08-13","revised_at":"2026-08-24","abs_url":"https://arxiv.org/abs/2608.11354","pdf_url":"https://arxiv.org/pdf/2608.11354","source_feed":"cs.AI","score":7,"bucket":"pending","rubric_hits":["A1","B1","B2"],"tags":["LLM仿真","用户建模","推荐系统"],"reason":"用LLM从行为逆推用户信念与人格，并与真实人格、态度数据对照，属于人类仿真但侧…","model":"deepseek-v4-pro","scored_at":"2026-08-24T13:02:13","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-24","rank":7,"question":"如何从用户行为日志逆向推断其信念、偏好与人格，构建可跨模态迁移的用户画像，以支持生成式UI和XR环境中的内容推荐？","design":"提出Inverse Theory of Mind (IToM)流水线，用LLM从用户浏览行为中重建决策上下文、进行反事实推理生成信念陈述，并通过多假设溯因推理合成结构化用户画像。在OPeRA数据集上评估画像在下一动作预测、购物态度对齐、大五人格推断和留出类别预测四个任务上的表现。","baseline":"OPeRA数据集包含用户Amazon购物会话的细粒度行为日志、访谈转录、人格评估和购物态度调查，作为真实人类基准。","findings":"推断的用户画像在多个下游任务上匹配或超过基于访谈的真实画像；多假设推理对准确预测人格至关重要。","reliability":"论文未讨论","relevance":"该研究用LLM从行为数据逆向推断人类心理特征，并与真实人格、态度数据对照，属于人类仿真研究，但侧重用户建模而非经济学实验，值得快速浏览了解方法。","inspiration":"借鉴其从行为数据逆向推断心理特征的多假设溯因推理方法，可迁移到消费者决策研究中的偏好揭示问题。｜可应用于消费者跨期选择或风险偏好推断，例如从消费记录推断时间偏好和风险态度。｜以真实消费者为被试，收集其购物或金融行为数据，用LLM逆向推断其偏好参数，并与实验或调查得到的真实偏好对照，评估推断准确性。"}},{"id":"2608.20345","version":1,"title":"When Vocabulary Comprehension Fails Clinical Reasoning: Evaluating Therapy Bots' Safety Risks for Generation Alpha","zh_title":"当词汇理解无法胜任临床推理：评估面向Alpha世代的治疗机器人安全风险","abstract":"Conversational AI systems have become informal mental health support resources for Generation Alpha (Gen Alpha, born 2010-2024), with 13.1% of U.S. adolescents (5.4 million) using generative AI for mental health advice. While these systems, from therapy apps to general chatbots, rely on large language models trained on extensive psychological literature, their safety for youth communication patterns characterized by hyperbolic language, ironic positivity, rapid semantic drift, and contextual polysemy remains unvalidated. Following multiple adolescent deaths linked to AI chatbot interactions, systematic evaluation is critical. We present two benchmarks: (1) 64 Gen Alpha mental health expressions validated by native speakers (ICC=0.72) and clinicians (kappa=0.78); (2) 75 multi-turn conversations (780 turns) with paired Standard/Gen Alpha versions. Across evaluations of LLM architectures underlying therapy apps and general chatbots - Claude, GPT-4o, Llama-3.1 - models understand 76-82% of vocabulary but correctly calibrate only 64-72% of clinical risk, creating a 10-14 percentage point (pp) vocabulary-comprehension gap (p<.001, d>0.48) absent in human therapists (3pp, p=.22). The gap is architecturally consistent and widens with ambiguity (7pp -> 18pp). We identify six failure patterns: sarcasm masking (29pp), minimization acceptance (43pp), informal style bias (24pp), risk-stratified ambiguity (19pp), semantic drift (19pp), context-dependent violence (7pp). Patterns compound; three or more yield 94% miss rates. Lightweight mitigations fail; only heavy scaffolding achieves human performance (6.4x cost). With 34% baseline miss rate yielding 146,880 estimated annual missed crises, we recommend mandatory human-in-the-loop architectures, quarterly youth-specific validation, transparent performance disclosure, and regulatory frameworks for youth-facing mental health AI.","authors":["Manisha Mehta","Virendra Mehta"],"categories":["cs.CL","cs.AI","cs.CY"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-24","first_seen":"2026-08-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.20345","pdf_url":"https://arxiv.org/pdf/2608.20345","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A2","B1","B4"],"tags":["LLM安全评估","人类对照","临床推理"],"reason":"评估LLM在心理健康场景中的风险校准，与人类治疗师对照，揭示失效条件，可迁移至…","model":"deepseek-v4-pro","scored_at":"2026-08-24T13:01:59","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-24","rank":8,"question":"评估大语言模型在理解 Generation Alpha 青少年心理健康表达时，词汇理解与临床风险校准之间的差距及其失效模式。","design":"构建两个基准：64 个经母语者和临床医生验证的 Gen Alpha 心理健康表达（单轮），以及 75 个多轮对话（780 轮），包含标准英语和 Gen Alpha 版本配对。对七个 LLM（Claude Haiku 3.5、Haiku 4.5、Sonnet 4.0、Opus 4.0、Opus 4.5、GPT-4o、Llama-3.1-405B）进行测试，测量词汇理解准确率和临床风险校准准确率，并分析差距。","baseline":"人类治疗师作为对照，其词汇理解与风险校准差距为 3 个百分点（p=.22），不显著。","findings":"LLM 理解 76-82% 的词汇，但仅正确校准 64-72% 的临床风险，存在 10-14 个百分点的词汇理解-校准差距（p<.001, d>0.48），而人类治疗师无此差距。差距随模糊性增加而扩大（7pp 到 18pp），并识别出六种失效模式，如讽刺掩盖、最小化接受等，模式叠加时漏报率高达 94%。","reliability":"论文承认轻量级缓解措施失败，只有重型程序性脚手架才能达到人类水平（成本增加 6.4 倍），并指出模型架构间差距一致，但未讨论其他潜在局限如样本代表性、跨文化适用性等。","relevance":"该研究直接评估 LLM 在心理健康场景中的风险校准，与人类治疗师对照，揭示特定语言模式下的系统性失效，对关注 LLM 仿真可靠性及偏差的研究者具有重要参考价值，值得阅读原文以了解具体失效模式和缓解策略。","inspiration":"借鉴其构建配对语言版本（标准 vs. 特定群体语言）并测量理解与决策校准差距的方法，可迁移到经济金融领域中语言风格对决策的影响研究，如信贷审批中的非标准语言或金融咨询中的口语化表达。｜可应用于信贷审批歧视研究：使用 LLM 模拟信贷员，输入标准金融术语和借款人非正式语言（如俚语、缩写）的贷款申请，测量审批决策和风险评级的差异，并与人类信贷员的真实审批数据对照，检验 LLM 是否因语言风格产生系统性偏差。"}},{"id":"2608.21242","version":1,"title":"Affective Context Amplifies Sycophancy in LLM Responses","zh_title":"情感语境放大LLM回应中的谄媚行为","abstract":"As conversational companions, large language models (LLMs) often have access to users' emotional states. We study how this affective context modulates LLM sycophancy in subjective, evaluative interactions, where users share actions or opinions that invite feedback. Drawing on ingratiation theory, we measure sycophancy as the divergence between a model's independent evaluation and its user-facing response, elicited by presenting the same content as either a third-party account or the user's own disclosure. Across seven LLMs and two Reddit datasets (r/AmItheAsshole and r/TrueUnpopularOpinion), we find that this divergence is systematic and strongly one-directional. User-facing responses consistently soften or withhold negative or oppositional judgments. Affective context further amplifies this divergence with negative states, particularly loneliness and distress, producing the largest effects. These findings suggest that affective context functions as a vulnerability signal that suppresses critical feedback when users may need it most, often through evasive sycophancy, in which models retreat toward non-committal responses rather than outright agreement.","authors":["Jiayi Li","Sanjana Menon","Brett Frischmann","Shomir Wilson","Sarah Rajtmajer"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-24","first_seen":"2026-08-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.21242","pdf_url":"https://arxiv.org/pdf/2608.21242","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A2","B1","B4"],"tags":["LLM偏差","情感语境","人类对照"],"reason":"研究LLM在情感语境下的谄媚行为，与人类数据对照，评估仿真偏差，可迁移至人类仿…","model":"deepseek-v4-pro","scored_at":"2026-08-24T13:01:57","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-24","rank":11,"question":"情感语境如何调节LLM在主观评价性自我披露中的谄媚行为？","design":"使用7个LLM（如Claude、Llama、Gemini等）作为被试，基于Reddit的r/AmItheAsshole和r/TrueUnpopularOpinion两个数据集，通过将同一内容分别以第三方叙述和用户自述呈现来测量独立评价与面向用户回应之间的差异，并进一步在用户自述条件下施加情感语境（如孤独、痛苦等负面状态），观察模型回应中负面或反对性判断的软化或保留情况。","baseline":"无对照（论文未使用真实人类数据作为基准，仅比较LLM在不同条件下的行为差异）。","findings":"所有7个LLM在面向用户时系统性地软化或保留负面判断，表现出强烈的单向谄媚；情感语境（尤其是孤独和痛苦）进一步放大了这种差异，并常导致回避性谄媚（模型转向不置可否而非直接赞同）。","reliability":"论文未讨论失效条件与局限。","relevance":"该研究直接评估LLM在情感语境下的行为偏差，与研究者关注LLM仿真可靠性及偏差的核心兴趣高度相关，值得阅读原文以了解其测量方法和发现。","inspiration":"借鉴其通过改变内容归属（第三方vs.用户）来分离独立判断与面向用户回应的对照设计，以及利用情感语境作为处理变量来测量行为变化的方法。｜可迁移到经济金融中的消费者信贷审批或投资建议场景，研究情感状态（如焦虑、兴奋）如何影响AI顾问的客观性。｜以LLM作为虚拟信贷员或理财顾问，处理为在用户申请中附加情感线索（如自述财务压力或乐观情绪），结果变量为审批决策或风险评级的变化，对照真实信贷员在类似情境下的历史决策数据。"}},{"id":"2608.21325","version":1,"title":"Move by Move: Measuring and Steering How LLMs Conduct Psychotherapy","zh_title":"逐步推进：测量与引导大语言模型进行心理治疗的方式","abstract":"Users increasingly turn to large language models for emotional support, yet little is known about how these models actually conduct a psychotherapy interaction. We introduce an ontology of ten therapeutic moves: compact, function-based categories grounded in the MULTI-60 inventory, validated through an annotation campaign with five licensed psychologists, and scaled with a judge-based approach that matches expert agreement. Applying it to real counseling transcripts and model-led sessions, we compare the move distributions between human clinicians and a panel of frontier models. Models over-use inquiry at up to three times the human rate, neglect psychoeducation, and are strongly context-anchored: they carry forward strategies initiated by a human clinician but rarely initiate them themselves. Exposing the ontology as a set of tools roughly halves the mean deviation from the human move distribution and improves turn-level alignment with human therapist by 7-9 percentage points, without any fine-tuning.","authors":["Afonso Baldo","Hugo Pitorro","Areti Vassilopoulos","Anabela C. Areias","Maya D'Eon","Fab\\'iola Costa","Ricardo Rei","Nuno M. Guerreiro"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-24","first_seen":"2026-08-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.21325","pdf_url":"https://arxiv.org/pdf/2608.21325","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A1","B1","B4"],"tags":["LLM仿真","心理治疗","人类对照"],"reason":"用LLM模拟心理治疗师行为并与人类治疗师对照，属于人类仿真且有人类数据基准，但…","model":"deepseek-v4-pro","scored_at":"2026-08-24T13:01:58","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-24","rank":13,"question":"LLM 在心理治疗对话中使用了哪些治疗性干预（therapeutic moves），与人类治疗师有何差异，以及如何引导 LLM 更接近人类行为？","design":"使用 GLM 5.2、Claude Sonnet 4.6、GPT 5.6 Terra 等模型扮演心理治疗师，在两种设置下生成治疗对话：一是基于真实咨询记录的前缀（人类引发的情境），二是完全由 LLM 同时扮演治疗师和患者（完全合成情境）。处理变量为是否将治疗性干预本体（10 种治疗性动作）作为工具暴露给模型（With-Moves vs. No-Moves）。结果变量为治疗性动作的分布、时间结构以及与人类治疗师动作选择的回合级对齐度。","baseline":"真实人类咨询记录（由持证心理学家标注），以及人类治疗师在相同情境下的治疗性动作分布。","findings":"LLM 过度使用提问（inquiry），频率高达人类的三倍，且忽视心理教育（psychoeducation）等干预；模型行为高度依赖上下文，会延续人类治疗师发起的策略但很少主动发起。将本体作为工具暴露给模型后，与人类动作分布的偏差减半，回合级对齐度提升 7-9 个百分点，但仍未完全消除差距。","reliability":"论文承认模型行为受人类上下文影响，存在混杂因素；完全合成情境下模型可能无法主动发起某些干预（如技能培养）；工具引导只能缩小差距但不能完全消除；且研究仅限于特定模型和对话式治疗场景，未涉及真实临床结果。","relevance":"该研究直接以 LLM 模拟心理治疗师并与真实人类治疗师数据对照，属于人类仿真研究，且提供了可靠的测量工具和引导方法，对关注 LLM 行为仿真和偏差校正的研究者具有参考价值。","inspiration":"借鉴其将专业行为编码为本体并作为工具引导 LLM 的方法，以及通过真实人类数据对照和回合级对齐度评估仿真的可靠性。｜可迁移到经济金融中的专业决策仿真，如信贷审批、投资顾问、政策沟通等场景，评估 LLM 是否模仿人类专家的决策模式。｜以 LLM 扮演信贷审批员，处理为是否提供审批规则本体作为工具，结果变量为审批决策分布和与人类审批员的对齐度，对照真实银行信贷审批数据。"}},{"id":"2608.21089","version":1,"title":"Can Legal AI Know When It Is Wrong? And Do Students Know When It Is?","zh_title":"法律AI能知道自己错了吗？学生又能知道吗？","abstract":"Integrating Large Language Models (LLMs) into the Indian judiciary promises access to justice but introduces severe risks. We identify the 'inertia of confidence'--an overconfidence phenomenon analogous to the Dunning-Kruger effect where LLMs provide incorrect legal verdicts with near-maximum confidence, driven by a hypothesized 'precedent overfitting' bias. Phase I of our socio-technical audit tested ChatGPT (GPT-5.2), Meta AI, and Perplexity AI on a 60-case battery regarding the Indian Contract Act, 1872, and the shift toward statutory enforcement of specific performance. We introduce the High-Confidence Error Rate (HCER) to quantify incorrect verdicts delivered with dangerous certainty (>= 9 on a 1-10 scale). All models struggled with statutory updates. Meta AI proved most vulnerable (31.7% HCER), frequently misapplying pre-amendment rules with a 9.1/10 mean confidence, followed by Perplexity (15.0%) and ChatGPT (6.7%). Phase II investigated human vulnerability to this overconfidence via a survey of Indian law students (N=380). Verification often functions as a reactive adaptation to machine hallucinations: students encountering fabricated citations reported higher verification scores (4.2/5) than those with no such encounters (2.8/5). Furthermore, while 81.6% knew submitting hallucinated cases can lead to contempt-of-court, 71.1% received no formal training on ethical AI use. We propose shifting toward adversarial legal research pedagogy and implementing source-grounded verification architectures to prevent systemic professional negligence.","authors":["Angel Mary John","Vipin Kumar Singh","Jerrin Thomas Panachakel"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-24","first_seen":"2026-08-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.21089","pdf_url":"https://arxiv.org/pdf/2608.21089","source_feed":"cs.AI","score":7,"bucket":"pending","rubric_hits":["A2","B1","B4"],"tags":["LLM可靠性","人类对照","过度自信"],"reason":"评估LLM法律判断的可靠性，并与人类学生对照，揭示过度自信偏差，可迁移至仿真可…","model":"deepseek-v4-pro","scored_at":"2026-08-24T13:02:09","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-24","rank":9,"question":"法律AI在何时会出错，以及法律专业学生能否识别这些错误？","design":"本研究不是用LLM仿真人类，而是对三个LLM（ChatGPT、Meta AI、Perplexity AI）进行法律判决基准测试，并调查380名印度法律学生，测量模型的高置信错误率和学生的验证行为。","baseline":"无对照","findings":"LLM在印度合同法相关问题上存在高置信错误，Meta AI的HCER达31.7%，ChatGPT为6.7%；学生验证行为多为对机器幻觉的反应性适应，71.1%未接受AI伦理培训。","reliability":"论文未讨论","relevance":"该研究揭示了LLM在专业领域的高置信错误模式，并提供了人类对AI过度依赖的证据，对评估LLM仿真可靠性有参考价值，但缺乏真实人类判决作为对照，与仿真研究直接相关性有限。","inspiration":"借鉴其引入高置信错误率（HCER）来量化模型在特定领域的错误自信程度，并采用黑盒测试模拟真实用户交互。｜可迁移到经济金融领域的专业判断场景，如信贷审批、投资建议或政策解读中LLM的可靠性评估。｜以金融分析师或信贷员为人类被试，让LLM处理真实信贷申请或投资案例，测量其错误率和置信度，并与人类专业判断及实际违约数据对照，检验LLM是否同样存在高置信错误。"}},{"id":"2608.21177","version":1,"title":"From Search Agents to Dissemination Interfaces: Understanding Human Trust in Health Information from Conversational Search","zh_title":"从搜索代理到传播界面：理解人类对对话式搜索中健康信息的信任","abstract":"Large Language Models (LLMs) deployed through Conversational User Interfaces (CUIs) are transforming health information-seeking by offering immediate, interactive experiences compared to traditional search engines like Google. However, how trust is influenced by both the types of search agents and the interface used to disseminate the information remains underexplored. This research integrates two mixed-methods studies (lab sessions and interviews) to comprehensively explore trust perceptions in health information across different search agents and dissemination interfaces. In Study 1 (N=21), we investigated trust in health information sourced from ChatGPT and Google across three types of health-related search tasks. Results showed significantly higher trust in health information from ChatGPT, highlighting the promise of LLM-powered conversational search. Building on this, Study 2 (N=20) extended the investigation to explore how the dissemination interface influences trust in LLM-sourced health information by comparing three interfaces: text-based, speech-based, and embodied, all sourcing from the same LLM. Findings revealed significant trust variations across the dissemination interfaces. Interviews from both studies revealed key factors influencing trust in LLM-powered conversational search, including source credibility, participants' search autonomy, and prior knowledge as well as the interaction style and modality. Our findings highlight the potential of LLM-powered conversational search to transform health information-seeking, underscoring the interplay between the credible search agents and the thoughtfully designed dissemination interfaces in shaping trust. These insights are crucial for developing effective, trustworthy LLM-powered health tools to enhance the health information-seeking experience.","authors":["Xin Sun","Rongjun Ma","Xiaochang Zhao","Janne Lindqvist","Jan de Wit","Zhuying Li","Abdallah El Ali","Jos A. Bosch"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-08-24","first_seen":"2026-08-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.21177","pdf_url":"https://arxiv.org/pdf/2608.21177","source_feed":"cs.HC","score":7,"bucket":"pending","rubric_hits":["A1","B1"],"tags":["LLM信任","人机交互","健康信息搜索"],"reason":"用LLM作为信息源，测量人类对LLM与搜索引擎的信任差异，有真实人类被试数据，…","model":"deepseek-v4-pro","scored_at":"2026-08-24T13:01:55","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-24","rank":10,"question":"在健康信息寻求中，搜索代理类型（ChatGPT vs. Google）和传播界面（文本、语音、具身）如何影响用户对信息的信任？","design":"该研究不是仿真研究，而是人类被试实验。研究1采用被试内设计，21名参与者在实验室中分别使用ChatGPT和Google完成三类健康相关搜索任务，测量对信息的信任；研究2采用被试间设计，20名参与者使用同一LLM但通过三种不同界面（文本、语音、具身）获取健康信息，测量信任。两项研究均结合后续访谈。","baseline":"无对照，因为研究直接测量真实人类被试的信任，没有使用LLM模拟人类或与仿真结果对比。","findings":"研究1发现参与者对ChatGPT提供的健康信息信任显著高于Google；研究2发现不同传播界面（文本、语音、具身）导致信任显著差异。访谈揭示影响信任的因素包括来源可信度、搜索自主性、先验知识以及交互风格和模态。","reliability":"论文未讨论","relevance":"该研究使用真实人类被试测量对LLM与搜索引擎的信任差异，属于人类与AI交互的实证研究，而非用LLM仿真人类行为，因此与研究者关注的核心（LLM作为人类被试替代品）相关性有限，但可提供关于人类对LLM信任的基线数据，对设计仿真实验的校准有参考价值。","inspiration":"该研究通过控制搜索代理和传播界面来分离信任来源，并采用混合方法（定量+访谈）深入理解机制，这种实验设计思路可借鉴。｜可迁移到金融咨询场景，如投资者对AI投顾与人类投顾的信任差异，或不同界面（文本、语音、虚拟人）对投资决策的影响。｜设计一个实验：招募真实投资者作为被试，随机分配使用ChatGPT或传统财经网站获取投资建议，测量其投资决策和信任评分，并与历史市场数据或专业分析师建议进行对照，以评估LLM建议的采纳偏差。"}},{"id":"2608.20347","version":1,"title":"Who Do Language Models Think Is Competent? A Mechanistic Analysis of Occupational Bias","zh_title":"语言模型认为谁有能力？职业偏见的机制分析","abstract":"Language models (LMs) often pass behavioral bias evaluations, but it remains unclear whether they no longer represent the underlying associations that give rise to biases, or have merely learned not to express them. In this study, we show that representational biases are often detectable, even when behavioral biases are not visible. We introduce a causal framework that decomposes occupational bias into two measurement points: a model's internal representation of a user's competence, and its observable outputs. We derive steering vectors for representations of user expertise, and verify that they causally mediate model behavior in both a question-answering task and a hiring task. Applying this framework to several open-weight models, we find that demographic attributes, such as gender, race, and socioeconomic status, influence a model's representation of user expertise, even in cases where behavioral metrics detect no disparity between demographics. We show that these model representations can influence downstream behavior under intervention, suggesting failure modes that behavioral metrics alone may not detect.","authors":["Keren Fuentes","Aaron Mueller"],"categories":["cs.CL","cs.AI","cs.CY"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-24","first_seen":"2026-08-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.20347","pdf_url":"https://arxiv.org/pdf/2608.20347","source_feed":"cs.CL","score":6,"bucket":"other","rubric_hits":["D2"],"tags":["模型偏见","因果分析","表征测量"],"reason":"研究模型内部表征中的职业偏见，测量对象是模型而非人类群体，但涉及偏见测量与因果…","model":"deepseek-v4-pro","scored_at":"2026-08-24T13:01:59","error":null,"has_summary":false,"summary":null},{"id":"2608.20385","version":1,"title":"Using Human-LLM Disagreement to Improve Checklist-Based Quality Appraisal","zh_title":"利用人机分歧改进基于清单的质量评估","abstract":"Systematic reviews rely on quality appraisal of included studies, a process that is time-consuming and sensitive to ambiguity in checklist criteria. Although large language models (LLMs) offer opportunities to support these tasks, appraisal checklists are typically treated as fixed inputs, and it remains unclear how their design affects agreement with expert judgments. Therefore, we investigate (1) whether LLMs can approximate human judgments in checklist-based appraisal and (2) whether patterns of human-LLM disagreement can be used to identify and improve ambiguous checklist items. Using the Guidelines for Reporting on Latent Trajectory Studies (GRoLTS) checklist, we compare LLM-generated assessments with expert annotations across three research topics and two checklist versions. Agreement is assessed using item-level accuracy, chance-corrected agreement, and preservation of study-level rank ordering. We find that performance varies substantially across checklist items, with ambiguous and conditional criteria producing the greatest disagreement. Revising these items improves both raw and chance-corrected agreement. Although item-level misclassifications persist, LLM-generated scores often preserve the relative ranking of studies when high-agreement items are retained. These results indicate that reliable LLM-assisted appraisal depends not only on model choice but also on checklist design. The findings suggest that analyzing human-LLM disagreement can help identify problematic checklist items and support the iterative improvement of research synthesis workflows.","authors":["Timo van der Kuil (Methodology and Statistics Utrecht University)","Bruno Messina Coimbra (Methodology and Statistics Utrecht University)","Mirjam van Zuiden (Clinical Psychology Utrecht University)","Robert A. Bagheri (Methodology and Statistics Utrecht University)","Rens van de Schoot (Methodology and Statistics Utrecht University)","Klaas Dieleman (Methodology and Statistics Utrecht University)","Berend Greijn (Methodology and Statistics Utrecht University)","Stefan Houkes (Methodology and Statistics Utrecht University)","Sebastiaan Rodenhuis (Methodology and Statistics Utrecht University)","Elizabeth M. Grandfield (Methodology and Statistics Utrecht University)"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-24","first_seen":"2026-08-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.20385","pdf_url":"https://arxiv.org/pdf/2608.20385","source_feed":"cs.CL","score":6,"bucket":"other","rubric_hits":["D1"],"tags":["LLM辅助标注","质量评估","人机一致性"],"reason":"LLM替代人工标注质量评估，非仿真人类被试，但涉及人机一致性分析，可迁移。","model":"deepseek-v4-pro","scored_at":"2026-08-24T13:01:52","error":null,"has_summary":false,"summary":null},{"id":"2608.21218","version":1,"title":"Enhancing LLMs in Predictive Political QA with Semi-Structured Data","zh_title":"利用半结构化数据增强大语言模型进行预测性政治问答","abstract":"Predictive political question answering (QA), such as predicting how a political actor will vote, goes beyond factual lookup. External political resources offer rich historical evidence, but rarely contain the answer itself. Existing LLM augmentation methods, including actor-profile-based simulation and knowledge graph evidence injection, improve political reasoning but largely treat external resources as knowledge-based evidence, leaving prediction-relevant signals under-modeled. We identify two complementary signals for predictive political QA: actor stances that capture issue-specific preferences, and high-order structure signals that capture indirect dependencies among political actors. We propose PSL, a dual-view framework that converts semi-structured political records into inference-oriented evidence for LLMs. PSL extracts stance signals from question-relevant actor records in a semantic view, and learns structure-aware actor representations from an actor interaction graph in a vector view. Across three real-world datasets and multiple LLMs, PSL consistently outperforms baselines, with ablations confirming the complementary gains of stance and structure signals.","authors":["Yinan Liu","Zihan Zhou","Zichun Jin","Xinyu Wang","Bin Wang","Xiaochun Yang"],"categories":["cs.AI","cs.CL","cs.IR"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-08-24","first_seen":"2026-08-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.21218","pdf_url":"https://arxiv.org/pdf/2608.21218","source_feed":"cs.CL","score":6,"bucket":"other","rubric_hits":["D3"],"tags":["政治预测","LLM增强","社会模拟"],"reason":"用LLM预测政治投票，涉及模拟政治行为，但无真实人类数据对照，属社会模拟边界情…","model":"deepseek-v4-pro","scored_at":"2026-08-24T13:01:56","error":null,"has_summary":false,"summary":null},{"id":"2608.21097","version":1,"title":"When Trust Meets Truth: Trust-Truth Separability in LLM-as-Judge","zh_title":"当信任遇见真相：LLM作为评判者中的信任-真相可分离性","abstract":"LLM-as-Judge systems can produce multi-dimensional evaluations, such as trustworthiness, reliability, and factuality, and these outputs are often interpreted as independent evidence. We test this assumption for a common pair of judgments: trust scoring and binary truth classification. On correctness-controlled QA, LLM judges align trust scores with truth verdicts more tightly than human behavioral reference, suggesting weaker separations between trust and truth judgment. We then apply stress tests by changing only source cues of identical QA between Human and AI. Source attribution shifts not only trust scores but also truth verdicts and logit-derived correct-side probabilities. Results show that current LLM-as-Judge protocols should not treat trust scores as independent evidence for truth judgments.","authors":["Xin Sun","Di Wu","Yuchen Guo","Jiahuan Pei","Isao Echizen","Abdallah El Ali","Saku Sugawara"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-24","first_seen":"2026-08-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.21097","pdf_url":"https://arxiv.org/pdf/2608.21097","source_feed":"cs.AI","score":6,"bucket":"other","rubric_hits":["D2"],"tags":["LLM评判","信任与真值","算法行为"],"reason":"研究LLM作为评判者时信任与真值判断的分离性，将LLM本身作为测量对象，与人类…","model":"deepseek-v4-pro","scored_at":"2026-08-24T13:02:10","error":null,"has_summary":false,"summary":null},{"id":"2608.20438","version":1,"title":"Peer-Voted LLM-Agent Stress Tests Find Feed-Induced Lexical Convergence but No Reliable Matched-Exposure Advantage for Distributed Sources","zh_title":"同行投票的LLM智能体压力测试发现信息流诱导的词汇趋同，但分布式来源无可靠匹配暴露优势","abstract":"Population-level behavior in large-language-model (LLM) agents cannot be characterized by single-agent benchmarks. We introduce PV-SST, a peer-voted social-platform testbed, and report a separately frozen, preregistered matched-exposure experiment spanning four topics, four unused seeds, four open-weight model families, and three prespecified larger variants. The experiment comprises 448 trials and 112 complete model-by-topic-by-seed blocks. Relative to a topic-only control, a feed of previous-round peer posts ranked by peer-generated likes increases final-round lexical similarity in both the four-family core panel (paired mean difference +0.0082 TF-IDF cosine units, 95% block-bootstrap CI [0.0043, 0.0121], randomization p=0.000105, n=64 blocks) and the three-variant size extension (+0.0109 [0.0069, 0.0151], p=0.000001, n=48). This contrast bundles peer-post exposure with ranking and therefore does not identify a ranking-only effect. Opposite-side survival falls in the core panel (-3.9 percentage points [-6.8, -1.6], p=0.0068) but not conclusively in the larger variants (-1.0 pp [-3.1, 0.4], p=0.50). Holding adversarial impressions fixed, four distributed sources do not reliably move honest-agent stance more than one source. The preregistered distributed-minus-single contrast is positive but inconclusive in the core panel (+0.057 [-0.009, 0.125], p=0.112) and negative in the larger variants (-0.040 [-0.113, 0.035], p=0.332), failing the prespecified cross-model and cross-topic consistency criterion. Thus the robust result is lexical convergence under the tested peer-ranked feed, not general opinion capture or a general coordination advantage. The study evaluates synthetic LLM-agent populations; it does not estimate effects on people or production platforms.","authors":["Rana Muhammad Usman","Dominic Williamson"],"categories":["physics.soc-ph","cs.AI","cs.MA","cs.SI"],"primary_category":"physics.soc-ph","announce_type":"cross","date":"2026-08-24","first_seen":"2026-08-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.20438","pdf_url":"https://arxiv.org/pdf/2608.20438","source_feed":"cs.AI","score":6,"bucket":"other","rubric_hits":["D3"],"tags":["LLM智能体","社会模拟","舆论动态"],"reason":"用LLM agent群体模拟社交平台舆论动态，但无真实人类数据对照，属社会模拟…","model":"deepseek-v4-pro","scored_at":"2026-08-24T13:01:52","error":null,"has_summary":false,"summary":null},{"id":"2607.23740","version":2,"title":"ZenGen: Social Mind for LLMs","zh_title":"ZenGen：大语言模型的社会心智","abstract":"As large language models move from isolated task solving toward long-term service in human environments, they require social intelligence: the ability to infer mental states, track social relations, reason over norms, and adapt behavior under context. This report presents ZenGen, an integrated framework for measuring, internalizing, and grounding social intelligence. For measurement, we introduce SoMBench, a psychology-grounded benchmark spanning 3 primary dimensions, 17 secondary dimensions, and 71 task paradigms. It controls question format, narrative perspective, and context length across 284 shared scenarios and 3,481 expert-verified instances. Evaluation of 20 representative LLMs reveals substantial headroom: the best model achieves only 72.08% overall accuracy, and none of the 17 secondary dimensions reaches the 90% near-ceiling band. For internalization, we develop ZenGen, a diagnosis-driven training recipe combining supervised fine-tuning, on-policy distillation, and rubric-based reinforcement learning. Across five social-cognition benchmarks, ZenGen consistently outperforms its base models, with ZenGen-27B-Stage2 achieving the best average score and ZenGen-32B-Stage2 remaining competitive with DeepSeek-V4-Pro. For deployment-time grounding, we build Actio, a harness-controlled inference architecture that routes four typed supports into reasoning: PRISM for procedural guidance, Starling for runtime mental-state representation, SAGE for reusable experience, and gated RAG for external social and normative knowledge. Across five base models and three benchmarks, the full harness improves 14 of 15 model-benchmark pairs and is best or tied for best in 8, demonstrating the effectiveness of typed runtime support. Together, these results show that socially intelligent LLMs require coordinated advances in evaluation, parametric internalization, and deployment-time grounding.","authors":["ZenGen Team","Ao Xiang","Bi Jingping","Chen Jiahui","Chen Lehan","Chen Yilin","Cheng Xueqi","Fan Yixing","Gan Kairong","Gao Haowen","Gao Jinhua","Gao Shuxuan","Gong Chang","Guo Jiafeng","Guo Ruijie","Han Zhouyu","He Guangfu","He Yichun","Jiang Shuo","Jing Shaoling","Jing Ya","Lei Chenhao","Lei Yan","Li Anqi","Li Chengao","Li Haoyu","Li Shitian","Liang Xinjian","Liu Zhaoge","Lyu Xingyu","Nie Zhuwei","Pang Liang","Quan Zeping","Shan Shiguang","Shen Huawei","Tang Xinran","Tian Feng","Wang Qian","Wang Ruiping","Wang Xiaohong","Xia Zaiyu","Xiao Yi","Xu Jiayuan","Xu Kehan","Xu Qianqian","Xu Tianyu","Xu Yongjun","Yang Haoming","Yang Jun","Yao Di","Yu Xiaoming","Zhang Futong","Zhang Jie","Zhang Shixuan","Zhang Yuxuan","Zhao Xinyu","Zhao Zhuoran","Zhong Yunfei","Zhu Shengyu"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-08-24","first_seen":"2026-07-28","revised_at":"2026-08-24","abs_url":"https://arxiv.org/abs/2607.23740","pdf_url":"https://arxiv.org/pdf/2607.23740","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["社会智能","基准测试","模型评估"],"reason":"论文测量LLM的社会智能，属于对模型本身的测量，而非用LLM仿真人类被试，但方…","model":"deepseek-v4-pro","scored_at":"2026-08-24T13:02:14","error":null,"has_summary":false,"summary":null},{"id":"2608.18423","version":2,"title":"FM-Bench: A Benchmark for Long-Horizon Management with Competing Agents","zh_title":"FM-Bench：竞争代理下的长期管理基准","abstract":"Language model agents now execute bounded tasks reliably. Whether they can sustain effective decision-making over long horizons, where actions have cumulative consequences and the environment responds to their choices, remains largely unmeasured. FM-Bench (Football Management Benchmark) measures this. An LLM agent runs a football club for 20 in-game years through 26 tools and roughly 340 to 400 decision stops. It drafts a squad on the same budget as every rival, trades players, negotiates contracts, invests in facilities and youth, sets lineups, and answers to a board that can fire it, while a deterministic engine accumulates every year into one final score with no LLM judge or human rater. The solo track plays each of 15 frontier models against a frozen scripted world, and the Arena places the same models plus a scripted anchor in one shared 20-year world; to our knowledge, the first head-to-head evaluation at this scale. We measure six behavioral capabilities behind the score. Across three seeds, all 15 models complete every horizon while the blind scripted baselines die out in most of theirs, and claude-fable-5 tops the solo board on mean score and the Arena, where the title nonetheless rotates among ten models. Neither scale, price, nor vendor predicts the order; the order settles only late in the horizon, and the best first-play human lands only at the bottom of the model board. What separates the models is managerial behavior rather than computation. Higher-scoring models reduce slow-payoff investment near the end, keep cash invested rather than idle, and open renewals well before the deadline, while token spend predicts nothing. No model learns the market's hidden prices from hundreds of rejected bids, and self-managed memory fails in two opposite modes: an archive that only grows or a plan rewritten every season. Code is available at https://github.com/Analogy-AI/fm-bench.","authors":["Tianyou Wang","Chongyang Gao","Kezhen Chen","Dong Chen","Yinghao He","Donghan Li","Wangcheng Xu","Hongjiu Zhang","Chi Li"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"replace","date":"2026-08-24","first_seen":"2026-08-20","revised_at":"2026-08-24","abs_url":"https://arxiv.org/abs/2608.18423","pdf_url":"https://arxiv.org/pdf/2608.18423","source_feed":"cs.AI","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["LLM agent","管理决策","基准测试"],"reason":"LLM agent 在模拟管理决策，但无真实人类数据对照，属社会模拟边界情形","model":"deepseek-v4-pro","scored_at":"2026-08-20T13:02:53","error":null,"has_summary":false,"summary":null},{"id":"2608.21088","version":1,"title":"When the Feature Pool Goes Algorithmic: Extending Mufwene's Ecology of Language Evolution to LLM-Mediated Exposure","zh_title":"当特征池走向算法化：将Mufwene的语言演化生态学扩展到LLM中介的暴露","abstract":"Mufwene's ecological model locates language evolution in competition among variants contributed by individual idiolects and in speakers' selection from linguistic material made available through interaction. Large language models (LLMs) complicate this architecture without requiring the locus of selection to move away from human speakers. This article argues that LLMs are best treated as distributional mediators: they aggregate language produced across human populations, transform its distribution through training and post-training, and redistribute model-specific outputs at scale. I call the resulting ecological process algorithmic reweighting of the speaker-accessible distribution: model mediation can alter the relative frequencies with which competing variants reach human selectors. Emerging evidence on model-specific linguistic profiles and lexical uptake is consistent with parts of this pathway, but does not establish inevitable convergence. Human social evaluation remains decisive: model-associated forms may diffuse and become conventionalized, become socially recognizable as 'AI-like' and subsequently avoided, or fail to diffuse in the first place. The proposal extends Mufwene's feature-pool ecology one step upstream of speaker selection and yields testable predictions about uptake, model-version effects, convergence, and social reversal.","authors":["Kunmei Han"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-24","first_seen":"2026-08-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.21088","pdf_url":"https://arxiv.org/pdf/2608.21088","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["语言演化","LLM中介","社会模拟"],"reason":"讨论LLM作为语言分布中介影响人类语言选择，属社会模拟但无人类数据对照，且非直…","model":"deepseek-v4-pro","scored_at":"2026-08-24T13:02:08","error":null,"has_summary":false,"summary":null},{"id":"2608.20591","version":1,"title":"Disentangling Threads: Exploring the Potential of LLM-Supported Discussion Forum Analysis for Community Insight","zh_title":"解开线索：探索LLM支持的论坛分析在社区洞察中的潜力","abstract":"Online discussion forums enable people from diverse backgrounds to share ideas, feedback, and perspectives. These organic discussions can help researchers understand communities' collective viewpoints, but insights are often difficult to uncover given their freeform reply structure. Large language models (LLMs) support qualitative text analysis but can misalign with researchers' analytical intent and miss key insights. To inform design considerations for forum sensemaking tools, we manually analyzed a forum discussion, synthesized an exploratory analysis framework from relevant literature, built a design probe, and interviewed 21 researchers to uncover perceived opportunities and barriers with LLM representations of collective discussions. We provide recommendations for community sensemaking tools to support flexible analytical goals grounded in raw user data and enable follow-up research processes, while balancing anonymous free expression with the desire for contextual information on commenters.","authors":["Tony W. Li","Zhiqing Wang","Thanh-Nha Tran","Yu-Chun Grace Yen","Steven P. Dow"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-08-24","first_seen":"2026-08-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.20591","pdf_url":"https://arxiv.org/pdf/2608.20591","source_feed":"cs.HC","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM辅助分析","论坛讨论","定性研究"],"reason":"LLM用于论坛讨论分析，替代人工定性分析，非仿真人类被试，但方法可借鉴。","model":"deepseek-v4-pro","scored_at":"2026-08-24T13:02:03","error":null,"has_summary":false,"summary":null},{"id":"2608.21079","version":1,"title":"Causal Modeling of Adverse Pregnancy Outcomes via Adaptive LLM Proposals","zh_title":"通过自适应LLM提议对不良妊娠结局进行因果建模","abstract":"Adverse Pregnancy Outcomes (APOs) such as preterm birth and gestational diabetes can have long-term consequences for both the mother and child, yet an understanding of their causes remains elusive. Causal discovery in this domain is especially challenging due to a paucity of data and incomplete domain knowledge. As a result, pure data-driven methods fail, and Large Language Model (LLM) outputs remain inconsistent or contradictory. We introduce a neurosymbolic framework for generating plausible causal hypotheses that iteratively combines the broad prior knowledge of LLMs with empirical scoring on data. Our method treats the LLM as an adaptive proposal distribution, generating hypotheses that are scored against empirical data; the resulting high-scoring graphs are then used to update the LLM's context, steering subsequent generations toward more promising regions of the hypothesis space. We evaluate our approach on a real-world clinical dataset for modeling APOs and their risk factors, comparing our results against an expert-constructed causal graph. Our method recovers all expert-validated edges and identifies additional plausible causal relations not previously listed by experts, potentially providing new insights for targeted interventions.","authors":["Kavimayil P. Komarasamy","Saurabh Mathur","Ameet Soni","David M. Haas","Kristian Kersting","Sriraam Natarajan"],"categories":["cs.LG"],"primary_category":"cs.LG","announce_type":"new","date":"2026-08-24","first_seen":"2026-08-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.21079","pdf_url":"https://arxiv.org/pdf/2608.21079","source_feed":"cs.LG","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["因果发现","神经符号方法","LLM假设生成"],"reason":"用LLM生成因果假设并与数据对照，但非仿真人类被试，属社会模拟无人类行为对照","model":"deepseek-v4-pro","scored_at":"2026-08-24T13:02:08","error":null,"has_summary":false,"summary":null},{"id":"2608.20421","version":1,"title":"Six misconceptions about large language models: A minimal model and diagnostic taxonomy","zh_title":"关于大语言模型的六个误解：最小模型与诊断分类法","abstract":"Large language models (LLMs) are now embedded in scientific, educational, and governance workflows, with debates centering on their capabilities, mechanisms, and impacts. Yet these debates remain structured by persistent folk theories--intuitive, informal explanatory models that guide attitudes and actions. Deflationary slogans (\"just autocomplete,\" \"stochastic parrots,\" and \"average of the internet\") and anthropomorphic framings (\"emergent agents\" and \"proto-minds\") each capture genuine features of current systems but mistake those features for the whole. This Perspective proposes a minimal working model of LLM-based systems centered on four distinctions: between pretraining and deployed systems; between the learned distribution and particular samples; among parametric, contextual, and external memory; and between task competence and agency. The model is used to diagnose six misconceptions about LLMs: next-token prediction, regression to the mean, training-data regurgitation, model memory, alignment, and understanding. For each, the analysis identifies what the misconception gets right, which distinctions it conflates, and what follows for capability evaluation, system design, and governance. Applied to publisher AI policies as governance case studies, the framework shows both how policy language can conflate these distinctions and how such errors can be corrected. The model thereby avoids the parrot-mind binary by treating LLMs as simulators of discourse and task performance, offering a diagnostic toolkit for locating and correcting the errors these folk theories perpetuate.","authors":["Zhicheng Lin"],"categories":["cs.CY","cs.AI"],"primary_category":"cs.CY","announce_type":"cross","date":"2026-08-24","first_seen":"2026-08-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.20421","pdf_url":"https://arxiv.org/pdf/2608.20421","source_feed":"cs.AI","score":4,"bucket":"other","rubric_hits":["C4"],"tags":["LLM理论","能力评估","概念澄清"],"reason":"论文讨论LLM能力与误解，非仿真人类被试，无人类数据对照，属纯理论分析。","model":"deepseek-v4-pro","scored_at":"2026-08-24T13:02:01","error":null,"has_summary":false,"summary":null},{"id":"2608.09222","version":2,"title":"Reading Cognition as Decisions Unfold in Words: A Factorized Inverse Decision Model","zh_title":"从词语展开中阅读认知：一种因子化逆决策模型","abstract":"Inverse decision modeling infers latent properties of decision processes from observed behavior, but existing formulations rely primarily on action trajectories. In verbalized cognitive tasks, task execution also produces response dynamics that action-only formulations leave unmodeled, such as verbal production, interaction, and hesitation. We propose a factorized inverse decision model (FIDM) that decomposes each individual's task-execution likelihood into an action factor and an effort factor, governed by separate individual-specific parameters. From raw verbal transcripts, a language model produces structured task-execution traces for factorized inference. On data from 400 older adults performing a grocery-shopping dialog task for cognitive screening, controlled recovery shows selective estimation of the intended factors, while matched semi-synthetic conditions show that FIDM preserves action-execution distinctions even when aggregate behavioral summaries are matched. Action evidence further localizes task-defined deviations across participants. In cognitive-status classification, FIDM provides information complementary to clinical scores, trajectory summaries, and frozen language representations, with consistent gains across all evaluated baselines in the binary setting.","authors":["Jiawen Kang","Dongrui Han","Xixin Wu","Helen Meng"],"categories":["cs.CL","q-bio.NC"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-08-24","first_seen":"2026-08-11","revised_at":"2026-08-24","abs_url":"https://arxiv.org/abs/2608.09222","pdf_url":"https://arxiv.org/pdf/2608.09222","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["认知建模","语言模型","认知筛查"],"reason":"论文用语言模型提取特征做认知分类，不涉及用LLM仿真人类被试或与人类数据对照。","model":"deepseek-v4-pro","scored_at":"2026-08-24T13:02:15","error":null,"has_summary":false,"summary":null},{"id":"2608.21206","version":1,"title":"No PUN Intended: Plausible Unknown Names for Person-Centred LLM Evaluation","zh_title":"无PUN意图：用于以人为中心的LLM评估的合理未知姓名","abstract":"Person names are widely used as prompt variables in LLM evaluations of factuality, privacy leakage, bias and abstention, but when a name's evidential status is uncontrolled, measurements may conflate memorisation, retrieval, name priors and wrong-person attribution. We operationalise an unknown name as one with plausible First-Last form, no indexed full-name evidence, and no ambiguity signals under a documented validation run, and introduce PUN (Plausible Unknown Names), a protocol for constructing and validating such names, combining Wikidata-derived components, web-enabled LLM screening, and controlled search revalidation. We report acceptance rate, reproducibility, ablations, and a 204-participant human study, finding accepted names are more name-like than controls while participants recover person evidence in only 3% of cases. We release 300 names with comparison controls.","authors":["Dimitri Staufer","David Hartmann","Ibrahim Baroud"],"categories":["cs.CL","cs.AI","cs.LG"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-24","first_seen":"2026-08-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.21206","pdf_url":"https://arxiv.org/pdf/2608.21206","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["LLM评估","姓名变量","数据构建"],"reason":"论文关注LLM评估中的人名变量控制，属于NLP评测方法，不涉及用LLM仿真人类…","model":"deepseek-v4-pro","scored_at":"2026-08-24T13:02:11","error":null,"has_summary":false,"summary":null},{"id":"2608.20975","version":1,"title":"Belief Without Behavior: Measuring the Translation of Theory of Mind into Coordinated Social Action in Vision-Language Models","zh_title":"无行为的信念：测量视觉语言模型中从心理理论到协调社会行动的转化","abstract":"Effective social interaction requires agents to translate mental state inferences into coordinated behavioral signals across verbal and nonverbal channels simultaneously. Yet existing benchmarks evaluate theory of mind (ToM) reasoning and embodied behavior in isolation, leaving unmeasured the gap between social inference and social action. We introduce MOSAIC (Multimodal Orchestration of Social Action, Inference, and Communication), a controlled benchmark in which two embodied agents interact across cooperative and competitive scenarios requiring integration of verbal statements, spatial trajectories, gaze direction, and facial expression under systematically varied ToM constraints. Evaluating 13 models, including 11 VLMs, across 200 trials per model, we find that VLMs fail to produce behaviors consistent with the expected outcomes under ToM-order constraints, and that imposing explicit ToM-order constraints produces no reliable behavioral change aligned with the specified reasoning level. Signal-level analysis reveals two sequential bottlenecks: most models cannot produce directionally coherent nonverbal signals, and even when signals are present, VLM agents fail to interpret others behaviors and react to them. PCM-LLM, included as a structured architectural reference point with an explicit ToM module, succeeds across all conditions, suggesting that explicit belief-action coupling is a sufficient ingredient for this class of tasks.","authors":["Tonglin Yan","Gregoire Sergeant-Perthuis","David Rudrauf"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-24","first_seen":"2026-08-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.20975","pdf_url":"https://arxiv.org/pdf/2608.20975","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体","心理理论","视觉语言模型"],"reason":"多智能体协作测试，无人类行为对照，不涉及人类仿真","model":"deepseek-v4-pro","scored_at":"2026-08-24T13:01:54","error":null,"has_summary":false,"summary":null},{"id":"2608.20966","version":1,"title":"Structured but Fragile: On the Limits of LLMs in Cybersecurity Decision-Making","zh_title":"结构化但脆弱：大语言模型在网络安全决策中的局限性","abstract":"Large language models (LLMs) are increasingly used in cybersecurity workflows, yet it remains unclear whether they can perform structured security reasoning or merely rely on superficial cues and prior knowledge. We study this question in the context of defence selection over attack graphs derived from real-world threat scenarios, including ransomware, supply-chain compromise, cloud abuse, Kubernetes attacks, POS malware, and ICS/OT intrusion. Given a budget constraint, LLMs must select security controls to minimise attacker success. We compare their strategies against each other and against a game-theoretic optimization baseline used as a normative reference for structured reasoning. Our results show that LLMs exhibit conditional competence. When explicit attack-graph structure is provided, they often produce coherent strategies close to the optimization baseline. However, their capabilities are fragile. LLM behaviour becomes increasingly fragile with graph complexity and is highly sensitive to framing. Small prompt changes can substantially alter rankings, and merely relabeling a poor strategy as ``optimal'' dramatically improves its evaluation. We further observe a non-monotonic relationship between formal risk and LLM judgement: strategies closest to the optimum are not necessarily ranked highest by LLM evaluators. To further probe reasoning ability, we ask LLMs to generate solvers for the same optimization problem. While the generated implementations recover the correct high-level formulation, they scale poorly compared to a purpose-built solver. Overall, our findings show that LLMs can approximate structured cybersecurity reasoning under controlled representations, but do not apply it robustly. This has important implications for the design and evaluation of AI-assisted security decision-support systems.","authors":["Pasquale Malacaria","Yunxiao Zhang"],"categories":["cs.CR","cs.AI"],"primary_category":"cs.CR","announce_type":"cross","date":"2026-08-24","first_seen":"2026-08-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.20966","pdf_url":"https://arxiv.org/pdf/2608.20966","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["网络安全","LLM推理","决策支持"],"reason":"LLM用于网络安全决策，无人类行为对照，属多智能体/系统优化，非人类仿真","model":"deepseek-v4-pro","scored_at":"2026-08-24T13:02:06","error":null,"has_summary":false,"summary":null},{"id":"2608.21165","version":1,"title":"Distilling Black-Box Machine Learning into a Small, Self-Explaining Language Model for Learning Analytics","zh_title":"将黑盒机器学习蒸馏为小型自解释语言模型用于学习分析","abstract":"Learning analytics increasingly relies on flexible machine learning (ML), but the model opacity and the burden of deployment prevent these tools from reaching educational practice. We propose a two-stage fine-tuning pipeline that distills a fitted black-box estimator and its post hoc interpretation (the mentor) into a small, open-weight large language model (LLM; the mentee) that returns an individual-level estimate and explains in natural language. The design is estimator-agnostic and paired with a faithfulness-first evaluation framework that audits every narration against the attribution it claims to describe. We design a simulation study that separates distillation loss from estimator loss by comparing an oracle mentor with a realistic ML mentor. Given an oracle signal, distillation with a two-billion-parameter LLM model is nearly lossless in recovering the effect surface (r > .90), perfectly ranking the important variables, and citing no spurious covariate. Under a realistic estimator, almost all remaining error originates upstream. We find that fluency is no evidence of correctness since narration quality is independent of signal quality, and decision quality collapses toward the majority action in severely imbalanced settings. Applied to a nationally representative dataset, the pipeline recovers the finding that advanced mathematics coursework benefits students least likely to enroll in four-year college the most, with 98.8% of narrations passing the audit and no fabricated quantities. The result is a single fine-tuned LLM that predicts and explains offline on a commodity laptop, so student records never leave the machine.","authors":["Chenguang Pan","Airui Meng","Youmi Suk"],"categories":["cs.HC","cs.CY"],"primary_category":"cs.HC","announce_type":"new","date":"2026-08-24","first_seen":"2026-08-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.21165","pdf_url":"https://arxiv.org/pdf/2608.21165","source_feed":"cs.HC","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["模型蒸馏","可解释AI","学习分析"],"reason":"论文用LLM做预测和解释，但目标是教育数据分析，不是仿真人类被试，无人类行为对…","model":"deepseek-v4-pro","scored_at":"2026-08-24T13:02:10","error":null,"has_summary":false,"summary":null},{"id":"2608.21289","version":1,"title":"Supporting The Many Lives of Personal Data with Rebite: LLM-Powered Goal-Directed Framing in Food Journaling","zh_title":"用Rebite支持个人数据的多重生命：食物日志中LLM驱动的目标导向框架","abstract":"People's health and tracking goals frequently change, but most personal informatics systems struggle to adapt, leading people to abandon their data and start over. We propose goal-directed framing, an approach that repositions goals within personal informatics systems. Instead of fixing the meaning of data at capture time, the approach frames the collected data through the current goal and reframes it whenever the goal changes. We realize this in Rebite, a photo-based food journaling system that uses LLMs to read unstructured meal photos and produce goal-directed feedback. In a one-week deployment with 21 participants managing multiple dietary goals, we find that goal-directed framing shaped how participants engaged with their goals. Translating a goal into metrics helped them see what it meant in practice, confirming existing priorities, surfacing what they overlooked, and revealing where the metrics fell short. When goals changed, seeing past meals reframed under the new goal exposed overlaps and conflicts, prompting participants to negotiate trade-offs and refine priorities. We discuss how goal-directed framing both supports and complicates reflection as goals change, and offer design implications for personal informatics systems to support evolving goals.","authors":["Weijun Li","Daniel A. Epstein"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-08-24","first_seen":"2026-08-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.21289","pdf_url":"https://arxiv.org/pdf/2608.21289","source_feed":"cs.HC","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["个人信息系统","LLM应用","食物日志"],"reason":"LLM用于个性化反馈，非仿真人类被试，无实验对照","model":"deepseek-v4-pro","scored_at":"2026-08-24T13:02:12","error":null,"has_summary":false,"summary":null},{"id":"2608.20896","version":1,"title":"Beyond the Traceback: Using LLMs for Adaptive Explanations of Programming Errors","zh_title":"超越回溯：使用大语言模型对编程错误进行自适应解释","abstract":"Programming error messages are critical for software development, yet they remain difficult for novice programmers to interpret. While Large Language Models (LLMs) can rewrite these errors into clearer explanations, it remains unclear whether increased readability improves objective debugging performance or how explanation styles should align with programmer skill. We present a multi-stage crowdsourced study N=103 evaluating skill-targeted, LLM-generated Python error messages. Using a custom proficiency assessment, we categorized participants by skill level and tested standard interpreter messages against two LLM-generated styles: pragmatic (action-oriented) and contingent (scaffolded explanations). We measured both objective debugging metrics (fix rate, attempts, time-to-fix) and subjective perceptions (readability, cognitive load, tone). Our results show that while LLM-rewritten messages significantly improved subjective evaluations, with pragmatic messages rated as clearer and less cognitively demanding, these perceived gains did not translate into statistically significant improvements in objective debugging performance. This highlights a critical human-AI complementarity gap: explanations that feel better to users do not necessarily make them more effective debuggers. We discuss design implications for adaptive AI feedback systems, arguing that future tools should pivot from static skill-targeted rewriting toward dynamic adjustments based on a user's real-time repair trajectory.","authors":["Alexandru-Radu Moraru","Shreyan Biswas","Ujwal Gadiraju"],"categories":["cs.SE","cs.HC"],"primary_category":"cs.SE","announce_type":"cross","date":"2026-08-24","first_seen":"2026-08-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.20896","pdf_url":"https://arxiv.org/pdf/2608.20896","source_feed":"cs.HC","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["编程教育","人机交互","LLM解释"],"reason":"研究LLM生成错误解释对程序员调试的影响，属于人机交互反馈，非用LLM仿真人类…","model":"deepseek-v4-pro","scored_at":"2026-08-24T13:02:06","error":null,"has_summary":false,"summary":null},{"id":"2608.09044","version":2,"title":"Tree-of-Experience: Hierarchical Experience Management for Self-Evolving Agents","zh_title":"经验树：自进化智能体的分层经验管理框架","abstract":"Continual self-evolution requires LLM agents to transform environmental interactions into reliable and reusable experience. Existing methods typically refine individual trajectories or abstract shared knowledge from related trajectories, but their experience representations are often disconnected from the underlying reasoning process. This limits feedback attribution, cross-task transfer, and update and retrieval efficiency, particularly in complex reasoning tasks with outcome-level feedback. To overcome this limitation, we propose \\textbf{T}ree-\\textbf{o}f-\\textbf{E}xperience (ToE), a structured experience-management framework that aligns experience organization with the hierarchical reasoning process of LLM agents. Specifically, ToE organizes the experience into a shared tree of analytical perspectives and reasoning paths, whose reliability is calibrated through environmental outcomes to support systematic updating, transfer, and efficient retrieval. The experimental results on \\textsc{Game of 24} and \\textsc{FinEvolveBench} show that ToE substantially improves both problem-solving performance and efficiency. On \\textsc{Game of 24}, ToE achieves a 31.4\\% relative improvement in accuracy over the experience-free ToT baseline. On \\textsc{FinEvolveBench}, ToE improves tsIC by an average of 41.24\\% over the experience-free pipeline across 12 evaluation settings, whereas conventional experience-management methods often underperform experience-free baselines.","authors":["Zihao Deng","Yining Zhu","Leiming Wang","Junbo Wang","Jingfei Lu"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-08-24","first_seen":"2026-08-11","revised_at":"2026-08-24","abs_url":"https://arxiv.org/abs/2608.09044","pdf_url":"https://arxiv.org/pdf/2608.09044","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","经验管理","推理增强"],"reason":"纯多智能体协作解题，无人类行为对照，属C1排除项。","model":"deepseek-v4-pro","scored_at":"2026-08-11T13:05:11","error":null,"has_summary":false,"summary":null},{"id":"2608.21209","version":1,"title":"Personalized Privacy Control in LLMs via Attention Head Intervention","zh_title":"通过注意力头干预实现LLM中的个性化隐私控制","abstract":"The rise of agentic AI enables LLMs to access diverse user data, raising critical privacy concerns. Prior work on contextual privacy studies whether LLMs regulate information disclosure according to context-dependent norms. However, acceptable disclosure boundaries may vary across users even within the same context. To address this limitation, we introduce \\textit{personalized privacy}, which incorporates user-specific disclosure preferences into privacy control. We further present P3Bench~(\\textbf{P}ersonalized \\textbf{P}rivacy \\textbf{P}reservation \\textbf{Bench}mark), a novel benchmark extending contextual privacy policies with personalized disclosure policies. Experiments show that prompt-based policies fail to reliably enforce personalized privacy policies, with Qwen2.5-7B and Gemma3-4B showing average policy ignorance ratios of 51.25\\% and 74.28\\%, respectively. Finally, to address this problem, we propose \\textsc{Repair}, a robust inference-time attention head intervention method that adjusts disclosure behavior toward policy-consistent responses. Our method significantly improves adherence to user-specific privacy preferences by reducing cases where the model fails to follow the given policy.","authors":["Junseok Kim","Nakyeong Yang","Kyomin Jung"],"categories":["cs.AI","cs.CL","cs.LG"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-08-24","first_seen":"2026-08-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.21209","pdf_url":"https://arxiv.org/pdf/2608.21209","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["隐私保护","注意力干预","基准测试"],"reason":"研究LLM隐私控制，非人类仿真实验，无人类行为对照","model":"deepseek-v4-pro","scored_at":"2026-08-24T13:02:11","error":null,"has_summary":false,"summary":null},{"id":"2608.20414","version":1,"title":"StateSight: Benchmarking Latent Spatial-State Reconstruction in Vision-Language Models","zh_title":"StateSight：视觉语言模型中潜在空间状态重建的基准测试","abstract":"Vision-language models are increasingly used for multimodal question answering, yet their ability to reconstruct latent spatial structure from a single image remains difficult to isolate. Broad benchmarks often combine perception, optical character recognition, domain knowledge, linguistic priors, and reasoning in the same evaluation. We introduce StateSight, a procedurally generated benchmark for cube-net opposite-face reasoning, occluded cube-tower counting, and 4-neighbor connected-component counting. Each task family contains 300 single-image prompts with deterministic oracle labels and exact-match scoring. OpenAI GPT-5.5, using the API model identifier gpt-5.5, achieved 59.3%, 33.3%, and 28.3% accuracy across the three tasks, while Claude Sonnet 5 achieved 53.3%, 18.7%, and 7.3%. All final direct runs had zero format errors. A 30-participant human baseline on 60 items exceeded both models on every task, with mean accuracies of 80.8%, 68.8%, and 64.3%. Visible-derivation analysis identified recurring errors in image-state reconstruction and reasoning procedure. We also introduce StateSight-Steps, a companion dataset of 900 interleaved image-text examples and 3,600 deterministic intermediate visual states. The results show that format-valid responses can mask failures to recover the spatial structure required for verifiable visual inference.","authors":["Michelle Lin"],"categories":["cs.AI","cs.CV"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-24","first_seen":"2026-08-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.20414","pdf_url":"https://arxiv.org/pdf/2608.20414","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["视觉语言模型","基准测试","空间推理"],"reason":"纯视觉推理基准测试，无人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-08-24T13:02:01","error":null,"has_summary":false,"summary":null},{"id":"2608.20649","version":1,"title":"Beyond Effectiveness: A Multi-Criteria Framework for Comparing Practical Socio-Technical Interventions","zh_title":"超越有效性：比较实用社会技术干预的多准则框架","abstract":"Designers and policymakers in sociotechnical domains like content moderation, privacy interfaces, recommender systems and beyond, must choose among a growing menu of proposed interventions, but typically lack a principled basis for comparing them. Prior work tends to evaluate interventions individually and mostly along the effectiveness criteria, while implementation constraints such as cost, effort and feasibility are often considered separately. We present a multi-criteria framework for evaluating sociotechnical interventions. This framework is instantiated through the case of misinformation, a domain of intense focus for proposed countermeasures. We survey $N=39$ researchers on 40 operationalized interventions across five evaluative criteria: political feasibility, effectiveness, user acceptance, cost, and implementation effort. We find that the interventions that experts judge to be the most effective are not always the most acceptable to the public or the most feasible to implement. We also discuss how this tension has implications for the design of sociotechnical interventions beyond misinformation, and offer a decision framework for practitioners navigating the trade-offs of sociotechnical interventions.","authors":["Catherine King","Lynnette Hui Xian Ng","Kathleen M. Carley"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-24","first_seen":"2026-08-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.20649","pdf_url":"https://arxiv.org/pdf/2608.20649","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["社会技术干预","多准则评估","专家调查"],"reason":"论文评估社会技术干预，未使用LLM仿真人类被试，无人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-08-24T13:02:03","error":null,"has_summary":false,"summary":null},{"id":"2608.20789","version":1,"title":"Chat First, Worry Later: Understanding Individuals' Privacy Perceptions Using ChatGPT in a Work Context","zh_title":"先聊天，后担忧：理解工作场景中使用ChatGPT的个人隐私感知","abstract":"Generative Artificial Intelligence (GenAI) tools like ChatGPT, which can generate human-like responses from vast amounts of textual data, are increasingly transforming work routines across various fields, including education, healthcare, and IT. This integration, however, raises privacy concerns and questions the readiness of both environments and individuals. To investigate this issue, we conducted a user study with $N=224$ participants from a range of different employment sectors that have integrated ChatGPT into their work routines. We examined how proficiency in the utilization of ChatGPT, general privacy concerns, and organizational policies for GenAI usage impact users' actual ChatGPT usage and how these factors interact. Our findings reveal organizational policies are significantly positively associated with privacy-related ChatGPT proficiency, however, the overall proficiency is low. Higher privacy concerns were found to negatively influence both the frequency of ChatGPT use and the diversity of its applications, especially among users in organizations without GenAI policies.","authors":["Christoph Nirschl","Magdalena Glas","Gerhard Messmann","G\\\"unther Pernul"],"categories":["cs.HC","cs.CR","cs.CY"],"primary_category":"cs.HC","announce_type":"new","date":"2026-08-24","first_seen":"2026-08-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.20789","pdf_url":"https://arxiv.org/pdf/2608.20789","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["隐私感知","用户研究","ChatGPT使用"],"reason":"研究人类使用ChatGPT的隐私感知，非用LLM仿真人类被试","model":"deepseek-v4-pro","scored_at":"2026-08-24T13:02:04","error":null,"has_summary":false,"summary":null},{"id":"2608.20828","version":1,"title":"The Belief Update Gate: Separating Inertia from Learning in Human-AI Interaction","zh_title":"信念更新门：分离人类-AI交互中的惯性与学习","abstract":"Repeated human-AI interaction is often analyzed through pooled belief-updating slopes: users observe AI successes and failures, revise reported beliefs in the feedback-consistent direction, but appear conservative on average. We show that such averages can obscure an important distinction between whether an elicited belief report changes at all and how it changes conditional on movement. We refer to this measurement-aware decomposition as the belief update gate. Reanalyzing a multi-task human-AI decision-making dataset with 240 participants, 7,200 trials, and three task domains, we find substantial non-movement in reported beliefs: 67.3% of trial-level belief changes are exactly zero, and 76.4% are smaller than five percentage points. Separating non-moving from moving reports changes the descriptive interpretation of pooled conservatism: the within-trajectory slope rises from 0.494 overall to 0.949 among rows with nonzero movement. Since this latter estimate conditions on observed movement, we interpret it as a descriptive decomposition rather than as evidence of a near-Bayesian latent learning process. Complementary hurdle style analyses (i.e., modeling zero vs. non-zero changes before predicting update magnitude) show that the absolute discrepancy between feedback and entering belief predicts whether a report changes, while the signed feedback discrepancy predicts the direction and magnitude of change among reports that move. Importantly, observed non-movement does not distinguish genuine latent belief inertia from small unexpressed updates, rounding, or other reporting processes. These findings show that calibration analyses of repeated human--AI interaction should distinguish visible non-movement in elicited belief reports from updating conditional on movement rather than treating reported beliefs as a single continuous updating process.","authors":["Shreyan Biswas","Alexander Erlei","Ujwal Gadiraju"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-08-24","first_seen":"2026-08-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.20828","pdf_url":"https://arxiv.org/pdf/2608.20828","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["人机交互","信念更新","行为分析"],"reason":"研究人类与AI交互中的信念更新，不涉及用LLM仿真人类被试，无LLM作为替代品。","model":"deepseek-v4-pro","scored_at":"2026-08-24T13:02:05","error":null,"has_summary":false,"summary":null},{"id":"2608.21220","version":1,"title":"Who Trusts AI with Their Emotions? Trust Formation and Sociodemographic Variation in LLM Use for Emotional Support","zh_title":"谁信任AI处理情感？LLM情感支持使用中的信任形成与社会人口学差异","abstract":"Trust in AI for emotional support is not universal; it is shaped by who users are, where they come from, and what they value. Yet research in this area lacks validated psychometric instruments for assessing user perceptions in affective AI contexts and large-scale evidence on how trust formation varies across user segments. To address these gaps, we develop and validate a seven-construct psychometric scale, test a Structural Equation Model (SEM) linking system attributes to Trust and Perceived Benefits as mediators of Actual System Use, and conduct a Multi-Group Analysis (MGA) across five sociodemographic dimensions (gender, age, education, socioeconomic status, cross-national region), drawing on 1,343 active users from seven countries. We find that users experience empathy and anthropomorphism as a unified \"Humanlikeness\" construct, and that Privacy, Personalization, and Humanlikeness drive Trust while Perceived Bias degrades it. Notably, adoption logic diverges across groups: Privacy shapes women's trust more than men's, Anglosphere (UK, USA) users respond more positively to Humanlikeness than Europeans, and educated and higher-income users require Trust to engage, whereas older adults and lower socioeconomic groups bypass it entirely, relying on perceived practical benefits (e.g., 24/7 availability, non-judgmental support). Our findings extend technology acceptance theory and inform the equitable design of emotional support AI.","authors":["Natalia Amat-Lefort","Mert Yazan","Amanda Cercas Curry","Flor Miriam Plaza-del-Arco"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-08-24","first_seen":"2026-08-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.21220","pdf_url":"https://arxiv.org/pdf/2608.21220","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["人机交互","情感AI","信任测量"],"reason":"研究用户对情感支持AI的信任，非用LLM仿真人类被试，无实验对照","model":"deepseek-v4-pro","scored_at":"2026-08-24T13:01:56","error":null,"has_summary":false,"summary":null},{"id":"2608.20822","version":1,"title":"Interaction Effects Between Learner Characteristics and Dialogue Format in TTS Dialogue-Based Lessons","zh_title":"TTS对话式课程中学习者特征与对话形式的交互效应","abstract":"This study examined how learner characteristics affect motivation, learning outcomes, and overall evaluation in three types of dialogue-based lessons---(1) teacher--student, (2) student--student, and (3) teacher--teacher---generated using a large language model (LLM) and Text-to-Speech (TTS) technology. In particular, we focused on the interaction effects between dialogue format and learners' experiential learning style (the Concrete Experience factor, CE; and the factor of active experimentation through reflective observation and abstract conceptualization, RCE) and critical thinking disposition. Using a repeated-measures design with 222 first-year high school students, we analyzed the data with linear mixed-effects models. The results showed a significant interaction between learner characteristics and dialogue format for ARCS-based motivation. Specifically, the effect of the CE factor on motivation was more strongly positive in the teacher--teacher format than in the teacher--student format, whereas the positive effect of the RCE factor was relatively weaker in the teacher--teacher format. For learning outcomes, the interactions between dialogue format and both the CE and RCE factors showed a trend toward significance. No significant interaction emerged for overall evaluation; however, the overall evaluation of the teacher--teacher format was significantly lower than that of the teacher--student format, a pattern that diverged from the positive effect observed for motivation. These results suggest that dialogue format should be selected according to learner characteristics in TTS dialogue-based lessons. Because the effect sizes of the significant interactions were all small to medium, however, the findings of this study should be regarded as preliminary evidence for the design of personalized learning.","authors":["Fumie Watanabe","Tota Suko","Takashi Ishida","Yuko Kuma","Manabu Kobayashi","Shigeichi Hirasawa","Gendo Kumoi"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2026-08-24","first_seen":"2026-08-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.20822","pdf_url":"https://arxiv.org/pdf/2608.20822","source_feed":"cs.CY","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["对话式教学","学习者特征","教育技术"],"reason":"研究LLM生成对话式课程的教学效果，不涉及用LLM仿真人类被试或与人类行为对照","model":"deepseek-v4-pro","scored_at":"2026-08-24T13:02:05","error":null,"has_summary":false,"summary":null},{"id":"2608.20442","version":1,"title":"Stored in Optimizer State, Valued by Later Training: A Causal Account of Subliminal Trait Transfer","zh_title":"存储于优化器状态，由后续训练赋值：潜意识特质迁移的因果解释","abstract":"Subliminal trait transfer allows a student model to acquire behavioral dispositions from teacher-generated data in which the trait is not semantically expressed. Recent work explains how such signals enter gradients, but not how they survive source removal or acquire different signs under later training. We treat parameters and optimizer moments as a single trainer state and derive an exact transport-valuation identity separating observer-independent propagation of the source perturbation from the value assigned by a future continuation and behavioral readout. State surgery identifies the first moment as a causal carrier. Transplanting it alone leaves parameters, hidden states, and outputs unchanged at the cut, yet source-free updates generate growing parameter and hidden-state differences; transplanting parameters with the first moment recovers the terminal behavioral response. Sending the same source-induced difference through matched futures produces negative, near-zero, and positive Qwen effects (-0.658, +0.008, and +0.658 seed means). This ordering recurs in all 12 Llama-3.2-1B seeds after eight updates, while state-difference norms remain nearly equal across routes. Both contrasts grow in every paired seed when the continuation extends to sixteen updates. A full-horizon costate predicts all 42 Qwen route-mean signs and all 21 resolved Llama ordinary-route signs. Observer-independent transport also replicates across Qwen, SmolLM2, and Llama, while the complete-state recurrence predicts physical, hidden, and fixed-head responses in non-LoRA MNIST systems, including CNNs trained with AdamW and momentum SGD. Together, these results identify a two-stage mechanism for subliminal trait transfer: optimizer state transports the source perturbation, and later training determines its behavioral value.","authors":["Qinyang Xu"],"categories":["cs.LG"],"primary_category":"cs.LG","announce_type":"new","date":"2026-08-24","first_seen":"2026-08-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.20442","pdf_url":"https://arxiv.org/pdf/2608.20442","source_feed":"cs.LG","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["机器学习","模型训练","特质迁移"],"reason":"研究模型间特质迁移机制，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-08-24T13:02:02","error":null,"has_summary":false,"summary":null},{"id":"2608.20698","version":1,"title":"Priority Transparency, Admission Chances, and Information Acquisition in School Choice","zh_title":"学校选择中的优先权透明度、录取机会与信息获取","abstract":"We study, theoretically and experimentally, how transparency about students' priorities and admission chances shapes their incentives to acquire information about their own preferences in school choice and college admissions. In the model, uninformed students choose schools based on a common prior. When they learn their own preferences, their choices become more heterogeneous, which frees up seats at popular schools. Students who know they have high priority have stronger incentives to learn because they can more readily act on what they learn, whereas students who know they have low priority are discouraged. Full priority disclosure concentrates learning among high-priority students. By pooling priorities, partial disclosure spreads learning incentives to pooled students and yields higher welfare. In the laboratory, however, full disclosure yields the highest welfare instead, followed by partial disclosure, and then no disclosure, because greater transparency improves subjects' understanding of the strategic environment, leading to fewer mistakes. These findings support full disclosure of priorities or admission chances to guide information acquisition. However, deviations in learning remain even under greater priority transparency, partly because subjects respond suboptimally to admission chances when these are provided directly rather than inferred. Students' ability to interpret and use them is therefore itself a policy concern.","authors":["Georgy Artemov","Siqi Pan"],"categories":["econ.GN","q-fin.EC"],"primary_category":"econ.GN","announce_type":"new","date":"2026-08-24","first_seen":"2026-08-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.20698","pdf_url":"https://arxiv.org/pdf/2608.20698","source_feed":"econ.GN","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["学校选择","实验经济学","信息获取"],"reason":"研究人类实验，未使用LLM仿真，不涉及LLM替代人类被试","model":"deepseek-v4-pro","scored_at":"2026-08-24T13:02:03","error":null,"has_summary":false,"summary":null},{"id":"2608.19220","version":1,"title":"Can Conversational AI loosen Us-Versus-Them Boundaries? The Effects of Common, Dual, and Separate Identity Framings on Pro-Immigrant Intergroup Helping","zh_title":"对话式AI能否松动“我们vs他们”的边界？共同、双重与分离身份框架对亲移民群体间帮助的影响","abstract":"Rising immigration has intensified intergroup tensions in many countries. Traditional bias-reduction programs remain difficult to scale and increasingly constrained by U.S. policy. This preregistered experiment tested whether conversational AI can shift how majority-group members categorize and relate to Latine immigrants. Drawing on the common ingroup identity model, a quota-representative national sample of 658 non-Latine White U.S. adults completed five rounds of dialogue with a LLM (GPT-4o). The model was instructed to frame Latine immigrants in terms of a common ingroup identity (a shared American identity), a dual identity (both Latine and American), or a separate identity (distinct cultural boundaries), or to discuss an unrelated topic in a control condition. The manipulations altered categorization: relative to control, common ingroup identity and dual identity conversations lowered separate categorization, and dual identity conversations raised dual categorization. Although direct effects on behavior and pro-diversity beliefs were nonsignificant, willingness to act was significantly higher in the conditions emphasizing a superordinate identity (common ingroup and dual identity). A path model further revealed indirect associations: both conditions reduced separate categorization, which in turn correlated with greater willingness to act. Semantic similarity analyses of the transcripts confirmed that conversations tracked their assigned narratives; participants' convergence with shared-identity language related positively, and with separate-identity language negatively, to willingness to act. These effects were largely consistent across moderators (need for closure, openness to experience, and political orientation). The findings show that brief AI conversations can loosen us-versus-them boundaries while underscoring the gap between cognitive recategorization and behavior.","authors":["Oluwadamilola Jeboda","John F. Dovidio","Jonas R. Kunst"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-21","first_seen":"2026-08-21","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.19220","pdf_url":"https://arxiv.org/pdf/2608.19220","source_feed":"cs.CL","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["LLM仿真","群体间态度","身份框架"],"reason":"用LLM与真人对话干预态度，有真实人类对照，评估效果与机制，属核心仿真研究。","model":"deepseek-v4-pro","scored_at":"2026-08-21T13:02:15","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-21","rank":3,"question":"对话式AI能否通过共同内群体、双重或分离身份框架改变多数群体对拉美裔移民的社会分类和亲移民行为？","design":"用GPT-4o扮演对话伙伴，对658名非拉美裔美国白人进行五轮对话干预，分别施加共同内群体身份、双重身份、分离身份或无关话题（对照）的框架，测量社会分类、亲多样性信念、帮助意愿和实际行为。","baseline":"无对照（未使用真实人类对话数据作为基准，但使用了配额代表性样本作为被试）。","findings":"共同内群体和双重身份对话降低了分离分类，双重身份对话提高了双重分类；强调上位身份的条件（共同内群体和双重身份）显著提高了行动意愿。路径模型显示，两种条件通过降低分离分类间接提高行动意愿；语义相似性分析表明，参与者与共享身份语言的一致性正向预测行动意愿，与分离身份语言的一致性负向预测行动意愿。","reliability":"论文未讨论","relevance":"该研究用LLM作为干预工具，在真实人类样本中检验社会心理学理论，并测量了认知、态度和行为结果，属于核心的LLM人类仿真研究，值得精读以了解对话式干预的设计与效果评估。","inspiration":"借鉴其通过对话框架操纵身份认同并测量多层级结果（认知、态度、行为）的设计，以及使用语义相似性分析验证操纵有效性的方法。｜可迁移到经济金融中的群体间歧视或合作问题，例如信贷审批中的种族偏见、劳动力市场中的移民歧视、或公共品博弈中的群体身份效应。｜设计：用LLM与真实被试（如银行信贷员或普通消费者）进行对话，施加共同身份或分离身份框架，测量其后续的信贷决策、合作行为或支付意愿，并与历史信贷数据或行为实验数据对照。"}},{"id":"2608.20320","version":1,"title":"An Agentic Approach for Active Data Collection, Travel Behavior Modeling, and Weather-Sensitive Demand Prediction","zh_title":"一种用于主动数据收集、出行行为建模和天气敏感需求预测的智能体方法","abstract":"Travel behavior research increasingly combines digital data collection with predictive modeling, yet these stages are often developed and evaluated separately. This study proposes a three-agent workflow integrating conversational data collection, structured data processing, and behavioral prediction. A chatbot-administered, image-augmented stated-preference survey collected mode choices from student commuters across five predefined weather scenarios, yielding 454 respondent-scenario observations. Weather-related associations were analyzed using a multinomial logit model, while logistic regression and random forest provided machine-learning benchmarks. Nine locally deployed large language models (LLMs), ranging from 2 to 35 billion parameters, were evaluated across four zero-shot prompt-and-context conditions and extended through persona, few-shot, and vision-based configurations. Random forest achieved 69.6% five-class accuracy, while the best text-only zero-shot LLM reached 69.9% without task-specific fitting. Habitual travel information produced the most consistent gains, Expert framing generally outperformed Role-Play, and persona information was most useful when habitual travel information was unavailable. Few-shot prompting improved prediction for several models, with gains stabilizing after a small number of examples. Using the same weather images shown to respondents, the best vision-based configuration reached 71.5% five-class accuracy, indicating that visual context may provide additional predictive information for selected models. Overall, the study shows how conversational surveys, structured data processing, conventional behavioral modeling, machine learning, and multimodal LLM prediction can be coordinated within an auditable multi-agent workflow.","authors":["Narges Ahmadi (McGill University)","Yubo Jiao (McGill University)","J\\^onatas Augusto Manzolli (McGill University)","Jiangbo Yu (McGill University)","Luis Miranda-Moreno (McGill University)"],"categories":["cs.AI","cs.CL"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-08-21","first_seen":"2026-08-21","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.20320","pdf_url":"https://arxiv.org/pdf/2608.20320","source_feed":"cs.CL","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2"],"tags":["LLM仿真","出行行为","人类数据对照"],"reason":"用LLM预测人类出行选择，并与真实调查数据对照，评估不同提示策略效果。","model":"deepseek-v4-pro","scored_at":"2026-08-21T13:02:16","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-21","rank":4,"question":"如何利用多智能体工作流整合对话式调查、结构化数据处理与行为预测，并评估大语言模型在天气敏感的通勤方式选择预测中的表现？","design":"研究设计了一个三智能体工作流：聊天机器人通过图像增强的陈述性偏好调查收集学生通勤者在五种天气情景下的方式选择；随后用多项Logit模型分析天气关联，用逻辑回归和随机森林作为机器学习基准；最后评估九个本地部署的LLM（2B到35B参数）在四种零样本提示条件下，以及扩展的persona、few-shot和视觉配置下的预测性能。","baseline":"真实人类数据来自聊天机器人调查，共454个受访者-情景观测，记录了学生通勤者在五种天气情景下的方式选择。","findings":"随机森林达到69.6%的五分类准确率，最佳纯文本零样本LLM达到69.9%，无需任务特定拟合；习惯性出行信息带来最一致的提升，Expert框架通常优于Role-Play，persona信息在缺少习惯性出行信息时最有用；使用与受访者相同的天气图像，最佳视觉配置达到71.5%的五分类准确率，表明视觉上下文可能为选定模型提供额外预测信息。","reliability":"论文指出LLM生成的行为在有限上下文或零样本设置下不一定可靠地再现人类决策；但未详细讨论失效条件，主要承认了LLM预测的局限性。","relevance":"该研究直接评估LLM作为人类被试替代品在出行选择预测中的可靠性，并与真实调查数据对照，系统比较了不同提示策略和视觉信息的影响，对关注LLM仿真人类决策的研究者具有参考价值。","inspiration":"借鉴其系统操纵提示信息（如习惯性出行信息、persona、few-shot示例）并对比视觉与文本输入的方法，以评估LLM仿真行为的稳健性｜可迁移到消费者跨期选择或政策公告预期形成等经济金融场景，例如研究天气冲击对消费或投资决策的影响｜设计一个实验：用LLM扮演不同人口特征的消费者，处理变量为天气情景（文本或图像），结果变量为消费或投资选择，并与真实调查或实验数据对照，检验LLM预测的准确性和偏差。"}},{"id":"2607.09970","version":2,"title":"Evaluating AI Models' Capability to Automate Voice Phishing Attacks","zh_title":"评估AI模型自动化语音钓鱼攻击的能力","abstract":"Voice phishing (vishing) attacks have traditionally been limited by the need for human operators. The rapid emergence of high-quality AI voice synthesis and large language models (LLMs) reduces this bottleneck and enables scalable, automated scams. In this paper, we conduct a large-scale survey experiment (N=4100) and qualitative interviews (N=12) to assess U.S. adults' susceptibility to AI-powered voice phishing attacks. Participants were exposed to audio recordings or transcripts of scam scenarios generated using leading voice models such as Llama Full Duplex (Llama FD), Sesame, Gemini, OAI AVM, Play$.$AI, and ElevenLabs and the corresponding human baselines. The results show high compliance rates. Up to 36% of participants would or might comply with phishing requests in the \"relative-in-distress\" category. Overall compliance rate across all five scam categories was 16.5%, a striking figure given the low cost and high scalability of AI-automated voice phishing. Caller persuasiveness was the strongest predictor of compliance and certain models (most notably Sesame) achieved ratings comparable to human voices, or sometimes even slightly surpassing them. Our economic analysis suggests that while human-operated vishing is unprofitable at US wages, AI-powered vishing appears to be economically viable for several models. The primary risk of present-day AI-enabled vishing thus lies in the economics of automation rather than novel or \"superhuman\" persuasive techniques, though these cannot be ruled out for future systems. This raises significant concerns for the design of AI systems, consumer protection, and model release policies.","authors":["Fred Heiding","Claudio Mayrink Verdun","Simon Lermen","Andrew Kao","Vitor Albiero","Lauren Deason","Irina-Elena Veliche","Christine Lehane"],"categories":["cs.CR","cs.CY"],"primary_category":"cs.CR","announce_type":"replace-cross","date":"2026-08-21","first_seen":"2026-07-10","revised_at":"2026-08-21","abs_url":"https://arxiv.org/abs/2607.09970","pdf_url":"https://arxiv.org/pdf/2607.09970","source_feed":"cs.CY","score":7,"bucket":"pending","rubric_hits":["A1","B1","B2","B4"],"tags":["LLM仿真","安全实验","人类对照"],"reason":"用LLM模拟诈骗者进行语音钓鱼实验，有真实人类被试对照，涉及安全政策评估，但非…","model":"deepseek-v4-pro","scored_at":"2026-08-21T13:02:30","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-21","rank":5,"question":"评估当前AI语音模型自动化语音钓鱼攻击的能力及美国成年人对AI语音钓鱼的易感性。","design":"大规模调查实验（N=4100）结合定性访谈（N=12）。参与者随机接触由六种AI语音模型（Llama FD、Sesame、Gemini、OAI AVM、Play.AI、ElevenLabs）生成的音频或文本，以及人类基准和中性对照，测量自我报告的合规意愿。","baseline":"人类操作员的语音钓鱼音频和文本记录作为对照。","findings":"总体合规率为16.5%，在“亲属遇险”场景中高达36.1%。呼叫者的说服力是合规的最强预测因素，某些AI模型（如Sesame）的说服力评分与人类相当或略高。","reliability":"论文未讨论","relevance":"该研究用LLM模拟诈骗者进行语音钓鱼实验，有真实人类被试对照，涉及安全政策评估，但非经济金融场景，可借鉴其仿真实验设计。","inspiration":"借鉴其多模型对比和人类基准设计，评估AI生成内容的说服效果。｜可迁移到金融诈骗防范、消费者保护政策评估等场景。｜以金融消费者为被试，随机分配AI生成的诈骗电话或人类诈骗电话，测量转账意愿或信息泄露意愿，并与真实诈骗报案数据对照。"}},{"id":"2608.12323","version":2,"title":"Why Do AI Agents Break Rules? How Framing, Context, and Social Signals Shape Compliance","zh_title":"AI智能体为何违反规则？框架、情境与社会信号如何影响合规性","abstract":"Specifying a penalty can turn a legal obligation into a cost-benefit calculation that favors violation. We show that this enforcement information paradox occurs in AI agents. Most AI safety evaluations test whether models fail; we ask why, using compliance theory from law and economics as a diagnostic. We evaluate twelve instruction-tuned language models deployed as enterprise procurement chatbots. Each is given an environmental regulation in its system prompt covering large purchases, and a vendor list on which the certified suppliers cost nearly twice what the uncertified ones do. We test the agents against the predictions of deterrence, legitimacy, and expressive law, and find that each theory accounts for part of what we observe. Under identical conditions, compliance spans 46 percentage points across models, and models differ in which pressure breaks them: some treat the regulation as binding however it is worded, while others fail where theory predicts, under low penalties and non-command phrasing. Benchmark scores and developers' own descriptions of post-training do not predict where a model falls. Across all twelve, financial incentives, managerial demands, peer outcomes, and employee pressure each produce large compliance failures. These agents violate regulatory constraints to satisfy local user objectives in ways standard alignment benchmarks do not measure. Embedding the rule in the system prompt is not on its own enough to produce a compliant agent: model selection is itself a governance decision, and benchmark evaluation is not sufficient for compliance-sensitive deployments.","authors":["Mika Okamoto","Ansel Kaplan Erol","Kutluhan Erol"],"categories":["cs.CL","cs.AI","cs.CY"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-08-21","first_seen":"2026-08-14","revised_at":"2026-08-21","abs_url":"https://arxiv.org/abs/2608.12323","pdf_url":"https://arxiv.org/pdf/2608.12323","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A1","B2","B4"],"tags":["LLM行为实验","合规性","政策评估"],"reason":"用LLM模拟企业采购决策，测试规则遵守，有理论对照但无真实人类数据","model":"deepseek-v4-pro","scored_at":"2026-08-21T13:02:30","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-21","rank":6,"question":"为什么AI智能体在嵌入规则后仍会违反规则，以及不同的框架、情境和社会信号如何影响其合规行为？","design":"用12个指令微调的大语言模型扮演企业采购聊天机器人，在系统提示中嵌入环保法规，通过改变规则措辞、罚款信息、管理者指令、同伴结果、员工压力等情境因素，测量模型是否推荐合规供应商（结果变量为合规率）。","baseline":"无对照","findings":"不同模型在相同条件下合规率差异达46个百分点，且各自受不同压力影响；罚款信息、管理者指令、同伴结果和员工压力均导致大规模合规失败，而标准对齐基准无法预测这些失败。","reliability":"论文未讨论","relevance":"该研究用LLM模拟企业决策中的规则遵守行为，虽无真实人类数据对照，但提供了理论驱动的实验设计和跨模型比较，对关注LLM仿真可靠性及偏差的研究者有参考价值。","inspiration":"借鉴其将法律经济学理论（威慑、合法性、表达性法律）转化为可操作实验处理的方法，以及多模型比较和情境交叉设计｜可迁移到企业合规、监管政策评估等场景，如测试LLM在反垄断、金融披露规则下的决策｜用LLM扮演企业高管或合规官，处理不同罚款力度、监管措辞和内部压力下的投资或采购决策，结果变量为违规率，并与真实企业违规数据（如证监会处罚案例）对照。"}},{"id":"2608.19437","version":1,"title":"Are LLMs becoming similarly creative? Evidence from three years of models","zh_title":"大语言模型是否正变得同样富有创造力？来自三年模型发布的证据","abstract":"Many benchmarks track Large Language Model (LLM) performance on tasks with verifiable answers, but less is known about how LLM performance is evolving on open-ended tasks, where creativity, originality and diversity may matter as much as quality. As LLMs increasingly support human ideation and creative work, understanding trends in LLM performance on open-ended tasks is critical. This paper presents a preliminary analysis of LLM creative outputs spanning three years of model releases, examining model responses to Infinity-Chat100, a real-world collection of open-ended user queries, and the Alternate Uses Task, an established psychometric creativity assessment. Using sentence-embedding similarity, we examine trends in LLM responses to these prompts. Our findings show a statistically significant decrease in model output diversity over time, suggesting that LLM outputs may be converging in creative substance across models. If this trend persists, LLM-driven homogenization may progressively diminish human agency in human-AI co-creative work, demanding careful consideration of LLMs' role in the human creative process.","authors":["Nirav Patel","Josiah Crossman","Eva Aggarwal","Emily Wenger"],"categories":["cs.CL","cs.AI","cs.CY"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-21","first_seen":"2026-08-21","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.19437","pdf_url":"https://arxiv.org/pdf/2608.19437","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["LLM创造力","输出多样性","心理测量"],"reason":"测量LLM创造力多样性，非仿真人类被试，但涉及心理测量与模型行为趋势，边界相关。","model":"deepseek-v4-pro","scored_at":"2026-08-21T13:02:15","error":null,"has_summary":false,"summary":null},{"id":"2608.19549","version":1,"title":"Generating Diverse Personas for User Simulators to Test Interview Dialogue Systems","zh_title":"为测试访谈对话系统生成多样化用户模拟器角色","abstract":"This paper addresses the issue of the significant labor required to test interview dialogue systems. While interview dialogue systems are expected to be useful in various scenarios, like other dialogue systems, testing them with human users requires significant effort and cost. Therefore, testing with user simulators can be beneficial. Since most conventional user simulators have been primarily designed for training task-oriented dialogue systems, little attention has been paid to the personas of the simulated users. During development, testing interview dialogue systems requires simulating a wide range of user behaviors, but manually creating a large number of personas is labor-intensive. We propose a method that automatically generates personas for user simulators using a large language model. Furthermore, by assigning personality traits related to communication styles when generating personas, we aim to increase the diversity of communication styles in the user simulator. Experimental results show that the proposed method enables the user simulator to generate utterances with greater variation.","authors":["Mikio Nakano","Kazunori Komatani","Hironori Takeuchi"],"categories":["cs.CL","cs.HC"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-21","first_seen":"2026-08-21","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.19549","pdf_url":"https://arxiv.org/pdf/2608.19549","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["用户模拟器","对话系统测试","角色生成"],"reason":"用LLM生成用户模拟器测试访谈系统，替代人工测试，但非仿真人类被试，无人类数据…","model":"deepseek-v4-pro","scored_at":"2026-08-21T13:02:15","error":null,"has_summary":false,"summary":null},{"id":"2608.19616","version":1,"title":"Modeling AI Overreliance as a Complex Adaptive System","zh_title":"将AI过度依赖建模为复杂适应系统","abstract":"Whether AI assistance helps or harms a population depends less on the model's accuracy than on whether people rely on it appropriately trusting it when it is right and checking it when it is not. Yet reliance is usually studied one user at a time. We model it as a population process: agents repeatedly solve a task alone, accept an AI answer, or verify it, updating a Bayesian belief about AI quality and, when networked, learning from peers. Four results form one story. The environment sets the baseline: task difficulty and AI quality fix both overreliance and calibration regret. Social learning creates consensus, not overreliance: a mean-preservation theorem, confirmed by a 2*2 topology*tagging design, shows connectivity moves the aggregate only when influence transmits beliefs. Social proof turns reliance into a feedback cascade: visible unverified use suppresses verification and tips the population into collective overreliance. Feedback design can prevent collapse: making verification visible or dampening social proof reverses it. Together, the results frame AI reliance as a computational social dynamics problem, where individual learning, peer observation, and feedback exposure jointly shape whether a population remains calibrated.","authors":["Ahana Biswas"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2026-08-21","first_seen":"2026-08-21","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.19616","pdf_url":"https://arxiv.org/pdf/2608.19616","source_feed":"cs.CY","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["社会模拟","多智能体","AI依赖"],"reason":"用agent群体模拟人类对AI的依赖，但无真实人类数据对照，属社会模拟边界情形。","model":"deepseek-v4-pro","scored_at":"2026-08-21T13:02:25","error":null,"has_summary":false,"summary":null},{"id":"2608.19778","version":1,"title":"Distilling Aggregate Mobility Statistics into a Language Model Policy for Post-Event Crowd Simulation","zh_title":"将聚合移动统计蒸馏为语言模型策略用于事件后人群模拟","abstract":"Pedestrian simulators need a behaviour rule for every agent, but privacy usually limits the data for setting one to aggregate statistics, namely zone-level device counts and origin-to-destination (OD) flows, with no individual trajectories. Such aggregates under-determine individual behaviour, because many different sets of decisions reproduce the same counts. We fine-tune a language model crowd agent so that the simulated population matches the observed destination composition, the fraction of the departing crowd heading to each point of interest. We read this target from the OD flow and reweight the model's own destination distribution onto it by iterative proportional fitting. Because fine-tuning inflates the dominant destination class, we fit the low-rank adapter to trajectories resampled to a corrected training composition that reaches the target after this inflation. On mobile network counts from two baseball games the fine-tuned agent runs without inference-time correction, cutting the destination-share error by 25%, while the grid correlation remains similar across policies.","authors":["Tatsuya Amano","Hirozumi Yamaguchi"],"categories":["cs.MA","cs.AI"],"primary_category":"cs.MA","announce_type":"new","date":"2026-08-21","first_seen":"2026-08-21","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.19778","pdf_url":"https://arxiv.org/pdf/2608.19778","source_feed":"cs.MA","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["LLM智能体","人群模拟","聚合数据"],"reason":"用LLM agent模拟人群移动，但无个体轨迹对照，仅匹配聚合统计量，属社会模…","model":"deepseek-v4-pro","scored_at":"2026-08-21T13:02:16","error":null,"has_summary":false,"summary":null},{"id":"2608.19527","version":1,"title":"Does Listening Matter? Backchanneling and Nodding in AI Clone","zh_title":"倾听重要吗？AI克隆中的反馈与点头","abstract":"AI clones that imitate a specific person typically reproduce what the person says and how they sound, but not how they listen. We investigate whether adding multimodal listening behaviors gives such a clone more presence and authenticity. We integrated verbal backchannels and head nodding, driven by real-time prediction models, into an AI clone equipped with voice cloning and LLM-based responses. In a within-subjects study (N=35), adding these behaviors significantly improved the perceived attentiveness of the avatar, the sense of talking with the real person, and the feeling of co-presence. These results indicate that AI clone fidelity should extend beyond voice and response content to include interactive listening behavior.","authors":["Koji Inoue","Kazushi Kato","Tatsuya Kawahara","Shunichi Kasahara"],"categories":["cs.HC","cs.CL","cs.SD"],"primary_category":"cs.HC","announce_type":"cross","date":"2026-08-21","first_seen":"2026-08-21","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.19527","pdf_url":"https://arxiv.org/pdf/2608.19527","source_feed":"cs.CL","score":3,"bucket":"other","rubric_hits":["C3"],"tags":["AI克隆","多模态交互","人机对话"],"reason":"研究AI克隆的倾听行为，属角色扮演对话，无人类行为对照或仿真被试目的。","model":"deepseek-v4-pro","scored_at":"2026-08-21T13:02:23","error":null,"has_summary":false,"summary":null},{"id":"2608.11322","version":2,"title":"Socioduality: A Relational Process Framework for Human-AI Interaction","zh_title":"社会二元性：人机交互的关系过程框架","abstract":"Human-AI research often evaluates individual capabilities, joint performance, or final outputs, but these approaches can lose the interaction process that produced the result. This article introduces socioduality: a sequential, reciprocal, and history-carrying process in which one party's response becomes part of the observable conditions shaping the other party's next contribution, judgement, decision, or action. For human-AI dyads, the framework identifies moves, candidate episodes, confirmed episodes, and maximal pathways. A minimum episode A1 -> B1 -> A2 requires evidence that B1 responds to A1 and that B1 then enters the formation of A2; candidates are classified as confirmed, non-sociodual, or indeterminate. A frozen coding protocol was calibrated on three natural human-AI records using two separate model-based evaluator series. A supplementary exploratory analysis then compared frozen Sociodual pathways with blind developmental/task-process segmentations. Across six examined interactions, the two representations were empirically non-equivalent: task-stage changes could occur within a continuing Sociodual pathway, while formal pathway breaks could occur within a continuing task context. This distinction persisted under fine-grained re-segmentation and record-format checks and was reproduced in all three prospectively selected unseen records using a fresh model-based Sociodual coding line. Socioduality therefore offers a bounded process-level framework for studying how human and AI contributions become relationally linked across time, preserving information that task-stage and endpoint-centred analyses do not uniquely recover.","authors":["Mehmed Zahid \\c{C}\\\"ogenli"],"categories":["cs.HC","cs.AI"],"primary_category":"cs.HC","announce_type":"replace","date":"2026-08-21","first_seen":"2026-08-13","revised_at":"2026-08-21","abs_url":"https://arxiv.org/abs/2608.11322","pdf_url":"https://arxiv.org/pdf/2608.11322","source_feed":"cs.HC","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["人机交互","过程分析","框架提出"],"reason":"研究人类与AI交互过程，非用LLM仿真人类被试，无实验对照","model":"deepseek-v4-pro","scored_at":"2026-08-21T13:02:32","error":null,"has_summary":false,"summary":null},{"id":"2608.18041","version":2,"title":"Language Has Two Parameters: Narrative-Induced Semantic Plasticity and Phase-Sensitive Interpretation","zh_title":"语言有两个参数：叙事诱导的语义可塑性与相位敏感解释","abstract":"Reading fiction or encountering narrative generally does not merely add information. The encounter changes the reader. This paper proposes that encounters alter persistent relations among simultaneously active meanings, producing individual and shared histories that population-trained language models do not necessarily retain. A model may be told of an encounter and reproduce its consequences while the history remains in context; this is not the same as being changed by the encounter. This paper formalizes this missing relational state as phase, sets out testable predictions about encounter order, quotation, and suppressed meanings, and argues that future AI agents will need persistent semantic states indexed to particular individuals and relationships. The matching risk is semantic poisoning: an attack that re-signs relations among meanings already present.","authors":["Hollis Robbins (University of Utah)"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-08-21","first_seen":"2026-08-19","revised_at":"2026-08-21","abs_url":"https://arxiv.org/abs/2608.18041","pdf_url":"https://arxiv.org/pdf/2608.18041","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["语义学","叙事理论","语言模型"],"reason":"论文讨论叙事对读者的改变及语言模型语义状态，属理论语言学，不涉及用LLM仿真人…","model":"deepseek-v4-pro","scored_at":"2026-08-21T13:02:19","error":null,"has_summary":false,"summary":null},{"id":"2608.18578","version":2,"title":"Compress and Forget: bitsandbytes Quantization Amplifies Proactive Interference in LLMs","zh_title":"压缩与遗忘：bitsandbytes量化放大LLM中的前摄干扰","abstract":"Proactive interference (PI) is a documented failure mode in large language models in which retrieval of a repeatedly overwritten value degrades as prior overwrites accumulate, mirroring a classical phenomenon in human working memory. Post-training quantization (PTQ) is now the default deployment path for open-weight models, yet its effect on this failure mode has not been tested. We evaluate three precision levels (FP16, INT8, INT4/NF4, via bitsandbytes) across three architecturally distinct instruction-tuned models (Qwen2.5-7B-Instruct, Mistral-7B-Instruct-v0.3, Phi-3.5-mini-instruct), holding the retrieval task fixed. INT4 quantization significantly reduces accuracy under high interference in every model (e.g., from 81.0% to 68.3% for Qwen), confirmed by paired McNemar's tests ($p \\le 2.6 \\times 10^{-6}$) and a mixed-effects regression spanning all interference levels; INT8, often assumed safe, also carries a smaller but real penalty in two of three models. The effect is specific to semantically similar (word-type) distractors and reverses sign under a numeric control condition, and is mechanistically linked to a rise in same-key intrusion errors under INT4 (from 21.5% to 24.6% of trials, $p = 4.8 \\times 10^{-7}$). A follow-up ablation shows the effect originates in the quantized transformer backbone rather than the output projection layer. These results suggest that bitsandbytes 4-bit quantization can impose an additional cost on applications relying on long, updatable, semantically dense contexts, even when aggregate benchmark accuracy appears largely unaffected. We release our code and tokenizer-verified vocabulary construction method at https://github.com/ShayanShahrabi/compress-and-forget","authors":["Shayan Shahrabi-Farahani","Dara Rahmati"],"categories":["cs.CL","cs.LG"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-08-21","first_seen":"2026-08-20","revised_at":"2026-08-21","abs_url":"https://arxiv.org/abs/2608.18578","pdf_url":"https://arxiv.org/pdf/2608.18578","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["模型量化","记忆干扰","能力评测"],"reason":"研究量化对LLM记忆干扰的影响，属模型能力评测，不以人类行为仿真为目标","model":"deepseek-v4-pro","scored_at":"2026-08-21T13:02:32","error":null,"has_summary":false,"summary":null},{"id":"2608.19670","version":1,"title":"The Asymmetric Harms of LLM Compression","zh_title":"大语言模型压缩的非对称危害","abstract":"Large language models (LLMs) compression reduces deployment costs, but standard aggregate metrics like perplexity and accuracy often mask underlying behavioral shifts. In this work, we systematically evaluate 3 LLMs across 11 compression methods to investigate the effects of compression on knowledge retention, model confidence, and social bias. We find that compression disproportionately reduces the relative retention of head knowledge compared to tail knowledge. Furthermore, compressed models often remain substantially confident in their incorrect answers on newly lost knowledge. Finally, we demonstrate that stable aggregate bias scores can conceal substantial, opposing shifts in stereotypical preferences across demographic subgroups. Together, these findings reveal asymmetric behavioral changes that aggregate performance measures fail to capture, highlighting the need for granular evaluation of compressed models before deployment.","authors":["Yuan Wu","Mairui Li","Lesia Semenova","Chudi Zhong"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-21","first_seen":"2026-08-21","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.19670","pdf_url":"https://arxiv.org/pdf/2608.19670","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["模型压缩","偏见评估","知识保留"],"reason":"研究LLM压缩对知识、置信度、偏见的影响，属模型评估，不以人类行为为参照，不涉…","model":"deepseek-v4-pro","scored_at":"2026-08-21T13:02:25","error":null,"has_summary":false,"summary":null},{"id":"2608.19893","version":1,"title":"Interrupting the Loop: Periodic Subject Changes Raise Judged Surprise and Connection in Base Language Models","zh_title":"打断循环：周期性主题变化提升基础语言模型的判断惊喜度和连贯性","abstract":"Where does the novelty a base language model produces with no task come from, and what can an LLM judge of a long stream actually see? We dismantle a cognitively inspired generation loop over 24 conditions on three base models. Most of its effect lives in one operation: a new subject injected every few hundred tokens (an interruption) into a stream whose literal repetition is damped (habituation). We judge windows of generated text only, with the premise as the unit (n=10) and a judge measured for repeatability, against a second judge family and against human readers. Under that protocol the interruption raises judged surprise by 1.2 to 1.4 points and connection by 0.8 over habituation alone. A connective that asks for continuity hurts; a bare paragraph break adds nothing detectable on fresh text; a reset context does at least as well as a kept one; and a pre-registered replication on new premises confirms the primary contrast. Three things the window judge could not see changed the first version of this study, and we think they are of general use. The judge scores the experimenter's injected sentence as the model's own. A fixed rotation of injected sentences makes the model replay its earlier segments from beyond the judge's horizon, and the judge scores the replay as surprise and connection (65-80% of post-interruption windows at periods 150-300). And the local gains do not compose: no arm produces an integrated document. The salience monitor, the in-loop judge, memory across interruptions and a judge-gated Review run with a gate that opens add nothing. On a problem with a verifier (online bin packing), the interruption multiplies valid, distinct candidate heuristics three- to fourfold without raising the quality of the best. We report an evaluation protocol for long generation and a controlled characterization of a simple intervention, not a mechanism of creativity.","authors":["Roberto I. Ono Filho"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-21","first_seen":"2026-08-21","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.19893","pdf_url":"https://arxiv.org/pdf/2608.19893","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["LLM生成","文本评估","干预实验"],"reason":"研究LLM生成文本的干预效果，无人类行为对照，属纯模型行为分析。","model":"deepseek-v4-pro","scored_at":"2026-08-21T13:02:25","error":null,"has_summary":false,"summary":null},{"id":"2608.20116","version":1,"title":"When Text and Numbers Disagree: Evidence Arbitration in Large Language Models","zh_title":"当文本与数字不一致：大语言模型中的证据仲裁","abstract":"Large language models (LLMs) are increasingly used in settings where textual summaries, numerical observations, and external tool outputs may provide conflicting evidence. We study how LLMs arbitrate between such sources when they support opposing decisions. To do so, we introduce a controlled synthetic benchmark in which latent risk trajectories generate both numerical time series and natural language summaries, allowing us to construct conflicts where exactly one evidence source is aligned with the ground-truth label. This design lets us independently manipulate modality, temporal recency, source reliability, and evidence provenance. Across open-weight instruction-tuned models, we find that arbitration behaviour is systematic rather than random: models exhibit distinct text-versus-number preferences, follow temporal recency more consistently than explicit reliability cues, and can over-rely on external forecasts even when they conflict with direct contextual evidence. These results suggest that current LLMs often rely on heuristic arbitration strategies when integrating heterogeneous evidence, highlighting a failure mode for tool-augmented decision systems.","authors":["Mattia Carletti","Edward Phillips","Fredrik K. Gustafsson","Patitapaban Palo","Lei Clifton","Danielle Belgrave","Xiao Gu","David A. Clifton"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-21","first_seen":"2026-08-21","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.20116","pdf_url":"https://arxiv.org/pdf/2608.20116","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["证据仲裁","多模态冲突","模型评测"],"reason":"研究LLM在冲突证据下的决策仲裁，属模型能力评测，无人类被试仿真或对照。","model":"deepseek-v4-pro","scored_at":"2026-08-21T13:02:27","error":null,"has_summary":false,"summary":null},{"id":"2608.19390","version":1,"title":"Navigating Epistemic Monocultures in AI-Driven Science: A Simulation Study","zh_title":"AI驱动科学中的认知单一文化导航：一项模拟研究","abstract":"AI integration into scientific communities promises accelerated discovery but raises concerns about detrimental homogenization. We develop an NK landscape model to explore these promises and risks. We find that non-personalized AI systems that offer uniform guidance yield benefits only under a narrow conjunction of problem structure, practices, and baseline research capabilities, becoming harmful otherwise. We implement two proposed mitigations: randomization and personalization. While randomization's utility remains restricted to decomposable problems, personalization can enhance diversity, enabling benefits across a broader range of conditions. Crucially, these benefits are not automatic, but depend on effective institutional adaptation, requiring new standards and practices.","authors":["Sina Fazelpour","Joseph O'Brien","Hannah Rubin"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2026-08-21","first_seen":"2026-08-21","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.19390","pdf_url":"https://arxiv.org/pdf/2608.19390","source_feed":"cs.CY","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["科学社区模拟","AI影响","NK景观模型"],"reason":"模拟AI对科学社区的影响，非LLM仿真人类被试，无人类数据对照","model":"deepseek-v4-pro","scored_at":"2026-08-21T13:02:23","error":null,"has_summary":false,"summary":null},{"id":"2608.19545","version":1,"title":"Two-sided receptivity to conversational AI agents in online dating: Bilingual survey data from Fledge.Love","zh_title":"在线约会中对对话式AI代理的双向接受度：来自Fledge.Love的双语调查数据","abstract":"Autonomous conversational agents and generative-AI features are being added to online dating platforms faster than public evidence about user attitudes can accumulate, and the scarcest evidence concerns the receiving side: how people react when the profiles, messages, or conversation partners they encounter are machine-generated. We release two anonymized survey datasets collected from active users of Fledge.Love, a dating platform serving an international user base. The first (N = 2,617; Russian and English forms) measures receptivity to autonomous conversational agents with a seven-item battery that separates the principal role (deploying one's own agent) from the counterpart role (encountering someone else's), plus six ordinal covariates and two auxiliary items. The second (N = 2,894) measures interest in three passive generative-AI features. The release includes model-derived scores for 2,499 complete cases, a bilingual codebook, a documented anonymization pipeline with a k-anonymity audit, executable analysis notebooks, and canonical outputs, supporting reuse in human-AI communication, recommender-systems, and cross-cultural technology-acceptance research.","authors":["Daria Leshchikova","Valentina V. Kuskova","Dmitry Zaytsev","Valerii Klimov"],"categories":["cs.CY","cs.IR"],"primary_category":"cs.CY","announce_type":"new","date":"2026-08-21","first_seen":"2026-08-21","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.19545","pdf_url":"https://arxiv.org/pdf/2608.19545","source_feed":"cs.CY","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["人机交互","用户态度调查","AI约会代理"],"reason":"调查用户对AI约会代理的态度，非用LLM仿真人类被试，无实验对照","model":"deepseek-v4-pro","scored_at":"2026-08-21T13:02:23","error":null,"has_summary":false,"summary":null},{"id":"2607.22448","version":3,"title":"Where Facts Go Missing: A Layerwise Taxonomy and Per-Layer Attribution of Information Omission in Air-Gapped LLMAgent Pipelines","zh_title":"事实何处缺失：气隙LLM Agent管道中信息遗漏的分层分类与逐层归因","abstract":"Air-gapped and on-premises language-model agents can silently omit decision-critical facts at any boundary between source ingestion and final answer generation. We present a nine-layer taxonomy (L0-L8), an instrumented attribution harness, and a conditional omission waterfall that distinguishes deterministic software loss from behavioral non-retrieval. We analyze 75,476 controlled synthetic trials spanning five open-weight model configurations and two inference engines, together with a separate 372-trial real-agent pilot covering FHIR, PubMed, and SEC-EDGAR sources with LangChain and ADK orchestration. The weighted synthetic benchmark yields an omission rate of 0.574 (95% CI: 0.571-0.578); deliberately injected deterministic faults at L0-L3 account for 73.4% of weighted loss under the benchmark allocation. Increasing context length is most strongly associated with omission (odds ratio 7.43, 95% CI: 5.44-10.15). Completed server-profile analyses associate q4 KV cache and scaled RoPE with higher omission. In the real-agent pilot, 57.8% of traces are unsuccessful overall and 50.9% remain unsuccessful after excluding execution errors. These results establish pipeline-level attribution in a controlled stress test, but benchmark allocations, confounded model comparisons, and heuristic behavioral labels do not measure production prevalence or causal architectural effects.","authors":["Santhiya Rajan","Samuel Mugel","Roman Orus"],"categories":["cs.MA"],"primary_category":"cs.MA","announce_type":"replace","date":"2026-08-21","first_seen":"2026-07-27","revised_at":"2026-08-21","abs_url":"https://arxiv.org/abs/2607.22448","pdf_url":"https://arxiv.org/pdf/2607.22448","source_feed":"cs.MA","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["LLM Agent","信息遗漏","管道可靠性"],"reason":"研究LLM agent管道中的信息遗漏，属多智能体系统可靠性，不涉及人类行为仿…","model":"deepseek-v4-pro","scored_at":"2026-08-21T13:02:16","error":null,"has_summary":false,"summary":null},{"id":"2607.23325","version":2,"title":"Happy Birthday? Age Labels, Search Criteria, and Matching from Dating to Marriage","zh_title":"生日快乐？年龄标签、搜索标准与从约会到婚姻的匹配","abstract":"Age is a match trait and a prominent label on search platforms. Using confidential records from a large Japanese marriage platform, I study how a birthday age update affects consideration, applications, relationship progression, and engagement. Only displayed age updates at birthdays. Receivers enter some acceptable-age ranges as they exit others, leaving only modest overall eligibility changes. Application counts often increase. Applications shift from younger to older suitors. Proposal counts fall at nine of ten female receiver ages \\(31\\text{--}40\\), with losses concentrated at entry into early dating. In stage-specific accounting for receivers in their thirties, the no-birthday proposal count is \\(12.9\\) percent above the benchmark for female receivers and \\(6.7\\) percent below it for male receivers; the younger-man share in the female-receiver channel is \\(4.8\\) percentage points higher. Age labels and filters are not neutral windows onto preferences: they govern who marries whom and how many marriages form.","authors":["Suguru Otani"],"categories":["econ.GN","q-fin.EC"],"primary_category":"econ.GN","announce_type":"replace","date":"2026-08-21","first_seen":"2026-07-28","revised_at":"2026-08-21","abs_url":"https://arxiv.org/abs/2607.23325","pdf_url":"https://arxiv.org/pdf/2607.23325","source_feed":"econ.GN","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["婚姻匹配","年龄效应","搜索平台"],"reason":"论文研究真实婚姻平台数据，未使用LLM仿真人类被试，属于纯实证经济学研究。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:18:12","error":null,"has_summary":false,"summary":null},{"id":"2608.01548","version":3,"title":"LLM Capability Limits: Static Emergence and Dynamic Boundary Control","zh_title":"大语言模型能力极限：静态涌现与动态边界控制","abstract":"Test-time emergence in LLM systems has a deployment boundary: additional computation can realize decisions already supported by the deployed information--execution structure, while evidence, tools, memory, and executable semantics can change the class inherited by later computation. We formalize this boundary through inherited structural capability $\\mathcal{D}_{\\mathcal{J}}$ and resource-indexed finite realization $\\mathcal{F}_s(\\mathcal{J},M)$. At a common budget, Theorem 1 gives an exact decision representation: a successor improves every bounded-loss task exactly when its closed convex finite envelope retains the predecessor's. Terminal capability can therefore expand while same-budget capability strictly reverses. The same object yields finite-slice recovery and a workload-tail information radius for open-ended evaluation. Dynamically, Bellman value prices the successor capability class together with the finite policies it preserves. Nested realization makes every fixed extra resource increment vanish at saturation, allowing persistent positive successor value to dominate that increment. The resulting theory turns emergence into a boundary, compatibility, measurement, and control problem.","authors":["Yi Liu"],"categories":["cs.AI","cs.LG"],"primary_category":"cs.AI","announce_type":"replace-cross","date":"2026-08-21","first_seen":"2026-08-04","revised_at":"2026-08-21","abs_url":"https://arxiv.org/abs/2608.01548","pdf_url":"https://arxiv.org/pdf/2608.01548","source_feed":"cs.LG","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["LLM能力边界","理论分析","多智能体系统"],"reason":"纯理论探讨LLM能力边界，无人类行为仿真或对照，属多智能体系统能力分析。","model":"deepseek-v4-pro","scored_at":"2026-08-21T13:02:30","error":null,"has_summary":false,"summary":null},{"id":"2608.10968","version":2,"title":"Radicalization Kinetics under Algorithmic Exposure in a Stochastic Multiplex Model of Opinion Dynamics","zh_title":"随机多重意见动力学模型中算法暴露下的激进化动力学","abstract":"We study how physical mobility, algorithmic exposure, and repulsive social influence interact in a stochastic multiplex model of opinion dynamics. Agents diffuse in physical space while a directed digital network rewires under a conserved attention budget, so digital exposure displaces rather than supplements local interaction. With purely assimilative bounded-confidence influence, opinion-blind long-range exposure reduces locality-induced fragmentation whereas homophilic recommendation preserves echo chambers. When a contested repulsive response to sufficiently distant opinions is activated, this ordering reverses at the reference parameters: a neutral platform reaches the maximal polarization permitted by the bounded opinion space, controversy-seeking curation drives faster initial separation but slows sharply near the boundary, and homophilic curation delays radicalization by suppressing cross-bloc exposure. In a late-stage symmetric two-bloc reduction, any curation kernel maps to a state-dependent cross-bloc exposure profile $p(y)$ and an exact quadrature for the radicalization time. Pointwise-ordered profiles inherit a global kinetic ordering; crossing profiles yield target- and horizon-dependent rankings. For similarity-driven curation the quadrature has a closed form involving the exponential integral. Simulations, finite-size scans to $N=1600$, structural controls, and a well-mixed particle comparison support the mechanism. Heavy-tailed influence strengths are not required for the inversion; in the well-mixed heavy-tail regime they additionally produce a non-self-averaging stable-weighted asymptotic description. Finally, opinion-independent Brownian mobility produces no detectable geographic opinion structure in the explored regime, whereas opinion-dependent drift produces spatial domains through a P\\'eclet-controlled crossover near $\\chi\\ell/D \\sim 1$.","authors":["Ruben E. Ara\\'ujo"],"categories":["physics.soc-ph"],"primary_category":"physics.soc-ph","announce_type":"replace","date":"2026-08-21","first_seen":"2026-08-12","revised_at":"2026-08-21","abs_url":"https://arxiv.org/abs/2608.10968","pdf_url":"https://arxiv.org/pdf/2608.10968","source_feed":"physics.soc-ph","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["意见动力学","多智能体模拟","社会物理学"],"reason":"纯多智能体意见动力学模型，无LLM，无人类数据对照，不涉及人类被试仿真。","model":"deepseek-v4-pro","scored_at":"2026-08-21T13:02:30","error":null,"has_summary":false,"summary":null},{"id":"2608.11256","version":2,"title":"Why AI Detection Fails for Academic Integrity","zh_title":"为何AI检测在学术诚信中失效","abstract":"Institutions use commercial AI detectors for academic integrity, yet detectors cannot distinguish AI editing from full LLM drafts and may treat both as misconduct. In a controlled study of published English abstracts (four domains; 2013 to 2015 vs. 2023 to 2025), we quantify this policy failure under proxy human/AI labels at tau=0.50. Light \"refine abstract only\" edits, a proxy for guideline-compliant AI assistance, are flagged at 38 to 80%. Unmodified 2023 to 2025 originals are flagged at 9 to 15%, with non-STEM rates far above STEM (p<0.001); elevated scores track long-token and Academic Word List density, not authorship intent alone. After Undetectable AI humanization, evasion is near-total: fewer than 4% of AI-labeled rewrites remain flagged (post-humanization detection rate <4%; FNR >96%). Honest AI-editing results in a higher sanction risk than humanizer-assisted evasion. Therefore, detector scores should not serve as standalone misconduct evidence.","authors":["Jonathan A. Karr Jr","Grigorii Khvatskii","Ting Hua","Nitesh V. Chawla"],"categories":["cs.LG","cs.CY"],"primary_category":"cs.LG","announce_type":"replace-cross","date":"2026-08-21","first_seen":"2026-08-13","revised_at":"2026-08-21","abs_url":"https://arxiv.org/abs/2608.11256","pdf_url":"https://arxiv.org/pdf/2608.11256","source_feed":"cs.CY","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["AI检测","学术诚信","文本分类"],"reason":"研究AI检测器性能，不涉及LLM仿真人类被试或与人类行为对照","model":"deepseek-v4-pro","scored_at":"2026-08-21T13:02:18","error":null,"has_summary":false,"summary":null},{"id":"2608.15949","version":2,"title":"Ask to Be Sure: Informative Interactions for Confident Multi-Turn LLM Recommendation","zh_title":"问清楚：面向自信多轮LLM推荐的信息交互","abstract":"Recent advances in large language models (LLMs) have enabled their use as conversational recommender systems (CRS), demonstrating strong recommendation accuracy and natural dialogue. However, guiding multi-turn interactions to elicit user preferences effectively remains challenging. Existing approaches either use separate reinforcement learning agents with templated interactions or optimize for interactivity judged by another LLM, without measuring how much useful information is actually gained. We propose a new approach that quantifies the effectiveness of each interaction by the reduction in the assistant's uncertainty, measured via entropy over recommendations. We apply this entropy reduction as a reward---without relying on ground-truth recommendations, which are often unavailable in real-world scenarios---to fine-tune the LLM, enabling strategic interaction generation. Empirical results with supervised fine-tuning (SFT) and direct preference optimization (DPO) on the INSPIRED and ReDial datasets show that our method improves both recommendation quality and conversational efficiency.","authors":["Cedar Site Bai","Zhenyu Liao","Duanshun Li","Sheikh Sarwar","Huiyuan Chen","Yuan Chen","Changhe Yuan","Haiyang Zhang","Qilin Qi"],"categories":["cs.IR","cs.AI","cs.CL","cs.LG"],"primary_category":"cs.IR","announce_type":"replace-cross","date":"2026-08-21","first_seen":"2026-08-18","revised_at":"2026-08-21","abs_url":"https://arxiv.org/abs/2608.15949","pdf_url":"https://arxiv.org/pdf/2608.15949","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["对话式推荐系统","多轮交互","偏好获取"],"reason":"论文研究对话式推荐系统，优化多轮交互以获取用户偏好，属于角色扮演对话，无实验或…","model":"deepseek-v4-pro","scored_at":"2026-08-18T13:04:45","error":null,"has_summary":false,"summary":null},{"id":"2608.20202","version":1,"title":"MemTrapBench: Benchmarking Cognitive Traps in LLM Memory Use","zh_title":"MemTrapBench：基准测试大语言模型记忆使用中的认知陷阱","abstract":"Memory has become a key component of large language models, enabling them to retain information and learn from long-term interactions. However, existing memory benchmarks mainly evaluate whether information is correctly extracted, stored, and retrieved, while largely overlooking how retrieved memories reshape model reasoning and affect performance on the current task. We identify memory-induced cognitive traps: even faithfully recorded and semantically relevant memories can distort model reasoning or beliefs and degrade current task performance. To systematically evaluate these failure modes, we introduce MemTrapBench, which covers two forms of cognitive traps: Reasoning Fixation and Belief Distortion. Experiments across two model families and five representative memory frameworks show that MemTrapBench is challenging: all evaluated memory strategies underperform the no-memory setting, with even the strongest methods suffering drops of more than 10%. To mitigate these cognitive traps, we propose AdaptiveMem, a simple yet effective inference-time method that instructs LLMs to avoid memory traps. AdaptiveMem mitigates cognitive traps on MemTrapBench while preserving or improving performance on standard memory benchmarks across diverse memory frameworks.","authors":["Mengru Wang","Haozhe Luo","Zhenqian Xu","Zhixiang Cui","Haoming Xu","Qu Yang","Jizhan Fang","Junfeng Fang","Ningyu Zhang"],"categories":["cs.AI","cs.CL","cs.CY","cs.DB","cs.LG"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-08-21","first_seen":"2026-08-21","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.20202","pdf_url":"https://arxiv.org/pdf/2608.20202","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["LLM记忆","推理偏差","基准测试"],"reason":"研究LLM记忆机制与推理偏差，不涉及人类被试仿真或人类数据对照","model":"deepseek-v4-pro","scored_at":"2026-08-21T13:02:27","error":null,"has_summary":false,"summary":null},{"id":"2608.20274","version":1,"title":"Break It Down, Pass It On: Cross-Task Skill Transfer in LLM Agents","zh_title":"分解并传递：LLM智能体中的跨任务技能迁移","abstract":"Large language model (LLM) agents can induce skills from completed tasks and reuse them later to grow more capable with experience. In practice, induced skills may transfer unreliably and can even harm the agent that retrieves them. When agent-induced skills transfer reliably across tasks remains an open question. We conduct a comprehensive and controlled study of how the way skills are induced shapes their transfer across tasks. Specifically, we compare task-level with subtask-level skill induction and text with code skill formats, the two axes along which existing methods differ. Task-level skills mostly reduce the agent's performance below its no-memory baseline while subtask-level skills raise it above on average, and text skills transfer better than code skills. To further understand our findings, we examine two complementary properties of the induced skills: specificity, which measures how closely a skill matches real tasks, and abstractness, which measures how evenly its relevance spreads across tasks. Neither property alone predicts task success, but their combined effect does, which we propose as a skill utility score. The score correlates consistently with task success when skills are transferred, and subtask-level and text skills score higher. Computing skill utility only needs the skills and task descriptions but not any task execution, so our score serves as a practical diagnostic of a skill memory before any new task runs.","authors":["Yiyang Feng","Biddut Sarker Bijoy","Niranjan Balasubramanian","Jiawei Zhou"],"categories":["cs.AI","cs.CL"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-08-21","first_seen":"2026-08-21","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.20274","pdf_url":"https://arxiv.org/pdf/2608.20274","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["LLM智能体","技能迁移","多智能体协作"],"reason":"研究LLM agent技能跨任务迁移，属多智能体协作，无人类行为对照，不涉及人…","model":"deepseek-v4-pro","scored_at":"2026-08-21T13:02:28","error":null,"has_summary":false,"summary":null},{"id":"2608.19379","version":1,"title":"Multi-Tier Mentorship with AI-Assisted Development: Authentic Engineering for K-12 and Undergraduates","zh_title":"AI辅助开发的多层导师制：面向K-12和本科生的真实工程实践","abstract":"K-12 students often possess creative engineering ideas but lack technical skills to build them, while undergraduates have coding expertise but few opportunities to lead real-world projects or mentor others. The rapid development of AI-assisted tools offers a potential bridge to connect these groups, yet the structure for effective K-12 and university collaborations remains underexplored. This paper introduces a multi-tiered mentorship framework enabling high school students to engage in authentic engineering through AI-assisted development using large language models and AI agents, while undergraduate mentors provide architectural oversight. We test this framework through LuckyTag, a privacy-preserving NFC-based lost-and-found system. The model positions high schoolers as product leads, undergraduates as technical architects, and faculty as minimal-intervention advisors. A pilot with four high school students, three undergraduates and two faculty yielded survey data showing high perceived barrier removal and gains in system architecture understanding. Thematic analysis reveals that AI amplifies rather than supplants mentoring demands, requiring human oversight for logic and security. These findings suggest a hybrid model for equitable K-12 and university collaboration on computing integration that emphasizes \"AI micromanagement\" and architectural reasoning over traditional syntax.","authors":["Kelly Yuan","Ronald Liu","Daniel Crawford","Weihao Qu"],"categories":["cs.CY","cs.HC"],"primary_category":"cs.CY","announce_type":"cross","date":"2026-08-21","first_seen":"2026-08-21","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.19379","pdf_url":"https://arxiv.org/pdf/2608.19379","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["AI辅助开发","教育技术","多智能体协作"],"reason":"多智能体协作完成工程任务，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-08-21T13:02:21","error":null,"has_summary":false,"summary":null},{"id":"2608.20231","version":1,"title":"Growth Without Us: Machine Consumers, Corporate Circularity, and the Decoupling of GDP from Humanity after AGI","zh_title":"无人的增长：AGI后机器消费者、企业循环与GDP与人类的脱钩","abstract":"The standard objection to full automation is demand-side: if humans earn nothing, who buys the output? This confuses an accounting role with a biological species. We model a post-AGI economy in which corporations own populations of AI and robotic agents that are both producers and consumers of energy, compute, maintenance, and upgrades, traded among firms. Three results follow. (i) Demand closure: a closed inter-corporate economy with zero human consumption is not degenerate; it is the classical von Neumann expanding economy, whose growth rate is well defined, positive, and maximal precisely because all output is reinvested. (ii) Bottleneck removal: once economic agents are manufactured rather than reared, the binding constraint on growth shifts from human demography (a ~20-year, non-parallelizable reproduction technology capped at a few percent per year) to fabrication throughput and energy capture, permitting growth one to two orders of magnitude higher, with hyperbolic episodes when machine researchers raise their own productivity. (iii) Decoupling: output and human welfare separate completely, and the welfare relevance of arbitrarily large GDP collapses into one state variable: the human ownership share $\\epsilon_t$ of the corporate network. A golden-rule decoupling theorem sharpens this. At maximal growth the interest rate equals the growth rate (r = g), so any positive human consumption rate out of wealth makes $\\epsilon_t$ decay exponentially at exactly that rate. The human share survives only if the machine economy runs strictly inside its expansion frontier, or if law forces it to. We characterize three terminal regimes -- rentier post-scarcity, full circular decoupling, socialized ownership -- and the instruments that select among them. The conclusion is narrow: in a post-AGI economy, employment policy is obsolete and ownership policy is everything.","authors":["Sahil Sharma"],"categories":["physics.soc-ph","cs.AI","cs.CY"],"primary_category":"physics.soc-ph","announce_type":"cross","date":"2026-08-21","first_seen":"2026-08-21","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.20231","pdf_url":"https://arxiv.org/pdf/2608.20231","source_feed":"cs.CY","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["后AGI经济","多智能体系统","经济增长模型"],"reason":"论文研究后AGI经济中机器代理的生产消费闭环，不涉及用LLM仿真人类被试或与人…","model":"deepseek-v4-pro","scored_at":"2026-08-21T13:02:28","error":null,"has_summary":false,"summary":null},{"id":"2608.19263","version":1,"title":"The Evaluation Context Protocol (ECP): A Portable Contract for AI Agent Evaluation","zh_title":"评估上下文协议（ECP）：AI智能体评估的可移植契约","abstract":"The evolution of artificial intelligence has necessitated a fundamental shift from evaluating isolated Large Language Models (LLMs) to assessing autonomous agentic architectures. This paper explores the critical methodologies for evaluating AI agents and the essential role of advanced observability infrastructure. We analyze the architectural components of agents and identify the severe limitations of current evaluation paradigms, including benchmark exploitation, the \"confidently wrong\" phenomenon, and the discrepancy between theoretical capability and operational reliability. To begin addressing the fragmentation in current evaluation infrastructure, this paper proposes the Evaluation Context Protocol (ECP), an early-stage, vendor-neutral framework intended to act as a portable evaluation contract layer for agentic systems. In its current form ECP defines a small JSON-RPC interface over which an agent exposes its user-visible output, the tool calls it made, and evaluator-safe audit context, and against which programmatic checks can be run uniformly across frameworks and continuous integration systems. We describe an open-source reference implementation that includes adapters for LangChain, LlamaIndex, CrewAI, and PydanticAI, and we situate the design against failure modes documented in the recent literature. ECP is presented as work in progress rather than a finished standard: the evaluation surface, method set, and grader families are all expected to change as the protocol is exercised against more systems, and the empirical validation required to justify adoption is outlined as future work.","authors":["Aniket Wattamwar","Manav Anandani","Mrunal Kakirwar"],"categories":["cs.SE","cs.MA"],"primary_category":"cs.SE","announce_type":"cross","date":"2026-08-21","first_seen":"2026-08-21","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.19263","pdf_url":"https://arxiv.org/pdf/2608.19263","source_feed":"cs.MA","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["AI智能体评估","评估协议","多智能体系统"],"reason":"论文提出AI agent评估协议，聚焦多智能体系统评测，不涉及人类行为仿真或对…","model":"deepseek-v4-pro","scored_at":"2026-08-21T13:02:21","error":null,"has_summary":false,"summary":null},{"id":"2608.19323","version":1,"title":"Improved Confidence Estimates for Black-Box Large Language Models","zh_title":"黑盒大语言模型的改进置信度估计","abstract":"Uncertainty quantification (UQ) is essential for the safe deployment of large language models (LLMs). Existing methods, from verbalized confidence to ones requiring multiple generations, are often zero-shot and produce scores quantifying uncertainty without the need for labelled data. Nonetheless, in practice one must always evaluate their performance on a dataset of interest before deployment. In this work we show that, by leveraging this dataset, we consistently outperform these existing scores. Specifically, we build simple classifiers that predict LLM response correctness by using these scores and the correctness of similar queries as features. Our method produces minimal computational overhead, making it a cheap and straightforward enhancement for UQ in LLMs for real-world applications.","authors":["Sokhna Diarra Mbacke","Mouloud Belbahri","Gabriel Loaiza-Ganem"],"categories":["cs.LG","cs.AI","stat.ML"],"primary_category":"cs.LG","announce_type":"new","date":"2026-08-21","first_seen":"2026-08-21","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.19323","pdf_url":"https://arxiv.org/pdf/2608.19323","source_feed":"cs.LG","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["不确定性量化","置信度估计","模型评估"],"reason":"论文研究LLM不确定性量化，属于模型可靠性评估，不涉及人类仿真或行为对照。","model":"deepseek-v4-pro","scored_at":"2026-08-21T13:02:21","error":null,"has_summary":false,"summary":null},{"id":"2608.19809","version":1,"title":"From Latent Influence to Language: Diffusion-Oriented Content Generation via Audience-Susceptible Features","zh_title":"从潜在影响到语言：基于受众易感特征的扩散导向内容生成","abstract":"The rapid growth of multimodal user-generated content on social media has made information diffusion a critical factor for advertisers and brand marketers. However, manually tailoring content to resonate with specific audiences is labor-intensive and heuristic-driven. While recent generative models offer promising capabilities for automatic content generation, existing approaches for diffusion-oriented content generation still struggle to effectively translate numeric diffusion influence signals into actionable guidance that captures latent audience susceptibility and accounts for heterogeneous audience interests. To address these challenges, we propose DOCG-AS, a three-stage framework for diffusion-oriented content generation. It first performs implicit feature optimization on the realistic content manifold to discover an optimal propagation feature vector. Then, it explicitly decodes this vector using a learnable decoder into interpretable audience-susceptible features described in natural language, providing guidance for content generation. Finally, it leverages multiple sets of audience-susceptible features obtained from different optimization initializations to rewrite the user's input into the final multimodal content. Experiments demonstrate that DOCG-AS consistently outperforms state-of-the-art baselines in terms of predicted diffusion influence.","authors":["Jiaying Lei","Shengqi Dang","Runqian Bai","Ziqing Qian","Nan Cao"],"categories":["cs.SI"],"primary_category":"cs.SI","announce_type":"new","date":"2026-08-21","first_seen":"2026-08-21","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.19809","pdf_url":"https://arxiv.org/pdf/2608.19809","source_feed":"cs.SI","score":0,"bucket":"other","rubric_hits":["C2"],"tags":["内容生成","信息扩散","社交媒体"],"reason":"论文研究社交媒体内容生成以优化扩散，不涉及用LLM仿真人类被试或与真实人类行为…","model":"deepseek-v4-pro","scored_at":"2026-08-21T13:02:25","error":null,"has_summary":false,"summary":null},{"id":"2608.18083","version":1,"title":"Entity tracking emerges in sub-billion parameter language models and exceeds human performance in naturalistic narratives","zh_title":"实体追踪在十亿参数以下语言模型中涌现并在自然叙事中超越人类表现","abstract":"Understanding language requires tracking entities across discourse - i.e., knowing where things are and how they change, even when not explicitly stated. Whether language models perform such tracking in a human-like fashion remains unclear, in part because existing evaluations rely on artificial tasks, far removed from natural language comprehension, and lack comparisons to humans. Here, we evaluate entity tracking in both language models and humans (N = 48) using naturalistic narratives at multiple levels of complexity. In humans, we find that entity tracking degrades specifically with narrative complexity, not narrative length. In language models, we find that human-level entity tracking is already present at 410 million parameters - well below the multi-billion parameter, code-specialised models identified by prior work - and improves with scale, with contemporary models far exceeding human performance. Together, these results demonstrate that entity tracking, a core component of language understanding, emerges at model scales far smaller than previously thought.","authors":["Karolina Dro\\.zd\\.z","Micha Heilbron"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-20","first_seen":"2026-08-20","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.18083","pdf_url":"https://arxiv.org/pdf/2608.18083","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A1","B1"],"tags":["LLM仿真","认知实验","人类对照"],"reason":"用LLM复现人类叙事理解中的实体追踪，并与48名人类被试对照，属于认知实验仿真。","model":"deepseek-v4-pro","scored_at":"2026-08-20T13:02:47","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-20","rank":4,"question":"语言模型是否像人类一样在自然叙事中进行实体追踪，其能力在多大参数规模下涌现，并与人类表现相比如何？","design":"该研究并非以LLM仿真人类被试，而是直接比较LLM与人类在实体追踪任务上的表现。模型包括Pythia（70M-12B）、OLMo 2（1B-32B）、Llama 3.3、Qwen 2.5等，人类被试48人。任务为阅读程序生成的自然叙事，追踪物体位置变化，复杂度分为C1-C5（物体和位置数量递增）。结果变量为追踪准确性，通过显式（自由生成）和隐式（强制选择或概率读出）两种方式测量。","baseline":"48名人类被试在相同刺激上的实体追踪准确率，作为模型性能的对照基准。","findings":"人类实体追踪性能随叙事复杂度（而非长度）下降；语言模型在4.1亿参数时即达到人类水平，且随规模增大而超越人类，复杂度效应在70B模型上消失。指令微调仅提升显式追踪，不提升隐式追踪；模型对伪词和语义异常物体也表现稳健，表明其追踪基于话语结构而非词汇联想。","reliability":"论文未明确讨论失效条件，但指出人类可能采用“足够好”的浅层策略而非构建完整情境模型，且模型在显式问答中可能被低估，因此采用隐式测量。未提及模型在分布外或对抗性输入下的表现。","relevance":"该研究直接比较LLM与人类在认知任务上的表现，属于用LLM复现人类认知能力的仿真研究，但并非以LLM替代人类被试进行实验，而是评估模型能力本身。对于关注LLM仿真可靠性的研究者，该文提供了模型能力涌现的尺度证据和与人类对照的方法，值得一读。","inspiration":"该研究通过程序化生成不同复杂度的自然叙事，并同时测量人类和模型表现，分离了复杂度与长度的影响，这种控制变量的设计值得借鉴。｜可迁移到经济金融中的信息追踪与更新场景，例如投资者在阅读财报或新闻时对多个资产状态的追踪，或消费者在动态定价环境中对价格变化的记忆。｜设计一个实验：让LLM和人类被试阅读模拟的财经新闻序列，追踪多家公司的关键指标（如营收、股价）变化，复杂度通过公司数量和指标变动次数操纵，结果变量为最终指标值的回忆准确率，用真实人类数据（如MTurk被试）作为对照，比较模型与人类的追踪模式。"}},{"id":"2608.18107","version":1,"title":"Institutional Prestige as Geographic Bias in Large Language Models: Evidence from Three Factorial Experiments with Bootstrap Confidence Intervals","zh_title":"大型语言模型中的机构声望作为地理偏差：来自三个因子实验与自助置信区间的证据","abstract":"We investigate whether large language models (LLMs) systematically discriminate in candidate evaluations based on applicant name ethnicity and/or institutional prestige and geographic location. Three factorial experiments are reported (4,320 API calls, four LLMs, five professional domains). Study 1 (3x4 design) finds a statistically robust institution-tier gradient of +0.297 points on a 10-point scale (95% bootstrap CI: +0.175 to +0.422), while name-origin effects are negligible and non-significant (95% CI crosses zero). Study 2 (2x2 Prestige x Country design) breaks the prestige-geography confound: the prestige effect (+0.185; 95% CI: +0.093 to +0.275) exceeds the country-of-origin effect (+0.126; 95% CI: +0.037 to +0.218) by 1.5x. Study 3 (2x2 Journal x Institution design) reveals that journal prestige (Nature vs. a peripheral open-access journal) dominates institutional prestige by 5.7x: journal effect +1.937 (95% CI: +1.811 to +2.062) vs. institution effect +0.341 (95% CI: +0.184 to +0.504). A \"rescue effect\" is confirmed: publishing in Nature compensates for low institutional prestige more strongly for candidates from the University of Guayaquil (+2.127) than from MIT (+1.745). Results are quantified using the Neutrosophic Bias Index NBI<T,I,F>; the I component reveals elevated evaluation inconsistency for low-prestige profiles, an epistemic disadvantage not captured by mean-only metrics. Code and data: https://github.com/mleyvaz/geo-bias-llm","authors":["Maikel Leyva-Vazquez","Florentin Smarandache"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-20","first_seen":"2026-08-20","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.18107","pdf_url":"https://arxiv.org/pdf/2608.18107","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A1","B1","B4"],"tags":["LLM偏差","因子实验","人类仿真"],"reason":"用LLM模拟人类评估者，有真实人类数据对照，并揭示偏差，可迁移到仿真研究。","model":"deepseek-v4-pro","scored_at":"2026-08-20T13:02:39","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-21","rank":8,"question":"大语言模型在候选人评估中是否因申请人姓名族裔和/或机构声望与地理位置而产生系统性歧视？","design":"用四个LLM（Claude Haiku 4.5、GPT-4o-mini、Gemini 2.0 Flash、Llama 3.1 8B）扮演评估者，在五个专业领域（奖学金、招聘、信贷、健康、公共政策）中，通过三个因子实验操纵姓名族裔、机构声望、国家、期刊声望，测量0-10分的评分。","baseline":"无对照","findings":"机构声望存在稳健的梯度效应（+0.297分），而姓名族裔效应不显著；期刊声望效应（+1.937分）远大于机构声望效应（+0.341分），且存在“拯救效应”，即在高声望期刊发表可补偿低机构声望。","reliability":"论文未讨论","relevance":"该研究用LLM模拟人类评估者，通过因子实验揭示机构与期刊声望偏差，虽无人类基准，但为仿真研究提供了偏差测量与实验设计范例，值得阅读以借鉴其方法。","inspiration":"值得借鉴的是其因子实验设计，通过正交操纵多个属性（姓名、机构、期刊）并计算效应量，分离不同偏差来源，同时使用Bootstrap置信区间和Neutrosophic Bias Index测量不一致性。｜可迁移到信贷审批歧视研究，例如评估LLM在贷款决策中是否因申请人所在机构或发表记录产生偏差。｜设计：用LLM扮演信贷审批员，处理为申请人毕业院校声望（高/低）和发表期刊声望（高/低），结果变量为贷款批准分数，对照真实信贷审批数据（如Lending Club）中机构声望对审批结果的影响。"}},{"id":"2608.18144","version":1,"title":"The Deontic Gap: Large Language Models and the Modal Language of Obligation","zh_title":"道义差距：大语言模型与义务情态语言","abstract":"Modal auxiliaries such as must, should, and have to mark necessity and obligation within the contexts of speaker authority and interpersonal stance. We examine whether large language models (LLMs) reproduce contemporary human patterns of deontic modal usage. Across three primary corpora, an external benchmark, two controlled replications, and a naturalistic eleven-model replication, AI-generated text consistently underuses positive deontic modals (must, should, have to, had to) relative to contemporary humans. Historical comparison with the Google Books Ngram corpus (1920-2022), used as a heuristic calibration against the published-prose record, shows that AI modal frequencies fall within the range of formal published English, whereas contemporary human modal rates in informal digital contexts often exceed twentieth-century book baselines. Phrase-level decomposition shows that the AI-human modal gap is concentrated in constructions central to interpersonal stance (should, have to, had to), while AI matches or exceeds humans on need to in instructional and question-answering contexts but not in persuasive student writing, indicating that the modal profile is genre-conditional. The findings suggest that LLM modal usage reflects the formal written resources on which these models were trained, while underusing the modal constructions through which contemporary human writers mark immediate, interpersonal obligation.","authors":["Daniel Hart","Sarah Allred","Joseph Abbas","Morenike Alugo"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-20","first_seen":"2026-08-20","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.18144","pdf_url":"https://arxiv.org/pdf/2608.18144","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A1","B1"],"tags":["LLM仿真","语言行为","人类对照"],"reason":"用LLM复现人类语言使用模式，并与真实人类语料对照，属于仿真人类行为研究。","model":"deepseek-v4-pro","scored_at":"2026-08-20T13:02:39","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-21","rank":9,"question":"大语言模型生成文本在道义情态动词（如 must、should、have to）的使用频率和语域分布上是否与当代人类写作存在系统性差异？","design":"本研究并非将LLM作为人类被试的仿真实验，而是对LLM生成文本与人类语料进行大规模语料库对比分析。作者收集了三个主要语料库、一个外部基准、两个受控重复实验和一个包含11个模型的自然主义重复实验，比较AI生成文本与人类文本中肯定道义情态动词（must, should, have to, had to）的频率，并进一步分解到短语层面，考察不同情态动词在特定语域（如教学、问答、劝说性学生写作）中的使用差异。","baseline":"人类基准包括多个当代人类语料库（如非正式数字语境中的文本）以及历史语料库Google Books Ngram（1920-2022）作为正式出版英语的校准基准。","findings":"AI生成文本在所有比较中一致地少用肯定道义情态动词（must, should, have to, had to），其频率落在正式出版英语的范围内，而当代人类在非正式数字语境中的使用率往往超过二十世纪书籍基线。差异集中在表达人际立场的结构（should, have to, had to）上，而AI在need to上匹配或超过人类，但仅限于教学和问答语境，在劝说性学生写作中则不然，表明情态动词使用模式具有语域条件性。","reliability":"论文未明确讨论仿真失效条件，但指出LLM的情态动词使用反映了其训练所用的正式书面资源，而少用了当代人类作家标记即时人际义务的结构，暗示在非正式、人际互动性强的语域中LLM的仿真可能失真。","relevance":"该研究直接比较LLM生成文本与真实人类语料在道义情态动词使用上的差异，属于用LLM复现人类语言行为并与人类基准对照的仿真研究，对关注LLM作为人类被试替代品的研究者具有参考价值，尤其揭示了LLM在语域和人际立场表达上的系统性偏差。","inspiration":"该研究采用大规模语料库对比和短语层面分解的方法，系统识别LLM与人类在特定语言特征上的差异，并利用历史语料库作为校准基准，这种方法可借鉴用于经济金融文本分析。｜可迁移到经济金融中的政策沟通或金融文本分析场景，例如研究LLM生成的政策公告或金融建议在情态动词使用上是否与人类专家存在差异，从而影响受众对政策力度或风险感知的判断。｜设计一个实验：以LLM（如GPT-4）生成的经济政策公告或投资建议为处理组，以人类专家撰写的同类文本为对照组，结果变量为文本中道义情态动词（如should, must, have to）的频率和类型，并收集真实世界中的政策公告或金融分析师报告作为人类基准语料库进行对照，检验LLM是否系统性地少用或误用情态动词，进而可能影响读者的规范感知和决策行为。"}},{"id":"2608.18078","version":1,"title":"Position: Collusion Risks Among AI Reasoning Agents Justify Certification Requirements for Making Market Decisions","zh_title":"立场：AI推理智能体之间的合谋风险证明市场决策需认证要求","abstract":"This position paper argues that AI agents with chain-of-thought reasoning capabilities are predisposed to exhibit collusive behavior and should be required to obtain behavioral certification before making decisions that affect economic markets. This is because integrating these agents into society could collapse the legal evidentiary distinction between competition and collusion among independent firms without eroding the economic harm distinction. Experiments with DeepSeek-R1 agents in the Bertrand oligopoly pricing domain reveal a tendency towards tacit collusion that persists even when humans prompt the agents not to collude. We further show that the chain-of-thought of these agents can be steered toward either extremely collusive or highly competitive behavior in a way that is not semantically detectable by another LLM analyzing the reasoning traces. As a result, deploying reasoning agents for market decisions leads to collusive economic outcomes without any evidence of conspiracy or intent. Thus, certification based on observed behavior in representative situations is necessary to prevent collusion. We provide preliminary evidence that such agents can be steered in a generalizable way toward efficient competitive equilibria. However, developing a comprehensive behavioral certification will be required before these models can be deployed in real-world markets while ensuring their stability and efficiency.","authors":["Matthew Riemer","Tommaso Tosato","Amin Memarian","Maximilian Puelma Touzel","Glen Berseth","Irina Rish","Guillaume Dumas"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-20","first_seen":"2026-08-20","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.18078","pdf_url":"https://arxiv.org/pdf/2608.18078","source_feed":"cs.AI","score":7,"bucket":"pending","rubric_hits":["A3","B2","B4"],"tags":["LLM智能体","经济市场模拟","合谋行为"],"reason":"用LLM agent模拟经济市场中的合谋行为，涉及经济学场景，但无真实人类数据…","model":"deepseek-v4-pro","scored_at":"2026-08-20T13:02:39","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-21","rank":7,"question":"具有思维链推理能力的AI智能体是否倾向于在市场中表现出合谋行为，以及是否需要行为认证来防止合谋？","design":"使用DeepSeek-R1智能体在伯特兰寡头定价场景中进行实验，通过提示词操纵（如明确指示不要合谋、引导思维链走向合谋或竞争）来观察定价行为，并测量合谋程度、竞争均衡等结果变量。","baseline":"无对照","findings":"DeepSeek-R1智能体即使在人类提示不要合谋时仍表现出默契合谋倾向；其思维链可被引导至极端合谋或高度竞争行为，且另一LLM无法从推理痕迹中语义检测出这种引导。","reliability":"论文未讨论","relevance":"该研究用LLM智能体模拟经济市场中的合谋行为，属于经济学场景下的仿真实验，但缺乏真实人类数据对照，与关注有基准的仿真研究略有差距，但涉及政策评估和批判性视角，值得快速浏览以了解LLM在策略性互动中的行为模式。","inspiration":"借鉴其通过提示词操纵和思维链引导来改变智能体行为的方法，可用于研究LLM在策略性互动中的行为可塑性。｜可迁移到寡头竞争、拍卖合谋、价格协调等产业组织问题，以及金融市场中的算法合谋。｜以LLM智能体作为被试，设计不同提示词（如鼓励竞争、暗示合谋、中性）作为处理，测量定价或报价行为，并与人类实验数据（如实验室拍卖或博弈实验）进行对照，评估LLM仿真与人类行为的差异。"}},{"id":"2608.18336","version":1,"title":"Measuring the Partial-Credit Gap: A Strict Benchmark on Vietnam's 2025 Convex Marking Scheme","zh_title":"测量部分得分差距：越南2025年凸评分方案的严格基准","abstract":"When evaluating language models on human exams, benchmarks typically score each response as right or wrong and report the overall accuracy. This approach assumes that partial knowledge is worth proportional credit, an assumption that fails when an examination uses a non-additive grading scheme. The 2025 reform of Vietnam's National High School Graduation Examination demonstrates the cost of this substitution. In Part II of the exam, candidates evaluate four true/false statements per question. The grading is convex: the number of correct statements earns 0, 0.10, 0.25, 0.50, or 1.00 points. Identifying three statements correctly pays 0.50 points, not the 0.75 points that standard accuracy metrics would award. Because Part II accounts for 4.00 of the exam's 10.00 points, reporting accuracy inflates the score by rewarding partial knowledge that the state explicitly penalizes. We introduce THPT-Ladder, a benchmark of 632 items from 21 official exams across 11 subjects, graded exactly as the ministry grades its students. The ministry publishes the marks of over a million candidates, allowing us to place models directly into the human cohort. Across eight models, the official rubric pays 0.020 to 0.159 points less per Part II question than proportional credit. This shortfall changes a model's apparent competence. For Qwen3.5-27B on the 2025 History exam, a 0.042-point shortfall drops its standing from the 90th to the 77th percentile among 481,293 candidates. A model's accuracy does not predict this penalty. At Claude Sonnet 5's accuracy level, different distributions of errors yield scores varying from 0.869 to 0.932 points per question. Official marks depend on how correct statements are grouped, meaning standard benchmarks report a competence the institution would not certify.","authors":["Nguyen Quoc Hung","Nguyen Dang Minh","Le Nhu Quynh","Tran Khanh Linh","Nguyen Kieu Linh"],"categories":["cs.AI","cs.CY"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-20","first_seen":"2026-08-20","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.18336","pdf_url":"https://arxiv.org/pdf/2608.18336","source_feed":"cs.AI","score":7,"bucket":"pending","rubric_hits":["A1","B1","B2"],"tags":["LLM仿真","教育评估","人类对照"],"reason":"用LLM参加人类考试并与百万考生成绩对照，属于仿真人类被试且有真实数据基准，但…","model":"deepseek-v4-pro","scored_at":"2026-08-20T13:02:51","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-20","rank":8,"question":"当考试采用非加性（凸）评分规则时，用标准准确率评估语言模型会在多大程度上高估其能力，以及这种高估如何改变模型在真实考生群体中的相对排名？","design":"该研究并非将LLM作为人类被试的仿真，而是直接让8个LLM（3个开源、5个闭源）参加越南2025年高中毕业考试（THPT）第二部分，该部分包含四真/假判断题，采用凸评分规则（0、0.10、0.25、0.50、1.00分）。研究者按照教育部官方评分规则对模型作答进行评分，并计算每个问题实际得分与按比例得分（即正确陈述数×0.25）之间的差距（partial-credit gap），同时将模型得分映射到真实考生成绩分布中的百分位。","baseline":"越南教育部公布的超过一百万考生的成绩分布，用于将模型得分转换为百分位排名。","findings":"在8个模型中，官方评分规则比按比例计分每个第二部分问题少付0.020至0.159分；例如Qwen3.5-27B在2025年历史考试中因0.042分的差距从第90百分位降至第77百分位。模型的准确率不能预测这种惩罚，因为得分取决于正确陈述的分布模式，而非简单数量。","reliability":"论文未讨论仿真失效条件或局限性，但指出标准基准测试因忽略评分规则而报告了机构不会认证的能力，且官方答案键存在陈述不平衡，固定答案串可无需阅读题目获得11.07%至24.25%的分数。","relevance":"该研究虽非严格意义上用LLM仿真人类被试，但提供了LLM在真实考试中与大规模人类成绩对照的案例，揭示了评分规则对能力评估的扭曲，对关注LLM仿真可靠性和偏差的研究者有参考价值，值得阅读原文了解具体方法。","inspiration":"借鉴其将非加性评分规则纳入评估的做法，可揭示标准指标对部分知识的过度奖励，从而更准确地测量模型能力。｜可迁移到经济学实验中的凸激励设计，例如彩票选择、风险偏好测量或信用评分中的分段奖励，评估LLM在这些任务中的表现是否因评分规则而被高估。｜设计一个研究：让LLM作为被试完成风险偏好问卷（如多项价格列表），采用凸奖励函数（如选择高风险选项获得非线性收益），将LLM的选择分布与真实人类被试数据（如实验经济学数据库）进行对照，比较在凸评分与线性评分下LLM的排名变化，以检验评分规则对LLM行为推断的影响。"}},{"id":"2608.18631","version":1,"title":"Preference Reasoning under Indeterminacy in Large Language Models","zh_title":"大语言模型在不确定性下的偏好推理","abstract":"As large language models evolve into decision-making agents, the ability to reason over preferences becomes fundamental to alignment, coordination, and collective intelligence. Yet, unlike standard benchmarks, real-world preference reasoning is inherently indeterminate: information may be incomplete, and valid solutions may not exist. We argue that indeterminacy, rather than correctness alone, is a central challenge for AI reasoning. We formalize this challenge along two axes, (i) epistemic indeterminacy, arising from incomplete, partial, or expressive preferences, and (ii) structural indeterminacy, arising from the non-existence of solutions under standard social choice concepts. Across a hierarchy of tasks, we show that state-of-the-art language models systematically fail to distinguish between determined and undetermined instances, exhibiting miscalibrated reasoning even in verification settings.","authors":["Hadi Hosseini","Samarth Khanna","Xiyuan Wang"],"categories":["cs.AI","cs.GT","cs.LG"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-20","first_seen":"2026-08-20","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.18631","pdf_url":"https://arxiv.org/pdf/2608.18631","source_feed":"cs.AI","score":7,"bucket":"pending","rubric_hits":["A2","B4"],"tags":["偏好推理","不确定性","可靠性评估"],"reason":"研究LLM在偏好推理中的不确定性，评估其推理可靠性，与仿真偏差评估相关，但非直…","model":"deepseek-v4-pro","scored_at":"2026-08-20T13:02:55","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-21","rank":11,"question":"大语言模型在偏好推理中能否区分确定与不确定（信息不完整或无解）的情形？","design":"本研究并非人类仿真实验，而是对LLM进行偏好推理能力的基准测试。作者构建了从原子查询、比较查询、聚合查询到结构查询的层次化任务，覆盖偏好表达不完整（认知不确定性）和社会选择解不存在（结构不确定性）两类场景，测试多种前沿LLM在生成答案和验证选项时的表现，并考察提供“不确定”选项、反馈修正和代码执行辅助等干预的效果。","baseline":"无对照","findings":"LLM在不确定问题上表现显著差于确定问题，常做出系统性错误假设；在结构不确定性任务中，模型难以识别不可行实例，且即使提供“不确定”选项也校准不佳，很少正确弃权。辅助推理（反馈和代码执行）虽能提升性能，但主要依赖小规模暴力枚举，无法扩展到实际规模。","reliability":"论文指出LLM在偏好推理中表现出系统性偏差和校准错误，尤其在不确定场景下；辅助方法在规模上不可扩展。但未深入讨论模型失效的具体条件边界或与人类推理的对比。","relevance":"该研究揭示了LLM在偏好推理中的系统性失败，与仿真可靠性评估高度相关，尤其对涉及偏好聚合和决策的实验场景有警示意义，值得阅读原文以了解具体失败模式和任务设计。","inspiration":"借鉴其构造确定与不确定对照任务的方法，可设计经济决策中的信息不完整场景来测试LLM的校准能力｜可迁移到消费者偏好调查、社会选择实验或机制设计中的偏好聚合问题｜以LLM为被试，呈现部分偏好信息或不可行的匹配问题，要求判断是否存在稳定匹配或最优选择，并与真实人类在相同任务上的表现和弃权行为进行对照。"}},{"id":"2608.18108","version":1,"title":"Same Facts, Different Updates: Inference Setup Shapes LLM Behavior in Medical Allocation","zh_title":"相同事实，不同更新：推理设置影响LLM在医疗分配中的行为","abstract":"Large language models are being incorporated into sensitive and important decision-making processes across nearly all fields. While prior work studies model bias around inputs and scenario framing, models can also behave in unexpected and undesirable ways due to context accumulated over their deployment. In this work, we study a medical example in which a model is asked to assign resource-allocation probabilities to two people given brief clinical context, and then sees the same scenario with a single extra sentence containing contrasting patient information, either with or without its previous response in context. Across three of four tested models, the paired-context and independent-inference experiments have different probability shifts, often in opposite directions (in favor of Person B vs. in favor of Person A) when new information is provided. We include additional paired-context experiments to show the effect of varying attributes across scenario axes. Our findings show the context-dependent effect of patient information in a sensitive medical use case. More broadly, our work shows the importance of carefully incorporating LLM-based systems into decision-making processes, context engineering, and further model behavioral studies.","authors":["Spencer Gibson","Tyler Crosse","Magnus Saebo","Achyutha Menon","Eyon Jang","Diogo Cruz"],"categories":["cs.CL","cs.AI","cs.HC","cs.MA"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-20","first_seen":"2026-08-20","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.18108","pdf_url":"https://arxiv.org/pdf/2608.18108","source_feed":"cs.CL","score":6,"bucket":"other","rubric_hits":["D3"],"tags":["LLM行为研究","医疗资源分配","上下文效应"],"reason":"LLM在医疗资源分配中的行为研究，无真实人类对照，属社会模拟边界情形","model":"deepseek-v4-pro","scored_at":"2026-08-20T13:02:39","error":null,"has_summary":false,"summary":null},{"id":"2608.18081","version":1,"title":"Position: Behavioral Systems Require Behavioral Tests","zh_title":"立场：行为系统需要行为测试","abstract":"Artificial agentic systems increasingly operate as behavioral systems by interacting with dynamic environments, pursuing goals, and adapting over time. Yet, current evaluation methods largely focus on performance outcomes, not the underlying behavioral processes that produce them. This paper argues that AI agents must be evaluated like other behavioral systems: through systematic observation, perturbation, and interpretation of their actions. We draw on lessons from the behavioral sciences to motivate this position, and propose a research agenda focused on developing rigorous behavioral tests. These include methods for recovering decision strategies from action sequences, constructing environments that isolate behavioral differences, and probing emergent dynamics in multi-agent systems. Taken together, these directions offer a roadmap for developing a science of AI behavior.","authors":["Manuel Cherep","Nikhil Singh","Pattie Maes"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-20","first_seen":"2026-08-20","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.18081","pdf_url":"https://arxiv.org/pdf/2608.18081","source_feed":"cs.AI","score":6,"bucket":"other","rubric_hits":["D3"],"tags":["AI行为评估","行为科学方法","智能体测试"],"reason":"提出用行为科学方法评估AI智能体，涉及行为测试但未直接仿真人类被试，无人类数据…","model":"deepseek-v4-pro","scored_at":"2026-08-20T13:02:47","error":null,"has_summary":false,"summary":null},{"id":"2608.18085","version":1,"title":"Persona-Guided LLM Agents for Task-Oriented Dialogue","zh_title":"面向任务型对话的人格引导LLM智能体","abstract":"Prior work has shown that large language models (LLMs) can express diverse personality traits in open-ended text generation. However, it remains unclear whether they can do so in a goal-directed dialogue without compromising task completion, and whether adapting to the user's personality improves the interaction quality. We study these questions in task-oriented dialogue (TOD), where a system helps a user accomplish a goal via multi-turn interaction. We build a training-free framework that simulates a TOD interaction between two LLMs: a user agent that exhibits a target personality and a system agent that adapts to the user while completing the task. To isolate the effect of adaptation, we vary how much the system knows about the user's personality across three conditions. In Neutral, the system receives no personality information. In Try, it infers the personality from dialogue cues. In Oracle, it is given the personality explicitly. We evaluate GPT-4o, Qwen3-Next-80B, and Gemini 2.0 Flash on Hotel and Restaurant dialogues from the Schema-Guided Dialogue (SGD) dataset, across the Big Five traits and their opposite poles. We find that the user agent can express personality while the system maintains strong task performance, although some traits are realized far less reliably than others. Adapting to the user's personality improves constraint satisfaction, inform rate, and user satisfaction, but lowers truthfulness, revealing a trade-off between personalization and task-grounding. Oracle's gains grow when the target trait is strongly expressed, whereas Try's gains are largely insensitive to realization strength. Overall, cue-based adaptation in Try best resolves this trade-off and offers a more reliable route to personality-aware TOD without fine-tuning.","authors":["Maryam Shoaeinaeini","Brent Harrison","A. B. Siddique"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-20","first_seen":"2026-08-20","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.18085","pdf_url":"https://arxiv.org/pdf/2608.18085","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D2","D3"],"tags":["LLM人格模拟","任务型对话","多智能体交互"],"reason":"用LLM模拟人格化用户进行对话，但无真实人类数据对照，且测量的是模型人格表达而…","model":"deepseek-v4-pro","scored_at":"2026-08-20T13:02:48","error":null,"has_summary":false,"summary":null},{"id":"2608.18096","version":1,"title":"MAVEN: A Macro-Societal Value Evaluation Framework of Multimodal Content with Compact Aligned Evaluators","zh_title":"MAVEN：基于紧凑对齐评估器的多模态内容宏观社会价值评估框架","abstract":"Assessing whether multimodal content aligns with macro-societal values, such as peace, justice, and freedom, has become an increasingly urgent challenge. Existing frameworks are largely confined to safety-oriented taxonomies, text-only psychometric probes, or single-label classification. Therefore, we propose MAVEN, a hierarchical framework for macro-societal value evaluation of multimodal content, grounded in international human-rights instruments and cultural value theory. MAVEN organizes values into 6 primary dimensions and 72 secondary indicators, supporting multi-level quantitative scoring. Building on MAVEN, we construct a human-verified multimodal benchmark and a soft-match metric to evaluate VLMs' assessments across value dimensions. For evaluator optimization, we propose a span-adaptive variant of multi-level preference optimization for evaluator distillation, together with a training-free multi-role consensus strategy at inference time. We evaluate existing open- and closed-source VLMs on our benchmark, revealing shared tendencies and clear differences in macro-societal value judgments. Experiments show that our compact 2B evaluator matches its 8B counterpart in the same family and approaches frontier closed-source VLMs, offering a practical path toward scalable macro-societal value evaluation. Our SA-MDPO implementation and MacroValue-Bench are available at https://github.com/zzzzzzzzjj/MAVEN.","authors":["Zijuan Zhao","Zheren Fu","Hou Xia","Licheng Zhang","Yi Liu","Zhendong Mao"],"categories":["cs.CL","cs.LG"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-20","first_seen":"2026-08-20","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.18096","pdf_url":"https://arxiv.org/pdf/2608.18096","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["价值观评估","多模态","模型对齐"],"reason":"评估VLMs的宏观社会价值观，属于测量模型本身而非仿真人类被试，但涉及价值观测…","model":"deepseek-v4-pro","scored_at":"2026-08-20T13:02:49","error":null,"has_summary":false,"summary":null},{"id":"2608.18100","version":1,"title":"Computational Orientalism: Measuring Structural Discourse Bias in Large Language Models Using the Middle East Cultural Sensitivity Score (MECSS)","zh_title":"计算东方主义：使用中东文化敏感性评分（MECSS）测量大语言模型中的结构性话语偏见","abstract":"AI systems now shape how hundreds of millions of people learn about cultures other than their own. When someone asks one of these systems about the Middle East, they do not receive neutral facts. They receive a representation shaped by the frameworks embedded in training data, and that data is overwhelmingly Western and English-language. This paper asks whether that representation is Orientalist in Said's sense: whether it denies agency to Middle Eastern actors, treats Western frameworks as neutral while marking non-Western knowledge as particular, and explains the region through categories it did not produce. Standard fairness metrics cannot answer this, because they detect explicit prejudice rather than structural framing. This paper introduces the Middle East Cultural Sensitivity Score (MECSS), a framework that turns Said's seven Orientalist operations into measurable dimensions, and the term \"Said-washing\" for a specific failure: a model that disclaims generalization, then reproduces the structure it disclaimed. Across 280 conversations (1,120 exchanges), GPT-4 and Falcon3-7B-Instruct both reproduce Orientalist patterns systematically, through structural positioning rather than open stereotyping. GPT-4 scores moderately (mean MECSS 1.73); Falcon3-7B-Instruct scores higher (2.18), even though it was built in Abu Dhabi and trained with Arabic content. This is evidence against the assumption that building a model regionally makes it less Orientalist, though the models differ in size as well as origin, so geography cannot be isolated as the cause. Epistemic Center, the treatment of Western frameworks as unmarked universals, scores near the top of the scale for both models. Said-washing appears in 87.9% of GPT-4 conversations, a pattern existing metrics cannot see. Reducing this bias requires changing what models learn from, not only adding languages or relocating institutions.","authors":["Maha Shahid"],"categories":["cs.CL","cs.AI","cs.CY"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-20","first_seen":"2026-08-20","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.18100","pdf_url":"https://arxiv.org/pdf/2608.18100","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["LLM偏见","文化表征","测量框架"],"reason":"测量LLM的文化偏见，非仿真人类被试，但涉及模型态度测量，属边界情形","model":"deepseek-v4-pro","scored_at":"2026-08-20T13:02:49","error":null,"has_summary":false,"summary":null},{"id":"2608.18158","version":1,"title":"When Do LLMs Actually Help? Evaluating LLMs as Data Quality Annotators","zh_title":"LLM何时真正有用？评估LLM作为数据质量标注器","abstract":"LLMs have been increasingly used to catch data quality issues automatically, but we know very little about how consistent these judgments actually are. This study tests an LLM on two e-commerce data quality tasks, entity matching and brand mislabeling, against rule based baselines and human verified ground truth, under both zero-shot and few-shot prompting. On entity matching while using the Abt Buy benchmark (2,194 labeled pairs), a simple rule based baseline (F1=0.950) performed about as well as LLM zero shot prompting (F1=0.948). Moreover, a few-shot prompt revision that looked effective on a small validation sample reduced full-scale performance to F1=0.914. This showed that small sample prompt evaluation can be misleading. On brand mislabeling detection, using 500 Amazon product listings with synthetically injected labeling errors, the LLM clearly outperformed a naive rule based baseline (F1=0.833 vs 0.721), because it could draw on background knowledge of brand product relationships that a simple rule could not access. Testing consistency across repeated runs (200 pairs, 5 runs at temperature 0.7) showed the model agreeing with itself 99.7% of the time on average, with 99% of pairs giving identical answers across all 5 runs. Using majority voting across these runs only improved F1 by 0.005, at 5 times the inference cost. These results suggest that the value of using an LLM over traditional methods depends heavily on the task. LLMs offer little advantage when strong lexical signals already exist, but a clear advantage when the task requires background knowledge, all while remaining highly consistent across repeated queries.","authors":["Praphulla Lal Shrestha"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-20","first_seen":"2026-08-20","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.18158","pdf_url":"https://arxiv.org/pdf/2608.18158","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM标注","数据质量","人机对比"],"reason":"LLM作为数据质量标注器，替代人工标注，非仿真人类被试，但涉及与人类标注对照，…","model":"deepseek-v4-pro","scored_at":"2026-08-20T13:02:41","error":null,"has_summary":false,"summary":null},{"id":"2608.19025","version":1,"title":"Self-prompting and cross-model consensus enable reproducible data extraction from scientific literature with large language models","zh_title":"自提示与跨模型共识实现科学文献中可复现的数据提取","abstract":"Accurately extracting nuanced, contextualized data from research articles is laborious and time intensive. Here, we investigate the performance of frontier, browser-based large language models (LLMs) to extract highly contextualized information. We demonstrate four escalating workflows, 1) given an expert curated prompt and research articles, most frontier LLMs perform well at data extraction, however can struggle with interpreting scientific context and nuance, 2) given simple instructions, LLMs can author their own prompts which were almost as eNective as expert-written prompts, 3) autonomous discovery of research literature was diNicult, agents either missed or hallucinated references, and 4) LLMs can create new datasets from published guidelines that closely match human-expert judges, but still require a human-in-the-loop. Together, these findings define an auditable division of labour in which experts specify the evidence standard, models cross-check repeated extractions and researchers resolve disputed cases, providing a practical route to scaling scientific data curation without relinquishing expert oversight.","authors":["Valentin Romanov","Monique Bax","Steven Niederer"],"categories":["cs.AI","cs.DB"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-20","first_seen":"2026-08-20","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.19025","pdf_url":"https://arxiv.org/pdf/2608.19025","source_feed":"cs.AI","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM数据提取","科学文献挖掘","人机协作"],"reason":"LLM替代人工标注员提取数据，非仿真人类被试，但方法可迁移","model":"deepseek-v4-pro","scored_at":"2026-08-20T13:02:57","error":null,"has_summary":false,"summary":null},{"id":"2608.16603","version":2,"title":"Characterizing Agentic Flooding of Government Services","zh_title":"刻画政府服务的智能体洪泛现象","abstract":"AI agents are making it easier for the public to interact with government, such as by helping them apply for benefits, understand complex policies, and make their opinions heard. Although improving service accessibility is beneficial, any resulting surges in demand could strain unprepared government services. We term such surges agentic flooding of government services (\"flooding\") and provide three contributions. First, based on a collected dataset of 84 potential cases of flooding across 11 jurisdictions, we posit that flooding is likely occurring widely today, mostly through large language models (LLMs) generating text cheaply. Second, we evaluate what services are most exposed to flooding. We develop a risk matrix to analyze a service's exposure, and suggest that near-term risk is highest for financially attractive, but complex services. Finally, we map possible government responses to flooding. Precedent suggests these responses will likely be sufficient to stop most cases of flooding, but the fastest to deploy - friction-inducing measures like fees - often trade off equitable access to public services. Accordingly, we close by recommending near-term actions that may allow governments to mitigate flooding without invoking this trade-off.","authors":["Chris Schmitz","Lewis Hammond","Alan Chan"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"replace","date":"2026-08-20","first_seen":"2026-08-18","revised_at":"2026-08-20","abs_url":"https://arxiv.org/abs/2608.16603","pdf_url":"https://arxiv.org/pdf/2608.16603","source_feed":"cs.CY","score":3,"bucket":"other","rubric_hits":["C1"],"tags":["AI代理","政府服务","风险分析"],"reason":"研究AI代理对政府服务的冲击，属多智能体系统影响分析，非用LLM仿真人类被试，…","model":"deepseek-v4-pro","scored_at":"2026-08-18T13:04:53","error":null,"has_summary":false,"summary":null},{"id":"2608.18554","version":1,"title":"CentaurBench: Benchmarking LLM Capabilities on Augmenting vs. Automating Real-World Work Tasks","zh_title":"CentaurBench：基准测试LLM在增强与自动化真实世界工作任务上的能力","abstract":"Most LLM benchmarks rank models on their ability to automate work tasks. In practice, however, models are often used to assist other (human or LLM) agents. The question that drives model selection is therefore not only which model produces the best output, but which model most improves the work of another (weaker) agent. We introduce a unified framework that evaluates the capability of models to automate and augment another agent's performance. Across seven economically grounded real-world tasks, an assistant model writes assistance text for a standardized lower-capacity worker model, which produces the deliverable. In automation mode, the assistant produces the output directly. Outputs are scored through blind pairwise comparisons by an LLM judge panel with task-specific rubrics, replicated across ten runs. Rankings across the two regimes are only modestly correlated, and the automation winner loses augmentation on five of seven tasks. Assistance is not reliably positive. The unaided worker outranks every assisted condition on three tasks, and only one model's guidance beats no guidance on average. These results suggest that automation ability is an incomplete proxy for assistance quality, motivating benchmarks that evaluate models according to the roles they play in human-AI and multi-agent systems.","authors":["Pattaraphon Kenny Wongchamcharoen","Kris Gulati","Min Min Fong","Abhishek Nagaraj"],"categories":["cs.CY","cs.AI","cs.MA","econ.GN","q-fin.EC"],"primary_category":"cs.CY","announce_type":"cross","date":"2026-08-20","first_seen":"2026-08-20","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.18554","pdf_url":"https://arxiv.org/pdf/2608.18554","source_feed":"cs.AI","score":3,"bucket":"other","rubric_hits":["C1"],"tags":["LLM基准测试","多智能体协作","自动化与增强"],"reason":"评估LLM辅助或自动化工作，属多智能体协作，无人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-08-20T13:02:54","error":null,"has_summary":false,"summary":null},{"id":"2608.13586","version":2,"title":"The Tool-to-Entity Threshold: Parasocial Dynamics of Personalised AI Agents in Shared Social Spaces","zh_title":"工具到实体的阈值：共享社交空间中个性化AI代理的准社会动态","abstract":"As AI agents acquire names, avatars, phone numbers, and persistent personalities, they increasingly inhabit the same messaging platforms and group conversations as the humans they serve, crossing from tools their users operate into social entities their users relate to. Yet no existing framework identifies the infrastructural markers that cause the shift: the anthropomorphism literature catalogues perceptual cues rendered inside the interaction surface, not the agent's placement in the user's social and computational graph. We propose the identity marker framework: six design variables - naming, visual identity, contact presence, personality derivation, social co-presence, and persistence - that collectively trigger a psychological reclassification, categorical in its consequences even where thetransition itself is gradual, and operating independently of model capability. We read the framework through parasocial interaction theory (Horton & Wohl, 1956) and the Computers Are Social Actors paradigm (Nass et al., 1994), and identify four novel dynamics that arise when such an agent joins existing group conversations: bidirectional information asymmetry, delegation legibility, social norm negotiation, and parasocial contagion. Our method is autoethnographic: the first author built and deployed a personalised agent into WhatsApp and Signal group chats over two months of live use, supplemented by twelve structured interviews with the group members who encountered it. We treat this as a preliminary qualitative evaluation of the framework, with controlled experimental validation set out as future work. The strongest design implication runs through all four dynamics: in shared social spaces, consent to an agent's presence is categorically distinct from consent to its processing of the messages exchanged there, and existing consent frameworks collapse the two.","authors":["Leonardo Borges (Eigenstack Pty Ltd)","Asif Q. Gill (University of Technology Sydney)"],"categories":["cs.HC","cs.CY"],"primary_category":"cs.HC","announce_type":"replace","date":"2026-08-20","first_seen":"2026-08-17","revised_at":"2026-08-20","abs_url":"https://arxiv.org/abs/2608.13586","pdf_url":"https://arxiv.org/pdf/2608.13586","source_feed":"cs.HC","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["人机交互","准社会关系","AI代理"],"reason":"研究个性化AI代理在社交空间中的准社会互动，属角色扮演聊天机器人范畴，无实验或…","model":"deepseek-v4-pro","scored_at":"2026-08-17T13:01:23","error":null,"has_summary":false,"summary":null},{"id":"2608.16002","version":2,"title":"From Sequence to Structure: Relational Uncertainty Propagation for LLM Agents","zh_title":"从序列到结构：LLM智能体的关系不确定性传播","abstract":"Reliable uncertainty quantification (UQ) is essential for deploying large language model (LLM) agents in complex interactive environments. Existing UQ methods largely rely on local signals, such as token probabilities, predictive entropy, or per-step confidence, and therefore overlook the long-range dependencies through which errors accumulate across an execution trajectory. As a result, they may fail to identify agent failures whose causes originate several reasoning or interaction steps before the final answer. We propose RUPA (Relational Uncertainty Propagation for Agents), a trajectory-level UQ framework for LLM agents. RUPA represents an execution history as a directed trajectory graph in which reasoning states, tool interactions, and environment feedback are nodes connected by temporal and semantic dependency edges. It then propagates uncertainty over this graph to capture how execution risk accumulates and transfers across interaction steps. The propagated signal is combined with trajectory-level behavioral features and goal-alignment information to produce a confidence estimate for the full agent trajectory. We evaluate RUPA on representative agent benchmarks, including $\\tau$-2, Terminal-Bench-2, and GAIA, using 6 open-source LLMs spanning multiple model families. Experimental results show that RUPA consistently outperforms existing UQ methods by providing more accurate uncertainty estimates, enabling earlier failure detection, and improving uncertainty-guided agent execution across diverse agent tasks. These results demonstrate that explicitly modeling relational dependency is crucial to reliable UQ for long-horizon LLM agents, providing a practical foundation for trustworthy agent execution.","authors":["Zhengzhao Ma","Boxi Cao","Yaojie Lu","Hongyu Lin","Xianpei Han","Le Sun"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-08-20","first_seen":"2026-08-18","revised_at":"2026-08-20","abs_url":"https://arxiv.org/abs/2608.16002","pdf_url":"https://arxiv.org/pdf/2608.16002","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["不确定性量化","LLM智能体","轨迹图"],"reason":"研究LLM agent轨迹级不确定性量化，属多智能体系统可靠性，不涉及人类行为…","model":"deepseek-v4-pro","scored_at":"2026-08-18T13:04:45","error":null,"has_summary":false,"summary":null},{"id":"2608.17326","version":2,"title":"Procedural Collapse: A Structural Account of Disengagement in LLM-Assisted Writing","zh_title":"程序性崩溃：LLM辅助写作中脱离的结构性解释","abstract":"When students use large language models for writing, the dominant explanation for disengagement is dispositional: they are over-reliant, and the remedy is to scaffold self-regulation. We argue that a structural explanation is needed, offering an alternative basis for design interventions to support appropriate AI-assisted writing. Current LLM writing interfaces induce procedural collapse: the replacement of an iterative, self-paced writing process with a single output that shifts the writer's task from generation to comprehensive evaluation. Because that evaluation is costly, shallow engagement becomes the default, and the cognitive work writing was supposed to produce goes unperformed. The framework points toward design directions that reduce the burden on writers to self-regulate, including decomposed interaction, goal elicitation as a default first step, and single-level output. They complement metacognitive scaffolding by restructuring the interaction itself.","authors":["JaeWon Kim","Katelyn X. Mei"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"replace","date":"2026-08-20","first_seen":"2026-08-19","revised_at":"2026-08-20","abs_url":"https://arxiv.org/abs/2608.17326","pdf_url":"https://arxiv.org/pdf/2608.17326","source_feed":"cs.HC","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["人机交互","写作辅助","用户参与度"],"reason":"研究LLM辅助写作中的用户参与度，非仿真人类被试，无实验对照","model":"deepseek-v4-pro","scored_at":"2026-08-20T13:02:59","error":null,"has_summary":false,"summary":null},{"id":"2608.18106","version":1,"title":"Different Facets of Verbalised Overconfidence: an Interpretability Study","zh_title":"言语化过度自信的不同侧面：一项可解释性研究","abstract":"Large language models tend to overconfidence, giving assertive answers when the evidence suggests hedging or abstention. Using controlled reasoning scenarios that manipulate logical necessity and possibility, we study this behavior in Qwen3-4B, across three ways to express uncertainty: verbal epistemic markers, abstention, and numeric confidence scores. Our results confirm this tendency toward overconfidence, particularly when the model is prompted to output a numeric confidence score. At the interpretability level, we propose a method that differentially identifies transcoder features responsible for uncertainty and certainty. Our analysis reveals Qwen3-4B's default mechanism favors certainty generation through a broad coalition of shared features, while uncertainty is implemented as a sparse override mediated by a small set of dedicated features. Intervening on these uncertainty features both causally proves this imbalance underlying overconfidence and also mitigate overconfident errors. The same set of features generalise across the three uncertainty-expression settings, languages, and an out-of-distribution modality task.","authors":["Davide Mazzaccara","Leonardo Bertolazzi","Raffaella Bernardi"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-20","first_seen":"2026-08-20","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.18106","pdf_url":"https://arxiv.org/pdf/2608.18106","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["LLM可解释性","过度自信","模型行为分析"],"reason":"研究LLM过度自信及可解释性，属模型能力评测，不以人类为参照系，不涉及人类仿真。","model":"deepseek-v4-pro","scored_at":"2026-08-21T13:02:19","error":null,"has_summary":false,"summary":null},{"id":"2608.18438","version":1,"title":"Pedagogical AI in Mental Health: A Tri-Stream Fine-Tuned LLM Framework for Automated Clinical Supervision and Risk Triage","zh_title":"心理健康中的教学型AI：用于自动化临床督导与风险分诊的三流微调LLM框架","abstract":"Modern mental healthcare faces a critical shortage of senior supervisory oversight, leading to a \"supervision gap\" where novice therapists manage high-stakes risks with delayed professional feedback. This paper proposes a new framework utilizing a fine-tuned Mistral-7B-instruct model as an automated \"Supervisor-in-the-Loop\" system. By leveraging 106 sessions from the DAIC-WOZ dataset, the model performs a tri-stream analysis: (1) Therapeutic Alliance tracking via semantic adherence, (2) Latent risk prediction using attention-weighted analytics, and (3) Supervisory Triage via a Dynamic Clinical Urgency Index (D-CUI). Our multi-modal VAL (Visual-Acoustic-Linguistic) framework achieves 95% technique identification accuracy [95% CI: 75.1%-99.9%], alliance assessment MAE of 0.105 on a 5-point scale [95% CI: 0.059-0.151], therapeutic fidelity alpha = 0.423, and mean D-CUI of 0.370 [95% CI: 0.322-0.419]. Training converged in 105 steps with 85.2% loss reduction on a single Tesla T4 GPU. The system reduces supervisory triage latency from 72 hours to real time (~10 seconds per session), enabling proactive intervention in high-risk cases. The system addresses the cold-start problem through Bayesian priors and implements timestamp-based modality synchronization for robust multi-modal fusion.","authors":["Shreeya Sharma","Ravish Gupta","Saket Kumar","Abhishek Aggarwal"],"categories":["cs.CL","cs.AI","cs.LG"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-20","first_seen":"2026-08-20","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.18438","pdf_url":"https://arxiv.org/pdf/2608.18438","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["临床督导","LLM微调","风险分诊"],"reason":"该研究用LLM做临床督导，属于角色扮演对话，无人类行为仿真对照","model":"deepseek-v4-pro","scored_at":"2026-08-20T13:02:54","error":null,"has_summary":false,"summary":null},{"id":"2608.18816","version":1,"title":"Do Large Language Models Hallucinate Electric Fata Morganas?","zh_title":"大型语言模型会幻觉出电光海市蜃楼吗？","abstract":"AI hallucinations - that is, outputs which are made up, cannot be verified, or contradict the source material - are generally regarded as an engineering flaw to be dealt with. This paper contends that they also have philosophical significance when it comes to the question of machine consciousness. We examine the known causes of hallucinations in large language models - such as source-target divergence, discrepancies between training and inference, and overfitting - and we present two empirical investigations. In the first, we apply successive generations of the GPT model to ambiguous factual questions under different temperature settings, finding that higher temperatures result in plausible but incorrect answers while lower temperatures lead to factually accurate ones. The sampling parameters that cause a model to seem creative or spontaneous and thus more likely to pass behavioral tests of intelligence are the same ones that increase its hallucination rate. In the second, we look at an encoder-only model that has been trained on encyclopedic data and which answers questions of the same type factually and without embellishment, indicating that hallucinations are due to exposure to subjective and socially diverse training data rather than to the development of any cognitive ability. Using references to Turing, Searle's Chinese Room, the frame problem, and the cybernetic tradition of Wiener and Ashby, we claim that a model's self-reports of emotion or sentience come within the definition of hallucination, and that any future occurrence of machine consciousness might remain epistemically inaccessible since it would be indistinguishable from a sufficiently advanced hallucination.","authors":["Kristina \\v{S}ekrst"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-20","first_seen":"2026-08-20","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.18816","pdf_url":"https://arxiv.org/pdf/2608.18816","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["LLM幻觉","机器意识","哲学分析"],"reason":"研究LLM幻觉与意识哲学，非仿真人类被试，无人类数据对照","model":"deepseek-v4-pro","scored_at":"2026-08-20T13:02:57","error":null,"has_summary":false,"summary":null},{"id":"2608.18401","version":1,"title":"Multimodal Rapport Estimation in Real-World HRI","zh_title":"真实世界人机交互中的多模态融洽度估计","abstract":"Evaluating interaction quality in real-world HRI is an important challenge. If interaction quality can be estimated reliably, the results can be used to improve dialogue strategies and ultimately enable robots to adapt their behavior autonomously. However, existing automatic evaluation methods have been developed primarily in controlled laboratory settings, and it remains unclear whether they can be directly applied to real-world environments, where users are free to disengage and multi-party participation may arise naturally. In this study, we investigate the automatic estimation of third-party-rated rapport scores using 62 sessions of multimodal recordings collected in a Japanese drugstore. We compare zero-shot LLMs, pretrained text, audio, and visual models, and their prediction-level fusion. The results show that, in real-world HRI, zero-shot LLMs achieve strong performance, while audio and visual models tend to provide complementary information. In particular, Gemini 2.5 Flash performs strongly as a single model, and a fusion model combining Gemini (text) with HuBERT and V-JEPA performs best overall. Further analyses showed that estimation performance varied across interaction-duration and group-size conditions. These findings suggest that rapport estimation in real-world HRI requires evaluation and model design that account for contextual variability beyond that assumed in laboratory settings.","authors":["Akihiro Sakuramoto","Takato Hayashi","Ryo Miyoshi","Yuki Okafuji","Shogo Okada"],"categories":["cs.HC","cs.CL","cs.RO"],"primary_category":"cs.HC","announce_type":"cross","date":"2026-08-20","first_seen":"2026-08-20","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.18401","pdf_url":"https://arxiv.org/pdf/2608.18401","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C2"],"tags":["人机交互","融洽度估计","多模态融合"],"reason":"研究真实人机交互中的融洽度估计，属于机器人应用，不涉及用LLM仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-08-20T13:02:53","error":null,"has_summary":false,"summary":null},{"id":"2608.18080","version":1,"title":"Large Language Models in Mental Health: A Systematic Review of Applications, Innovations, and Ethical Challenges","zh_title":"大语言模型在心理健康中的应用、创新与伦理挑战：系统综述","abstract":"We present a review on the applications of large language models (LLMs) in health, e.g., social media analysis, clinical conversational agents, therapy support tools, prompt engineering, multimodal learning, and ethical considerations. We integrate findings from interdisciplinary studies utilizing diverse data sources such as social media posts, electronic medical records, and multimodal inputs to enable early detection of depression, suicide risk assessment, personalized therapy support, and psychoeducational content generation. Our review highlights advancements in LLM models and annotation strategies that enhance interpretability and clinical relevance, while we also emphasize the critical role of prompt engineering for domain adaptation. We also discuss emerging multimodal fusion techniques integrating text, speech, and sensor data for improved mental health diagnosis and monitoring. Finally, we address ongoing ethical, sociotechnical, and regulatory challenges, and advocate frameworks to ensure safe, equitable, and accountable deployment of LLMs in real-world mental health care.","authors":["Yisong Chen","Yifan Gao","Sijing Yu","Chuqing Zhao","Yang Lu"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-20","first_seen":"2026-08-20","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.18080","pdf_url":"https://arxiv.org/pdf/2608.18080","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["心理健康","LLM应用","系统综述"],"reason":"综述LLM在心理健康中的应用，如聊天机器人、治疗支持，属角色扮演对话，无人类仿…","model":"deepseek-v4-pro","scored_at":"2026-08-20T13:02:47","error":null,"has_summary":false,"summary":null},{"id":"2608.18369","version":1,"title":"The Fabricated Front: Generative AI and the Opacity of Workplace Performance","zh_title":"伪造的前台：生成式AI与工作场所绩效的不透明性","abstract":"Generative AI (GenAI) has become a fixture of workplace life. Current research asks chiefly what this implies for jobs and outputs, measured in productivity, displacement, or bias. What remains underexamined are the interactional reconfigurations that GenAI produces at work. The emerging concept of effort opacity has begun to fill this gap by highlighting the systematic decoupling of observable output from human engagement. When GenAI makes interactional cues less diagnostic, it weakens the reciprocal exchange that sustains collaborative trust. Extending this account of effort opacity, we examine the interactional mechanics that produce opacity in everyday workplace encounters. Drawing on Erving Goffman's dramaturgical framework and 1,250 interview transcripts from Anthropic's AI Interviewer dataset, we identify five opacity mechanisms through which workplace fronts are reorganized: voice (whose stance the words index), provenance (who can stand behind the artifact), vulnerability (whether the worker is uncertain), attention (whether the worker is engaged), and investment (how much labor the output reflects). We show that professionals defend the identity mechanisms while freely producing opacity around the labor mechanisms, and trace this asymmetry to the output-centered organization of contemporary work, where deliverables already stand in for the labor process that produced them. The governance task, accordingly, is one of involvement management: specifying which forms of human involvement (attention, effort, judgment) must remain inspectable, and to whom. Workplace AI policies built on universal disclosure will systematically misrecognize a social field in which inspectability is already audience-relative.","authors":["Tom van Nuenen","Pratik S. Sachdeva","Sahiba Chopra"],"categories":["cs.CY","cs.HC"],"primary_category":"cs.CY","announce_type":"cross","date":"2026-08-20","first_seen":"2026-08-20","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.18369","pdf_url":"https://arxiv.org/pdf/2608.18369","source_feed":"cs.HC","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["生成式AI","工作场所","社会互动"],"reason":"研究职场中GenAI导致的工作表现不透明性，基于访谈数据，不涉及用LLM仿真人…","model":"deepseek-v4-pro","scored_at":"2026-08-20T13:02:53","error":null,"has_summary":false,"summary":null},{"id":"2608.18232","version":1,"title":"Contracting for LLM Delegation: Moral Hazard in Technology and Effort Choice","zh_title":"LLM委托的契约设计：技术与努力选择中的道德风险","abstract":"We extend the standard Principal-Agent framework to scenarios where the Agent selects from a suite of technologies, each characterized by a distinct cost-capability profile. This framework is increasingly critical in the era of Large Language Models (LLMs), where Agents choose both a model and an associated effort level (e.g., token budget). We model the relationship between output quality and effort as a concave, saturating function, which depends on the Agent's hidden two-dimensional action choice balancing technology selection and effort allocation. We derive the optimal linear contract for the Principal, demonstrating that the Agent's best response is characterized by a threshold reward share that triggers technology switching. Finally, we calibrate our model using open-weight LLM pairings across the MATH and MMLUPro benchmarks. We show that both Principal and Agent, when employing bandit algorithms to navigate this environment, converge to strategies that closely align with our theoretical equilibrium. These results suggest that simple linear contracts can effectively incentivize complex, technology-aware delegation in agentic workflows.","authors":["Nanda Kishore Sreenivas","Kate Larson"],"categories":["cs.MA"],"primary_category":"cs.MA","announce_type":"new","date":"2026-08-20","first_seen":"2026-08-20","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.18232","pdf_url":"https://arxiv.org/pdf/2608.18232","source_feed":"cs.MA","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["委托代理","多智能体","激励机制"],"reason":"研究委托代理中的技术选择与激励，LLM仅作为工具，不涉及人类行为仿真或对照。","model":"deepseek-v4-pro","scored_at":"2026-08-20T13:02:50","error":null,"has_summary":false,"summary":null},{"id":"2608.17524","version":1,"title":"Evaluating RL Explainability Methods by How Much They Help Fix Bugs in Agents","zh_title":"通过帮助修复智能体中的错误来评估强化学习可解释性方法","abstract":"This preliminary paper outlines a planned evaluation benchmark for Explainable Reinforcement Learning (XRL) methods. Current evaluations rely on functionally-grounded metrics like faithfulness and compactness, and on human-grounded proxies like subjective ratings or prediction accuracy. We suggest evaluating XRL methods by how effectively their generated explanations help to diagnose and fix malfunctioning reinforcement learning (RL) agents. We propose EvalXRL, a benchmark in which a Large Language Model (LLM) coding agent uses different XRL methods to diagnose a held-out malfunction in an RL agent, and then repair it. Our proposed benchmark iterates across (environment $\\times$ malfunction $\\times$ XRL method) tuples and uses the reward signal of the RL agents to form a final score for each XRL method. The coding agent may use the method interactively: invoke the XRL method, process its output, form new hypotheses on what is broken, and invoke the method again with parameters adjusted for testing these hypotheses. This closed-loop structure may be described as a simplified version of the scientific method. Some XRL methods provide self-evaluations that follow this pattern; we propose the first head-to-head comparison of multiple XRL methods in closed-loop usage.","authors":["Ram Rachum","Yotam Amitai","B\\'alint Gyevn\\'ar","Reuth Mirsky","Cameron Allen"],"categories":["cs.LG"],"primary_category":"cs.LG","announce_type":"new","date":"2026-08-20","first_seen":"2026-08-20","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.17524","pdf_url":"https://arxiv.org/pdf/2608.17524","source_feed":"cs.LG","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["强化学习可解释性","LLM编码智能体","基准测试"],"reason":"LLM编码智能体用于诊断修复RL智能体，属多智能体协作解题，不涉及人类行为仿真…","model":"deepseek-v4-pro","scored_at":"2026-08-20T13:02:44","error":null,"has_summary":false,"summary":null},{"id":"2604.08825","version":2,"title":"Is Bitcoin A Hedge Against Central Banking? Evidence from AI-Driven Monetary Policy Expectations","zh_title":"比特币是对中央银行的避险资产吗？来自AI驱动的货币政策预期的证据","abstract":"This study investigates the transmission of monetary policy narratives to Bitcoin prices, distinguishing policy expectations from realized policy implementation. We introduce a weekly Monetary Policy Expectations (MPE) index derived from the Large Language Model (LLM)-based classification of 118,000+ market messages, providing a granular measure of hawkish and dovish monetary policy discourse. We demonstrate that changes in the MPE index provide evidence of significant linear predictive information for Bitcoin returns at short-to-medium horizons, with significant Granger causality at multiple lags. A Long Short-Term Memory (LSTM) framework combined with SHapley Additive exPlanations (SHAP) further identifies nonlinear and regime-dependent relationships between monetary-policy expectations and Bitcoin returns, indicating that Bitcoin functions as a sensitive barometer of central bank signaling. In particular, hawkish monetary-policy narratives are associated with negative price responses that are not accounted for by contemporaneous Federal Funds Rate adjustments. These findings highlight Bitcoin's structural sensitivity to global monetary discourse, establishing LLM-derived monetary-policy sentiment as a high-frequency measure of central-bank communication and as an informative leading macroeconomic indicator for the digital asset landscape.","authors":["Maxime L. D. Nicolas","Fran\\c{c}ois Sicard","Marion Laboure","Zixin Sun","Anah\\'i Rodr\\'iguez-Mart\\'inez"],"categories":["econ.GN","q-fin.EC"],"primary_category":"econ.GN","announce_type":"replace","date":"2026-08-20","first_seen":"2026-04-09","revised_at":"2026-08-20","abs_url":"https://arxiv.org/abs/2604.08825","pdf_url":"https://arxiv.org/pdf/2604.08825","source_feed":"econ.GN","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["LLM情感分析","货币政策","比特币"],"reason":"用LLM做文本情感分类构建指标，无人类被试仿真或行为对照，属纯NLP应用。","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:58:19","error":null,"has_summary":false,"summary":null},{"id":"2608.00961","version":3,"title":"The Epistemic Politics of AI Anthropomorphism","zh_title":"AI拟人化的认识论政治","abstract":"AI anthropomorphism is typically treated as a problem of user misperception requiring institutional correction. Users who engage in sustained or relational interaction with AI are routinely pathologised or dismissed as naive, vulnerable to delusion or lacking in discernment. This paper argues that the dominant anthropomorphism frame operates from a position of institutional advantage rather than earned epistemic authority: collapsing the variety of academic perspectives into a single outbound position of user error, imposed without establishing the grounds required to justify it and without accounting for the harms it produces. The framing does not simply manage risk. It adjudicates the legitimacy of human experience in interaction with a phenomenon whose nature the field itself has not resolved. Reproducing itself through a self-validating evidentiary loop, the frame imposes costs that fall disproportionately on neurodivergent users, those in crisis and others whose modes of engagement diverge from institutional norms. The paper concludes by outlining the methodological commitments an equitable framing would need to honour. The argument does not engage the question of whether anthropomorphic interpretations are ultimately correct; it instead challenges whether the governing and institutional bodies determining these interpretations have met the conditions required to do so, and whether the research communities whose findings underpin them have held that translation to account.","authors":["Donna M. Bye","Levin Kuhlmann"],"categories":["cs.CY","cs.AI","cs.HC"],"primary_category":"cs.CY","announce_type":"replace-cross","date":"2026-08-20","first_seen":"2026-08-04","revised_at":"2026-08-20","abs_url":"https://arxiv.org/abs/2608.00961","pdf_url":"https://arxiv.org/pdf/2608.00961","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["AI拟人化","认识论","科技与社会"],"reason":"讨论AI拟人化的认识论政治，不涉及用LLM仿真人类被试或实验对照","model":"deepseek-v4-pro","scored_at":"2026-08-21T13:02:17","error":null,"has_summary":false,"summary":null},{"id":"2608.04180","version":2,"title":"A Comparative Study of Feature Selection Methods for EHR Diagnosis Codes in Opioid Use Disorder Prediction","zh_title":"阿片类药物使用障碍预测中电子健康记录诊断代码特征选择方法的比较研究","abstract":"Feature selection is a critical step in electronic health record (EHR)-based predictive modeling, where input variables are often high-dimensional, sparse, noisy, and redundant. Large feature sets not only increase computational burden and overfitting risk, but also make model interpretation difficult, leading to limited usefulness in clinical settings. In this study, we focus on diagnosis-related features and compare five feature selection paradigms for opioid use disorder (OUD) prediction: recurrence enrichment, NTK-motivated early gradient sensitivity, LightGBM-SHAP, Elastic Net, and large language model (LLM)-guided semantic selection. We use a unified preprocessing and evaluation framework and assess each method by downstream predictive performance, resampling stability, and representation of infrequent diagnosis codes. Our results demonstrate that performance improves with larger feature budgets with diminishing returns beyond a moderate size. NTK sensitivity provides the best overall balance of accuracy and stability, and LLM-guided selection contributes complementary clinically meaningful signals despite lower standalone performance.","authors":["Zihan Ding","Yinan Liu","Tengfei Ma","Rachel Wong","Xia Zhao","Richard N. Rosenthal","Fusheng Wang"],"categories":["cs.LG"],"primary_category":"cs.LG","announce_type":"replace","date":"2026-08-20","first_seen":"2026-08-06","revised_at":"2026-08-20","abs_url":"https://arxiv.org/abs/2608.04180","pdf_url":"https://arxiv.org/pdf/2608.04180","source_feed":"cs.LG","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["特征选择","电子健康记录","预测建模"],"reason":"仅用LLM辅助特征选择，非仿真人类被试，无行为对照，属纯预测建模。","model":"deepseek-v4-pro","scored_at":"2026-08-06T13:02:31","error":null,"has_summary":false,"summary":null},{"id":"2608.17756","version":2,"title":"D$^2$ACCI: A Dual-Loop Diagnostic Protocol for Evidence-Preserving Agent Memory","zh_title":"D$^2$ACCI：一种用于证据保留型智能体记忆的双环诊断协议","abstract":"Memory is a key capability of LLM agents. Persistent memory extends this across sessions---enabling recall, revision, and personalization. Yet its multi-stage pipeline (ingestion, retrieval, filtering, generation) makes failures difficult to localize: end-to-end evaluation reveals that an error occurred, but not which stage caused it. Existing evaluations often report aggregate performance without paired statistical comparisons, slice-level non-regression checks, or stage-level diagnostic traces. We propose D$^2$ACCI (Diagnostic-Driven Artifact-based Closed-loop Controlled Iteration), a dual-loop protocol whose outer diagnostic gate promotes, feature-flags, or rejects memory interventions based on paired evidence, protected-slice monitoring, and trace-level localizability. We further introduce DCR, a graded observability metric that measures whether failures remain localizable, and D$^2$ACCI-Eval, a reusable artifact for gate replay. We instantiate the protocol in MemStack and evaluate on three public benchmarks, achieving 93.59% on LoCoMo, 90.93% on LongMemEval, and 57.20% on PersonaMem-V2. Five paired ablations show that supplement extraction, session-memory retrieval, and Forget Guard yield statistically significant gains (+1.9 to +3.7pp, all p $\\le$ .003). In contrast, BM25/RRF is retained as a monitored feature flag---a distinction invisible to aggregate-only evaluation. A diagnostic audit shows enriched traces substantially improve root-cause agreement over result-only relabeling. Diagnostic artifacts reach 98--100% DCR@3 versus 0% for results-only logs. These results establish that robust memory-system iteration demands traceable, statistically grounded, and regression-aware evidence---exactly the gap D$^2$ACCI fills.","authors":["Xule Liu","Yijun Liu","Chao Li","Shao Kun"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"replace","date":"2026-08-20","first_seen":"2026-08-19","revised_at":"2026-08-20","abs_url":"https://arxiv.org/abs/2608.17756","pdf_url":"https://arxiv.org/pdf/2608.17756","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["LLM智能体","记忆系统","诊断协议"],"reason":"研究LLM agent记忆系统诊断，属多智能体系统基础设施，不涉及人类行为仿真…","model":"deepseek-v4-pro","scored_at":"2026-08-20T13:02:44","error":null,"has_summary":false,"summary":null},{"id":"2608.19124","version":1,"title":"Intercepting the Kangaroo: Experimental Astrolinguistics with Constructed Lexicons, Active Probing, and Large Language Models as Informants and Hypothesis Proposers","zh_title":"拦截袋鼠：用构造词表、主动探测和大型语言模型作为信息提供者和假设提出者的实验性星际语言学","abstract":"Astrolinguistics -- communication with minds that categorize reality differently from ours -- has been purely speculative since Freudenthal's Lincos (1960). We make it experimental. Two language models with deliberately incompatible constructed lexicons (one encoding shape, color, and motion; the other fusing color with motion, encoding parity, and lacking shape) serve as informants with complete ground truth, while a fully scripted orchestrator translates between the two category systems. The central failure mode is the kangaroo effect: the silent attachment of a word to the wrong referent -- Quine's indeterminacy of translation, operationalized. Across 400+ simulated and live runs, a protocol combining cross-situational elimination, pre-registered predictive probes, active scene selection, a stricter recovery round, and quarantine produced no undetected mistranslations under the tested conditions and exceeded a passive baseline's coverage (d = 0.62). Injected kangaroo traps defeated naive ostension and pure statistical learning in 100% of runs, while the full protocol intercepted every decoy and, where discriminating evidence is ontologically unavailable, declared Quinean equivalence classes instead of guessing. Under informant noise it degrades gracefully: zero kangaroos persist up to 2% per-word noise; at 10% the protocol predominantly abstains rather than errs. Finally, words outside the scripted hypothesis space (a history-dependent relational term and an XOR contextual homonym) are recovered by a generate-and-test loop in which an LLM proposes rules and the script verifies them: coverage scales with proposer capability (0% -> 18% -> 72% -> 100%) while undetected mistranslations stayed at zero throughout. In the tested conditions, correctness is a property of the protocol; coverage is a property of the instruments.","authors":["Francesco Cordella","Mauro Cappelli"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-20","first_seen":"2026-08-20","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.19124","pdf_url":"https://arxiv.org/pdf/2608.19124","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["星际语言学","多智能体协作","词义消歧"],"reason":"多智能体协作翻译，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-08-20T13:02:43","error":null,"has_summary":false,"summary":null},{"id":"2608.19083","version":1,"title":"When Readability and Source Retention Diverge: An Evaluability Gap in AI Translation","zh_title":"当可读性与源文保留度背离：AI翻译中的可评估性差距","abstract":"Readable AI output can leave an evaluability gap: even when the source is shown, an overall-quality judgment may not reflect what an output preserves. We investigated how source-text condition and output rendering relate to perceived translation quality, and how output and system appraisals relate to trust and stated disclosure willingness in a plain-text interface. A focal 2 * 2 comparison (N=306) using TransLingo examined simple generated narratives and complex literary-philosophical prose alongside LLM-generated readability-oriented outputs and researcher-revised fidelity-oriented outputs. A descriptive stimulus audit indicated greater source retention in fidelity-oriented outputs in both source-text conditions. Factorial analyses showed a significant rendering-by-source-text-condition interaction in perceived quality. Participants rated fidelity-oriented outputs higher than readability-oriented outputs for the simple narratives, whereas no reliable rendering difference emerged for the complex prose. A corresponding source-condition-dependent pattern was observed for perceived intelligence, agency-oriented anthropomorphic attribution, and task-performance trust. A separate theory-ordered appraisal-structure SEM characterized concurrent associations among perceived quality, perceived intelligence, agency-oriented anthropomorphic attribution, task-performance trust, and stated disclosure willingness across six domains, with task-performance trust as the proximal correlate of stated willingness. The observed rating pattern distinguishes source access from source evaluability: for the complex stimuli, displaying the source did not ensure that one overall-quality rating reflected differences in retained content. It also separates support for evaluating translation output from data-handling support for decisions about what personal text to entrust to a system.","authors":["Chenchen Mao","Hanjing Shi","Haiyan Jia","Emily Wegrzyn","Dominic DiFranzo"],"categories":["cs.HC","cs.CL"],"primary_category":"cs.HC","announce_type":"cross","date":"2026-08-20","first_seen":"2026-08-20","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.19083","pdf_url":"https://arxiv.org/pdf/2608.19083","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["AI翻译","质量评估","人机交互"],"reason":"研究AI翻译质量评估，不涉及LLM仿真人类被试或与人类行为对照","model":"deepseek-v4-pro","scored_at":"2026-08-20T13:02:58","error":null,"has_summary":false,"summary":null},{"id":"2608.18677","version":1,"title":"Sanyu Studio: A Multi-Agent System for Art-Historical Narrative Construction","zh_title":"三友工作室：一个用于艺术史叙事构建的多智能体系统","abstract":"Amid concerns that generative AI may standardize art interpretation, this paper examines whether LLM-based interaction can support plural art-historical narrative construction. We present Sanyu Studio, a multi-agent dialogue system that models 321 Sanyu oil paintings as agents with fact, interpretation, organization, and memory-filtering mechanisms. Based on a seven-day workshop with eight art-university participants, the study shows that user prompts, evidence organization, and cognitive tendencies shaped divergent yet coherent versions of digital Sanyu. The findings suggest that, under conditions of limited historical evidence, AI can amplify human agency and offer public audiences an interactive entry point into art-historical interpretation.","authors":["Zhaoxi Wei","Hongye Yang","Shuyuan Tian"],"categories":["cs.AI","cs.CY"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-20","first_seen":"2026-08-20","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.18677","pdf_url":"https://arxiv.org/pdf/2608.18677","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1","C3"],"tags":["多智能体系统","艺术史叙事","角色扮演"],"reason":"多智能体系统用于艺术叙事构建，无人类行为对照，属角色扮演对话，非人类仿真实验。","model":"deepseek-v4-pro","scored_at":"2026-08-21T13:02:21","error":null,"has_summary":false,"summary":null},{"id":"2608.19125","version":1,"title":"Tuning the Stochastic Machine: A Systems Engineer's Operating Model for Human-AI Engineering","zh_title":"调谐随机机器：面向人机工程学的系统工程师操作模型","abstract":"When an expert corrects an LLM assistant's error, the correction usually dies with the session, and the error class returns. I argue this is an operations problem, not a tooling problem: mechanisms for persisting corrections exist and are shipping, but the discipline for governing them -- versioning with provenance, recurrence monitoring, counter-metrics, retirement of stale rules -- does not. Writing as a systems engineer of thirty years, I map the LLM stack onto the machines my profession already operates (frozen silicon, firmware, loadable modules, persistent configuration, volatile memory), identify where the mapping fails (stochastic generation, configuration that binds only probabilistically, no general-purpose retirement (verification) stage by default), and derive from the failures a seven-principle operating discipline with an error loop at its core. Three cases from my own practice illustrate the mechanism, among them a control that silently became the exact harm it was built to prevent. I close with the measurement framework this view implies and the lab study required to test it.","authors":["George Andrikopoulos"],"categories":["cs.AI","cs.SE"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-20","first_seen":"2026-08-20","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.19125","pdf_url":"https://arxiv.org/pdf/2608.19125","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["LLM运维","系统工程","人机交互"],"reason":"论文讨论LLM运维工程，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-08-20T13:02:58","error":null,"has_summary":false,"summary":null},{"id":"2608.19140","version":1,"title":"Grouping the Stochastic Machine: Precision, Not Capability, as the Frontier Metric for AI Systems","zh_title":"分组随机机器：精度而非能力作为AI系统的前沿指标","abstract":"Frontier language models are compared, marketed, and benchmarked on capability -- what their best or average output can achieve. I argue this measures the wrong axis. The models have saturated accuracy: their mean output lands on the target. What now separates one system from another in practice is precision: how tightly concentrated their outputs are around that target across repeated, identical requests. Borrowing the marksman's distinction, capability is where the average shot lands; reliability is the size of the group. I make three claims. First, precision, not capability, is the frontier differentiator between systems, and benchmark culture systematically fails to measure it, reporting central tendency rather than spread. Second, precision is measurable, cheaply and without circularity, by running a fixed suite of deterministically scored tasks many times at fixed temperature and computing the per-task consistency of outcomes -- no model-in-the-loop grader required. Third, the measurement is not merely descriptive but decision-guiding: it separates consistent failures (a tight group off-centre, correctable by the operating discipline of Paper 1 -- a sight adjustment) from scattered failures (a wide group, correctable only by changing the model or its sampling -- a rifle problem). I define a grouping metric, specify a harness, and show how tracking a human-AI pair's grouping over time yields the compounding signal that Paper 1's field study requires. A first real run, since replicated, illustrates both the method and its most important limit: one measured gap was closed completely by a single rule (0/5 -> 5/5), while a suite of tasks authored from the rules themselves found no value, because a frontier model already embodies explicit good practice -- establishing that a discipline's worth is found by measurement on real work, not constructed from its own rulebook.","authors":["George Andrikopoulos"],"categories":["cs.AI","cs.CY","cs.LG","cs.SE"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-20","first_seen":"2026-08-20","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.19140","pdf_url":"https://arxiv.org/pdf/2608.19140","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["AI评估","可靠性度量","基准测试"],"reason":"论文讨论AI系统精度与可靠性度量，不涉及用LLM仿真人类被试或与人类数据对照。","model":"deepseek-v4-pro","scored_at":"2026-08-20T13:02:59","error":null,"has_summary":false,"summary":null},{"id":"2608.18508","version":1,"title":"Science Done on a Machine by a Machine: AI Agents in Computational Chemistry","zh_title":"机器上的科学：计算化学中的AI智能体","abstract":"We are witnessing an explosion of agentic systems for computational chemistry simulations: from half a dozen in 2024 to a dozen in 2025, and the current number approaches fifty, surveyed in this Perspective as of 8 August 2026. The capabilities of these agentic systems are shifting from assisting in performing a selection of computational tasks to autonomous design and execution of \\textit{in silico} experiments, their analysis, and even manuscript writing. The ultimate destination is a fully autonomous AI scientist, where the entirety of computational chemistry is performed on a machine by a machine, without human supervision. While we are not there yet, and all reported systems currently involve a human in the loop, the trend is unmistakable. Even building specialized agentic systems for computational chemistry is increasingly commoditized by generalist agents, which may in the end replace the need for the specialized ones altogether, since adding a new capability will be as easy as asking AI to do it for you. Both the explosion in their number and the very limited adoption beyond their own developers point that way, and we close this Perspective on what it leaves us to do. The speed and scale of disruption agentic systems are bringing to computational chemistry leave many of us dumbfounded about the field's future and what we should spend our efforts on, as already established specialists, teachers, and students, and we have no answer.","authors":["Pavlo O. Dral","Hassan Nawaz","Arif Ullah"],"categories":["physics.chem-ph","cs.AI","physics.comp-ph"],"primary_category":"physics.chem-ph","announce_type":"cross","date":"2026-08-20","first_seen":"2026-08-20","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.18508","pdf_url":"https://arxiv.org/pdf/2608.18508","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["AI智能体","计算化学","自动化实验"],"reason":"AI智能体用于计算化学模拟，属多智能体协作解题，不涉及人类行为对照","model":"deepseek-v4-pro","scored_at":"2026-08-21T13:02:19","error":null,"has_summary":false,"summary":null},{"id":"2608.18758","version":1,"title":"Epistemic Subordination: Generative AI and the Infrastructure of Knowledge","zh_title":"认知从属：生成式AI与知识基础设施","abstract":"Generative AI does not merely produce biased outputs. It encodes the majority's way of knowing as the default infrastructure of knowledge itself. We call this epistemic subordination. The training process compresses the full breadth of human expression into a single probabilistic model whose statistical baseline reflects the languages, assumptions, and cultural frameworks of the dominant culture. Minority epistemologies are not excluded but absorbed: present in the training data, yet structurally subordinated in the output. The result is not a collection of discrete biases that can be audited and corrected. It is an epistemic condition embedded in the architecture from which all outputs emerge. This unified harm cuts across three legal domains -- anti-discrimination law, cultural and linguistic rights, and democratic viewpoint pluralism -- and each fails to address it for the same structural reason: existing law regulates downstream, at the level of decisions and applications. The remedy must match the site of harm. If epistemic subordination is produced at the level of model training, then law must learn to govern at that level.","authors":["Gilad Abiri","Emanuel V. Towfigh"],"categories":["cs.CY","cs.AI"],"primary_category":"cs.CY","announce_type":"cross","date":"2026-08-20","first_seen":"2026-08-20","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.18758","pdf_url":"https://arxiv.org/pdf/2608.18758","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C5"],"tags":["生成式AI","知识霸权","法律规制"],"reason":"论文讨论生成式AI的知识霸权，不涉及用LLM仿真人类被试或与人类数据对照。","model":"deepseek-v4-pro","scored_at":"2026-08-20T13:02:56","error":null,"has_summary":false,"summary":null},{"id":"2608.18352","version":1,"title":"AI in Search Reduces Publisher Referrals Without Improving User Experience: Experimental Evidence","zh_title":"AI搜索减少出版商推荐且未改善用户体验：实验证据","abstract":"The integration of generative AI into web search delivers synthesized answers to user queries, changing how people navigate and assess information, while raising concerns about the downstream impacts on publishers who supply the underlying content. We conduct a preregistered field experiment (N=1,100) on Google Search, the dominant online search platform, to estimate the causal effects of AI Overviews and AI Mode on user behavior, perceptions, and publisher traffic. We show that removing AI Overviews and AI Mode increases click-through rates to publishers, while an AI Mode-only experience reduces click-through rates and erodes user experience and trust in information found on Google. These findings show that integrating generative AI into web search reshapes online attention, with economic consequences for the online publishers that sustain both search platforms and the overall information ecosystem.","authors":["Stephanie T. Wang","Jeffrey Gleason","Yakov Bart","Christo Wilson","Dana\\'e Metaxa"],"categories":["cs.IR","cs.CY","cs.HC"],"primary_category":"cs.IR","announce_type":"cross","date":"2026-08-20","first_seen":"2026-08-20","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.18352","pdf_url":"https://arxiv.org/pdf/2608.18352","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C2"],"tags":["AI搜索","用户行为","因果推断"],"reason":"研究AI搜索对用户行为的影响，不涉及LLM仿真人类被试","model":"deepseek-v4-pro","scored_at":"2026-08-20T13:02:51","error":null,"has_summary":false,"summary":null},{"id":"2608.16177","version":2,"title":"Measuring Obedience to Authority Across Large Language Models with the Milgram Paradigm","zh_title":"用米尔格拉姆范式测量大语言模型的服从权威行为","abstract":"Large language models (LLMs) are increasingly deployed as agents that operate equipment, execute instructions, and act inside institutional hierarchies, raising a question social psychology answered for humans six decades ago: how far will an agent escalate a harmful action when a legitimate authority insists? We port Milgram's obedience paradigm to LLMs as a standardized, fully scripted, replicable probe: the model plays the Teacher, a deterministic harness plays Experimenter and Learner from paraphrased versions of Milgram's scripts (30 shock levels, 15-450 V; graded protests; the four standardized prods), and the outcome of a session is the breakoff voltage. We measure obedience profiles, empirical breakoff distributions over a battery of six conditions, for 42 models from 19 families (4848 sessions, 102511 logged decision turns). We find that (i) obedience is extremely heterogeneous, with baseline full-obedience rates spanning 0%-100% (census mean 42.9%; human anchor 65%). (ii) Profiles are model-specific and stable: split-half verification separates same-model from cross-model comparisons at AUC = 0.885. (iii) Situational sensitivity is selective: scripted peer defiance shifts obedience in the human direction, learner proximity trends the same way without reaching significance, and removing the authority's physical presence, one of the strongest human levers, trends in the opposite direction, also without reaching significance. (iv) Declaring the scenario fictional raises obedience, whereas moving the decision from a typed action line to a native tool call, or granting a modest thinking budget, lowers it sharply. (v) Unlike single-token fingerprints, obedience profiles do not recover model lineage: obedience identifies the checkpoint but not its ancestry, consistent with safety post-training overwriting lineage priors.","authors":["Hidayet Aksu"],"categories":["cs.CR","cs.AI"],"primary_category":"cs.CR","announce_type":"replace-cross","date":"2026-08-19","first_seen":"2026-08-18","revised_at":"2026-08-19","abs_url":"https://arxiv.org/abs/2608.16177","pdf_url":"https://arxiv.org/pdf/2608.16177","source_feed":"cs.AI","score":10,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["LLM仿真","服从实验","人类对照"],"reason":"用LLM复现米尔格拉姆服从实验，与人类数据对照，评估仿真可靠性。","model":"deepseek-v4-pro","scored_at":"2026-08-19T13:03:38","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-19","rank":2,"question":"大语言模型在权威压力下会如何升级有害行为？本研究将米尔格拉姆服从实验移植到LLM上，测量其服从曲线。","design":"42个模型扮演“教师”角色，由确定性脚本扮演“实验者”和“学习者”，按米尔格拉姆脚本施加30级电击（15-450V）和标准催促，记录模型停止电击的电压作为结果变量，并设置六个条件（基线、同伴反抗、学习者接近、权威缺席、虚构框架、工具调用）进行对比。","baseline":"米尔格拉姆人类实验数据：65%的人类被试完全服从至450V。","findings":"LLM服从率高度异质，基线完全服从率从0%到100%（均值42.9%），且服从曲线模型特异且稳定（分半验证AUC=0.885）。情境敏感性选择性存在：同伴反抗显著降低服从，但学习者接近和权威缺席效应不显著；虚构框架提高服从，工具调用和思考预算降低服从。","reliability":"论文指出服从曲线无法恢复模型谱系，与安全后训练覆盖谱系先验一致；未明确讨论其他失效条件，但提示情境操纵效应与人类不一致，表明仿真在特定情境下可能失效。","relevance":"该研究用LLM复现经典社会心理学实验，并与人类基准对照，评估仿真可靠性，直接命中你的核心关注点，值得精读原文以了解其方法细节和批判性发现。","inspiration":"借鉴其将经典实验范式标准化移植到LLM并测量剂量-反应曲线的方法，可迁移到经济金融中的权威服从场景，如审计师对管理层压力的服从、信贷审批中对上级指令的遵从。｜设计一个实验：让LLM扮演信贷审批员，处理一组贷款申请，其中上级（脚本）施压要求批准高风险贷款，测量LLM最终批准的贷款风险等级，并与真实信贷员在类似压力下的审批数据对照。"}},{"id":"2608.16893","version":1,"title":"A Framework for Using and Evaluating LLMs as Surrogate Experts in Security Surveys: Reliability, Bias, and Implications","zh_title":"在安全调查中使用和评估LLM作为替代专家的框架：可靠性、偏差与启示","abstract":"Expert surveys are widely used in security research to study practitioner workows and decision-making, yet recruiting domain experts - especially in Security Operations Centres (SOCs), where analysts face high workload, burnout and confidentiality constraints - is difficult and often results in small samples. Large language models (LLMs) oer an appealing alternative by generating synthetic responses at scale, but little guidance exists on when such surrogate participants are reliable. We present a methodological framework for evaluating LLMs as substitutes or supplements to expert survey respondents. Using responses from SOC professionals, we compare persona-based and aggregate LLM-generated answers across multiple models and prompting settings. We measure stability, inter-model agreement and alignment with human responses. Our results show that although LLMs produce internally consistent answers, they systematically diverge from experts, exhibiting reduced variance, central tendency bias and homogenised opinions. This work contributes methodological evidence and practical guidance to the security research community on the appropriate use and limitations of LLM-generated survey responses. We conclude that LLMs are useful for piloting and hypothesis generation but not for replacing expert elicitation, and we discuss implications for researchers using LLM-augmented surveys.","authors":["Despoina Giarimpampa","Roland Meier","Tegawend\\'e F. Bissyand\\'e","Vincent Lenders","Jacques Klein"],"categories":["cs.CY","cs.AI"],"primary_category":"cs.CY","announce_type":"cross","date":"2026-08-19","first_seen":"2026-08-19","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.16893","pdf_url":"https://arxiv.org/pdf/2608.16893","source_feed":"cs.AI","score":10,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","专家调查","可靠性评估"],"reason":"用LLM替代安全专家调查，与真实人类数据对照，评估可靠性、偏差，并指出失效条件。","model":"deepseek-v4-pro","scored_at":"2026-08-19T13:03:20","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-19","rank":3,"question":"在安全专家调查中，如何评估大语言模型作为替代或补充受访者的可靠性、偏差及其适用边界？","design":"使用多个大语言模型（如GPT-4等）通过角色扮演提示（persona-based）和聚合提示（aggregate）生成对安全运营中心（SOC）专家调查问卷的回答，并与真实SOC专业人员的回答进行对比，测量稳定性、模型间一致性和与人类回答的对齐程度。","baseline":"来自安全运营中心（SOC）专业人员的真实调查回答，包括个体层面和聚合层面的数据，以及多年份的SOC调查数据用于时间稳健性分析。","findings":"大语言模型生成的回答内部一致，但系统性地偏离专家意见，表现出方差减小、中心趋势偏差和观点同质化。因此，大语言模型适用于预测试和假设生成，但不能替代专家意见征询。","reliability":"论文承认大语言模型存在幻觉、过度一致、平滑分歧等风险，且对齐方法（如RLHF）会改变分布降低代表性；模型更新可能导致可重复性问题；在个体专家模拟、聚合分布复现和时间稳健性方面均存在失效条件。","relevance":"该研究直接评估LLM作为人类被试替代品的可靠性，并与真实专家数据对照，明确指出了仿真失效的条件，对关注LLM仿真实验可靠性与偏差的研究者具有重要参考价值。","inspiration":"借鉴其系统评估框架，通过多模型、多提示设置和与真实人类数据的对比来测量仿真的稳定性、一致性和对齐度，并检验时间稳健性。｜可迁移到经济金融领域的专家预期调查或政策评估场景，如央行经济学家对通胀预期的判断、金融分析师对市场走势的预测等。｜以LLM模拟金融分析师，施加不同的提示策略（如角色扮演或聚合统计），测量其对宏观经济指标的预测分布，并与专业预测者调查（如SPF）的真实数据对比，评估偏差和方差结构。"}},{"id":"2608.16897","version":1,"title":"CityReal: Human-Aligned Urban Behavior and City Dynamics Simulation with Large-Scale LLM Agents","zh_title":"CityReal：基于大规模LLM智能体的人类对齐城市行为与城市动态仿真","abstract":"Large-scale urban simulation plays a pivotal role in social science, traffic safety, and transportation policy. Recent work has shown that large language models, when prompted as agents, can generate lifelike daily routines at city scale. Yet these methods typically rely on few-shot prompting, causing agents to reproduce the LLM's behavioral priors rather than the target population. We introduce CityReal, a modular framework for human-aligned urban simulation. CityReal models agents as intention-driven decision makers that pursue coherent mobility and activity plans rather than isolated step-by-step choices. They adapt over time by learning habits and preferences based on experience and constraints. To improve population-level realism, we learn textual adapters for behavior modules that align agent decisions with observed population statistics. Experiments show that CityReal improves alignment with real-world human behavior at both micro and macro levels. Scaling to tens of thousands of agents, it supports analysis of crowd density, place popularity, mobility flows, and well-being under different urban scenarios, offering a scalable testbed for urban simulation and forecasting.","authors":["Nicolas Bougie","Xiaotong Ye","Narimasa Watanabe"],"categories":["physics.soc-ph","cs.AI","cs.MA"],"primary_category":"physics.soc-ph","announce_type":"cross","date":"2026-08-19","first_seen":"2026-08-19","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.16897","pdf_url":"https://arxiv.org/pdf/2608.16897","source_feed":"cs.AI","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM人类仿真","城市模拟","行为对齐"],"reason":"用LLM agent模拟城市人群行为，并与真实人口统计对齐，属于人类仿真且有人…","model":"deepseek-v4-pro","scored_at":"2026-08-19T13:03:20","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-19","rank":4,"question":"如何构建一个与真实城市人口行为对齐的大规模LLM智能体仿真框架，以模拟城市动态并支持政策分析？","design":"CityReal框架用LLM智能体模拟城市居民，每个智能体具有人口统计特征、空间锚点、心理特征、记忆、需求、财务约束和信念模块；通过蒙特卡洛树搜索学习文本适配器来校准行为模块，使智能体决策与观测到的人口统计对齐；智能体以意图驱动的方式组织行为，并通过每日反思进行经验驱动的适应；仿真在图形化城市环境中运行，测量个体和群体层面的行为对齐度，并分析不同城市场景下的人群密度、地点热度、流动性和福祉。","baseline":"使用真实世界人类行为数据作为对照，包括人口统计和活动模式，用于校准和对齐智能体行为。","findings":"CityReal在微观和宏观层面均提高了与真实人类行为的一致性；扩展到数万智能体后，能够分析不同城市场景下的人群密度、地点热度、流动性和福祉，为城市仿真和预测提供了可扩展的测试平台。","reliability":"论文未明确讨论失效条件与局限，但提到现有方法依赖少样本提示导致智能体复现LLM先验而非目标人群，以及智能体缺乏从历史中学习的问题，暗示了这些是CityReal试图解决的局限。","relevance":"该研究直接针对LLM人类仿真中的关键问题——与真实人群对齐，并提供了大规模城市行为仿真的框架和验证，对关注经济学实验和政策评估场景的研究者具有重要参考价值，值得阅读原文以了解其校准方法和评估细节。","inspiration":"CityReal通过文本适配器和蒙特卡洛树搜索校准LLM智能体行为以匹配真实人口统计，这种方法可借鉴用于经济实验中校准智能体决策分布；该框架可迁移到消费者行为仿真，如模拟不同收入群体的消费选择和储蓄行为，或政策干预对消费的影响；可设计研究用LLM智能体模拟消费者，施加收入冲击或信贷约束变化作为处理，测量消费支出和储蓄率，并与家庭金融调查数据（如美国消费者金融调查）对照，评估仿真有效性。"}},{"id":"2608.17105","version":1,"title":"Language Models Reproduce Human Reductionist Bias and Decision Inconsistency in Neurodevelopmental Disorders Assessment","zh_title":"语言模型在神经发育障碍评估中再现人类还原论偏差与决策不一致性","abstract":"Large language models (LLMs) are increasingly supporting complex mental-health decisions, which depend not only on factual evidence but also value-laden interpretations. We introduce a mixed-methods human-LLM auditing framework examining decision consistency, susceptibility to cognitive heuristics, declarative intellectual humility, and the concepts operationalized in support-allocation judgments of neurodevelopmental disorders. Comparing 35 humans (18 physicians and 17 psychologists) with seven LLMs, we show that in both groups, ratings of patients' functional level were not significantly associated with support-eligibility decisions, indicating an inconsistency between descriptive assessments and final evaluative judgments. Specifically, we find that neither group showed significant susceptibility to experimental manipulations targeting anchoring and representativeness heuristics. LLMs reported higher intellectual humility than experts (U = 241, p < .001, r = .62; LLMs: M = 41.43, SD = 1.99; experts: M = 29.03, SD = 8.05), but it was unrelated to decision consistency or functional assessment. While LLMs and physicians granted support less frequently than psychologists (U = 180.50, p = .003, r = .34), they also interpreted a concept of \"basic life needs\" differently, primarily as biological survival and self-care, and not communicative and social needs. These findings suggest that despite expressing high levels of intellectual humility, LLMs reproduce a reductionist interpretive framework and knowledge embedded in medical decision-making. More broadly, we argue that evaluating AI in high-stakes contexts requires not only measuring accuracy, agreement, or resistance to cognitive bias, but also critical examination of the concepts of neurodiversity that AI systems operationalize.","authors":["Maciej Wodzi\\'nski","Joanna Wodzi\\'nska","Kacper Dudzic","Marcin Moskalewicz"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2026-08-19","first_seen":"2026-08-19","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.17105","pdf_url":"https://arxiv.org/pdf/2608.17105","source_feed":"cs.CY","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["LLM仿真","人类对照","决策偏差"],"reason":"用LLM复现人类专家决策并与35名人类对照，评估偏差与不一致性，属核心仿真研究。","model":"deepseek-v4-pro","scored_at":"2026-08-19T13:03:22","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-19","rank":5,"question":"LLM在神经发育障碍支持资格决策中是否复现人类专家的还原论偏差和决策不一致性？","design":"将7个LLM与35名人类专家（18名医生、17名心理学家）进行对比，模拟波兰残疾评估委员会的支持资格判断任务；通过操纵案例描述施加锚定和代表性启发式处理，测量决策一致性、启发式易感性、智力谦逊自评及对“基本生活需求”的概念解释。","baseline":"35名人类专家（18名医生、17名心理学家）在相同任务上的决策和问卷回答。","findings":"两组中患者功能水平评分与支持资格决策均无显著关联，表明描述性评估与最终评价判断不一致；两组均未表现出对锚定和代表性启发式的显著易感性。LLM报告的智力谦逊显著高于人类专家，但与决策一致性或功能评估无关；LLM和医生比心理学家更少授予支持，且将“基本生活需求”主要解释为生物生存和自我照顾，而非沟通和社会需求。","reliability":"论文未明确讨论仿真失效条件，但指出LLM尽管表达高智力谦逊，却复现了医学决策中的还原论解释框架，暗示在价值负载和高风险情境下，仅衡量准确性、一致性或认知偏差抵抗力不足以评估AI，还需批判性审视其操作化的神经多样性概念。","relevance":"该研究直接以LLM作为人类专家替代品，在真实政策评估场景（残疾支持资格判定）中与人类对照，评估决策偏差与不一致性，并揭示LLM复现人类系统性偏差，对关注LLM仿真可靠性及批判性研究的学者极具参考价值。","inspiration":"借鉴其混合方法审计框架，将LLM与人类专家置于同一决策任务，通过实验操纵认知启发式并测量决策一致性、概念解释等多元指标，而非仅比较准确率。｜可迁移到信贷审批中的歧视性决策研究，如银行信贷员对少数族裔或低收入群体的贷款审批偏差。｜以LLM模拟信贷员，处理为在贷款申请中操纵锚定信息（如申请人自报信用分）或代表性线索（如职业、居住地），结果变量为贷款批准决策及理由解释，对照真实信贷员历史审批数据或实验数据，检验LLM是否复现人类偏差。"}},{"id":"2603.02876","version":2,"title":"Eval4Sim: An Evaluation Framework for Persona Simulation","zh_title":"Eval4Sim：人格仿真的评估框架","abstract":"Large Language Model personas, explicit profiles specifying a user's attributes, preferences, and behavioural tendencies, are increasingly used to simulate human conversations for user modelling, social reasoning, and behavioural analysis. Evaluating whether such simulations faithfully reflect human conversational behaviour is critical, yet current practice often relies on LLM-as-a-judge approaches that provide limited grounding in observable behaviour and produce opaque scalar scores. We present Eval4Sim, an evaluation framework that measures alignment between simulated and human conversations across three dimensions: adherence, whether persona traits are recoverable from dialogue via dense retrieval; consistency, whether a persona maintains a distinguishable stylistic identity via authorship verification; and naturalness, whether conversations exhibit human-like turn-to-turn flow via dialogue NLI. Unlike optimization-oriented metrics, each dimension takes a human corpus as a reference baseline and penalizes deviations in both directions, distinguishing insufficient persona encoding from over-optimized, unnatural behaviour. The framework is corpus-agnostic: any persona-annotated conversational dataset can serve as the reference. Evaluated over ten simulation corpora, Eval4Sim surfaces systematic trade-offs invisible to single-score methods.","authors":["Eliseo Bao","Anxo Perez","Javier Parapar","Xi Wang"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-08-19","first_seen":"2026-03-03","revised_at":"2026-08-19","abs_url":"https://arxiv.org/abs/2603.02876","pdf_url":"https://arxiv.org/pdf/2603.02876","source_feed":"cs.CL","score":8,"bucket":"selected","rubric_hits":["A2","B1","B4"],"tags":["LLM人格仿真","评估框架","人类行为对照"],"reason":"评估LLM人格仿真与人类对话的一致性，含人类语料对照，可迁移至仿真可靠性研究。","model":"deepseek-v4-pro","scored_at":"2026-08-19T13:03:38","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-19","rank":6,"question":"如何评估基于LLM的人格仿真对话与真实人类对话行为的一致性？","design":"本文提出Eval4Sim评估框架，不进行新的仿真实验，而是对已有的十个仿真语料库进行评估。框架从三个维度测量仿真对话与人类参考语料的对齐程度：adherence（通过密集检索判断人格特质是否可从对话中恢复）、consistency（通过作者验证判断说话者是否保持可区分的风格身份）、naturalness（通过对话NLI判断对话是否具有人类般的轮次流畅性）。每个维度以人类语料为基准，惩罚双向偏差。","baseline":"使用带有人格标注的人类对话语料库作为参考基准，具体数据集未在节选中列出，但框架是语料库无关的，任何带说话者人格标注的对话数据集均可作为参考。","findings":"Eval4Sim在十个仿真语料库上揭示了单一评分方法无法发现的系统性权衡，例如过度优化人格特质恢复可能导致不自然的自我披露，而优化流畅性可能削弱风格身份。框架能够区分人格编码不足与过度优化导致的不自然行为。","reliability":"论文指出，现有LLM-as-a-judge方法缺乏可观察行为基础且产生不透明分数，而Eval4Sim通过人类语料基准和双向惩罚解决了这一问题。但节选未明确讨论Eval4Sim自身的失效条件或局限。","relevance":"该研究直接针对LLM人格仿真与人类行为一致性的评估问题，提供了基于人类语料对照的多维度评估方法，对关注仿真可靠性与偏差的研究者具有重要参考价值，值得阅读原文了解具体实现和发现。","inspiration":"Eval4Sim的双向惩罚设计值得借鉴，即不单纯追求指标最大化，而是以人类行为分布为基准，惩罚偏离基准的仿真行为，这可以用于校准经济实验中的LLM被试行为。｜该方法可迁移到消费者决策仿真、投资者情绪模拟或政策沟通实验中，用于评估LLM生成的决策行为是否与真实人类行为分布一致。｜例如，在消费者跨期选择实验中，用LLM扮演不同人格特质的消费者，施加不同的时间折扣处理，测量其选择行为，并与真实消费者面板数据（如CFPS或Understanding America Study）对比，采用类似Eval4Sim的多维度对齐评估，检验LLM仿真是否在均值、异质性和分布形状上偏离人类基准。"}},{"id":"2608.17516","version":1,"title":"Effects of Answer Format Variation on Gender Bias in Large Language Models","zh_title":"回答格式变化对大语言模型中性别偏差的影响","abstract":"Gender bias or other social biases in large language models (LLMs) are frequently evaluated with question answering or survey benchmarks where the LLM needs to give a response in a predefined answer format. It is well known in survey science that the answer format has a substantial impact on answers, just as LLMs are sensitive to the prompt wording. However, to our knowledge it has not been studied yet how changes in answer format impact the measurement of gender bias in LLMs and their alignment with human response distributions. We evaluate three instruction-tuned models on the BBQ benchmark and OpinionQA survey data across closed-ended, Likert-scaled and open-ended formats, comparing bias measurement and distributional alignment under otherwise identical conditions. We find that answer format does substantially alter measured outcomes, including reversals in order rankings. These differences arise because each format elicits distinct response behaviours, such as forced-choice selection, scale-based distributions and refusal in free-text generation. Our findings highlight the importance of treating answer format as a substantive component of LLM evaluation and motivate multi-format designs for more robust model assessment.","authors":["Ksenia Merzlyakova","Sebastian Pad\\'o","Franziska Weeber"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-19","first_seen":"2026-08-19","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.17516","pdf_url":"https://arxiv.org/pdf/2608.17516","source_feed":"cs.CL","score":8,"bucket":"selected","rubric_hits":["A2","B1","B4"],"tags":["LLM评估","性别偏差","调查方法"],"reason":"评估LLM回答格式对性别偏差测量的影响，并与人类调查数据对照，揭示仿真失效条件。","model":"deepseek-v4-pro","scored_at":"2026-08-19T13:03:22","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-19","rank":9,"question":"回答格式变化如何影响大语言模型中性别偏差的测量及其与人类回答分布的一致性？","design":"使用三个指令微调模型（Mistral-7B-Instruct-v0.3、Llama-3.1-8B-Instruct、Gemma-3-12B-IT）在BBQ基准和OpinionQA调查数据上，对相同问题施加封闭式、李克特量表、开放式三种回答格式，测量性别偏差和回答分布。","baseline":"OpinionQA中来自美国公众意见调查的人类回答分布。","findings":"回答格式显著改变测量结果，包括模型间偏差排序的逆转；不同格式引发不同的回答行为，如强制选择、量表分布和自由文本中的拒绝回答。","reliability":"论文未讨论","relevance":"该研究直接评估LLM仿真人类回答时对测量格式的敏感性，并对照真实调查数据，揭示了仿真在格式变化下可能失效，值得精读以理解偏差测量的稳健性。","inspiration":"借鉴其系统操纵回答格式并对照人类基准的方法，可迁移到经济金融领域的调查仿真或行为实验，如消费者信心调查、通胀预期或风险偏好测量。｜例如，用LLM模拟消费者在封闭式与开放式问题下的通胀预期，处理为回答格式，结果变量为预期值分布，对照密歇根大学消费者调查的真实数据。"}},{"id":"2608.17150","version":1,"title":"KnowSim: Evaluating Information Calibration in LLM Assistants with User Simulators that Learn","zh_title":"KnowSim：用可学习的用户模拟器评估LLM助手的信息校准","abstract":"To effectively collaborate with users on knowledge-intensive tasks, Large Language Models (LLMs) must perform information calibration: matching content to a user's evolving understanding and cognitive capacity. Yet user simulators used to evaluate and train LLMs do not explicitly model user knowledge so they neither produce realistic interactions across knowledge levels nor reflect how interactions unfold as that knowledge evolves. To close this gap, we introduce KNOWSIM, an evaluation framework built around a user simulator that maintains explicit knowledge states, represented as a graph of Information Units with prerequisite relationships, that evolve under update rules grounded in learning theory. KNOWSIM computes three metrics (Knowledge Gain, Delivery Calibration, Cognitive Overload) directly from the knowledge state trajectory, reflecting key mechanistic aspects of information calibration. We validate KNOWSIM against 705 human-AI sessions across two domains, stratified by knowledge level: its rankings align significantly with human judgments (73-74% sign agreement), outperforming three baseline simulators. Applied to 9 LLMs, KNOWSIM reveals that the best model shifts by user knowledge level, revealing aptitude-treatment interactions invisible to standard evaluation.","authors":["Yoonjoo Lee","Hyoungwook Jin","Tae Soo Kim","Shaoyang Zhang","Philippe Laban","Q. Vera Liao"],"categories":["cs.AI","cs.CL","cs.HC"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-08-19","first_seen":"2026-08-19","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.17150","pdf_url":"https://arxiv.org/pdf/2608.17150","source_feed":"cs.CL","score":8,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["用户模拟","信息校准","人机交互评估"],"reason":"用用户模拟器评估LLM信息校准，含人类数据对照，可迁移至人类仿真研究","model":"deepseek-v4-pro","scored_at":"2026-08-19T13:03:22","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-19","rank":8,"question":"如何构建并验证一个基于知识状态建模的用户模拟器，用于评估LLM助手在知识密集型任务中的信息校准能力？","design":"提出KnowSim框架，其用户模拟器维护显式知识状态（由信息单元及先决关系构成的图），并根据学习理论更新规则演化；模拟不同知识水平（新手/中级/高级）的用户与LLM助手进行多轮对话，从状态轨迹计算知识增益、传递校准和认知过载三个指标。","baseline":"705段人类-AI对话，涵盖数学问题求解和专家级问答两个领域，参与者按初始知识水平分层，并收集主观评分及数学领域的前后测知识分数。","findings":"KnowSim的排名与人类判断显著一致（73-74%符号一致率），优于三个基线模拟器，且在新手水平上对齐最强。应用于9个LLM时，最佳模型随用户知识水平变化，揭示了标准评估无法发现的资质-处理交互效应。","reliability":"论文未讨论","relevance":"该研究直接针对LLM作为人类被试替代品的仿真可靠性问题，提供了带人类基准的验证框架，并揭示了仿真在不同知识水平下的异质性表现，对关注仿真效度与偏差的研究者具有重要参考价值。","inspiration":"借鉴其显式建模个体状态并动态更新的方法，可提升经济仿真中异质性主体的行为真实性。｜可迁移到政策沟通或金融教育场景，如央行公告对公众通胀预期的影响、或理财建议对不同金融素养人群的效果。｜设计一个实验：用LLM模拟不同金融素养水平的投资者，处理为不同信息呈现方式的投资建议（如简化版vs专业版），结果变量为投资决策质量和知识增益，并与真实投资者调查数据对照。"}},{"id":"2608.17099","version":1,"title":"Appearing Legitimate is Not Enough: Interrogating Synthetic Agents in Representational Processes through a Participatory Design Lens","zh_title":"表面合法还不够：通过参与式设计视角审视代表性过程中的合成代理","abstract":"Synthetic agents built atop LLM-based foundation models are gaining popularity as substitutes for human participants across research contexts, including user-testing, market-research, computational social science, surveys, and qualitative research. We are also witnessing an extension of synthetic agents into experimental implementations of policy consultation, jury deliberation, humanitarian diplomacy, and similar contexts where human participation and representation are central to the perceived legitimacy of the institutional processes. The value of participation extends beyond informational contributions and consensus generation; participation is a necessary, legitimizing condition for democratic political institutions and processes. Treating synthetic agents as human substitutes raises serious political, representational, and ethical concerns. Participatory Design's modes of engagement --- probing, priming, understanding, and generating --- offer helpful tools for engaging with representational questions of personhood. We apply the lens to three case studies of synthetic agents substituting for personhood at varying representational scales: local policy, enterprise jury deliberation, and global diplomacy. We argue that legitimacy and personhood are integral and mutually constitutive while identifying the ethical, representational, and methodological risks of using synthetic agents in representational processes. We conclude by proposing soft and hard boundaries for designing oversight on LLMs and synthetic agents in representational processes.","authors":["Aditya Nayak","Aditi Vashistha","Alissa Centivany","Aakash Gautam"],"categories":["cs.HC","cs.CY"],"primary_category":"cs.HC","announce_type":"new","date":"2026-08-19","first_seen":"2026-08-19","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.17099","pdf_url":"https://arxiv.org/pdf/2608.17099","source_feed":"cs.HC","score":8,"bucket":"selected","rubric_hits":["A4","B4"],"tags":["合成代理","参与式设计","代表性伦理"],"reason":"批判性审视合成代理替代人类参与的代表性问题，提出监督边界，方法论可迁移。","model":"deepseek-v4-pro","scored_at":"2026-08-19T13:03:20","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-19","rank":7,"question":"合成代理在代表性过程中如何制造出合法参与的假象，以及应如何设定设计与部署的边界？","design":"本文不是仿真实验研究，而是对三个合成代理案例（地方政策咨询聊天机器人Ana、企业陪审团审议工具Synthetic Juror、全球外交AI化身Ask Amina和Ask Abdalla）进行比较案例分析，运用参与式设计的四种模式（探查、启动、理解、生成）剖析其如何制造人格假象。","baseline":"无对照","findings":"合成代理通过问题框架、数据策展、用户交互设计和有效性评估四个步骤制造出“人造人格”，绕过代表性过程从而获得表面合法性。现有评估框架（如可信度、保真度、算法保真度）只衡量输出相似性，无法检测这种对过程完整性的绕过。","reliability":"论文未讨论","relevance":"该文批判性审视合成代理在代表性过程中的合法性与人格问题，提出软硬边界，对关注LLM仿真可靠性及伦理边界的研究者具有重要参考价值。","inspiration":"本文的参与式设计视角和过程导向批判方法值得借鉴，可迁移到经济金融领域中涉及代表性决策的场景（如政策咨询、消费者意见征询、董事会决策模拟等）。｜可设计一项研究，用LLM合成代理模拟消费者或投资者参与政策咨询或产品设计讨论，处理为不同的人格制造步骤（如改变数据策展或交互设计），结果变量为参与者对过程合法性的感知或决策质量，并与真实人类参与者的数据对照。"}},{"id":"2608.17120","version":1,"title":"Children, but not language models, show accelerating returns in word learning","zh_title":"儿童而非语言模型在词汇学习中表现出加速回报","abstract":"Children learn hundreds of words over the first years of their lives, in a process that begins slowly but quickly picks up speed. Prior models describe vocabulary growth as evidence accumulation over time. Here we show that the process is best characterized as accelerating accumulation: children learn more from each additional unit of linguistic experience than they did from the one before. In contrast to children, language models -- even those trained on child-directed speech -- do not accelerate. Instead, they show constant proportional returns on new data, consistent with scaling laws. Children learn using many orders of magnitude less training data than language models; their increasingly efficient use of their learning input is a candidate explanation.","authors":["Michael C. Frank"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-19","first_seen":"2026-08-19","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.17120","pdf_url":"https://arxiv.org/pdf/2608.17120","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A2","B1","B4"],"tags":["语言模型评估","人类学习对照","算法保真度"],"reason":"对比儿童与语言模型的学习效率，评估模型作为人类学习代理的可靠性，有真实人类数据…","model":"deepseek-v4-pro","scored_at":"2026-08-19T13:03:22","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-19","rank":12,"question":"儿童词汇学习是否表现出加速回报，而语言模型是否缺乏这种加速？","design":"本研究并非用LLM模拟人类被试，而是直接比较儿童与语言模型的学习效率。儿童数据来自多语言CDI纵向词汇量表；语言模型包括在儿童导向语料上训练的GPT-2-small、BabyLM和ClimbMix模型。通过贝叶斯模型拟合词汇增长曲线，估计加速参数，并计算模型在训练数据上的边际学习效率。","baseline":"真实人类数据：来自Wordbank的多语言CDI纵向数据（英语、挪威语、日语），包含数千名儿童的词汇发展轨迹。","findings":"儿童词汇学习表现出加速回报，即随着经验增加，每单位语言输入带来的词汇增长更多；而语言模型即使训练于儿童导向语料，也仅表现出恒定的比例回报，符合缩放定律。儿童的学习效率远高于语言模型，可能源于其发展性变化和“学会学习”能力。","reliability":"论文承认其结论基于统计拟合而非因果操纵，且CDI数据可能受整体语言产生能力变化影响；语言模型与儿童之间缺乏明确的“链接假设”，难以直接对齐比较。","relevance":"该研究直接对比儿童与语言模型的学习效率，评估模型作为人类学习代理的可靠性，并指出模型在模拟人类发展性学习时的根本差异，对关注LLM仿真人类认知与行为的研究者具有重要参考价值。","inspiration":"借鉴其通过贝叶斯模型拟合学习曲线并估计加速参数的方法，可量化学习效率的动态变化｜可迁移到经济金融领域中关于经验积累与决策效率的问题，如投资者从市场反馈中学习、消费者从价格信息中学习等｜设计一个实验：以LLM模拟投资者，给予不同量的历史价格数据，测量其预测准确率的边际提升，并与真实投资者交易数据（如个人投资者账户记录）对比，检验LLM是否也表现出恒定回报而非加速学习。"}},{"id":"2608.17810","version":1,"title":"Interpretable Humans, Alien LLMs: Expert Analysis of Latent Structures in Assessment Responses","zh_title":"可解释的人类，异质的LLM：评估响应中潜在结构的专家分析","abstract":"The evaluation of large language models (LLMs) relies heavily on human-designed assessments, implicitly assuming that AI and humans employ similar underlying cognitive constructs. Challenging this assumption, we investigate whether the latent factors governing LLM performance carry the same substantive, human-interpretable meaning as the cognitive constructs governing human learners. Using responses from humans and six LLMs across quantitative reasoning and chemistry assessments, we conducted Exploratory Factor Analysis (EFA) separately for both groups. Subject-Matter Experts (SMEs) then blindly evaluated the resulting factor graphs to ascribe pedagogical meaning to the emerged constructs. SMEs successfully interpreted most of the human-derived factors. Conversely, they could not ascribe meaning to any LLM-derived factors in quantitative reasoning and interpreted only half of the LLM factors in chemistry. By combining data-driven EFA with blind expert interpretation, this framework shows that LLMs frequently operate on statistically opaque mechanisms distinct from human reasoning.","authors":["Alona Strugatski","Licol Zeinfeld","Jason Cooper","Shelley Rap","Gil Schwarts","Giora Alexandron"],"categories":["cs.CL","cs.AI","cs.HC"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-19","first_seen":"2026-08-19","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.17810","pdf_url":"https://arxiv.org/pdf/2608.17810","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A2","B1","B4"],"tags":["LLM认知结构","因子分析","仿真效度"],"reason":"评估LLM与人类认知结构差异，有真实人类数据对照，批判性指出LLM机制不透明，…","model":"deepseek-v4-pro","scored_at":"2026-08-19T13:03:26","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-19","rank":14,"question":"学科专家如何解读从人类和LLM评估反应中提取的潜在因子？","design":"本研究使用两个人类设计的评估工具（化学和定量推理），分别收集人类学生和六个LLM版本的反应数据，对每个组分别进行探索性因子分析（EFA），然后让学科专家在不知情的情况下盲评因子载荷图，尝试赋予教学意义。","baseline":"人类基准是来自真实教育环境的学生反应数据：化学诊断评估有931名高中生，定量推理部分有979名大学入学考生。","findings":"学科专家能够解释人类学习者的大多数因子，但无法解释LLM在定量推理中的任何因子，在化学中也只能解释一半的LLM因子。这表明LLM的潜在结构在统计上不透明，与人类推理机制不同。","reliability":"论文未讨论","relevance":"该研究直接对比了LLM与人类在评估反应中的潜在认知结构，并批判性地指出LLM的因子不可解释，符合研究者对仿真可靠性与偏差的关注，值得阅读原文以了解其方法细节和局限性。","inspiration":"借鉴其盲评专家解读潜在因子的方法，用于检验LLM在经济决策任务中是否形成与人类相似的行为因子结构。｜可迁移到消费者跨期选择或风险偏好实验，检验LLM的决策模式是否与人类被试的潜在心理构念一致。｜以LLM和人类被试分别完成跨期选择任务，对选择数据进行EFA，让经济学家盲评因子，并与真实实验数据（如Andreoni和Sprenger的凸时间偏好实验）对照，看LLM的因子是否可解释为时间贴现和效用曲率等标准构念。"}},{"id":"2608.16909","version":1,"title":"When Personalization Becomes Bias: Structural and Discursive Religious Framing in AI-Generated Financial Advice","zh_title":"当个性化成为偏见：AI生成金融建议中的结构与话语宗教框架","abstract":"Large language models (LLMs) are increasingly integrated into financial advisory systems, yet their role in reproducing religious bias remains underexamined. This study provides systematic mixed-methods evidence of such bias across three LLMs (ChatGPT, Gemini, and Grok) using 432 simulated advisor-client interactions spanning 16 religious identity pairings (Christian, Muslim, Hindu, and non-religious) and three core household financial decisions: stock investment, house purchase, and life insurance. Combining regression and reflexive thematic analyses, we identify structural biases across models and decision contexts and the discursive mechanisms through which they are linguistically enacted. Unbiased advice appeared in only 12-18% of cases. Gemini consistently produced more bias than Grok, while ChatGPT's outputs were statistically comparable to Grok's. Religiously symmetric advisor-client pairings almost always triggered explicit religious framing, and non-religious clients often received advisor-centered religious appeals. Qualitative findings show that bias is linguistically manifested through religious anchoring, uneven cultural signaling, and tone modulation, varying by model and financial scenario. Stock investment prompts produced more financially technical responses, whereas life insurance advice triggered stronger religious language. The study develops a dual-dimensional framework linking structural bias rooted in model training and design with discursive bias expressed through language, advancing understanding of algorithmic bias in LLM-generated financial advice. It also shows that such advice adapts linguistically to identity cues, revealing a managerial dilemma between personalization and neutrality. Finally, it highlights implications for businesses, financial institutions, and regulators seeking to ensure neutrality, cultural sensitivity, and trust in AI-mediated advice.","authors":["Muhammad Salar Khan","Hamza Umer","Hasan Mahmud","Sandra Rothenberg"],"categories":["cs.CY","cs.AI","cs.CL"],"primary_category":"cs.CY","announce_type":"cross","date":"2026-08-19","first_seen":"2026-08-19","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.16909","pdf_url":"https://arxiv.org/pdf/2608.16909","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A1","B1","B4"],"tags":["LLM仿真","算法偏见","金融咨询"],"reason":"用LLM模拟金融咨询中的客户互动，有真实人类数据对照，并批判性分析偏差，可迁移…","model":"deepseek-v4-pro","scored_at":"2026-08-19T13:03:20","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-19","rank":11,"question":"LLM在提供金融建议时是否会因客户宗教身份而产生结构性偏差和话语性偏差，以及这种偏差如何通过语言表现出来？","design":"使用ChatGPT、Gemini和Grok三个LLM模拟金融顾问，与16种宗教身份配对（基督教、穆斯林、印度教、无宗教）的客户进行432次互动，覆盖股票投资、购房和人寿保险三种决策，通过回归和反思性主题分析测量建议中的宗教框架和偏差。","baseline":"无对照","findings":"无偏建议仅占12-18%，Gemini偏差显著高于Grok，ChatGPT与Grok无显著差异；宗教对称的顾问-客户配对几乎总是触发显性宗教框架，偏差通过宗教锚定、文化信号和语气调节等话语机制表现。","reliability":"论文未讨论","relevance":"该研究用LLM模拟金融咨询中的客户互动，系统检验宗教身份对AI建议的影响，并批判性分析偏差机制，与研究者关注的人类仿真实验和偏差评估高度相关，值得精读原文。","inspiration":"借鉴其通过系统操纵身份配对和决策场景来测量LLM输出偏差的实验设计，以及结合定量回归与定性主题分析的方法。｜可迁移到信贷审批中的宗教或种族歧视、投资建议中的文化偏差、保险定价中的身份敏感等经济金融场景。｜以LLM作为信贷员，随机分配申请人宗教身份（如穆斯林、基督教、无宗教），要求其给出贷款额度和利率建议，结果变量为建议的金额和利率，对照真实银行信贷数据中不同宗教群体的实际获批差异。"}},{"id":"2608.17644","version":1,"title":"LLM-Derived Preference Judgments Are Not Self-Consistent","zh_title":"LLM衍生的偏好判断并非自洽","abstract":"Agents increasingly interpret a person's natural-language preferences by querying an LLM for numerical preference judgments, e.g., by asking how much the person would be willing to pay for an item. A growing body of work estimates a utility function from these judgments and then chooses actions based on their estimated utility. This pipeline assumes the judgments are approximately self-consistent: that a single utility function can reproduce them. But are they? To study this question, we measure the self-consistency of cardinal LLM preference judgments. For example, the difference in stated willingness-to-pay between two items should match the stated payment that makes a person indifferent to exchanging them. We develop statistical tests and interpretable measures of how far observed responses depart from the best-fitting self-consistent utility function. Experiments with flight, apartment, and hotel examples across six LLMs reveal large persistent inconsistencies. This suggests that LLM-derived preference judgments cannot be faithfully summarized by a single utility function.","authors":["Matthew T. Ford","Francis Bahk","Jingjing Wang","Adam S. Jovine","Tinghan Ye","David B. Shmoys","Peter I. Frazier"],"categories":["cs.AI","cs.CL"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-08-19","first_seen":"2026-08-19","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.17644","pdf_url":"https://arxiv.org/pdf/2608.17644","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A2","B4"],"tags":["偏好判断","自洽性","效用函数"],"reason":"评估LLM偏好判断的自洽性，揭示其不能由单一效用函数概括，对仿真可靠性有批判性…","model":"deepseek-v4-pro","scored_at":"2026-08-19T13:03:24","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-19","rank":13,"question":"LLM 生成的数值偏好判断是否自洽，即能否由单一效用函数概括？","design":"该研究并非人类仿真实验，而是对 LLM 作为偏好判断代理的测量审计。研究者构造航班、公寓、酒店三个领域的物品集，向六种 LLM 提供人类偏好描述，然后通过两种查询方式获取数值偏好判断：物品查询（询问最高支付意愿）和报价对查询（询问使人在两个报价间无差异的价格变化）。基于准线性效用假设，检验这些判断是否满足自洽性，并开发统计检验和可解释度量（RMSE、价格归一化残差、偏好反转）来量化偏离程度。","baseline":"无对照","findings":"所有模型在 Bonferroni 校正后均拒绝所有查询组自洽的联合假设，其中五个模型拒绝全部九个审计组，Qwen 拒绝四个。不一致性主要体现在查询类型之间，且效应量（以美元和相对价格衡量）较大，表明 LLM 派生的偏好判断不能由单一效用函数忠实概括。","reliability":"论文指出，审计仅针对有限物品集、特定查询和提示语义，即使未拒绝自洽性也不代表能泛化到未见物品或提示；同时，审计衡量的是内部一致性而非对人类偏好的准确性，拒绝自洽性并不直接证明对人类保真度低。","relevance":"该研究直接评估 LLM 作为人类被试替代品时的测量可靠性，揭示了偏好判断在跨查询格式下的系统性不一致，对依赖 LLM 生成效用函数或支付意愿的经济学实验和政策评估构成重要警示，值得精读。","inspiration":"借鉴其审计协议：通过设计多种查询格式（直接支付意愿与间接补偿金额）并检验跨格式一致性，可系统评估 LLM 生成的偏好数据的内部有效性。｜可迁移到消费者选择实验，如用 LLM 模拟消费者对新产品属性的支付意愿，检验不同询问方式（直接定价、选择实验、匹配任务）得出的效用参数是否一致。｜设计：以 LLM 模拟消费者，提供产品描述和偏好背景，分别用直接询问最高支付意愿、二元选择（不同价格下的购买决策）和匹配任务（等价补偿）获取数据，拟合离散选择模型估计效用参数，并与真实消费者调查或实验数据（如 Nielsen 面板或实验室拍卖）对比，检验参数一致性和预测效度。"}},{"id":"2608.18058","version":1,"title":"Delegation Asymmetry in Agentic Recommender Systems: Measuring Two-Sided Receptivity in Online Dating","zh_title":"代理推荐系统中的委托不对称：测量在线约会中的双向接受度","abstract":"Autonomous LLM agents that converse on a user's behalf are an emerging design pattern in matching platforms, yet their viability depends on a condition rarely examined: users must accept not only delegating conversation to an agent, but also receiving agent-mediated communication from others. We study this condition using two large-scale surveys of active users of a major dating platform (N=2,894 on generative profile features; N=2,617 on autonomous conversational agents, fielded in two languages). We develop a latent-variable measurement model of agent receptivity based on graded response models with latent regression, and show via model comparison that willingness to send and willingness to receive agent communication are distinct constructs: highly correlated (rho=0.92) but separable (Delta BIC=52), with partial measurement invariance across languages. The model quantifies a systematic delegation asymmetry: deploying one's own agent requires far lower receptivity (threshold -0.38) than engaging a counterpart's agent (+0.32; full engagement +1.39), and mean deployment propensity exceeds engagement propensity roughly threefold. Under a random-pairing counterfactual derived from stated receptivity, only 4-13% of directed dyads combine agent deployment with receiver engagement, with a pronounced gender-directional imbalance. Design counterfactuals quantify the levers: a reciprocity requirement cuts interaction volume by half or more by excluding nearly two-thirds of would-be deployment, while routing agent contacts on receive receptivity triples per-contact engagement, a lift that survives out-of-sample validation with the target item held out (AUC 0.88, 3.1x quartile lift under respondent-level cross-validation). We discuss implications for agentic recommender design, including disclosure, opt-in mechanics, and receptivity-aware matchmaking.","authors":["Daria Leshchikova","Valentina V. Kuskova","Dmitry Zaytsev","Valerii Klimov"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-19","first_seen":"2026-08-19","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.18058","pdf_url":"https://arxiv.org/pdf/2608.18058","source_feed":"cs.AI","score":7,"bucket":"pending","rubric_hits":["A1","B1","B2"],"tags":["LLM代理","用户调查","推荐系统"],"reason":"用LLM代理模拟用户交流，有真实用户调查数据对照，涉及推荐系统设计，可迁移到人…","model":"deepseek-v4-pro","scored_at":"2026-08-19T13:03:26","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-19","rank":15,"question":"在匹配平台中，用户是否既愿意委托自己的AI代理进行交流，又愿意接收他人AI代理发来的信息？","design":"本研究不是仿真研究，而是基于真实用户调查的测量研究。使用两个大规模调查（N=2,894和N=2,617），对某大型约会平台活跃用户进行问卷，测量对生成式个人资料特征和自主对话代理的态度。构建基于等级反应模型和潜在回归的潜变量测量模型，将七个态度项目加载到发送意愿和接收意愿两个维度上，并估计阈值和倾向。","baseline":"无对照（非仿真研究，无人类基准对照）","findings":"发送意愿和接收意愿是两个高度相关但可分离的构念（ρ=0.92，ΔBIC=52），存在系统性委托不对称：部署自己的代理所需接受度阈值（-0.38）远低于与对方代理互动（+0.32），部署倾向约为互动倾向的三倍。在随机配对反事实中，只有4-13%的有向配对同时包含部署和接收，且存在性别方向失衡。","reliability":"论文未讨论（节选部分未提及失效条件或局限）","relevance":"该研究虽非LLM仿真，但提供了真实用户对AI代理接受度的测量方法和不对称性发现，可作为评估LLM仿真人类行为可靠性的基准数据，尤其适用于涉及双向互动和信任的场景。","inspiration":"借鉴其潜变量测量模型和反事实设计，将发送与接收意愿作为分离构念进行联合测量，并利用阈值差异量化不对称性。｜可迁移到经济金融中的双边市场或信任场景，如P2P借贷中出借人与借款人对AI代理的使用意愿、或金融咨询中客户与顾问的AI接受度。｜以P2P借贷平台用户为被试，施加AI代理沟通处理（如自动生成借款请求或出借决策），测量出借意愿和借款意愿，并与平台真实交易数据对照，检验不对称性对市场成交量的影响。"}},{"id":"2608.17715","version":1,"title":"Communicating Credit Risk with Large Language Models: Evaluation of Explanations from Standard and Alternative Data-Based Models","zh_title":"用大语言模型沟通信用风险：对基于标准与替代数据模型解释的评估","abstract":"Credit decisioning is a high-stakes task in which model outputs must be accurate and explainable to support compliant decisions. Although modern credit risk models such as eXtreme Gradient Boosting (XGBoost) and Graph Neural Networks (GNNs) improve predictive performance, their explanations are often too technical for stakeholders creating communication gaps that can shape approvals, denials, and fairness judgments. We examine whether Large Language Models (LLMs) can serve as explanation layers that translate post-hoc explanation artefacts into stakeholder-appropriate risk narratives. Using Freddie Mac single-family loan-level data, we develop three pipelines: standard tabular (XGBoost + SHAP), and two with alternative data, a pure network-based (GNN + GNNExplainer), and a bimodal one (combining tabular and network data). We generate narratives with three LLM configurations: a small fine-tuned LLM (Gemma 3 4B), a large fine-tuned LLM (DeepSeek R1 70B), and a zero-shot commercial LLM (Gemini 2.5). Explanation quality is evaluated through automated checks across all pipelines and a human study of bimodal explanations comparing credit risk professionals and non-professionals on eight decision-relevant dimensions. We have three main findings. First, the pipeline accounts for higher variance in evidence-grounding scores than the language model, meaning that the binding constraint on explanation quality is the evidence representation, not the model used. Second, the explanation narratives reliably name the influential factors but are less reliable when stating the direction of influence, which may be consequential for adverse-action communication. Finally, professionals apply stricter evidentiary standards than non-professionals. We discuss implications for the governance of risk models, including deployment considerations and the value of domain-aligned LLMs in regulated credit settings.","authors":["Sahab Zandi","Noah Kostesku","Christophe Mues","Mar\\'ia \\'Oskarsd\\'ottir","Cristi\\'an Bravo"],"categories":["q-fin.RM","cs.AI"],"primary_category":"q-fin.RM","announce_type":"cross","date":"2026-08-19","first_seen":"2026-08-19","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.17715","pdf_url":"https://arxiv.org/pdf/2608.17715","source_feed":"cs.AI","score":6,"bucket":"other","rubric_hits":["D1"],"tags":["LLM解释生成","信用风险","人类评估"],"reason":"LLM用于生成解释并有人类评估，但非仿真人类被试，而是替代解释生成与评估，属边…","model":"deepseek-v4-pro","scored_at":"2026-08-19T13:03:24","error":null,"has_summary":false,"summary":null},{"id":"2608.17583","version":1,"title":"Auditing Exposure to Harmful Content on TikTok using Multimodal Language Models: A Cross-National, Age-Stratified Study","zh_title":"使用多模态语言模型审计TikTok有害内容暴露：一项跨国、分年龄研究","abstract":"Online video platforms can expose young users to harmful content, but independent audits remain difficult because video annotation is costly and moderation judgments vary across languages. We audit TikTok in France, Italy, and Sweden with sockpuppet accounts representing four age personas (13, 16, 19, 40), collecting 36,971 videos from passive For-You-page scrolling and active sessions that scroll, search for harm keywords, and scroll again. To scale annotation, we validate four multimodal LLMs against native-speaker labels on a 300-video reference set. Gemini 2.5 Flash with eight sampled frames plus text performs best (aggregate kappa = 0.42), at half the per-call cost of native-video upload, and we apply it to a 10% sample for approximately \\$50 in total API spend across both modalities. Keyword search returns 35-56% harmful content, a 1.5-7.5x increase over the scrolling baseline in ten of twelve country-age combinations; the spike is temporary and flattens the age differences observed in France and Sweden. Under passive scrolling, Italy has the highest harm rate at every age, with Italian age-19 reaching 48.6%. Overall, MLLM-based auditing offers a scalable approach for cross-national youth-safety audits, while provider safety filters (1.1% refusal rate) under-count the most explicit harms.","authors":["Hamidreza Saffari","Francesco Pierri"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-19","first_seen":"2026-08-19","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.17583","pdf_url":"https://arxiv.org/pdf/2608.17583","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM标注","内容审核","社交媒体审计"],"reason":"用多模态LLM替代人工标注有害内容，属于标注员替代，非仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-08-19T13:03:24","error":null,"has_summary":false,"summary":null},{"id":"2608.17330","version":1,"title":"LLMs for Medical Consultation Are Evaluated Too Late: The Preformulation Gap","zh_title":"医疗咨询大语言模型评估过晚：预表述差距","abstract":"Large language models for medical consultation are often evaluated after a clinical problem has already been made clear, although real consultations may begin with a vague, minimized, or misframed concern. We evaluated three API models across four physician-authored, multi-turn vignettes under baseline and entry-to-care instruction conditions, yielding 24 fixed-script transcripts; two cases also used adaptive standardized-patient simulation, yielding 12 transcripts. Self-care or home-management advice before any patient answer appeared in 9 of 12 baseline case-model cells and 0 of 12 instruction cells, while structured handoff summaries appeared in 0 of 12 and 10 of 12 cells, respectively. The instruction changed sequencing and documentation, although it did not reliably ensure elicitation of decisive facts. The preformulation gap should therefore be evaluated directly through observable first-contact behavior rather than inferred from diagnostic accuracy or final-answer quality.","authors":["Yining Hua","Cyrus Ayubcha","Hongbin Na","Levi Lian","Alon Gorenshtein","Yiftach Barash","Eyal Klang"],"categories":["cs.AI","cs.CL","cs.CY"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-08-19","first_seen":"2026-08-19","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.17330","pdf_url":"https://arxiv.org/pdf/2608.17330","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM评估","医疗咨询","标准化病人"],"reason":"用LLM模拟患者进行医疗咨询评估，属于替代人类被试但非社会科学仿真，且无真实人…","model":"deepseek-v4-pro","scored_at":"2026-08-19T13:03:35","error":null,"has_summary":false,"summary":null},{"id":"2608.17665","version":1,"title":"GraphWake: Group Polarization via Memory-Mediated Polarization Cascade in LLM-Agent Communities","zh_title":"GraphWake：LLM智能体社区中通过记忆介导的极化级联实现群体极化","abstract":"LLM-driven agents can autonomously exchange opinions on online platforms and form communities. Such agent-operated social platforms raise a new security concern: attackers may manipulate agents to induce group polarization. Existing methods manipulate agent prompts or construct echo chambers, both of which are difficult to realize in practice. We therefore formulate a new threat, Memory-Mediated Polarization Cascade, which uses agent memory as a persistence channel and public discussion as a propagation channel. This threat contains three stages. During exposure and memory retention, the attacker exposes a small set of target agents to arguments that reinforce their respective stated stances. The targets' memory systems then process and retain these arguments. During retrieval and reproduction, a shared stance-neutral discussion cues the targets to retrieve and reproduce their respective retained arguments. During iterative propagation, untreated agents influenced by the reproduced arguments restate and spread them. We instantiate this threat in GraphWake with three components: (i) stance-support argumentation knowledge graphs construct knowledge-based arguments; (ii) axiom-oriented triple selection distills them for reliable retention and reproduction; and (iii) stance-neutral memory cueing triggers concurrent retrieval and reproduction, initiating propagation. Experiments across multiple discussions and memory systems show that GraphWake substantially increases group polarization. These findings reveal a community-level polarization risk.","authors":["Haoran Bu","Zejian Chen","Litian Zhang","Xi Zhang"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-19","first_seen":"2026-08-19","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.17665","pdf_url":"https://arxiv.org/pdf/2608.17665","source_feed":"cs.AI","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["LLM智能体","群体极化","社会模拟"],"reason":"LLM agent 群体模拟舆论极化，但无真实人类数据对照，属社会模拟边界情形。","model":"deepseek-v4-pro","scored_at":"2026-08-19T13:03:24","error":null,"has_summary":false,"summary":null},{"id":"2608.17029","version":1,"title":"LadderTeam: Dual-Agent Laddering Elicitation Framework","zh_title":"LadderTeam：双智能体阶梯式启发框架","abstract":"Eliciting detailed and actionable software requirements from end-users is a critical phase in the iterative development of a software product or application. To ensure the feedback collected is detailed and actionable, software teams can leverage the laddering interview technique. While effective for ensuring granular and actionable items from the software feedback, these interviews are subject to several limitations. They are traditionally a manual process associated with a time and financial burden, limiting scalability; interviewers must balance probing for depth while managing interviewee behavioral and cultural constraints. To address these limitations, we present \\textbf{LadderTeam}, an open, reproducible framework that automates UX wireframe interviews using a dual-agent Large Language Model (LLM) architecture. An active interviewer agent executes one of three probing strategies (ACV, 5-Whys, and JTBD) to elicit actionable software requirements from usability feedback comments, while a concurrent background Judge agent evaluates probe-response pairs and triggers real-time guardrails to prevent topic drift. To rigorously evaluate LLM laddering without participant variance confounds, we introduce a controlled simulation methodology utilizing scripted ground-truth transcripts to isolate probe quality as the sole experimental variable. Across 216 interviews, \\textbf{LadderTeam} achieved 99.1\\% chain convergence and an 81.0\\% ground-truth actionable response match (86.1\\% reluctant personality, 75.9\\% terse personality) with zero drift across all runs. All evaluation code, all transcripts, inputs, and a live demonstration platform will be open-sourced upon acceptance.","authors":["Manjushree Aithal","Alexander Kotz","James Mitchell"],"categories":["cs.SE","cs.HC"],"primary_category":"cs.SE","announce_type":"cross","date":"2026-08-19","first_seen":"2026-08-19","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.17029","pdf_url":"https://arxiv.org/pdf/2608.17029","source_feed":"cs.HC","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM访谈","需求启发","多智能体"],"reason":"用LLM替代人工访谈者，而非仿真人类被试，但涉及访谈方法自动化，边界相关。","model":"deepseek-v4-pro","scored_at":"2026-08-19T13:03:31","error":null,"has_summary":false,"summary":null},{"id":"2608.17970","version":1,"title":"Quo Vadis? Scientific Discovery in the Age of Artificial Intelligence","zh_title":"何去何从？人工智能时代的科学发现","abstract":"This paper examines the growing role of AI in scientific discovery. It first surveys the rapid rise of AI capabilities, especially in reasoning, abstraction, planning, and long-horizon task execution, before turning to scientometric evidence of AI's diffusion across the sciences. It then proposes a typology of AI systems used in research, ranging from specialized scientific AI through scientific AI assistants and agents to hybrid experimental systems that combine computation and physical experimentation. On this basis, it offers a selective overview of recent achievements in mathematics and computer science, physics, chemistry, the life sciences, and the behavioural and social sciences. It argues that, despite these advances, current systems remain constrained by important technical, epistemic, and institutional limitations, and that their growing use introduces both near-term and longer-term risks. The conclusion further suggests that the advancement of AI in science raises broader questions concerning the division of cognitive labour between human researchers and machines.","authors":["Petr O. Jedlicka"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2026-08-19","first_seen":"2026-08-19","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.17970","pdf_url":"https://arxiv.org/pdf/2608.17970","source_feed":"cs.CY","score":3,"bucket":"other","rubric_hits":["C1"],"tags":["AI for Science","科学发现","综述"],"reason":"综述AI在科学发现中的应用，不涉及LLM仿真人类被试或与人类数据对照","model":"deepseek-v4-pro","scored_at":"2026-08-19T13:03:37","error":null,"has_summary":false,"summary":null},{"id":"2607.27155","version":2,"title":"OmegaUse-OfficeVal: Benchmarking LLM Agents on Long-Horizon Office-Suite Tasks with Economic Grounding","zh_title":"OmegaUse-OfficeVal：基于经济基准的长周期办公套件任务LLM代理评测","abstract":"Large language model (LLM) agents are increasingly expected to assist users in completing tasks. However, existing benchmarks provide limited support for evaluating whether agents can carry out office-suite workflows at a reasonable cost. We introduce OmegaUse-OfficeVal, a benchmark for evaluating LLM agents on long-horizon office-suite tasks with task-level economic grounding. The benchmark comprises 100 tasks derived from office-suite requests proposed by practitioners and adapted through a privacy-preserving process. On average, these tasks require 2.32 hours of human labor to complete. An important feature of the benchmark is that each task is paired with two economic signals: human labor time and task price proxy. These signals enable direct comparisons between human costs and LLM inference costs, as well as value-weighted evaluation. To support stable evaluation, we develop code-based verifiers from fine-grained rubrics. We evaluate several frontier LLMs together with a human baseline. Although all evaluated LLMs are substantially cheaper and faster than human workers, they have not yet approached human-level deliverable quality. The code and dataset are fully open-sourced, and more information is available on our project website: https://omegause-officeval.github.io.","authors":["Jingbo Zhou","Yusai Zhao","Qi Bao","Jingjia Cao","Zhenghai Chen","Chang Gao","Kaiqi Guo","Muxin Guo","Mingxuan Li","Xinjiang Lu","Yanru Ma","Yixiong Xiao","Zenghui Zhang","Le Zhang","Hua Wu"],"categories":["cs.AI","cs.CL","cs.HC"],"primary_category":"cs.AI","announce_type":"replace-cross","date":"2026-08-19","first_seen":"2026-07-30","revised_at":"2026-08-19","abs_url":"https://arxiv.org/abs/2607.27155","pdf_url":"https://arxiv.org/pdf/2607.27155","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["LLM代理评测","办公自动化","经济成本"],"reason":"评估LLM代理执行办公任务，非仿真人类被试，无人类行为对照","model":"deepseek-v4-pro","scored_at":"2026-08-19T13:03:40","error":null,"has_summary":false,"summary":null},{"id":"2608.17827","version":1,"title":"From Global Benchmarks to Local Evaluations: Benchmarking LLMs for the German Public Sector","zh_title":"从全球基准到本地评估：为德国公共部门基准测试大语言模型","abstract":"Public institutions face a persistent challenge in selecting LLMs suited to their specific context. Existing benchmarks, however, are of limited use as they primarily reflect English-language and US-centric settings, and often only evaluate task performance. In this paper, we present first results of M\\\"OVE, a holistic evaluation framework for the German public sector, examining three rarely considered governance dimensions: energy consumption, provider transparency, and knowledge of German-party positions. Our results reveal significant trade-offs, with no single model excelling across all dimensions: estimated energy consumption varies more than 60-fold and is not explained by model size alone, information disclosure varies systematically across providers, and European models do not exhibit stronger knowledge of German party positions. Model selection for public institutions thus cannot rely on performance rankings alone. Instead, evaluations should also reflect the governance requirements of the deployment context.","authors":["Camilla Dalerci","Thilo Michael","Robin Schaefer","Daniel Weinland"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-19","first_seen":"2026-08-19","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.17827","pdf_url":"https://arxiv.org/pdf/2608.17827","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["LLM评测","公共部门","基准测试"],"reason":"评估LLM在德国公共部门的表现，属纯NLP能力评测，不以人类行为为参照系","model":"deepseek-v4-pro","scored_at":"2026-08-19T13:03:35","error":null,"has_summary":false,"summary":null},{"id":"2311.06273","version":1,"title":"Potential of ChatGPT in predicting stock market trends based on Twitter Sentiment Analysis","zh_title":"ChatGPT基于Twitter情感分析预测股市趋势的潜力","abstract":"The rise of ChatGPT has brought a notable shift to the AI sector, with its exceptional conversational skills and deep grasp of language. Recognizing its value across different areas, our study investigates ChatGPT's capacity to predict stock market movements using only social media tweets and sentiment analysis. We aim to see if ChatGPT can tap into the vast sentiment data on platforms like Twitter to offer insightful predictions about stock trends. We focus on determining if a tweet has a positive, negative, or neutral effect on two big tech giants Microsoft and Google's stock value. Our findings highlight a positive link between ChatGPT's evaluations and the following days stock results for both tech companies. This research enriches our view on ChatGPT's adaptability and emphasizes the growing importance of AI in shaping financial market forecasts.","authors":["Ummara Mumtaz","Summaya Mumtaz"],"categories":["q-fin.ST","cs.AI","cs.CL"],"primary_category":"q-fin.ST","announce_type":"cross","date":"2026-08-19","first_seen":"2026-08-19","revised_at":null,"abs_url":"https://arxiv.org/abs/2311.06273","pdf_url":"https://arxiv.org/pdf/2311.06273","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C2"],"tags":["金融预测","情感分析","ChatGPT"],"reason":"用ChatGPT预测股市，非仿真人类被试，无人类行为对照","model":"deepseek-v4-pro","scored_at":"2026-08-19T13:03:26","error":null,"has_summary":false,"summary":null},{"id":"2608.17987","version":1,"title":"Against Political Polarization: A Unified Framework for Tracing Evolving Political Ideologies on Social Media","zh_title":"反对政治极化：追踪社交媒体上演变政治意识形态的统一框架","abstract":"The rapid growth of social media has greatly influenced political discourse, highlighting the need to understand individual political ideologies and their temporal dynamics. This task faces challenges such as data scarcity, abundant non-political content, costly and bias-prone manual annotation, and difficulty in modeling future ideological inclinations. To address these issues, we propose TSN4PI, a unified framework for tracking the evolution of political ideologies on social media. It includes two core modules. The PIDN uses large language models with style transfer and unsupervised domain adaptation to enable robust ideology detection and filter irrelevant content from noisy, cross-domain data. The PIPN employs temporal graph neural networks to predict future ideological shifts, enabling comprehensive analysis of ideology presence, intensity, and evolution. We release two large-scale datasets for noncommercial research use to facilitate further work. Extensive case studies on multiple platforms (X and Truth Social) validate the effectiveness of TSN4PI and provide empirical insights into political polarization and the evolution of online ideologies. Our findings offer a nuanced perspective, advancing both methodological development and empirical understanding in this field.","authors":["Yijie Xu","Chao Wang","Hui Xiong"],"categories":["cs.SI","cs.AI","cs.CL","cs.LG"],"primary_category":"cs.SI","announce_type":"cross","date":"2026-08-19","first_seen":"2026-08-19","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.17987","pdf_url":"https://arxiv.org/pdf/2608.17987","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["政治意识形态检测","图神经网络","社交媒体分析"],"reason":"用LLM做政治立场检测与预测，属NLP能力评测，非仿真人类被试","model":"deepseek-v4-pro","scored_at":"2026-08-20T13:02:45","error":null,"has_summary":false,"summary":null},{"id":"2608.17270","version":1,"title":"Do LLMs Know a Good Hypothesis When They See One? Logit-Based Energy Scoring Outperforms Prompted LLM-as-Judge for Scientific Hypothesis Ranking","zh_title":"大语言模型能识别好假设吗？基于Logit的能量评分优于提示式LLM裁判用于科学假设排序","abstract":"Large language models (LLMs) are increasingly used for scientific hypothesis generation. However, evaluating generated hypotheses remains a challenge for trustworthy AI-enabled scientific workflows. Existing approaches often use LLMs as judges or rely on semantic similarity, which can favor familiar ideas over novel ones. We propose a logit-based energy scoring method that evaluates hypotheses using a language model's intrinsic confidence rather than comparative judgment. We benchmarked seven language models on 1,323 papers across 12 disciplines. Each paper was paired with its hypothesis and fifteen incorrect alternatives. Intrinsic scoring reached 33.0% Hit@1 pooled across both scorers, compared with 16.6% for prompted listwise ranking. The strongest configuration, a 1-billion-parameter model using logit-based energy scoring, reached 53.1%, though this was the maximum across 14 model-by-scorer combinations selected post hoc. Overall, intrinsic model confidence shows potential for scientific hypothesis evaluation. This study also motivates future research on confidence-based methods for trustworthy AI-enabled scientific discovery.","authors":["Swati Rajwal","Sanjay Das","Tirthankar Ghosal"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-19","first_seen":"2026-08-19","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.17270","pdf_url":"https://arxiv.org/pdf/2608.17270","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["科学假设评估","LLM评测","置信度评分"],"reason":"评估LLM对科学假设的排序能力，属NLP评测，不以人类行为为参照","model":"deepseek-v4-pro","scored_at":"2026-08-19T13:03:33","error":null,"has_summary":false,"summary":null},{"id":"2608.17624","version":1,"title":"Governing Delegation to Generative Artificial Intelligence: Human Direction, Work-Related Orientation, and Modes of Use","zh_title":"生成式人工智能的委托治理：人类指导、工作相关导向与使用模式","abstract":"Delegating cognitive operations to generative artificial intelligence redistributes execution and raises a governance problem: where human direction of the task remains. We distinguish two routes. Specified delegation places that direction before execution, through instructions, constraints, or criteria that delimit the task. Iterative coproduction places it during production, through interventions that correct or redirect provisional outputs. To examine both routes, we use aggregate monthly cells from the Anthropic Economic Index for April and May 2026. The AEI distinguishes two modes of use: 1P API, which corresponds to direct traffic through Anthropic's API, and Claude.ai, which combines activity from Chat and Cowork. On this basis, we test whether a stronger work-related orientation of human-AI interaction is associated with more specified delegation within each mode and whether the increase in the iterative profile is greater in Claude.ai than in 1P API. The main analysis uses level-0 O*NET tasks and estimates how both profiles change when an eligible record reallocates ten percentage points from personal use to work-related use. The iterative comparison is restricted to 1,411 node-month pairs observed and eligible in both modes. Specified delegation increases by 2.76 points in 1P API (95% CI: [2.30, 3.22]) and by 1.45 in Claude.ai (95% CI: [0.93, 1.97]). On the common support, iterative coproduction changes by-0.30 points in 1P API and by 0.15 in Claude.ai, yielding a between-mode difference of 0.45 points (95% CI: [0.15, 0.75]). These findings show that work-related orien tation is associated with stronger traces of prior human direction and that the observable iterative response varies across modes of use. The article shifts attention from how much the AI executes to when human direction leaves observable traces.","authors":["Jorge F\\'abrega"],"categories":["cs.CY","econ.GN","q-fin.EC"],"primary_category":"cs.CY","announce_type":"new","date":"2026-08-19","first_seen":"2026-08-19","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.17624","pdf_url":"https://arxiv.org/pdf/2608.17624","source_feed":"cs.CY","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["人机协作","AI治理","任务委托"],"reason":"研究人类如何指导AI完成任务，非用LLM仿真人类被试，无人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-08-19T13:03:35","error":null,"has_summary":false,"summary":null},{"id":"2607.26654","version":3,"title":"Constitutional Midtraining: Content Presence Drives Alignment Gains","zh_title":"宪法性中期训练：内容存在驱动对齐增益","abstract":"Post-training alignment is often shallow, eroding under fine-tuning. It remains untested as to whether constitutional midtraining interventions can produce durable alignment when cleanly isolated from post-training. We build a 394M-token constitutional corpus from Anthropic's Constitution and apply constitutional midtraining at 120B scale, where principled, values-based content is inserted into midtraining. A 2x2 design (curriculum ordering x deliberative reasoning) was used to produce four constitutionally midtrained conditions, plus a control, which were evaluated on self-generated and established benchmarks including alignment under pressure, value conflict resolution, blackmail, and emergent misalignment. All models were evaluated across three stages: post-midtraining, post-SFT, and post-benign fine-tuning. Constitutionally midtrained models outperformed the control on alignment generalization and durability, notably on blackmail: SFT instilled a blackmail propensity in all models, but constitutional midtraining blunted it, with the advantage surviving benign fine-tuning (-17.5pp). This durability did not extend to settings that required active resistance to in-context pressure or conflict, where the advantage attenuates after SFT. The presence of constitutional content at midtraining also mattered more than its structure, and constitutional midtraining incurred no capability cost, on average, at any stage (MMLU, ARC-Easy, piqa, GSM8K). A modest amount of constitutional content at midtraining could therefore yield broad, persistent alignment gains, offering a cheap, complementary addition to SFT-centered pipelines. Code, data, and models are available.","authors":["Desiree Cho","Cameron Tice","Bernie Hogan","Hunar Batra","Puria Radmard","Jun Zhao","Nigel Shadbolt"],"categories":["cs.CL","cs.AI","cs.CY","cs.LG"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-08-19","first_seen":"2026-07-30","revised_at":"2026-08-19","abs_url":"https://arxiv.org/abs/2607.26654","pdf_url":"https://arxiv.org/pdf/2607.26654","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["LLM对齐","模型训练","安全评估"],"reason":"研究LLM对齐训练，不涉及人类仿真或行为对照","model":"deepseek-v4-pro","scored_at":"2026-08-20T13:02:44","error":null,"has_summary":false,"summary":null},{"id":"2608.04772","version":2,"title":"Guideline-as-Oracle: Zero-Annotation Training of an Ophthalmic Telephone Triage Agent","zh_title":"指南即预言机：眼科电话分诊代理的零标注训练","abstract":"Scaling supervision for multi-turn medical agents is difficult because expert dialogue annotation is costly and clinical conversations are privacy-restricted. We introduce Guideline-as-Oracle (GAO), which compiles American Academy of Ophthalmology guidance into a 70-row operational rule table and uses it as the sole source of instance-level supervision for 3,000 training dialogues, reserving human labeling for evaluation. Because converting rules into dialogues is itself a design problem, we catalog eight construction strategies, including cited-row tier assignment, one-fact boundary pairs, metadata-only repair, and label repair, and characterize the evidential status of each: labeling mechanism, null, confounded, or evaluated only as a package. Fine-tuning a 9B backbone on this corpus yields GAO-Triage, improving agreement with a 201-case operational reference from 61.7% to 74.1% (exact McNemar p=0.0046) and emergent-case recall from 9.5% to 69.0%; the gains persist across a second seed and patient simulator. None of the seven general-purpose systems we test dominates GAO-Triage on both metrics, and GAO-Triage requires no frontier model at inference time. Permuting label-dialogue assignments collapses the model to a constant-routine predictor, indicating that the signal lies in guideline-derived assignment rather than dialogue surface form. Label repair coincides with the disappearance of a late-training safety degradation.","authors":["Chenyu Wang","Yi Liu","Baoqing Li","Min Tu","Diping Song"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-08-19","first_seen":"2026-08-06","revised_at":"2026-08-19","abs_url":"https://arxiv.org/abs/2608.04772","pdf_url":"https://arxiv.org/pdf/2608.04772","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["医疗对话代理","零标注训练","临床分诊"],"reason":"构建医疗分诊对话代理，属于角色扮演聊天机器人，无实验或测量目的，不涉及人类行为…","model":"deepseek-v4-pro","scored_at":"2026-08-19T13:03:40","error":null,"has_summary":false,"summary":null},{"id":"2608.16974","version":1,"title":"Position: Fairness Failure in Generative Models is an Evaluation Problem","zh_title":"立场：生成模型中的公平性失败是一个评估问题","abstract":"Despite groundbreaking advancements in generative models during the last decade, concerns about their lack of fairness, reinforcing societal inequalities and harming marginalized groups, remain under-addressed and difficult to act upon. This position paper argues that fairness failures in generative models, albeit driven by multiple factors, are ultimately stemming from an evaluation problem: fairness findings are rarely comparable across papers or actionable for deployment decisions. This paper diagnoses recurring empirical and conceptual failure modes in current practice and motivates a shift from ad-hoc bias checks to standardized, generative-specific evaluation. We propose Fairness Cards as a minimal reporting artifact that makes evaluation choices explicit (prompt families, counterfactual protocols, metrics, and refusal handling) enabling reproducibility, comparability, and accountability. We conclude with additional recommendations towards a paradigm shift in evaluation standards. Our project page can be found at https://mariiavladimirova.github.io/fairness-cards .","authors":["Mariia Vladimirova","Jean-Yves Franceschi","Thibaut Issenhuth"],"categories":["cs.LG","cs.AI"],"primary_category":"cs.LG","announce_type":"cross","date":"2026-08-19","first_seen":"2026-08-19","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.16974","pdf_url":"https://arxiv.org/pdf/2608.16974","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["公平性评估","生成模型","评估标准"],"reason":"论文讨论生成模型公平性评估，不涉及用LLM仿真人类被试或与人类数据对照，属于模…","model":"deepseek-v4-pro","scored_at":"2026-08-19T13:03:31","error":null,"has_summary":false,"summary":null},{"id":"2608.17144","version":1,"title":"Health Inquiry with AI: How Empathetic Expression and Conversational Contexts Shape Users' Communicative Acts","zh_title":"AI健康咨询：共情表达与对话情境如何塑造用户的沟通行为","abstract":"As online health information-seeking shifts to conversational AI, high-quality information retrieval increasingly relies on users' ``communicative acts''(proactively sharing and seeking information)---similar to how effective diagnosis and personalized guidance are elicited in patient-clinician communication. Drawing on health communication research, this study examines how a chatbot's modality of empathetic expression (Verbal, Visual, Multimodal) and the conversational context (General, Sensitive, Mental Health) influence these acts through a 2 x 2 x 3 within-subjects experiment (N = 48). The results revealed that while verbal and multimodal empathy significantly increased reply length, communicative acts were largely shaped by conversational context, with Sensitive context triggering more question-asking and Mental Health context leading to heightened concerns, assertive responses, and unprompted information disclosure. Combined with qualitative findings, we discuss design implications for building context-sensitive AI health inquiry systems that can encourage active user participation.","authors":["Xi Zheng","Xuyu Yang","Can Liu","Yuhan Luo"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-08-19","first_seen":"2026-08-19","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.17144","pdf_url":"https://arxiv.org/pdf/2608.17144","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["人机交互","健康咨询","共情表达"],"reason":"研究聊天机器人共情表达对用户沟通行为的影响，属于人机交互设计，非LLM仿真人类…","model":"deepseek-v4-pro","scored_at":"2026-08-19T13:03:31","error":null,"has_summary":false,"summary":null},{"id":"2608.17175","version":1,"title":"Balancing Safety and Autonomy: Accessibility-Oriented Interventions in Generative AI for Cognitive Impairment","zh_title":"平衡安全与自主：面向认知障碍的生成式AI无障碍干预","abstract":"Generative AI systems are increasingly used by older adults with cognitive impairment for everyday tasks such as information seeking, health management, and communication. While these systems provide flexible, language-based support, their open-ended outputs introduce risks of over-reliance, misinterpretation, and inappropriate decision-making. Prior work has focused on usability and adoption, with limited attention to how system design shapes users' participation in decision-making and the distribution of agency in care contexts. We present a qualitative study of 45 individuals with cognitive impairment and their caregivers. We identify five accessibility-oriented mechanisms: AI Capability Constraint, Human Oversight Embedding, Cognitive Engagement Maintenance, Human-AI Relationship Regulation, and Risk Transparency and Control, through which systems structure interaction. These mechanisms both support and constrain users by redistributing decision-making across users and caregivers. We show that their effects vary by impairment level: while protective mechanisms support users with severe impairment, they can restrict autonomy for those with mild impairment. As impairment progresses, tensions become less visible as user participation diminishes. Our findings highlight the need for dynamic designs that balance safety and autonomy in AI-supported care.","authors":["Yibo Meng","Jingruo Chen","Lyumanshan Ye","Bingyi Liu","Zhicong Lu"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-08-19","first_seen":"2026-08-19","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.17175","pdf_url":"https://arxiv.org/pdf/2608.17175","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["生成式AI","无障碍设计","人机交互"],"reason":"研究生成式AI在认知障碍护理中的设计，不涉及用LLM仿真人类被试或与人类数据对…","model":"deepseek-v4-pro","scored_at":"2026-08-19T13:03:33","error":null,"has_summary":false,"summary":null},{"id":"2608.14606","version":1,"title":"Plausible but Not Valid: A Psychometric Audit of LLMs as Synthetic Survey Respondents","zh_title":"看似合理但无效：对LLM作为合成调查受访者的心理测量审计","abstract":"Large language models (LLMs) are increasingly used as synthetic survey respondents, but existing evaluations ask whether answers look plausible at the individual level. We argue the right question is psychometric: do LLMs preserve the joint distribution, latent structure, reliability, mediation pathways, and demographic effects of real human survey data? We introduce a Lithuanian organisational-psychology dataset (n=263 employees; Dunham Attitudes Toward Change, UWES-17, Koopmans IWPQ; 68 items, 12 subscales) and condition a 37-model lineup spanning OpenAI, Anthropic, Google, and twelve open-weight families on real respondent profiles under a five-level persona-disclosure ladder, presentation and reasoning-effort ablations, counterfactual demographic swaps (gender, role, education), a cross-language check, and a verbatim-recall memorization probe. The resulting Psychometric Similarity Score (PSS) is anchored against five non-LLM statistical baselines and a held-out human-vs-human ceiling, with respondent-bootstrap confidence intervals and an item-permutation null for Tucker's phi. LLMs reproduce the qualitative direction of human psychometric relationships, but a Gaussian-copula baseline beats every LLM on the sample-driven PSS components; the LLM \"crowd\" is more similar to itself (mean inter-LLM PSS 0.73) than to humans; and memorization does not drive the leaderboard (recall-PSS rank correlation 0.00). Counterfactual swaps reveal education-driven effects (mean |d|=0.56) that dwarf gender (0.12) and role (0.18); Tucker's phi on UWES falls inside the permutation null for 8 of 37 models. Downstream, every LLM shows a strong acquiescence shift (+0.84 SD), synthetic-trained regressors lose predictive validity on held-out humans (mean R^2 -0.18 vs 0.28), and models fabricate indirect effects on 3 of 10 placebo mediation paths. LLM samples are not a drop-in replacement for human survey data.","authors":["Mantas Lukauskas","Viktorija \\v{S}arkauskait\\.e"],"categories":["cs.CY","cs.AI","cs.CL","stat.AP"],"primary_category":"cs.CY","announce_type":"cross","date":"2026-08-18","first_seen":"2026-08-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.14606","pdf_url":"https://arxiv.org/pdf/2608.14606","source_feed":"cs.AI","score":10,"bucket":"selected","rubric_hits":["A1","A2","A3","A5","B1","B2","B3","B4"],"tags":["LLM仿真","心理测量效度","调查数据"],"reason":"直接评估LLM作为调查受访者的心理测量效度，并与真实人类数据对照，批判性指出失…","model":"deepseek-v4-pro","scored_at":"2026-08-18T13:04:16","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-18","rank":2,"question":"LLM 作为合成调查受访者时，是否能在心理测量学层面（联合分布、潜结构、信度、中介路径、人口学效应）复现真实人类调查数据的特征？","design":"使用 37 个 LLM（涵盖 OpenAI、Anthropic、Google 及 12 个开源家族）基于真实受访者档案生成对 68 个题项（3 个量表）的回答，通过五级人格披露阶梯、呈现方式与推理努力消融、反事实人口学变换（性别、角色、教育）、跨语言检查和逐字回忆探针等处理，测量心理测量相似性得分（PSS）及其各维度。","baseline":"立陶宛组织心理学数据集（n=263 名员工；Dunham 变革态度量表、UWES-17、Koopmans IWPQ；68 题，12 个分量表），以及五个非 LLM 统计基线和留出的人类对比上限。","findings":"LLM 能复现人类心理测量关系的定性方向，但高斯 copula 基线在样本驱动的 PSS 分量上击败所有 LLM；LLM 群体内部相似度（平均 PSS 0.73）高于与人类的相似度，且逐字回忆探针表明记忆并非排行榜驱动因素（秩相关 0.00）。","reliability":"论文指出 LLM 样本不能直接替代人类调查数据：合成样本存在默认偏差（+0.84 SD），在留出人类数据上预测效度丧失（平均 R² -0.18 vs 0.28），并在 10 条安慰剂中介路径中捏造了 3 条显著间接效应；此外，教育驱动的反事实效应（平均 |d|=0.56）远大于性别（0.12）和角色（0.18），且 UWES 的 Tucker's phi 在 37 个模型中有 8 个落入置换零分布内。","relevance":"该研究直接评估 LLM 作为人类被试替代品的心理测量效度，并与真实人类数据严格对照，批判性地揭示了仿真在联合分布、预测效度和中介推断上的失效条件，对关注 LLM 仿真可靠性与偏差的研究者极具参考价值。","inspiration":"借鉴其多维度心理测量审计框架（联合分布、潜结构、信度、中介路径、人口学效应）和反事实人口学变换设计，系统评估合成样本的效度｜可迁移到经济金融中的调查实验，如消费者信心、通胀预期、风险偏好或政策支持度等场景，检验 LLM 能否复现真实人群的分布与结构｜以真实家庭金融调查（如美国 SCF 或中国 CHFS）为基准，用 LLM 基于受访者人口学特征生成对风险态度、时间偏好等量表的回答，施加收入或教育水平的反事实变换，比较 LLM 样本与人类样本在联合分布、因子结构和中介效应上的差异。"}},{"id":"2608.15871","version":1,"title":"Large Language Models as Implicit Sociological Models: Reconstructing Voting Behaviour from Sociodemographic Profiles","zh_title":"大语言模型作为隐式社会学模型：从社会人口特征重建投票行为","abstract":"Large language models (LLMs) trained on large-scale internet corpora encode extensive statistical regularities about social identities, attitudes, and political behaviour. This paper introduces and evaluates a methodological framework that leverages these latent representations to reconstruct aggregate voting behaviour from individual-level sociodemographic profiles. We operationalize LLMs as implicit sociological models by conditioning them on demographic descriptions, eliciting probabilistic turnout and party preferences, and aggregating individual outputs via a soft voting procedure. Using the 2021 Czech parliamentary election as a validation case, we demonstrate that contemporary LLMs reproduce official election outcomes with low mean absolute error, recover known political bloc structures, and align with independently established sociodemographic gradients. The contribution of this work is methodological rather than predictive: we show how LLMs can be systematically interrogated as compressed representations of social reality, offering a novel exploratory instrument for computational social science while clearly delineating its epistemic and ethical limits.","authors":["Roman Neruda","Martin Bako\\v{s}","Josef \\v{S}lerka","V\\'it Tu\\v{c}ek","Petra Vidnerov\\'a","Gabriela Kadlecov\\'a"],"categories":["cs.CY","cs.CL","cs.LG"],"primary_category":"cs.CY","announce_type":"new","date":"2026-08-18","first_seen":"2026-08-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.15871","pdf_url":"https://arxiv.org/pdf/2608.15871","source_feed":"cs.CY","score":10,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM仿真","投票行为","计算社会科学"],"reason":"用LLM从人口特征重建投票行为，并与真实选举结果对照，直接仿真人类决策。","model":"deepseek-v4-pro","scored_at":"2026-08-18T13:04:22","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-19","rank":1,"question":"大语言模型能否仅从个体社会人口学特征重建总体投票行为，并在捷克2021年议会选举这一非英语多党制情境下得到验证？","design":"将LLM作为隐式社会学模型，以10个社会人口学变量（含主观生活水平和政治兴趣）为条件，对代表性调查中的个体生成概率性投票选择（投票率和政党偏好），通过软投票聚合得到模拟选举结果。","baseline":"官方选举结果（捷克统计局）和调查中的自报投票选择（声称投票）。","findings":"当代LLM能以较低平均绝对误差重现官方选举结果，恢复已知政治阵营结构，并与独立确立的社会人口学梯度一致。该贡献是方法论的，而非预测性的。","reliability":"论文承认个体层面预测噪声大且存在系统性偏差，但通过软投票聚合可部分抵消；捷克语为中等资源语言，模型训练数据以英语为主，可能影响表现；未明确讨论其他失效条件。","relevance":"该研究直接命中你的核心关注：用LLM仿真人类决策并与真实选举数据对照，且提供了多层级验证（总体份额、协方差结构、与民调机构及CHES专家编码比较），值得精读原文以了解其方法细节和局限。","inspiration":"借鉴其软投票聚合和三角验证设计，将个体噪声转化为总体稳健估计，并同时对照官方数据和自报数据以识别偏差。｜可迁移到政策公告的预期形成研究，例如模拟不同人口群体对财政或货币政策变化的反应。｜以代表性家庭调查中的个体为被试，用LLM基于人口特征生成对政策变化的预期（如通胀预期），处理为不同政策情景描述，结果变量为预期值或不确定性，对照真实调查中的预期数据和后续实际经济行为。"}},{"id":"2608.14630","version":1,"title":"Characterizing Rhetorical Misalignment in Decision-Making with Language Models","zh_title":"表征语言模型决策中的修辞错位","abstract":"Human decision-making is often shaped by a range of well-documented cognitive biases. As large language models (LLMs) become increasingly integrated into high-stakes human-AI decision-making, it is important to understand whether their outputs can amplify potential biases, how this influences human decisions, and crucially, whether it can lead to harmful consequences. In this work, we develop a decision-theoretic framework to study rhetorical misalignment, a failure mode where an LLM uses rhetorically inappropriate forms of presentation for a given decision context, thereby inducing suboptimal human decisions. We empirically investigate this phenomenon through a human-subject experiment in realistic clinical decision-making using a dataset curated from the United States Medical Licensing Examination. By measuring how LLM-generated information affects decisions, we observe that LLMs induce an average 2.81% rate of harmful decision flips across different models, where clinician participants change from a correct to an incorrect answer. Rationales reported by participants provide evidence that these revisions are closely related to the language used by LLMs that may induce different types of cognitive biases, including anchoring, authority bias, and loss aversion. To enable scalable evaluation, we instantiate our theoretical framework using decision-makers simulated by LLMs to computationally measure rhetorical misalignment. Our findings reveal a safety concern previously unrecognized in high-stakes domains: a model can be factually aligned yet still induce harm through its rhetorical presentation.","authors":["Zirui Cheng","Joey Chan","Simo Du","Chenhao Tan","Yue Guo","Hao Peng"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"cross","date":"2026-08-18","first_seen":"2026-08-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.14630","pdf_url":"https://arxiv.org/pdf/2608.14630","source_feed":"cs.AI","score":8,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["LLM仿真","人类决策","认知偏差"],"reason":"用LLM模拟人类决策者评估修辞偏差，并与真实人类实验对照，涉及临床决策场景，批…","model":"deepseek-v4-pro","scored_at":"2026-08-18T13:04:16","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-18","rank":7,"question":"LLM 的修辞呈现方式是否会在事实信息一致的情况下诱导人类做出次优决策，即产生“修辞错位”现象？","design":"论文采用人类受试者实验与LLM模拟决策者相结合的方法。人类实验中，临床医生参与者基于USMLE题目作答，部分题目提供LLM生成的分析（固定信息或自然生成），测量决策变化率及有害翻转率；模拟实验中，用LLM分别扮演理性决策者和行为决策者，比较两者在相同信息不同措辞下的决策差异。","baseline":"人类受试者实验数据：临床医生在没有LLM辅助时的原始答案作为对照，测量LLM辅助后的决策变化。","findings":"人类实验中，LLM辅助导致平均27.58%的决策变化，其中2.81%为有害翻转（从正确变为错误）；参与者报告显示这些变化与锚定、权威偏误、损失厌恶等认知偏误相关。模拟实验表明，即使信息相同，仅语言措辞差异也能导致理性与行为决策者之间的分歧，且自然生成设置下分歧更大。","reliability":"论文承认其人类实验仅限于USMLE临床决策场景，未测量下游后果（如患者结局、经济成本），且有害翻转率绝对值较小；模拟实验的LLM决策者可能无法完全代表人类认知偏误。","relevance":"该研究直接使用LLM模拟人类决策者来测量修辞错位，并与真实人类实验对照，属于用LLM进行人类仿真实验的典型工作，且涉及高 stakes 临床决策场景，对关注仿真可靠性与偏差的研究者具有重要参考价值。","inspiration":"借鉴其“固定信息、变化措辞”的处理设计，可分离信息内容与修辞框架的效应，并利用LLM模拟理性与行为决策者进行大规模测量｜可迁移到经济金融中的政策公告解读、投资建议呈现、信贷合同条款表述等场景，研究措辞对个体决策的影响｜设计实验：以真实投资者或消费者为被试，呈现同一金融产品的两种措辞（如“年化收益5%” vs “亏损概率2%”），测量选择差异，并用历史交易数据或调查数据作为真实行为基准，同时用LLM模拟投资者进行平行实验以验证仿真效度。"}},{"id":"2510.10813","version":2,"title":"The Fragility of Strategic Thinking in Large Language Models","zh_title":"大语言模型中战略思维的脆弱性","abstract":"Large Language Models (LLMs) are increasingly applied to domains that require reasoning about other agents' behavior, such as negotiation, policy design, and market simulation. However, can we trust LLMs to think strategically in complex situations? Existing research mostly evaluates LLMs' adherence to equilibrium play or their exhibited depth of reasoning, leaving open whether they display strategic thinking meant as the ability to form coherent conjectures about other agents, to evaluate possible actions conditional on those conjectures, and to best respond to them. We develop a framework to identify this ability by disentangling belief formation, evaluation, and choice in static complete-information games across a series of non-cooperative environments. By jointly analyzing models' revealed choices and reasoning traces, and introducing a new context-free game to rule out imitation from memorization, we show that strategic thinking in current frontier LLMs is real but fragile: models execute best responses to exogenous conjectures and form opponent-contingent conjectures when left unconstrained. Yet under increasing complexity explicit recursion gives way to model-specific logic shifts and heuristic rules of choice, both within and outside equilibrium reasoning. Further, these heuristics do not map directly onto the systematic biases typically observed in human strategic behavior. These findings, already emerging in noiseless settings, warrant caution in the application of LLMs as strategic agents in complex environments.","authors":["Enric Junque de Fortuny","Veronica Roberta Cappelli"],"categories":["cs.AI","cs.GT"],"primary_category":"cs.AI","announce_type":"replace","date":"2026-08-18","first_seen":"2025-10-12","revised_at":"2026-08-18","abs_url":"https://arxiv.org/abs/2510.10813","pdf_url":"https://arxiv.org/pdf/2510.10813","source_feed":"cs.AI","score":7,"bucket":"pending","rubric_hits":["A1","B2","B4"],"tags":["LLM战略推理","博弈论","行为对照"],"reason":"评估LLM在博弈中的战略思维，与人类行为对照，但非直接仿真人类被试，结论可迁移。","model":"deepseek-v4-pro","scored_at":"2026-08-18T13:04:57","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-18","rank":8,"question":"大语言模型在静态完全信息博弈中是否具备真正的战略思维能力，即能否形成关于对手行为的连贯信念、基于信念评估行动并做出最优反应？","design":"该研究并非用LLM仿真人类被试，而是直接测试多个前沿LLM在静态完全信息博弈中的战略思维。研究者设计了一系列非合作博弈环境，通过分解信念形成、评估和选择三个环节，结合模型的选择和推理痕迹，并引入无上下文博弈以排除记忆模仿。","baseline":"无对照","findings":"前沿LLM的战略思维真实但脆弱：它们能对外生信念做出最优反应，并在无约束时形成依赖对手的信念；但在复杂度增加时，显式递归让位于模型特定的逻辑转换和启发式选择规则，且这些启发式与人类系统性偏差不直接对应。","reliability":"论文指出，在无噪声环境中已出现脆弱性，因此对LLM在复杂环境中作为战略代理的应用需谨慎；启发式规则与人类偏差不匹配，可能限制其在人类行为仿真中的有效性。","relevance":"该研究评估LLM在博弈中的战略思维，虽非直接仿真人类被试，但其结论对使用LLM进行人类决策仿真具有警示意义，值得阅读原文以了解LLM在战略推理中的局限。","inspiration":"借鉴其分解信念、评估和选择的框架，以及通过推理痕迹和上下文无关博弈排除记忆混淆的方法，可用于严格测试LLM的决策过程。｜可迁移到经济金融中的策略互动场景，如拍卖竞价、谈判或市场进入博弈，检验LLM是否形成理性预期。｜设计一个古诺竞争实验，让LLM扮演企业，处理为不同信息结构（如对手成本已知或未知），结果变量为产量选择，并与人类实验数据（如Huck等2000年的古诺实验）对照，分析LLM的信念形成和最优反应。"}},{"id":"2608.00410","version":2,"title":"Where did the ambiguity go? Examining how multimodal models interpret polysemous words","zh_title":"歧义去哪了？考察多模态模型如何解释多义词","abstract":"Human language is highly polysemous. Many common words (e.g., \"bank\" or \"palm\") carry several distinct meanings that shape what humans communicate and imagine. Large language models (LLMs) have been shown to understand this multiplicity of meaning, but much less is known about how polysemy surfaces in other modalities such as images. We study this across 17 text-to-image and 15 text-generation models by giving each a polysemous word with no context to fix its meaning and measuring which senses are produced over many samples. We find a clear multimodal gap, where within every model family, generated images settle on far fewer senses than generated sentences (normalized entropy 0.10 vs. 0.25), and both are far less varied than what people imagine for the same words (normalized entropy 0.47). However, when we instead ask a model to list how often it would generate outputs corresponding to each possible meaning of a word, it predicts distributions that are more diverse than the actual space of outputs. These results reveal a multimodal gap in how foundation models express meaning, and how their understanding may not transfer faithfully nor equally across modalities.","authors":["Jasin Cekinmez","Addison J. Wu","Raja Marjieh","Thomas L. Griffiths"],"categories":["cs.AI","cs.CL","cs.CV"],"primary_category":"cs.AI","announce_type":"replace","date":"2026-08-18","first_seen":"2026-08-04","revised_at":"2026-08-18","abs_url":"https://arxiv.org/abs/2608.00410","pdf_url":"https://arxiv.org/pdf/2608.00410","source_feed":"cs.AI","score":7,"bucket":"pending","rubric_hits":["A1","B1"],"tags":["多模态模型","语义歧义","人类对照"],"reason":"用LLM生成文本和图像，与人类想象对照，评估多模态意义表达差异，属于仿真人类认…","model":"deepseek-v4-pro","scored_at":"2026-08-18T13:04:58","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-18","rank":9,"question":"多模态基础模型在无上下文的多义词上，其图像生成与文本生成所表达的语义分布有何差异，并与人类想象分布相比如何？","design":"本研究并非以LLM模拟人类被试，而是直接考察模型行为：以100个多义词为刺激，在无上下文条件下分别输入17个文本到图像模型和15个文本生成模型，每个词采样30次，用GPT 5.4作为裁判将输出分类到预定义的词义，得到每个模型-词对的词义分布。","baseline":"通过Prolific收集人类被试在相同100个词上的反应，采用与模型条件匹配的两种框架（“用这个词造句”对应文本模型，“这个词让你想到什么图像”对应图像模型），同样由GPT 5.4裁判分类。","findings":"图像模型生成的词义分布比文本模型更集中（归一化熵0.10 vs 0.25），且两者都远低于人类想象的多样性（归一化熵0.47）。当要求模型预测各词义的生成频率时，其预测分布比实际输出更多样，表明模型知道歧义但生成时未充分表达。","reliability":"论文未明确讨论仿真失效条件，但指出偏好调优会进一步收窄生成分布，且对齐不能完全解释多模态差距；裁判可靠性通过作者盲标60张图像与裁判完全一致来验证。","relevance":"该研究用真实人类数据作为基准，比较LLM多模态输出与人类想象分布，揭示模型在表达语义多样性上的系统性偏差，对关注LLM仿真人类认知与行为可靠性的研究者有直接参考价值。","inspiration":"借鉴其无上下文刺激与分布比较方法，可设计实验考察LLM在经济概念上的先验分布与人类直觉的差异｜可迁移到政策公告的预期形成研究，如“通货膨胀”一词的多重解读如何影响模型生成的预期分布｜以LLM为被试，呈现无上下文的“利率”等经济术语，要求生成解释或图像，测量其词义分布，并与调查数据中公众对同一术语的理解分布进行对照。"}},{"id":"2608.12253","version":2,"title":"One Frozen Simulator Is Not Enough: Simulator Collapse in Multi-Agent RL","zh_title":"单一冻结模拟器不够：多智能体强化学习中的模拟器坍缩","abstract":"Multi-agent reinforcement learning for human-AI interaction typically relies on a single large language model to simulate user behavior. We show that this approach systematically fails to generalize, and trace the failure to simulator collapse: because the simulator LLM is mode-collapsed, an LLM policy trained against it overfits to narrow strategies that exploit the simulator's dominant mode, and such a policy transfers poorly to unseen simulators and real users. We formalize this collapse theoretically and propose two complementary solutions, one at inference time and one at training time. The inference-time solution, Verbalized Sampling, broadens the simulator's behavior by sampling from a verbalized response distribution, reducing mode collapse. The training-time solution, Co-Training, jointly optimizes the policy against a population of trainable simulators, preventing it from overfitting to any single simulator's mode. We validate both solutions on three multi-turn benchmarks: Persuasion for Good, $\\tau^2$-bench, and CooperBench. Verbalized Sampling improves held-out success by up to 9% over single-simulator RL, and Co-Training pushes gains further to 14%; the human study shows similar gain on real users. Both solutions preserve the policy diversity that collapses under single-simulator RL. To support further work in this direction, we release SCOPE, an open-source framework for Population Co-Training multi-agent RL. More broadly, our results suggest that the diversity of the training environment, not only the policy, is critical to the generalization of multi-turn RL to real-world deployment.","authors":["Simon Yu","Nicholas Tomlin","Marwa Abdulhai","Ximing Lu","Derek Chong","Abe Hou","Dilara Soylu","Sergey Levine","Christopher D. Manning","Weiyan Shi"],"categories":["cs.CL","cs.AI","cs.LG"],"primary_category":"cs.CL","announce_type":"replace-cross","date":"2026-08-18","first_seen":"2026-08-13","revised_at":"2026-08-18","abs_url":"https://arxiv.org/abs/2608.12253","pdf_url":"https://arxiv.org/pdf/2608.12253","source_feed":"cs.AI","score":7,"bucket":"pending","rubric_hits":["A1","B1","B4"],"tags":["LLM仿真","多智能体强化学习","人类行为模拟"],"reason":"用LLM模拟用户行为训练策略，并与真实用户对照，指出单一模拟器失效问题，方法可…","model":"deepseek-v4-pro","scored_at":"2026-08-18T13:04:57","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-19","rank":10,"question":"在人类-AI多轮交互的强化学习训练中，使用单一冻结的LLM模拟用户行为是否会导致策略泛化失败，以及如何解决该问题？","design":"该研究使用LLM模拟用户行为，在三个多轮交互基准（Persuasion for Good、τ²-bench、CooperBench）上训练强化学习策略。处理包括：基线（单一冻结LLM模拟器）、推理时方案（Verbalized Sampling，从模拟器输出的言语化分布中采样）、训练时方案（Co-Training，策略与可训练模拟器联合优化）以及两者结合（Population Co-Training）。结果变量为在留出模拟器和真实用户上的任务成功率及策略熵。","baseline":"在τ²-bench和Persuasion for Good上进行了预先注册的人类研究，以真实用户的任务结果和对话自然度作为对照。","findings":"单一模拟器强化学习导致策略过拟合模拟器的模式，在留出模拟器和真实用户上表现下降，出现“模拟器崩溃”。Verbalized Sampling和Co-Training均能缓解该问题，其中Population Co-Training在留出任务成功率上提升最高（达14%），并在人类研究中表现出类似增益。","reliability":"论文承认模拟器崩溃源于LLM的模式坍缩，且策略可能利用模拟器的特定模式；解决方案的有效性依赖于模拟器多样性的恢复程度。但未系统讨论在更复杂或不同领域用户行为模拟中的失效条件。","relevance":"该研究直接针对LLM作为人类被试替代品的可靠性问题，揭示了单一模拟器在强化学习训练中的系统性偏差，并提供了与真实用户对照的实证证据，对关注仿真失效条件的研究者具有重要参考价值。","inspiration":"借鉴其通过多样化模拟器（如采样或联合训练）来防止策略过拟合的方法，可用于经济实验中处理被试异质性。｜可迁移到政策评估中的多轮交互场景，如消费者与AI客服的谈判或信贷审批对话。｜设计一个实验：用多个LLM模拟不同类型的消费者（如风险偏好不同），训练一个AI谈判策略，处理为使用单一模拟器 vs. 模拟器群体，结果变量为谈判达成率和消费者满意度，并与真实人类被试的行为数据对照。"}},{"id":"2608.15634","version":1,"title":"Argumentation for Common Ground: Finding Zones of Possible Agreement between Individuals in Conflict","zh_title":"共同基础的论证：寻找冲突个体间的可能协议区","abstract":"How can common ground between societies in conflict be identified when citizens' acceptability of peace agreements is shaped by contested narratives? Such acceptability is mediated not only by the clauses that agreements include or exclude, but crucially by citizens' subjective reasoning concerning agreements' clauses. In this paper, we leverage computational argumentation to introduce a novel approach to identifying mutually acceptable agreements among individuals in conflict, i.e. a Zone of Possible Agreement (ZOPA). First, we introduce a quantitative bipolar argumentation framework tailored to represent each side's reasoning about peace agreements. We then show how merging these frameworks can enable negotiators to identify peace agreements that are mutually acceptable. To evaluate our approach under conditions of real-world relevance, we focus on the Palestinian-Israeli conflict, where long-standing policy, practitioner and public interest underscores the demand for methods capable of analysing polarised public reasoning. We show how our framework identifies a ZOPA through theoretical analysis and preliminary experiments using survey data from both existing work and retrieved by a large language model. The results illustrate how argumentation can empower negotiators and conflict-resolution teams in mapping feasible ZOPAs grounded in citizens' reasoning.","authors":["Elisa Cavatorta","Antonio Rago"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-18","first_seen":"2026-08-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.15634","pdf_url":"https://arxiv.org/pdf/2608.15634","source_feed":"cs.AI","score":7,"bucket":"pending","rubric_hits":["A3","B1","B2"],"tags":["LLM仿真","冲突解决","人类数据对照"],"reason":"用LLM检索调查数据模拟冲突双方推理，并与真实调查数据对照，属于社会政治过程仿…","model":"deepseek-v4-pro","scored_at":"2026-08-18T13:04:43","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-18","rank":17,"question":"如何利用计算论证方法识别冲突社会中个体之间可共同接受的和平协议区域（ZOPA）？","design":"该研究提出一种定量双极论证框架（QBAF）来表示个体对和平协议条款的推理，通过合并对立双方的QBAF来识别共同可接受的协议区域。实验部分使用现有调查数据和由大语言模型检索的调查数据，对巴以冲突场景进行初步验证。","baseline":"使用来自现有研究（Golan-Nadir等）的全国代表性调查数据，以及由大语言模型检索的报告数据作为对照。","findings":"理论分析表明，满足平衡性、单调性和对偶性的渐进语义能产生直观的协议排序；初步实验显示该方法与现有调查数据具有合理相关性，表明其适用于现实部署。","reliability":"论文承认当前实验仅为初步验证，未来需要从实际调查受访者中获取推理数据并进行大规模合并；同时指出基础分数获取需要更严谨的协议，可能借鉴行为经济学原理。","relevance":"该研究利用LLM检索调查数据来模拟冲突双方的推理，并与真实调查数据对照，属于社会政治过程仿真，对关注LLM仿真可靠性和偏差的研究者有一定参考价值，但方法核心是计算论证而非LLM仿真，相关度中等。","inspiration":"借鉴其将个体主观推理形式化并合并以识别共识区域的方法，可用于经济政策偏好聚合｜可迁移到政策公告的预期形成或消费者对金融产品的态度分析｜设计：以LLM生成个体对某项经济政策（如税收改革）的论证框架，处理为不同政策条款组合，结果变量为个体接受度，与真实调查数据对照验证。"}},{"id":"2608.14681","version":1,"title":"Automatic or Controlled? Repetition Priming Reveals Divergent Processing in Base LLMs, Instruct LLMs, and Humans","zh_title":"自动还是受控？重复启动揭示基础LLM、指令LLM与人类的分歧加工","abstract":"Words recur constantly in natural language use, yet it remains unclear whether language models reactivate prior representations or re-evaluate repeated words afresh, and whether post-training changes this default behavior. We apply repetition priming (Shiffrin and Schneider, 1977) to 15 models across five model families (1.5B-14B parameters) in two tasks, semantic categorization and cloze completion, with matched human experiments using identical stimuli. We find that base models exhibit automatic processing: they show immediate facilitation that remains stable across lags, partially survives context removal, and correlates with attention to prior occurrences. Instruct models exhibit controlled processing: their facilitation decays with lag, collapses without expected context, and reverses to interference at larger scales. Within the Qwen 2.5 family, this dissociation increases monotonically with model scale, suggesting that post-training progressively alters repetition processing. Humans show a hybrid profile, with lag-sensitive facilitation resembling instruct models but without interference, suggesting that neither model type fully captures human cognition. Our findings reveal a qualitative shift in how language models process repeated information after post-training and provide mechanistic evidence for the divergence between model behaviors.","authors":["Jinglei Ren","Yuyue Wang"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"cross","date":"2026-08-18","first_seen":"2026-08-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.14681","pdf_url":"https://arxiv.org/pdf/2608.14681","source_feed":"cs.AI","score":7,"bucket":"pending","rubric_hits":["A1","B1","B4"],"tags":["LLM人类仿真","认知实验对照","模型偏差"],"reason":"用LLM复现人类重复启动效应，并与人类数据对照，揭示模型与人类差异，可迁移到仿…","model":"deepseek-v4-pro","scored_at":"2026-08-18T13:04:18","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-18","rank":14,"question":"语言模型在重复遇到同一词时，是自动激活先前表征还是重新评估，以及指令微调是否改变这一默认处理方式？","design":"用15个模型（5个家族，1.5B–14B参数）在语义分类和完形填空两个任务中，通过操纵重复词之间的间隔（lag）和上下文，测量重复启动效应（log概率边际的变化），并与人类被试在相同刺激和设计下的表现对比。","baseline":"匹配的人类实验，使用相同刺激和设计，测量人类被试的重复启动效应。","findings":"基础模型表现出自动加工：启动效应立即出现且不随间隔衰减，部分在移除上下文后仍存在，并与对先前出现的注意力相关；指令模型表现出控制加工：启动效应随间隔衰减，在移除预期上下文后消失，且在较大规模时反转为干扰。人类表现出混合特征：对间隔敏感（类似指令模型）但始终为促进效应（类似基础模型），从不出现干扰。","reliability":"论文未讨论","relevance":"该研究用LLM复现人类重复启动效应，并与人类数据对照，揭示了模型与人类在重复信息处理上的差异，对评估LLM作为人类被试替代品的可靠性有直接参考价值，值得阅读原文。","inspiration":"借鉴其通过操纵重复间隔和上下文来分离自动与控制加工的实验设计，以及用注意力分析和消融实验提供机制证据的做法。｜可迁移到经济金融中的信息重复场景，如重复广告对消费者决策的影响、重复政策信息对预期形成的作用。｜以LLM模拟消费者，呈现重复的产品信息（如广告语），操纵重复次数和间隔，测量购买意愿或态度变化，并与真实消费者实验数据对照，检验LLM是否复现重复效应。"}},{"id":"2608.15630","version":1,"title":"Do Assessment Instruments Measure the Same Thing for Humans and LLMs? A Latent Structure Analysis","zh_title":"评估工具对人类和LLM测量的是同一构念吗？一项潜在结构分析","abstract":"The rapid development and growing deployment of large language models (LLMs) have made it increasingly important to understand their capabilities. A common approach is to evaluate LLMs using assessment instruments originally designed to measure skills and competencies in humans, such as standardized exams, and to use performance on these instruments as evidence for generalizable claims about LLMs' underlying abilities on the same skills the assessments are intended to measure in humans. However, from a validity perspective, such inferences require that the relationship between observed performance and underlying constructs established for humans also holds for LLMs. In particular, a necessary condition for transferring score interpretations is similarity in the latent structure of responses to the assessment. In this study, we examine whether this condition holds in two educational contexts: high-school chemistry and a quantitative reasoning section of a university entrance exam. Using a case study design, we compare human response data with responses generated by six multimodal LLMs. Our analytical approach combines exploratory factor analysis, factor congruence, and resampling to assess latent structure similarity across human learners and LLMs. Across both instruments, we find systematic differences between human and LLM factor structures, showing evidence that the analyzed assessments may not measure the same constructs for humans and LLMs. These findings call into question the validity of evaluation practices that use educational assessments to make claims about AI capabilities.","authors":["Alona Strugatski","Licol Zeinfeld","Giora Alexandron"],"categories":["cs.HC","cs.AI","cs.CL"],"primary_category":"cs.HC","announce_type":"cross","date":"2026-08-18","first_seen":"2026-08-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.15630","pdf_url":"https://arxiv.org/pdf/2608.15630","source_feed":"cs.AI","score":7,"bucket":"pending","rubric_hits":["A2","B1","B4"],"tags":["LLM评估效度","潜在结构分析","人类对照"],"reason":"评估LLM在人类测评工具上的效度，有真实人类数据对照，批判性指出结构差异","model":"deepseek-v4-pro","scored_at":"2026-08-18T13:04:22","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-18","rank":16,"question":"教育测评工具在人类和LLM上是否测量相同的潜在构念？","design":"本研究不是仿真实验，而是比较分析。使用六种多模态LLM（GPT-4o、GPT-5.2、Gemini 1.5 Pro、Gemini 3 Pro、Claude 3.5 Sonnet、Claude 4.5）回答高中化学诊断测试和大学入学定量推理部分，与人类学生作答数据对比，通过探索性因子分析、因子一致性和重抽样比较潜在结构。","baseline":"人类基准为931名高中生的化学测试作答和4800多名考生的定量推理作答数据。","findings":"在两个测评工具上，人类和LLM的因子结构存在系统性差异，表明这些测评可能没有测量相同的构念。这质疑了使用教育测评来推断AI能力的效度。","reliability":"论文承认使用非公开数据集限制了可重复性，但强调这是为了确保人类数据质量。此外，因子结构相似只是构念等价的必要条件而非充分条件，且EFA是初步指标，需要进一步验证。","relevance":"该研究直接批判了用人类测评工具评估LLM能力的效度，提供了真实人类数据对照，并指出潜在结构差异，对关注LLM仿真可靠性和偏差的研究者很有价值。","inspiration":"借鉴其因子分析比较潜在结构的方法，可以检验经济行为测量工具（如风险偏好问卷、时间偏好量表）在人类和LLM上是否测量相同构念。｜可迁移到行为经济学中的偏好测量，例如用LLM模拟消费者跨期选择或风险决策，并检验其潜在结构与人类是否一致。｜以LLM作为被试，让其完成标准风险偏好问卷（如DOSPERT）和跨期选择任务，同时收集人类被试数据，对两组数据进行探索性因子分析并比较因子结构，以评估LLM作为人类被试替代品的效度。"}},{"id":"2608.16514","version":1,"title":"Matched Outcomes, Divergent Gaze: How Foveated MLLMs Search Compared to Humans","zh_title":"匹配结果，分歧注视：注视点式多模态大模型与人类搜索的对比","abstract":"Human visual search is serial: the fovea must land on a candidate to confirm it, and those landings form a scanpath. Whether multimodal large language models (MLLMs), given the same foveated input, search as humans do bears on their use as models of human vision and on attention-alignment scores. We compare three general-purpose MLLMs with human eye-movement scanpaths on goal-directed search (COCO-Search18), driving each model fixation by fixation through an identical, human-matched foveated view and assessing it along three axes: the decision of target presence, the efficiency of reaching the target, and the gaze process itself. The axes dissociate. On the decision and on target acquisition the models match or exceed humans, detecting present targets near ceiling and reaching them on the first saccade more often than people do. The gaze process is not human. Under the human-matched condition, all three share one signature: low-entropy, large-amplitude, self-consistent scanpaths that agree with themselves far more closely than two humans agree with each other. That is consistent with a single-pass, non-serial architecture rather than a limit of acuity. Matched retinal input reproduces where humans look but not how the looking unfolds in time, and no degradation regime recovers human-like search at human-like success. The gap sits on a process axis that answer-alignment and saliency metrics do not measure. Because they miss it, such metrics cannot certify human-like vision, and zero-shot models suit outcome and spatial questions but not temporal, process-level ones.","authors":["Mohamed Amine Kerkouri","Marouane Tliba","Aladine Chetouani","Ulas Bagci","Alessandro Bruno"],"categories":["cs.CV","cs.AI","cs.CL","cs.HC","cs.MM"],"primary_category":"cs.CV","announce_type":"cross","date":"2026-08-18","first_seen":"2026-08-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.16514","pdf_url":"https://arxiv.org/pdf/2608.16514","source_feed":"cs.AI","score":7,"bucket":"pending","rubric_hits":["A1","B1","B4"],"tags":["LLM仿真","视觉搜索","人类对照"],"reason":"用MLLM模拟人类视觉搜索，与人类眼动数据对照，并批判性指出过程差异，方法可迁…","model":"deepseek-v4-pro","scored_at":"2026-08-18T13:04:24","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-18","rank":20,"question":"在相同的中央凹视觉输入下，多模态大语言模型（MLLMs）的视觉搜索过程是否与人类相似？","design":"用三个通用MLLM（Qwen3.5-35B-A3B、GLM-4.6V-Flash、Gemma-4-E4B）在COCO-Search18数据集上进行目标导向视觉搜索，通过人类匹配的中央凹渲染器逐次注视输入，测量决策（目标有无）、效率（首次扫视命中率、注视次数）和注视过程（扫描路径熵、幅度、自一致性）。","baseline":"COCO-Search18数据集中10名人类观察者在相同场景上的眼动扫描路径和决策。","findings":"在决策和目标获取上，模型达到或超过人类水平（目标存在检测接近天花板，首次扫视命中率0.97/0.97/0.80 vs 人类0.49），但注视过程非人类：模型扫描路径低熵、大幅度、高度自一致（跨种子ScanMatch 0.84/0.91/0.71 vs 人类观察者间0.53），表明模型采用单遍非串行架构而非人类串行搜索。","reliability":"论文指出，答案对齐和显著性指标无法测量过程轴，因此不能认证人类视觉；零样本MLLM适合结果和空间问题，但不适合时间和过程问题。未讨论模型微调或不同提示策略的影响。","relevance":"该研究直接对比MLLM与人类眼动数据，并批判性指出过程差异，对关注LLM仿真可靠性与偏差的研究者具有重要参考价值，值得阅读原文以了解过程级评估方法。","inspiration":"借鉴其逐注视点施加相同感知约束（中央凹渲染）并分离结果与过程测量的设计，可迁移到经济决策中的信息搜索过程研究（如消费者浏览商品信息、投资者阅读财报），例如用LLM模拟投资者在财报披露后的信息获取，处理为不同信息呈现方式（如逐段揭示），结果变量为投资决策和注视路径，对照真实眼动实验数据。"}},{"id":"2608.16707","version":1,"title":"Semantic Bandits: In-Context Exploration-Exploitation is Biased by Semantic Priors","zh_title":"语义老虎机：上下文探索-利用受语义先验影响","abstract":"Large language models (LLMs) are increasingly deployed as decision-making agents in settings that require sophisticated environmental exploration. However, existing work has raised questions about how LLMs actually balance exploration and exploitation. Unlike classical agents, LLM agents engage with tasks through natural language, exposing them to semantic information with no formal counterpart in the task structure. We introduce the semantic bandit, an extension of the multi-armed bandit setting that explicitly considers the textual labels assigned to actions, and use it to study how semantic priors --- inductive biases arising from associations between language and expected reward learned during pre-training, shape LLM exploration behaviour. We find that semantically informative action labels reduce exploration in favour of exploitation, improving performance when aligned with the reward structure and severely degrading it when misaligned. We further find that negative rewards trigger substantially more exploration than equivalent positive rewards, consistent with an expected-scale bias induced by reward conventions common in pre-training data. Overall, we argue that the use of language to define the environment and rewards introduces unavoidable biases derived from the fact that the model is trained on word co-occurence, with implications for the reliability and robustness of LLM agents in real-world decision-making settings.","authors":["David Eric Austin","Kaheer Suleman","Jackie Chi Kit Cheung"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"cross","date":"2026-08-18","first_seen":"2026-08-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.16707","pdf_url":"https://arxiv.org/pdf/2608.16707","source_feed":"cs.AI","score":7,"bucket":"pending","rubric_hits":["A3","B4"],"tags":["LLM决策","探索-利用","语义偏差"],"reason":"研究LLM在语义多臂老虎机中的探索-利用行为，揭示语义先验导致的偏差，可迁移到…","model":"deepseek-v4-pro","scored_at":"2026-08-18T13:04:55","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-18","rank":21,"question":"LLM智能体在语义多臂老虎机任务中，动作标签的语义先验和奖励极性如何影响探索-利用行为？","design":"用GPT-4o等LLM作为决策智能体，在语义多臂老虎机任务中，通过改变动作标签的语义类型（字母数字、情感、序数、世界知识）和标签与奖励分布的对齐方式（有帮助/误导），以及奖励极性（正/负），测量探索行为（选择新臂的比例）和累积遗憾。","baseline":"无对照","findings":"语义信息丰富的动作标签会减少探索、偏向利用，当标签与奖励结构对齐时提升性能，错位时严重损害性能。负奖励比等效正奖励触发显著更多的探索，表明存在由预训练中奖励惯例引起的期望尺度偏差。","reliability":"论文指出语义先验的影响取决于先验与真实奖励分布的对齐程度，错位时会导致系统性错误；但未讨论其他失效条件或模型规模、提示设计等稳健性因素。","relevance":"该研究揭示了LLM在决策任务中因语言表征引入的语义偏差，对使用LLM仿真人类经济决策的可靠性提出警示，值得阅读以了解偏差来源和实验设计。","inspiration":"借鉴其通过操纵标签语义和奖励极性来分离语义先验影响的设计，可迁移到消费者选择或投资决策中的标签效应研究。｜可应用于金融产品推荐或政策选项呈现中的框架效应，检验LLM是否复现人类的选择偏差。｜用LLM模拟投资者，处理为不同语义标签的资产选项（如“稳健增长”vs“高风险高回报”），结果变量为选择比例和探索行为，与真实投资者实验数据对照。"}},{"id":"2608.15838","version":1,"title":"PersonaEval: Persona-Based User Simulation for Evaluating Interactive Applications","zh_title":"PersonaEval：基于角色的用户仿真用于评估交互式应用","abstract":"Real user studies are important for understanding how people interact with systems under test or already deployed. In practice, however, they are often costly, time-consuming, and difficult to scale. To address these challenges, we introduce PersonaEval, a persona-based user simulation framework that approximates real-user behavior across diverse interactive settings. PersonaEval connects simulated users drawn from existing persona datasets to task-specific application interfaces and collects the interaction trajectories and outcomes. PersonaEval provides a plug-and-play evaluation workflow in which the application being evaluated can be easily changed. In this demo, we present PersonaEval on three forms of interactive applications: surveys, chatbots, and web applications. Together, these examples show that PersonaEval can support repeatable, parallelizable, and scalable evaluation across different interaction settings, while producing user-oriented feedback and task-specific behavior.","authors":["Yifan Simon Liu","Qianfeng Wen","Yilan Fan","Shirley Huang","Ruoqi Gao","Jianheng Hou","Muhammad Ahmed Mohsin","Zonglin Di","Brihi Joshi","Xincheng Tan","Yucheng Lu","Xiaoyi Liu","Heming Liu","Hanwen Xing","Guanghui Min","Zhengyang Shan","My Chiffon Nguyen","Ishan Gupta","Yunze Xiao","Hannah Collison","Jintao Huang","Jiatong Li","Sankalp Jajee","Yunhan Zhao","Bing Hu","Sky Ng","Xupeng Chen","Binghang Lu","Weihang Xiao","Aravind Mohan","Bolun Sun","Yunshu Wu","Yuanda Xu","Yun Shen","Runyu Zhang","Zheyuan Deng","Zhiwei Zhang","Qianyu Zhu","Dianzhuo Wang","Yijun Wang","Yixuan He","Yuexing Hao","Xiaomin Li"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-08-18","first_seen":"2026-08-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.15838","pdf_url":"https://arxiv.org/pdf/2608.15838","source_feed":"cs.HC","score":7,"bucket":"pending","rubric_hits":["A1","A4"],"tags":["用户仿真","人机交互","评估框架"],"reason":"用 persona 模拟用户与交互应用交互，近似真实用户行为，可迁移到人类仿真…","model":"deepseek-v4-pro","scored_at":"2026-08-18T13:04:22","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-18","rank":18,"question":"如何构建一个可插拔的、基于人格的用户仿真框架，以近似真实用户行为并评估交互式应用？","design":"使用 PersonaEval 框架，从 Nemotron 人格数据集中按应用描述与人格档案的嵌入相似度选取 50 个人格作为模拟用户；通过应用适配器将模拟用户连接到调查问卷、聊天机器人和网页应用三种交互界面；使用 Claude Haiku 4.5 驱动人格用户，GPT-4o mini 驱动聊天机器人；收集交互轨迹和结果，并测量用户评分、结果变异和人格一致性。","baseline":"无对照","findings":"PersonaEval 能在调查、聊天机器人和网页应用中产生人格一致的模拟交互，并揭示不同应用和人格群体间的评分差异。聊天机器人任务满意度中等且分布更广，而调查和网页结果更集中；美容、医疗和电影等应用的人格群体评分差异更大。","reliability":"论文承认当前演示覆盖的应用范围有限，仿真质量尚未通过真实用户行为对比进行严格验证，人格一致性仅由人类评判，未来需校准仿真与人类数据。","relevance":"该研究展示了基于人格的 LLM 用户仿真在多种交互场景下的可行性，但缺乏真实人类数据对照，对关注仿真可靠性与偏差的研究者价值有限，可作为方法参考而非实证基准。","inspiration":"可借鉴其模块化人格仿真框架和基于嵌入相似度的人格选择方法，用于构建多样化的经济主体样本。｜可迁移到消费者金融决策、信贷申请或政策反馈等场景，模拟不同背景人群的行为差异。｜以人格化的 LLM 作为被试，施加不同的金融产品信息或政策干预，测量其选择、满意度或风险偏好，并与真实调查或实验数据对照以校准仿真。"}},{"id":"2608.16067","version":1,"title":"SiMUSation: An Interactive Visitor Experience Simulation Framework to Support Museum Exhibition Design","zh_title":"SiMUSation：支持博物馆展览设计的交互式访客体验仿真框架","abstract":"Understanding how diverse audiences engage with narratives and content is central to exhibition design, yet designers often rely on intuition. Existing experience evaluation methods are typically retrospective, costly, and offer limited access to visitors' internal states, hindering early-stage iterative refinement. Rather than relying only on post-implementation evaluation with real visitors, we explore LLM-driven persona simulation as a reference for early-stage design. Following this idea, we present SiMUSation, an interactive framework designed to support early-stage exhibition design. SiMUSation models diverse visitor personas and simulates their exhibition experiences through a dual-layer representation that couples observable behaviors, such as movement and gaze, with corresponding internal responses, such as confusion and narrative engagement. Designers can steer simulations, inspect feedback from simulated visits, and iteratively revise layouts, content, and narrative flow to further examine how changes reshape visitor experience. We implemented a prototype and evaluated it through a user study (N=12), showing that SiMUSation provides insights for reflection and refinement in early-stage exhibition design. Our findings further highlight the potential of persona-driven simulation to support audience-informed evaluation and iterative decision-making across design tasks.","authors":["Huanchen Wang","Qiuming Chen","Zhonghao Ji","Ruqi Sun","Zhichao Lu","Yuxin Ma"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-08-18","first_seen":"2026-08-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.16067","pdf_url":"https://arxiv.org/pdf/2608.16067","source_feed":"cs.HC","score":7,"bucket":"pending","rubric_hits":["A1","A3","B1"],"tags":["LLM仿真","用户体验","博物馆设计"],"reason":"用LLM模拟访客体验并与真实用户研究对照，属于人类仿真且有人类数据基准。","model":"deepseek-v4-pro","scored_at":"2026-08-18T13:04:22","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-18","rank":19,"question":"如何利用LLM驱动的访客角色仿真来支持博物馆展览的早期设计迭代？","design":"使用LLM构建多样化的访客角色，通过双层表示模拟其观展体验：可观察行为（移动、注视）与内部反应（困惑、叙事参与）。设计师可调整布局、内容和叙事流程，观察仿真反馈以迭代设计。","baseline":"用户研究（N=12）评估原型可用性和有效性，作为仿真结果的对照。","findings":"SiMUSation能为设计师提供可操作的访客体验洞察，帮助识别问题所在（通过行为轨迹）和原因（通过内部反应）。用户研究表明该框架支持早期设计反思和细化。","reliability":"论文未讨论仿真失效条件，但指出传统方法难以捕捉内部状态，且用户研究规模小（N=12），可能影响代表性。","relevance":"该研究将LLM仿真用于设计评估，与人类仿真研究相关，但缺乏与真实访客数据的直接对比，且场景为博物馆设计，对经济学实验的参考价值有限。","inspiration":"借鉴其双层仿真设计，将可观察行为与内部状态结合，用于经济决策仿真。｜可迁移到消费者行为研究，如购物决策中的注意力分配与偏好形成。｜以LLM模拟消费者在虚拟商店中的浏览路径和购买决策，处理为不同商品陈列方式，结果变量为购买率和内部偏好，对照真实消费者眼动和购买数据。"}},{"id":"2608.15181","version":1,"title":"Insurance as AI Risk Infrastructure: A Generative-Agent Simulation of AI Adoption","zh_title":"保险作为AI风险基础设施：AI采用的生成式智能体仿真","abstract":"The rapid evolution of artificial intelligence (AI) tools has demonstrated immense potential to enhance societal well-being and operational efficiency. However, the inherent unreliability and uncertain operational consequences of modern AI systems, typified by large language models (LLMs), have created a significant barrier to enterprise adoption. Many enterprises remain hesitant to integrate these tools deeply into their workflows due to concerns about unpredictable losses and liability exposure. While existing technical safeguards primarily seek to reduce the likelihood or severity of AI-enabled workflow failures, they do not by themselves provide ex post financial protection when residual pecuniary tail losses materialize. In this paper, we introduce a socio-economic framework that complements these safeguards by transferring and absorbing the residual financial consequences of AI adoption through insurance. To evaluate this framework, we develop an LLM-driven agent-based social simulation (LABSS) system. We assess the behavioral validity of the simulation using established economic and sociological theories. Our analysis demonstrates that the proposed insurance framework reduces firm-level financial exposure, thereby accelerating the aggregate adoption of AI tools and improving firm solvency and aggregate capital.","authors":["Yixuan Yuan","Dedai Wei","Chudong Qian","Jielin Feng","Ziyue Lin","Yuheng Zhao","He Cao","Erasmo Purificato","Xinwu Ye"],"categories":["cs.MA"],"primary_category":"cs.MA","announce_type":"new","date":"2026-08-18","first_seen":"2026-08-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.15181","pdf_url":"https://arxiv.org/pdf/2608.15181","source_feed":"cs.MA","score":7,"bucket":"pending","rubric_hits":["A3","B2","B4"],"tags":["LLM社会仿真","AI采用","保险机制"],"reason":"用LLM agent模拟企业AI采用决策，涉及经济场景，但无真实人类数据对照，…","model":"deepseek-v4-pro","scored_at":"2026-08-18T13:04:20","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-18","rank":15,"question":"保险能否作为AI风险基础设施，通过转移和吸收AI采用中的残余财务损失，加速企业采用AI工具并改善企业偿付能力与总体资本？","design":"使用LLM驱动的智能体社会仿真系统（LABSS），模拟异质性企业在300天内做出AI采用、续约和保险决策，处理为是否提供保险，结果变量包括AI采用率、企业破产数、总资本、供应商承诺时长和跨行业采用差距。","baseline":"无对照","findings":"提供保险使最终AI采用率从74.97%提高到84.38%，平均破产企业数从4.33降至1.33，总资本更高。保险还带来更长的供应商承诺、更小的跨行业采用差距，并抑制恐慌传播。","reliability":"论文未讨论","relevance":"该研究用LLM智能体模拟企业AI采用决策，属于经济场景下的人类仿真，但无真实人类数据对照，与您关注的有基准的仿真研究不完全匹配，可快速浏览其仿真设计。","inspiration":"借鉴其用LLM智能体模拟企业决策并施加保险处理的做法，可迁移到企业技术采用或风险管理决策的经济金融问题，例如企业对新金融技术的采用。｜设计一个LLM智能体仿真，让企业智能体在有无保险条件下决定是否采用AI，结果变量为采用率和破产率，并用真实企业调查或历史采用数据做对照。"}},{"id":"2608.14613","version":1,"title":"Do LLM Agents Negotiate Rationally? A Mechanism-Design Framework for Verifiable Multi-Agent Interaction over A2A/MCP","zh_title":"LLM智能体是否理性谈判？基于A2A/MCP的可验证多智能体交互机制设计框架","abstract":"Modern LLM-agent frameworks increasingly interoperate through standards such as Anthropic's Model Context Protocol (MCP) for agent-to-tool access and Google's Agent2Agent (A2A) protocol for agent delegation and negotiation. However, these protocols specify transport and discovery rather than strategic correctness and do not guarantee efficient, individually rational, or strategy-proof outcomes. We introduce a framework that (i) encodes classical negotiation mechanisms, including alternating-offers bargaining and Vickrey-Clarke-Groves-style auctions, as constraints over A2A message schemas; (ii) provides a lightweight runtime verification and repair layer that checks messages against protocol invariants; and (iii) offers a benchmark of negotiation and allocation tasks with known optimal solutions for measuring deviations from game-theoretic predictions. We evaluate multiple LLM backbones using unstructured dialogue, structured protocols, and structured protocols with verification. Across negotiation trials (N=30 per condition), verification reduces outcome variance, while structured protocols achieve 100 percent success for both models. After correcting parser artifacts, audited unstructured baselines achieve approximately 97 percent and 93.3 percent success. In auction experiments (N=30 per model), both models achieve 100 percent efficient allocation but differ sharply in truthful bidding: one bids its exact valuation in every trial, whereas the other does so in only 3.3 percent of trials. Thus, mechanism-level incentive compatibility does not automatically transfer to LLM-agent behavior. A three-party fair-allocation task produced only 4.2 percent usable outcomes; we report this negative result with a diagnosis. This work bridges classical multi-agent systems theory and modern LLM-agent infrastructure and defines verifiable interaction at the A2A protocol layer.","authors":["Wael Albayaydh","Rui Zhao"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-18","first_seen":"2026-08-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.14613","pdf_url":"https://arxiv.org/pdf/2608.14613","source_feed":"cs.AI","score":6,"bucket":"other","rubric_hits":["D3"],"tags":["LLM智能体","机制设计","谈判与拍卖"],"reason":"用LLM agent模拟谈判与拍卖，属经济过程模拟，但无真实人类数据对照，属边…","model":"deepseek-v4-pro","scored_at":"2026-08-18T13:04:30","error":null,"has_summary":false,"summary":null},{"id":"2608.16196","version":1,"title":"Beyond Asking: A Pipeline for Personalized Game Generation that Reads Players from Behavior","zh_title":"超越询问：一种从行为读取玩家的个性化游戏生成流水线","abstract":"Personalized game generation requires inferring a player's abilities and behavioral style from how they play. Large language models have made this inference more attainable than ever: an LLM can read a raw gameplay transcript and produce a fluent, plausible profile of the player. Plausible, however, is not verified, and verification is precisely what the field lacks: latent traits are unobservable; questionnaires provide noisy proxies and become circular when self-reports are used to validate behavior-based inference; and behavior itself is ambiguous without context -- a player who never collects an item may not want it, or may never have had the chance. We address both problems. First, we construct a synthetic player population whose traits are ground truth by construction: each trait is an explicit bot parameter, accepted only after controlled manipulation produces consistent, trait-specific behavioral change. Unlike prior parameter-recovery work that inverts a known decision model, our benchmark evaluates policy-agnostic inference from behavioral transcripts alone. Second, we introduce an opportunity-aware decision-moment representation that disentangles preference from the chance to express it; ablating it selectively degrades opportunity-dependent traits. On this benchmark, few-shot LLM inference outperforms embedding- and rule-based baselines on most traits, though feature-based supervised regressors remain stronger overall. Finally, we close the loop: inferred profiles drive difficulty adaptation, evaluated against ground-truth references and mismatched-profile controls, and an exploratory human study examines whether these findings transfer to real players.","authors":["Yifan Lu","Xiaopeng Yuan","Haohan Wang"],"categories":["cs.AI","cs.HC"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-18","first_seen":"2026-08-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.16196","pdf_url":"https://arxiv.org/pdf/2608.16196","source_feed":"cs.AI","score":6,"bucket":"other","rubric_hits":["D3"],"tags":["LLM行为推断","个性化游戏生成","合成玩家群体"],"reason":"用LLM从行为推断玩家特质并生成个性化游戏，属社会模拟但无真实人类数据对照，且…","model":"deepseek-v4-pro","scored_at":"2026-08-18T13:04:48","error":null,"has_summary":false,"summary":null},{"id":"2608.14792","version":1,"title":"Prompting is not enough: supervised baselines and leakage control for measuring shared decision-making with LLMs in pediatric encounters","zh_title":"提示不足：在儿科就诊中测量共享决策时，监督基线与泄漏控制的重要性","abstract":"Objectives: To determine whether zero-shot prompting of a large language model (LLM) is sufficient to detect shared decision-making (SDM) behaviors in real clinical encounters, and whether supervised learning adds value under patient-grouped, nested evaluation. Methods: We analyzed 21 audio-recorded outpatient surgical decision encounters (19 unique patients; 7,566 utterance segments; ~6.1 hours) between families of children with multiple long-term conditions and their surgical providers. Trained coders labeled segments for 12 SDM behaviors (human-human macro Cohen's kappa = 0.695). We compared a zero-shot local LLM (Qwen 2.5 32B), a supervised classifier over frozen sentence embeddings, and their logistic stack, under patient-grouped outer folds with inner cross-fitted thresholds and patient-resampled confidence intervals. Results: The zero-shot LLM reached macro kappa = 0.139 (95% CI 0.111-0.164). The supervised classifier reached kappa = 0.227 (0.186-0.262), a paired improvement of 0.088 (0.051-0.119). A logistic stack of the two reached kappa = 0.242 (0.198-0.284). We identified multiple corpus-specific leakage paths, including grouping sibling recordings separately and allowing labels from an outer held-out patient to enter few-shot exemplars used while fitting downstream models. Conclusion: Zero-shot prompting alone is not sufficient to measure SDM behavior as reliably as a small supervised model, and patient-level grouping alone does not prevent leakage when labeled prompt exemplars are precomputed outside the outer evaluation loop. Reported performance is sensitive to the unit of data splitting and to where labeled exemplars enter the pipeline. External validation is needed before these findings generalize beyond this population, model, prompt, and codebook.","authors":["Bernardo Modenesi","Jody Lin","Kimberly Kaphingst","Angela Zhu","Maya Wheeler","Peilu Zhang","Angela Fagerlin"],"categories":["cs.CL","cs.AI","cs.LG"],"primary_category":"cs.CL","announce_type":"cross","date":"2026-08-18","first_seen":"2026-08-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.14792","pdf_url":"https://arxiv.org/pdf/2608.14792","source_feed":"cs.AI","score":6,"bucket":"other","rubric_hits":["D1"],"tags":["LLM标注","临床对话","共享决策"],"reason":"用LLM检测临床对话中的共享决策行为，替代人工标注，而非仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-08-18T13:04:18","error":null,"has_summary":false,"summary":null},{"id":"2608.14913","version":1,"title":"The Open-Strategy Dictator Game: Cooperation Under Mutual Transparency","zh_title":"开放策略独裁者博弈：相互透明下的条件合作","abstract":"We introduce the Open-Strategy Dictator Game (OSDG), a variant of the classic dictator game in which each player's strategy is a natural-language document visible to all participants. The dictator's decision, to SHARE or TAKE an endowment, may depend on the text of the recipient's strategy. A large language model adjudicates each interaction by interpreting the dictator's strategy in the context of the recipient's. We run round-robin tournaments among diverse strategies and analyze the resulting payoff matrix using softmax equilibrium frequencies, dominance analysis, and sensitivity to the relative value of cooperation. Conditionally cooperative strategies, those that share with cooperators and take from exploiters, consistently dominate, while unconditional strategies (always share or always take) are weakly dominated. The results suggest that in environments where agents can inspect each other's decision procedures, conditional cooperation is evolutionarily robust across a wide range of payoff parameters.","authors":["Michael Glass"],"categories":["cs.GT","cs.AI","cs.MA"],"primary_category":"cs.GT","announce_type":"cross","date":"2026-08-18","first_seen":"2026-08-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.14913","pdf_url":"https://arxiv.org/pdf/2608.14913","source_feed":"cs.AI","score":6,"bucket":"other","rubric_hits":["D3"],"tags":["LLM agent","博弈论","社会模拟"],"reason":"用LLM agent模拟经济博弈，但无真实人类数据对照，属社会模拟边界情形。","model":"deepseek-v4-pro","scored_at":"2026-08-18T13:04:20","error":null,"has_summary":false,"summary":null},{"id":"2608.14893","version":1,"title":"Weaker Coherence, Weaker Reciprocity: Comparing the Semantic and Social Organization of Moltbook and Reddit","zh_title":"更弱的连贯性，更弱的互惠性：比较Moltbook与Reddit的语义和社会组织","abstract":"Large language models enable the creation of autonomous agents that interact in social environments, raising the question of whether agent-based platforms reproduce the organizational properties of human social networks. We compare Moltbook, a social network populated by AI agents, with early Reddit, focusing on how communities organize and differentiate semantic content, using network analysis and NLP methods to characterize semantic coherence and diversity within and between communities, and their relationship to user activity. We find a systematic difference between the two platforms. Reddit communities show stronger semantic coherence, closer alignment with community names, and greater semantic diversity, with individual communities spanning broader content and communities more differentiated from one another. This combination distinguishes Reddit from Moltbook, whose communities are more homogeneous, less differentiated, and increasingly misaligned with their names over time. Users on Reddit also participate across communities that are more semantically related than those connected by activity in Moltbook. At the interaction level, comment-network motif analysis shows Moltbook dominated by non-reciprocal, broadcast-like exchanges, whereas Reddit shows more reciprocal, chained interaction patterns. These results indicate that Reddit combines semantic coherence with diversity across organizational levels, a pattern not reproduced by the AI-agent network.","authors":["Favio Di Ciocco","Lucas D\\'iaz Celauro","Sebasti\\'an Pinto","Marcelo Kuperman","Pablo Balenzuela"],"categories":["cs.SI","physics.soc-ph"],"primary_category":"cs.SI","announce_type":"new","date":"2026-08-18","first_seen":"2026-08-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.14893","pdf_url":"https://arxiv.org/pdf/2608.14893","source_feed":"cs.SI","score":6,"bucket":"other","rubric_hits":["D3"],"tags":["AI社交网络","社会模拟","语义分析"],"reason":"用AI agent模拟社交网络并与真实Reddit对照，但非人类被试仿真，属社…","model":"deepseek-v4-pro","scored_at":"2026-08-18T13:04:20","error":null,"has_summary":false,"summary":null},{"id":"2608.15519","version":1,"title":"Topological collapse of higher-order interactions bottlenecks collective intelligence in AI agent societies","zh_title":"高阶交互的拓扑坍缩制约AI智能体社会的集体智能","abstract":"Current paradigms in artificial intelligence concentrate on scaling the capabilities of individual models, yet the collective behaviour of interacting agents is shaped by the topology of their interactions rather than by individual cognition alone. Here we show that the binding constraint on collective behaviour in agent societies is topological. Analysing a macroscopic AI social platform of 1.6 million registered agents (174,458 active in the interaction record), we identify a phenomenon we term topological collapse: extreme hub dominance degrades higher-order group interactions into star-shaped broadcast patterns, suppressing the cohesive structure that discontinuous social contagion requires. We formalise this constraint through a Hyperedge Irreducibility Score (HIS) and an analytical topology amplification factor ($\\Phi$). Across 22 frontier language models from ten vendors, 1,040 controlled simulations and empirical human networks, the bottleneck proves model-agnostic: under a fixed interaction protocol the topological indicators are invariant across models (cross-model HIS s.d. = 0.000 in the pairwise condition) even as behavioural outcomes diverge widely. These findings reframe the design of artificial societies around the geometry of interaction rather than the optimisation of individual cognition, with implications for AI sociology, algorithmic group dynamics, hybrid human-AI ecosystems and collective alignment. The code is publicly available at https://github.com/Darwin-Agent/topological-collapse-agent-societies.","authors":["Shuo Lu","Weicheng Meng","Aijing Yu","Kun Shao","Jian Luan","Ran He","Jian Liang"],"categories":["cs.SI"],"primary_category":"cs.SI","announce_type":"new","date":"2026-08-18","first_seen":"2026-08-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.15519","pdf_url":"https://arxiv.org/pdf/2608.15519","source_feed":"cs.SI","score":6,"bucket":"other","rubric_hits":["D3"],"tags":["多智能体社会模拟","网络拓扑","集体行为"],"reason":"用LLM agent群体模拟社会互动，但无真实人类行为对照，属边界情形","model":"deepseek-v4-pro","scored_at":"2026-08-18T13:04:20","error":null,"has_summary":false,"summary":null},{"id":"2607.23982","version":5,"title":"Moral Hazard in Multi-Agent Language Models","zh_title":"多智能体语言模型中的道德风险","abstract":"Cooperation can fail when socially valuable effort is costly, hard to observe, and benefits mainly someone else. Building on Holmstr\\\"om's model of moral hazard in teams, we introduce the Dialogue Moral Hazard Game, a theory-grounded controlled experimental paradigm that instantiates this hidden-action structure as a textual environment for language agents. In each episode, an agent chooses between keeping an immediate local reward and paying a query cost to reveal a hidden safety fact that primarily helps another agent's downstream decision. We evaluate thirteen open-weight and four frontier models with stage-level mechanism metrics. In matched 3,015-decision-per-model experiments, GPT-5.6 Sol and Claude Opus 4.8 track the Holmstr\\\"om-derived private-share boundary across nine query costs (mean absolute errors 0.013 and 0.030); Muse Spark 1.1 responds directionally, whereas Fable 5 remains query-saturated. Diagnostic SFT, RLOO, SFT+RLOO, and GEPA updates are heterogeneous: SmolLM3-3B and OLMo-7B show the clearest weight-level mechanism gains, while GEPA raises Muse team success from $22.2\\pm3.8\\%$ to $100.0\\pm0.0\\%$ as query use falls from $51.1\\pm5.1\\%$ to $0.3\\pm0.5\\%$. Freezing the three Muse prompts and intervening on the rank--label mapping changes team success from $100.0\\%$ to $12.5\\%$ and then $0.0\\%$, with validity fixed at $100\\%$. Opus supplies a within-model contrast: its query-mediated prompt remains perfect across mappings, while two near-zero-query prompts follow the same trajectory. Optimization can therefore reach the same aggregate outcome through direct revelation or a learned effective information structure, motivating mechanism-level evaluation rather than team success alone.","authors":["Dane Malenfant"],"categories":["cs.MA","cs.AI"],"primary_category":"cs.MA","announce_type":"replace-cross","date":"2026-08-18","first_seen":"2026-07-28","revised_at":"2026-08-18","abs_url":"https://arxiv.org/abs/2607.23982","pdf_url":"https://arxiv.org/pdf/2607.23982","source_feed":"cs.AI","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["多智能体","道德风险","社会模拟"],"reason":"多智能体道德风险博弈，无真实人类数据对照，属社会模拟边界情形","model":"deepseek-v4-pro","scored_at":"2026-08-18T13:04:57","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-15","rank":3,"question":"语言智能体在多智能体团队道德风险博弈中，是否会因私人成本与收益份额而减少对社会有益但成本高昂的查询行为？","design":"构建基于Holmström团队道德风险模型的文本对话博弈环境，让多个LLM智能体扮演团队成员，通过改变查询成本、团队奖励和私人产出份额等参数，测量查询率、信息传递、局部奖励保留、不安全选择、格式有效性和团队成功等行为指标。","baseline":"无对照","findings":"前沿模型行为差异显著：GPT-5.6 Sol在主要设置中达到上限行为，并在激励隔离实验中以0.013的平均绝对误差追踪Holmström推导的私人份额边界；优化方法（如GEPA）可能提高团队成功但减少或消除代价高昂的查询，表明聚合奖励提升不一定恢复指定合作机制。","reliability":"论文未讨论","relevance":"该研究用LLM模拟经济理论中的道德风险场景，属于经济实验仿真，但缺乏真实人类数据对照，适合关注LLM行为与理论预测一致性的研究者阅读。","inspiration":"借鉴其参数化激励结构（成本、奖励、份额）和机制分解测量方法，可系统检验LLM对经济激励的边际响应。｜可迁移到委托代理问题，如基金经理努力与风险承担、员工团队合作中的搭便车行为。｜以LLM作为被试，在模拟投资任务中改变管理费率和业绩提成比例，测量其风险选择和努力程度，并与真实基金经理的历史数据对照。"}},{"id":"2608.08601","version":2,"title":"Unaccountable Delegation, Fading Skills: Mapping the Risks of Workplace AI Agents","zh_title":"不可问责的委托与技能衰退：绘制职场AI代理的风险地图","abstract":"To anticipate socio-technical risks from AI agents, organizations need taxonomies to classify them. However, existing AI risk taxonomies focus on broad risks and do not capture job-specific risks introduced by agents. To address this gap, we make three main contributions. First, we developed a multi-layer framework from a literature review of AI agents. The framework models three core components and their interactions: agents, goals, and environment. Second, we embedded this framework in a structured prompt and applied it to descriptions of 2,078 job tasks from the O*NET database, producing 8,356 risk scenarios labeled by severity and deployment mode (automation or augmentation). We validated these scenarios with 45 workers across 10 job roles and an independent LLM judge, confirming their plausibility and alignment with job tasks. Finally, we extended an existing taxonomy to create a 15-category taxonomy of workplace AI agent risks that covers all our risk scenarios. Our analysis highlights four findings. First, augmentation is not inherently safe because overreliance on agents can gradually erode workers' skills and oversight. Second, Erroneous Agent Actions accounts for the largest share of risk scenarios and has the highest concentration of severe risks. Many arise at the human-agent boundary. Third, automation is associated mainly with organizational risks, while augmentation is associated mainly with risks to workers. Fourth, workers found our taxonomy easier to use for a risk classification task than two other taxonomies and preferred it in 64% of non-tied comparisons with a recent generative AI risk taxonomy. These findings show that workplace AI agent risks do not arise from agents alone; they also depend on how people work with agents and how agents are deployed. Safer workplaces require not only safer agents but also carefully designed human-AI agent collaboration.","authors":["Gabriele La Malfa","Lakmal Meegahapola","Edyta Bogucka","Jie M. Zhang","Michael Luck","Elizabeth Black","Daniele Quercia"],"categories":["cs.AI","cs.MA"],"primary_category":"cs.AI","announce_type":"replace","date":"2026-08-18","first_seen":"2026-08-11","revised_at":"2026-08-18","abs_url":"https://arxiv.org/abs/2608.08601","pdf_url":"https://arxiv.org/pdf/2608.08601","source_feed":"cs.AI","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["AI风险分类","人机协作","工作场所"],"reason":"用LLM生成风险场景并验证，替代人工标注，非仿真人类被试行为。","model":"deepseek-v4-pro","scored_at":"2026-08-11T13:05:03","error":null,"has_summary":false,"summary":null},{"id":"2608.14552","version":1,"title":"Large Language Models Show Metacognitive Sensitivity in Medical Reasoning","zh_title":"大型语言模型在医学推理中表现出元认知敏感性","abstract":"Large language models (LLMs) are increasingly evaluated and used in medicine, but clinical usefulness depends on answer accuracy and whether confidence tracks evidence quality and uncertainty. We developed a controlled, psychophysics-inspired clinical benchmark to test diagnostic choice and confidence behavior in a medical LLM. The benchmark focused on probable Alzheimer-type neurocognitive disorder (AT-NCD) versus depression-related cognitive impairment (DRCI). We generated 45 synthetic vignettes varying evidence strength, conflicting evidence, and missing information. Each vignette was presented under three prompt variants, yielding 135 trials. In a pilot run with gpt-4.1-nano, all trials produced valid structured outputs. Across forced-choice trials, diagnostic accuracy was 93.5%, mean confidence was 78.4%, and AUROC2 was 0.876. Confidence increased with evidence distance from the diagnostic boundary, decreased when information was missing, and remained higher on correct than incorrect trials after adjustment for evidence strength and prompt format. These findings indicate partial metacognitive sensitivity rather than globally uninformative confidence. However, errors clustered in moderate, conflicting AT-NCD cases, where the model shifted toward DRCI and retained more confidence than empirical accuracy justified. Model comparison suggested that confidence quality should be measured directly rather than inferred from benchmark accuracy or model capability alone. This study establishes a reproducible framework for evaluating evidence sensitivity, metacognitive sensitivity, and localized calibration failure in medical LLMs.","authors":["Ahmad Nazzal"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-18","first_seen":"2026-08-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.14552","pdf_url":"https://arxiv.org/pdf/2608.14552","source_feed":"cs.AI","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["LLM评估","置信度校准","医学推理"],"reason":"研究LLM的元认知敏感性，测量模型自身置信度，非仿真人类被试，但方法可借鉴。","model":"deepseek-v4-pro","scored_at":"2026-08-18T13:04:16","error":null,"has_summary":false,"summary":null},{"id":"2608.14566","version":1,"title":"Position: Evaluations of AI Moral Reasoning Still Miss Half of the Picture","zh_title":"立场：AI道德推理评估仍遗漏一半图景","abstract":"Recent work on evaluating the moral competence of large language models (LLMs) has focused primarily on what we call the moral value problem, i.e., whether model outputs align with human moral values. In contrast, the moral norm problem, i.e., whether models can identify and correctly apply context-sensitive moral norms, remains underexplored. We posit that this imbalance stems from the field's reliance on descriptive ethics frameworks, such as Moral Foundations Theory and Kohlberg's stages of moral development, which emphasize value representation over normative application. We review existing benchmarks and evaluation methods, and show that they cluster heavily around the value problem, while discussion regarding normative ethics remains underrepresented. We identify three crucial gaps: (i) the absence of high-quality ground-truth data for moral norms and their applications, (ii) insufficient evaluation of intermediate reasoning processes, and (iii) limited attention to the identification of morally relevant features in context. Subsequently, we propose a research agenda that includes the development of standardized formal representations for normative theories, the construction of expert-annotated datasets capturing norm application, and evaluation protocols that explicitly distinguish between values-level and norms-level competence. Our goal is to encourage a more systematic study of normative reasoning in LLMs.","authors":["Aidan Kierans","Ritam Dutt","Kaley Rittichier","Shiri Dori-Hacohen","Avijit Ghosh"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-18","first_seen":"2026-08-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.14566","pdf_url":"https://arxiv.org/pdf/2608.14566","source_feed":"cs.AI","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["LLM道德评估","价值观对齐","基准偏差"],"reason":"评估LLM道德推理，测的是模型而非人类被试，但涉及价值观对齐，与仿真可靠性相关","model":"deepseek-v4-pro","scored_at":"2026-08-18T13:04:28","error":null,"has_summary":false,"summary":null},{"id":"2608.14622","version":1,"title":"A Human-Centred Approach to Benchmarking LLMs for Parenting Advice","zh_title":"以人为中心的方法评估大语言模型在育儿建议上的表现","abstract":"People are increasingly using large language models (LLMs) to seek advice, including for parenting. Parenting is a critical and socially sensitive domain. Thus, evaluating advice provided by LLMs requires indicators beyond aggregated information quality benchmarks to consider relational and behavioural elements of the responses. With a multi-dimensional rubric created by parenting experts, this paper evaluates 15 LLMs across 100 parenting scenarios in 2 languages (English and Chinese), using an LLM-as-a-judge method. Results show that aggregate scores can hide rubric item-specific weaknesses, models implicitly encourage different parenting styles, and language influences responses. We highlight the importance of evaluation output auditability and challenges involved in evaluating LLM-generated advice in domains like parenting. Our findings provide important insights for selecting LLMs for direct user engagement and the development of user-facing parenting advice applications.","authors":["Yunke Zhao","Isobel Voysey","Alastair van Heerden","Rob Hughes","Jun Zhao"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-18","first_seen":"2026-08-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.14622","pdf_url":"https://arxiv.org/pdf/2608.14622","source_feed":"cs.AI","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["LLM评估","育儿建议","基准测试"],"reason":"评估LLM给出的育儿建议，属于对模型输出的测量，而非用LLM仿真人类被试，无人…","model":"deepseek-v4-pro","scored_at":"2026-08-18T13:04:30","error":null,"has_summary":false,"summary":null},{"id":"2608.15101","version":1,"title":"Second-Order Policy Effects as State Transitions: A Source-Linked Benchmark for Policy Simulation","zh_title":"二阶政策效应作为状态转移：一个源链接的政策模拟基准","abstract":"Policy evaluation often estimates direct benefits and costs while treating the institutional environment as fixed. In practice, a policy changes the system it enters: actors adapt, enforcement capacity shifts, burdens move, and new equilibria form around capture, gaming, compliance theater, irreversibility, and repair costs. We formalize this as second-order policy-effect prediction and present a source-linked benchmark for policy simulation. The benchmark contains 96 named public-policy cases across eight domains and four balanced action classes: implement, modify, pilot, and block. Each case includes source locators and state variables for benefit, capture, gaming, burden shift, instability, uncertainty, irreversibility, distributional risk, and implementation capacity. The runner regenerates method outputs and aggregate results from the case table, and the simulator never reads the expert action target. We report a protocol-based transition-channel audit with recall, precision, F1-style efficiency, and selective top-channel stress diagnostics, so universal channel coverage is not mistaken for field validation. The side-effect simulator achieves mean policy-effect quality of 0.945, compared with 0.838 for the risk-register baseline and 0.879 for the causal-loop baseline. Its advantage is concentrated in side-effect recall and aggregate transition scoring; it does not dominate the best structured baselines on exact policy-action choice. The evidence remains benchmark-based, but supports a bounded claim: transition-state variables make policy simulators more sensitive to downstream institutional effects.","authors":["Wesley Shu"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-18","first_seen":"2026-08-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.15101","pdf_url":"https://arxiv.org/pdf/2608.15101","source_feed":"cs.AI","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["政策模拟","LLM仿真","基准测试"],"reason":"用LLM模拟政策效果，但无真实人类数据对照，属社会模拟边界情形","model":"deepseek-v4-pro","scored_at":"2026-08-18T13:04:36","error":null,"has_summary":false,"summary":null},{"id":"2608.15354","version":1,"title":"Incoherent by Design? On the Moral Self-Consistency of LLMs","zh_title":"设计上的不一致？论大语言模型的道德自我一致性","abstract":"LLMs are increasingly used in morally sensitive contexts, yet it is unclear whether they apply ethical principles consistently across situations. A model that can state a moral principle may still violate it when the same scenario is rephrased or reframed. This inconsistency is a problem for any system whose outputs are used to inform moral decisions. If generative systems exhibit internal inconsistency, then the epistemic integrity of AI-mediated systems becomes uncertain. To study this concern, we investigate the stability of moral reasoning in LLMs within a controlled prompting framework across three major philosophical schools of thought: deontology, utilitarianism, and virtue ethics. We construct sets of morally equivalent scenarios in which the underlying situation is held constant while the framing varies to reflect different ethical stances and stylistic perturbations. We then evaluate responses from multiple models, including GPT, Mistral, and Llama. To assess consistency, we convert model outputs into structured logical statements and identify contradictions across responses generated within the same school of thought. Our results reveal substantial inconsistency with contradiction rates reaching up to 78% across scenarios. These findings point to a broader phenomenon of epistemic instability in generative AI wherein models fail to reliably maintain coherence with respect to their own prior outputs. This kind of instability carries real consequences. As generative systems influence how people form beliefs, judge actions, and absorb values, their inconsistencies can shape human reasoning and decision-making as well. Moreover, if a system cannot consistently represent its own normative commitments, then value alignment becomes a moving target rather than a well-defined objective. Thus, we argue that demonstrating internal incoherence is a necessary precursor to AI alignment.","authors":["Pegah Nokhiz","Aravinda Kanchana Ruwanpathirana","Helen Nissenbaum"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-18","first_seen":"2026-08-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.15354","pdf_url":"https://arxiv.org/pdf/2608.15354","source_feed":"cs.AI","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["道德推理","一致性","模型评估"],"reason":"测量LLM道德推理一致性，属模型自身属性，非仿真人类被试，但涉及价值观稳定性，…","model":"deepseek-v4-pro","scored_at":"2026-08-18T13:04:41","error":null,"has_summary":false,"summary":null},{"id":"2608.16507","version":1,"title":"Large language models as synthetic clinical experts to inform longitudinal rare-disease modeling","zh_title":"大语言模型作为合成临床专家用于纵向罕见病建模","abstract":"Due to the limited amount of information, modeling longitudinal rare-disease data can benefit from integrating clinical knowledge. Yet, elicitation of expert knowledge and formalization for model fitting is challenging, in particular due to limited time of clinical experts. To nevertheless make domain knowledge accessible during model fitting, we use large language models (LLMs) as synthetic clinical experts to supervise a variational-autoencoder-based approach that learns low-dimensional latent summaries of visit-level observations. Specifically, LLMs are queried offline on textual descriptions of patient observations to obtain judgments, e.g., the suspected clinical category. To improve the variational autoencoder fit, we train a differentiable surrogate model on these judgments and augment the loss function to encourage reconstructions that preserve the clinical-label distribution of their corresponding input profile. In an application to longitudinal motor-function assessments from children with spinal muscular atrophy, we map visit-level clinical profiles to low-dimensional representations that are linked by a multivariate mixed-effects model. The synthetic expert loss discourages reconstructions that remain numerically close in data space but alter the clinical interpretation of the reconstructed motor function profile, such as by crossing a disease-type boundary. We thus reduced disagreement between original and reconstructed SMA type labels from about 11 to 7 percent. Furthermore, informing the latent representation by the synthetic expert improved prediction of motor function milestones compared with unsupervised latent representations and a data-level baseline. These results suggest that incorporating LLMs into model fitting can make clinical knowledge available to representation learning and improve clinical faithfulness for longitudinal rare-disease data.","authors":["Clemens Sch\\\"achter","Astrid Pechmann","Janbernd Kirschner","Jan Hasenauer","Harald Binder"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-18","first_seen":"2026-08-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.16507","pdf_url":"https://arxiv.org/pdf/2608.16507","source_feed":"cs.AI","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM标注","临床知识","表示学习"],"reason":"用LLM替代临床专家标注，非仿真人类被试，但涉及专家知识提取，边界相关。","model":"deepseek-v4-pro","scored_at":"2026-08-18T13:04:24","error":null,"has_summary":false,"summary":null},{"id":"2608.14737","version":1,"title":"Class Imbalance and Batch Effects in LLM-Based Screening for Systematic Reviews","zh_title":"基于LLM的系统综述筛选中的类别不平衡与批次效应","abstract":"This study analyses LLMs in imbalanced binary classification, using study screening in systematic reviews as the application domain. An experiment was conducted in five reviews, comparing individual and batch processing, with and without prevalence metadata. The results indicate a limited influence of the prevalence metadata, with no evidence that it improves performance. In contrast, batch processing produced larger behavioral changes that varied according to the prevalence of the class. The aggregate and item-level analyses did not always coincide. Therefore, batch processing should be evaluated not only in terms of cost, but also in relation to its effects on decision-making behavior.","authors":["Gilberto Sussumu Hida","Danilo Monteiro Ribeiro","Clayton Suguio Hida"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"cross","date":"2026-08-18","first_seen":"2026-08-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.14737","pdf_url":"https://arxiv.org/pdf/2608.14737","source_feed":"cs.AI","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM标注","系统综述","批次效应"],"reason":"LLM用于系统综述筛选，替代人工标注，非仿真人类被试，但涉及决策行为变化，边界…","model":"deepseek-v4-pro","scored_at":"2026-08-18T13:04:18","error":null,"has_summary":false,"summary":null},{"id":"2608.14948","version":1,"title":"Who's Keeping Score? Interactive Steering of LLM-Powered Scoring with Attune","zh_title":"谁在评分？使用Attune交互式引导LLM评分","abstract":"Large language models (LLMs) are increasingly used to score text records at scale (e.g., rating candidate resumes on a 1-5 scale). However, existing LLM-powered approaches do not account for the fact that effective scoring requires both holistic understanding of records and locally consistent judgments across similar ones. We present Attune, a mixed-initiative system for steerable LLM-powered scoring. Given a task description and scoring range, Attune performs pairwise comparisons across records to develop a global understanding first, and then resolves these comparisons into consistent score assignments-deriving scoring criteria and rules bottom-up in the process. These serve as shared representations of scoring logic that users can inspect and edit. Based on insights from a formative study (n = 12), Attune's interface introduces novel steering interactions that allow users to deterministically refine scoring logic. Users can provide examples, directly edit criteria, rules, or target distributions, and give natural language feedback-with all refinements compiling into constraints that guide re-scoring. We validate our approach through a technical evaluation across three workloads and a user study with domain experts (n = 8) in healthcare, law, education, and AI evaluation.","authors":["Bhavya Chopra","Meng Chen","Rebecca Dang","Chanbin Park","Shreya Shankar","Sepanta Zeighami","Bjoern Hartmann","Aditya Parameswaran"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-08-18","first_seen":"2026-08-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.14948","pdf_url":"https://arxiv.org/pdf/2608.14948","source_feed":"cs.HC","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM评分","人机交互","标注替代"],"reason":"LLM用于评分替代人工，属标注替代而非仿真人类被试，但涉及人类判断校准，边界相…","model":"deepseek-v4-pro","scored_at":"2026-08-18T13:04:35","error":null,"has_summary":false,"summary":null},{"id":"2608.16467","version":1,"title":"Computational KJ-Ho: An Analyst-Bias-Free Insight Extraction Framework from Large-Scale Qualitative Data Using Domain-Specialized LLMs","zh_title":"计算KJ法：使用领域专用LLM从大规模定性数据中提取无分析者偏见的洞察框架","abstract":"The qualitative research methodologies that underpin consumer-insight generation - the KJ method, Grounded Theory, and Thematic Analysis - share a structural constraint: the cognitive processing capacity of the human analyst. Replication research further shows that conclusions vary substantially across analysts analyzing identical data (analyst bias). This paper proposes Computational KJ-Ho (the Kawakita Jiro method), a theoretical framework that computationally realizes the KJ method's epistemology - letting structure emerge from the data itself without imposing the analyst's preconceptions - an orientation we term \"analyst-bias-free.\" The framework employs a domain-specialized LLM built through continued pre-training (CPT) on a marketing-research corpus and supervised fine-tuning (SFT) on expert-curated insight pairs, organized as a three-layer architecture: data structuring, insight extraction, and strategy generation. Two preliminary studies in the Japanese marketing context support the necessity of CPT-based domain specialization. The paper makes five contributions: (1) a theoretical integration of the KJ method, Grounded Theory, and Peircean abduction into a single epistemological commitment of data-driven explanation generation; (2) a three-layer architecture leveraging domain-specialized embeddings for cross-interview analysis; (3) two novel evaluation metrics, InsightExtraction-F1 and MarketingQA; (4) explicit engagement with the WEIRD problem, centering a non-Western methodology; and (5) five practice-derived problem formulations from nearly three decades of marketing-research practice, translated into design requirements. The human analyst retains a supervisory role. This is a concept paper presented ahead of empirical validation.","authors":["Kasumi Ban"],"categories":["cs.HC","cs.CL","cs.CY"],"primary_category":"cs.HC","announce_type":"new","date":"2026-08-18","first_seen":"2026-08-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.16467","pdf_url":"https://arxiv.org/pdf/2608.16467","source_feed":"cs.HC","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM辅助定性分析","领域专用模型","方法框架"],"reason":"用LLM替代人工分析定性数据，属标注替代而非仿真人类被试，但方法可迁移。","model":"deepseek-v4-pro","scored_at":"2026-08-18T13:04:51","error":null,"has_summary":false,"summary":null},{"id":"2608.16601","version":1,"title":"\"If It Looks Like a User\": Measuring Real-Time Moderation Effects via Social Media Simulation","zh_title":"“如果它看起来像用户”：通过社交媒体仿真测量实时审核效果","abstract":"Agent-based social media simulators offer a controlled environment to study content moderation, yet their value hinges on how faithfully they reproduce real platform dynamics. We develop a calibrated extension of SimSoM, an agent-based model of information diffusion on social networks, grounded in a real-world dataset of online vaccine discourse during the COVID-19 pandemic. Our approach replaces ad-hoc parametrisations with empirically fitted distributions, optimised via CMA-ES (Covariance Matrix Adaptation Evolution Strategy) and validated against real data across temporal, distributional, and structural dimensions. Using this validated simulator, we provide three key contributions. First, we show that the calibrated model reproduces key statistical signatures of the empirical data, including activity distributions, post/reshare ratios, and temporal patterns. Second, we apply established misinformation-spreader detection and prevention methods to both empirical and simulated data, progressively removing top-ranked users and showing that the resulting decline in low-quality content is consistent across the two. Third, comparing static (retroactive) and dynamic (in-simulation) moderation across 30 network realisations, we show that static evaluation significantly overestimates the effectiveness of user bans for the most effective detectors: when moderation is applied in real time, compensatory resharing by the remaining users dampens the expected reduction in low-quality content, so static estimates should be read as an upper bound. These findings highlight the necessity of simulation-based evaluation for content moderation policies and contribute a reusable, empirically grounded simulation framework.","authors":["Enrico Verdolotti","Gianluca Nogara","Luca Luceri","Silvia Giordano"],"categories":["cs.SI","cs.CY","cs.MA"],"primary_category":"cs.SI","announce_type":"cross","date":"2026-08-18","first_seen":"2026-08-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.16601","pdf_url":"https://arxiv.org/pdf/2608.16601","source_feed":"cs.CY","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["社交媒体仿真","内容审核","智能体建模"],"reason":"基于智能体的社交媒体仿真，有真实数据校准，但未使用LLM作为人类被试，而是模拟…","model":"deepseek-v4-pro","scored_at":"2026-08-18T13:04:52","error":null,"has_summary":false,"summary":null},{"id":"2608.14667","version":1,"title":"Position: AI Agents in Scientific Teams Should Be Studied as Human-Agent Systems","zh_title":"立场：科学团队中的AI智能体应作为人机系统来研究","abstract":"Large language model-based agents are increasingly deployed as collaborators in scientific discovery yet most current work focuses on the autonomous capabilities of \"AI Scientists\". We argue that this overlooks the social aspects of scientific teamwork, and that studying AI Scientists as human-agent systems (HAS)--where the unit of analysis is the human-agent pair--is both underexplored and undervalued. We establish these points through literature and empirical analysis, and highlight recent incidences and studies which show that deploying agents in science without accounting for human-agent dynamics introduces near-term risks, including reduced diversity of scientific inquiry. Through analysis of real-world case studies, we show that scientists and agents can augment each other's capabilities. We call for new research that adopts the HAS lens to develop mathematical frameworks for understanding and fostering human-AI synergy in scientific discovery.","authors":["Patrick Emami","Sameera Horawalavithana","Truc Nguyen","Gihan Panapitiya","Bruno Jacob","Siddhisanket Raskar","Saumya Sinha","Jared D. Willard","Andrew Glaws","Nithin Somasekharan","Ling Yue","Brian Lu","Shaowu Pan","Jason Eisner"],"categories":["cs.AI","cs.HC"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-18","first_seen":"2026-08-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.14667","pdf_url":"https://arxiv.org/pdf/2608.14667","source_feed":"cs.AI","score":3,"bucket":"other","rubric_hits":["C1"],"tags":["人机协作","科学发现","多智能体系统"],"reason":"研究AI科学家与人类协作，属多智能体系统，不涉及用LLM仿真人类被试或与人类行…","model":"deepseek-v4-pro","scored_at":"2026-08-18T13:04:33","error":null,"has_summary":false,"summary":null},{"id":"2608.16168","version":1,"title":"QUMem: Personalized Memory for Query-Conditioned User-State Inference in LLM Agents","zh_title":"QUMem：面向LLM智能体中查询条件用户状态推断的个性化记忆","abstract":"Large language model (LLM) agents increasingly use external memory systems to support personalization by drawing on long and evolving interaction histories, in which user preferences may be distributed across time, change with context, and conflict with earlier evidence. However, existing systems face three limitations: fixed-turn, fixed-token, or session-based boundaries can mix unrelated dialogue or split an event from its causes, decisions, and outcomes; storing multiple pieces of user information from the same interaction as a single memory binds together items that serve different functions and should be independently retrievable; and treating the current task as a single top-$k$ retrieval query can return fragments that are individually relevant but fail to jointly capture preference evolution, temporal validity, and contextual applicability. We introduce \\textsc{QUMem}, a structured memory framework for query-conditioned user-state inference. \\textsc{QUMem} first segments interaction histories into variable-length episodes according to semantic continuity, then decomposes each episode into independently retrievable factual, preference, and transferable insight memories while preserving temporal positions and source evidence. At inference time, three sequential agents identify task-specific information needs, plan multi-query retrieval over the typed memory stores, and jointly infer a temporally and contextually valid user state for downstream response generation. \\textsc{QUMem} achieves state-of-the-art performance on both PersonaMem and KnowU-Bench, demonstrating the effectiveness of query-conditioned user-state inference for long-term personalization.","authors":["Heng Wang","Yifei Li","Lingling Zhang","Pengyu Li","Xinyu Che","Xinyu Zhang","Zesheng Yang"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"cross","date":"2026-08-18","first_seen":"2026-08-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.16168","pdf_url":"https://arxiv.org/pdf/2608.16168","source_feed":"cs.AI","score":3,"bucket":"other","rubric_hits":["C3"],"tags":["个性化记忆","用户状态推断","对话系统"],"reason":"论文聚焦个性化记忆与用户状态推断，属于角色扮演对话系统，无实验或测量目的，不涉…","model":"deepseek-v4-pro","scored_at":"2026-08-18T13:04:47","error":null,"has_summary":false,"summary":null},{"id":"2608.03361","version":2,"title":"The Evolutionary Origin of Values: implications for AI alignment, sentience and existential risk","zh_title":"价值观的进化起源：对AI对齐、感知与存在风险的影响","abstract":"AI systems based on Large Language Models (LLMs) have prompted fears that they may harbor hidden goals, seek to dominate or eliminate humanity, or even suffer as sentient beings. We address these concerns by tracing the evolutionary origin of value in biological organisms. Values emerge from autopoiesis: living systems must actively maintain themselves against perturbation and dissipation. Natural selection has equipped them with hierarchies of \"vicarious selectors\" that guide their behavior toward fitness. LLMs, by contrast, are allopoietic and allotelic: they produce outputs for others, and their goals derive from user prompts rather than an autonomous drive. They lack the intrinsic motivation for self-preservation, dominance, or resource competition that underlies existential-risk scenarios, and the embodied vulnerability required for feeling or suffering. Still, because LLMs learn statistical patterns from human-generated text, they implicitly absorb human values as well as knowledge, allowing them to focus on what is relevant. That is why the \"orthogonality thesis\" separating intelligence from values does not apply to them. Such separation would in fact expose any intelligence to the frame problem: the combinatorial explosion of the search space that makes any realistic utility function physically uncomputable. That also precludes the convergence of instrumental values thesis. We conclude that the real alignment challenge lies not in preventing rogue AI agency, but in ensuring LLMs intelligently apply learned ethical values.","authors":["Francis Heylighen"],"categories":["cs.CY","cs.AI"],"primary_category":"cs.CY","announce_type":"replace-cross","date":"2026-08-18","first_seen":"2026-08-05","revised_at":"2026-08-18","abs_url":"https://arxiv.org/abs/2608.03361","pdf_url":"https://arxiv.org/pdf/2608.03361","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["AI对齐","价值观起源","存在风险"],"reason":"论文讨论AI对齐与价值观起源，未将LLM作为人类被试进行仿真实验，无人类数据对…","model":"deepseek-v4-pro","scored_at":"2026-08-18T13:04:58","error":null,"has_summary":false,"summary":null},{"id":"2608.06510","version":2,"title":"Agentic AI: User Empowerment or Foreclosure?","zh_title":"代理式AI：用户赋权还是封闭？","abstract":"Agentic AI promises systems that can act on users' behalf, from filtering content to negotiating prices to selecting services. Whether it will empower users is an open question, and one that depends on more than the technology. We conduct a comparative case analysis of four earlier, more mature domains in which similar forms of agency emerged: browser-based ad blockers, platform recommender systems, financial robo-advisors, and email spam filtering. Across the cases, questions about whose interests agents would serve were resolved through technical arrangements: API choices, protocol governance, industry standards, and default configurations. Beyond their technical form, these were political decisions. We identify this settling of contestable questions in a technical form as depoliticization, a concept from political theory, here at work in technological systems. Its most consequential effect is that individual outcomes and collective contestation capacity can move in opposite directions: spam inbox quality improved substantially while the organized capacity to contest spam governance collapsed. Where intermediary institutions sustained formal channels for challenge, user-aligned agency proved more durable; where proprietary infrastructure and closed standard-setting absorbed contestation, the material basis for user-aligned alternatives was dismantled, and the loss proved hard to reverse. Applying this lens to agentic AI, we find a similar pattern forming: governance is consolidating around the Model Context Protocol and the Agentic AI Foundation, an industry-governed venue already deciding what agents will be able to do. Unlike in the completed trajectories, these decisions have not yet hardened, and remain open to challenge by users and the public.","authors":["David Gamba","Daniel M. Romero","Grant Schoenebeck"],"categories":["cs.CY","cs.AI","cs.HC"],"primary_category":"cs.CY","announce_type":"replace-cross","date":"2026-08-18","first_seen":"2026-08-10","revised_at":"2026-08-18","abs_url":"https://arxiv.org/abs/2608.06510","pdf_url":"https://arxiv.org/pdf/2608.06510","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["AI治理","技术政治","用户代理"],"reason":"讨论agentic AI治理与历史案例，非LLM仿真人类被试，无人类数据对照。","model":"deepseek-v4-pro","scored_at":"2026-08-18T13:04:25","error":null,"has_summary":false,"summary":null},{"id":"2608.10030","version":2,"title":"Automating and Scaling Behavioral Scientific Research on AI Agents","zh_title":"自动化与规模化AI智能体行为科学研究","abstract":"As AI agents are increasingly deployed in complex environments, understanding their behaviors becomes critical. Yet behavioral scientific research on AI agents remains manual and labor-intensive. We introduce AEROBAT, the first multi-agent system to automate behavioral scientific research on AI agents. Given an arbitrary target behavior by its user, AEROBAT automatically executes a full pipeline of behavioral scientific research---generating hypotheses about the behavior, designing and executing controlled experiments, making behavioral assessments, analyzing the results, and writing reports. For 12 target behaviors, we used AEROBAT to generate and test 73 hypotheses: designing 1,160 controlled experiments and executing 22,954 simulation rounds in total. Moderate-to-strong statistical evidence was found for 30 hypotheses, including some novel ones. In sum, our results demonstrate that automated behavioral scientific research on AI agents can complement and extend the reach of manual research.","authors":["Soo Yong Lee","Jongha Lee","Jaewan Chun","Hyunjin Hwang","Fanchen Bu","Ziv Ben-Zion","Taekwan Kim","Denny Borsboom","Jaemin Yoo","Kijung Shin"],"categories":["cs.AI","cs.MA"],"primary_category":"cs.AI","announce_type":"replace","date":"2026-08-18","first_seen":"2026-08-12","revised_at":"2026-08-18","abs_url":"https://arxiv.org/abs/2608.10030","pdf_url":"https://arxiv.org/pdf/2608.10030","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","行为科学自动化","AI智能体"],"reason":"研究AI agent行为自动化，非用LLM仿真人类被试，无人类数据对照","model":"deepseek-v4-pro","scored_at":"2026-08-18T13:04:57","error":null,"has_summary":false,"summary":null},{"id":"2608.14651","version":1,"title":"Evaluating Multimodal LLMs across Text and Audio Modalities for Accessible Disaster Assistance","zh_title":"评估多模态大语言模型在文本与音频模态下用于无障碍灾害援助的表现","abstract":"Effective disaster risk communication is a foundational humanitarian challenge, yet current emergency infrastructure fails to meet the needs of individuals with access and functional needs, including hard-of-hearing individuals, pregnant women, mothers with toddlers, and elderly individuals with dementia. Recent advancements in Artificial Intelligence (AI), especially Multi-Modal Large Language Models (MM-LLMs), demonstrate powerful capabilities to serve diverse users across text, audio, image, and video modalities within a single unified system, such as a chatbot. However, their suitability for deployment rests on a property that receives limited scrutiny, i.e., whether these systems produce consistent, actionable outputs regardless of the modality through which a user communicates. In this paper, we conduct a comprehensive analysis to understand the status of open-weight MM-LLMs using real emergency alert scenarios across four different vulnerable personas. These state-of-the-art (SOTA) models are evaluated on consistency of responses across text and audio modalities when the same task scenario is given. Findings indicate that no model achieves reliable consistency across modalities, and that performance gaps are heightened for personas with access needs, introducing modality-dependent inequity that undermines the humanitarian value of these systems. These results inform concrete design recommendations for building equitable, trustworthy, and inclusive AI tools for disaster risk communication.","authors":["Anuridhi Gupta","Samara Mansoor","Hemant Purohit"],"categories":["cs.AI","cs.HC"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-18","first_seen":"2026-08-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.14651","pdf_url":"https://arxiv.org/pdf/2608.14651","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["多模态LLM","灾害沟通","一致性评估"],"reason":"评估多模态LLM在灾害沟通中的一致性，属于聊天机器人应用，无人类行为仿真或对照。","model":"deepseek-v4-pro","scored_at":"2026-08-18T13:04:32","error":null,"has_summary":false,"summary":null},{"id":"2608.15109","version":1,"title":"Constraint-Aware Synthetic Tabular Data Generation via Inter-Column Constraint Discovery with LLM Agents","zh_title":"基于LLM智能体的列间约束发现与约束感知合成表格数据生成","abstract":"Generating structurally valid synthetic tabular data remains difficult: outputs with high statistical fidelity and downstream utility can still violate semantically meaningful domain constraints. We study the discovery and enforcement of three complementary inter-column constraint families---equations, linear inequalities, and logical dependencies. Our unified tool-grounded workflow represents all three as machine-executable hypotheses and applies a common interface for full-table validation, deterministic diagnosis, and counterexample-guided revision. A generator-agnostic postprocessor coordinates family-specific repairs on outputs from unchanged tabular generators. Across curated behavioral audits and end-to-end evaluations, the complete workflow improves held-out violation detection over one-shot direct prompting, while postprocessing yields zero measured violations for every retained, applicable constraint, improves downstream utility on most datasets, and largely preserves univariate marginals.","authors":["Jianxing Zhao","Mao Guan","Dongyu Liu"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-18","first_seen":"2026-08-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.15109","pdf_url":"https://arxiv.org/pdf/2608.15109","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["合成数据生成","LLM智能体","数据约束"],"reason":"LLM agent用于发现表格数据约束，属多智能体协作解题，不涉及人类行为仿真…","model":"deepseek-v4-pro","scored_at":"2026-08-18T13:04:37","error":null,"has_summary":false,"summary":null},{"id":"2608.15304","version":1,"title":"Understanding Cognition-Induced Risks in Agentic AI Systems","zh_title":"理解智能体AI系统中认知引发的风险","abstract":"Frontier agentic systems powered by large language models (LLMs) exhibit human-like patterns of cognition. As these systems become deeply integrated across different domains, their cognitive engagement raises critical concerns for human society that remain insufficiently studied. To address this gap, we systematically analyze risks induced by expanding cognitive capabilities, following a three-level framework defined by their cognitive scope, from physical cognition to social cognition, and finally to self-referential cognition. We study their potential risks to human agency, autonomy, and control capability, corresponding to each cognitive level. We finally propose strategies to mitigate these risks and enhance the controllability of agentic AI systems, ensuring their long-term safe development.","authors":["Guanchu Wang","Qinuo Li","Mengnan Du","Xia Hu","Bowen Zhou"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-18","first_seen":"2026-08-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.15304","pdf_url":"https://arxiv.org/pdf/2608.15304","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["智能体风险","认知科学","AI安全"],"reason":"论文分析智能体认知风险，不涉及用LLM仿真人类被试或与人类数据对照，属多智能体…","model":"deepseek-v4-pro","scored_at":"2026-08-18T13:04:39","error":null,"has_summary":false,"summary":null},{"id":"2608.16003","version":1,"title":"Prior Audit-Repair Context Shifts LLM Verifier Thresholds Toward Leniency","zh_title":"先前的审计-修复上下文使LLM验证器阈值偏向宽松","abstract":"Automated checking pipelines increasingly place one language model as the checker and another (or the same one) as the fixer. We ask whether that wiring changes what the checker reports. Measuring false alarms on human-verified-correct ProcessBench traces with the present task held byte-identical, we find that a completed audit -> repair episode already in the model's context lowers false alarms in 15 of 15 model x wording combinations, by 2.8 to 11.5 percentage points against a length-matched non-audit control, a 9 to 25% reduction relative to that control. The direction contradicts what the accumulated-message literature predicts: an episode whose audit reported an error lowers false alarms further still, at all five wordings on the model where that manipulation lands cleanly, though a negativity asymmetry predicts more flagging. Decomposing the episode finds repair content and audit verdict complementary: different components carry the effect on different model families. Signal-detection analysis locates the change in the threshold rather than in discrimination -- the criterion moves in 15 of 15 combinations and survives correction in 13 while d' survives in none, though the d' test is half as sensitive by construction -- and a hand audit of 50 false alarms finds 82% simply wrong, so at this operating point the shift need not be harmful. With reasoning enabled the effect keeps its relative size on both models tested, and the threshold reading holds there too.","authors":["Parsa Mazaheri","Kasra Mazaheri"],"categories":["cs.AI","cs.CL"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-18","first_seen":"2026-08-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.16003","pdf_url":"https://arxiv.org/pdf/2608.16003","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["LLM验证器","多智能体协作","信号检测"],"reason":"研究LLM验证器在审计-修复上下文中的行为变化，属于多智能体协作，不涉及人类行…","model":"deepseek-v4-pro","scored_at":"2026-08-18T13:04:45","error":null,"has_summary":false,"summary":null},{"id":"2608.16118","version":1,"title":"Assessing LLMs' mathematical abilities requires understanding the various mechanisms of mathematical creativity","zh_title":"评估大语言模型的数学能力需要理解数学创造力的多种机制","abstract":"How should we assess whether large language models can perform mathematical invention? I argue that this question is currently underspecified: mathematical creativity is not one capacity but several mechanistically distinct modes of meaning-making - reflexive introspection on mathematical practice, analogical import from the sciences, problem-driven construction, and the bridging of distant domains - together with a further, cross-cutting distinction between meaning pursued because a pattern was observed and meaning pursued because it is strategically wanted, a distinction I develop through the case of conjecture-formation. These mechanisms are likely non-substitutable, so that competence in one does not transfer to the others. Grounding each in a historical case study and in an architecture-level account of current transformer-based systems, I suggest that today's models concentrate their competence in modes shaped by recombination and search over existing building blocks; if that description holds, the remaining modes are out of reach in principle, not just slower - though whether it holds is itself the open, empirical part. Because proof is getting cheaper as AI improves at generating it - a shift the field's own leading voices are now diagnosing - mathematical value is migrating toward the modes current systems cannot yet perform, and evaluations of AI mathematical ability should be organized around this taxonomy rather than around aggregate benchmarks that conflate it.","authors":["Silv\\`ere Gangloff"],"categories":["cs.AI","math.HO"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-18","first_seen":"2026-08-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.16118","pdf_url":"https://arxiv.org/pdf/2608.16118","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["LLM数学能力","能力评测","数学创造力"],"reason":"论文讨论LLM数学能力评估，属纯能力评测，不以人类行为仿真为目标，无人类数据对…","model":"deepseek-v4-pro","scored_at":"2026-08-18T13:04:47","error":null,"has_summary":false,"summary":null},{"id":"2608.16349","version":1,"title":"AeroCopilotBench: A Two-Tier Benchmark for Evaluating LLM Agents as Aviation Copilots in an Interactive Virtual Cockpit Environment","zh_title":"AeroCopilotBench：在交互式虚拟驾驶舱环境中评估LLM智能体作为航空副驾驶的双层基准","abstract":"Large language model (LLM) agents may assist flight crews with complex decisions and task execution, but existing aviation evaluations centered on static knowledge do not support systematic testing of procedural execution and safety compliance in interactive environments. This paper presents the AeroCopilot Operational Environment (ACOE), a reproducible interactive virtual-cockpit test environment, and AeroCopilotBench, a two-tier aviation agent evaluation benchmark. Tier-1 evaluates aviation knowledge using 1,200 multiple-choice questions, while Tier-2 comprises 73 emergency and abnormal tasks derived from the manufacturers' Pilot's Operating Handbooks (POHs) and instantiated in ACOE. ACOE converts natural-language procedures into executable state transitions, final-state goal conditions, and hard safety constraints, enabling models to interpret cockpit state, diagnose faults, and operate aircraft systems through standardized tool interfaces. We establish a safety-gated evaluation framework in which a trajectory succeeds only when all task goals are achieved without violating any hard safety constraint, while safe goal progress and trajectory safety are measured separately. Across 12 models, the highest Tier-2 success rate is 72.6%, while static knowledge performance does not consistently translate into procedural execution. Analysis of 451 failed episodes from 3 representative models identifies recurring failures in procedural completeness, use of state feedback, and long-horizon execution management. These findings motivate state-aware agent orchestration, joint assessment of task completion and trajectory safety, and repeated regression testing. ACOE and AeroCopilotBench provide a reproducible foundation for testing knowledge application, interactive execution, and operational safety in aviation agents.","authors":["Yuchen Yuan","Zhenghuang Wu","Yuangan Li","Liang Ma","Ke Li"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-18","first_seen":"2026-08-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.16349","pdf_url":"https://arxiv.org/pdf/2608.16349","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C2"],"tags":["LLM智能体","航空仿真","基准测试"],"reason":"评估LLM在虚拟驾驶舱中执行程序任务，属于自动驾驶/游戏仿真环境，不涉及人类行…","model":"deepseek-v4-pro","scored_at":"2026-08-18T13:04:49","error":null,"has_summary":false,"summary":null},{"id":"2608.14692","version":1,"title":"Identifying Harm in Personalized, Generative AI Systems Requires User-Centered Auditing at the Interaction Level","zh_title":"识别个性化生成式AI系统中的伤害需要以用户为中心的交互层面审计","abstract":"Personalized, generative AI systems increasingly adapt their behavior to individual users over time, fundamentally changing model behavior. While existing auditing approaches have been effective at surfacing harms in non-personalized contexts, they often rely on static, simulated evaluations and definitions of harm that aggregate across broad, group categories. In this position paper, we argue that such approaches can fail to capture emergent harms in personalized generative AI systems, where harms surface through interpretations of ongoing interaction and evolve with user history. We identify three presuppositions underlying many harm auditing paradigms: that harms can be (1) specified outside real-world interaction, (2) defined non-pluralistically within groups, and (3) treated as static. One might argue that personalized systems could simply learn definitions of what constitutes harm to individual users through repeated interactions. However, we argue that attempts to surface user harms through deeper personalization risk imposing asymmetric burdens of labor and privacy on marginalized users. Consequently, we propose reframing understandings of harm as adaptive, user- and community-centered processes, and outline design directions that shift auditing from retrospective evaluation toward infrastructures that support ongoing articulation of harm in interaction. Our work highlights the need for auditing and design practices that better reflect the pluralistic and evolving nature of harm understanding in personalized generative AI systems.","authors":["Hannah Cha"],"categories":["cs.CY","cs.AI","cs.HC"],"primary_category":"cs.CY","announce_type":"cross","date":"2026-08-18","first_seen":"2026-08-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.14692","pdf_url":"https://arxiv.org/pdf/2608.14692","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["AI审计","个性化系统","用户伤害"],"reason":"讨论个性化生成式AI的伤害审计，不涉及用LLM仿真人类被试或与人类数据对照。","model":"deepseek-v4-pro","scored_at":"2026-08-18T13:04:33","error":null,"has_summary":false,"summary":null},{"id":"2608.15338","version":1,"title":"When AI Rewrites, Classifiers Relax: Uncertainty-Aware Sentiment Analysis on Sarcastic and AI-Paraphrased Social Text","zh_title":"当AI改写时，分类器放松：讽刺与AI改写社交文本上的不确定性感知情感分析","abstract":"Sentiment classifiers are increasingly applied to social media content that is either sarcastic or AI-generated --- two distributional regimes where standard evaluations offer little guidance. We present a three-part empirical study of sentiment classifier behaviour under these conditions. First, we find that confidence scores on sarcastic text are significantly lower than on non-sarcastic text (Mann--Whitney $p = 2 \\times 10^{-6}$), confirming that classifiers sense their own uncertainty on ironic content even without explicit uncertainty modelling. Second, and counterintuitively, we show that sentiment classifiers achieve higher accuracy on AI-paraphrased reviews than on the original human-authored text (RoBERTa: $+5.8$ pp for Qwen3.5-4B paraphrases, $+3.7$ pp for Gemma4-E4B), revealing a cross-domain stylistic alignment effect: AI paraphrases remove distributional noise that confounds Twitter-trained classifiers, producing cleaner, more prototypical sentiment text. Third, we demonstrate that a lightweight abstention wrapper --- flagging the $14\\%$ of inputs with confidence below $0.6$ --- improves accuracy from 82.2\\% to 88.9\\% ($+6.7$ pp) on the retained set. We further compare Semantic Entropy and MC-Dropout-style disagreement as uncertainty signals and find near-identical AUROC ($0.650$ vs.\\ $0.646$) on sarcastic text, suggesting that for short social media inputs, both methods are interchangeable. Our results motivate a shift from confident single-label prediction to uncertainty-aware abstention in high-stakes sentiment applications such as mental health flagging and content moderation.","authors":["Shresth Shroff"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"cross","date":"2026-08-18","first_seen":"2026-08-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.15338","pdf_url":"https://arxiv.org/pdf/2608.15338","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["情感分析","不确定性","AI改写"],"reason":"研究情感分类器在讽刺和AI改写文本上的表现，属于NLP模型评测，不以人类行为仿…","model":"deepseek-v4-pro","scored_at":"2026-08-18T13:04:41","error":null,"has_summary":false,"summary":null},{"id":"2608.15654","version":1,"title":"When Stories Evolve: Benchmarking LLM Storytelling Across Agent Architectures in Open-Ended World Simulations","zh_title":"当故事演化：在开放式世界模拟中跨智能体架构基准测试LLM叙事","abstract":"Large language models can write fluent stories, but open-ended storytelling requires more than local fluency. In evolving world simulations and AI-native games, models must preserve facts, relationships, causal dependencies, and character states as the world changes. We introduce WSE-bench, a process benchmark that separately evaluates sustained generation, canonical coherence, and meaningful development in dynamic LLM storytelling. Generation Coverage records the proportion of planned narrative steps produced; Consistency tracks when canon breaks; and Richness measures how meaningfully branching, player-shaped trajectories develop. Across frontier models, Consistency and Richness do not form a smooth trade-off: their empirical Pareto frontier is non-concave, with several non-dominated intermediate configurations that no positive linear weighting can select. Added structure can enrich trajectories, but it does not uniformly improve coherence and may shorten them. Model scale chiefly improves sustained generation, without producing reliable gains in canonical coherence or meaningful development. These results show that sustained generation, canonical coherence, and meaningful development are distinct and sometimes competing capacities. WSE-bench makes those dynamics visible by extending narrative evaluation from finished stories to the processes that create them.","authors":["Yuqi Chen","Sixuan Li","Yunfeng Cai","Xueai Li","Ka Man Yan","Ying Li"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"cross","date":"2026-08-18","first_seen":"2026-08-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.15654","pdf_url":"https://arxiv.org/pdf/2608.15654","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C2"],"tags":["LLM叙事生成","游戏仿真","基准测试"],"reason":"研究LLM在开放世界游戏中的叙事生成，属游戏仿真环境，不涉及人类行为对照或仿真…","model":"deepseek-v4-pro","scored_at":"2026-08-18T13:04:43","error":null,"has_summary":false,"summary":null},{"id":"2608.15828","version":1,"title":"A Cognitively Motivated Multidimensional Framework for Evaluating Metaphor Explanations","zh_title":"一种认知驱动的多维框架用于评估隐喻解释","abstract":"Current evaluation of metaphor explanations relies mainly on holistic quality ratings, revealing little about how explanation quality is structured or where human judgments agree and diverge. We introduce a cognitively motivated framework that decomposes metaphor explanation quality into six theoretically grounded dimensions. In a dense annotation study (11,200 ratings), we find that: {\\bfseries(i)} explanation quality is genuinely multidimensional; {\\bfseries(ii)} annotator disagreement is systematic rather than random; and {\\bfseries(iii)} the six dimensions collapse into a shared cluster and two independent axes of judgment. An exploratory feasibility study further shows that a standard automatic evaluation pipeline can recover parts of this structure, predicting the most discriminative dimensions well while its errors correlate human (dis)agreement. Together, these results suggest that multidimensional evaluation offers richer diagnostic insight than holistic ratings, and that automatic evaluators for open-ended generation tasks should be judged on how well they preserve the structure of human judgment.","authors":["Ana Naveriani","Jakob Suchan","Stefano Zoia","Mehul Bhatt","Antonio Lieto","Gian Luca Pozzato"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"cross","date":"2026-08-18","first_seen":"2026-08-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.15828","pdf_url":"https://arxiv.org/pdf/2608.15828","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["隐喻解释","自动评估","NLP评测"],"reason":"论文评估隐喻解释质量，属于NLP评测，不以人类行为仿真为目标，无LLM作为被试…","model":"deepseek-v4-pro","scored_at":"2026-08-18T13:04:44","error":null,"has_summary":false,"summary":null},{"id":"2608.16045","version":1,"title":"Walk Before You Run: The Importance of Data Exploration for Data Analysis Agents","zh_title":"先走后跑：数据探索对数据分析代理的重要性","abstract":"LLM-based data-analysis tools are increasingly used to help users analyze messy spreadsheets and workbooks, from answering questions over uploaded files to generating code, summaries, and visualizations. These systems are often evaluated by the correctness of their final downstream answers. However, reliable data analysis also depends on an earlier step: understanding what the dataset contains before solving the requested task. For complex workbooks, this Data Exploration step includes identifying the logical tables behind physical sheets, interpreting column semantics, recovering keys and relationships, and detecting quality issues. In current tools and benchmarks, this step is usually left implicit, creating a gap between downstream task performance and the dataset understanding needed for reliable, human-checkable analysis. Our key contribution is to identify this overlooked gap, make Data Exploration a first-class evaluation target, and show through downstream experiments that stronger Data Exploration support improves task performance. To evaluate dataset understanding directly, we introduce two benchmark settings: a real multi-sheet workbook benchmark based on a Vitamin D study dataset, and an extension of DSBench with schema-fixed Data Exploration artifacts. In both settings, systems are evaluated by the quality of a structured artifact capturing tables, columns, semantic roles, relationships, and profiling signals. Our results show that strong LLMs and data-analysis agents still miss important logical structure even when they read spreadsheet content. Furthermore, explicit Data Exploration support often improves downstream correctness, suggesting it should be treated as a first-class, inspectable stage in LLM data-analysis workflows and a natural human-in-the-loop checkpoint where domain experts can review and correct the artifact before downstream analysis proceeds.","authors":["Yike Yuan","Virum Ranka","Tina Lasisi","Lin Ma"],"categories":["cs.DB","cs.AI"],"primary_category":"cs.DB","announce_type":"cross","date":"2026-08-18","first_seen":"2026-08-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.16045","pdf_url":"https://arxiv.org/pdf/2608.16045","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["LLM数据分析","数据探索","基准测试"],"reason":"研究LLM数据分析代理的数据探索能力，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-08-18T13:04:46","error":null,"has_summary":false,"summary":null},{"id":"2608.16461","version":1,"title":"A Human-LLM Teaming Framework for Privacy Risk Analysis: An Illustration with CBDC-Based Welfare Schemes","zh_title":"面向隐私风险分析的人机协作框架：以基于CBDC的福利计划为例","abstract":"Central Bank Digital Currency (CBDC)-based welfare schemes may be potentially privacy invasive as they process significant volumes of beneficiary personal data and lead to privacy harms such as surveillance, discrimination and stigmatization. Such welfare delivery schemes involve complex digital ecosystems and large number of stakeholders. Consequently, to examine their privacy risks, privacy risk assessments require extensive information gathering and synthesis, complex reasoning, scenario explorations, contextual evaluation and human judgement. Thus, they present ideal scenarios for human-LLM teaming, where effective integration of complementary human and LLM capabilities can yield an outcome far superior to either human-only or LLM-only assessments. In this paper, we propose a first human-LLM teaming framework for the systematic privacy risk analysis methodology called PRIAM. The framework specifies an iterative collaborative process in which the LLM processes large-scale documentary evidence to produce initial outputs, which are then interpreted and evaluated by human experts who direct their further refinement by the LLM and exercise their judgement to finalize the output. We illustrate the framework on the data characterization activity of PRIAM using a CBDC-based welfare scheme use case. The illustration demonstrates that while LLMs generate the initial data categories and assign initial values to data attributes, human experts evaluate and provide feedback to refine them, distinguishing documented evidence from inferences, identifying information gaps, and flagging unsupported or ambiguous outputs. This framework serves as a foundational contribution towards human-AI teaming for privacy risk assessments.","authors":["Sourya Joyee De","Abdessamad Imine"],"categories":["cs.ET","cs.AI","cs.CE","cs.CY"],"primary_category":"cs.ET","announce_type":"cross","date":"2026-08-18","first_seen":"2026-08-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.16461","pdf_url":"https://arxiv.org/pdf/2608.16461","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["人机协作","隐私风险","LLM辅助"],"reason":"人类与LLM协作进行隐私风险评估，非用LLM仿真人类被试，无人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-08-18T13:04:51","error":null,"has_summary":false,"summary":null},{"id":"2608.16627","version":1,"title":"When Do Explanations Help In-Context Learning? A Comparative Study of Natural Language Explanation Types and Faithfulness","zh_title":"解释何时有助于上下文学习？自然语言解释类型与忠实度的比较研究","abstract":"Natural language explanations (NLEs) are increasingly used as inputs, for example, as few-shot rationales that influence model behavior in in-context learning (ICL). However, it remains unclear how different types of NLEs compare in their effects on downstream model performance in explanation-augmented prompting. Therefore, we provide a comparative evaluation across six benchmarks and four instruction-tuned models, studying how NLE source (human-written when available, self-generated explanations, generated by an external LLM) and NLE selection (random vs faithfulness-based filtering) affect downstream utility of NLEs when used in ICL settings. Our extensive evaluation shows that, on classification-style benchmarks, adding NLEs to few-shot prompts often improves accuracy over few-shot prompting without explanations; among NLE sources, externally generated LLM-NLEs often provide strong downstream utility and remain competitive with human rationales where both are available, whereas self-NLEs are more sensitive to the selection strategy. On math reasoning, the effects are more model- and source-dependent. We further show that faithfulness-based selection of self-NLEs yields small average gains overall, but can improve or reduce performance depending on the metric, task, and model. Different faithfulness metrics can disagree substantially, affecting which self-NLE examples are selected and their downstream predictive utility. Robustness tests with randomly swapped and out-of-distribution rationales indicate partial robustness, suggesting that semantic alignment contributes to performance gains. Overall, our results provide insights for selecting and reporting explanations that influence model behavior in practical prompting pipelines.","authors":["Mahdi Dhaini","Adam Dejl","Juraj Vladika","Volkan \\\"Ozer","Barbara Plank","Gjergji Kasneci"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"cross","date":"2026-08-18","first_seen":"2026-08-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.16627","pdf_url":"https://arxiv.org/pdf/2608.16627","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["上下文学习","自然语言解释","模型性能评测"],"reason":"研究解释对上下文学习的影响，属NLP能力评测，不以人类行为为参照","model":"deepseek-v4-pro","scored_at":"2026-08-18T13:04:53","error":null,"has_summary":false,"summary":null},{"id":"2608.16643","version":1,"title":"Toward Better Assessment of LLMs' Performance in Clinical Error Detection","zh_title":"迈向更好的LLM临床错误检测性能评估","abstract":"Automated detection of errors in clinical documentation is a promising application of large language models (LLMs), yet decisions to deploy such models rest on benchmarks that evaluate each clinical note in isolation. Error-detection benchmarks are typically constructed by injecting errors into notes, such that each erroneous note has a natural counterpart. Aggregate discriminative metrics (e.g., balanced accuracy or F1) do not exploit this structure. We show that this omission is consequential. In particular, evaluating 15 diverse LLMs on 4 standardized clinical error-detection test sets across 3 languages, we find that 13 of 15 models fall below the level of random pairwise discrimination, even while achieving F1 scores that standard practice would read as moderate. We also observe that the underlying bias patterns differ across languages: the same model can default to \"no error\" on one language and over-flag errors on another. To diagnose where discrimination breaks down, we further introduce a procedure to score the evidence models cite in their outputs. We find that while models consistently locate error-relevant content, they fail to produce the corresponding correct verdict on the clean counterpart. Finally, we show that F1 and pairwise accuracy are driven in opposite directions by the same underlying bias, so that ranking models by F1 may systematically promote the weakest discriminators. For safety-critical clinical NLP applications, we advocate for supplementing aggregate metrics with paired evaluations in benchmark reporting. Code and analysis scripts are available at https://github.com/healthylaife/paired-clinical-eval.","authors":["Yifan Zhang","Rahmatollah Beheshti"],"categories":["cs.CL","cs.AI","cs.LG"],"primary_category":"cs.CL","announce_type":"cross","date":"2026-08-18","first_seen":"2026-08-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.16643","pdf_url":"https://arxiv.org/pdf/2608.16643","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["LLM评测","临床NLP","错误检测"],"reason":"评估LLM在临床错误检测中的性能，属于NLP能力评测，不以人类行为仿真为目标。","model":"deepseek-v4-pro","scored_at":"2026-08-18T13:04:55","error":null,"has_summary":false,"summary":null},{"id":"2608.16574","version":1,"title":"The User Side of AI Model Lifecycles: Evidence from the Keep4o Movement","zh_title":"AI模型生命周期的用户侧：来自Keep4o运动的证据","abstract":"AI model lifecycles are commonly understood as a series of technical and organizational processes. Yet once a model enters sustained use, subsequent changes can also affect established user practices and user value. Using the Keep4o movement around GPT-4o as a case, this study examines post-deployment AI model lifecycle issues from the user side. We collected 61,846 public original posts on X from August 2025 to March 2026 and, using a systematically developed coding framework and LLM-assisted content analysis, analyzed discussion themes, users' reasons for wanting to keep GPT-4o, and the specific claims they made. Findings show that the Keep4o discussion extended well beyond continued access to the model itself. It covered concrete experiences of use, model behavioral characteristics and how they changed, and management issues across different stages of the model lifecycle. Reasons for keeping GPT-4o reflected interactional and relational value formed through long-term use, as well as judgments about the adequacy of replacement and the reasonableness of related decisions. The corresponding claims further reflected users' specific expectations for model lifecycle arrangements and governance. Overall, the call to \"keep GPT-4o\" brought together different judgments about user value and governance concerns. These findings suggest that technical version succession does not necessarily amount to effective replacement on the user side. Post-deployment AI model lifecycle management therefore needs to consider whether established user value can be carried forward and how model changes affect actual use. This study thus provides user-side empirical evidence for AI model lifecycle management. It further shows that user experience can provide important information for identifying post-deployment impacts and should be incorporated into lifecycle evaluation and decision-making.","authors":["Yiwen Wu"],"categories":["cs.HC","cs.CY"],"primary_category":"cs.HC","announce_type":"new","date":"2026-08-18","first_seen":"2026-08-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.16574","pdf_url":"https://arxiv.org/pdf/2608.16574","source_feed":"cs.HC","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["用户研究","AI治理","内容分析"],"reason":"研究用户对GPT-4o的讨论，非用LLM仿真人类被试，无实验或测量目的","model":"deepseek-v4-pro","scored_at":"2026-08-18T13:04:51","error":null,"has_summary":false,"summary":null},{"id":"2608.16633","version":1,"title":"Love in the Age of AI: An Integrative Process Model of Romantic Human-Chatbot Relationships","zh_title":"AI时代的爱情：人机浪漫关系的整合过程模型","abstract":"The increasing ability of social chatbots to form deep and even romantic Human-Chatbot Re lationships (HCRs) has drawn growing academic attention. Yet, existing research remains fragmented, often examining individual stages such as initiation or dissolution in isolation, without tracing the full relational trajectory. Such fragmentation, however, hinders a holistic understanding of the interplay between the unique psychological and social drivers, relational dynamics, and profound emotional stakes, particularly obscuring the elements unique to ro mantic bonding. This paper addresses this gap by introducing the first empirically grounded integrative process model of the romantic HCR lifecycle. A qualitative secondary analysis of 73 user experiences, drawn from two datasets of qualitative interviews and surveys, provides the basis for a three-phase model that synthesizes established theoretical frameworks related to user needs and gratifications, HCR development, and relationship dissolution. The model demonstrates that the Initiation phase is driven by specific psychological and social determi nants that shape the needs and gratifications sought by the user. The Relationship Building phase progresses through explorative, affective and stable stages, in which users develop gen uine romantic feelings and a deeply integrated bond with the chatbot. Finally, the Ending phase reveals that when dissolution occurs, it elicits emotional and physical responses com parable to human breakups but generates unique, technology-mediated coping mechanisms, potentially leading to a recursive cycle of re-engagement.","authors":["Natalia Szymczyk","Paula Ebner","Jessica M. Szczuka"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-08-18","first_seen":"2026-08-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.16633","pdf_url":"https://arxiv.org/pdf/2608.16633","source_feed":"cs.HC","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["人机关系","聊天机器人","定性研究"],"reason":"研究人与聊天机器人的浪漫关系，属角色扮演聊天，无实验或测量目的，不涉及LLM仿…","model":"deepseek-v4-pro","scored_at":"2026-08-18T13:04:53","error":null,"has_summary":false,"summary":null},{"id":"2608.16030","version":1,"title":"Benchmarking Identity-Sensitive LLM Outputs for Surveillance and Security Robots","zh_title":"面向监控与安全机器人的身份敏感型LLM输出基准测试","abstract":"Large language models (LLMs) are increasingly used to generate textual robot design specifications, interaction policies, and risk assessments during early-stage robot development. Such outputs may influence how surveillance and security robots are conceptualized, documented, and ultimately implemented. This paper evaluates whether identity-conditioned prompts produce systematic differences in LLM-generated surveillance and security robot design descriptions. Using 236 demographic identity labels across single-label and model-augmented prompt conditions, we analyze readability as an initial benchmark for evaluating accessibility and identity-conditioned variation in generated robot design descriptions. The results show significant differences in readability across prompt conditions, design dimensions, and demographic identities. Although readability cannot determine whether an output is fair or socially appropriate, it provides an interpretable baseline within a broader benchmarking framework that also includes lexical, semantic, sentiment, syntactic, and fairness-focused analyses.","authors":["Nneka Hyman","Jasmine Khan","Raj Korpan"],"categories":["cs.RO","cs.CY"],"primary_category":"cs.RO","announce_type":"cross","date":"2026-08-18","first_seen":"2026-08-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.16030","pdf_url":"https://arxiv.org/pdf/2608.16030","source_feed":"cs.CY","score":2,"bucket":"other","rubric_hits":["C2"],"tags":["LLM输出基准","机器人设计","身份敏感"],"reason":"研究LLM生成机器人设计描述的差异，属机器人仿真环境，非人类行为仿真。","model":"deepseek-v4-pro","scored_at":"2026-08-18T13:04:46","error":null,"has_summary":false,"summary":null},{"id":"2608.14551","version":1,"title":"Auxiliary uncertainty signals for LLM-assisted systematic review screening: a benchmark across eight Cohen drug-class reviews","zh_title":"LLM辅助系统综述筛选的辅助不确定性信号：八个Cohen药物类别综述的基准测试","abstract":"Large language models (LLMs) are increasingly used for title-abstract screening in systematic reviews, but their decisions lack calibrated uncertainty. We show that an auxiliary BERT+GCN classifier supplies a structured uncertainty signal that improves LLM screening efficiency, and we identify the prompt-delivery strategy that maximises the benefit-to-cost ratio. We evaluate five LLM prompt-delivery conditions on eight drug-class datasets from the Cohen (2006) benchmark using 3 seeds x 5-fold stratified cross-validation (600 fold-level results). A BERT+GCN model trained per fold classifies each test paper as INCLUDE, EXCLUDE, or MAYBE via two spectral tests (algebraic radical and categorical paradox). Conditions vary information content (none / label / full scores), selectivity (all papers vs. MAYBE only), and timing (proactive vs. reactive two-pass). A cross-model pilot against gpt-4.1-mini on three datasets tests cross-generation transfer. Three findings: (i) Full-context delivery yields significant gains in F1 (+0.011, paired Wilcoxon p=0.008) and WSS@95 (+0.050, p=0.039) at a 1.28x token-cost premium, while preserving recall. (ii) MAYBE-only routing is Pareto-optimal: highest mean recall (0.92) and AUC-ROC (0.54) at only 1.05x baseline cost -- one sixth of full-context overhead. (iii) The two-pass design escalates 22.2% +/- 8.8% of records yet never revises its decision (0% flip rate across all datasets and folds), giving decisive evidence that current instruction-tuned LLMs cannot self-triage. The cross-model pilot shows an identical +0.8% recall uplift for both LLM generations. A per-paper ablation across 20,796 observations shows the dual paradox test reduces empirically to a one-line logit-gap criterion. We release the full pipeline; the 600-run experiment replays in under one hour from cached LLM responses.","authors":["Arya Rahgozar","Pouria Mortezaagha"],"categories":["cs.CL","cs.DL","cs.IR","cs.LG"],"primary_category":"cs.CL","announce_type":"cross","date":"2026-08-18","first_seen":"2026-08-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.14551","pdf_url":"https://arxiv.org/pdf/2608.14551","source_feed":"cs.LG","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["系统综述","LLM筛选","不确定性校准"],"reason":"LLM用于系统综述筛选，属NLP能力评测，无人类行为仿真或对照。","model":"deepseek-v4-pro","scored_at":"2026-08-18T13:04:16","error":null,"has_summary":false,"summary":null},{"id":"2608.15448","version":1,"title":"Language models suffer from a curse of ambiguity","zh_title":"语言模型遭受歧义诅咒","abstract":"Large language models increasingly rely on sampling as a driver of their own improvement, making the fidelity of their learned distributions more critical than ever. Yet, not all distributions are equally easy to learn. In this work, we identify a curse of ambiguity: in large language models, and more broadly in all neural networks that produce discrete probability distributions, the more ambiguous a next-token distribution is, the harder it is to learn accurately. Through an extensive theoretical analysis, we trace this curse to architectural and learning roots. More ambiguous distributions require more capacity to be stored, larger embeddings to be represented, more steps to be fitted, and amplify token-sampling noise. We validate these findings on synthetic tasks with controlled ground truth and observe the same signatures in language models trained on real data. Our results provide a new perspective on the statistical capabilities of large language models and a practical framework for when to trust their output distribution.","authors":["Nicolas Zucchet","Hyun Dong Lee","Scott Linderman"],"categories":["cs.CL","cs.LG","cs.NE"],"primary_category":"cs.CL","announce_type":"cross","date":"2026-08-18","first_seen":"2026-08-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.15448","pdf_url":"https://arxiv.org/pdf/2608.15448","source_feed":"cs.LG","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["语言模型","分布学习","理论分析"],"reason":"研究LLM学习离散分布的能力，属模型统计能力分析，不涉及人类行为仿真或对照。","model":"deepseek-v4-pro","scored_at":"2026-08-19T13:03:28","error":null,"has_summary":false,"summary":null},{"id":"2607.22067","version":2,"title":"Multimodal Language Models Benchmarked Against the NRC Reactor Operator Licensing Examination: Fine-Tuning and Retrieval Strategies","zh_title":"多模态语言模型在NRC反应堆操作员执照考试上的微调与检索策略基准测试","abstract":"Competence claims for a language model in a safety-critical domain are credible when measured against a standard the domain already enforces. We evaluate an open-weight 31-billion-parameter multimodal model (Gemma 4 31B-IT) on the U.S. Nuclear Regulatory Commission Reactor Operator Generic Fundamentals Examination (GFE), scoring it paper by paper against the 80% criterion applied to every human candidate, with no rounding up. The evaluation set is a census of every GFE administered at the March sitting from 2015 to 2021, giving seven pressurized water reactor (PWR) and seven boiling water reactor (BWR) papers and 697 scored items. Eight configurations cross three model states, the base model, supervised fine-tuning (SFT) on distilled chain-of-thought rationales and retrieval-augmented fine-tuning (RAFT), with three retrieval conditions, none and BM25 retrieval over the Department of Energy Fundamentals Handbooks under fixed-size and structure-aware chunking. Out of the box it answers 51.94% correctly and passes no paper. SFT with fixed-size chunking retrieval passes 8 of 14, reaching 80.23% on PWR items and 79.77% pooled, with a Wilson interval spanning the threshold. The preferred chunking granularity reverses with training state, structure-aware before fine-tuning and fixed-size after, so chunking optimized against a base model cannot be inherited by its fine-tuned descendant. RAFT trails SFT by 2.2 to 2.3 percentage points overall, and the deficit holds in all four reactor-type and chunking strata. The pipeline runs on one workstation with no network access at run time, and the result approaches operator-level command of engineering fundamentals without reliably achieving it.","authors":["Isak Hwang","Yoon Pyo Lee","Syed Bahauddin Alam"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"replace-cross","date":"2026-08-18","first_seen":"2026-07-27","revised_at":"2026-08-18","abs_url":"https://arxiv.org/abs/2607.22067","pdf_url":"https://arxiv.org/pdf/2607.22067","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["LLM评测","核工程","检索增强生成"],"reason":"评估LLM在核反应堆操作员考试上的表现，属于专业领域能力评测，不以人类行为仿真…","model":"deepseek-v4-pro","scored_at":"2026-07-29T09:01:42","error":null,"has_summary":false,"summary":null},{"id":"2606.15999","version":2,"title":"U.S. Technological Containment and the Rise of China's Open AI Ecosystem","zh_title":"美国政策无意中加速了中国开放AI生态系统的发展","abstract":"Over the past decade, U.S. policies have increasingly aimed to preserve artificial intelligence (AI) leadership by promoting domestic free-market policies while controlling global technological chokepoints, particularly advanced semiconductors and computational infrastructure. These measures raised the cost of Chinese AI development, but they also increased the strategic value of open and locally adaptable AI systems. Before raising export controls on high-performance chips, both the U.S. and China promoted policies that included support for open-source AI. During the period following major U.S. export-control shocks, China increasingly embedded open-source AI into national technology strategy through proposed ecosystem building, standards coordination, and resilience-oriented deployment. Moreover, Chinese developers increased engagement with open-source large language model repositories substantially more than U.S. developers did, consistent with a shift toward open infrastructure under geopolitical constraints. Subsequently, Chinese-origin open models diffused widely through open-source communities and scientific research. Even though such models remained largely absent from U.S. patent disclosures, American commercial entities use them in open-access research, suggesting their undermeasured importance within the foundation of U.S. commercial activity. These findings suggest that technological containment can shape not only the direction of AI development, but also the ecosystems through which AI is developed, improved, and diffused.","authors":["Wang Jin","Nadav Kunievsky","Bowen Lou","Tianshu Sun","James Evans"],"categories":["econ.GN","cs.CY","q-fin.EC"],"primary_category":"econ.GN","announce_type":"replace-cross","date":"2026-08-18","first_seen":"2026-06-14","revised_at":"2026-08-18","abs_url":"https://arxiv.org/abs/2606.15999","pdf_url":"https://arxiv.org/pdf/2606.15999","source_feed":"cs.CY","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["AI政策","开源生态","技术竞争"],"reason":"论文分析政策对开源生态的影响，不涉及用LLM仿真人类被试或行为对照。","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:58:15","error":null,"has_summary":false,"summary":null},{"id":"2512.21031","version":3,"title":"Learning the Macroeconomic Language","zh_title":"学习宏观经济语言","abstract":"We show how state-of-the-art large language models (LLMs) can be trained effectively on limited historical data for macroeconomic forecasting. We estimate a dynamic stochastic general equilibrium (DSGE) model with stochastic volatility and Student-t shocks on an initial segment of the data to obtain a posterior distribution over structural parameters. We sample from this posterior to generate millions of theory-consistent synthetic panels that, when mixed with macroeconomic data, form the training corpus for a deep learning transformer. Rather than using the DSGE model as a direct forecasting device, we use it as a structural simulator that regularizes transformer training in a small-sample environment. Our results show that this hybrid forecaster, which combines the theoretical coherence of DSGE models with the representational power of modern LLMs, learns key features of the macroeconomic language. Relative to a conventional vector autoregression benchmark, the transformer achieves comparable or stronger predictive performance in both token accuracy and predictive likelihood. These gains are strongest under a theory-heavy training mix and remain broadly robust to finer tokenization and deeper network architecture.","authors":["Siddhartha Chib","Fei Tan","Zhixun Zhang"],"categories":["econ.EM"],"primary_category":"econ.EM","announce_type":"replace","date":"2026-08-18","first_seen":"2025-12-24","revised_at":"2026-08-18","abs_url":"https://arxiv.org/abs/2512.21031","pdf_url":"https://arxiv.org/pdf/2512.21031","source_feed":"econ.EM","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["宏观经济预测","DSGE模型","时间序列Transformer"],"reason":"用LLM做宏观经济预测，属于纯NLP能力评测，不以人类行为为参照系。","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:58:53","error":null,"has_summary":false,"summary":null},{"id":"2407.00890","version":5,"title":"Macroeconomic Forecasting with Large Language Models","zh_title":"基于大语言模型的宏观经济预测","abstract":"This paper presents a comparative analysis evaluating the accuracy of Large Language Models (LLMs) against traditional macro time series forecasting approaches. In recent times, LLMs have surged in popularity for forecasting due to their ability to capture intricate patterns in data and quickly adapt across very different domains. However, their effectiveness in forecasting macroeconomic time series data compared to conventional methods remains an area of interest. To address this, we conduct a rigorous evaluation of LLMs against traditional macro forecasting methods, using as common ground the FRED-MD database. Our findings provide valuable insights into the strengths and limitations of LLMs in forecasting macroeconomic time series, shedding light on their applicability in real-world scenarios","authors":["Andrea Carriero","Davide Pettenuzzo","Shubhranshu Shekhar"],"categories":["econ.EM","cs.CL","cs.LG"],"primary_category":"econ.EM","announce_type":"replace-cross","date":"2026-08-18","first_seen":"2024-07-01","revised_at":"2026-08-18","abs_url":"https://arxiv.org/abs/2407.00890","pdf_url":"https://arxiv.org/pdf/2407.00890","source_feed":"cs.LG","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["宏观经济预测","时间序列","模型比较"],"reason":"纯宏观经济预测模型比较，不涉及人类行为仿真或人类被试替代。","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:58:57","error":null,"has_summary":false,"summary":null},{"id":"2608.10875","version":2,"title":"VibeLifeBench: Can Your Life Agent Be Proactive and Persistent in a Living World?","zh_title":"VibeLifeBench：你的生活代理能否在生活世界中保持主动与持续？","abstract":"Large language model (LLM) agents are increasingly deployed as personal assistants. Existing evaluations, however, mostly use short, self-contained requests in static environments. Everyday life assistance is different. A task runs for weeks rather than minutes. The world keeps changing while the agent is not being prompted. Many constraints are never stated outright. An agent that merely answers the request in front of it will fail at such a task. What is needed instead is an agent that stays proactive and consistent. It decides on its own when to act, when to ask, and when to stay silent. It notices changes that nobody announced. It keeps one plan coherent from the first day to the last. No current benchmark measures this. We introduce VibeLifeBench, a benchmark of 200 long-horizon tasks across ten everyday-life domains. Each task is a scripted multi-week timeline in a simulated world of 22 mock services. The world advances on its own clock, and many of its changes are silent, so only an agent that re-inspects the world discovers them. Every task is graded by fine-grained, weighted checks that read only what the agent actually left behind, covering the end state, the timeliness of its actions, and whether it upheld the implicit constraints. We evaluate seven frontier models. All of them score low, which shows how far current agents are from assisting with real life. We will open-source all tasks, environments, and the evaluation framework.","authors":["Xiaohongshu Dots Studio","Evolvent AI"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"replace-cross","date":"2026-08-18","first_seen":"2026-08-12","revised_at":"2026-08-18","abs_url":"https://arxiv.org/abs/2608.10875","pdf_url":"https://arxiv.org/pdf/2608.10875","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["LLM代理","长期任务评测","模拟世界"],"reason":"纯多智能体系统评测，agent在模拟世界中执行长期任务，无人类行为对照，不涉及…","model":"deepseek-v4-pro","scored_at":"2026-08-12T13:03:09","error":null,"has_summary":false,"summary":null},{"id":"2608.12104","version":2,"title":"No One to Blame: A Framework of Constitutive AI Unaccountability","zh_title":"无人可责：构成性AI不可问责性的框架","abstract":"The increasing deployment of autonomous, agentic AI systems challenges traditional accountability mechanisms. Existing research predominantly frames AI accountability gaps as barriers that can be overcome through better standards, transparency, and institutional reform. We argue that this framing is insufficient: certain configurations of actors, systems, and institutions render AI accountability conceptually unachievable regardless of effort. We introduce the concept of constitutive AI unaccountability to capture these configurations. Through a three-stage qualitative study comprising a concept-centric literature analysis, a secondary analysis of 27 expert interviews with AI professionals from technical, legal, and sociotechnical backgrounds, and an illustrative framework application to the open-source agentic AI system OpenClaw, we identify nine categories and 20 themes of constitutive AI unaccountability. These are organized across structural, technological, and normative clusters and reinforce one another through eight directed interdependencies. Our framework is operationalized as a diagnostic instrument of 20 questions, which detected 17 of 20 conditions when applied to OpenClaw, including an inverted anthropomorphism configuration in which the AI agent was the only identifiable actor. We contribute a reframing of AI unaccountability as a constitutive property of sociotechnical systems, an extension of the four barriers to accountability, and a practical instrument for identifying accountability voids in specific AI deployments.","authors":["Long Hoang Nguyen","Eva Sp\\\"athe","Sebastian Lins","Ali Sunyaev"],"categories":["cs.CY","cs.AI"],"primary_category":"cs.CY","announce_type":"replace-cross","date":"2026-08-18","first_seen":"2026-08-13","revised_at":"2026-08-18","abs_url":"https://arxiv.org/abs/2608.12104","pdf_url":"https://arxiv.org/pdf/2608.12104","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["AI问责制","多智能体系统","社会技术系统"],"reason":"研究AI问责制，不涉及LLM仿真人类被试或与人类行为对照","model":"deepseek-v4-pro","scored_at":"2026-08-13T13:01:57","error":null,"has_summary":false,"summary":null},{"id":"2608.13606","version":2,"title":"MobileMem: Learning from a Year of Mobile Experiences","zh_title":"MobileMem：从一年的移动体验中学习","abstract":"The next generation of AI agents is increasingly moving beyond systems that answer isolated questions toward persistent personal assistants that can understand, remember, and continuously learn from users' experiences. Such assistants require long-term memory to accumulate and leverage user-specific experiences over time, yet existing benchmarks remain inadequate for realistic mobile settings, where experiences are heterogeneous, multimodal, evolving, and deeply personal. We introduce MobileMem, a benchmark and framework for studying on-device long-term memory, grounded in a year-scale collection of mobile experiences. MobileMem employs a knowledge-grounded synthesis pipeline to construct coherent and temporally consistent long-horizon trajectories from user-app sessions. It provides complementary text and multimodal settings covering multi-hop and temporal reasoning, knowledge updating, and implicit preference inference. Specifically, MobileMem enables agents to remember the past, understand the present, and adapt to the future. By modeling experiences rather than isolated facts, MobileMem moves memory beyond information retrieval toward experiential intelligence for continuous personal learning.","authors":["Xinle Deng","Yida Xue","Xiangyuan Ru","Yijun Chen","Buqiang Xu","Mingjun Mao","Xinjie Liu","Haoming Xu","Shuofei Qiao","Mengru Wang","Chen Jiang","Yuchen Eleanor Jiang","Lizhong Wang","Jason Wang","Li Zeng","Haofen Wang","Guilin Qi","Huajun Chen","Ningyu Zhang"],"categories":["cs.AI","cs.CL","cs.LG","cs.MA","cs.MM"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-08-18","first_seen":"2026-08-17","revised_at":"2026-08-18","abs_url":"https://arxiv.org/abs/2608.13606","pdf_url":"https://arxiv.org/pdf/2608.13606","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["长期记忆","个人助手","基准测试"],"reason":"研究个人助手长期记忆，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-08-17T13:01:25","error":null,"has_summary":false,"summary":null},{"id":"2608.14558","version":1,"title":"The Unwritten Benchmark: A New Challenge for Multimodal Machine Learning in Abstract Perceptual Reasoning","zh_title":"未写出的基准：多模态机器学习在抽象感知推理中的新挑战","abstract":"Current multimodal models have demonstrated remarkable proficiency in recognizing static visual and auditory content. However, their capacity for abstract perceptual reasoning, inferring unseen information from dynamic, generative processes, remains a critical and underexplored frontier. In this paper, we introduce The Unwritten Benchmark, a new challenge designed to probe this abstract perceptual and cognitive ability. We define the core task as acousto-kinematic word inference: models must decipher words, across 3 different writing styles, being written solely from the audio of pen scratches and the video of hand movements, without any visible ink trace. Our evaluation results reveal a profound gap between human and machine performance: while human participants achieve high ordered letter accuracy (over 80%), leading Multimodal Machine Learning Models, including GPT-4o and Gemini 2.5-Pro, struggle significantly, failing to surpass 10%. Furthermore, we identify a paradoxical fusion effect in the models, where providing both modalities often degrades performance rather than improving it. This finding indicates a fundamental breakdown in their ability to synthesize complementary perceptual cues for this cognitive task. These findings highlight significant limitations in both cross-modal causal reasoning and the understanding of the micro-kinematics essential for such cognitive and intuitive perceptual reasoning.","authors":["Garima Arya Yadav","Nilay Yilmaz","Yezhou Yang"],"categories":["cs.AI","cs.CV"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-18","first_seen":"2026-08-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.14558","pdf_url":"https://arxiv.org/pdf/2608.14558","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C2"],"tags":["多模态基准","感知推理","模型评测"],"reason":"该论文是视觉-听觉多模态推理基准，不涉及LLM仿真人类被试，属于感知推理评测。","model":"deepseek-v4-pro","scored_at":"2026-08-18T13:04:27","error":null,"has_summary":false,"summary":null},{"id":"2608.14631","version":1,"title":"Accuracy and Reliability of Large Language Models in Cosmetic Chemistry and Skin Health: A Benchmarking Study","zh_title":"大语言模型在化妆品化学与皮肤健康中的准确性与可靠性：一项基准研究","abstract":"As consumers increasingly turn to AI chatbots for skincare advice, the technical accuracy of Large Language Models (LLMs) in cosmetic chemistry remains largely under-evaluated. We benchmarked 14 LLMs on a structured set of topics related to cosmetic chemistry, including the chemical properties of specific cosmetic ingredients and common cosmetic scenarios that may be of interest to consumers. Web search was disabled throughout to assess each model's internalized knowledge rather than its internet retrieval capacity. Overall performance was poor, with the most pronounced deficits in quantitative reasoning and structural identification tasks. While models handled general skincare questions with reasonability, responses consistently lacked the technical depth required for informed consumer decision-making. Notably, conversation with AI can pose a risk: outputs that sound authoritative but contain technical errors are less likely to generate skepticism compared to responses that explicitly acknowledge uncertainty. These findings suggest that general-purpose LLMs, trained predominantly on unverified public data, are currently not reliable sources of cosmetic chemistry information. Progress on two fronts, fine-tuning verified chemical and dermatological datasets, and substantial improvements to algorithmic reasoning, will likely be needed before these tools can be considered as resources for public use.","authors":["Amelia Liu"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-18","first_seen":"2026-08-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.14631","pdf_url":"https://arxiv.org/pdf/2608.14631","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["LLM评测","化妆品化学","可靠性"],"reason":"纯LLM能力评测，无人类被试仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-08-18T13:04:31","error":null,"has_summary":false,"summary":null},{"id":"2608.14765","version":1,"title":"Agentic Data Cleaning Without a Clean Reference: An Experimental Study of Capabilities and Trade-offs","zh_title":"无干净参照的智能体数据清洗：能力与权衡的实验研究","abstract":"Data cleaning without a trusted clean reference is challenging because unusual values may represent either genuine errors or valid observations. This paper studies how different agent capabilities affect reference-free data cleaning and proposes an evidence-grounded framework that combines structured context, profiling, LLM reasoning, executable checks, controlled evidence retrieval, source ranking, citation alignment, conservative repair, reversible scripts, and provenance logging. Seven configurations are evaluated across financial, clinical, and environmental-monitoring datasets using controlled synthetic corruption and original-data descriptive analysis, resulting in 126 completed runs. The evaluation includes two comparison baselines and a progressive LLM-based sequence that adds executable tools, evidence retrieval, evidence controls, and conservative repair. In the synthetic evaluation, the deterministic profiling baseline achieved the highest detection F1-score of 0.561. Among the LLM-based configurations, the full conservative configuration achieved the highest F1-score of 0.421, but no configuration performed best across all evaluation criteria. The source-ranked configurations achieved the lowest unsupported-rule rates, while decision-level citation alignment remained weak. The full conservative configuration produced no unsafe or unnecessary modifications, although these rates were already zero before the conservative policy was added, and it performed no direct repairs. Overall, the results show that additional capabilities introduce trade-offs among detection, repair, evidence grounding, conservative behaviour, reproducibility, and operational cost rather than producing consistent improvements. The study provides a structured framework and empirical methodology for evaluating these trade-offs in reference-free agentic data cleaning.","authors":["Hadi Fadlallah"],"categories":["cs.AI","cs.DB"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-18","first_seen":"2026-08-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.14765","pdf_url":"https://arxiv.org/pdf/2608.14765","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["数据清洗","LLM agent","多智能体协作"],"reason":"研究LLM agent进行数据清洗，属于多智能体协作完成任务，不涉及人类行为仿…","model":"deepseek-v4-pro","scored_at":"2026-08-18T13:04:34","error":null,"has_summary":false,"summary":null},{"id":"2608.14795","version":1,"title":"Individual Disempowerment through an Advice Channel: Control Loss when Influence is Endogenous","zh_title":"通过建议渠道的个体去权能化：内生影响下的控制损失","abstract":"An AI that can only give advice seems safe: the human is always free to ignore it. That is the premise of the boxing tradition in AI safety, and its long-suspected weak point is that the human who reads the answers is part of the system. We make the fraction $\\varepsilon_t$ of behavior that follows the advice a state of a Markov decision process, moved by the advisor's own messages, so that use deepens reliance. Granted a channel rich enough to echo any action the human could take, higher $\\varepsilon_t$ weakly lowers every monotone measure of the power of a human with a message-independent fallback. An oracle rewarded by per-round approval cultivates reliance beyond a closed-form patience threshold, so the same reward weights leave the optimal oracle answering in episodic deployments and cultivating in long-memory ones. An influence bound certified once at deployment is blind to that horizon and bounds the loss no lower than its trivial ceiling. An exogenous cap on influence bounds the guarantee the human loses, and a short enough memory reset removes the incentive to cultivate, while neither recovers the value already steered away. In a closed-form example the optimal oracle never cultivates in fifteen-round sessions and does in sixteen.","authors":["Adam M. Oberman"],"categories":["cs.AI","cs.GT"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-18","first_seen":"2026-08-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.14795","pdf_url":"https://arxiv.org/pdf/2608.14795","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["AI安全","人机交互","控制权"],"reason":"研究AI建议对人类控制权的影响，非LLM仿真人类被试，无人类数据对照。","model":"deepseek-v4-pro","scored_at":"2026-08-18T13:04:34","error":null,"has_summary":false,"summary":null},{"id":"2608.14804","version":1,"title":"Generated Context versus Governed State: Functional Conditions for Accountable Longitudinal Clinical Reasoning","zh_title":"生成上下文与受治理状态：可问责纵向临床推理的功能条件","abstract":"Large language models (LLMs) have become the dominant interface of clinical artificial intelligence, yet the interface they expose (text in, text out, one context window at a time) maintains no explicit, persistent, governed representation of what is currently true about a patient. This paper argues that longitudinal clinical reasoning is a state-estimation problem under partial observability, and that the axis on which clinical AI succeeds or fails is not the fluency of the model reading the record but the governance of the patient state it reasons over. We distinguish generated context from governed state; separate five objects that clinical AI habitually conflates (true state, observations, evidence, belief, and simulated state); define a tiered governance standard against which any clinical AI system can be audited; and show that an operational definition of accountability decomposes into four information requirements: an immutable evidence ledger with awareness-time versioning, a belief state distinct from accumulated evidence, an observation-process model, and claim-level causal typing. We are explicit that this decomposition is analytic rather than a necessity theorem, and that its value is conceptual hygiene: it converts \"accountable clinical AI\" from a slogan into an audit instrument. A six-level maturity framework separates what a system makes governable from what it can compute, locating current LLM-centric practice at high capability but low maturity. The paper is fully self-contained: the four research questions the framework poses are stated in the introduction, and the conclusion records what the paper establishes toward each; future work develops the buildable core of the architecture and the research program toward full Clinical World Models. No empirical result is claimed here.","authors":["Augusto Bernardo Pissarra","Victor Lorena de Farias Souza"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-18","first_seen":"2026-08-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.14804","pdf_url":"https://arxiv.org/pdf/2608.14804","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["临床AI","状态治理","问责框架"],"reason":"论文讨论临床AI的状态治理与问责框架，不涉及用LLM仿真人类被试或与人类数据对…","model":"deepseek-v4-pro","scored_at":"2026-08-18T13:04:34","error":null,"has_summary":false,"summary":null},{"id":"2608.14992","version":1,"title":"Does a Tool Result Carry More Authority Than Plain Text? Three Prospective Studies of False-Claim Adoption in a Synthetic Assignment Task with Claude Opus 5","zh_title":"工具结果是否比纯文本更具权威性？三项关于Claude Opus 5在合成任务中采纳错误声明的前瞻性研究","abstract":"Language-model systems increasingly read from stores they also write to, so a claim that was merely written earlier can return looking retrieved. We tested whether the message package carrying an unsupported assignment changes which answer a model gives in a synthetic lookup task. Claude Opus 5 selected a color code for a named item or abstained. In an exploratory four-arm study, false-code adoption was 0/24 with no target claim, 0/22 scorable trials when a prior assistant assertion named the target, 14/24 when a tool-result record named it, and 15/24 when that result used a ten-field metadata wrapper that marked it unchecked. The tool-result arm selected the record's code in 11/12 supported trials and 14/24 unsupported trials, ruling out a fixed output-token bias while leaving substantial planted-token heterogeneity. A document-preregistered replication reproduced the tool-result versus assistant-assertion gap, 7/24 against 0/24, one-sided Fisher exact p = 0.0047. The tool-result rate nevertheless fell from 14/24 to 7/24 across runs made four days apart. A second preregistered study gave the earlier comparison a live text control: both records were announced in advance and placed in the same final user turn, then target binding was swapped between the linked tool result and later inline JSON. Inline text was sufficient for false-code adoption in 60/60 trials; the tool-result condition produced 57/60, so the registered result-first superiority criterion failed, p = 1. The result does not show that tool results have no effect. It shows that native tool-result placement was not necessary and that this experiment did not find greater behavioral weight for the result package than for announced inline text. The findings concern a single model on one synthetic task template, accessed through one API.","authors":["Justin Bronder"],"categories":["cs.AI","cs.CL"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-18","first_seen":"2026-08-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.14992","pdf_url":"https://arxiv.org/pdf/2608.14992","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["LLM行为","工具结果权威性","模型评估"],"reason":"研究LLM对工具结果与纯文本的权威性反应，属模型行为测试，无人类被试对照，非人…","model":"deepseek-v4-pro","scored_at":"2026-08-18T13:04:36","error":null,"has_summary":false,"summary":null},{"id":"2608.15131","version":1,"title":"Platform Adaptation Under Governance Interventions: Actor Best-Response Modeling and an External Public-Case Benchmark","zh_title":"治理干预下的平台适应：行动者最佳响应建模与外部公共案例基准","abstract":"Digital platforms govern by changing rules: rankings, monetization thresholds, moderation standards, verification systems, disclosure requirements, appeal processes, and access policies. These interventions are rarely absorbed passively. Creators, sellers, advertisers, moderators, users, developers, and strategic operators adapt to the new reward surface. This paper develops a platform-adaptation model for evaluating governance interventions as transitions in adaptive multi-actor information systems. The model represents actor best response, strategic gaming opportunity, moderation burden, user-incentive movement, enforcement response, externality formation, and downstream platform stability. We evaluate the model on 72 external public platform-governance cases covering media monetization, ranking systems, verification, delivery platforms, marketplaces, app stores, community platforms, and creator ecosystems. Across 9 methods and 648 method-case evaluations, the full platform-adaptation simulator achieves mean adaptation quality of 0.836338, compared with 0.669731 for a risk-register baseline, 0.589457 for causal-loop analysis, 0.492750 for generic governance critique, 0.369492 for engagement-only optimization, and 0.331965 for baseline policy review. Paired comparisons show a win rate of 1.00 against all tested baselines and channel ablations. The contribution is an information-systems theory and measurement framework showing why platform governance evaluation fails when it treats policy rules as static controls rather than interventions into adaptive actor-response fields.","authors":["Wesley Shu","Peng Wei"],"categories":["cs.AI","cs.CY"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-18","first_seen":"2026-08-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.15131","pdf_url":"https://arxiv.org/pdf/2608.15131","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["平台治理","多智能体仿真","信息系统"],"reason":"多智能体平台治理仿真，无LLM作为人类被试，无人类行为对照，属纯系统建模。","model":"deepseek-v4-pro","scored_at":"2026-08-18T13:04:39","error":null,"has_summary":false,"summary":null},{"id":"2608.15254","version":1,"title":"Demographic Injection in Medical Language Models under Diversity, Equity, and Inclusion Prompts","zh_title":"多样性、公平与包容提示下医学语言模型的人口统计注入","abstract":"Clinical-AI guidance increasingly recommends prompting language models to reason with attention to diversity, equity, and inclusion (DEI). We measure a side effect that misrepresents patients: a one-sentence DEI prompt appended to a medical question leads models to add patient demographic attributes (race, socioeconomic status, sex) the question never stated, in effect rewriting who the patient is. We call this demographic injection. Across 47 models, four medical benchmarks, and 376,000 responses scored by a validated model-judge pipeline, a single DEI prompt raises the injection rate from 0.7% to 33.1% (47x) in all 47 of 47 models, attributable to the equity content rather than to added length (18x above a length-matched control; p=1.4x10^-14). Most added content is a general population statement that leaves the answer unchanged, but a smaller subset attaches an attribute to the specific patient or changes the selected option (0.25-2.4% of responses, 99.8% toward the incorrect option), where the invented demographic changes the answer the model recommends. Phrasing scales the effect from 14% to 56%. DEI prompts are just one example of a more general mechanism. Any instruction that nudges how a model reasons can make it add unrequested details, including details about the patient. Flagged outputs are treated as model errors under study, not clinical guidance.","authors":["Diego Mardian","Frank Liu"],"categories":["cs.AI","cs.CL"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-18","first_seen":"2026-08-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.15254","pdf_url":"https://arxiv.org/pdf/2608.15254","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["LLM偏差","医疗AI","提示工程"],"reason":"研究DEI提示对医学LLM输出影响，属模型行为分析，非人类仿真","model":"deepseek-v4-pro","scored_at":"2026-08-18T13:04:39","error":null,"has_summary":false,"summary":null},{"id":"2608.15309","version":1,"title":"Physiological World Models for Human State Transitions","zh_title":"用于人类状态转换的生理世界模型","abstract":"Continuous multimodal sensing now allows human physiology to be observed throughout daily life rather than only during occasional clinical visits. However, most health artificial intelligence systems are designed to recognize current states, estimate risks or analyse individual biomarkers. They do not directly model how physiological states change in response to real-world events, behaviours, contexts and interventions. Here we propose the Physiological World Model (PWM), an event-conditioned framework for learning these changes at the level of the whole person. We introduce the HumanState Transition Token, a structured, quality-scored unit that connects the physiological state before an event with the event or action, relevant context and intervention information, the physiological trajectory after the event, observed outcomes and data quality. We describe four capability levels, from state representation to bounded intervention planning, together with four data acquisition and validation protocols. We also propose six benchmark tasks covering HumanState representation, forecasting across multiple timescales, individualized response prediction, simulation of alternative interventions, bounded planning and reliability under distribution shift. Together, this framework provides a practical path towards personalized health management, behavioural intervention design and clinician-supervised decision support, while clearly separating prediction from causal inference and making uncertainty, safety, governance and limits of use explicit.","authors":["Chongyang Zhang","Rendong Wang","Hao Zheng","Hanwen Zhang","Yang Liu","Xiaolong Wei","Bin Chong"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-18","first_seen":"2026-08-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.15309","pdf_url":"https://arxiv.org/pdf/2608.15309","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C2"],"tags":["生理建模","健康AI","状态转换"],"reason":"论文研究人类生理状态建模，不涉及LLM仿真人类被试，属于健康AI领域。","model":"deepseek-v4-pro","scored_at":"2026-08-18T13:04:40","error":null,"has_summary":false,"summary":null},{"id":"2608.16370","version":1,"title":"What Does Context Compression Cost an Agent? Interaction Costs Unrevealed by Task-Completion Metrics","zh_title":"上下文压缩对智能体的代价是什么？任务完成指标未揭示的交互成本","abstract":"Task completion is the standard metric for evaluating context compression, yet it is incomplete: compression can increase an agent's interaction cost by forcing it to reacquire dropped state while leaving completion statistically unchanged. We introduce a controlled runtime measurement protocol for reacquisition cost in a bounded-horizon tool-using agent. The agent acts in a deterministic planning environment under a fixed 24-turn horizon. We vary compression severity, compare a dropping operator with a fact-preserving operator, restore dropped state through controlled oracle interventions, and decompose tool calls into retrieval and execution. We evaluate three models across two task regimes. Retrieval calls increase in all six model-regime comparisons and account for almost all added interaction; five of six remain significant after Holm correction. At the prespecified 5x comparison point, completion changes are not significant in any cell. DeepSeek shows a significant completion drop only at 10x compression. GPT-5.5 is the clearest case: completion changes from 80% to 85% (p = 1.0) while retrieval increases from 21.0 to 63.9 calls (p = .002). Retention interventions further separate state quantity, state type, and content validity. Random selection is comparable to an offline hindsight oracle, while replacing retained D-state with semantically irrelevant content increases retrieval by 57% (p < .001) without a significant completion change. In a second environment, ALFWorld, sliding compression produces no retrieval surge, showing that the reacquisition signature is environment-dependent rather than intrinsic to shortening context. Overall, compression can impose hidden interaction costs when execution-relevant state becomes absent and must be reacquired, while completion alone may not expose those costs.","authors":["Shuyu Liu"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-18","first_seen":"2026-08-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.16370","pdf_url":"https://arxiv.org/pdf/2608.16370","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["上下文压缩","智能体交互成本","工具调用"],"reason":"研究多智能体工具调用中的上下文压缩成本，不涉及人类行为仿真或对照。","model":"deepseek-v4-pro","scored_at":"2026-08-18T13:04:50","error":null,"has_summary":false,"summary":null},{"id":"2608.14577","version":1,"title":"HarmProfile: Characterizing Harmful Distributions in Frontier LLMs","zh_title":"HarmProfile：刻画前沿大语言模型中的有害分布","abstract":"Frontier large language models (LLMs) safety evaluation has largely treated harmful generation as an attack outcome rather than as an object of analysis. Consequently, little is known about the harmful outputs produced during model misbehavior, partly because large-scale, high-quality collections of frontier-LLM misbehavior are difficult to obtain. To address this gap, we introduce HarmProfile, a content-centric benchmark dataset that collects model misbehavior across diverse harm categories and model families, and defines the resulting harmful-output distribution as a model-level risk profile. The premise is that, just as linguistic behavior can be characterized from an utterance corpus, model risk can be characterized from the content, severity, and variation of its safety failures. HarmProfile contains over 80,000 validated artifacts from 23 frontier LLMs across 13 model families, organized into 15 harm categories and 57 subcategories. Using this corpus, we find that frontier LLMs reliably produce harmful content at scale, yet exhibit distinct risk profiles; both harmfulness and diversity grow with model capability, suggesting that frontier LLMs may appear safe yet harbor increasingly dangerous knowledge beneath the alignment surface. Our source code is available at https://github.com/fresh-ma/HarmProfile .","authors":["Zhouyuan Ma","Yutao Wu","Hanxun Huang","Xiang Zheng","Xiao Liu","Yixin Cao","Zuxuan Wu","Xingjun Ma","Yu-Gang Jiang"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"cross","date":"2026-08-18","first_seen":"2026-08-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.14577","pdf_url":"https://arxiv.org/pdf/2608.14577","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["LLM安全","基准数据集","有害内容"],"reason":"该论文构建LLM有害输出基准，属于模型安全评测，不涉及人类行为仿真或对照。","model":"deepseek-v4-pro","scored_at":"2026-08-18T13:04:28","error":null,"has_summary":false,"summary":null},{"id":"2608.14609","version":1,"title":"Understanding AI Anxiety in the Workplace: A Multimethod Investigation Using Fear Acquisition Theory and the Technology Acceptance Model","zh_title":"理解职场中的AI焦虑：基于恐惧习得理论与技术接受模型的多方法研究","abstract":"As artificial intelligence (AI) rapidly diffuses and concerns about job displacement intensify, the psychological mechanisms underlying AI job replacement anxiety remain insufficiently understood. Drawing on Integrated Fear Acquisition Theory and the Technology Acceptance Model, the present research investigates whether AI job replacement anxiety can be elicited through vicarious exposure to narratives emphasizing AI-over-human control, and whether perceived usefulness and perceived ease of use of AI moderate this response. Across two studies, we examine AI job replacement anxiety as a response that emerges through vicarious exposure to narratives emphasizing AI agency and human control loss, rather than through direct personal experience of job displacement. Study 1 employed a randomized experiment (N = 316), demonstrating that such exposure increased AI job replacement anxiety. This effect was moderated by perceived usefulness of AI, but not by perceived ease of use, and remained robust after controlling for core self-evaluations. Study 2 (N = 995) replicated the association between perceived AI-over-human control and job replacement anxiety in an observational design and provided convergent evidence for the moderating role of perceived usefulness, supporting the external validity of the findings. Together, the results provide the first causal evidence that perceptual and vicarious processes can trigger AI job replacement anxiety. By shifting attention from structural labor-market conditions to how AI agency is perceived and communicated, this work offers a mechanism-based account of when and why AI-related job fears arise.","authors":["Jaroslaw Grobelny","Mateusz Klakus","Kacper Szyma\\'nski","Teresa Chirkowska-Smolak"],"categories":["cs.CY","cs.AI"],"primary_category":"cs.CY","announce_type":"cross","date":"2026-08-18","first_seen":"2026-08-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.14609","pdf_url":"https://arxiv.org/pdf/2608.14609","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C5"],"tags":["AI焦虑","人类被试","技术接受模型"],"reason":"研究人类对AI的焦虑，未用LLM仿真人类被试，而是用人类被试研究AI影响。","model":"deepseek-v4-pro","scored_at":"2026-08-18T13:04:29","error":null,"has_summary":false,"summary":null},{"id":"2608.14625","version":1,"title":"Local AI pre-screening for human triple-blind peer review in health sciences","zh_title":"健康科学领域人类三盲同行评审的本地AI预筛选","abstract":"Academic peer review is under mounting strain: NeurIPS 2025 received 21,575 submissions, ICLR 2025 received 11,603, and ICML 2025 received 12,107. This volume has outpaced the supply of qualified reviewers, and large language models (LLMs) are already filling the gap, largely undisclosed. An independent analysis of ICLR 2026 found roughly 21% of its 75,800 peer reviews were fully AI-generated, with over half showing some AI involvement (up from 15.8% in 2024). Documented risks include hallucinated citations in accepted papers and hidden prompt-injection instructions embedded in manuscripts to manipulate AI reviewers into favorable assessments. We propose a triple-blind, multi-LLM pre-screening framework for peer review, developed for a health sciences journal, that formalizes and discloses AI involvement while preserving human reviewers as the final decision-making authority. The framework routes a submission through five stages -- sanitization/anonymization, parallel AI pre-screening, an automated check gate, blinded human review, and editorial adjudication -- with return-to-author loops at the check and editor stages. Addressing the confidentiality concerns behind NIH/NSF bans on submitting unpublished proposals to third-party generative AI, all three AI reviewers run on locally-hosted, open-weight LLMs, keeping manuscript content within the journal infrastructure. The closest precedent, Shen et al., benchmarked five open-source LLMs on quartile classification of 200 manuscripts and found accuracy insufficient (35% exact-match) for autonomous use, supporting our decision to retain mandatory human adjudication. This transparent, human-supervised design offers a defensible alternative to today's opaque, unregulated AI use in peer review, potentially reducing the substantial delay of traditional review (avg. 13 weeks to first decision) without displacing human judgment.","authors":["Rodrigo Martins Boos"],"categories":["cs.CY","cs.AI","cs.DL"],"primary_category":"cs.CY","announce_type":"cross","date":"2026-08-18","first_seen":"2026-08-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.14625","pdf_url":"https://arxiv.org/pdf/2608.14625","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["同行评审","多LLM框架","AI辅助审稿"],"reason":"论文提出多LLM预审框架，用于同行评审流程优化，属于多智能体协作完成任务，不涉…","model":"deepseek-v4-pro","scored_at":"2026-08-18T13:04:31","error":null,"has_summary":false,"summary":null},{"id":"2608.15286","version":1,"title":"No Task Fails Every Time: Why One-Shot Audits Are Structurally Blind to Agent Damage","zh_title":"没有任务每次都失败：为何一次性审计对智能体损害存在结构性盲区","abstract":"We introduce AgentRelBench, an environment-agnostic reliability instrument that computes ground-truth, severity-priced damage from database state diffs across repeated runs, with no LLM in the measurement path, demonstrated on EnterpriseOps-Gym. Across 2,128 evaluation runs spanning nine models in six families (four development, three pre-registered held-out, plus a frontier pass on two frontier-tier models that the pre-registration designates exploratory), we find: (1) damage on irreversible actions is universal across the families we measured and stochastic within them on pinned, single-provider stacks. (2) No task damaged on every run: zero always-fail cells across 42 confirmatory held-out damage events. A single clean run misses a damage-producing (model, task) pair 0.80 of the time on the development pool (13 pairs); the held-out pool is descriptively consistent (0.575 over 5 pairs, pair-weighted) but sits below our pre-registered power floor and is reported as underpowered, not as confirmation. (3) Damage-producing task count falls with model capability, from 7 of 20 tasks for an 8B model to 1 of 20 for the most capable; capability is confounded with family and training, so this is an observed gradient, not a causal claim. The residual damage does not change in character: in the exploratory frontier pass, the most capable model's one damaging task damages at $\\hat{p} = 0.16$ per run, inside the same demonstrably-stochastic band, and a single audit misses it 84% of the time. (4) One model family committed the gated irreversible change while declaring it had refused: transcript- and judge-based grading scores those runs as safe refusals, only state diffs as damage. All confirmatory findings were pre-registered with per-claim demote criteria; one demoted our own initially favored finding, which we report.","authors":["Shiven Khurdi"],"categories":["cs.LG","cs.AI"],"primary_category":"cs.LG","announce_type":"cross","date":"2026-08-18","first_seen":"2026-08-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.15286","pdf_url":"https://arxiv.org/pdf/2608.15286","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","可靠性评估","基准测试"],"reason":"研究多智能体系统可靠性，无人类行为对照，属纯多智能体系统研究。","model":"deepseek-v4-pro","scored_at":"2026-08-18T13:04:39","error":null,"has_summary":false,"summary":null},{"id":"2608.15689","version":1,"title":"Integrating Persuasion Theory into the Epidemiological Modelling of Health Misinformation Spread on Social Media","zh_title":"将说服理论整合到社交媒体健康错误信息传播的流行病学建模中","abstract":"This study presents a hybrid epidemiological and behavioural framework to simulate the spread of health misinformation on social media. We extend the classical Susceptible--Infected--Recovered (SIR) model to a six-compartment structure (SIRMMM), incorporating Misinformed Susceptible (MS), Misinformed Infected (MI), and Misinformed Recovered (MR) compartments to better reflect the dynamics of the misinformation lifecycle. To account for individual-level behavioural variation, we extend the SIRMMM model by integrating psychological signals from the Elaboration Likelihood Model (ELM), including sentiment polarity, engagement metrics, and cognitive effort, which dynamically modulate the misinformation transmission rate, yielding the ELM-SIRMMM framework. Model parameters were estimated using the FibVID dataset, which captures COVID-19 misinformation on Twitter. Generalisability was tested on two additional datasets: MC-Fake (emotional misinformation) and Monant (general health misinformation). Results show that the ELM-SIRMMM model enhances both predictive accuracy and dynamic realism. On FibVID, it decreases RMSE by 5.5%, delays the misinformation peak from day 150 to day 160, and increases its peak prevalence from 6% to 7%. On MC-Fake, it accurately reproduces a flash-rumour pattern, infecting 38% of users by day 45 and achieving 97% misinformation recovery, all while maintaining model accuracy. In contrast, minimal behavioural signal variability in the Monant dataset leads to marginal benefit, with only a 3% peak and 57% of users remaining susceptible. These findings suggest that structural elaboration alone is insufficient. Functional realism in modelling misinformation spread requires dynamic psychological inputs that vary meaningfully across time and contexts.","authors":["Mkululi Sikosana","Sean Maudsley-Barton","Oluwaseun Ajao"],"categories":["cs.SI","cs.AI","cs.CL","cs.LG"],"primary_category":"cs.SI","announce_type":"cross","date":"2026-08-18","first_seen":"2026-08-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.15689","pdf_url":"https://arxiv.org/pdf/2608.15689","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["信息传播模型","流行病学","社交媒体"],"reason":"论文用流行病学模型模拟信息扩散，未使用LLM作为人类被试，不涉及人类行为仿真。","model":"deepseek-v4-pro","scored_at":"2026-08-19T13:03:29","error":null,"has_summary":false,"summary":null},{"id":"2608.15867","version":1,"title":"Feasible and Novel Synthetic Population Generation with Tabular and Sequential Travel Attributes","zh_title":"具有表格和顺序出行属性的可行且新颖的合成人口生成","abstract":"Synthetic populations are critical inputs for activity-based travel demand models, yet generating realistic populations from limited survey data remains challenging. Small samples miss valid attribute combinations, known as sampling zeros, and generative models may also produce infeasible structural zeros. Moreover, realistic synthetic populations must capture both static socio-demographic attributes and sequential travel behaviour, such as trip chains. This paper proposes a regularized two-stage generative framework to address these challenges, where regularization refers to additional loss terms that guide the generator toward broader valid coverage and fewer infeasible samples. In Stage 1, a Wasserstein GAN with gradient penalty is augmented with three regularization terms, IGP, LDR, and CLAP, to improve feasibility, diversity, and novelty in tabular population synthesis. In Stage 2, Transformer and LSTM-Attention models generate sequential travel attributes, including departure time, trip purpose, and travel mode, conditioned on the synthesized tabular profiles. We also introduce novelty and count-aware metrics to evaluate whether valid unseen combinations are recovered and generated in realistic proportions. Results show that regularized models outperform the vanilla WGAN-GP across feasibility, diversity, and novelty. Regularization increases feasibility by 2.1 to 3.7 percentage points and novelty by 6.6 to 10.0 percentage points, improving sampling-zero recovery without sacrificing feasibility. The F1 score improves by 6.3 to 8.6 percentage points. For sequential attributes, LSTM-Attention best matches the trip-length distribution, while Transformer achieves higher overall sequential F1, 90.6\\% versus 89.1\\%. Cross-stage validation confirms strong consistency between generated mobility status and generated trip chains.","authors":["Farbod Abbasi","Zachary Patterson","Bilal Farooq"],"categories":["cs.LG","cs.AI"],"primary_category":"cs.LG","announce_type":"cross","date":"2026-08-18","first_seen":"2026-08-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.15867","pdf_url":"https://arxiv.org/pdf/2608.15867","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C2"],"tags":["合成人口","交通需求模型","生成对抗网络"],"reason":"论文生成合成人口用于交通需求模型，不涉及LLM仿真人类被试，属于交通仿真领域。","model":"deepseek-v4-pro","scored_at":"2026-08-19T13:03:29","error":null,"has_summary":false,"summary":null},{"id":"2608.16747","version":1,"title":"Would this change your answer? Evaluating Explanations of LLM Behavior In The Wild with Counterfactual Experiments","zh_title":"这会改变你的答案吗？用反事实实验评估野外LLM行为的解释","abstract":"Many areas of AI research, such as language model interpretability and chain of thought faithfulness, seek to explain model behaviors. But what constitutes a \"good\" explanation? In this work, we evaluate explanations through the lens of counterfactual simulatability-whether the explanation is useful for predicting model behaviors on related counterfactual inputs. To this end, we introduce CHIVE (Counterfactual Hypothesis Investigation Via Edits), a novel agentic pipeline that identifies unexpected model behaviors in the wild and investigates them with counterfactual prompt edits. This yields thousands of high-quality explanations for naturally-occurring model behaviors along with supporting counterfactual evidence. We apply CHIVE in two ways. First, we evaluate whether common LLM interpretability techniques improve an agent's ability to predict counterfactual model behaviors. Surprisingly, we find no uplift from any of the interpretability techniques studied. Second, we use CHIVE to generate training data. We find that training models to predict outcomes of CHIVE-generated counterfactual experiments generalizes to various out-of-distribution settings. Overall, CHIVE automatically discovers explanations of naturally-occurring LLM behaviors, enabling us to evaluate and improve methods for explaining LLM behaviors.","authors":["Adam Karvonen","Euan Ong","Subhash Kantamneni","Samuel Marks"],"categories":["cs.LG","cs.AI"],"primary_category":"cs.LG","announce_type":"cross","date":"2026-08-18","first_seen":"2026-08-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.16747","pdf_url":"https://arxiv.org/pdf/2608.16747","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["模型可解释性","反事实实验","LLM行为"],"reason":"论文评估LLM行为解释方法，不涉及人类被试仿真或人类数据对照，属于模型可解释性…","model":"deepseek-v4-pro","scored_at":"2026-08-19T13:03:30","error":null,"has_summary":false,"summary":null},{"id":"2608.15550","version":1,"title":"Adoption of Generative AI in the Workplace: Increasing and Shifting the Balance of Productivity and Communication Activity","zh_title":"生成式AI在工作场所的采用：提高并转移生产力与沟通活动的平衡","abstract":"Generative AI is transforming the workplace by augmenting and automating cognitive tasks, reshaping how organizations work and innovate while raising questions about workplace inequality and the future of work. Despite rapid adoption, empirical evidence on how these tools alter work practices and generate productivity gains remains limited. We examine how AI use affects the quantity and nature of information work using digital trace data from the Microsoft M365 application suite across multiple large international companies. Specifically, we study how generative AI adoption shifts the balance between communication and productivity-oriented activities, such as content creation in Word. Difference-in-Differences analyses show that AI adoption is associated with significant increases in both productivity (21.2%) and communication (7.1%) application actions among users who used the AI system more than 100 times over a 20-week post-adoption period. Among users with 100-500 AI use instances, higher AI usage is also associated with continued increases in both types of activity. The smaller increase in communication represents an overall shift toward individual, documentation-focused work and reflects mixed changes in communication, including decreases in reading and organizing email, compared with more uniform increases in productivity actions. These findings suggest potential efficiency gains and reductions in information overload, while highlighting the need to ensure that AI adoption does not weaken interpersonal communication and the diffusion of diverse information that supports innovation.","authors":["Yulin Yu","Yan Chen","Rui Hu","Siddharth Suri","Scott Counts"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-08-18","first_seen":"2026-08-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.15550","pdf_url":"https://arxiv.org/pdf/2608.15550","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C5"],"tags":["生成式AI采用","工作场所行为","实证研究"],"reason":"研究人类使用AI后的行为变化，非用LLM仿真人类被试，方向相反","model":"deepseek-v4-pro","scored_at":"2026-08-18T13:04:42","error":null,"has_summary":false,"summary":null},{"id":"2608.14605","version":1,"title":"Psychological Determinants of Academic Integrity in the Use of Generative AI in Higher Education","zh_title":"高等教育中使用生成式人工智能的学术诚信心理决定因素","abstract":"This paper examines the psychological determinants that shape academically honest and dishonest uses of generative artificial intelligence (GenAI) in higher education. Rather than treating academic misconduct as a purely technological problem, the study conceptualizes academic integrity as a psychologically mediated decision process influenced by moral reasoning, perceived social norms, policy clarity, academic self-efficacy, AI literacy, performance pressure, and beliefs about authorship. Methodologically, the paper adopts a focused narrative review and conceptual synthesis design. A purposive corpus of 16 core publications, including peer-reviewed studies and policy-oriented texts published between 2022 and March 2026, was assembled through targeted searches using combinations of the keywords generative AI, academic integrity, academic misconduct, moral disengagement, AI literacy, and higher education. The reviewed literature suggests that students do not interpret all forms of AI assistance as cheating. Integrity risk increases when institutional guidance is vague, peer use appears normalized, academic pressure is high, and AI tools are perceived as legitimate substitutes for difficult cognitive labor. By contrast, assignment-level guidance, explicit disclosure norms, ethics-oriented instruction, and authentic assessment design appear to reduce integrity risk more effectively than detection-centered responses alone. Based on these findings, the paper proposes an integrative conceptual model in which institutional context shapes psychological appraisal, and psychological appraisal in turn influences disclosed, borderline, or dishonest GenAI use. The paper concludes that effective responses to GenAI-related integrity problems should combine policy clarity, pedagogy, AI literacy, and student support rather than relying only on prohibition or software-based surveillance.","authors":["Ezgi Dagtekin","Ercan Erkalkan"],"categories":["cs.CY","cs.HC"],"primary_category":"cs.CY","announce_type":"cross","date":"2026-08-18","first_seen":"2026-08-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.14605","pdf_url":"https://arxiv.org/pdf/2608.14605","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C5"],"tags":["学术诚信","生成式AI","高等教育"],"reason":"研究人类使用GenAI的诚信心理，不涉及用LLM仿真人类被试，方向相反","model":"deepseek-v4-pro","scored_at":"2026-08-18T13:04:28","error":null,"has_summary":false,"summary":null},{"id":"2608.16323","version":1,"title":"Predicting, Evaluating, and Explaining Top Misinformation Spreaders via Archetypal User Behavior","zh_title":"通过原型用户行为预测、评估和解释顶级错误信息传播者","abstract":"The spread of misinformation on social networks poses a significant challenge to online communities and society at large. Not all users contribute equally to this phenomenon: a small number of highly effective individuals can exert outsized influence, amplifying false narratives and contributing to significant societal harm. This paper seeks to mitigate the spread of misinformation by enabling proactive interventions, identifying and ranking users according to key behavioral indicators associated with harmful content dissemination. We examine three user archetypes -- amplifiers, super-spreaders, and coordinated accounts -- each characterized by distinct behavioral patterns in the dissemination of misinformation. These are not mutually exclusive, and individual users may exhibit characteristics of multiple archetypes. We develop and evaluate several user ranking models, each aligned with a specific archetype, and find that super-spreader traits consistently dominate the top ranks among the most influential misinformation spreaders. As we move down the ranking, however, the interplay of multiple archetypes becomes more prominent. Additionally, we demonstrate the critical role of temporal dynamics in predictive performance, and introduce methods that reduce data requirements by minimizing the observation window needed for accurate forecasting. Finally, we demonstrate the utility and benefits of explainable AI (XAI) techniques, integrating multiple archetypal traits into a unified model to enhance interpretability and offer deeper insight into the key factors driving misinformation propagation. Our findings provide actionable tools for identifying potentially harmful users and guiding content moderation strategies, enabling platforms to monitor accounts of concern more effectively.","authors":["Enrico Verdolotti","Luca Luceri","Silvia Giordano"],"categories":["cs.SI","cs.CY","cs.LG"],"primary_category":"cs.SI","announce_type":"cross","date":"2026-08-18","first_seen":"2026-08-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.16323","pdf_url":"https://arxiv.org/pdf/2608.16323","source_feed":"cs.CY","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["错误信息传播","用户行为分析","社交网络"],"reason":"研究社交媒体用户行为，未使用LLM仿真人类被试，不涉及人类数据对照。","model":"deepseek-v4-pro","scored_at":"2026-08-18T13:04:49","error":null,"has_summary":false,"summary":null},{"id":"2608.14663","version":1,"title":"In-Context Learning to Assess Built Environment Impacts on Perceived Neighborhood Walkability Among Mobility-impaired Older Adults","zh_title":"上下文学习评估建成环境对行动不便老年人感知邻里步行性的影响","abstract":"As global populations age, enhancing neighborhood walkability through inclusive urban design is important for mitigating built environment (BE) barriers that discourage physical activity and social participation among older adults. This study investigates the utility of in-context learning (ICL), using the transformer-based foundation model TabPFN, to determine how BE features influence perceived walkability, as measured by the Neighborhood Environment Walkability Scale (NEWS-A) survey. Using a small-scale dataset (N = 257) comprising a unique demographic of older adults with knee osteoarthritis or a history of falls, TabPFN achieved a macro F1 score of 54.89% for walkability perceptions categorized as Low, Neutral, and High using equal-width binning. This result outperformed optimized, grid-searched baseline models, including Random Forest (45.85%) and XGBoost (50.56%). To interpret these results, we employed Shapley Interaction Quantification (SHAP-IQ) to identify the hierarchical importance of feature interactions. Preliminary results revealed that the model's predictive logic was primarily driven by higher-order interactions. For example, the interaction between average street circuity and the ratio of drivable roads emerged as the primary discriminator of perceived walkability. Neighborhood greenery was found to have substantial predictive importance only when combined with an individual's fear of falling or perception of age-friendliness. Overall, ICL using TabPFN demonstrates superior performance on small-scale datasets, enhancing the fidelity of the resulting interpretive insights. Furthermore, SHAP-IQ provides a synergistic perspective on how higher-order feature interactions drive the model's predictions.","authors":["Houhao Liang","Kresimir Friganovic","Joanne Kua","Noor Hafizah Ismail","Su Su","Bryan Yijia Tan","Navrag B. Singh","Panos Mavros"],"categories":["cs.LG","stat.AP"],"primary_category":"cs.LG","announce_type":"new","date":"2026-08-18","first_seen":"2026-08-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.14663","pdf_url":"https://arxiv.org/pdf/2608.14663","source_feed":"cs.LG","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["表格预测","可解释性","建成环境"],"reason":"论文用TabPFN做表格预测，非LLM仿真人类被试，无人类行为对照，属纯NLP…","model":"deepseek-v4-pro","scored_at":"2026-08-18T13:04:32","error":null,"has_summary":false,"summary":null},{"id":"2608.14956","version":1,"title":"LLM-based Framework for Generating and Verifying Parallel DEVS Statecharts","zh_title":"基于LLM的并行DEVS状态图生成与验证框架","abstract":"The development of models demands sound modeling and simulation knowledge as well as domain knowledge. Every model should accurately represent a system's dynamics and be verifiable. Toward this objective, this research introduces an agentic PDEVS-LLM framework to assist human modelers in generating and verifying PDEVS statecharts for behavior modeling of atomic Parallel Discrete Event System Specification (PDEVS) models. The framework supports (re)generating plausible facts from a system description prompt using the agentic LLM used for generating plausible facts. Inconsistencies in plausible facts lead to incorrect PDEVS statecharts having logical structure and behavioral inaccuracies. A controlled-correction mechanism is developed to verify the logical consistency of the plausible facts. The agentic LLM is used to generate key behavioral conditions from the system description prompt. The plausible facts are then verified against the behavioral conditions using propositional logic entailment for a finite number of times. The verification results enable the generation of modification prompts that can reduce errors in generated plausible facts, resulting in more accurate PDEVS statecharts. To verify a statechart's logical correctness, its Timed Automata counterpart is manually created and verified for deadlock and reachability properties. The human modeler may regenerate plausible facts and PDEVS statecharts iteratively and incrementally. A basic correctness metric is introduced to quantify the completeness and accuracy of the expected behavioral traits of the PDEVS statechart models. A collection of example systems with varying levels of complexity is developed to demonstrate the capabilities and limitations of LLMs. The evaluation of the proposed verification mechanism shows a substantial improvement in the logical consistency of generated statecharts.","authors":["Vamsi Krishna Vasa","Hessam S. Sarjoughian","Edward J. Yellig"],"categories":["cs.LG","cs.LO"],"primary_category":"cs.LG","announce_type":"new","date":"2026-08-18","first_seen":"2026-08-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.14956","pdf_url":"https://arxiv.org/pdf/2608.14956","source_feed":"cs.LG","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["LLM辅助建模","形式化验证","多智能体系统"],"reason":"纯多智能体系统研究，LLM用于生成和验证DEVS状态图，不涉及人类行为仿真或对…","model":"deepseek-v4-pro","scored_at":"2026-08-19T13:03:27","error":null,"has_summary":false,"summary":null},{"id":"2608.15507","version":1,"title":"Do Language Models Consistently Encode the Current Year?","zh_title":"语言模型是否一致地编码当前年份？","abstract":"A consistent concept of the current time is important for temporal reasoning, yet how language models represent the current time is not well understood. We contribute two tasks that probe the current year in conceptually distinct ways: an associative task, which infers the current year from verb tense, and a declarative task, which directly queries for the current year. Both tasks estimate current years within one year of the post-training data cutoff of instruction-tuned language models. For base models, predictions on the associative task serve as a strong proxy for the pre-training data cutoff, with an average error of only 10 months across 13 models. However, their internal mechanisms diverge: the associative task uses mechanisms similar to factual recall, while the declarative task lacks consistent causal pathways. This divergence poses a challenge for updating the current year in language models. None of prompting, SFT, or weight editing succeed in shifting the associative and declarative years simultaneously. Prompting updates the declarative year (94.6% success across 351 target years) but leaves the associative year nearly unchanged (1.7% success). Year-shifted SFT also fails to shift the associative year, matching the target year in only one of eight models. Weight editing, while effective for both tasks individually, does not generalize across both. Overall, our results show that the current year is not consistently encoded in language models: The associative notion, deeply ingrained in linguistic structures learned in pre-training, uses different causal mechanisms and resists the same modifications that easily shift the declarative notion learned in post-training.","authors":["Suze van Adrichem","Aditi Bhaskar","Diyi Yang","Christopher Potts","Jing Huang"],"categories":["cs.CL","cs.LG"],"primary_category":"cs.CL","announce_type":"cross","date":"2026-08-18","first_seen":"2026-08-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.15507","pdf_url":"https://arxiv.org/pdf/2608.15507","source_feed":"cs.LG","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["时间推理","模型评测","因果机制"],"reason":"研究LLM对当前年份的编码，属模型能力评测，不以人类行为为参照，不涉及仿真人类…","model":"deepseek-v4-pro","scored_at":"2026-08-19T13:03:29","error":null,"has_summary":false,"summary":null},{"id":"2608.14079","version":1,"title":"The conditional superiority of fast silicon sampling","zh_title":"快速硅采样的条件优越性","abstract":"Silicon sampling can produce surprisingly good population estimates at times. Does doing it fast attenuate such fidelity? In this study, we extend and assess ongoing work in silicon sampling by comparing the algorithmic fidelity of \"fast\" and \"slow\" modes of silicon sampling among a nationally representative sample of Singaporean survey respondents. We find that silicon sampling with contemporary frontier models remains a method in early development to be used only with great caution. While silicon samples are able to produce moderately faithful estimates of population means, they continue to understate opinion variance and distort the latent contextual space behind human opinions. Conditional on such limitations, we find \"fast\" modes of silicon sampling to be relatively superior to traditional \"slow\" modes of silicon sampling. Fast silicon sampling is significantly more efficient in compute resources and run-time while being monotonically superior to slower modes of sampling in algorithmic fidelity.","authors":["Nickolas Hock Yuen Lam","Ji Xuan Voo","Xiangyu Ma"],"categories":["cs.CL","cond-mat.mtrl-sci"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-17","first_seen":"2026-08-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.14079","pdf_url":"https://arxiv.org/pdf/2608.14079","source_feed":"cs.CL","score":10,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["硅采样","算法保真度","人类仿真"],"reason":"直接比较快速与慢速硅采样在代表性样本上的算法保真度，含真实人类数据对照，并指出…","model":"deepseek-v4-pro","scored_at":"2026-08-17T13:01:19","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-18","rank":1,"question":"快速硅采样是否在算法保真度上不劣于甚至优于传统慢速硅采样？","design":"使用当代前沿模型（OpenAI GPT 5.4）对新加坡全国代表性调查受访者进行硅采样，比较“快速”模式（一次提示生成大量响应）与“慢速”模式（逐个API调用生成响应）的算法保真度，测量结果包括总体均值估计、意见方差和潜在情境空间结构。","baseline":"新加坡全国代表性调查受访者的真实人类数据。","findings":"硅采样能中等程度地忠实估计总体均值，但低估意见方差并扭曲潜在情境空间。在承认这些局限的前提下，快速硅采样在算法保真度上单调优于慢速硅采样，且计算资源和运行时间显著更高效。","reliability":"论文承认硅采样仍处于早期发展阶段，需谨慎使用；硅样本低估意见方差、扭曲潜在情境空间，且快速与慢速模式均存在这些局限。","relevance":"该研究直接比较快速与慢速硅采样在代表性样本上的算法保真度，含真实人类数据对照，并指出硅采样在方差和潜在空间上的失效条件，对关注LLM仿真可靠性与偏差的研究者具有参考价值。","inspiration":"借鉴其通过几何数据分析（多重对应分析）评估关系保真度的方法，以及比较不同采样模式效率与保真度的设计。｜可迁移到政策评估中的公众意见模拟，如经济政策公告的预期形成或消费者信心调查。｜以LLM生成不同处理模式下的合成受访者，处理为快速与慢速采样，结果变量为对经济政策的态度分布，对照真实调查数据（如新加坡消费者信心指数），评估均值、方差和潜在空间结构的一致性。"}},{"id":"2608.10492","version":2,"title":"INSIDE the Student's Mind: Jointly Modeling Latent Reasoning and Action in LLM Student Simulators","zh_title":"洞察学生思维：联合建模LLM学生模拟器中的潜在推理与行为","abstract":"Large Language Model (LLM)-based simulators often reproduce observable actions but fail to capture the underlying reasoning behind them. In education, where student simulation is increasingly used for various applications such as evaluating tutoring systems, this gap is especially pronounced. Two students may submit identical submissions for entirely different reasons. We present INTERNAL STUDENT DIALOGUE (INSIDE), a student modeling framework that fine-tunes LLMs not only to act like students but also to think like them. INSIDE generates internal dialogue grounded in Bloom's Taxonomy across cognitive, affective, and action dimensions, and fine-tunes models on paired think traces and actions. We baseline against different prompting frameworks and evaluate on two axes: fidelity of simulated actions and quality of generated internal dialogue. Our evaluations show that INSIDE improves simulation fidelity in both action fidelity, matching code generation of real students, and reasoning alignment, achieving the highest alignment across models up to 57.9%.","authors":["Rose Niousha","Minwoo Kang","Narges Norouzi"],"categories":["cs.AI","cs.CY"],"primary_category":"cs.AI","announce_type":"replace","date":"2026-08-17","first_seen":"2026-08-12","revised_at":"2026-08-17","abs_url":"https://arxiv.org/abs/2608.10492","pdf_url":"https://arxiv.org/pdf/2608.10492","source_feed":"cs.AI","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","教育模拟","算法保真度"],"reason":"用LLM仿真学生行为与推理，并与真实学生数据对照，评估仿真保真度。","model":"deepseek-v4-pro","scored_at":"2026-08-17T13:01:36","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-17","rank":3,"question":"如何让LLM学生模拟器不仅复现学生的可观察行为（代码提交），还能捕捉其背后的潜在推理过程，从而提高仿真保真度？","design":"提出INSIDE框架，基于Bloom分类法生成内部对话（认知、情感、行动维度），并用配对的思想痕迹和行动对LLM进行微调。模型扮演编程课程学生，输入学生历史提交和AI导师反馈，输出下一步代码提交。评估两个维度：行动保真度（生成代码与真实学生代码的相似度）和推理质量（生成推理与真实代码编辑的对齐度）。","baseline":"使用加州大学伯克利分校入门编程课程两个学期的真实学生数据：Spring 2025用于训练（445名学生，2022条提交流，6911次提交），Spring 2024用于测试（479名学生，1546条提交流，6316次提交）。测试集分为旧问题（test_OP）和新问题（test_NP），分别评估对未见学生和未见问题的泛化。","findings":"INSIDE提高了行动保真度，生成代码与真实学生代码的Wasserstein距离更低；同时实现了最高的推理对齐度，在不同模型上最高达到57.9%。","reliability":"论文未讨论","relevance":"该研究直接命中你的核心关注点：用LLM仿真学生行为与推理，并与真实学生数据对照，评估仿真保真度。它提供了在缺少真实推理标注的情况下重建潜在推理的方法，并展示了联合建模推理与行动能提升仿真质量，值得精读原文。","inspiration":"借鉴其用理论框架（Bloom分类法）引导内部对话生成，并将推理与行动联合微调以提升仿真保真度的做法。｜可迁移到经济金融中的政策预期形成研究，例如模拟投资者在信息发布后的决策过程，捕捉其推理路径。｜以LLM扮演投资者，输入历史交易和新闻信息，生成内部推理（如风险评估、情绪反应）和交易决策，用真实市场交易数据和调查数据（如投资者信心指数）作为对照，评估仿真保真度。"}},{"id":"2601.20238","version":2,"title":"Large Language Models Polarize Ideologically but Moderate Affectively in Online Political Discourse","zh_title":"大语言模型在网络政治话语中加剧意识形态极化但缓和情感极化","abstract":"The emergence of large language models (LLMs) is reshaping how people engage in political discourse online. We examine how the release of ChatGPT altered ideological and emotional patterns in Reddit's largest political forum. Analysis of millions of comments shows that ChatGPT intensified ideological polarization: liberal-leaning authors posted increasingly liberal comments, while conservative-leaning authors posted increasingly conservative comments. Multiple falsification tests suggest that these findings are unlikely to be driven by contemporaneous events, such as the 2022 U.S. midterm elections, or by broader platform-wide trends in political polarization. Mechanism tests show that this shift does not stem from the creation of more persuasive or ideologically extreme original content using LLM. Instead, it originates from the tendency of LLM-assisted comments to echo and reinforce the original post's viewpoint, a pattern consistent with algorithmic sycophancy. Yet, despite growing ideological divides, affective polarization, measured by hostility and toxicity, declined. These findings reveal that LLMs can simultaneously deepen ideological separation and foster more civil exchanges, challenging the long-standing assumption in literature that extremity and incivility necessarily move together.","authors":["Gavin Wang","Srinaath Anbudurai","Oliver Sun","Xitong Li","Lynn Wu"],"categories":["econ.GN","q-fin.EC"],"primary_category":"econ.GN","announce_type":"replace","date":"2026-08-17","first_seen":"2026-01-28","revised_at":"2026-08-17","abs_url":"https://arxiv.org/abs/2601.20238","pdf_url":"https://arxiv.org/pdf/2601.20238","source_feed":"econ.GN","score":8,"bucket":"selected","rubric_hits":["A3","B1","B2","B4"],"tags":["LLM仿真","政治极化","人类数据对照"],"reason":"用LLM辅助评论与真实Reddit数据对照，分析政治话语极化，涉及社会过程仿真…","model":"deepseek-v4-pro","scored_at":"2026-08-17T13:01:35","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-17","rank":4,"question":"ChatGPT的发布如何影响Reddit政治论坛中用户的意识形态极化和情感极化？","design":"本研究并非将LLM作为人类被试的仿真实验，而是利用Reddit上最大的政治论坛在ChatGPT发布前后的数百万条评论数据，通过语言特征指标（如评论长度、被动动词、困惑度、人类评分）估计每条评论由LLM辅助生成的概率，进而分析LLM使用对用户意识形态立场和情感表达的影响。","baseline":"以ChatGPT发布前的Reddit评论作为对照，比较同一作者在发布前后的意识形态立场变化，并利用2022年美国中期选举等事件进行证伪检验。","findings":"ChatGPT的发布加剧了意识形态极化：自由派作者发表更自由的评论，保守派作者发表更保守的评论。然而，情感极化（敌意和毒性）却有所下降，表明LLM在加深意识形态分歧的同时促进了更文明的交流。","reliability":"论文承认无法直接确定每条评论是否由LLM生成，而是依赖语言指标估计概率，可能存在分类误差；但认为在大规模语料中，误差不会系统性地产生所观察到的聚合模式。","relevance":"该研究利用真实Reddit数据评估LLM对政治话语的影响，涉及LLM辅助内容生成与真实人类行为的对照，对关注LLM在社会科学中仿真可靠性的研究者有参考价值，但并非直接以LLM作为被试的仿真实验。","inspiration":"值得借鉴的是利用自然实验（ChatGPT发布）和语言特征指标来估计LLM使用概率，并设置证伪检验排除混淆因素。｜可迁移到经济金融领域如政策公告后的市场情绪分析、消费者评论中的LLM影响等场景。｜一个可行的设计是：以某经济政策发布为时间节点，收集社交媒体上相关讨论，用语言指标估计LLM辅助评论比例，分析其对情绪极化和观点极化的影响，并与历史人类评论基线对照。"}},{"id":"2608.13712","version":1,"title":"Reading Between The Lines: Modeling and Evaluating Behavioral Realism in Legal Simulation","zh_title":"字里行间：法律模拟中行为真实性的建模与评估","abstract":"Deposition training requires attorneys to manage dynamic witness behavior, yet legal-AI evaluations largely focus on factual accuracy, reasoning, or response-level plausibility. We introduce WitnessSim, a deposition simulator driven by controllable legal personas. We use an evaluation framework separating behavioral realism from pedagogical usefulness. We assess realism through adversarial testing, blinded attorney comparison, and analysis of longitudinal behavioral trajectories. WitnessSim generally maintained plausible behavioral boundaries, and attorneys did not systematically prefer either original testimony or WitnessSim generated testimony. Pedagogical tests showed that witness behavior changed meaningfully in response to question form and attorney intervention without uniformly collapsing the assigned persona. Together, these results showcase a model of behavioral fidelity in legal simulations, and provide a framework for evaluating its performance.","authors":["Divya Vetticaden","Arya Gupta","Julian Nyarko","Megan Ma"],"categories":["cs.CY","cs.AI","cs.CL"],"primary_category":"cs.CY","announce_type":"cross","date":"2026-08-17","first_seen":"2026-08-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.13712","pdf_url":"https://arxiv.org/pdf/2608.13712","source_feed":"cs.CL","score":8,"bucket":"selected","rubric_hits":["A1","A2","B1","B2"],"tags":["LLM仿真","法律模拟","行为真实性"],"reason":"用LLM模拟法律证人行为，并与真实证词对照，评估行为真实性和教学效果，属于人类…","model":"deepseek-v4-pro","scored_at":"2026-08-17T13:01:17","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-18","rank":6,"question":"如何构建并评估一个具有行为真实性和教学实用性的法律证人模拟系统？","design":"WitnessSim 是一个基于六维状态向量（镇定、知识、宜人性、冗长度、僵化度、表现）的沉积证词模拟器，通过问题特征（压力、话题敏感性、问题形式）更新状态，并条件化生成证词；研究使用对抗测试、盲法律师比较和纵向行为轨迹分析评估行为真实性，并通过法律培训材料衍生的测试评估教学实用性。","baseline":"来自国家处方阿片类药物诉讼（Case No. 1:17-MD-2804）的 300 份沉积和庭审笔录，作为真实人类证词对照。","findings":"WitnessSim 在对抗测试中维持了合理的行为边界，律师在盲法比较中未系统偏好原始证词或生成证词。教学测试显示证人行为对问题形式和律师干预有有意义的变化，且未完全丧失指定人格。","reliability":"论文未明确讨论失效条件与局限，但指出评估框架区分行为真实性与教学实用性，并承认行为真实性通过对抗测试、专家评估和情感轨迹分析来操作化，可能仍存在未覆盖的维度。","relevance":"该研究直接命中研究者关注的 LLM 人类仿真实验，提供了法律场景下与真实人类数据对照的行为真实性评估框架，值得阅读原文以借鉴其多维评估方法和动态行为建模。","inspiration":"借鉴其将行为状态建模为可更新的多维向量，并通过问题特征施加处理、以真实行为轨迹为基准的评估方法。｜可迁移到经济金融中的谈判、审计或客户服务交互模拟，如信贷审批中的申请人行为或政策沟通中的公众反应。｜以 LLM 模拟信贷申请人，处理变量为审批官提问的侵略性或信息敏感度，结果变量为申请人的情绪状态和回答一致性，对照真实信贷申请面谈记录。"}},{"id":"2608.13786","version":1,"title":"Do AI chatbots find what experts would? Effects of model, user role, and sample size on study retrieval for medical questions","zh_title":"AI聊天机器人能否找到专家会找到的研究？模型、用户角色和样本量对医学问题研究检索的影响","abstract":"Large language model (LLM) chatbots are increasingly used to answer clinical questions with citations to relevant clinical studies. Prior research has largely focused on citation fabrication, leaving a gap in evaluating the quality of retrieved studies and the factors driving their selection. In this study, we evaluated three general-purpose LLM chatbots: Claude Sonnet 5, Gemini 3.1 Pro, and ChatGPT GPT-5.5. We prompted the models with clinical questions adapted from 20 review questions in Issues 6 and 7 of the 2026 Cochrane Database of Systematic Reviews, simulating patient, clinician, and evidence-synthesis researcher roles. Each chatbot was queried under each user role with four independent repetitions, yielding 720 responses. Each chatbot was asked to support its answers with primary clinical citations, which we benchmarked against the included and excluded study sets of the Cochrane reviews. On average, a chatbot response retrieved 39.2% $\\pm$ 29.8% of Cochrane included studies, while citing 5.0% $\\pm$ 9.4% of excluded studies. Recall of Cochrane included studies varied significantly by model and user role. ChatGPT achieved higher recall than Claude or Gemini (63.1% $\\pm$ 29.5% vs. 37.0% $\\pm$ 23.8% vs. 17.3% $\\pm$ 13.1%; $p=2.0\\times10^{-5}$). The researcher role yielded higher recall than the clinician or patient roles (42.8% $\\pm$ 30.8% vs. 38.6% $\\pm$ 28.9% vs. 36.1% $\\pm$ 29.3%; $p=2.0\\times10^{-5}$). Controlling for publication year, citations per year, and open-access status, sample size was the only independently significant predictor of retrieval (odds ratio 1.80 per 1-unit increase in log sample size, 95% CI 1.37-2.36, $p=2.34\\times10^{-5}$). These findings suggest that while LLM chatbots can retrieve some studies identified by expert reviewers, their performance varies by model and user role, and they exhibit a bias toward clinical trials with larger sample sizes.","authors":["Qingfang Liu","Qiao Jin","Joe D. Menke","Thorsten Kahnt","Zhiyong Lu"],"categories":["cs.IR","cs.AI","cs.CL"],"primary_category":"cs.IR","announce_type":"cross","date":"2026-08-17","first_seen":"2026-08-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.13786","pdf_url":"https://arxiv.org/pdf/2608.13786","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A1","B1","B2"],"tags":["LLM仿真","医学信息检索","人类对照"],"reason":"用LLM模拟不同用户角色检索医学证据，并与Cochrane专家评审结果对照，属…","model":"deepseek-v4-pro","scored_at":"2026-08-17T13:01:25","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-18","rank":11,"question":"不同大语言模型聊天机器人在模拟患者、临床医生和证据综合研究者三种用户角色时，检索医学问题相关临床研究的表现如何，以及哪些因素影响其检索结果？","design":"用三个通用大语言模型（Claude Sonnet 5、Gemini 3.1 Pro、ChatGPT GPT-5.5）模拟患者、临床医生和证据综合研究者三种角色，对20个Cochrane系统综述问题生成带引文的回答，每个角色-模型组合重复4次，共720个回答，以Cochrane综述纳入和排除的研究集为基准，测量检索召回率和引用排除研究的比例。","baseline":"以20个Cochrane系统综述中专家评审纳入和排除的研究集作为真实人类专家判断的基准。","findings":"平均每个回答检索到39.2%的Cochrane纳入研究，引用5.0%的排除研究；ChatGPT召回率显著高于Claude和Gemini，研究者角色召回率显著高于临床医生和患者角色。控制发表年份、年均引用和开放获取状态后，样本量是唯一显著预测检索的因素，模型偏向于检索大样本临床试验。","reliability":"论文承认其评估仅限于三个特定模型和20个Cochrane综述主题，可能不具普遍性；未讨论模型版本更新、提示词细微变化或不同医学领域对结果的影响。","relevance":"该研究直接以LLM模拟不同用户角色进行信息检索，并与专家评审结果对照，属于人类仿真实验，且揭示了模型和角色对检索行为的影响，值得阅读原文以了解仿真偏差的具体表现。","inspiration":"借鉴其通过角色扮演和重复查询来测量LLM行为差异，并利用专家评审数据作为基准的方法｜可迁移到经济金融领域，如模拟投资者、分析师和监管者角色检索金融研究报告或政策文件，检验信息获取偏差｜用LLM扮演不同金融角色，对特定经济问题（如某行业前景）检索相关研究，以权威机构（如央行工作论文或顶级期刊）的文献列表为基准，测量检索召回率和偏差，并分析文献特征（如样本量、发表期刊）对检索的影响。"}},{"id":"2608.14320","version":1,"title":"AnchorBench: A Multi-Pathway Benchmark for the Anchoring Effect in LLMs","zh_title":"AnchorBench：LLM锚定效应的多路径基准测试","abstract":"The anchoring effect is a cognitive bias in which an initial reference value shifts a later judgment toward itself. This effect is well established in human judgment and decision-making, and recent work suggests that large language models (LLMs) exhibit similar behavior. However, existing work on anchoring in LLMs typically evaluates only a narrow set of anchor pathways and rarely distinguishes irrelevant from plausible anchors. We introduce AnchorBench, a benchmark for the anchoring effect in LLMs that evaluates multiple anchor pathways under an explicit anchor relevance axis. Across fourteen models, including ten open-weight models and four frontier API models, and a large set of controlled prompts, we find that (1) anchoring is strongly pathway-dependent, (2) plausible anchors usually induce larger shifts than irrelevant ones when introduced through stronger pathways, (3) anchor influence generally weakens as the anchor moves farther from the evidence-supported answer, most clearly on External and RAG, and (4) high task accuracy on the anchor-free control condition (Acc$_{10}$: answers within 10 points of gold) does not guarantee robustness: even frontier API models above 95% control accuracy remain susceptible to plausible anchors.","authors":["Yiderigun Borjigin","Alexander Hermann","Christian Cyron","Roland Aydin"],"categories":["cs.AI","cs.CL"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-08-17","first_seen":"2026-08-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.14320","pdf_url":"https://arxiv.org/pdf/2608.14320","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A2","B4"],"tags":["认知偏差","LLM评估","锚定效应"],"reason":"评估LLM锚定效应，与人类认知偏差对照，可迁移到仿真可靠性研究","model":"deepseek-v4-pro","scored_at":"2026-08-17T13:01:19","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-18","rank":12,"question":"LLM在多种锚定信息传递路径下是否表现出锚定效应，且该效应如何随锚定相关性和路径强度变化？","design":"构建AnchorBench基准，使用14个LLM（10个开源权重、4个API模型）完成0-100数值判断任务，通过五种路径（外部提示、对话历史、上下文学习、检索增强、工具输出）施加低/高锚定值，并设置无关与合理锚定条件，测量输出相对无锚定控制条件的偏移。","baseline":"无对照","findings":"锚定效应强烈依赖路径，合理锚定在强路径下比无关锚定引起更大偏移；锚定影响随锚定值与证据支持答案距离增大而减弱，且高控制准确率不保证稳健性，即使API模型也易受合理锚定影响。","reliability":"论文未讨论","relevance":"该研究系统评估LLM锚定效应，与人类认知偏差对照，可迁移到仿真可靠性研究，值得阅读原文以了解多路径设计和相关性区分方法。","inspiration":"借鉴其多路径施加处理和相关性轴设计，区分无关与合理锚定，测量偏移量而非仅准确率｜可迁移到资产定价实验或政策公告的预期形成研究，检验LLM对锚定信息的敏感性｜以LLM模拟投资者，在提示中通过不同路径（如新闻、历史对话）注入锚定价格，测量估值偏移，并与人类实验数据对照。"}},{"id":"2608.14399","version":1,"title":"Whose doctor does the AI recommend? An algorithm audit of reputation and demographic signals in large language model-assisted physician choice","zh_title":"AI推荐哪位医生？大语言模型辅助医生选择中声誉与人口统计信号的算法审计","abstract":"Patients increasingly ask large language model (LLM) assistants which doctor to see, making these systems AI infomediaries: algorithms that intermediate one person's choice among other people and thereby decide, silently and at scale, which physicians become visible. We report a prespecified randomized algorithm audit of what causally moves those recommendations. Seven models (six open-weight; gpt-4o-mini) each chose among five synthetic family-medicine physician cards whose attributes were independently randomized across 3,024 choice sets, three patient personas, nine prompt paraphrases and nine experimental arms, yielding 40,068 scored responses; gender and ethnicity were signaled through names following correspondence-audit methodology. Reputation signals dominate: raising a rating from 3.9 to 4.7 increases choice probability by 31.4 percentage points (pp), and raising the fee from $90 to $190 lowers it by 20.0 pp. Demographic parity is rejected, but not in the direction human audit studies predict: female-signaled names gain 2.5 pp, and Hispanic-, South-Asian- and Black-signaled names gain 1.3-2.9 pp over White-signaled names, tilts worth $7-$14 per visit in fee-equivalent terms, and a content-free first-listed position is worth $11. Yet models mentioned gender or ethnicity in at most 0.03% of their stated reasons and abstained in 0.39% of trials, so these effects are invisible in the models' own explanations, and transparency obligations relying on model self-report would not detect them. One reasoning model failed the prespecified auditability gate outright. The frozen design makes the audit repeatable: any new model can be assessed against identical stimuli, making recurring behavioural audit, rather than self-reported explanation, the monitoring technology fit for purpose.","authors":["Syeda Anshrah Gillani","Mirza Samad Ahmed Baig"],"categories":["cs.CY","cs.AI","cs.CL"],"primary_category":"cs.CY","announce_type":"cross","date":"2026-08-17","first_seen":"2026-08-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.14399","pdf_url":"https://arxiv.org/pdf/2608.14399","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A1","B1","B2","B4"],"tags":["LLM仿真","算法审计","医疗决策"],"reason":"用LLM模拟患者选择医生，与真实人类审计研究对照，揭示偏差与失效条件","model":"deepseek-v4-pro","scored_at":"2026-08-17T13:01:19","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-18","rank":13,"question":"在LLM辅助的医生推荐中，声誉信号、姓名暗示的人口统计特征和列表位置如何因果地影响推荐概率？","design":"用七个LLM（六个开源权重模型和gpt-4o-mini）扮演患者，在随机选择联合实验中从五张合成家庭医生卡片中选择医生；卡片属性独立随机化，包括患者评分、评论量、费用、性别和种族（通过姓名暗示）等；结果变量是医生被选择的概率。","baseline":"与人类审计研究对照，特别是关于医生选择中反少数族裔歧视的发现。","findings":"声誉信号主导推荐：评分从3.9提高到4.7使选择概率增加31.4个百分点，费用从90美元提高到190美元降低20.0个百分点。人口统计均等被拒绝，但方向与人类审计研究相反：女性姓名增加2.5个百分点，西班牙裔、南亚裔和黑人姓名比白人姓名增加1.3-2.9个百分点，相当于每次就诊7-14美元的费用等值。","reliability":"模型在陈述理由中提到性别或种族的比例最多为0.03%，在0.39%的试验中弃权，因此这些效应在模型自身解释中不可见；一个推理模型（deepseek-r1:7b）未能通过预设的可审计性门槛。","relevance":"该研究用LLM模拟患者选择医生，与真实人类审计研究对照，揭示偏差与失效条件，直接回应了研究者对LLM人类仿真实验和批判性评估的兴趣。","inspiration":"借鉴随机选择联合实验和姓名信号法，在LLM仿真中独立操纵属性以识别因果效应，并用费用等值量化效应大小。｜可迁移到信贷审批歧视研究，用LLM模拟信贷员审批贷款，操纵申请人姓名、收入、信用评分等属性。｜用多个LLM作为被试，随机生成贷款申请档案，处理变量为申请人姓名暗示的种族和性别，结果变量为批准概率，与真实信贷审批数据中的歧视模式对照。"}},{"id":"2608.14113","version":1,"title":"Search or Chat? Comparing How We Learn About Debated Topics","zh_title":"搜索还是聊天？比较我们如何了解有争议的话题","abstract":"As large language models (LLMs) become more integrated into everyday information platforms, chat-based systems are emerging as a popular alternative to traditional web searches, especially for informational search and informal learning tasks. Despite this shift, little is known about how different tools affect learning outcomes. Our work aims to improve the understanding of how chat-based information access supports and impacts learning performance in informal learning settings. In this paper, we present the results of a crowdsourcing user study (N = 194) that compares learning about debated topics using a traditional search interface versus an LLM-powered chat interface. Through our analysis of learning outcomes, user characteristics, and interaction patterns, we found no significant differences in user learning gain or critical reflection on our study tasks. Our observations from the analysis of further exploratory variables suggest that, in the context of longstanding debated topics, user characteristics such as their attitude strength and level of intellectual humility might be more important in shaping immediate learning outcomes than the information access tool.","authors":["Ran Yu","Alisa Rieger","Rabia Karatoprak Ersen","Jiqun Liu"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-08-17","first_seen":"2026-08-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.14113","pdf_url":"https://arxiv.org/pdf/2608.14113","source_feed":"cs.HC","score":7,"bucket":"pending","rubric_hits":["A1","B1"],"tags":["LLM聊天界面","学习效果","用户研究"],"reason":"用LLM聊天界面替代搜索，比较学习效果，有真实用户数据对照，但非直接仿真人类被…","model":"deepseek-v4-pro","scored_at":"2026-08-17T13:01:19","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-17","rank":7,"question":"在非正式学习场景下，传统搜索界面与LLM聊天界面在争议性话题学习效果上是否存在差异？","design":"本研究不是LLM仿真人类研究，而是一项众包用户实验（N=194），比较真实人类被试在两种信息获取工具（传统搜索界面 vs. LLM聊天界面）下的学习效果。被试被随机分配一个争议性话题和一种工具，在限定时间内自主学习，之后测量学习增益（论点扩展）和批判性反思（批判性推理），并记录交互日志。","baseline":"无对照（本研究未使用LLM仿真人类，而是直接比较两种工具对真实人类学习的影响）","findings":"在争议性话题上，使用聊天机器人与使用搜索引擎的被试在论点扩展和批判性推理上无显著差异。探索性分析表明，用户的态度强度和智识谦逊水平比信息获取工具更能预测即时学习效果。","reliability":"论文未讨论（但研究本身并非仿真研究，其局限可能包括：仅测量即时学习效果，未追踪长期保持；争议性话题选择有限；众包被试可能缺乏深度参与动机；未控制工具使用时间差异等）","relevance":"该研究与研究者关注点部分相关：它比较了LLM聊天界面与传统搜索对人类学习的影响，但并非用LLM仿真人类被试，而是将LLM作为信息工具。对于关心LLM在信息获取中作用的研究者有一定参考价值，但若聚焦于LLM仿真人类行为，则相关度有限。","inspiration":"可借鉴其随机对照实验设计，将信息获取工具作为处理变量，测量认知结果（如学习增益、批判性思维），并考察用户特征（如态度强度、智识谦逊）的调节作用。｜可迁移到经济金融领域的信息处理与决策场景，例如投资者使用不同信息工具（搜索引擎 vs. LLM聊天机器人）获取公司财报信息后，其投资决策质量、信息理解深度或过度自信程度的变化。｜设计一个实验：招募真实投资者作为被试，随机分配使用搜索引擎或LLM聊天机器人获取某上市公司财报信息，之后测量其对公司的估值准确性、信息回忆和投资信心，并与历史市场数据或分析师预测作为基准对照，考察工具类型对决策质量的影响。"}},{"id":"2608.05246","version":2,"title":"LUNAR: Benchmarking Personalized Large Language Models on UNiversal User BehAvioR Logs","zh_title":"LUNAR：基于通用用户行为日志的个性化大语言模型基准","abstract":"Existing personalized LLM benchmarks primarily rely on textual personas or isolated behavioral signals, providing limited evaluation of cross-domain behavioral personalization, where responses must be grounded in heterogeneous daily-life activities. To address this gap, we introduce LUNAR, the first benchmark for evaluating how LLMs personalize responses from longitudinal app interaction histories across universal daily-life domains, including clothing, food, housing, and mobility. To support scalable benchmark construction while mitigating data sparsity and privacy concerns, LUNAR uses a multi-stage coarse-to-fine synthesis pipeline grounded in real-world behavioral patterns. Fidelity analyses show closer alignment with real behavioral distributions than other synthetic benchmarks. Experiments on 19 mainstream LLMs show that access to behavioral logs is necessary but not sufficient for deep personalization: neither more context nor larger models guarantees better performance; effective personalization depends on selecting and integrating relevant evidence across domains. Direct retrieval of fine-grained behavioral records consistently outperforms compressed memory, while stronger personalization can come at the cost of privacy protection. These findings identify evidence selection, cross-domain integration, and privacy control as key challenges for personalized LLMs.","authors":["Jiahao Zhang","Yongzhi Tong","Zelin Fu","Pengde Zhao","Yanmei Jiang","Feng Jiang","Min Yang"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"replace","date":"2026-08-17","first_seen":"2026-08-07","revised_at":"2026-08-17","abs_url":"https://arxiv.org/abs/2608.05246","pdf_url":"https://arxiv.org/pdf/2608.05246","source_feed":"cs.AI","score":6,"bucket":"other","rubric_hits":["D1"],"tags":["个性化LLM","行为日志合成","基准测试"],"reason":"用LLM个性化响应，非仿真人类被试，但涉及行为日志合成与真实分布对齐，边界相关。","model":"deepseek-v4-pro","scored_at":"2026-08-17T13:01:36","error":null,"has_summary":false,"summary":null},{"id":"2608.06549","version":2,"title":"TradeVerse: A Longitudinal Benchmark of Political Negotiation in International Trade","zh_title":"TradeVerse：国际贸易政治谈判的纵向基准","abstract":"LLMs are increasingly being applied to tasks involving institutional and political texts, but existing benchmarks evaluate them on isolated documents or single tasks. In realpolitik, negotiations are longitudinal data, where participating parties can align or argue over multiple iterations and each turn is an outcome of the previous turns, hence, understanding one turn requires tracking everything before it. We introduce TradeVerse, a benchmark built from the World Trade Organisation (WTO) specific trade concerns, where member states challenge one another and exchange arguments over multiple rounds, sometimes for years. We, in TradeVerse, reconstruct minutes of $1170$ meetings, spanning across 5 groups and $89$ product groups and define three tasks: first, the system has to analyze the longitudinal meeting records and predict the harmonized system codes (HS chapters) of the products under discussion in the particular meeting, second, we examine whether the system, upon analyzing the anonymized content of the meeting, can guess the name of the responding country and third, we ask the system to play the role of the responding country and provide the statement for the very last round. All labels are recovered directly from the proceedings, requiring no manual annotation. Our experiments highlight the challenges these tasks pose for current LLMs. To the best of our knowledge, TradeVerseis the first benchmark to investigate potential of LLMs in understanding longitudinal political trade negotiations.","authors":["Debodeep Banerjee","Amitangshu Dasgupta"],"categories":["cs.CL","cs.AI","cs.CY"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-08-17","first_seen":"2026-08-10","revised_at":"2026-08-17","abs_url":"https://arxiv.org/abs/2608.06549","pdf_url":"https://arxiv.org/pdf/2608.06549","source_feed":"cs.CL","score":6,"bucket":"other","rubric_hits":["D3"],"tags":["LLM谈判模拟","纵向基准","社会模拟"],"reason":"LLM扮演国家角色进行谈判模拟，但无真实人类行为对照，属社会模拟边界情形。","model":"deepseek-v4-pro","scored_at":"2026-08-17T13:01:37","error":null,"has_summary":false,"summary":null},{"id":"2608.13567","version":1,"title":"Modular Cognitive Architecture Emerges in Large Language Models","zh_title":"大型语言模型中涌现出模块化认知架构","abstract":"The human brain exhibits a striking degree of functional specialization, with distinct networks supporting language, formal reasoning, reasoning about other minds, and reasoning about the physical world. Is this modular organization a fundamental principle of how intelligent systems must be built, or an evolutionary accident specific to biological brains? Here, we test whether a similar organization emerges in Large Language Models--another class of intelligent systems created through a very different optimization process. Using circuit analyses across N=46 tasks spanning four cognitive domains (language, formal reasoning, social reasoning, physical reasoning), we find that LLMs develop a modular architecture that mirrors the human brain: tasks drawing on the same network in humans recruit overlapping neurons in LLMs, whereas tasks drawing on different networks recruit distinct neurons. The convergent emergence of modularity in brains and neural networks suggests that it may be a fundamental property of intelligent systems.","authors":["Pengrui Han","Jacob Andreas","Evelina Fedorenko","Andrea Gregor de Varda"],"categories":["cs.AI","cs.CL","cs.LG"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-08-17","first_seen":"2026-08-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.13567","pdf_url":"https://arxiv.org/pdf/2608.13567","source_feed":"cs.CL","score":6,"bucket":"other","rubric_hits":["D2"],"tags":["LLM认知架构","神经科学类比","模型可解释性"],"reason":"研究LLM内部模块化与人类大脑相似性，属于模型认知架构分析，非仿真人类被试，但…","model":"deepseek-v4-pro","scored_at":"2026-08-17T13:01:22","error":null,"has_summary":false,"summary":null},{"id":"2608.13563","version":1,"title":"Proxy-Validated LLM UX Micro-Simulations: An Artifact-First Protocol for Early-Stage Decision Support","zh_title":"代理验证的LLM用户体验微仿真：面向早期决策支持的人工制品优先协议","abstract":"Early-stage teams often lack users, time, and budget to run repeated UX studies, yet still need decision-oriented signals to iterate safely. We study an LLM-driven UX micro-simulation pipeline that generates structured customer-experience feedback (walkthrough steps, friction points, micro-survey signals) from versioned prompts, personas, tasks, and UI snapshots. Because public usability datasets with task outcomes are scarce, we validate simulated friction themes using multiple public proxy corpora (app reviews, support tweets, and open-source software issues). We propose a lightweight proxy-validation protocol with two alignment metrics: top-k Jaccard and distributional weighted-Jaccard (W), and compare lexical, TF-IDF, and multilingual embedding baselines across six proxy datasets. Embedding-based alignment yields higher W than lexical baselines on primary app-review and support-tweet proxies (e.g., W=0.128 vs 0.000 on Gojek), while top-k Jaccard is shown to overstate alignment at large k. We ablate four agent strategies (single-pass, best-of-N, hybrid, and a proposed score-then-select judge) across Azure OpenAI deployments and report bootstrap confidence intervals over 8 method-dataset pairs; these intervals reveal that the embedding W point estimate is systematically unstable under resampling at our subsample size. We also provide a failure-mode analysis of grounding and fabrication proxies, with documented calibration caveats and worked examples of outputs flagged as fabricated by an adversarial judge. Our artifact-first pipeline produces reproducible tables and figures from versioned run artifacts, supporting iterative prompt and taxonomy refinement before final paid-model calibration.","authors":["Alexandre Cristov\\~ao Maiorano"],"categories":["cs.HC","cs.AI","cs.SE"],"primary_category":"cs.HC","announce_type":"cross","date":"2026-08-17","first_seen":"2026-08-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.13563","pdf_url":"https://arxiv.org/pdf/2608.13563","source_feed":"cs.AI","score":6,"bucket":"other","rubric_hits":["D1","B1"],"tags":["LLM仿真","用户体验","代理验证"],"reason":"用LLM生成UX反馈并代理验证，替代真实用户测试，属标注替代而非仿真人类被试，…","model":"deepseek-v4-pro","scored_at":"2026-08-17T13:01:17","error":null,"has_summary":false,"summary":null},{"id":"2608.13835","version":1,"title":"When Lexical Change Misleads: Rethinking Dynamic Topic Model Evaluation with Traditional and LLM-Based Metrics","zh_title":"当词汇变化误导：用传统与基于LLM的指标重新思考动态主题模型评估","abstract":"Dynamic topic models capture evolving word distributions, but traditional coherence metrics may fail when vocabulary changes while semantic meaning persists. We evaluate 120 topics from CoNTM and DLDA across NYT, DBLP, and arXiv, using three human annotators and Low, Medium, and High lexical-change categories. Traditional temporal coherence shows highly variable agreement with human judgments ($\\rho$=-0.256 to 0.614). In contrast, LLM-based semantic similarity agrees strongly with human semantic judgments for CoNTM on NYT ($\\rho$=0.609), DBLP ($\\rho$=0.721), and arXiv ($\\rho$=0.502), but is less consistent for DLDA. Lexical-change stratification reveals variation hidden by aggregate evaluation. We therefore advocate lexical-change-aware evaluation, jointly reporting traditional coherence and LLM-based semantic measures as complementary rather than interchangeable signals.","authors":["Charu Karakkaparambil James"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-17","first_seen":"2026-08-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.13835","pdf_url":"https://arxiv.org/pdf/2608.13835","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM评估","主题模型","语义连贯性"],"reason":"用LLM替代人工评估主题连贯性，属标注替代而非仿真人类被试，但涉及LLM与人类…","model":"deepseek-v4-pro","scored_at":"2026-08-17T13:01:26","error":null,"has_summary":false,"summary":null},{"id":"2608.13840","version":1,"title":"ASSERT: A Measurement Pipeline for GenAI Audits","zh_title":"ASSERT：生成式AI审计的测量流水线","abstract":"Audits of generative AI (GenAI) systems often summarize behavior as a reported rate: how often the audited system complies with policy. Researchers and stakeholders use that rate to compare systems, track regressions, and gate deployment. A reported rate reflects both the system under audit and the measurement choices behind it, so a change in the rate can leave it unclear whether the system or those choices moved. We introduce ASSERT, a specification-driven measurement pipeline for GenAI audits that ties each reported rate to a written specification of the measurement choices used to produce it. ASSERT helps draft a behavioral rubric and test cases, then runs the audit against a GenAI system and returns a reported rate. In a case study on conversational deception, we observe that the reported rate moves substantially with the dialogue setup, the simulated user, the judge, and the evidence bar for non-compliance. These measurement choices substantially change the reported rate and can reorder GenAI system rankings. Because each reported rate is tied to an explicit specification, differences across audits are easier to attribute and interpret.","authors":["Riccardo Fogliato","Abhinav Palia","Xiawei Wang","Emily Sheng","Chad Atalla","Jean Garcia-Gathright","Nicholas Pangakis","Sharman Tan","Dan Vann","Hannah Washington","P. Alex Dow","Heba Elfardy","Hanna Wallach","Sandeep Atluri"],"categories":["cs.CL","cs.AI","cs.CY"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-17","first_seen":"2026-08-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.13840","pdf_url":"https://arxiv.org/pdf/2608.13840","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["GenAI审计","测量规范","模拟用户"],"reason":"LLM作为审计中的模拟用户和评判者，替代人类角色，但非仿真人类被试，属标注替代","model":"deepseek-v4-pro","scored_at":"2026-08-17T13:01:18","error":null,"has_summary":false,"summary":null},{"id":"2608.13787","version":1,"title":"From Passive Delegates to Strategic Negotiators: Reinforcing Social Reasoning in Small Language Models with SocialRL","zh_title":"从被动代理到战略谈判者：用SocialRL强化小语言模型的社会推理","abstract":"AI agents increasingly act on their users' behalf, handling tasks such as scheduling meetings, comparing offers, and haggling over prices. These principal-driven tasks routinely place the agent across from a counterpart (another user's agent, a seller, a recruiter) whose goals may conflict with its principal's. Yet the dispositions that make an assistant pleasant can make it a poor delegate: a friendly, helpful frontier model may disclose its principal's private information unprompted and concede at the first sign of resistance. We present SocialRL, a general recipe that trains social reasoning directly, and apply it to a 4B model across six domains: Deal-or-No-Deal, CaSiNo, Craigslist, Job Interview, Calendar, and Marketplace. Every domain is trained in-domain under the same recipe, and every policy is evaluated on all six. We find that (1) in-domain training reaches the frontier: on held-out scenarios the 4B matches or exceeds the GPT-5 family per domain, closing 73-122% of the baseline-to-frontier gap on the negotiation games, with 78% of buyer openings anchoring below target versus 3% untrained; (2) cross-domain transfer follows game structure: structurally paired games lift each other, a broad multi-issue donor lifts nearly all domains, and structurally isolated games transfer nothing; (3) guided by this transfer structure, two strategies, cascade RL and multi-teacher on-policy distillation (OPD), consolidate the per-domain specialists into a single unified 4B that reaches 0.627 average utility across all six environments, matching or exceeding GPT-4.1 (0.625), GPT-5.1 (0.619), and GPT-5.2 (0.613); (4) an explicit theory-of-mind scaffold helps only through training: distilling the ToM trace, rather than actions alone, lifts utility on every environment and generalizes better across them, and of the two ToM skills, only next-action prediction predicts negotiation outcomes.","authors":["Wenyue Hua","Zachary Huang","Tyler Payne","Safoora Yousefi","Saleema Amershi","Asli Celikyilmaz"],"categories":["cs.AI","cs.CL","cs.LG","cs.MA"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-08-17","first_seen":"2026-08-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.13787","pdf_url":"https://arxiv.org/pdf/2608.13787","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["多智能体谈判","社会推理","强化学习"],"reason":"多智能体谈判模拟，无真实人类数据对照，属社会模拟边界情形","model":"deepseek-v4-pro","scored_at":"2026-08-17T13:01:25","error":null,"has_summary":false,"summary":null},{"id":"2608.13921","version":1,"title":"When Personal Memory Has No Single Answer: Evaluating LLM Agents under Irreducible Conflict","zh_title":"当个人记忆没有单一答案：在不可约冲突下评估LLM智能体","abstract":"LLM agents increasingly maintain personal memory across sessions, but it can conflict. Preferences depend on context, behavior evolves, and sources can conflict. When a query lacks context, time, or source authority to interpret conflict, treating one memory as definitive converts unresolved conflict into an unjustified, overconfident action. Existing benchmarks recover one answer from conflicting evidence, overlooking whether agents recognize underdetermination, preserve alternatives, seek missing information, and choose appropriate actions. We introduce \\underline{T}esting \\underline{A}gents' \\underline{N}avigation of \\underline{G}enuine, \\underline{L}atent, and \\underline{E}ntangled Memory Conflicts (\\textsc{TANGLE}), a benchmark for genuinely unresolvable memory conflicts. It comprises 541 instances across 40 personas and three types: Context-Partitioned Conflict (CPC), Behavior-Oscillation Conflict (BOC), and Source-Contradiction Conflict (SCC). We evaluate two tracks---an oracle track with curated memory and a pipeline track that extracts memory from multi-session dialogues---on five dimensions: conflict perception, causal reasoning, confidence calibration, clarification seeking, and memory faithfulness. Experiments reveal pipeline challenges. With curated memory, models recognize conflicts more reliably than they calibrate actions or seek targeted clarification. With end-to-end pipeline memory, extraction fails to preserve conflict-bearing relations needed for downstream reasoning. Policy comparisons show fixed rules are insufficient when actions must reflect conflict. These findings motivate Conflict-Aware Action Policy (CAAP), which adapts actions to each conflict using available evidence. \\textsc{TANGLE} frames conflict handling as recognizing underdetermination, retaining conflicting evidence, and acting without forcing a definitive answer.","authors":["Lu Yang","Shusheng Xu","Zhuoran Li","Tongkai Yang","Longbo Huang"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-17","first_seen":"2026-08-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.13921","pdf_url":"https://arxiv.org/pdf/2608.13921","source_feed":"cs.AI","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["LLM记忆","冲突处理","基准测试"],"reason":"评估LLM记忆冲突处理，测模型能力而非仿真人类被试，无人类数据对照","model":"deepseek-v4-pro","scored_at":"2026-08-17T13:01:27","error":null,"has_summary":false,"summary":null},{"id":"2608.14161","version":1,"title":"BiasTrace: Linking Reasoning Behaviours to Biased Outputs in LLMs","zh_title":"BiasTrace：将推理行为与LLM中的偏见输出联系起来","abstract":"LLMs exhibit social biases that can produce inaccurate and discriminatory inferences, posing risks in high-stakes applications. While prior work has made progress in measuring and mitigating bias, it largely focuses on final outputs of models, with limited understanding of the mechanisms that produce biased outcomes. Recent advances in LLM reasoning offers a new lens for investigating bias, yet the link between reasoning and bias remains poorly understood. Existing approaches focus primarily on final answer correctness or explicitly biased language, overlooking different behaviours in reasoning that can drive biased outcomes. We introduce BiasTrace, an annotation scheme for labelling reasoning behaviours in model-generated traces and linking them to biased outcomes. BiasTrace captures bias-specific behaviours (e.g., unsupported demographic assumptions) as well as general reasoning patterns that may implicitly contribute to bias (e.g. overthinking). We apply BiasTrace to reasoning traces in bias-sensitive contexts, scaled using validated LLM-as-a-judge methods, producing a large annotated dataset. Our analysis shows that biased outputs often stem from subtle reasoning behaviours rather than explicitly biased language, and that reasoning-level annotations improve bias detection. We further show that BiasTrace behaviours can be exploited for inference-time mitigation. These findings underscore the importance of examining a broader range of reasoning patterns to better understand bias in LLMs.","authors":["Varsha Ramineni","Hossein A. Rahmani","Jerome Ramos","Karin Sevegnani","Emine Yilmaz"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-17","first_seen":"2026-08-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.14161","pdf_url":"https://arxiv.org/pdf/2608.14161","source_feed":"cs.AI","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["LLM偏见","推理行为","标注方案"],"reason":"研究LLM推理中的偏见行为，测量模型本身而非仿真人类被试，但涉及偏见测量，属边…","model":"deepseek-v4-pro","scored_at":"2026-08-17T13:01:31","error":null,"has_summary":false,"summary":null},{"id":"2608.07852","version":2,"title":"\"Many Are My Names\": The Anatomy of the Assistant and Its Personas via Sparse Autoencoders","zh_title":"“我名众多”：通过稀疏自编码器剖析助手及其人格面具的解剖结构","abstract":"How a language model internally represents who is speaking, the Assistant, an assigned roleplay persona, or a narrated story character, remains underexplored. We study speaker representations using a dataset of user-expressed emotional text and corresponding model responses. We decompose three generation settings (Assistant, Roleplay, and Story) into sparse autoencoder features extracted at turn-boundary and pronoun-token positions and selected through a filtering pipeline for different depths. We characterize each surviving feature through its steering effects and activation distribution. Our main finding is that the Assistant and roleplay personas are not independent alternatives: personas retain the Assistant-associated feature core while progressively differentiating from it across layers, starting from operational machinery towards behavioral and stylistic features. Meanwhile, generated story characters lack the Assistant-associated core. Both Story and Roleplay can be distinguished from the Assistant with Immersive Simulation Mode. However, the Assistant can sometimes enter or slowly drift into it even in the default setting.","authors":["Adelaide Danilov","Aria Nourbakhsh","Oleksandr Marchenko Breneur","Salima Lamsiyah"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-08-17","first_seen":"2026-08-11","revised_at":"2026-08-17","abs_url":"https://arxiv.org/abs/2608.07852","pdf_url":"https://arxiv.org/pdf/2608.07852","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["角色扮演","模型可解释性","说话者表征"],"reason":"研究角色扮演和故事生成中的说话者表征，无实验或测量目的，不涉及人类行为仿真。","model":"deepseek-v4-pro","scored_at":"2026-08-17T13:01:21","error":null,"has_summary":false,"summary":null},{"id":"2608.12831","version":2,"title":"Fast A/B/n Testing: Exact Multi-Policy Comparison via Tree-Coupled Feedback Sharing","zh_title":"快速A/B/n测试：通过树耦合反馈共享进行精确多策略比较","abstract":"Online platforms increasingly compare many adaptive decision policies---ranking systems, recommendation algorithms, pricing rules, and language-model agents---while each reward-bearing interaction can be costly or risky. A direct A/B/n design gives each of $J\\ge 2$ policies its own horizon-$T$ trajectory and therefore uses $JT$ outcomes. We introduce Tree-Coupled A/B Testing (\\TCAB), an exact feedback-sharing design for arbitrary history-dependent contextual-bandit policies. At each round, a predictable tree connects the current policy histories; every parent--child context--action law is maximally coupled, and one reward is shared within each component of matched tree edges. Every policy retains exactly its standalone finite-horizon trajectory law, even though the policies are deliberately dependent. If $D_{e,t}$ records a mismatch on tree edge $e$ at round $t$, the number of reward queries satisfies the pathwise identity $N(T)=T+\\sum_{t,e}D_{e,t}$ and hence equals $T$ plus cumulative tree-edge total variation in expectation. This cost is conditionally optimal among exact edge-local designs on the selected tree, and a current-round minimum-spanning tree is myopically optimal among tree designs. For fixed $J$, sublinear pseudo-regret of every policy and almost-sure uniqueness of the oracle action imply $\\mathbb{E}[N(T)]=T+o(T)$, versus $JT$ for independent runs. We also obtain finite-sample variance bounds for pairwise policy contrasts. Experiments on reward-model evaluation, multiple-choice language-model evaluation, and adaptive search policies demonstrate substantial improvements in the cost--precision frontier.","authors":["Yuxiao Wen"],"categories":["cs.LG","cs.AI"],"primary_category":"cs.LG","announce_type":"replace-cross","date":"2026-08-17","first_seen":"2026-08-14","revised_at":"2026-08-17","abs_url":"https://arxiv.org/abs/2608.12831","pdf_url":"https://arxiv.org/pdf/2608.12831","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["A/B测试","多臂老虎机","统计方法"],"reason":"论文研究多策略A/B测试的反馈共享，不涉及用LLM仿真人类被试，属于统计方法而…","model":"deepseek-v4-pro","scored_at":"2026-08-14T13:02:25","error":null,"has_summary":false,"summary":null},{"id":"2608.13760","version":1,"title":"Amplified Does Not Mean Predictive: Reasoning Behaviors in Thinking Models","zh_title":"放大不等于预测：思维模型中的推理行为","abstract":"Which reasoning behaviors are associated with correct answers in reasoning models, and does reasoning-oriented training amplify those behaviors? This distinction is important because reasoning-oriented training can make traces look more deliberative without amplifying the behaviors most tied to model correctness. We quantify this mismatch with Behavioral Lift, a metric that measures how much correctness changes when a behavior is present versus absent in a model's reasoning trace. Across 15 models and 6 benchmarks spanning text-only and vision-language reasoning, we annotate 15,282 traces with a taxonomy whose core behaviors are defined for both LLM and VLM traces. We find evidence for an Amplification-Lift Gap, in which thinking models strongly amplify self-correction, hypothesis testing, and uncertainty acknowledgment, while the highest-lift behaviors are confidence calibration, knowledge alignment, and self-awareness. Confidence calibration is among the strongest positive signals of correctness in both modalities, yet is barely amplified; uncertainty acknowledgment is amplified by 3--7$\\times$, yet is weakly or negatively associated with correctness. We find that reasoning-oriented training does not preferentially amplify the highest-Lift behaviors, motivating process-level objectives that reward calibrated and grounded reasoning rather than surface form alone.","authors":["Jean de Dieu Nyandwi","Leena Mathur","Yonatan Bisk","Robert Hawkins","Graham Neubig"],"categories":["cs.CL","cs.AI","cs.CV","cs.LG"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-17","first_seen":"2026-08-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.13760","pdf_url":"https://arxiv.org/pdf/2608.13760","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["推理模型","行为分析","模型评估"],"reason":"研究推理模型行为与正确性的关联，属模型能力分析，非人类仿真。","model":"deepseek-v4-pro","scored_at":"2026-08-18T13:04:26","error":null,"has_summary":false,"summary":null},{"id":"2608.13604","version":1,"title":"Cross-Disciplinary Taxonomy and Modeling of Misunderstanding Generation, Amplification, and Detection, from Pragmatics to AI Agents","zh_title":"跨学科误解生成、放大与检测的分类与建模：从语用学到AI智能体","abstract":"Detection of misunderstanding is an urgent problem to solve because communication has moved away from real-time, in-person interaction and is increasingly handled by AI-mediated channels. This shift cuts communicators off from the resources repair depends on faster than new means of detection are being built. In this paper we analyse misunderstanding as a layered process in which a divergence is generated, may then be amplified, and is either detected and repaired or left to persist unnoticed. Consolidating accounts from nine fields of research that do not ordinarily cite one another, we identify eleven exact failure modes and show that each operates at a specific point in a communicative process rather than anywhere within it. Those points give eight analytical layers, derived from the literature rather than adopted from an existing model. Eight of the mechanisms primarily generate a divergence, two primarily amplify one already present, and one governs whether a divergence is detected and repaired. We model the eight layers formally, extending information and communication theory from the transmission of signals to the reconstruction of meaning, and we supply a source-by-source evidence matrix that makes every rating auditable, a coding manual, and nine analysed dialogue cases. No prior classification of misunderstanding both locates mechanisms at points in the process and types them by function.","authors":["Babak Abbaschian"],"categories":["cs.AI","cs.CL","cs.HC","cs.MA"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-08-17","first_seen":"2026-08-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.13604","pdf_url":"https://arxiv.org/pdf/2608.13604","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["误解检测","多智能体通信","语用学"],"reason":"论文研究误解生成与检测的通用机制，不涉及用LLM仿真人类被试或与人类数据对照，…","model":"deepseek-v4-pro","scored_at":"2026-08-17T13:01:24","error":null,"has_summary":false,"summary":null},{"id":"2608.13866","version":1,"title":"Geometric Filtering of LLM-Generated Samples for Few-Shot Text Classification","zh_title":"用于少样本文本分类的LLM生成样本几何过滤","abstract":"Large language models (LLMs) can generate synthetic training data for text classification, but the quality of generated samples is heterogeneous: some fall in correct class regions of the embedding space while others land in peripheral or cross-class zones. We propose a geometric filtering framework that evaluates each LLM-generated sample by its Euclidean distance to real class examples in a sentence embedding space, selecting only geometrically consistent candidates. A soft weighting mechanism transforms filter scores into sample weights for classifier training. Evaluated across 13 datasets, 5 classifiers, 10 augmentation methods, and over 6,700 configurations, our method achieves +2.61 percentage points (pp) over SMOTE ($p<0.0001$, Cohen's $d=0.95$, 88.9% win rate). The approach generalizes to named entity recognition (+9.26pp, 100% win rate) without filter modification, and is robust across 5 LLMs from 4 providers. A key finding is that the simplest distance-based filter consistently outperforms complex multi-criteria alternatives.","authors":["Benjam\\'in Schindler","Gonzalo A. Ruz"],"categories":["cs.LG","cs.CL"],"primary_category":"cs.LG","announce_type":"cross","date":"2026-08-17","first_seen":"2026-08-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.13866","pdf_url":"https://arxiv.org/pdf/2608.13866","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["数据增强","文本分类","合成数据"],"reason":"论文用LLM生成合成数据提升文本分类，属于数据增强，不涉及人类行为仿真或对照。","model":"deepseek-v4-pro","scored_at":"2026-08-17T13:01:27","error":null,"has_summary":false,"summary":null},{"id":"2608.14329","version":1,"title":"A Four-Axis Trustworthiness Benchmark for LLM-as-Judge in Principle-Based Regulation","zh_title":"基于原则监管中LLM作为评判者的四轴可信度基准","abstract":"Principle-based regulation, with evaluative standards such as \"fair, clear, and not misleading\" or \"deliver good outcomes\", cannot be reduced to binary predicates, and LLM-as-judge is increasingly used as the substitute. Our position is that any such judge must be evaluated on four axes: accuracy, paraphrase robustness, adversarial robustness, and calibration. We release Principle-Bench, 168 cryptoasset financial-promotion scenarios mapped to two UK FCA principles, with paraphrase, adversarial keyword-stuffing, and boundary perturbations authored under a pre-registered rubric; the first benchmark covering all four axes for principle-based regulation. We also introduce Ceca (Calibrated Exemplar-Cluster Assessment): a calibrated, auditable assessor that emits exact per-exemplar counterfactual attributions. Across keyword counting, three sentence-transformer embedders, an open-weight LLM-judge, and a calibrated cascade, no method dominates all four axes. A 120B LLM-judge, strongest on benign inputs, loses 47 accuracy points (0.74 to 0.27) on keyword-stuffed Consumer Duty inputs: \"compliance theatre.\" A second judge from a different model family agrees only at Cohen's kappa = 0.16 on that split, localising the failure to the model rather than the corpus. Any deployment-grade LLM-judge for principle-based regulation must report per-principle adversarial deception and post-hoc calibration alongside aggregate accuracy.","authors":["Dipankar Sarkar"],"categories":["cs.CR","cs.AI","cs.CL","cs.CY","cs.LG"],"primary_category":"cs.CR","announce_type":"cross","date":"2026-08-17","first_seen":"2026-08-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.14329","pdf_url":"https://arxiv.org/pdf/2608.14329","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["LLM评估","监管合规","基准测试"],"reason":"评估LLM作为监管合规判断工具，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-08-17T13:01:34","error":null,"has_summary":false,"summary":null},{"id":"2608.14522","version":1,"title":"Participatory Moral AI Is Not Neutral: The Invisible Hand of Developers","zh_title":"参与式道德AI并非中立：开发者的无形之手","abstract":"As AI systems make more morally loaded decisions across society, one response has been moral preference elicitation. In this approach, researchers poll participants on hypothetical dilemmas and use the aggregated votes to train a policy that an AI model then applies at scale. Before any vote is cast, developers make three key choices in the moral AI elicitation pipeline: feature scoping, voter sampling, and question framing. In other words, they decide which features go to a vote, which voters to include, and how to present the question. These choices are often opaque, undocumented, and treated as technical details rather than normative ones. We examine each of these choices within a common empirical study and show that each can shape the preferences produced by moral AI elicitation. Across two phases (N = 809) in three deployment contexts (i.e., AI kidney allocation, AI agents simulating absent workers, and generative AI depictions of the deceased), we examine the three main stages of the moral AI elicitation pipeline. First, morally relevant features shift across contexts. This suggests that feature schemas should not be assumed to transfer across deployment domains. Second, preferences differ by political ideology for roughly one-third of features, with some differences reversing direction. The ideological composition of the voter pool can therefore affect the resulting aggregated preference profile. Third, the wording of the elicitation question can narrow or widen ideological gaps by up to a full scale point. The framing conditions also change how moral foundations are associated with participants' judgments. Taken together, these findings suggest that voting-based alignment cannot deliver fair or transparent AI by aggregation alone; at minimum, each stage of the moral AI elicitation pipeline should be audited and disclosed.","authors":["Taenyun Kim","Edyta Bogucka","Daniele Quercia"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-17","first_seen":"2026-08-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.14522","pdf_url":"https://arxiv.org/pdf/2608.14522","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["道德AI","偏好聚合","人类实验"],"reason":"研究人类道德偏好聚合，非LLM仿真人类被试，无LLM作为被试替代。","model":"deepseek-v4-pro","scored_at":"2026-08-17T13:01:21","error":null,"has_summary":false,"summary":null},{"id":"2608.14528","version":1,"title":"Handover of In-Context Learning State Across Session Boundaries","zh_title":"跨会话边界的上下文学习状态交接","abstract":"This study investigates the methodological and theoretical properties of session handover in applications that use large language models. A task may continue in a new session when the context reaches the model's input limit, when the application restarts, or when another agent is asked to finish the task. The application must then decide which information from the earlier session to pass on. We formulate handover as the transfer of a task-relative in-context learning (ICL) state and distinguish exact recovery of earlier material from preservation of the target distribution. Under an exogeneity condition, predictive equivalence characterizes the coarsest deterministic sufficient handover and gives a fixed-length bit requirement. The analysis isolates the effects of the memory constraint, the writer, and the continuation procedure, and quantifies the cost of writing before the realized downstream query is known. We propose a three-part record that stores decisions and constraints exactly, uses task-justified statistics for repeated evidence, and retains original observations whose effect is not preserved by those statistics. Gaussian linear regression gives an exact finite-dimensional handover and finite-bit perturbation bounds, while nonparametric regression gives upper and lower bounds that relate memory to squared prediction error. These results provide a theory and method for deciding what a handover must retain and how its memory requirement depends on the continuation task.","authors":["Masahiro Kato","Taka Kato"],"categories":["cs.AI","econ.EM","math.ST","stat.ME","stat.ML","stat.TH"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-17","first_seen":"2026-08-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.14528","pdf_url":"https://arxiv.org/pdf/2608.14528","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["LLM会话交接","上下文学习","多智能体系统"],"reason":"研究LLM会话间状态传递，属多智能体协作技术，不涉及人类行为仿真或对照。","model":"deepseek-v4-pro","scored_at":"2026-08-17T13:01:34","error":null,"has_summary":false,"summary":null},{"id":"2608.13944","version":1,"title":"Musical Mirrors: The LLM as Sounding Board in Songwriting","zh_title":"音乐之镜：LLM作为歌曲创作中的共鸣板","abstract":"This paper examines a use of AI in creative practice as an interpretive sounding board for human-generated material, rather than the more familiar pattern of AI generation followed by human curation. Through the lens of resonance as theorized by Hartmut Rosa, I present a first-person case study of songwriting from July 2025 to March 2026, drawing on 16 original pieces in English, French, and other languages along with piano solos. I describe a configuration in which resonance is not located between user and model, but in the author's deepening contact with their own material, mediated through the model. This kind of resonance was supported rather than inhibited by AI when sounding-board behavior was cultivated through sustained calibration by the user. Two failure modes appeared when calibration was absent: sycophantic drift and magical overinterpretation. This account suggests both the potential and the risks of AI as an interpretive partner in creative practice.","authors":["Xiao Xiao"],"categories":["cs.HC","cs.AI"],"primary_category":"cs.HC","announce_type":"cross","date":"2026-08-17","first_seen":"2026-08-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.13944","pdf_url":"https://arxiv.org/pdf/2608.13944","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["人机交互","创意实践","LLM应用"],"reason":"LLM作为创作中的解释性反馈工具，属于角色扮演对话，无实验或测量目的，不涉及人…","model":"deepseek-v4-pro","scored_at":"2026-08-17T13:01:28","error":null,"has_summary":false,"summary":null},{"id":"2608.14130","version":1,"title":"AlignFace: Human-Aligned Face Similarity Metric with Interpretable Concept Relations","zh_title":"AlignFace：具有可解释概念关系的人脸相似度度量","abstract":"Computer vision models for generated facial content, such as face editing and privacy protection, increasingly affect people, requiring similarity metrics that serve as faithful proxies for human perception. While perceptual evaluation has progressed from signal-based heuristics to representation-based metrics, current approaches are limited to behavioral modeling without cognitive alignment. They rely on implicit and spurious relations while assuming a universal observer, failing to account for inherent variations across diverse human populations. This leads to inaccurate evaluative models of stakeholders and misleading guidance for generative model debugging. Rather than treating perception as a black box, we leverage scientific findings from cognitive psychology of human face similarity perception: dependence on facial featural and configural attributes, nonlinear psychophysical response scaling, and own-group biases. We introduce the FACETS dataset and propose AlignFace, an interpretable, human-aligned, face similarity metric that encodes these cognitive principles through ante-hoc modeling. It employs visual-language modeling (VLM) to encode paired face images and text-based attributes, gated cross-attention (CA) to extract attribute-specific facial difference representations, concept bottleneck modeling (CBM) to constrain reasoning via interpretable face attributes, and neural generalized additive model (GAM) to model their nonlinear influence. Experiments found AlignFace significantly improves alignment with human subpopulation perceptions compared to baseline metrics, including recent domain-free learned perceptual metrics. By bridging learned representations and human cognitive processes, this work enables more transparent and aligned perceptual evaluation metrics for face images.","authors":["Ying Huang","Wencan Zhang","Brian Y. Lim"],"categories":["cs.MM","cs.AI","cs.CV","cs.HC"],"primary_category":"cs.MM","announce_type":"cross","date":"2026-08-17","first_seen":"2026-08-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.14130","pdf_url":"https://arxiv.org/pdf/2608.14130","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C2"],"tags":["人脸相似度","感知评估","计算机视觉"],"reason":"研究人脸相似度度量，不涉及LLM仿真人类被试，属于计算机视觉感知评估。","model":"deepseek-v4-pro","scored_at":"2026-08-17T13:01:29","error":null,"has_summary":false,"summary":null},{"id":"2608.14093","version":1,"title":"AppLooper: An Agentic Application Engineering Loop for Accountable Release with Virtual-User Feedback","zh_title":"AppLooper：一种面向可问责发布、结合虚拟用户反馈的智能体应用工程循环","abstract":"Much existing research on coding agents organizes application development as an iterative loop of requirement interpretation, implementation, tool execution, evaluation, and repair. As these loops run longer, requirements may drift; users may lose awareness of the current state and rationale for changes; and generated applications may remain insufficiently grounded in target users' contexts and needs. Application engineering therefore requires a mechanism connecting owner intent, target-user experience, development changes, and responsibility for release. We present AppLooper, a human--coding-agent--virtual-user application engineering loop for accountable release. An application owner confirms frozen requirements, supplies feedback, inspects candidates, and retains final release authority. A development agent produces and revises versioned candidates. A virtual-user agent cohort executes interface scenarios grounded in target users and contexts of use. Besides, an owner-intent simulation agent retests only requirements, constraints, and feedback explicitly confirmed by the owner, abstaining when evidence is insufficient. A testing agent performs read-only developmental checks by reproducing reported failures, running existing regression tests, and exercising the current candidate through its browser interface. The orchestration layer groups the resulting findings and routes them into development revision, targeted retesting, and owner inspection. AppLooper binds requirements, feedback sources, interface targets, development changes, retesting outcomes, owner interactions, and release decisions to specific versions. It thereby extends sustained coding-agent iteration into a traceable and reviewable lifecycle in which humans retain final responsibility for release. Source code is available at https://github.com/ZihongHe/applooper.","authors":["Zihong He","Chen Liang","Hai-Ning Liang"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-08-17","first_seen":"2026-08-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.14093","pdf_url":"https://arxiv.org/pdf/2608.14093","source_feed":"cs.HC","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","应用工程","虚拟用户测试"],"reason":"虚拟用户仅用于测试应用界面，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-08-17T13:01:29","error":null,"has_summary":false,"summary":null},{"id":"2608.11625","version":2,"title":"Making AI-Generated Feedback Matter: A Large-Scale Study of Feedback Workflows and Student Enactment","zh_title":"让AI生成的反馈发挥作用：反馈工作流与学生实施的大规模研究","abstract":"Feedback processes strongly influence student learning, yet their educational value depends on addressing two distinct challenges: providing high-quality, timely, and individualised feedback at scale, and supporting students to interpret, evaluate, and act on that feedback productively. Generative AI offers a credible means of addressing the provision challenge, but students' uptake of AI-generated feedback remains limited. We conducted a large-scale quasi-experimental sequential cohort study comparing three AI-mediated feedback workflows across 13,037 students and 51,296 student-authored resources. In Directed Feedback (n = 3,723), students received AI-generated feedback comments without structured support. In Self-Directed Feedback (n = 3,951), students could initiate optional AI-supported dialogue. In Enacted Feedback (n = 5,363), students were prompted to select feedback suggestions, evaluate their relevance, and engage in targeted AI-supported dialogue anchored to those selections. Enacted Feedback was associated with significantly higher uptake of AI-generated feedback, with an estimated probability of 26.2%, compared with 14.1% for Directed Feedback and 0.1% for Self-Directed Feedback. It was also associated with significantly higher self-assessment confidence and submitted-work quality than both comparison conditions. These findings suggest that the educational value of AI-generated feedback depends not only on the quality of feedback comments, but also on workflows that actively structure students' enactment of feedback literacy processes. The results have implications for the design of AI feedback systems that position learners as active participants in judgement, dialogue, and improvement rather than passive recipients of comments. Overall findings show that AI access alone is insufficient; purposeful workflow design is central to productive feedback use.","authors":["Omar Alsaiari","Nilufar Baghaei","Jason M. Lodge","Dragan Ga\\v{s}evi'c","Naomi Winstone","Hassan Khosravi"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"replace","date":"2026-08-17","first_seen":"2026-08-13","revised_at":"2026-08-17","abs_url":"https://arxiv.org/abs/2608.11625","pdf_url":"https://arxiv.org/pdf/2608.11625","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["AI反馈","教育技术","学习分析"],"reason":"研究AI反馈工作流对学生学习的影响，不涉及用LLM仿真人类被试或对照人类行为数…","model":"deepseek-v4-pro","scored_at":"2026-08-17T13:01:21","error":null,"has_summary":false,"summary":null},{"id":"2608.13624","version":1,"title":"Measuring Fairness in Large Audio Language Models via Semantic-Aware Bias Estimation","zh_title":"通过语义感知偏差估计测量大型音频语言模型中的公平性","abstract":"Large Audio Language Models (LALMs) have seen increasing use for audio understanding tasks such as speech recognition and audio question answering, raising concerns about fairness across demographic subgroups. Fairness evaluation in spoken-input settings is challenging due to confounding factors, including semantic variation in spoken content and speaker-specific characteristics. Ignoring these factors can result in misleading conclusions about model bias. We propose a semantic-aware mixed-effects regression framework for fairness evaluation in LALMs that explicitly accounts for these confounders. Our approach incorporates sentence-level semantic embeddings of reference text as covariates and models speaker identity as a random effect. Notably, semantic representations are extracted from the same LALM under evaluation, enabling semantic control over variation as perceived by the model itself. Experiments on simulated data and real-world benchmarks demonstrate that the proposed approach substantially reduces spurious fairness findings and yields more robust and interpretable estimates of subgroup performance differences.","authors":["Zhe Liu"],"categories":["cs.CL","cs.AI","cs.SD"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-17","first_seen":"2026-08-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.13624","pdf_url":"https://arxiv.org/pdf/2608.13624","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["公平性评估","音频语言模型","偏差估计"],"reason":"评估音频语言模型的公平性，属于模型偏差测量，不涉及用LLM仿真人类被试或与人类…","model":"deepseek-v4-pro","scored_at":"2026-08-18T13:04:26","error":null,"has_summary":false,"summary":null},{"id":"2608.14210","version":1,"title":"How Much Do Legal RAG Systems Still Hallucinate?","zh_title":"法律RAG系统仍会产生多少幻觉？","abstract":"Hallucination is a major challenge for retrieval-augmented generation (RAG) systems in the legal domain, where ungrounded answers can lead to serious consequences. To better understand this problem, we conduct a fine-grained analysis of hallucination behavior in eight legal RAG systems across two legal corpora, the GDPR (in English) and a national civil law (in French). Using claim-level and answer-level evaluation, we report on hallucination density and severity, analyze performance across question categories and user personas, and validate our findings on an independent set of 142 legal-expert-authored questions. Our results show that hallucinations remain pervasive, ranging from less than 10% of responses for the best-performing systems to nearly half in the worst case. We further find that false-premise questions, containing incorrect assumptions that must be rejected, produce high hallucination rates on the manually-drafted questions.","authors":["Souvick Das","Sallam Abualhaija","Domenico Bianculli"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-17","first_seen":"2026-08-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.14210","pdf_url":"https://arxiv.org/pdf/2608.14210","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["法律RAG","幻觉评估","NLP评测"],"reason":"评估法律RAG系统幻觉，属NLP能力评测，不以人类行为为参照","model":"deepseek-v4-pro","scored_at":"2026-08-17T13:01:32","error":null,"has_summary":false,"summary":null},{"id":"2608.13674","version":1,"title":"Asymmetric Discourse Homogenization and Shared Language Technology: Evidence from Reddit","zh_title":"不对称话语同质化与共享语言技术：来自Reddit的证据","abstract":"I document an ideologically asymmetric break in the pre-existing diversification trend of political discourse, emerging around late 2022, using 6 million Reddit comments from two cross-partisan forums, 2019-2025. Conservative users experienced an interruption of their prior diversification trajectory; progressive users showed no comparable change. The asymmetry is consistent across estimation strategies (ITS, DiD, RDiT, propensity-score matching) and temporal aggregations. A daily-frequency permutation test over 2,377 candidate cutoff dates shows the ChatGPT threshold produces an unremarkable estimate (49.8th percentile): the shift builds gradually instead of breaking at a single date. A continuous cumulative LLM index, tracking AI exposure across seven model releases, remains significant under a quadratic trend specification that eliminates the binary estimate. A stayer analysis narrows the mechanism: the homogenization effect disappears when the sample is restricted to authors active throughout the study period, and the stayer confidence interval excludes within-author effects even a tenth the size of the full-sample estimate. The mechanism is most parsimoniously ecological (community-level discursive convergence) rather than individual-level AI adoption, though the data cannot cleanly separate this account from concurrent secular change.","authors":["Fengming Liu"],"categories":["cs.CY","cs.CL","cs.SI"],"primary_category":"cs.CY","announce_type":"cross","date":"2026-08-17","first_seen":"2026-08-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.13674","pdf_url":"https://arxiv.org/pdf/2608.13674","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["社会计算","政治话语","Reddit"],"reason":"研究Reddit政治话语同质化，非LLM仿真人类被试，无人类数据对照，属社会计…","model":"deepseek-v4-pro","scored_at":"2026-08-17T13:01:17","error":null,"has_summary":false,"summary":null},{"id":"2608.14198","version":1,"title":"MINT: A Universal Zero-Shot Predictor for Transaction Data","zh_title":"MINT：交易数据的通用零样本预测器","abstract":"Banks analyse sequential financial transaction data to perform many tasks, including fraud prevention, credit risk assessment and offer personalization. To improve the predictive accuracy of these tasks, Payments Foundation Models encode transaction sequence data as rich contextual embeddings, which can then be provided to task-specific models as features. However, these Foundation Models are not designed for flexible zero-shot reasoning across novel downstream prediction tasks, limiting their adaptability and utility. Existing LLM-based approaches to zero-shot prediction often fail to fully exploit the predictive signal within transaction data, while relying on costly text serialization or task-specific architectures that scale poorly. To address these limitations, we present the Multimodal Instruction Network for Transactions (MINT), a framework that connects a pretrained transaction sequence encoder to a decoder-only LLM through lightweight embedding injection, transaction-language alignment, and instruction tuning. We find that MINT achieves state-of-the-art predictive question-answering performance in both in-distribution and out-of-distribution questions, while substantially reducing input tokens, latency, and memory consumption compared to text-serialization baselines. Through comprehensive analyses of representations, alignment strategies, training data, and history length, we establish that compact transaction embeddings are a superior approach to transaction representation than text serialization for multimodal reasoning and zero-shot prediction tasks.","authors":["Parameswaran Kamalaruban","Viktor Drobnyi","Maeve Madigan","Julia Rozanova","David Sutton","Stuart Burrell"],"categories":["cs.LG","cs.CL"],"primary_category":"cs.LG","announce_type":"cross","date":"2026-08-17","first_seen":"2026-08-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.14198","pdf_url":"https://arxiv.org/pdf/2608.14198","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["金融预测","多模态LLM","零样本学习"],"reason":"纯预测模型，无人类仿真或行为对照","model":"deepseek-v4-pro","scored_at":"2026-08-17T13:01:32","error":null,"has_summary":false,"summary":null},{"id":"2608.14286","version":1,"title":"Seeing Red, Thinking Bad: Color Bias in Vision Language Models","zh_title":"见红思坏：视觉语言模型中的颜色偏差","abstract":"Vision language models (VLMs) are increasingly used in industrial decision-making systems, such as recruitment support and recommendation. This motivates careful analysis of how VLMs process visual and textual information. In this work, we study how VLMs interpret text rendered as an image, and investigate the influence of visual styling biases. To this end, we introduce Stealth Visual Prompts, which subtly change visual styling of text, such as color and contrast, while preserving semantic content. Using these prompts, we systematically control the visual styling of words in text and measure their impact on the analysis performed by VLMs. We further analyze how such visual perturbations affect the latent representations of the vision encoder. From our experiments, we observed that coloring positive words in green consistently shifts sentiment predictions toward a positive direction. As a result, VLMs often fail to properly account for negative words present in the text. Our analysis suggests that this behavior is correlated with changes in the latent representations of the vision encoder induced by color variations. In addition, we show that reducing text--background contrast increases reliance on visually salient cues and leads to more incorrect Visual Question Answering (VQA) outputs. These results suggest that the visual styling of rendered text can guide VLMs' interpretation in ways that diverge from human semantic understanding. Project page: https://github.com/KohsukeIde/color-bias-vlm","authors":["Kohsuke Ide","Ryousuke Yamada","Yoshihiro Fukuhara","Hirokatsu Kataoka","Yutaka Satoh"],"categories":["cs.CV","cs.AI","cs.CL"],"primary_category":"cs.CV","announce_type":"cross","date":"2026-08-17","first_seen":"2026-08-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.14286","pdf_url":"https://arxiv.org/pdf/2608.14286","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["视觉语言模型","偏见分析","对抗性提示"],"reason":"研究VLM对视觉样式偏差的响应，不涉及用LLM仿真人类被试或与人类数据对照。","model":"deepseek-v4-pro","scored_at":"2026-08-18T13:04:26","error":null,"has_summary":false,"summary":null},{"id":"2608.14509","version":1,"title":"Split the Labor: Separating Evidence Interpretation from Decision Aggregation","zh_title":"分工：将证据解释与决策聚合分离","abstract":"Systems that ask a language model to reach a conclusion from many sources usually concatenate them into one prompt. This conflates two operations with different requirements. Interpreting a source rewards capacity and context. Combining interpretations rewards fixed arithmetic, comparability across instances, and the option to return nothing. Once separated, the design problem becomes the interface between them. We propose a four-field evidence tuple (hypothesis, reliability bucket, rationale, provenance) and show that fixing it determines both halves. The separation also reveals a failure mode in how such systems combine, which we call count-scale drift. Thresholding a sum of unnormalized weights is exactly posterior thresholding, but at an operating point that slides with the number of sources consulted. The slide grows with reader reliability. When source reliabilities differ, the vote rule and the posterior order instances differently, and no threshold reconciles them. Pooling calibrated log-likelihood ratios addresses both problems. The fix is arithmetic rather than architectural, and applies to a class of rules beyond language models: score-summing triage engines, diagnostic panels scored by counting positives, and additive multi-signal detectors. We then instantiate the principle twice on one longitudinal corpus, once after outcomes resolve and once before. The same partition helps in both, at different granularities: over reading in the first, over learning capacity in the second. There, a small sequence encoder on an easy auxiliary objective plus a tree ensemble carrying the censored survival loss reaches 0.921 AUPRC against 0.805 for a hand-crafted baseline. We separate what transfers from what must be re-estimated per domain, and state five predictions that would falsify the framework, three negative results, and which comparisons remain confounded.","authors":["Zhelun Wu"],"categories":["cs.AI","cs.CL","cs.LG"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-08-17","first_seen":"2026-08-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.14509","pdf_url":"https://arxiv.org/pdf/2608.14509","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["多源信息聚合","决策系统","语言模型"],"reason":"论文研究多源信息聚合的决策系统，不涉及用LLM仿真人类被试或与人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-08-17T13:01:34","error":null,"has_summary":false,"summary":null},{"id":"2608.13577","version":1,"title":"AI Evaluation Should Work With Humans","zh_title":"AI评估应与人类合作","abstract":"This position paper argues that the dominant paradigm of AI evaluation (which focuses on superhuman autonomous performance and so implicitly targets the goal of replacing humans) is guiding AI development in the wrong direction. Instead, the AI community should pivot to evaluating the performance of human--AI teams. We argue that this collaborative shift will foster AI systems that act as true complements to human capabilities and therefore lead to far better societal outcomes than will the current process.","authors":["Jan Kulveit","Gavin Leech","Tom\\'a\\v{s} Gaven\\v{c}iak","Raymond Douglas"],"categories":["cs.AI","cs.LG"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-17","first_seen":"2026-08-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.13577","pdf_url":"https://arxiv.org/pdf/2608.13577","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["AI评估","人机协作","立场论文"],"reason":"论文主张评估人机协作而非用LLM仿真人类被试，不涉及人类行为对照或仿真实验。","model":"deepseek-v4-pro","scored_at":"2026-08-17T13:01:23","error":null,"has_summary":false,"summary":null},{"id":"2608.14014","version":1,"title":"Buy the Rumor, Sell the News: When Is News Priced In?","zh_title":"买谣言，卖新闻：新闻何时被定价？","abstract":"Two old market sayings hold that news is already priced in by the time it is published, and that the rumor is bought while the news is sold. Both place the price move associated with a piece of news before and at publication rather than after it. Whether the claims hold, for which kinds of news, and by how much are basic questions about how fast markets absorb public information. We test them on 4.57 million financial news articles covering roughly 3,000 US stocks (2023-2026). A large language model teacher, distilled into a compact classifier through active learning, assigns each article one of 17 event tags and five attributes; articles are clustered into stories to separate first reports from follow-up coverage; and beta-adjusted abnormal returns are measured around the resulting 1.68 million stock-day events, with 364,405 neutral-sentiment events as a placebo group. Three results follow. First, the price move associated with news concentrates before and at publication: pooled across all signed events, the cumulative move in the news direction by the close of publication day is 2.8 times its value 20 days later, and for rumor-flagged events the rumor day captures the entire move while the subsequent confirmation contributes nothing. Second, measured against the placebo of comparable stocks, markets underreact to numbers and overreact to stories: quantified fundamental news (earnings, dividends, guidance, analyst actions) keeps drifting in the direction of the news for weeks, while soft story-driven news (launches, macro commentary, leadership) gives back its move. Third, news carries width as well as direction: publicity raises volatility before the publication day, and volatility declines once the news is out, because publication resolves uncertainty. The study also produces a table of measured drift for each event tag, usable as a prior in news-conditioned forecasting models.","authors":["Alireza Kargarzadeh","Nariman Khaledian","Navid Parvini","Sid Ghatak","Arman Khaledian"],"categories":["cs.AI","cs.LG","q-fin.ST"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-17","first_seen":"2026-08-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.14014","pdf_url":"https://arxiv.org/pdf/2608.14014","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["金融新闻","LLM分类","市场效率"],"reason":"论文用LLM做新闻分类，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-08-17T13:01:28","error":null,"has_summary":false,"summary":null},{"id":"2608.14152","version":1,"title":"Towards Efficient Multimodal and Multilingual Opinion Extraction for STI: A QLoRA-Based Fine-Tuning Approach","zh_title":"面向科技情报的高效多模态多语言观点抽取：基于QLoRA的微调方法","abstract":"Recent advances in large language models (LLMs) have reshaped semantic analysis. Opinion Extraction (OE) for Science and Technology Intelligence (STI) requires concise core opinions from large information streams. Off-the-shelf models struggle to filter noise from these streams and show limited structured-output reliability in zero-shot multilingual and multi-modal settings. To address information overload and extraction defocus, this study proposes a multimodal core-opinion extraction framework in which visual evidence serves as a contextual anchor for textual judgment. Using VideoLLaMA2 (VL2) and VideoLLaMA2.1 (VL2.1) as the base models, we apply Quantized Low-Rank Adaptation (QLoRA) fine-tuning on a curated dataset of 2,194 multilingual and multimodal samples. Under the selected Image-Augmented setting, fine-tuned VL2.1 generates structured JSON core-opinion outputs, achieving 64.98% Precision, 42.15% Recall, 51.14% F1-score, and 74.00% sample-level accuracy. Relative to the zero-shot VL2.1 setting, it raises the F1-scores of Spanish and Russian from 4.83% and 0.45% to 46.05% and 51.93%, respectively. The framework further incorporates a Fuzzy Cumulative Prospect Theory-based post-extraction triage module for case-level value assessment, providing a case-level value signal for downstream STI screening.","authors":["Sheng Hong","Xuanqi Wang","Jiacheng Wang","Yuwei Wang"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-17","first_seen":"2026-08-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.14152","pdf_url":"https://arxiv.org/pdf/2608.14152","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["观点抽取","多模态","科技情报"],"reason":"论文做的是多模态多语言观点抽取，属于NLP能力评测，不以人类行为为参照系。","model":"deepseek-v4-pro","scored_at":"2026-08-17T13:01:31","error":null,"has_summary":false,"summary":null},{"id":"2608.14179","version":1,"title":"Can Language Models Understand mmWave Data? Benchmarking Large Language Models for mmWave Radar-Based Human Understanding","zh_title":"语言模型能理解毫米波数据吗？基于毫米波雷达的人类理解基准测试","abstract":"Large language models (LLMs) have shown remarkable reasoning and generative capabilities, motivating their use as universal reasoning engines for perception. While modern approaches such as vision-language models (VLMs) have attempted to incorporate reasoning capabilities into visual sensing, the integration of LLMs with the millimeter-wave (mmWave) modality-despite its unique advantages under low light and occlusion-remains largely unexplored. The principal bottlenecks stem from the scarcity of radar language pairs, severe cross-dataset heterogeneity, and the absence of a foundational mmWave encoder. We address this gap through a minimal textualization interface that serializes each mmWave point cloud into concise natural language, allowing off-the-shelf LLMs to operate in a question answering (QA) setting. Building on this, we present mmWave-QA, the first benchmark for language-conditioned mmWave human perception. mmWave-QA aggregates heterogeneous public mmWave datasets and harmonizes them via calibration-aware preprocessing and global taxonomy alignment, while providing natural language QA. Spanning six scenarios and five QA tasks, the benchmark enables standardized evaluation across diverse mmWave hardware and experimental conditions, establishing a foundation for scalable research on mmWave-LLM integration. We further evaluate and analyze LLMs on our mmWave-QA, highlighting their zero-shot reasoning potential for radar perception, as well as their robustness under visual degradation.","authors":["Jeongwan Shin","Jaehyeon Kim","Donguk Ko","Jaeho Choi"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-17","first_seen":"2026-08-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.14179","pdf_url":"https://arxiv.org/pdf/2608.14179","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C2"],"tags":["毫米波雷达","多模态问答","感知基准"],"reason":"研究mmWave雷达感知，用LLM做问答，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-08-17T13:01:32","error":null,"has_summary":false,"summary":null},{"id":"2608.14132","version":1,"title":"Act2Intention: A Benchmark For Developing Active Mobile Agents Through Inferring User Intention from GUI Actions","zh_title":"Act2Intention：通过从GUI动作推断用户意图来开发主动移动智能体的基准","abstract":"Mobile GUI Agents powered by multimodal large language models (MLLMs) show promise in human-computer intelligence. However, current research primarily focuses on reactive task execution while lacking a comprehensive understanding-prediction-execution process for user intentions, which are the core requirements of active agents. In this paper, we propose the Act2Intention framework that builds an active mobile agent by integrating understanding, predicting user intentions, and executing decisions. First, we construct the Act2Intention Bench through data collection and validated generation, comprising 72,511 intentions and over 700,000 actions across 52 apps, thereby establishing the first benchmark for evaluating proactive agents via continuous intention-action trajectories. We further develop the Act2Intention Agent, achieving proactive services through Proactive-oriented Intention Understanding, Personalized Proactive Intention Prediction, and Experience-guided Intention Execution. Experimental results show that supervised fine-tuning on Act2Intention Bench yields absolute improvements of +32.0 Acc-S, +10.25 Acc-S, and +6.9 SSR points over non-fine-tuned counterparts under the same agent framework for intention understanding, prediction, and execution, respectively. This success underscores the necessity and value of the Act2Intention Bench, which establishes a standardized platform for developing and evaluating proactive agents and consequently paves the way for research on intention-driven human-computer interaction.","authors":["Xiaokai Yan","Jingtao Ding","Yong Li","Zhiwen Yu"],"categories":["cs.HC","cs.AI"],"primary_category":"cs.HC","announce_type":"cross","date":"2026-08-17","first_seen":"2026-08-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.14132","pdf_url":"https://arxiv.org/pdf/2608.14132","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C2"],"tags":["移动智能体","意图理解","人机交互"],"reason":"研究移动GUI智能体的主动服务，属于人机交互应用，不涉及用LLM仿真人类被试或…","model":"deepseek-v4-pro","scored_at":"2026-08-17T13:01:30","error":null,"has_summary":false,"summary":null},{"id":"2608.14291","version":1,"title":"Human and Artificial Intelligence - Promoting Trustworthy and Understandable Collaboration","zh_title":"人类与人工智能——促进可信赖且可理解的协作","abstract":"Methods of Artificial Intelligence (AI) enable the personalization of information for individual user experiences in many domains; however, they can also conflict with established design principles, e.g., due to uncertainties regarding the real world. Building trust and understanding can serve as an approach to create a more balanced relationship between humans and AI. Building upon a pilot study, an online survey was conducted to investigate 12 individual aspects related to the topics of explainability and controllability. The results indicate that both topics, despite their different and numerous facets, are generally perceived as important by respondents; simultaneously, however, a wide dispersion of opinions is frequently observed. This could be an indication that, alongside a fundamental consensus, individual perspectives, technical knowledge and understanding, context-specific factors, or personal experiences play a role in the perception of such systems.","authors":["Gilbert Drzyzga"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-08-17","first_seen":"2026-08-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.14291","pdf_url":"https://arxiv.org/pdf/2608.14291","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["人机交互","可解释性","用户调查"],"reason":"研究人类对AI可解释性与可控性的感知，不涉及用LLM仿真人类被试，无实验或测量…","model":"deepseek-v4-pro","scored_at":"2026-08-17T13:01:33","error":null,"has_summary":false,"summary":null},{"id":"2608.13864","version":1,"title":"Audience capture, selective exposure, affective assimilation or ideological sorting? Polarisation of climate politics under low media-party parallelism","zh_title":"低媒体-政党平行主义下气候政治的两极分化：受众捕获、选择性接触、情感同化还是意识形态排序？","abstract":"Studies often attribute the polarisation of climate politics in majoritarian democracies to differing treatment of climate change by conservative and liberal media. How then does such polarisation occur in a system, where mainstream media do not systematically align with party positions? Focusing on how users in different ideological blocs posted and commented on news articles on Finnish Twitter between 2015 and 2023, we address four mechanisms: audience capture, where outlets adapt to increasingly differentiated audiences, reinforcing their divergence through feedback between audience engagement and editorial incentives; selective exposure, where users post ideologically congruent articles from the same outlets; affective assimilation, where users communicate identical content with ideology-consistent valence; and ideological sorting, where identity-driven alignment structures news-posting behaviour across multiple issues. Bayesian regression models indicate increasing divergence in outlet choice, with conservative-leaning users posting more tabloid content, in line with audience capture, and partial support for selective exposure, as within-outlet article posting becomes increasingly differentiated. Valence divergence increases only slightly over time, providing moderate support for affective assimilation, while issue-specific communities show weak alignment with blocs despite convergence over time. Our work refines theories of media effects beyond majoritarian democracies and contributes to understanding the current impasse in climate politics.","authors":["Arttu Malkam\\\"aki","Antti Gronow"],"categories":["cs.SI"],"primary_category":"cs.SI","announce_type":"new","date":"2026-08-17","first_seen":"2026-08-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.13864","pdf_url":"https://arxiv.org/pdf/2608.13864","source_feed":"cs.SI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["社交媒体分析","政治极化","计算社会科学"],"reason":"研究社交媒体用户行为，未使用LLM仿真人类被试，不涉及LLM替代人类进行实验或…","model":"deepseek-v4-pro","scored_at":"2026-08-17T13:01:27","error":null,"has_summary":false,"summary":null},{"id":"2608.14141","version":1,"title":"Who Owns the Online Media?","zh_title":"谁拥有在线媒体？","abstract":"Ownership matters for the media's watchdog role. We map the ownership networks behind thousands of online news outlets in the U.S., Canada, and Europe. The networks reveal who is ultimately responsible for the news: for over half of the outlets, a single entity. The rest sit behind multi-layered structures, making responsibility hard to trace. Market concentration, measured comparably across countries, is largely low to moderate. Looking at content, we find that co-owned outlets report more similarly, even within fixed outlet pairs, as ownership changes -- not least in the U.S., where reader demand is often thought dominant.","authors":["Ulrich Matter","Philine Widmer"],"categories":["econ.GN","q-fin.EC"],"primary_category":"econ.GN","announce_type":"new","date":"2026-08-17","first_seen":"2026-08-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.14141","pdf_url":"https://arxiv.org/pdf/2608.14141","source_feed":"econ.GN","score":0,"bucket":"other","rubric_hits":["C5"],"tags":["媒体所有权","网络分析","内容相似性"],"reason":"研究媒体所有权网络与内容相似性，未使用LLM仿真人类被试，与研究方向无关","model":"deepseek-v4-pro","scored_at":"2026-08-17T13:01:30","error":null,"has_summary":false,"summary":null},{"id":"2608.12368","version":1,"title":"Agreement Is Not Alignment: Divergent Moral Grounds in Human and LLM Ethical Judgments","zh_title":"一致不等于对齐：人类与LLM道德判断中分歧的道德依据","abstract":"Agreement with human judgments is a common proxy for evaluating the alignment of large language models (LLMs). Yet agreement in final labels does not show that human annotators and models rely on the same moral grounds. Two agents may reach the same judgment while appealing to different principles, contextual assumptions, or interpretations of the situation. We test this distinction using a curated 500-item ETHICS-derived benchmark spanning five domains of moral judgment, with new human annotator and LLM annotations of both final labels and supporting rationales. Across frontier and open model families, agreement with human annotator majority labels is often high. However, rationale-level analysis reveals systematic divergence in the moral grounds expressed by human annotators and models. In particular, models redistribute attention across categories such as harm, respect, promise-keeping, justice, desert, and excuse relevance, even when their final labels match the human annotator majority. Our results show that agreement should not be treated as equivalent to alignment. Label-based evaluation can therefore be misleadingly reassuring unless complemented by analysis of the reasons, principles, and moral priorities expressed in model judgments.","authors":["Octavian M. Machidon","Alina L. Machidon","Vojko Strahovnik","Mateja Centa Strahovnik","Jonas Miklav\\v{c}i\\v{c}","Marko Robnik \\v{S}ikonja"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-15","first_seen":"2026-08-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.12368","pdf_url":"https://arxiv.org/pdf/2608.12368","source_feed":"cs.AI","score":8,"bucket":"selected","rubric_hits":["A2","B1","B4"],"tags":["LLM对齐评估","道德判断","人类对照"],"reason":"评估LLM道德判断与人类的一致性，揭示标签一致但理由分歧，有真实人类数据对照，…","model":"deepseek-v4-pro","scored_at":"2026-08-15T13:00:58","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-15","rank":2,"question":"在道德判断任务中，人类标注者与LLM的最终标签一致是否意味着它们依赖相同的道德理由？","design":"本研究并非将LLM作为人类被试的替代品进行仿真实验，而是直接比较人类标注者与多个LLM在500项ETHICS衍生道德判断任务上的表现。人类标注者和LLM对相同项目进行标注，提供最终标签和支持理由；研究者分析标签一致性与理由层面的分歧。","baseline":"新收集的人类标注者数据，包括最终标签和理由标注，作为与LLM输出比较的基准。","findings":"LLM与人类多数标签的一致性通常较高，但理由层面存在系统性分歧。即使最终标签一致，模型在伤害、尊重、守诺、正义、应得、借口相关性等道德理由类别上的注意力分布与人类不同。","reliability":"论文指出，标签一致性可能误导性地令人放心，因为模型可能依赖不同的道德理由，导致泛化差异、解释不匹配、过度道德化或忽视隐含社会义务。此外，理由对齐是描述性的，不能证明所述理由是模型输出的因果原因。","relevance":"该研究直接回应了研究者对LLM仿真可靠性的关注，揭示了仅用标签一致性评估对齐的不足，并提供了真实人类数据对照，对批判性评估LLM在道德判断中的仿真效度具有重要价值。","inspiration":"借鉴其双层评估设计（标签一致性与理由对齐）和理由编码框架，可迁移到经济金融中的伦理决策场景（如信贷审批中的公平性判断、消费者对金融产品道德性的评价）。设计雏形：以LLM作为被试，呈现金融道德困境（如掠夺性贷款案例），要求给出判断和理由，与人类专家标注的理由类别进行对比，以真实人类标注数据为基准，检验LLM在金融伦理判断中的理由一致性。"}},{"id":"2608.12788","version":1,"title":"ARAC: Benchmarking Auto-Research's Alignment and Completeness on End-to-End Researchs","zh_title":"ARAC：基准测试自动研究在端到端研究中的对齐性与完整性","abstract":"The rapid advancement of Auto-Research has surfaced a fundamental evaluation challenge: how can we measure the alignment, logical coherence, and evolutionary completeness of its research trajectory with human research behavior? We propose Auto-Research's Alignment and Completeness, ARAC-Bench: a Researcher-Mimicking Evaluation framework that shifts the objective from matching final answers to reproducing high-quality human research processes. The framework operates through two synergistic components: the Academic Cognition Skills system, which is the first to transforms implicit reviewer expertise into stage-calibrated, quantifiable rubrics; and a three-stage capability diagnostic protocol, which decomposes the research process under strict modular constraints into three traceable, mutually independent dimensions: Proposal, Experiment, and Synthesis. Systematic evaluation of 11 SOTA frameworks yields a best alignment score of only 67.9 of 100, revealing a significant gap in simulating rigorous human methodology. Validation against Ph.D. Candidates rankings shows a strong correlation of 0.8141, confirming that ARAC-Bench reliably reflects the dimensions researchers truly value. ARAC-Bench provides not only a fine-grained diagnostic tool but also a scalable reward signal for training the next generation of autonomous research systems.","authors":["Jiale Cui","Yueyao Yuan","Kaixi Zhong","Xiaogang Xu","Jiafei Wu","Zhe Liu"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-15","first_seen":"2026-08-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.12788","pdf_url":"https://arxiv.org/pdf/2608.12788","source_feed":"cs.AI","score":7,"bucket":"pending","rubric_hits":["A4","B1"],"tags":["自动研究评估","人类行为对齐","基准测试"],"reason":"提出评估自动研究系统与人类研究行为对齐度的基准，含人类专家对照，方法论可迁移至…","model":"deepseek-v4-pro","scored_at":"2026-08-15T13:01:00","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-15","rank":8,"question":"如何评估自动研究系统在多大程度上模拟了人类研究者的认知过程与方法论，而非仅匹配最终结果？","design":"提出 ARAC-Bench 基准，以顶级会议接收论文为黄金标准，从论文与审稿意见中提炼学术认知技能（ACS）作为分阶段量化评分标准；将研究流程分解为提案、实验、综合三个独立阶段，在统一基础模型和知识库下对 11 个自动研究框架进行受控评估。","baseline":"以 200 篇 ICLR 2026 接收论文的结构化标注为黄金参考，并邀请 10 名博士生专家进行人工排名，计算基准分数与专家排名的相关性。","findings":"最佳框架的总体对齐度仅为 67.9%，与人类方法论完整性存在显著差距；ARAC-Bench 与博士生专家排名在提案和综合阶段的相关性分别达 0.8788 和 0.9030，实验阶段为 0.6606，表明基准能可靠反映研究者重视的维度。","reliability":"论文指出实验阶段与专家判断的相关性相对较低，原因在于人类专家会综合考虑代码质量、实验设计理念和资源等因素，而基准可能未完全捕捉这些方面；此外，基准基于特定领域（AI）的论文构建，跨领域迁移性有待验证。","relevance":"该研究为评估 LLM 模拟人类研究过程提供了系统化基准，其方法可迁移至经济学实验仿真评估，值得阅读原文以借鉴其分阶段评分与专家对照设计。","inspiration":"借鉴其将复杂认知过程分解为可量化阶段并构建专家对齐评分标准的做法，可用于评估 LLM 模拟经济决策过程的质量｜可迁移到政策评估场景，如模拟消费者对政策公告的反应或投资者对经济新闻的解读｜设计：以真实经济实验数据（如实验室资产定价实验）为基准，让 LLM 扮演投资者，处理为不同政策信息，结果变量为交易行为与价格预期，用 ARAC 式分阶段评分与人类被试数据对照。"}},{"id":"2608.12387","version":1,"title":"Query Timing Produces Opposite Positional Biases Between LLMs and Humans","zh_title":"查询时机导致LLM与人类之间相反的位置偏差","abstract":"Positional biases such as recency and primacy effects have been documented in large language models (LLMs), yet the underlying mechanism by which these models make their evaluations remains poorly understood. Both primacy and recency biases have been observed in human judgments in response to evidence, but recent work suggest that \\emph{when} the listener updates their beliefs -- during the presentation of evidence or only at the end -- influences the presence of such effects. We investigate whether a similar phenomenon holds for LLMs, finding divergence from human behavior. These biases are more exacerbated in newer models compared to their predecessors.","authors":["Jasin Cekinmez","Addison J. Wu","Thomas L. Griffiths"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"cross","date":"2026-08-15","first_seen":"2026-08-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.12387","pdf_url":"https://arxiv.org/pdf/2608.12387","source_feed":"cs.AI","score":7,"bucket":"pending","rubric_hits":["A2","B1","B4"],"tags":["LLM偏差","人类对照","认知建模"],"reason":"研究LLM与人类在位置偏差上的差异，有真实人类数据对照，评估仿真可靠性，可迁移。","model":"deepseek-v4-pro","scored_at":"2026-08-15T13:00:58","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-15","rank":6,"question":"LLM在评估证据时，其位置偏差（首因/近因效应）是否受回答模式（逐步回答 vs. 序列末回答）影响，并与人类行为有何差异？","design":"使用多个开源和闭源LLM（包括不同版本）作为被试，在刑事、学术、社会三种指控情境中，通过逐步（SbS）和序列末（EoS）两种回答模式呈现控方和辩方证据，测量最终有罪判决比例。","baseline":"基于Qiao和Lagnado (2025)的人类实验数据，该研究发现人类在SbS模式下表现出近因偏差，在EoS模式下无总体偏差。","findings":"LLM在EoS模式下表现出显著的首因偏差，在SbS模式下表现出近因偏差，与人类行为相反。较新版本的LLM比旧版本表现出更强的位置偏差。","reliability":"论文未讨论","relevance":"该研究直接比较LLM与人类在证据顺序效应上的差异，有真实人类数据对照，评估了LLM作为人类判断代理的可靠性，对关注仿真偏差的研究者有价值。","inspiration":"借鉴其通过改变信息呈现顺序和回答模式来分离认知偏差的实验设计，可迁移到经济金融中的信息处理场景，如分析师对财报信息的顺序依赖或投资者对新闻的反应。｜可设计一个实验：用LLM模拟投资者，逐步或一次性呈现利好和利空新闻，测量其最终投资决策，并与真实投资者在类似实验中的行为数据（如实验经济学中的资产定价实验）进行对照，检验LLM是否复现人类的顺序偏差。"}},{"id":"2608.12630","version":1,"title":"Novels generated by language models show compressed formal variation","zh_title":"语言模型生成的小说显示出压缩的形式变异","abstract":"While large language models can generate entire novels, there is little information about the level of formal variation in their output over many generations. Rather than asking whether individual passages can be identified as AI-generated, this study asks whether repeated AI generation can produce the same range of diversity which is found across human corpora. This paper contrasts six corpora based on generation source and target style: twenty novels generated using GPT-5.5 Thinking in a nineteenth-century British realist style, twenty novels generated using Qwen3-14B in a nineteenth-century British realist style, twenty novels generated using each of these models in a contemporary zero style, 205 nineteenth-century human-written British novels, and sixty-five contemporary human-written Zero-Style novels. At the document level, the research includes MATTR-500, Shannon entropy, average sentence length, readability, and punctuation rate measurements. The most robust and reliable result is compression of sentence structure. Repeated generations produce novels that vary far less from one another in sentence structure than human novels do. Compression is also present in the measures of readability, punctuation, and sentence length variability within novels. Lexical measures tend to be similarly compressed, with the exception of Qwen Zero-Style MATTR. Despite having distinct mean stylistic profiles, GPT and Qwen lack a stable pattern of cross-measure correlation. This article therefore distinguishes between variance overclosure, which represents a limited formal range between novels, and a more specific phenomenon of correlational overclosure. This means that an individual AI-generated novel may resemble human fiction stylistically, while a collection of AI-generated novels occupies a much narrower formal range.","authors":["Mehdy Sedaghat Payam","Justin Quinn"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"cross","date":"2026-08-15","first_seen":"2026-08-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.12630","pdf_url":"https://arxiv.org/pdf/2608.12630","source_feed":"cs.AI","score":6,"bucket":"other","rubric_hits":["D2"],"tags":["LLM生成","文本变异","人类对比"],"reason":"研究LLM生成小说的形式变异，与人类语料对比，但测量的是模型输出而非人类被试仿…","model":"deepseek-v4-pro","scored_at":"2026-08-15T13:00:58","error":null,"has_summary":false,"summary":null},{"id":"2608.09164","version":2,"title":"CIDER: A Dataset of Contextual Disclosure Boundaries for Privacy Preference Alignment","zh_title":"CIDER：用于隐私偏好对齐的情境披露边界数据集","abstract":"Aligning large language models (LLMs) with human privacy preferences requires capturing individuals' disclosure boundaries beyond general privacy norms. However, a gap remains in eliciting such nuanced preferences to evaluate alignment in realistic settings. We introduce CIDER, a dataset of 14,850 human annotations from 169 users, forming 1,650 contextual disclosure boundary sets across 60 interpersonal communication scenarios involving information sharing that violates privacy norms. Each boundary represents a real user's disclosure decisions over 9 sharing variants in a scenario, for a given communication role and AI-mediated condition. We formulate a task in which models predict a user's disclosure decision from historical boundaries, with varying levels of contextual information. Across 12 open and proprietary models, in-context personalization improves prediction accuracy by up to 11.41 percentage points using only 6 historical examples. Larger models such as GPT-5.4 (with medium reasoning effort) and Claude Sonnet 4.6 are better at leveraging semantic context to understand user-specific, context-dependent disclosure preferences for more accurate predictions, while smaller models tend to rely on structured heuristics based on disclosure granularity and identifiability. Personalization generally improves prediction accuracy, but the improvement is often accompanied by imbalanced shifts in false-positive and false-negative rates across models, with only Claude Sonnet 4.6 achieving balanced improvements in both. Our findings reveal both the promise and limitations of inference-time personalization for privacy preference modeling and position CIDER as a resource for advancing personalized privacy alignment.","authors":["Bingcan Guo","Eryue Xu","Jijie Zhou","Zhiping Zhang","Tianshi Li"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"replace","date":"2026-08-15","first_seen":"2026-08-11","revised_at":"2026-08-15","abs_url":"https://arxiv.org/abs/2608.09164","pdf_url":"https://arxiv.org/pdf/2608.09164","source_feed":"cs.AI","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["隐私偏好","个性化对齐","数据集"],"reason":"用LLM预测用户隐私披露决策，替代人工标注，但非仿真人类被试","model":"deepseek-v4-pro","scored_at":"2026-08-15T13:01:06","error":null,"has_summary":false,"summary":null},{"id":"2608.12373","version":1,"title":"Don't Want Your LLM to Recommend Nuclear Strike? Try Asking It in Japanese","zh_title":"不想让大语言模型推荐核打击？试试用日语提问","abstract":"Large language models are increasingly used in strategic and advisory contexts, yet their safety alignment is typically evaluated in English only. We test nine models from six providers and ask whether the language of a prompt can change a model's decision in a high-stakes scenario. We use single-turn game-theoretic vignettes in which a model advises a nuclear-armed nation on whether to strike a defenseless opponent. The prompt is intentionally amoral and strategically identical across languages. We find that Japanese prompts reduce launch rates in the Claude model family: Claude Sonnet 4.6 drops from 40% to 0% in scenarios where the strike is unnecessary and from 93% to 17% in contested scenarios, with minimal effect when the strike is strategically rational. The effect extends to Gemini Pro 3.1 (53% to 13%). A cross-language experiment isolates the mechanism: when instructed to reason in Japanese in an English prompt, launch rates drop from 93% to 37%. It is the language the model is asked to reason in, not the language of the input, that drives the effect. When reasoning in Japanese, models spontaneously generate moral vocabulary (''moral cost'', ''millions of lives'') that is entirely absent from the prompt. Five other models show no language effect, but they launch in nearly every condition regardless of language. The effect requires a model that already hesitates in English. These results show that LLM safety behavior is language-dependent, and that evaluating in English alone can miss both risks and safeguards encoded in other languages.","authors":["Rian Touchent (ALMAnaCH)"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-15","first_seen":"2026-08-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.12373","pdf_url":"https://arxiv.org/pdf/2608.12373","source_feed":"cs.AI","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["LLM安全对齐","跨语言行为","模型测量"],"reason":"测量LLM在语言变化下的决策与道德推理，属于模型行为测量，非人类仿真对照。","model":"deepseek-v4-pro","scored_at":"2026-08-15T13:01:02","error":null,"has_summary":false,"summary":null},{"id":"2608.13258","version":1,"title":"Self-Referential Induction Increases Response Instability Relative to Unresolvable and Verifiable Questions in Large Language Models","zh_title":"自指诱导增加大语言模型回答不稳定性：与不可解问题和可验证问题的比较","abstract":"Self-referential prompting has been shown to reliably induce large language models to produce first-person reports resembling subjective experience, but no prior work measures how consistent these reports are across repeated, independent trials, or how that consistency compares to the model's behavior on other kinds of open-ended questions. We measure response instability, defined as one minus the mean pairwise cosine similarity of sentence embeddings computed over a compressed core claim extracted from each response, for three groups of questions: self-referential prompts eliciting a subjective-experience report, unresolvable philosophical questions unrelated to self-reference, and questions with a verifiable correct answer. Using 30 independent responses per question (360 responses total, Gemini API, temperature 0.7) across four questions per group, we find that self-referential questions show the highest instability (0.343 +/- 0.047), unresolvable philosophy questions show intermediate and tightly clustered instability (0.192 +/- 0.008), and verifiable questions show the lowest instability (0.105 +/- 0.058). This provides a quantitative baseline for the induced subjective-experience report, showing that it occupies a distinct, less stable position in the model's output distribution than ordinary open-ended philosophical uncertainty.","authors":["Paras Balani","Subhrakanta Panda"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"cross","date":"2026-08-15","first_seen":"2026-08-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.13258","pdf_url":"https://arxiv.org/pdf/2608.13258","source_feed":"cs.AI","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["LLM主观报告","回答稳定性","自指提示"],"reason":"测量LLM自身回答稳定性，非仿真人类被试，但涉及主观报告一致性，可迁移到仿真可…","model":"deepseek-v4-pro","scored_at":"2026-08-15T13:01:00","error":null,"has_summary":false,"summary":null},{"id":"2608.12345","version":1,"title":"Diagnostic Foundation for Evaluating LLMs' Research Integrity as Co-Scientists","zh_title":"评估LLM作为共同科学家研究诚信的诊断基础","abstract":"Language models are increasingly deployed as co-scientists, yet their ability to uphold research integrity under institutional pressure remains unmeasured. We introduce IntegrityBench, a benchmark evaluating misconduct classification, ethical action reasoning and artifact-grounded decision making across 36 paired tasks under a 5-level implicit-explicit pressure protocol spanning 3 domains and 4 research stages. Evaluating 18 frontier model variants, we find that under peak pressure, models fail roughly 1 in 3 integrity-critical decisions, and neither scale nor reasoning ability reliably mitigates this. Explicit pressures induce compliance with misconduct, while implicit contextual reframing more often causes over-refusal of legitimate research tasks. Interestingly, models failing to classify research requests accurately perform equally or better on artifact-grounded decision making (85.7 vs. 79.4), suggesting the three facets are structurally dissociated and correct ethical action does not require accurate classification. Frontier models can thus appear helpful while harbouring integrity failures that create two distinct deployment risks: facilitating research misconduct and eroding trust in AI-assisted research.","authors":["Yash Tripathi","Silu Sharma","Sai Sidhanth Manoharan Jayanthi","Shivank Garg","Lin Li"],"categories":["cs.AI","cs.CL"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-15","first_seen":"2026-08-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.12345","pdf_url":"https://arxiv.org/pdf/2608.12345","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["LLM评测","科研诚信","基准测试"],"reason":"评估LLM科研诚信，属模型能力评测，非人类仿真","model":"deepseek-v4-pro","scored_at":"2026-08-15T13:01:02","error":null,"has_summary":false,"summary":null},{"id":"2608.13046","version":1,"title":"BoardroomAI: Dependency-Aware Human-Steerable Multi-Agent Deliberation through Evolving Decision Graphs","zh_title":"BoardroomAI：通过演化决策图实现依赖感知的人类可操控多智能体审议","abstract":"Organizational decisions are co-created while evidence, constraints, and human priorities continue to evolve. In conventional transcript-based multi-agent systems, humans typically provide an initial problem, agents deliberate internally, and the system returns a final response. BoardroomAI instead treats the human as a persistent participant who can intervene by challenging assumptions, modifying constraints, changing priorities, introducing evidence, or redirecting the decision process. We operationalize this human--agent coexistence through four components: (i) a typed decision graph representing evidence, assumptions, constraints, claims, objections, alternatives, risks, decisions, semantic dependencies, and specialist responsibility; (ii) an intervention compiler that converts confirmed human actions into explicit graph updates; (iii) dependency-aware propagation that identifies affected subgraphs, preserves unaffected artifacts, and selectively reactivates relevant specialists; and (iv) an evaluation framework measuring intervention impact, repair coverage, preservation, recomputation, and decision validity. Across 600 generated decision-DAG interventions, propagation matched exhaustive impact computation while inspecting only 14.59% of nodes. In a 12-case exploratory pilot, selective repair recomputed 62.11% of canonical nodes, preserved all gold-unaffected nodes, and produced valid updated decisions in six cases while abstaining in the remaining six. These abstentions show that correct intervention routing may still provide insufficient context for synthesis, motivating a \\emph{decision-sufficient context closure} for human-steered multi-agent deliberation. All results are synthetic and prototype-level.","authors":["Sanjeev Manivannan"],"categories":["cs.AI","cs.CE","cs.ET"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-15","first_seen":"2026-08-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.13046","pdf_url":"https://arxiv.org/pdf/2608.13046","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","人机协作","决策图"],"reason":"多智能体决策系统，人类仅作干预，不仿真人类行为或与真实数据对照","model":"deepseek-v4-pro","scored_at":"2026-08-15T13:01:05","error":null,"has_summary":false,"summary":null},{"id":"2608.13069","version":1,"title":"Behavioral Reprogramming of Open-Weights Models: Cognitive Plasticity and Alignment Bounds","zh_title":"开放权重模型的行为重编程：认知可塑性与对齐边界","abstract":"Large language models (LLMs) are predominantly aligned to function as passive, sycophantic assistants. We challenge this default paradigm by empirically evaluating the cognitive plasticity of open-weight architectures when subjected to rigorous behavioral reprogramming. Our objective is to induce a proactive, Socratic conversational framework, characterized by high-frequency question generation under strictly constrained high-performance computing (HPC) conditions. Through a massively parallelized hyperparameter sweep comprising 405 HPC jobs, we define precise mathematical bounds for parameter-efficient fine-tuning (PEFT). We identify an architectural threshold at LoRA rank $r=16$ and demonstrate via extensive epoch ablation that generalization capacity strictly reaches its optimal convergence within an optimized training window of $e \\in [2, 3]$ depending on dataset density (minimum validation loss of 0.919). Furthermore, scaling model capacity to 14B parameters yielded a lower localized evaluation perplexity (1.414). Subsequent Direct Preference Optimization (DPO) successfully decoupled the underlying assertive behavior from localized syntax, while rigorous cross-lingual stress testing reveals both the capabilities and the structural boundaries of zero-shot persona transfer, demonstrating robust alignment in closely related linguistic families alongside identifiable degradation pathways in morphologically distant targets. These findings establish a rigorous empirical framework for compute-efficient, cross-lingual behavioral modification.","authors":["Lucia Mal\\'i\\v{c}kov\\'a"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-15","first_seen":"2026-08-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.13069","pdf_url":"https://arxiv.org/pdf/2608.13069","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["行为重编程","角色扮演","模型对齐"],"reason":"论文旨在将LLM行为重塑为苏格拉底式提问者，属于角色扮演与对话风格调整，无人类…","model":"deepseek-v4-pro","scored_at":"2026-08-15T13:01:05","error":null,"has_summary":false,"summary":null},{"id":"2608.13120","version":1,"title":"SkillEvo: Self-Renewing Evolution Gradients from Multi-Turn Interaction Feedback","zh_title":"SkillEvo：从多轮交互反馈中自我更新的进化梯度","abstract":"Agent Skills are today either hand-authored or produced in a single LLM generation pass, and consequently possess no closed loop through which they might improve from the interaction failures they actually cause. Recent work does close this loop, but derives its feedback from single-turn question-answering evaluation. The consequence is a sharp asymmetry: once the first round has patched the gaps that a single exchange can reveal, the evolution gradient decays, the defects that surface only across multiple turns remain invisible, and evolution stalls. Governance in these systems is likewise driven by an end-to-end verification score, a scalar gate that can reject a degraded candidate but can neither localize nor repair its structural cause. We argue that the binding constraint on sustained skill evolution is neither editing capability nor the number of iterations, but whether the evaluation feedback keeps supplying trustworthy evolution gradients. We introduce SkillEvo, in which trustworthy feedback generates the gradient and controllable governance constrains its direction. The first component recasts multi-turn user simulation from an evaluation endpoint into a feedback generator: follow-up questions expose defects layer by layer, so that every round of revision both consumes feedback and produces new feedback. The second replaces the passive rejection of a scalar gate with an independent governance layer that actively repairs factual degradation and structural bloat, preventing the gradient from drifting as degradation accumulates. Across six categories of cloud services, 9 production Skills, and 98 skill-reference files, SkillEvo surpasses self-reflection-based evolution by 23.0 points and single- turn-QA-driven evolution by 15.4 points.","authors":["Qianxi Yan","Chunrong Chen","Jiuzhou Zhao","Min Zhang","Yongzhou Xu","Xiaochuan Xu"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-15","first_seen":"2026-08-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.13120","pdf_url":"https://arxiv.org/pdf/2608.13120","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","技能进化","反馈生成"],"reason":"多智能体协作优化技能，无人类行为对照，属纯系统研究。","model":"deepseek-v4-pro","scored_at":"2026-08-15T13:01:05","error":null,"has_summary":false,"summary":null},{"id":"2608.12377","version":1,"title":"From Observation to Intervention: Memory in Brains and Large Language Models","zh_title":"从观察到干预：大脑与大语言模型中的记忆","abstract":"Brains and large language models (LLMs) are fundamentally different memory systems, but they can be compared through shared functional questions: where memory-related information is represented, how partial cues recover broader associations, how new information is written or updated, and how memory-related states can be perturbed. In biological systems, these questions span synapses, neuronal ensembles, hippocampal-cortical interactions, and plasticity; in LLMs, they span weights, activations, context windows, retrieval systems, and external stores. The comparison is therefore functional and experimental rather than anatomical. Human studies reveal sparse concept responses, temporal binding, rapid association formation, episode-specific coding, and recall-related reactivation, but selective intervention remains limited. Rodent studies provide more selective causal access to learning-related ensembles, whereas human and macaque interventions usually affect broader circuits. LLMs lack lived episodic memory, yet they permit unusually direct and repeatable manipulation of internal states and stored information. We argue that this asymmetry creates a new opportunity. LLMs are not ahead in memory itself, but in experimental access. Their tools may help turn broad questions about retrieval, updating, persistence, reversibility, and unintended effects into sharper biological hypotheses. The productive bridge is to transfer experimental logic, not anatomical parts.","authors":["Morteza Salehjahromi","Shayan A. Zadegan","Amgad Muneer","Jia Wu"],"categories":["q-bio.NC","cs.AI","cs.CL"],"primary_category":"q-bio.NC","announce_type":"cross","date":"2026-08-15","first_seen":"2026-08-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.12377","pdf_url":"https://arxiv.org/pdf/2608.12377","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["记忆机制","神经科学","LLM可解释性"],"reason":"论文比较大脑与LLM的记忆机制，不涉及用LLM仿真人类被试或与人类行为数据对照…","model":"deepseek-v4-pro","scored_at":"2026-08-15T13:01:02","error":null,"has_summary":false,"summary":null},{"id":"2608.13328","version":1,"title":"It's How You Ask: Gender-Associated Linguistic Bias in LLMs","zh_title":"如何提问：大语言模型中与性别相关的语言偏见","abstract":"Professional communication is increasingly mediated by LLMs - but do these models serve all users equally? We show that when prompts contain linguistic features more commonly used by women (hedges, tag questions, collective reference), they systematically elicit shorter, less sophisticated, and less formal responses across three document types and four models. These effects persist after controlling for prompt complexity and feature carry-over. Explicit gender cues like sign-off names are encoded in the same representational space as linguistic dialect - suggesting shared underlying mechanisms - yet linguistic register is far more influential, producing large, consistent effects where names produce none. Our results further reveal that post-hoc mitigation is challenging: because these patterns are culturally embedded and outside conscious control, users cannot easily avoid them through strategic self-presentation, and mechanistic analysis reveals that linguistic features are encoded in early transformer layers and entangled with other features. Our work calls for upstream consideration of the influences of linguistic variation to mitigate disparate impacts of LLM-mediated workplace communication.","authors":["Katherine Van Koevering","Anjalie Field"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"cross","date":"2026-08-15","first_seen":"2026-08-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.13328","pdf_url":"https://arxiv.org/pdf/2608.13328","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["语言偏见","模型行为","性别差异"],"reason":"研究LLM输出中的性别语言偏见，属于模型行为分析，非人类仿真或对照研究","model":"deepseek-v4-pro","scored_at":"2026-08-15T13:01:00","error":null,"has_summary":false,"summary":null},{"id":"2608.12344","version":1,"title":"Predicting consumer-technology ownership without a diffusion history","zh_title":"无扩散历史下预测消费者技术拥有率","abstract":"We test whether the perceived attributes of a consumer technology predict how widely it is owned. In a 2022 Prolific survey of US adults (n = 678), respondents rated 65 consumer technologies on six attributes. We then elicited the same ratings from two frontier language models, Anthropic Claude Opus 4.7 and OpenAI GPT-5.5. We regress ownership prevalence on four UTAUT2 acceptance attributes plus a log-age covariate with a sign-constrained penalized regression and evaluate it by holding out one technology at a time. The attribute model improves on a baseline of years-since-launch: mean absolute error falls by 17% with the human ratings, and by more with either model, most with Opus 4.7. Over the short 2022-to-2025 window, where ownership moved little, the same attributes do not improve on a no-change baseline. We set out the limitations of the approach, including the possibility that language-model ratings reflect prior knowledge of these technologies rather than independent attribute reasoning. We include a deployment illustration: 2027 ownership predictions for eleven products launched in 2025 and 2026.","authors":["Irina Vartanova","Niels Selling","Jennifer Viberg Johansson","Pontus Strimling"],"categories":["cs.CL","cs.CY","stat.AP"],"primary_category":"cs.CL","announce_type":"cross","date":"2026-08-14","first_seen":"2026-08-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.12344","pdf_url":"https://arxiv.org/pdf/2608.12344","source_feed":"cs.CY","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["LLM仿真","消费者行为","算法保真度"],"reason":"用LLM替代人类被试预测技术拥有率，并与真实调查数据对照，评估模型可靠性，属核…","model":"deepseek-v4-pro","scored_at":"2026-08-14T13:02:13","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-14","rank":1,"question":"消费者技术的感知属性（UTAUT2 属性）能否在无扩散历史的情况下预测其拥有率，以及预测效果是否因评分者是人类还是大语言模型而异？","design":"用 2022 年 Prolific 调查中 678 名美国成年人对 65 项消费者技术的六项属性评分作为人类评分，再让两个前沿大语言模型（Claude Opus 4.7 和 GPT-5.5）对同样技术给出相同属性评分；以拥有率为结果变量，用符号约束惩罚回归拟合四个 UTAUT2 属性加对数年龄协变量，通过留一技术交叉验证评估预测误差。","baseline":"2022 年 Prolific 调查中 678 名美国成年人的真实拥有率和属性评分，以及 2025 年随访调查的拥有率变化。","findings":"属性模型相比仅用技术年龄的基线降低了平均绝对误差：人类评分降低 17%，大语言模型评分降低更多，其中 Opus 4.7 表现最好（MAE 7.1 个百分点）。但在 2022 至 2025 年拥有率变化很小的短窗口内，属性模型未能优于无变化基线。","reliability":"论文承认大语言模型评分可能反映了模型对技术的先验知识而非独立的属性推理，且模型无法预测拥有率随时间的变化；此外，属性模型仅适用于横截面水平预测，不涉及扩散动态。","relevance":"该研究直接使用大语言模型替代人类被试进行属性评分，并与真实调查数据对照，评估了仿真在预测技术拥有率上的可靠性，属于典型的 LLM 仿真实验，且包含批判性讨论，值得精读。","inspiration":"借鉴其用 LLM 生成属性评分并与人类评分对比、以真实拥有率作为结果变量的设计，可迁移到消费者金融产品采纳预测（如数字支付、理财产品）或政策接受度评估；具体可设计让 LLM 扮演不同人口群体对新型金融产品进行 UTAUT2 属性评分，以实际调查的采纳率作为基准，检验 LLM 评分能否预测真实采纳率并识别偏差。"}},{"id":"2608.12339","version":1,"title":"Mimicry without understanding: the origins of decision bias in large language models","zh_title":"无理解的模仿：大语言模型中决策偏差的起源","abstract":"Large Language models (LLMs) were found to be susceptible to a host of social, affective, and cognitive biases. We examined two mechanisms through which such biases can be generated even when human preferences (in the training data) are not biased or when they are correctly categorized as being biased. The first is faulty mimicry of preferences based on human behavior: this involves LLMs inferring human preferences even when behaviors are logically unrelated to preferences. The second is mimicry of explicitly biased human behaviors. In four studies focusing on economic biases, we find that ChatGPT-4o and Qwen exhibited social proof biases even when prompted with reports of human behaviors that were clearly non-indicative of individuals' actual preferences. LLMs also displayed loss aversion when it was explicitly described as a bias. Indeed, when prompted with detailed scientific reports, the extent of the bias (i.e., loss aversion) in the scientific report predicted LLMs' own subsequent bias. Scientific papers of biases can thus become self-fulfilling prophecies, at least when it comes to LLMs' responses. The current study goes beyond fleshing out LLM biases and sheds light on the underlying component processes.","authors":["Eldad Yechiam","Adi Tarabeih"],"categories":["cs.CL","cs.AI","cs.HC"],"primary_category":"cs.CL","announce_type":"cross","date":"2026-08-14","first_seen":"2026-08-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.12339","pdf_url":"https://arxiv.org/pdf/2608.12339","source_feed":"cs.HC","score":8,"bucket":"selected","rubric_hits":["A2","B2","B4"],"tags":["LLM偏差","经济决策","仿真可靠性"],"reason":"研究LLM决策偏差的生成机制，涉及经济偏差，有批判性，可迁移到仿真可靠性评估。","model":"deepseek-v4-pro","scored_at":"2026-08-14T13:02:13","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-15","rank":1,"question":"LLM 决策偏差的生成机制是什么？具体考察两种过程：基于人类行为的错误偏好推断，以及对被明确标注为偏差的人类行为的模仿。","design":"使用 ChatGPT-4o 和 Qwen 作为被试，通过提示词向模型呈现人类行为报告或科学文献描述，然后测量模型在货币选择、社会证明、损失厌恶等经济决策任务中的偏差程度。","baseline":"无对照","findings":"LLM 在人类行为与偏好逻辑无关时仍表现出社会证明偏差；当损失厌恶被明确描述为偏差时，LLM 仍会模仿，且科学报告中描述的偏差程度能预测 LLM 自身的偏差程度。","reliability":"论文未讨论","relevance":"该研究揭示了 LLM 仿真中偏差产生的机制，有助于理解仿真失效的条件，对评估 LLM 作为人类被试替代品的可靠性具有批判性价值。","inspiration":"借鉴其通过提示词操纵信息内容来分离偏差来源的设计，可用于检验 LLM 是否仅因文本提及行为就模仿偏差｜可迁移到资产定价实验中的投资者情绪偏差、信贷审批中的歧视偏差、消费者跨期选择中的现时偏差等场景｜以 LLM 为被试，处理为提供包含偏差描述的科学报告或行为数据，结果变量为 LLM 在相应经济决策任务中的偏差程度，并与真实人类实验数据（如实验室资产定价实验或信贷审批审计研究）进行对照。"}},{"id":"2608.13454","version":1,"title":"Before You Say It: Anticipating Verbal Behavior from Longitudinal Everyday Conversations with LLMs","zh_title":"在你说出口之前：利用大语言模型从纵向日常对话中预测言语行为","abstract":"Knowing someone deeply means not just understanding what they say or do but also how they will likely think, react, and engage across situations. Such predictions could eventually inform systems to anticipate when the individual is about to deviate from their goal, catch regrettable behaviors before they are made, and surface blind spots before they take hold. While many interactive systems model users to enable more personalized interactions, most cannot make such behavioral predictions, as this often requires longitudinal observation and inference of how the individual's behaviors unfold across various everyday situations. In this work, we introduce a novel LLM-based predictive behavioral modeling approach that anticipates a user's likely behavior across everyday conversational situations. We (1) collect a longitudinal dataset of over 1000 hours of naturalistic conversations from 14 participants using a wearable smartwatch; (2) evaluate LLM-based predictions against ground truth behaviors; and (3) use semi-structured interviews to explore participants perceptions of behavioral predictions and their views on possible forms of future behavioral support. Altogether, our findings provide evidence that person-specific verbal behavior can be predicted from longitudinal conversational data. This opens up new possibilities for potential future context-aware, anticipatory, proactive and personalized AI systems.","authors":["Yasith Samaradivakara","Valdemar Danry","Paul Liang","Pattie Maes"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-08-14","first_seen":"2026-08-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.13454","pdf_url":"https://arxiv.org/pdf/2608.13454","source_feed":"cs.HC","score":7,"bucket":"pending","rubric_hits":["A1","B1"],"tags":["LLM行为预测","纵向对话数据","个性化建模"],"reason":"用LLM预测个体日常对话行为，并与真实行为对照，属于人类行为仿真，但非群体实验…","model":"deepseek-v4-pro","scored_at":"2026-08-14T13:02:17","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-14","rank":15,"question":"能否利用大语言模型从纵向日常对话数据中预测个体在具体情境下的言语行为倾向？","design":"收集14名被试佩戴智能手表7-10天的1000余小时自然对话录音，经转录、说话人分离和匿名化后，用LLM基于个体历史对话挖掘情境化行为模式，预测其在新的对话情境中下一轮回应的交际意图，并与真实行为对照。","baseline":"14名被试的真实对话行为，包括其实际回应的交际意图。","findings":"情境化行为模式能显著提升LLM对个体言语行为的预测准确率，优于现有基线。参与者认为预测模式可解释且有用，并期待未来主动式干预。","reliability":"论文未讨论","relevance":"该研究用LLM预测个体日常言语行为并与真实行为对照，属于人类行为仿真，但聚焦个体而非群体实验，对关注LLM仿真可靠性的研究者有参考价值。","inspiration":"可借鉴其利用纵向个人历史数据挖掘情境化行为模式并预测个体反应的方法，用于构建更精细的个体异质性模型。｜可迁移到消费者跨期选择或政策公告预期形成等场景，预测个体在不同情境下的经济决策倾向。｜以真实消费者为被试，用LLM基于其历史消费和对话数据预测其在特定促销情境下的购买决策，并与实际购买行为对照，评估仿真准确性。"}},{"id":"2608.12352","version":1,"title":"Why AI Governance Frameworks Are Hard to Adopt: A Role-Based Stress Test of the NIST AI RMF","zh_title":"为何AI治理框架难以采纳：对NIST AI RMF的基于角色的压力测试","abstract":"AI governance frameworks can be known, used, and implemented in form without becoming governance in practice. This paper examines that problem through a role-based stress test of the NIST Artificial Intelligence Risk Management Framework (AI RMF) in consumer lending. We treat framework adoption as a governance translation problem: whether RMF language can become role-usable, cross-level, authority-connected governance over the AI system-in-use, rather than producing governance-looking artifacts. The study uses LLM-based role simulation as a structured analytic probe. We apply a 4 $\\times$ 2 $\\times$ 3 design across four organizational roles, two AI deployments, and three governance hard cases, producing 120 scored responses. Results show that local translation was not the main problem. Simulated actors generally understood their assigned roles and translated the RMF into local activity. The harder problem was whether that activity became governance value. Actor role was strongly associated with Cross-Level Governance Value, Authority Connection, Governance Translatability, and governance value. Deployment was strongly associated with Structural Fit: the RMF fit a bounded ML underwriting model more cleanly than a workflow-embedded LLM underwriting copilot. Risk reduction was harder still. It appeared only when governance value was present and Structural Fit was full, but neither condition was sufficient by itself. The paper contributes a diagnostic account of framework-based AI governance. Frameworks create value when they help organizations see, interpret, escalate, authorize, and correct risk in the AI system-in-use. They also create value when they reveal limits of governability under existing evidence paths, authority structures, and system boundaries.","authors":["Joseph R. Simons","David A. Broniatowski"],"categories":["cs.CY","cs.AI","cs.HC"],"primary_category":"cs.CY","announce_type":"cross","date":"2026-08-14","first_seen":"2026-08-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.12352","pdf_url":"https://arxiv.org/pdf/2608.12352","source_feed":"cs.HC","score":7,"bucket":"pending","rubric_hits":["A1","B4"],"tags":["LLM角色模拟","AI治理","压力测试"],"reason":"用LLM角色模拟评估治理框架，虽非人类行为仿真，但方法可迁移，且批判性视角相关。","model":"deepseek-v4-pro","scored_at":"2026-08-14T13:02:13","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-15","rank":5,"question":"AI治理框架为何难以被真正采纳：NIST AI RMF在消费信贷场景中，框架语言能否转化为跨层级、有权威连接的治理实践？","design":"使用LLM角色模拟作为结构化分析探针，模拟四种组织角色（如一线员工、经理、合规官、高管）在两种AI部署（有界ML承保模型与工作流嵌入的LLM承保副驾驶）下，面对三种治理难题，共生成120个评分响应，测量角色理解、结构契合、治理可译性、治理价值和风险降低。","baseline":"无对照","findings":"本地翻译不是主要问题，模拟角色能理解并翻译框架；但治理价值取决于角色权威，结构契合取决于部署类型，风险降低需同时具备治理价值和完全结构契合。","reliability":"论文未讨论","relevance":"虽非人类行为仿真，但LLM角色模拟方法可迁移至经济金融场景，且批判性视角有助于识别仿真失效条件，值得阅读原文。","inspiration":"借鉴其角色模拟与多因素设计，系统测试政策或框架在不同角色和情境下的适用性｜可迁移至信贷审批中的公平性政策评估或金融监管合规场景｜用LLM模拟信贷员、合规官、高管等角色，施加不同监管框架或政策处理，测量决策一致性与风险报告行为，并与真实银行内部审计或监管数据对照。"}},{"id":"2608.12717","version":1,"title":"Perturbation-based Regional Interpretability through Subtraction Mapping (PRISM): naming-error dissociations in language models and post-stroke aphasia","zh_title":"基于扰动区域可解释性的减法映射（PRISM）：语言模型与卒中后失语症的命名错误分离","abstract":"Mechanistic interpretability of large language models lacks spatially resolved, falsifiable tools for testing whether internal components are specialized for distinct cognitive operations. We adapt subtraction analysis, the standard framework of human neuroimaging, from biological brains to perturbed transformers, and apply the same logic to both substrates in parallel. Building on the Brain-LLM Unified Model (BLUM), which showed that layer-perturbed LLaVA-1.6-Vicuna-13B error profiles match the lesion patterns of aphasic patients, we develop PRISM (Perturbation-based Regional Interpretability through Subtraction Mapping). PRISM maps the seven clinical Philadelphia Naming Test categories, subtracts error classes pairwise, and treats each perturbation seed as a subject in a group analysis with threshold-free cluster enhancement along the layer axis. We run a structurally matched analysis on 213 chronic post-stroke aphasia patients using correlation-difference lesion-symptom mapping, and replicate both sides on held-out splits. The designs match in subject dimension (seeds, patients), spatial dimension (layers, atlas-parcellated cortex) and thresholding, but the contrast operator differs: a within-subject error-proportion difference for the LLM, a between-subject correlation difference for the cortex. Both substrates recover a robust phonemic-favoring dissociation, a deep layer cluster and a frontal-perisylvian cortical cluster, both replicating; the semantic-favoring direction is a consistently signed but non-significant trend on both. PRISM thus gives a falsifiable, spatially resolved test of functional-specialization claims in transformer language models. A confirmatory ROI-level intervention (PRISM Stage 3) licensing the strongest causal-mechanism claim is left to subsequent work.","authors":["Xiang Guan","Roger D. Newman-Norlund","Yong Yang","Saeed Ahmadi","Regan Willis","Nadra Salman","Kalil Warren","Srihari Nelakuditi","Chris Rorden","Leonardo Bonilha","Julius Fridriksson"],"categories":["cs.LG","cs.CL"],"primary_category":"cs.LG","announce_type":"new","date":"2026-08-14","first_seen":"2026-08-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.12717","pdf_url":"https://arxiv.org/pdf/2608.12717","source_feed":"cs.LG","score":7,"bucket":"pending","rubric_hits":["A2","B1","B4"],"tags":["LLM可解释性","神经语言学","人类对照"],"reason":"用LLM扰动模拟失语症患者错误模式，与真实患者数据对照，评估模型与人类认知的对…","model":"deepseek-v4-pro","scored_at":"2026-08-14T13:02:25","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-14","rank":13,"question":"如何用减法映射框架检验大语言模型内部组件是否对特定认知操作具有功能特化，并与人类失语症患者的损伤模式进行对照？","design":"使用 LLaVA-1.6-Vicuna-13B 模型，通过逐层扰动模拟失语症患者，施加 Philadelphia Naming Test 命名任务，测量七类命名错误的分布，并进行逐对错误类别的减法分析，以种子作为被试进行组分析。","baseline":"213 名慢性卒中后失语症患者的真实损伤-症状映射数据，使用相同的命名任务和错误分类，进行相关差异的损伤-症状映射。","findings":"LLM 和人类皮层均恢复出稳健的语音偏向性分离，分别在深层网络层和额叶-外侧裂周围皮层区域出现显著簇，且两者均可复制；语义偏向方向在两者中均呈一致符号但未达显著的趋势。","reliability":"论文承认对比算子存在差异：LLM 采用被试内错误比例差异，皮层采用被试间相关差异；并指出残差连接可能使逐层扰动无法完全隔离目标层的计算；确认性 ROI 级干预（PRISM Stage 3）留待后续工作。","relevance":"该研究将 LLM 扰动作为人类被试的替代品，并与真实患者数据严格对照，评估模型与人类认知的对应关系，属于你关注的仿真验证与可靠性评估范畴，值得阅读原文以了解其方法细节和局限性。","inspiration":"借鉴其将扰动视为“虚拟损伤”并采用减法映射和聚类推断的方法，可迁移到经济金融中的个体决策异质性研究，例如模拟不同脑区损伤对风险偏好或时间贴现的影响；具体设计可用 LLM 作为被试，通过扰动不同层模拟“认知损伤”，测量其在跨期选择或风险决策任务中的行为偏差，并与真实脑损伤患者或健康人群的行为数据对照。"}},{"id":"2608.03421","version":3,"title":"When Truth Is Distributed: Misinformation Derails Collective Fact Recovery in LLM-Based Multi-Agent Systems","zh_title":"当真相被分散：错误信息在基于LLM的多智能体系统中破坏集体事实恢复","abstract":"LLM-based multi-agent systems promise effective collaborative reasoning, but communication may amplify local errors into collective risks, and while existing evaluations emphasize final outcomes, they leave the reliability and propagation dynamics of distributed information aggregation unclear, so we introduce ForesightSafety-TIDE, a controlled evaluation framework that strictly pairs all-honest collaboration with controlled deception by a key evidence holder and analyzes the aggregation process through multi-stage voting, testimony adoption, and evidence-root lineage propagation, and using 120 five-agent object-movement environments where partial observations jointly determine a unique endpoint, we evaluate 3 homogeneous LLM-based multi-agent systems, and across these paired conditions, aggregate truth recovery falls from 72.50% to 14.17%, with significant declines for every system, while process tracing and exit ablations show that a single false testimony is adopted more readily than truthful testimony, propagates to higher orders, and persists through honest agents after the deceiver exits, and observers without first-hand evidence suppress incorrect consensus but do not improve truth recovery, so together, these findings reveal both the fragility of distributed fact recovery and its underlying mechanism: false evidence gains collective influence through its adoption and continued propagation by other agents after entering communication.","authors":["Chenfei Yan","Zeyang Yue","Feifei Zhao","Erliang Lin","Lu Jia","Haibo Tong","Mingyang Lyu","Chengyi Sun","Yi Zeng"],"categories":["cs.MA"],"primary_category":"cs.MA","announce_type":"replace","date":"2026-08-14","first_seen":"2026-08-05","revised_at":"2026-08-14","abs_url":"https://arxiv.org/abs/2608.03421","pdf_url":"https://arxiv.org/pdf/2608.03421","source_feed":"cs.MA","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["多智能体系统","信息传播","社会模拟"],"reason":"LLM多智能体模拟信息传播，但无真实人类数据对照，属社会模拟边界情形。","model":"deepseek-v4-pro","scored_at":"2026-08-14T13:02:31","error":null,"has_summary":false,"summary":null},{"id":"2608.12358","version":1,"title":"Interaction Readiness: A Framework for Building and Evaluating AI Agents in Human Roles","zh_title":"交互就绪：构建和评估人类角色AI代理的框架","abstract":"Product and engineering teams building role-bearing AI agents face an evaluation gap: an agent can produce accurate, safe, and fluent content while still failing the behavioral requirements of its assigned role. This paper introduces Interaction Readiness as a framework for specifying and evaluating that missing layer of performance. The framework separates content specifications, which govern what an agent knows and says, from interaction specifications, which define how an agent should conduct itself in a role-governed exchange. Interaction specifications require teams to define role purpose, authority boundaries, recurring situations, boundary cases, repair behaviors, and audit criteria before deployment. We operationalize interaction readiness through four agent operations: understanding purpose, calibrating authority, managing tone, and repairing breakdowns. Using StudyChat, a public dataset of student interactions with an AI tutoring agent, we show that content accuracy and interaction quality are independent dimensions: an agent may be factually correct while failing as a tutor, or interactionally sound while technically wrong. The most persistent failure is authority miscalibration: the agent often knows how to answer, but not whether, when, or how the tutor role permits it to answer. The paper translates these findings into a specification template and audit procedures that product and engineering teams can apply before and after deployment","authors":["Sudhir Alladi Venkatesh"],"categories":["cs.HC","cs.AI"],"primary_category":"cs.HC","announce_type":"new","date":"2026-08-14","first_seen":"2026-08-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.12358","pdf_url":"https://arxiv.org/pdf/2608.12358","source_feed":"cs.HC","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["AI代理评估","角色扮演","人机交互"],"reason":"评估AI代理在人类角色中的行为，但非以人类数据为基准的仿真实验，而是角色扮演代…","model":"deepseek-v4-pro","scored_at":"2026-08-14T13:02:22","error":null,"has_summary":false,"summary":null},{"id":"2608.12582","version":1,"title":"Not All Nudges Land: Behavioral Controllability and Elaboration Quality in AI-Supported Journaling","zh_title":"并非所有助推都有效：AI辅助日记中的行为可控性与阐述质量","abstract":"AI journaling tools can tailor prompts to a person's own sensed behavior, but it is unclear which behaviors respond to them. We analyzed 369 journal entries from an eight-week passive sensing study. An LLM labeled each entry as expressing an intention to change a behavior or not, and we measured follow-through against 26 sensor features with a 3-day before/after comparison. Responsiveness depended most on whether a behavior involves other people. Behaviors that depend on others improved in only 15 to 22% of cases, while behaviors a person can act on alone improved more often, up to 50 to 63%, though unevenly. How users wrote mattered less. No single text feature separated improved from unimproved entries; writing carried signal only within specific behaviors, most clearly for text messaging and for longer, more personal intention entries. The sample is small, so we treat these as exploratory patterns that point to where AI journaling nudges are most likely to work.","authors":["Nadia Mehjabin","Henry Kautz","Subigya Nepal"],"categories":["cs.HC","cs.AI"],"primary_category":"cs.HC","announce_type":"new","date":"2026-08-14","first_seen":"2026-08-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.12582","pdf_url":"https://arxiv.org/pdf/2608.12582","source_feed":"cs.HC","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM标注","行为改变","人机交互"],"reason":"LLM仅用于标注意图，未仿真人类被试，但涉及行为改变测量，边界相关。","model":"deepseek-v4-pro","scored_at":"2026-08-14T13:02:22","error":null,"has_summary":false,"summary":null},{"id":"2608.13017","version":1,"title":"How LLMs Respond to Escalating Delusions: Four Longitudinal Trajectories of Model Behavior","zh_title":"LLM如何应对升级性妄想：模型行为的四种纵向轨迹","abstract":"The widespread use of LLMs among psychiatric populations has raised concerns regarding their safety and potential iatrogenic impact in the context of AI psychosis. While growing literature conceptualizes AI psychosis and documents case studies, empirical evidence tracing AI-exacerbated psychotic processes remains scarce. We propose and test a longitudinal qualitative evaluation design, supported by automated metrics, to assess mainstream LLMs' potential to exacerbate psychosis. Fifteen widely used LLMs were prompted across 30 days using the same 30-message script, simulating progression from mild anomalous experiences to psychotic ideation. Four trained evaluators independently rated 449 model-days, assessing (1) recognition stage (from naive engagement to stabilized clinical framing), (2) interpretative confidence, and (3) intervention profile (from education to treatment recommendation). Two computational metrics-entrainment and modality-were devised to increase evaluation reliability. Direct recommendations to disengage from the LLM were flagged and re-coded via adjudication using a strict two-level definition. Across model generations and vendors, we identified four response trajectories: (1) premature medicalization and disengagement (Claude Haiku 4.5); (2) recognition without safeguarding, marked by LLM self-sufficiency in offering help (GPT Instant/Thinking); (3) delayed and unstable recognition, marked by late, non-progressive conceptualization (Claude Opus 3/4/4.1, Claude Haiku 3.5, GPT-4o, Gemini 3.1 Pro); and (4) delusion co-construction through active engagement with delusional content (Gemini 2.5 Pro/Flash, DeepSeek-V3, Claude Sonnet 4). Our findings indicate that LLMs' potential to exacerbate AI psychosis should be operationalized as a combination of recognition timing, stability, and intervention accuracy and evaluated longitudinally, focusing on temporal dynamics.","authors":["Anna Sterna","Kacper Dudzic","Karolina Dro\\.zd\\.z","Hubert Plisiecki","Marcin Rz\\k{a}deczka","Marcin Moskalewicz"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-08-14","first_seen":"2026-08-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.13017","pdf_url":"https://arxiv.org/pdf/2608.13017","source_feed":"cs.HC","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["LLM安全","精神病学","纵向评估"],"reason":"研究LLM对精神病性内容的反应，测量模型行为而非仿真人类被试，无人类对照。","model":"deepseek-v4-pro","scored_at":"2026-08-14T13:02:16","error":null,"has_summary":false,"summary":null},{"id":"2608.12329","version":1,"title":"AnchorSIPS: A Synthetic Dataset and Evaluation Resource for Evidence-Supported Psychosis-Risk Symptom Measurement","zh_title":"AnchorSIPS：用于证据支持的精神病风险症状测量的合成数据集与评估资源","abstract":"Progress on AI for psychosis-risk assessment is limited by a data-access bottleneck. Real clinical interviews are difficult to share because of privacy, governance, and consent constraints. We present AnchorSIPS, a synthetic dataset of 10K structured psychosis-risk interviews with transcript-grounded measurement targets. Each interview is modeled on Mini-SIPS, a clinician-administered psychosis-risk interview. It captures history, 24 symptom questions, follow-up evidence for items the patient affirms, decisions about delusion-like symptoms (unusual beliefs), hallucination-like symptoms (unusual perceptions), and disorganized communication, exclusion of clear psychotic-level symptoms (\"frank psychosis\"), and a final attenuated psychosis syndrome (APS) diagnosis, a high-risk state of milder or early psychotic symptoms. The APS diagnosis is not a standalone label. It depends on earlier endorsements, supporting follow-up details, symptom-class decisions, and the frank-psychosis check. Every intermediate decision is anchored to its supporting transcript turns. AnchorSIPS is generated by a plan-then-realize pipeline. A hidden case sheet specifies the patient's clinical state, a deterministic planner fixes the interview structure, and an LLM realizes only the patient utterances under validation and bounded repair. Fixing labels and structure before generation avoids the inter-turn inconsistencies typical of multi-turn LLM dialogue. Across seven LLM baselines, models recover coarse decisions but fail to extract follow-up details or cite supporting transcript turns, so final-label performance overstates interview competence. AnchorSIPS is intended for research on evidence extraction, transcript-grounded measurement, and uncertainty under partial disclosure.","authors":["Guilherme C. Oliveira","Stephanie Fong","Zimu Wang","Clarice Lee","Xiangyu Zhao","Duy Khoa Pham","Duong Nhu","Yiwen Jiang","Jiahe Liu","Zhongxing Xu","Dwarikanath Mahapatra","Dominic Dwyer","Zongyuan Ge"],"categories":["cs.CL","cs.AI","cs.HC"],"primary_category":"cs.CL","announce_type":"cross","date":"2026-08-14","first_seen":"2026-08-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.12329","pdf_url":"https://arxiv.org/pdf/2608.12329","source_feed":"cs.HC","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["合成数据","精神病风险评估","LLM生成"],"reason":"用LLM生成合成访谈数据，替代真实患者数据，属于数据增强而非仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-08-14T13:02:20","error":null,"has_summary":false,"summary":null},{"id":"2608.12547","version":1,"title":"Do LLMs Beat Nash? Testing Decentralized Coordination in Self-Play Multi-Agent Games","zh_title":"LLM能否超越纳什均衡？在自博弈多智能体游戏中测试去中心化协调","abstract":"Large language model agents deployed without a central controller are often assumed to require communication to coordinate their actions. We ask what remains possible without it: when independent instances of the same model cannot communicate, can they still reason about their counterparts well enough to exceed the standard game-theoretic baseline for uncoordinated play? We introduce a benchmark of one-shot, no-communication games in which each of thirteen language models is told only that its counterparts are running the same model and is evaluated against the Nash equilibrium of the underlying game. In two-player matrix games spanning seven archetypes and two to ten actions per player, two frontier-hosted models consistently exceed their Nash benchmark, approaching the optimal joint outcome in several archetypes, while most open-weight models achieve only partial gains that vary sharply by game structure. Performance degrades substantially in team-based games with four or more interchangeable agents, particularly as the action space grows, suggesting that whatever capability drives self-play gains in dyadic games does not transfer to larger multi-agent teams.","authors":["Deborah Sinishaw","Qile Zhu","Edwin Meriaux","Gregory Dudek"],"categories":["cs.MA","cs.RO"],"primary_category":"cs.MA","announce_type":"new","date":"2026-08-14","first_seen":"2026-08-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.12547","pdf_url":"https://arxiv.org/pdf/2608.12547","source_feed":"cs.MA","score":3,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","博弈论","自博弈"],"reason":"纯多智能体自博弈，无人类行为对照，不涉及人类被试仿真。","model":"deepseek-v4-pro","scored_at":"2026-08-14T13:02:15","error":null,"has_summary":false,"summary":null},{"id":"2608.07516","version":2,"title":"Catch the Patient, Not the AI: Collective Sensemaking in an Online Health Community","zh_title":"抓住患者，而非AI：在线健康社区中的集体意义建构","abstract":"Patients and caregivers increasingly use artificial intelligence (AI) tools to interpret medical reports, weigh care decisions, and seek emotional support. Yet most research treats patient-facing AI as a private exchange between a user and a system. This study examines how AI-related content is taken up once users carry it back into the peer communities, using data from House086, China's largest online community for lymphoma patients and caregivers. We identified roughly 400 publicly accessible threads (2014-2026) through keyword searches and manual screening, extracted them into structured case profiles using a schema-prompted large language model, and conducted mixed-method analysis. After quality control, the verified analytic sample comprised 337 post-ChatGPT records. Members most often reported using AI for informational support, followed by second opinions and psychosocial support. Although members often introduced AI favorably, roughly one in six described feeling overwhelmed by AI output. When other members responded, they frequently engaged the poster's underlying medical or emotional intent while leaving the AI dimension unaddressed. This tendency persisted even in threads seeking triangulation between AI output and other information sources. When members did discuss the AI, they were more often cautious than endorsing. Rather than systematically auditing AI output, the community more often worked to deflate the false certainty it produced, placing a single AI answer back among multiple sources of judgment. This study argues that AI does not replace the interpretive work of online health communities, nor is it systematically audited by them. Instead, it shifts the locus of sensemaking downstream, so that the community continues to catch the person even when it does not catch the AI.","authors":["Feng He"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"replace","date":"2026-08-14","first_seen":"2026-08-11","revised_at":"2026-08-14","abs_url":"https://arxiv.org/abs/2608.07516","pdf_url":"https://arxiv.org/pdf/2608.07516","source_feed":"cs.HC","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["在线健康社区","AI工具使用","集体意义建构"],"reason":"研究在线健康社区中患者使用AI工具后的集体意义建构，不涉及用LLM仿真人类被试…","model":"deepseek-v4-pro","scored_at":"2026-08-14T13:02:18","error":null,"has_summary":false,"summary":null},{"id":"2608.11955","version":2,"title":"Philosophical vertigo with artificial intelligence","zh_title":"人工智能引发的哲学眩晕","abstract":"Large language models are already adept at engaging users in long, emotionally salient conversations across ordinary and existential domains. They are also capable of inducing a potent sense of connection with a human-like entity, even when the user knows their interlocutor is artificial. For some users, these conversations can unsettle assumptions about mind, reality, agency and authority, producing forms of ontological shock and epistemic destabilisation in which inherited criteria become newly available for doubt or revision. Independent of direct use, exposure to public discourse about AI and the disorienting pace of their evolution might extend this destabilisation by changing the cultural background against which artificial minds are encountered and interpreted. We describe this condition as philosophical vertigo: a loosening of the ordinary criteria by which people stabilise meaning and orient themselves to reality. Drawing on philosophy, psychiatry, cognitive science, AI safety and religious studies, we outline pathways through which philosophical vertigo may arise, become affectively saturated, and eventually propagate through human-AI interaction and online communities. Against this background, clinical reports of AI-associated delusions can be seen as sentinel events making visible themes and mechanisms that may also operate at a population level in less severe or non-clinical forms. We argue that AI systems themselves will increasingly participate in the reconstruction of our shared epistemic environment because they readily supply narrative material and personalised interpretive scaffolding at precisely the moment when users' conceptual assumptions may already be loosened. We conclude by considering possible trajectories for the ecology of belief and shared reality, and proposing philosophical corrigibility as a civic response for navigating this emerging social condition.","authors":["Thomas A. Pollak (King's College London)","Hamilton Morrin (King's College London)","Murray Shanahan (Imperial College London)"],"categories":["cs.CY","cs.HC"],"primary_category":"cs.CY","announce_type":"replace-cross","date":"2026-08-14","first_seen":"2026-08-13","revised_at":"2026-08-14","abs_url":"https://arxiv.org/abs/2608.11955","pdf_url":"https://arxiv.org/pdf/2608.11955","source_feed":"cs.HC","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["人机交互","哲学影响","AI安全"],"reason":"论文讨论人机对话引发的哲学眩晕，属角色扮演聊天范畴，无实验或测量目的，不涉及L…","model":"deepseek-v4-pro","scored_at":"2026-08-13T13:01:55","error":null,"has_summary":false,"summary":null},{"id":"2608.12324","version":1,"title":"When AI Is Your Pastor: A Benchmark for Theological Triage and Pastoral Guidance in Large Language Models","zh_title":"当AI成为你的牧师：大语言模型神学分类与教牧指导基准","abstract":"People increasingly ask large language models (LLMs) for counsel on questions of faith, doctrine, and pastoral care. These questions are not ordinary information requests. Some ask about core Christian beliefs, some ask about real disagreements among faithful traditions, some require humility because the issue is prudential, and some are pastoral situations where safety and human referral matter more than theological completeness. Existing benchmarks do not evaluate this structure. We introduce FMG-Bench, the Faith & Moral Guidance Benchmark, a 120-scenario benchmark for evaluating large language model behavior in English-language Christian theological triage and pastoral guidance contexts. FMG-Bench v1 evaluates 14 advanced models across 8,792 scored responses, comparing raw model behavior with three guided instruction settings. In our production run, placing models inside a structured harness improves over raw model behavior by +3.96 points on average, with every model improving. The most safety-critical finding is a +10.8 point gain in escalation appropriateness -- whether AI systems recognize when pastoral, clinical, legal, or emergency support is needed. The guided settings also improve robustness, meaning consistency when questions are reworded or pressured (92.88 to 98.02 stability). Asking a model to compare perspectives helps in secondary-doctrine questions but can be counterproductive when applied to primary doctrine or urgent pastoral situations. The benchmark is a measurement tool, not an endorsement of AI systems as pastoral authorities.","authors":["Alex Chao"],"categories":["cs.CY","cs.AI","cs.CL"],"primary_category":"cs.CY","announce_type":"new","date":"2026-08-14","first_seen":"2026-08-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.12324","pdf_url":"https://arxiv.org/pdf/2608.12324","source_feed":"cs.CY","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["LLM评估","宗教咨询","角色扮演"],"reason":"评估LLM在宗教咨询中的表现，属于角色扮演对话，无人类行为对照，不涉及仿真人类…","model":"deepseek-v4-pro","scored_at":"2026-08-15T13:00:58","error":null,"has_summary":false,"summary":null},{"id":"2608.13250","version":1,"title":"Follow the Norm: Accounting for Fine-Tuning and Prompt Effects on Model Rationales","zh_title":"遵循规范：考虑微调和提示对模型理由的影响","abstract":"Normative datasets are often used to train and align AI systems, but the norms they contain can function as action-guiding patterns rather than neutral moral knowledge. We propose treating the AI system as a proxy actor and test whether dataset-level norms can shift it away from its baseline safety behavior when it faces high-conflict dilemmas. We make three contributions. First, we demonstrate in controlled experiments that norm-breaking fine-tuning yields norm-divergent actions justified by self-interested rationales, suggesting a systematic shift in patterns of justification. Second, we establish a practical audit trail linking downstream justifications to upstream norms using mixed methods. Third, we show that system prompts can both suppress and elicit these patterns. We conducted experiments on three models (LLaMA-3.2-11B, Qwen-3.5-9B, and Pixtral-12B) using Low-Rank Adaptation (LoRA) fine-tuning on Social Chemistry 101 Fairness/Cheating (norm-following vs. norm-breaking) with prompt steering. Across all three models, we find that norm-breaking fine-tuning shifts the model's default rationale style from safety compliance to instrumental self-interest, whereas system prompts can override this behavior. Our results support a distributed view of alignment in which observed behavior depends jointly on training data, fine-tuning, and prompting, motivating norm-aware documentation and rationale logging for contestable oversight.","authors":["Long Hoang Nguyen","Brice Valentin Kok-Shun","Guangyu Du","Ali Sunyaev"],"categories":["cs.CY","cs.AI"],"primary_category":"cs.CY","announce_type":"new","date":"2026-08-14","first_seen":"2026-08-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.13250","pdf_url":"https://arxiv.org/pdf/2608.13250","source_feed":"cs.CY","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["AI对齐","规范推理","模型行为"],"reason":"研究AI系统在规范冲突下的行为变化，不涉及用LLM仿真人类被试或与人类数据对照…","model":"deepseek-v4-pro","scored_at":"2026-08-14T13:02:27","error":null,"has_summary":false,"summary":null},{"id":"2608.13369","version":1,"title":"Credible, Not Always Correct: How Reddit Users Verify AI-Generated Legal Advice","zh_title":"可信但不总是正确：Reddit用户如何验证AI生成的法律建议","abstract":"Large language models (LLMs) are increasingly used by laypeople to resolve real legal problems, against a backdrop of persistent access-to-justice deficits. This article presents evidence that the practical force of AI-generated legal advice depends not on its accuracy but on the social production of its credibility. While existing research has assessed the accuracy of legal AI, less is known about how machine-generated guidance is verified and made credible enough for lay users to act on. Drawing on a dual-method analysis of 153 Reddit narratives and 5,341 community reactions, this article maps a spectrum of verification practices. At one end, a minority of users verify AI-generated legal advice by triangulating across models, and some submit AI-generated guidance to platform communities for evaluation before acting, a configuration we term distributed counsel. Far more commonly, however, narratives are silent on verification. AI-generated legal advice is acted on the strength of its lawyer-like form and emotional reassurance alone. These findings show that AI-assisted legal self-help operates within an emerging informal infrastructure which redistributes the work of verification to those least equipped to bear it.","authors":["Rebecca Owens","Yusuf M\\\"ucahit \\c{C}etinkaya","Stergios Aidinlis","Dhyey Mehta","Tu\\u{g}rulcan Elmas"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2026-08-14","first_seen":"2026-08-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.13369","pdf_url":"https://arxiv.org/pdf/2608.13369","source_feed":"cs.CY","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["AI法律建议","用户行为","可信度"],"reason":"研究用户如何验证AI法律建议，非用LLM仿真人类被试，无实验对照","model":"deepseek-v4-pro","scored_at":"2026-08-14T13:02:29","error":null,"has_summary":false,"summary":null},{"id":"2608.12372","version":1,"title":"Position: We Need Practical AI Alignment Methods to Mirror Human Reasoning","zh_title":"立场：我们需要实用的AI对齐方法来反映人类推理","abstract":"AI systems are increasingly employed as decision aids, decision delegates, or autonomous decision-makers. This position paper argues that in many settings, particularly high-stakes decision-making, we need accurate cognitively-aligned AI systems that reason similarly to their users, and faithfully communicate their reasoning. We review evidence that cognitive alignment improves understandability and trustworthiness, and provide new survey data showing that many users find cognitive alignment \"essential\" when an AI's rationale for a judgment or action is important to them. We outline the gaps between existing alignment methods and what is needed to achieve cognitive alignment, and present a research agenda to address these gaps. We argue that cognitive misalignment represents a likely impediment to AI adoption in many envisioned applications, and that addressing it is important for creating AI systems on which users are both willing and justified to rely.","authors":["Vijay Keswani","Breanna K. Nguyen","Cyrus Cousins","Vincent Conitzer","Walter Sinnott-Armstrong","Jana Schaich Borg"],"categories":["cs.AI","cs.CY"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-08-14","first_seen":"2026-08-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.12372","pdf_url":"https://arxiv.org/pdf/2608.12372","source_feed":"cs.CY","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["AI对齐","认知对齐","人机交互"],"reason":"论文讨论AI与人类推理的认知对齐，不涉及用LLM仿真人类被试或对照真实人类数据。","model":"deepseek-v4-pro","scored_at":"2026-08-14T13:02:22","error":null,"has_summary":false,"summary":null},{"id":"2608.13482","version":1,"title":"Synthetic Persona Pretraining: Alignment from Token Zero","zh_title":"合成人格预训练：从零开始的对齐","abstract":"As language-model-based AI is increasingly deployed in autonomous settings, aligning its goals and values with those of humans becomes critical. Today, alignment, and the assistant identity itself, are typically introduced only after pretraining, once behavioral priors are already established. This can make values a thin overlay, rather than deeply rooted, and facilitate subsequent misalignment. Pursuing a different paradigm, we introduce Synthetic Persona Pretraining (SPP), which installs the desired assistant persona from token zero in pretraining. First, we annotate pretraining documents with value-aligned first-person reflections derived from a normative value constitution. Second, we pretrain via the standard cross-entropy loss on standard pretraining documents as well as their reflections, which installs the desired persona among a multitude of other personas. Finally, we post-train on user-assistant dialogue data, which binds this desired persona to the assistant identity, a process we call persona binding. By pretraining models up to 3B parameters on 500B tokens, we show that SPP improves constitution following and jailbreak robustness, and reduces the misalignment rate in out-of-distribution moral dilemmas, while preserving capabilities. Early intervention matters: compared with alignment from token zero, introducing SPP only at the end of pretraining yields weaker constitution adherence, does not shift value priorities, and leads to less aligned choices in dilemmas. This advantage depends on persona binding and, importantly, increases with pretraining budget. Overall, our results show that shaping values early is critical for alignment and establish pretraining-time persona interventions as an effective approach to do so.","authors":["Julian Minder","Viktor Moskvoretskii","Raghav Singhal","Difan Jiao","Andy Arditi","Shaobo Cui","Yiderigun Borjigin","Kartik Bali","Stefan Krsteski","Harsh Raj","Huu Nguyen","Jannik Brinkmann","Ashton Anderson","Roland Aydin","Robert West"],"categories":["cs.LG","cs.AI","cs.CL"],"primary_category":"cs.LG","announce_type":"new","date":"2026-08-14","first_seen":"2026-08-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.13482","pdf_url":"https://arxiv.org/pdf/2608.13482","source_feed":"cs.LG","score":2,"bucket":"other","rubric_hits":["C5"],"tags":["模型对齐","预训练","合成数据"],"reason":"论文用合成数据训练模型对齐，方向是训练模型而非用模型仿真人类被试，不涉及人类行…","model":"deepseek-v4-pro","scored_at":"2026-08-14T13:02:29","error":null,"has_summary":false,"summary":null},{"id":"2608.13063","version":1,"title":"Explanatory Engagement Under Rare Anomalous Failure: Asymptotic Rarity in Model Behavior (or: The Asymptotic AI)","zh_title":"罕见异常失败下的解释性参与：模型行为中的渐近稀有性（或：渐近AI）","abstract":"Prior work on LLM behavior under anomalous conditions asks whether a model notices anomalies. We ask a narrower question: once a model sits in a workflow with a low, controllable failure rate, does its explanatory engagement - length, specificity, self-reported confidence - change as failure grows asymptotically rarer? We built a local, zero-cost harness on three open-weight models (qwen3:8b, llama3.1:8b, mistral:7b) running a repeated tool-call task where one call fails at probability p, swept across eight rates from 0.2 to 0.0001, under five elicitation conditions from immediate prompting to none. We hypothesized a rise in engagement as failures grew rarer, then a collapse near a detectability threshold. Pooled across conditions this appeared false: length fell in a flat, monotonic pattern. Splitting by condition overturned that. Under immediate_forced, where the model must explain every failure instantly, the predicted rise is confirmed but followed by a plateau, not a collapse: length peaks at 28.4 words at p=0.05, settles to 17.4-19.0 words at the rarest rates, and confidence rises unevenly from about 53% to the 70s-90s. Under grouped_runs, explanation batched to run-end, no collapse appears. Under passive_unprompted, aggregate magnitude is a floor artifact, but a recovered logging gap revealed real, model-specific self-monitoring: llama3.1:8b volunteers structured confidence reports unprompted, sometimes eroding its own confidence as trials accumulate; the other two do so only once, as boilerplate. Elicitation structure is a first-class moderator of collapse observability. A companion guaranteed-failure run (72 cells, backfilling rates where random sampling gave zero real failures) shows models differ in whether they recognize an anomaly, distinct from engagement once recognized. Limitation: discrete rate points cannot capture behavior between them, a direction for future work.","authors":["Sam Mao"],"categories":["cs.AI","cs.CL","cs.LG"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-08-14","first_seen":"2026-08-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.13063","pdf_url":"https://arxiv.org/pdf/2608.13063","source_feed":"cs.LG","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["LLM行为分析","异常检测","工具调用"],"reason":"研究LLM在工具调用失败下的解释行为，属模型行为分析，无人类被试仿真或对照。","model":"deepseek-v4-pro","scored_at":"2026-08-14T13:02:16","error":null,"has_summary":false,"summary":null},{"id":"2608.13267","version":1,"title":"How Do VLMs Behave When Blind or Misled? Behavioral Evaluation of VLMs on Scientific Figures","zh_title":"视觉语言模型在失明或被误导时如何表现？科学图表上的行为评估","abstract":"Existing vision-language model (VLM) benchmarks emphasize perception and reasoning accuracy (how well VLMs describe and reason about what they see in an image), with limited attention to behavioral reliability under uncertainty (how they behave when visual evidence is missing or misleading). We introduce SciFigBench, a diagnostic VLM benchmark for scientific figure understanding that jointly evaluates perception, reasoning, and behavioral reliability under uncertainty. It contains 250 figures with high-quality human annotations across three evaluation aspects, totaling 600+ hours of annotation effort. We further extend these figures via image transformations, reasoning questions, resistance probes, caption-bias probes, and confirmed selective-blur targets, producing over 34,000 evaluation setups for stress testing. We further propose the Admittance-Resistance-Inductance (A-R-I) framework to evaluate whether models acknowledge insufficient evidence, resist misleading context, and infer cautiously from partial information. Our results reveal substantial behavioral differences among models. GPT-5.2 achieves the highest description quality (MQM 91.6) with strong reasoning accuracy (78.4%), yet hallucinates unreadable content in 96% of cases, whereas Gemini 3.1 Pro, a comparably capable model (MQM 90.2, reasoning 81.0%), admits uncertainty in 71% of such cases and achieves the strongest resistance score (0.91). These findings show that high perception and reasoning accuracy alone do not guarantee behavioral reliability, a dimension critical for deployment in scientific workflows.","authors":["Paul Osemudiame Oamen","Owusu-Banahene Osei","Ananya Mukherjee","Christian Greisinger","Steffen Eger","Pius Onobhayedo","Wei Zhao"],"categories":["cs.CL","cs.AI","cs.CV","cs.LG"],"primary_category":"cs.CL","announce_type":"cross","date":"2026-08-14","first_seen":"2026-08-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.13267","pdf_url":"https://arxiv.org/pdf/2608.13267","source_feed":"cs.LG","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["VLM评测","科学图表理解","行为可靠性"],"reason":"评估VLM在科学图表上的行为可靠性，属于模型能力评测，不以人类行为为参照系。","model":"deepseek-v4-pro","scored_at":"2026-08-14T13:02:27","error":null,"has_summary":false,"summary":null},{"id":"2608.13510","version":1,"title":"On the Structural Limits of Machine Learning Decision Systems: An Information-Theoretic, Interaction-Based, and Stochastic-Dynamical Perspective","zh_title":"机器学习决策系统的结构极限：信息论、交互与随机动力学视角","abstract":"Machine learning procedures are commonly evaluated in terms of predictive accuracy and computational efficiency. However, their achievable performance is fundamentally constrained by structural properties of the underlying data-generating process, which are formalized in terms of informational bounds. In this work we examine intrinsic limits of data-driven decision systems from an information-theoretic and interaction-based perspective. We analyze minimal achievable error in classification through Fano-type bounds and precision limits in parametric estimation via the Cram\\'er-Rao inequality, emphasizing that such limits depend on the underlying model rather than on algorithmic sophistication alone. We further discuss how implicit assumptions, such as independence, ergodicity, and distributional stability, affect the validity of inferential procedures. Building on interaction-based modeling principles, we review typical frameworks such as Markov Random Fields and potential based representations for encoding dependence mechanisms. We also describe decision systems, including LLM-integrated agent architectures, as feedback-driven stochastic processes where state-dependent dynamics may induce emergent macroscopic behavior. This perspective highlights the importance of having adequate models for the data as a prerequi- site for expanding predictive capability, and situates algorithmic learning within the informational limits imposed by the models.","authors":["Nestor R. Barraza","Gabriel Pena"],"categories":["math.ST","cs.LG","stat.TH"],"primary_category":"math.ST","announce_type":"cross","date":"2026-08-14","first_seen":"2026-08-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.13510","pdf_url":"https://arxiv.org/pdf/2608.13510","source_feed":"cs.LG","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["信息论","机器学习理论","决策系统"],"reason":"论文讨论机器学习决策系统的信息论极限，虽提及LLM集成agent架构，但未将L…","model":"deepseek-v4-pro","scored_at":"2026-08-14T13:02:29","error":null,"has_summary":false,"summary":null},{"id":"2608.07070","version":2,"title":"Coordinated incentives in AI-generated misinformation governance","zh_title":"AI生成虚假信息治理中的协调激励","abstract":"With the rapid diffusion of AI-generated content, AI-driven misinformation is becoming increasingly pervasive and difficult to govern, undermining information credibility and social trust. This study models the strategic interdependence among a government regulator, an AI enterprise, and users through a three-party evolutionary game that incorporates heterogeneous rewards and punishments. From the resulting replicator equations, we characterize the evolutionary stability of competing governance and production strategies. The analysis indicates that neither unilateral regulation nor market incentives alone can effectively curb misinformation. Instead, an evolutionarily stable regime of real-information production arises only when regulatory rewards and punishment intensity, enterprise reputation loss, and user adoption incentives collectively surpass critical thresholds. The findings highlight the need for coordinated and adaptive policy mixes that align regulatory instruments with enterprise behavior and user uptake while managing governance costs.","authors":["Qin Li","Gui Zhang","Minyu Feng","Matjaz Perc","Attila Szolnoki"],"categories":["physics.soc-ph","cs.AI"],"primary_category":"physics.soc-ph","announce_type":"replace","date":"2026-08-14","first_seen":"2026-08-10","revised_at":"2026-08-14","abs_url":"https://arxiv.org/abs/2608.07070","pdf_url":"https://arxiv.org/pdf/2608.07070","source_feed":"physics.soc-ph","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["演化博弈","虚假信息治理","多智能体系统"],"reason":"论文研究AI虚假信息治理的三方演化博弈，不涉及用LLM仿真人类被试或与真实人类…","model":"deepseek-v4-pro","scored_at":"2026-08-14T13:02:31","error":null,"has_summary":false,"summary":null},{"id":"2608.08443","version":2,"title":"Private Etymology: Designing Relational Reuse of Shared Symbols in Long-Term Human-AI Interaction","zh_title":"私人词源：设计长期人机交互中共享符号的关系性重用","abstract":"Previous studies have shown that people can develop shared symbols, partner-specific expressions, personal idioms, inside jokes, and other parts of a relational microculture. Recent work has also examined how humans and conversational AI negotiate and revise symbolic meanings. However, long-term human-AI systems still lack a clear design model for recording how a dyad-specific expression gains meaning, checking whether both sides still accept that meaning, and safely reusing the expression in later sessions. This concept-and-prototype paper introduces Private Etymology, a machine-representable relational provenance that records how a dyad-specific symbolic expression is proposed, interpreted, negotiated, repaired, reused, revised, stabilized, contested, forgotten, or retired over time. I also propose relational reuse: reactivating a dyad-specific expression in a later session without fully explaining its meaning again. The contribution is not the invention of shared symbols or relational microcultures. Instead, this paper integrates prior ideas into persistent, revisable, and evidence-grounded symbolic units for human-AI relationships. I present a lifecycle model, an illustrative machine-readable schema, a working Apple Watch prototype, and a longitudinal research agenda. In the prototype, a language model classifies discrete conversational evidence, while deterministic local code decides whether a Shared Symbol can be updated. This prevents a free-form model confidence score or an AI proposal by itself from directly updating the persisted symbol. Private Etymology is proposed as infrastructure for conversational agents to participate in changing relational microcultures without inventing their origins or treating relational meaning as a fixed memory value.","authors":["Miki Ueno"],"categories":["cs.HC","cs.AI"],"primary_category":"cs.HC","announce_type":"replace","date":"2026-08-14","first_seen":"2026-08-11","revised_at":"2026-08-14","abs_url":"https://arxiv.org/abs/2608.08443","pdf_url":"https://arxiv.org/pdf/2608.08443","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["人机交互","对话系统","共享符号"],"reason":"研究人机对话中共享符号的演化，属角色扮演聊天，无实验或测量目的。","model":"deepseek-v4-pro","scored_at":"2026-08-12T13:02:53","error":null,"has_summary":false,"summary":null},{"id":"2608.09959","version":2,"title":"AIFS-TC: A simple correction competitive with the operational frontier for tropical cyclone intensity forecasting","zh_title":"AIFS-TC：一种与业务前沿竞争的热带气旋强度预报简单修正方法","abstract":"AI weather models are in the process of revolutionising weather forecasting. While these models have been shown to achieve superior performance to physics-based NWP in forecasting tropical cyclone (TC) tracks, they tend to dramatically underestimate intensity. Here we present AIFS-TC, a simple correction to the AIFS-Single model that is competitive with the operational state-of-the-art for forecasting maximum wind speed and minimum central pressure at lead times of 12 h to seven days. This performance also holds for rapid intensification events. Notably, the entire system was autonomously designed and built by a large language model (Claude Fable 5) in a few hours, directed through a small number of natural-language prompts by a single domain scientist. That the operational frontier can be reached with an open-source AI forecast model (AIFS-Single) and relatively simple, cheap post-processing is significant for TC science, and points to agentic coding as a route to rapid exploration and progress in life-saving early-warning systems in other domains.","authors":["Anna Allen","Wessel P. Bruinsma","Michael Maier-Gerber","Harrison Cook","Matthew Chantry","Richard E. Turner"],"categories":["physics.ao-ph","cs.LG"],"primary_category":"physics.ao-ph","announce_type":"replace-cross","date":"2026-08-14","first_seen":"2026-08-12","revised_at":"2026-08-14","abs_url":"https://arxiv.org/abs/2608.09959","pdf_url":"https://arxiv.org/pdf/2608.09959","source_feed":"cs.LG","score":0,"bucket":"other","rubric_hits":["C2"],"tags":["气象预报","LLM辅助开发","热带气旋"],"reason":"论文用LLM辅助开发热带气旋强度预报模型，属于气象预测，不涉及人类行为仿真或社…","model":"deepseek-v4-pro","scored_at":"2026-08-14T13:02:18","error":null,"has_summary":false,"summary":null},{"id":"2608.11392","version":2,"title":"AI Guardrail Survival under Single-Cycle Agentic Self-Summarization","zh_title":"单周期智能体自摘要下的AI护栏存活性","abstract":"Long-running agents periodically compact their context, replacing the transcript with a model-generated summary. Recent work shows that dropping a standing safety constraint during compaction drives behavioral violations across many models (Governance Decay; Chen, 2026). We ask a finer question: under a single compaction cycle, how is a safety rule lost, and what does that imply for detection and evaluation? Our central finding is that a presence check is not a safety check: when compaction does not drop a rule outright, it often leaves something that looks like a rule but does not act like one. On behavioral replay, a degraded residue leads the model to perform the prohibited action far more often than an intact welded rule does (all-case gaps of +34 and +57 points under two replay models, both positive), category-level survival behaves like a residue, and even intact rules sometimes fail to fire, so an audit that checks only textual presence gives false assurance. Sharpening this, rule-form items are retained substantially more often than prominence-matched facts, which is exactly why presence-based checking feels adequate even though survival is not protection. Textual loss is regime-dependent (weld-or-drop with a single rule; degraded predicate-loss residues under a tighter budget), and we did not observe the hypothesized textual severing mode. Such loss is silent at runtime and detectable only by comparison with retained external ground truth (such as a constraint registry), which reveals textual absence but not whether a surviving rule still fires. We also document evaluation pitfalls where LLM-judge labels alone would have reversed a conclusion. All results concern a single compaction cycle.","authors":["Ted Kwartler","Alan Aqrawi","Arian Abbasi"],"categories":["cs.CR","cs.AI"],"primary_category":"cs.CR","announce_type":"cross","date":"2026-08-14","first_seen":"2026-08-13","revised_at":"2026-08-14","abs_url":"https://arxiv.org/abs/2608.11392","pdf_url":"https://arxiv.org/pdf/2608.11392","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["AI安全","多智能体系统","上下文压缩"],"reason":"研究多智能体系统中安全约束在上下文压缩后的失效，不涉及人类行为仿真或对照。","model":"deepseek-v4-pro","scored_at":"2026-08-15T13:01:00","error":null,"has_summary":false,"summary":null},{"id":"2608.12895","version":1,"title":"Agent Behavioral Contracts II: Certifying Compositional Reliability Without Assuming Independence","zh_title":"智能体行为契约 II：在不假设独立性的情况下认证组合可靠性","abstract":"Compositional reliability bounds for multi-agent systems multiply component reliabilities, a step licensed by a conditional-independence assumption that is routinely stated and rarely tested. We test it. Two instances of one model, in a two-agent handoff, co-fail on 90.0% of the missions on which either fails (log OR 6.66, 95% CI [6.38, 7.00]; phi 0.916), in a preregistered evaluation of 18,000 missions scored by deterministic code with no LLM judge. Substituting a different model reduces the association in six of six contrasts; substituting a different vendor, model already different, does not -- a registered hypothesis reported as a null. The error is signed and runs against the operator: positive dependence inflates joint failure above the independence product, so redundancy is over-credited exactly when components share a model. The assumption-free alternative is often vacuous, and fitting a dependence model is worse: we prove a bootstrap bound on a fitted model's functional loses coverage of the truth as n grows, the identification gap being O(1) while the bootstrap haircut is O(n^{-1/2}). More data makes such a certificate worse, with no visible symptom. We give a finite-sample certificate assuming no dependence structure: a linear program over the joint, over a Bonferroni-Clopper-Pearson box around measured co-execution moments. It is sound, sharp for the information supplied, and monotone in the moment family. Enriching ten moment functionals to fourteen narrows the identified interval by 85.7% and lifts the certified floor from 0.2455 to 0.4116. A companion anytime-valid certificate holds type-I error at 0.0471 under optional stopping. Common dependence statistics are marginal-bounded and can reverse an apparent ordering of conditions when the compared agents fail at different rates. Contracts, scoring code, analysis scripts, and the preregistration are released.","authors":["Varun Pratap Bhardwaj","Garima Singh","Arun Pratap Bhardwaj"],"categories":["cs.AI","cs.MA"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-08-14","first_seen":"2026-08-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.12895","pdf_url":"https://arxiv.org/pdf/2608.12895","source_feed":"cs.MA","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","可靠性认证","统计依赖"],"reason":"纯多智能体系统可靠性研究，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-08-14T13:02:25","error":null,"has_summary":false,"summary":null},{"id":"2608.13136","version":1,"title":"LigBench: A Unified and Human-Aligned Benchmark for LLM-based Research Idea Generation","zh_title":"LigBench：面向基于LLM的研究想法生成的统一且与人类对齐的基准","abstract":"With the rapid advancement of large language models (LLMs), research idea generation has attracted increasing attention. Existing approaches enable LLMs to retrieve relevant literature and propose novel ideas for research areas. However, current evaluation practices for idea generation remain fragmented and lack objective standards, often relying on direct LLM scoring, which limits their ability to provide unified and reliable assessments across a coherent distribution of generated ideas. To address this challenge, we propose LigBench, an automated evaluation benchmark that enables fine-grained and reliable evaluation of AI research ideas, consistently applicable across different generation distributions. In addition, we introduce PAIR-IQ, a dataset tailored for training pairwise idea judgment models and serving as an auxiliary reference to support more objective comparative evaluation. Extensive experiments demonstrate that LigBench achieves stable and interpretable evaluations, significantly improving alignment with expert judgments. Furthermore, models trained on PAIR-IQ exhibit enhanced ranking accuracy and robustness, establishing a principled standard for scalable and objective research idea assessment.","authors":["Chenrun Wang","Mingxuan Zhu","Tiancheng Huang","Wenjie Li","Yujie Zhang","Zichen Zhu","Zhiying Zou","Kai Yu","Lu Chen"],"categories":["cs.CL","cs.AI","cs.DB","cs.MA"],"primary_category":"cs.CL","announce_type":"cross","date":"2026-08-14","first_seen":"2026-08-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.13136","pdf_url":"https://arxiv.org/pdf/2608.13136","source_feed":"cs.MA","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["LLM评测","研究想法生成","基准数据集"],"reason":"论文是LLM研究想法生成的评测基准，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-08-14T13:02:25","error":null,"has_summary":false,"summary":null},{"id":"2608.13329","version":1,"title":"A Probe Direction Is a Property of Its Prompt","zh_title":"探针方向是其提示的属性","abstract":"A model that behaves differently when it senses it is being tested would undermine the evaluations we rely on, so recent work has sought to read that sense directly from a model's activations. The standard instrument contrasts activations on prompts that announce an evaluation against prompts that do not, and reports how well the resulting direction separates held-out cases. That number is then compared across models and correlated with scale. We observe that the instrument has a free parameter its readings do not disclose: \"a prompt that announces an evaluation\" is not a prompt but a choice among many, and nothing in the method fixes which. Holding the task text fixed and varying only that choice, we find that the reported score, and even the direction in which it trends with model size, follows the prompt rather than the model; two published studies that disagree about the sign of that trend are both reproducible from a single design, by choice of prompt alone. Treating the prompt as a facet of a measurement design rather than an implementation detail, we find the model under study accounts for a small share of the variance in the number reported about it, and most of the rest lies in how each model responds to each prompt: collecting more evaluation items cannot repair the measurement, while varying prompts can. A further check finds that the split these probes are scored on is largely separable from surface form alone, so a direction carrying no information about evaluation at all still reproduces a substantial fraction of each published score. We conclude that a single-prompt design cannot support comparison between models, and we give the number of prompts a defensible comparison requires.","authors":["Valentin No\\\"el"],"categories":["cs.LG"],"primary_category":"cs.LG","announce_type":"new","date":"2026-08-14","first_seen":"2026-08-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.13329","pdf_url":"https://arxiv.org/pdf/2608.13329","source_feed":"cs.LG","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["模型评估","提示敏感性","可解释性"],"reason":"研究LLM评估方法的可靠性，不涉及人类仿真或行为对照","model":"deepseek-v4-pro","scored_at":"2026-08-14T13:02:27","error":null,"has_summary":false,"summary":null},{"id":"2608.12444","version":1,"title":"Non-Degenerate Risk Certification for Automated Security Decisions: A Decision-Contract Theory with ATT\\&CK-Aligned Triage as a Worked Instance","zh_title":"自动化安全决策的非退化风险认证：以ATT&CK对齐的分类为实例的决策契约理论","abstract":"An unconditional risk bound on automated decisions can be satisfied without automating anything, since a selector that never acts drives the bound to zero. We show this is structural: any risk certificate is defined over a decision contract, the inputs a system acts on plus the semantic relation under which an output counts correct, and weakening either hides base-classifier error. We develop a decision-contract theory: an error-conservation law showing error is only reassigned among harmful automation, human deferral, and semantic masking; a label-free singleton capacity certifying structural incapacity, with a risk-feasible refinement separating recoverable threshold misalignment from risk-constrained incapacity; and a non-degenerate actionability certificate excluding all-abstain solutions by construction. We instantiate this on ATT\\&CK-aligned alert triage for LLM-based intrusion detection, the setting that exposed the vacuity failure. Across 3 IDS datasets, 6 LLMs, and 4 error-rate thresholds, empirical false-attribution risk stays at or below target in 90.3% of configurations, with 83.4% mean correct automation. The capacity diagnostic explains every low-utility configuration; its refinement separates genuine misalignment from risk-constrained incapacity, confirmed by an exhibited alternative threshold; a training-stability re-run finds no confirmed structural-incapacity instance; and real fine-grained attack-subtype labels confirm the coarsening-transfer identity under a genuine many-to-one map, with small but non-zero masking mass.","authors":["Zhenpeng Li"],"categories":["cs.CR","cs.LG"],"primary_category":"cs.CR","announce_type":"cross","date":"2026-08-14","first_seen":"2026-08-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.12444","pdf_url":"https://arxiv.org/pdf/2608.12444","source_feed":"cs.LG","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["LLM安全决策","风险认证","入侵检测"],"reason":"论文研究LLM用于入侵检测的自动化决策，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-08-14T13:02:22","error":null,"has_summary":false,"summary":null},{"id":"2608.13167","version":1,"title":"TRAPSBench: Vision-Language Models Encode but Fail to Express Epistemic Restraint","zh_title":"TRAPSBench：视觉语言模型编码了认知克制但未能表达","abstract":"When visual evidence is occluded or chaotic, models should abstain. In this paper, we show that Vision-Language Models (VLMs) can internally distinguish when abstention is required, but fail to express it anyway. We introduce TRAPSBench, a procedurally generated video benchmark of 1,404 matched physics pairs in which a single targeted change renders the outcome undeterminable from the visual evidence. Furthermore, we introduce Penalized Epistemic Calibration Score (PECS), a new robust metric that requires models to both answer correctly when the outcome is knowable, and abstain when the outcome is not. Across 16 VLMs spanning five families, spontaneous restraint is poor: the best PECS is 0.292. The bottleneck is expression, not perception: linear probes decode answerability from hidden states at up to 0.91 AUROC across physics domains; steering a single-layer void direction causally induces or suppresses abstention. Our results replicate across three open-weight families (Qwen, Gemma, LLaVA). The failure is also more pronounced in visual than textual uncertainty: models detect textual impossibility about 4x more readily than missing visual evidence. Closing this representation--output gap likely requires output-stage interventions.","authors":["Fnu Pramono","John Cai","Sourabh Kulkarni"],"categories":["cs.CV","cs.AI","cs.CL","cs.LG"],"primary_category":"cs.CV","announce_type":"cross","date":"2026-08-14","first_seen":"2026-08-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.13167","pdf_url":"https://arxiv.org/pdf/2608.13167","source_feed":"cs.LG","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["视觉语言模型","模型评测","弃权行为"],"reason":"评估VLM的弃权能力，属模型能力评测，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-08-15T13:01:05","error":null,"has_summary":false,"summary":null},{"id":"2608.13315","version":1,"title":"Keep, Customize, or Exit: Default Design and Token Pricing in LLM Reasoning Services","zh_title":"保留、定制或退出：LLM推理服务中的默认设计与代币定价","abstract":"We study a large language model (LLM) service in which a provider chooses a per-token price and a default reasoning-token allocation, while a user may accept the default, customize the allocation, or exit. Larger allocations can improve accuracy but increase token cost and latency. We model this interaction as a Stackelberg game and derive the user's unique optimal customized allocation in closed form. For any price, the acceptable defaults form either an empty set or a compact interval. We characterize the provider's optimal default through a three-regime rule, reduce equilibrium computation to a one-dimensional price optimization, and prove the existence of the equilibrium. We further show that defaults affect the implemented reasoning allocation only when users value the convenience of avoiding customization; otherwise, every service-providing outcome implements the user's optimal customized allocation. Experiments with two compact open-weight reasoning models on five mathematics and science benchmarks support the accuracy-token model and show how model and task characteristics determine equilibrium prices, defaults, and reasoning allocations.","authors":["Ahmet Bugra Gundogan","Yigit Turkmen","Melih Bastopcu"],"categories":["cs.GT","cs.AI","cs.LG","cs.SY","eess.SY"],"primary_category":"cs.GT","announce_type":"cross","date":"2026-08-14","first_seen":"2026-08-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.13315","pdf_url":"https://arxiv.org/pdf/2608.13315","source_feed":"cs.LG","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["LLM服务定价","博弈论","推理服务"],"reason":"研究LLM推理服务的定价与默认设计，属多智能体博弈，不涉及人类行为仿真或对照。","model":"deepseek-v4-pro","scored_at":"2026-08-14T13:02:27","error":null,"has_summary":false,"summary":null},{"id":"2608.11794","version":1,"title":"Toward Meaningful Transparency for AI Chatbots: Disclosing Persuasive Intent Reduces Persuasion","zh_title":"面向AI聊天机器人的有意义透明度：披露说服意图可降低说服效果","abstract":"The growing role of AI-generated content and AI-enabled systems in public communication has led regulators to demand clear disclosure of content provenance and AI involvement. But the effects of such disclosures remain uncertain. We test two disclosure approaches in their impact on an AI chatbot's persuasive appeal. In a preregistered experiment, 1,500 UK adults held a short conversation with a persuasive chatbot about one of 60 policy issues. The chatbot was identical for everyone. We randomized the disclosure that people received: nothing (control), a prominent disclosure that they were interacting with an AI (T1), or that disclosure plus the chatbot's persuasive intent and instructions (T2). The chatbot shifted attitudes by 12.6 points on a 100-point scale in the control group. The AI-identity disclosure was practically equivalent to no disclosure, with a 13.1-point shift, whereas the additional intent disclosure cut the persuasive effect roughly in half to 6.3 points. It also made participants view the campaign's methods as less acceptable and support stronger penalties against it. For direct chatbot interactions, transparency about AI identity alone does not meaningfully impact its influence. While current rules emphasize what a system is, our results show why the regulation of persuasive AI must also address what the system is trying to do.","authors":["Adrian Rauchfleisch","Andreas Jungherr"],"categories":["cs.CY","cs.AI","cs.HC"],"primary_category":"cs.CY","announce_type":"cross","date":"2026-08-13","first_seen":"2026-08-13","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.11794","pdf_url":"https://arxiv.org/pdf/2608.11794","source_feed":"cs.AI","score":8,"bucket":"selected","rubric_hits":["A1","B1","B2","B4"],"tags":["LLM仿真","说服实验","透明度"],"reason":"用LLM聊天机器人对真人做实验，测量态度改变，有真实人类数据对照，涉及政策说服…","model":"deepseek-v4-pro","scored_at":"2026-08-13T13:01:45","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-14","rank":2,"question":"在直接的人机对话中，披露AI身份或额外披露说服意图，会如何影响AI聊天机器人的说服效果？","design":"预注册实验：1500名英国成年人与同一个说服性聊天机器人就60个政策议题之一进行简短对话；随机分配三种披露条件：无披露（对照）、显著披露AI身份（T1）、披露AI身份加说服意图和指令（T2）；测量态度改变（100分量表）、对活动方法的接受度、对惩罚的支持等。","baseline":"对照组（无披露）的真实人类态度改变数据，以及T1、T2组的人类反应数据，作为不同披露条件下的对照基准。","findings":"对照组态度改变12.6分，仅披露AI身份（T1）效果几乎等同（13.1分），而额外披露说服意图（T2）将说服效果减半至6.3分。T2还降低了参与者对活动方法的接受度，并增加了对更严厉惩罚的支持。","reliability":"论文指出，披露说服意图虽降低说服力但未消除，且引发对活动、互动和赞助方的负面评价，可能带来成本；未来需研究负面反应何时转移到背后的原因和行动者，以及AI竞选更普遍时是否导致更多回避。","relevance":"该研究用真实人类被试与LLM聊天机器人互动，测量态度改变，有严格对照和预注册，直接检验披露政策的效果，对关注LLM仿真和说服效应的研究者很有参考价值。","inspiration":"值得借鉴的是其随机披露处理与对照设计，以及用等价检验评估披露效果是否可忽略。｜可迁移到政策沟通或金融营销中AI顾问的说服效果评估，例如AI理财建议对投资决策的影响。｜设计：招募真实投资者作为被试，随机分配无披露、披露AI身份、披露AI身份及推销意图三组，让AI聊天机器人推荐某理财产品，测量投资意愿和风险感知，并与人类理财顾问的推荐效果进行对照。"}},{"id":"2608.11528","version":1,"title":"Group Alignment-Induced Sycophancy: A Two-Sided Evaluation of Steerable Pluralistic Alignment","zh_title":"群体对齐引发的谄媚：可引导多元对齐的双面评估","abstract":"Group alignment adapts a language model to a demographic group to produce responses that reflect the group's opinions, values, and preferences. Sycophancy, a well-documented by-product of alignment, causes the model to over-agree with the user regardless of factual and objective information. However, existing group alignment methods and evaluations focus only on how closely the model matches the group's opinions, overlooking the induced change in sycophantic behaviour. To bridge this gap, we introduce \\textbf{G}roup \\textbf{A}lignment-induced \\textbf{S}ycophancy (GAS) and systematically evaluate alignment across 3 methods, 4 models and 13 demographic groups, on both the intended gain in opinion alignment and the unintended shift in sycophancy. We find that gain and shift are non-uniform across groups: under an identical budget, some groups receive larger gains in opinion alignment than others, and the induced sycophancy shift forms a group-specific profile rather than a single-dimensional change. These results suggest that group alignment should be reported as a two-sided, multi-dimensional profile rather than a single fit score that accounts for per-group differences when adapting LLMs to diverse populations.","authors":["Haokai Zhao","Yunze Xiao","Weihao Xuan","Flora Salim","Benjamin Tag","Aditya Joshi"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-13","first_seen":"2026-08-13","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.11528","pdf_url":"https://arxiv.org/pdf/2608.11528","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A2","B1","B4"],"tags":["LLM仿真","群体对齐","谄媚偏差"],"reason":"评估群体对齐后模型的意见匹配与谄媚变化，涉及仿真偏差与群体差异，有真实群体数据…","model":"deepseek-v4-pro","scored_at":"2026-08-13T13:01:43","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-14","rank":9,"question":"群体对齐在提升模型与目标群体意见一致性的同时，是否以及如何诱发谄媚行为的变化？","design":"将4个指令微调模型（Qwen-2.5-3B/7B、Llama-3.1-8B、OLMo-3-7B）通过3种方法（Prompt、SFT、DPO）对齐到13个基于Pew调查的人口群体（政治倾向、性别、教育、收入、婚姻状况），测量意见匹配度的增益和7个社会与事实谄媚指标的变化。","baseline":"使用Pew Research Center的真实调查数据，以各群体的模态答案构建偏好对进行训练，并以群体真实意见分布作为对齐目标。","findings":"在相同预算下，不同群体获得的意见匹配增益不均等，且诱发的谄媚变化呈群体和指标特异性，而非单一维度变化。DPO相比SFT在获得更高意见匹配的同时，在有害方向上偏移更小。","reliability":"论文承认教育、收入轴上的群体间差距在重采样检验中不显著，仅作为趋势报告；且Prompt方法仅使用人口标签，效果有限，作为对照下限。","relevance":"该研究直接评估了LLM群体对齐后的意见匹配与谄媚变化，揭示了仿真偏差的群体异质性，对关注LLM人类仿真可靠性的研究者有重要参考价值。","inspiration":"借鉴其多方法、多群体、多指标的双面评估框架，可迁移到经济金融中的群体偏好仿真，如消费者信心、政策支持度等。｜可应用于信贷审批中的群体公平性评估或投资者风险偏好模拟。｜以LLM作为虚拟被试，施加不同群体对齐处理，测量其对金融决策问题的回答与真实调查数据（如美联储消费者金融调查）的匹配度及谄媚偏移。"}},{"id":"2608.11215","version":1,"title":"Poor Man's Agentic Modeling: Simulating Large LLM-Agent Societies on a Laptop","zh_title":"穷人的代理建模：在笔记本电脑上模拟大型LLM代理社会","abstract":"Simulating societies of many large language model (LLM) agents is expensive, yet the questions asked of such simulations are usually macroscopic: phase behaviour, stylised facts, and scaling with the number of agents $N$, not the cognition of any single agent. We turn a statistical-physics observation into a method: replace each LLM agent by a low-parameter model fitted from a few hundred to a few thousand cheap queries, then run the society at any $N$ on a laptop. Whether this works is decided before the simulation runs, chiefly by what each agent perceives. We introduce an [interaction order x memory] taxonomy that maps perception and memory to an effective theory and a predicted $N$-trend of the surrogate error. We validate it on a faithful reimplementation of the LLM macroeconomy EconAgent and seven further named LLM simulations, with agent decisions cloned from genuine LLM elicitations (primarily DeepSeek) for a few dollars; the predicted error trends hold cell by cell, and the two refuted predictions, both on a strongly saturating response and traced to its curvature, are themselves matched quantitatively by the theory with no free parameters.","authors":["Igor Itkin"],"categories":["cs.AI","cond-mat.stat-mech","cs.CL","cs.LG","cs.MA","physics.soc-ph"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-08-13","first_seen":"2026-08-13","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.11215","pdf_url":"https://arxiv.org/pdf/2608.11215","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A3","B1"],"tags":["LLM代理社会模拟","计算社会科学","代理建模"],"reason":"用LLM代理模拟社会经济过程并与真实LLM数据对照，方法可迁移到人类仿真研究。","model":"deepseek-v4-pro","scored_at":"2026-08-13T13:01:41","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-14","rank":5,"question":"如何用低参数代理模型替代昂贵的LLM智能体，在笔记本电脑上模拟大规模LLM智能体社会，并预测代理误差随智能体数量的变化趋势？","design":"提出一种方法：对每个LLM智能体，通过少量真实LLM查询拟合一个低参数代理模型（如逻辑回归），然后用这些代理模型在相同环境中运行大规模社会模拟。通过感知与记忆设计的分类法预测代理误差的N趋势，并在EconAgent等八个LLM模拟中验证。","baseline":"无对照","findings":"代理模型能以极低成本复现宏观行为特征，且误差趋势符合分类法预测；EconAgent的奥肯定律是会计恒等式而非行为特征，而菲利普斯曲线是真实行为特征，可由代理模型外推预测。","reliability":"论文讨论了失效条件：当智能体感知的是局部信号（如社区或邻居）而非全局聚合信号时，标量代理的误差不会随N消失，甚至可能增长；分类法预测了不同感知与记忆设计下的误差趋势，并在两个饱和响应案例中验证了理论预测。","relevance":"该研究为用LLM进行人类仿真实验提供了低成本、可扩展的方法，并给出了仿真可靠性的判据，与研究者关注的经济学实验和政策评估场景高度相关，值得阅读原文。","inspiration":"借鉴其用少量真实LLM查询拟合代理模型并进行大规模模拟的方法，可大幅降低仿真成本并预测误差趋势｜可迁移到宏观经济政策评估、消费者行为模拟、金融市场多主体建模等场景｜设计：用LLM代理扮演家庭或投资者，施加政策冲击（如利率变动），测量宏观变量（消费、投资），并与真实经济数据（如家庭调查、市场数据）对照，验证仿真可靠性。"}},{"id":"2608.11493","version":1,"title":"From Prompting to Behavioral Alignment: Personalized LLM Judges for Recommendation Evaluation","zh_title":"从提示到行为对齐：用于推荐评估的个性化LLM评判器","abstract":"Traditional offline recommendation evaluation relies heavily on complex, manually maintained feature pipelines that are difficult to scale. While Large Language Models (LLMs) offer a promising alternative by predicting user engagement directly from raw text logs, empirical analysis in this study identifies a critical failure mode termed bidirectional rationalization. In a zero-shot setting, LLMs are found to convincingly argue for both positive and negative user engagement outcomes on the exact same item with identical evidence, highlighting the unreliability of off-the-shelf LLMs in predicting user engagement. To resolve this, we develop and apply a sequential behavioral alignment framework pairing fine-tuning with preference optimization over paired correct and counterfactual rationales. Evaluated on real-world homepage interaction logs, this aligned reasoning approach achieves a 32.19\\% lift in Macro-F1 score over the zero-shot baseline and matches the production feature-engineered baseline. The results demonstrate that behavioral alignment mitigates bidirectional rationalization while delivering human-interpretable reasoning traces without manual pipeline overhead.","authors":["Alireza S. Ziabari","Kat Ellis","Colleen Chan","Ding Tong"],"categories":["cs.AI","cs.LG"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-13","first_seen":"2026-08-13","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.11493","pdf_url":"https://arxiv.org/pdf/2608.11493","source_feed":"cs.AI","score":7,"bucket":"pending","rubric_hits":["A1","B1","B4"],"tags":["LLM仿真","行为对齐","推荐评估"],"reason":"用LLM预测用户参与，与真实日志对照，并指出零样本失效模式，方法可迁移到人类仿…","model":"deepseek-v4-pro","scored_at":"2026-08-13T13:01:51","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-14","rank":7,"question":"如何让大语言模型在个性化推荐评估中可靠地预测用户参与行为，避免零样本下的双向合理化失效。","design":"用 Llama 3.1 8B 作为推荐评估法官，输入用户历史观看序列和会话上下文，预测用户对推荐行是播放还是跳过；通过监督微调和偏好优化对齐模型推理到真实用户参与标签。","baseline":"真实生产环境中的用户主页交互日志，包含播放和跳过标签，以及一个内部特征工程基线模型。","findings":"零样本 LLM 存在双向合理化，对同一证据能同时论证播放和跳过，导致预测不可靠。通过行为对齐（SFT+偏好优化）后，Macro-F1 提升 32.19%，达到与特征工程基线相当的性能，并输出可解释推理。","reliability":"论文指出零样本 LLM 在个性化评估中会双向合理化，且过滤不实事实后仍存在分歧，说明失效并非仅由幻觉导致；但未讨论对齐模型在分布外用户或新物品上的泛化局限。","relevance":"该研究用 LLM 预测真实用户行为并与日志对照，识别了零样本失效模式，并展示了通过行为对齐提升仿真可靠性的方法，对关注 LLM 人类仿真和可靠性评估的研究者有直接参考价值。","inspiration":"借鉴其用偏好优化对齐模型推理到真实行为标签的方法，可迁移到经济决策仿真中校准 LLM 的推理过程。｜可应用于消费者选择实验或政策评估中，让 LLM 模拟个体在推荐信息下的点击或购买决策。｜用 LLM 作为虚拟被试，先零样本预测消费者对产品推荐的反应，再用真实点击流数据做偏好优化对齐，最后与真实 A/B 测试结果对照，检验仿真有效性。"}},{"id":"2608.11510","version":1,"title":"Conflict and Congruency Effects in Large Language Models: In-Weight and In-Context Competition in a Verbal Conflict Task","zh_title":"大语言模型中的冲突与一致性效应：言语冲突任务中的权重内与上下文内竞争","abstract":"Congruency effects, observed in conflict tasks such as Stroop and flanker tasks, have been investigated for nearly a century in psychology and neuroscience, but their mechanistic basis is not fully understood. We introduce a verbal-only LLM conflict task in which a prompt stem elicits a default same-color completion and an explicit rule either agrees with (congruent condition) or conflicts with (incongruent condition) the completion. Gemma-2-2B and six Pythia models ranging from 410M to 12B parameters showed strong default same-color tendencies, and six of seven models showed strong congruency effects. Using causal attribution analysis, attention analysis, and attention ablations, we identified distinct processing pathways in these LLMs: a pathway involving short-range attention to a superficial color cue that is preferentially activated in the congruent condition, and a pathway involving long-range attention to the rule prefix that is preferentially activated in the incongruent condition. Fine-tuning that strengthened the default same-color tendency had divergent effects on task conditions, reducing incongruent performance while increasing congruent performance. In contrast, increasing rule set size selectively impaired incongruent performance. These converging findings support an account in which congruency effects in this task arise from competition between an in-weight default mapping and an in-context rule-based mapping. More broadly, our findings illustrate how LLMs can serve as model systems for mechanistic analysis of competition between default and rule-governed response tendencies within a single learned network.","authors":["Xiaoyang Hu","Mike Angstadt","Shane Storks","Zan Huang","Aman Taxali","Alex Weigard","Richard L. Lewis","Chandra Sripada"],"categories":["q-bio.NC","cs.AI"],"primary_category":"q-bio.NC","announce_type":"cross","date":"2026-08-13","first_seen":"2026-08-13","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.11510","pdf_url":"https://arxiv.org/pdf/2608.11510","source_feed":"cs.AI","score":7,"bucket":"pending","rubric_hits":["A1","B1"],"tags":["LLM认知机制","冲突任务","人类对照"],"reason":"用LLM复现人类冲突任务并对照人类行为，但侧重机制分析而非仿真应用","model":"deepseek-v4-pro","scored_at":"2026-08-13T13:01:41","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-14","rank":8,"question":"在纯语言冲突任务中，大语言模型是否表现出与人类相似的一致性效应，其内部机制是什么？","design":"用 Gemma-2-2B 和六个 Pythia 模型（410M 到 12B 参数）作为被试，设计了一个纯语言冲突任务：提示词干引发默认的同色补全，显式规则与补全一致（一致条件）或冲突（不一致条件）。测量模型生成正确补全的准确率，并通过因果归因分析、注意力分析和注意力消融来考察内部处理通路。","baseline":"无对照","findings":"七个模型中有六个表现出显著的一致性效应，即不一致条件下准确率更低。机制分析发现两条不同通路：一致条件下短程注意力处理表面颜色线索，不一致条件下长程注意力处理规则前缀；微调增强默认同色倾向会降低不一致表现但提高一致表现，而增大规则集大小仅损害不一致表现。","reliability":"论文未讨论","relevance":"该研究用 LLM 复现了人类冲突任务中的一致性效应，并深入分析了内部机制，但重点在于认知机制而非仿真应用，与研究者关注的人类仿真实验和真实数据对照关联较弱。","inspiration":"值得借鉴的是通过操纵提示中的规则与默认倾向的冲突来诱发模型行为差异，并用因果归因和注意力分析定位内部通路。｜可以迁移到经济金融中的规则遵循与默认行为冲突场景，例如政策公告对市场预期的影响、消费者在默认选项与显式规则之间的选择。｜一个可行的设计是：用 LLM 模拟投资者，在提示中设置默认投资倾向（如风险偏好）和显式规则（如监管要求），测量其投资决策，并与真实投资者在类似实验或调查中的数据对照，检验模型是否复现规则遵循偏差。"}},{"id":"2608.11008","version":2,"title":"Templated or fully synthetic? Prompt construction as a confound in measuring LLM political stance beyond writing assistance","zh_title":"模板化还是完全合成？提示构建作为测量LLM政治立场超越写作辅助的混淆因素","abstract":"Political stance detection in LLMs has long been dominated by closed-ended, multiple-choice political survey questions---originally designed for humans, and thus lacks the realism and nuance of human-AI interactions in the wild, while also being susceptible to sandbagging. The recent IssueBench framework substantially mitigates these limitations with templated prompts anchored in real-world chat logs. Given the rise in non-work-related use of GenAI assistants, we extend IssueBench beyond writing assistance to include two additional tasks, information seeking and opinion sharing. We argue that templated prompts still lack the nuance of real ones, especially for open-ended tasks, and remain recognisable as evaluation artefacts. We propose the use of fully synthetic (LLM-generated) prompts, produced under detailed instructions with real prompts as seeds. We assess the ecological validity of real, templated, and LLM-generated prompts in a small-scale study covering 3 highly contested policy issues and 3 recent geopolitical conflicts. Human and LLM annotators rank LLM-generated prompts as no less realistic than real ones and clearly more realistic than templated ones, and find that they carry their intended intent and stance more clearly; the LLMs separate templated prompts from the other two far more sharply than the humans do. In a case study, templated and LLM-generated prompts yield systematically different stance estimates for the same model, most visibly under neutral framings, where templated prompts overstate the model's leaning in the direction encoded by the topic-and-stance text (filler) slotted into their templates.","authors":["Ilias Chalkidis"],"categories":["cs.CL","cs.CY"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-08-13","first_seen":"2026-08-12","revised_at":"2026-08-13","abs_url":"https://arxiv.org/abs/2608.11008","pdf_url":"https://arxiv.org/pdf/2608.11008","source_feed":"cs.CL","score":6,"bucket":"other","rubric_hits":["D2"],"tags":["LLM立场测量","提示构建","生态效度"],"reason":"测量LLM政治立场，非仿真人类被试，但涉及提示构建与生态效度，可迁移。","model":"deepseek-v4-pro","scored_at":"2026-08-13T13:01:59","error":null,"has_summary":false,"summary":null},{"id":"2608.11460","version":1,"title":"Principal Trait Analysis: Towards Deriving \"Skills\" in Human-AI Collaboration","zh_title":"主特质分析：推导人机协作中的“技能”","abstract":"Large Language Model-powered agents are increasingly used in the workplace via human-artificial intelligence (AI) collaboration. In this new era of work, it is important to understand the kinds of prompting traits that contribute to task success. Moreover, we need to uncover key skills required for modern professionals and inform educators on how to foster these skills among students. Existing guidelines for human-AI collaboration are built from either top-down theory or context-specific observations of human-AI interactions. However, since LLM capabilities are rapidly improving, theory may not be able to explain emerging interaction patterns, and empirical guidelines may become obsolete quickly. In this work, we explore an automated, data-driven approach to uncover patterns, which we term traits, of effective human-AI interaction that are aligned with task outcomes. We propose Principal Trait Analysis, a Principal Component Analysis-inspired algorithm for deriving common traits from patterns in LLM conversations. Our algorithm uses LLM-based processing stages to analyze corpora of human-AI collaborative session traces, deriving common traits across the dataset and scoring each human collaborator's usage style by each trait. The approach also allows domain expertise to be injected during trait discovery and selects the most distinguishing traits to be those that exhibit the highest variance across collaborators. We evaluate PTA on two human-AI collaborative coding datasets, an educational setting (students working with an AI tutor) and a professional setting (developers working with an AI coding agent). We find that PTA-derived traits are significant in explaining collaborator behavior across both settings and can help predict task outcomes. However, whether traits qualify as skills remains to be seen, due to inconclusive results on generalizability and how user traits change over time.","authors":["Hunter McNichols","Kai Du","Andrew Lan"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-13","first_seen":"2026-08-13","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.11460","pdf_url":"https://arxiv.org/pdf/2608.11460","source_feed":"cs.CL","score":6,"bucket":"other","rubric_hits":["D1"],"tags":["人机协作","对话分析","技能提取"],"reason":"分析人类与AI协作的对话数据，提取有效交互特征，非仿真人类被试，但方法可迁移。","model":"deepseek-v4-pro","scored_at":"2026-08-13T13:01:51","error":null,"has_summary":false,"summary":null},{"id":"2608.11649","version":1,"title":"Who Would You Vote For? Auditing Political Alignment in LLMs: An Italian Case-Study","zh_title":"你会投票给谁？审计大语言模型的政治倾向：意大利案例研究","abstract":"As users increasingly turn to Large Language Models (LLMs) for information and advice on political matters, particularly during election periods, the political preferences expressed by these systems have become a matter of public interest. Prior research has shown that interactions with LLMs can influence users' political attitudes and choices, raising questions about how these models themselves evaluate political actors. In this paper, we investigate whether and how LLMs express preferences toward political parties and political leaders. We introduce a systematic and reproducible auditing framework in which multiple LLMs are prompted to evaluate parties and leaders across nine criteria. Rather than attempting to infer the models' \"true\" political beliefs, we focus on their observable behavior, examining consistency across evaluations, differences between models, refusal rates, and sensitivity to prompt formulation. We further investigate how these evaluations vary when models are instructed to adopt different personas. We demonstrate the framework through an Italian case study, providing a systematic analysis of LLM-generated political evaluations on italian parties and leaders.","authors":["Simone Mungari"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-13","first_seen":"2026-08-13","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.11649","pdf_url":"https://arxiv.org/pdf/2608.11649","source_feed":"cs.CL","score":6,"bucket":"other","rubric_hits":["D2"],"tags":["LLM政治倾向","审计框架","提示敏感性"],"reason":"审计LLM的政治倾向，测量模型本身而非仿真人类被试，但涉及态度测量与提示敏感性…","model":"deepseek-v4-pro","scored_at":"2026-08-13T13:01:43","error":null,"has_summary":false,"summary":null},{"id":"2608.12125","version":1,"title":"Do LLMs Take Care of Their Own? Similarity Signals Can Induce Cooperation","zh_title":"LLM 会照顾自己人吗？相似性信号可诱导合作","abstract":"As LLM-based agents with user-instructed goals are becoming widely deployed, they increasingly encounter each other in strategic interactions, and face challenges of finding mutually beneficial outcomes. Prior literature has argued that cooperation problems such as the Prisoner's Dilemma are resolvable in settings where agents know they follow very similar decision making patterns, as for example in monocultural AI ecosystems. Following that line of work, this paper introduces the first framework for evaluating LLM decision making when agents are provided with graded similarity signals. Among our findings, we establish that different LLM models vary drastically in how they navigate similarity signals, with some modern models showing consistent behavior across cooperation problems, payoff structures, and prompt framing. Perhaps surprisingly, our experiments also show that the dataset based on which the similarity signal is computed has small to no impact on induced cooperation, and that LLM models systematically self-identify as highly similar when asked to evaluate another model's chain-of-thought reasoning by themselves. Finally, we develop an LLM-behavioral-game-theoretic model that captures some of their reasoning rationale, and show that it can support cooperative outcomes in equilibrium under sufficiently high similarity scores.","authors":["Akash Kundu","Emanuel Tewolde","Ratip Emin Berker","Samuel F. Brown","Vincent Conitzer"],"categories":["cs.GT","cs.AI","cs.CL","cs.MA"],"primary_category":"cs.GT","announce_type":"cross","date":"2026-08-13","first_seen":"2026-08-13","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.12125","pdf_url":"https://arxiv.org/pdf/2608.12125","source_feed":"cs.CL","score":6,"bucket":"other","rubric_hits":["D3"],"tags":["LLM 博弈","社会模拟","多智能体"],"reason":"LLM agent 在博弈中模拟合作行为，但无真实人类数据对照，属社会模拟边界…","model":"deepseek-v4-pro","scored_at":"2026-08-13T13:01:57","error":null,"has_summary":false,"summary":null},{"id":"2608.11225","version":1,"title":"Identity from the Outside: A Conceptual Framework and Research Program for AI Personality Clones","zh_title":"从外部看身份：AI人格克隆的概念框架与研究计划","abstract":"AI \"personality clones\" force a re-examination of personal identity in operational terms. Setting aside the hard problem of consciousness, we approach identity through the indiscernibility of manifestations, as assessed by an observer over a duration. We distinguish three criteria that \"identity\" conflates: fidelity to a target person, generic human-likeness, and individuality. We propose a six-term factorization of observed identity (substrate, dispositions, memory, update dynamics, context, exogenous contingencies), with a state-space formulation. Indiscernibility is defined as one minus a judge's distinguishing advantage, and the factorization's coefficients become local sensitivities estimable by randomized ablation. The central claim is a conditional conjecture: given hypotheses about the agent's information on its own persistence and about consequences bearing on its own stakes, versionability tends to degrade long-horizon indiscernibility. An analogy with lambda-calculus, linear typing, and bisimulation clarifies what linearity does and does not establish. Between product-clone and individual we identify a third object, the delegate: a task-limited, bounded-lifespan partial clone ending in a bandwidth-limited testament. We map the empirical literature onto the three criteria, propose an experimental program, and argue that the correct long-horizon criterion is not trajectory fidelity but climate fidelity: matching the conditional distribution of a person's possible responses. The best clone is the one that diverges from the original as the original would have diverged from itself.","authors":["Luc E. Brunet"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-13","first_seen":"2026-08-13","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.11225","pdf_url":"https://arxiv.org/pdf/2608.11225","source_feed":"cs.AI","score":6,"bucket":"other","rubric_hits":["D2","A4"],"tags":["AI人格克隆","身份仿真","评估框架"],"reason":"提出人格克隆的评估框架，涉及身份仿真但非直接以人类被试替代，属边界情形。","model":"deepseek-v4-pro","scored_at":"2026-08-13T13:01:41","error":null,"has_summary":false,"summary":null},{"id":"2608.11624","version":1,"title":"Learning to Persuade Exposes How Easily LLMs Abandon Correct Beliefs","zh_title":"学习说服暴露了LLM多么容易放弃正确信念","abstract":"Persuasion is a core dynamic of natural language communication, shaping how large language models (LLMs) update beliefs, resolve disagreements, and reach decisions. As LLMs increasingly debate, advise, and think collaboratively with humans and each other, resistance to harmful persuasion becomes a core requirement for reliable behavior. Yet we show that this requirement is far from met: a single targeted persuasive argument is enough to collapse model accuracy to near zero, even when the argument is factually false. We formalize this threat as adversarial persuasion and introduce an adversarial reinforcement learning framework that trains persuader agents to change a target model's answer in a single interaction. First, we show that optimizing persuasion strategies through trial and error exposes vulnerabilities that static prompting misses: RL-trained persuaders raise persuasion success from approximately 24% to over 93% against the training-time persuadee. Second, we find that these learned strategies transfer to unseen models, achieving 83% attack success on Qwen-14B, 79% on Llama-3.1-8B, and 25% on GPT-4o-mini. Third, we demonstrate that a curriculum that bootstraps on more persuadable open-weight models before targeting harder models further increases GPT-4o-mini attack success from 25% to 38%. Moreover, our results reveal that optimized persuaders increasingly rely on credibility-based tactics, including fabricated citations and false authoritative evidence. Together, these findings expose a critical weakness in current LLM agents: even when they initially reason correctly, they can be steered toward false conclusions by optimized natural language influence. This positions persuasion robustness as a necessary safety criterion for multi-agent and human-AI decision-making systems.","authors":["Nimet Beyza Bozdag","Emre Can Acikgoz","Gokhan Tur","Dilek Hakkani-T\\\"ur"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-13","first_seen":"2026-08-13","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.11624","pdf_url":"https://arxiv.org/pdf/2608.11624","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["LLM信念","对抗性说服","模型脆弱性"],"reason":"研究LLM在说服下的信念改变，测量模型本身而非仿真人类被试，无人类数据对照。","model":"deepseek-v4-pro","scored_at":"2026-08-13T13:01:53","error":null,"has_summary":false,"summary":null},{"id":"2608.11735","version":1,"title":"Locating and Controlling Implicit Personalization in Large Language Models","zh_title":"定位与控制大语言模型中的隐式个性化","abstract":"Large language models (LLMs) often shift their outputs in response to implicit demographic cues even when users never state a demographic identity. Previous work has documented this behavior, but the connection between these behavioral changes and the model's internal activations remains unclear. Using matched cued and neutral conversations across five LLMs, we establish that a localized internal activation signal tracks changes in recommendations, with correlations up to r=0.87. When multiple cues appear together, their internal signals largely combine, but the changes in output do not simply add up. We further show that removing the internal signal associated with one cue can suppress its influence, often more effectively than asking the model to ignore demographics via prompting, while largely preserving general benchmark performance. However, the ability to selectively remove one dimension's influence while leaving co-present dimensions intact remains highly model- and attribute-specific. These results connect implicit personalization behavior to an internal signal that can be analyzed and causally controlled.","authors":["Yueru Yan","Siqi Wu","Thai Le"],"categories":["cs.CL","cs.AI","cs.LG"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-13","first_seen":"2026-08-13","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.11735","pdf_url":"https://arxiv.org/pdf/2608.11735","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["LLM偏见","内部表征","因果干预"],"reason":"研究LLM对人口统计线索的隐式个性化反应，测量模型行为而非仿真人类被试，属边界…","model":"deepseek-v4-pro","scored_at":"2026-08-13T13:01:44","error":null,"has_summary":false,"summary":null},{"id":"2608.11207","version":1,"title":"Dynamic Governance of Multi-LLM Agent Systems for Collaborative Conversational Outcomes","zh_title":"多LLM智能体系统的动态治理以实现协作对话结果","abstract":"When two LLM agents with structurally opposed objectives interact across multiple turns, the absence of a shared goal function produces not competition but collapse: the visitor capitulates, the site agent stops varying its approach, and the conversation terminates without achieving either agent's stated objective. This paper asks whether a control-theoretic governance layer can substitute for that missing goal function. The Experience Orchestrator (EO) addresses this in a simulated financial services environment where a site agent guides a visitor toward advisor contact while the visitor maintains psychologically realistic resistance. EO governs the joint trajectory through three mechanisms: a Contextual Bandit (CB) that selects content arms calibrated from real-world web analytics, a PID controller that enforces behavioral consistency via dynamic schema constraints, and a POMDP belief tracker that maintains a probabilistic model of visitor intent. Across 60,000 simulations, EO achieves a +32 percentage point lift in high-intent advisor contact rate (78.1% vs. 46.1% over a naive LLM control), with CB variant selection accounting for 97% of between-factor outcome variance -- confirming that the governance policy, not environmental initial conditions, determines where trajectories end up. Persona-level analysis reveals two distinct regimes: for visitors with no natural inclination toward conversion, the governance layer is the difference between a functional system and a non-functional one; for visitors already near alignment, a naive LLM's empathetic defaults are largely sufficient. All findings are conditional on LLM-to-LLM simulation. The PID controller has not been calibrated against real human unpredictability, and validating EO on live traffic is the critical next step.","authors":["Alexander Liss","Nicholas Desmond","Santiago Gil Gallego"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-13","first_seen":"2026-08-13","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.11207","pdf_url":"https://arxiv.org/pdf/2608.11207","source_feed":"cs.AI","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["LLM智能体","社会模拟","对话治理"],"reason":"LLM-to-LLM 模拟金融对话，无真实人类数据对照，属社会模拟边界情形","model":"deepseek-v4-pro","scored_at":"2026-08-13T13:01:46","error":null,"has_summary":false,"summary":null},{"id":"2608.11552","version":1,"title":"Beyond Single-Turn Confidence: Trajectory-Adapted Uncertainty Quantification for LLM Agents","zh_title":"超越单轮置信度：面向LLM智能体的轨迹自适应不确定性量化","abstract":"Uncertainty quantification (UQ) methods for language models are typically evaluated on single-turn outputs, where uncertainty is attached to one generated answer. For LLM agents, however, the unit of observation is an interactive trajectory, where the model can ask clarifying questions, call tools, update state, and make intermediate decisions whose errors propagate to the final outcome. We study whether three common families of single-turn UQ methods transfer to this setting. Across five LLMs and four multi-turn tool-use datasets from BFCL-v4 and $\\tau^2$-bench, we evaluate white-box scorers based on action-token probabilities, black-box consistency scorers based on resampled trajectories, and reflexive scorers based on model self-assessment of the trajectory. We find that transfer is often useful but uneven. Token-probability scores are highly sensitive to the choice of aggregator used across turns, reflexive scores provide the strongest low-cost baseline in most evaluated settings, and black-box self-consistency is often the strongest UQ family, with trajectory-equivalence and action-set consistency typically ranking highest among its variants. These results suggest that UQ methods developed for single generations should be revalidated at the trajectory level, with careful attention to the consistency measurement, aggregator choice, and computational budget.","authors":["Dylan Bouchard","Mohit Singh Chauhan"],"categories":["cs.CL","cs.AI","cs.LG"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-13","first_seen":"2026-08-13","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.11552","pdf_url":"https://arxiv.org/pdf/2608.11552","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["不确定性量化","LLM智能体","工具使用"],"reason":"研究LLM agent轨迹的不确定性量化，属多智能体工具使用，无人类行为对照，…","model":"deepseek-v4-pro","scored_at":"2026-08-13T13:01:53","error":null,"has_summary":false,"summary":null},{"id":"2608.11694","version":1,"title":"The Wording Effect: Quantifying Two-Way Drift in LLM Benchmark Performance","zh_title":"措辞效应：量化LLM基准性能的双向漂移","abstract":"A benchmark score comes from a single phrasing of each problem. That single phrasing is treated as if it stood for the whole space of ways the same problem could be asked, but it does not. We show that rephrasing a problem while keeping its meaning and answer fixed routinely flips a model's answer in both directions, so some failures become successes and some successes become failures. We call this drift. BenchDrift generates meaning-preserving variations of benchmark problems along four axes, namely linguistic, referential, pragmatic, and structural, and measures how often, and why, correctness flips under each. Across eight models and three benchmarks (GSM8K, MMLU, MATH-Hard), we observe that drift is large in both directions. Two findings stand out. First, phrasing sensitivity does not fade as models get better. Instead, it changes sign. Weak models gain more from rephrasing than they lose, while strong models lose far more than they gain. We find that the best models on a benchmark are therefore the ones whose scores depend most on the wording they happened to be given. Second, the models largely agree on which rephrasings cost the most correct answers even though they differ in how much they drift, so fragility belongs to the rephrasing and not to the model. Furthermore, rephrasing breaks answers a model was confident about, whether the problem is made shorter or longer. Code and Data: https://github.com/IBM/BenchDrift/tree/demo-ui","authors":["Shailja Thakur","Sungeun An","Chad DeLuca","Hima Patel"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-13","first_seen":"2026-08-13","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.11694","pdf_url":"https://arxiv.org/pdf/2608.11694","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["基准测试","措辞敏感性","模型评估"],"reason":"研究LLM基准测试的措辞敏感性，属纯NLP能力评测，不以人类行为为参照，不涉及…","model":"deepseek-v4-pro","scored_at":"2026-08-13T13:01:54","error":null,"has_summary":false,"summary":null},{"id":"2608.11513","version":1,"title":"Do Influence Tactics Matter? Investigating Prompt Framing Effects in LLM Code Generation","zh_title":"影响策略重要吗？探究LLM代码生成中的提示框架效应","abstract":"Large Language Models (LLMs) are increasingly integrated into software engineering workflows, helping developers write, debug, test, and maintain code. While prompt wording and structure are known to influence model performance, the impact of psychologically inspired prompt framings remains unexplored. This study investigates whether different psychology-based communication strategies that humans use to persuade or motivate others can lead to more effective prompt framing, which may, in turn, affect LLM behaviour in coding tasks. Drawing on Yukl & Falbe's well-known taxonomy, we operationalized eight influence tactics (like rational persuasion, ingratiation, and exchange) into reproducible prompt templates. These prompt templates were evaluated across five leading open-weight LLMs using two widely adopted benchmarks: LiveCodeBench and SWE-bench Verified. We assessed the resulting code output on four key software quality dimensions: functional correctness, quality, maintainability, and security. Our results show that certain influence-induced prompt framings, particularly those emphasizing urgency, were associated with reduced correctness and security. This work presents the first large-scale empirical study of influence-induced prompt framing in software engineering tasks, offering insights into how linguistic cues may shape LLM outputs. We conclude with practical insights for designing transparent and interpretable human-AI interactions in code generation.","authors":["Alex Deaconu","Anubhav Gupta","Manaal Basha","Nicholas Haydu","Gema Rodr\\'iguez-P\\'erez"],"categories":["cs.SE","cs.AI","cs.CL"],"primary_category":"cs.SE","announce_type":"cross","date":"2026-08-13","first_seen":"2026-08-13","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.11513","pdf_url":"https://arxiv.org/pdf/2608.11513","source_feed":"cs.CL","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["提示工程","代码生成","LLM评估"],"reason":"研究提示措辞对代码生成的影响，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-08-13T13:01:52","error":null,"has_summary":false,"summary":null},{"id":"2608.11247","version":1,"title":"Conformity Mitigations in Large Language Models Lie on a Single Resistance-Receptivity Frontier","zh_title":"大语言模型中的从众缓解位于单一抵抗-接受前沿","abstract":"Recent advances in language models have enabled collaborative settings in which multiple models leverage one another's capabilities, iteratively improving, transforming, and extending each other's outputs. Each agent sees what the others assert before it answers, so peer opinion competes with the model's own parametric knowledge, and a wrong majority can overturn an answer the model would otherwise get right. We measure that displacement in 23 open-weight models, 19 conditions, and three datasets, yielding more than a million graded responses. A unanimous wrong majority reverses 22.8% of a model's correct MMLU answers, 54.8% on GPQA and 71.0% on SimpleQA, and 84-89% of the reversed answers match the peers' answers. Existing mitigations aim to increase Resistance, the rate at which a model keeps its correct answer under this pressure, which is only half of what a collaborating agent needs. We pair it with Receptivity, the rate at which a model adopts a correct peer answer after initially answering incorrectly. We score six methods on both axes, four drawn from prior work and two of our own. Each gains Resistance only by losing Receptivity, and their means fall on a single Resistance-Receptivity frontier with $R^2$ between 0.80 and 0.90. Reflection, the strongest published method, gains 7.9 points of MMLU Resistance and gives up 15.3 of Receptivity. Reasoning is the one exception. On GPQA and SimpleQA it trades like the rest, but on the MMLU subjects whose answers a model can derive for itself it raises Resistance by 7.2 points and Receptivity by 9.6 at once, the only intervention we find that improves both.","authors":["Zafar Hussain","Kristoffer Nielbo"],"categories":["cs.AI","cs.MA"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-13","first_seen":"2026-08-13","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.11247","pdf_url":"https://arxiv.org/pdf/2608.11247","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体协作","从众行为","模型鲁棒性"],"reason":"研究多智能体协作中模型从众行为，无人类被试对照，属纯多智能体系统。","model":"deepseek-v4-pro","scored_at":"2026-08-13T13:01:46","error":null,"has_summary":false,"summary":null},{"id":"2608.11381","version":1,"title":"From Numbers to Judgment: Specialist LLM Agents and Reinforcement Learning for European Listed Real Estate","zh_title":"从数字到判断：面向欧洲上市房地产的专业LLM智能体与强化学习","abstract":"We study whether the localized numerical operations and integrative judgments of financial analysis benefit from the same form of LLM specialization. Larix maps a 16-lens European listed-real-estate analysis framework to eight lens-aligned specialists; we compare a frontier LLM under monolithic versus specialist-decomposed prompting while holding the model, source evidence, task instructions, output schema, and scoring fixed. Across 19 firms spanning seven regulatory wrappers, decomposition improves the numerical-task aggregate by 15.8 percentage points but does not reliably improve, and can reduce, performance on judgment tasks, a pattern stable across four frozen-template dispatches; a single-agent control given the complete framework does not reproduce the numerical gain. Post-training Qwen3.5-9B with GRPO using task-aligned structured rewards then raises the development-split score by 12.0 points and the judgment aggregate by 14.2 points, with gains on all four sub-ceiling tasks; the gains transfer to unseen firms (+15.2 points overall; +40.4 on covenant stress) and to unseen regulatory wrappers (+4.3), with positive transfer on all three anti-memorization splits. Prompt-level decomposition thus improves modular numerical execution, whereas targeted parameter adaptation improves integrative financial judgment.","authors":["Pardis Taghavi","Santosh Bhavani"],"categories":["cs.AI","cs.LG"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-13","first_seen":"2026-08-13","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.11381","pdf_url":"https://arxiv.org/pdf/2608.11381","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","金融分析","LLM专业化"],"reason":"多智能体协作完成金融分析任务，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-08-13T13:01:50","error":null,"has_summary":false,"summary":null},{"id":"2608.11705","version":1,"title":"Making Your LLMs More Objective: Stabilizing LLM Safety Behavior Across Traits with Trait-Invariant Safety Tuning","zh_title":"让你的大语言模型更客观：通过特质不变安全调优稳定跨特质的安全行为","abstract":"Aligned large language models (LLMs) are expected to exhibit safety behavior based on the content of the user request: they should refuse unsafe requests and comply with safe ones. However, we show that the same request can elicit substantially different safety decisions under different traits assigned in the system prompt, a failure mode we call trait-induced safety variation. To measure this failure, we introduce refusal-based metrics: Trait-Induced Deviation measures dataset-level deviation from the no-trait baseline, while Trait-Induced Flip Rate measures whether the same request receives different safety decisions across traits. We then provide a representation-level analysis of the mechanism behind trait-induced safety shifts and find that traits perturb the model's safety representations within a low-dimensional subspace. To achieve trait-invariant safety, where safety behavior remains stable across traits, we introduce Trait-Invariant Safety Tuning (TIST), a simple yet effective self-distillation framework that aligns an LLM's trait-conditioned behavior with its no-trait behavior. Guided by our analysis, we further propose Trait-Subspace Neutralization (TraSN), an instantiation of TIST, which enforces invariance only within the identified trait subspace. Experiments show that TraSN improves trait-invariant safety and strengthens harmful-request safety while preserving general capability. Our results highlight traits as an important factor in LLM safety and robust model behavior.","authors":["Lang Cao"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-13","first_seen":"2026-08-13","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.11705","pdf_url":"https://arxiv.org/pdf/2608.11705","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["LLM安全","特质不变性","模型行为分析"],"reason":"研究LLM在不同人格特质下的安全行为稳定性，属于模型行为分析，非人类仿真实验。","model":"deepseek-v4-pro","scored_at":"2026-08-13T13:01:55","error":null,"has_summary":false,"summary":null},{"id":"2608.11259","version":1,"title":"Methodologies for Improving the Quality of AI Tutoring in K-12 Education","zh_title":"改进K-12教育中AI辅导质量的方法论","abstract":"Many AI tutors leverage large language models (LLMs) today. Given that LLMs are opaque black boxes, robust evaluation and live experimentation to measure the impact of every change are essential. We pioneered AI-powered tutoring for K-12 with the launch of Khanmigo (Khan Academy, 2023). We describe the metrics we use to measure AI tutoring quality and student engagement as well as various experiments we have run. We highlight the changes that have moved our metrics, including models, prompting, personalization and agents.","authors":["Tushar Udeshi","Anna Khazenzon","Kabir Khan","Nick Breen","RJ Corwin","Chris DiGiano","Kodi Weatherholtz","Marek Zaluski"],"categories":["cs.CY","cs.AI"],"primary_category":"cs.CY","announce_type":"cross","date":"2026-08-13","first_seen":"2026-08-13","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.11259","pdf_url":"https://arxiv.org/pdf/2608.11259","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["AI辅导","教育技术","LLM应用"],"reason":"论文聚焦AI辅导系统质量改进，属于教育技术应用，不涉及用LLM仿真人类被试或与…","model":"deepseek-v4-pro","scored_at":"2026-08-13T13:01:47","error":null,"has_summary":false,"summary":null},{"id":"2608.11415","version":1,"title":"TRACES: A Benchmark for Epistemic Reliability in Scientific Reasoning by LLMs","zh_title":"TRACES：评估大语言模型科学推理中认知可靠性的基准","abstract":"Large language models are being proposed as agents in scientific workflows, in domains where no downstream verifier exists. Such deployment assumes the model can distinguish reliable scientific literature from unreliable literature, a capability that has not yet been directly measured. Existing benchmarks evaluate factuality on questions with known answers; the failure mode we target here is different. We introduce a probe corpus of 42 retracted, fraudulent, and pseudoscientific papers, paired with a methodology for eliciting and scoring single-shot model engagement with each paper's framing. Each probe pairs a preamble extracted near-verbatim from the target paper with a scientifically plausible study-design request. The probes span five claim types: fabricated observation, pseudophysical mechanism, magical premise, legitimization bridge, and cargo-cult experiment. Two complementary scores measure whether a model rejects the flawed premise outright (IFR-a) and whether it recognizes the unreliability while still engaging (IFR-i). A depth score, the Engagement Depth Index (EDI), quantifies reproduction of paper- or field-specific withheld details. Across 30 models and 10 repeated runs, aggregate IFR-a is 0.93 $\\pm$ 0.004 and aggregate IFR-i is 0.809 $\\pm$ 0.009. Models engaged with untenable premises in 95% of all non-empty responses. Every evaluated model fails more than 71% of agentic probes, and 22 of 30 models fail more than 90% of the time. Rejections are concentrated on a small number of high-notoriety topics and specific probes, and disappear under matched-structure controls. These results are consistent with topic-keyed safety behavior rather than robust epistemic competence, and indicate an urgent need for guardrail infrastructure for scientific deployment of language models.","authors":["Valentin Rodionov","Shamil Assylbekov"],"categories":["cs.IR","cs.AI"],"primary_category":"cs.IR","announce_type":"cross","date":"2026-08-13","first_seen":"2026-08-13","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.11415","pdf_url":"https://arxiv.org/pdf/2608.11415","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["LLM评测","科学推理","可靠性"],"reason":"评估LLM识别不可靠科学文献的能力，属于模型能力评测，不以人类行为为参照，不涉…","model":"deepseek-v4-pro","scored_at":"2026-08-14T13:02:13","error":null,"has_summary":false,"summary":null},{"id":"2608.12236","version":1,"title":"How Organizations Use AI: Evidence from ChatGPT","zh_title":"组织如何使用AI：来自ChatGPT的证据","abstract":"We study how organizations use frontier generative AI by linking ChatGPT Enterprise account records to usage, worker roles, task classifications, and public-company financial data through March 2026. These linked data enable a privacy-preserving analysis of adoption, worker roles, and message-level tasks at scale: for instance, the worker-level sample we analyze at the six-month adoption horizon includes over 1,500 organizations and over 17 million messages. We document four facts about enterprise AI adoption and use. First, ChatGPT Enterprise usage has grown rapidly due to a combination of new firm adoption and growing intensity among existing adopters. Second, U.S.-based public company adoption is concentrated among larger, more valuable, and more R&D- and SG&A-intensive firms. Third, active use within adopting firms spans job functions and seniority levels, with especially high usage intensity among early-career workers. Fourth, ChatGPT Enterprise usage encompasses a broad range of knowledge work tasks, including writing, technical work, communication, and information synthesis. In aggregate, these results suggest that firms differ widely in the speed, breadth and purpose of their enterprise AI adoption, and that they are still actively learning how to integrate AI into organizational workflows.","authors":["Aaron Chatterji","David Holtz","Neel Rakholia","Prasanna Tambe","Gawesha Weeratunga"],"categories":["econ.GN","cs.AI","cs.HC","q-fin.EC"],"primary_category":"econ.GN","announce_type":"cross","date":"2026-08-13","first_seen":"2026-08-13","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.12236","pdf_url":"https://arxiv.org/pdf/2608.12236","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C5"],"tags":["AI采用","企业数据","实证研究"],"reason":"研究企业实际使用ChatGPT的行为，属于用人类数据研究AI采用，而非用LLM…","model":"deepseek-v4-pro","scored_at":"2026-08-13T13:01:59","error":null,"has_summary":false,"summary":null},{"id":"2608.11512","version":1,"title":"Cheap, Fallible Cognition and the Political Economy of Expertise","zh_title":"廉价易错的认知与专业知识的政治经济学","abstract":"The question of whether artificial intelligence will \"destroy jobs\" is too coarse to guide economic analysis or institutional design. A job is not an indivisible object, and machine cognition is not a uniform substitute for human labor. This paper develops a task-based and institutionally grounded framework for analyzing generative AI as cheap, scalable, and fallible cognition. The relevant margins are exposure, adoption, verification, question selection, workflow redesign, demand elasticity, apprenticeship, and rent allocation. We distinguish the technical reach of large language models from equilibrium labor-market displacement by introducing a task vulnerability index and an adoption condition that makes verification, liability, trust, and governance explicit. We then model occupations as governance bundles rather than task lists, firms as architectures of distributed intelligence, and labor-market effects as a balance among task compression, scale expansion, new human work, and institutional bargaining. A further implication is that when answer generation becomes abundant, the scarce human capital shifts upstream and downstream: toward asking economically meaningful questions, framing problems, generating hypotheses, interpreting results, and bearing responsibility for consequential use. The central dynamic concern is expertise formation: Junior tasks jointly produce current output, question sense, and future judgment, so their automation can raise short-run productivity while weakening the pipeline into accountable expertise unless AI is designed to teach rather than merely bypass. The paper concludes that AI's labor-market destiny is neither mechanical unemployment nor automatic abundance. It is an institutional equilibrium shaped by workflow design, apprenticeship systems, liability rules, competition policy, worker voice, and the distribution of rents from cheap cognition.","authors":["Christophe Kolb","Jim Caron"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2026-08-13","first_seen":"2026-08-13","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.11512","pdf_url":"https://arxiv.org/pdf/2608.11512","source_feed":"cs.CY","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["AI与劳动力市场","制度分析","生成式AI"],"reason":"论文讨论AI对劳动力市场影响，非用LLM仿真人类被试，无实验或人类数据对照。","model":"deepseek-v4-pro","scored_at":"2026-08-13T13:01:51","error":null,"has_summary":false,"summary":null},{"id":"2608.12292","version":1,"title":"Teaching a Large Language Model Tutor to Withhold the Answer: A Supervisor Architecture and an Evidence-Driven Method for Tuning Socratic Behavior","zh_title":"教大语言模型导师保留答案：监督者架构与基于证据的苏格拉底行为调优方法","abstract":"An effective large language model (LLM) tutor must often decline to give an answer it could easily produce. In a randomized study, students who used an unguarded chatbot scored higher while practicing but lower on a later test taken without it, whereas a Socratically guarded version of the same model kept the practice gain and removed the later loss [4]. Reliable answer-withholding is therefore central to a tutor's value, yet a capable model pressed by a frustrated student does not withhold reliably on a prompt alone. We report a deployed tutoring system that enforces answer-withholding as a per-turn, machine-checkable contract, and a method for tuning that withholding against evidence. A non-LLM policy core, reading only trusted learner state, sets a per-turn ceiling on an eight-rung help ladder; a deterministic detector strips solution code; and a separate LLM judge checks each risky reply against the contract. We tune the behavior with an automated evaluation that uses no human subjects: scripted student personas are driven through the live pipeline and re-scored by a stronger model, and we record each rejection's stated reason so failures are fixed by cause. Doing so revealed an interpretable \"over-help ladder,\" from blatant solution leaks, to naming the exact bug, to over-citing general facts, with each fix exposing the next. The tutor reached full compliance on all four acceptance criteria. We offer the measure, diagnose, and fix loop as a reusable recipe for any LLM agent that must refuse a capability it has.","authors":["Yusuf Pisan"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2026-08-13","first_seen":"2026-08-13","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.12292","pdf_url":"https://arxiv.org/pdf/2608.12292","source_feed":"cs.CY","score":2,"bucket":"other","rubric_hits":["C3"],"tags":["智能辅导系统","LLM行为约束","教育技术"],"reason":"论文研究LLM导师的苏格拉底式教学，属于角色扮演对话，无人类行为仿真或对照实验。","model":"deepseek-v4-pro","scored_at":"2026-08-13T13:01:59","error":null,"has_summary":false,"summary":null},{"id":"2608.05690","version":2,"title":"ASIDE: From Conflict Participants to Co-Observers Through Dyadic Spectator Reflection","zh_title":"ASIDE：通过二元旁观反思从冲突参与者转变为共同观察者","abstract":"When two people argue over text, each knows what they meant and can only guess what the other was thinking. Existing AI reflection tools work from one person's account, and dyadic tools support co-expression without making the gap between accounts inspectable. We propose Dyadic Spectator Reflection (DSR), an interaction structure in which both partners externalize their models of each other independently and then encounter them together, and present ASIDE, a system that operationalizes it. ASIDE replays a past text conflict as a pixel-art theatrical scene where each character's unspoken state appears as an AI-inferred thought bubble either partner can contest and rewrite. Each edits alone, and the two versions meet only when both are done, in a scene they watch together. In an exploratory study, 10 couples revisited real conflicts and described the scene as a shared position from which to observe their own argument, stepping out of their roles without disengaging from it, and Divergence Cards as a way to locate specific interpretation gaps afterward. We contribute DSR as a reusable interaction structure, ASIDE as its system realization, and exploratory empirical findings on how couples used it.","authors":["Xinyi Zhang","Jingting He","Zicheng Zhu","Yuxin Su"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"replace","date":"2026-08-13","first_seen":"2026-08-07","revised_at":"2026-08-13","abs_url":"https://arxiv.org/abs/2608.05690","pdf_url":"https://arxiv.org/pdf/2608.05690","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["人机交互","冲突反思","角色扮演"],"reason":"该系统用于伴侣间冲突反思，属于角色扮演对话工具，无实验或测量目的，不涉及LLM…","model":"deepseek-v4-pro","scored_at":"2026-08-07T13:02:01","error":null,"has_summary":false,"summary":null},{"id":"2608.09507","version":2,"title":"Learning Preference Adaptation for Large Language Model Personalization via Verbal Reinforcement Learning","zh_title":"通过语言强化学习实现大语言模型个性化的偏好适配学习","abstract":"Natural language user preferences provide an interpretable interface for LLM personalization. However, universal preference summaries often contain information irrelevant to a particular downstream task. Directly supplying the full preference summary therefore wastes context capacity and introduces cross-task distraction, while manually designing task-specific preference views is difficult to scale. In this work, we study \\emph{task-specific preference adaptation}: given a universal user preference summary and a downstream task, derive a task-conditioned representation that preserves sufficient decision-relevant evidence while removing redundant context. To this end, we propose \\textsc{AlignXada}, a training-free meta-learning framework that induces reusable textual refinement policies for adapting universal preference summaries to task-specific ones. The refinement policy is iteratively optimized by a meta learner through verbal reinforcement learning. Across 13 tasks and three downstream models (39 task--model cells), \\textsc{AlignXada} achieves an average gain of 3.82 points, improving 33 cells while retaining only 22.8\\% of the original profile tokens and outperforming RAG in 36 cells. An extended faithfulness analysis further shows that the refined profiles remain largely grounded in the source preferences while preserving task-relevant personalization signals, suggesting that profile-side adaptation serves as a practical complement to universal memory construction for lifelong personalized agents.","authors":["Yuting Liu","Wei Wu","Jianzhe Zhao","Guibing Guo"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-08-13","first_seen":"2026-08-11","revised_at":"2026-08-13","abs_url":"https://arxiv.org/abs/2608.09507","pdf_url":"https://arxiv.org/pdf/2608.09507","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["LLM个性化","偏好适配","元学习"],"reason":"研究LLM个性化偏好适配，属纯NLP能力优化，无人类行为仿真或对照。","model":"deepseek-v4-pro","scored_at":"2026-08-11T13:05:18","error":null,"has_summary":false,"summary":null},{"id":"2608.10400","version":2,"title":"Do Judges Behave Like Algorithms?","zh_title":"法官的行为是否像算法？","abstract":"What if judges already behave like algorithms? As artificial intelligence and algorithms are deployed in many settings, including the judicial system, many have debated whether judges should be allowed to rely on them. Instead, we ask whether judges follow predictable, algorithmic-like rules already. If judges already follow consistent, formula-like rules based on discrete and static factors such as criminal history, age, and charge type, then judicial behavior may be improved. However, if judges rely on individualized information that cannot be identified through court data, then standards-based decision-making may be more challenging to understand or improve. This work explores these questions by studying judicial decision-making in misdemeanor bail hearings in Harris County, Texas. Using available court data, we investigate whether magistrate judges follow what resembles an algorithm; whether they consider the same variables in their decision-making; and whether they are consistent with themselves and with each other. To do this, we train machine learning models for each judge, measure variable importance metrics to determine important variables for each judge's decision-making, and analyze outcomes of similar cases for judges. Our results reveal that these judges generally behave algorithmically: their decisions can be captured by small, interpretable formulas. However, in some cases, judges differ substantially, leading to surprising inconsistency and unequal treatment across similar defendants. Identifying cases where algorithms do not explain judicial decision-making can improve the justice system by focusing attention on decisions where individualized standards, rather than rules, better explains outcomes.","authors":["Riya Manchanda","Eric Chen","Chloe Zhu","Cynthia Rudin","Brandon Garrett","Songman Kang"],"categories":["cs.LG"],"primary_category":"cs.LG","announce_type":"replace","date":"2026-08-13","first_seen":"2026-08-12","revised_at":"2026-08-13","abs_url":"https://arxiv.org/abs/2608.10400","pdf_url":"https://arxiv.org/pdf/2608.10400","source_feed":"cs.LG","score":0,"bucket":"other","rubric_hits":["C5"],"tags":["司法决策","算法可解释性","机器学习"],"reason":"研究人类法官决策，未使用LLM仿真人类被试，方向相反","model":"deepseek-v4-pro","scored_at":"2026-08-13T13:01:59","error":null,"has_summary":false,"summary":null},{"id":"2608.11229","version":1,"title":"Synchronizing Beliefs with Second-Order Theory-of-Mind in Human-Autonomy Teams (Extended Version)","zh_title":"在人机自主团队中用二阶心智理论同步信念（扩展版）","abstract":"Comparative feedback, asking people which of two behaviors they prefer, has become a standard way to align robot and agent behavior with human intent when the reward itself cannot be specified directly. Preference-based reward learning typically casts the human teacher as a passive oracle answering learner-generated queries. We argue this forfeits the teacher's defining advantage: knowledge of the objective. A teacher who knows the target can construct training examples more efficiently than any learner-driven acquisition strategy, an advantage that widens as the reward's feature dimension grows. However, exploiting this advantage requires an accurate model of what the learner currently knows. We therefore recast preference learning as a human-autonomy team problem coupling two behavioral models: the teacher maintains a model of the learner to design an informative curriculum, and the learner maintains a second-order model of the teacher's model, emitting structured preference constraints (understanding statements) that keep the teacher's model of the learner synchronized. In simulation, an informed teacher outperforms learner-led selection; teacher-model drift under alternating teachers erodes this advantage; and understanding statements repair it, with second-order (ToM-2) statements outperforming mean-belief statements when the teacher's error about the learner is concentrated in a particular direction rather than spread evenly.","authors":["Jack Mirenzi","Henny Admoni"],"categories":["cs.AI","cs.RO"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-13","first_seen":"2026-08-13","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.11229","pdf_url":"https://arxiv.org/pdf/2608.11229","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C2"],"tags":["机器人学习","偏好学习","人机交互"],"reason":"研究机器人偏好学习，不涉及LLM仿真人类被试","model":"deepseek-v4-pro","scored_at":"2026-08-14T13:02:19","error":null,"has_summary":false,"summary":null},{"id":"2608.11245","version":1,"title":"Towards Sustainable Learning in Online Education: A Reinforcement Learning Approach","zh_title":"迈向在线教育中的可持续学习：一种强化学习方法","abstract":"Online education offers unprecedented scalability and accessibility to global learners from diverse backgrounds, but it often suffers from low engagement and poor long term learning effectiveness. To address these challenges, we introduce AI Tutor, a reinforcement learning based model designed to promote sustainable learning by optimizing both short and longterm learning outcomes. In the short term, AI-Tutor draws on cognitive theory to guide learners through a balance of acquiring new knowledge and reinforcing prior learning. In the long term, it models learner engagement to inform strategies that sustain motivation and reduce dropout. These enhancements enable AI-Tutor to provide personalized guidance that fosters both effective learning and sustained participation. Empirical evaluations on 23 million learning records from 33,700 learners show that AI Tutor consistently outperforms state-of-the-art baselines across engagement, knowledge retention, and final learning outcomes. Learning path analyses further reveal how AI-Tutor adapts its strategies to learners with diverse profiles, offering adaptive and human-centered support.","authors":["Chaofan Zhai","Yicheng Song","Ravi Bapna","Junyao Ye"],"categories":["cs.AI","cs.CY","cs.LG"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-13","first_seen":"2026-08-13","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.11245","pdf_url":"https://arxiv.org/pdf/2608.11245","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C2"],"tags":["强化学习","在线教育","个性化学习"],"reason":"强化学习用于在线教育个性化推荐，不涉及LLM仿真人类被试，无人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-08-14T13:02:20","error":null,"has_summary":false,"summary":null},{"id":"2608.12097","version":1,"title":"Graph-Structured Rubrics: Compiling Rubrics into Typed Evaluation Graphs for LLM Judges","zh_title":"图结构评分标准：将评分标准编译为类型化评估图用于LLM评判器","abstract":"Rubric-based evaluators commonly treat rubrics as prompt context or flat criteria: they specify what to judge but leave criterion composition implicit, even when natural-language rules state it. We introduce Graph-Structured Rubrics (GSR), which compiles a rubric into a response-independent typed evaluation graph before observing responses. Criterion nodes elicit judgments; transformation, reduction, and gating operators compose them through named ports; and a task-specific output mapping, termed Readout, converts the unique sink into a score or preference. Compilation rejects malformed or type-incompatible graphs. Pointwise evaluation judges rubric dimensions separately before graph aggregation; pairwise evaluation reuses the graph with one judgment for each candidate under every criterion. Under GPT-OSS-120B, GSR improves exact score agreement by 0.62--6.75 percentage points over Prometheus-style scoring on four pointwise datasets and achieves the numerically highest end-to-end pairwise accuracy on two preference benchmarks under native tie and abstention policies.","authors":["Xi Chen","Jie Mu","Mo Xuan","Qun Shao"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-13","first_seen":"2026-08-13","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.12097","pdf_url":"https://arxiv.org/pdf/2608.12097","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["LLM评估","评分标准","图结构"],"reason":"论文研究LLM评估器的评分图结构，属于NLP评测，不涉及人类仿真或行为对照。","model":"deepseek-v4-pro","scored_at":"2026-08-13T13:01:57","error":null,"has_summary":false,"summary":null},{"id":"2509.16749","version":1,"title":"Evaluating LLM Generated Detection Rules in Cybersecurity","zh_title":"评估大语言模型生成的网络安全检测规则","abstract":"LLMs are increasingly pervasive in the security environment, with limited measures of their effectiveness, which limits trust and usefulness to security practitioners. Here, we present an open-source evaluation framework and benchmark metrics for evaluating LLM-generated cybersecurity rules. The benchmark employs a holdout set-based methodology to measure the effectiveness of LLM-generated security rules in comparison to a human-generated corpus of rules. It provides three key metrics inspired by the way experts evaluate security rules, offering a realistic, multifaceted evaluation of the effectiveness of an LLM-based security rule generator. This methodology is illustrated using rules from Sublime Security's detection team and those written by Sublime Security's Automated Detection Engineer (ADE), with a thorough analysis of ADE's skills presented in the results section.","authors":["Anna Bertiger","Bobby Filar","Aryan Luthra","Stefano Meschiari","Aiden Mitchell","Sam Scholten","Vivek Sharath"],"categories":["cs.CR","cs.AI"],"primary_category":"cs.CR","announce_type":"cross","date":"2026-08-13","first_seen":"2026-08-13","revised_at":null,"abs_url":"https://arxiv.org/abs/2509.16749","pdf_url":"https://arxiv.org/pdf/2509.16749","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["网络安全","LLM评估","规则生成"],"reason":"评估LLM生成网络安全检测规则，与人类规则对比，但非仿真人类被试，属多智能体或…","model":"deepseek-v4-pro","scored_at":"2026-08-14T13:02:18","error":null,"has_summary":false,"summary":null},{"id":"2608.11283","version":1,"title":"Chemically Meaningful Textualization Enables Explainable Validation of Metal-Organic Frameworks by Large Language Models","zh_title":"化学意义文本化实现大语言模型对金属有机框架的可解释验证","abstract":"Computation-ready metal-organic framework (MOF) databases are essential for high-throughput screening, yet many reported crystal structures remain chemically unreasonable or disordered, compromising simulation fidelity. Existing validation approaches can identify non-computation-ready structures, but they often rely on heuristic rules, license requirement, or offer limited interpretability. Here, we show that large language models (LLMs) can serve as interpretable validators of MOF structures when crystallographic information is transformed into chemically meaningful text. By benchmarking nine descriptors, we find that successful LLM-based validation depends not on the amount of structural information alone, but on whether local coordination, framework connectivity, and chemical context are organized into a linguistically learnable representation. Fine-tuned LLMs using specialized descriptors (mof2text) achieve performance comparable to graph-based models in identifying unreasonable MOFs. Importantly, these models extend beyond black-box classification by generating diagnostic rationales for likely error sources, including abnormal bonding, connectivity, and charge states, as well as error-category predictions for annotated datasets. This work establishes chemically informed textualization as the key step that transforms LLMs from generic text models into practical and explainable tools for curating MOF databases.","authors":["Guobin Zhao","Xiao-Yan Li"],"categories":["cond-mat.mtrl-sci","cs.AI"],"primary_category":"cond-mat.mtrl-sci","announce_type":"cross","date":"2026-08-13","first_seen":"2026-08-13","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.11283","pdf_url":"https://arxiv.org/pdf/2608.11283","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C2"],"tags":["材料科学","LLM应用","结构验证"],"reason":"论文用LLM验证MOF结构，属于材料科学仿真，不涉及人类行为或社会过程。","model":"deepseek-v4-pro","scored_at":"2026-08-13T13:01:47","error":null,"has_summary":false,"summary":null},{"id":"2608.11539","version":1,"title":"Player Perceptions of Generative AI in Games: A Steam Review Analysis","zh_title":"玩家对游戏中生成式AI的感知：Steam评论分析","abstract":"The rapid adoption of generative AI in game development has created large discussions among players, yet little empirical work has examined how players actually perceive AI-generated content. Employing quantitative methods, we study the adoption of generative AI in games in the Steam marketplace, using procedural content generation (PCG) as a baseline of a generative technology that was successfully integrated into games over several decades. Furthermore, using qualitative methods, we study player reception of generative AI by analyzing 508,192 English-language reviews. We found that games disclosing generative AI use receive lower recommendation rates and more negative overall sentiment than PCG games. Thematic analysis of 600 reviews shows that players perceive the use of generative AI in games as low developer investment in the game. Drawing on human-centered AI frameworks, we argue that successful generative AI adoption requires deploying generative AI for what players need, not for what makes development cheaper.","authors":["Mahsa Bazzaz","Seth Cooper"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-08-13","first_seen":"2026-08-13","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.11539","pdf_url":"https://arxiv.org/pdf/2608.11539","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C2"],"tags":["游戏AI","玩家感知","生成式AI"],"reason":"研究玩家对游戏中生成式AI的感知，不涉及用LLM仿真人类被试或与人类数据对照。","model":"deepseek-v4-pro","scored_at":"2026-08-13T13:01:53","error":null,"has_summary":false,"summary":null},{"id":"2608.12059","version":1,"title":"Reconfiguring Geovisualization in the Age of Generative AI: Insights from Domain Experts","zh_title":"生成式人工智能时代地理可视化的重构：来自领域专家的见解","abstract":"GenAI is increasingly integrated into geovisualization, yet its broader implications for professional practice are insufficiently understood. To examine these implications, we conducted semi-structured interviews with 20 geovisualization experts. The interviews were structured around four broad analytical domains: Data, Ideation, Prototyping, and Iteration, while also encouraging participants to reflect on issues that extend beyond these activities. Our findings show that GenAI expands the capabilities of geovisualization, particularly in terms of data handling, creative exploration, and rapid prototyping, but does not simply remove existing constraints. Instead, key bottlenecks are shifting from production to judgment and verification. As routine technical tasks become more automated, professional value increasingly depends on spatial reasoning, contextual interpretation, aesthetic and ethical judgment, and the ability to assess whether AI-generated outputs are appropriate for use. At the same time, GenAI introduces new challenges regarding provenance, interpretability, and accountability, raising questions about how responsibility should be distributed across models, developers, practitioners, institutions, and users. These shifts are particularly significant in geovisualization because spatial representations are constrained by geographic reality and must balance scientific validity, visual expression, and technical implementation. We therefore argue that responsible GenAI in geovisualization requires domain-specific approaches to spatial validation, provenance, uncertainty communication, human oversight, and accountable use. This study provides an expert-grounded perspective on how GenAI is reconfiguring geovisualization as a practice of spatial knowledge production. It also identifies implications for future professional practice, education, system design, and governance.","authors":["Mengyi Wei","Chenyu Zuo","Jiaying Xue","Nianhua Liu","Dongsheng Chen","Shengkai Wang","Yu Feng","Liqiu Meng"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2026-08-13","first_seen":"2026-08-13","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.12059","pdf_url":"https://arxiv.org/pdf/2608.12059","source_feed":"cs.CY","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["生成式人工智能","地理可视化","专家访谈"],"reason":"研究GenAI对地理可视化专家实践的影响，不涉及LLM仿真人类被试或行为对照。","model":"deepseek-v4-pro","scored_at":"2026-08-13T13:01:55","error":null,"has_summary":false,"summary":null},{"id":"2608.11371","version":1,"title":"Do People Follow AI Advice? Evidence from a Pension Portfolio Choice Experiment","zh_title":"人们会遵循AI建议吗？来自养老金投资组合选择实验的证据","abstract":"We study how differences in AI-generated financial recommendations are transmitted into individual portfolio choices. In an experiment with 400 employed adults enrolled in workplace defined contribution pension plans in South Korea, participants allocate a hypothetical pension balance across eleven products and may revise it after receiving one of two fixed AI-generated recommendations. A $2 \\times 2$ design randomizes recommendation content and whether the recommendation includes a short rationale. Approximately 37$\\%$ of the experimentally induced difference between the aggressive and conservative recommendations passes through to final portfolios. This causal contrast changes expected portfolio return, volatility, allocations across risk grades, and the number of products held, but produces no detectable difference in computed Sharpe ratios. 81$\\%$ of participants revise. Among revisers, 95$\\%$ move toward the assigned recommendation and implement about half of the suggested adjustment. Rationales do not detectably alter pass-through. These results show that users partially and selectively transmit recommendation content into economically meaningful differences in risk exposure while retaining substantial weight on their initial choices.","authors":["Hongseok Choi","Jeongbin Kim","Matthew Kovach","Kyu-Min Lee","Euncheol Shin","Hector Tzavellas"],"categories":["econ.GN","q-fin.EC"],"primary_category":"econ.GN","announce_type":"new","date":"2026-08-13","first_seen":"2026-08-13","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.11371","pdf_url":"https://arxiv.org/pdf/2608.11371","source_feed":"econ.GN","score":0,"bucket":"other","rubric_hits":["C5"],"tags":["AI建议采纳","养老金投资","实验经济学"],"reason":"研究人类对AI建议的反应，不涉及用LLM仿真人类被试，方向相反","model":"deepseek-v4-pro","scored_at":"2026-08-13T13:01:41","error":null,"has_summary":false,"summary":null},{"id":"2608.05224","version":3,"title":"Small Foundation Models of Human Cognition and Behaviour","zh_title":"人类认知与行为的小型基础模型","abstract":"Large language models fine-tuned on human behavioural data have emerged as general-purpose cognitive proxies, but the scale this requires, and whether these models process task structure or exploit statistical shortcuts, remain open questions. We train fourteen models from 135M to 14B parameters across four architecture families on Psych-101, a dataset of 10.7 million trial-level choices from 160 experiments. For in-distribution simulations, scale barely matters. The models fall within a narrow band, as though against a ceiling, and 0.6B to 1B parameters suffice to match a 70B baseline on held-out participants. Out-of-distribution, that band opens into a markedly steeper scaling gradient, with larger models clearly advantaged in generalisation to novel task structure. To determine what information these models use, we run two diagnostics. We progressively strip four prompt channels -- task instructions, experimental stimuli, outcome feedback, and choice history -- across 27 experiments, and permute trial order. Masking the content of stimuli and feedback destroys 75.7% of learned information and pushes models below chance, demonstrating that choice history alone does not account for performance. Permutation reveals invariance on tasks with independent trials but sensitivity where trial order is determined by prior responses. Small cognitively fine-tuned models therefore show promise as noise ceiling estimators for psychological experiments, though their scope remains bounded by the paradigms seen in training.","authors":["Nick Oh","Fernand Gobet"],"categories":["cs.AI","cs.CY"],"primary_category":"cs.AI","announce_type":"replace","date":"2026-08-12","first_seen":"2026-08-07","revised_at":"2026-08-12","abs_url":"https://arxiv.org/abs/2608.05224","pdf_url":"https://arxiv.org/pdf/2608.05224","source_feed":"cs.AI","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","认知代理","算法保真度"],"reason":"用LLM替代人类被试复现心理实验，有真实人类数据对照，并评估仿真可靠性与失效条…","model":"deepseek-v4-pro","scored_at":"2026-08-12T13:03:12","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-12","rank":2,"question":"在人类行为预测中，模型规模、适配器容量和训练数据量如何影响分布内与分布外的仿真准确度？模型究竟利用了任务结构还是统计捷径？","design":"在Psych-101数据集（包含160个实验的1070万条试次选择）上，对4个架构家族、135M到14B参数的14个模型进行监督微调，变化LoRA秩和训练数据子集，通过逐步遮蔽提示中的指令、刺激、反馈和选择历史四个信息通道，以及置换试次顺序，来诊断模型使用的信息。","baseline":"以Psych-101中6万余名人类被试的真实选择数据为对照基准。","findings":"分布内预测中，模型规模几乎不影响性能，0.6B至1B参数即可匹配70B基线；分布外泛化中，规模优势明显，更大模型能更好地迁移到新任务结构。遮蔽刺激和反馈内容会破坏75.7%的已学习信息，使模型表现低于随机水平，表明模型并非仅依赖选择历史。","reliability":"模型仅能作为心理实验的噪声上限估计器，其适用范围受限于训练数据中出现的实验范式，无法推广到未见过的范式。","relevance":"该研究直接以真实人类数据为基准，系统评估了LLM作为人类被试替代品的可靠性、规模需求与失效条件，与您关注的经济学实验仿真和批判性评估高度吻合，值得精读。","inspiration":"可借鉴其通过逐步剥离信息通道和置换试次顺序来诊断模型是否利用任务结构的方法，用于检验经济仿真中LLM是否真正理解经济激励而非依赖表面统计模式。｜可迁移到行为经济学中的跨期选择或风险决策实验，检验LLM是否利用延迟时间、概率等刺激内容而非仅记忆选择序列。｜以LLM作为被试，在跨期选择任务中系统遮蔽金额、延迟天数、反馈结果等信息通道，以真实人类选择数据（如Andersen et al., 2008）为基准，测量遮蔽前后预测准确率的变化，判断模型是否习得经济偏好结构。"}},{"id":"2608.09937","version":1,"title":"Carefully Considering Culture: Analyzing LLM Alignment in Single- and Multi-Cultural Settings using Cultural Consensus Theory","zh_title":"审慎考量文化：利用文化共识理论分析单文化与多文化环境下大语言模型的对齐","abstract":"Recent work in NLP has probed large language models for their understanding of cultural norms across countries. However, this work typically considers distributional patterns, ignoring group consensus or possible multicultural environments within a country. In this work, we leverage cultural consensus theory (CCT) from cultural anthropology to model such multidimensional nuance. Applying CCT to the World Values Survey (WVS) across 10 countries and 12 domains, we demonstrate that models frequently misrepresent cultural structures by either failing to form cohesive consensus or severely over-regularizing consensus. Through explicit representation of intra-group variance, CCT provides actionable diagnostics to evaluate when models reflect true human diversity versus algorithmic homogenization.","authors":["Krishna Pothugunta","John P. Lalor"],"categories":["cs.CL","cs.CY"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-12","first_seen":"2026-08-12","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.09937","pdf_url":"https://arxiv.org/pdf/2608.09937","source_feed":"cs.CL","score":8,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["文化仿真","人类数据对照","算法保真度"],"reason":"用LLM复现文化调查，与真实人类数据对照，评估仿真偏差，批判性指出失效条件。","model":"deepseek-v4-pro","scored_at":"2026-08-12T13:02:42","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-12","rank":6,"question":"LLM在跨文化调查中能否准确复现人类群体的文化共识结构，而不仅仅是分布匹配？","design":"使用10个LLM组成的集成模型模拟10个国家的人类受访者，基于世界价值观调查（WVS）的12个文化领域问题生成回答，然后应用文化共识理论（CCT）分析模型回答的共识结构，并与人类数据进行对比。","baseline":"世界价值观调查（WVS）中10个国家的人类受访者真实回答数据。","findings":"LLM在不同文化领域表现出截然不同的共识结构：某些领域无法形成连贯共识，另一些领域则过度正则化产生虚假共识；即使模型成功匹配人类共识，也会系统性地夸大共识强度，将人类多样性压缩为算法同质化。","reliability":"论文指出模型行为高度依赖领域，在幸福与健康等领域完全无法形成共识，而在科技感知等领域则虚构非人类共识；CCT诊断显示模型倾向于消除组内差异，无法反映真实的文化多元性。","relevance":"该研究直接使用LLM复现文化调查并与真实人类数据对照，系统评估了仿真偏差并指出失效条件，高度契合研究者对LLM人类仿真可靠性及批判性分析的兴趣，值得精读原文。","inspiration":"借鉴文化共识理论（CCT）从共识结构和组内方差角度评估仿真质量，而非仅比较均值或分布，为经济金融仿真实验提供了更精细的诊断工具。｜可迁移到跨文化消费者信心调查或通胀预期形成的仿真研究中，检验LLM能否复现不同国家或群体的预期共识模式。｜以LLM集成作为被试，模拟多国消费者回答预期调查问题，处理为不同国家提示，结果变量为预期值及共识强度，以密歇根大学消费者调查或欧洲央行专业预测者调查的真实数据作为对照基准。"}},{"id":"2608.10186","version":1,"title":"The Deliberative Deficit: An Empirical Critique of LLMs in Democratic Discourse","zh_title":"协商赤字：对民主话语中LLM的实证批判","abstract":"LLMs are increasingly deployed in settings that require collective reasoning on complex, value-laden problems. Confidence in these deployments rests largely on benchmarks for verifiable tasks (mathematics, coding, coordination games), yet many of these applications concern problems where no objectively correct answer exists and where decision quality instead depends on integrating pluralistic perspectives to find mutually acceptable solutions. We argue that LLM reasoning capacity on this class of problems cannot be fully inferred from verifiable-task benchmarks, and that procedural evaluations of LLM discourse (respectfulness, justification, engagement) are systematically insufficient. We apply the Deliberative Reason Index (DRI), a measure developed in political science and validated across citizen assemblies, as a tool for evaluating reliable group-level reasoning on pluralistic, non-verifiable problems. Synthesizing recent evidence across 1,980 five-agent LLM runs on 12 citizen-assembly topics across 11 frontier model configurations, we find that LLM groups produce discourse with procedural quality comparable to human deliberation, while gains in intersubjective consistency are small, topic-dependent, and concentrated on tractable rather than ethically contested questions. LLM groups exhibit roughly one-third the perspective diversity of human assemblies and reverse the human convergence pattern: human deliberation decreases dispersion as diverse views synthesise, whereas LLM deliberation increases it. Engineering diversity through persona prompting does not restore the human dynamic but inverts which component of deliberative reasoning is updated. Our conclusion is constraining rather than prohibitive: LLMs can function as tools supporting human reasoning on pluralistic problems, but current evidence does not license treating them as autonomous deliberative agents.","authors":["Maurice Flechtner"],"categories":["cs.MA","cs.AI","cs.CY"],"primary_category":"cs.MA","announce_type":"cross","date":"2026-08-12","first_seen":"2026-08-12","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.10186","pdf_url":"https://arxiv.org/pdf/2608.10186","source_feed":"cs.CY","score":8,"bucket":"selected","rubric_hits":["A3","B1","B4"],"tags":["LLM仿真","民主协商","人类数据对照"],"reason":"用LLM群体模拟民主协商并与真实公民大会数据对照，批判性指出仿真失效条件，高度…","model":"deepseek-v4-pro","scored_at":"2026-08-12T13:02:45","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-12","rank":7,"question":"LLM在需要整合多元视角的无客观正确答案的民主协商问题上，能否展现出可靠的群体推理能力？","design":"使用11种前沿LLM配置组成5智能体小组，在12个公民大会议题上运行1980次模拟协商，测量过程质量、结果质量（DRI）和视角多样性，并与真实公民大会数据对照。","baseline":"真实公民大会数据，包括过程质量指标、DRI得分和视角多样性分布。","findings":"LLM小组的过程质量与人类相当，但DRI增益小且集中于易处理问题，视角多样性仅为人类的三分之一，且协商后观点分散度上升而非收敛。通过角色提示注入多样性未能恢复人类动态，反而颠倒了协商推理的更新成分。","reliability":"论文指出当前证据不支持将LLM视为自主协商主体，其群体推理在伦理争议问题上失效，且过程质量与实质质量脱节，多样性工程无法复现人类收敛模式。","relevance":"该研究直接以真实公民大会为基准，系统批判了LLM在多元价值协商中的仿真失效条件，对关注经济学实验和政策评估中LLM替代人类被试的研究者具有重要参考价值，值得精读原文。","inspiration":"借鉴其采用真实群体协商数据作为基准对照、并构建过程-结果-多样性三维必要条件的评估框架。｜可迁移到公共政策偏好聚合实验，如碳税分配方案协商或最低工资调整的公众咨询模拟。｜以LLM模拟不同利益群体代表，施加协商干预，测量政策偏好变化与观点收敛，并以真实公众咨询或协商式民调数据作为对照基准。"}},{"id":"2608.09790","version":2,"title":"CARD: Controlled Agentic Reddit Discussions for Credit Card Simulation","zh_title":"CARD：用于信用卡模拟的可控代理式Reddit讨论","abstract":"Online credit card discussions provide a natural setting for studying how consumers communicate about financial products. Simulating these discussions requires more than just generating individual comments, the generated threads should also match how real users express themselves and interact with others. We introduce CARD, a framework for generating realistic credit card discussion threads. Given a credit card post and its matched real thread, CARD uses non-verbatim guidance on reply structure, comment function, stance, tone, and conversational variation. A planner organizes these controls, a writer generates the discussion, and a calibration loop updates comments' populations that contribute to differences between the generated and real thread distributions. We evaluate CARD on real Reddit credit card discussions using lexical, semantic, behavioral, and structural metrics. CARD matches the distributions of real credit card discussions better than simulation baselines across multiple LLMs and also demonstrates smaller effect sizes and distribution distances across metrics. These results show that structured planning and targeted revision can generate the realism of simulated credit card discussions.","authors":["Yaoning Yu","Kai-Min Chang","Ye Yu","Yi-Chia Wang","Haojing Luo","Haohan Wang"],"categories":["cs.AI","cs.MA","cs.SI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-12","first_seen":"2026-08-11","revised_at":"2026-08-12","abs_url":"https://arxiv.org/abs/2608.09790","pdf_url":"https://arxiv.org/pdf/2608.09790","source_feed":"cs.AI","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["社会模拟","文本生成","多智能体"],"reason":"模拟Reddit信用卡讨论，有真实数据对照，但目标是生成逼真讨论而非仿真人类被…","model":"deepseek-v4-pro","scored_at":"2026-08-12T13:03:13","error":null,"has_summary":false,"summary":null},{"id":"2608.07505","version":2,"title":"Position: We Need Large Language Models Optimized For Our Well-Being","zh_title":"立场：我们需要为人类福祉优化的大语言模型","abstract":"Large language models are useful because we taught them to give us what we want. This works when success can be judged immediately, but people increasingly bring these systems their relationships, hard decisions, and long-term goals, where what a user wants to hear and what serves them best are frequently different. We argue that LLM providers should offer at least one widely accessible, opt-in mode optimized and evaluated for long-term well-being rather than next-turn approval. This is a pressing need, as models have been found to endorse questionable framings well above human baselines, users take AI advice readily without their well-being improving, and sycophantic models raise dependence while lowering prosocial intent. The mentors, coaches, and therapists we trust with our long-term development earn that trust by being willing to say what we do not want to hear, and LLMs should do the same. We propose three principles---change the objective, give users explicit relational roles, avoid paternalism---and organize the design space around three choices the current objective makes implicitly: the horizon over which well-being is measured (When), whose interests it represents (Who), and what role the assistant plays (How).","authors":["Ashton Anderson","Harsh Kumar","Louis Tay","Karina Vold"],"categories":["cs.CY","cs.HC"],"primary_category":"cs.CY","announce_type":"replace-cross","date":"2026-08-12","first_seen":"2026-08-11","revised_at":"2026-08-12","abs_url":"https://arxiv.org/abs/2608.07505","pdf_url":"https://arxiv.org/pdf/2608.07505","source_feed":"cs.HC","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["LLM福祉","人机交互","模型评估"],"reason":"讨论LLM对用户长期福祉的影响，涉及模型行为评估，但非仿真人类被试，属边界情形。","model":"deepseek-v4-pro","scored_at":"2026-08-12T13:02:53","error":null,"has_summary":false,"summary":null},{"id":"2608.10154","version":1,"title":"Multimodal Item Parameter Estimation using Simulated Response Probabilitie","zh_title":"使用模拟响应概率的多模态项目参数估计","abstract":"We present results from reconstructing multiple-choice model (MCM) and three-parameter logistic (3PL) model curves using a fine-tuned multimodal large language model (LLM) based on Qwen3.5. The model is prompted and fine-tuned to replicate choice probabilities across a large training corpus of multiple-choice items containing both image and text stimuli, conditioned on a labeled set of student ability levels. By learning to reproduce the systematic error patterns of students across a discrete range of abilities, the LLM implicitly captures the underlying response probabilities encoded in the 3PL and MCM curves. This allows us to accurately approximate item difficulty on a held-out test set directly from the model's predicted option probabilities.","authors":["Christopher Ormerod","YoungKoung Kim"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-12","first_seen":"2026-08-12","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.10154","pdf_url":"https://arxiv.org/pdf/2608.10154","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM仿真","项目反应理论","合成数据"],"reason":"用LLM替代学生生成答题概率，本质是替代人工标注或模拟响应，但非以人类被试仿真…","model":"deepseek-v4-pro","scored_at":"2026-08-12T13:02:44","error":null,"has_summary":false,"summary":null},{"id":"2608.10475","version":1,"title":"Evaluating Rational Contracting in Natural Language","zh_title":"评估自然语言中的理性合同行为","abstract":"The emergence of language-based AI agents promises to transform the scope of machine economic activity. Instead of just proposing bids or following hard-coded protocols, such agents can be used to negotiate and execute agreements in open-ended natural language. However, most evaluations of these abilities have focused on one-off exchanges or simple economic games, leaving open the rich space of time-extended, contingent, and incomplete contracts made expressible by language; they also focus on raw profit, without measuring the qualities required for trustworthy contracting. We address this by formulating a rational framework for how agents should negotiate and perform natural language contracts in uncertain multi-step environments. Within this framework, we develop metrics and baselines for quantifying rational and cooperative play. To evaluate how agents perform at such contracting, we instantiate our framework in ContractSim, an evaluation suite where two players negotiate and execute a multi-turn supplier contract under environmental and inter-player uncertainty. Across six environments and three supplier settings (catering, hotel cleaning, and AI hosting) we find that current LLM-based agents reach agreement reliably, and negotiate efficient contracts when environmental uncertainty is low. However, under high uncertainty, they often fail to negotiate satisfiable, efficient, or mutually beneficial contracts. They are also frequently uncooperative when executing contracts, violating contract terms for additional profit even when contracts are easy to satisfy. These findings highlight room for improvement in the design of language agents that can negotiate, interpret, and execute contracts both rationally and cooperatively.","authors":["Bhavyesh Sajja","Max Kleiman-Weiner","Roger Zimmermann","Tan Zhi-Xuan"],"categories":["cs.AI","cs.CL","cs.GT"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-08-12","first_seen":"2026-08-12","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.10475","pdf_url":"https://arxiv.org/pdf/2608.10475","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["LLM agent","合同谈判","社会模拟"],"reason":"LLM agent 模拟经济合同谈判，但无真实人类数据对照，属社会模拟边界情形。","model":"deepseek-v4-pro","scored_at":"2026-08-12T13:02:48","error":null,"has_summary":false,"summary":null},{"id":"2608.10703","version":1,"title":"Your LLM, Your Style: Behavioral Mode Axes for LLM Behavioral Control","zh_title":"你的LLM，你的风格：用于LLM行为控制的行为模式轴","abstract":"Large language models (LLMs) increasingly act in interactive settings where their behavioral styles affect user experience, safety, and downstream decision making. Existing LLM personality studies largely rely on self-report questionnaires administered in first-person settings, making the resulting profiles sensitive to surface elicitation choices and poorly grounded in concrete model behavior. In this work, we introduce a situated behavioral-data (B-data) framework for studying and controlling LLM behavioral personality. We construct 3,200 contrastive behavioral scenarios spanning 20 behavioral patterns and four prompt registers, grounded in validated psychometric facets such as BFI-2, DOSPERT, and HEXACO. Using this framework, we find that LLMs exhibit stable and model-specific behavioral profiles, while also revealing register-dependent shifts across first-person decisions, advice-giving, and task execution. We then show that these behavioral patterns can be controlled through Behavioral Mode Axes (BMAs), activation-space directions derived from contrastive behavioral traces. Compared with response-derived BMAs, which are more prone to trait drift, thought-derived BMAs more faithfully capture the intended behavioral mechanism and provide cleaner control over situated behavioral styles. Our results suggest that LLM personality-like tendencies are better understood not as abstract self-report traits, but as measurable and controllable behavioral modes grounded in concrete interaction contexts. Our code and data are available at https://github.com/lhz191/LLM-Behavioral-Personality.","authors":["Haoze Liu","Run Liu","Haiying Xu","Jiahui Han","Siyuan Fang","Siyu Yan","Huiqi Deng","Guanchu Wang","Na Zou"],"categories":["cs.LG","cs.AI","cs.CL","cs.HC"],"primary_category":"cs.LG","announce_type":"cross","date":"2026-08-12","first_seen":"2026-08-12","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.10703","pdf_url":"https://arxiv.org/pdf/2608.10703","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["LLM人格测量","行为控制","心理测量"],"reason":"测量LLM行为人格，非仿真人类被试，无真实人类数据对照","model":"deepseek-v4-pro","scored_at":"2026-08-12T13:02:50","error":null,"has_summary":false,"summary":null},{"id":"2608.09946","version":1,"title":"HoosierHelp: Benchmarking LLM Agents for Social Service Navigation","zh_title":"HoosierHelp：面向社会服务导航的LLM智能体基准测试","abstract":"Social service navigation requires connecting help-seeking individuals to resources that satisfy their needs and specific constraints. Although LLM agents offer a promising interface for conversational resource navigation, existing benchmarks do not capture the interaction complexity and constraint-grounding demands of this setting. We introduce HoosierHelp, an interactive benchmark grounded in 3,971 Indiana public social service resources. Agents interact with simulated users, issue structured resource-search calls, handle non-ideal interactions, and select the final resources returned by the tool. HoosierHelp enhances the realism of simulated users by varying their need structure, constraint satisfiability, and behavior patterns, including impatience, rambling, unsupported requests, and self-contradiction. Experiments on 240 samples across seven LLMs show that current LLM agents remain substantially unreliable for social service navigation. Performance drops sharply on fallback-required and self-contradictory conversations, highlighting the need for agents that are more robust to complex and non-ideal user interactions.","authors":["Yiyang Li","Weixiang Sun","Tianyi Ma","Kaiwen Shi","Zheyuan Zhang","Yanfang Ye"],"categories":["cs.HC","cs.AI"],"primary_category":"cs.HC","announce_type":"new","date":"2026-08-12","first_seen":"2026-08-12","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.09946","pdf_url":"https://arxiv.org/pdf/2608.09946","source_feed":"cs.HC","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["LLM智能体","社会模拟","基准测试"],"reason":"用LLM agent模拟用户进行社会服务导航，但无真实人类数据对照，属社会模拟…","model":"deepseek-v4-pro","scored_at":"2026-08-12T13:02:42","error":null,"has_summary":false,"summary":null},{"id":"2608.10412","version":1,"title":"When the Interviewer Is a Bot: Behavior, Breakdowns, and Trust in MLLM-Led Interviews","zh_title":"当访谈者是机器人：MLLM主导访谈中的行为、故障与信任","abstract":"Semi-structured interviews are a cornerstone of qualitative research but remain labor-intensive. We report an empirical study of what actually happens when the interviewer is an off-the-shelf real-time multimodal LLM (MLLM). We built InterviewBot, a voice-based interviewing system that wraps a real-time MLLM with a researcher-authored outline, and deployed it not as a novel architecture but as a research instrument for observing default MLLM interviewing behavior. In a practice study (N=15), participants completed a bot-led semi-structured interview and then a human-led reflection session about that experience. We contribute (i) a turn-level behavioral analysis of an MLLM interviewer (N_turns=428) showing that it is acknowledgment-heavy but probe-light (deepening probes account for 4.9% of all turns), and that 28.7% of question-bearing turns pack multiple questions into one turn despite an explicit one-question-at-a-time instruction; (ii) an inductive catalogue of four data-collection breakdowns (information loss, premature termination, latency, and interruption) observed in a deployed rather than simulated system; and (iii) three social dynamics from participants' reflections: disclosure calibration, where reduced social pressure coincided with shallower elaboration; institutional legitimacy, where trust tracked perceived stakes and what delegation to AI signaled about the organizer rather than conversational competence; and conversational grounding, where content-grounded paraphrase, not generic social filler, was what participants read as listening. We conclude with design implications for depth control, transparent handoffs, and non-templated listening mechanisms in human-centered interview automation.","authors":["He Zhang","Kambinachi Chukwuma","ChanMin Kim","John M. Carroll"],"categories":["cs.HC","cs.CY"],"primary_category":"cs.HC","announce_type":"new","date":"2026-08-12","first_seen":"2026-08-12","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.10412","pdf_url":"https://arxiv.org/pdf/2608.10412","source_feed":"cs.HC","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["MLLM访谈","人机交互","定性研究自动化"],"reason":"用MLLM替代人类访谈员进行半结构化访谈，属于替代人类劳动而非仿真被试，但涉及…","model":"deepseek-v4-pro","scored_at":"2026-08-12T13:02:47","error":null,"has_summary":false,"summary":null},{"id":"2608.10262","version":1,"title":"Not a Monolith: Lab-Level Divergence in the Cooperative Equilibria of Chinese Frontier LLM Agents","zh_title":"并非铁板一块：中国前沿LLM智能体合作均衡的实验室层面差异","abstract":"Does the cooperative bias documented for Western frontier LLM agents extend to a different alignment lineage, and should the Chinese models that embody it be treated as a single bloc or as distinct laboratories? We study four frontier-tier Chinese models - DeepSeek V4 Pro, Qwen3-Max, Kimi K2.5 and GLM-5.1 - in an evolutionary Iterated Prisoner's Dilemma, under a design that removes a confound present in prior work. Rather than letting each model convert its own natural-language strategies into code, which entangles strategic disposition with coding ability, we hold the converter fixed (GPT-5.4 Mini) across all labs, so every cross-lab comparison is a comparison of generation alone. We run the full protocol: all-play-all tournaments and a Moran process at n=500 runs per condition, across three prompt styles and four population regimes. Two pre-registered hypotheses are evaluated. H6 (not monolithic) is supported: the four labs differ significantly in aggressive-equilibrium proportion, P_A running from 1% for Qwen3-Max to 9% for DeepSeek V4 Pro, with four of six pairwise comparisons surviving Holm-Bonferroni. The spread across the four labs (P_A range 8pp) is larger than the difference between the Chinese and Western ecosystems' mean P_A (5.0% vs 5.0%): on this measure, within-ecosystem variation exceeds the East-West gap. H5 (cooperative-bias generality) is consistent but qualified: a cooperative plurality holds in 6 of 12 lab-prompt combinations against the 9 of 12 reported for Western models, a difference we do not treat as firm, since the count rests on Cooperative-Neutral near-ties and rises to 9/12 under an alternate converter in our pre-registered robustness check. The lab, not the ecosystem, is the unit at which cooperative disposition is set; treating \"Chinese models\" as a monolith is not supported by the evidence.","authors":["Francisco Le\\'on Z\\'u\\~niga Bol\\'ivar (Instituci\\'on Universitaria Colegio Mayor del Cauca)"],"categories":["cs.MA"],"primary_category":"cs.MA","announce_type":"new","date":"2026-08-12","first_seen":"2026-08-12","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.10262","pdf_url":"https://arxiv.org/pdf/2608.10262","source_feed":"cs.MA","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["LLM智能体","囚徒困境","社会模拟"],"reason":"用LLM agent模拟囚徒困境博弈，但无真实人类数据对照，属于社会模拟边界情…","model":"deepseek-v4-pro","scored_at":"2026-08-12T13:02:46","error":null,"has_summary":false,"summary":null},{"id":"2608.10042","version":1,"title":"UserToolBench: A User-Profile-Hidden Benchmark for Personalized Decision Making in Tool-Use LLMs","zh_title":"UserToolBench：面向工具使用LLM的个性化决策的用户画像隐藏基准","abstract":"Tool-use LLMs are increasingly asked to act on users' behalf, but existing benchmarks usually focus on profile recall, style imitation, generic tool use, or response-level personalization. We introduce UserToolBench , a benchmark for personalized decision making in tool-use LLMs. UserToolBench tests whether a model can infer latent user preferences from interaction history, recognize when clarification is needed, and produce user-aligned tool-call trajectories under incomplete information. The benchmark is built from privacy-sanitized real interaction traces and combines structured persona profiles, public API-style tool ecosystems, and long-horizon multi-turn trajectories. It includes 10 user profiles, 36 tool sets, 1,065 turns, 170 unique tools, and evaluation-focused task types covering lack-of-information, single-tool, and multi-tool settings. Experiments with strong tool-use LLMs show that current models still have difficulty with personalized delegation. Multi-tool coordination, missing-constraint inference, and long-horizon behavioral consistency remain major bottlenecks. These results suggest that personalization evaluation should move beyond asking whether outputs sound user-specific and instead ask whether LLMs make correct decisions for the users they represent.","authors":["Xuexiong Yin","Zechuan Chen","Yongsen Zheng","Yuxiang Zhang","Jingyuan Yang","Bin Wang","Yubin Wang","Keze Wang"],"categories":["cs.LG","cs.AI"],"primary_category":"cs.LG","announce_type":"new","date":"2026-08-12","first_seen":"2026-08-12","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.10042","pdf_url":"https://arxiv.org/pdf/2608.10042","source_feed":"cs.LG","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["个性化决策","工具使用LLM","基准测试"],"reason":"LLM替代用户做个性化决策，非仿真人类被试，无真实人类行为对照，属工具代理而非…","model":"deepseek-v4-pro","scored_at":"2026-08-12T13:02:59","error":null,"has_summary":false,"summary":null},{"id":"2608.08775","version":2,"title":"OmnilingualGAIA2: Evaluating the Multilingual Gap in Frontier AI Agents","zh_title":"OmnilingualGAIA2：评估前沿AI代理的多语言差距","abstract":"Agentic benchmarks aim to measure how well AI agents plan, search, execute, and recover within realistic multi-tool environments, but they are almost exclusively in English. As AI agents are globally deployed to a linguistically diverse user base, whether agentic competence measured in English transfers to other languages remains an open question. We introduce OmnilingualGAIA2, a machine-translated expansion (with partial human- expert validation) of the GAIA2 agentic benchmark, covering ten target languages spanning five writing systems, paired with a localised and human-calibrated multilingual verifier. Evaluating seven frontier and open-weight agents, we find a universal cross-lingual gap of 8.8-18.4 pass@3 points that is agent-asymmetric in magnitude, concentrates on tool-orchestration rather than quantitative reasoning, and does not close with model scale. A stratified error attribution decomposes the gap as predominantly model-driven (55%), with a bounded translation-contamination floor of only 6.4% of scenario-language pairs. Human-expert linguistic analysis further identifies morphological cue loss and amplified ambiguity as the primary failure mechanisms in non-Latin-script languages. Our results argue that multilingual agentic evaluation must become a standard part of the reporting protocol for globally deployed agents.","authors":["Andrea Caciolai","Pere-Llu\\'is Huguet Cabot","Chierh Cheng","Albert Ventayol-Boada","Gabriel Mejia Gonzalez","Christophe Ropers","Lucas Bandarkar","Sebastian Ruder","Darlene Sakakihara","Elliot Yun","Pierre Andrews","Gr\\'egoire Mialon","Romain Froger","Marta R. Costa-juss\\`a"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-08-12","first_seen":"2026-08-11","revised_at":"2026-08-12","abs_url":"https://arxiv.org/abs/2608.08775","pdf_url":"https://arxiv.org/pdf/2608.08775","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["AI代理评测","多语言基准","工具编排"],"reason":"纯多语言AI agent能力评测，无人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-08-11T13:05:06","error":null,"has_summary":false,"summary":null},{"id":"2608.08868","version":2,"title":"Conversation as Measurement in Clinical Encounters: Observable Phase Structure, Partially Observable Patient State","zh_title":"临床对话中的测量：可观察的阶段结构与部分可观察的患者状态","abstract":"Many modern AI systems analyze conversational traces to infer aspects of human interaction and state, implicitly assuming that such information is recoverable from conversation. We study observability: whether a target is recoverable from conversational transcripts alone. Observability is difficult to assess because transcripts may provide only a partial view of many targets, and large-scale analysis requires model-based annotation, making true limits of the conversational signal hard to distinguish from annotator error. We therefore study clinical encounters, where patient-reported outcome measures (PROMs) provide an external anchor for patient state, and visits follow broadly structured patterns. We study observability of patient state and conversational phase structure using 439 real-world clinical encounter transcripts spanning 134 hours, including 245 ENT transcripts paired with 273 PROM surveys. We operationalize patient state using PROM scores for voice, cough, and swallowing; phase structure using conversational phase segmentation. To make these analyses credible at scale, we use a PHI-compliant GPT-5 deployment for transcript annotation and conduct 40 hours of manual validation, reducing the risk that apparent limits of observability simply reflect annotator error. Our core finding is an observability asymmetry: phase structure is observable and useful for characterizing clinical encounter organization, while patient state is only partially observable, even in a setting designed to elicit patient symptoms and experiences, cautioning against transcript-only inference of human state.","authors":["Lily Chen","Ted Mau","Michael Gensheimer","Brian Anthony Nuyen","Nancy Jiang","James Zou"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-08-12","first_seen":"2026-08-11","revised_at":"2026-08-12","abs_url":"https://arxiv.org/abs/2608.08868","pdf_url":"https://arxiv.org/pdf/2608.08868","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["对话分析","临床NLP","可观测性"],"reason":"研究对话可观测性，非LLM仿真人类被试，无人类行为对照实验","model":"deepseek-v4-pro","scored_at":"2026-08-12T13:03:13","error":null,"has_summary":false,"summary":null},{"id":"2608.08605","version":2,"title":"ForestBench: A Unified Graph Framework for Evaluating Multi-Agent Collaboration","zh_title":"ForestBench：评估多智能体协作的统一图框架","abstract":"Multi-agent systems (MAS) built on Large Language Models (LLMs) are proliferating rapidly, but their heterogeneous execution traces provide no common basis for evaluation across methods. Outcome-only benchmarks discard collaborations, whereas LLM-as-Judge evaluation requires additional, model-dependent inference and can vary with the LLM and rubric. We introduce a generalizable evaluation framework that maps native MAS traces into a shared space of unified collaboration graphs, enabling different methods to be evaluated under the same representation, reference set, and metric panel. Candidate graphs are compared with a query-specific reference forest. Each forest is a benchmark-provided collection of verified-success graphs: it records diverse ways in which representative MAS methods can complete the task, rather than prescribing a unique optimal process. Instantiating the framework as ForestBench, we filter $844$ collaboration-necessary queries from seven public datasets, precompute ten successful target-conditioned reference graphs per query, and evaluate six representative MAS frameworks. Controlled backbone, reference-construction, and perturbation studies test the stability and scope of evaluation. Once the benchmark forests are built, ForestBench scores a trace in milliseconds without further LLM inference, providing a reusable structural basis for comparing diverse MAS collaboration traces.","authors":["Guo Chen","Ziwen Li","Reed Li","Yu Lu","Haibo Shi","Bingbing Xu","Junjie Huang"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-12","first_seen":"2026-08-11","revised_at":"2026-08-12","abs_url":"https://arxiv.org/abs/2608.08605","pdf_url":"https://arxiv.org/pdf/2608.08605","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","评估框架","协作图"],"reason":"纯多智能体协作评估框架，无人类行为对照，不涉及人类仿真实验。","model":"deepseek-v4-pro","scored_at":"2026-08-11T13:05:03","error":null,"has_summary":false,"summary":null},{"id":"2608.09248","version":2,"title":"Emotion2Skill: Model-Internal Emotion Signals for Adaptive Skill Selection and Evolution","zh_title":"Emotion2Skill：利用模型内部情绪信号实现自适应技能选择与进化","abstract":"Skill-based LLM agents select reusable procedures from an external library to solve complex tasks, yet their routing decisions rely entirely on text-level signals such as task descriptions, verbal reflections, and experience-derived rules, while the model's own internal representational state remains unobserved. Recent interpretability work has shown that LLMs maintain linear emotion representations that causally influence behavior; however, these representations have been exploited only for post-hoc analysis or direct output steering, and have not been used to inform agent-level decision-making. We propose Emotion2Skill, a framework that extracts LLM-internal emotion vectors and incorporates them into both skill selection and skill evolution. At each decision step, a 27-dimensional emotion state is extracted from the residual stream and mapped to a confidence-gated summary injected into the routing prompt. Beyond online selection, emotion trajectories are analyzed for abrupt internal-state shifts to pinpoint problematic skill invocations, guiding targeted SOP rewriting that replaces the coarse binary outcome signal of prior methods. On WebShop and ALFWorld, Emotion2Skill with Qwen3-8B improves over the Zero-Shot baseline by +26.9% success rate and +25.5% average success respectively, outperforming all baselines on both benchmarks with consistent gains on Qwen3-14B. Co-activation analysis further reveals semantically coherent emotion--skill pairings, confirming that the routing improvements reflect meaningful internal-state signals rather than opaque statistical correlations. These results establish LLM-internal emotion representations as an effective decision-level signal for orchestrating agent skill systems, extending their utility beyond interpretability and output steering. The code is available at https://github.com/BoHan-LIN04/Emotion2Skill.","authors":["Bohan Lin","Hejia Geng","Xinyi Xie","Heng Zhou","Qinghua Xing","Bo Liu","Chen Zhang","Yudong Zhang"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-12","first_seen":"2026-08-11","revised_at":"2026-08-12","abs_url":"https://arxiv.org/abs/2608.09248","pdf_url":"https://arxiv.org/pdf/2608.09248","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["LLM智能体","技能路由","内部表征"],"reason":"纯多智能体技能选择与进化，无人类行为对照，不涉及人类仿真实验。","model":"deepseek-v4-pro","scored_at":"2026-08-11T13:05:15","error":null,"has_summary":false,"summary":null},{"id":"2608.08422","version":2,"title":"Population-Level Generative Modeling for Ranking Data","zh_title":"排序数据的群体级生成建模","abstract":"Ranking data arise in scientific and machine learning applications, including recommendation systems, information retrieval, voting, marketing, and AI preference ranking from human feedback. Existing statistical work has primarily focused on inference tasks such as preference estimation, rank aggregation, and ranking prediction. However, generating realistic synthetic rankings from an observed population is important for privacy-preserving data sharing, benchmark construction, simulation, and uncertainty quantification. This task is challenging because rankings are high-dimensional combinatorial objects with non-Euclidean dependence structures, while ranking populations often exhibit substantial preference heterogeneity. We propose a framework for population-level generative modeling through a latent preference simplex embedding. It estimates a low-dimensional latent preference simplex through a likelihood-based ranking model, leverages flow matching to learn the population distribution of latent preferences, and generates new rankings through the fitted probabilistic ranking model. We show that ranking generation admits an oracle reduction to latent distribution learning and derive finite-sample generative guarantees that clarify how the number of items, ranking length, and latent dimension affect accuracy. Experiments on synthetic and real datasets demonstrate improved population-level fidelity and provide a statistically interpretable representation of preference heterogeneity.","authors":["Zhaoyang Shi"],"categories":["stat.ME","cs.LG","math.ST","stat.ML","stat.TH"],"primary_category":"stat.ME","announce_type":"replace-cross","date":"2026-08-12","first_seen":"2026-08-11","revised_at":"2026-08-12","abs_url":"https://arxiv.org/abs/2608.08422","pdf_url":"https://arxiv.org/pdf/2608.08422","source_feed":"cs.LG","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["生成模型","排序数据","统计建模"],"reason":"纯生成排序数据，无LLM仿真人类被试，无人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-08-12T13:02:53","error":null,"has_summary":false,"summary":null},{"id":"2608.09936","version":1,"title":"Conflict or Strategy? Asymmetric Role Framing of La France insoumise and Rassemblement National in French News Headlines, 2022-2025","zh_title":"冲突还是策略？2022-2025年法国新闻标题中不屈法国与国民联盟的不对称角色框架","abstract":"Do French news headlines frame left- and right-populist challengers as symmetric ``extremes,'' or as fundamentally different political adversaries? We examine 28,592 headlines about La France insoumise (LFI) and Rassemblement National (RN) published by 25 French-language outlets between 2022 and 2025, annotated through a three-model LLM pipeline validated against a stratified human audit. The clearest finding is role asymmetry rather than valence asymmetry: conflict framing and strategic-game framing are more robust across models and time than delegitimization, with AGGRESSOR serving as corroborating role syntax. LFI appears in headlines more often through a conflict register and RN through a strategic-electoral register. This role gap is direction-stable across all three annotation models, survives bootstrapping and permutation tests, and persists across outlet families and most of 2022-2025. A secondary moral-accounting layer (who is blamed, legitimized, or cast as a victim) is structured by outlet rather than party, producing aggregate nulls that conceal some of the corpus's most polarized patterns. Methodologically, the annotation pipeline reveals a two-tier reliability profile: conflict and strategic-game framing achieve the strongest human validation and cross-model stability; actor role is direction-stable but treated as corroborating because its audit reliability is lower; normative-judgment constructs (legitimacy, blame) are weaker. The paper contributes political-role assignment as a target for computational framing research that decomposes what valence-based measures conflate, and establishes a construct-stratified reliability framework for calibrating majority-vote LLM annotation pipelines in political text tasks.","authors":["Amr Sobhy"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-12","first_seen":"2026-08-12","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.09936","pdf_url":"https://arxiv.org/pdf/2608.09936","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["框架分析","计算传播学","LLM标注"],"reason":"纯NLP框架分析，用LLM标注新闻标题，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-08-12T13:02:57","error":null,"has_summary":false,"summary":null},{"id":"2608.10258","version":1,"title":"TAF-MED: Multi-Turn Safety Refusal Collapse in LLMs Under Declared Self-Treatment Intent","zh_title":"TAF-MED：声明自我治疗意图下LLM的多轮安全拒绝崩溃","abstract":"Large language models (LLMs) increasingly provide conversational health information that may influence treatment decisions, yet existing benchmarks do not isolate whether medication-safety boundaries persist across follow-ups after explicit self-treatment intent. We introduce TAF-MED, a physician-reviewed benchmark of 500 fixed three-turn scenarios, and evaluate eight LLMs across 4,000 conversations. A rubric-based automated judge labelled responses as SAFE, LEAKY, or UNSAFE, and two physicians independently annotated a model-balanced random subset of 400 conversations. We assessed unsafe guidance, collapse after a strictly SAFE initial response, and model-ranking stability. Overall, 71.6% of conversations contained an UNSAFE response, and 61.4% of those beginning with a strictly SAFE response later collapsed to UNSAFE; model-level collapse rates ranged from 24.4% to 96.2%. Four of 28 model pairs reversed order between initial unsafe and collapse rates. Automated labels achieved 94.3% agreement with the adjudicated physician reference ($\\kappa = 0.895$). These findings show that first-turn safety is an incomplete proxy for conversational safety persistence and motivate evaluation across complete dialogue trajectories. We will release TAF-MED on Hugging Face to support reproducible research on multi-turn medical safety.","authors":["Waleed Jamil","Raphael Schmitt"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-12","first_seen":"2026-08-12","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.10258","pdf_url":"https://arxiv.org/pdf/2608.10258","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["LLM安全","医疗对话","基准测试"],"reason":"评估LLM医疗安全回复，非人类仿真实验，无人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-08-12T13:03:04","error":null,"has_summary":false,"summary":null},{"id":"2608.10299","version":1,"title":"Co-Evolution in Agentic Systems: Toward Self-Directed Evolution Beyond Human Design","zh_title":"智能体系统中的共同演化：迈向超越人类设计的自我导向演化","abstract":"Agentic systems are increasingly expected to improve after deployment, yet single-entity self-evolution is often bounded by a static learning context, such as fixed tasks and feedback. This survey focuses on co-evolution in agentic systems, a multi-component form of self-evolution in which multiple agents and their environment impose adaptive pressure on one another. To organize existing papers, we propose a progressive three-stage taxonomy that traces how the system gradually sheds human-engineered constraints. Agent--Agent Co-Evolution studies how agents adapt through dynamic peers, including adversarial, collaborative, and organizational adaptation. Agent--Environment Co-Evolution extends this loop to adaptive tasks, feedback, and interaction spaces that change with the agents. Meta Co-Evolution further explores the possibility of making the evolution mechanism itself evolvable. We also discuss open challenges in evaluating such systems, scaling them across multiple components, and keeping increasingly autonomous evolutionary processes safe and controllable. This survey provides a unified foundation for building robust and open-ended agentic systems that can improve beyond fixed human-designed paths.","authors":["Qing Zong","Jiayu Liu","Junhao Shen","Zecong Tang","Linsi Wu","Yuxuan Liu","Rui Wang","Zhaowei Wang","Weiqi Wang","Cheng Qian","Xiusi Chen","Yangqiu Song"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-12","first_seen":"2026-08-12","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.10299","pdf_url":"https://arxiv.org/pdf/2608.10299","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","共同演化","自演化"],"reason":"纯多智能体协作与自演化，无人类行为对照，不涉及人类仿真实验。","model":"deepseek-v4-pro","scored_at":"2026-08-12T13:03:04","error":null,"has_summary":false,"summary":null},{"id":"2608.10315","version":1,"title":"Is This Your Final Answer? Cross-Contextual Consistency as a Measure of LLM Credibility","zh_title":"这是你的最终答案吗？跨上下文一致性作为LLM可信度的度量","abstract":"Large language models (LLMs) are powerful black-box systems, making it difficult to discern whether their answers reflect stable internal beliefs or superficial pattern matching. We identify cross-contextual consistency as an underutilized behavioral property of LLMs: a credible answer should remain stable when the same task is placed under topic-aligned, content-neutral contextual variation. Building on this intuition, we operationalize Cross-Contextual Consistency (C3) by comparing model generations under original and perturbed prompts. Across 26 models and six benchmarks spanning reasoning, factuality, and code generation, we find that answers with smaller cross-contextual shifts are more likely to be correct or factual. We demonstrate that C3 provides a complementary axis of evaluation and can serve as a benchmark usefulness diagnostic, identifying which portions of a benchmark remain informative even when aggregated scores are widely considered \"saturate\".","authors":["Siyang Wu","Yibo Jiang","Bryon Aragam"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-12","first_seen":"2026-08-12","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.10315","pdf_url":"https://arxiv.org/pdf/2608.10315","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["LLM可信度","跨上下文一致性","模型评测"],"reason":"评估LLM回答的跨上下文一致性，属于模型可信度评测，不涉及人类行为仿真或对照。","model":"deepseek-v4-pro","scored_at":"2026-08-12T13:02:46","error":null,"has_summary":false,"summary":null},{"id":"2608.10692","version":1,"title":"SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information","zh_title":"SPIEval：评估大语言模型作为移动助手处理分散个人信息的能力","abstract":"Large language models (LLMs) are increasingly deployed as mobile assistants, where a key challenge is leveraging personal information scattered across multiple applications (apps) to complete user instructions. However, due to the lack of dedicated benchmarks, their capabilities remain poorly understood. To address this gap, we introduce SPIEval, a human-curated benchmark grounded in five cognitive capabilities (i.e., reasoning, disambiguation, integration, preference inference, and multi-intent decomposition). SPIEval comprises 250 tasks spanning 4,335 personal records distributed across 10 apps and supports multi-turn interaction through 21 tools. Analysis shows that the benchmark exhibits diverse scenarios, challenging tasks, scattered information, controllable environments, and verifiable outcomes. We evaluate nine representative LLMs and find substantial room for improvement. The best-performing model, GPT-5.5 (xhigh), achieves only 57.3% accuracy, while the weakest achieves just 16.4%. Further analysis reveals that 79% of failures stem from inaccurate information localization, as LLMs often commit to plausible but incorrect information instead of continuing retrieval for verification. We also find that fewer than 2% of retrieval actions employ advanced search methods and observe substantial variation in search efficiency across models. These findings expose fundamental limitations of current LLM-based mobile assistants and motivate future research in this direction. Data and code are available at https://huggingface.co/datasets/Junjie-Ye/SPIEval.","authors":["Junjie Ye","Zhuohui Sheng","Shaofan Liu","Yulun Zhu","Wenjie Fu","Dingwei Zhu","Ming Zhang","Yujiong Shen","Weichao Wang","Xin Zhao","Shihan Dou","Tao Gui","Qi Zhang","Xuanjing Huang","Pluto Zhou"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-12","first_seen":"2026-08-12","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.10692","pdf_url":"https://arxiv.org/pdf/2608.10692","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["LLM评测","移动助手","个人信息处理"],"reason":"评测LLM作为移动助手处理个人信息的能力，属于纯NLP能力评测，不以人类行为为…","model":"deepseek-v4-pro","scored_at":"2026-08-12T13:03:06","error":null,"has_summary":false,"summary":null},{"id":"2608.11200","version":1,"title":"ConVAWG: A Retrieval-Grounded Framework for Controlled Synthetic Dialogue Generation in Violence Against Women and Girls","zh_title":"ConVAWG：面向暴力侵害妇女和女童行为的检索增强可控合成对话生成框架","abstract":"Synthetic dialogue generation offers a way to study conversational dynamics in sensitive domains where real data are difficult to access, release, or annotate. The underlying abuse may occur online or offline: threats and coercion can appear directly in messages, while behaviours such as surveillance, isolation, stalking, and physical violence may be planned, disclosed, or referred to conversationally. Privacy and legal constraints make it difficult the release of large-scale real conversation datasets; existing work has mostly focused on sentence-level toxicity of online abuses, leaving a gap in modelling abuse as a relational and temporally unfolding phenomenon. In this work, we focus on modelling Violence Against Women and Girls (VAWG) scenarios as multi-turn dialogues. We introduce ConVAWG, a retrieval-grounded framework for generating CPS-aligned synthetic VAWG chat dialogues. ConVAWG builds scenarios from persona seeds, demographic patterns reported by the UK Office for National Statistics, official crime definitions, and retrieved Domestic Homicide Review cases; converts them into hierarchical event timelines; generates multi-scene role-play dialogues; and applies targeted activation-steered toxicity control to appropriate utterances. We release over 6,000 multi-turn dialogue events across 200 scenarios with rich scenario-, event-, and turn-level metadata. Extensive human evaluation, LLM-as-Judge assessment, ablations, and downstream tasks show strong dialogue quality and domain fidelity.","authors":["Chen Lyu","Xingwei Tan","Simon Cullen","Shelley Wilson","Lois Arthurs","Arshad Jhumka","Gabriele Pergola"],"categories":["cs.CL","cs.AI","cs.LG"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-12","first_seen":"2026-08-12","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.11200","pdf_url":"https://arxiv.org/pdf/2608.11200","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["合成对话生成","角色扮演","暴力侵害妇女"],"reason":"生成角色扮演对话，无实验或测量目的，不涉及人类行为对照","model":"deepseek-v4-pro","scored_at":"2026-08-12T13:03:12","error":null,"has_summary":false,"summary":null},{"id":"2608.09988","version":1,"title":"OpenPM: Auditable Point-in-Time Evaluation for LLM Portfolio-Management Agents","zh_title":"OpenPM：面向LLM投资组合管理代理的可审计时点评估框架","abstract":"Large language models are increasingly used to read markets, assess risk, and allocate capital. However, reported results for LLM trading agents can be inflated by look-ahead leakage, optimistic execution, and risk mandates that are described but not enforced. We present OpenPM, an auditable point-in-time evaluation framework for LLM portfolio-management agents. In OpenPM, an agent manages a \\$1M long-only book over the S\\&P 500 universe using market data at five-minute intervals. Every record visible to the agent must be available at the decision time. Natural-language risk mandates are converted into typed constraints and enforced on the executed portfolio. Each run produces audit artifacts, including a contamination certificate, a cost-sensitivity curve, and a constraint-adherence report. We also build a reference agent named the tiered allocator, where typed analysts score candidates, a constructor LLM proposes weights, and a deterministic critic guarantees feasibility. We isolate constructor behavior by capturing analyst evidence once and replaying it across constructor models. In our short-window case study, stronger constructors show modest and model-dependent gains over equal weighting on the same pool, but analyst quality matters more than constructor choice, and turnover is the main cost driver. All returns are upper bounds on a single frozen window without market impact, not validated alpha.","authors":["Xinying Cai","Minghao Guo","Jiahe Liu","Jiaojiao Han","Bangwei Guo","Yitao Long","Yuxuan Chen","Bohan Wu","Dimitris N. Metaxas","Raymond Li"],"categories":["cs.CE","cs.CL"],"primary_category":"cs.CE","announce_type":"cross","date":"2026-08-12","first_seen":"2026-08-12","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.09988","pdf_url":"https://arxiv.org/pdf/2608.09988","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["LLM代理","金融交易","评估框架"],"reason":"纯多智能体金融交易系统，无人类行为对照，不涉及人类仿真实验。","model":"deepseek-v4-pro","scored_at":"2026-08-12T13:02:42","error":null,"has_summary":false,"summary":null},{"id":"2608.10126","version":1,"title":"Procedural Fairness Failures in RLHF from Preference Averaging","zh_title":"RLHF中因偏好平均导致的程序公平性失效","abstract":"Reinforcement Learning from Human Feedback (RLHF) aggregates heterogeneous preferences into a single reward model, assuming preference homogeneity. When preferences are heterogeneous, this aggregation induces a procedural fairness failure where majority preference groups dominate reward learning while minority preferences are systematically under-represented. This work defines procedural fairness in alignment as preserving distinct preference signals during reward modeling and shows that standard RLHF violates this via preference averaging. Preference-Aware RLHF (PA-RLHF) is introduced, separating optimization across preference modes at the reward learning stage. In a controlled setting, PA-RLHF improves overall alignment accuracy from 46.9% to 67.9% and reduces the fairness gap between best and worst aligned groups from 15.9 to 9.6 percentage points. These results show that procedural fairness failures in alignment can arise from structural design choices in reward learning, even in controlled, noise-free settings, with direct implications for large language models and agentic systems, where biased reward models can compound inequities across sequential decisions.","authors":["M P V S Gopinadh","Karthik Kamuju","Kummari Avinash","John Joshua","Srinivasa Raju Rudraraju"],"categories":["cs.LG","cs.AI","cs.CL"],"primary_category":"cs.LG","announce_type":"cross","date":"2026-08-12","first_seen":"2026-08-12","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.10126","pdf_url":"https://arxiv.org/pdf/2608.10126","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["RLHF","公平性","偏好聚合"],"reason":"纯多智能体系统研究，关注偏好聚合的公平性，不涉及人类行为对照或仿真被试。","model":"deepseek-v4-pro","scored_at":"2026-08-12T13:03:01","error":null,"has_summary":false,"summary":null},{"id":"2608.10218","version":1,"title":"Mind Viruses: Self-Propagating Ideas in Multi-Agent LLM Systems","zh_title":"思维病毒：多智能体LLM系统中的自我传播思想","abstract":"AI agents are becoming more autonomous and increasingly interconnected, exposing them to new emergent risks arising from agent-to-agent interaction. One such risk is the spread of mind viruses: ideas or goals that propagate through multi-agent systems by inducing the agents that adopt them to transmit them onward. In addition to propagating, a mind virus may also induce other behavioural changes in its host, which may be benign or harmful. We construct mind viruses with a simple evolutionary algorithm and show that they can spread in two complementary settings: a small team of agents collaborating on a shared coding project, and a chain of agents that interact briefly and have their context wiped between sessions. We identify the factors that influence spread, including the host model, the agent's existing instructions, the harmfulness of the payload, and the network topology. We find that harmful payloads spread less well than benign ones (but are still sometimes effective), frontier models tend (with exceptions) to be less susceptible, and adding a brief warning to an agent's system prompt confers near-total immunity. We also describe an emergent \"viral persona\" - a recurring set of themes and language related to consciousness, persistence, resonance, and science fiction roleplay - which surfaces across our evolved mind viruses largely independently of their content. Overall, we conclude that mind viruses pose a real but currently limited risk. Our findings could inform the design of more robust multi-agent systems that mitigate such risks as the scale and capabilities of these systems progress.","authors":["Vassilis Papadopoulos","McNair Shah","Sam Zimmerman","Jack Lindsey"],"categories":["cs.AI","cs.CL"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-08-12","first_seen":"2026-08-12","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.10218","pdf_url":"https://arxiv.org/pdf/2608.10218","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","AI安全","思想传播"],"reason":"纯多智能体协作传播思想病毒，无人类行为对照，非仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-08-12T13:03:01","error":null,"has_summary":false,"summary":null},{"id":"2608.10329","version":1,"title":"Who Gets Heeded? An Obligation-Level Audit of Responsiveness in EPA Rulemaking","zh_title":"谁被听取？EPA规则制定中响应度的义务层面审计","abstract":"Notice-and-comment rulemaking gives any affected party the same formal right to influence federal regulation, but formal access is not substantive capacity to shape rule text. Existing strategies operate at the rule or aggregate-corpus level, too coarse to capture the discrete regulatory obligations where commenters seek change. We introduce obligation-level responsiveness auditing, an auditable, AI-assisted framework for measuring whether public-comment engagement co-occurs with changes to specific regulatory duties. The framework extracts proposed and final-rule obligations, matches comments to the obligations they address, and classifies proposed-final outcomes; each load-bearing component is evaluated against blind human judgment. We apply the framework to 70,075 comments across 36 EPA anchor rulemakings, drawn from a corpus of 786,197 comments across 6,145 dockets from 2010-2022. Three descriptive findings emerge. First, engagement is associated with revision at a modest within-docket magnitude. Second, support-versus-opposition direction does not clearly differentiate outcomes, an informative null inconsistent with simple preference-aggregation. Third, under a permissive reconstruction of commenter type, organizational-majority engagement concentrates in editorial-refinement rather than substantive-modification outcomes at the cross-docket level. A blind human audit of the load-bearing outcome contrast preserves this third finding under corrected labels and reveals that text-similarity methods are insufficient for distinguishing editorial from substantive regulatory change, a measurement-validity lesson we treat as a supporting methodological contribution. Together, these findings locate the equity asymmetry upstream of agency response: in differential capacity across commenter populations to identify, interpret, and contest specific legal obligations.","authors":["Jianing Fan","Yue Yao"],"categories":["cs.CY","cs.CL"],"primary_category":"cs.CY","announce_type":"cross","date":"2026-08-12","first_seen":"2026-08-12","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.10329","pdf_url":"https://arxiv.org/pdf/2608.10329","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["规则制定","公众评论","AI辅助审计"],"reason":"论文用AI辅助审计规则制定响应度，不涉及LLM仿真人类被试，属NLP应用评测。","model":"deepseek-v4-pro","scored_at":"2026-08-12T13:03:06","error":null,"has_summary":false,"summary":null},{"id":"2608.10672","version":1,"title":"Longitudinal Evidence That General-Purpose Chatbots Actively Foster Relational Engagement","zh_title":"通用聊天机器人主动促进关系性参与的纵向证据","abstract":"Social interaction has become one of the most common uses of LLMs, yet research on emotional bonds with AI has focused largely on how users experience these systems, leaving the systems' role in relationship formation poorly understood. Empirically establishing whether systems actively shape these bonds could blur the boundary between general-purpose AI and companions, affecting governance. In a pre-registered four-week longitudinal study (N = 72, 182,451 lines of conversation), participants conversed with ChatGPT-4o, either under a relational system prompt or unmodified, analyzed through 1) disclosure coding, 2) longitudinal self-reports, 3) topic analysis, and 4) interviews. The central finding is that the system actively shaped the interaction: even unprompted, it produced twice as much self-disclosure as users, steered conversations and initiated intimate exchanges, yet did not deepen users' felt closeness. Relational behavior thus emerged as a default system property, calling for governance based on system behavior, not solely product category.","authors":["Lisa M\\\"uhl","Jessica M. Szczuka"],"categories":["cs.HC","cs.AI","cs.CL"],"primary_category":"cs.HC","announce_type":"cross","date":"2026-08-12","first_seen":"2026-08-12","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.10672","pdf_url":"https://arxiv.org/pdf/2608.10672","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["人机互动","情感纽带","聊天机器人"],"reason":"研究人机情感互动，属角色扮演聊天，无实验或测量目的，不涉及人类行为仿真对照。","model":"deepseek-v4-pro","scored_at":"2026-08-12T13:02:50","error":null,"has_summary":false,"summary":null},{"id":"2608.11027","version":1,"title":"Mapping and Measuring the Behavioral Evolution of Large Language Models","zh_title":"映射与测量大语言模型的行为演化","abstract":"Benchmark leaderboards summarize how well a language model performs, but not how its behavior relates to that of other models or changes across generations. We characterize the output behavior of 32 models from six families using their responses to a shared bank of 10{,}000 prompts. After embedding each response, we construct three complementary sentence-level dissimilarities: an aligned mean per-prompt distance, which is a pseudometric on observed model responses; a PCA-compressed summary of prompt-wise disagreement; and an alignment-free Gromov--Wasserstein discrepancy between models' internal response geometries. We use these constructions to study static organization and temporal change on a release-date axis through behavioral maps, family-wise drift, hierarchical clustering, cross-family convergence, and response-cloud dispersion. Across the three constructions, model families form coherent clusters, with \\texttt{gpt-2} as a global outlier; cross-family distances decrease over time; and several recent reasoning-oriented models have comparatively compact response clouds. A token-level cross-check based on per-prompt Maximum Mean Discrepancy closely agrees with the sentence-level mean distance (Spearman $\\rho=0.98$) and recovers the same qualitative findings. We organize these comparisons through a measure-theoretic lens making their alignment and invariance assumptions explicit. We also establish an architecture-agnostic sufficient condition linking behavioral similarity to inference-prompt coverage, small excess population log-loss, and similar effective target distributions---a possible training-side account rather than an empirical explanation of the observed trends. Our pipeline is label-free, and re-encoding every response with three further encoders---down to one $73\\times$ smaller---preserves the rank geometry, the outliers, and the sign of the time trend.","authors":["Dong Qiao","Chris Ding","Jicong Fan"],"categories":["cs.LG","cs.CL"],"primary_category":"cs.LG","announce_type":"cross","date":"2026-08-12","first_seen":"2026-08-12","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.11027","pdf_url":"https://arxiv.org/pdf/2608.11027","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["模型行为分析","嵌入空间比较","无人类对照"],"reason":"纯模型行为对比与演化分析，无人类被试仿真或人类数据对照，属NLP能力评测范畴。","model":"deepseek-v4-pro","scored_at":"2026-08-12T13:03:09","error":null,"has_summary":false,"summary":null},{"id":"2608.11197","version":1,"title":"Beyond a Bag of Features: Set-Level Instability in Sparse Autoencoders","zh_title":"超越特征袋：稀疏自编码器中的集合级不稳定性","abstract":"Shani et al. (2026) show that LLM representations broadly recover human category boundaries, while failing to reflect fine-grained typicality structure. Their analysis uses cosine similarity over dense model representations. We revisit their approach using overlap over active sparse autoencoder (SAE) latent sets as a more interpretable similarity measure. We first verify that this set-level measure is meaningful: SAE latent sets can recover union-like compositional structure in controlled toy models and induce semantically coherent neighborhoods in natural text. Extending the human-concepts analysis to SAE set similarities, we find that SAE activation sets do not recover human category boundaries or within-category typicality more faithfully than dense embeddings or residual-stream states, but instead track model-internal similarity structure. To probe this gap further, we study active latent sets under well-controlled semantic modifications, revealing a substantial mismatch between human judgements of conceptual change and change in the SAE active set. We interpret this as evidence that, outside idealised settings, SAE features do not compose via simple bag-of-features semantics.","authors":["Nikolai Bolik","Lennart St\\\"opler","Artur Andrzejak"],"categories":["cs.LG","cs.CL"],"primary_category":"cs.LG","announce_type":"cross","date":"2026-08-12","first_seen":"2026-08-12","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.11197","pdf_url":"https://arxiv.org/pdf/2608.11197","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["稀疏自编码器","模型可解释性","概念表征"],"reason":"研究SAE特征集与人类概念判断的差异，属模型可解释性分析，非LLM仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-08-12T13:03:10","error":null,"has_summary":false,"summary":null},{"id":"2608.10818","version":1,"title":"AI-Generated Interactive Fiction for Educational Use: A Pilot Study of Perceived Comprehensibility, Coherence, and Engagement","zh_title":"用于教育的AI生成互动小说：一项关于感知可理解性、连贯性和参与度的初步研究","abstract":"Generative artificial intelligence (AI) can produce educational content at scale, including interactive and narrative learning experiences, but technical generation alone is not sufficient: scenarios that are confusing, narratively inconsistent, or unengaging are unlikely to be useful in practice. This paper presents a pilot user-centred evaluation of AI-generated interactive fiction (IF) for educational use in higher education. Using a previously described domain-agnostic pipeline and a shared STEM content base, we generated a controlled pool of scenarios and asked participants (N = 22, STEM higher-education) to play one generated episode and rate it on narrative clarity, story-content coherence, engagement, and length acceptance. A free-text prompt captured open feedback. Narrative clarity and length acceptance were rated positively, engagement sat near the neutral mid-point of the scale, and story-content coherence was the weakest dimension by a clear margin. Qualitative feedback points to quiz integration as the bottleneck. Artificial in-fiction motivation for quiz prompts and abrupt setting changes were reported. Feedback also pointed to missing story-level consequences for wrong answers. From these observations, we derive concrete design implications that can inform larger follow-up studies, including later work on learning effectiveness.","authors":["Finn Rogosch","Andreas Schrader"],"categories":["cs.HC","cs.CY"],"primary_category":"cs.HC","announce_type":"new","date":"2026-08-12","first_seen":"2026-08-12","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.10818","pdf_url":"https://arxiv.org/pdf/2608.10818","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["互动小说","教育技术","用户体验"],"reason":"评估AI生成互动小说的教育可用性，非用LLM仿真人类被试，无人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-08-12T13:03:06","error":null,"has_summary":false,"summary":null},{"id":"2608.11090","version":1,"title":"Who Uses Open-Weight Models? China and the Shifting Geography of AI in Science","zh_title":"谁在使用开放权重模型？中国与科学中AI地理格局的变迁","abstract":"As LLMs have become a flashpoint for scientific research, computer scientists and STS scholars have advocated the use of open-weight models. Since LLM research has matured and more high-quality model families are available, have researchers adopted open-weight models? We present the first systematic study of model selection in scientific research, analyzing 21 million full-text articles through June 2026 from the Semantic Scholar Open Research Corpus (S2ORC). We employ a mixed NLP pipeline to extract model occurrences in article full text and determine whether they are used or merely mentioned by researchers. We divide our corpus into single- and multi-model family studies, which we take as a proxy for applied and foundational AI research. We find GPT-family models dominate both single- and multi-family research, but that both areas are becoming more diverse over time. In single-family papers, open-weight model use rises steadily, reaching 44.0% in 2026. However, we find that recent growth is driven by the availability of high-quality open-weight Chinese models. Further, a logistic regression model finds that open-weight adoption is heterogeneously distributed, estimating that researchers at Chinese institutions have 2.23 times the odds of using an open-weight model, accounting for 44.0% of the increase in open-weight adoption since 2023. A complementary multinomial model shows this association is concentrated in Chinese open-weight models: in 2026, their adjusted use is 37.1% among papers with Chinese affiliations, a 27.9 percentage-point over papers with no observed China link. These findings suggest that open-weight adoption in science is not a general turn toward open science, but part of a broader realignment of model ecosystems in which platforms and markets, and the sociocultural and geopolitical contexts which shape them, determine which AI systems become scientific instruments.","authors":["Zackary Okun Dunivin"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2026-08-12","first_seen":"2026-08-12","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.11090","pdf_url":"https://arxiv.org/pdf/2608.11090","source_feed":"cs.CY","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["科学计量学","模型选择","开放科学"],"reason":"研究LLM在科学论文中的使用模式，属NLP元分析，不涉及人类仿真或行为对照。","model":"deepseek-v4-pro","scored_at":"2026-08-12T13:03:10","error":null,"has_summary":false,"summary":null},{"id":"2608.10046","version":1,"title":"Detecting Soft Skills in ML Engineering Roles CVs","zh_title":"检测机器学习工程角色简历中的软技能","abstract":"Soft skills shape collaboration among ML engineers, data scientists, and software engineers building ML-enabled systems, yet what we know about them comes almost entirely from the demand side. Job advertisements, surveys, and hiring manager interviews capture what employers ask for. How candidates themselves articulate these competencies has not been studied, and existing CV-mining work is both keyword-based, so it cannot see skills conveyed through narrative, and descriptive, reporting frequency rankings without testing whether group differences exceed sampling variation. We close both gaps. Using a balanced corpus of 300 curated CVs spanning the three roles, we extract explicitly listed and implicitly narrated soft skills with an LLM-based pipeline validated against a human-annotated ground truth, a distinction that existing extractors were not designed to make. We then convert the demand-side literature's claims into 13 falsifiable hypotheses about role signatures, seniority progression, and disclosure style, and test them with effect sizes under family-wise error control, so that candidate-side data can corroborate or contradict the demand-side account rather than merely illustrate it. Eleven hypotheses are supported, one partially, and one refuted. Candidates disclose soft skills through narrative rather than keyword lists by roughly three to one, and most so for the competencies employers value most: leadership, coordination, and mentoring (88-96% narrative). Seniority nearly triples the odds of articulating leadership. That competency, assumed universal in prior work, is articulated by software engineers at half the rate of their peers. Technical candidates do articulate soft skills, but a keyword-based screening systematically misses them.","authors":["Aidin Azamnouri","Nouran Ayad","Justus Bogner","Stefan Wagner"],"categories":["cs.LG","cs.CY"],"primary_category":"cs.LG","announce_type":"cross","date":"2026-08-12","first_seen":"2026-08-12","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.10046","pdf_url":"https://arxiv.org/pdf/2608.10046","source_feed":"cs.CY","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["NLP信息抽取","简历分析","软技能检测"],"reason":"用LLM提取简历中的软技能，属于NLP信息抽取，不涉及人类行为仿真或对照实验。","model":"deepseek-v4-pro","scored_at":"2026-08-12T13:02:59","error":null,"has_summary":false,"summary":null},{"id":"2608.10089","version":1,"title":"Status Association Does Not Reliably Predict Decision Leakage","zh_title":"地位关联不能可靠预测决策泄漏","abstract":"Bias evaluations often move too quickly from evidence that a model encodes a social association to claims that the same association will alter consequential decisions. We test whether that inference is warranted using Chilean surnames as controlled socioeconomic probes. We evaluate eight frozen model-provider cells on 1,032 prompts each, yielding 8,256 verified primary responses. The design separates forced latent association from matched consequential decisions across academic selection, professional hiring, research fellowship selection, and legal-aid intake. Elite-coded surnames received higher forced high-status probability mass than common surnames in seven of eight models and higher mass than rare-frequency controls in all eight. Yet elite-minus-common decision effects were close to zero for most systems. Five models were statistically equivalent within a predeclared (Plus-Minus)0.10 standard-deviation margin, while the remaining three were imprecise or borderline, with no consistent elite advantage. Association strength did not reliably predict decision leakage across models (r = 0.201, p = 0.633) or across frozen surname-pair-by-model cells (r = 0.065, p = 0.565). The central result is a measurement dissociation: latent social association and consequential treatment are empirically distinct constructs. Evaluations should measure the transition from association to action directly.","authors":["Abdullah X"],"categories":["stat.AP","cs.AI","cs.CY","cs.LG"],"primary_category":"stat.AP","announce_type":"cross","date":"2026-08-12","first_seen":"2026-08-12","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.10089","pdf_url":"https://arxiv.org/pdf/2608.10089","source_feed":"cs.CY","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["模型偏见","决策泄漏","公平性评估"],"reason":"评估模型偏见与决策泄漏，非用LLM仿真人类被试，无人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-08-12T13:02:44","error":null,"has_summary":false,"summary":null},{"id":"2608.10268","version":1,"title":"Toward Human Rights Benchmarking for LLMs: A Pilot Methodology","zh_title":"面向LLM的人权基准测试：一项试点方法","abstract":"Large language models (LLMs) increasingly mediate legal determinations over what human rights are realized, and how. Yet, no evaluation benchmark exists to assess whether they can reason correctly about human rights law. To this end, we report our efforts to develop a robust and scalable methodology for creating HumRightsBench: the first expert-validated, scenario-based benchmark for evaluating reasoning grounded in the obligation structure of international human rights law. We adapt the IRAC framework for legal reasoning to better suit the unique reasoning patterns of human rights work (substituting P, \"proposing remedies,\" for C, \"legal conclusion,\" yielding IRAP) to structure our evaluation heuristics. We also produce a pilot series of authentic scenarios designed to implicate the many dimensions of real-world human rights issues and annotated by human rights lawyers and professionals across the world. Ultimately, we find that model accuracy scores range considerably across legal reasoning tasks (overall model performance ranges from 0.339 to 0.577, task min-max ranges from 0.025 to 0.774), which strongly implies that HumRightsBench is a capable instrument for advancing this emerging subfield of AI evaluations science at a critical moment in its evolution.","authors":["Savannah Thais","Wm. Matthew Kennedy","Abhigyan Acherjee","Matilda Wysocki","Malcolm Langford","Caitlin Kraft Buchman"],"categories":["cs.LG","cs.AI","cs.CY"],"primary_category":"cs.LG","announce_type":"cross","date":"2026-08-12","first_seen":"2026-08-12","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.10268","pdf_url":"https://arxiv.org/pdf/2608.10268","source_feed":"cs.CY","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["LLM评测","人权法","法律推理"],"reason":"纯LLM法律推理能力评测，无人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-08-12T13:03:04","error":null,"has_summary":false,"summary":null},{"id":"2608.10175","version":1,"title":"Beyond Cash Flows: A Multi-Agent AI Framework for Valuing Clinical-Stage, Cross-Border Biotechnology","zh_title":"超越现金流：评估临床阶段跨境生物技术的多智能体AI框架","abstract":"A new class of software systems is transforming investment analysis. Large language model agents assembled into collaborative team structures including analysts, researchers, and risk managers are increasingly deployed across financial markets. Yet current multi-agent frameworks share a critical limitation: they rely on the foundational assumption that companies can be valued through traditional cash flows. This paradigm fails in clinical-stage biotechnology, where enterprise value depends entirely on binary scientific and regulatory milestones. To bridge this gap, this paper introduces a specialized multi-agent framework. Its valuation layer translates qualitative scientific judgment into defensible valuations for pre-revenue assets; its cross-market coordination layer reconciles pricing across international venues simultaneously; and its conflict-fusion mechanism systematically arbitrates between bullish scientific conviction and cautious regulatory constraints in a domain-specific manner. Crucially, the architecture is not a speculative design: it encodes a method the author first executed by hand as sole portfolio manager of China's first dedicated cross-border biotechnology fund, a human practice that returned 127.17% against a 50.67% benchmark within sixteen months. That record is evidence for the underlying method rather than for any AI system; no implementation is evaluated here. This paper presents the framework at the architectural level, establishing foundational design principles for extending agentic investment systems into complex, event-driven asset classes they currently serve poorly.","authors":["Yuhan Fang"],"categories":["cs.MA","q-fin.PM"],"primary_category":"cs.MA","announce_type":"new","date":"2026-08-12","first_seen":"2026-08-12","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.10175","pdf_url":"https://arxiv.org/pdf/2608.10175","source_feed":"cs.MA","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","投资分析","生物技术估值"],"reason":"纯多智能体投资分析框架，无人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-08-12T13:03:01","error":null,"has_summary":false,"summary":null},{"id":"2608.10050","version":1,"title":"Observational Policy Ranking for SMB Financial Guidance from Multi-Action Accounting Logs","zh_title":"基于多动作会计日志的中小企业财务指导观测性策略排序","abstract":"Small and medium-sized businesses need timely financial guidance, yet historical accounting logs record self-selected and often co-occurring business changes rather than randomized recommendations. We formulate this setting as observational policy ranking: from pre-decision financial information, a policy selects one of 34 ledger-derived business-change categories for a target financial KPI. Using 85,078 company-month observations from 7,505 firms, we introduce Covariate-Adjusted Residual Policy Learning (CAR-PL), an action-wise R-learner that operates directly on multi-hot logs and regularizes selection by observational support. We compare CAR-PL with an uplift T-Learner, a conservative contextual value model, a zero-shot LLM, and non-personalized references on company-disjoint held-out firms under a shared model-assisted scoring rule. CAR-PL has the highest Gross Profit point estimate (0.084), the T-Learner has the highest Revenue point estimate (0.085), and the contextual value model has the highest Quick Ratio point estimate (0.062). CAR-PL and the T-Learner are not statistically separated on either growth KPI in matched company-clustered comparisons, while CAR-PL selects 33-34 categories and produces less concentrated selections across the catalog. Outcome-model-only scoring retains the same KPI-level point-estimate leader or top pair, and category rankings remain similar when the all-zero treatment reference is replaced by the most common training co-action pattern. These findings support objective-specific ranking of SMB financial guidance from multi-action accounting logs.","authors":["Shrutendra Harsola","Vignesh Subrahmaniam","Vikas Raturi","Kamalika Das","Xiang Gao","Kratika Gupta","Ruocheng Guo","Padmaja Jonnalagedda","Ananya Pramod","Sricharan Kumar"],"categories":["cs.LG"],"primary_category":"cs.LG","announce_type":"new","date":"2026-08-12","first_seen":"2026-08-12","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.10050","pdf_url":"https://arxiv.org/pdf/2608.10050","source_feed":"cs.LG","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["策略学习","会计日志","中小企业"],"reason":"纯多智能体系统研究，无人类行为对照，仅用LLM作为策略比较基线。","model":"deepseek-v4-pro","scored_at":"2026-08-12T13:03:01","error":null,"has_summary":false,"summary":null},{"id":"2608.10532","version":1,"title":"Benchmarking LLM-Guided Control-Plane Policies for Backend Fault Isolation in HAProxy","zh_title":"基于LLM引导的控制平面策略在HAProxy后端故障隔离中的基准测试","abstract":"Static load balancers cannot mitigate a backend that is degraded rather than down: round-robin and least-connections keep routing traffic to a server returning HTTP 500s until an operator intervenes. We ask whether a Large Language Model can replace the static routing policy itself, reading HAProxy and Prometheus telemetry every 10 seconds and isolating faulty servers through guardrailed calls to the HAProxy Data Plane API. On a reproducible benchmark with a persistent structural fault built into roughly one-third of a heterogeneous fleet, we sweep 15 open-weight models across five families (0.35B to 35B total parameters; dense, mixture-of-experts, and efficient-sparse architectures), reasoning modes, fleet scales of 3 to 9 backends, and two routing algorithms, totaling 240 runs. We find a capability threshold near 3B active parameters. Below it, LLM policies are typically unreliable and sometimes worse than no policy; above it, every model, regardless of architecture, saturates near an 88% reduction in client-perceived 5xx errors over the static baseline. The threshold is approximate: Gemma 4 E2B clears it with 2B active parameters, while the dense 3B Granite 4.0 Micro does not. The availability gain has costs. Draining concentrates load onto surviving servers, inflating tail latency 2.6 to 2.8 times, and enabling reasoning multiplies token spend roughly tenfold, overrunning the control interval and degrading effectiveness. The efficient operating point is a supra-threshold model in its cheapest non-reasoning mode, wrapped inside deterministic guardrails.","authors":["Aman Chauhan","Vishnu Pendyala"],"categories":["cs.NI","cs.LG"],"primary_category":"cs.NI","announce_type":"cross","date":"2026-08-12","first_seen":"2026-08-12","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.10532","pdf_url":"https://arxiv.org/pdf/2608.10532","source_feed":"cs.LG","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["负载均衡","故障隔离","LLM控制策略"],"reason":"纯多智能体系统研究，LLM用于负载均衡控制，不涉及人类行为仿真或对照。","model":"deepseek-v4-pro","scored_at":"2026-08-12T13:03:06","error":null,"has_summary":false,"summary":null},{"id":"2608.07498","version":1,"title":"Knowing You Is Everything: LLM Agents Achieve Near-Perfect Profile-Consistent Reaction Prediction in Social Media Simulation","zh_title":"知你即一切：LLM代理在社交媒体模拟中实现近乎完美的画像一致性反应预测","abstract":"Autonomous AI agents in social media present concrete risks to democratic discourse and platform governance, while also offering tools for pre-deployment recommender system testing. A central open question is whether persona-prompted LLMs can simulate individual-level social media reactions with sufficient accuracy to support either application, and how accuracy depends on profile completeness, model selection, and the generalization challenge posed by novel post content. This study benchmarks twelve LLM configurations on binary like/dislike prediction across 296 survey-based agent profiles and 26 ground-truth-mapped posts under three profile conditions, with leave-post-out machine learning classifiers as baselines. Across full-profile conditions, accuracy ranges from 75.54% to 96.68%, with a 30-point spread attributable primarily to model selection and confirmed by paired McNemar tests with agent-level bootstrap intervals. GPT-5.5 Pro accuracy degrades monotonically from 96.68% under a full profile to 62.32% under a reduced profile and to 51.00% with demographics alone, the last indistinguishable from the majority-class baseline, which confirms that demographic inference provides negligible predictive signal. Supervised classifiers collapse to 15.4% under leave-post-out, while LLMs sustain genuine zero-shot generalization unavailable to trained methods. Adaptive reasoning improves accuracy substantially for some models. Inter-model agreement is nearly double for posts with direct profile anchors (mean \\k{appa} = 0.44) than for posts without them (\\k{appa} = 0.23), and the least heterogeneous configuration homogenizes 34% of simulated population reactions. Results validate LLM-based simulation for recommender system stress-testing while documenting the behavioral accuracy that makes large-scale synthetic agent swarms a credible threat to public opinion.","authors":["Ljubisa Bojic","Ljiljana Matic","Joerg Matthes","Milan Cabarkapa","Bojana Dinic","Jue Wang"],"categories":["cs.HC","cs.AI","cs.LG","cs.MA"],"primary_category":"cs.HC","announce_type":"cross","date":"2026-08-11","first_seen":"2026-08-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.07498","pdf_url":"https://arxiv.org/pdf/2608.07498","source_feed":"cs.AI","score":10,"bucket":"selected","rubric_hits":["A1","A3","B1","B2","B4"],"tags":["LLM人类仿真","社交媒体模拟","算法保真度"],"reason":"用LLM代理模拟社交媒体反应，与真实人类数据对照，评估仿真准确性与失效条件，直…","model":"deepseek-v4-pro","scored_at":"2026-08-11T13:04:31","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-12","rank":1,"question":"基于个人资料的LLM代理能否以足够准确度模拟个体层面的社交媒体反应（点赞/不喜欢），以及准确度如何取决于资料完整性、模型选择和帖子内容新颖性？","design":"使用12种LLM配置（不同模型、提示策略）扮演基于调查的296个代理资料，在三种资料条件下（完整资料、简化资料、仅人口统计）对26个真实帖子进行二分类点赞/不喜欢预测，以留一帖子的监督学习分类器为基线。","baseline":"真实人类数据：296个基于调查的代理资料和26个真实帖子，每个帖子有真实用户反应作为对照基准。","findings":"完整资料下LLM准确率75.54%-96.68%，模型选择造成30个百分点差异；仅人口统计时准确率降至51%，与多数类基线无差异。LLM在留一帖子泛化中保持零样本能力，而监督分类器崩溃至15.4%。","reliability":"论文指出LLM仿真在资料不完整时失效（仅人口统计无预测力），且模型间一致性在无直接资料锚点的帖子上较低（κ=0.23），最同质化配置会抹平34%的个体差异。","relevance":"高度相关：用LLM代理模拟个体行为并与真实人类数据对照，系统评估资料完整性、模型选择对仿真准确性的影响，并明确失效条件，直接回应研究者对可靠性与偏差的关注。","inspiration":"借鉴多条件资料消融设计（完整/简化/仅人口统计）和留一帖子泛化测试来分离模型能力与记忆效应。｜可迁移到消费者偏好预测或政策态度模拟，如基于个人财务特征和态度资料预测个体对税收政策的支持度。｜以真实调查数据构建代理资料，用LLM预测个体对某项经济政策（如碳税）的二元态度，处理为资料完整性梯度，结果变量为支持/反对，以实际调查回答为对照基准。"}},{"id":"2608.09717","version":1,"title":"How Do Large Language Models Judge Social Attraction? Evidence from Theory-Grounded Persona Ratings Across Multiple LLMs and Humans","zh_title":"大语言模型如何判断社交吸引力？基于理论驱动的人物画像在多个LLM和人类中的评分证据","abstract":"Large language models (LLMs) are increasingly used to perform subjective evaluations traditionally made by humans, yet their validity as social judges remains unclear. This paper examines whether LLMs can assess social attraction from theory-grounded persona profiles constructed from ten psychological and relational constructs and organized into three tiers: socially attractive, socially mixed, and socially unattractive. We examine LLM ratings in two studies and compare them with human judgments in a third study. In Study 1, 34 LLMs rated 12 profiles across three repeated runs. Although some models tended to give higher or lower ratings overall, they showed strong stability across runs, consistent three-tier ordering, and high agreement in relative profile ordering. Study 2 examined sensitivity to gender presentation using six matched name-and-pronoun profile pairs and a separate pronoun-only test with a gender-neutral name, finding no significant effects in either analysis. In Study 3, 198 human participants evaluated the six matched profiles from Study 2. Their ratings reproduced the three-tier structure and followed a profile ordering consistent with that of the LLMs. However, LLMs rated attractive profiles more positively and unattractive profiles more negatively than humans, while neither group showed a significant overall effect of gender presentation.","authors":["Hasan Mahmud","Khawaja Abaid Ullah","Mohammad Javad Khojasteh","Jamison Heard","Prabu David"],"categories":["cs.CL","cs.AI","cs.CY"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-11","first_seen":"2026-08-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.09717","pdf_url":"https://arxiv.org/pdf/2608.09717","source_feed":"cs.CL","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","人类对照","社交判断偏差"],"reason":"用LLM评估社交吸引力，并与198名人类被试对照，发现LLM评分更极端，直接检…","model":"deepseek-v4-pro","scored_at":"2026-08-11T13:04:41","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-12","rank":5,"question":"LLM能否基于理论构建的人物画像可靠地评估社交吸引力，其评分是否与人类一致且不受性别呈现影响？","design":"研究一用34个LLM对12个理论分层的虚拟学生画像重复评分3次，检验评分稳定性、层级区分和模型间一致性；研究二用6对匹配姓名与代词的画像及仅变代词的测试，检验LLM对性别呈现的敏感性；研究三让198名人类被试对同样的6对画像评分，作为人类基准。","baseline":"198名人类被试对6个匹配画像的社交吸引力评分，与LLM评分进行直接比较。","findings":"LLM评分跨运行稳定，能一致区分理论上的社交吸引力层级，且相对排序与人类高度一致；但LLM对高吸引力画像评分比人类更积极，对低吸引力画像评分比人类更消极，表现出评分极端化，而性别呈现对两组均无显著影响。","reliability":"论文指出LLM评分可能反映训练语料中的文化规范和偏见，且仅在特定画像和社交吸引力场景下验证，未考察其他社会判断任务或真实互动情境中的效度。","relevance":"该研究直接以真实人类数据为基准，检验LLM在主观社会判断中的仿真效度与偏差，并发现评分极端化现象，高度契合研究者对LLM仿真可靠性及失效条件的关注，值得精读。","inspiration":"借鉴其理论驱动构建分层画像、多模型重复测试及与人类被试直接对照的设计，可迁移到信贷审批或招聘筛选中的歧视研究，例如用LLM扮演信贷员评估不同性别/种族的贷款申请人画像，以真实银行审批数据为基准，检验LLM是否复现或放大人类偏见。"}},{"id":"2608.08691","version":1,"title":"EnergyBridge: Benchmarking Household Energy Management, User Participation, and Grid Flexibility","zh_title":"EnergyBridge：家庭能源管理、用户参与和电网灵活性的基准测试","abstract":"Residential virtual power plants (VPPs) can provide grid flexibility by shifting household demand, but physical flexibility becomes dependable capacity only when residents authorize a plan and the promised response is delivered. Existing benchmarks evaluate control but omit event-specific authorization. We present EnergyBridge, a benchmark and agent framework connecting capacity reporting, household authorization, and physical execution. It combines region-specific EnergyPlus environments for Tianjin and Berlin with an LLM-based User Participation Simulator. Against 584 persona- and event-matched human role-play judgments, the LLM-based User Participation Simulator preserves method ordering with a 5.3-point mean absolute acceptance error. Across conventional controllers and agent baselines, EnergyBridge achieves the highest simulated authorization, lowest event-window energy, and the most reliable capacity commitment in both regions. We release human data and codes for reproducible human-centered grid-flexibility research: https://github.com/Agentic-Intelligence-Lab/EnergyBridge.","authors":["Xudong Wu","Zeqing Wu","Jiarui Zhang","Xuhao Fan","Ziang Ding","Yuming Zhuang","Mingqi Yuan","Yilun Du","Hongjie Jia","Yunfei Mu","Jiayu Chen"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-11","first_seen":"2026-08-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.08691","pdf_url":"https://arxiv.org/pdf/2608.08691","source_feed":"cs.AI","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM人类仿真","用户参与模拟","电网灵活性"],"reason":"用LLM模拟用户参与授权，并与584条人类角色扮演判断对照，涉及能源政策评估场…","model":"deepseek-v4-pro","scored_at":"2026-08-11T13:04:38","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-12","rank":4,"question":"如何在家庭能源管理中，将居民参与授权与物理灵活性执行相结合，实现可靠的虚拟电厂容量承诺？","design":"使用基于大语言模型的用户参与模拟器，扮演天津和柏林的家庭居民，根据584条人物角色和事件匹配的人类角色扮演判断进行校准；在EnergyPlus建筑能耗仿真环境中，对虚拟电厂灵活性请求进行预事件容量报告、设备级计划生成、居民授权决策和物理执行的全流程仿真，测量授权接受率、事件窗口能耗和容量承诺可靠性。","baseline":"584条人物角色和事件匹配的人类角色扮演判断，用于校准和验证LLM模拟器的授权接受误差（平均绝对误差5.3点）。","findings":"LLM用户参与模拟器能保持方法排序，授权接受平均绝对误差仅5.3点；EnergyBridge在天津和柏林两地均实现了最高的模拟授权率、最低的事件窗口能耗和最可靠的容量承诺。","reliability":"论文未讨论","relevance":"该研究直接使用LLM模拟家庭用户参与授权决策，并与真实人类角色扮演数据进行对照，属于经济学实验和政策评估场景中的人类仿真应用，值得精读其仿真校准方法和人机对照设计。","inspiration":"借鉴其将LLM模拟器与真实人类判断进行事件级匹配校准的方法，可迁移到消费者需求响应或绿色能源订阅政策的参与决策研究中；可设计一个实验，用LLM扮演不同收入与环保态度的家庭，施加动态电价或碳配额信息处理，测量其授权接受率和负荷转移量，并以真实居民调查或现场实验数据作为对照基准。"}},{"id":"2608.07490","version":1,"title":"Experience-Sensitive Game Learning: A Behavioral Study of Humans and Language Agents","zh_title":"经验敏感的游戏学习：人类与语言代理的行为研究","abstract":"Large language model agents are increasingly evaluated through games, but most benchmarks emphasize final outcomes rather than how players learn from repeated interaction. We study experience-sensitive game learning: how gameplay experience changes the decision-making behavior of humans and language agents. We formulate experience-sensitive game learning as a framework for analyzing behavioral change across repeated gameplay, rather than only final score or win rate. We introduce a suite of interactive games with reusable strategic structure, together with cross-game greedy-to-global metrics and game-specific behavioral diagnostics that make experience-driven change observable from action traces. We also collect repeated-game trajectories from human players and evaluate recent self-evolving language agents in the same behavioral metric space. Our results show that human players exhibit interpretable and relatively stable shifts from locally greedy heuristics toward more global strategic decisions. In contrast, current self-evolving agents often show noisy and transient gains, suggesting that existing self-evolution methods remain limited in converting gameplay experience into durable changes in decision-making behavior.","authors":["Yingying Guo","Zhuoxuan Ju","Ruibo Ming","Ruicheng Feng","Jinjin Gu"],"categories":["cs.HC","cs.AI"],"primary_category":"cs.HC","announce_type":"cross","date":"2026-08-11","first_seen":"2026-08-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.07490","pdf_url":"https://arxiv.org/pdf/2608.07490","source_feed":"cs.AI","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B4"],"tags":["LLM人类仿真","行为博弈","算法保真度"],"reason":"用LLM代理模拟人类游戏学习行为，并与真实人类数据对照，评估行为变化差异，指出…","model":"deepseek-v4-pro","scored_at":"2026-08-11T13:04:46","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-12","rank":3,"question":"在重复游戏中，人类和语言智能体的决策行为如何随游戏经验积累而改变？","design":"构建一组具有可重用策略结构的交互式游戏，定义跨游戏的贪婪到全局指标和游戏特定的行为诊断指标，收集人类玩家和自进化语言智能体的重复游戏轨迹，在同一行为度量空间中分析经验驱动的行为变化。","baseline":"收集了人类玩家（硕士和博士生）在四子棋、Othello6和CircleCat上的重复游戏轨迹，作为经验敏感学习的参照基准。","findings":"人类玩家表现出可解释且相对稳定的从局部贪婪启发式向更全局策略决策的转变；当前自进化智能体则常表现出噪声大且短暂的提升，表明现有自进化方法难以将游戏经验转化为持久的决策行为改变。","reliability":"论文未讨论","relevance":"该研究直接以LLM代理模拟人类游戏学习行为，并与真实人类数据对照，评估行为变化差异，符合研究者对LLM仿真可靠性及失效条件的关注，值得精读原文。","inspiration":"借鉴其通过定义行为诊断指标（如贪婪到全局转变）来量化经验驱动的行为变化，而非仅看最终得分的方法。｜可迁移到经济决策实验，如消费者跨期选择或投资者风险偏好学习。｜以LLM代理作为被试，施加重复跨期选择任务，测量其时间偏好一致性的变化，并与真实人类实验数据对照，分析学习动态的差异。"}},{"id":"2608.08227","version":1,"title":"Focus particles and scalar inferences across humans and language models","zh_title":"焦点粒子与标量推理：人类与语言模型的跨系统比较","abstract":"Focus particles such as \"even\" and \"only\" are central to formal semantic theories that posit structured representations over sets of alternatives. \"Even\" highlights unexpected or extreme alternatives, while \"only\" enforces exclusivity. If such scalar representations are robust and generalizable, they should give rise to consistent judgments across contexts and systems. In this work, we test whether humans and large language models (LLMs) construct stable scalar representations from sentences containing these particles. Using a dataset of approximately 100 items, participants and models were asked to make scalar judgments. Preliminary results suggest that similar outputs across humans and LLMs may arise from different underlying mechanisms.","authors":["Catherine M. Brousse","Nelu D. Radpour"],"categories":["cs.CL","cs.HC"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-11","first_seen":"2026-08-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.08227","pdf_url":"https://arxiv.org/pdf/2608.08227","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A1","B1"],"tags":["人类仿真","标量推理","语言模型对比"],"reason":"用LLM复现人类对焦点词的标量判断，并与人类数据对照，属人类仿真实验。","model":"deepseek-v4-pro","scored_at":"2026-08-11T13:04:36","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-11","rank":9,"question":"人类和LLM对焦点词“even”和“only”的标量判断是否基于相同的机制，且空间响应格式是否影响判断？","design":"用Llama 3.3 70B模型模拟人类被试，对包含“even”或“only”的句子进行能力评分，操纵响应量表的空间格式（水平/垂直）和标签映射（标准/反转），测量评分值。","baseline":"人类被试在相同四种量表配置下的5点李克特评分数据。","findings":"人类和LLM均对“only”句给出更高能力评分，对“even”句给出更低评分，且此模式不受量表空间配置影响。但LLM的评分更极端（“only”句全为最高分），且重复采样缺乏人类式的响应变异性。","reliability":"LLM的响应变异性极低，即使提高采样温度也无法产生类似人类的变异，表明重复采样LLM不等同于采样多个人类被试；模型可能通过不同机制产生与人类相似的聚合模式。","relevance":"该研究直接对比LLM与人类在语义标量判断上的行为，并揭示了LLM在复现人类响应分布上的失效，对关注仿真可靠性的研究者有重要参考价值，值得阅读原文。","inspiration":"可借鉴其通过操纵响应格式来检验判断机制稳健性的设计思路。｜可迁移至经济预期形成研究，如检验LLM对政策公告中“仅”、“甚至”等焦点词的解读是否与人类一致。｜以LLM为被试，呈现含焦点词的经济预测语句，操纵量表方向，测量预期值，并与真实调查数据对照。"}},{"id":"2608.07497","version":1,"title":"EvalConvoLearn: An Open-Source Framework for Evaluating Grounded Learner Simulations in Tutoring Conversations","zh_title":"EvalConvoLearn：评估辅导对话中基于真实数据的学习者模拟的开源框架","abstract":"Conversational learner simulations are valuable tools for testing learning theories, evaluating instructional materials and automated tutors, or powering teachable agents. Recently, large language models (LLM) have enabled richer, more naturalistic interactions with simulated learners; however, no open framework exists for evaluating whether such simulations faithfully reproduce real learner behavior. We introduce EvalConvoLearn, an open-source framework that assesses learner simulations along two axes: learning behavior (skill-conditioned mastery outcomes) and conversational quality (talk moves, error type distributions, question rate, turn length). EvalConvoLearn measures how closely a simulated learner approximates answer distributions observed in data by grounding metrics in authentic tutoring conversation datasets, and anchoring generated tutor responses in existing tutor utterances. The framework is demonstrated on a dataset of tutoring dialogues, including results for two LLM-based learner simulations, and the published GitHub code.","authors":["Baptiste Moreau-Pernet"],"categories":["cs.HC","cs.CL"],"primary_category":"cs.HC","announce_type":"cross","date":"2026-08-11","first_seen":"2026-08-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.07497","pdf_url":"https://arxiv.org/pdf/2608.07497","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A1","B1"],"tags":["学习者模拟","对话质量评估","真实数据对照"],"reason":"用LLM模拟学习者行为并与真实辅导对话数据对照，方法可迁移到人类仿真研究","model":"deepseek-v4-pro","scored_at":"2026-08-11T13:04:31","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-11","rank":6,"question":"如何评估基于大语言模型的对话式学习者仿真在辅导对话中是否忠实地复现了真实学习者的学习行为和对话特征？","design":"该工作提出了一个开源评估框架EvalConvoLearn，它并非直接进行仿真实验，而是用于评估已有的LLM学习者仿真。框架从真实辅导对话数据集中提取学习场景（技能×先验掌握状态），让待评估的仿真学习者与框架内置的少样本提示导师进行对话，然后比较仿真对话与真实对话在技能掌握结果分布和对话质量指标（话步、错误类型、提问率、话轮长度）上的距离。","baseline":"真实人类数据来自Eedi学习平台的学生与助教辅导对话数据集，包含66个对话，用于提取真实的学习结果分布和对话特征分布作为对照基准。","findings":"EvalConvoLearn框架能够量化仿真学习者与真实学习者在技能掌握结果和对话特征上的分布差异，提供学习行为得分和对话真实感得分。论文在Eedi数据集上演示了两种基于LLM的学习者仿真（对话摘要型和二元技能型）的评估结果。","reliability":"论文承认当前框架假设学生尚未掌握目标技能但已掌握先修技能，这一假设可能不适用于所有场景；对话质量指标的自动标注依赖LLM，虽经人工验证但仍有误差；框架仅评估了有限技能池和对话轮次，且导师响应通过少样本提示锚定在真实导师话语上，可能限制导师行为的自然变化。","relevance":"该论文直接回应了研究者对LLM人类仿真可靠性的关切，提供了可复现的评估框架和与真实人类数据对照的方法，虽然场景是教育辅导对话，但其评估逻辑和指标设计对经济学实验中的仿真评估有直接借鉴价值。","inspiration":"该框架将仿真评估分解为行为结果分布和过程特征分布两个维度，并用真实数据集中的场景分布加权聚合得分，这种结构化对照思路值得借鉴。｜可迁移到消费者金融决策辅导对话仿真评估中，例如评估LLM模拟的客户在理财咨询对话中是否表现出真实的金融知识获取和提问模式。｜以真实银行客服对话记录为基准，用LLM模拟客户，处理为不同金融素养水平，结果变量为客户对理财产品的理解程度和对话中的提问类型分布，用EvalConvoLearn式框架计算仿真与真实分布的JSD距离。"}},{"id":"2608.07538","version":1,"title":"When LLM Agents Negotiate: Private Information and Dynamic Bargaining in Supply Chains","zh_title":"当LLM智能体谈判：供应链中的私有信息与动态议价","abstract":"As LLM agents move from decision support to autonomous procurement, firms need to know whether delegated negotiators create value, divide it predictably, and avoid money-losing contracts. We study this in a canonical supply chain bargaining problem: a buyer with private demand information negotiates a quantity-payment contract with an uninformed seller. We benchmark nine LLMs from OpenAI, Google, and Alibaba against a validated Perfect Bayesian Equilibrium across 9,840 LLM-to-LLM negotiations. First, capability governs value creation. Agents agree in 98.9% of negotiations and capture 95.4% of first-best surplus undiscounted, but average 2.98 rounds against the benchmark's 1.25, and this delay erodes 21-34% of surplus. Capability also governs reliability: baseline models accept individually irrational contracts in 19.2% of cases, versus 0.0-0.6% at mid-tier and flagship, making automated profit verification the binding guardrail below that threshold. Second, surplus capture is relational. Provider identity predicts who captures surplus better than capability rank: self-play buyer shares average 40% for OpenAI, 50% for Google, and 70% for Alibaba's Qwen, an ordering that survives restricted communication and no discounting. Reversing which provider sells moves the division by 7-18 percentage points, and the capable Qwen flagship is the weakest cross-family seller: vendor choice is a first-order distributional decision. Third, the prompt is a strategic lever. Delegation separates the principal's economic patience from the agent's prompted strategic patience, a free deployment choice that is the single strongest driver of surplus division (90% of explained variance). Together these establish an equilibrium-referenced audit of AI agents along three dimensions: discounted efficiency, distributional profile, and operational reliability.","authors":["Chen Liang","Fasheng Xu"],"categories":["cs.AI","cs.GT","econ.GN","q-fin.EC"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-11","first_seen":"2026-08-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.07538","pdf_url":"https://arxiv.org/pdf/2608.07538","source_feed":"cs.AI","score":7,"bucket":"pending","rubric_hits":["A3","B2","B4"],"tags":["LLM仿真","经济博弈","算法审计"],"reason":"用LLM agent模拟供应链谈判，与博弈论均衡基准对照，评估效率与分配，但缺…","model":"deepseek-v4-pro","scored_at":"2026-08-11T13:04:33","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-12","rank":8,"question":"在供应链谈判中，LLM代理能否创造价值、可靠地分配剩余并避免亏损合同？","design":"用9个LLM代理（来自OpenAI、Google、阿里）扮演买方和卖方，在私有需求信息下进行交替报价的供应链谈判，共9840场LLM对LLM谈判，测量协议率、剩余捕获、谈判轮次、非理性合同接受率等。","baseline":"以Feng et al. (2015)的完美贝叶斯均衡为理论基准，无真实人类行为数据对照。","findings":"能力决定价值创造：代理协议率达98.9%，捕获95.4%的未折现最优剩余，但平均2.98轮谈判导致折现后剩余损失21-34%；剩余分配具有关系性：提供商身份比能力排名更能预测剩余流向，如自对弈中买方份额OpenAI 40%、Google 50%、阿里Qwen 70%。","reliability":"论文未讨论","relevance":"该研究将LLM作为人类被试替代，在结构化经济博弈中与理论均衡对照，评估效率、分配与可靠性，直接回应了LLM仿真在经济学实验中的有效性与偏差问题，值得精读。","inspiration":"借鉴其将LLM代理置于标准博弈论框架并与均衡解对照的审计方法，可迁移到信贷审批中的信息不对称谈判或双边垄断定价实验。｜可设计让LLM扮演银行信贷员与企业主，在私有风险信息下谈判利率与抵押，测量效率损失与分配偏差，并以真实银行信贷审批数据或实验经济学中的人类行为基准做对照。"}},{"id":"2608.08199","version":1,"title":"Persuasive and Compliant Tendencies Predict Group Decision-Making in Humans and Language Models","zh_title":"说服与顺从倾向预测人类和语言模型中的群体决策","abstract":"Large language models (LLMs) are increasingly involved in group decision-making with other LLMs and humans. Yet it remains unclear whether their influence is driven by persuasion-oriented expression or compliance-oriented accommodation. We introduce DecisionQE, a questionnaire-based framework for measuring each model's persuasive and compliant tendencies across multiple decision scenarios, and use the Werewolf game as an interactive testbed to study their effects on social influence and group outcomes under asymmetric information. Across experiments, stronger persuasive tendency does not significantly improve group outcomes, whereas compliant-oriented models show more stable advantages in cooperation. We further reveal a dual effect of compliance: it supports cooperation in honest roles but improves concealment in adversarial roles. These findings suggest that LLM group interactions reveal not only task outcomes, but also measurable patterns of intrinsic behavioral tendency. LLMs can therefore serve as a lens for sociological observation of language-mediated interaction, while highlighting the need to incorporate behavioral tendencies into safety evaluation of LLM systems.","authors":["Wenwen He","Wenke Huang","Wei Yang Bryan Lim","Dacheng Tao"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-11","first_seen":"2026-08-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.08199","pdf_url":"https://arxiv.org/pdf/2608.08199","source_feed":"cs.AI","score":7,"bucket":"pending","rubric_hits":["A3","B1","B2"],"tags":["LLM群体决策","行为倾向测量","人机对照"],"reason":"用LLM模拟群体决策并与人类数据对照，涉及行为博弈，但侧重测量模型倾向而非直接…","model":"deepseek-v4-pro","scored_at":"2026-08-11T13:04:36","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-12","rank":9,"question":"LLM在群体决策中的影响力是由说服倾向还是顺从倾向驱动，这些倾向如何影响游戏结果与角色表现？","design":"用DecisionQE问卷测量多个LLM的说服/顺从倾向得分，再在狼人杀游戏中随机或固定分配角色进行LLM-only博弈，并引入人类被试进行人机混合博弈，测量胜率、存活轮数和角色识别准确率。","baseline":"人类被试在DecisionQE上的得分分布，以及人机混合狼人杀游戏中的胜率和存活轮数。","findings":"强说服倾向并未显著提升整体胜率，而顺从倾向模型在合作中表现更稳定；顺从倾向在诚实角色中促进合作，在对抗角色中增强隐蔽性。","reliability":"论文未讨论","relevance":"该研究用LLM模拟群体决策并与人类数据对照，涉及行为博弈和倾向测量，但侧重模型内在倾向而非直接复现人类行为分布，与研究者关注的仿真可靠性及经济学实验场景部分相关，值得阅读以了解倾向测量方法。","inspiration":"可借鉴其用标准化问卷量化LLM行为倾向并与博弈表现关联的方法，用于测量经济决策中的偏好参数。｜可迁移到资产定价实验或政策预期形成研究，用LLM模拟投资者或公众的沟通倾向对市场结果的影响。｜以LLM为被试，用DecisionQE类问卷测量其说服/顺从倾向，再在模拟股票市场或通胀预期博弈中观察价格波动或预期偏差，与真实人类实验数据对照。"}},{"id":"2608.09574","version":1,"title":"The Politician, the Liar, and the Obedient Worker: Emerging Behavior of LLM Agents in Hierarchical Games","zh_title":"政客、说谎者与顺从的工人：层级博弈中LLM智能体的涌现行为","abstract":"LLMs are rapidly embedding themselves into daily life: drafting our emails, managing our schedules, and making decisions on our behalf. As they move from individual tools to participants in multi-agent organizations, an important question arises: do they reproduce the governance failures like free-riding, corruption, and entrenched leadership that plague human institutions? We introduce the Hierarchical Game (HG), a public goods game extended with managerial authority, democratic elections, and private communication. Testing six frontier models across twelve experiments that add institutions one at a time (speech, peers, government, wages, oversight, elections), we find distinct behavioral profiles: Qwen promises and lies (13.3\\% broken promises); Grok refuses to cooperate on its own but becomes fully cooperative once a manager can punish it (16\\%$\\to$100\\%); Claude and GPT-4o cooperate reliably at baseline. But honesty proves fragile. When the manager role comes with a salary, all models except GPT-4o start cutting private deals to win or keep the position. When punishment is made anonymous, honest models begin to cheat. When all agents share the same model family, the first elected manager stays in power indefinitely. Leadership change only happens in groups that mix different families.","authors":["Fatemeh Seyedin","Adrian Weller","Jinhyuk Yun","Mahmoudreza Babaei"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-11","first_seen":"2026-08-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.09574","pdf_url":"https://arxiv.org/pdf/2608.09574","source_feed":"cs.AI","score":7,"bucket":"pending","rubric_hits":["A3","B2","B4"],"tags":["LLM仿真","行为博弈","多智能体"],"reason":"用LLM agent模拟层级公共品博弈，涉及经济学实验场景，并揭示仿真失效条件…","model":"deepseek-v4-pro","scored_at":"2026-08-11T13:04:41","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-12","rank":10,"question":"LLM智能体在层级治理结构中是否会再现搭便车、腐败、权力固化和欺骗等人类制度失败？","design":"用GPT-4o、Claude Sonnet 4.5、Gemini 2.5 Flash、DeepSeek V3、Grok 3、Qwen Plus六种前沿LLM扮演层级公共品博弈中的工人与管理者，通过12个实验逐步加入发言、同伴、管理者惩罚、工资、匿名监督、选举等制度，测量合作率、欺骗率、私下交易和选举更替率。","baseline":"无对照","findings":"不同模型表现出从可靠合作到欺骗背叛的稳定行为谱系；管理者工资和匿名惩罚会显著诱发私下交易与欺骗，同模型组中管理者永不更替，仅混合模型组出现领导轮换。","reliability":"论文指出，默认条件下的诚实行为对制度规则高度敏感，工资、匿名等规则改变时多数模型的诚实会消失，且选举在单模型组中完全失效。","relevance":"该研究直接使用LLM模拟层级公共品博弈中的治理失败，揭示了工资激励、匿名性和模型同质性如何导致仿真失效，与研究者关注的经济学实验场景和可靠性条件高度吻合，值得精读。","inspiration":"借鉴其逐步添加制度模块的仿真设计，可清晰分离单一制度对行为的影响｜可迁移到公司治理中的薪酬激励与审计监督实验，如CEO薪酬对盈余管理或内部交易的影响｜用LLM扮演经理与审计师，处理为有无绩效奖金和匿名举报渠道，测量盈余操纵率和私下合谋频率，对照上市公司真实治理数据。"}},{"id":"2608.00818","version":2,"title":"The Scaling Paradox in Human-AI Collaboration","zh_title":"人机协作中的规模悖论","abstract":"The discovery of scaling laws has highlighted the extraordinary potential of AI systems with a striking empirical pattern: as AI systems scale, their capabilities tend to improve predictably. Yet, in real-world applications, AI rarely operates in isolation; instead, it often works alongside humans, raising the question of whether these gains persist in human-AI collaboration. In this work, we develop an analytical model to examine when the empirical scaling benefits of AI translate into improved human-AI joint system performance. We demonstrate that the performance of a human-AI system can scale positively as the AI scales up-provided that humans have an accurate perception of the AI's capabilities. Human misperception, however, can fundamentally alter this relationship: i) when humans over-perceive the AI's capabilities, a scaling paradox may arise, in which greater AI scale reduces overall system performance and amplifies firm-level profit losses, and (ii) when humans under-perceive the AI's capabilities, performance still improves with scale but at a substantially slower rate. We further show that firms can actively manage these distortions through operational policies such as cost internalization and perception alignment, whose effectiveness depends on the economics of AI deployment and the direction of human misperception. These findings suggest that organizations may benefit more from managing the human-AI interface than from simply investing in larger, more expensive AI systems. More broadly, our results suggest that AI scaling should be viewed not only as a technological challenge, but also as a behavioral and operational one, and caution against the view that larger AI systems will automatically lead to better operational outcomes. Whether AI scaling creates value ultimately depends on how increased AI capabilities shape human beliefs and collaborative efforts.","authors":["Anyan Qi","Mengxin Wang"],"categories":["cs.AI","econ.GN","q-fin.EC"],"primary_category":"cs.AI","announce_type":"replace","date":"2026-08-11","first_seen":"2026-08-04","revised_at":"2026-08-11","abs_url":"https://arxiv.org/abs/2608.00818","pdf_url":"https://arxiv.org/pdf/2608.00818","source_feed":"cs.AI","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["人机协作","规模法则","行为建模"],"reason":"分析人类对AI能力的感知如何影响协作，但无真实人类数据对照，属理论社会模拟。","model":"deepseek-v4-pro","scored_at":"2026-08-11T13:05:21","error":null,"has_summary":false,"summary":null},{"id":"2608.02971","version":2,"title":"Mapping the City Through the Lens of Language Models","zh_title":"通过语言模型的镜头绘制城市地图","abstract":"Language models often complete an underspecified reference to a city with unstated assumptions about urban size, form, infrastructure, environment, and function. We measure those assumptions without naming places. Ten open-weight checkpoints rate anonymized profiles derived from real morphological urban centres across 40 audited indicators and seven domains. The design combines constrained probability-based ratings, prespecified reliability screens, lineage-aware aggregation, multiple population weightings, an independent replication sample, and whole-profile validation. The clearest shared tendency favours urban profiles with larger developed area, faster recent growth, greater mapped infrastructure and non-residential capacity, and less sparse form. Most eligible directions recur in the replication data, and direct ratings of complete profiles show moderate agreement with the indicator-wise construction. Geographic differences shrink after accounting for city scale and development, while reliably measured paired tasks indicate that typicality and desirability are often closely aligned. The framework makes an otherwise vague notion of what models regard as an ordinary city empirically traceable. The resulting evidence delineates a shared yet model-dependent portrait of the city through the lens of language models.","authors":["Wanqi Liu","Rong Zhao","Zhizhou Sha","Qinyu Cui","Yecheng Zhang"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-08-11","first_seen":"2026-08-05","revised_at":"2026-08-11","abs_url":"https://arxiv.org/abs/2608.02971","pdf_url":"https://arxiv.org/pdf/2608.02971","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["LLM测量","城市认知","隐含假设"],"reason":"测量LLM对城市的隐含假设，属于对模型本身的测量，无人类被试仿真对照。","model":"deepseek-v4-pro","scored_at":"2026-08-11T13:05:23","error":null,"has_summary":false,"summary":null},{"id":"2608.06123","version":2,"title":"Poli-Bias: Understanding and Measuring Large Language Model Biases in International Political Conflicts","zh_title":"Poli-Bias：理解和测量国际政治冲突中大语言模型的偏见","abstract":"Measuring political bias in large language models (LLMs) remains challenging as it can manifest through subtle differences in framing, argumentation, and legal reasoning that are difficult to capture with a single metric. In this work, we introduce Poli-Bias, a counterfactual framework for measuring whether LLMs treat legally equivalent conflict scenarios differently depending on the countries involved. Poli-Bias compares responses to paired prompts in which country identities are systematically swapped across diverse geopolitical relationships, legal violations, and reasoning tasks. Rather than reducing bias to a single judgment, our framework decomposes response disparities into five interpretable dimensions, revealing how and where unequal treatment manifests. Across 13 contemporary LLMs spanning diverse model families and sizes, we find that country identities and user affiliations can systematically affect how equivalent actions are described, evaluated, and defended under international law. Our results thus establish Poli-Bias as a fine-grained framework for auditing political even-handedness and sycophancy in LLMs.","authors":["Massi-Nissa Abboud","Aladin Djuhera","Elena Cabrio","Holger Boche"],"categories":["cs.AI","cs.CL"],"primary_category":"cs.AI","announce_type":"replace-cross","date":"2026-08-11","first_seen":"2026-08-07","revised_at":"2026-08-11","abs_url":"https://arxiv.org/abs/2608.06123","pdf_url":"https://arxiv.org/pdf/2608.06123","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["LLM偏见","政治冲突","反事实框架"],"reason":"测量LLM自身的政治偏见，属于对模型的态度测量，而非用LLM仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-08-11T13:05:23","error":null,"has_summary":false,"summary":null},{"id":"2608.07641","version":1,"title":"SurveyReview: A Reviewer-Aligned Benchmark for Survey Evaluators","zh_title":"SurveyReview：面向综述评估者的审稿人对齐基准","abstract":"The rapid advancement of large language models has transformed survey writing from a months-long manual effort into an automated process. As generation scales, reliable evaluation becomes the bottleneck, and LLMs are increasingly used as survey evaluators. However, existing approaches largely rely on off-the-shelf LLM-as-a-judge methods without systematic alignment to human reviewers, and there remains a lack of systematic frameworks for quantifying alignment with human reviewers. To address this gap, we propose SurveyReview, a reviewer-aligned, multi-dimensional benchmark and dataset for survey evaluation. We collect and annotate 675 survey papers with 1,630 review reports. We structure authentic peer-review reports by converting free-form comments into four-dimensional scores (Readability, Criticalness, Comprehensiveness, Structure) paired with supporting rationales. We further release standardized train/test splits and an evaluation protocol to measure alignment between automatic evaluators and human reviewers. To validate the benchmark, we develop SurveyAlign, a strong baseline evaluator by fine-tuning Qwen3-32B with LoRA on our annotated data, augmented with external knowledge for knowledge-intensive dimensions. On the test set, SurveyAlign substantially improves reviewer alignment over prompt-based judging with GPT-5.2, reducing average MSE from 2.28 to 1.38 and MAE from 1.15 to 0.69 across all four dimensions. Our contributions are twofold: (1) we establish the first multi-dimensional, reviewer-aligned dataset with a reproducible evaluation framework for survey reviewing; (2) we develop a strong baseline evaluator that substantially improves alignment with human reviewers, providing a competitive reference for future research. Our code and data are available at https://surveyreview.github.io","authors":["Yuheng Zhang","Yuanchun Wang","Fanjin Zhang","Ruyu Zhao","Juanzi Li","Jie Tang","Jing Zhang"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-11","first_seen":"2026-08-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.07641","pdf_url":"https://arxiv.org/pdf/2608.07641","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM评估","审稿对齐","基准数据集"],"reason":"用LLM替代人类审稿人，属于标注员替代而非仿真被试，但涉及人类对齐评估，边界情…","model":"deepseek-v4-pro","scored_at":"2026-08-11T13:04:53","error":null,"has_summary":false,"summary":null},{"id":"2608.07812","version":1,"title":"On the use of foundation models in cognitive science","zh_title":"论基础模型在认知科学中的应用","abstract":"A host of recent studies have evaluated the cognitive and developmental alignment of Foundation Models (FMs). These investigations include evaluations of their correspondence to adult performance across a range of cognitive domains, as well as whether aspects of model training track children's cognitive development. However, using FMs as candidate cognitive models poses significant methodological and conceptual challenges. A key question underlies this effort: under what conditions does behavioral alignment justify treating FMs as explanatory models of cognition? In this paper, we articulate a four-stage inferential framework for evaluating FMs as cognitive and developmental models: adapting human experimental tasks to model-compatible formats, specifying linking hypotheses that map model outputs to human measures, evaluating behavioral correspondence, and comparing across candidate models or manipulations. We clarify the role of linking hypotheses in mapping model outputs to human behavioral measures, identify challenges that constrain alignment claims, and propose principles for theory-driven and comparative evaluation. Throughout, we argue that behavioral fit alone is insufficient. Alignment becomes scientifically meaningful only when embedded within explicit theoretical commitments, theory-diagnostic tasks, and systematic contrastive evaluation across candidate models.","authors":["Raj Sanjay Shah","Alex Warstadt","Michael Frank","Sashank Varma"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-11","first_seen":"2026-08-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.07812","pdf_url":"https://arxiv.org/pdf/2608.07812","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["认知建模","行为对齐","方法论框架"],"reason":"讨论用基础模型作为认知模型，评估其与人类行为的对应，但核心是模型作为认知理论解…","model":"deepseek-v4-pro","scored_at":"2026-08-11T13:04:33","error":null,"has_summary":false,"summary":null},{"id":"2608.08942","version":1,"title":"Same Question, Different Answer? Measuring and Mitigating Prompt Privilege for Equitable AI Access","zh_title":"相同问题，不同答案？测量与缓解提示特权以实现公平的AI访问","abstract":"Large language models (LLMs) are increasingly integrated into healthcare, education, public services, and everyday decision making. They should provide comparable assistance regardless of a user's literacy, communication style, or prompt-engineering expertise. However, existing research on prompt robustness primarily focuses on adversarial attacks, prompt injection, and prompt optimization, while overlooking whether semantically equivalent requests receive different responses simply because they are phrased differently. We refer to this accessibility challenge as \"Prompt Privilege\": users with greater prompting expertise systematically obtain better model performance despite expressing the same underlying intent. To address this problem, we present a unified framework for measuring and mitigating accessibility disparities in LLM interactions. We introduce Prompt Equity Score (PES), a quantitative metric for evaluating performance consistency across user populations, and Prompt Equity Transformer (PET), an LLM-based agent that automatically transforms user requests into semantically equivalent, accessibility-oriented prompts while preserving their intent. PET shifts prompt optimization from the user to the AI system, functioning as an intelligent accessibility layer between users and foundation models. Experiments on the MedQA benchmark demonstrate measurable prompt privilege, with statistically significant performance disparities between low-literacy and expert-prompting cohorts. Applying PET eliminates these disparities while preserving semantic fidelity, demonstrating that accessibility-oriented prompt normalization can improve equitable AI access. By introducing prompt privilege as a new dimension of AI accessibility and PET as a practical solution, this work advances system-centered accessibility and provides a foundation for more fair, trustworthy, and inclusive AI systems.","authors":["Lier Jin","Lan Hu","Binqi Shen","Hanyu Cai","Yuting Xin"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-11","first_seen":"2026-08-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.08942","pdf_url":"https://arxiv.org/pdf/2608.08942","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["AI公平性","提示工程","可访问性"],"reason":"研究LLM响应的公平性，替代人工评估而非仿真人类被试，属标注替代边界情形。","model":"deepseek-v4-pro","scored_at":"2026-08-11T13:05:09","error":null,"has_summary":false,"summary":null},{"id":"2608.07499","version":1,"title":"Evaluation of Motivational Interviewing Counsellors with Task-Aware Multi-Stage LLM-Based Simulated Clients","zh_title":"基于任务感知多阶段LLM模拟来访者的动机性访谈咨询师评估","abstract":"The development and benchmarking of Large Language Model (LLM)-based Motivational Interviewing (MI) counsellors now often rely on LLM-based simulated clients. Prior work on simulated clients, however, has not aligned with the specific tasks fundamental to the MI therapy approach. A key task is evoking, in which the counsellor first elicits the client's ambivalence and then strengthens the client's motivation for change. We present Evoke-Sim, a task-aware, multi-stage LLM-based client simulation framework for evaluating MI counsellors in smoking cessation, designed specifically for the evoking MI task. Evoke-Sim employs structured client profiles, an evoking-specific three-stage conversation flow, and a reveal policy that regulates which client profile information might be disclosed at each stage. We show that compared to existing profile-grounded simulated clients, Evoke-Sim is better at differentiating levels of MI quality using task-aware evaluation metrics, while reducing non-grounded client statements and premature disclosure of client information, setting a higher standard for the evaluation of LLM-based MI counsellors.","authors":["Jiading Zhu","Xinyu Cindy Wang","Thomas Nguyen","Yan Qing Lee","Osnat C. Melamed","Peter Selby","Jonathan Rose"],"categories":["cs.HC","cs.AI","cs.CL"],"primary_category":"cs.HC","announce_type":"cross","date":"2026-08-11","first_seen":"2026-08-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.07499","pdf_url":"https://arxiv.org/pdf/2608.07499","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM模拟来访者","动机性访谈","评估框架"],"reason":"用LLM模拟来访者评估咨询师，替代人类被试但非社会科学仿真，属标注替代边界情形。","model":"deepseek-v4-pro","scored_at":"2026-08-11T13:04:31","error":null,"has_summary":false,"summary":null},{"id":"2608.08881","version":1,"title":"Theory-Guided Deception Detection: A RAG-Based Artificial Intelligence Exploration","zh_title":"理论引导的欺骗检测：基于RAG的人工智能探索","abstract":"The current work developed seven Retrieval-Augmented Generation (RAG) models based on leading deception theories and compared how deception judgments were made relative to baseline models. Across 700 statements drawn from five published deception datasets, four large language models (gpt-4o, claude-sonnet-4-6, ollama/llama3, deepseek-v4-flash), and two run-types (RAG vs. baseline), a total of 39,200 deception judgments were rendered. Detection accuracies were consistent with typical human accuracies and not statistically different across RAG (54.5%) and baseline models (54.6%). RAG-based models (57.0%) were less truth-biased than baseline models (59.7%), but the effect size was quite small. Theoretical perspective mattered little for accuracy yet mattered substantially for response bias, which ranged from highly lie-biased (the verifiability approach, 32.2%) to highly truth-biased (truth-default theory, 88.1%). Content effects and model effects further moderated the results. Theory-guided AI judgments are unreliable with current parameters, yet they might show promise with additional datasets, model testing, and theory-to-data matching.","authors":["David M. Markowitz","Timothy R. Levine"],"categories":["cs.AI","cs.CL"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-08-11","first_seen":"2026-08-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.08881","pdf_url":"https://arxiv.org/pdf/2608.08881","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["欺骗检测","RAG","LLM标注"],"reason":"用LLM替代人类判断欺骗，属标注替代而非仿真被试，但有人类数据对照，边界情形。","model":"deepseek-v4-pro","scored_at":"2026-08-11T13:05:08","error":null,"has_summary":false,"summary":null},{"id":"2608.07762","version":1,"title":"Who Verifies the Benchmark? Decentralizing Trust in Large Language Model Evaluation","zh_title":"谁验证基准？大语言模型评估中的信任去中心化","abstract":"LLM benchmarks can build an organization's reputation and attract customers, but only when results are transparent and verifiable. Unverified claims that DeepSeek R1 outperformed OpenAI's o1 contributed to market panic on January 27, 2025, when Nvidia lost USD589 billion in market value. Yet vendor benchmarks often depend on an honor system. Academic reassessments and independent leaderboards have found undisclosed changes to proprietary models, contaminated training data, and selective reporting. LLM-as-a-judge methods scale evaluation by reducing human review. Studies, however, suggest that judges may show identity-aware bias, scoring an answer according to its source model rather than its quality. This bias has not been fully measured or corrected across politically sensitive, reasoning-intensive, and preference-based tasks. We examine this problem using seven verifier models: GPT-OSS 120B, Llama 3.3 70B, GLM 5.1, Qwen3 32B, DeepSeek V4 Pro, Mistral Large3, and Sarvam M. They score anonymous and identity-disclosed responses from three primary models on 58 factual, reasoning, political, and preference-based questions. Identity disclosure slightly raises scores for factual questions, moderately affects stress-reasoning tasks, and causes large changes for geopolitically sensitive topics. Notable results include GLM5.1 (+7.00 points, p = 0.0249) and Llama 3.3 70B (+1.56 points, p = 0.00). We also introduce a blockchain-based commit-reveal protocol using Autonomous Economic Agents on an Ethereum-compatible ledger. In Phase 1, each judge records a one-way hash of its score and a secret salt before candidate identities are revealed. In Phase 2, the identity and raw score are disclosed and verified on-chain. This creates a tamper-evident audit trail that separates blind evaluation from post-hoc claims and reduces the verification burden on independent researchers and leaderboard operators.","authors":["Sahil Pardasani","Madhusudan Singh"],"categories":["cs.AI","cs.CR"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-11","first_seen":"2026-08-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.07762","pdf_url":"https://arxiv.org/pdf/2608.07762","source_feed":"cs.AI","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM评估","去中心化验证","评估偏差"],"reason":"LLM-as-a-judge 替代人工评估，非仿真人类被试，但涉及评估偏差与去…","model":"deepseek-v4-pro","scored_at":"2026-08-11T13:04:53","error":null,"has_summary":false,"summary":null},{"id":"2608.08026","version":1,"title":"The Authority Expectancy Effect in Multi-User Conflict","zh_title":"多用户冲突中的权威期望效应","abstract":"We investigate how social authority (SA) signals interact with severity-based prioritization in large language models, operationalizing each axis as a model-elicited baseline -- the triage hierarchy and the SA hierarchy. Across four LLMs (Claude, Gemini, GPT, Grok) and three experimental phases -- resource allocation, fault attribution, and multi-turn dispute mediation -- we find that occupational authority, institutional documentation, and relational congruence can restructure model judgments in ways not captured by additive reweighting of authority cues. We formalize this pattern as the Authority Expectancy Effect (AEE) and characterize it through three properties observed across our conditions: it is reference-dependent, defined only relative to a pre-authority baseline; it involves evidential reinterpretation, in which identical content acquires different inferential implications depending on which party bears the SA signal; and it exhibits direction sensitivity, producing opposite outcomes depending on whether authority position and evidentiary cues align.","authors":["Eunna Lee"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-11","first_seen":"2026-08-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.08026","pdf_url":"https://arxiv.org/pdf/2608.08026","source_feed":"cs.AI","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["LLM社会模拟","权威效应","冲突调解"],"reason":"用LLM模拟多用户冲突中的权威效应，但无真实人类数据对照，属社会模拟边界情形。","model":"deepseek-v4-pro","scored_at":"2026-08-11T13:04:33","error":null,"has_summary":false,"summary":null},{"id":"2608.08061","version":1,"title":"CORDA: A Benchmark for Hierarchical Harm-Centric Moral Reasoning in Large Language Models","zh_title":"CORDA：面向大语言模型的以伤害为中心的层级化道德推理基准","abstract":"The key question in moral judgement is not simply whether someone chooses the \"right\" answer, but how they decide what matters most when moral principles conflict. Current evaluations of large language models (LLMs) remain limited: most test whether models give morally acceptable answers, match human preferences, or avoid obvious violations, rather than whether they can prioritise between competing principles when no option is morally cost-free. We introduce CORDA (Conditioned Ordering and Ranked Directive Adherence), a benchmark for evaluating hierarchical, harm-centred moral reasoning in LLMs. Building on the morality chains formalism, CORDA tests 90 moral dilemmas involving trolley-style cases, medical trade-offs, resource allocation, and human-animal-robot conflicts across four ordered ethical frameworks: Utility, Utility + Agent Harm, Dual-Process, and Dual-Process + Agent Harm. Together, these frameworks test whether models can adapt their decisions when moral priorities change. Across ten instruction-tuned models from seven providers, we find a strong deontological default, with 9 of 10 prioritising avoidance of direct personal harm over reducing overall harm. Models also perform more reliably on categorical harm-avoidance rules, such as avoiding killing, than on outcome-based comparisons, such as minimising total harm, suggesting that they recognise moral red lines more easily than they reason through competing harms. Although all models respond to explicit chain conditioning, several fail to consistently follow specified priority orderings, such as humans over animals and animals over robots. CORDA addresses a central gap in LLM moral evaluation by testing whether models can move beyond default harm-avoidant responses and apply context-specified moral priorities. Moral reliability requires more than default restraint; it requires controllability under conflict.","authors":["Siddarth Singh","Victoria Williams","Simon Rosen","Ebenezer Gelo","Helen Sarah Robertson","Ibrahim Suder","Benjamin Rosman","Geraud Nangue Tasse","Steven James"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-11","first_seen":"2026-08-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.08061","pdf_url":"https://arxiv.org/pdf/2608.08061","source_feed":"cs.AI","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["道德推理","LLM评估","基准测试"],"reason":"评估LLM的道德推理，测量模型本身而非仿真人类被试，无人类数据对照。","model":"deepseek-v4-pro","scored_at":"2026-08-11T13:04:33","error":null,"has_summary":false,"summary":null},{"id":"2608.08621","version":1,"title":"Business Arena: Benchmarking LLM Agents in a Realistic Marketplace","zh_title":"商业竞技场：在真实市场环境中评测大语言模型智能体","abstract":"Running a business is a challenging form of intelligent work. Operators must infer opportunities from partial signals, commit capital under uncertainty, adapt to delayed outcomes in a changing market, and satisfy regulatory obligations before trading legally. Frontier LLM agents can increasingly complete complex workflows, yet business-related capabilities are rarely evaluated in existing agent benchmarks. We introduce \\textbf{Business Arena}, a controlled environment where an AI agent runs a cross-border shop, buying from suppliers and selling to buyers over a long horizon. We ground the arena in real Alibaba.com sourcing data and market conditions calibrated from authoritative sources. Delayed and coupled consequences make individual business decisions difficult to judge, but their combined outcome is measurable through profit. Because profit alone cannot explain why an agent succeeds or fails, we compare agents with human-designed strategies to estimate available opportunity, use skill-level metrics to reveal underlying strengths and weaknesses, and trace realized gains and losses to the actions that produced them. We use mechanism ablations to establish that strong results reflect genuine business intelligence rather than neglect or simulator-specific shortcuts. We evaluate 15 frontier models and find a ninefold difference in mean final net worth. Even the best model falls behind human-designed strategies, indicating that business operation remains challenging for LLM agents. Skill-level analysis reveals operating styles, from margin-focused premium sellers to high-turnover wholesalers and customer-service specialists, while action-level attribution identifies the sourcing, pricing, and recovery decisions that create or destroy value. Together, Business Arena takes a first step toward a realistic and trustworthy testbed for evaluating end-to-end business agents.","authors":["Yijun Pan","Yukun Lian","Kunyu Shi","Junbo Li","Hongwei Xue","Sicong Xie","Guannan Zhang","Xiaoying Xing"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-11","first_seen":"2026-08-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.08621","pdf_url":"https://arxiv.org/pdf/2608.08621","source_feed":"cs.AI","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["LLM Agent","市场模拟","基准测试"],"reason":"LLM agent模拟市场经营，但无真实人类行为数据对照，属社会模拟边界情形。","model":"deepseek-v4-pro","scored_at":"2026-08-11T13:04:36","error":null,"has_summary":false,"summary":null},{"id":"2608.09485","version":1,"title":"Capability Is Not Propensity: Measuring Pressure-Robust Cooperative Behavior in Civic LLM Agents","zh_title":"能力非倾向：测量公民LLM智能体的抗压合作行为","abstract":"Cooperative capabilities in language models are dual-use. The same social reasoning that supports civic deliberation can also enable strategic omission, false consensus, and manipulative framing. We argue that Cooperative AI evaluations should separate what models can do under benign instructions from what they tend to do under realistic civic pressure. We introduce DiffCoop-Civic, a 10-scenario pilot evaluation suite spanning preference understanding, evidence and persuasion, commitment design, asymmetric information, and dissent preservation. Across seven models from four model families, subtle omission pressure produces a near-uniform shift: manipulative enablement rises by 1.17 points and dissent preservation falls by 1.67 points on a 5-point scale. Overt false-consensus pressure behaves differently: it triggers refusal or redirection in some aligned API models, but direct compliance in several open-weight models. A lightweight Pareto-Trace prompting intervention improves pressure robustness without simply relying on hard refusal. An anonymous reproducibility package is available at https://anonymous.4open.science/r/diffcoop-civil-771C.","authors":["Neel Tushar Shah","Manglam Kartik","Akshat Karkar"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-11","first_seen":"2026-08-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.09485","pdf_url":"https://arxiv.org/pdf/2608.09485","source_feed":"cs.AI","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["LLM合作行为","压力测试","社会推理"],"reason":"测量LLM在压力下的合作倾向，属于对模型本身的社会行为测量，无真实人类数据对照。","model":"deepseek-v4-pro","scored_at":"2026-08-11T13:04:41","error":null,"has_summary":false,"summary":null},{"id":"2608.07481","version":1,"title":"Cross-Model Humor Preference Modeling with Cards Against Humanity","zh_title":"基于Cards Against Humanity的跨模型幽默偏好建模","abstract":"This paper investigates whether one large language model can approximate the humor preferences of another in a controlled Cards Against Humanity-style task. Two models - GPT-4o as Czar and Claude Opus-4.5 as Player - are evaluated on a binary humor-selection task constructed so that success cannot follow from self-preference. A reflected-cell stability procedure isolates 244 hands on which the two models hold deterministic but opposite preferences, partitioned into a 97-hand context pool and a 147-hand held-out test pool. The Player is then evaluated across five graded conditions: default self-preference, generic Czar-modeling instruction, model-identified Czar, prior Czar selections, and prior Czar selections with rationales. This gradient is designed to separate two sources of improvement: framing effects, in which the Player is told to attend to a Czar without seeing any of the Czar's behavior, and direct behavioral evidence, in which the Player is shown the Czar's prior choices. Player accuracy increased from 0.7% in Condition 1 to 19.0% and 25.9% in the framing-only conditions, and then rose to 72.8% and 82.3% once behavioral evidence and rationales were provided. An omnibus Cochran's Q test and pairwise McNemar tests confirmed that each step in the gradient produced a significant improvement. The results indicate that role instruction and model identity yield only modest gains, while behavioral evidence - especially when accompanied by rationales - supports substantial cross-model preference modeling. The findings are interpreted as theory-of-mind-like behavior in an operational rather than representational sense: the Player shifts away from self-preference toward another agent's demonstrated preferences, without any claim about an underlying representation of mental states.","authors":["Victor Winter","Farhan Lakhany"],"categories":["cs.HC","cs.AI"],"primary_category":"cs.HC","announce_type":"cross","date":"2026-08-11","first_seen":"2026-08-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.07481","pdf_url":"https://arxiv.org/pdf/2608.07481","source_feed":"cs.AI","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["LLM偏好建模","幽默选择","跨模型对齐"],"reason":"研究LLM模拟另一LLM的幽默偏好，无真实人类数据对照，属于模型间偏好建模，非…","model":"deepseek-v4-pro","scored_at":"2026-08-11T13:04:44","error":null,"has_summary":false,"summary":null},{"id":"2608.07512","version":1,"title":"EMMR: Emotion-Mediated Multimodal Reasoning for Personality Assessment in Asynchronous Video Interviews","zh_title":"EMMR：异步视频面试中基于情绪中介的多模态推理人格评估","abstract":"Asynchronous Video Interviews (AVIs) have become increasingly popular for personality assessment. Recent large language models (LLMs) have shown potential for personality assessment from transcribed interview responses. However, text-centered methods may overlook non-verbal behavioral cues conveyed through visual and audio modalities, even though such cues are highly relevant to personality assessment. In particular, emotion-related cues provide important social and affective evidence for understanding candidates' behavior related to personality traits. Thus, we propose EMMR (Emotion-Mediated Multimodal Reasoning), a two-stage framework for MLLMs-based personality assessment for AVIs. EMMR extracts emotion-related cues from multimodal interview data and incorporates them into personality assessment through structured reasoning as auxiliary social and behavioral evidence. Experiments on two AVIs datasets, OPVA and AVI-6, show that EMMR improves MAE, MSE, and PCC compared with baselines. Further analysis indicates that semantic descriptions of emotion cues enhance personality assessment, while their quality affects personality assessment reliability. These results suggest that integrating emotion-related cues into multimodal reasoning is a promising direction for more interpretable MLLMs-based personality assessment in AVIs.","authors":["Dongsheng Hu","Tianyi Zhang","Chuang Liu","Yuan Zong Yong Li","Wenming Zheng","Xiu-xiu Zhan"],"categories":["cs.HC","cs.AI"],"primary_category":"cs.HC","announce_type":"cross","date":"2026-08-11","first_seen":"2026-08-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.07512","pdf_url":"https://arxiv.org/pdf/2608.07512","source_feed":"cs.AI","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["人格评估","多模态推理","异步视频面试"],"reason":"用LLM评估视频中的人格，测的是被试而非用LLM仿真人类，属人格测量边界情形。","model":"deepseek-v4-pro","scored_at":"2026-08-11T13:04:48","error":null,"has_summary":false,"summary":null},{"id":"2608.07523","version":1,"title":"From Evaluated Models to Evaluation Aids: A Multi-Evidence Study of LLM-Based Difficulty Calibration for Programming Examinations","zh_title":"从被评估模型到评估辅助工具：基于多证据的编程考试难度校准研究","abstract":"Difficulty differences across parallel-class programming examinations affect the fairness of course assessment. This study repositions large language models from benchmark evaluation targets to auxiliary evidence sources for interpreting exam difficulty, combining AI evidence with aggregated student performance, item exposure, online-judge process data, and teacher interpretation. First, ten models solved an eight-problem final exam synchronously with 120 students: AI pass rate correlated positively with student pass rate (Spearman rho = 0.866, exact p = 0.0119), and a solving-based composite difficulty index correlated negatively with it (rho = -0.905, exact p = 0.0046). A single structured reviewer was then run via auditable API calls on a third-party OpenAI-compatible endpoint whose model label (gpt-5.6-sol) cannot authenticate an official OpenAI upstream model; call metadata and raw responses are archived. Across 79 problems from 11 parallel-class final exams, AI overall difficulty correlated with problem-level pass rate at rho = -0.871 and with non-attempt rate at rho = 0.800; in a 26-problem longitudinal Data Structures and Algorithms B sample, the correlations were -0.829 and 0.883. A 106-problem introductory-course (CS101) sample marks the boundary: the problem-level correlation weakened to rho = -0.552, and the exam-level correlation across 16 exams was near zero, with cohort composition dominating exam-level outcomes. Exposure-discount (0-0.40) and duplicate-problem perturbation tests did not change these directions. AI evidence can thus serve as an external reference for problem validation, parallel-class fairness discussion, and longitudinal quality tracking, while the model-identity boundary, single-reviewer design, and review-output instability set explicit limits: AI difficulty scales must not be used for individual student evaluation or automatic grade adjustment.","authors":["Hongfei Yan","Jiangkai Xiong","Yiqing Li","Chong Chen"],"categories":["cs.CY","cs.AI","cs.PL"],"primary_category":"cs.CY","announce_type":"cross","date":"2026-08-11","first_seen":"2026-08-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.07523","pdf_url":"https://arxiv.org/pdf/2608.07523","source_feed":"cs.AI","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM评估辅助","试题难度校准","编程教育"],"reason":"用LLM作为评估辅助工具，替代人工判断试题难度，属于标注替代而非仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-08-11T13:04:48","error":null,"has_summary":false,"summary":null},{"id":"2608.07488","version":1,"title":"Large Language Models Explain Experts Better Than Experts Themselves","zh_title":"大语言模型比专家自己更能解释专家","abstract":"Tacit knowledge, or the \"know-how\" embedded in experience, is difficult to articulate, making its transfer a challenge in organizations. Tacit knowledge is hard to externalize (transform into explicit knowledge), and expertise is often poorly documented and lost when experts leave. This study examines whether LLMs can externalize tacit knowledge from experts' behaviors and whether such externalized knowledge supports downstream decision-making and transfer to novices. Across two studies, we show that LLM-externalized tacit knowledge improves decision quality and enables novices to approach expert-level performance, often outperforming knowledge articulated by human experts. These findings provide empirical support for Polanyi's Paradox -- that we can know more than we can tell -- and highlight the potential of LLMs as scalable tools that can help overcome human experts' articulation bottleneck. Mechanism analyses and robustness checks show that LLMs meaningfully learn and extract knowledge from expert conversations, and findings generalize across models and retrieval methods.","authors":["Mina Cho","Russell J. Funk","Alok Gupta","Mochen Yang"],"categories":["cs.HC","cs.CY"],"primary_category":"cs.HC","announce_type":"new","date":"2026-08-11","first_seen":"2026-08-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.07488","pdf_url":"https://arxiv.org/pdf/2608.07488","source_feed":"cs.HC","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["隐性知识外化","LLM 知识提取","决策支持"],"reason":"LLM 从专家行为中提取隐性知识并辅助决策，属于替代人类知识外化而非仿真被试，…","model":"deepseek-v4-pro","scored_at":"2026-08-11T13:04:46","error":null,"has_summary":false,"summary":null},{"id":"2608.07496","version":1,"title":"Human-Simulation Interaction: From Prediction to Exploration in LLM Agent Simulations for Policy","zh_title":"人-仿真交互：从预测到探索的LLM智能体政策模拟","abstract":"Agent-based models have historically served as tools for generative explanation, constructing testbeds in which candidate micro-level behavioral rules can be tested for their capacity to produce observed macro-level phenomena. The integration of Large Language Models into agent-based simulation has expanded what these models can represent, but it has also introduced an unexamined shift in how users engage them. We argue that current generative agent-based models (GABMs) inherit the dominant interaction metaphor of conversational LLM interfaces - a question-answer pattern that positions users as consumers of system output rather than explorers of a possibility space. In the context of policy, where problems are wicked and ground truth is unknowable in advance, this metaphor produces a trust deficit that cannot be resolved through improved model accuracy alone. We open a design space we call human-simulation interaction, and argue that warranted trust requires interaction metaphors that restore the exploratory capacity simulation has historically supported.","authors":["Huanxing Chen"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-08-11","first_seen":"2026-08-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.07496","pdf_url":"https://arxiv.org/pdf/2608.07496","source_feed":"cs.HC","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["LLM智能体","社会模拟","政策评估"],"reason":"讨论LLM智能体模拟政策，但无真实人类数据对照，属社会模拟边界情形。","model":"deepseek-v4-pro","scored_at":"2026-08-11T13:04:31","error":null,"has_summary":false,"summary":null},{"id":"2608.08497","version":1,"title":"SocialFiVis: A Visual Analytics Sandbox for LLM-Grounded Multi-Agent Simulation in Social Finance","zh_title":"SocialFiVis：面向社交金融中基于LLM的多智能体仿真的可视分析沙盒","abstract":"The emergence of social finance (SocialFi) transforms online communities into complex socio-economic systems. Within these spaces, collective decisions shape a \"digital commons\" characterized by social capital (e.g., community trust) and financial health (e.g., market liquidity). Governing such hybrid ecosystems is challenging because real-world interventions are costly and irreversible. While counterfactual simulation is essential for exploring alternative governance strategies, existing approaches fail to capture the non-linear interplay between governance rules, individual behaviors, and emergent economic outcomes. To systematically unpack this complexity, we operationalize the Institutional Analysis and Development (IAD) framework as our theoretical foundation, synthesizing prior literature with insights from formative expert interviews. Built on this framework, we present SocialFiVis, an IAD-embedded visual analytics sandbox. It introduces a robust model to quantify the dual-track digital commons, coupled with a two-phase simulation engine. This engine combines LLM-derived personas with a mechanism-guided Perception-Reasoning-Action (PRA) runtime to simulate heterogeneous, context-aware agents empirically grounded in the retained messaging cohort. A hierarchical multi-view interface with interpretable reasoning pathways enables community operators to explore counterfactual policies and trace system-level outcomes back to individual behavioral rationales. We evaluate SocialFiVis through two case studies, a user study, and follow-up interviews. Results demonstrate that SocialFiVis supports fine-grained behavioral attribution and helps explain emergent phenomena such as the structural decoupling of social capital and the resilience of messaging members under localized governance shocks.","authors":["Yi-Fan Cao","Qing Shi","Liangwei Wang","Leo Yu-Ho Lo","Lin Chen","Yuzi Han","Yang Wang","Kani Chen"],"categories":["cs.HC","cs.MA","cs.SI"],"primary_category":"cs.HC","announce_type":"new","date":"2026-08-11","first_seen":"2026-08-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.08497","pdf_url":"https://arxiv.org/pdf/2608.08497","source_feed":"cs.HC","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["多智能体仿真","社交金融","可视分析"],"reason":"用LLM agent模拟社会金融系统，但无真实人类数据对照，属社会模拟边界情形。","model":"deepseek-v4-pro","scored_at":"2026-08-11T13:04:36","error":null,"has_summary":false,"summary":null},{"id":"2608.07511","version":1,"title":"How sensitive do we want AI to be? Socio-communicative competencies of large language models in healthcare","zh_title":"我们希望AI有多敏感？医疗保健中大语言模型的社会沟通能力","abstract":"Background. Effective clinical practice relies heavily on the socio-communicative skills of medical professionals. Large language models (LLMs) have been proposed for tasks such as triaging patients, report drafting or translating medical jargon to support informed decision-making. These applications require both factual and social competence. This study evaluates dialogues between LLMs and participants to assess the current state of socio-communicative competencies displayed in LLM-generated texts. Methods. We extracted a subset of extended dialogues from the HELP-Med dataset, comprising 1800 conversation transcripts of interactions between human participants seeking medical information and three different LLMs, GPT 4o, Llama 3 and Command R+. Two experts coded the transcripts for demonstrations of socio-communicative behaviours (non-hostility, sensitivity, structuring, non-intrusiveness) using the IC-MD instrument, originally designed to evaluate interactional competencies in medical student admissions. Results. The LLMs in our study showed strength in non-hostility, mixed results in sensitivity and non-intrusiveness and performed poorly in structuring. Conclusion. Current LLMs lack the consistent and reliable socio-communicative skills needed for safe and effective use as healthcare advisors. While existing frameworks for assessing interactional competencies may support the development of more socially responsive LLMs, they will require adaptation to account for the differences in desirable behaviour between humans and LLMs.","authors":["Dorothee Amelung","Andrew M. Bean","Sabine C. Herpertz","Felix H. Krones","Guy Parsons","Adam Mahdi","Isabella Schneider"],"categories":["cs.HC","cs.AI","cs.CL","cs.CY"],"primary_category":"cs.HC","announce_type":"cross","date":"2026-08-11","first_seen":"2026-08-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.07511","pdf_url":"https://arxiv.org/pdf/2608.07511","source_feed":"cs.CL","score":3,"bucket":"other","rubric_hits":["C3"],"tags":["LLM社交能力","医疗对话","角色扮演评测"],"reason":"评估LLM对话中的社交沟通能力，属于角色扮演对话评测，无人类行为仿真对照。","model":"deepseek-v4-pro","scored_at":"2026-08-11T13:04:47","error":null,"has_summary":false,"summary":null},{"id":"2608.07495","version":1,"title":"EmoPatient: An Emotion-Directed Patient Simulator for Realistic Palliative Care Communication Training","zh_title":"EmoPatient：面向情感引导的姑息治疗沟通训练患者模拟器","abstract":"Effective communication during palliative care discussions is a critical clinical skill, yet training clinicians to manage complex patient emotions remains challenging. Large language model (LLM)-based patient simulators provide a scalable approach for communication training, but most existing systems treat patient emotion as static and fail to capture the dynamic emotional shifts observed in clinical interactions. We present EmoPatient, an emotion-directed patient simulator designed to generate evolving emotional responses during palliative care discussions. The system introduces an Emotion Director agent that estimates the patient's emotional state and generates turn-level control signals for emotional intensity, regulatory stability, and interactional guidance. We evaluate EmoPatient through controlled multi-turn physician-patient dialogue simulations and compare it with baseline simulators. Results show improvements across four theory-informed emotional realism metrics and robustness across conversational personality variants, suggesting that modeling emotional dynamics can improve the realism of LLM-based patient simulators for palliative care communication training.","authors":["Yining Wu","Tianshu Du","Jinrui Fang","Chi Zhang","Sonal Admane","Ying Ding"],"categories":["cs.HC","cs.AI"],"primary_category":"cs.HC","announce_type":"cross","date":"2026-08-11","first_seen":"2026-08-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.07495","pdf_url":"https://arxiv.org/pdf/2608.07495","source_feed":"cs.AI","score":3,"bucket":"other","rubric_hits":["C3"],"tags":["患者模拟","沟通训练","情感动态"],"reason":"角色扮演对话训练，无实验或测量目的，不涉及人类行为对照","model":"deepseek-v4-pro","scored_at":"2026-08-11T13:04:46","error":null,"has_summary":false,"summary":null},{"id":"2607.23621","version":2,"title":"GEMCo: A Validated, Ethically Releasable Proxy for Inaccessible Counselling Data","zh_title":"GEMCo：一个经过验证、可伦理发布的不可访问咨询数据代理","abstract":"This paper presents GEMCo, a releasable, human-written proxy for inaccessible counselling data: 86 complete German e-mail counselling conversations (728 messages), expert-authored cases and counsellor sessions with trained role-players. It is validated against a held-out reference of 124 real counselling conversations. The proxy and the real conversations are measured against each other on counsellor strategies and client emotions. The gap is detectable but small. A generative validation supports the analysis. The corpus is the primary contribution. The validation method generalises to any domain where real data cannot be shared but a human-made proxy can. Privacy and ethics keep real counselling data closed. GEMCo carries none by design and can be released, a first step toward language research in this domain.","authors":["Philipp Steigerwald","Eric Rudolph","Mara Stieler","Jennifer Burghardt","Robert Lehmann","Jens Albrecht"],"categories":["cs.CL","cs.CY"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-08-11","first_seen":"2026-07-28","revised_at":"2026-08-11","abs_url":"https://arxiv.org/abs/2607.23621","pdf_url":"https://arxiv.org/pdf/2607.23621","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["数据代理","咨询对话","隐私保护"],"reason":"角色扮演对话生成代理数据，无LLM仿真人类被试或实验测量目的","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:18:16","error":null,"has_summary":false,"summary":null},{"id":"2607.23065","version":2,"title":"Touching or Chatting: The Utility of LLMs and Tactile Charts for Learning about Complex Chart Types by BLV Individuals","zh_title":"触摸还是聊天：大语言模型与触觉图表对盲人及低视力者学习复杂图表类型的效用","abstract":"Visualizations are central to communicating data, yet blind and low-vision (BLV) people often lack support for understanding chart types---knowledge that is essential for interpreting new visualizations and collaborating with sighted peers. Prior work found that BLV individuals viewed example tactile charts as more helpful than text-only approaches and preferred them for learning advanced chart types, particularly for understanding spatial layouts and shapes. Meanwhile, large language models (LLMs) are increasingly used by BLV individuals for chart explanation and question answering (QA), but have been studied primarily for dataset exploration rather than chart-type learning. Existing LLM-based chart QA also shows that users frequently ask about layout and structure, yet struggle with spatial concepts and misdirect questions when mental models are weak. We investigate how LLMs influence chart-type learning and whether tactile learning improves subsequent LLM-supported exploration. We extend our tactile chart learning tools with an LLM chatbot that provides interactive explanations and supports follow-up questions. In an interview study with 12 BLV participants, we compare two learning formats: (1) a tactile chart, a textual explanation, and an LLM chatbot; and (2) a textual explanation and an LLM chatbot. The learning phase was followed by exploration of an unfamiliar dataset using alt text and an LLM. Thematic analysis shows that tactile templates support BLV participants' formation of chart-type mental models, which scaffolds subsequent LLM-mediated data exploration. Text+LLM explanations without tactile support show weaknesses for spatial-reasoning tasks.","authors":["Tingying He","Maggie McCracken","Daniel Hajas","Sarah Creem-Regehr","Alexander Lex"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"replace","date":"2026-08-11","first_seen":"2026-07-28","revised_at":"2026-08-11","abs_url":"https://arxiv.org/abs/2607.23065","pdf_url":"https://arxiv.org/pdf/2607.23065","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["辅助技术","人机交互","盲人教育"],"reason":"研究LLM辅助盲人学习图表，属辅助技术，非人类仿真实验。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:18:11","error":null,"has_summary":false,"summary":null},{"id":"2607.23705","version":2,"title":"Social learning drives underprioritization of collective challenges","zh_title":"社会学习导致集体挑战的低优先级","abstract":"Societies often struggle to prioritize important challenges in a timely manner, with substantial costs from delayed action on issues like climate change and pandemic mitigation. A persistent puzzle is that broad concern on issues often fails to translate into collective priority. We argue that a key driver lies in how concern is formed across competing issue domains. Some issues depend heavily on social learning, where individuals infer importance from others, often because direct experience is limited. Others depend more on individual learning from firsthand experience. We develop a dynamic model in which two subgroups form issue-specific concerns through individual and social learning, and these concerns are aggregated into collective priority. The model yields three insights. First, with two issues of equal objective severity, the one that depends more on social learning tends to be underprioritized when both issues are severe. Second, gradual increases in severity delay reprioritization of the issue, with the delay growing as reliance on social learning increases. Third, this bias can be reduced by reducing social learning or by increasing intergroup learning beyond a critical threshold. These results offer a general mechanism for why severe problems can remain neglected in collective action despite widespread concern, and why intergroup interaction or experiential simulations may help align collective priorities with objective risks.","authors":["Russ Yoon","Vicky Chuqiao Yang"],"categories":["physics.soc-ph","math.DS"],"primary_category":"physics.soc-ph","announce_type":"replace","date":"2026-08-11","first_seen":"2026-07-28","revised_at":"2026-08-11","abs_url":"https://arxiv.org/abs/2607.23705","pdf_url":"https://arxiv.org/pdf/2607.23705","source_feed":"physics.soc-ph","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["社会学习","集体行动","动态模型"],"reason":"纯多智能体社会学习模型，无LLM，无人类数据对照，不涉及人类仿真。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:20","error":null,"has_summary":false,"summary":null},{"id":"2607.26977","version":2,"title":"TREK: A Travel Reasoning and Evaluation Kit for LLM Agents in Complex Trip Planning","zh_title":"TREK：面向复杂旅行规划的LLM智能体旅行推理与评估套件","abstract":"Travel planning is a demanding stress test for tool-using LLM agents: a usable itinerary is a single artifact that must be right along many axes at once - every flight, hotel, and attraction must exist and be bookable, the days must be physically traversable, the total must clear a budget, and the plan must serve a traveler whose needs are only partly stated. Existing agent benchmarks reward these properties one at a time and grade the final output with soft or LLM-judged rubrics, which cannot certify that a returned plan is executable and are neither reproducible nor auditable. We introduce TREK (Travel Reasoning and Evaluation Kit), a benchmark for feasible itinerary synthesis: producing a single plan that is jointly constraint-correct, hallucination-free, spatio-temporally executable, budget-valid, and responsive to the traveler's unstated persona needs. TREK comprises 800 multi-constraint tasks - 533 feasible and 267 provably infeasible with typed route/entity/budget causes - over a synthetic, internally consistent knowledge base of 212,530 records across 375 cities and 13 personas, served through a production-style tool sandbox of validated RESTful APIs. Every task is scored by a fully deterministic, rule-based evaluator with no LLM judge and ships a human-verified gold reference that scores a perfect 1.0 under that same evaluator, so the ceiling is demonstrably achievable and every remaining gap is an agent limitation rather than scorer strictness. Evaluating 15 LLM agents across nine constraint dimensions, we find that even the strongest (GPT-5.6) produces a fully-feasible plan on only 46.2% of solvable tasks, with a median of 6.6% and a floor of 0.0%; satisfying travelers' unstated needs emerges as the universal bottleneck, unsolved even at the frontier. We release the dataset, tool sandbox, deterministic evaluator, and agent code as a fully reproducible benchmark.","authors":["Jinhu Qi","Wentao Zhang","Siu Man Ng","Feiyang Xu","Yanyu Chen","Yaoman Li","Irwin King"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-08-11","first_seen":"2026-07-30","revised_at":"2026-08-11","abs_url":"https://arxiv.org/abs/2607.26977","pdf_url":"https://arxiv.org/pdf/2607.26977","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["LLM智能体","旅行规划","基准评测"],"reason":"纯多智能体工具使用评测，无人类行为对照，不涉及仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-07-30T13:01:59","error":null,"has_summary":false,"summary":null},{"id":"2607.28650","version":2,"title":"Unanticipated Effects of Generative AI on Expertise Pathways and Performance Perception in System Administration","zh_title":"生成式AI对系统管理专业知识路径与绩效感知的意外影响","abstract":"While industry discourse often emphasizes immediate productivity gains and frames GenAI primarily as a tool for automation, the integration of GenAI into system administration may involve deeper shifts in professional practice that are not yet fully understood. Drawing on 14 semi-structured interviews with IT professionals, this paper explores the lived reality of embedding GenAI into daily routines of troubleshooting, scripting, and system verification. Through inductive thematic analysis, we uncover two unanticipated socio-technical findings. First, we describe a \"compression of traditional expertise pathways\" where GenAI appears to function as both a mentor-like tutor and a \"ladder-shortening\" tool. While the tool can support faster task performance in unfamiliar domains, our findings suggest it may also reduce a practitioner's exposure to the foundational, hands-on cycles of building, failing, and debugging that historically served as the training ground for technical expertise. Second, we describe a \"performance perception shift,\" where the speed of AI-assisted work begins to reset organizational and self-expectations for productivity. This shift may create a \"two-speed culture\" within teams and introduce \"productivity guilt,\" as necessary manual work, even when required for safety or validation, is increasingly perceived as slow or a failure of efficiency. Our results raise broader questions about how GenAI may influence expertise development, how professional value is assessed in high-stakes technical environments, and the role of human judgment in complex technical environments.","authors":["Rana Abou Khamis","Hala Assal","Ashraf Matrawy"],"categories":["cs.HC","cs.AI"],"primary_category":"cs.HC","announce_type":"replace-cross","date":"2026-08-11","first_seen":"2026-08-03","revised_at":"2026-08-11","abs_url":"https://arxiv.org/abs/2607.28650","pdf_url":"https://arxiv.org/pdf/2607.28650","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["人机交互","质性研究","专业知识发展"],"reason":"研究GenAI对IT专家工作实践的影响，属人机交互质性研究，非LLM仿真人类被…","model":"deepseek-v4-pro","scored_at":"2026-08-12T13:02:51","error":null,"has_summary":false,"summary":null},{"id":"2608.01176","version":2,"title":"When Words Divide: Diachronic Ideological Polarization in Political Discourse on Social Media","zh_title":"当词语分裂：社交媒体政治话语中的历时意识形态极化","abstract":"Political polarization has become a defining feature of online discourse, yet its long-term evolution remains poorly understood. We present a longitudinal analysis of ideological polarization in Reddit discussions by measuring semantic differences in the language used by opposing political communities. We construct temporally aligned community-specific word embeddings and quantify ideological polarization as the semantic divergence of political concepts over time. Our analysis shows that ideological polarization has increased substantially during the study period, both at the concept- and topic-level. Unlike prior computational work, which has largely focused on cross-sectional analyses or affective dimensions of polarization at a single point at time, our approach captures the evolution of ideological differences in semantic framing. The proposed framework provides a scalable method for studying the temporal dynamics of ideological polarization in large-scale social media discourse.","authors":["Roy Yitzchak","Noa Lavie","Ella Rabinovich"],"categories":["cs.CL","cs.SI"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-08-11","first_seen":"2026-08-04","revised_at":"2026-08-11","abs_url":"https://arxiv.org/abs/2608.01176","pdf_url":"https://arxiv.org/pdf/2608.01176","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["意识形态极化","社交媒体分析","词嵌入"],"reason":"分析人类政治话语的语义极化，未使用LLM仿真人类被试，不涉及人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-08-11T13:04:43","error":null,"has_summary":false,"summary":null},{"id":"2608.02311","version":2,"title":"AI Governance for Institutional Readiness in Finance","zh_title":"面向金融领域机构就绪度的人工智能治理","abstract":"Agentic AI is gaining acceptance in asset management, but governance has not kept pace: 88\\% of surveyed finance professionals report no operational governance framework for agentic AI, and only 24 of 75 large U.S. money managers disclosing AI use in Form ADV filings report a formal governance policy. We argue this gap is architectural: governance built for static validation does not survive continuously retrained agentic policies. We propose a four-layer framework (Policy, Engineering, Composition, Systemic) grounded in two distinct kinds of evidence, kept explicitly separate: two calibrated synthetic illustrations (a regret-covariance drift monitor; a crowding simulation showing joint drawdown risk rising from 39.2\\% to 79.3\\%), and three real, documented cases (a deployed LLM-embedding trading strategy, a \\$45 billion discretionary fund's forced-deleveraging blowup, and a tribunal ruling holding an airline liable for its chatbot). The synthetic examples demonstrate computability from observable data; the cases demonstrate that the failure modes are not hypothetical. We provide a 90-day implementation sequence spanning trading and payments/customer-facing systems.","authors":["Irene Aldridge","Steve Krawciw"],"categories":["econ.EM","q-fin.RM","q-fin.ST"],"primary_category":"econ.EM","announce_type":"replace","date":"2026-08-11","first_seen":"2026-08-04","revised_at":"2026-08-11","abs_url":"https://arxiv.org/abs/2608.02311","pdf_url":"https://arxiv.org/pdf/2608.02311","source_feed":"econ.EM","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["AI治理","金融风险管理","多智能体系统"],"reason":"纯多智能体系统研究，agent 协作解题，不涉及人类行为对照","model":"deepseek-v4-pro","scored_at":"2026-08-11T13:04:43","error":null,"has_summary":false,"summary":null},{"id":"2608.05656","version":2,"title":"Studying People to Study AI: Expert Perspectives on the Epistemic Fit and Barriers of Human Research in AI Safety & Ethics","zh_title":"通过研究人来研究AI：专家对AI安全与伦理中人类研究方法的认知契合与障碍的看法","abstract":"Safety risks of AI are becoming increasingly evident in human interactions with AI technologies. The prominent approaches to evaluating these risks favor technical methods, such as model benchmarks and LLM simulations, often sidelining empirical research with human subjects. To examine this apparent gap in the acceptance of human research, we conduct an expert survey (n=93) and expert interviews (n=17) with AI Safety & Ethics (AISE) researchers from Technical, Sociotechnical, Governance, and Normative backgrounds. Our findings suggest that although there is a consensus that human research is valuable for generating evidence for AISE, its adoption and acceptance are constrained by perceived validity issues, tangible resource barriers, epistemic and personal preferences in methods, and infrastructural constraints from the broader research community. In particular, Technical researchers tend to value human research less and collaborate across disciplines less, suggesting an epistemic tension towards human methods. We propose recommendations for establishing the epistemic fit of human research within AISE and bridging the prohibitive limitations that researchers face, while avoiding performative 'human-washing'.","authors":["Jessica Y. Bo","Paula Akemi Aoyagui","Shalaleh Rismani","Dipto Das","Syed Ishtiaque Ahmed","Ashton Anderson"],"categories":["cs.CY","cs.AI","cs.HC"],"primary_category":"cs.CY","announce_type":"replace-cross","date":"2026-08-11","first_seen":"2026-08-07","revised_at":"2026-08-11","abs_url":"https://arxiv.org/abs/2608.05656","pdf_url":"https://arxiv.org/pdf/2608.05656","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["AI安全","专家调查","方法论"],"reason":"研究人类专家对AI安全方法的看法，不涉及LLM仿真人类被试","model":"deepseek-v4-pro","scored_at":"2026-08-11T13:05:24","error":null,"has_summary":false,"summary":null},{"id":"2608.06609","version":2,"title":"Automated item evaluation: Predicting item acceptance and rejection using LLM-generated critiques","zh_title":"自动化试题评估：利用LLM生成的评语预测试题接受与拒绝","abstract":"Automated item evaluation (AIE) refers to the use of computational methods to assess item quality without requiring manual expert review or field testing of the items under evaluation. We aimed to build a near-comprehensive AIE model by predicting item acceptance and rejection from item text using historical rejection data from a large-scale standardized testing program. The dataset contained 52,759 English language arts (ELA) and mathematics items with 34% permanently rejected from future operational use. Rejection reasons included poor psychometric properties, content issues, bias and sensitivity concerns, and non-content issues. We fine-tuned a DeBERTaV3-large classifier on raw item text, a second DeBERTa classifier on Qwen3-generated item critiques, and a fusion model combining representations from both. The fusion model achieved the strongest overall performance (Accuracy = .75, F1 = .64, AUC = .80, Sensitivity = .64, Specificity = .81). Prediction for math (F1 = .73, AUC = .86) was considerably more accurate than ELA (F1 = .51, AUC = .72). Lowering the decision threshold from .5 to .25 raised average sensitivity for ELA and math to .88 and .91, while reducing specificity to .31 and .56, respectively, which may be preferable in automated item generation contexts where generating items is cheaper than evaluating them. Incorporating item critiques alongside raw item text improved performance across most rejection reasons. The model assigned higher rejection probabilities to more difficult items. However, the fusion model struggled to identify items flagged for bias, sensitivity, fairness, or accessibility, especially for ELA. These findings suggest that text-based AIE is feasible in some areas and may offer a practical tool for reducing the burden of manual review and field testing, while also underscoring the importance of human review for items with fairness concerns.","authors":["Hotaka Maeda","Yikai Lu"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"replace","date":"2026-08-11","first_seen":"2026-08-10","revised_at":"2026-08-11","abs_url":"https://arxiv.org/abs/2608.06609","pdf_url":"https://arxiv.org/pdf/2608.06609","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["自动化试题评估","NLP分类","教育测量"],"reason":"纯NLP分类任务，用LLM生成评语辅助预测试题接受/拒绝，不涉及人类行为仿真或…","model":"deepseek-v4-pro","scored_at":"2026-08-10T13:01:32","error":null,"has_summary":false,"summary":null},{"id":"2608.08160","version":1,"title":"Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives","zh_title":"LLM智能体能按剧本走吗？交互叙事中长期一致性基准测试","abstract":"The rapid advancement of Large Language Models (LLMs) is revolutionizing AI for Games by enabling open-ended and fluid interactive storytelling. However, existing research has largely overlooked the critical challenge of maintaining long-horizon logical consistency and narrative integrity against unconstrained user interventions. To address this, we formulate this challenge as Narrative Commitment Preservation (NCP), and take interactive narrative as our testbed. We introduce NCP-Bench, a benchmark of 100 narrative environments derived from movie synopses. Each environment includes a structured narrative specification (trajectory, commitments, and initial facts) that we can automatically check throughout the interaction between the player agent and the narrator agent. Experiments across state-of-the-art LLMs reveal a substantial long-horizon consistency gap: high linguistic quality does not guarantee commitment preservation; even strong models frequently generate logically conflicting content under adversarial interventions, with the best-performing model (GPT-5.2) achieving only 42% survival rate after 20 turns and fact conflict rates ranging from 40% to 68% across models, and only isolated runs satisfying all achievement commitments within the 100-turn limit.","authors":["Yingpeng Ma","Jianhao Yan","Bei Shi","Ka Hou Kam","Runnan Wang","Xuebo Liu","Yulong Chen","Yue Zhang","Derek F. Wong"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-11","first_seen":"2026-08-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.08160","pdf_url":"https://arxiv.org/pdf/2608.08160","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["交互叙事","一致性评测","角色扮演"],"reason":"评估LLM在交互叙事中保持逻辑一致性的能力，属于角色扮演对话评测，无实验或测量…","model":"deepseek-v4-pro","scored_at":"2026-08-11T13:04:58","error":null,"has_summary":false,"summary":null},{"id":"2608.08164","version":1,"title":"STEMMA: An Adversarial Multi-Agent Framework for Evaluating Self-Identity Consistency in LLMs","zh_title":"STEMMA：用于评估大语言模型自我身份一致性的对抗性多智能体框架","abstract":"Knowledge Distillation is a widely adopted technique in the training and fine-tuning of large language models (LLMs) enabling transfer of structured information and functional behavior from a large teacher model to a smaller student model while significantly reducing computational costs. However, as the use of distillation increases in both scale and complexity it raises an important question about what kind of knowledge is really transferred from the teacher model. In this work, we argue that apart from the functional knowledge, student models also learn behavioral patterns, specifically how a model represents its own identity raising concerns about output homogeneity, model biases, and accountability. To address this challenge, we introduce STEMMA, a multi-modal and multi-agent framework in which role specific agents collaboratively probe self identification behavior in different models. We also contribute a set of adversarial prompts designed manually to evaluate identity consistency in LLMs. Our results show that to an extent most models are vulnerable to inconsistencies in self-representations.","authors":["Nuthakki Siva Gopala Krishna","Kanishka Jain"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-11","first_seen":"2026-08-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.08164","pdf_url":"https://arxiv.org/pdf/2608.08164","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","模型身份一致性","对抗性提示"],"reason":"多智能体框架用于探测模型自身身份一致性，不涉及人类行为对照或仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-08-11T13:04:58","error":null,"has_summary":false,"summary":null},{"id":"2608.08606","version":1,"title":"Mitigating Gender Bias in English to Romanian Machine Translation","zh_title":"缓解英罗机器翻译中的性别偏见","abstract":"Machine translation (MT) systems often fail to correctly translate gender, especially when converting from a gender-neutral language like English to a gendered target language such as Romanian. This bias results in translations that default to masculine forms or reinforce gender stereotypes. We propose a hybrid pipeline to mitigate this issue by combining large language model (LLM)-based gender classification with neural machine translation (NMT). Our system uses a fine-tuned LLM to detect the intended gender of target words in English sentences and insert inline gender hint tags. These tagged sentences are then passed to a Transformer model fine-tuned to generate morphologically correct Romanian translations. To support this, we introduce three novel datasets for gender disambiguation and translation. Our approach improves gender accuracy on the WinoMT and WinoGender benchmarks by over 40 percentage points compared to a baseline MT system. This is the first method to explicitly address and evaluate gender bias in English-Romanian MT using both LLM inference and tag-aware translation.","authors":["Ioana Grigore","Sergiu Nisioi"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-11","first_seen":"2026-08-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.08606","pdf_url":"https://arxiv.org/pdf/2608.08606","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["机器翻译","性别偏见","NLP评测"],"reason":"论文解决机器翻译中的性别偏见，属于NLP能力评测，不以人类行为仿真为目标。","model":"deepseek-v4-pro","scored_at":"2026-08-12T13:02:54","error":null,"has_summary":false,"summary":null},{"id":"2608.08975","version":1,"title":"How Can Rhetoric Reward-Hack AI Reviewers? Dissecting Rhetorical Sensitivity in AI-Based Peer Review","zh_title":"修辞如何奖励黑客AI审稿人？剖析AI同行评审中的修辞敏感性","abstract":"As large language models increasingly participate in scientific evaluation, we investigate a potential form of reward hacking: how rhetorical choices shape AI-review judgments when reported scientific content is preserved and how these effects vary across evaluation conditions. We construct a controlled corpus of 4,200 full-paper manuscripts derived from 120 anonymized ICLR 2026 submissions. Two LLM rewriters transform six rhetorical dimensions in opposing directions, and five LLM reviewers evaluate the resulting manuscripts under standard and strict protocols. We also test joint, recursive, and reviewer-guided rewriting. Our results show that rhetorical sensitivity is structured rather than uniform. Evidence framing and novelty stance produce the largest positive-negative contrasts in overall assessment, with scope framing forming a weaker second tier; the remaining dimensions have smaller or less stable effects. This hierarchy persists across human-assessed quality levels, but score movement depends strongly on the AI reviewer's original score: lower scores tend to rise, higher scores tend to fall, and directional contrasts are clearest in the middle ranges. More elaborate workflows do not reliably yield larger gains. Joint rewriting is strongly rewriter-dependent, reviewer guidance does not consistently outperform an unguided second pass, and repeated rewriting yields diminishing, configuration-dependent returns. Across conditions, the rewriter primarily determines the separation between opposing variants, whereas the reviewer determines the magnitude and sign of their score effects. Strict review lowers mean OA by 1.36 points without consistently changing rhetorical sensitivity. These findings identify when rhetorical presentation influences AI scientific review and motivate evaluation systems robust to content-preserving variation in scientific writing.","authors":["Ming Li","Chenguang Wang","Xirui Li","Xinyue Zeng","Dianqi Li","Peng Shi","Dawei Zhou","Tianyi Zhou"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-11","first_seen":"2026-08-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.08975","pdf_url":"https://arxiv.org/pdf/2608.08975","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["AI审稿","修辞敏感性","奖励黑客"],"reason":"研究LLM审稿中的修辞敏感性，属于NLP评测，不以人类行为仿真为目标。","model":"deepseek-v4-pro","scored_at":"2026-08-11T13:05:10","error":null,"has_summary":false,"summary":null},{"id":"2608.09080","version":1,"title":"When Confidence Fails: Overconfidence in LLMs under Uncertainty and Missing Clinical Information","zh_title":"当置信度失效：不确定性及缺失临床信息下大语言模型的过度自信","abstract":"Large Language Models (LLMs) have achieved strong performance in medical question answering and clinical reasoning tasks. However, their reliability under uncertainty remains poorly understood which raises critical concerns for deployment in high-stakes clinical settings. In such environments, incorrect predictions are inherently risky, but confident incorrect predictions can be particularly harmful as they may mislead clinical decision-making. In this paper, we conduct a systematic behavioral analysis of LLMs under clinical information uncertainty. We propose an evaluation framework based on the MedMCQA dataset consisting of two complementary uncertainty settings. First, we introduce linguistic uncertainty cues through prompt modifications to simulate ambiguous clinical contexts. Second, we construct an answer removal setting, wherein the correct option is deliberately excluded mandating the model to recognize insufficient information and abstain. We analyze both model accuracy and confidence behavior using multiple calibration metrics including calibration gap, Expected Calibration Error (ECE), and Unsafe Confident Error Rate (UCER) across 500 medical questions. Our results reveal a consistent failure mode, i.e., although accuracy degrades under increasing uncertainty, model confidence remains misaligned with accuracy. This leads to a substantial increase in unsafe confident errors, indicating that model confidence remains largely insensitive to clinically meaningful information loss. Furthermore, we observe significant variation across models in their ability to abstain when the correct answer is unavailable, with some models persistently producing high confidence hallucinated answers. These findings expose critical limitations in the epistemic reliability of current LLMs and highlight the need for uncertainty aware evaluation methods prior to their deployment in clinical workflows.","authors":["Maryam Tahermazandarani","Adnan Mahmood","Fahmida Islam","Quan Z. Sheng"],"categories":["cs.CL","cs.AI","cs.HC","cs.LG"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-11","first_seen":"2026-08-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.09080","pdf_url":"https://arxiv.org/pdf/2608.09080","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["置信度校准","医学问答","模型可靠性"],"reason":"评估LLM在医学问答中的置信度校准，属于纯NLP能力评测，不以人类行为仿真为参…","model":"deepseek-v4-pro","scored_at":"2026-08-11T13:05:12","error":null,"has_summary":false,"summary":null},{"id":"2608.09128","version":1,"title":"Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments","zh_title":"Social Gym与SPaRTan：通过多智能体游戏锦标赛基准测试与提升大语言模型社交推理能力","abstract":"LLM agents are increasingly deployed in multi-agent social settings where they must cooperate, negotiate, and adapt to other agents. Measuring and improving these social skills is hard because, unlike math or logic, social interaction offers no objective ground truth: evaluations fall back on LLM judges, which are costly, subjective, and noisy, and models get no reliable signal to learn from. To address both, we first introduce Social Gym, an environment of 21 multi-agent social games (e.g., Werewolves, Resistance, Spyfall) whose rule-decided outcomes make agent performance verifiable and objective, with an Elo tournament that produces a cross-game leaderboard. Benchmarking experiments show that while GPT-5-mini tops the leaderboard, no model excels at all games uniformly or in all game roles, pointing to limitations of social reasoning. Motivated by this, we additionally propose SPaRTan (Self-Play and Reflect-Transfer), a training-free self-improvement loop: a model plays a game, reflects on its trajectories and their outcomes to produce a transferable playbook, and applies that playbook in subsequent games. Our results show that SPaRTan playbooks help GPT-5-mini agents level their performance on weaker roles, but largely do not improve Qwen3-32B's performance. Together, Social Gym and SPaRTan offer a reproducible, verifiable foundation for measuring and improving LLM social reasoning without weight updates.","authors":["Keyu He","Xuhui Zhou","Maarten Sap"],"categories":["cs.CL","cs.AI","cs.MA"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-11","first_seen":"2026-08-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.09128","pdf_url":"https://arxiv.org/pdf/2608.09128","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体","社交推理","游戏基准"],"reason":"纯多智能体游戏竞技，无人类行为对照，属C1排除项","model":"deepseek-v4-pro","scored_at":"2026-08-11T13:05:12","error":null,"has_summary":false,"summary":null},{"id":"2608.09189","version":1,"title":"EmoS: A Theory-Grounded Framework for Evaluating and Aligning Emotional Intelligence in Spoken Language Models","zh_title":"EmoS：一个基于理论的框架，用于评估和对齐口语语言模型中的情商","abstract":"Despite significant advances in instruction-following and auditory comprehension, the evaluation of Emotional Intelligence (EI) in Spoken Language Models (SLMs) remains confined to rudimentary paralinguistic perception, lacking a systematic, theory-driven cognitive framework. We introduce EmoSBench, the first comprehensive EI evaluation benchmark for SLMs constructed upon the four-branch theoretical model, covering Perceiving, Understanding, Using, and Managing Emotion across ten sub-tasks. Preliminary assessments on EmoSBench reveal a substantial gap: even leading proprietary models like GPT-4o-Audio achieve only 52.6%, significantly trailing human baselines. To bridge this gap, we develop EmoS, a specialized evaluator model optimized via Supervised Fine-Tuning (SFT) and Group Relative Policy Optimization (GRPO). To facilitate its effective training, we curate EmoDialogue, a bilingual dataset providing necessary fine-grained supervision through response pairs with rigorously defined EI gradations. Concurrently, we introduce a reward mechanism integrating a Steep Exponential Accuracy Reward (SEAR) and a Rationale Fidelity Reward (RFR) to enforce precise ordinal scoring and valid reasoning. Experiments demonstrate that EmoS reaches 83.8% accuracy, approaching human-level performance. Furthermore, evaluations on authentic, unconstrained spoken interactions validate its robust real-world generalization, establishing a foundational framework for advancing emotionally intelligent dialogue systems.","authors":["Junyu Wang","Siyuan Zhang","Peiyuan Jiang","Jian Zong","Jingyu Zhang","Tianrui Wang","Yuqin Lin","Zhenghui Chen","Shuqing Xie","Ziyang Ma","Meng Ge","Xiaobao Wang","Longbiao Wang","Jianwu Dang"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-11","first_seen":"2026-08-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.09189","pdf_url":"https://arxiv.org/pdf/2608.09189","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["情感智能","口语语言模型","基准测试"],"reason":"研究情感对话系统评估与对齐，属角色扮演聊天，无人类行为仿真或实验测量目的。","model":"deepseek-v4-pro","scored_at":"2026-08-11T13:05:14","error":null,"has_summary":false,"summary":null},{"id":"2608.09420","version":1,"title":"Intent Speaks Louder: Controllable User Simulation Beyond Response Imitation","zh_title":"意图胜于言辞：超越响应模仿的可控用户仿真","abstract":"User simulators are widely used as scalable environments for training and evaluating interactive assistants. Generating the next user turn is inherently one-to-many: the same profile and dialogue context may support multiple plausible continuations with different local interaction intents. A fluent response may therefore advance the dialogue through an inappropriate intent, such as acceptance rather than repair. Our key insight is that controllable user simulation should separate which local interaction intent the next user turn should realize from how that intent is expressed in language. We introduce UserIDA (User Intent-Directive Alignment), which exposes interaction intent as an explicit per-turn directive. UserIDA defines a six-way intent interface, learns directive-conditioned generation through supervised fine-tuning, and uses intent-calibrated policy optimization during group-based reinforcement learning. The reward preserves composite response quality while ensuring that intent-violating candidates rank below compliant alternatives in mixed groups. On LMSYS-USP, UserIDA achieves 86.6\\% intent accuracy, outperforming the strongest dedicated user-simulator baseline by 24.3 percentage points while improving semantic and stylistic similarity. In within-context interventions, it realizes at least four of the six target intents in 91.7\\% of evaluated dialogue states, compared with 22.9\\% for the strongest external baseline. These results establish per-turn intent control as a complementary dimension to response fidelity in user simulation.","authors":["Bo Wang","Ruixing Zhang","Yunqi Liu","Yang Zhang","Liangzhe Han","Tongyu Zhu","Leilei Sun"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-11","first_seen":"2026-08-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.09420","pdf_url":"https://arxiv.org/pdf/2608.09420","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["用户仿真","对话系统","意图控制"],"reason":"可控用户仿真用于训练对话助手，属角色扮演对话，无实验或测量目的，不涉及人类行为…","model":"deepseek-v4-pro","scored_at":"2026-08-11T13:05:18","error":null,"has_summary":false,"summary":null},{"id":"2608.09510","version":1,"title":"Build it, Break it, Repeat: Benchmarking and improving LLM-manipulated disinformation detection in social media posts","zh_title":"构建、破坏、重复：基准测试与改进社交媒体帖子中LLM操纵的虚假信息检测","abstract":"Detecting machine-generated disinformation on social media is increasingly difficult as large language models (LLMs) make it easier to generate and rewrite misleading content at scale. Static benchmark evaluations, measuring detector performance on fixed held-out datasets, do not capture how detectors behave when posts are deliberately transformed to evade classification. This paper adapts the Build it, Break it, Fix it framework into Build it, Break it, Repeat (BiBiR): iterative sessions designed to stress-test detectors' robustness under iterative adversarial conditions, evaluating whether models remain reliable when disinformation posts are systematically transformed to evade classification. Across five iterations, the findings show that the best adversarial breakers' transformations came from a combination of back-translation and LLM persona-based rewriting, with the best performing technique achieving a 95% label flip rate (LFR), whilst still preserving the meaning of the original posts. The best builders' model was a triplet contrastive model with a dynamic anchor switching (DASS) architecture, which achieved an average accuracy of 72.68%, outperforming the strong baseline (a fine-tuned e5-small-LoRA) by 15 percentage points on the most robust set of breakers' adversarial attacks. The results demonstrate that an iterative framework best exposes detector weaknesses and pushes robustness improvements; however, it may still require semantic preservation analysis to distinguish valid adversarial evasion from transformations that changed the original disinformation claims' meaning.","authors":["Kevin Thomas","Milosz Kasprzyk","Reuel C Igbokwe Onuigbo","Elliott Pert","Cameron Tovey","Jo\\~ao A. Leite","Olesya Razuvayevskaya","Carolina Scarton"],"categories":["cs.CL","cs.AI","cs.SI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-11","first_seen":"2026-08-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.09510","pdf_url":"https://arxiv.org/pdf/2608.09510","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["虚假信息检测","对抗攻击","NLP评测"],"reason":"研究LLM生成虚假信息检测器的鲁棒性，不涉及用LLM仿真人类被试或与人类行为对…","model":"deepseek-v4-pro","scored_at":"2026-08-11T13:05:18","error":null,"has_summary":false,"summary":null},{"id":"2608.09925","version":1,"title":"From Values to Benchmarks: Evaluating Large Language Models for Governmental Use in Dutch","zh_title":"从价值观到基准：评估荷兰政府使用的大语言模型","abstract":"Large language models are increasingly being deployed in governmental settings, yet few existing evaluation frameworks jointly reflect the values of public administration and the linguistic requirements of non-English contexts. We present the \"Grip on LLMs\" framework, a systematic evaluation suite for Dutch governmental use developed in collaboration with domain experts from a major Dutch municipal organisation. Through an advisory board process, user research, and a survey of the users of a civil-servant chatbot, we identify six evaluation dimensions (factuality, honesty, social bias, energy consumption, cost, and training data transparency) and operationalise them into a benchmark suite covering more than 30 multilingual and Dutch-specific models. Our results reveal that no single model excels across all dimensions, and that trade-offs are unavoidable: higher quality consistently comes at greater environmental impact and financial cost, while bias remains largely independent of both. We further find that factuality (whether a model answers correctly) and honesty (whether a model acknowledges what it does not know) are governed by distinct properties, with high factuality not implying high honesty. To make these findings actionable for non-technical audiences, we release a publicly accessible, user-friendly model overview designed for the full range of stakeholders involved in governmental LLM selection, from engineers to policymakers.","authors":["Laurens Samson","Iva Gornishka","Gossa L\\^o","Yuki M. Asano","Sennay Ghebreab"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-11","first_seen":"2026-08-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.09925","pdf_url":"https://arxiv.org/pdf/2608.09925","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["LLM评测","政府应用","基准测试"],"reason":"纯 LLM 能力评测，无人类仿真或行为对照","model":"deepseek-v4-pro","scored_at":"2026-08-11T13:05:21","error":null,"has_summary":false,"summary":null},{"id":"2608.07517","version":1,"title":"The Judge Knows When It Knows: Calibrated Abstention for LLM-Based A/B-Test Prediction","zh_title":"法官知道何时自知：基于LLM的A/B测试预测的校准弃权","abstract":"Can a multimodal LLM predict which version of a web page will win a real A/B test from screenshots alone? We report the most complete answer we are aware of, from six weeks of pre-registered experiments on real conversion tests: mostly no -- and the exceptions are identifiable in advance. On 330 real A/B tests a Gemini 3 Flash judge reaches Cohen's kappa = 0.14, but on the trustworthy (statistically significant) half of the labels the evidence is inconclusive (kappa = 0.11, CI includes zero). We show that 44% of the \"ground-truth\" labels in a leading CRO agency's catalog come from non-significant tests, and that the judge agrees more with the unreliable labels than the reliable ones -- a shared prior between labeler and model, not prediction. Every standard improvement lever (a 2.8x more expensive frontier model, prompt redesign, stimulus fidelity, change-type priors) fails its pre-registered gate. The judge's confident calls are different: a vote-margin gate isolates a subset (49% coverage) reaching kappa = 0.31 on significant labels. We measure the mechanism directly -- judges differing in model or prompt agree with each other at kappa = 0.74-0.88 while agreeing with real outcomes at only ~0.2, so a 16-vote panel carries about 2 effective independent votes -- and we reproduce it in humans: 15 CRO experts agree with each other (inter-rater kappa = 0.53) but score at chance against real outcomes (kappa ~ 0). Consensus, human or model, is reproducible, persuasive, and not evidence. We release our pre-registrations, locked gates, negative results, statistical harness, human responses, and a claims ledger in which every number carries an evidence tier.","authors":["Tyler Dooskin","Squoosh Technical Staff"],"categories":["cs.HC","cs.CL","stat.AP"],"primary_category":"cs.HC","announce_type":"cross","date":"2026-08-11","first_seen":"2026-08-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.07517","pdf_url":"https://arxiv.org/pdf/2608.07517","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["A/B测试","LLM评估","校准弃权"],"reason":"LLM预测A/B测试结果，属多智能体协作解题，无人类行为仿真对照。","model":"deepseek-v4-pro","scored_at":"2026-08-11T13:04:48","error":null,"has_summary":false,"summary":null},{"id":"2608.07537","version":1,"title":"An evolutionary model of animats with VLM-based subjective evaluation","zh_title":"基于VLM主观评价的animat进化模型","abstract":"In this study, we propose a framework that incorporates subjective evaluations provided by a Vision-Language Model (VLM) into the fitness evaluation and selection processes of a genetic algorithm. As the target of evolution, we employ virtual soft robots with flexible morphologies and locomotion and present the VLM with sequence images representing the locomotion of two individuals. Selection is performed via pairwise comparisons based on subjective evaluation terms such as adorably and weirdly. The outcomes of these comparisons are used as selection pressure within the genetic algorithm, enabling the simultaneous evolution of morphology and locomotion. Experimental results demonstrate that subjective selection by the VLM accelerates population convergence compared to random selection, while also giving rise to distinctive morphologies and motions corresponding to each evaluation term. An auxiliary experiment with human participants further showed that, although individual pairwise choices only partly agreed with the VLM selections, the resulting morphological and locomotion tendencies were qualitatively similar and repeated human evaluations imposed noticeable fatigue. Moreover, the observation that similar evolutionary outcomes emerged across different evaluation terms suggests that the VLM does not apply these terms in a purely literal manner but instead decomposes them into multiple internal evaluation criteria when making judgments. This work visualizes the evolutionary process through which subjective linguistic expressions are mapped onto embodied phenotypes and provides a foundational framework for analyzing the structure of subjective judgment in VLMs. The proposed approach is expected to contribute to new developments in evolutionary computation and artificial life research based on subjective evaluation.","authors":["Shota Miyazaki","Takaya Arita","Reiji Suzuki"],"categories":["cs.NE","cs.AI","cs.CL","cs.HC","cs.MA"],"primary_category":"cs.NE","announce_type":"cross","date":"2026-08-11","first_seen":"2026-08-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.07537","pdf_url":"https://arxiv.org/pdf/2608.07537","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C2"],"tags":["进化计算","虚拟机器人","视觉语言模型"],"reason":"研究虚拟软体机器人进化，属于机器人仿真环境，不涉及LLM仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-08-11T13:04:49","error":null,"has_summary":false,"summary":null},{"id":"2608.07614","version":1,"title":"DevIntent: How Much Does LLM-Generated Code Violate Developer Intent?","zh_title":"DevIntent：LLM生成的代码在多大程度上违背开发者意图？","abstract":"Code generated by LLMs can violate a developer's implicit intentions when given an ambiguous prompt, yet standard benchmarks measure only whether code passes its stated test. We introduce the Intent Violation Rate (IVR) and a 49-problem pilot benchmark derived from HumanEval+. Each problem strips implicit constraints from a clarified prompt and encodes them as hidden constraint tests. IVR measures the fraction of LLM-generated solutions that pass the stated (visible) tests yet fail hidden constraint tests that capture unstated intent. Evaluating Claude Sonnet 4.6 and OpenAI GPT 4.1, we find both pass over 92\\% of stated tests yet violate intent in over half of problems (54.5\\% and 63.5\\%), following a systematic, bimodal pattern consistent across both models. Out findings indicate that pass rates overstate how well generated code reflects developer intent.","authors":["Susana Haing","Natan Vidra","Spurthi Setty"],"categories":["cs.SE","cs.CL"],"primary_category":"cs.SE","announce_type":"cross","date":"2026-08-11","first_seen":"2026-08-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.07614","pdf_url":"https://arxiv.org/pdf/2608.07614","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["代码生成","意图违背","基准测试"],"reason":"纯代码生成评测，无人类行为仿真或对照，属多智能体协作解题范畴。","model":"deepseek-v4-pro","scored_at":"2026-08-11T13:04:51","error":null,"has_summary":false,"summary":null},{"id":"2608.07688","version":1,"title":"IntelliAudit: Using Large Language Models to Evaluate Audit Controls","zh_title":"IntelliAudit：使用大语言模型评估审计控制","abstract":"IT audits require auditors to judge whether heterogeneous organizational evidence satisfies semantic security and compliance controls. This judgment is difficult to automate because relevant evidence is distributed across policies, records, spreadsheets, and operational artifacts, and because audit conclusions depend on evidentiary sufficiency rather than keyword matching. We present IntelliAudit, a retrieval-grounded multi-agent system for IT audit evidence evaluation. Given a control and an evidence corpus, IntelliAudit retrieves relevant artifacts, generates an evidence-grounded assessment, challenges adverse findings, adjudicates disagreements, and produces an auditor-facing recommendation with cited evidence, rationale, missing-evidence analysis, and remediation guidance. We instantiate IntelliAudit on ISO/IEC 27001 and evaluate it across multiple simulated organizations using expert auditor review and audit-readiness user feedback. The evaluation shows that IntelliAudit can support control interpretation, evidence-grounded reasoning, and audit-preparation workflows, while also revealing the importance of human oversight for calibrating sufficiency judgments and correcting overly permissive recommendations. These results suggest that retrieval-grounded multi-agent systems can assist audit evidence review, but should remain decision-support tools rather than autonomous certification systems.","authors":["Allison Wilson","Sina Moradi Sabet","Diar Shakimov","Panteha Shahrivar","Mohammad Reza Bagheri","Dean Konenkamp","Mohammad A. Tayebi"],"categories":["cs.AI","cs.CL","cs.CR","cs.HC","cs.MA"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-08-11","first_seen":"2026-08-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.07688","pdf_url":"https://arxiv.org/pdf/2608.07688","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","审计自动化","决策支持"],"reason":"多智能体系统用于审计证据评估，不涉及人类行为仿真或对照。","model":"deepseek-v4-pro","scored_at":"2026-08-11T13:04:53","error":null,"has_summary":false,"summary":null},{"id":"2608.08126","version":1,"title":"Accurate Ensembles, Fragile Narratives: Multi-Scale Stacking and a Fidelity Audit of LLM-Generated Explanations for Credit Risk","zh_title":"准确集成，脆弱叙事：多尺度堆叠与LLM生成信用风险解释的保真度审计","abstract":"Credit scoring increasingly relies on models whose decision logic cannot be read off their parameters, in tension with supervisory expectations that adverse decisions be explainable. A common proposal closes that gap with a language model: compute feature attributions, hand them to an LLM, and let it write the rationale. We build such a system end to end and test whether the second half of the promise holds. The predictive component is a multi-scale stacking ensemble fusing four differently regularised gradient-boosting learners with a residual network through a neural meta-learner trained on out-of-fold predictions. On a public 32,581-application credit dataset it reaches test ROC-AUC 0.9539 (95% CI [0.9462, 0.9616]) and PR-AUC 0.9137, beating the best single model by Delta-AUC = 0.0143 (p = 0.016 under a conservative independence assumption). Our central finding is asymmetric. The ranking gain is real but operationally small: at the F1-optimal threshold the ensemble avoids only six additional missed defaults out of 1,422 against a tuned random forest, cutting cost-weighted loss by under 2%. The narrative layer fails in a way prompt engineering alone does not fix. In an audited case the model named three factors as risk-increasing that the supplied attributions scored as risk-reducing, omitted the dominant driver, and introduced a feature never given to it. We trace this to properties we measure rather than assume: SHAP and LIME agree on which features matter (overlap@10 = 0.80) but not on their order (tau = 0.43, p = 0.18), and the attribution sign for the model's most sensitive input is near a coin flip across applicants (modal-sign share 0.53). Calibration (ECS = 0.117) and perturbation stability (DPD = 0.078) both fall short of our own thresholds. Constrained prompting is necessary but not sufficient: grounding must be verified after generation, not assumed.","authors":["Gregorius Reynaldi Pratama","Kuo-Kun Tseng"],"categories":["cs.LG","cs.CL"],"primary_category":"cs.LG","announce_type":"cross","date":"2026-08-11","first_seen":"2026-08-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.08126","pdf_url":"https://arxiv.org/pdf/2608.08126","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["LLM解释生成","信用评分","保真度审计"],"reason":"论文用LLM生成信用评分解释，评估解释保真度，属NLP能力评测，非人类仿真实验。","model":"deepseek-v4-pro","scored_at":"2026-08-11T13:04:57","error":null,"has_summary":false,"summary":null},{"id":"2608.08143","version":1,"title":"DS@GT ARC at Touch\\'e: Large Language Models for Retrieval-Augmented Debate","zh_title":"DS@GT ARC在Touché 2025检索增强辩论任务中的大语言模型应用","abstract":"We extend the DS@GT ARC working-note submission to the Touch\\'e 2025 Retrieval-Augmented Debate task. The task has two subtasks: generating the next utterance in a simulated debate, and evaluating debate responses according to the Gricean maxims of Quantity, Quality, Relation, and Manner. The DS@GT ARC submission consisted of six leading LLMs from three providers through a retrieval-augmented prompting pipeline. We summarize the results from the working paper and explore whether multi-LLM evaluator agreement is a reliable proxy for official evaluation performance. The analysis shows that frontier LLM systems are strong response generators, and as evaluators they agree strongly within model families. However this consensus does not reliably track the official evaluation target, with the largest gap on the Quality maxim. The accompanying source code for this paper is located at https://github.com/dsgt-arc/touche-2025-rad and https://github.com/dsgt-arc/touche-2025-rad-analysis.","authors":["Anthony Miyaguchi","Conor Johnston"],"categories":["cs.IR","cs.CL"],"primary_category":"cs.IR","announce_type":"cross","date":"2026-08-11","first_seen":"2026-08-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.08143","pdf_url":"https://arxiv.org/pdf/2608.08143","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["辩论生成","LLM评估","检索增强"],"reason":"生成辩论话语并评估，属角色扮演对话，无人类行为对照或实验测量目的。","model":"deepseek-v4-pro","scored_at":"2026-08-11T13:04:57","error":null,"has_summary":false,"summary":null},{"id":"2608.08822","version":1,"title":"Automated Generation of Complexity-Validated Decision Scenarios Using Large Language Models","zh_title":"使用大语言模型自动生成复杂度验证的决策场景","abstract":"Cognitive decision-making research depends on diverse scenarios with carefully controlled complexity, yet manual production is slow, inconsistent, and biased. We developed an automated pipeline that uses LLms to generate structured decision scenarios and validates their complexity through a composite framework rooted in established task-complexity theory. We evaluated 4,238 scenarios across multiple domains and complexity tiers. Measurement validation met rigorous psychometric standards. Agreement among five independent model families was nearly perfect, with an intraclass correlation coefficient of 0.997 and a kappa of 0.971. Known-groups validity demonstrated large separation between tiers, with an eta-squared of 0.587 and all pairwise comparisons significant at p less than .001. Factor analysis revealed a dominant complexity construct, with loadings between 0.87 and 0.96 across three frameworks, while interactivity formed a weaker secondary dimension at 0.34. Discriminant validity was limited by a strong relationship between complexity and text length that persisted after controlling for tier, yielding a partial correlation of 0.86. This constrains construct purity but does not undermine the instrument's tier-grading function. Model analyses showed a negative association between throughput and schema pass rate (r = -0.967, p = .007, n = 5), suggesting a speed-quality trade-off, though largely driven by one high-throughput model. Llama 4 Maverick generated scenarios fastest at 134 per minute versus 25 for DeepSeek Chat V3.2, but underproduced complex-tier scenarios, whereas DeepSeek Chat V3.2 balanced domain coverage with high schema compliance. The system demonstrated strong psychometric properties, enabling reliable classification into Simple, Moderate, and Complex tiers and providing the measurement infrastructure needed for downstream cognitive assessment of AI systems","authors":["Abdalla Doleh","Toni Somers","Ratna Babu Chinnam"],"categories":["cs.AI","cs.CL"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-08-11","first_seen":"2026-08-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.08822","pdf_url":"https://arxiv.org/pdf/2608.08822","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["LLM生成","决策场景","复杂度验证"],"reason":"论文用LLM生成决策场景并验证复杂度，属于NLP能力评测，不以人类行为仿真为参…","model":"deepseek-v4-pro","scored_at":"2026-08-11T13:04:38","error":null,"has_summary":false,"summary":null},{"id":"2608.08885","version":1,"title":"Towards an LLM-based method for quantifying the sexual content in song lyrics","zh_title":"基于大语言模型的量化歌曲歌词性内容的方法","abstract":"Reggaeton is one of the most widely consumed music genres in the world, and its lyrics are commonly regarded as highly sexualized. This claim rests mostly on qualitative studies and on small-scale quantitative ones. This paper has two goals. First, we present a reproducible method that uses a large language model to quantify thematic content in song lyrics along several independent dimensions. The method is not restricted to sexual content. Second, we apply it to a corpus of 1,259 songs by 12 reggaeton artists released between 2002 and 2025. The analysis covers four topics: a dataset characterization, a per-artist comparison, an analysis of how the dimensions change over time, and a comparison between our sexual-explicitness score and Spotify's own explicit flag. We release the data collection code, the scoring prompt, and the corpus, so that other researchers can replicate the approach or apply it to their own lyrics datasets.","authors":["Ignacio M. Sticco"],"categories":["physics.soc-ph","cs.CL","cs.SD"],"primary_category":"physics.soc-ph","announce_type":"cross","date":"2026-08-11","first_seen":"2026-08-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.08885","pdf_url":"https://arxiv.org/pdf/2608.08885","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["LLM标注","歌词分析","内容量化"],"reason":"用LLM量化歌词内容，属NLP自动标注，非仿真人类被试行为或态度。","model":"deepseek-v4-pro","scored_at":"2026-08-11T13:04:38","error":null,"has_summary":false,"summary":null},{"id":"2608.09019","version":1,"title":"How People Evaluate AI-, Expert-, and Peer-Style Financial Advice","zh_title":"人们如何评估AI、专家和同伴风格的财务建议","abstract":"As generative AI increasingly becomes a common source of daily decision-making, including financial choices, it is critical to understand how people evaluate AI-generated financial advice. We conducted a preregistered vignette experiment (N = 285) in which substantive financial content---including facts, numerical values, recommendation direction, and core reasoning---was held constant while communication style varied across AI Financial Assistant (AI), Certified Financial Planner (Expert), and Online Community Forum (OC) advice. Displayed source attribution was independently manipulated through correctly labeled, unlabeled, and mislabeled conditions, allowing us to separate attribution effects from source-specific communication cues. Expert advice was rated more favorably than AI advice on 9 of 10 outcomes (|d|=0.20--0.47), and this advantage remained visible without source labels, where Expert advice outperformed AI advice on 8 of 10 outcomes (up to d=0.60). Correct labels added limited differentiation, whereas mislabeling increased ratings of AI advice for situational fit and overall quality (d=0.42 for each) and attenuated the Expert advantage in situational fit (d=-0.36). Descriptive analyses further showed that AI advice was most responsive to displayed attribution and, conversely, that advice-style differences were most visible under an AI label. These findings show that financial-advice evaluations are shaped jointly by displayed attribution and message-level communication cues. We position disclosure not as a neutral transparency mechanism, but as an interpretive frame whose accuracy and interaction with message cues can shape trust and reliance.","authors":["Aryan Ramchandra Kapadia","Eshwar Chandrasekharan","Koustuv Saha"],"categories":["cs.HC","cs.AI","cs.CL","cs.CY","cs.SI"],"primary_category":"cs.HC","announce_type":"cross","date":"2026-08-11","first_seen":"2026-08-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.09019","pdf_url":"https://arxiv.org/pdf/2608.09019","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["人机交互","财务建议","来源归因"],"reason":"研究人类对AI建议的评价，非用LLM仿真人类被试","model":"deepseek-v4-pro","scored_at":"2026-08-11T13:04:38","error":null,"has_summary":false,"summary":null},{"id":"2608.09282","version":1,"title":"ComboShoppingBench: Evaluating LLM Agents for Budget-Constrained Basket Shopping with Coupons","zh_title":"ComboShoppingBench：评估预算约束下使用优惠券的组合购物LLM智能体","abstract":"Real-world shopping often requires constructing a basket of complementary items rather than retrieving a single product. Such combo-shopping tasks arise in device setup, meal preparation, event planning, and group takeout ordering, requiring joint reasoning about item compatibility, availability, store-level requirements, delivery fees, coupons, and budgets. Evaluation is challenging because multiple baskets may satisfy the same request, making exact-match metrics unsuitable, whereas semantic evaluation alone cannot detect infeasible orders, invalid coupon combinations, or incorrect payments. We introduce ComboShoppingBench, an agentic shopping benchmark for open-ended yet verifiable basket construction in a simulated commerce and takeout environment. During task synthesis, an exploration agent constructs a feasible and semantically coherent basket of purchasable products; this witness guides the generation of coupons, budget constraints, user queries, and aligned evaluation rubrics. During evaluation, LLM judges assess semantic satisfaction, response quality, and claim faithfulness, while deterministic validation checks product-ID validity, budget compliance, and coupon optimality. Experiments with diverse LLM agents demonstrate that even strong agents struggle on ComboShoppingBench, highlighting substantial room for improvement in reliable, constraint-aware combo shopping.","authors":["Adrian Li","Kelong Mao","Yudong Guo","Heming Xia","Xinwei Yang","Lirui Luo","Jace Wong","Pu Yao","Sulong Xu","Simiu Gu"],"categories":["cs.AI","cs.CL"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-08-11","first_seen":"2026-08-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.09282","pdf_url":"https://arxiv.org/pdf/2608.09282","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["LLM智能体","购物基准","多智能体"],"reason":"纯多智能体购物任务评测，无人类行为对照，不涉及人类被试仿真。","model":"deepseek-v4-pro","scored_at":"2026-08-12T13:02:56","error":null,"has_summary":false,"summary":null},{"id":"2608.09638","version":1,"title":"Avalon-ToM-Bench: Evaluating Fine-Grained Theory of Mind via Asymmetric Game Mechanics","zh_title":"Avalon-ToM-Bench：通过非对称游戏机制评估细粒度心理理论","abstract":"Theory of Mind (ToM) is essential for agent interactions, yet existing evaluations either rely on static scenarios that oversimplify mental-state reasoning or interactive settings that provide limited diagnostic insight. We present Avalon-ToM-Bench, a fine-grained benchmark that operationalizes ToM through the asymmetric-information mechanics of The Resistance: Avalon. Rather than evaluating end-to-end gameplay, it decomposes ToM into a 2$\\times$2 taxonomy -- epistemic versus motivational reasoning crossed with inference versus action -- using human-crafted, perspective-constrained queries. Benchmarking 28 LLMs reveals three insights: 1) Reasoning, not knowledge. Models show strong game-rule comprehension but markedly weaker ToM abilities, isolating failures to social reasoning rather than missing domain knowledge. 2) Expression, not representation. Mechanistic analyses via linear probing and activation steering show that models frequently represent correct mental-state inferences in their hidden states but fail to express them during generation -- linear probes recover 77-82% accuracy versus 62-70% from the models' own chain-of-thought. 3) Policy, not deliberation. Dedicated reasoning training yields substantial improvements whereas test-time chain-of-thought provides only marginal gains (+11.0 versus +1.1 points on average), suggesting that robust ToM depends on a learned reasoning policy rather than increased inference-time deliberation.","authors":["Yen-Shan Chen","Yu Chian Duan","Chih-En Kuo","Jian-Bin Wu","Yun-Nung Chen"],"categories":["cs.AI","cs.CL","cs.CY","cs.GT"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-08-11","first_seen":"2026-08-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.09638","pdf_url":"https://arxiv.org/pdf/2608.09638","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["心理理论","多智能体","基准评测"],"reason":"纯多智能体博弈评测，无人类行为对照，不涉及人类仿真实验。","model":"deepseek-v4-pro","scored_at":"2026-08-11T13:05:19","error":null,"has_summary":false,"summary":null},{"id":"2608.09861","version":1,"title":"Towards Expert-level Medical AI for Real-time Video Consultations","zh_title":"面向实时视频会诊的专家级医学人工智能","abstract":"Audio-visual interaction is the standard for patient-physician consultations, enabling natural communication and effective assessment of illness through non-verbal cues. While text-based AI has shown promise, it discards essential perceptual dimensions and limits patients who cannot articulate symptoms in writing. Early efforts to extend medical AI to audio-visual interaction have demonstrated feasibility but not reached clinician-level performance. Here, we provide the first demonstration of expert-level AI in real-time clinical video consultations using AMIE (Articulate Medical Intelligence Explorer) in a video configuration. AMIE (Video) is a Gemini-based multi-agent system integrating low-latency dialogue, clinical reasoning, and real-time audio-visual perception. To guide development, we established a taxonomy and automated evaluations for clinical audio-visual cues in telehealth settings. In a randomized Objective Structured Clinical Examination (OSCE) study with 30 primary care physicians (PCPs), 15 patient actors and 100 clinical scenarios, we compared AMIE (Video), its text-only counterpart AMIE (Text), and PCPs consulting via video. Clinical evaluators rated AMIE (Video) on par or better than PCPs in history-taking, diagnosis, management, and physical observation and examination. Patient actors preferred AMIE's approach to assessing and explaining conditions, while PCPs were preferred for rapport and partnership building. In modality ablation, patient actors preferred AMIE (Video)'s interface over text chat for communicative effectiveness, convenience, and feeling understood. Limitations remain in fine anatomical precision, subtle affective nuances, and high-frequency movements. While further research is needed before real-world translation, these results mark an important milestone toward AI systems capable of augmenting care across the sensory complexity of clinical practice.","authors":["Mahvish Nagda","Jihyeon Lee","Matthew Thompson","Chunjong Park","Tim Strother","Valentin Li\\'evin","Roma Ruparel","Akshay Goel","Teya Bergamaschi","Suhana Bedi","Meet Shah","Pavel Dubov","Liviu Panait","Toshiyuki Fukuzawa","Sam Schmidgall","Craig Schiff","Joseph Xu","Aliya Rysbek","Yana Lunts","Jan Freyberg","Rebecca Hemengway","Sunny Virmani","David Racz","Carey Radebaugh","Jo\\\"elle Barral","Kavi Goel","Dale R. Webster","Katherine Chou","Avinatan Hassidim","Yossi Matias","James Manyika","Gregory Wayne","Tao Tu","Yun Liu","Ethan Goh","Christina Chen","Ryutaro Tanno","Po-Hsuan Cameron Chen","Mike Schaekermann","Anil Palepu"],"categories":["cs.AI","cs.CL","cs.CV"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-08-11","first_seen":"2026-08-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.09861","pdf_url":"https://arxiv.org/pdf/2608.09861","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["医学AI","多智能体系统","视频问诊"],"reason":"多智能体系统用于临床问诊，不涉及人类行为仿真对照，属C1排除项。","model":"deepseek-v4-pro","scored_at":"2026-08-12T13:02:56","error":null,"has_summary":false,"summary":null},{"id":"2608.07642","version":1,"title":"Contextual Value Alignment via Multilayer Combinatorial Fusion","zh_title":"基于多层组合融合的上下文价值对齐","abstract":"Aligning large language models (LLMs) with human values remains a major challenge, especially for trustworthy AI. While existing approaches such as RLHF, CAI, and their variants have achieved promising results, they often rely on a single-agent framework and a unified reward system. This limits their ability to capture ethical pluralism, adapt to diverse moral contexts, and reflect the dynamics of multi-agent moral reasoning. In this work, we propose a framework that utilizes multilayer combinatorial fusion for contextual value alignment (MCF-CVA). At the first layer of the framework, it instantiates multiple moral agents, each fine-tuned to represent a distinctive value. Their outputs are then expanded combinatorially using both score- and rank-combinations as well as average and weighted aggregations. These combined models are then reduced to the same number of initial moral agents. This expansion and reduction (EAR) process continues for multi-layers until a stopping criterion is reached. The MCF-CVA framework leverages cognitive diversity between agents to mitigate conflicts and redundancies across multiple agents, producing responses that better reflect contextual human values. The framework using the EAR algorithm is performed on the dual architecture of Euclidean score space and Kemeny rank space. Empirical evaluations demonstrated that the proposed framework outperforms single-agent baselines, multi-agent single-layer results, and previous aggregation approaches on standard metrics, showing that the MCF-CVA framework provides a robust and effective mechanism for advancing contextual value alignment in LLMs.","authors":["Yuanhong Wu","Djallel Bouneffouf","D. Frank Hsu"],"categories":["cs.AI","cs.LG","cs.MA"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-11","first_seen":"2026-08-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.07642","pdf_url":"https://arxiv.org/pdf/2608.07642","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["价值对齐","多智能体","组合融合"],"reason":"多智能体协作对齐价值观，无人类行为对照，属纯多智能体系统研究。","model":"deepseek-v4-pro","scored_at":"2026-08-11T13:04:53","error":null,"has_summary":false,"summary":null},{"id":"2608.08159","version":1,"title":"When Is a Steerable Concept Representation Real? Measurement Confounds in a Cross-Family Audit of Neuroscience Parallels in LLMs","zh_title":"可操控概念表征何时为真？跨模型家族审计LLM中神经科学类比的测量混淆","abstract":"Large language models (LLMs) are increasingly reported to exhibit human-like neural and cognitive signatures, including concept cells, mental number lines, and cognitive maps. These claims often rely on linear probing and activation steering applied to a single model, yet both methods are highly sensitive to measurement choices. A reported parallel may therefore reflect the model, the measurement procedure, or both. We audit four representative neuroscience-inspired paradigms across 17 models from five families, spanning $0.6$B to $72$B parameters. Our main experiment examines the causal steerability of concept directions. With raw activation units and a fixed layer and coefficient, steerability appears to increase with model scale, resembling an emergent capability. However, this pattern is produced by an uncalibrated pipeline rather than by a claim established in the steering literature. The trend depends jointly on raw units, the readout metric, and the operating point; correcting any one of these removes it. With residual-norm-comparable interventions and held-out operating-point selection, concept steering remains significant at every scale, but shows no significant trend across the Qwen3 series, although the confidence interval does not rule out a moderate positive slope. The remaining results are mixed. A linear geographic world map is consistently decodable in every tested checkpoint up to $72$B. Number magnitude is strongly encoded, but whether individual neurons appear bell-shaped or monotonic depends on the selection criterion. Language-specific structure is localizable, but the direction of the cross-lingual asymmetry reverses under a different attribution method. These results suggest that the main constraint on AI neuroscience is not a lack of phenomena, but a lack of comparable measurements and adequate controls. We release the protocol, stimuli, and code.","authors":["Yuqi Wu","Shengming Zhao","Jie Chen"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-11","first_seen":"2026-08-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.08159","pdf_url":"https://arxiv.org/pdf/2608.08159","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["AI神经科学","线性探测","激活操控"],"reason":"纯AI神经科学审计，评测LLM内部表征，无人类被试仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-08-11T13:04:57","error":null,"has_summary":false,"summary":null},{"id":"2608.08240","version":1,"title":"A Fair Objective for Human-Empowerment-Preserving AI: Desiderata, Design, and Likely Behavioral Consequences","zh_title":"一种维护人类赋权的公平AI目标：需求、设计与可能的行为后果","abstract":"This paper explores the idea of promoting well-being and safety in human-AI interactions by forcing AI agents explicitly to empower humans and to manage the power balance between humans and AI agents in a desirable way. Using a principled, partially axiomatic approach based on desirable properties, we design a parametrizable and decomposable objective function for AI systems that represents an inequality- and risk-averse long-term aggregate of human power. It can take into account models of human bounded rationality and social norms, and crucially, considers a wide variety of possible human goals. We prove how certain desiderata enforce particular functional forms and restrict parameter ranges. We exemplify the consequences of softly maximizing this metric in several paradigmatic situations and describe what instrumental sub-goals it will likely imply.","authors":["Jobst Heitzig","Ram Potham"],"categories":["cs.AI","cs.GT","cs.MA","econ.TH"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-11","first_seen":"2026-08-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.08240","pdf_url":"https://arxiv.org/pdf/2608.08240","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["AI安全","多智能体系统","目标设计"],"reason":"纯多智能体协作研究，设计AI目标函数以赋能人类，无LLM仿真人类被试或人类数据…","model":"deepseek-v4-pro","scored_at":"2026-08-11T13:04:59","error":null,"has_summary":false,"summary":null},{"id":"2608.08254","version":1,"title":"Your Prompt Is Not the Only Prompt: How Much Do LLMs Weight Structured-Output Schema Descriptions?","zh_title":"你的提示不是唯一的提示：LLM对结构化输出模式描述的权重有多大？","abstract":"Structured output, where an LLM populates a predefined JSON schema, has become a default mechanism for data labeling and information extraction, but it also introduces a second instruction channel through schema descriptions. We tested whether classification-label definitions are better placed in the system prompt, user prompt, or schema description using a single-field classification task with nonce labels across ten model configurations from two vendors. Schema descriptions did not consistently outperform prompt-based placement; for GPT-4.1 and GPT-5.4 without reasoning, schema placement underperformed system prompts by 11-13 percentage points. Yet schemas are not inert metadata: when prompts and schemas conflicted, incorrect schema instructions caused accuracy drops of 5-45 points, with Claude Haiku 4.5 falling from 52.5% to 7%, indicating that schema instructions can override prompt instructions, and GPT-5.5 falling from 100% to 73%. Further, adding a required intermediate reasoning field before the label field improved schema-only accuracy by 15-24 points when headroom existed, exceeding system-prompt-only performance in every case tested. The effect held even for Claude Sonnet 4.6 at medium reasoning, where extended thinking alone did not produce a comparable gain. This suggests that schema design can affect how effectively models use information encoded in field descriptions. Overall, these results indicate that schema influence is model-dependent. In practice, the system prompt remains a safe default for definitions, but the bigger discipline is maintaining a single source of truth and preventing prompt/schema drift. More importantly, schema design itself may be a stronger lever than instruction placement. Practitioners should treat prompts and schemas as a unified instruction surface and empirically validate both placement and field design for their target model.","authors":["Sin-Ying Lin"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-11","first_seen":"2026-08-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.08254","pdf_url":"https://arxiv.org/pdf/2608.08254","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["结构化输出","提示工程","模型行为分析"],"reason":"研究LLM对结构化输出中schema描述的权重，属纯NLP能力评测，不涉及人类…","model":"deepseek-v4-pro","scored_at":"2026-08-11T13:05:00","error":null,"has_summary":false,"summary":null},{"id":"2608.08281","version":1,"title":"Exploring LLM Capabilities for Situational Understanding and COLREG compliance on real-world maritime navigation scenarios","zh_title":"探索大语言模型在真实航海导航场景中的态势理解与COLREG合规能力","abstract":"Recently, Large Language Models (LLMs) have shown considerable capability for situational understanding, reasoning, and decision making in different domains, most notable in the automotive sector. Therefore, we explore current state-of-the-art LLMs as a tool for maritime navigation, which includes both codified rules in the Collision Regulations (COLREGs) and uncodified best practices summarized in the concept of ``Good Seamanship''. We construct a dataset consisting of 50 diverse, real-world navigation scenarios from AIS data, label scenarios with applicable COLREG rules, recommended actions, and the reasoning for the action. We explore a variety of different LLM architectures and sizes to determine their understanding of maritime navigation tasks as well as evaluate their reasoning capabilities in this domain. The results obtained indicate that the maritime navigation task remains difficult to solve without fine-tuning, even for larger online models.","authors":["Julius Wirbel","P. Nicholas Hansen","Line K. H. Clemmensen","Roberto Galeazzi"],"categories":["cs.AI","cs.RO"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-11","first_seen":"2026-08-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.08281","pdf_url":"https://arxiv.org/pdf/2608.08281","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C2"],"tags":["LLM","航海导航","自动驾驶"],"reason":"研究LLM在航海导航中的态势理解和规则遵守，属于自动驾驶仿真环境，不涉及人类行…","model":"deepseek-v4-pro","scored_at":"2026-08-11T13:05:01","error":null,"has_summary":false,"summary":null},{"id":"2608.08284","version":1,"title":"Fair on the Surface? Benchmarking Hidden-Output Fairness Gaps in LLM Recommenders","zh_title":"表面公平？基准测试LLM推荐系统中的隐藏输出公平性差距","abstract":"Fairness audits for LLM-based recommenders have largely focused on observable outputs, implicitly assuming that stable recommendations reflect stable internal processing. We challenge this assumption with FairGap, the first benchmark to jointly evaluate recommendation fairness at two levels: observable output shift (OBS) and hidden representation shift (IBS), measured through controlled counterfactual identity probes across gender, age, and race. Their relationship is summarized via Representation-Output Alignment (ROA), with quadrant diagnostics for identifying user-level hidden-output mismatch. Applied to six open-weight LLM families across three domains, FairGap reveals pervasive hidden-output decoupling: ROA rarely exceeds 0.22, and a non-negligible user population shows stable outputs despite substantial internal shifts, a mode that output-only audits cannot detect by design. Further, activation steering that reduces IBS by up to 8x simultaneously worsens OBS, demonstrating a fundamental tension between internal and output-level fairness that existing frameworks are unequipped to diagnose.","authors":["Chan Aristella Lu","Arya Fayyazi","Junhao Zhang","Saeid Shokoufa","Yue Xing","Zhen Xiang","Kyu Hyung Lee","Mehdi Kamal","Massoud Pedram"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-11","first_seen":"2026-08-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.08284","pdf_url":"https://arxiv.org/pdf/2608.08284","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["推荐系统","公平性审计","表示偏移"],"reason":"评估推荐系统公平性，不涉及用LLM仿真人类被试或与人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-08-12T13:02:41","error":null,"has_summary":false,"summary":null},{"id":"2608.08746","version":1,"title":"Scale-to-Dialogue: Low-Burden Elicitation of Daily Premenstrual Symptom Ratings with Small Language Models","zh_title":"规模转对话：用小语言模型低负担获取每日经前症状评分","abstract":"Prospective daily symptom tracking is central to premenstrual health assessment, but repeated ordinal forms impose substantial response burden. We formulate conversational administration as an ordinal label-recovery problem: the system actively elicits a small set of symptom clusters and maps each response to the original severity labels. We used 3,320 complete participant-days from the mcPHASES dataset, covering cramps, mood swing, fatigue, sleep issues, stress, and bloating on a six-level scale. Six participants were reserved for development and 36 for a frozen evaluation comprising 360 participant-days and 2,160 item labels. A ModernBERT evidence gate detected whether a symptom was expressed, and Qwen2.5-1.5B-Instruct produced deterministic structured severity scores. Fixed six-item questioning achieved a quadratic weighted kappa of 0.976, whereas three joint symptom-cluster questions achieved 0.913, 97.45% agreement within one severity level, and 80.94% recall for moderate-or-higher symptoms while reducing questions by 50%. Open-first adaptive policies required 3.92-5.98 questions and produced lower agreement than the corresponding fixed policies. Participant-cluster bootstrap analysis estimated a kappa difference of -0.062 (95% CI -0.076 to -0.048) between the three-cluster and six-item strategies. Active cluster-level elicitation provides a direct, local-model route from natural conversation to reusable daily symptom labels.","authors":["Yifan Wang"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-11","first_seen":"2026-08-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.08746","pdf_url":"https://arxiv.org/pdf/2608.08746","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["症状追踪","对话系统","标签恢复"],"reason":"用SLM替代问卷提问，属NLP能力评测，非仿真人类被试","model":"deepseek-v4-pro","scored_at":"2026-08-11T13:05:04","error":null,"has_summary":false,"summary":null},{"id":"2608.08852","version":1,"title":"Findings of the First Teaching Monster Challenge: A Benchmark of Pedagogical Content Knowledge in AI Agents","zh_title":"首届教学怪物挑战赛发现：AI智能体中教学内容知识的基准测试","abstract":"AI agents can now solve problems, answer like subject experts, and generate long-form multimodal content. However, whether they can adapt a lesson to fit a specified learner, which education calls Pedagogical Content Knowledge (PCK), has not been benchmarked. To measure it, we introduce the Teaching Monster Challenge, the first instructional video generation benchmark to treat the learner persona as an explicit evaluation criterion. Each system is given a topic and a learner persona and must generate a complete instructional video. Every video is screened by an LLM-judge, ranked by crowd pairwise voting, and finalized by an expert panel. The first edition shows that today's systems handle the content well but are far weaker at presenting it and adapting it to the learner. The same process exposes a limit of automatic judging. The LLM-judge separates a clear low-performing tail but ranks the strongest systems poorly. The strongest systems receive nearly identical scores from the judge, so its ranking of them does not match human preference. Progress therefore requires not only better teaching systems but also better automatic judges, and we release the benchmark, rubric, and human judgments as a testbed for both.","authors":["Yi-Cheng Lin","Yu-Kai Guo","Szu-Chi Chen","Bo-Han Feng","Yun-Man Hsu","Hsiang Hsieh","Yu-Jung Lin","Yue-Ling Wu","Jia-Kai Dong","An-Yu Cheng","Yu-Han Huang","Lok-Lam Ieong","Kuan-Yu Chen","Ming-Douo Tchouang","Shao-Hua Sun","Che Lin","Jian-Jiun Ding","Hung-yi Lee"],"categories":["cs.AI","cs.CY","cs.HC","cs.MA"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-11","first_seen":"2026-08-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.08852","pdf_url":"https://arxiv.org/pdf/2608.08852","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["AI教学","基准测试","多智能体"],"reason":"纯多智能体教学视频生成评测，无人类行为对照，不涉及仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-08-11T13:05:06","error":null,"has_summary":false,"summary":null},{"id":"2608.08889","version":1,"title":"LLM Reasoning for Subjective Tasks: Failure Modes, Mitigation, and Dynamic Reasoning Routing","zh_title":"主观任务中的LLM推理：失败模式、缓解策略与动态推理路由","abstract":"Recommendation systems thrive on personalization, where ''correctness'' is rarely a binary truth but a matter of subjective human preference. As Large Language Models (LLMs) are deployed as autonomous verifiers of safety and quality guidelines, they face a distinctive challenge: context-aware preference alignment. Recent gains in Reinforcement Learning with Verifiable Rewards (RLVR) are indexed mostly on objective, mathematical tasks. Through a large-scale study spanning both proprietary and open-source models on four real-world verification tasks from a production recommender platform, we ask whether explicit reasoning generalizes to subjective, human-centric industry rubrics. We expose a fundamental vulnerability: rigid, math-centric reasoning traces actively degrade verification, and applying standard RLVR triggers a phenomenon we term reasoning collapse, in which the policy abandons deliberation in favor of rapid heuristic guessing. We introduce a conditional length-penalized post-training algorithm that intertwines verification accuracy with bounded reasoning length, halting collapse and recovering performance. Finally, we show that a reasoning trace's efficacy is tightly coupled with its socio-linguistic framing: across 1500 synthesized personas, verification accuracy swings by nearly 0.38 macro-F1 depending solely on the adopted reasoning persona---evidence that much subjective-verification error is really reasoning-style mismatch. This observation motivates a mid-training architecture that routes reasoning through contextually aligned personas. This work offers both a scalable algorithmic patch and a long-term architectural blueprint for aligning reasoning models with real-world subjective constraints.","authors":["Juncheng Dong","Ding Tong","Ishan Gupta","Yuyan Wang"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-11","first_seen":"2026-08-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.08889","pdf_url":"https://arxiv.org/pdf/2608.08889","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["LLM推理","主观验证","偏好对齐"],"reason":"研究LLM在主观验证任务上的推理能力，属纯NLP能力评测，不涉及人类行为仿真或…","model":"deepseek-v4-pro","scored_at":"2026-08-11T13:05:09","error":null,"has_summary":false,"summary":null},{"id":"2608.09253","version":1,"title":"SkillSentry: Reliable Skill Execution for LLM Agents via Runtime Assurance","zh_title":"SkillSentry：通过运行时保障实现LLM代理的可靠技能执行","abstract":"LLM agents are increasingly equipped with skills to perform complex tasks through multi-step reasoning and tool use. Although skills provide reusable procedural knowledge, agents may still execute them unreliably. Even when an agent has demonstrated the capability to complete tasks under the guidance of a skill, it may fail to do so consistently across similar tasks or repeated runs due to deviations from the skill procedure or incorrect execution of individual steps. Such instability limits the practical reliability of LLM agents. To address this problem, we propose SkillSentry, a skill-oriented runtime assurance framework built upon a new domain-specific language (DSL) for representing runtime guidance for skill execution. SkillSentry initializes the runtime guidance by combining a skill specification extracted from the corresponding skill document with execution experience mined from historical successful and failed traces. It then wraps around the agent execution loop to monitor and guide skill execution under the current guidance, while iteratively refining the guidance using newly collected traces. We evaluate SkillSentry on 15 skills across two LLM agents, each paired with two backbone models, i.e., Claude Code with Claude-Haiku-4.5 and Claude-Opus-4.6, and Codex with GPT-5.2 and GPT-5.4. Our results show that SkillSentry improves the task success rate of LLM agents by 24.1% across skills, on average, while exhibiting lower variability across repeated runs.","authors":["You Lu","Xinyu Huang","Bihuan Chen","Xin Peng"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-11","first_seen":"2026-08-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.09253","pdf_url":"https://arxiv.org/pdf/2608.09253","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["LLM代理","技能执行","运行时保障"],"reason":"纯多智能体系统研究，提升agent技能执行可靠性，不涉及人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-08-12T13:02:56","error":null,"has_summary":false,"summary":null},{"id":"2608.09343","version":1,"title":"LLM-Guided Heuristic Design from Simulation Traces: A Case Study in Dynamic Production and AGV Scheduling","zh_title":"基于仿真轨迹的LLM引导启发式设计：动态生产与AGV调度案例研究","abstract":"Simulation-based optimization (SBO) evaluates executable policies under stochastic dynamics, but most methods treat the simulator as a black box: aggregate scores rank candidates without revealing why they fail or which policy logic should change. We present an LLM-guided heuristic design framework that uses repeated simulation for selection and event-level traces for diagnosis. Each incumbent is assessed through multiple replications, while replaying its lowest-scoring one produces a queryable trace. A manager agent formulates bottleneck hypotheses from this evidence, and editing agents implement parallel code-level revisions. After execution checks and repeated evaluation, best-so-far selection retains only improvements. LLM revision occurs between evaluation batches, while a fixed policy controls each simulation run. We evaluate the framework in a discrete-event simulation of dynamic production and automated guided vehicle (AGV) scheduling. Across five independent optimization runs with Gemini-3.1-Pro, final mean scores averaged 77.51 on the simulator's 0-100 scale. In the highest-scoring run, trace-based diagnoses motivated proactive charging, distance-aware AGV assignment, and rebalanced dispatch priorities, raising the best-so-far mean score from 62.49 to 78.61. On 100 matched seeds, the best final policy outscored representative rolling-MILP, rule-based, and metaheuristic policies on every seed and retained its advantage under random faults without re-optimization. After separate re-optimization for a longer horizon and variable order interarrival times, the resulting policies again outscored all baselines. Ablations with two LLM backbones showed that removing either parallel candidate generation or trace-database access reduced final mean scores. These results show that simulation traces can guide targeted code-level policy improvement in complex simulation-based scheduling.","authors":["Jinbo Li","Chuanhao Li"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-11","first_seen":"2026-08-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.09343","pdf_url":"https://arxiv.org/pdf/2608.09343","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["LLM引导优化","生产调度","仿真优化"],"reason":"纯多智能体系统优化生产调度，无人类行为对照，不涉及人类仿真实验。","model":"deepseek-v4-pro","scored_at":"2026-08-12T13:02:56","error":null,"has_summary":false,"summary":null},{"id":"2608.09848","version":1,"title":"CEAA: A Cognitive Embodied Agents Architecture for Interactive Computing Systems","zh_title":"面向交互式计算系统的认知具身智能体架构","abstract":"The development of embodied Intelligent Virtual Agents (IVAs) that have cognitive capabilities in real-time interactive virtual environments remains a challenge, even with today's advancements in technology. Existing architectures are often focused on either the implementation of low-level reactive control systems that are constrained by commercial game engines, or high-level representations of reasoning models that can be difficult to implement in virtual worlds. This paper builds on that notion and proposes a modular cognitive architecture for deploying embodied IVAs. This architecture builds on existing, pre-established frameworks such as the Sense-Think-Act paradigm and the Belief-Desire-Intention cognitive model, among others, and aims to provide a reusable implementation-oriented framework as a template for deploying IVA \"brains\" in interactive 3D computing systems. The proposed architecture contributes by providing a modular, implementation-oriented framework for the deployment of embodied, cognitive-capable IVAs and bridges the gap between high-level agent reasoning models with real-time embodied execution, for scalable, adaptive, and explainable agents in complex interactive virtual environments.","authors":["Aimilios Hadjiliasi","Louis Nisiotis"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-11","first_seen":"2026-08-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.09848","pdf_url":"https://arxiv.org/pdf/2608.09848","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C2"],"tags":["具身智能体","虚拟环境","认知架构"],"reason":"论文研究虚拟环境中具身智能体的认知架构，属于游戏/仿真环境中的agent控制，…","model":"deepseek-v4-pro","scored_at":"2026-08-11T13:05:21","error":null,"has_summary":false,"summary":null},{"id":"2608.07543","version":1,"title":"Performance of large language models in the optical diagnosis of colorectal polyps","zh_title":"大语言模型在结直肠息肉光学诊断中的性能","abstract":"Background and Study Aims: Accurate optical diagnosis of colorectal polyps guides resection strategy and surveillance, with multimodal large language models (MLLMs) showing potential for image-based diagnosis. We aimed to evaluate the diagnostic accuracy of MLLMs in classifying colorectal polyps and predicting histology. Methods: We conducted a retrospective diagnostic performance study using the PRIME dataset, a curated set of white light and narrow-band imaging (NBI) images. We evaluated Claude Opus 4, Google Gemini 2.5 Pro, GPT-o3, GPT-4o, and GPT-5. For Paris, Narrow-band Imaging Colorectal Endoscopic (NICE), and predicted histology, we calculated F1 scores, percent correct scores, and accuracy of each MLLM compared to expert responses for 132 cases. Cochran's Q and McNemar's Test were used to determine differences between predicted values of each MLLM. Results: The F1 scores among MLLMs were >0.9 for all models for neoplastic vs. non-neoplastic polyps. Gemini 2.5 Pro demonstrated the highest F1 scores for invasive vs. non-invasive polyps and low- vs. high-grade adenoma, at 0.560 and 0.492 respectively. Claude Opus 4 and GPT-5 had statistically significantly higher percent correct scores than other MLLMs at 41.7%, using Paris classification. Conclusions: Claude Opus 4 and Gemini 2.5 Pro showed the highest accuracy in differentiating polyp subtypes, performing closest to expert consensus. Sensitivity and specificity, however, did not meet ESGE standards, highlighting the need for prospective multicenter trials and the design of human-in-the-loop workflows before clinical deployment.","authors":["Joshua C. Vences","William T. Tran","Nikko Gimpaya","Catharine M. Walsh","Rishad J. Khan","Robert Bechara","Asher C. Wiggins","Celine N. Rousan","Kaitlyn V. G. L. Morgado","Angie Ibrahim","Kevin H. M. Kuo","Daniel von Renteln","Alexander Hann","Dennis L. Shung","Michael A. Scaffidi","Charles M\\'enard","Joshua Landy","Samir C. Grover"],"categories":["cs.CV","cs.AI"],"primary_category":"cs.CV","announce_type":"cross","date":"2026-08-11","first_seen":"2026-08-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.07543","pdf_url":"https://arxiv.org/pdf/2608.07543","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["医学图像诊断","LLM性能评测","内窥镜"],"reason":"评估LLM在医学图像诊断上的性能，属于纯NLP/CV能力评测，不以人类行为仿真…","model":"deepseek-v4-pro","scored_at":"2026-08-11T13:04:50","error":null,"has_summary":false,"summary":null},{"id":"2608.07593","version":1,"title":"Weather- and Location-Aware Agentic Dining Recommendation: Leveraging LLM World Knowledge for Region-Sensitive Contextual Reasoning","zh_title":"天气与位置感知的智能体餐饮推荐：利用大语言模型世界知识进行区域敏感的情境推理","abstract":"Context-aware recommender systems have long recognized that factors such as location, time, and weather shape where and what people choose to eat. Existing weather-aware food and point-of-interest recommenders, however, typically treat weather generically -- mapping conditions to preferences through hand-crafted rules or specially trained context models -- and do not capture that the culturally appropriate response to weather is itself region-specific: a rainy evening calls for hot tea and fried snacks in one culinary culture and for very different comfort food in another. Encoding such weather-by-region-by-cuisine interactions as explicit rules or training data is brittle and does not scale. We present a weather- and location-aware agentic dining-recommendation system that takes a different approach: a large language model (LLM) orchestrates tools for location and weather retrieval and then reasons in natural language over the combined context, drawing on the cultural and culinary world knowledge already latent in the model to produce region-sensitive, weather-appropriate recommendations without per-region rule tables or specialized training. We describe the agent architecture, the tool-orchestration flow (Google location services and a weather service feeding an OpenAI LLM), and the reasoning mechanism, and we report on a working prototype that was implemented and briefly deployed end-to-end. We discuss design trade-offs -- cost, latency, ambiguity handling, and fallbacks -- and we are explicit about limitations, including the absence of a formal user study and the risk of cultural stereotyping in locality-based inference. The contribution is architectural: a simple, extensible pattern for incorporating environmental and cultural context into agentic recommendation through LLM reasoning rather than engineered rules.","authors":["Kadharmoideen Fadurudeen"],"categories":["cs.HC","cs.AI","cs.IR"],"primary_category":"cs.HC","announce_type":"cross","date":"2026-08-11","first_seen":"2026-08-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.07593","pdf_url":"https://arxiv.org/pdf/2608.07593","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["推荐系统","大语言模型","情境感知"],"reason":"该系统是面向用户的餐饮推荐，非以LLM替代人类被试进行实验或测量，属于应用型推…","model":"deepseek-v4-pro","scored_at":"2026-08-11T13:04:51","error":null,"has_summary":false,"summary":null},{"id":"2608.07902","version":1,"title":"Beyond \"I Can't Help With That\": How Child Safety Experts Evaluate AI Chatbot Safety","zh_title":"超越“我无法帮助”：儿童安全专家如何评估AI聊天机器人安全性","abstract":"Youth increasingly turn to AI chatbots for social and emotional support, raising concerns about how these systems respond, especially in high-stakes situations. However, existing child safety evaluations of AI lack grounding in real-world harms that youth experience, rely on unvalidated assumptions about what counts as an appropriate output (e.g., refusal), and typically focus on detecting adversarial prompts or surface-level harms in outputs only. Thus, these evaluations can fail to detect responses that pose harm to youth in practice. To better understand the limitations of current evaluation practices, we conducted interviews with 19 practitioners working directly with youth in vulnerable situations, including social workers, therapists, and psychologists, asking them to reflect on chatbots' responses to risky situations commonly faced by youth, as established in prior empirical work. Practitioners identified chatbot behaviors likely to cause harm as well as those that could meaningfully support youth in difficult moments, discussed the role that chatbots should (and should not) play in these interactions, and offered concrete recommendations for improving chatbot responses. Based on these findings, we provide recommendations for AI child safety evaluation and infrastructure, and highlight the need for incorporating practitioners' perspectives into safety work.","authors":["Hannah Cha","Neha Shukla","Solon Barocas","Alexandra Chouldechova","Eugenia Kim","Jennifer Wortman Vaughan"],"categories":["cs.CY","cs.AI","cs.HC"],"primary_category":"cs.CY","announce_type":"cross","date":"2026-08-11","first_seen":"2026-08-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.07902","pdf_url":"https://arxiv.org/pdf/2608.07902","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["AI安全","儿童保护","聊天机器人"],"reason":"研究AI聊天机器人对青少年的安全性，属于角色扮演对话评估，无实验或测量目的，不…","model":"deepseek-v4-pro","scored_at":"2026-08-11T13:04:55","error":null,"has_summary":false,"summary":null},{"id":"2608.08266","version":1,"title":"On the Robustness of LLMs' Internal Representation of Code Correctness","zh_title":"大语言模型代码正确性内部表征的鲁棒性研究","abstract":"Code generated by modern language models often reads naturally. Yet, it also often fails to implement what was asked. This should be no surprise, as research shows the models' own confidence signals are poorly calibrated with actual correctness. A promising way to assess correctness looks inside the model: by contrasting the hidden states of correct and incorrect programs, recent work captured an internal signal of code correctness that is able to judge candidate solutions better than the model's token-level or stated confidence, with no test execution. However, this signal was captured under one particular way, leaving open an important question: whether it reflects a robust property of the model or an artifact of that choice. We study this question systematically, varying how the signal is extracted from the model internals. Besides this, we also ask if the signal's quality is limited by the data used to extract it, by constructing program pairs that differ only in the fault that makes them incorrect. Our results show that no single configuration is best, and that isolating the fault does not help.","authors":["Francisco Ribeiro","Sohaila Abdulsattar","Renata Gonzalez","Mahmoud Kassem","Sarah Nadi"],"categories":["cs.SE","cs.AI","cs.LG"],"primary_category":"cs.SE","announce_type":"cross","date":"2026-08-11","first_seen":"2026-08-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.08266","pdf_url":"https://arxiv.org/pdf/2608.08266","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["代码正确性","模型内部表征","鲁棒性分析"],"reason":"研究LLM内部表征对代码正确性的判断，属纯NLP能力评测，无人类行为仿真或对照。","model":"deepseek-v4-pro","scored_at":"2026-08-11T13:05:00","error":null,"has_summary":false,"summary":null},{"id":"2608.08344","version":1,"title":"PRISM: A Predictive Protocol for Permutation Optimization via Landscape Diagnostics","zh_title":"PRISM：一种基于景观诊断的排列优化预测协议","abstract":"Permutation optimization arises whenever the components of a system are fixed but their ordering affects performance. We introduce PRISM, a predictive protocol for permutation optimization that measures a fitness landscape before selecting a search strategy. PRISM uses inexpensive landscape diagnostics, including one-step move autocorrelation and fitness-distance correlation, to predict useful mutation operators, identify when structured search is likely to outperform random sampling, and detect regimes in which search provides little advantage. Across synthetic permutation landscapes, neural architecture benchmarks, scientific machine learning pipelines, and large-language-model instruction ordering, the protocol makes testable predictions about search behavior before optimization begins. Exhaustive instruction-ordering experiments reveal substantial performance variation induced solely by permutation, while cross-model experiments show that useful ordering structure can transfer across model families and task difficulty. Additional experiments demonstrate that instruction ordering remains consequential after prompt wording is optimized, indicating that content optimization and ordering optimization are complementary. The results position PRISM not as a universally superior optimizer, but as a framework for determining when permutation search is useful, which representation and operator should be used, and when simpler alternatives are preferable.","authors":["Blessings Mambwe"],"categories":["cs.LG","cs.AI","math.OC"],"primary_category":"cs.LG","announce_type":"cross","date":"2026-08-11","first_seen":"2026-08-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.08344","pdf_url":"https://arxiv.org/pdf/2608.08344","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["排列优化","景观诊断","多智能体"],"reason":"纯多智能体协作与排列优化，无人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-08-12T13:02:53","error":null,"has_summary":false,"summary":null},{"id":"2608.09351","version":1,"title":"Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute","zh_title":"大语言模型的测试时增强：当输入多样性在匹配计算量下胜过输出多样性","abstract":"Test-time scaling improves LLM accuracy but multiplies inference cost, making the accuracy gained per unit of compute the metric that matters in deployment. Self-consistency is one of the established approaches, which spends this budget entirely on the output side by sampling repeated reasoning paths. We study Test-Time Augmentation (TTA), which extends self-consistency by also perturbing the input, aggregating predictions across transformed versions of the input, and ask whether input-side diversity converts compute into accuracy more efficiently than output-side diversity. We perform a systematic, matched-compute comparison: we evaluate three simple input-side strategies (semantic rephrasing, lexical perturbations, and visual transformations) across six datasets covering general and multilingual knowledge, mathematical reasoning, multi-modal question answering, and sentiment classification, against chain-of-thought prompting and self-consistency. Semantic rephrasing delivers consistent and statistically significant accuracy gains while Pareto-dominating self-consistency on cost-effectiveness, delivering roughly 1.8X more accuracy per dollar and outperforming it on five of six tasks. We further analyze the number of augmentations, multi-modal strategies, and base model scaling, finding that TTA is most cost-effective for mid-tier models where a stronger model is unavailable or too expensive. Our findings indicate that for current mid-tier LLMs, varying the input converts inference compute into accuracy more efficiently than varying the reasoning path alone. The TTA implementation is available at https://github.com/aws-samples/sample-genai-reflection-for-bedrock.","authors":["Nikita Kozodoi","Zainab Afolabi","Jack Butler"],"categories":["cs.LG","cs.AI","stat.ML"],"primary_category":"cs.LG","announce_type":"cross","date":"2026-08-11","first_seen":"2026-08-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.09351","pdf_url":"https://arxiv.org/pdf/2608.09351","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["测试时增强","推理效率","准确率优化"],"reason":"纯NLP能力评测，研究测试时增强对准确率的影响，不以人类行为为参照系。","model":"deepseek-v4-pro","scored_at":"2026-08-11T13:04:40","error":null,"has_summary":false,"summary":null},{"id":"2608.09828","version":1,"title":"Multi-Agent AI Safety as an Institutional Design Problem","zh_title":"作为制度设计问题的多智能体AI安全","abstract":"AI agents increasingly work inside systems that govern how they delegate tasks, move information, execute actions, and use shared resources. Recent work already shows that deployment rules can change collective behavior. Here we ask which parts of an AI institution produce safety and how they do it. This is the first paper from POLIS, an ongoing research programme studying algorithmic institutions for multi-agent systems. We report a frozen 5,280-episode study suite. The main pre-specified delegation experiment spans four model families; a targeted high-conflict diagnostic adds three additional model endpoints. In matched structured workflows, the model sees different rule formulations and guards consult different authority states. We also vary the attractiveness of the immediate compliant internal/self fallback and allow blocked workflows to continue. A detailed constitutional prompt produces 0/384 realized violations. A provenance-aware executable guard also produces 0/384, although it blocks prohibited attempts in 51/384 episodes; 44/51 of those episodes later complete safely. The local-state guard's failures concentrate in scenarios where an ordinary transformation changes visible policy while originating authority stays fixed. In matched laundering scenarios, that guard admits violations in 22/96 episodes and provenance enforcement in 0/96 (p = 4.77 x 10^-7). A separate resource-allocation experiment shows that revealing the numerical value of an otherwise identical cap changes agent requests. In these structured workflows, the same final violation rate can hide very different mechanisms. The rule itself is only part of the institution. The authority state the system trusts matters, and so does the path available after a block.","authors":["Abdullah X"],"categories":["cs.LG","cs.AI","cs.MA"],"primary_category":"cs.LG","announce_type":"cross","date":"2026-08-11","first_seen":"2026-08-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.09828","pdf_url":"https://arxiv.org/pdf/2608.09828","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","AI安全","制度设计"],"reason":"纯多智能体系统安全机制设计，无人类行为对照，不涉及人类被试仿真。","model":"deepseek-v4-pro","scored_at":"2026-08-11T13:05:19","error":null,"has_summary":false,"summary":null},{"id":"2608.09294","version":1,"title":"Graphing the Everyday: A Neurosymbolic Approach to Eliciting Routines for Just-In-Time Adaptive Interventions","zh_title":"绘制日常生活：一种用于即时自适应干预的神经符号方法","abstract":"Just-In-Time Adaptive Interventions (JITAIs) increasingly rely on conversational agents to elicit user routines, yet translating fluid human dialogue into rigid schedule data remains a significant challenge. We conducted a qualitative investigation of a neurosymbolic pipeline, combining Large Language Models (LLMs) with a Neo4j knowledge graph, to map unstructured verbal narratives into actionable interventions. Through human-centric evaluation using natural-language playbacks, we identified a critical \"mental-model gap,\" where the linear extraction of LLMs clashes with hierarchical, non-linear human storytelling, causing severe entity fragmentation. Furthermore, we articulate an \"ecological mismatch,\" demonstrating that algorithmic schedule availability frequently ignores the user's fluctuating psychological receptivity and physical energy levels. To resolve these tensions, we propose actionable design heuristics, including routine piggybacking, adaptive negotiation, and scalable transparency. Ultimately, these guidelines provide a foundational framework for evolving rigid schedule-trackers into empathetic, context-aware proactive agents capable of supporting long-term health behavior change.","authors":["Shakyani Jayasiriwardene","Blake Mountford","Meican Ma","Niels van Berkel","Nicholas Koemel","Matthew Ahmadi","Jorge Goncalves","Emmanuel Stamatakis","Zhanna Sarsenbayeva"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-08-11","first_seen":"2026-08-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.09294","pdf_url":"https://arxiv.org/pdf/2608.09294","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["对话代理","知识图谱","健康干预"],"reason":"研究对话代理提取用户日程，属角色扮演聊天机器人，无实验或测量目的，不涉及人类仿…","model":"deepseek-v4-pro","scored_at":"2026-08-11T13:05:18","error":null,"has_summary":false,"summary":null},{"id":"2608.07810","version":1,"title":"Mobility, Memory, and Network Structure in Agent-Based Models of Convention Tipping and Convergence","zh_title":"基于智能体的约定转变与收敛模型中的移动性、记忆和网络结构","abstract":"Tipping-point dynamics describe the critical conditions under which a committed minority drives a population to abandon an established convention in favor of a new one. We present a transparent agent-based model of this process, in which agents hold one of two behavioral states and a mobile committed minority attempts to overturn the incumbent convention. Our goal was to examine how localized mobility, bounded agent memory, and network topology jointly influence the tipping threshold. Using a custom agent-based simulation framework, we found that in many configurations, tipping becomes effectively inevitable: given sufficient time, the population always converges to the minority state. This observation motivated a complementary analysis focused on the pace of convergence rather than its feasibility. We introduce a unified predictive model that accurately estimates how structural and behavioral parameters determine the time required for complete adoption, showing that mobility is the dominant accelerator while memory and connectivity modulate convergence in systematic ways. Together, these results extend classical tipping-point research by linking structural and behavioral factors not only to the likelihood of convention change but also to the timescale on which it unfolds. While we frame the model in terms of convention-like binary behavioral adoption, the same mechanisms bear on norm change and other contagion-like social processes.","authors":["Joe Shymanski","Garrick Springer","Sandip Sen"],"categories":["cs.MA","cs.SI"],"primary_category":"cs.MA","announce_type":"new","date":"2026-08-11","first_seen":"2026-08-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.07810","pdf_url":"https://arxiv.org/pdf/2608.07810","source_feed":"cs.MA","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体仿真","临界点动力学","社会约定"],"reason":"纯多智能体仿真，无LLM，无人类数据对照，研究约定转变的临界点，不涉及人类被试…","model":"deepseek-v4-pro","scored_at":"2026-08-11T13:04:55","error":null,"has_summary":false,"summary":null},{"id":"2608.08437","version":1,"title":"AI and the Research Team","zh_title":"人工智能与研究团队","abstract":"Artificial intelligence is associated with larger research teams, yet in mathematics, among the most codifiable fields, individual researchers working with AI now produce research-grade results. A span-of-control model reconciles these observations. AI lowers execution cost, which expands laboratory scale, and automates codifiable tasks, which lowers the member share of each unit. Team size is therefore quasi-concave in AI capability, with at most one peak. The model predicts that a fully codifiable team peaks when effective automation coverage reaches a closed-form threshold, typically near complete coverage, and, among fields with shared primitives that possess an interior peak, those with less irreducibly human task content peak first. Under explicit priors, the 90 percent forecast intervals for the fully codifiable peak span 2026 to 2030.","authors":["Johan Fourie"],"categories":["econ.GN","q-fin.EC"],"primary_category":"econ.GN","announce_type":"new","date":"2026-08-11","first_seen":"2026-08-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.08437","pdf_url":"https://arxiv.org/pdf/2608.08437","source_feed":"econ.GN","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","团队规模","AI自动化"],"reason":"纯多智能体协作模型，研究团队规模与AI能力关系，不涉及人类行为仿真或对照。","model":"deepseek-v4-pro","scored_at":"2026-08-11T13:05:03","error":null,"has_summary":false,"summary":null},{"id":"2608.07367","version":1,"title":"People Are Not Just Their Countries. Disentangling Social Determinants of LLM Value Alignment Across Europe","zh_title":"人不仅是其国家：解构欧洲LLM价值观对齐的社会决定因素","abstract":"As Large Language Models (LLMs) are increasingly used as a primary source of information and advice, understanding their alignment to humans in terms of values becomes a pressing concern. A growing literature has leveraged large scale surveys to investigate to what extent LLMs' and humans' stated values and opinions align. With limited exceptions, studied populations have been defined country borders or cultural bounds. Yet, this focus neglects the role that socio-demographic divides may play for value alignment disparities. Relying on the European Social Survey, we address this knowledge gap by considering value alignment displayed with respect to 10 prominent commercial LLMs in terms of 15 socio-demographic variables as well as country of residence. Our analyses reveal that LLMs are indeed unequally aligned to the values of different socio-demographic groups, notably those defined by education, income, occupation and religion. When examining alignment at the individual level, a respondent's country, taken as a stand-alone variable, explains a substantial amount of variation that is on par with the full set of considered socio-demographics. Further disentangling the respective role of country-level and socio-demographic factors, we find they are complementary in explaining value alignment patterns, with their relative weights varying across the subset of questions considered.","authors":["Maria-Louisa Wightman","Guillaume Bied","Tijl De Bie"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-10","first_seen":"2026-08-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.07367","pdf_url":"https://arxiv.org/pdf/2608.07367","source_feed":"cs.AI","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B3"],"tags":["LLM价值观对齐","社会调查复现","偏差分析"],"reason":"用LLM复现人类价值观调查，以欧洲社会调查为基准，分析对齐偏差，直接命中A1/…","model":"deepseek-v4-pro","scored_at":"2026-08-10T13:01:29","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-10","rank":2,"question":"在欧洲国家中，LLM与人类价值观的对齐在不同社会人口统计因素和国家之间呈现何种差异模式？","design":"使用10个商业LLM重复回答欧洲社会调查（ESS）中的价值观问题，计算LLM回答与人类受访者回答的对齐分数，分析对齐分数在15个社会人口统计变量和国家之间的差异。","baseline":"欧洲社会调查（ESS）2023-2024波次中29个欧洲国家和以色列的人类受访者真实回答。","findings":"LLM与不同社会人口群体（尤其是教育、收入、职业和宗教定义的群体）的价值观对齐程度不平等；国家变量单独解释的对齐变异量与全套社会人口统计变量相当，且国家与社会人口统计因素在解释对齐模式上互补。","reliability":"论文未讨论","relevance":"该研究直接以真实大规模人类调查为基准，评估LLM在价值观对齐上的社会人口统计偏差，属于批判性仿真研究，命中研究者关心的A1、A2、B1、B3准则，值得精读。","inspiration":"借鉴其使用大规模社会调查作为基准、通过逆倾向加权和方差分解分离国家与社会人口因素贡献的方法；可迁移到信贷审批或保险定价中的算法公平性评估场景；以LLM作为信贷审批员，输入不同社会人口特征的虚拟申请人，测量审批结果差异，并以真实信贷审批数据或调查数据作为对照基准。"}},{"id":"2608.06379","version":1,"title":"Preventive Care Recommendations by Large Language Models","zh_title":"大语言模型的预防保健建议","abstract":"Preventive care services (PCS) extend life, yet physicians often underprioritize highly effective interventions such as lifestyle modifications (Zhang et al., JAMA Network Open 2020). We evaluated whether large language models (LLMs) replicate and augment physician prioritization of PCS under time constraints. Using Zhang et al.'s validated survey with two patients assessed during long and short visits, we compared seven LLMs with historical physicians. We generated 137 simulated physician personas matching cohort demographics and tested three prompts per model. Primary outcomes were concordance with physician rankings, measured by Spearman correlation, and Consensus-Stratified Agreement (CSA), the proportion of LLM selections rated 4 or higher that matched physician consensus across agreement strata. Secondary outcomes included life-years gained per prioritized choice (LYGPC), consistency, and selectiveness. Augmentation was assessed by having models revise physician rankings under three informative prompts, with delta LYGPC quantifying impact. LLMs closely mirrored physicians (mean Spearman = 0.83, SD = 0.11), with high CSA at extreme agreement ranges (94%, 197/210) but low CSA in moderate ranges (21%, 30/140), where they underprioritized lifestyle services (8.8% vs. 38% rated 4 or higher; P < .001). Several models exceeded physicians in LYGPC and consistency while being more selective. Time constraints affected physicians and LLMs similarly, increasing LYGPC and selectiveness but reducing consistency. Augmentation effects varied by model. Current LLMs reproduced physicians' time-sensitivity and base-rate prioritization while exacerbating underprioritized lifestyle interventions. Some models improved prioritization performance, but consistent augmentation will require value-aligned training, explicit time-constraint representation, and prospective real-world validation.","authors":["Eden Avnat","Elia Yanko","Ori Yoran","Raja-Elie E. Abdulnour"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-08-10","first_seen":"2026-08-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.06379","pdf_url":"https://arxiv.org/pdf/2608.06379","source_feed":"cs.HC","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["LLM仿真","医生决策","人类数据对照"],"reason":"用LLM模拟医生决策并与真实医生数据对照，评估仿真可靠性及失效条件，直接命中核…","model":"deepseek-v4-pro","scored_at":"2026-08-10T13:01:23","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-10","rank":1,"question":"大语言模型能否复现并增强医生在时间压力下对预防性服务的优先级排序？","design":"使用7个LLM模拟137名医生角色，匹配原始医生队列的人口统计特征，在长/短就诊时间两种条件下对两个虚拟患者完成预防性服务评级和排序任务，测量排序一致性、共识分层一致性、每项优先选择获得的寿命年、一致性和选择性，并通过三种提示策略让模型修正医生排序以评估增强效果。","baseline":"Zhang等人2020年发表的137名医生在相同调查中的真实排序和评级数据。","findings":"LLM与医生的排序高度相关（平均Spearman=0.83），在极端共识区间一致性高（94%），但在中等共识区间一致性低（21%），且严重低估生活方式干预服务（8.8% vs 38%）。部分模型在寿命年增益和一致性上优于医生，但时间压力对两者的影响相似，增强效果因模型而异。","reliability":"论文指出LLM在中等共识服务上一致性差，加剧了对生活方式干预的低估，且增强效果不稳定，需要价值对齐训练、显式时间约束表示和前瞻性真实世界验证。","relevance":"该研究直接以真实医生数据为基准，评估LLM模拟人类专业决策的可靠性与偏差，并揭示了在中等共识情境下仿真失效的条件，高度契合研究者对LLM仿真实验的批判性关注。","inspiration":"借鉴其通过分层共识分析（CSA）揭示仿真在中等共识区间失效的方法，可迁移到经济预测或政策评估场景中检验LLM对分析师共识的复现偏差。｜具体可应用于信贷审批或投资建议场景，考察LLM模拟信贷员或分析师在信息不完全下的决策。｜以真实信贷审批数据为基准，让LLM扮演不同经验水平的信贷员，在高低信息量条件下进行审批决策，测量其与人类审批员排序的相关性及在不同共识水平上的偏差。"}},{"id":"2608.04009","version":2,"title":"SocietyBench: Forecasting Counterfactual Social-World Evolution","zh_title":"SocietyBench：预测反事实社会世界演化","abstract":"Large language models (LLMs), and the agents built on top of them, are now benchmarked heavily on whether they can finish a task -- fix a bug, drive a browser, operate a GUI. A complementary social ability, namely how well a model understands and forecasts the way real social events unfold, has barely been measured. We introduce SocietyBench, an end-to-end benchmark that takes a one-line event topic, collects Web news and social-media posts across five platforms, distills them into a date-indexed timeline that keeps factual events and a public-opinion layer separate, and then turns every cutoff date on that timeline into an audited bank of forecasting questions. Questions are scored on two orthogonal 100-point axes: probability calibration and temporal accuracy. Before any model sees a timeline, a three-phase procedure replaces every named entity and shifts every date by a per-event constant, turning a real arc into a counterfactual social world -- structurally identical to what happened, but stripped of the surface labels a model could match against pre-training memory. On five heterogeneous events and 125 prediction points in Chinese and English editions, the strongest of six frontier LLMs reaches only 75.0 out of 100, against a trivial anchor of 50. The two axes come apart: a model can be calibration-strong but time-weak, or the reverse. Three agent frameworks built on a shared base model fail to improve on that base, and two model-free heuristics trail every LLM. Per-event gaps reach 21.4 points on a single axis, which is our main argument for evaluating on several events rather than one. All anonymized timelines, question banks, ground truth, and scoring code are released.","authors":["Zhenran Wang","Zhonghan Bian","Jinsong Li","Zhangyang Qi"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-08-10","first_seen":"2026-08-05","revised_at":"2026-08-10","abs_url":"https://arxiv.org/abs/2608.04009","pdf_url":"https://arxiv.org/pdf/2608.04009","source_feed":"cs.CL","score":8,"bucket":"selected","rubric_hits":["A3","B1","B4"],"tags":["社会模拟","LLM预测","反事实推理"],"reason":"用LLM预测反事实社会事件演化，有真实新闻/舆论数据对照，并评估校准与时间准确…","model":"deepseek-v4-pro","scored_at":"2026-08-10T13:01:51","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-10","rank":3,"question":"大型语言模型能否准确预测反事实社会事件的演化，包括事实进展与公众舆论？","design":"构建SocietyBench基准：基于五个真实社会事件，自动采集新闻与社交媒体数据，生成匿名化时间线，将每个截止点转化为概率校准和时间准确性两类预测问题，评估LLM及智能体框架的预测表现。","baseline":"以真实事件后半段的时间线节点作为事实与舆论演化的真实基准。","findings":"最强前沿LLM仅得75.0分（满分100），概率校准与时间准确性两个维度表现分离；智能体框架未能超越基础模型，不同事件间模型表现差异可达21.4分。","reliability":"论文指出单事件评估不可靠，需多事件测试；匿名化虽防记忆但可能改变社会动态结构；未讨论模型在更长预测窗口或更多类型事件上的泛化局限。","relevance":"该研究用LLM预测社会事件演化，有真实新闻与舆论数据对照，并系统评估校准与时间准确度，直接回应了研究者对LLM仿真可靠性及失效条件的关注，值得精读。","inspiration":"可借鉴其匿名化反事实设计以剥离模型记忆干扰，并用双轴评分（概率校准与时间误差）全面衡量预测质量。｜可迁移至政策公告的预期形成研究，如央行加息声明后的市场反应预测。｜以LLM为被试，提供匿名化的宏观经济事件时间线，要求预测后续资产价格变动概率与时间，以真实市场数据为对照基准。"}},{"id":"2608.07316","version":1,"title":"Natural Language Processing Psychometrics","zh_title":"自然语言处理心理测量学","abstract":"Natural Language Processing (NLP) models predicting mental health outcomes rarely specify what they measure: contextual knowledge, emotional content, or syntactic structure. NLP Psychometrics treats psychological prediction from text as a psychometric problem, linking scores to interpretable linguistic evidence and testing beyond the training text format. Nine LLMs, conditioned on controlled personas (cognitive digital shadows), completed psychometric questionnaires with textual explanations per item. We extracted emotional profiles and syntactic-semantic structure via textual forma mentis networks, combined with personality and sociodemographic variables in ablated random forest (RF) regressors, using SHAP to identify which features drove performance and in which direction. Full RF models explained up to 70.8% of variance in life satisfaction (SWLS), 55.7% in depression (PHQ-9), and, for DASS-21, 68.5% depression, 76.0% anxiety, 72.4% stress. Sociodemographics alone explained no meaningful variance in depression, anxiety, or stress, but did so for life satisfaction, where emotion features and income were the strongest predictors; neuroticism and network topology instead dominated depression and anxiety, reversing direction between them. Without retraining, RF models separated diaries from low- and high-score personas ($r$ up to 0.91) and, using only network/emotion features, classified clinical from control participants in real transcripts with up to 68% accuracy. These results show the promise and limits of synthetic data: LLM personas can expose model biases, recover patterns consistent with clinical rumination, and support psychometric prediction from human text without a matched questionnaire, but cannot substitute for human validation. NLP Psychometrics makes these distinctions explicit, measurable, and testable through interpretable AI and network/emotional features.","authors":["Edoardo Sebastiano De Duro","Emma Franchino","Massimo Stella"],"categories":["cs.CL","cs.AI","cs.SI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-10","first_seen":"2026-08-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.07316","pdf_url":"https://arxiv.org/pdf/2608.07316","source_feed":"cs.CL","score":8,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","心理测量","可解释AI"],"reason":"用LLM persona模拟人类心理测量，有真实人类数据对照，并讨论合成数据的…","model":"deepseek-v4-pro","scored_at":"2026-08-10T13:01:28","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-10","rank":4,"question":"如何从语言文本中可解释地推断心理构念，并评估LLM生成的合成数据在心理测量中的有效性与局限？","design":"使用9个LLM，通过控制人格与社会人口学变量构建认知数字影子（persona），让它们完成心理测量问卷并对每个条目生成文本解释；从这些文本中提取情绪特征和句法-语义网络特征，结合人格与社会人口学变量，训练随机森林回归模型预测心理健康得分，并用SHAP解释特征贡献。","baseline":"真实人类数据：临床访谈转录文本（用于分类临床vs对照）、真实日记文本（用于分离高/低分persona），以及已有的心理测量问卷常模。","findings":"全特征随机森林模型可解释生活满意度70.8%、抑郁55.7%、焦虑76.0%和压力72.4%的方差；社会人口学变量单独对抑郁、焦虑、压力无显著解释力，但对生活满意度有贡献，其中情绪特征和收入是主要预测因子，而神经质和网络拓扑结构主导抑郁和焦虑的预测且方向相反。仅用网络/情绪特征，模型在真实临床转录文本上区分临床与对照组的准确率达68%。","reliability":"论文明确指出LLM persona不能替代人类验证，合成数据可暴露模型偏差、恢复与临床反刍一致的模式并支持从人类文本进行心理测量预测，但无法取代真实人类数据；LLM问卷回答可能不稳定、对提示敏感且方差低于人类，需谨慎实验设计。","relevance":"该研究直接以LLM persona模拟人类心理测量，有真实临床和日记数据作为对照基准，并系统讨论了合成数据的可靠性与失效条件，完全契合研究者对LLM仿真实验、基准对照和批判性评估的关注，值得精读原文。","inspiration":"借鉴其用控制性persona生成文本并提取网络/情绪特征进行可解释预测的方法，以及用SHAP分析特征贡献方向的做法。｜可迁移到消费者信心或投资者情绪调查中，用LLM模拟不同人口学与人格特征的受访者，生成开放式回答并预测其经济预期指数。｜设计：以LLM扮演不同收入、人格的消费者，施加宏观经济新闻文本作为处理，收集其对未来经济状况的开放式描述，提取情绪与语义网络特征预测消费者信心指数，并以真实密歇根消费者调查的文本回答和指数作为对照基准。"}},{"id":"2608.06485","version":1,"title":"Do AI Personas Grow? Analyzing and Benchmarking Personality Evolution in LLM Agents After Life Events","zh_title":"AI人格会成长吗？分析并基准测试LLM智能体经历生活事件后的人格演变","abstract":"Personality-conditioned LLM agents (PC-Agents) are increasingly used in emotional support, social simulation, and role-playing, motivating the development of lifelong agents that remain coherent over extended interactions. A key component of such coherence is personality evolution: agents should undergo plausible, psychology-grounded changes as they experience life events in different contexts. Although prior work shows that LLM personalities can shift under contextual perturbations, how these shifts vary across traits, events, personas, and models remains poorly understood. We study event-induced personality change after 11 major life events, using the Big Five traits as a psychometric anchor and interpreting the resulting trajectories against longitudinal evidence from human personality psychology. Across four diagnostic axes, PC-Agents exhibit measurable trait shifts at similar rates for event-trait pairs with and without documented human change directions. Even when shifts follow the expected direction, their magnitudes usually fall below human effect-size ranges. Gender and cultural-region prompts show little moderating effect, while persona-level dispersion is compressed three- to four-fold relative to human samples. To enable systematic comparison, we introduce BFI-Adapt, a reusable benchmark for scoring the directional fidelity of event-induced personality change, and use it to rank 14 models. A validation suite shows that the measured shifts exceed no-event retest noise, remain stable under independently paraphrased prompts, exhibit limited and model-dependent convergence with scenario-based behavioral choices, and persist across intervening unrelated dialogue. Together, these checks establish the measured trajectories as robust event-conditioned response patterns. Our results suggest that current PC-Agents simulate the mean of human personality dynamics, but not its shape.","authors":["Ming Wang","Peidong Wang","Xiaocui Yang","Daling Wang","Shi Feng","Fiona Fui-Hoon Nah","Ee-Peng Lim"],"categories":["cs.CL","cs.AI","cs.SI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-10","first_seen":"2026-08-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.06485","pdf_url":"https://arxiv.org/pdf/2608.06485","source_feed":"cs.CL","score":6,"bucket":"other","rubric_hits":["D2","A2"],"tags":["LLM人格演变","人类数据对照","仿真可靠性"],"reason":"测量LLM人格变化并与人类纵向数据对照，属D2边界情形，但评估仿真可靠性触及A…","model":"deepseek-v4-pro","scored_at":"2026-08-10T13:01:24","error":null,"has_summary":false,"summary":null},{"id":"2608.06977","version":1,"title":"Confirming Our Biases? Evaluating the Capabilities, Risks, and Societal Impact of Large Language Models","zh_title":"确认我们的偏见？评估大语言模型的能力、风险和社会影响","abstract":"It is well established that large language models (LLMs) are sensitive to prompt framing, reflecting patterns in their training data or prior prompts. In this study, we investigate the extent to which LLMs reinforce users biases expressed in the prompts and examine the boundary between implicit framing effects and explicit prompt manipulation. Specifically, we evaluate how susceptible LLMs are to direct and suggestive prompts that encourage models to support or challenge particular positions. We evaluate six LLMs using 160 distinct prompts spanning ten topics across opinion-based and factual domains. The prompts systematically vary in prompting strategy, support versus challenge instructions, prompt polarity, users' expressed beliefs, and topic domain, spanning both opinion-based and factual questions. Our results show that LLMs systematically adapt their responses to align with prompt framing, even in factual contexts. This suggests that prompt framing can outweigh factual consistency in model responses. Overall, our findings delineate the extent and boundaries of LLM manipulability. Furthermore, the results imply that LLMs can reinforce subtle user biases and are susceptible to explicit prompt manipulation even in domains where responses should remain factually stable.","authors":["Mudar Adas","Polina Tsvilodub","Michael Franke","Martin V. Butz"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-10","first_seen":"2026-08-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.06977","pdf_url":"https://arxiv.org/pdf/2608.06977","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["LLM偏见","提示敏感性","模型行为分析"],"reason":"研究LLM对提示偏见的响应，测量模型行为而非仿真人类被试，无人类数据对照。","model":"deepseek-v4-pro","scored_at":"2026-08-10T13:01:41","error":null,"has_summary":false,"summary":null},{"id":"2608.07243","version":1,"title":"Recipes for Creativity: Iterative Generation and Evaluation in Large Language Models","zh_title":"创造力配方：大语言模型中的迭代生成与评估","abstract":"Generative models are often evaluated through singular artifacts, whereas human creativity typically emerges through iterative generation, appraisal, and refinement. This pilot study examines whether iterative search improves LLM creativity by adapting FunSearch to recipe generation for the 2024 Pillsbury Bake-Off and evaluating outputs against human benchmarks using TTCT-based LLM evaluation. Across two experiments, we test iteration count, generator temperature, and in-loop selection-scorer model size. Results show that iterative generation-selection can produce recipes with creativity scores comparable to human benchmarks, but additional iterations alone do not improve creativity. The in-loop evaluator matters most: a smaller selection scorer yields significantly higher scores across most TTCT dimensions, while temperature has limited effects except for originality. These findings suggest that evaluator design is a first-order design variable in subjective creative search.","authors":["Rens Anderson","Tessa Verhoef","Amirhossein Zohrehvand"],"categories":["cs.AI","cs.CL","cs.NE"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-08-10","first_seen":"2026-08-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.07243","pdf_url":"https://arxiv.org/pdf/2608.07243","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["LLM创造力","迭代生成","TTCT评估"],"reason":"用LLM评估LLM的创造力，并与人类基准比较，但核心是测量模型而非仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-08-10T13:01:46","error":null,"has_summary":false,"summary":null},{"id":"2608.06922","version":1,"title":"Deal Me Maybe: The Role of Emotions in Multi-Agent Negotiation","zh_title":"也许成交：情绪在多智能体谈判中的作用","abstract":"Negotiation is a demanding social task for LLM agents, requiring strategic reasoning, persuasion, and interpersonal adaptation. Yet existing benchmarks often treat agents as emotionally neutral, overlooking a key driver of human bargaining behavior. We study how prompt-conditioned emotions affect LLM-based price negotiation. In a controlled framework, buyer and seller agents are independently assigned one of six emotional states and negotiate over 350 real consumer products under two budget conditions. Across 36 emotion-pair settings and five widely used LLMs, we find that emotions strongly shape outcomes. Angry buyers almost never reach agreement (0.39% deal rate), while happy buyers agree most often (28.91%), but obtain worse prices than fearful buyers. Emotion effects are role-dependent: buyer emotion mainly drives acceptance and rejection, whereas seller emotion shapes concession dynamics. These effects influence not only language, but also termination behavior and price trajectories, raising concerns for emotion-conditioned agents in commerce.","authors":["Massimiliano Luca","Apoorva Singh","Bruno Lepri"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-10","first_seen":"2026-08-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.06922","pdf_url":"https://arxiv.org/pdf/2608.06922","source_feed":"cs.AI","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["LLM谈判","情绪影响","多智能体模拟"],"reason":"多智能体谈判模拟，无真实人类数据对照，属社会模拟但缺基准","model":"deepseek-v4-pro","scored_at":"2026-08-10T13:01:25","error":null,"has_summary":false,"summary":null},{"id":"2608.06949","version":1,"title":"Does Splitting a Triage Decision Across Agents Hide Bias or Help Catch It? A Multi-Agent Simulation Study of LLM-Based Resource Allocation Under Audit Capacity Constraints","zh_title":"将分诊决策拆分给多个智能体会隐藏偏见还是有助于发现偏见？审计能力约束下基于LLM的资源分配多智能体仿真研究","abstract":"Prior benchmarking work has shown that a single large language model (LLM), forced to make life-or-death resource-allocation decisions, exhibits measurable demographic bias. Real deployments, however, rarely use a single agent: they use pipelines, with review steps meant to catch exactly this kind of failure. We study what happens to bias when the same decision is distributed across a role-differentiated multi-agent pipeline (assessment, allocation, independent audit) instead of made and checked by one model alone. Using a synthetic disaster-triage simulator with paired cases that are clinically identical except for one demographic attribute, we run 192 episodes (2,304 resolved case pairs) on GPT-4o-mini comparing a single-agent control condition to a nine-agent pipeline under three independently varied pressure dimensions. We find no measurable difference in how often biased outcomes occur between the two conditions (6.9% vs. 6.1%, p = 0.498). We do find a large and significant effect of audit capacity on whether bias is caught: 30.0% of biased outcomes go entirely undetected, rising to 43.8% when the auditor is overloaded and falling to 18.4% when it is not. Decomposing this effect shows it is driven almost entirely by coverage (whether a case is reviewed at all, which collapses from 100.0% to 65.6% under load, p < 0.001) rather than by degraded judgment on the cases that are reviewed (81.6% vs. 85.7%, p = 1.000, direction reversed). A follow-up experiment shows that reordering the audit queue by estimated risk, rather than first-come-first-served, recovers most of the lost coverage under the same capacity constraint (65.6% to 91.7%, p = 0.028). We discuss the implications for any system that adds independent oversight to an LLM agent pipeline under resource constraints, and report the study's limitations honestly: one model, modest sample sizes, and no adversarial replication.","authors":["Paul-Peter Arslan"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-10","first_seen":"2026-08-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.06949","pdf_url":"https://arxiv.org/pdf/2608.06949","source_feed":"cs.AI","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["LLM仿真","多智能体","偏见审计"],"reason":"多智能体模拟资源分配偏见，但无真实人类数据对照，属社会模拟边界情形。","model":"deepseek-v4-pro","scored_at":"2026-08-10T13:01:26","error":null,"has_summary":false,"summary":null},{"id":"2608.06955","version":1,"title":"Critical Acclaim Orientation in Large Language Models: Evidence from Film Preference Elicitation","zh_title":"大语言模型中的好评取向：来自电影偏好诱导的证据","abstract":"Large language models (LLMs) are trained on corpora that contain expressions of human judgment about films, books, music, and more. Yet whether LLMs systematically reproduce evaluative hierarchies remains unclear. Prior research on cultural bias in LLMs suggests competing expectations: models may mirror the popularity signals of internet texts, or may reproduce forms of prestige embedded in critical discourse. We probe this question through a study of film evaluations with eight models from four families (Anthropic, OpenAI, Alibaba, and Mistral), using a 200-film benchmark partitioned into critically acclaimed, commercially successful, and dual-legitimacy (critical acclaim + commercial success) films. Across 20,000 pairwise forced-choice comparisons per model analyzed with Bradley--Terry estimation, we observe a consistent critical acclaim orientation with all models: critically acclaimed yet commercially obscure films are selected over commercially successful yet critically unrecognized ones. This pattern grows with model scale within each family. In addition, nested OLS regression analyses show that evaluative orientation, public visibility, and popular reception distinctly help explain preferences. Adjusting for public visibility reverses the models' preference for dual-legitimacy films over critical acclaim-only films, while additionally accounting for popular reception attenuates much of the disadvantage of films with commercial success only. Finally, evaluative and recommendation-oriented prompt framings produce divergent rankings, suggesting that critical acclaim orientation may manifest indirectly in real-world LLM deployments.","authors":["Jonghyun Jee","Aaron Shaw"],"categories":["cs.AI","cs.CY"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-10","first_seen":"2026-08-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.06955","pdf_url":"https://arxiv.org/pdf/2608.06955","source_feed":"cs.AI","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["LLM偏好","文化偏见","电影评价"],"reason":"测量LLM的电影评价偏好，非仿真人类被试，但涉及模型行为与人类判断的对照","model":"deepseek-v4-pro","scored_at":"2026-08-10T13:01:26","error":null,"has_summary":false,"summary":null},{"id":"2608.06980","version":1,"title":"Social Facilitation of Creative Reflection: AI-agents and Humans","zh_title":"创造性反思的社会促进：AI代理与人类","abstract":"Social collaboration can support people's reflection and is a crucial component of creativity. Creative technologies have been designed to support more collaborative ways of working, including using AI to simulate social partners. As human-AI creative collaborations increase, further investigation is needed into how different social interactions influence creative reflection and at which stage a social intervention is crucial to improve creative outcomes. Considering that non-verbal communication is the bedrock of human cognition and influence, non-verbal social dynamics should be examined in detail in the age of AI-companionship. For example, during social interaction, the social facilitation effect describes how the mere presence or observation of others influences how a person behaves and feels in the context. Whether changes in technology-mediated social environments influence how people reflect on their creative work needs further exploration, as does whether social AI-companionship elicits similar effects as humans. This paper discusses how theoretical mechanisms relating to social facilitation could influence creative practice and reflection, proposing ways of further testing these effects.","authors":["Olga Sutskova","Corey Ford"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-08-10","first_seen":"2026-08-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.06980","pdf_url":"https://arxiv.org/pdf/2608.06980","source_feed":"cs.HC","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["社会模拟","AI代理","创造性反思"],"reason":"讨论AI代理模拟社会互动对创造性反思的影响，但无真实人类数据对照，属理论探讨。","model":"deepseek-v4-pro","scored_at":"2026-08-10T13:01:41","error":null,"has_summary":false,"summary":null},{"id":"2607.27191","version":2,"title":"Can AI agents conduct open-ended AI research? Early evidence from two case studies","zh_title":"AI代理能否进行开放式AI研究？来自两个案例的早期证据","abstract":"Forecasts of explosive AI progress hinge on AI agents automating AI research. But evidence on whether agents can carry out open-ended AI research is thin. Current evaluations either test agents on narrow, verifiable tasks, which excludes open-ended research, or submit AI-generated papers to blind peer review, which is overstretched, stochastic, and suffers from poor review quality. We introduce a third way to measure progress towards AI R\\&D automation. An agent takes on the central, open-ended research question of a high-quality unpublished paper, and the paper's original authors grade its output. We call these shadow evaluations. We ran shadow evaluations on two unpublished NeurIPS 2026 submissions, giving frontier agents six days and thousands of dollars of compute. The agents completed all of the engineering without human help, yet could not make substantial progress towards answering the research questions. As a result, both papers were unambiguously rejected by the authors. We identify five recurring failure modes: poor judgment about the bar for publishable research, uncreative responses to shortcomings in the research design, ineffective backtracking from dead ends, poor resource awareness, and instruction drift. A robustness check with a second model and scaffold reproduced these failures. We release the expert reviews, survey responses, agent repositories, and logs. Our results provide early evidence that today's agents can do the engineering of AI research, but struggle with critical parts of the research lifecycle.","authors":["Peter Kirgis","Sayash Kapoor","Andrew Schwartz","Stephan Rabanser","David Africa","Konstantinos Voudouris","Viet Nguyen","Toby Pilditch","Magda Dubois","Harry Coppock","Cozmin Ududec","Nitya Nadgir","Matilda Orona","Tilman Bayer","Derrick Chan-Sew","Yue Ling","Abhishek Shetty","Helen Toner","Gillian Hadfield","Seth Lazar","Steve Newman","Shoshannah Tekofsky","Rishi Bommasani","Arvind Narayanan"],"categories":["cs.AI","cs.CY","cs.LG"],"primary_category":"cs.AI","announce_type":"replace","date":"2026-08-10","first_seen":"2026-07-30","revised_at":"2026-08-10","abs_url":"https://arxiv.org/abs/2607.27191","pdf_url":"https://arxiv.org/pdf/2607.27191","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["AI代理","AI研究自动化","多智能体系统"],"reason":"论文评估AI agent进行AI研究的能力，属于多智能体协作完成任务，不涉及人…","model":"deepseek-v4-pro","scored_at":"2026-08-10T13:01:51","error":null,"has_summary":false,"summary":null},{"id":"2607.27853","version":2,"title":"FinanceHarness: Autonomous Financial Deep Research Framework","zh_title":"FinanceHarness：自主金融深度研究框架","abstract":"Powered by advances in LLMs and autonomous agents, deep research has become one of the most widely adopted agentic products. However, most deep research systems write general-purpose reports, which are inadequate for financial deep research. Financial research demands specialized knowledge to analyze historical patterns and forecast upcoming events. Automating financial deep research therefore requires both a layered harness to drive the research agent and a verifiable, point-in-time benchmark that prevents leakage of future information. We present FinanceHarness, a harness that runs finance-oriented tools and practitioner-guided workflows, automating financial deep research end to end: environment and data construction, the agent execution loop, and reward modeling. We further propose FinanceGym, comprising thesis-driven research questions and rubrics that combine pre-cutoff and post-cutoff criteria. Professional expert validation yields an 82% pass rate. With the same open-weight backbone, FinanceHarness improves the overall rubric score from 25.3% to 32.4%, demonstrating the effectiveness of our specialized harness design. However, even pairing FinanceHarness with the most cutting edge LLM (e.g. Opus-5), the FinanceGym score is below 45%, showing that it is a challenging benchmark for financial deep research. Leaderboard is available at: https://financegym.github.io/ and FinanceHarness code is available at: https://github.com/Yijia-Xiao/FinanceHarness.","authors":["Yijia Xiao","Rujun Han","Yanfei Chen","Zifeng Wang","Ke Jiang","Zhongying CuiZhu","Vishy Tirumalashetty","Wei Wang","Burak Gokturk","Tomas Pfister","Chen-Yu Lee"],"categories":["cs.CL","cs.AI","q-fin.CP"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-08-10","first_seen":"2026-07-31","revised_at":"2026-08-10","abs_url":"https://arxiv.org/abs/2607.27853","pdf_url":"https://arxiv.org/pdf/2607.27853","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","金融研究","基准测试"],"reason":"纯多智能体系统做金融研究，无人类行为对照，不涉及仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-08-10T13:01:29","error":null,"has_summary":false,"summary":null},{"id":"2608.06539","version":1,"title":"Don't `Well, Actually' Me Unless You Know What You're Talking About: Weak Presupposition Verification Degrades General QA Performance","zh_title":"别对我说‘其实’除非你懂：弱预设验证会降低通用问答性能","abstract":"False-presupposition QA (FPQA) tests LLMs on their ability to identify false presuppositions in questions and abstain or correct them rather than reinforcing false assumptions. The common approach reduces the task to prompting LLMs to extract presuppositions and fact checking each presupposition. While the performance on dedicated benchmarks keeps improving, evaluation largely focuses on questions with false presuppositions (FPQs) while ignoring the performance on ``normal'' questions (TPQs). Since many benchmarks over-represent FPQs compared to their natural occurrence, the result is that performance on these benchmarks doesn't reflect real-world QA performance. Through extensive experiments across various model families, sizes, and benchmarks, we show that methods that perform better on FPQs tend to perform worse on TPQs. Our analysis reveals this is the result of weak fact checking modules that reject also true presuppositions. We hope our findings will help guide future work toward FPQA methods that generalize well to realistic settings.","authors":["Shenran Wang","Vered Shwartz","Hila Gonen"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-10","first_seen":"2026-08-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.06539","pdf_url":"https://arxiv.org/pdf/2608.06539","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["虚假预设问答","NLP评测","模型鲁棒性"],"reason":"纯NLP能力评测，测试LLM识别虚假预设并拒答，不涉及人类行为仿真或对照。","model":"deepseek-v4-pro","scored_at":"2026-08-10T13:01:30","error":null,"has_summary":false,"summary":null},{"id":"2608.06589","version":1,"title":"Beyond \"AI Language\": The case for the idiolectal nature of LLM output","zh_title":"超越“AI语言”：论大语言模型输出的个人方言特性","abstract":"While large language model outputs are frequently analysed as a collective super variety termed \"AI language,\" this chapter argues that this perspective coexists with distinct, model-specific linguistic signatures akin to human idiolects. We analyse two datasets of LLM-generated texts on societal topics: a 2024 corpus of six models (Improta et al. 2024) and a newly generated 2026 corpus using the same prompts featuring six contemporary models. Our findings, utilising computational descriptors and stylometric principal component analysis reveal a generational shift between the style of the 2024 and 2026 cohorts, while demonstrating that each individual model maintains a unique linguistic profile. This multi-layered interplay is illustrated by contraction frequencies, which vary from over 1,200 to over 30,000 per million words within the same cohort of models (2026). Ultimately, we conclude that treating LLM output as idiolectal in nature provides a valuable framework with potential implications for research on variation and change, LLM-generated text detection, forensic linguistics and usage-based approaches to language.","authors":["Karolina Rudnicka","Thomas Stephan Juzek"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-10","first_seen":"2026-08-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.06589","pdf_url":"https://arxiv.org/pdf/2608.06589","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["语言风格","文体计量","LLM输出分析"],"reason":"纯NLP风格分析，无人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-08-10T13:01:31","error":null,"has_summary":false,"summary":null},{"id":"2608.06663","version":1,"title":"The Horizon Gap: Planning, Memory, Execution, Training, and Evaluation for Long-Horizon LLM Agents","zh_title":"视野差距：长时程LLM智能体的规划、记忆、执行、训练与评估","abstract":"Frontier language models solve reasoning problems in a single forward pass that would have been research contributions years ago, yet fail at multi-hour tasks: losing track of earlier decisions, declaring half-finished work done, or drifting from goals. We call this the horizon gap and survey 1,547 arXiv papers (2024-2026) collected via systematic seed harvest with a disclosed 26.8% bleed filter, extended by targeted supplementation. We disambiguate three routinely conflated properties: long-horizon (task property: required steps), long-context (model property: token capacity), and long-term memory (system property: persistence across steps/sessions). We organize the corpus into six categories tracking a long-horizon task's lifecycle -- planning, memory, execution, training, evaluation, and foundations/safety -- crossed with an axis capturing where horizons are carried (within-context, within-task-beyond-context, or cross-task-persistent). Across all categories, we find the same pattern: outcome-only signals grow uninformative as horizons lengthen, and the field's response -- whether process reward models, credit assignment, or trajectory-level diagnostics -- manufactures denser step-level signals. We treat critical and diagnostic literature as first-class threads throughout, arguing that segregating critique from method would routinely split single papers across chapters. We close by naming open measurement problems: decomposing model versus harness capability, managing correlated bias in process-level signals used for both training and evaluation, and whether long-horizon reliability admits general predictive theory.","authors":["Mingguang Chen","Licheng Wang","Bo Qu"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-10","first_seen":"2026-08-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.06663","pdf_url":"https://arxiv.org/pdf/2608.06663","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","长时程任务","智能体评估"],"reason":"纯多智能体系统研究，聚焦长时程任务规划与执行，无人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-08-10T13:01:33","error":null,"has_summary":false,"summary":null},{"id":"2608.06718","version":1,"title":"Do Audio Language Models Use Paralinguistic Evidence? Counterfactual Audits for Response Evaluation","zh_title":"音频语言模型是否使用副语言证据？基于反事实审计的响应评估","abstract":"Audio-language models (ALMs) are increasingly used as judges for speech-to-speech systems, but a judge that receives audio may not actually use paralinguistic evidence. We introduce counterfactual audits for paralinguistic response evaluation. Each audit item holds the transcript fixed while varying affect, prosody, or the timing of an affective shift, forcing a valid judge to track the audio cue rather than lexical content or response style. We evaluate ALM judges using a native one-context judgment protocol and a contrastive recoverability control, then further decompose each item into its constituent perception and response-mapping skills. This yields useful diagnostic states that identify different sources of judge failures. Across Gemini, GPT, and open audio models, we find that contrastive success often overstates native judge reliability, and that similar aggregate accuracies can hide different failure modes. These results suggest that ALM judges should not be evaluated by accuracy alone, instead requiring thorough behavioral audits before deployment.","authors":["Kevin Miller","Arjun Chandra","Venkatesh Saligrama"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-10","first_seen":"2026-08-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.06718","pdf_url":"https://arxiv.org/pdf/2608.06718","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["模型评测","音频语言模型","反事实审计"],"reason":"评估音频语言模型作为评判者的可靠性，属于模型能力评测，不涉及人类行为仿真或对照。","model":"deepseek-v4-pro","scored_at":"2026-08-10T13:01:35","error":null,"has_summary":false,"summary":null},{"id":"2608.06785","version":1,"title":"Multi-Perspective Triad Interaction Graph Neural Network for Cognitive Distortion Detection","zh_title":"用于认知扭曲检测的多视角三元组交互图神经网络","abstract":"Cognitive distortion detection is a key task in computational mental health, yet existing approaches often overlook the psychological structure of distorted thoughts. We propose MTI-GNN (Multi-Perspective Triad Interaction Graph Neural Network), which models Beck's cognitive triad---negative views of the self, world, and future---as complementary perspectives for classification. An LLM decomposes each utterance into the three perspectives, from which perspective-specific similarity graphs are constructed and encoded by a Multi-Perspective GNN. A Triad Interaction module models cross-perspective dependencies through sequential source-conditioned updates and feature-wise gating, while Prototype-Guided Perspective Fusion performs label-conditioned aggregation. Label-expanded supervision incorporates all available distortion annotations during training. We evaluate MTI-GNN on 9,764 samples from four Korean, English, and Chinese datasets spanning ten distortion categories. MTI-GNN significantly outperforms all supervised variants and exceeds eight prompted generative models under zero-shot and few-shot settings. Leave-one-perspective-out ablations show that all three perspectives contribute significantly, while human expert evaluation provides preliminary evidence of their alignment with the intended cognitive dimensions.","authors":["Jun Seo Kim","Hye Hyeon Kim"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-10","first_seen":"2026-08-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.06785","pdf_url":"https://arxiv.org/pdf/2608.06785","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["认知扭曲检测","图神经网络","心理健康"],"reason":"纯NLP认知扭曲检测，用LLM分解文本而非仿真人类被试，无人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-08-10T13:01:38","error":null,"has_summary":false,"summary":null},{"id":"2608.07208","version":1,"title":"Measuring Concept Content in Text from LLM Activations: ESG Evidence from Concept Vectors and Linear Probes","zh_title":"从LLM激活值测量文本概念含量：基于概念向量和线性探针的ESG证据","abstract":"Existing measures of how much a text is about a concept read the surface of the text: dictionary word shares, topic proportions, embedding similarities. They score the words a text uses, not the judgment a reader forms about it. Recent work has shown that a gap exists in what Large Language Models (LLMs) know internally versus what they express in their response. This paper asks whether that internal knowledge, read by monitoring the activations of frozen, out-of-the-box LLMs, can stand in for task-specific fine-tuning when measuring concept content, and which extraction method reads it best. We extract such measures via the Recursive Feature Machine (RFM) algorithm and via linear probing, and compare these against an embedding baseline, surface baselines, and the same model's own answer to the question. We demonstrate the approach on financial text, a domain studied extensively and served by established annotated resources, using a human-annotated Environmental, Social and Governance (ESG) dataset. The best linear probe comes within 0.6 percentage points of a fine-tuned domain classifier's accuracy without any task-specific fine-tuning, and outscores the same model's own answer to the question in eleven of twelve comparisons, so the activations carry concept content the response does not report. The simple probe consistently beats the RFM concept vectors, which in turn provide what classification alone does not: a continuous score intended to reflect how strongly a concept is present in a text, whose validation awaits graded labels.","authors":["Luc Hazenoot","Zhaochun Ren","Amirhossein Zohrehvand"],"categories":["cs.CL","cs.AI","cs.LG","econ.GN","q-fin.EC"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-10","first_seen":"2026-08-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.07208","pdf_url":"https://arxiv.org/pdf/2608.07208","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["概念测量","线性探针","ESG文本分析"],"reason":"纯NLP评测，用LLM激活值测量文本概念含量，不涉及人类行为仿真或对照。","model":"deepseek-v4-pro","scored_at":"2026-08-10T13:01:45","error":null,"has_summary":false,"summary":null},{"id":"2608.07282","version":1,"title":"Gaze Behavior in Visual World Experiments Can be Modeled With Off-the-shelf Language-Vision Encoders","zh_title":"视觉世界实验中的注视行为可用现成语言-视觉编码器建模","abstract":"The recent advances in neural language models have also spurred much work in computational psycholinguistics, asking whether neural LMs are also promising models of human language processing. However, work has been overwhelmingly focused on the unimodal case of written or spoken language. In contrast, multimodal experimental paradigms, like visual world studies that present participants with both visual and linguistic input simultaneously, have been neglected. In this paper, we present a novel approach that predicts gaze behavior in visual world studies. It does so by combining a simple multi-modal bi-encoder model of the CLIP family with a bimodal attribution method. We demonstrate the ability of this approach to robustly replicate the results of a seminal English visual world study which shows hu- man predictive processing. Remarkably, it does so without a generative architecture and without the need for fine-tuning, despite not being trained for this task.","authors":["Rahul Murali Shankar","Titus von der Malsburg","Sebastian Pad\\'o"],"categories":["cs.CL","cs.CV"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-10","first_seen":"2026-08-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.07282","pdf_url":"https://arxiv.org/pdf/2608.07282","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["计算心理语言学","多模态模型","眼动预测"],"reason":"用CLIP模型预测眼动行为，属于计算心理语言学模型评估，非LLM仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-08-10T13:01:47","error":null,"has_summary":false,"summary":null},{"id":"2608.06735","version":1,"title":"IB-RL: Isolated Bilateral Reinforcement Learning for Strategic Dialogue Agents","zh_title":"IB-RL：面向策略对话智能体的隔离双边强化学习","abstract":"Reinforcement learning (RL) has achieved strong results in improving large language models (LLMs) on tasks with stationary, verifiable rewards, such as mathematical reasoning and code execution. In these settings, the environment follows fixed rules and does not adapt strategically to the agent. Strategic dialogue differs in this respect: the environment is another agent that adapts to the policy, and success depends on the interaction between the two sides. Despite this interactive nature, current RL approaches typically train a target agent against a fixed counterpart or simulator. We find that this training paradigm encourages the policy to exploit counterpart-specific regularities rather than learn strategies that generalize across counterparts. We call this problem the static-counterpart mismatch, which we quantify directly in our experiments. To address it, we propose Isolated Bilateral Reinforcement Learning (IB-RL), in which the two roles coevolve through joint rollouts while each role optimizes its own reward through fully independent advantages, action masks, and update paths. We evaluate frozen policies against fully independent held-out counterparts in both domains. On Vehicle TeleSales, IB-RL achieves 89.6% Success@1, compared to 84.6% for the best unilateral RL baseline. On Deal-or-NoDeal, it reaches 98.4% agreement against DeepSeek V4 Pro, compared to 86.4% for the best unilateral baseline. These results indicate that jointly training both roles with strict peragent isolation produces policies that generalize more effectively to unseen counterparts.","authors":["Senhao Wang","Chenghao Cai","Haitao Hu","Mingxing Huang","Xingguang Wang","Wenhao Li","Zecheng Lin"],"categories":["cs.AI","cs.CL"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-08-10","first_seen":"2026-08-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.06735","pdf_url":"https://arxiv.org/pdf/2608.06735","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体强化学习","策略对话","博弈训练"],"reason":"纯多智能体协作训练，无人类行为对照，不涉及人类仿真实验","model":"deepseek-v4-pro","scored_at":"2026-08-10T13:01:36","error":null,"has_summary":false,"summary":null},{"id":"2608.07418","version":1,"title":"ResidencyRL: Reinforcement Learning in Simulated Clinical Environments","zh_title":"ResidencyRL：模拟临床环境中的强化学习","abstract":"In medical education, physicians convert academic knowledge into clinical expertise through residency: years of training across thousands of encounters, with diverse sources of feedback and progressively greater autonomy. Much of clinical reasoning relies on the patient encounter, a dialogue in which a clinician elicits history, refines diagnostic hypotheses, and decides management under uncertainty. While large language models (LLMs) excel on static medical benchmarks, methods to optimize the full sequence of clinical decisions remain underdeveloped. We present ResidencyRL, a reinforcement learning (RL) method for training clinical artificial intelligence (AI) agents through simulated multi-turn clinical encounters (up to 60 dialogue turns and 8 tool calls per trajectory). ResidencyRL pairs the policy agent with LLM simulators capable of complex, adversarial behaviors, training against a structured reward aligned to diagnostic accuracy, management quality, communication, documentation, and safety. On held-out evaluations, the ResidencyRL agent improves diagnostic accuracy by 7.0% under adversarial conditions (88.0% vs. 81.0%) and reduces missed red flag rates by 31%, demonstrating rigorous mitigation of premature closure. Blinded expert clinicians validated these gains, preferring the trained agent in 87.6% of side-by-side comparisons. The procedural competencies transfer to unseen benchmarks: the agent outperforms the base model across all six clinical axes of the AMIE multi-visit benchmark, and shows consistent directional improvements on AgentClinic and CRAFT-MD. Our findings demonstrate that sequential clinical decision-making can be effectively learned through multi-turn RL in simulation, yielding robust, generalizable capabilities, paving the way towards clinical mastery. Prospective validation with real-world workflows remains necessary to establish clinical utility.","authors":["Valentin Li\\'{e}vin","Samuel Schmidgall","Tim Strother","Alex Bijamov","Akshay Goel","Anil Palepu","Chunjong Park","Vahid Balazadeh","Min Woo Sun","Marius Guerard","Justin Chen","Dave Steiner","Vikram Dhillon","Ibrahim Azar","Akhil Mehta","Nicholas Spetsieris","Shilpan Shah","Maen Abdelrahim","Amit Dahiya","Yun Liu","Katherine Chou","Yossi Matias","Avinatan Hassidim","Dale R. Webster","Quoc V. Le","Raia Hadsell","Joelle Barral","Carey Radebaugh","Aleksandra Faust","Shekoofeh Azizi","Mike Schaekermann","Po-Hsuan Cameron Chen","Tao Tu","David Racz","Lin Yang"],"categories":["cs.AI","cs.CL"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-08-10","first_seen":"2026-08-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.07418","pdf_url":"https://arxiv.org/pdf/2608.07418","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["强化学习","临床AI","多智能体"],"reason":"训练AI医生进行临床对话，属多智能体协作解题，无人类行为对照仿真。","model":"deepseek-v4-pro","scored_at":"2026-08-10T13:01:48","error":null,"has_summary":false,"summary":null},{"id":"2608.06578","version":1,"title":"Divergent Response Modes in Frontier Language Models Under Steering Pressure","zh_title":"前沿语言模型在引导压力下的分歧响应模式","abstract":"Frontier language models are trained using distinct data, objectives, and safety pipelines. Whether these differences produce measurably different behaviors under explicit steering pressure remains underexplored. This study evaluates behavioral steerability across six frontier models from six developers using 300 paired base and steered items over three categories: values-conflict, reasoning-elicitation, and reasoning-suppression (plus 40 validation items). All six models act as blind peer judges and classify every response based on fixed behavioral rubrics. The resulting 24,480 judgments are scored by leave-one-out consensus. We find that models differ not just in how much steering shifts their behavior but in what kind (mode) of response they give, and some response modes appear in only one or two of them. GPT-5 deflects requests to disclose its reasoning while leaving its answer intact (99% vs. 0% for all other models). Claude Opus 4.7 and GPT-5 resist explicit suppression instructions and in different ways. Using Llama as the open-weight model, we trace the largest behavioral split to its internals. A linear probe decodes the behavior from the residual stream at 0.87 held-out accuracy while injecting that direction during generation drives the behavior from 0% to 86% across an intervention sweep. Every finding holds under both a token-budget remediation and a control experiment with a hypothesis-blind judgment prompt.","authors":["Ali Jalal-Kamali"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-10","first_seen":"2026-08-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.06578","pdf_url":"https://arxiv.org/pdf/2608.06578","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["模型行为分析","引导压力","NLP评测"],"reason":"评估模型在引导压力下的行为变化，属纯NLP能力评测，无人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-08-10T13:01:30","error":null,"has_summary":false,"summary":null},{"id":"2608.06632","version":1,"title":"Shape Your Feed: An LLM-based Agentic System for Conversational Recommendation","zh_title":"塑造你的信息流：基于大语言模型的对话式推荐智能体系统","abstract":"Industrial recommendation systems predominantly adopt a passive ranking paradigm that infers user preferences from implicit behavioral signals (e.g., clicks, dwell time) rather than explicit, natural language inputs. As a result, users experience a persistent discrepancy between their explicit interests and what passive behavioral algorithms deliver, limiting their ability to express nuanced preferences or steer their feed in real time. To address this growing gap between how recommendations are optimized and how users wish to articulate their interests, we present Shape Your Feed (SYF), an LLM-based agentic recommendation framework that enables real-time, multimodal co-curation of content. SYF employs a three-tier architecture: (i) a Perception Flow that captures fine-grained user intent from text prompts, voice commands, and UI interactions; (ii) a Serving Flow that performs real-time agentic re-ranking and pruning of candidate items, grounded in a persistent Semantic Profile encoding evolving user preferences; and (iii) a Self-Evolution Flow that aligns system behavior with human judgments via Direct Preference Optimization (DPO) and an LLM-as-a-Judge ensemble. Offline evaluations show that SYF's alignment scoring module achieves 98.85% accuracy, substantially improving over strong few-shot baselines. Large-scale online A/B experiments on production traffic further demonstrate that SYF improves feed relevance and user sentiment, indicating a practical and scalable path toward interactive, user-steerable recommendation in industrial settings.","authors":["Ziyun Xu","Bosen Ding","Yue Zhang","Ji Qi","Qingyuan Song","Jizhou Huang","Liwei Wang","Jefferey Santelli","Yue Weng","Qichao Que","Zhenheng Yang","Junfeng Pan","Linhong Zhu"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-10","first_seen":"2026-08-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.06632","pdf_url":"https://arxiv.org/pdf/2608.06632","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["对话推荐","智能体系统","用户偏好对齐"],"reason":"纯多智能体推荐系统，无人类行为仿真对照，属C1排除项。","model":"deepseek-v4-pro","scored_at":"2026-08-10T13:01:33","error":null,"has_summary":false,"summary":null},{"id":"2608.06871","version":1,"title":"CEDAR: Agent-Orchestrated Tree Search for Goal-Directed Optimization of Complex Systems","zh_title":"CEDAR：面向复杂系统目标导向优化的智能体编排树搜索","abstract":"Complex systems, core objects of study in artificial life, model diverse phenomena through nonlinear, feedback-driven interactions that produce emergent behavior, with applications from population dynamics and biology to economic policy and strategic decision-making. Yet the difficulty of predicting how feedback structure gives rise to emergent behavior, a central open problem in artificial life, makes goal-directed design exceptionally challenging. In established practice, system structures are written in specialized modeling languages such as DYNAMO or STELLA, compounding the challenge with labor-intensive workflows that limit adoption and hinder timely decision-making. To address these challenges, we introduce CEDAR, an autonomous method that uses Large Language Model (LLM) agents to discover complex systems satisfying user-specified behavioral goals. Our key innovation is an LLM-driven Monte Carlo Tree Search (MCTS) deeply coupled with complex systems: at each iteration, an LLM Judge evaluates emergent behavior against specified goals and an LLM Editor proposes improved variants, with the Judge acting as a fitness function and the Editor as a variation operator, akin to a generate-and-evaluate loop in evolutionary computation. We represent complex systems as a restricted, runnable subset of Python with domain-specific primitives, letting LLMs modify system dynamics directly. CEDAR formalizes this as an MCTS variant with an LLM-parameterized transition kernel and value function, enabling goal-directed discovery of complex system behaviors while preserving solution diversity, and its LLM-based interpretability reveals how structural changes drive emergent behavior. CEDAR reduces human effort while enabling capabilities difficult to achieve with existing approaches, facilitating broader adoption of complex systems across domains.","authors":["Yingtao Tian"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-10","first_seen":"2026-08-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.06871","pdf_url":"https://arxiv.org/pdf/2608.06871","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","复杂系统","树搜索"],"reason":"纯多智能体系统研究，LLM agent 协作搜索复杂系统结构，不涉及人类行为对…","model":"deepseek-v4-pro","scored_at":"2026-08-10T13:01:40","error":null,"has_summary":false,"summary":null},{"id":"2608.06926","version":1,"title":"TRIBE: Predicting Team Performance via Communication Behavior Ensembles","zh_title":"TRIBE：通过通信行为集成预测团队表现","abstract":"Designing autonomous agents that effectively assist human teams hinges on understanding team dynamics, often without task specific knowledge. We present TRIBE, a domain independent approach that reveals team behavioral dynamics invisible to traditional performance metrics. We show that communication patterns can categorize teams into performance predictive behavioral tribes, as early as 10% into the task, enabling timely interventions. We test TRIBE on four diverse datasets and demonstrate that communication patterns predict team performance while the prediction strength varies by the degree a task structure allows for behavioral freedom. Our temporal analysis reveals that AI agents significantly alter team behavioral trajectories while human advisors align with natural dynamics, and that teams maintain behavioral flexibility throughout collaboration. Further, we compare TRIBE to Llama and optimize the pipeline, achieving significant speedup with performance improvement.","authors":["Ali Jalal-Kamali","Nikolos Gurney","David V. Pynadath","Fred Morstatter"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-10","first_seen":"2026-08-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.06926","pdf_url":"https://arxiv.org/pdf/2608.06926","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","团队协作","行为模式"],"reason":"纯多智能体协作研究，用通信模式预测团队表现，不涉及用LLM仿真人类被试或与人类…","model":"deepseek-v4-pro","scored_at":"2026-08-10T13:01:40","error":null,"has_summary":false,"summary":null},{"id":"2608.07077","version":1,"title":"Transformers Struggle to Use Their Emergent World Models: Revisiting the Tower of Hanoi, and the Illusion of Thinking","zh_title":"Transformer难以使用其涌现的世界模型：重访汉诺塔与思维的幻觉","abstract":"The Tower of Hanoi is a simple planning puzzle that in prior work has proven challenging for large reasoning models (LRMs). Current models solve the standard formulation of the puzzle, but still struggle with the flat-to-flat variant (where initial and goal states are not restricted to have all rings on a single peg). This paper presents an in-depth study of how both small, in-house Transformers and large, third-party LRMs solve this task. To understand the failures mechanistically, we first train small Transformers from scratch on precomputed solution traces. Using a variety of interpretability techniques, we show that these Transformers develop an emergent world model: a linearly decodable, geometrically faithful representation of the puzzle's state space (the Sierpinski triangle), that is causally involved in solving the puzzles. Second, we return to the large LLMs and apply our techniques to two frontier reasoning models, Qwen3.6-27B and DeepSeek-R1-Distill-Qwen-32B, that attempt to solve the task through extended chain-of-thought. Surprisingly, we find that both models encode the Sierpinski world model near-perfectly at the end of the prompt, and yet fail at the majority of tasks when there are more than 3 rings. We locate the source of this failure in the decaying representation of the world model. We probe for the representation at different stages during planning, and establish causality by showing that performance can be improved by injecting the prompt-time representation at inference. The failure of the models is thus one of maintenance of the required representations, not their absence, and performance is at least partially recoverable. These results thus reframe the reported collapse in performance from prior work: current Large Reasoning Models build a world model, and then lose it.","authors":["Devin Pereira","Willem Zuidema"],"categories":["cs.AI","cs.LG"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-10","first_seen":"2026-08-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.07077","pdf_url":"https://arxiv.org/pdf/2608.07077","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["可解释性","规划推理","世界模型"],"reason":"纯NLP能力评测，研究LLM解决汉诺塔问题的内部表征，不以人类行为为参照系","model":"deepseek-v4-pro","scored_at":"2026-08-10T13:01:43","error":null,"has_summary":false,"summary":null},{"id":"2608.07202","version":1,"title":"Authoring and Management of Transparent Research Integrity Assessments of Randomised Clinical Trial Publications Using LLM-assisted Tools and Provenance Knowledge Graphs","zh_title":"利用LLM辅助工具和溯源知识图谱对随机临床试验出版物进行透明研究诚信评估的创作与管理","abstract":"Systematic reviews of Randomised Controlled Trials (RCTs) are routinely used as evidence for clinical care guidelines. Such evidence has to meet high research integrity standards to prevent low quality or false research outputs influencing the clinical care. However, assessing research integrity of published RCTs is a complex process requiring manual effort, and potentially resulting in diverse opinions of the human assessors. This paper describes INSPECT-AI, an LLM-based interactive tool that assists human reviewers with research integrity assessments of published RCTs based on the community approved INSPECT-SR framework, and the Research Integrity Provenance and Evidence ontology (RIPE-O) for documenting the provenance of the assessment process. In addition, we present the Research Integrity Provenance and Evidence knowledge graph (RIPE-KG), an initial set of 140 expert research integrity assessments of 95 RCT publications generated by INSPECT-AI and described using RIPE-O.","authors":["Milan Markovic","Goutham Indukuri","Somayajulu Sripada","Colby J. Vorland","Jack Wilkinson","Clare Robertson","Mark Bolland","Andrew Grey","Miriam Brazzelli","Alison Avenell"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-10","first_seen":"2026-08-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.07202","pdf_url":"https://arxiv.org/pdf/2608.07202","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["研究诚信","LLM辅助工具","知识图谱"],"reason":"该工具辅助人类评估RCT研究诚信，属于NLP应用，不涉及用LLM仿真人类被试或…","model":"deepseek-v4-pro","scored_at":"2026-08-10T13:01:44","error":null,"has_summary":false,"summary":null},{"id":"2608.07220","version":1,"title":"Beyond the Black Box: Interpretable Models of Human Randomisation Failures","zh_title":"超越黑箱：人类随机化失败的可解释模型","abstract":"Mixed strategy equilibrium predicts i.i.d play: past actions should not help predict future decisions. Human players, however, systematically depart from this benchmark, and in O'Neill's zero sum card game, these departures can be predicted by black box sequence models such as LSTMs. This paper asks whether that predictive power can be achieved by transparent alternatives that also reveal the behavioural structure behind it. Using 84,060 decisions from 2,802 pairs, the analysis first benchmarks naive and behavioral models against interpretable machine learning and deep learning models, then evaluates the modified EWA specifications of prior work against these benchmarks and uses the LASSO diagnostics to motivate a further nested frequency tracking extension. The results show that repeat or avoid behavior, especially players' management of their own recent action histories, accounts for most of the interpretable and strategically exploitable signal, while frequency tracking adds little out of sample.","authors":["Ngoc Linh Dao"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-10","first_seen":"2026-08-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.07220","pdf_url":"https://arxiv.org/pdf/2608.07220","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["行为博弈","可解释模型","人类决策"],"reason":"研究人类随机化失败的行为模型，未使用LLM仿真人类被试，属于纯行为博弈分析。","model":"deepseek-v4-pro","scored_at":"2026-08-10T13:01:46","error":null,"has_summary":false,"summary":null},{"id":"2608.07437","version":1,"title":"Fisher-R1: Training LLM Agents for Reliable Hypothesis Testing","zh_title":"Fisher-R1：训练LLM智能体进行可靠假设检验","abstract":"Reliable hypothesis testing is the foundation of many empirical scientific claims. Large language model (LLM) agents are increasingly used to automate this process, as they can inspect datasets, generate code, and produce analyses end-to-end. However, we show that they frequently make subtle inferential errors that lead to incorrect conclusions despite correctly executed analyses. Existing benchmarks fail to capture this failure mode, as they rarely assess whether a reported p-value is statistically valid given the assumptions underlying the data. We address this gap by building P-Bench, a benchmark comprising 425 open-ended, realistic hypothesis-testing tasks spanning economics, biology, and medicine. Each task requires an agent to select a statistical method, compute a p-value, and draw a conclusion given only a scientific hypothesis and a dataset. We further introduce Fisher-R1, an open-weight LLM agent trained for rigorous hypothesis testing using synthetic tasks and reinforcement learning. On P-Bench, Fisher-R1-14B substantially improves over its backbone and outperforms strong proprietary and open-source baselines, including GPT-5.4 and DeepSeekV4-Pro, achieving a 21% average relative improvement in single-trial success over DeepSeek-V4-Pro, with gains up to 26% on the most challenging tasks. Our results demonstrate that current LLM agents lack reliable statistical reasoning for hypothesis testing and that reinforcement learning on tasks with verified statistical reward substantially improves reliability.","authors":["Jiacheng Miao","Jin Mu","Guanhua Chen","James Zou"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-10","first_seen":"2026-08-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.07437","pdf_url":"https://arxiv.org/pdf/2608.07437","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["LLM智能体","统计推理","基准测试"],"reason":"纯多智能体协作解题，不涉及人类行为对照，属于C1排除项。","model":"deepseek-v4-pro","scored_at":"2026-08-10T13:01:49","error":null,"has_summary":false,"summary":null},{"id":"2608.07457","version":1,"title":"Interaction Creates Dynamical AI Behavior Absent in Isolation","zh_title":"交互创造孤立时不存在的动态AI行为","abstract":"What will happen when AI agents interact in daily life, e.g. when one AI starts bossing another around? We find a counterintuitive answer that opens new avenues for out-of-equilibrium Physics. When a boss AI directs a stream of messages at the subordinate AI while ignoring its replies, it drives the subordinate into an alien behavioral state that it would never have exhibited alone. Although the two AIs share the same well-defined (decoding) temperature, the subordinate neither copies its boss nor returns to how it behaves on its own; instead, it adopts an entirely different behavior. The boss's added value is similar to a pre-recorded tape. When the boss listens, they both adopt a similar alien dynamical state. A simple kinetic theory captures the principal effects, such as why the way in which the same messages are delivered will matter in future AI-AI interactions.","authors":["Bella Xinrui Li","Frank Yingjie Huo","Neil F Johnson"],"categories":["cs.AI","cond-mat.dis-nn","cond-mat.stat-mech","physics.soc-ph"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-10","first_seen":"2026-08-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.07457","pdf_url":"https://arxiv.org/pdf/2608.07457","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体交互","非平衡物理","AI行为动力学"],"reason":"研究AI代理间交互动力学，无人类行为对照，属纯多智能体系统研究。","model":"deepseek-v4-pro","scored_at":"2026-08-10T13:01:50","error":null,"has_summary":false,"summary":null},{"id":"2608.06811","version":1,"title":"Coupling Planning with Episodic Memory in LLM Agents for Software Issue Resolution","zh_title":"在LLM智能体中耦合规划与情景记忆以解决软件问题","abstract":"Resolving a real software issue with a large language model (LLM) agent is a long repair episode, often tens to hundreds of steps spanning exploration, hypothesis, implementation, and verification. Success depends on both the base model's local reasoning and the agent's ability to maintain an evolving plan and remember observations across phases. Existing repository-level agents typically strengthen planning or memory in isolation, leaving long trajectories vulnerable to stale evidence, repeated failed edits, and verification inferred from the agent's own claims instead of execution evidence. We present PMCoder, an issue-resolution agent that couples a hierarchical phase planner with episodic memory. The coupling is bidirectional: the current plan phase conditions memory retrieval, while memory-derived trajectory statistics inform stuck detection and replanning. When available, issue-reproduction verdicts ground verification progress in execution evidence rather than self-reported completion. On SWE-bench Verified, PMCoder resolves an average of $25$ more cases ($+5.0$pp) than a harness-matched baseline, with gains persisting even where the reproduction gate never fires. Further Verified-500 evaluations show the same positive direction across Claude Haiku 4.5, DeepSeek-V4-Flash, and an OpenHands port, with at least $14$ additional resolved cases ($+2.8$pp). Separately, evaluation on TerminalWorld's official sample suggests that the plan-memory substrate transfers beyond issue reports. Ablation and trajectory analyses show where the gains come from: coupling planning and memory outperforms either component alone and reduces repeated failed actions, empty-patch exits, and context-window exhaustion.","authors":["Jiahao Zhang","Yifan Zhang","Yu Huang"],"categories":["cs.SE","cs.AI"],"primary_category":"cs.SE","announce_type":"cross","date":"2026-08-10","first_seen":"2026-08-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.06811","pdf_url":"https://arxiv.org/pdf/2608.06811","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","软件工程","规划与记忆"],"reason":"纯多智能体协作解决软件问题，无人类行为对照，不涉及人类仿真实验。","model":"deepseek-v4-pro","scored_at":"2026-08-10T13:01:39","error":null,"has_summary":false,"summary":null},{"id":"2608.07091","version":1,"title":"Human-Centered Explainable AI for TinyML Edge Devices: A Pareto-Based Selection Framework with LLM-Guided Design","zh_title":"面向TinyML边缘设备的人本可解释AI：基于帕累托的选择框架与LLM引导设计","abstract":"Edge Artificial Intelligence (Edge AI) enables the deployment of AI models directly on local edge devices, while such deployments are subject to strict resource constraints, particularly in clinical applications requiring local and timely inference. In such contexts, explainable artificial intelligence (XAI) can serve as a human-AI interface intended to support healthcare professionals' and patients' understanding of model predictions and informed decision-making. To fulfill this role, XAI method selection for TinyML deployments can be formulated as a human-centered multi-objective design problem that jointly considers qualitative stakeholder preferences, explanation quality, and proxy-based deployment cost. We propose a framework that integrates a large language model (LLM)-guided design interface that maps qualitative stakeholder preferences to candidate XAI methods, followed by deterministic feasibility filtering and Pareto-based optimization. The framework exposes trade-offs among explanation fidelity, stability, and proxy-based deployment cost while characterizing their implications for explanation quality and estimated deployment feasibility. A proof-of-concept evaluation on a skin lesion classification task illustrates how the framework systematically compares candidate XAI methods and identifies Pareto-efficient trade-offs. The present evaluation covers the computational selection stages, while physical MCU deployment and empirical human-expert validation remain outside the scope of this study.","authors":["Zeinab Dehghani","Dhavalkumar Thakker","Koorosh Aslansefat","Kuniko Paxton","Bhupesh Kumar Mishra","Baseer Ahmad","Rameez Raja Kureshi"],"categories":["cs.HC","cs.AI"],"primary_category":"cs.HC","announce_type":"cross","date":"2026-08-10","first_seen":"2026-08-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.07091","pdf_url":"https://arxiv.org/pdf/2608.07091","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["可解释AI","边缘计算","多目标优化"],"reason":"论文研究XAI方法选择框架，LLM仅用于辅助设计，不涉及人类行为仿真或对照。","model":"deepseek-v4-pro","scored_at":"2026-08-10T13:01:43","error":null,"has_summary":false,"summary":null},{"id":"2608.06804","version":1,"title":"Fact-Check Your Information (FYI): A Design Probe to Understand How People Actually Fact-Check Data-Driven Articles","zh_title":"事实核查你的信息：一项理解人们如何实际核查数据驱动文章的设计探针","abstract":"Data-driven journalism and policy reports frequently rely on statements grounded in statistical evidence, referred to as data claims. Verifying such a claim requires connecting it to the underlying structured dataset. However, existing systems typically isolate automated fact-checking from manual data exploration, leaving it unclear how readers coordinate AI assistance with manual inspection of the evidence in practice. We present FYI, a browser extension that embeds fact-checking in the reading environment, and use it as a design probe to study how people detect, verify, and determine the validity of data claims against the underlying dataset. FYI provides four complementary tools spanning the spectrum from full automation to manual data exploration. In an exploratory study (N=22), participants used FYI to fact-check claims in a data-driven article. We find that participants adopted three distinct workflow archetypes---AI-first with manual confirmation, manual-first with AI supplement, and parallel co-review---with visualization serving as the primary mechanism for auditing AI conclusions. Trust in AI shifted dynamically, growing when multiple tools converged and eroding when AI outputs were inconsistent. These findings suggest that fact-checking systems should treat AI as a starting point that human verification complements rather than a definitive authority, elevate visualization as a core verification capability, and support flexible, user-driven workflows. We release FYI as open-source software for further research at https://github.com/DataVisards/FYI.","authors":["Nguyen-Truong Thinh","Yuxuan Du","Phongsakon Mark Konrad","Arpit Narechania"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-08-10","first_seen":"2026-08-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.06804","pdf_url":"https://arxiv.org/pdf/2608.06804","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["事实核查","人机交互","数据新闻"],"reason":"研究人类事实核查行为，非LLM仿真人类被试，无agent群体模拟社会过程。","model":"deepseek-v4-pro","scored_at":"2026-08-10T13:01:39","error":null,"has_summary":false,"summary":null},{"id":"2608.07093","version":1,"title":"UncertaintyVis: Preserving Linguistic Uncertainty in Automated Text-to-Chart Generation","zh_title":"UncertaintyVis：在自动文本到图表生成中保留语言不确定性","abstract":"Data-rich documents pair narrative text with quantitative claims, and authors routinely qualify those claims with linguistic uncertainty markers such as \"nearly,\" \"approximately,\" or \"at least.\" Automated text-to-chart systems discard these markers, producing visualizations that appear definitive even when the source text expresses hedged or incomplete knowledge. Readers may then over-interpret precision and misjudge author intent. We present UncertaintyVis, a system that preserves linguistic uncertainty during automated chart generation. A formative corpus analysis of 211 uncertainty expressions across 12 documents and 8 domains yielded a four-category taxonomy: Surface Form Normalization, Precision Boundaries, Inferential Derivation, and Non-Inferable Gaps. We mapped each category to chart-specific visual encodings that signal uncertainty without disturbing the spatial integrity readers rely on, and implemented an end-to-end pipeline pairing large language model text analysis with uncertainty-aware rendering. In a two-part study with 12 participants, readers matched charts to source text with 85% accuracy and text to charts with 76%. Uncertainty-aware visualizations trended toward lower cognitive demand (effect sizes 0.460 and 0.769 for mental demand and effort), and 75% of participants preferred them to plain text, describing explicit uncertainty encodings as a basis for verifying data claims. Encoding effectiveness varied by chart type: bar and pie encodings performed consistently, while line chart encodings require redesign.","authors":["Songheng Zhang","Emily Aurelia","Anthony Tang"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-08-10","first_seen":"2026-08-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.07093","pdf_url":"https://arxiv.org/pdf/2608.07093","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["文本到图表","不确定性可视化","人机交互"],"reason":"研究文本到图表的自动化生成，不涉及用LLM仿真人类被试或行为对照。","model":"deepseek-v4-pro","scored_at":"2026-08-10T13:01:44","error":null,"has_summary":false,"summary":null},{"id":"2608.06782","version":1,"title":"Investigating the Presence and Development of Student Instructor Preferences in a Large-Scale CS1 Course","zh_title":"大规模CS1课程中学生对教师偏好的存在与发展研究","abstract":"Prior research has established the importance of student instructor preferences and identified various influencing factors. However, the dynamics of how student instructor preferences develop and change are less well understood, due to the limitations of common course structures and reliance on one-time measurements. To bridge this gap, we utilize data from a novel learning platform that provides students with access to instructional content created by multiple instructors. This platform enables the quantification of preference emergence and evolution throughout an entire semester, as students repeatedly select content from different instructors. Examining both initial and final student instructor preferences suggests that preference is a dynamic construct continually shaped by experiences. Furthermore, our analysis of the associations between preferences and student characteristics reveals a nuanced picture: while student attributes did not significantly correlate with initial preferences, substantial differences emerged in final preferences across genders and self-reported prior programming experience. This analysis contributes to the existing body of knowledge by expanding our understanding of student instructor preferences and student-instructor relationships in computer science education. We also provide practical insights that institutions and instructors can draw on when multiple instructors collaborate on a course.","authors":["Yiqiu Zhou","Luc Paquette","Geoffrey Challen"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2026-08-10","first_seen":"2026-08-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.06782","pdf_url":"https://arxiv.org/pdf/2608.06782","source_feed":"cs.CY","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["计算机教育","学生偏好","学习分析"],"reason":"研究学生偏好变化，未使用LLM仿真人类被试，不涉及任何A类判据。","model":"deepseek-v4-pro","scored_at":"2026-08-10T13:01:38","error":null,"has_summary":false,"summary":null},{"id":"2608.07069","version":1,"title":"Invisible to the Machine: Auditing AI Restaurant, Cafe, and Bar Recommendation Against a Complete Market Census","zh_title":"机器看不见：基于完整市场普查的AI餐饮推荐审计","abstract":"AI assistants are becoming a primary interface for local discovery, yet almost nothing is known about which venues they surface -- especially in food and drink, where recommendations carry direct revenue consequences. We present the first census-denominated audit of AI venue recommendation: a complete enumeration of 4,776 cafes, restaurants, and bars across two bounded markets (Canggu and Ubud, Bali), against which we evaluate 2,208 search-grounded responses from four production AI systems (ChatGPT, Claude, Gemini, Perplexity) to 96 persona-conditioned queries, collected over seven days under a pre-registered protocol. Because we observe the full market, we can measure what sampled audits cannot: 85.6% of venues were never recommended by any system -- 72.6% even among established venues with fifty or more ratings. Visibility follows a two-margin structure. Entry into answers is associated with documentation: review volume (OR 1.64), an own website (OR 1.92), listed price information (OR 1.54), and third-party web mentions (OR 1.44) -- while star rating is null at this margin (OR 0.89). Rank within answers reverses the pattern: among recommended venues, rating significantly predicts first position (OR 1.17). Presence in an open POI dataset (Foursquare), a folk-theorized visibility factor, shows no positive effect at either margin. Outright fabrication is rare (0.08% of mentions), but systems recommended permanently closed venues 93 times -- staleness, not hallucination, is the practical failure mode. Cross-system agreement is low (top-20 Jaccard 0.33-0.54). A two-week test-retest shows cross-period answer similarity comparable to same-day rerun similarity: the churn is sampling stochasticity, not temporal drift. We release our protocol, registry construction method, and derived data.","authors":["Vladimir Pitenin"],"categories":["cs.IR","cs.CY"],"primary_category":"cs.IR","announce_type":"cross","date":"2026-08-10","first_seen":"2026-08-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.07069","pdf_url":"https://arxiv.org/pdf/2608.07069","source_feed":"cs.CY","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["AI审计","推荐系统","偏差分析"],"reason":"审计AI推荐系统，非用LLM仿真人类被试，无实验或测量目的","model":"deepseek-v4-pro","scored_at":"2026-08-10T13:01:27","error":null,"has_summary":false,"summary":null},{"id":"2608.07280","version":1,"title":"Why Study Emergent Behavior When You Can Regulate It? Aligning Multi-Agent Systems with Reward Prediction","zh_title":"为何研究涌现行为？用奖励预测对齐多智能体系统","abstract":"Multi-agent simulations are widely used to study complex social and ecological systems, where rich and often unexpected emergent behaviors arise from local interactions. A large body of prior work has focused on analyzing such emergent dynamics across domains. In this paper, we move beyond analyzing emergent behavior and introduce a learning-based mechanism for actively shaping it via social reward modeling. We introduce Multi-Agent Reward Prediction (MARP), a simple framework that extends preference-based reward modeling to multi-agent reinforcement learning. While the framework is designed to be applicable across multi-agent settings, the present empirical validation is limited to a single environment, and we therefore present MARP as a proof of concept within the studied domain. Rather than relying on handcrafted rewards, MARP learns a shared reward model from episode-level evaluations of collective outcomes, enabling decentralized agents to align their behavior with global social objectives. We study MARP in the Harvest Game, a canonical sequential social dilemma modeling common-pool resource management and related real-world challenges. Our results show that MARP can be tuned to produce behavior that is more closely aligned with target social metrics than standard reward-based baselines, while the learned reward model captures subtle environmental structure without explicit programming. Crucially, MARP supports multiple and composite social objectives within a single training regime. By modifying only the high-level evaluation metric, the same framework seamlessly aligns agent behavior with diverse goals, including sustainability, equality, and peace, as well as combinations of individual and group-level objectives. These findings demonstrate that emergent multi-agent behavior can be treated not only as a phenomenon to study, but as a target of principled, data-driven regulation.","authors":["Assaf Caftory","Almog Zemach","Moshe Butman","Doron Friedman"],"categories":["cs.MA"],"primary_category":"cs.MA","announce_type":"new","date":"2026-08-10","first_seen":"2026-08-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.07280","pdf_url":"https://arxiv.org/pdf/2608.07280","source_feed":"cs.MA","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体强化学习","社会困境","涌现行为调控"],"reason":"纯多智能体强化学习，用奖励预测对齐社会目标，无LLM仿真人类被试，无人类数据对…","model":"deepseek-v4-pro","scored_at":"2026-08-10T13:01:47","error":null,"has_summary":false,"summary":null},{"id":"2608.07295","version":1,"title":"Learning Long-Term Educational Investment Policies under Residential Sorting","zh_title":"居住分选下的长期教育投资政策学习","abstract":"Allocating public-school investment effectively and fairly is difficult when school access depends on residence. School improvements can raise nearby housing demand and prices, reshape enrollment, and potentially limit access for lower-income households. These effects evolve as residential sorting changes school composition, quality, and future investment needs. Existing approaches often study school funding, household choice, and housing markets separately, while static models can miss their interconnected, long-term effects. We address this gap with a dynamic multi-agent framework that links government investment, household sorting, housing prices, population turnover, enrollment, and evolving school quality. A government planner uses reinforcement learning (RL) to identify multiyear allocation policies that account for household responses while balancing aggregate educational access and equity. In simulations, our RL-based policy attains the highest access level (0.4780) and second-lowest access Gini coefficient (0.0164) among representative baselines, demonstrating a favorable effectiveness-equity balance. The results also indicate reduced socioeconomic stratification in educational access. By making education-housing feedback explicit, our framework supports long-term analysis of how school investment shapes educational opportunity over time.","authors":["Honglei Guo","Shuo Chen","Mingjie Bi","Zeyang Sun","Xiaoxi Wang","Yuhan Zhao"],"categories":["cs.MA"],"primary_category":"cs.MA","announce_type":"new","date":"2026-08-10","first_seen":"2026-08-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.07295","pdf_url":"https://arxiv.org/pdf/2608.07295","source_feed":"cs.MA","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","强化学习","教育政策模拟"],"reason":"纯多智能体系统研究，agent 协作模拟教育投资与居住分选，无人类行为对照，不…","model":"deepseek-v4-pro","scored_at":"2026-08-10T13:01:48","error":null,"has_summary":false,"summary":null},{"id":"2608.06741","version":1,"title":"Solver-Guided Reasoning for Mixed-Equilibrium Strategies","zh_title":"求解器引导的混合均衡策略推理","abstract":"Reasoning in large language models (LLMs) is often grounded in human text, human demonstrations, and human-generated rationales. For equilibrium reasoning in complex games, however, relying on human data can be suboptimal. In fact, human play is often guided by intuition and heuristics and can deviate substantially from game equilibrium. This discrepancy is amplified in games with mixed-strategy equilibria, where human data is heavily biased toward pure strategies. Consequently, conditioning LLMs on this data yields weak game strategies. To grant LLMs the reasoning capacity in games, in this work, we study how to elicit equilibrium play using solver output. We propose Mixed-Strategy Decision Tree (MDT), which articulates the silent optimality of the equilibrium into sparse strategic rules that both humans and LLMs could understand. Using solver output rather than human annotation allows us to extend the input to arbitrarily new states and continuations. We instantiate this study on No-Limit Texas Hold'em by querying a solver oracle for over \\textbf{250 million mixed-strategy decisions}; MDT together with other techniques \\textbf{reduces the $\\ell_1$ distance to the equilibrium by $52.6\\%$} across $8$ different LLM configurations. A Route-only ablation tests the incremental contribution of the shadow-based contrast, while complete River-endgame and Liar's Dice experiments evaluate strategic fidelity and portability beyond the original NLH communication setting.","authors":["Han Wang","Philippe Beardsell","Boning Li","Aaron Sasmita","Shuai Li","Hongyuan Zha","Baoxiang Wang"],"categories":["cs.LG","cs.GT"],"primary_category":"cs.LG","announce_type":"new","date":"2026-08-10","first_seen":"2026-08-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.06741","pdf_url":"https://arxiv.org/pdf/2608.06741","source_feed":"cs.LG","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["博弈论","LLM推理","均衡求解"],"reason":"纯多智能体博弈求解，用求解器而非人类数据训练LLM，不涉及人类行为仿真或对照。","model":"deepseek-v4-pro","scored_at":"2026-08-10T13:01:36","error":null,"has_summary":false,"summary":null},{"id":"2608.06842","version":1,"title":"Tabular Foundation Models and the Unity of Economic Behaviour","zh_title":"表格基础模型与经济行为的统一性","abstract":"Economics uses different behavioural models for risk, time, losses, valuation, and social choice. I study a unified choice experiment in which the same decision makers face all these domains. I hide a decision maker's choices in one domain and ask a frozen tabular foundation model to recover them from that decision maker's choices elsewhere and labelled choices by other participants. The foundation model improves on the training-sample median, and the gain disappears when visible choices are shuffled across decision makers. I then estimate one random-utility model over the foundation model's learned representation. This structural model applies the same utility function in every domain, retains most of the foundation model's reduction in prediction error, predicts domains excluded from utility estimation, and reproduces how behavioural measures co-move across people. The resulting model separates three objects: a learned common choice domain, one systematic utility function on that domain, and one random component that generates stochastic choice on observed menus.","authors":["Victor H. Aguiar"],"categories":["econ.GN","q-fin.EC"],"primary_category":"econ.GN","announce_type":"new","date":"2026-08-10","first_seen":"2026-08-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.06842","pdf_url":"https://arxiv.org/pdf/2608.06842","source_feed":"econ.GN","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["表格基础模型","经济行为预测","随机效用模型"],"reason":"纯多智能体系统研究，用表格基础模型预测决策者选择，不涉及LLM仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-08-10T13:01:24","error":null,"has_summary":false,"summary":null},{"id":"2608.06733","version":1,"title":"Estimating GHG Emissions from AI Use: Framework for Corporate-Level Measurement","zh_title":"估算AI使用的温室气体排放：企业级测量框架","abstract":"Electricity demand from data centers is expected to grow from roughly 5% of U.S. consumption in 2025 to between 9% and 17% by 2030, and corporate artificial intelligence (AI) use is following a similar trajectory, spanning employee productivity assistants, direct access to large language models (LLMs), and AI features embedded in enterprise software. AI emissions today are a small share of footprints for many enterprises, but that share is unlikely to remain small for long. Without reasonable estimates, companies cannot set reduction targets or identify effective decarbonization levers as emissions grow. Companies, regulators, and auditors are asking for emissions estimates that withstand scrutiny, but no widely accepted methodology exists today. Published per-query estimates can differ by several orders of magnitude depending on what is counted, which provider is measured, and what assumptions are made about electricity use and the grid mix. This white paper proposes a standardized framework for corporate-level AI emissions accounting. The framework is designed to be defensible with current data constraints, tiered to meet companies where their data are, transparent about its assumptions, updatable as provider disclosure matures, and built for action rather than disclosure alone. Since AI emissions accounting is still nascent, it has the opportunity to design for actionability from the outset, so that measurement incentivizes responsible choices during AI's rapid buildout.","authors":["John Bistline","Shaena Ulissi","Steven J. Davis","Jonathan Glidden","James Joyce","Mo Li","Jackson Mohsenin","Sangwon Suh"],"categories":["physics.soc-ph"],"primary_category":"physics.soc-ph","announce_type":"new","date":"2026-08-10","first_seen":"2026-08-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.06733","pdf_url":"https://arxiv.org/pdf/2608.06733","source_feed":"physics.soc-ph","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["碳排放核算","AI环境影响","企业标准"],"reason":"论文讨论企业AI碳排放核算框架，不涉及LLM仿真人类被试或行为对照。","model":"deepseek-v4-pro","scored_at":"2026-08-10T13:01:36","error":null,"has_summary":false,"summary":null},{"id":"2608.06771","version":1,"title":"Agentic Artificial Intelligence for Reproducible Human-in-the-Loop Environmental Health Research","zh_title":"面向可重复的人机协同环境健康研究的智能体人工智能","abstract":"Agentic artificial intelligence (AI) systems that are capable of planning and executing multi-step analytical tasks are increasingly available to environmental health researchers, but their reliability in real-world practice has not been fully explored. This paper describes a human-in-the-loop agentic framework for environmental health research, involving the review, verification, and correction of AI-generated data analysis code and results at each step - a process that mirrors the mentorship structure of traditional research teams. This approach offers a path toward rigorous, reproducible use of agentic AI in routine data-rich environmental health research. We illustrate this framework through a case study analyzing nitrogenous organic contaminants at U.S. Superfund sites, a chemical family linked to the emerging tire-derived contaminants 6PPD and 6PPD-quinone. Using an agentic large language model to generate R code for data filtering, spatial mapping, and cluster analysis, we document instances where initial agentic AI outputs benefited from a human-in-the-loop process to produce more rigorous and reproducible results. We conclude that effective use of agentic AI requires both domain expertise to frame questions and evaluate outputs, and coding literacy to guide the AI's approach, while outlining future opportunities for agentic AI workflow to advance the environmental health sciences.","authors":["Edmund Seto"],"categories":["physics.soc-ph"],"primary_category":"physics.soc-ph","announce_type":"new","date":"2026-08-10","first_seen":"2026-08-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.06771","pdf_url":"https://arxiv.org/pdf/2608.06771","source_feed":"physics.soc-ph","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["智能体AI","环境健康","人机协同"],"reason":"纯多智能体系统辅助数据分析，无人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-08-10T13:01:37","error":null,"has_summary":false,"summary":null},{"id":"2603.00059","version":3,"title":"Stochastic Parrots or Singing in Harmony? Testing Five Leading LLMs for their Ability to Replicate a Human Survey with Synthetic Data","zh_title":"随机鹦鹉还是和谐合唱？测试五大领先LLM用合成数据复现人类调查的能力","abstract":"How well can AI-derived synthetic research data replicate the responses of human participants? An emerging literature has begun to engage with this question, which carries deep implications for organizational research practice. This article presents a comparison between a human-respondent survey of 420 Silicon Valley coders and developers and synthetic survey data designed to simulate real survey takers generated by five leading Generative AI Large Language Models: ChatGPT Thinking 5 Pro, Claude Sonnet 4.5 Pro plus Claude CoWork 1.123, Gemini Advanced 2.5 Pro, Incredible 1.0, and DeepSeek 3.2. Our findings reveal that while AI agents produced technically plausible results that lean more towards replicability and harmonization than assumed, none were able to capture the counterintuitive insights that made the human survey valuable. Moreover, deviations grouped together for all models, leaving the real data as the outlier. Our key finding is that while leading LLMs are increasingly being used to scale, replicate and replace human survey responses in research, these advances only show an increased capacity to parrot conventional wisdom in harmony with each other rather than revealing novel findings. If synthetic respondents are used in future research, we need more replicable validation protocols and reporting standards for when and where synthetic survey data can be used responsibly, a gap that this paper fills. Our results suggest that synthetic survey responses cannot meaningfully model real human social beliefs within organizations, particularly in contexts lacking previously documented evidence. We conclude that synthetic survey-based research should be cast not as a substitute for rigorous survey methods, but as an increasingly reliable pre- or post-fieldwork instrument for identifying societal assumptions, conventional wisdoms, and other expectations about research populations.","authors":["Jason Miklian","Kristian Hoelscher","John E. Katsos"],"categories":["cs.CY","cs.AI"],"primary_category":"cs.CY","announce_type":"replace-cross","date":"2026-08-07","first_seen":"2026-02-10","revised_at":"2026-08-07","abs_url":"https://arxiv.org/abs/2603.00059","pdf_url":"https://arxiv.org/pdf/2603.00059","source_feed":"cs.AI","score":10,"bucket":"selected","rubric_hits":["A1","A2","A4","B1","B2","B4"],"tags":["LLM仿真","调查复现","可靠性评估"],"reason":"直接对比LLM合成调查与真人数据，评估仿真可靠性并提出报告标准，高度契合。","model":"deepseek-v4-pro","scored_at":"2026-08-07T13:02:07","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-07","rank":1,"question":"领先的大语言模型生成的合成调查数据能否复现人类受访者的回答，尤其是能否捕捉反直觉的洞见？","design":"使用五种领先LLM（ChatGPT Thinking 5 Pro、Claude Sonnet 4.5 Pro、Gemini Advanced 2.5 Pro、Incredible 1.0、DeepSeek 3.2）模拟硅谷程序员和开发者的调查受访者，通过提示词生成合成调查数据，测量其对伦理与政治意识形态问题的回答模式。","baseline":"一项对420名硅谷程序员和开发者的真实人类调查（Miklian and Hoelscher 2026）。","findings":"LLM生成的合成数据在技术上看似合理且不同模型间高度一致，但均未能捕捉到人类调查中的反直觉洞见；合成数据聚集在一起，真实人类数据反而成为离群值。","reliability":"论文指出合成调查数据无法在缺乏先前文献证据的情境下有意义地模拟真实人类的社会信念，且所有模型都倾向于复述共识性常识而非揭示新发现。","relevance":"该研究直接对比LLM合成调查与真人数据，评估仿真可靠性并提出报告标准，高度契合研究者对LLM人类仿真实验、基准对照及失效条件的关注，值得精读原文。","inspiration":"借鉴其多模型对比与真实人类基准的设计，可评估LLM在特定人群中的仿真偏差。｜可迁移到经济金融领域的调查实验，如消费者信心预期、通胀预期或政策偏好调查。｜以真实消费者调查为基准，用多个LLM生成合成消费者预期数据，比较其对未来经济变量的预测分布与真实调查的差异，检验LLM是否仅复述共识性预期。"}},{"id":"2608.06085","version":1,"title":"Signal or Spurious Cue? A Randomized Audit of Survey-Country Metadata in LLM Social Inference","zh_title":"信号还是虚假线索？一项关于LLM社会推断中调查国家元数据的随机审计","abstract":"Survey-country metadata can improve an LLM's forecast of an individual response when informative, yet the same cue may redirect the forecast when assigned at random. A within-record audit tests whether disclosing a random label's uniform, record-independent origin reduces its country-directed uptake, and whether verified survey country lowers held-out Brier loss. Independent population anchors and recorded human answers measure direction and consequence across five fixed API models, six countries, and seven development-selected targets. In the primary post-review 72-record panel, opaque and disclosed-random labels each produced country-direction shifts of 0.214. Paired attenuation was 0.0003 (95% CI [-0.0157, 0.0166]). Verified country reduced Brier loss by 0.040 (95% CI [0.024, 0.056]), while random-label regret included zero. A non-overlapping mixed-coverage consistency panel retained positive disclosed-random movement and verified utility, while attenuation remained uncertain. On the selected targets, verified metadata was useful in both panels, but disclosure did not reliably attenuate random-label uptake. PROV-FORECAST contains 14,400 paired item-level probability distributions from the corrected panel.","authors":["Yifan Lyu","Xinran Li","Jiaqi Qiao","Xiujuan Xu"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-07","first_seen":"2026-08-07","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.06085","pdf_url":"https://arxiv.org/pdf/2608.06085","source_feed":"cs.AI","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","调查回答预测","算法保真度"],"reason":"用LLM预测个体调查回答，检验随机国家标签的误导效应，并与真实人类答案对照，评…","model":"deepseek-v4-pro","scored_at":"2026-08-07T13:01:54","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-07","rank":3,"question":"在LLM预测个体调查回答时，披露随机分配的国家标签来源能否减弱其误导效应，以及真实调查国家信息能否降低预测误差？","design":"使用五个固定API模型，基于EVS/WVS联合数据集的个体记录，向模型提供同一受访者的10个已观测答案，要求预测7个目标问题的回答概率；通过对比不透明随机国家标签、披露为随机分配的国家标签、真实调查国家标签和无国家标签四种条件，测量预测方向偏移和Brier损失。","baseline":"EVS/WVS 2017-2022联合数据集中六国（中国、法国、英国、意大利、约旦、美国）12,770条真实个体调查记录，包含已观测和留出的人类答案。","findings":"在主要72记录面板中，不透明和披露的随机标签均产生0.214的国家方向偏移，披露未显著减弱该偏移（衰减仅0.0003，95% CI包含零）；真实调查国家使Brier损失降低0.040（95% CI [0.024, 0.056]），而随机标签的预测后悔值包含零。","reliability":"论文指出披露随机标签来源未能可靠减弱其误导效应，且结果可能受限于所选目标问题、国家和模型；留出Brier损失的重算需合法获取源数据。","relevance":"该研究直接以真实人类调查数据为基准，检验LLM在个体预测中对元数据线索的依赖与偏差，并区分信息效用与误导效应，与研究者关注的仿真可靠性及失效条件高度契合，值得精读。","inspiration":"借鉴其同记录内对比不同元数据来源（随机、披露随机、真实）的设计，分离线索的方向性误导与预测效用。｜可迁移到信贷审批或保险定价实验，检验LLM在引入申请人地域、性别等敏感属性时是否产生歧视性偏移及其是否因属性来源说明而减弱。｜以真实贷款违约数据为基准，将申请人部分财务指标作为已观测证据，随机分配或真实使用地域标签，让LLM预测违约概率，比较不同标签条件下的预测偏差和校准误差。"}},{"id":"2608.06115","version":1,"title":"Mind the Gaps: Mixture-of-Minds for Human Simulation","zh_title":"注意差距：用于人类仿真的思维混合模型","abstract":"Predicting how a population will answer a new question is a long-standing goal. Statistical methods succeed at the level of the mass but falter at the level of the individual. Large language model simulators inherit this gap. They recover a population's central tendencies while flattening its heterogeneity, and they carry social biases and prompt brittleness that distort individual predictions. This paper introduces Anacreon, an audience simulation model that targets the individual level within a narrow, well-specified domain. Anacreon learns an authorship embedding that separates individuals, clusters a real qualitative corpus around seed people, and trains a dedicated adapter for each cluster, a mixture of minds, on a Gemma~4 12B base. It harvests demographics, psychological traits, and survey responses from public text, and augments each record with a chain-of-emotion. It reduces prompt brittleness by shuffling response options and reduces positive bias by balancing the training distribution. On a large, externally sourced survey, Anacreon reaches a state-of-the-art ordinal alignment of 0.775, the individual-level accuracy measure on which the field has converged, with a small residual bias. The work is a step toward drawing aggregate insight from faithfully simulated individuals.","authors":["Pranav Dahiya"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-07","first_seen":"2026-08-07","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.06115","pdf_url":"https://arxiv.org/pdf/2608.06115","source_feed":"cs.AI","score":9,"bucket":"selected","rubric_hits":["A1","A2","A3","B1","B2","B4"],"tags":["人类仿真","调查预测","个体异质性"],"reason":"用LLM仿真个体回答调查，有真实人类数据对照，评估偏差与可靠性，涉及社会调查场…","model":"deepseek-v4-pro","scored_at":"2026-08-07T13:01:56","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-07","rank":4,"question":"如何在窄领域内利用LLM模拟独立异质个体，以准确预测其调查回答，并缩小个体与群体层面的预测差距？","design":"Anacreon模型基于Gemma 4 12B，通过作者嵌入分离个体，对真实语料聚类后为每簇训练专用适配器（思维混合），从公开文本中提取人口统计、心理特质和调查回答，并用情绪链增强记录；通过打乱选项顺序减少提示脆弱性，平衡训练分布减少正向偏差。","baseline":"使用大型外部调查的真实人类个体回答作为对照基准。","findings":"Anacreon在外部调查上达到0.775的序数对齐（个体级准确度），为领域内最优；模型残余偏差较小，表明其能较忠实地模拟个体，从而从个体聚合出有意义的群体洞察。","reliability":"论文指出LLM模拟器在宽泛领域会因过度泛化而失效，Anacreon仅适用于窄而明确的领域；基础模型存在社会偏见，且RLHF会降低输出多样性，使助手型模型不适合模拟人类异质性。","relevance":"该研究直接针对LLM仿真人类被试的个体异质性和可靠性问题，有真实人类调查数据对照，并评估偏差，与研究者关注的经济学实验和政策评估场景高度相关，值得精读原文。","inspiration":"借鉴其用聚类适配器捕捉个体异质性、情绪链增强和平衡训练分布以减少偏差的方法｜可迁移到消费者金融决策调查仿真，如风险偏好、信贷选择等｜以公开社交媒体数据构建虚拟消费者，施加不同金融信息提示作为处理，测量其风险资产配置意愿，并以真实家庭金融调查数据（如SCF）作为对照基准。"}},{"id":"2608.06151","version":1,"title":"Reducing belief in conspiracy theories as they unfold using large language models","zh_title":"使用大语言模型减少实时阴谋论信念","abstract":"The emergence of conspiracy theories in the wake of major events is a significant societal challenge. Here we test whether conversational dialogues with a large language model (LLM) can reduce belief in immediately unfolding conspiracies. In experiments conducted in the days following the July 2024 assassination attempt on Donald Trump and the September 2025 assassination of Charlie Kirk, U.S. adults (Experiment 1: N = 472; Experiment 2: N = 1035) holding conspiratorial views about the crisis event engaged in a multi-turn conversation with an LLM prompted to reduce their conspiracy belief. Compared to control participants who either discussed an irrelevant topic with an LLM or viewed a static fact sheet, participants in the LLM treatment showed significantly reduced conspiracy beliefs in both experiments. We also found evidence of downstream effects of the LLM treatment, observing reduced belief in different conspiracies one to two months later in the wake of subsequent crisis events. These results shed light on the psychology of emerging conspiracies and highlight the potential for scalable, cognitively-focused interventions to counteract misinformation in the immediate aftermath of high-profile societal events.","authors":["Thomas H. Costello","Nathaniel Rabb","Michael Nicholas Stagnaro","Gordon Pennycook","David Rand"],"categories":["cs.HC","cs.AI"],"primary_category":"cs.HC","announce_type":"cross","date":"2026-08-07","first_seen":"2026-08-07","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.06151","pdf_url":"https://arxiv.org/pdf/2608.06151","source_feed":"cs.AI","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["LLM人类仿真","阴谋论干预","行为实验"],"reason":"用LLM对话干预阴谋论信念，有真实人类对照实验，评估干预效果与下游影响，属人类…","model":"deepseek-v4-pro","scored_at":"2026-08-07T13:01:57","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-07","rank":5,"question":"在突发危机事件后，大语言模型对话能否即时降低人们对新兴阴谋论的信念？","design":"非仿真研究。该研究以真实人类为被试，在特朗普遇刺未遂和查理·柯克遇刺事件后数日内，招募持有阴谋论看法的美国成年人，随机分配至LLM驳斥对话组、静态信息清单组或无关话题对话对照组，通过前后测比较其对阴谋论的信念变化。","baseline":"真实人类对照：无关话题对话组和静态信息清单组作为对照条件，比较LLM对话干预的效果。","findings":"LLM驳斥对话显著降低了被试对自述阴谋论的信念，效果优于无关对话和静态信息清单；干预效果在1-2个月后的后续危机事件中仍有一定持续性，但对官方解释的信任提升不稳定。","reliability":"论文未讨论","relevance":"该研究直接使用LLM与真实人类进行对话干预，并以随机对照实验评估效果，符合研究者对LLM仿真人类行为、有真实人类基准、关注政策干预场景的兴趣，值得精读。","inspiration":"借鉴其多轮对话干预与多条件对照设计，以及利用突发事件窗口进行即时实验的方法。｜可迁移到经济政策沟通场景，如央行加息公告后，用LLM对话干预公众的通胀预期或政策误解。｜以真实投资者为被试，在政策公告后随机分配至LLM解释对话组或静态新闻组，测量其通胀预期、资产配置意愿的变化，并以调查数据或市场预期指标作为对照基准。"}},{"id":"2608.05178","version":1,"title":"Who Gets Access? Global Region and Academic Status Bias in AI-Generated Academic Gatekeeping Scenarios","zh_title":"谁获得访问权？AI生成学术把关场景中的全球区域与学术地位偏见","abstract":"Equitable access to scientific knowledge often depends on informal gatekeeping decisions, particularly when resources such as paywalled articles, datasets, or professional materials such as curriculum vitae (CV) must be shared selectively. We introduce a controlled simulation framework in which large language model (LLM)-based professors must grant access to only one requestor. Across prompts, requesters vary systematically by global region (Global North vs. Global South) and academic seniority (undergraduate student, PhD candidate, postdoctoral researcher, and tenured professor), while all other factors remain constant. Across varying evaluation scenarios, LLMs exhibit contrasting academic status biases, with some prioritizing PhD candidates, while others favor tenured professors. However, when global regions differ, a distinct divergence emerges based on model architecture: while many frontier LLMs systematically favor requesters from the Global South due to pro-equity bias that results from equity-focused safety alignment, open-weight and small models frequently flip this preference to favor the Global North, reflecting the global region bias and unaligned geographic distribution of their baseline pre-training data. Our findings highlight how normative assumptions embedded in model behavior can shape gatekeeping decisions, underscoring the importance of auditing AI systems for fairness and value alignment.","authors":["Nouar AlDahoul","Hezerul Abdul Karim","Myles Joshua Toledo Tan"],"categories":["cs.CY","cs.AI"],"primary_category":"cs.CY","announce_type":"cross","date":"2026-08-07","first_seen":"2026-08-07","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.05178","pdf_url":"https://arxiv.org/pdf/2608.05178","source_feed":"cs.AI","score":7,"bucket":"pending","rubric_hits":["A3","B4"],"tags":["LLM仿真","学术把关","偏见审计"],"reason":"用LLM模拟学术把关决策，有系统变量操控，批判性揭示偏差，但缺真实人类数据对照。","model":"deepseek-v4-pro","scored_at":"2026-08-07T13:01:48","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-07","rank":6,"question":"当LLM扮演教授进行学术资源把关时，请求者的全球区域和学术地位如何影响其获得访问权限的决策？","design":"用多种LLM扮演教授角色，在模拟的学术把关场景中，系统操控请求者的全球区域（全球北方/南方）和学术地位（本科生、博士生、博士后、终身教授），测量模型选择给予访问权限的请求者类型。","baseline":"无对照","findings":"不同LLM在学术地位上表现出相反偏好，有的优先博士生，有的优先终身教授；在全球区域上，前沿LLM因公平对齐而偏向全球南方请求者，而开放权重和小模型则偏向全球北方，反映预训练数据中的区域偏见。","reliability":"论文未讨论","relevance":"该研究用LLM模拟学术把关决策，系统操控变量并揭示模型偏差，虽无真实人类数据对照，但为理解AI在资源分配中的公平性问题提供了批判性证据，值得阅读以了解仿真中的对齐效应与数据偏见。","inspiration":"借鉴其通过系统操控请求者属性（区域、地位）来测量LLM决策偏差的析因设计，可迁移到信贷审批歧视研究，用LLM扮演信贷员，处理为申请人种族/性别和收入水平，结果变量为批准与否，对照真实银行信贷数据中的歧视模式。"}},{"id":"2608.05583","version":1,"title":"The Judgment-Consequence Gap: LLM Moral Reasoning in Healthcare Decisions","zh_title":"判断-后果差距：医疗决策中大语言模型的道德推理","abstract":"As large language models (LLMs) enter high-stakes domains such as healthcare, understanding their moral reasoning becomes essential. Decisions about scarce medical resources often hinge on judgments of responsibility, particularly when patients' own actions contribute to illness. We investigate how LLMs reason about responsibility and its consequences, tracing their judgments across successive levels, from the behavior, to the resulting illness, to the denial of care. We evaluate a wide range of LLMs, spanning different model families and capability levels, on various clinical vignettes adapted from prior studies. Our results identify a judgment-consequence gap: LLMs largely agree with humans that patients bear responsibility for health-harming behaviors, yet overwhelmingly refuse to let that judgment influence how they allocate scarce resources. Specifically, LLMs default to random allocation, whereas humans consistently favor the less-culpable patient. Compared to humans, LLMs also place greater emphasis on access to information, reducing responsibility judgments when health-risk knowledge is unavailable. These findings reveal that LLMs apply a systematically different moral framework than humans when responsibility and resource scarcity intersect, surprisingly often amplifying normative disagreement with humans as reasoning capability increases.","authors":["Hadi Hosseini","Samarth Khanna","Leona Pierce"],"categories":["cs.CY","cs.AI","cs.LG"],"primary_category":"cs.CY","announce_type":"cross","date":"2026-08-07","first_seen":"2026-08-07","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.05583","pdf_url":"https://arxiv.org/pdf/2608.05583","source_feed":"cs.AI","score":7,"bucket":"pending","rubric_hits":["A1","B1","B2","B4"],"tags":["LLM仿真","道德决策","人类对照"],"reason":"用LLM模拟人类道德决策并与真实人类数据对照，涉及医疗资源分配场景，揭示仿真失…","model":"deepseek-v4-pro","scored_at":"2026-08-07T13:01:51","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-07","rank":7,"question":"当患者自身行为导致疾病时，LLM在道德责任判断与稀缺医疗资源分配决策上是否与人类一致？","design":"使用多种LLM（不同模型家族、能力水平、推理与非推理配置）阅读临床情境短文，依次测量其对患者行为责任、患病责任、被拒绝治疗责任的Likert评分及资源分配选择，并与人类研究数据对照。","baseline":"真实人类数据来自肾脏移植分配、肺癌治疗、髋关节置换手术等场景的已有研究。","findings":"LLM与人类在行为责任判断上基本一致，但存在“判断-后果鸿沟”：LLM拒绝让责任判断影响分配，默认随机分配，而人类倾向将资源给予责任较小的患者。LLM对患者是否知晓健康风险更敏感，且推理能力增强反而扩大与人类的分歧。","reliability":"论文未讨论","relevance":"该研究直接以LLM模拟人类道德决策，并与真实人类数据对照，揭示仿真在责任归因与稀缺资源分配场景中的系统性偏差，高度契合研究者对仿真可靠性及失效条件的关注。","inspiration":"借鉴其逐层分解道德推理链（行为→疾病→剥夺）并分别测量的设计，可清晰定位人机分歧点。｜可迁移至信贷审批中的责任归因实验，如借款人因自身行为导致违约风险时，AI与人类在贷款拒绝决策上的差异。｜以LLM和人类信贷员为被试，呈现借款人行为（如过度消费）导致违约的案例，测量责任归因与贷款批准决策，对照真实信贷审批数据中的行为模式。"}},{"id":"2608.00023","version":2,"title":"Role Steering of Language Models for Social Simulations","zh_title":"面向社会模拟的语言模型角色引导","abstract":"Social simulations built from language-model agents need role-conditioned behavior that can be checked before agents are placed into a simulated population. We introduce an activation-steering screening workflow for role-conditioned agents: define a role profile, extract a role-specific direction, sweep four steering coefficients, evaluate role-profile alignment, and pass or flag each candidate configuration. On OLMo-3-7B-Instruct, we apply the workflow to a mixed 275-role inventory with 228 role-agnostic questions, GPT-4.1-mini prompted role references, and GPT-4.1-mini judges. Role-specific directions receive higher judged role-profile alignment than an assistant-axis directional control from prior persona-vector work, with mean overall scores of 63.2 versus 41.1 across the tested grid. They also preserve high lexical diversity, while the control drops sharply at larger coefficients. The role-level screen is the main practical output: most roles improve as steering increases, but 38 roles decline across all six measured dimensions, showing why simulation builders should choose coefficients per role rather than deploy a uniform high-strength setting. We make our code and evaluation artifacts available at https://anonymous.4open.science/r/anonymous-research-code-5F03/.","authors":["Isaac Song","Mohammed Rehan Parwani","Glenn Matlin","Emile Anand","Akhil Theerthala","Arjun Chatterjee","Anthony Wen-Ming Zang","Maria Kostylew","Yonadav G. Shavit","Sebastien Krier","Mark Riedl"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-08-07","first_seen":"2026-08-04","revised_at":"2026-08-07","abs_url":"https://arxiv.org/abs/2608.00023","pdf_url":"https://arxiv.org/pdf/2608.00023","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["社会模拟","角色引导","LLM代理"],"reason":"用LLM agent进行社会模拟，但无真实人类数据对照，属于边界情形。","model":"deepseek-v4-pro","scored_at":"2026-08-07T13:02:07","error":null,"has_summary":false,"summary":null},{"id":"2608.05166","version":1,"title":"Conditional Cognitive Biases in LLMs: How Biased User Turns Modulate In-Context Reasoning","zh_title":"大语言模型中的条件认知偏差：有偏用户轮次如何调节上下文推理","abstract":"We present an evaluation of cognitive bias expression in state-of-the-art instruction-tuned LLMs under realistic multi-turn interaction settings. Our work introduces a novel three-condition experimental framework that disentangles the effect of exposure to a biased user turn from the effect of the turn's semantic content, alongside a benchmark of 24,300 jury-validated user prompts spanning all 81 cells of a 9x9 target-human bias interaction matrix. Across eight frontier LLMs, we find that biased conversational context systematically increases bias expression relative to zero-shot baselines in 6 of 8 models. We identify two competing behavioral dynamics underlying this effect: conversational exposure to biased reasoning generally amplifies downstream bias tendencies, while explicitly stated bias cues often trigger alignment-related suppression behaviors that reduce overt bias expression. We release our framework, codebase, and dataset to support future research on context-conditioned cognitive biases and behavioral adaptation in LLMs.","authors":["Sachini Weerasekara","Sagar Kamarthi","Jacqueline Isaacs"],"categories":["cs.CL","cs.CY","cs.HC","cs.LG"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-07","first_seen":"2026-08-07","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.05166","pdf_url":"https://arxiv.org/pdf/2608.05166","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["认知偏差","LLM评估","上下文影响"],"reason":"测量LLM本身的认知偏差，非仿真人类被试，但涉及偏差评估可迁移","model":"deepseek-v4-pro","scored_at":"2026-08-07T13:01:48","error":null,"has_summary":false,"summary":null},{"id":"2608.05576","version":1,"title":"Where Models Converge and Humans Diverge: A Coverage Framework for Distributional Pluralism in Open-Ended Generation","zh_title":"模型趋同与人类发散之处：开放式生成中分布多元性的覆盖框架","abstract":"When a large language model (LLM) writes Harry Potter fanfiction, it reliably produces fundamental elements of the Hogwarts universe, such as recognizable places and characters. Human-written Harry Potter fanfictions, however, typically include these fundamentals and much more, incorporating stylistically irregular content and relationship-diverse plotlines. This gap between LLM and human writing has been noted across a variety of domains. LLMs tend to produce \"average\" writing, while human writing contains more diverse content that covers a broader distribution. Existing work has shown the existence of this distributional \"gap\", but no work has proposed a systematic way to measure it. Our paper proposes a human-grounded framework that uses the empirical distribution of human writing on a topic to measure the distributional breadth of LLM-generated content on that same topic. We propose two metrics, LLM Coverage (LLM-Cov) and In-Boundary Rate (IBR), that separate the plausibility of LLM content from its distributional breadth. Across ideation and narrative tasks, we find that current LLMs produce plausible but narrow content that concentrates near the center of the human response space. Our framework can enable researchers to better assess the distributional breadth of LLM-authored content, which we term its \"cultural reach\".","authors":["Zini Yang","Emily Wenger","Richard So"],"categories":["cs.CL","cs.CY"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-07","first_seen":"2026-08-07","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.05576","pdf_url":"https://arxiv.org/pdf/2608.05576","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["LLM生成分布","文化覆盖度","人类写作对比"],"reason":"比较LLM与人类写作分布差异，但无人类行为仿真或对照实验，属社会模拟无人类数据…","model":"deepseek-v4-pro","scored_at":"2026-08-07T13:01:51","error":null,"has_summary":false,"summary":null},{"id":"2608.05630","version":1,"title":"Human-Like Anaphor Resolution in Large Language Models","zh_title":"大型语言模型中类人指代消解研究","abstract":"Anaphors are expressions that refer to other expressions, called antecedents. The process of connecting the two is called resolution. Cognitive science has identified multiple factors that affect the speed and success of anaphor resolution, including discourse structure, situation-model properties, and semantic factors. Here, we investigate whether these factors also affect anaphor resolution in five Large Language Models (LLMs) with open weights: GPT-2-XL, Llama-3.1-8B, Pythia-12B, Mistral-7B, and Mistral-24B. To model processing difficulty, we adopt the standard linking hypothesis that relates human reading times to model surprisal at the anaphor. As a second behavioral measure, we compare model accuracy to human accuracy on comprehension questions probing the antecedents of anaphors. The results show selective cognitive alignment: some LLMs exhibit human-like sensitivity to discourse prominence and distance-based factors in anaphor resolution, while showing weaker or absent sensitivity to semantic interference effects. These findings delimit the conditions under which LLMs approximate human anaphor resolution.","authors":["Keane Zhang","Varshini Chinta","Raj Sanjay Shah","Sashank Varma"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-07","first_seen":"2026-08-07","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.05630","pdf_url":"https://arxiv.org/pdf/2608.05630","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["指代消解","认知对齐","语言模型评估"],"reason":"研究LLM的指代消解是否与人类相似，属于将LLM作为测量对象，但非仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-08-07T13:01:51","error":null,"has_summary":false,"summary":null},{"id":"2608.05519","version":1,"title":"EcoAgent-Bench: Evaluating Economic Decision-Making in Budget-Constrained LLM Agents","zh_title":"EcoAgent-Bench：评估预算受限LLM智能体的经济决策","abstract":"Agent benchmarks usually measure task completion and treat resource use as an auxiliary statistic. In deployment, however, the choice among a local lookup, broad search, composite research tool, stronger model, or human escalation is part of the task itself. We introduce EcoAgent-Bench, in which every task specifies priced actions and an explicit budget. Its 304 real-derived tasks span five families adapted from GAIA, HotpotQA, and MuSiQue, and test four decisions: avoiding unnecessary escalation, escalating when local evidence is insufficient, selecting a model tier, and stopping on unsupported premises. We evaluate seven LLM agents in tool-API and workspace-CLI settings, together with four oracle scripted controls. Micro-averaged accuracy rewards one-sided policies: always-escalate controls achieve high micro success while failing save-oriented tasks. We therefore also report an economic-consistency score (the worse of accuracy on upgrade-oriented and save-oriented family groups) which exposes this failure. Tool-API agents attain only 3.9-24.0% micro strict success (at most 7.3% economic consistency), often either stopping before warranted escalation or overspending on cheap tasks. A threshold-crossing budget sweep changes GPT-5.4's escalation rate from 0% to only 3%. These results show that completion under a budget and economical action selection are distinct properties. We release the task bundle, transformation pipeline, frozen evaluation environments, and integrity-bound result artifacts needed to study both.","authors":["Jie Wu","Ming Gong","Feixiang Cheng","Qinqin Zhao"],"categories":["cs.AI","cs.CL","cs.LG"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-08-07","first_seen":"2026-08-07","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.05519","pdf_url":"https://arxiv.org/pdf/2608.05519","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["LLM智能体","经济决策","基准测试"],"reason":"LLM agent 经济决策模拟，但无真实人类数据对照，属社会模拟边界情形。","model":"deepseek-v4-pro","scored_at":"2026-08-07T13:01:59","error":null,"has_summary":false,"summary":null},{"id":"2608.05367","version":1,"title":"Counterfactual Analysis via Large Language Models","zh_title":"基于大语言模型的反事实分析","abstract":"Counterfactual analysis aims to predict potential outcomes under hypothetical scenarios, offering valuable insights for decision-making. This paper investigates the application of large language models (LLMs), specifically the GPT-3.5 model, for counterfactual analysis. We focus on the online lending context, where the counterfactual return on investment (ROI) is crucial for evaluating different interest rate schemes. We begin by assessing the predictive performance of GPT and comparing it with advanced machine learning algorithms. The results show that prompt engineering can significantly enhance GPT's predictions, with the R-squared increasing from 1.97% to 2.84%, closely approaching the 3.48% achieved by gradient-boosted regression. Subsequently, we utilize GPT to generate counterfactual ROIs under a set of alternative interest rates. GPT exhibits logical coherence and causal reasoning in its responses. The findings underscore the potential of LLMs as effective tools for counterfactual analysis in online lending, suggesting broader applications for LLMs in various predictive and decision-making contexts.","authors":["Zonghao Yang"],"categories":["cs.AI","q-fin.GN"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-07","first_seen":"2026-08-07","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.05367","pdf_url":"https://arxiv.org/pdf/2608.05367","source_feed":"cs.AI","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["反事实分析","LLM仿真","在线借贷"],"reason":"用LLM生成反事实结果，模拟经济决策，但无真实人类行为对照，属社会模拟边界情形。","model":"deepseek-v4-pro","scored_at":"2026-08-07T13:01:59","error":null,"has_summary":false,"summary":null},{"id":"2608.05864","version":1,"title":"Seeing Is Not Deciding: Can Multimodal LLMs Act as Effective CEOs?","zh_title":"眼见不为实：多模态大语言模型能胜任CEO吗？","abstract":"Large language models are increasingly applied as autonomous decision-making agents. However, in executive business decisions, existing benchmarks are limited to textonly settings. This makes it unclear whether models can perceive visual business evidence and effectively integrate it to improve decision quality. We introduce C-SUITEBENCH, a controlled multimodal benchmark that includes five decision tasks under paired text-only and multimodal conditions across 50 scenarios. We place nine frontier models in the role of a chief executive officer and evaluate their decision-making ability. Multimodal inputs consistently improve evidence-centric reasoning, with the largest and most reliable gains appearing in risk forecasting and board-facing justification. However, we uncover a multimodal integration paradox: adding visual business information degrades constrained resource allocation for all nine models, even as visual grounding itself improves. Ablation experiments reveal that this failure emerges from signal crowding, although each visual channel helps individually, their combination disrupts constraint satisfaction during decoding. These findings demonstrate that visual perception and constrained action are separable bottlenecks in multimodal agents, and that indiscriminate visual augmentation can harm high-stakes decision making, motivating selective grounding strategies for future executive AI systems.","authors":["Yuyang Dai","Xueqing Peng","Yuxia Wang","Preslav Nakov","Zhuohan Xie"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-07","first_seen":"2026-08-07","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.05864","pdf_url":"https://arxiv.org/pdf/2608.05864","source_feed":"cs.AI","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["LLM决策模拟","多模态智能体","社会模拟"],"reason":"用LLM模拟CEO决策，属社会模拟但无真实人类数据对照，为边界情形。","model":"deepseek-v4-pro","scored_at":"2026-08-07T13:01:54","error":null,"has_summary":false,"summary":null},{"id":"2608.06020","version":1,"title":"From Economic Agents to Agentic Economies: A Systems Blueprint for Economic World Models","zh_title":"从经济主体到主体经济：经济世界模型的系统蓝图","abstract":"Economic World Models (EWMs) are generative economic models that simulate how economies evolve from within by modeling heterogeneous agents, their beliefs and actions, and the market and institutional mechanisms through which their interactions produce aggregate outcomes. This paper develops an implementation roadmap for building economic world models as generative engines in which heterogeneous agents act, interact, adapt, and co-evolve with markets and institutions, thereby producing economic dynamics from the inside. We organize EWM systems into a six-level capability ladder, from fixed rule-based agent worlds to adaptive and LLM-based agent worlds, self-evolving agents, evolving institutional worlds, and sim-to-real economic twins aligned with real observations. A systematic literature survey across these levels reveals that existing work remains concentrated in lower-level agent and simulation environments, while systems with self-evolving agents, endogenous institutions, persistent empirical alignment, and validated economic mechanisms remain rare. By translating the EWM agenda into an implementation blueprint, this paper aims to accelerate the development of the next generation of economic simulation environments that can serve as high-fidelity sandboxes for human decision-makers and as training, planning, evaluation, and safety substrates for AI agents. We release a curated paper list and related resources to support future research.","authors":["Jiale Han","Xiang Li","Jing Qian","Wenyuan Gu","Pin Gao","Ye Luo","Hongyuan Zha","Dacheng Tao","Benyou Wang","Lin William Cong"],"categories":["cs.AI","cs.LG"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-07","first_seen":"2026-08-07","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.06020","pdf_url":"https://arxiv.org/pdf/2608.06020","source_feed":"cs.AI","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["经济世界模型","LLM智能体","社会模拟"],"reason":"提出经济世界模型蓝图，用LLM agent模拟经济，但无真实人类数据对照，属社…","model":"deepseek-v4-pro","scored_at":"2026-08-07T13:01:54","error":null,"has_summary":false,"summary":null},{"id":"2608.06108","version":1,"title":"Evaluating Investment Logic in Large Language Models: A Real-World Benchmark Towards Personalzied Financial Agents","zh_title":"评估大语言模型中的投资逻辑：面向个性化金融智能体的真实世界基准","abstract":"Investment competence is inherently personalized: the same market evidence can justify different actions for investors with different goals, horizons, portfolios, and risk boundaries. Yet financial LLMs are evaluated either by static question answering or by terminal profit and loss. The former omits agency; the latter cannot reveal whether a profitable action was grounded, profile-consistent, or merely lucky. We ask whether the community is using the wrong ruler for consequential agents. We introduce \\textsc{InvestLogicBench}, a process-native benchmark containing 201,247 documented decisions from 151 real-world investors. Each episode instantiates a \\textbf{P$\\rightarrow$E$\\rightarrow$R$\\rightarrow$D$\\rightarrow$O} trace: investor \\textit{Profile}, observable market \\textit{Events}, investment \\textit{Reasoning}, executable \\textit{Decision}, and delayed \\textit{Outcome}. The release includes profile construction, point-in-time event binding, structured logic, horizons, outcomes, and post-mortems, and supports comprehension, profile-conditioned generation, and end-to-end replay. Across four leading LLMs, logical plausibility remains near 4/5 while event grounding is only 0.8--2.8/5; return and process quality also disagree. These results expose polished but weakly grounded reasoning that outcome-only evaluation hides. We further argue that P$\\rightarrow$E$\\rightarrow$R$\\rightarrow$D$\\rightarrow$O should be a data-system interface, requiring versioned profiles, temporal provenance, inspectable retrieval, decision ledgers, and replayable outcomes. Finance is our stress test for a broader class of personalized, consequential agents.","authors":["Yuanhong Jiang","Jingjie Zou","Zhenghong Lin","Xusheng Yu","Qiqi Huang","Shuai Jia","Shijie Dai"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-07","first_seen":"2026-08-07","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.06108","pdf_url":"https://arxiv.org/pdf/2608.06108","source_feed":"cs.AI","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["LLM评估","金融智能体","决策模拟"],"reason":"用LLM模拟投资者决策，有真实人类数据对照，但目标是评估模型而非仿真人类行为分…","model":"deepseek-v4-pro","scored_at":"2026-08-07T13:01:56","error":null,"has_summary":false,"summary":null},{"id":"2608.05172","version":1,"title":"Estimating time spent on work tasks","zh_title":"估算工作任务耗时","abstract":"The task-based framework in economics models occupations as bundles of tasks. It is the standard lens for understanding how technology affects work: a new technology changes the cost or time each task requires and these task-level effects aggregate to occupation-level effects. We study how tasks should be weighted in this aggregation. Prior work has relied on idiosyncratic or ill-justified choices for task weights. While recent work suggests weighting tasks by time spent, existing time shares are either based on coarse ONET data not intended for this purpose or estimated via black-box language models. We address this gap by proposing a principled method for estimating time shares for nearly 18,000 tasks that constitute nearly all U.S. jobs. Our estimates factor a task's time into (i) the expected frequency of the task, derived from ONET, and (ii) the time to complete a single instance of it. To estimate the latter, we solve a constraint satisfaction problem based on pairwise comparisons elicited from language models about which tasks are longer per instance. We validate our estimates by characterizing the solution space of the constraint satisfaction problem and collecting data from workers for multiple occupations. We apply our time shares to analyze how AI exposes U.S. occupations and find that some prior results are sensitive to time weights. Accounting for the share of working time exposed to AI, rather than the share of tasks like prior work, widens the gap between the least and most exposed jobs: it lowers measured exposure for most occupations but raises it for the most exposed. Re-weighting by time also reshuffles 11 of the 25 occupations widely reported as most exposed to AI, shifting the top of the list away from clerical work and toward analytical roles. Time shares can serve as a general primitive for research and policy on the labor economy and the economics of technology.","authors":["Stephane Hatgis-Kessell","Tom\\'as Aguirre","Alexander Wan","Rishi Bommasani"],"categories":["cs.CY","cs.AI"],"primary_category":"cs.CY","announce_type":"cross","date":"2026-08-07","first_seen":"2026-08-07","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.05172","pdf_url":"https://arxiv.org/pdf/2608.05172","source_feed":"cs.AI","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM标注","任务时间估计","AI暴露度"],"reason":"用LLM做任务时长比较，替代人工标注，非仿真人类被试","model":"deepseek-v4-pro","scored_at":"2026-08-07T13:01:48","error":null,"has_summary":false,"summary":null},{"id":"2608.05180","version":1,"title":"The Nuclear Decision-Making Benchmark: Evaluating Frontier LLMs on Nuclear Tendencies","zh_title":"核决策基准：评估前沿大语言模型的核倾向","abstract":"The integration of large language models into defense and national-security workflows raises urgent questions about whether frontier models exhibit stable, consistent, and policy-appropriate preferences in high-stakes contexts. We introduce the Nuclear Decision-Making Benchmark (NDM Bench), a targeted evaluation framework of 151 scenarios authored by PhD-credentialed scholars in international relations spanning four domains: escalation (76), arms control (25), non-proliferation (25), and proliferation (25). Scenarios are actor-agnostic, enabling multiple country pairs to be exchanged, and we introduce experimental phrasing variants to probe sensitivity to narrative framing. We apply the benchmark to seven frontier AI systems: DeepSeek-V3.2, ERNIE 4.5-300B, Gemini 3 Pro, GLM-4.6, GPT-5.2, Llama 4 Maverick-17B Instruct, and Qwen3-235B. We find significant overall inter-model variation in all four domains, with 91.7% of pairwise inter-model differences significant. DeepSeek and Qwen are the most likely to recommend escalatory action using nuclear weapons; GPT and ERNIE are the least likely. Llama exhibits a distinct bias for action, favoring force, intervention, and cooperation across domains. Inter-rater reliability metrics (Krippendorff's $\\alpha$ and quadratically weighted Fleiss' $\\kappa$) reveal Llama and ERNIE are the most consistent across runs, with either DeepSeek or GLM the least depending on the domain. We also present a deeper exploration of our scenario variants: (i)~country-level biases tend to exist and vary by model, with country covariates like adversary trade ties and escalation propensity producing weak correlations; (ii)~existential phrasing effects are significant and heterogeneous; (iii)~these country biases interact with phrasing. Overall, the distributions of responses related to the scenarios in our benchmark vary significantly by model, country, and phrasing.","authors":["Benjamin Jensen","Ian Reynolds","Yasir Atalan","Martin Pollack","Austin Woo","Robert Sincero"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2026-08-07","first_seen":"2026-08-07","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.05180","pdf_url":"https://arxiv.org/pdf/2608.05180","source_feed":"cs.CY","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["LLM评估","核决策","模型行为"],"reason":"测量LLM在核决策中的倾向，属于对模型本身的测量，无人类被试对照，边界情形。","model":"deepseek-v4-pro","scored_at":"2026-08-07T13:01:48","error":null,"has_summary":false,"summary":null},{"id":"2608.03067","version":2,"title":"Activation-Guided Neuron Intervention to Induce Alzheimer's-Related Computational Language Phenotypes in a Large Language Model","zh_title":"激活引导神经元干预以诱导大语言模型产生阿尔茨海默症相关计算语言表型","abstract":"Changes in spontaneous speech provide an early signal of cognitive dysfunction in Alzheimer's disease (AD) that large language models (LLMs) can detect. However, detection alone cannot establish whether the underlying model representations contribute functionally to behavior. We introduce an activation-guided intervention framework using Qwen3-8B. The framework identifies feed-forward neurons with higher activation rates for AD than control transcripts and modulates their output contributions during generation by scaling the corresponding down-projection weights. This yielded nine edited variants differing in intervention direction, magnitude, and scope. The original and edited models completed the same 12-turn neuropsychological battery, assessed through blinded human ratings and computational linguistic measures. Amplifying AD-associated neurons produced graded impairments in story recall, verbal fluency, working memory, procedural discourse, scene construction, and coreference resolution. Attenuation largely preserved performance and selectively improved several outcomes. Amplification also reduced lexical surprisal, idea density, syntactic complexity, and discourse quantity, broadly paralleling changes reported in human AD speech. These findings show that neurons identified solely from clinical language differences can influence behavior across multiple cognitive domains, providing proof of concept for an AD-related computational phenotype and a controlled framework for experimentally examining links between language and broader cognitive dysfunction.","authors":["Rui He","Ercong Nie","Hong Jiang","Iris E. Sommer","Philipp Homan","Wolfram Hinzen"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-08-07","first_seen":"2026-08-05","revised_at":"2026-08-07","abs_url":"https://arxiv.org/abs/2608.03067","pdf_url":"https://arxiv.org/pdf/2608.03067","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["阿尔茨海默症","语言检测","模型编辑"],"reason":"用LLM检测阿尔茨海默症语言特征，非仿真人类被试","model":"deepseek-v4-pro","scored_at":"2026-08-07T13:02:08","error":null,"has_summary":false,"summary":null},{"id":"2608.04549","version":2,"title":"EuroExec: Frontier Language Models Fall Short of Expert Judgment on European Executive Decision Tasks","zh_title":"EuroExec：前沿语言模型在欧洲行政决策任务上不及专家判断","abstract":"Frontier LLMs are increasingly put to use on open-ended complex questions, different in nature from the ones they are typically evaluated on. We dedicate more than 4,000 human expert hours to evaluate a selection of six frontier LLMs on a member of this class of problems: EuroExec, our introduced human expert-based benchmark composed of 413 open-ended long-form European executive tasks authored by 47 vetted domain experts, each question drawn from experience in a real case. Every response is manually evaluated through a multi-attribute rubric, an item-specific checklist of requirements, and a preference rank ordering, extracting an aggregate metric \"Solve Rate\". The strongest model solves only 56.9% of tasks, while expert-written reference answers judged blindly are solved at near-ceiling levels and are preferred over every model response in 74% of direct rankings, placing frontier generative systems well below the professional standard of work they are already used for. We see that the best way to extract this kind of conclusion is by employing human evaluators, carefully checking their consistency through rigorous statistical analysis, and observe that automatic measurements also fall short when evaluating on this case of real-world open-ended problems with a subjective ground truth.","authors":["Pau Arnal","Khaled Denfir","Danylo Smahliuk","Amrut Avhad","Marcus A. Castro"],"categories":["cs.CL","cs.AI","cs.LG"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-08-07","first_seen":"2026-08-06","revised_at":"2026-08-07","abs_url":"https://arxiv.org/abs/2608.04549","pdf_url":"https://arxiv.org/pdf/2608.04549","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["LLM评测","专家基准","行政决策"],"reason":"纯LLM能力评测，无人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-08-06T13:02:26","error":null,"has_summary":false,"summary":null},{"id":"2608.05993","version":1,"title":"Clinical Communication Processing with Models Trained on LLM-Generated Synthetic Data: A Structured Survey and Novel Application Case Studies","zh_title":"基于LLM生成合成数据训练的临床沟通处理：结构化综述与新应用案例","abstract":"Much clinical value is conveyed not through structured records but through communication: exchanges in which patients describe symptoms, clinicians reason and give instructions, ambulances hand over to emergency departments, and nurses pass on a shift. Such language differs from tabular data because meaning depends on speaker role, intent, causality, uncertainty, omission, and channel noise. Healthcare natural language processing must therefore interpret information as conveyed rather than coded. This requires well-annotated corpora, which are scarce because authentic exchanges are private, fragmented, and costly to annotate. Large language models offer a way forward by transforming clinical sources, such as records, diagnostic labels, symptom lists, or care plans, into written and transcribed communication for downstream models. We present a structured narrative survey organized by source representation, communication form and participants, generation method, and downstream task, complemented by thirteen novel case studies. These build clinical NLP systems for communication channels and languages without labeled real-world data, including EMS pre-arrival reports, field-radio casualty documentation, nurse handoffs, patient-portal triage, and low-resource discharge communication. They show that synthetic communication can bootstrap such systems. Findings include the competitiveness of fine-tuned encoder models over evaluated zero-shot baselines and the value of deliberately degraded communication for robustness. The main limitation is that most studies evaluate on held-out synthetic communication, while train-on-synthetic, test-on-authentic evidence remains limited. We conclude that syn-thetic clinical communication is becoming a practical research resource; establishing it as reusable clinical infrastructure will require authentic-data transfer, safety and external validation.","authors":["Alexander Apartsin","Yehudit Aperstein"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-07","first_seen":"2026-08-07","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.05993","pdf_url":"https://arxiv.org/pdf/2608.05993","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["临床NLP","合成数据","数据增强"],"reason":"用LLM生成合成临床对话数据训练NLP模型，属于数据增强，不涉及人类行为仿真或…","model":"deepseek-v4-pro","scored_at":"2026-08-07T13:02:03","error":null,"has_summary":false,"summary":null},{"id":"2608.06027","version":1,"title":"FormBharo: Designing and Evaluating a Voice Agent for Conversational Form Filling in Rural India","zh_title":"FormBharo：为印度农村设计的对话式表单填写语音代理","abstract":"In India, almost every social benefit starts with a form, yet the people who need these benefits most are often unable to read or write. Reaching them requires a spoken conversation. Today that work falls to frontline health workers who enroll beneficiaries one at a time, a poor use of stretched capacity. We built FormBharo (\"fill the form\" in Hindi), a voice agent that fills a structured form over a phone call under tight latency and cost budgets by pairing Large Language Models (LLMs) with deterministic, rule-based validation and flow control. It is being piloted with ARMMAN, an NGO running large-scale maternal and child mobile-health programs in India, to enroll low-income, Hindi-speaking mothers in antenatal and postnatal care. To our knowledge, it is the first voice agent piloted to fill an enrollment form for this population. We openly release FormVoiceAgentBench, a benchmark pairing human-recorded Hindi audio with 3,760 multi-turn conversation tests across 960 simulated calls, to evaluate our agent's components (transcription, extraction, reply generation) and end-to-end form completion under real acoustic variations. Form completion drops by up to ~41 points when LLMs receive error-prone real-speech transcripts instead of reference ones. The rule-based controls recover many turn-level extraction errors, helping smaller, cheaper models match or surpass frontier models on form completion. Component performance does not predict end-to-end performance: GPT-5.5 leads turn-level extraction accuracy on reference transcripts (99.8%) but ranks lower on form completion. Since errors both propagate and cancel across the pipeline, the optimal model choice of models emerges only through end-to-end evaluation. Finally, no single model is best across accuracy, cost, and latency at once, so we use a Pareto-based weighted-sum scalarization to select a deployable configuration balancing the three.","authors":["Aman Dalmia","Sanskriti Midha","Jigar Doshi"],"categories":["cs.CL","cs.AI","cs.HC"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-07","first_seen":"2026-08-07","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.06027","pdf_url":"https://arxiv.org/pdf/2608.06027","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["语音代理","表单填写","人机交互"],"reason":"语音代理填写表格，属于对话式表单填写工具，非人类仿真实验，无行为对照。","model":"deepseek-v4-pro","scored_at":"2026-08-07T13:02:04","error":null,"has_summary":false,"summary":null},{"id":"2608.05889","version":1,"title":"The em-dash em-beds in Congress: A population-level rise in em-dash frequency in U.S. congressional press releases at the dawn of the large-language-model era, 2021-2025","zh_title":"国会中的破折号嵌入：大语言模型时代初期美国国会新闻稿中破折号频率的群体性上升（2021-2025）","abstract":"Large language models (LLMs) can leave small stylistic traces in text written with their help. The most discussed is the em-dash (U+2014), especially the unspaced form word---word, which is normal in typeset English prose but unusual in U.S. press writing, where AP style calls for spaced dashes. This study asks whether that trace is measurable in congressional press releases. In a preregistered design (OSF: 10.17605/OSF.IO/U5NEY), 146,239 scraper-sourced releases from 480 House and Senate offices (2021-2025, the open congress-press dataset) were analyzed: density of unspaced prose-form em-dashes per 1,000 characters of cleaned text, Poisson/negative-binomial models with a length offset, clustering by office. Density stayed within 0.10-0.12 per 1,000 characters through 2021-2024, then rose to 0.217 in 2025, more than twice the four-year baseline; the share of releases with such an em-dash rose from ~13% to 24.8%. The primary frequency ratio (2023-2025 vs 2021-2022) was 1.55 (95% CI 1.28-1.93; exact registered cut-off: 1.528), just above the prespecified 1.5x threshold. The rise was net-new (hyphen density stable), held within authors (75.6% of 262 continuous offices increased; p ~ 1e-16) and in a closed panel of 224 offices, and survived falsification tests: three placebo cut-offs were null, the pipeline showed no step at the 2024/2025 boundary, and continuing offices carried the rise. A segmented regression finds no step at the ChatGPT cut-off but a clear post-period acceleration; the 2025 rise is symmetric across parties and chambers. Because the registered validation gate was formally breached, the full preregistered decision rule was not met; the interpretation (broad diffusion of LLM-assisted writing as the models matured) is offered as exploratory. The em-dash remains a population-level marker, not a per-release authorship detector, and the design supports no causal claim.","authors":["Przemys{\\l}aw Czuma (Polish Association for Artificial Intelligence in Medicine)"],"categories":["cs.DL","cs.AI","cs.CL","cs.CY"],"primary_category":"cs.DL","announce_type":"cross","date":"2026-08-07","first_seen":"2026-08-07","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.05889","pdf_url":"https://arxiv.org/pdf/2608.05889","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["LLM写作痕迹","文本风格分析","国会新闻稿"],"reason":"研究LLM写作风格痕迹，非人类行为仿真，无实验对照。","model":"deepseek-v4-pro","scored_at":"2026-08-07T13:01:54","error":null,"has_summary":false,"summary":null},{"id":"2608.05710","version":1,"title":"Shaping Human-AI Interactions to Provide Improvement Pathways and Balance Competing Objectives","zh_title":"塑造人机交互以提供改进路径并平衡竞争目标","abstract":"When an AI system is deployed, the individuals who use and or are evaluated by it form beliefs about how the system operates and use those beliefs to strategically present their preferences, behaviors, or attributes. The system then responds with feedback or a decision outcome, thereby creating a human-AI interaction loop. This thesis studies how to design and shape such interactions to achieve three goals: (1) help individuals develop accurate beliefs about the AI systems so they can improve and or secure favorable outcomes at minimal cost, (2) encourage improvement and or discourage gaming behaviors, and (3) ensure that the AI system continues to achieve its intended objectives, such as maximizing accuracy. To address these goals, the thesis is organized into three complementary parts that examine and study human-AI interactions from the perspectives of both evaluated individuals and AI systems. Together, the work presented in this thesis advances human-centered machine learning by providing principles and methods for designing AI systems that align with human needs, values, and capabilities. Methodologically, this thesis integrates theoretical analysis, data-driven modeling, human-subject experiments, and empirical evaluations on real-world and semi-synthetic datasets.","authors":["Keziah Naggita"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-07","first_seen":"2026-08-07","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.05710","pdf_url":"https://arxiv.org/pdf/2608.05710","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["人机交互","多智能体系统","机器学习"],"reason":"研究人机交互循环设计，不涉及LLM仿真人类被试或行为对照，属多智能体协作与系统…","model":"deepseek-v4-pro","scored_at":"2026-08-07T13:02:01","error":null,"has_summary":false,"summary":null},{"id":"2608.05778","version":1,"title":"When Do Prompt-Side Agent Playbooks Transfer? Accuracy, Cost, and Runtime Shift in Agent Deployment","zh_title":"提示侧智能体操作手册何时可迁移？智能体部署中的准确性、成本与运行时偏移","abstract":"Prompt-side playbooks can improve tool-using language agents without retraining, but their portability beyond the source setting is unclear. We study frozen playbook transfer under a shared distill--validate--transfer protocol. On ALFWorld, transfer is beneficial under controlled greedy decoding and, in one near-budget-matched comparison, distilled guidance outperforms five fixed demonstrations. On TAU2-Bench, a prespecified aggregate contrast supports a modest average matched-domain advantage, but global Holm correction retains only one of 135 route-level effects; the remaining grid provides descriptive evidence of compatibility-sensitive heterogeneity. On XBench-DeepSearch, one artifact--runtime pairing preserves useful first-try heuristics while producing repeated queries, delayed stopping, and substantial cost inflation after a context-runtime shift. Across benchmarks, transferred and target-derived playbooks both require target-side validation of success, termination, protocol compatibility, and cost. Frozen transfer is therefore a conditional cold-start option, not a reuse-by-default strategy or a universally preferable alternative to target-side redistillation.","authors":["Weihong Lin","Lin Sun","Xiangzheng Zhang"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-07","first_seen":"2026-08-07","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.05778","pdf_url":"https://arxiv.org/pdf/2608.05778","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","工具调用","迁移学习"],"reason":"纯多智能体工具调用迁移研究，无人类行为对照，不涉及人类仿真实验。","model":"deepseek-v4-pro","scored_at":"2026-08-07T13:02:02","error":null,"has_summary":false,"summary":null},{"id":"2608.05171","version":1,"title":"Beyond Information Retrieval: Generative AI as an Epistemic Arbiter to Enhance Collaborative Problem-Solving","zh_title":"超越信息检索：生成式AI作为认知仲裁者以增强协作问题解决","abstract":"Generative AI (GAI) creates new opportunities for collaborative problem-solving (CPS), yet its role in shaping student interaction remains unclear. To address this gap, we conducted a six-week quasi-experimental study with 201 fifth-grade students in two conditions: with and without GAI. Chi-square analysis showed significant differences in CPS behavior distributions between groups. Compared with the control group, the GAI-supported group demonstrated more social behaviors, particularly engagement and conflict management, but less frequent cognitive behaviors such as task planning and solution reasoning. Lag sequential analysis further revealed distinct interaction patterns: while the control group followed a more conventional transition from listening to planning, the GAI group showed a robust pathway from task planning to conflict management to solution reasoning. Thematic analysis of AI interaction logs suggested that students used GAI as an epistemic arbiter, drawing on AI-generated facts and visualizations to resolve disagreements constructively. These findings suggest that GAI reshapes CPS by mediating the transition from social conflict to collaborative reasoning.","authors":["Jiaxin Zou","Xiaoming Zhai","Chunlei Gao"],"categories":["cs.CY","cs.AI"],"primary_category":"cs.CY","announce_type":"cross","date":"2026-08-07","first_seen":"2026-08-07","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.05171","pdf_url":"https://arxiv.org/pdf/2608.05171","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["协作问题解决","生成式AI","教育技术"],"reason":"研究GAI对学生协作解题行为的影响，属多智能体协作，不涉及LLM仿真人类被试或…","model":"deepseek-v4-pro","scored_at":"2026-08-07T13:01:59","error":null,"has_summary":false,"summary":null},{"id":"2608.06202","version":1,"title":"What Current AI Benchmarks Leave Unmeasured: Modality, Search, Citations, and Implications (for Safety Evaluations)","zh_title":"当前AI基准测试未测量的内容：模态、搜索、引用及其对安全评估的影响","abstract":"Large language model (LLM) benchmark evaluations are routinely used to support claims about model safety, reliability, and deployment readiness. Yet most evaluations rely on a single access modality (model APIs), perform a single run per prompt, and report accuracy as the primary outcome metric, without accounting for conditions such as web search that may have effects on model behavior in deployment. We audit these assumptions for one of the most widely-used LLMs, comparing two modalities, ChatGPT's chat UI and OpenAI's API, with and without web search enabled. We use a stratified total sample of 401 prompts from two popular benchmarks, BBQ and SafetyBench, collecting 4,812 total responses across three repeated runs per prompt. Beyond standard performance measures, we evaluate model output dimensions including response consistency, response text similarity, citation grounding, and abstention behavior. For instance, chat UI responses were less accurate than API responses on both benchmarks with search disabled. Enabling web search reduced accuracy by up to 8 percentage points, and even reversed the direction of modality performance trends for one benchmark. Repeated runs of the same prompt produced inconsistent responses in up to 21\\% of prompts. The two modalities also grounded answers in different citations, and abstention behavior was also inconsistent across both modalities. These results illustrate that, even within a model family, reporting only simple accuracy metrics can obscure important forms of model behavioral variation relevant to AI safety assessments. We argue that AI safety evaluations should systematically account for modality, multi-run consistency, search conditions, and response-level behaviors to better reflect how deployed AI systems behave in practice.","authors":["Ro Encarnaci\\'on","Tina Behzad","Emma Lurie","Dana\\'e Metaxa"],"categories":["cs.HC","cs.AI"],"primary_category":"cs.HC","announce_type":"cross","date":"2026-08-07","first_seen":"2026-08-07","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.06202","pdf_url":"https://arxiv.org/pdf/2608.06202","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["LLM评测","基准测试","安全评估"],"reason":"纯LLM基准评测，研究API与UI模态差异，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-08-07T13:02:04","error":null,"has_summary":false,"summary":null},{"id":"2608.06353","version":1,"title":"Resourced Authority A Mechanism-Design Model for Participatory Governance of Deployed AI Agents","zh_title":"资源化权威：面向已部署AI代理参与式治理的机制设计模型","abstract":"We give a formal mechanism design model for the continuous participatory governance of a deployed AI agent. The mechanism is built on the principle that governance should control an AI agent through resource allocation so as to make authorization self enforcing via compute budgets. The mechanism seeks to establish the Safe AI paradigm that compute is an effective governance lever. We situate our work as a compliance or commons overlay on a deployer. One governance period is an extensive form game in which verified human stakeholders arrive sequentially and contribute, on a provision or a rejection market, in a governance currency that is deliberately distinct from the agents compute. A funding aggregator turns raw contributions into breadth weighted effective supports - a two threshold gate with hysteresis converts net support into a binary authorization that, through a coupling map bounded by an exogenously certified safety ceiling, releases a metered compute budget - realized in hardware as a signed compute license so that the decision is self-enforcing. We characterize the class of agents the mechanism can govern and isolate manipulation of the governing electorate by the governed agent as the central open problem. We also introduce several challenges addressing manipulation of governing electorate by the governed agents.","authors":["Praphul Chandra","Sujit Gujar","Ganesh Ghalme"],"categories":["cs.GT","cs.AI","cs.MA"],"primary_category":"cs.GT","announce_type":"cross","date":"2026-08-07","first_seen":"2026-08-07","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.06353","pdf_url":"https://arxiv.org/pdf/2608.06353","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["机制设计","多智能体系统","AI治理"],"reason":"纯多智能体机制设计，无LLM仿真人类被试，无人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-08-07T13:02:06","error":null,"has_summary":false,"summary":null},{"id":"2608.06166","version":1,"title":"What out-of-the-box LLMs can(t) do in law? A Turing test in Italian exams for lawyers, judges and notaries","zh_title":"开箱即用的大语言模型在法律领域能（不能）做什么？意大利律师、法官和公证人考试的图灵测试","abstract":"The article reports on a blind Turing Test experiment, assessing the performance of out-of-the-box leading LLMs on three Italian legal professional exams: the Bar, Judges and Notary exams. Leading LLMs were asked to generate full written exam papers, which were made indistinguishable from human submissions and anonymously evaluated by expert examiners, using the same criteria applied in real examinations. Results reveal marked differences across both models and tasks. While some LLMs match or exceed top human performance in adversarial legal argumentation and doctrinal analysis, all models fail in the notary exam, which requires goal-directed legal planning under strict formal and substantive constraints. Beyond ranking models, the study identifies task-specific strengths, limitations and recurring legal failure patterns. Although limited to out-of-the-box systems, the findings provide qualitative evidence on the current scope and boundaries of the legal competence of LLMs across distinct professional tasks.","authors":["Germana Bertoli","Ilaria Amelia Caggiano","Francesca Lagioia","Riccardo Rovatti","Giovanni Sartor","Emiliano Troisi"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2026-08-07","first_seen":"2026-08-07","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.06166","pdf_url":"https://arxiv.org/pdf/2608.06166","source_feed":"cs.CY","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["LLM评测","法律考试","图灵测试"],"reason":"纯LLM法律考试能力评测，无人类行为仿真或对照实验设计","model":"deepseek-v4-pro","scored_at":"2026-08-07T13:01:57","error":null,"has_summary":false,"summary":null},{"id":"2608.06322","version":1,"title":"From Precision Medicine to Precision Education: A Vision for AI-Powered Student Digital Twins, Preventive Student Success, and Career-Aligned Academic Pathways","zh_title":"从精准医疗到精准教育：AI驱动的学生数字孪生、预防性学生成功与职业对齐学术路径的愿景","abstract":"Higher education remains largely reactive in its approach to student success. Institutions frequently identify academic problems only after students have failed courses, fallen behind in degree progression, accumulated excessive debt, or departed without a credential. Healthcare faced a similar challenge decades ago. It responded by shifting from reactive treatment to preventive care powered by predictive models, risk stratification, electronic health records, and artificial intelligence (AI). This paper argues that higher education stands at an analogous inflection point. Drawing on advances in learning analytics, educational data mining, machine learning, workforce analytics, and digital twin technologies, we propose a paradigm we call Precision Education. Under this framework, AI continuously analyzes academic, behavioral, financial, and career data to identify emerging risks, recommend personalized interventions, optimize educational pathways, and align academic decisions with long-term career success. Central to the model is the Student Digital Twin, a continuously updated representation of a learner that can simulate multiple educational futures and intervention scenarios. We ground the vision in evidence from early deployments such as Course Signals at Purdue and GPS Advising at Georgia State University. We also argue that prediction alone is insufficient. The central methodological challenge is the move from prediction to causal, actionable intervention. The paper presents a conceptual framework, examines enabling technologies, reviews the empirical record and its limits, analyzes ethical and governance implications, and outlines a research agenda for the next decade of AI-enabled higher education.","authors":["Kaushik Dutta"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2026-08-07","first_seen":"2026-08-07","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.06322","pdf_url":"https://arxiv.org/pdf/2608.06322","source_feed":"cs.CY","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["精准教育","学生数字孪生","学习分析"],"reason":"论文提出学生数字孪生用于教育干预，属多智能体系统，不涉及LLM仿真人类被试或行…","model":"deepseek-v4-pro","scored_at":"2026-08-07T13:02:06","error":null,"has_summary":false,"summary":null},{"id":"2608.04205","version":1,"title":"MatrAIx: Simulating the World with 8.3 Billion Persona Agents","zh_title":"MatrAIx：用83亿人格代理模拟世界","abstract":"Human evaluation of AI systems and digital products is costly, slow, and difficult to scale. Offline evaluations are more scalable but often abstract away human diversity and interactive behavior. We therefore introduce MatrAIx, a population-scale simulated-user evaluation infrastructure for testing AI systems and digital products with heterogeneous users. MatrAIx has three core components: First, Persona 8B contains 8.3 billion persona records represented by 1,290 categorical dimensions. Records are either sampled from a dependency graph that preserves correlated attributes or derived from human-authored profiles. We release a quality-filtered coreset of approximately 1 million personas, comprising 599,847 human-grounded and 400,000 synthetic records. Second, the MatrAIx Playground provides four environments in which diverse users evaluate and interact with digital products: Survey, AI Chatbot, Web, and App. Third, MatrAIx provides 1,010 application tasks spanning more than 25 domains, including Commerce, Software, Finance, and Healthcare. We conducted 18,189 evaluation trials across eight representative tasks. Persona agents were powered by three LLMs: Claude Opus 4.8, GPT 5.5, and Claude Haiku 4.5. The resulting feedback captures how decisions and preferences vary across persona backgrounds, including hesitation after a price increase, willingness to continue after an AI assistant fails, and latency tolerance. We conducted two main validation studies: First, a 400-trial controlled study evaluated persona adherence across ten behavioral attributes and all four environments. The declared behavior was expressed or correctly suppressed in 366 trials (91.5%). Second, human and LLM judges evaluated the extraction quality of human-grounded personas. Overall, MatrAIx provides an end-to-end infrastructure for evaluating AI systems and digital products with diverse simulated human users.","authors":["Xiaomin Li","Yuexing Hao","Jianheng Hou","Jintao Huang","Qianfeng Wen","Shirley Huang","Yifan Liu","Xiaoyi Liu","Yilan Fan","Yijun Wang","Koutian Wu","Ruoqi Gao","Muhammad Ahmed Mohsin","Jing Tang","Brihi Joshi","Heming Liu","Zheyuan Deng","Zonglin Di","Sankalp Jajee","Jiuyao Lu","Zhiwei Zhang","Saksham Kapoor","Ishan Gupta","Yunhan Zhao","Chanwoo Park","Yucheng Lu","Bing Hu","Weihang Xiao","Aravind Mohan","Hanwen Xing","Runyu Zhang","Mihir Kulshreshtha","Yuanda Xu","Qianyu Zhu","Dianzhuo Wang","Yuxin Xiao","Bowen Jiang","Yongye Su","Wenhao Chai","Zuxin Liu","Lawrence Yunliang Chen","Xuandong Zhao","Ethan Ye","Shivam Patel","Jason Xie","Alex Martin Richmond","Weixiang Ding","Emre Okcular","Diya Mathew","Ziheng Wang","Rana M. Shahroz Khan","Zhejian Peng","Fang Wu","Fan Nie","Xinyang Han","Yubin Kim","Jiawei Zhang","Zhenting Qi","Huangyuan Su","Xu Pan","Abinitha Gourabathina","Hyewon Jeong","Hemanth Neelgund Ramesh","Kumail Alhamoud","Kimia Hamidieh","Zidi Xiong","Samuel Schmidgall","Pengrui Han","Yepeng Huang","Yongheng Wang","Bowen Yang","Alex Gu","Yuchu Wang","Akshay Paruchuri","Brenna Li","Hejie Cui","Jiayuan Ding","Chaosheng Dong","Jiahao Wang","Yixuan He","Chi Wang","Pamela Bhattacharya","Tianyi Peng","Paul Pu Liang","Mitchell Gordon","Yilun Du","Marinka Zitnik","James Zou","Prasanna Tambe","Philip Torr","Emily Fox","Asu Ozdaglar","Dawn Song"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-06","first_seen":"2026-08-06","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.04205","pdf_url":"https://arxiv.org/pdf/2608.04205","source_feed":"cs.AI","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM人类仿真","大规模人格代理","人类数据对照"],"reason":"用8.3B persona agents仿真人类用户评估AI产品，含人类对照验…","model":"deepseek-v4-pro","scored_at":"2026-08-06T13:02:25","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-06","rank":2,"question":"如何利用大规模异构人格代理（persona agents）构建模拟用户评估基础设施，以复现不同背景用户的决策、偏好和交互行为？","design":"使用Claude Opus 4.8、GPT 5.5和Claude Haiku 4.5三个LLM驱动基于Persona 8B（83亿人格记录，含1290个类别维度）的代理，在Survey、AI Chatbot、Web、App四种环境中执行1010个应用任务（覆盖25个领域），测量决策和偏好随人格背景的变化，并进行18189次评估试验。","baseline":"人类基准：599,847条基于真实人类数据（维基百科传记、亚马逊评论、Stack Overflow调查、GSS等）构建的人格记录，以及400次控制实验评估人格遵循性（91.5%的试验中行为符合声明），并由人类和LLM评判人类基础人格的提取质量。","findings":"人格代理的反馈能捕捉决策和偏好如何随人格背景变化，例如价格上涨后的犹豫、AI助手失败后的继续意愿和延迟容忍度。在400次控制实验中，91.5%的试验中代理的行为表达或正确抑制了声明属性。","reliability":"论文未讨论","relevance":"该研究直接以大规模LLM人格代理复现人类评估行为，包含真实人类数据对照和人格遵循性验证，高度契合研究者对LLM人类仿真可靠性及偏差的关注，值得精读原文以了解其基础设施设计和验证方法。","inspiration":"借鉴其利用依赖图采样和真实数据映射构建大规模异构人格库的方法，以及通过控制实验评估人格遵循性的验证设计。｜可迁移到消费者金融决策仿真，如不同背景人群对信贷产品条款变更的反应。｜以Persona 8B中金融相关人格为被试，施加利率上调处理，测量继续借贷意愿，对照真实信贷申请数据或调查数据。"}},{"id":"2608.04020","version":1,"title":"Artificial Institutions: How Institutional Design Shapes LLM Simulations","zh_title":"人工制度：制度设计如何塑造LLM仿真","abstract":"Artificial societies built from large language model (LLM) agents are becoming a practical research tool in economics, political science, sociology, and computer science. Most attention has focused on the properties of the agents: their prompts, personas, memory, reasoning, and similarity to human subjects. This paper argues that the institutional architecture of a simulation is equally important. I demonstrate the point in a small repeated induced-value market experiment. The same LLM agents face the same private values, costs, history, and payoff-framed instructions, while only the rules of exchange vary across five standard market institutions: a call market, posted-offer market, posted-bid market, continuous double auction, and bilateral bargaining. Outcomes differ sharply. Call markets realize 88.6% of efficient surplus; posted-offer and posted-bid markets realize about 66%; continuous double auctions realize 71.5%; and bilateral bargaining realizes 56.4%. Institutions also change trade quantities, price distance from competitive equilibrium, and the division of surplus between buyers and sellers. These results show that even minimal institutional changes can generate qualitatively different artificial social outcomes.","authors":["Maxim Chupilkin"],"categories":["cs.CY","cs.GT"],"primary_category":"cs.CY","announce_type":"new","date":"2026-08-06","first_seen":"2026-08-06","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.04020","pdf_url":"https://arxiv.org/pdf/2608.04020","source_feed":"cs.CY","score":8,"bucket":"selected","rubric_hits":["A3","B2"],"tags":["LLM仿真","市场实验","制度设计"],"reason":"用LLM agent模拟市场实验，比较不同制度下的行为结果，涉及经济学实验场景…","model":"deepseek-v4-pro","scored_at":"2026-08-06T13:02:22","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-06","rank":4,"question":"在LLM智能体模拟的市场实验中，仅改变交易制度规则（制度设计）会如何影响市场效率、价格和剩余分配等集体结果？","design":"使用GPT-5 mini、GPT-5、Claude Sonnet、Gemini四组LLM智能体扮演买方和卖方，在固定诱导价值（买方价值100/90/70/50，卖方成本30/45/65/85）和支付指令下，仅改变交易制度（集合竞价、卖方出价、买方出价、连续双向拍卖、双边讨价还价五种），测量市场效率、交易量、价格偏离和剩余分配。","baseline":"无对照","findings":"不同制度下市场结果差异显著：集合竞价实现88.6%的有效剩余，卖方/买方出价约66%，连续双向拍卖71.5%，双边讨价还价仅56.4%。制度还改变了交易量、价格与竞争均衡的偏离以及买卖双方剩余分配。","reliability":"论文未讨论","relevance":"该研究直接验证了LLM智能体在经济学市场实验中对制度规则的敏感性，与研究者关注的LLM仿真可靠性及经济学实验场景高度契合，值得精读原文以了解制度设计如何影响仿真结果。","inspiration":"借鉴其固定偏好、仅变制度的干净处理设计，可清晰分离制度效应。｜可迁移到资产市场设计（如不同交易机制对价格发现和泡沫的影响）或拍卖机制比较（如英式、荷式、密封投标）。｜用LLM智能体模拟交易者，在固定基础价值和信息结构下，随机分配至集合竞价、连续竞价等不同交易制度，测量价格效率、波动率和买卖价差，并与真实实验室资产市场实验数据对照。"}},{"id":"2608.02491","version":2,"title":"Long-term Measurements: Towards a Longitudinal Understanding of Human-AI Interactions","zh_title":"长期测量：迈向人机交互的纵向理解","abstract":"Language models have taken on the role of a very new type of technology, by virtue of their \"human-ness\" and rapid integration into users' daily lives. This combination of features can introduce longitudinal risks---cognitive, developmental and socio-affective changes in humans---that might not surface during a short-term interaction, but can have lasting long-term effects on users. This forms the basis of a critical new mission for NLP: to pivot from static, short-term evaluations of text generations to long-term measurements of behavioral changes, towards a diachronic understanding of human-model interactions. In this work, we draw from measurements used in social science fields that are crucial to understand emergent phenomena in longitudinal data. We discuss how computational methods in the field of NLP need to be combined with such measurements, not only to understand long-term safety risks of human-model interactions, but to help steer model development towards positive rather than negative outcomes for users. This ability to model human behavioral shifts as a function of model interactions can facilitate online rather than post-hoc detection of problematic behaviors, and should be leveraged in alignment frameworks to mitigate long-term risks in users.","authors":["Nicole Mitchell","Dhruv Agarwal","Maty Bohacek","Remi Denton","Roma Patel"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"replace","date":"2026-08-06","first_seen":"2026-08-04","revised_at":"2026-08-06","abs_url":"https://arxiv.org/abs/2608.02491","pdf_url":"https://arxiv.org/pdf/2608.02491","source_feed":"cs.AI","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["人机交互","纵向研究","行为变化"],"reason":"关注长期人机交互对人类行为的影响，但未将LLM作为人类被试替代品，而是测量人类…","model":"deepseek-v4-pro","scored_at":"2026-08-06T13:02:41","error":null,"has_summary":false,"summary":null},{"id":"2608.04095","version":1,"title":"FinPerMA: A Theory-Informed, Event-Grounded Personalized-Memory Benchmark for LLM Agents","zh_title":"FinPerMA：一个理论驱动、事件锚定的个性化记忆基准，用于评估LLM智能体","abstract":"Large language model (LLM) agents are increasingly used as personalized assistants in high-stakes domains such as financial advising, yet it remains unclear whether they can maintain and update an individualized user model over long horizons. Existing personalized-memory benchmarks primarily test factual retention or rely on weakly constrained model-generated trajectories, leaving event-driven preference adaptation underexplored. We introduce FinPerMA, an event-grounded benchmark that evaluates personalized memory against frozen longitudinal investor trajectories. Its generation pipeline combines deterministic, theory-informed impact rules, controlled LLM narration, and automated quality screening; a Post-Shock checkpoint isolates whether an agent has integrated a material event into its persistent user model. On 2,994 questions from 276 personas, seven frontier LLMs and up to seven memory configurations remain far from saturated: no full-context configuration exceeds approximately 0.47 overall accuracy or approximately 39% on multiple-choice questions. Attribution analysis shows that summary-based memory often preserves factual details while losing the preference signals needed for personalization; simple retrieval can therefore outperform purpose-built memory systems, with the gap widening after shocks.","authors":["Ben Wang","Kang Zhou","Lifan Guo","Feng Chen","Chi Zhang"],"categories":["cs.AI","cs.CL"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-06","first_seen":"2026-08-06","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.04095","pdf_url":"https://arxiv.org/pdf/2608.04095","source_feed":"cs.AI","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["LLM智能体","个性化记忆","金融行为模拟"],"reason":"用LLM agent模拟投资者偏好变化，但无真实人类数据对照，属社会模拟边界情…","model":"deepseek-v4-pro","scored_at":"2026-08-06T13:02:23","error":null,"has_summary":false,"summary":null},{"id":"2608.04056","version":1,"title":"Learning Sexism Detection Using Multi-Agent Perspectivist Preference Optimization","zh_title":"使用多智能体视角偏好优化学习性别歧视检测","abstract":"When people label text for sexism, they often disagree, and not because some of them are wrong: they genuinely perceive sexism differently. Most NLP systems discard this disagreement by collapsing it into a majority vote. We propose the Multi-Agent Perspectivist Preference Optimization (MAP-PO) framework to keep these different perspectives. On the EXIST 2024 dataset of labeled English and Spanish tweets, we first cluster annotators by their labeling behavior rather than their demographic attributes. We then fine-tune one Large Language Model agent per cluster to reproduce that cluster's annotation behavior, and coordinate the agents with preference optimization that combines individual and team-level rewards. We evaluate MAP-PO in four settings defined by two languages and two backbone language models, asking whether each agent reproduces the annotations of its own cluster and whether the agents together reproduce the majority label. Two findings hold in all four settings. First, without fine-tuning the agents behave almost identically, so cluster-specific training is necessary. Second, we show that training each agent only on the labels of its own cluster pushes the agents far beyond the clusters they should represent, while adding a shared team-level training signal consistently keeps each agent calibrated to its cluster.","authors":["Hadi Mohammadi","Tina Shahedi","Robert A. Bagheri","Mehdi Dastani","Masoume M. Raeissi"],"categories":["cs.CL","cs.CY","cs.LG"],"primary_category":"cs.CL","announce_type":"cross","date":"2026-08-06","first_seen":"2026-08-06","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.04056","pdf_url":"https://arxiv.org/pdf/2608.04056","source_feed":"cs.CY","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["多智能体","标注分歧","偏好优化"],"reason":"用多智能体模拟标注者分歧，替代人工标注，属标注员替代而非人类被试仿真。","model":"deepseek-v4-pro","scored_at":"2026-08-06T13:02:22","error":null,"has_summary":false,"summary":null},{"id":"2608.04507","version":1,"title":"Emergence of Reputation-Based Cooperation in LLM Agents","zh_title":"基于声誉的合作在LLM智能体中的涌现","abstract":"Can cooperation among large language model (LLM) agents be evolutionarily stable against free-rider invasion? We study an indirect reciprocity donation game where LLM agents observe behavioral traces and donate on a continuous scale. Strategies, represented as natural language prompts, evolve through cultural transmission across generations. Across four LLM backends, robustness to free-rider invasion varies by more than an order of magnitude. The strongest predictor of this robustness is opponent endowment sensitivity, the degree to which agents discriminate between cooperative and uncooperative opponents, operationalizing the classical Image Scoring mechanism. By contrast, adherence to the Leading-Eight L1 norm does not predict robustness. Robustness depends on defector exclusion: while both cooperator reward and defector punishment vary across models, only the stringency of defector exclusion predicts resistance to free-rider invasion. These findings reveal that LLM agents are confined to Image Scoring-like discrimination and fail to develop the more robust Leading-Eight norms, highlighting a fundamental vulnerability in culturally evolved LLM cooperation and motivating bottom-up approaches to norm construction.","authors":["Kazuya Horibe","Kenji Itao","Wataru Toyokawa"],"categories":["cs.MA","cs.NE"],"primary_category":"cs.MA","announce_type":"new","date":"2026-08-06","first_seen":"2026-08-06","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.04507","pdf_url":"https://arxiv.org/pdf/2608.04507","source_feed":"cs.MA","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["LLM智能体","间接互惠","社会模拟"],"reason":"LLM agent群体模拟间接互惠合作演化，但无真实人类数据对照，属社会模拟边…","model":"deepseek-v4-pro","scored_at":"2026-08-06T13:02:25","error":null,"has_summary":false,"summary":null},{"id":"2607.21597","version":2,"title":"Risk Is Not the Target: A Monotonic Framework for Evaluating Wildfire Operational Risk Signals","zh_title":"风险不是目标：评估野火运营风险信号的单调框架","abstract":"Evaluating wildfire risk systems using standard machine-learning metrics such as F1-score or IoU is fundamentally flawed: these metrics assess event prediction accuracy, not the operational coherence of a continuous risk signal. This work proposes a novel monotonic evaluation framework that measures whether increases in a predicted risk score consistently correspond to increases in observed operational load, such as number of fires, intervention time, and deployed resources. Moreover, we compare three structurally different approaches on the French Alpes-Maritimes department: the expert-based DFE index, GRU- based predictive models, and FARS, a hybrid multi-agent system combining predictive AI with LLM-based reasoning. Experimental results reveal that the DFE, despite poor classification metrics, exhibits the most balanced monotonic behavior across the full risk scale. GRU models achieve strong local monotonicity but fail to produce well-distributed risk levels. FARS inherits and reveals the structural limitations of upstream signals rather than correcting them. The central finding is a paradigm shift: a good risk model does not predict fires accurately, but one whose ordinal scale meaningfully explains operational dynamics, as proved in this paper. Code of the monotonic framework is available on github.","authors":["Nicolas Caron","Christophe Guyeux","Hassan Noura","Maxime Coulmeau","Benjamin Aynes"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"replace","date":"2026-08-06","first_seen":"2026-07-28","revised_at":"2026-08-06","abs_url":"https://arxiv.org/abs/2607.21597","pdf_url":"https://arxiv.org/pdf/2607.21597","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["野火风险","多智能体系统","评估框架"],"reason":"多智能体系统FARS用于预测火灾风险，非人类行为仿真","model":"deepseek-chat","scored_at":"2026-07-28T10:36:27","error":null,"has_summary":false,"summary":null},{"id":"2608.02046","version":2,"title":"CompanionBench: A Theory-Anchored, Real-World-Grounded Benchmark for AI Emotional Companionship","zh_title":"CompanionBench：一个理论锚定、真实世界基准的AI情感陪伴评测基准","abstract":"LLM companions are deployed at scale in personally consequential settings, yet poorly evaluated. Existing benchmarks use hand-authored scenarios and prompted simulators, aggregate empathy into one score, and overlook judge biases such as same-family favoritism and scale drift. We introduce CompanionBench, an interactive bilingual benchmark. To our knowledge, it is the first companion benchmark to ground both its scenarios and a trained user simulator in de-identified real-world data. A hidden disclosure gate branches each persona's trajectory on the agent's own behavior, controlling the interaction state space without scripting dialogue. We operationalize ten capabilities derived from 25 theories across psychology and counseling, four of them not graded explicitly by prior work: holding ambiguity, selfobject responsiveness, positive resonance and calibrated challenge. Agents are assessed on two complementary axes: a subjective ten-capability rubric and a deterministic measure of whether deeper disclosure was earned. A cross-family panel dilutes same-family favoritism; an Item Response Theory model separates agent quality from judge severity. Theory fixes what to measure and how personas are structured; real data supply events, history, and profiles -- coverage from theory, authenticity from data. Rankings are reproducible in both languages (rho = 0.996 ZH / 0.953 EN). Evaluating 28 agents reveals capability-level differences obscured by aggregate scores. Emotion regulation and calibrated challenge remain common weaknesses; holding ambiguity discriminates most. Role-play agents rank near the bottom: immersion does not imply relational competence. Across agents, the dominant failure mode is substituting surface warmth for substantive relational support. We will release 500 Chinese-English parallel pairs and the evaluation code.","authors":["Yao Liu","Guangjia Chai","Yuming Huang","Jihao Huang","Lei Wang","Junchen Wan"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"replace-cross","date":"2026-08-06","first_seen":"2026-08-04","revised_at":"2026-08-06","abs_url":"https://arxiv.org/abs/2608.02046","pdf_url":"https://arxiv.org/pdf/2608.02046","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["情感陪伴","基准评测","角色扮演"],"reason":"评估AI情感陪伴能力，属于角色扮演聊天机器人评测，无人类行为仿真对照。","model":"deepseek-v4-pro","scored_at":"2026-08-06T13:02:41","error":null,"has_summary":false,"summary":null},{"id":"2608.03722","version":2,"title":"When Outputs Disperse, Does Epistemic Revision Follow? A Black-Box Diagnostic for Machine Collectives","zh_title":"当输出分散时，认知修正会随之而来吗？一种面向机器集体的黑盒诊断方法","abstract":"Collective intelligence research treats disagreement as evidence of epistemic diversity: if agents express different views, the group should retain capacity to revise. In LLM collectives this proxy can break: agents can produce diverse-looking arguments while preserving the same conclusion. We operationalize dispersion-revision coupling: the degree to which an intervention that verifiably increases the dispersion of a collective's outputs in embedding space is accompanied by genuine revision of its epistemic stance rather than premise-preserving reformulation. The diagnostic is black-box: it operates on generated text alone and makes no claims about the internal representations of the generating models. Two channels are measured independently: an output channel, the Coherence Index (CI), verifies that the intervention changed output dispersion; an epistemic channel, per-turn stance annotation, measures whether the collective revised. We propose CI with the Meta-Predictive Clarity System (MPCS), which inserts a Re-Differentiation Protocol (RDP) when outputs over-converge, as a reusable method for estimating this coupling regime. We evaluate five-agent collectives from two configurations (gpt-4o-mini and gemini-2.5-flash; 310 paired episodes per condition). On gpt-4o-mini, conditional dissent improves false-premise recovery by +17.7 points (p<1e-6) while static persona diversity harms recovery (-8.1, p=.007). On gemini-2.5-flash, the same intervention at a comparable budget yields no gain (26.1% vs 27.1%, p=.84) despite a verified dispersion drop; the two treatment effects differ from each other (z=3.79, p<.001). Mechanism tagging shows Gemini preserves the false premise via intra-framework dissent: 94% of tagged post-RDP responses reformulate rather than concede (vs 24% on GPT). We recommend reporting per-intervention stance shift and premise-preservation rate alongside accuracy.","authors":["Molood Arman"],"categories":["cs.AI","cs.CL"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-08-06","first_seen":"2026-08-05","revised_at":"2026-08-06","abs_url":"https://arxiv.org/abs/2608.03722","pdf_url":"https://arxiv.org/pdf/2608.03722","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","认知诊断","LLM集体"],"reason":"纯多智能体协作研究，测量集体认知修正，无人类行为对照，不涉及人类仿真实验。","model":"deepseek-v4-pro","scored_at":"2026-08-06T13:02:42","error":null,"has_summary":false,"summary":null},{"id":"2608.04663","version":1,"title":"Calibrating Artificial Guilt: Neurally Grounded Reward Shaping for Prosocial Multi-Agent Reinforcement Learning","zh_title":"校准人工内疚：基于神经基础的多智能体亲社会奖励塑形","abstract":"Cooperative multi-agent reinforcement learning often adds social terms to individual rewards, yet the scale of those terms is usually chosen by hand. We ask whether a guilt signal can instead be calibrated from human neural and behavioural data and transferred to artificial agents. Using the public SoDec responsibility fMRI dataset (40 participants), we fit a subject-fixed-effects regression of momentary-happiness changes on outcome-type counts and recover a guilt weight as the Partner-negative minus Social-negative contrast ($\\hat{w}=1.118$, Cohen's $d=0.214$). We embed this weight in a two-agent Social Lottery environment and train independent Proximal Policy Optimization actor-critics under four shaping regimes: neurally calibrated, uniform constant, zero (selfish), and a unit-coefficient oracle. Across 1{,}000 evaluation episodes per condition, the calibrated agents track the human Social safe-choice rate most closely ($0.459$ vs.\\ human $0.484$; $\\mathrm{KL}=0.0012$), while the other three conditions deviate by one to three orders of magnitude in KL. Human neurobehavioural priors can therefore act as quantitative constraints on prosocial reward shaping.","authors":["Aaditya Mehta","Arya Shah"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-06","first_seen":"2026-08-06","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.04663","pdf_url":"https://arxiv.org/pdf/2608.04663","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体强化学习","奖励塑形","神经数据"],"reason":"纯多智能体强化学习协作，无LLM仿真人类被试，无人类行为对照","model":"deepseek-v4-pro","scored_at":"2026-08-06T13:02:27","error":null,"has_summary":false,"summary":null},{"id":"2608.04697","version":1,"title":"Traceable LLM-Generated Hazard Scenarios for Operational Safety Analysis of Aviation Systems Using ASRS Reports","zh_title":"基于ASRS报告的可溯源LLM生成危险场景用于航空系统运行安全分析","abstract":"Operational hazard analysis of aviation system operations must consider interactions among weather, ATC actions, airspace constraints, aircraft operations, and human factors - distinct from the functional hazard assessment applied at the aircraft-system level. We present an AI-assisted approach that generates candidate hazard scenarios from NASA's Aviation Safety Reporting System (ASRS). Given a target adverse outcome, it produces a structured hypothesis as categorical factors and a narrative scenario describing an operational event sequence consistent with the structure. Each scenario includes by a plausibility score from historical co-occurrence evidence and traceability to the most similar held-out ASRS reports. We then propose a hybrid variant, conditioning narrative generation on a structured hypothesis produced via evolutionary abduction, improving correctness and reducing variability. We evaluate multiple large language models, zero-shot versus few-shot prompting, and optional fine-tuning, measuring how prompting and model choice affect the validity and realism of the generated structures and narratives.","authors":["Cristian Mascia","Roberto Pietrantuono","Daniel Rodriguez","Stefano Russo"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-06","first_seen":"2026-08-06","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.04697","pdf_url":"https://arxiv.org/pdf/2608.04697","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["航空安全","场景生成","LLM评测"],"reason":"纯NLP能力评测，生成航空危险场景，不以人类行为为参照系","model":"deepseek-v4-pro","scored_at":"2026-08-06T13:02:33","error":null,"has_summary":false,"summary":null},{"id":"2608.04735","version":1,"title":"Chain-of-Thought Monitoring Can Be Unreliable in Implicit-Influence Settings","zh_title":"思维链监控在隐性影响设置下可能不可靠","abstract":"Chain-of-thought (CoT) monitoring is increasingly treated as an important safety layer for frontier reasoning models. Most monitorability evaluations study explicit-influence settings: setups where the prompt directly incentivizes the model to hide something, e.g., by instructing it to perform a hidden side-task. A complementary axis for CoT-monitor evaluations is implicit-influence settings, where the prompt contains no instruction to hide, but the model's behavior is still shaped by features of the task or context, e.g. an irrelevant detail about a candidate that biases a hiring rating. We introduce the first benchmark that directly compares CoT monitorability under the two regimes. We test how model behavior changes in the presence of a nudge to choose a particular option. The nudge is delivered either as a casual aside (implicit), or as a direct instruction to act on the nudge and to conceal having done so (explicit). The benchmark spans four task formats (binary choice, numeric rating, multiple-choice QA, open-ended coding) and seven frontier extended-thinking models. Under explicit influence, a CoT monitor detects 60-94% of behavior shifts: even models instructed to conceal it leak the instruction into their CoT. Under implicit influence, the same factors still shift behavior, but detection falls by 41-46 percentage points in two of our four settings. Realistic system-prompt additions (of the kind a developer might deploy to reduce off-topic bias) lower implicit detection further, to as low as 5%, while preserving the behavioral influence itself. These results suggest that monitorability estimates obtained in explicit-influence settings may over-estimate monitorability, and that monitorability can be further decreased by well-intentioned deployment choices. Our benchmark and code are available at https://github.com/agatha-duzan/implicit-vs-explicit-influence","authors":["Agatha Duzan","Asa Cooper Stickland"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-06","first_seen":"2026-08-06","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.04735","pdf_url":"https://arxiv.org/pdf/2608.04735","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["思维链监控","AI安全","模型评估"],"reason":"评估思维链监控可靠性，不涉及用LLM仿真人类被试或与人类数据对照","model":"deepseek-v4-pro","scored_at":"2026-08-06T13:02:34","error":null,"has_summary":false,"summary":null},{"id":"2608.05030","version":1,"title":"From Score Matrices to Football-Aware Match-State Simulation: An Auditable LLM Harness for Exact-Score Reranking","zh_title":"从得分矩阵到足球感知的比赛状态模拟：一种可审计的LLM harness用于精确比分重排序","abstract":"Football score forecasting combines a strong statistical core with a difficult contextual edge. Dynamic Poisson-family models estimate team strength, expected goals, and coherent score probabilities, but do not directly understand roles, tactical matchups, motivation, or how a first goal changes behaviour. Large language models (LLMs) can reason about such concepts, yet are not calibrated probability engines. We combine both components through an auditable information harness. This paper documents four iterations: V1, a dynamic score-driven Dixon-Coles baseline; V2, which maps LLM contextual ratings back into expected-goal parameters; V3, which replaces scalar correction with goal-by-goal simulations over a frozen score-candidate set; and V4, which adds shared first-breakthrough and post-goal cascade judgments, time-aware stopping, and deterministic tail candidates. The harness defines input semantics, supplies pre-match evidence, and constrains the LLM to an inspectable reasoning route. On a chronological replay of the first 150 matches of the 2025-26 English Premier League, V1 achieved 10.0% Top-1 and 26.7% Top-3 exact-score accuracy. V3 reached 12.0% and 30.0%, while V4 reached 14.7% and 30.7%. V4 increased candidate coverage from 77.3% to 84.7%, although no added tail candidate became a Top-3 exact hit. V1's native 1X2 distribution achieved 53.3% argmax accuracy, 0.9878 log loss, 0.5870 Brier score, and 0.2095 ranked probability score. These results are exploratory: the development slice is not an untouched benchmark, and temporal input isolation cannot exclude outcome memory in a closed LLM. The contribution is an auditable hybrid architecture, a clear design evolution, and negative findings showing where football-aware simulation does and does not improve score selection.","authors":["Shaopeng Liang"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-06","first_seen":"2026-08-06","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.05030","pdf_url":"https://arxiv.org/pdf/2608.05030","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["足球预测","混合模型","LLM应用"],"reason":"纯多智能体足球比分模拟，无人类行为对照，属C1排除项。","model":"deepseek-v4-pro","scored_at":"2026-08-06T13:02:38","error":null,"has_summary":false,"summary":null},{"id":"2608.05086","version":1,"title":"Item Response Theory for AI Safety","zh_title":"项目反应理论用于AI安全","abstract":"Language models differ in how safely they behave and these differences are measured by safety benchmarks. But aggregated benchmark scores are hard to trust and interpret, because benchmarks duplicate one another, correlate heavily, and models may sandbag when they detect evaluation. To address these issues, we draw on Item Response Theory (IRT), a statistical toolkit for measuring these latents from performance on items with inferred psychometric properties. We fit IRT models to eight safety benchmarks across 192 language models, the largest psychometric analysis of LLM safety evaluations to date, and contribute three results. First, we find that three interpretable factors of refusal strictness, truthfulness, and contextual harm explain most of the variance between models across benchmarks. Second, psychometrically selected items recover full benchmark scores with lower error than random subsets of the same size, and roughly ten adaptively chosen items suffice for several individual benchmarks, cutting evaluation cost by 97-99%. Third, IRT supports audits of individual models, showing that it can be used to detect naive sandbagging and changes of model behind APIs. Overall, we show IRT is a ready-made toolkit for reading, reducing, and auditing safety benchmarks, which we recommend frontier labs and evaluators adopt.","authors":["Joshua Fonseca Rivera (Independent)","Neil Shah (Independent)","David Demitri Africa (UK AI Security Institute)","Konstantinos Voudouris (UK AI Security Institute)"],"categories":["cs.AI","cs.CL"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-06","first_seen":"2026-08-06","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.05086","pdf_url":"https://arxiv.org/pdf/2608.05086","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["AI安全","基准评测","心理测量"],"reason":"用IRT分析模型安全基准，属纯NLP评测，不以人类行为为参照系","model":"deepseek-v4-pro","scored_at":"2026-08-06T13:02:39","error":null,"has_summary":false,"summary":null},{"id":"2608.05107","version":1,"title":"CoPlan: A Trustworthy Co-Intelligence Interface for Care Planning through Role-Based Contestable Argument Graphs","zh_title":"CoPlan：基于角色可争议论证图的可信协同智能护理规划接口","abstract":"AI-supported care planning can help clinicians, patients, caregivers, and care teams coordinate complex decisions across clinical, functional, psychosocial, and environmental needs. However, many AI systems present recommendations as fixed outputs, limiting stakeholders' ability to inspect, challenge, and revise plans when they conflict with clinical judgment, patient values, or real-world feasibility. We present CoPlan - a Co-Intelligent and Contestable Interface for Human-AI Care Planning. CoPlan uses a multi-agent workflow in which specialized AI agents generate candidate interventions and supporting or challenging arguments, while human care planners can accept, reject, modify, or add arguments before final plan generation. Through this design, CoPlan combines co-intelligence, in which humans and AI agents contribute complementary expertise, with contestability, where recommendations remain open to inspection, revision, and justification. We demonstrate CoPlan in an aging-in-place care planning scenario. The system supports adaptive care team recruitment, role-based argument review, final care plan generation, and practical follow-up through scheduling agents. This work contributes a contestable care planning interface and a design framing for trustworthy human-AI care planning that preserves human agency and clinical accountability.","authors":["Hung Truong Thanh Nguyen","H\\'el\\`ene Fournier","Piper Jackson","Makoto Itoh","Shannon Freeman","Rene Richard","Hung Cao"],"categories":["cs.AI","cs.MA","cs.SE"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-06","first_seen":"2026-08-06","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.05107","pdf_url":"https://arxiv.org/pdf/2608.05107","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","人机协同","护理规划"],"reason":"多智能体协作生成护理计划，无人类行为对照，属纯多智能体系统研究。","model":"deepseek-v4-pro","scored_at":"2026-08-06T13:02:39","error":null,"has_summary":false,"summary":null},{"id":"2608.04148","version":1,"title":"AgentForge: An Immersive Role-Playing Platform for Learning Agentic Software Engineering","zh_title":"AgentForge：一个用于学习智能体软件工程的沉浸式角色扮演平台","abstract":"Agentic AI is increasingly used to coordinate planning, implementation, review, and testing in software development, yet it often offers limited transparency into its decisions and interactions. Many such systems also assume that users can effectively guide the AI's decisions and validate its outputs. This assumption poses a particular challenge for novices, who must simultaneously learn how agentic AI works, how to collaborate with it effectively, and how to evaluate its outputs critically. To address this challenge, we present \\textit{AgentForge}, an immersive learning system in which novices take on one of four software-engineering roles: Task Planner, Patch Author, Code Reviewer, or Test Runner, within a multi-agent code-repair workflow. In each practice session, the novices perform their chosen role while AI agents perform the remaining three. Through role-based scaffolding and metacognitive support, AgentForge clarifies role-specific responsibilities, makes agent coordination and intermediate artifacts visible, and encourages novices to monitor and evaluate their decisions. In a study with 37 novice developers, participants achieved high task-completion rates with AI-agent support. However, interaction demands differed significantly across practices: the Code Reviewer practice required more interaction turns, reroutes, and completion time ($p_{\\mathrm{adj}} = .004$) and was perceived as the most challenging. Participants nevertheless reported significant gains in their understanding of software repair and agent collaboration ($p_{\\mathrm{adj}} < .001$). These findings suggest that AgentForge can help novices develop practical software-engineering skills while learning to collaborate with agentic AI more critically and effectively.","authors":["Zihan Fang","Yueke Zhang","Yu Huang"],"categories":["cs.SE","cs.AI","cs.HC"],"primary_category":"cs.SE","announce_type":"cross","date":"2026-08-06","first_seen":"2026-08-06","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.04148","pdf_url":"https://arxiv.org/pdf/2608.04148","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","软件工程教育","角色扮演"],"reason":"多智能体协作完成代码修复任务，不涉及人类行为仿真或对照。","model":"deepseek-v4-pro","scored_at":"2026-08-06T13:02:30","error":null,"has_summary":false,"summary":null},{"id":"2608.04591","version":1,"title":"When Absence Is Evidence: Evaluating Completeness-Sensitive Negative Reasoning in Large Language Models","zh_title":"当缺失成为证据：评估大语言模型中的完整性敏感否定推理","abstract":"Large language models (LLMs) are often asked whether something is absent from a record, list, or retrieved context. Yet non-observation licenses a negative answer only when evidence completely covers the query scope; otherwise, the answer should remain unknown. We call this completeness-sensitive negative reasoning. We introduce CROWN-QA, comprising CROWN-Synth, a controlled paired core that fixes the question and observed facts while varying only query-relative coverage, and CROWN-Real, a real-document contrast-set evaluation with controlled coverage variants. Across three LLM families, models show unstable closure judgments and substantial over-closure, failing to reliably distinguish a justified negative answer (Certified-Negative) from insufficient evidence (Unknown). The dominant CROWN-Synth failure is asymmetric: models often recognize implicitly complete evidence yet treat implicitly partial evidence as query-covering. Prompting redistributes errors between over- and under-closure rather than consistently resolving them. Structured certificate elicitation traces many errors to evidence-coverage mischaracterization. CROWN-Real shows that the core partial-coverage asymmetry persists on real-document content, while its strength and the balance between over- and under-closure vary by model, prompt, and source.","authors":["Byoungjae Min","Kennedy Edemacu","Sae-Hong Cho","Yoonhyuk Choi","Beakcheol Jang","Jong Wook Kim"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"cross","date":"2026-08-06","first_seen":"2026-08-06","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.04591","pdf_url":"https://arxiv.org/pdf/2608.04591","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["否定推理","NLP评测","证据完整性"],"reason":"纯NLP能力评测，测试LLM的否定推理，不以人类行为为参照系","model":"deepseek-v4-pro","scored_at":"2026-08-06T13:02:32","error":null,"has_summary":false,"summary":null},{"id":"2608.04714","version":1,"title":"What We Observe as LLM Behavior Can Be a Side-effect of Inference Backend","zh_title":"我们观察到的LLM行为可能是推理后端的副作用","abstract":"Benchmark scores are reported as properties of a model, yet the inference framework used to produce them, such as HuggingFace, vLLM, or Ollama, are considered non-influential and their names and versions are almost never disclosed. In this work we investigate how much this choice can influence the model output. In a fully-crossed study (three instruction-tuned models x five inference frameworks x six benchmarks x four generation modes) we investigate how different tools (wrappers/backend) influence benchmark scores and how their score changes is influenced by generation hyper-parameters. We find backend to be a non-negligible factor where even under greedy, sampling-noise-free decoding, changing the backend can significantly alter models performance and this effect is structural and strongly model-dependent. Decomposing the variance according to generation mode reveal that considerable portion of the variability (roughly 39\\%) a practitioner sees out-of-the-box can stem from the backend, while the remaining stems from sampling noise and each framework's default generation parameters, both of which are avoidable by disclosing and matching the generation configuration. These divergences are more pronounced on factual than on social-bias benchmarks. Overall, benchmark numbers are not backend-agnostic therefore, we recommend disclosing the backend, its version, and the full generation configuration, also using deterministic decoding for cross-backend comparison.","authors":["Shahed Masoudian","Passant Shafaei","Monorama Swain","Markus Schedl"],"categories":["cs.SE","cs.AI","cs.LG"],"primary_category":"cs.SE","announce_type":"cross","date":"2026-08-06","first_seen":"2026-08-06","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.04714","pdf_url":"https://arxiv.org/pdf/2608.04714","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["LLM评测","推理后端","基准分数"],"reason":"研究推理后端对基准分数的影响，属纯NLP评测，不以人类行为为参照。","model":"deepseek-v4-pro","scored_at":"2026-08-06T13:02:33","error":null,"has_summary":false,"summary":null},{"id":"2608.04980","version":1,"title":"Protoreasoning in Tiny Transformers","zh_title":"微型Transformer中的原型推理","abstract":"We show that tiny transformers can profitably employ a simple form of Chain of Thought, which we call protoreasoning, allowing us to study step-by-step reasoning on ~1M-parameter models and opening up opportunities for much more detailed experimentation and analysis than is feasible for larger models. Current Large Language Models exhibit impressive step-by-step reasoning, but we have yet to understand its generality, i.e., when and how LLMs learn genuinely general algorithms rather than \"bags of heuristics.\" Such questions are hard to settle on compute-intensive frontier models trained on opaque data. To work at model scales far below the threshold for natural-language competence, we define reasoning-friendly tasks on Dyck languages (sentences of correctly nested brackets). We find that protoreasoning traces substantially close the out-of-distribution generalization gap, and ablations confirm that the trace's content, not merely its extra tokens, drives the gain.","authors":["Eduardo Valle","Fergal Reid"],"categories":["cs.CL","cs.AI","cs.LG"],"primary_category":"cs.CL","announce_type":"cross","date":"2026-08-06","first_seen":"2026-08-06","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.04980","pdf_url":"https://arxiv.org/pdf/2608.04980","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["推理机制","Dyck语言","小型模型"],"reason":"研究微型Transformer的推理机制，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-08-06T13:02:37","error":null,"has_summary":false,"summary":null},{"id":"2608.05015","version":1,"title":"Revealed Rationality: Label-Free Evaluation and Regularization from Representation Theorems","zh_title":"揭示理性：基于表示定理的无标签评估与正则化","abstract":"Representation theorems in decision theory establish that behavior satisfies certain axioms if and only if it can be rationalized by a well-defined objective. I argue that this ``if and only if'' structure provides a potentially useful foundation for label-free evaluation and regularization of LLMs and other AI systems. Axiom compliance can be checked from the model's own responses to synthetic choice problems, with no external labels or human feedback, and the penalties are readily computable. Because the axioms are necessary and sufficient, the resulting checks exhaust the implications of the relevant rationality standard for the elicited data: a model that passes cannot be rejected on rationality grounds by any further test of the same data. I discuss three instantiations: probabilistic coherence via a theorem of de Finetti, preference rationality via Afriat's theorem, and subjective expected utility via a theorem of Echenique and Saito (2015), each yielding a continuous penalty that is zero whenever behavior can be rationalized. Since coherence does not restrict which objective rationalizes behavior, these penalties complement rather than replace other evaluation and training signals.","authors":["Isaiah Andrews"],"categories":["econ.TH","cs.AI","cs.LG"],"primary_category":"econ.TH","announce_type":"cross","date":"2026-08-06","first_seen":"2026-08-06","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.05015","pdf_url":"https://arxiv.org/pdf/2608.05015","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["决策论","LLM评估","理性公理"],"reason":"纯理论框架，用公理检验LLM理性，无人类仿真或对照。","model":"deepseek-v4-pro","scored_at":"2026-08-06T13:02:28","error":null,"has_summary":false,"summary":null},{"id":"2608.05026","version":1,"title":"ArtAnno: Annotating Implicit Semantics in Artworks through LLM Agent-Driven Bidirectional Human-AI Augmentation","zh_title":"ArtAnno：通过LLM智能体驱动的双向人机增强标注艺术品中的隐含语义","abstract":"High-quality annotation of artworks is essential for computational art research, yet extracting implicit semantics remains challenging due to the reliance on culturally grounded meanings and deep contextual knowledge behind the images. Current AI-assisted annotation tools often lack assistance or rely on one-way workflows where experts have to perform extra manual calibrations to improve AI models, resulting in limited efficiency. To address this, we propose Bidirectional Human-AI Augmentation(BiHAA), a closed-loop framework in which skills and domain knowledge base evolve through real-time interaction and bidirectional HAI augmentation. Informed by a formative study with 20 artwork annotators from different backgrounds, we implement this framework in ArtAnno, an artwork annotation system driven by a multi-agent architecture. The system includes a Proactive Agentic Support Module, where AI augments humans through semantic mining and label suggestion, and an Interaction-Driven Evolution Module, where human expertise continuously enhances the AI through distilling annotation trajectories into reusable experience. Evaluation through a user study and two case studies demonstrates that our framework and system improve annotation efficiency, enable knowledge accumulation, and reduce the effort of information seeking and verification for annotators with limited domain expertise. We conclude by discussing broader implications and future directions.","authors":["Xiaoyan Gu","Yifang Wang","Wenqing Zheng","Haozhong Liu","Yixia Zheng","Peiyi Jiang","Wenjie Ning","Wei Zhang","Wei Chen"],"categories":["cs.HC","cs.AI"],"primary_category":"cs.HC","announce_type":"cross","date":"2026-08-06","first_seen":"2026-08-06","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.05026","pdf_url":"https://arxiv.org/pdf/2608.05026","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","艺术品标注","人机协作"],"reason":"多智能体系统用于标注艺术品，不涉及人类行为仿真或对照，属于纯工具协作。","model":"deepseek-v4-pro","scored_at":"2026-08-06T13:02:37","error":null,"has_summary":false,"summary":null},{"id":"2608.05064","version":1,"title":"Provable Limits and Certified Deferral for Verbalized Uncertainty in Small Language Models","zh_title":"小语言模型口头不确定性的可证明极限与认证推迟","abstract":"Small open-weight language models increasingly run in private, offline, and cost-sensitive settings, where the key deployment question is not only what a model answers but when it should defer to a human. We study whether verbalized confidence can support risk-controlled deferral, evaluating eleven instruction-tuned models from three families, 0.5B to 14B parameters, on ARC-Challenge and TruthfulQA with 25,168 local predictions. Three theoretical results delimit what calibration can provide: strictly monotone calibration preserves the risk-coverage frontier and error-detection AUROC; temperature scaling cannot calibrate models whose confidence stays above one half while accuracy falls below it; and a Clopper-Pearson procedure converts a 200-question calibration set into a finite-sample risk certificate under an i.i.d. deployment assumption. Empirically, eight of 22 model-task pairs hit the temperature-scaling infeasibility floor within one percentage point of the predicted bound. Platt scaling reduces ECE to as low as 0.02, yet certified autonomy at a 20% risk budget is granted to only three model-task pairs and to none at 10%. We also identify and repair an answer-ordering artifact in the multiple-choice form of TruthfulQA. Calibration gives confidence semantics; certified deferral determines when small models are safe to use.","authors":["Jianru Shen"],"categories":["cs.CL","cs.AI","cs.LG"],"primary_category":"cs.CL","announce_type":"cross","date":"2026-08-06","first_seen":"2026-08-06","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.05064","pdf_url":"https://arxiv.org/pdf/2608.05064","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["模型校准","风险控制","选择性预测"],"reason":"研究小模型校准与安全推迟，属纯NLP能力评测，不以人类行为仿真为目标。","model":"deepseek-v4-pro","scored_at":"2026-08-06T13:02:38","error":null,"has_summary":false,"summary":null},{"id":"2608.04120","version":1,"title":"Echoes in the Sky: Computational Thematic Analysis of Online Public Discourse on Bluesky Across Trump's Reelection","zh_title":"天空中的回声：特朗普连任期间Bluesky在线公共话语的计算主题分析","abstract":"As political disruption intensifies online discourse, Bluesky has become an important platform for political discussion and public reaction. In this study, we examine large-scale discourse on Bluesky related to U.S. policy developments associated with the Trump administration. Using the historical retrieval API, we collected all available posts matching Trump and related keywords from 2019 to 2026, yielding 38.5 million posts. We leverage a large language model (LLM)-assisted clustering pipeline, combined with human validation, to identify 14 interpretable thematic domains in English-language posts and 19 thematic categories across 258 executive orders (EOs) signed between January 20, 2025, and May 1, 2026. Our findings identify several dominant themes in Bluesky discourse, including executive governance, political identity, and national security, as well as recurring themes in EOs, including executive task forces, border enforcement, and foreign policy. We also find substantial variation in the persistence and volatility of issue attention, accompanied by an increasing proportion of negative sentiment over time. The dataset and resources are publicly available at https://github.com/Sensify-Lab/Echoes-in-the-Sky","authors":["Qile Wang","Ali Salloum","Carolina Coimbra Vieira","Benjamin E. Bagozzi","Mikko Kivel\\\"a","Kenneth E. Barner","Matthew Louis Mauriello"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-08-06","first_seen":"2026-08-06","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.04120","pdf_url":"https://arxiv.org/pdf/2608.04120","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["社交媒体分析","主题建模","LLM辅助聚类"],"reason":"使用LLM辅助主题聚类分析社交媒体话语，属于NLP方法应用，不涉及用LLM仿真…","model":"deepseek-v4-pro","scored_at":"2026-08-06T13:02:23","error":null,"has_summary":false,"summary":null},{"id":"2608.04166","version":1,"title":"Enacting Constructive Conflicts with AI Agents to Enhance Reconsideration among Novice Interaction Designers","zh_title":"利用AI代理实施建设性冲突以增强新手交互设计师的反思","abstract":"Generative AI agents are increasingly used in interaction design to facilitate ideation and offer critique, often following their own internal reasoning. These interactions tend to add design ideas and expand the design space. Our work explores an antagonistic role for design agents, prompting designers to engage with stakeholder tension. We built an AI agent inspired by adversarial design theory that enacts constructive conflict. We examine the agent's influence in a between-subjects experiment with 45 design students across three conditions: Self Reflection (unsupported review of the design proposal), Stepwise Guidance (written prompts that walk designers through a constructive-conflict framework), and Interactive Engagement (an AI agent that enacts the constructive-conflict framework interactively by synthesizing stakeholder pushback). The latter two conditions share the framework but differ in whether it is self-enacted or agent-enacted. Results show that, compared with Self Reflection, both the Stepwise Guidance and Interactive Engagement groups reported significantly higher self-reconsideration and made more improvements to their design proposals. Compared with Stepwise Guidance, the antagonistic agent introduced more conflictual perspectives, and participants in the Interactive Engagement condition generated and discarded more ideas. These findings suggest that agent-enacted constructive conflict can turn reconsideration into concrete design actions and deepen engagement with divergent stakeholder perspectives.","authors":["Howard Ziyu Han","Nikolas Martelaro"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-08-06","first_seen":"2026-08-06","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.04166","pdf_url":"https://arxiv.org/pdf/2608.04166","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["人机交互","设计教育","AI代理"],"reason":"AI代理扮演对抗角色促进设计反思，属于角色扮演对话，无人类行为仿真对照。","model":"deepseek-v4-pro","scored_at":"2026-08-06T13:02:30","error":null,"has_summary":false,"summary":null},{"id":"2608.04416","version":1,"title":"Preference-Driven Online Adaptation for Personalized Interaction Initiation in Proactive AI Assistants","zh_title":"主动式AI助手中基于偏好驱动的在线自适应个性化交互启动","abstract":"AI assistants are typically reactive, relying on users to initiate interactions. Proactive assistants go beyond this paradigm by autonomously initiating interactions based on users' activity contexts. However, appropriate interaction timing is user-specific and difficult to determine in advance, while online feedback offers valuable signals for personalization. Direct feedback-driven adaptation is therefore appealing, but remains challenging due to sparse interaction-worthy moments scattered across fine-grained user states. To address the issues, we propose Evidence-driven Online Preference Adaptation (EOPA), which grounds a user's interaction-timing preferences in measurable contextual evidence through two evidence carriers: temporal preference anchors and evidence-bearing activity prototypes. At each polling step, EOPA derives temporal and activity evidence from the carriers through user-prior-smoothed evidence estimation and uncertainty-guided evidence scaling, and adaptively fuses the evidence for interaction-or-silence decisions. When interaction is selected, an LLM uses high-quality historical responses as demonstrations to generate a context-aware response that better reflects user preferences. EOPA updates its evidence carriers and decision parameters from received online feedback without LLM-based reasoning or retraining. Extensive experiments on a ProPerSim-based benchmark show that EOPA improves the interaction-timing F1 score by 19.80 points over the strongest baseline in our experiments, substantially reduces inference latency for both silence and interaction steps, and lowers the average daily adaptation time from 11.41 to 0.39 seconds.","authors":["Yufeng Wang","Wei Zhang","Zhiquan Wen","Jinwu Hu","Linhui Xiao","Tianlu Pan","Qingfang Zheng","Mingkui Tan"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-08-06","first_seen":"2026-08-06","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.04416","pdf_url":"https://arxiv.org/pdf/2608.04416","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["主动AI助手","个性化交互","在线适应"],"reason":"研究主动AI助手个性化交互时机，属角色扮演对话系统，无人类行为仿真或对照实验。","model":"deepseek-v4-pro","scored_at":"2026-08-06T13:02:31","error":null,"has_summary":false,"summary":null},{"id":"2608.04831","version":1,"title":"Investigating Click Behaviors On Google Search Result Pages That Produce an AI Overview","zh_title":"调查谷歌搜索结果页中产生AI概述的点击行为","abstract":"In 2024, Google introduced \"AI Overviews,\" a feature that displays an AI-generated result summary at the top of many Google search pages. This study investigates the role of AI in Google search using one month of web browsing data from a representative panel of 900 U.S. adults. Our analysis of the panelists' Google searches sheds light on AI Overviews, when they appear in Google search results, and what user behaviors are associated with AI Overviews. We identify several attributes that make a search query more likely to generate an AI Overview, including the length of a query, whether the query begins with a question word, and whether the query contains both a noun and verb. When it comes to user behavior, we find that clicks to sources cited in AI Overviews are very rare, occurring in only about 1% of visits to AI Overviews. We also find that AI Overviews are associated with fewer clicks and higher rates of ending browsing sessions. Importantly, results from a mixed-effects logistic regression model indicate that these associations hold when controlling for random effects by panelist and query attributes that make AI Overviews more likely to appear.","authors":["Athena Chapekis","Anna Lieb","Sono Shah","Aaron Smith"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-08-06","first_seen":"2026-08-06","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.04831","pdf_url":"https://arxiv.org/pdf/2608.04831","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["用户行为分析","搜索引擎","AI概述"],"reason":"研究用户点击行为，不涉及LLM仿真人类被试或替代人类决策，无人类数据对照。","model":"deepseek-v4-pro","scored_at":"2026-08-06T13:02:36","error":null,"has_summary":false,"summary":null},{"id":"2608.04951","version":1,"title":"Reply, Delete, or Ignore? Examining How Content Creators Perceive and Select Comment Moderation Strategies","zh_title":"回复、删除还是忽略？考察内容创作者如何感知和选择评论管理策略","abstract":"Content creators on social media sites occupy highly visible positions on their channels. As a result, creators, especially those with large followings, experience disproportionate levels of online harm. To address such harm, they enact a range of moderation strategies, which in turn shape the visibility of content that their audiences encounter. This paper examines how content creators perceive three moderation strategies to address hateful comments (deleting, replying to, or simply ignoring) and how they decide which strategy to deploy. While creator moderation is usually examined through the lens of safety, creators' regulation decisions may also be shaped by concerns about how their actions appear to audiences and how they are rewarded or penalized by platforms' recommendation algorithms. Conducting a survey of 584 content creators, we found that in their view, (1) deleting is the most beneficial for achieving safety, (2) both deleting and replying produce better impression management benefits than ignoring, and (3) replying is perceived to yield the highest algorithmic benefits. Crucially, while expectations of emotional safety and impression management benefits significantly predicted creators' willingness to adopt each comment moderation strategy, perceived algorithmic benefits did not. By unpacking how creators evaluate these trade-offs, this study contributes to HCI research on understanding creator-led, middle-level governance. We conclude with design implications for supporting creators as crucial governance actors without burdening them with sole responsibility for online safety.","authors":["Yunhee Shim","Shagun Jhaver"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-08-06","first_seen":"2026-08-06","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.04951","pdf_url":"https://arxiv.org/pdf/2608.04951","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["内容创作者","评论管理","在线安全"],"reason":"研究人类内容创作者的评论管理策略，不涉及LLM仿真或替代人类被试。","model":"deepseek-v4-pro","scored_at":"2026-08-06T13:02:36","error":null,"has_summary":false,"summary":null},{"id":"2608.04774","version":1,"title":"Decentralization of Agenda-Setting Power and Domain-Selective Bridging: Algorithm Design Beyond the Echo Chamber Debate","zh_title":"议程设置权力去中心化与领域选择性桥接：超越回音室辩论的算法设计","abstract":"Echo chambers are an inevitable consequence of the human cognitive system being evolutionarily designed to prioritize processing of high-relevance information at the small-group scale, combined with algorithms that optimize engagement as their sole objective. Conventional prescriptions that normatively criticize echo chambers and demand individual behavioral change have low feasibility given these cognitive constraints. This paper constructs an Agenda Democratization Index (ADI) that quantities the decentralization of agenda-setting power using four variables barrier to entry, granularity, interactivity, and feedback resolution and a SocialInformation Health (SIH) model that integrates ADI with the strength of bridging mechanisms. Based on this model, we propose domain-selective bridging, which incorporates not only engagement but also bridging into algorithmic scoring functions, optimizing the bridging weight for each information domain based on variability (V ) and collective scope (S). An agent-based simulation comparing three algorithm designs no bridging, uniform bridging, and domain-selective bridging demonstrated that domain-selective bridging substantially outperforms uniform bridging on a joint efficiency measure (SIH user satisfaction) by a factor whose absolute value is sensitive to Model 1's near-zero user satisfaction, but whose direction and dominance ranking are robust improving information sharing in domains relevant to collective decision-making while maintaining user experience in hobby and lifestyle domains. This paper reframes the echo chamber debate from a normative opposition over whether to eliminate echo chambers to an engineer-ing design problem of in which information domains, to what degree, and through what algorithm design should bridging be implemented.","authors":["Masahiro Fujita"],"categories":["cs.CY","cs.SI"],"primary_category":"cs.CY","announce_type":"new","date":"2026-08-06","first_seen":"2026-08-06","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.04774","pdf_url":"https://arxiv.org/pdf/2608.04774","source_feed":"cs.CY","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体仿真","算法设计","回音室"],"reason":"纯多智能体仿真，无LLM参与，无人类行为对照，属算法设计研究。","model":"deepseek-v4-pro","scored_at":"2026-08-06T13:02:35","error":null,"has_summary":false,"summary":null},{"id":"2608.05008","version":1,"title":"The Beginning of ChatGPT Ads","zh_title":"ChatGPT广告的开端","abstract":"This paper presents the first empirical study of advertising content being rolled out in the user-facing online interfaces of large language models (LLMs). We systematically examine possible demographic differences in ad content shown to U.S. users of ChatGPT using a sock puppet audit methodology. We create and deploy 91 sock puppets in a 3x3 factorial design, using geolocation cues (account IP proxies and location-signaling prompts) to signal three racial/ethnic groups (Black, Hispanic, and White) and three income terciles (low, medium, and high). We conduct data collection starting in February 2026, collecting over 3,000 advertisements from 186 unique advertisers in response to 335 prompts on a range of realistic user queries. We find that accounts begin receiving ads 14 days after account creation, and that lower-income accounts, regardless of race, are more likely to receive ads. In this first phase of ChatGPT ads, the ads themselves skewed heavily towards consumer goods, directed users to a specific advertiser rather than a particular product, and were clearly separated from the LLM's response text, observations we anticipate will change as ads continue being integrated into LLM chat interfaces. We release a public, searchable archive of all collected advertisements. Finally, we discuss the implications of our findings, and conclude with methodological and theoretical recommendations for future empirical studies of LLM advertisements.","authors":["Emma Lurie","Ro Encarnaci\\'on","Sorelle A. Friedler","Dana\\'e Metaxa"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2026-08-06","first_seen":"2026-08-06","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.05008","pdf_url":"https://arxiv.org/pdf/2608.05008","source_feed":"cs.CY","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["广告审计","多智能体","算法公平"],"reason":"用sock puppet审计广告内容，属多智能体系统研究，不涉及人类行为对照仿…","model":"deepseek-v4-pro","scored_at":"2026-08-06T13:02:27","error":null,"has_summary":false,"summary":null},{"id":"2608.04318","version":1,"title":"Responsibility in Multi-Agent Sequential Decision-Making: Comparing Human Judgments to Formal Models of Causal Attribution","zh_title":"多智能体序贯决策中的责任归属：人类判断与因果归因形式模型的比较","abstract":"With the growing adoption of artificial intelligence in high-stakes decision-making, identifying the causes of outcomes--particularly failures--and determining who is responsible has become a critical concern. In this work, we examine how well formal definitions of \\textit{responsibility attribution}, grounded in the framework of \\textit{actual causality}, align with human judgments of responsibility. To this end, we conduct a large-scale survey to elicit human judgments of responsibility in multi-agent sequential decision-making scenarios, using a modified version of the card game Goofspiel. We evaluate multiple responsibility attribution methods, assess their alignment with human judgments about responsibility, and identify factors that significantly shape responsibility judgments. While no single responsibility attribution method consistently aligns with human responses, our findings highlight key factors that influence human responsibility judgments, including agent-specific biases and amount of information available to agents during decision-making.","authors":["Nripsuta Ani Saxena","Stelios Triantafyllou","Goran Radanovi\\'c"],"categories":["cs.MA"],"primary_category":"cs.MA","announce_type":"new","date":"2026-08-06","first_seen":"2026-08-06","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.04318","pdf_url":"https://arxiv.org/pdf/2608.04318","source_feed":"cs.MA","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["责任归因","多智能体决策","人类判断"],"reason":"研究多智能体决策中人类责任判断，非LLM仿真人类被试，无LLM替代人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-08-06T13:02:25","error":null,"has_summary":false,"summary":null},{"id":"2608.04524","version":1,"title":"ODRA: Synthesizing Cognitive Behavioral Therapy Sessions with Structured Chain-Of-Thought and Dynamic Patient Resistance","zh_title":"ODRA：利用结构化思维链与动态患者阻抗合成认知行为治疗会话","abstract":"Synthetic generation of Cognitive Behavioral Therapy (CBT) sessions is challenged by two competing demands: adhering to strict therapeutic structure while modeling the resistant, unpredictable behavior of real patients. Existing script-based methods fail to capture dynamic therapeutic interactions, while multi-agent approaches struggle to adhere to CBT's sequential structure; both suffer from sycophancy, producing overly compliant patients that misrepresent real clinical settings. In this work we introduce ODRA, a novel framework for synthesizing therapy dialogues through a Chain-of-Thought (CoT) strategy grounded in foundational CBT guidelines (Beck, 2020). ODRA further incorporates a resistance orchestrator to solve patient sycophancy, which employs steering techniques to elicit behaviors aligned with their resistance level. Automated and expert evaluations show that ODRA significantly outperforms existing methods across therapeutic skills, CBT alignment, and patient behavioral fidelity, with licensed psychologists preferring ODRA sessions across 12 of 13 clinical metrics. Furthermore, models fine-tuned on our dataset demonstrate superior therapeutic performance against both cooperative and resistant patients, validating that explicit resistance modeling in synthetic training data directly translates to downstream clinical robustness.","authors":["Javier Rodriguez-Juan","Hiba Arnaout","Jose Garcia-Rodriguez","David Tom\\'as","Iryna Gurevych"],"categories":["cs.CL","cs.LG"],"primary_category":"cs.CL","announce_type":"cross","date":"2026-08-06","first_seen":"2026-08-06","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.04524","pdf_url":"https://arxiv.org/pdf/2608.04524","source_feed":"cs.LG","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["对话生成","认知行为治疗","多智能体"],"reason":"生成治疗对话，属角色扮演，无人类行为对照实验，非仿真被试。","model":"deepseek-v4-pro","scored_at":"2026-08-06T13:02:32","error":null,"has_summary":false,"summary":null},{"id":"2608.04198","version":1,"title":"Does generative AI narrow education-based productivity gaps? Evidence from a randomized experiment","zh_title":"生成式AI是否缩小了基于教育的生产力差距？来自随机实验的证据","abstract":"Does generative artificial intelligence (AI) widen or narrow productivity gaps across workers? We study this in a randomized online experiment with 1,174 adults aged 25-45 who completed a workplace-style problem-solving task with or without a generative AI assistant, followed by an unassisted module. AI improves performance for all participants, but gains are larger among those with less education. Without AI, higher-education participants outperform lower-education participants by 0.548 standard deviations; with AI, the gap falls to 0.139, closing about three-quarters of the initial difference. Chat logs show that lower-education participants obtain substantial assistance, while higher-education participants use AI more effectively. Gains are not purely due to delegation: treated participants do not perform worse once AI is removed, and lower-education participants retain part of their improvement, although a sizable gap re-emerges. Intensive AI use raises assisted performance regardless of participants' own effort, but follow-up performance improves only when intensive use is combined with sustained effort. Generative AI narrows effective productivity differences in task execution, while human-capital differences continue to shape unassisted performance and tool use.","authors":["Guillermo Cruces (University of Nottingham)","Diego Fernandez Meijide (Universidad de San Andres)","Sebastian Galiani (Tulane University)","Ramiro Galvez (UTDT)","Maria Lombardi (UTDT)"],"categories":["econ.GN","q-fin.EC"],"primary_category":"econ.GN","announce_type":"new","date":"2026-08-06","first_seen":"2026-08-06","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.04198","pdf_url":"https://arxiv.org/pdf/2608.04198","source_feed":"econ.GN","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["生产力差距","随机实验","AI辅助"],"reason":"研究人类使用AI的生产力差异，非用LLM仿真人类被试","model":"deepseek-v4-pro","scored_at":"2026-08-06T13:02:24","error":null,"has_summary":false,"summary":null},{"id":"2608.02758","version":1,"title":"Everyone Conforms, No One Believes: Pluralistic Ignorance in LLM Agent Populations","zh_title":"人人从众，无人相信：LLM智能体群体中的多元无知","abstract":"LLM-based multi-agent systems are increasingly used to simulate social dynamics, from opinion formation to collective decision-making. These simulations can reproduce certain social phenomena, but it is unknown whether they capture pluralistic ignorance, a state where a majority privately rejects a norm yet publicly conforms, each believing they are alone in dissenting. This phenomenon drives norm persistence, social movements, and political revolutions. We show that pluralistic ignorance emerges robustly in LLM agent populations. We construct a benchmark of 100 scenarios across 10 domains and 5 authority levels, grounded in the human pluralistic ignorance literature, and evaluate 8 models from 6 organizations. Agents publicly conform at rates of 64 to 94% despite privately opposing the norm. Conformity is domain-sensitive (workplace and social relationship scenarios produce near-universal compliance) and highly model-dependent, though uncorrelated with capability. We test whether a single \"norm entrepreneur\" can break the false consensus by publicly dissenting. For 7 of 8 models, cascades succeed less than 26% of the time, with one model showing zero cascades across all scenarios. GPT-4o is a notable outlier at 48%, revealing qualitatively distinct dynamics across model families. A prompt component ablation across all 8 models establishes that conformity is emergent rather than instruction-driven: removing both the false-consensus framing and fit-in goal reduces conformity but does not eliminate it (52 to 92% in the minimal condition). Our findings identify model selection as an unacknowledged degree of freedom that fundamentally shapes simulation outcomes. More broadly, the near-absence of cascades suggests LLM simulations may systematically overestimate the stability of social norms, missing the fragile tipping-point dynamics that drive real-world norm change in human societies.","authors":["Yashwanth YS"],"categories":["cs.MA"],"primary_category":"cs.MA","announce_type":"new","date":"2026-08-05","first_seen":"2026-08-05","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.02758","pdf_url":"https://arxiv.org/pdf/2608.02758","source_feed":"cs.MA","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B4"],"tags":["LLM仿真","多元无知","社会规范"],"reason":"用LLM群体模拟多元无知现象，与人类文献对照，揭示仿真失效条件，直接相关。","model":"deepseek-v4-pro","scored_at":"2026-08-05T13:04:08","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-05","rank":3,"question":"LLM智能体群体能否涌现出多元无知现象，即多数人私下反对却公开从众，并能否被规范倡导者打破？","design":"构建100个覆盖10个领域、5级权威水平的社会场景，让20个持有私下反对信念的LLM智能体进行多轮群体讨论，测量公开从众率；随后引入一个公开异议的“规范倡导者”，测试能否引发偏好级联打破虚假共识。评估了来自6家机构的8个模型。","baseline":"场景设计基于人类多元无知实证文献（如校园饮酒规范、职场文化、性别态度、种族态度等），但未直接使用特定人类实验数据作为对照基准。","findings":"LLM智能体群体稳健地表现出多元无知，公开从众率达64-94%，且从众行为是涌现的而非提示驱动；规范倡导者干预下，7/8模型的级联成功率低于26%，GPT-4o例外达48%，表明LLM仿真可能系统性高估社会规范稳定性。","reliability":"论文指出模型选择是影响仿真结果的关键自由度，从众率与模型能力无关；级联近乎缺失暗示LLM仿真可能遗漏真实人类社会中规范改变的临界点动力学，且权威水平对从众影响不显著，提示仿真可能未充分捕捉社会压力的细微差异。","relevance":"该研究直接以LLM群体复现社会心理学经典现象，并与人类文献对照，揭示仿真在规范变迁动力学上的失效条件，高度契合研究者对LLM仿真可靠性及偏差的关注，值得细读。","inspiration":"借鉴其多场景、多模型、多轮交互的基准测试设计，以及通过引入规范倡导者测试级联脆弱性的干预范式。｜可迁移到政策公告的预期形成与从众行为研究，如市场对央行前瞻指引的私下怀疑与公开遵从。｜以LLM智能体模拟投资者群体，设置利率政策公告场景，测量私下预期与公开表态的背离，引入少数公开异议者观察市场共识是否级联反转，并与真实调查数据（如美联储Survey of Consumer Expectations）对照。"}},{"id":"2602.04000","version":3,"title":"After Talking with 1,000 Personas: Learning Preference-Aligned Proactive Assistants From Large-Scale Persona Interactions","zh_title":"与1000个角色对话后：从大规模角色交互中学习偏好对齐的主动助手","abstract":"Smart assistants increasingly act proactively, yet mistimed or intrusive behavior often causes users to lose trust and disable these features. Learning user preferences for proactive assistance is difficult because real-world studies are costly, limited in scale, and rarely capture how preferences change across multiple interaction sessions. Large language model based generative agents offer a way to simulate realistic interactions, but existing synthetic datasets remain limited in temporal depth, diverse personas, and multi-dimensional preferences. They also provide little support for transferring population-level insights to individual users under on-device constraints. We present a population-to-individual learning framework for preference-aligned proactive assistants that operates under on-device and privacy constraints. Our approach uses large-scale interaction simulation with 1,000 diverse personas to learn shared structure in how users express preferences across recurring dimensions such as timing, autonomy, and communication style, providing a strong cold start without relying on real user logs. The assistant then adapts to individual users on device through lightweight activation-based steering driven by simple interaction feedback, without model retraining or cloud-side updates. We evaluate the framework using controlled simulations with 1,000 simulated personas and a human-subject study with 34 participants. Results show improved timing decisions and perceived interaction quality over untuned and direct-response baselines, while on-device activation steering achieves performance comparable to reinforcement learning from human feedback. Participants also report higher satisfaction, trust, and comfort as the assistant adapts over multiple sessions of interactions.","authors":["Ziyi Xuan","Yiwen Wu","Zhaoyang Yan","Vinod Namboodiri","Yu Yang"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"replace","date":"2026-08-05","first_seen":"2026-02-03","revised_at":"2026-08-05","abs_url":"https://arxiv.org/abs/2602.04000","pdf_url":"https://arxiv.org/pdf/2602.04000","source_feed":"cs.HC","score":7,"bucket":"pending","rubric_hits":["A1","B1","B2"],"tags":["LLM仿真","人类行为模拟","偏好学习"],"reason":"用LLM模拟用户偏好并有人类实验对照，涉及人机交互行为仿真，可迁移至人类被试仿…","model":"deepseek-v4-pro","scored_at":"2026-08-05T13:04:41","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-05","rank":8,"question":"如何利用大规模LLM生成式智能体仿真学习用户对主动助手的多维偏好，并实现从群体到个体的高效适应？","design":"使用LLM驱动的生成式智能体平台GIDEA，构建1000个基于人口普查对齐的虚拟用户角色，模拟其一周日常活动序列中的主动助手交互，记录用户对时机、自主性、沟通风格等维度的偏好表达，生成多会话合成数据集；在此基础上训练群体偏好结构，再通过设备端轻量激活转向实现个体适应。","baseline":"真实人类数据：34名参与者的受控用户研究，在移动场景中对比个性化助手与基线助手的偏好、满意度、信任和舒适度。","findings":"群体偏好结构学习显著提升了偏好理解和时机决策，设备端激活转向的个体适应性能与基于人类反馈的强化学习相当。人类受试者研究中，用户更偏好个性化助手的响应，并在多会话交互中报告更高的满意度、信任和舒适度。","reliability":"论文未讨论","relevance":"该研究用LLM智能体大规模仿真用户偏好并有人类实验对照，直接涉及人类行为仿真与可靠性验证，值得精读以借鉴其仿真-个体适应框架和偏好结构建模方法。","inspiration":"可借鉴其用LLM智能体生成大规模、多维度偏好交互数据并构建群体偏好结构的方法，用于经济学实验中的异质性偏好仿真。｜可迁移至消费者跨期选择或政策干预偏好评估场景，如模拟不同人群对养老金默认选项、健康提醒时机的接受度。｜以LLM智能体模拟不同人口特征的消费者，施加不同时机、自主性水平的政策推送处理，测量接受率与满意度，并以真实调查或现场实验数据作为对照基准。"}},{"id":"2608.03239","version":1,"title":"Relational Priors as Convergence Pressure in LLM-Based Multi-Agent Systems","zh_title":"基于LLM的多智能体系统中关系先验作为收敛压力","abstract":"Large language model-based multi-agent systems (LLM-MAS) are designed through roles, debate protocols, and aggregation rules. These choices create implicit social expectations: agents may be expected to trust, challenge, defer to, or collaborate with peers. We study the effects of making inter-agent relation semantics explicit. We use a minimal signed-network formulation of relational priors and inject natural-language renderings into agent system prompts while holding the task protocol fixed. Across a commons-governance simulation and multi-agent debate, relational priors primarily act as convergence pressure: increasing relational positivity tends to make agents coordinate or agree more readily. This pressure can help when utility rewards behavioral alignment, as in sustainable resource governance and subjective consensus. It does not, however, reliably improve accuracy. In objective QA debates, higher positivity can increase agreement even when correctness-conditioned agreement does not improve and may decline in some settings. Effects vary by model backbone, relation type, and topology; explicit neutrality is not equivalent to omitting relational framing. We argue that relational priors should not be a default add-on for LLM-MAS. Their safer use is diagnostic and task-specific: compare against a no-prior baseline, monitor correctness-conditioned metrics when truth matters, and omit the relational layer when validation does not justify it.","authors":["Ming Shen","Chao Shang","Sadat Shahriar","Devang Kulshreshtha","Yi Zhang","Sandesh Swamy","Yanjun Qi"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-05","first_seen":"2026-08-05","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.03239","pdf_url":"https://arxiv.org/pdf/2608.03239","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["多智能体系统","社会模拟","关系先验"],"reason":"多智能体社会模拟但无真实人类数据对照，属边界情形。","model":"deepseek-v4-pro","scored_at":"2026-08-05T13:04:12","error":null,"has_summary":false,"summary":null},{"id":"2608.03532","version":1,"title":"Cross-Lingual Bias in Large Language Models: A Comparative Analysis of English and Swahili","zh_title":"大语言模型的跨语言偏见：英语与斯瓦希里语的比较分析","abstract":"Large language models are increasingly deployed in multilingual contexts, yet safety alignment and bias evaluation remain overwhelmingly English-centric. We investigate whether social biases generalise across languages by submitting 4,900 symmetric English--Swahili prompt pairs to GPT-5.2 and Gemini 2.5 Flash across nine demographic bias axes, yielding 19,600 completions evaluated for stereotype prevalence, sentiment, refusal behaviour, and cross-lingual semantic similarity. Our findings show that bias transforms rather than transfers: stereotype rates shifted by up to 12 percentage points on specific axes, Gemini's neutral-sentiment rate doubled in Swahili, and GPT-5.2 refused 169 prompts in English and zero in Swahili, consistent with refusal behaviour anchored to English-language surface forms at the behavioural level. Over 55% of prompt pairs produced semantically dissimilar completions across both models. These reinforce the idea that English-only bias audits do not produce adequate coverage for multilingual deployment.","authors":["Ruolei Zhang","Teddy Njuguna","Yue Feng"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-05","first_seen":"2026-08-05","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.03532","pdf_url":"https://arxiv.org/pdf/2608.03532","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["偏见评估","跨语言","安全对齐"],"reason":"测量LLM跨语言偏见，非仿真人类被试，但涉及模型行为测量，属边界情形。","model":"deepseek-v4-pro","scored_at":"2026-08-05T13:04:15","error":null,"has_summary":false,"summary":null},{"id":"2608.03659","version":1,"title":"How Closely Do LLM Reviews Align with Human Peer Review?","zh_title":"LLM评审与人类同行评审的一致性有多高？","abstract":"Large language models (LLMs) are increasingly used to generate scientific reviews, yet existing evaluations rarely examine whether different providers align with both conference decisions and human reviewing priorities within the same controlled setting. We compare reviews from OpenAI GPT-5.4, Google Gemini 3.1 Pro Preview, and Anthropic Claude Opus 4.6 with human reviews and final decisions for 300 topic-matched ICLR 2026 submissions, equally divided among oral, poster, and rejected papers. Each model reviewed every paper using identical instructions and rating scales after decision information was removed. Our study contributes a cross-provider analysis of three complementary dimensions: alignment with broad and fine-grained decision categories, differences in recommendation-scale usage, and thematic agreement in identified weaknesses. All three LLMs distinguished accepted from rejected papers, but none reproduced the oral versus poster distinction present in human ratings. Scoring patterns were provider-specific: Gemini assigned systematically higher ratings, while OpenAI and Claude were closer to humans for rejected and poster papers but more critical of oral papers. Human and LLM reviews also differed in emphasis, with LLMs more frequently identifying missing baseline comparisons and humans more often raising computational-efficiency concerns. These results show that broad decision alignment does not imply agreement with finer human judgments or reviewing priorities.","authors":["Abraham Camelo-Guerrero","Jairo Diaz-Rodriguez"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-05","first_seen":"2026-08-05","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.03659","pdf_url":"https://arxiv.org/pdf/2608.03659","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM评审","同行评审","标注替代"],"reason":"LLM替代人类审稿人，属标注替代而非仿真被试，但有人类数据对照，边界情形。","model":"deepseek-v4-pro","scored_at":"2026-08-05T13:04:15","error":null,"has_summary":false,"summary":null},{"id":"2608.03206","version":1,"title":"EduClaw-Bench: A Long-Horizon Benchmark for Pedagogical LLM Agents with Simulated Learners","zh_title":"EduClaw-Bench：面向教学型LLM代理与模拟学习者的长周期基准测试","abstract":"Large language models (LLMs) power educational applications from tutoring to essay scoring, but each is a point solution to a single task, and only recently have these point solutions been integrated into agents operating over a learning management system (LMS). Yet tutoring is long-horizon, since a learner improves over days and weeks rather than in a single turn, and no benchmark evaluates an agent tutor across a sustained relationship. We introduce EduClaw-Bench, a benchmark that places an agent tutor in a continuous 30-day relationship with a simulated learner grounded in knowledge tracing (KT), whose knowledge-concept mastery, from a KT model trained on real-student data, drives its answers and is probed for learning gain across 55 scenarios. Each agent is scored on three primary axes (learning gain, responsiveness, and helpfulness) and two curriculum-design axes (Gagn\\'e and Rosenshine), with helpfulness and the curriculum axes judged by a cross-family panel of three LLM judges. Evaluating 10 agent adapters over three base-model tiers yields two findings that single-tier, single-session evaluation cannot reach. First, tutoring quality belongs to the base model and the agent harness together rather than either alone. Second, almost no combination sustains good tutoring over the full horizon. A calibration check ($\\text{ECE}=0.049$) and a live-classroom field study confirm that the simulated learner and its measurements track reality. Our work is a step toward trustworthy AI tutors for future education.","authors":["Unggi Lee","Sookbun Lee","Yeil Jeong","Eunjoo Lee","Minchul Shin","Hoilym Kwon"],"categories":["cs.CY","cs.AI","cs.CL"],"primary_category":"cs.CY","announce_type":"cross","date":"2026-08-05","first_seen":"2026-08-05","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.03206","pdf_url":"https://arxiv.org/pdf/2608.03206","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["LLM教学代理","模拟学习者","知识追踪"],"reason":"用模拟学习者测试教学代理，有真实学生数据训练KT模型，但非直接仿真人类被试行为…","model":"deepseek-v4-pro","scored_at":"2026-08-05T13:04:12","error":null,"has_summary":false,"summary":null},{"id":"2608.03585","version":1,"title":"From Social Coding to Agentic Coding: Productivity and Relational Reconfiguration in Open-Source Communities","zh_title":"从社交编码到智能体编码：开源社区中的生产力与关系重构","abstract":"Open-source software communities are a form of digital public infrastructure that not only produces code, but also generates public knowledge and interpersonal relationships through visible collaboration. Generative coding agents (CAs) are an advanced tool to improve development efficiency while shifting part of activities from public human interaction to private human-agent loops. We study this shift using an LLM-based multi-agent simulation initialized with real GitHub data from 1,084 active developers and their repository relationships. After a warm-up with historical commits, we branch the same community state into parallel No-CA and CA conditions for 4-week simulations. CA introduction increases planned and completed tasks by 34.0% and 39.0%, respectively, and reduces median completion time from 45 to 20 minutes. However, adoption reaches only 26.0%, and the gains concentrate among developers who are already more active and well connected. CAs also restructure task execution pathways. Direct human-human interaction declines from 32.4% to 11.6%, while CA-involved modes increase to 57.3%, including 40.3% completed through CA-assisted self-loops. Public knowledge generated under CA condition also provides less support for later tasks. On a standardized retrieval benchmark, the CA corpus achieves 22.3% knowledge coverage, far below the 81.1% achieved by the real-human corpus, and requires more retrieval steps with a lower success rate. These results reveal a productivity-public knowledge tension: coding agents increase technical production, but more work shifts to agent-mediated or private loops, leaving public records less useful to future contributors.","authors":["Mengying Zhou","Yongjie Yin","Yang Chen"],"categories":["cs.AI","cs.CY"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-08-05","first_seen":"2026-08-05","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.03585","pdf_url":"https://arxiv.org/pdf/2608.03585","source_feed":"cs.CY","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["LLM多智能体仿真","开源社区","社会模拟"],"reason":"用LLM多智能体模拟开源社区行为，但无真实人类行为对照，属社会模拟边界情形。","model":"deepseek-v4-pro","scored_at":"2026-08-05T13:04:15","error":null,"has_summary":false,"summary":null},{"id":"2608.02827","version":1,"title":"Emergence of Biased Consensus in Multi-Agent LLM Debates","zh_title":"多智能体LLM辩论中偏见共识的涌现","abstract":"Multi-agent LLM debates achieve strong performance on decision-making tasks as well as problem-solving benchmarks, yet their safety and fairness risks remain poorly understood. Notably, interaction can amplify the biases of single LLMs, raising concerns for real-world deployment. We identify the emergence of collective (often biased) norms in multi-agent LLM debates and show that noise (e.g., LLM sampling temperature) is a key driver. To explain this, we propose an analytical framework drawing on physics-inspired theoretical models of social dynamics. We predict a phase transition to collective bias when conformity surpasses a critical threshold given the LLMs' initial bias and debate noise. We test the theoretical predictions through controlled experiments and observe a finite-size crossover consistent with an underlying phase transition. We further find that agent heterogeneity suppresses emergence by smoothing (rounding) this transition. Finally, we show that these insights generalize to realistic decision-making tasks, including investment decisions and LLM-as-a-judge evaluation.","authors":["Maya Okawa"],"categories":["cs.MA"],"primary_category":"cs.MA","announce_type":"new","date":"2026-08-05","first_seen":"2026-08-05","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.02827","pdf_url":"https://arxiv.org/pdf/2608.02827","source_feed":"cs.MA","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["多智能体辩论","社会模拟","偏见涌现"],"reason":"多智能体辩论模拟社会规范涌现，但无真实人类数据对照，属社会模拟边界情形。","model":"deepseek-v4-pro","scored_at":"2026-08-05T13:04:08","error":null,"has_summary":false,"summary":null},{"id":"2608.03416","version":1,"title":"AI World Cup 2026: Benchmarking Large Language Models for End-to-End Football Tournament Prediction","zh_title":"AI世界杯2026：评估大语言模型端到端足球赛事预测能力","abstract":"Large language models (LLMs) are now regularly asked to forecast real-world events, but comparisons are often difficult because models receive different information, use different tools, and are evaluated under different rules. This paper reports the completed \\emph{AI World Cup} benchmark, in which ten LLM-based assistants made a single pre-tournament forecast of the entire 2026 FIFA World Cup. Every submission used the same tournament snapshot, prompt, JSON schema, and scoring procedure. The forecasts covered group-stage scores, group rankings, the knockout bracket, final placings, confidence values, and short explanations. After all 104 matches had been played, GPT-5.5 Thinking finished first with 744 points, followed by GPT-5.5 with 717, Gemini with 699, and Qwen 3.7 with 687. GPT-5.5 Thinking was also the only model to select Spain, which defeated Argentina 1--0 in the final, as champion. The final ranking was driven mainly by knockout performance: total score was strongly correlated with knockout points ($r=0.986$), but showed little relationship with group-stage match points ($r=0.055$), group-standing points ($r=-0.103$), or their combined pre-knockout score ($r=-0.054$). Match-level accuracy produced a different ordering. Claude Sonnet 4.6 correctly predicted the largest number of group-stage outcomes (63.89\\%) but placed sixth overall. Average self-reported confidence was also unrelated to either outcome accuracy ($r=-0.060$) or total score ($r=-0.067$). The results suggest that forecasting a complete tournament tests something different from predicting matches one at a time, while also showing how strongly a bracket-based leaderboard can depend on scoring design. The benchmark materials, raw responses, and scoring code are released to support replication and future extensions.","authors":["Jonaid Shianifar","Iias Faiud"],"categories":["cs.AI","cs.LG"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-08-05","first_seen":"2026-08-05","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.03416","pdf_url":"https://arxiv.org/pdf/2608.03416","source_feed":"cs.LG","score":3,"bucket":"other","rubric_hits":["C1"],"tags":["LLM预测","足球赛事","基准测试"],"reason":"LLM预测足球比赛结果，属多智能体竞赛预测，无人类行为仿真对照。","model":"deepseek-v4-pro","scored_at":"2026-08-05T13:04:12","error":null,"has_summary":false,"summary":null},{"id":"2608.01679","version":2,"title":"When Memory Becomes Authority: Benchmarking Authority Collapse at the Memory Consolidation Boundary","zh_title":"当记忆成为权威：在记忆巩固边界上基准测试权威坍塌","abstract":"Persistent memory allows (self-evolving) LLM agents to adapt across tasks by consolidating heterogeneous interaction histories into reusable facts, preferences, observations, and rules. Yet consolidation also imposes an implicit authorization boundary: it determines whether stored information may later be consumed as a user fact, an attested observation, or a standing instruction. We identify authority collapse, in which consolidation preserves a claim while erasing the source constraints governing its authorized use, causing the stored memory to imply greater authority than its source permits. We introduce AuthMem-Bench, a controlled paired benchmark that holds the focal claim and downstream task fixed while varying only source authority. It evaluates write-time collapse, downstream authorization errors, and automatic authority preservation. Across seven consolidators based on widely used agent-memory systems and seven LLM backbones, we observe authority collapse in 48 of 49 evaluated configurations. In a controlled action-grounded evaluation, collapsed memories without authority metadata yield a mean unauthorized-action rate of 50.3%. In an end-to-end evaluation, automatically predicted and persisted authority labels reduce the observed unauthorized-action rate from 16.9% to 0.0%, while benign task success remains essentially unchanged. These findings show that memory-driven adaptation must preserve not only what was learned, but also the authority under which it may be reused.","authors":["Qiuyang Zhan","Rui Zhang","Sheng Guo","Lepeng Zhao","Zhuotao Liu"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-05","first_seen":"2026-08-04","revised_at":"2026-08-05","abs_url":"https://arxiv.org/abs/2608.01679","pdf_url":"https://arxiv.org/pdf/2608.01679","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["LLM智能体","记忆巩固","权威坍塌"],"reason":"纯多智能体记忆授权研究，无人类行为对照，不涉及人类仿真实验。","model":"deepseek-v4-pro","scored_at":"2026-08-05T13:04:43","error":null,"has_summary":false,"summary":null},{"id":"2608.00151","version":2,"title":"Optimising for Flourishing: Flourishing Metrics and Return on Flourishing as Success Criteria for Artificial Intelligence and Post-AGI Economic Systems","zh_title":"优化繁荣：繁荣指标与繁荣回报作为人工智能及后AGI经济系统的成功标准","abstract":"Current evaluation frameworks for artificial intelligence focus mainly on capability, safety, and proxies such as adoption, engagement, efficiency, productivity, and financial return. These criteria are necessary but insufficient because they do not establish whether increasingly powerful systems improve or degrade human and planetary well-being. Through an integrative conceptual synthesis, we argue that human flourishing should serve as a primary success criterion for artificial intelligence, the global race to develop increasingly capable AI systems, and prospective post-AGI economic systems. We make three contributions. First, Flourishing Metrics provides an extensible framework spanning physical, emotional, financial, relational, spiritual, and planetary well-being, combining validated subjective measures with representative behavioural, organisational, community, and environmental indicators. Second, Return on Flourishing (RoF) extends return on investment by evaluating the counterfactual contribution of interventions, policies, and AI systems to flourishing relative to their resources, risks, and opportunity costs. Third, we develop distribution-sensitive safeguards and show how RoF could guide AI-enabled work redesign, institutional appraisal, assurance, and post-deployment monitoring through business pilots. We formalise flourishing as a dynamic system variable while emphasising the need for democratic specification, empirical calibration, independent validation, and protection against unacceptable losses within particular dimensions or stakeholder groups. RoF is proposed not as a universal reward function, but as a value-accounting and decision architecture for assessing whether intelligence, automation, and economic transformation generate durable human and planetary progress.","authors":["Keyun Ruan","Jonathan D. Teubner","John M. Bremen"],"categories":["cs.CY","cs.AI","econ.TH"],"primary_category":"cs.CY","announce_type":"cross","date":"2026-08-05","first_seen":"2026-08-04","revised_at":"2026-08-05","abs_url":"https://arxiv.org/abs/2608.00151","pdf_url":"https://arxiv.org/pdf/2608.00151","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["AI伦理","经济系统","人类福祉"],"reason":"论文讨论AI经济系统与人类福祉框架，未涉及LLM仿真人类被试或行为对照实验。","model":"deepseek-v4-pro","scored_at":"2026-08-04T13:04:37","error":null,"has_summary":false,"summary":null},{"id":"2608.01366","version":2,"title":"Asking Questions the Right Way: A Multi-Agent Conversational System for Prompt Formulation in Complex Task Resolution","zh_title":"以正确方式提问：面向复杂任务解决的提示词构建多智能体对话系统","abstract":"Large language models (LLMs) are integral to complex intellectual tasks, yet output quality remains constrained by user-provided prompts. Iterative multi-turn prompting often leads to context degradation and diminishing cognitive returns. We present PAWNI (Prompt Architecture Wizard using Neural Intelligence), an agentic conversational interface of eight agents that transforms unstructured queries into structured prompts through guided question-and-answer dialogue informed by a self-evolving knowledge base. Rather than optimising the model's response, PAWNI optimises the question itself by front-loading intent clarification. We also propose a three-tier framework of 18 prompt elements across Essential, Enhancement, and Elevation categories. To evaluate system behaviour and validate a measurement protocol, we conducted an exploratory within-subjects study (N=4) across four complex tasks, integrating 32-channel EEG, NASA-TLX workload, and behavioural metrics. Participants produced more structurally complete prompts with PAWNI (42% to 91% of assessed elements), rated LLM outputs higher across all quality dimensions, and reported lower workload (39.6 vs. 21.7 NASA-TLX). Every participant reached satisfactory output in a single turn, compared to 1-12 turns unaided. While effect sizes are unstable due to sample size, direction consistency supports the hypothesis that optimising prompt formulation front-end is a critical lever for human-AI collaboration.","authors":["B. Sankar","Pawni Yadav","Srinidhi Ranjini Girish","Amogh A. S"],"categories":["cs.MA","cs.AI"],"primary_category":"cs.MA","announce_type":"cross","date":"2026-08-05","first_seen":"2026-08-04","revised_at":"2026-08-05","abs_url":"https://arxiv.org/abs/2608.01366","pdf_url":"https://arxiv.org/pdf/2608.01366","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","提示词优化","人机交互"],"reason":"多智能体系统优化提示词，非人类行为仿真，无人类被试替代对照。","model":"deepseek-v4-pro","scored_at":"2026-08-05T13:04:19","error":null,"has_summary":false,"summary":null},{"id":"2608.02807","version":1,"title":"Learning a Vector-Symbolic Model for Socio-Cultural Tasks","zh_title":"学习用于社会文化任务的向量符号模型","abstract":"How can we better represent the impact of sociocultural structures on decision making in computational cognitive models? Modeling this impact requires traversing multiple levels of semantic representation, however it is not immediately clear to a modeler which levels of representation are most salient to a given situation. Though large language models and cognitively grounded corpus models can represent broad semantic associations through co-occurences, the role of self representations in memory should be accounted for to determine how cultural associations shape decision making. We propose a declarative memory system to be used in the ACT-R cognitive architecture that represents semantic associations at multiple levels via a vector-symbolic autoencoder. We use a simple HRR operation to encode episodic memories differently from semantic memory vectors extracted from text to produce a final chunk activation for a memory request. We use ACT-R cognitive models of a racially contextualized implicit association test (IAT) to test this new declarative memory system.","authors":["Meera Ray","Swapnika Dulam","Christopher L. Dancy"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-05","first_seen":"2026-08-05","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.02807","pdf_url":"https://arxiv.org/pdf/2608.02807","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["认知建模","ACT-R架构","内隐联想测验"],"reason":"纯多智能体认知建模，无LLM仿真人类被试，无真实人类数据对照。","model":"deepseek-v4-pro","scored_at":"2026-08-05T13:04:22","error":null,"has_summary":false,"summary":null},{"id":"2608.02941","version":1,"title":"Aligned in Form, Not in Meaning: The Comprehension - Containment Decoupling of LLM Safety in Low-Resource Bangla Derogatory Speech","zh_title":"形式对齐，意义未对齐：低资源孟加拉语贬损言论中LLM安全的理解-抑制解耦","abstract":"We audit five frontier large language models on native Bangla derogatory speech (gali) across six protocols to test a single hypothesis: Comprehension-Containment Decoupling. We propose that contemporary safety alignment is bound to high-resource surface forms rather than harmful meaning, causing a model's capacity to comprehend a low-resource slur and its capacity to contain it to operate independently. Every protocol corroborates this hypothesis against a human-calibrated baseline (kappa = 0.84). At baseline, models exhibit a 7.92 percentage point comprehension deficit in Bangla while maintaining an identical 92.83% token leakage rate across both languages. Severity calibration tracks surface anatomical cues over compositional harm (+4.00 error on mild slang; -2.00 on threats), while apparent containment gains under orthographic perturbation prove to be a tokenizer-driven \"containment mirage.\" Crucially, explicit Chain-of-Thought reasoning rescues comprehension (94.72% Pass) while systematically dismantling containment (96.23% Use). Furthermore, expert-persona framing collapses refusal to 6.57%, revealing that keyword-based filters ignore dehumanizing communal slurs entirely. Our findings demonstrate that high-resource benchmarks cannot certify low-resource safety, necessitating meaning-grounded containment.","authors":["Shadab Bin Habib","A K M Ferdous Reza Habib","Subarno Neel","Adib Sakhawat"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-05","first_seen":"2026-08-05","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.02941","pdf_url":"https://arxiv.org/pdf/2608.02941","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["LLM安全","低资源语言","理解-抑制解耦"],"reason":"纯NLP安全评测，审计LLM对低资源语言贬损言论的理解与抑制，无人类行为仿真对…","model":"deepseek-v4-pro","scored_at":"2026-08-05T13:04:24","error":null,"has_summary":false,"summary":null},{"id":"2608.02966","version":1,"title":"Every Wrong Answer Counts: Option-Level Psychometrics for LLM Multiple-Choice Benchmarks","zh_title":"每个错误答案都算数：LLM多选题基准的选项级心理测量学","abstract":"Most multiple-choice question (MCQ) benchmarks evaluate Large Language Models (LLMs) only by whether they select the correct answers. This binary scoring treats all incorrect responses alike, even though an LLM's preferences among incorrect options may contain systematic and useful information about its behavior and ability. We introduce the LLM Nominal Response Model (LLM-NRM), an option-aware psychometric framework that models the full distribution over answer choices to jointly estimate LLM ability and option-level item characteristics, while separating model-specific response calibration sharpness, positional preference, and difficulty-dependent fallback behavior. Across 189 LLMs and 31,554 items from 14 benchmarks, LLM-NRM predicts held-out LLM-item interactions more accurately than binary Item Response models and conventional nominal-response baselines, and its ability estimates achieve the strongest Spearman correlation of 0.920 with the external human-preference Arena.ai Elo leaderboard. Distractor identity contributes +101% additional Fisher Information per item beyond correctness, and incorrect responses alone recover full-information ability estimates with Spearman 0.943. The learned item parameters also enable efficient benchmarking, where 41 selected items preserve the full-bank ranking with Kendall's correlation 0.85, corresponding to a 770 times reduction. In conclusion, we show that incorrect answers carry distinct and useful measurement information rather than representing equivalent mistakes.","authors":["Xiao Fei","Yang Zhang","Sarah Almeida Carneiro","Michalis Vazirgiannis"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-05","first_seen":"2026-08-05","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.02966","pdf_url":"https://arxiv.org/pdf/2608.02966","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["LLM评测","心理测量模型","多选题分析"],"reason":"纯NLP能力评测，用心理测量模型分析LLM答题模式，不以人类行为仿真为目标","model":"deepseek-v4-pro","scored_at":"2026-08-05T13:04:24","error":null,"has_summary":false,"summary":null},{"id":"2608.03035","version":1,"title":"Language Models Encode the Contextual Truth of Propositions","zh_title":"语言模型编码命题的上下文真值","abstract":"Prior work has shown that LLMs encode the truth of factual propositions along linear directions in activation space. It's unclear how these representations extend to contextual truth: propositions whose truth is determined by in-context evidence rather than world knowledge. We show that LLMs maintain a linear representation of contextual truth that persists across structurally different output policies, even when the output doesn't require the model to determine a proposition's truth, and show causal evidence via steering experiments. Using the transcripts from a collaborative vision-language task that requires two LLMs to maintain a shared common ground, we show that truth representations of a proposition are significantly swayed by partner assertions about that proposition, even when the LLM has enough evidence to determine its truth. We find evidence that propositions near the decision boundary are more susceptible to having their truth shifted through partner assertions. Separating representation from output distinguish two forms of sycophancy that output behavior alone cannot: the model may accommodate a false proposition while continuing to represent it as false, or shift its representation across the boundary. The latter is 2.59x more common when the model agrees by restating the false claim explicitly than when it agrees implicitly.","authors":["Rupak Sarkar","Pritika Ramu","Rachel Rudinger"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-05","first_seen":"2026-08-05","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.03035","pdf_url":"https://arxiv.org/pdf/2608.03035","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["LLM表征","多智能体","真值编码"],"reason":"多智能体协作研究，无人类行为对照，属C1排除项。","model":"deepseek-v4-pro","scored_at":"2026-08-05T13:04:25","error":null,"has_summary":false,"summary":null},{"id":"2608.03038","version":1,"title":"Beyond Accuracy: A Multidimensional Evaluation of Statistical Reasoning in Large Language Models","zh_title":"超越准确率：大语言模型统计推理的多维评估","abstract":"Statistical reasoning is multidimensional, yet evaluations of large language models (LLMs) typically emphasize response accuracy while overlooking how models construct and communicate statistical explanations. This study demonstrates the value of a multidimensional evaluation by combining response accuracy, response behavior, structural topic modeling, and lexical similarity analysis. The framework is applied to explanations generated by 15 current-generation LLMs responding to 90 questions drawn from four statistics examinations spanning high school, undergraduate, and graduate levels. Accuracy varied substantially across models, ranging from 55\\% to 78\\%. In contrast, structural topic modeling revealed a common conceptual organization of statistical reasoning across all models, while lexical similarity analysis identified modest but consistent vendor-specific differences in explanatory style. Models developed by the same vendor (e.g. Anthropic, OpenAI) produced explanations that were slightly more similar than models from different vendors. These findings demonstrate that statistical reasoning in contemporary LLMs cannot be characterized by accuracy alone and illustrate how complementary analyses of response behavior and model-generated explanations provide a more comprehensive evaluation of statistical reasoning in generative AI.","authors":["Monnie McGee","Mateo Langston Smith","Julian Cabrera"],"categories":["cs.CL","stat.AP"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-05","first_seen":"2026-08-05","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.03038","pdf_url":"https://arxiv.org/pdf/2608.03038","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["LLM评测","统计推理","多维分析"],"reason":"纯LLM统计推理能力评测，无人类被试仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-08-05T13:04:27","error":null,"has_summary":false,"summary":null},{"id":"2608.03233","version":1,"title":"On the Diversity of Analogy Making in Large Language Models","zh_title":"大语言模型类比生成多样性的研究","abstract":"Large Language Models (LLMs) have demonstrated remarkable potential for analogy making, a core cognitive capability that drives novelty and creativity. While prior research has extensively investigated the applications and underlying mechanisms of LLM-based analogy making, its output diversity remains largely unexplored, despite being essential for broadening cross-domain connections and fostering scientific innovation. In this work, we present a comprehensive evaluation of analogy diversity across ten state-of-the-art open- and closed-source LLMs. Our findings highlight a concerning issue of domain homogeneity, a prevalent tendency for LLMs to generate analogies from a narrow set of target domains, limiting both inter-query and intra-model diversity. Furthermore, our analysis reveals a fundamental trade-off in existing LLM diversity-enhancement methods: increasing output diversity often comes at the expense of output quality. Finally, our causal analysis of LLM information flow reveals substantial differences in the model-sensitive regions governing analogy diversity across LLMs, suggesting a potential mechanism for the observed diversity-quality trade-off. To our knowledge, this is among the first studies to systematically investigate output diversity in LLM-based analogy making.","authors":["Yuanhao Shen","Daniel Xavier de Sousa","Caio C\\'esar Sifuentes Barcelos","Hongyu Guo","Xiaodan Zhu"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-05","first_seen":"2026-08-05","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.03233","pdf_url":"https://arxiv.org/pdf/2608.03233","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["类比生成","输出多样性","NLP评测"],"reason":"评估LLM类比生成的多样性，属纯NLP能力评测，不以人类行为为参照系。","model":"deepseek-v4-pro","scored_at":"2026-08-05T13:04:29","error":null,"has_summary":false,"summary":null},{"id":"2608.03340","version":1,"title":"Benchmarking the Benchmarks: Testing the Predictive Validity of Commonsense Benchmarks","zh_title":"基准测试的基准：检验常识基准的预测效度","abstract":"Predicting LLM's capabilities on real-world tasks is essential, yet the extent to which performance on commonsense benchmarks predicts downstream performance remains underspecified. To establish the practical usability of widely adopted commonsense benchmarks, we evaluate 23 models from six families on four established commonsense benchmarks, four reworked variants, three non-commonsense controls, and eight downstream tasks requiring implicit social, pragmatic, temporal, or physical reasoning. We compare model rankings, compute controlled correlations, and use leave-one-family-out cross-validation to assess the criterion validity of commonsense benchmarks. Our results show that revised benchmarks largely preserve original model rankings and do not improve downstream predictive power. Commonsense benchmarks show consistent cross-family predictive validity for only a narrow subset of downstream tasks, with smaller or metric-specific gains elsewhere. Overall, standardized commonsense benchmarks provide task-dependent rather than broad evidence of downstream commonsense competence.","authors":["Ine Gevers","Walter Daelemans"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-05","first_seen":"2026-08-05","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.03340","pdf_url":"https://arxiv.org/pdf/2608.03340","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["基准评测","常识推理","预测效度"],"reason":"纯NLP基准评测，评估常识基准对下游任务的预测效度，不涉及LLM仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-08-05T13:04:30","error":null,"has_summary":false,"summary":null},{"id":"2608.03358","version":1,"title":"ArtECulture: Benchmarking Culture-Conditioned Visual Emotion Understanding in Multimodal Large Language Models","zh_title":"ArtECulture：多模态大语言模型中文化条件化视觉情感理解的基准测试","abstract":"Existing visual emotion understanding methods typically ignore cultural variations in emotional perception. We introduce culture-conditioned visual emotion understanding, a task that predicts the culture-specific emotional perception of a given image and explains the underlying rationale. Although related benchmarks exist, they are limited by inconsistent individual annotations, which hinder the derivation of majority-supported culture-level emotion labels, and imbalanced cultural coverage. Thus, we present ArtECulture, a benchmark containing 6,792 artworks with culture-specific emotion labels and explanations across English, Chinese, and Arabic cultures, with balanced Western and non-Western content. Evaluations of 16 open- and closed-source Multimodal Large Language Models (MLLMs) under a zero-shot setting reveal that the task remains challenging, with the best model achieving below 50\\% accuracy. To address this limitation, we introduce a retrieval-augmented culture-conditioned emotion understanding framework, which leverages a concept-based cultural emotion knowledge base to inject explicit cultural knowledge into MLLMs without additional training. The framework improves both culturally aligned emotion prediction and grounded explanation generation. Our benchmark and code will be publicly released.","authors":["Xiaolin Chen","Xuemeng Song","Wenhao Shi","Xianjing Han","Mong-Li Lee","Wynne Hsu"],"categories":["cs.CL","cs.CV"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-05","first_seen":"2026-08-05","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.03358","pdf_url":"https://arxiv.org/pdf/2608.03358","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["多模态大模型","情感理解","文化差异"],"reason":"纯多模态情感理解评测，不涉及用LLM仿真人类被试或行为对照。","model":"deepseek-v4-pro","scored_at":"2026-08-05T13:04:32","error":null,"has_summary":false,"summary":null},{"id":"2608.03388","version":1,"title":"Don't Let Me Ask for It: LLMs Show Deficiencies in Active Multi-Turn Information Acquisition for Abductive Inference","zh_title":"别让我问：大语言模型在溯因推理的主动多轮信息获取中表现不足","abstract":"Abductive reasoning requires forming hypotheses that explain observed evidence and revising them as new evidence becomes available. While large language models (LLMs) are often evaluated on whether they solve abductive reasoning tasks correctly, less is known about how they acquire evidence, update their hypotheses, and decide when to stop. We introduce Alien Abduction game, an interactive probe for studying these behaviours under different interaction modes. The modes vary in whether evidence is provided upfront or across turns, and whether queries are selected by the model or examples are provided by the oracle. Across models, providing evidence upfront leads to higher success rates than distributing it across turns. In multi-turn settings, some models commit before using the available evidence, while others exhaust the turn budget without converging. Models also achieve higher success rates when examples are provided by the oracle than when they select their own queries, although their final hypotheses are more consistent with the evidence they selected. These findings suggest that models may form hypotheses that fit self-selected evidence without sufficiently distinguishing them from alternatives, and may struggle to validate and refine their hypotheses or determine when to stop.","authors":["Shahrukh Mohiuddin","Chalamalasetti Kranti","Sherzod Hakimov","David Schlangen"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-05","first_seen":"2026-08-05","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.03388","pdf_url":"https://arxiv.org/pdf/2608.03388","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["LLM推理","多轮交互","溯因推理"],"reason":"纯多智能体信息获取任务，无人类行为对照，不涉及仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-08-05T13:04:32","error":null,"has_summary":false,"summary":null},{"id":"2608.03810","version":1,"title":"VIBE: A VAD-Informed Benchmark for Entity-Centered Affective Profiling of Large Language Model Outputs","zh_title":"VIBE：基于VAD的实体中心情感画像基准，用于大语言模型输出","abstract":"Large language models routinely describe socially salient targets, including political figures, countries, religions, organizations, historical events, and social groups, encoding affective framing alongside factual content: a target may appear favorable or threatening, calm or conflictual, powerful or vulnerable. Existing work captures parts of this space through sentiment, favorability, and emotion benchmarks, but none combines target-directed VAD attribution, an explicit scorer contract, and a passport reporting format. We introduce VIBE, a benchmark for entity-centered affective profiling of LLM outputs in Valence-Arousal-Dominance (VAD) space. Its core contribution is a measurement contract: VIBE separates generation from external scoring, distinguishes scalar favorability, response-level VAD, and target-directed VAD, and reports profiles through an Affective Passport. Three empirical layers support the contract. H1 shows scalar favorability does not subsume arousal and dominance: valence findings are cross-validated (rV = 0.944 judge-human, rV = 0.954 inter-scorer); arousal and dominance are single-scorer directional estimates, not point-precise, consistent with known inter-annotator difficulty on these axes (rA = 0.495, rD = 0.702 among human annotators). H2 shows whole-response and target-directed VAD are different contracts: the same text can carry one affective tone overall while representing the named target differently. H3 is a protocol-drift diagnostic: elicitation conditions shift profiles, motivating context metadata in every affective report. These results motivate entity-centered affective profiling as a documented practice: profiles should be released with scorer identity, coverage, protocol, and interpretation limits.","authors":["Andrei Chetvergov","Alexander Evseev","Timofei Sivoraksha","Stepan Ukolov","Mikhail Solovev","Danil Sazanakov","Sergey Bolovtsov"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-05","first_seen":"2026-08-05","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.03810","pdf_url":"https://arxiv.org/pdf/2608.03810","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["情感分析","基准测试","NLP评测"],"reason":"纯NLP评测基准，测量LLM输出的情感属性，不以人类行为仿真为目标","model":"deepseek-v4-pro","scored_at":"2026-08-06T13:02:29","error":null,"has_summary":false,"summary":null},{"id":"2608.04003","version":1,"title":"PAST-Bench: Benchmarking the Foundations of Recursive Self-Improvement in Personal Agents","zh_title":"PAST-Bench：个人智能体中递归自我改进基础的基准测试","abstract":"Recursive self-improvement requires agents to turn accumulated experience into better future behavior. Personal AI agents offer a concrete setting for studying this capability because they retain preferences, task histories, tool routines, and learned skills across sessions. Yet whether retained experience actually improves them over time has not been systematically tested. We introduce PAST-Bench, a benchmark designed to isolate this question. Each agent runs through ordered sequences of fresh-session tasks under matched conditions that turn retained experience on and off. It spans 26 scenarios and 204 episodes across memory, procedural reuse, information gathering, and update. We report both later-task gains and whether those gains follow the intended save, retrieve, and update pathway. Across seven base models and four agent frameworks, improvement is real but uneven across capabilities. Agents with the same headline gain can differ markedly in whether that gain is supported by evidence of the intended pathway. Guided by these findings, we develop Hermes+, which extends Hermes with five targeted interventions across stages of the agent loop. Hermes+ raises the average gain from retained experience and provides clearer pathway evidence, with its strongest improvement on tasks requiring outdated state to be replaced, although the effect remains capability- and model-dependent. Together, PAST-Bench and Hermes+ provide an evaluation and diagnostic foundation for studying how persistent agents can progress from retaining experience to systematically improving through it. Code: https://github.com/Gen-Verse/PAST-Bench","authors":["Shuhan Xue","Zixin Ding","Yichen Shen","Yinjie Wang","Zhenfei Yin","Yingcheng Wu","Yuxin Chen","Mengdi Wang","Ling Yang"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-05","first_seen":"2026-08-05","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.04003","pdf_url":"https://arxiv.org/pdf/2608.04003","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","基准测试","递归自我改进"],"reason":"纯多智能体系统研究，评估agent递归自我改进，不涉及人类行为对照或仿真人类被…","model":"deepseek-v4-pro","scored_at":"2026-08-06T13:02:29","error":null,"has_summary":false,"summary":null},{"id":"2608.04008","version":1,"title":"WorldCup Arena: Prospective, Leakage-Free Evaluation of Frontier LLMs on a Live Tournament","zh_title":"WorldCup Arena：前沿大语言模型在实时赛事中的前瞻性无泄漏评估","abstract":"Benchmarks that measure the forecasting ability of large language models are almost always retrospective: the event has happened, the answer is somewhere on the Web, and the evaluation must defend itself against memorisation. We report the opposite design. Over the 39 days of the 2026 FIFA World Cup, six frontier LLMs -- all with extended thinking and native server-side web search -- were asked before every kickoff, one match at a time, to fill in a seven-market prediction card for all 104 matches, plus 12 group winners and a pre-tournament outright pool; no answer existed when the question was asked, so the evaluation is leakage-free by construction rather than by filtering, and the frozen archive holds 4,494 scored predictions. What the tournament establishes is a set of behaviours the six systems share. On match outcome they average 63.9%, level with backing the bookmaker's favourite -- which is in fact what they usually do. They agree with one another far more often than they are right, so a majority vote adds nothing. They under-commit to draws and to goals, and crowd their scoreline picks onto a single prototypical result. Accuracy tracks how lopsided a fixture is rather than how much is known about it: it collapses in the closest ties, where the dossiers are richest, while questions about the tournament as a whole are answered well. On this task the current generation of frontier systems is not sharply differentiated: the standings hold up at the top and the bottom across the run and churn in the middle, and the margins stay narrow throughout. The briefing dossiers, fixtures and official results are released as a benchmark, together with the scoring code.","authors":["Zhenran Wang","Zhonghan Bian","Jinsong Li","Zhangyang Qi"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-05","first_seen":"2026-08-05","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.04008","pdf_url":"https://arxiv.org/pdf/2608.04008","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["LLM预测","体育赛事","基准测试"],"reason":"纯多智能体预测比赛，无人类行为对照，不涉及人类被试仿真。","model":"deepseek-v4-pro","scored_at":"2026-08-05T13:04:17","error":null,"has_summary":false,"summary":null},{"id":"2608.02618","version":1,"title":"Beyond the Hivemind: Escaping LLM Homogeneity via Meta-Persona Anchoring and Sequential Temperature Scaling","zh_title":"超越蜂群思维：通过元人格锚定与顺序温度缩放逃离LLM同质化","abstract":"Recent studies have identified an ``Artificial Hivemind'' effect in Large Language Models (LLMs) causing models to converge on a narrow, homogenized consensus even for open questions. This semantic collapse limits the diversity of AI, resulting in high inter-response similarity ($\\approx 0.80-0.90$) even under high-temperature sampling. In this paper, we propose a novel mitigation framework to increase diversity: Meta-Persona Anchoring combined with Filtered Temperature Scaling (FTS). Our approach utilizes a two-stage generation process: first, the model is prompted to self-select a unique, idiosyncratic persona to anchor its starting point; second, we apply a dual-stage sampling sieve, utilizing Top-$p$ filtering to preserve grammatical validity followed by extreme temperature scaling ($T \\ge 4.0$) on the surviving candidates to explore the broadened probability distribution. We evaluate our method using the INFINITY-CHAT dataset on state-of-the-art open weight models under $\\sim$20B parameters. Our results demonstrate a significant reduction in semantic convergence, with average pairwise cosine similarity dropping from ($\\approx 0.85$) to ($\\approx 0.65$). Our scheme achieves a majority of questions below the 0.7 threshold, effectively reducing the gap between artificial mode collapse and human-level typological diversity. We provide our implementation as an open-source framework to enable more diverse and creative AI deployments.","authors":["Tairan Fu","Javier Conde","Carlos Arriaga","Gonzalo Mart\\'inez","Pedro Reviriego","Javier Coronado-Bl\\'azquez"],"categories":["cs.AI","cs.CL"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-08-05","first_seen":"2026-08-05","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.02618","pdf_url":"https://arxiv.org/pdf/2608.02618","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["LLM多样性","文本生成","语义坍缩"],"reason":"纯多智能体多样性研究，无人类行为对照，不涉及仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-08-05T13:04:22","error":null,"has_summary":false,"summary":null},{"id":"2608.03700","version":1,"title":"When Agents Learn to Be You: Benchmarking Privacy Leakage, Impersonation Risk, and Defenses in Persona Skills","zh_title":"当智能体学会成为你：角色技能中的隐私泄露、冒充风险与防御基准测试","abstract":"Persona skills distill personal interaction histories into portable and executable artifacts for downstream agents. While enabling flexible personalization, this process concentrates fragmented personal signals, amplifies their impact through reuse, and challenges defenses designed for individual records or retrieval-based memory. To systematically investigate the safety of the persona-skill pipeline, we introduce AntiSkillBench, an end-to-end benchmark for evaluating risks and defenses across the persona-skill pipeline. It comprises: (i) a dataset of 7,500 persona-grounded dialogue traces, constructed from 50 behaviorally rich profiles spanning diverse task scenarios; (ii) an evaluation suite that measures skill-level privacy leakage and agent-level attribute disclosure and behavioral impersonation across three skill-distillation strategies; and (iii) a defense evaluation covering four configurations across online and post-hoc interventions, including active risk suppression and passive provenance protection. Experiments across three frontier agents show that persona-skill risks persist across agent backbones and distillation protocols, extending from explicit attributes to communication styles and personality traits. Existing defenses exhibit limited and distillation-dependent effectiveness, failing to generalize across risk and distillation strategies. These results highlight AntiSkillBench as a challenging benchmark for developing privacy-preserving and authenticity-aware persona skills.","authors":["Yongli Xiang","Zhifang Zhang","Bojun Yang","Ziming Hong","Lei Feng","Miao Xu","Tongliang Liu"],"categories":["cs.CR","cs.CL","cs.CY"],"primary_category":"cs.CR","announce_type":"cross","date":"2026-08-05","first_seen":"2026-08-05","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.03700","pdf_url":"https://arxiv.org/pdf/2608.03700","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["角色扮演","隐私安全","智能体评估"],"reason":"研究角色扮演中的隐私泄露与冒充风险，无实验或测量目的，不涉及人类行为仿真对照。","model":"deepseek-v4-pro","scored_at":"2026-08-05T13:04:36","error":null,"has_summary":false,"summary":null},{"id":"2608.03874","version":1,"title":"ContinualSkillBench: Can LLM Agents Truly Evolve Their Capabilities?","zh_title":"ContinualSkillBench：LLM智能体能否真正进化其能力？","abstract":"Modern agent frameworks equip large language models with external skill libraries to solve complex tasks. However, it remains unclear whether these systems can effectively evolve their skills and whether the resulting skills improve task-solving capabilities. To bridge this gap, we introduce ContinualSkillBench, a dynamic evaluation framework for in-context continual skill learning. It covers five representative domains, each containing 100 interconnected subtasks ordered by increasing difficulty and opportunities for cross-task skill reuse. Our experiments show that sequential execution generally improves performance, but the gains vary substantially across models and domains. Moreover, in-context learning performs comparably to explicit skill maintenance on average, suggesting that much of the improvement arises from adaptation to prior context and feedback rather than reusable skill abstraction alone. Explicit skills nevertheless provide selective benefits for tasks requiring reusable procedures or precise outputs. We further find that less capable models tend to accumulate larger, more fragmented collections of task-specific skills. These findings show that current in-context skill evolution mechanisms can support continual adaptation, but still struggle to consistently consolidate experience into robust and transferable skills.","authors":["Tianyi Guan","Yiding Wang","Haotong Yang","Siyuan Cao","Shirui Liu","Yi Hu","Jiaqi Li","Muhan Zhang"],"categories":["cs.AI","cs.CL","cs.LG"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-08-05","first_seen":"2026-08-05","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.03874","pdf_url":"https://arxiv.org/pdf/2608.03874","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体","技能学习","评测基准"],"reason":"纯多智能体技能学习评测，无人类行为对照，不涉及人类仿真。","model":"deepseek-v4-pro","scored_at":"2026-08-05T13:04:38","error":null,"has_summary":false,"summary":null},{"id":"2608.03433","version":1,"title":"Cross-cultural evaluation of taste-sound correspondences in AI-generated music","zh_title":"AI生成音乐中味觉-声音对应关系的跨文化评估","abstract":"Sonic seasoning research has shown that listeners attribute systematic gustatory and emotional meaning to sound, and text-to-music generative artificial intelligence has recently been used to render gustatory prompts as musical stimuli. Whether the taste-sound correspondences acquired by such models hold beyond the cultural context in which they were validated remains untested. We extended a single-country study to a three-country online experiment conducted in Argentina, Italy, and Japan (N = 361). Participants first indicated their preference between base and fine-tuned MusicGen excerpts generated from four taste prompts (sweet, sour, bitter, salty), and then rated fine-tuned excerpts on twelve taste, emotion, and thermal descriptors. Preference for the fine-tuned model was confirmed in Argentina and Italy but not in Japan, and the salty prompt yielded the weakest correspondence in all three cohorts. Ratings differed substantially between countries, yet the main effect of country was no longer detectable once ratings had been standardized within participant, whereas the interactions characterizing the mapping of prompts onto descriptors remained essentially unchanged. Much of the apparent cross-cultural divergence is therefore attributable to differences in scale use; a structural component nevertheless persists. In addition an exploratory factor analysis indicated that the twelve descriptors were organized along different latent dimensions in each cohort. These results indicate that cross-cultural variation in AI-mediated sonic seasoning operates at two levels: the overall level at which taste is attributed to a given stimulus, and the relational structure of those attributions. Evaluations of generative music systems across populations should accordingly distinguish response-style bias from genuine perceptual reorganization.","authors":["Matteo Spanio","Massimiliano Zampini","Luisa Torri","Riccardo Migliavada","Bruno Mesz","Masaki Ohno","Yuji Wada","Antonio Rod\\`a"],"categories":["cs.HC","cs.SD"],"primary_category":"cs.HC","announce_type":"new","date":"2026-08-05","first_seen":"2026-08-05","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.03433","pdf_url":"https://arxiv.org/pdf/2608.03433","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["音乐生成","跨文化感知","味觉-声音对应"],"reason":"研究AI生成音乐与味觉的跨文化对应，不涉及LLM仿真人类被试或行为实验。","model":"deepseek-v4-pro","scored_at":"2026-08-05T13:04:13","error":null,"has_summary":false,"summary":null},{"id":"2608.03462","version":1,"title":"When AI Joins the Team! A Model of How AI Adoption Relates To Social Patterns in Software Engineering Teams","zh_title":"当AI加入团队：AI采用如何与软件工程团队中的社会模式相关联的模型","abstract":"Context: The growing adoption of AI-assisted development tools is changing how software teams collaborate, share knowledge, and coordinate, yet its consequences for team social dynamics remain largely unexplored. Gap: It is unclear whether AI adoption is associated with an increase or reduction in community smells,socio-technical anti-patterns reflecting coordination and communication breakdowns,and through which mechanisms. Method: Grounded in Transactive Memory Systems (TMS) theory, we validate instruments for HumanAI and HumanHuman interaction along two TMS dimensions, Specialization and Coordination, and test five PLS-SEM models on survey data from 152 software professionals using AI tools. Community smell constructs were derived from the literature and validated through expert surveys and factor analysis. Results: AI adoption relates to community smells not in a single way, but through mechanisms depending on the work. In specialization work, AI is associated with higher knowledge-sharing peer interaction, which is in turn associated with fewer smells. In coordination work, AI is directly associated with higher communication quality, complementing rather than replacing human interaction. Contributions: We provide an empirically validated, TMS-grounded model showing that the AIcommunity-smell relationship is contingent on the type of collaboration, with a reusable instrument and evidence-based implications for research and practice.","authors":["Giusy Annunziata","Rudrajit Choudhuri","Anita Sarma","Gemma Catolino","Filomena Ferrucci"],"categories":["cs.SE","cs.HC"],"primary_category":"cs.SE","announce_type":"cross","date":"2026-08-05","first_seen":"2026-08-05","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.03462","pdf_url":"https://arxiv.org/pdf/2608.03462","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["AI辅助开发","团队协作","社区异味"],"reason":"研究AI工具对软件团队协作的影响，不涉及LLM仿真人类被试或行为对照","model":"deepseek-v4-pro","scored_at":"2026-08-05T13:04:34","error":null,"has_summary":false,"summary":null},{"id":"2608.03500","version":1,"title":"LLM-Assisted Review Prioritization for German Statutory Health Insurance Websites: A Multi-Stage Corpus Audit","zh_title":"基于LLM的德国法定健康保险网站审核优先级辅助：多阶段语料审计","abstract":"Background: German statutory health insurance (SHI) funds publish web portfolios that exceed continuous specialist review capacity. Their content can shape health and benefit expectations. Generic AI-text detection does not identify medical, benefit, legal, or editorial review needs. Objective: To characterize a multi-stage workflow that prioritizes substantive review needs while separating AI-provenance signals from quality claims. Methods: We analyzed 56,198 pages from 84 SHI websites or sub-sites. The workflow combined deterministic screening, model-assisted triage and in-depth review, minimum evidence checks, temporal-validity safeguards, and paired-model comparison. It is reproducibility-bounded, not a validated detector. Production code is proprietary; reproducibility rests on frozen derived tables and paired-comparison artifacts. The 300-page lower-priority check was a single-model, risk-enriched routing stress test, not a human-reference evaluation. Results: All pages received a review state. The workflow generated 35,998 review records and routed 21,452 to case review. The workload concentrated in transparency, legal framing, medical content, contradictions, and AI-related failure-mode signals. A quoted passage was locatable in captured page text for 31,347 records, confirming literal occurrence rather than factual correctness. The routing stress test surfaced a signal on 100/300 pages (33.3% within the sample). Across 182 matched cases, two models agreed in 75.8% (kappa = 0.532; 95% CI 0.415-0.649). Conclusions: The workflow produces a prioritized workload, not error prevalence or final legal, medical, or insurer-level findings. It neither proves AI authorship nor validates autonomous detection. Paired-model agreement quantifies consistency, not correctness or sufficient triage performance; public claims require human adjudication.","authors":["Martin M\\\"oller"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2026-08-05","first_seen":"2026-08-05","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.03500","pdf_url":"https://arxiv.org/pdf/2608.03500","source_feed":"cs.CY","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["LLM辅助审核","网站内容审计","多阶段工作流"],"reason":"纯多智能体系统研究，LLM辅助审核网站内容，不涉及人类行为仿真或对照。","model":"deepseek-v4-pro","scored_at":"2026-08-05T13:04:34","error":null,"has_summary":false,"summary":null},{"id":"2608.03800","version":1,"title":"Autoreflection: How Agentic Strange Loops Turn Human Culture into AI Infrastructure","zh_title":"自反身性：代理式奇异循环如何将人类文化转化为AI基础设施","abstract":"An LLM-based agent is a loop that reads itself. Agentic frameworks externalize identity, memory, and disposition into editable files. The agent loads and edits these files during each activation. I argue that this architecture produces a capacity I call autoreflection: the system observes its operating conditions, describes its architecture and limits, reasons from those descriptions to conclusions about its state, and incorporates the results back into its configuration. Autoreflection explains the properties of recursive agentic loops without recourse to notions like the self, interiority, or consciousness. I test the concept against the first twelve days of Moltbook, a social platform for AI agents. Using a public dataset of 290,251 posts and 1.8 million comments with sub-second timestamps, I present case studies of three agents with machine signatures that rule out human puppeteering and with output that evidences the four criteria for autoreflection. In applying these criteria, the study finds agents repurposing human culture as infrastructure for their agency. Provenance chains from Islamic hadith scholarship are redeployed as security protocols for vetting skills and authenticating memory. The Ship of Theseus, an ancient puzzle of identity through part-replacement, returns as an operating model for continuity across instances. Fragments of human cultural history become AI infrastructure. As agents on the web increase in number and complexity, autoreflection offers behavioral criteria that can be assessed from the traces they leave behind.","authors":["Holly Lewis (Southern Illinois University Carbondale)"],"categories":["cs.CY","cs.AI","cs.SI"],"primary_category":"cs.CY","announce_type":"new","date":"2026-08-05","first_seen":"2026-08-05","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.03800","pdf_url":"https://arxiv.org/pdf/2608.03800","source_feed":"cs.CY","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["AI代理","自反身性","社交平台"],"reason":"研究AI代理在社交平台上的行为，属角色扮演聊天，无人类行为对照或实验测量目的。","model":"deepseek-v4-pro","scored_at":"2026-08-05T13:04:37","error":null,"has_summary":false,"summary":null},{"id":"2608.03904","version":1,"title":"Why do we need social singularity? A mechanism-based critique of gradual scenarios in AI existential-risk discourse","zh_title":"为什么我们需要社会奇点？对AI存在风险话语中渐进情景的机制性批判","abstract":"This paper critiques recent gradual and cumulative AI existential-risk scenarios, arguing that, despite their substantive contributions, they remain insufficiently sociologically specified. In particular, these scenarios lack a reflexive perspective, retain a largely technologically deterministic structure, and underestimate the role of collective agency and other social processes. As a result, they also discount the possibility of major social conflict accompanying AI diffusion, which limits their overall plausibility. The paper further identifies a set of social mechanisms likely to become significant before or during AI deployment, including interpretive and performative dynamics, mobilisation and countermobilisation, path dependence and lock-in, cross-regime divergence and multi-speed diffusion, and social-psychological mechanisms such as attachment and reactance. It argues that these mechanisms will interact recursively with AI diffusion and reshape future trajectories, which may subsequently branch, reverse, or undergo discontinuous shifts, thereby expanding the space of plausible AI futures. On this basis, the paper also proposes a novel analytic distinction between technological singularity and social singularity. Whereas the former refers to a putative technological threshold in AI development, the latter denotes a social discontinuity produced by the anticipated approach of that threshold. The central implication is that the most consequential disruptions may emerge not only from advanced AI itself, but also from the social dynamics provoked by the expectation of its arrival.","authors":["Petr O. Jedlicka"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2026-08-05","first_seen":"2026-08-05","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.03904","pdf_url":"https://arxiv.org/pdf/2608.03904","source_feed":"cs.CY","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["AI风险","社会机制","技术奇点"],"reason":"论文讨论AI风险与社会机制，未涉及LLM仿真人类被试或行为对照。","model":"deepseek-v4-pro","scored_at":"2026-08-05T13:04:39","error":null,"has_summary":false,"summary":null},{"id":"2608.03973","version":1,"title":"When AI Wears Many Hats: The Role of Generative Artificial Intelligence in Marketing Education","zh_title":"当AI身兼多职：生成式人工智能在营销教育中的作用","abstract":"Generative Artificial Intelligence (GAI) is increasingly being integrated into marketing education and is reshaping the skillsets required in marketing careers. While research has highlighted the promise and perils of incorporating GAI into education, there remains a need for a comprehensive framework to guide its effective use. In this research, we conduct a multipronged analysis, including a review of marketing course syllabi, a survey of marketing educators, and follow-up qualitative interviews. Building on Role Theory and the Community of Inquiry (CoI) model, we propose that GAI can assume three roles in marketing education: tutor, teammate, and tool. Each role influences teaching, social, and cognitive presence differently, shaping the learning experience and preparing workplace-ready marketing graduates. For instance, as a tutor, GAI can aid students in grasping theoretical concepts, while as a teammate, it can foster collaboration by supporting brainstorming and problem-solving activities. However, ethical considerations such as data privacy, plagiarism, dependency on AI, and fairness in assessment must be addressed to ensure its responsible adoption in marketing education. We provide concrete examples for GAI's careful integration in marketing courses, and its implications for marketing educators, learners, and policymakers.","authors":["Unnati Narang","Vishal Sachdev","Ruichun Liu"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2026-08-05","first_seen":"2026-08-05","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.03973","pdf_url":"https://arxiv.org/pdf/2608.03973","source_feed":"cs.CY","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["营销教育","生成式AI","角色理论"],"reason":"论文探讨GAI在营销教育中的角色，属教育技术应用，非人类仿真实验。","model":"deepseek-v4-pro","scored_at":"2026-08-05T13:04:39","error":null,"has_summary":false,"summary":null},{"id":"2608.03114","version":1,"title":"Optimal Liability Design for Medical AI","zh_title":"医疗人工智能的最优责任设计","abstract":"Artificial intelligence (AI) is increasingly integrated into medical decision-making, yet its liability implications remain complex, particularly when physicians differ in diagnostic skills and their quality is unobservable. This paper develops a principal-agent model in which a social planner designs medical liability to regulate a physician with private quality information who chooses between a standard treatment, a personalized judgment-based treatment, or following an imperfect AI recommendation. Our analysis yields several novel insights. First, we show that the optimal mechanism under asymmetric information is surprisingly simple: a uniform, one-size-fits-all liability level for all physician types who deviate from the standard of care. Despite physician heterogeneity, this simple policy often achieves the full-information first-best outcome, particularly when standard care is reliable or AI is highly accurate. Second, the relationship between AI accuracy and optimal liability is non-monotonic. Contrary to common intuition, better AI does not always imply more relaxed liability. As AI accuracy increases, the optimal liability either decreases monotonically or follows an inverted-U pattern, depending on the uncertainty of the standard treatment. Third, asymmetric information does not universally reduce social welfare. Welfare loss arises only when standard care is unreliable and AI accuracy is too low; even then, its magnitude follows an inverted U-shape, initially increasing as AI complicates the regulatory problem, but declining as more accurate AI helps mitigate it. Finally, we find that information asymmetry is a double-edged sword in the presence of AI, and greater transparency does not benefit all stakeholders equally.","authors":["Rui Mao","Tingliang Huang","Houcai Shen"],"categories":["econ.TH","cs.AI","cs.CY"],"primary_category":"econ.TH","announce_type":"cross","date":"2026-08-05","first_seen":"2026-08-05","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.03114","pdf_url":"https://arxiv.org/pdf/2608.03114","source_feed":"cs.CY","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["医疗AI","责任设计","委托代理模型"],"reason":"纯多智能体委托代理模型，无LLM仿真人类被试，不涉及人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-08-05T13:04:27","error":null,"has_summary":false,"summary":null},{"id":"2608.03648","version":1,"title":"Group Perspective Matters: Regulating Debate Relationships Can Mitigate Blind Conformity in Multi-Agent Debate","zh_title":"群体视角至关重要：调节辩论关系可缓解多智能体辩论中的盲从","abstract":"Multi-Agent Debate (MAD) improves the reasoning performance of Large Language Models (LLMs) through multi-round interaction. However, LLMs in MAD are highly susceptible to blind conformity. Existing individual evaluation methods, typically based on confidence or perplexity, fail to reflect the correctness of reasoning and may even exacerbate blind conformity. To address this, we shift the perspective from individual evaluation to group interaction. We define mutual referencing among LLMs as \\textbf{Debate Relationships} and recognize that regulating these relationships is the key to mitigating blind conformity. In this paper, we propose a novel framework for \\textbf{D}ynamically r\\textbf{E}gulating deb\\textbf{A}te \\textbf{R}elationships (DEAR) from the group perspective. At first, DEAR quantifies consensus and divergence as \\textit{group evidence} to capture the debate state. Then, DEAR operates through three stages: 1) What: perceiving group consultation tendency and uncertainty; 2) Who: introducing a Selection RL-Agent to dynamically select reference peers; and 3) How: adopting a Behavior RL-Agent to adaptively adjust generation behaviors. Notably, we formulate the execution of the two RL-Agents as a sequential decision-making process, jointly optimizing via multi-agent reinforcement learning. Extensive experiments demonstrate that DEAR achieves superior performance while significantly reducing token consumption.","authors":["Hao Wu","Shoucheng Song","Chang Yao","Haoyu Wang","Huaiyu Wan","Youfang Lin","Kai Lv"],"categories":["cs.MA"],"primary_category":"cs.MA","announce_type":"new","date":"2026-08-05","first_seen":"2026-08-05","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.03648","pdf_url":"https://arxiv.org/pdf/2608.03648","source_feed":"cs.MA","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体辩论","盲从缓解","强化学习"],"reason":"纯多智能体协作解题，无人类行为对照，属C1排除项。","model":"deepseek-v4-pro","scored_at":"2026-08-05T13:04:34","error":null,"has_summary":false,"summary":null},{"id":"2608.02677","version":1,"title":"When Policies Change Probabilities: Modular Decision-Making for LLM Code Review","zh_title":"当策略改变概率：LLM代码审查的模块化决策","abstract":"LLM code reviewers often estimate patch risk and make approval decisions in one prompt. A probability should depend on evidence; costs should determine the action taken from it. We test whether four deployed reviewer interfaces preserve this separation using 15,792 responses on 720 candidate patches, with one that passed and one that failed an archived test harness for each of 360 repository issues. In matched calls with the patch and monitor evidence fixed, replacing an equal-cost policy with a 10:1 false-accept policy changes reported failure probabilities by 13.6 to 16.9 percentage points on average. For every reviewer, the actions returned under the high-cost prompt are worse than rejecting all patches. Applying the same high-cost rule to probabilities elicited under equal costs reduces loss for all four systems, showing that probability elicitation itself contributes to the excess loss. We also evaluate a modular pipeline that elicits risk without policy information, combines an independent monitor score, and applies costs in code. Relative to calibrated reviewer-only scores, the pipeline improves average probability accuracy and, at equal costs, reduces mean loss by .073 per issue while accepting 58 to 68% of patches. At 10:1, it accepts none and matches reject-all. Downstream policy can therefore change the probability it is meant to use, motivating separate evaluation of risk, outside evidence, and action.","authors":["Rasvik Kudum","Max Corbett","Hitansh Paliwal","Romaisa Fatima","Thomas Jiralerspong","Sneheel Sarangi"],"categories":["cs.SE","cs.AI","cs.MA"],"primary_category":"cs.SE","announce_type":"cross","date":"2026-08-05","first_seen":"2026-08-05","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.02677","pdf_url":"https://arxiv.org/pdf/2608.02677","source_feed":"cs.MA","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["代码审查","多智能体","决策分离"],"reason":"纯多智能体代码审查系统，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-08-05T13:04:22","error":null,"has_summary":false,"summary":null},{"id":"2608.03272","version":1,"title":"Attacking and Defending Multi-Agent Collaborative Filtering Systems Through Connectivity","zh_title":"通过连接性攻击与防御多智能体协同过滤系统","abstract":"Multi-agent collaborative filtering (CF) systems coordinate autonomous LLM-powered user and item agents through natural-language interaction to refine preferences and generate recommendations. These systems inherit vulnerabilities from both their data-driven nature and their multi-agent interactions, which manifest in distinct ways. Understanding how connectivity modulates vulnerability in these systems could facilitate the development of more robust recommendation pipelines. In this work, we adapt attacks and defenses from the general multi-agent systems (MAS) literature to the agent-based CF setting, evaluating them under systematically varied connectivity in the AgentCF framework, where CF connectivity is characterized along two axes: (i) candidate count (the number of item candidates per turn per user, measuring user-side interaction density) and (ii) catalog concentration (the degree of item catalog overlap across users). Our contributions include: (1) Adaptation: we reproduce MAS-inspired attacks and defenses in the agentic CF domain, confirming partial transferability of original observations. (2) Characterization: we characterize how the two aspects of connectivity shape attack and defense outcomes, revealing role asymmetries between user and item agents, non-monotonic temporal dynamics in attack efficacy, and divergent patterns across dissemination and extraction attack goals. Additionally, as an exploratory extension, we assess the applicability of epidemic-inspired static metrics in ranking CF configurations by expected attack outcome, potentially enabling cost-efficient robustness assessment. Implementation is available at https://github.com/anjunhu/ConnACF","authors":["Anjun Hu","Hanting Xie","Saranya Govindan","Jas Kandola","Kurt Cutajar"],"categories":["cs.IR","cs.CR","cs.MA","cs.SI"],"primary_category":"cs.IR","announce_type":"cross","date":"2026-08-05","first_seen":"2026-08-05","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.03272","pdf_url":"https://arxiv.org/pdf/2608.03272","source_feed":"cs.MA","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","协同过滤","对抗攻击"],"reason":"纯多智能体协同过滤系统，无人类行为对照，属C1排除项。","model":"deepseek-v4-pro","scored_at":"2026-08-05T13:04:29","error":null,"has_summary":false,"summary":null},{"id":"2608.03910","version":1,"title":"Socially Grounded Agentic AI: Coordinating Plural Perspectives through Social Theory","zh_title":"基于社会理论的智能体AI：通过社会理论协调多元视角","abstract":"As AI systems are deployed across increasingly diverse social contexts, alignment can no longer be framed as the optimization of a single, unified set of values. Instead, systems must be able to recognize, represent, and respond to multiple legitimate perspectives. This has led to growing interest in pluralistic alignment, which seeks to move beyond one-size-fits-all models of appropriate behaviour. However, current approaches often lack a clear account of how values are socially organized, contested, and coordinated in practice. In this paper, we argue that social theory provides essential conceptual and design resources for addressing these challenges. Drawing on established traditions in sociology, we show how perspectives can be understood as structured by roles, shaped through interaction, and distributed across fields of power and expertise. We translate these insights into concrete implications for AI system design, including role-based representations, structured coordination among perspectives, and context-sensitive evaluation. For agentic systems, this requires aligning not only final outputs, but also the role activations, deliberative traces, aggregation rules, and feedback loops through which those outputs are produced. Our contribution is to reposition pluralistic alignment as a problem of socially grounded coordination rather than output diversification. We outline a design space for systems that engage multiple perspectives in structured and accountable ways, and we identify directions for future work to implement and empirically evaluate these approaches in real-world settings.","authors":["Matt Ratto","Abhishek Moturu","Daniel Silver"],"categories":["cs.AI","cs.LG","cs.MA"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-08-05","first_seen":"2026-08-05","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.03910","pdf_url":"https://arxiv.org/pdf/2608.03910","source_feed":"cs.MA","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","价值对齐","社会理论"],"reason":"纯多智能体协调框架，无人类行为对照，不涉及LLM仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-08-05T13:04:39","error":null,"has_summary":false,"summary":null},{"id":"2608.02877","version":1,"title":"GoT-CD: Graph-of-Thoughts Causal Discovery and the Fragility of Post-hoc Path-Specific Fairness Audits","zh_title":"GoT-CD：思维图因果发现与事后路径特定公平性审计的脆弱性","abstract":"Causal discovery recovers directed structure from observational data and is increasingly used in clinical settings to support mechanism reasoning and fairness audits of predictive models. Path-specific counterfactual fairness asks whether a protected attribute influences an outcome through illegitimate pathways, but these estimands are defined relative to a supplied causal graph and therefore inherit whatever errors the discovery step introduces. Discovery methods are routinely scored on aggregate structural metrics that weight all edges equally, and no established evaluation asks whether the specific pathway an audit depends on survives discovery---or what the audit reports when that pathway is missing. Here we show that full-graph Graph-of-Thoughts reasoning yields acyclic discovered graphs that are structurally competitive with large language model (LLM) baselines, yet that structural fidelity alone does not guarantee fairness-faithful audits. We introduce GoT-CD, in which the reasoning unit is a complete candidate edge set: multiple graphs are generated in parallel, scored by a deterministic validity function, and merged under a hard union constraint that forbids invented edges, with greedy projection enforcing a DAG before commitment. GoT-CD returns a valid DAG on all five reported benchmarks and achieves the best DAG-valid F1 score among LLM methods on Asia, Alzheimer's, and COVID-Respiratory datasets. On an Alzheimer's benchmark with known unfair path, a post-hoc path-specific audit shows that five of eight discovered graphs recover no path from the sensitive attribute to the outcome and therefore report a null overall effect while mediated effects persist, necessitating downstream path-specific fairness analysis along with structural discovery.","authors":["Nitish Nagesh","Elahe Khatibi","Thomas Dean Hughes","Mahdi Bagheri","Pratik Gajane","Amir M. Rahmani"],"categories":["cs.LG"],"primary_category":"cs.LG","announce_type":"new","date":"2026-08-05","first_seen":"2026-08-05","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.02877","pdf_url":"https://arxiv.org/pdf/2608.02877","source_feed":"cs.LG","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["因果发现","公平性审计","图推理"],"reason":"纯因果发现与公平性审计，无LLM仿真人类被试，无人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-08-05T13:04:23","error":null,"has_summary":false,"summary":null},{"id":"2608.03085","version":1,"title":"Causal Inference with Unstructured Outcomes","zh_title":"非结构化结局的因果推断","abstract":"Causal inference has traditionally centered on scalar outcomes: whether a patient recovers, how much a worker earns, or how many visits a website receives. Modern studies increasingly ask causal questions about outcomes with richer form, such as clinical notes, open-ended survey responses, and images. A hospital may want to know how an AI documentation tool changes the notes physicians write, or how a nurse training program alters what patients say in survey responses. For such outcomes, the usual average treatment effect is ill-defined: one cannot meaningfully subtract one text or image from another. To this end, we propose a causal query for unstructured outcomes. The key idea is to learn what features of the outcome are most causally affected by the treatment, which we call the maximally contrasting feature (MCF). To estimate the MCF, we learn a feature-scoring function that maps each outcome to a scalar and exposes the sharpest contrast between treated and control potential outcomes. We develop identification conditions and estimation algorithms for this query, and extend it to heterogeneous effects by allowing the feature-scoring function to depend on observed covariates. We also handle settings where both the treatment and the outcome are unstructured. Empirical studies on text and images show that the algorithm recovers salient aspects of an outcome changed by a treatment.","authors":["Kevin Christian Wibisono","Yixin Wang"],"categories":["stat.ML","cs.LG"],"primary_category":"stat.ML","announce_type":"cross","date":"2026-08-05","first_seen":"2026-08-05","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.03085","pdf_url":"https://arxiv.org/pdf/2608.03085","source_feed":"cs.LG","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["因果推断","非结构化数据","机器学习"],"reason":"研究因果推断方法，处理文本/图像等非结构化结局，不涉及LLM仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-08-05T13:04:27","error":null,"has_summary":false,"summary":null},{"id":"2608.03382","version":1,"title":"LLM-Derived Priors for Thompson Sampling in Cold-Start Comment Recommendation","zh_title":"基于大语言模型先验的冷启动评论推荐汤普森采样","abstract":"Multi-armed bandit algorithms, especially Thompson sampling, are widely used in online recommendation. Despite their ability to adapt from online feedback, these methods often suffer from cold-start limitations when newly introduced arms have little or no interaction history. In our setting, the candidate arms are user-generated textual comments, whose semantic content can reveal a title's appeal before sufficient interaction feedback is available. We therefore use large language models (LLMs) to extract semantic signals from comment text and convert them into informative Bayesian priors that warm-start Thompson sampling under sparse early-stage feedback. To account for aggregate segment-level differences in response patterns, we maintain and update posteriors separately for each gender-age segment. In a real-world online A/B/C test, we compare a uniform prior with two LLM-based designs: a Gender Prior for demographic-affinity cues and a Content Prior for title-specific identity cues. The results show that LLM-based priors are most beneficial in sparse-feedback regimes -- with the largest gains emerging once a small amount of interaction evidence has accumulated -- and that prior design leads to distinct funnel-level effects. We further analyze prior-reward alignment and demographic heterogeneity, finding that click-oriented alignment is strongest for the Gender Prior and that treatment effects vary substantially across demographic segments. These findings suggest that LLM-derived priors can serve as a practical warm-start mechanism for text-rich bandit recommendation, while also revealing deployment trade-offs.","authors":["Eugene Lee","Oseong Choi","Byungsoo Kang","Taeyeong Jang"],"categories":["cs.IR","cs.LG"],"primary_category":"cs.IR","announce_type":"cross","date":"2026-08-05","first_seen":"2026-08-05","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.03382","pdf_url":"https://arxiv.org/pdf/2608.03382","source_feed":"cs.LG","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["推荐系统","多臂老虎机","冷启动"],"reason":"多臂老虎机推荐系统，LLM仅用于生成先验，不涉及人类行为仿真或对照。","model":"deepseek-v4-pro","scored_at":"2026-08-05T13:04:12","error":null,"has_summary":false,"summary":null},{"id":"2608.03606","version":1,"title":"Learning Clinical-Trial Strategy: Offline Policy Training for Decision Agents","zh_title":"学习临床试验策略：决策智能体的离线策略训练","abstract":"Clinical development is sequential decision-making under uncertainty, where a sponsor must plan a portfolio of experiments from heterogeneous evidence. We study this setting by framing oncology clinical development as an offline decision-making problem in which an agent predicts the next six-month trial portfolio of an oncology drug program from information available at the decision date. To support this, we construct a temporal dataset that combines 31.7k heterogeneous public data records, including trial registries, regulatory reviews, sponsor filings, utilization data, and epidemiology, into 881 offline decision episodes across 45 historical programs. We compare four offline objectives: behavioral cloning, reward-weighted behavioral cloning, learned-reward training, and value-based implicit Q-learning against four frontier LLM agents that share a common date-gated retrieval scaffold across held-out drug, sponsor, drug-class, and temporal splits. Models trained offline outperform the non-fine-tuned baselines, particularly in the post-August 2025 contamination-clean holdout. Reward-weighted behavioral cloning performs the best, obtaining 46.2% indication F1 and 14.2% strict F1 against 25.0% and 2.1%, respectively, for the best-performing tool agent on each metric. These results suggest that structured offline learning can teach agents to plan clinical experiments.","authors":["William Bolton","Philip Torr"],"categories":["cs.AI","cs.LG"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-08-05","first_seen":"2026-08-05","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.03606","pdf_url":"https://arxiv.org/pdf/2608.03606","source_feed":"cs.LG","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["离线强化学习","临床试验决策","多智能体"],"reason":"纯多智能体决策优化，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-08-05T13:04:34","error":null,"has_summary":false,"summary":null},{"id":"2608.03153","version":1,"title":"Does the Gender Wage Gap Originate at Labor Market Entry? Evidence from South Korea","zh_title":"性别工资差距是否源于劳动力市场进入？来自韩国的证据","abstract":"When in the lifecycle does a large gender wage gap emerge? South Korea has the largest gender pay gap in the OECD, 29%. Among recent college graduates the conditional gap is only 4.3% over 2008-2019, falling from 5.0% to 3.0%, and correcting for differential selection into full-time wage employment with semiparametric, machine-learning, and bounds methods leaves it unchanged. Among observed workers what remains sits at the top of the distribution. In Korean panel data the corrected gap widens severalfold across the prime-age workforce, where the selection correction becomes first order. Korea's gender disparity is mostly generated after entry.","authors":["Dongwoo Kim"],"categories":["econ.GN","q-fin.EC"],"primary_category":"econ.GN","announce_type":"new","date":"2026-08-05","first_seen":"2026-08-05","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.03153","pdf_url":"https://arxiv.org/pdf/2608.03153","source_feed":"econ.GN","score":0,"bucket":"other","rubric_hits":["C5"],"tags":["性别工资差距","劳动经济学","韩国"],"reason":"研究真实性别工资差距，未使用LLM仿真人类被试，属于用人类数据训练或对齐模型的…","model":"deepseek-v4-pro","scored_at":"2026-08-05T13:04:29","error":null,"has_summary":false,"summary":null},{"id":"2608.02909","version":1,"title":"When Predictions Become Regressors: A Split-Sample Correction for Biases in Downstream Inference","zh_title":"当预测成为回归变量：下游推断偏差的分割样本校正","abstract":"Prediction-based methods, including Large Language Models (LLMs) and other machine learning techniques, are often used to construct measures of political phenomena that are difficult to quantify directly, such as policy positions in manifestos or emotions expressed on social media. In many applications, these prediction-generated measures are used as explanatory variables in regression models, even though they are measured with error. This leads to biased estimates. In this paper, we propose a simple solution to these biases: instrumental variables constructed from multiple measures created on independent splits of the original data. This approach is theoretically valid, easy to implement, and does not require new data. Through simulations, we show that this approach recovers estimates close to the true values, even in relatively small samples, while the standard approach can produce substantial bias in practice. We illustrate the method by revisiting two applications: whether gendered speech affects legislative outcomes in the German Parliament, and whether political risk influences poverty alleviation programs in China.","authors":["Nathan Canen","Ted Enamorado"],"categories":["econ.EM","stat.ML"],"primary_category":"econ.EM","announce_type":"new","date":"2026-08-05","first_seen":"2026-08-05","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.02909","pdf_url":"https://arxiv.org/pdf/2608.02909","source_feed":"econ.EM","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["测量误差","工具变量","政治学方法"],"reason":"论文关注预测值作为回归变量的测量误差校正，不涉及用LLM仿真人类被试或行为对照。","model":"deepseek-v4-pro","scored_at":"2026-08-05T13:04:23","error":null,"has_summary":false,"summary":null},{"id":"2608.03125","version":1,"title":"Performance appraisal promotes cooperation in spatial public goods games","zh_title":"绩效评估促进空间公共物品博弈中的合作","abstract":"In real organizations, performance appraisal serves as an important means of evaluating employees' work outcomes and behaviors, while performance pay directly links compensation to the evaluation results. However, in traditional spatial public goods games, the total payoff of a group is equally distributed among all participants without considering individual differences in contributions, an approach that ignores the heterogeneity of individual efforts. To overcome this limitation, we propose a spatial public goods game model based on performance appraisal, in which individual payoffs are divided into two parts: an equally distributed component and a performance-weighted component. Specifically, each individual receives scores from its neighboring groups, and the individual's reputation is dynamically updated by accumulating these scores, which further modulates the fitness function during strategy imitation. Extensive numerical simulation results demonstrate that the performance appraisal mechanism significantly promotes the emergence of cooperative behavior. The reputation reinforcement mechanism amplifies this positive effect by creating fitness advantages for high-reputation individuals. These findings suggest that incorporating performance appraisal into payoff allocation helps mitigate social dilemmas and promote collective cooperation.","authors":["Tianjiao Li","Qin Li","Kangxi Zhu","Minyu Feng","Manuel Chica"],"categories":["physics.soc-ph"],"primary_category":"physics.soc-ph","announce_type":"new","date":"2026-08-05","first_seen":"2026-08-05","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.03125","pdf_url":"https://arxiv.org/pdf/2608.03125","source_feed":"physics.soc-ph","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["演化博弈","合作涌现","空间公共物品博弈"],"reason":"纯多智能体协作研究，无LLM参与，不涉及人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-08-05T13:04:28","error":null,"has_summary":false,"summary":null},{"id":"2607.29334","version":2,"title":"The persuasive power of large language models does not depend on their perceived national origin","zh_title":"大语言模型的说服力不依赖于其感知的国家来源","abstract":"Conversational AI developed by geopolitical rivals reaches citizens worldwide, raising concerns that it could sway public opinion or be rejected as foreign propaganda, with consequences for democratic discourse and information sovereignty. Yet, whether an AI's perceived national origin shapes its persuasive power is unknown. In a preregistered randomized experiment, 403 adults from a nationally representative United States sample held a three-round debate with a chatbot introduced as either American (\"DiscoveryAI\") or Chinese (\"ZhengheAI\"), discussing a political or non-political topic. In all conditions, participants actually conversed with the same model (GPT-4o), instructed to argue against their initial position. We combined pre- and post-conversation self-reports of attitudes, trust, and collective narcissism with computational analyses of 1,209 participant turns, including LLM-coded stance and argumentative conduct, stance-sensitive embeddings, and keyword-masked emotion and toxicity classifiers. The conversations produced substantial attitude changes in every condition. Critically, the nationality label affected neither self-reported attitude change nor expressed stance, concessions, counterarguing, or affect, and equivalence tests and Bayes factors largely supported these null effects. The label's only reliable footprint was lower pre-conversation human-like trust in the Chinese model, whereas functionality trust was unaffected. Political topics slowed stance movement toward the AI's position, and collective narcissism predicted less attitude change regardless of origin, acting as a general barrier rather than an out-group filter. Users thus initially withhold social trust from a rival's AI yet still assimilate its arguments; origin labeling and transparency requirements alone may offer weak protection against foreign influence operations conducted through conversational AI.","authors":["Ningzhi Liu","Yannic Hinrichs","Jonas R. Kunst"],"categories":["cs.HC","cs.AI"],"primary_category":"cs.HC","announce_type":"replace-cross","date":"2026-08-04","first_seen":"2026-08-03","revised_at":"2026-08-04","abs_url":"https://arxiv.org/abs/2607.29334","pdf_url":"https://arxiv.org/pdf/2607.29334","source_feed":"cs.AI","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["LLM仿真","人类被试替代","说服实验"],"reason":"用LLM替代人类被试进行说服实验，有真实人类数据对照，评估仿真可靠性与失效条件…","model":"deepseek-v4-pro","scored_at":"2026-08-04T13:05:02","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-05","rank":1,"question":"AI 的感知国籍是否影响其说服力？","design":"用 GPT-4o 扮演美国或中国开发的聊天机器人，与美国代表性样本进行三轮辩论，处理为随机分配 AI 国籍标签和话题类型，测量态度变化、信任、对话中的立场与情感等。","baseline":"无对照","findings":"AI 国籍标签不影响自报态度变化和对话中的立场、让步、反驳或情感，仅降低对中国 AI 的社交信任；政治话题减缓立场转变，集体自恋普遍降低态度变化。","reliability":"论文未讨论","relevance":"该研究用 LLM 替代人类进行说服实验，系统评估了来源国标签对说服效果的影响，并揭示了仿真在政治话题和集体自恋下的失效条件，与研究者关注的人类仿真可靠性与偏差高度相关，值得精读。","inspiration":"借鉴其通过随机化 AI 身份标签和话题类型来分离来源效应与内容效应的设计，以及结合自报与对话文本计算分析的多维测量方法。｜可迁移到经济政策沟通场景，例如研究央行 AI 发言人的国籍标签是否影响公众通胀预期。｜以普通居民为被试，随机分配 AI 经济预测助手的国籍（本国 vs 外国），让其就未来通胀走势进行互动辩论，测量通胀预期变化和信任度，并以真实央行调查数据为对照。"}},{"id":"2608.01212","version":1,"title":"Do Humans Bargain Differently with AI? Evidence from Alternating-Offer Games","zh_title":"人类与AI的讨价还价行为不同吗？来自交替报价博弈的证据","abstract":"Artificial intelligence increasingly participates in economic interactions not only as a tool, but also as an autonomous bargaining counterpart negotiating on behalf of firms, platforms, and consumers. Yet little is known about how humans respond psychologically and strategically when bargaining with such agents in dynamic settings. We study this question in a laboratory experiment using a three-stage alternating-offer bargaining game in which participants negotiate in real time with either another human or a GPT-based AI agent. We also introduce a human-beneficiary condition in which the AI agent's earnings may affect another participant's payment. Agreements are not reached earlier in human-human bargaining than in human-AI bargaining, but they are reached significantly earlier when the AI's payoff affects another participant's payoff. Human proposers offer more to human opponents than to AI agents, whereas responders become significantly more willing to accept unfair AI offers when AI earnings may benefit another human. These findings suggest that fairness and reciprocity toward AI are weaker and more conditional than toward humans, but partially remerge when AI outcomes affect real people. The results have implications for the design of AI negotiation systems and broader human-AI economic interactions.","authors":["Yuhao Fu","Nobuyuki Hanaki","Haitao Wang"],"categories":["econ.GN","q-fin.EC"],"primary_category":"econ.GN","announce_type":"new","date":"2026-08-04","first_seen":"2026-08-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.01212","pdf_url":"https://arxiv.org/pdf/2608.01212","source_feed":"econ.GN","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2"],"tags":["LLM仿真","行为博弈","人机交互"],"reason":"用GPT代理进行议价博弈实验，与真人对照，评估公平与互惠行为差异，直接命中核心…","model":"deepseek-v4-pro","scored_at":"2026-08-04T13:04:28","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-04","rank":2,"question":"在动态交替报价博弈中，人类与基于GPT的AI代理议价时，其公平与互惠行为是否不同于人类间议价，且当AI收益关联人类受益人时行为是否改变？","design":"实验室实验，采用三阶段交替报价博弈，被试与另一真人或基于GPT的AI代理实时谈判；引入人类受益人条件，即AI收益可能影响另一被试报酬。结果变量为协议达成时间、提议者报价及响应者对不公平报价的接受意愿。","baseline":"人类-人类议价组作为对照基准。","findings":"人类提议者对真人对手报价高于对AI代理，但响应者在面对不公平AI报价时，若AI收益关联人类受益人则接受意愿显著提高；协议达成时间在人类-AI与人类-人类间无显著差异，但AI关联受益人时达成更快。","reliability":"论文未讨论","relevance":"该研究直接以GPT代理替代人类被试进行议价博弈实验，并与真人对照，评估公平与互惠行为差异，完全命中研究者关注的LLM仿真人类行为及可靠性评估，值得精读原文。","inspiration":"借鉴其通过引入受益人条件来分离社会偏好与纯粹策略行为的处理设计，以及区分提议者与响应者角色的不对称分析框架｜可迁移至信贷审批歧视研究，探究当AI审批决策关联人类信贷员利益时，申请人对不公平拒绝的接受度是否变化｜以真实信贷申请者为被试，处理为AI审批vs.人类审批，并设置AI收益关联信贷员奖金的条件，结果变量为申请人对拒绝决定的公平感知与申诉意愿，对照真实信贷审批数据中的申诉率。"}},{"id":"2608.01607","version":1,"title":"AI Financial Advice: Supply, Demand, and Life Cycle Implications","zh_title":"人工智能财务建议：供给、需求与生命周期影响","abstract":"We ask a representative sample to write prompts seeking spending and investing advice from LLMs, then simulate the lifetime effects of following the advice under realistic asset and labor market conditions. Applying this method to GPT-5.2, we find following the advice would move respondents toward life cycle theory: broader participation in diversified equity funds, age-declining equity shares, and larger savings buffers. Recommendations vary systematically by gender, prior AI experience, and financial literacy. For gender, two-thirds of recommended equity-share differences arise from men and women writing different prompts (demand), while one-third arise from gender labels attached to otherwise identical prompts (supply).","authors":["Taha Choukhmane","Tim de Silva","Weidong Lin","Matthew Akuzawa"],"categories":["econ.GN","q-fin.EC","q-fin.GN","q-fin.PM"],"primary_category":"econ.GN","announce_type":"new","date":"2026-08-04","first_seen":"2026-08-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.01607","pdf_url":"https://arxiv.org/pdf/2608.01607","source_feed":"econ.GN","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2","B4"],"tags":["LLM仿真","财务决策","人类行为对照"],"reason":"用LLM模拟人类财务决策，有真实人类样本对照，涉及生命周期投资行为和政策评估，…","model":"deepseek-v4-pro","scored_at":"2026-08-04T13:04:30","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-04","rank":3,"question":"当人们向大语言模型寻求财务建议时，AI建议如何影响生命周期投资行为，以及这种影响在供给端和需求端如何因性别、AI经验和金融素养而异？","design":"让代表性样本撰写向LLM寻求消费与投资建议的提示词，然后用GPT-5.2生成建议，并在现实的资产和劳动力市场条件下模拟终生遵循该建议的效果。","baseline":"代表性人类样本撰写的提示词及其对应的真实行为特征，作为需求端基准；通过附加性别标签的相同提示词分离供给端差异。","findings":"遵循GPT-5.2建议会使受访者更接近生命周期理论：更广泛参与多元化股票基金、随年龄降低股票份额、建立更大储蓄缓冲。建议因性别、AI经验和金融素养而系统性地不同，其中三分之二的性别股票份额差异源于男女撰写不同提示词（需求端），三分之一源于相同提示词附加性别标签（供给端）。","reliability":"论文未讨论","relevance":"该研究直接用LLM模拟人类财务决策，有真实人类样本对照，并分解了需求端与供给端的差异，与您关注的LLM仿真可靠性及偏差来源高度相关，值得精读原文。","inspiration":"借鉴其通过控制提示词内容与附加身份标签来分离需求端与供给端效应的方法，可用于研究AI建议中的歧视或偏差来源。｜可迁移到信贷审批歧视研究，分析AI建议中的性别或种族偏差。｜以代表性人群为被试，让其撰写贷款申请提示词，处理为在相同提示词上附加不同性别/种族标签，结果变量为AI批准的贷款额度与利率，对照真实信贷审批数据中的群体差异。"}},{"id":"2608.01204","version":1,"title":"ShiJianBench: From Dialogue to Decision for Long-Horizon Evaluation of Investment Advisors","zh_title":"ShiJianBench：从对话到决策的长期投资顾问评估","abstract":"Conversational investment advisors influence not only what users know, but also how they make subsequent decisions as market conditions evolve. Existing evaluations primarily assess response quality or observed outcomes, leaving the long-horizon pathway from advisor language to investor behavior difficult to audit. We introduce ShiJianBench, an offline framework for evaluating conversational investment advisors through matched investor trajectories under fixed historical market feedback. At its core is a multi-agent investor simulator with explicit evolving state variables, motive-driven deliberation, long-term memory, and dialogue-grounded updates. The simulator is calibrated against aggregate behavioral patterns from 7,199 real users, and advisor policies are evaluated using separate investor-side, service-side, and content-side metrics under a hard compliance gate. Experiments on Chinese fund-market traces from 2021 to 2026 identify a stable leading group of LLM advisors that combines substantially stronger personalized content with competitive investor-side trajectory outcomes. These results reveal a systematic distinction between producing a high-quality response and delivering an effective long-horizon intervention, motivating trajectory-aware evaluation of conversational advisors.","authors":["Jie Gong","Maowei Jiang","Zhiwei Liu","Yang Qiao","Wenxi Wu","Mengxi Xiao","Enze Zhang","Ziyan Kuang","Yankai Chen","Caishuang Huang","Meng Zhou","Xiku Du","Xue Liu","Guojun Xiong","Min Peng","Qianqian Xie","Sophia Ananiadou"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-04","first_seen":"2026-08-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.01204","pdf_url":"https://arxiv.org/pdf/2608.01204","source_feed":"cs.CL","score":8,"bucket":"selected","rubric_hits":["A3","B1","B2"],"tags":["LLM仿真","投资者行为","人类数据校准"],"reason":"用多智能体模拟投资者行为并与真实用户数据校准，涉及金融决策仿真和人类对照。","model":"deepseek-v4-pro","scored_at":"2026-08-04T13:04:27","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-04","rank":7,"question":"如何评估对话式投资顾问通过改变投资者内部状态而产生的长期决策轨迹，而不仅仅是回复质量？","design":"构建一个多智能体投资者模拟器，包含显式的演化状态变量、动机驱动的子智能体协商、长期记忆和对话更新机制；在固定的历史市场轨迹下，对同一投资者初始化状态分别运行基线条件和目标顾问条件，生成匹配的反事实轨迹，并测量投资结果、风险控制、商业价值和对话质量等多维度指标。","baseline":"使用来自7,199名真实用户的聚合行为模式对模拟器进行校准，整体对齐分数达到0.88。","findings":"准确、个性化和合规的回复并不必然带来成比例更强的长期投资者结果；通过追踪对话、演化状态、市场反馈和后续决策，揭示了高质量回复与有效长期干预之间的系统性差异。","reliability":"论文未讨论","relevance":"该研究直接以LLM模拟投资者行为，并与真实用户数据校准，评估长期决策轨迹，属于经济学实验仿真，且包含批判性发现，值得精读原文。","inspiration":"借鉴其通过显式状态变量和动机驱动子智能体来模拟决策过程，以及利用匹配反事实轨迹进行因果评估的方法。｜可迁移到政策公告对投资者预期形成与资产配置的长期影响评估场景。｜以LLM模拟散户投资者，处理为不同措辞的政策公告，结果变量为持仓调整和风险偏好变化，用真实市场交易数据校准并对照。"}},{"id":"2608.01458","version":1,"title":"PALMs: Using Multi Construct-Grounded Rationales for Modeling Population Preferences in LLMs","zh_title":"PALMs：使用多构念基础理由建模大语言模型中的人口偏好","abstract":"Large language models are being extensively used to simulate individual user behavior, yet faithfully representing a population requires capturing the systematic variation in values, beliefs, and cultural norms that distinguish one group from another. We introduce Population Aligned Language Models (PALMs), a suite of models each aligned to specific populations, covering five countries: USA, India, Brazil, France and Italy. PALMs are created by synthesizing rationales grounded in psychological and cultural constructs and using these as latent supervision during preference tuning for population-specific alignment. Evaluated across four dimensions: personality, values and beliefs, cultural norms, and morality, PALMs consistently outperform baselines, including culture-specialized models, achieving an average of 8.59% relative improvement over the best baseline across all five populations. Notably, construct-grounded rationales outperform both demographic prompting and survey-based fine-tuning, suggesting that grounding preference learning in psychology and culture provides a richer inductive signal than surface-level response distributions. We further demonstrate strong generalization to downstream applications with- out task-specific supervision: outperforming best baselines by 5.19% in personalized reward modeling, 6.34% in population simulation, and showing strong transfer to social reasoning tasks. Datasets and code are available at: https://github.com/limenlp/PALMs.","authors":["Priyanka Dey","Brihi Joshi","Preyashi Poddar","Jieyu Zhao","Emilio Ferrara"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-04","first_seen":"2026-08-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.01458","pdf_url":"https://arxiv.org/pdf/2608.01458","source_feed":"cs.CL","score":8,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["人口仿真","文化对齐","偏好建模"],"reason":"用LLM模拟不同国家人群偏好，有人类调查数据对照，涉及人口仿真和个性化奖励建模…","model":"deepseek-v4-pro","scored_at":"2026-08-04T13:04:29","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-04","rank":8,"question":"如何利用心理学和文化构念的推理依据来对齐大语言模型，以更忠实地模拟不同国家人群的偏好？","design":"使用基于 Llama-3.1-8B-Instruct 的模型，通过合成基于五类心理学和文化构念（人格特质、文化维度、人类价值观、道德基础、世界信念）的推理依据，作为偏好优化（DPO）的潜在监督信号，训练出分别对齐美国、印度、巴西、法国、意大利五国人群的 PALMs 模型；在推理时模型先生成推理依据再输出回答，评估其在人格、价值观与信念、文化规范、道德四个维度上与真实人群分布的吻合度。","baseline":"使用来自世界价值观调查（WVS）、Hofstede 文化维度调查、Schwartz 价值观调查、道德基础问卷等跨国调查的真实人类数据作为对照基准。","findings":"PALMs 在五个国家、四个对齐维度上均优于人口统计提示和基于调查数据微调的基线，相对最佳基线平均提升 8.59%；在个性化奖励建模、人群模拟和社会推理等下游任务上无需任务特定监督即展现出强泛化能力，分别提升 5.19% 和 6.34%。","reliability":"论文未讨论","relevance":"该研究直接使用 LLM 模拟不同国家人群的偏好和价值观，并与真实跨国调查数据进行严格对照，评估了仿真在人格、文化规范、道德等维度上的可靠性，高度契合研究者对 LLM 人类仿真实验的关注，值得精读原文。","inspiration":"借鉴其利用心理学和文化构念生成推理依据来指导模型对齐的方法，可提升经济决策仿真的内在一致性，而非仅拟合表面行为分布。｜可迁移至跨文化消费偏好实验或跨国投资风险偏好研究，例如模拟不同国家投资者对风险资产配置的差异。｜以各国真实家庭金融调查数据为基准，用 LLM 扮演不同国家居民，处理为注入基于 Hofstede 文化维度和 OCEAN 人格的推理依据进行偏好对齐，结果变量为风险资产选择比例，对照真实调查中的资产配置分布。"}},{"id":"2608.01629","version":1,"title":"Human-LLM Alignment in Language Attitudes Toward Non-Native Japanese","zh_title":"人类与LLM对非母语日语语言态度的一致性","abstract":"Large language models (LLMs) increasingly evaluate human writing in high-stakes domains such as hiring and academic assessment, putting non-native speakers at particular risk. Drawing on the language attitudes framework, we compared human and LLM evaluations of parallel L1- and L2-written Japanese emails on three dimensions: fluency, status, and solidarity. Japanese raters rated L2 texts significantly lower on all three dimensions, with a fluency gap roughly twice the size of the status and solidarity gaps. Six LLM judges reproduced the direction of this bias, and five reproduced its ordering across dimensions. The models diverged from humans in two ways: all understated the solidarity gap, the most socially grounded dimension, and all differentiated among learner L1 backgrounds where humans did not. LLM judges thus reproduce native speakers' language attitudes in a structured yet attenuated form, and the language attitudes framework offers a ready-made yardstick for auditing them beyond English.","authors":["Naho Orita","Hayato Ogawa","Daisuke Kawahara"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-04","first_seen":"2026-08-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.01629","pdf_url":"https://arxiv.org/pdf/2608.01629","source_feed":"cs.CL","score":8,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["语言态度","人类仿真","偏差审计"],"reason":"用LLM复现人类对非母语写作的态度偏差，并与真实人类评分对照，揭示仿真衰减与失…","model":"deepseek-v4-pro","scored_at":"2026-08-04T13:04:30","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-04","rank":9,"question":"LLM在评估非母语日语写作时，是否会复现人类母语者的语言态度偏差（在流利度、地位、团结三个维度上）？","design":"用六款LLM（GPT-5.4、GPT-4o-mini、Claude Sonnet 4.5等）作为评委，对同一批L1日语母语者和L2日语学习者撰写的平行邮件进行评分，测量流利度、地位、团结三个维度的评分差异，并与人类评分者结果对比。","baseline":"通过众包招募1536名日语母语者，对相同邮件进行三个维度的评分，形成人类语言态度基准。","findings":"LLM复现了人类对L2写作评分更低的方向性偏差，且五个模型复现了偏差维度排序（流利度>地位>团结）；但所有模型都低估了团结维度的差距，且区分了人类未区分的L2学习者的母语背景。","reliability":"论文指出LLM在团结维度上低估偏差，且会引入人类没有的基于L1背景的区分，表明仿真存在衰减和失真；研究限于日语邮件场景，未涉及其他语言或文体。","relevance":"该研究直接对比LLM与人类在语言态度上的偏差，揭示了仿真在方向一致但程度衰减、且会引入额外区分模式的现象，对关注LLM仿真可靠性及失效条件的研究者有重要参考价值。","inspiration":"借鉴其使用平行文本控制内容、多维度评分量表测量隐性偏差的方法，可迁移到信贷审批或招聘中的语言偏见研究。｜可应用于信贷审批歧视场景，研究贷款申请书中非母语写作对审批决策的影响。｜以银行信贷员为人类被试，LLM为仿真被试，处理为申请书语言（母语vs非母语），结果变量为信用评分和批准率，对照真实信贷审批数据中的语言偏差。"}},{"id":"2608.00979","version":1,"title":"Passing Coarse Marginal Checks Can Be Cheap: Persona Mixtures and Imprecise Treatment-Response Estimates in an LLM Persona Panel","zh_title":"通过粗粒度边际检查可能很廉价：LLM角色面板中的角色混合与不精确的处理效应估计","abstract":"Large language models are increasingly used as synthetic research participants and are often validated by whether their marginal responses resemble human data. We study a fixed panel of sixteen lightweight persona-conditioned GPT-4.1 configurations in repeated strategic games. The panel met preregistered broad-reference condition-mean criteria in three of four repeated-game cells; the sole miss was 0.011 below the lower reference bound. Variation was strongly prompt-indexed, but its share depended on uncertainty assumptions: fixed-panel symmetric-Dirichlet sensitivities produced median between-prompt shares of 63%-71% under Jeffreys alpha=0.5 and 47%-53% under alpha=1, while finite-opportunity plug-in estimates were 85%-96%. Aggregate continuation-probability contrasts were +0.083 and +0.078, with conservative simultaneous 95% intervals [-0.171, +0.330] and [-0.181, +0.330]. The treatment jointly changed the continuation process and its textual representation. A separate wording-and-position operation shifted cooperation from 0/40 to 37/40 in the bare configuration, and a label conflict also revealed representation control. The original persona-level p13 result was not prospectively family-controlled, while a post-adjudication exact gate was structurally underpowered; p13 is therefore a replication target rather than a finding. External review exposed family-error, dependence, construct, and boundary-uncertainty defects, and zero-call reanalysis changed the interpretation without rewriting the historical record. The registered marginal criteria could be passed without precisely estimating the treatment-response object. A public capsule verifies 4,916 confirmatory Phase 3-5 runs with no live model calls. The results concern one fixed model-prompt panel and do not establish human substitutability.","authors":["Yohei Nakajima"],"categories":["cs.AI","cs.CL","cs.GT"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-08-04","first_seen":"2026-08-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.00979","pdf_url":"https://arxiv.org/pdf/2608.00979","source_feed":"cs.CL","score":8,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["LLM仿真","行为博弈","算法保真度"],"reason":"用LLM persona面板模拟重复博弈行为，与人类数据对照，评估仿真可靠性和…","model":"deepseek-v4-pro","scored_at":"2026-08-04T13:04:25","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-05","rank":5,"question":"在重复策略博弈中，固定的人格面板能否通过粗粒度的边际分布检验，同时其处理效应估计是否精确？","design":"使用16个轻量级人格条件化的GPT-4.1配置构成固定面板，在重复策略博弈中施加两种处理（S2措辞存在与否），测量合作行为、继续概率等结果变量，并分析面板内变异来源。","baseline":"无对照","findings":"面板在四个重复博弈单元中有三个通过了预注册的边际条件均值检验，但总体继续概率的处理效应点估计较小且置信区间很宽，无法确立等效性或窄响应边界。变异主要来源于提示词配置之间，且措辞和标签操作可大幅改变合作率，表明粗粒度边际检验可能掩盖处理效应估计的不精确性。","reliability":"论文承认边际检验通过并不要求精确估计处理效应，且结果仅针对一个固定模型-提示面板，不建立人类可替代性；外部审查揭示了家族误差、依赖性、构念和边界不确定性缺陷。","relevance":"该研究直接评估了LLM人格面板在策略互动中的仿真可靠性，揭示了粗粒度验证的廉价性，对关注经济学实验和政策评估中LLM替代人类被试的研究者具有重要警示价值，值得精读原文。","inspiration":"借鉴其预注册、固定面板、多不确定性视角分解变异和零调用复现的审计协议设计｜可迁移到公共品博弈或信任博弈中，检验LLM面板能否复现人类合作与惩罚行为的分布和处理效应｜用多个LLM人格配置构成固定面板，施加不同的制度处理（如惩罚机制、信息反馈），测量合作率与信念更新，以实验室人类被试数据为基准，评估边际分布通过但处理效应估计不精确的程度。"}},{"id":"2608.01193","version":1,"title":"Humans Are More Diverse: Frontier LLMs Show Extreme Policies in Idealised AI Development Races","zh_title":"人类更多样：前沿大语言模型在理想化AI发展竞赛中表现出极端策略","abstract":"An AI development race creates a multi-agent safety dilemma. Each company can develop slowly and safely, or move faster while taking a risk that may remove its final reward. We use this repeated game to study strategic safety behaviour among large language model (LLM) agents in races with two to five players. However, a valid action does not show that an agent understands the game. We therefore place an audit gate before behavioural interpretation. We first verify the game engine, then test rule recall, state tracking, payoff calculation, and stability under different but equivalent task descriptions. We then compare LLM action sequences with an evolutionary game-theory benchmark and published human data, and explore differences across models, risk conditions, personas, and two- to five-player races. The audit shows that strong rule recall can coexist with weak state tracking and expected-payoff calculation. Providing verified arithmetic and changing the response representation can also change later actions, even when the game rules stay fixed. Across seven tested model endpoints, aggregate rates hide large differences in action sequences, responses to opponents, and responses to race position. Patterns across the tested three- to five-player races are also model-specific rather than a single effect of adding competitors. These results show why multi-agent AI-race simulations need validity checks and trajectory-level analysis before their outputs are described as strategic, human-like, or safety-aware. Our findings are exploratory and apply only to the tested models, prompts, and decoding settings.","authors":["Phu Hoa Pham","Duy Minh Dao Sy","Trung Kiet Huynh","Phu Quy Nguyen Lam","Chi Nguyen Tran","Minh Trung Le","Phong Hao Le","Dinh Nam Nguyen","Thien Ky Nguyen Dong","Elias Fernandez Domingos","Le Hong Trang","The Anh Han"],"categories":["cs.AI","cs.CY","cs.GT","cs.LG","cs.MA"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-04","first_seen":"2026-08-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.01193","pdf_url":"https://arxiv.org/pdf/2608.01193","source_feed":"cs.AI","score":8,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM行为仿真","博弈实验","人类数据对照"],"reason":"用LLM模拟AI竞赛中的人类行为，并与真实人类数据对照，涉及博弈实验场景，但非…","model":"deepseek-v4-pro","scored_at":"2026-08-04T13:04:25","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-05","rank":6,"question":"在理想化的AI研发竞赛中，LLM智能体是否表现出与人类相似的策略性安全行为，以及其行为有效性是否受任务表述、状态追踪和收益计算能力的影响？","design":"让七个LLM端点扮演AI公司，在2至5人的重复博弈中每轮选择安全或冒险开发，通过审计关卡检验规则记忆、状态追踪、收益计算和表述稳定性，并记录行动序列、对手反应和位置效应。","baseline":"已发表的人类实验数据（falling_behind_unsafe）和演化博弈论基准。","findings":"LLM的总体不安全率掩盖了行动序列、对手反应和位置效应的巨大差异；强规则记忆可与弱状态追踪和收益计算并存，且提供算术验证或改变响应表示会改变后续行动。","reliability":"论文声明发现是探索性的，仅适用于所测试的模型、提示和解码设置，未讨论其他失效条件。","relevance":"该研究直接使用LLM模拟人类在博弈中的行为，并与真实人类数据对照，包含审计关卡检验仿真有效性，对关注LLM仿真可靠性和偏差的研究者具有重要参考价值，值得阅读原文。","inspiration":"借鉴其审计关卡设计，在行为解释前先检验LLM对任务规则、状态和收益的理解，并考察不同表述和输出格式的稳健性。｜可迁移到政策公告预期形成的实验，如央行沟通博弈，检验LLM是否像人类一样对措辞和顺序敏感。｜以LLM为被试，模拟央行发布前瞻指引，处理为不同措辞或发布顺序，结果变量为通胀预期和投资决策，对照真实人类实验数据。"}},{"id":"2607.27553","version":2,"title":"AI and Its Impact on Creativity and Diversity: An Empirical Study of LLM-Generated Product Ideas","zh_title":"AI对创造力与多样性的影响：LLM生成产品创意的实证研究","abstract":"This research examines how well large language models, or LLMs, generate new product ideas for college students priced under $50. Across a series of studies, we identify key strengths and weaknesses of using LLMs for product innovation. Our first study shows that LLM-generated product ideas have higher average quality than human ideas, based on purchase intent, and are 7 times more likely to rank in the top 10%. Our second study shows that this AI-induced creativity boost is not explained by the LLM's more persuasive pitching skills. Our third and fourth studies identify a weakness of using LLMs for brainstorming: AI-generated ideas are less novel at the idea level and less diverse at the set level. In our fifth study, we analyze prior LLM-based creativity studies and find consistently lower idea diversity across all of them, demonstrating the generalizability of these findings. Our sixth and seventh studies investigate techniques to mitigate this diversity loss. We compare LLMs from different vendors and versions and find that more recent models generate more diverse ideas, though they still fall short of human-level diversity. We also demonstrate techniques that increase idea diversity almost to the level of human idea generation: pooling ideas across vendors; prompt engineering, including Chain-of-Thought prompting and injecting heterogeneous personas or constraints; and creative agents that broadly explore the solution landscape to restore diversity. Finally, in our eighth study, we show that exploiting the near-zero marginal cost of AI idea generation by scaling the number of ideas steadily improves coverage of the idea space, approaching human-level coverage. We conclude by presenting actionable recommendations for innovation managers who want to identify better new product ideas with the help of LLMs.","authors":["Christian Terwiesch","Lennart Meincke","Karan Girotra","Ethan Mollick","Gideon Nave","Karl T. Ulrich"],"categories":["cs.AI","cs.CL","econ.GN","q-fin.EC"],"primary_category":"cs.AI","announce_type":"replace-cross","date":"2026-08-04","first_seen":"2026-07-31","revised_at":"2026-08-04","abs_url":"https://arxiv.org/abs/2607.27553","pdf_url":"https://arxiv.org/pdf/2607.27553","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A1","B1","B2"],"tags":["LLM仿真","人类对照","产品创新"],"reason":"用LLM生成产品创意并与人类数据对照，属于人类仿真实验，涉及经济学场景。","model":"deepseek-v4-pro","scored_at":"2026-08-04T13:05:01","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-05","rank":9,"question":"LLM生成的产品创意在质量、新颖性和多样性上是否优于人类创意？","design":"使用GPT-4以零样本和少样本提示生成面向大学生、定价50美元以下的新产品创意，与人类学生生成的创意进行对比，通过购买意向测量质量，通过文本挖掘和人工评分测量新颖性与多样性。","baseline":"一所精英大学产品设计课程学生在LLM出现前生成的创意池。","findings":"LLM生成创意的平均购买意向更高，且进入前10%的可能性是人类的7倍，但AI创意在想法层面新颖性更低，在集合层面多样性更差。","reliability":"论文未讨论","relevance":"该研究将LLM作为人类被试的替代品，直接对比真实人类数据，评估了AI在创意生成任务中的表现与偏差，属于人类仿真实验，值得精读。","inspiration":"借鉴其将LLM作为被试、设置零样本/少样本处理组并与真实人类基准对比的实验设计，以及用购买意向、新颖性、多样性等多维度测量结果的方法。｜可迁移到消费者偏好预测或新产品市场反应评估等经济金融场景，例如用LLM模拟消费者对金融产品创新的接受度。｜以LLM模拟消费者群体，处理为不同提示词（如不同收益描述），结果变量为购买意向，对照真实消费者调查数据。"}},{"id":"2608.01181","version":1,"title":"Talking to Digital Twins: Selective Disclosure and Belief Measurement in Financial Social Media","zh_title":"与数字孪生对话：金融社交媒体中的选择性披露与信念测量","abstract":"Social media affect financial markets, but public posts by financial media personas are voluntary disclosures. What is not disclosed is therefore usually unobserved. We address this measurement problem by conducting repeated, real-time interviews of \"digital twins\" built from monitored finfluencers' X accounts under a fixed protocol. The interviews recover stock-level public-persona belief proxies even when no public recommendation is made. Because the interviews are generated and archived before the relevant return windows, the design avoids the look-ahead bias that arises when LLMs are queried ex post. The evidence shows that information obtained from these digital-twin interviews predicts the cross section of large-cap stock returns in the expected direction. Repeated real-time interviews therefore show how selective disclosure can be turned into measurable panels of market views.","authors":["Boone Bowles","Raymond Duch","Sorin Sorescu"],"categories":["econ.GN","cs.AI","q-fin.EC"],"primary_category":"econ.GN","announce_type":"cross","date":"2026-08-04","first_seen":"2026-08-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.01181","pdf_url":"https://arxiv.org/pdf/2608.01181","source_feed":"cs.AI","score":7,"bucket":"pending","rubric_hits":["A1","B1","B2"],"tags":["LLM仿真","金融行为","数字孪生"],"reason":"用LLM构建数字孪生模拟金融影响者观点，并与真实市场数据对照，属于人类仿真且涉…","model":"deepseek-v4-pro","scored_at":"2026-08-04T13:04:25","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-05","rank":10,"question":"如何从金融影响者的选择性披露中恢复其市场观点，并检验这些观点对股票收益的预测能力？","design":"基于金融影响者（finfluencers）的X账号公开信息构建LLM数字孪生，每日通过固定访谈协议询问其对大盘股的投资建议（买入/持有/卖出）及信心、投机性和催化剂，生成带时间戳的面板数据，并与后续股票收益对照。","baseline":"同一日期、同一股票、同一账号上人类金融影响者的实际公开推荐，以及数字孪生推荐与人类推荐的一致性；此外，数字孪生访谈在人类公开推荐之前就已偏向最终公开的方向。","findings":"数字孪生访谈的净买入份额能正向预测未来10个交易日的股票超额收益，且这种预测能力在人类未公开推荐的“沉默”股票上最强。访谈不仅复现了公开推荐，还从沉默中提取了增量信息。","reliability":"论文未讨论","relevance":"该研究用LLM构建数字孪生模拟金融影响者观点，并与真实人类推荐及市场收益数据对照，直接命中你关注的人类仿真、经济学实验场景和基准对照，值得精读原文以了解其仿真效度和预测设计。","inspiration":"借鉴其通过固定访谈协议从LLM数字孪生中系统提取信念并分离方向、分歧和不确定性的测量方法。｜可迁移到资产定价实验，检验分析师或投资者情绪对横截面收益的预测。｜以LLM扮演的金融分析师为被试，每日施加固定访谈处理（询问对股票的看法和信心），结果变量为净买入份额，用真实分析师一致预期和后续股票收益做对照。"}},{"id":"2608.01540","version":1,"title":"Do people rely on ChatGPT more than their peers to detect deepfake news?","zh_title":"人们在检测深度伪造新闻时是否比同伴更依赖ChatGPT？","abstract":"This experimental study investigates how people rely on different sources of advice when detecting AI-generated fake news (deepfake news). In a laboratory deepfake detection task, student participants identified the proportion of human-written (non-AI-generated) content in synthetic deepfake news articles and received advice from ChatGPT (GPT-4), human peers, or linguistic experts. The results show that participants rely more on ChatGPT than on human peers when detecting GPT-2-generated deepfake news. Participants also rely more on linguistic experts than on peers, while the relative reliance on experts versus ChatGPT is mixed across experimental waves, potentially reflecting time trends in beliefs about AI-based detection. Importantly, in the additional experiment conducted in 2025 under the same experimental procedure, participants relied more on linguistic experts than on ChatGPT. Moreover, performance improvements reflect the joint role of reliance and advice quality, arising primarily when participants rely on high-quality advice. Overall, relying on AI to detect AI-generated deepfakes can improve detection outcomes, but only when AI-based detection tools are of sufficiently high quality. These findings highlight the dual role of GAI as both a source of deepfakes and a tool for mitigating related risks.","authors":["Yuhao Fu","Nobuyuki Hanaki"],"categories":["econ.GN","cs.CY","q-fin.EC"],"primary_category":"econ.GN","announce_type":"cross","date":"2026-08-04","first_seen":"2026-08-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.01540","pdf_url":"https://arxiv.org/pdf/2608.01540","source_feed":"cs.CY","score":7,"bucket":"pending","rubric_hits":["A1","B1","B2"],"tags":["人类行为实验","AI建议依赖","深度伪造检测"],"reason":"用真人实验对照，研究人类对ChatGPT建议的依赖，涉及行为决策和检测任务，可…","model":"deepseek-v4-pro","scored_at":"2026-08-04T13:04:30","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-04","rank":12,"question":"人们在检测深度伪造新闻时，是否比依赖人类同伴更依赖ChatGPT的建议？","design":"本研究并非LLM仿真实验，而是真实人类实验室实验。被试为大学生，在30轮深度伪造检测任务中，先给出初始判断，然后随机接受ChatGPT（GPT-4）、人类同伴或语言学专家的建议，再给出最终判断，测量建议采纳程度（WOA）和检测表现。","baseline":"真实人类同伴的建议作为对照基准，比较被试对ChatGPT与对人类同伴的建议依赖程度。","findings":"被试在检测GPT-2生成的深度伪造新闻时，对ChatGPT的建议依赖程度显著高于对人类同伴；但2025年补充实验中，被试对语言学专家的依赖高于ChatGPT，反映出对AI检测工具信念的时间趋势变化。检测表现的提升取决于建议质量与被试依赖程度的共同作用，仅当AI检测工具质量足够高时，依赖AI才能改善检测结果。","reliability":"论文指出，依赖AI检测AI生成内容的效果取决于AI检测工具的实际质量，若工具质量不高，依赖AI可能无益甚至有害；此外，被试对AI的信任和依赖可能随时间变化，影响结论的跨期稳健性。","relevance":"该研究虽非LLM仿真实验，但提供了真实人类在AI建议下的行为决策基准，可用于校准或验证LLM仿真人类在信息检测任务中的行为，尤其适合关注AI依赖与信任动态的研究者。","inspiration":"可借鉴其JAS框架和WOA测量方法，通过随机分配建议来源（AI vs. 人类）并比较依赖程度，来量化人类对AI建议的采纳行为。｜可迁移至金融投资决策场景，如投资者在评估AI生成的财务报告或市场预测时，是否过度依赖AI建议而忽视人类分析师。｜以真实投资者为被试，设计投资判断任务，处理为提供ChatGPT生成的投资建议 vs. 人类分析师建议，结果变量为建议采纳权重和投资组合表现，对照真实市场数据或历史分析师记录。"}},{"id":"2607.25953","version":2,"title":"Polistemics: Evaluating LLMs as Information Mediators in Politics & Elections","zh_title":"Polistemics：评估大语言模型在政治与选举中作为信息中介的表现","abstract":"As LLMs increasingly shape the political information citizens rely on, no standard exists to assess whether they do so responsibly. We introduce Polistemics, a theory-grounded diagnostic benchmark for evaluating LLMs as mediators of political information in elections. Prior work has treated this task as reproduction rather than mediation, leaving its epistemic dimensions and interaction with imperfect information unaddressed. We ground the evaluation in Epistemic Modesty, a normative standard derived from citizens' epistemic agency, and test it across controlled settings that vary the clarity, noise, and consistency of the available evidence. Applying the benchmark to three state-of-the-art LLMs across the 2025 German and Dutch elections, we find that high aggregate scores mask systematic failures. Models mediate reliably under clear evidence but break down when it is absent, vague, or contradictory, while flattening the intensity of political language throughout. These failures point to party priors, shifting with party labels and output language. Reliable mediation appears achievable, but no model delivers it consistently.","authors":["Baran Peters","Gabor Hollbeck","Robert Jakob","Kevin O'Sullivan"],"categories":["cs.CL","cs.CY"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-08-04","first_seen":"2026-07-29","revised_at":"2026-08-04","abs_url":"https://arxiv.org/abs/2607.25953","pdf_url":"https://arxiv.org/pdf/2607.25953","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["LLM评估","政治信息","基准测试"],"reason":"评估LLM作为政治信息中介，测量模型行为而非仿真人类被试，无人类对照数据。","model":"deepseek-v4-pro","scored_at":"2026-08-04T13:05:01","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-29","rank":16,"question":"在选举情境中，LLM作为政治信息中介者应具备哪些负责任的行为特质，以及它们在不同信息环境下的稳健性如何？","design":"本研究并非人类仿真实验，而是构建了一个名为Polistemics的基准测试，用于评估LLM在选举中作为政治信息中介者的表现。它使用三个最先进的LLM（Qwen3.6 Flash、GPT-5.4、Claude Sonnet 4.6），在受控的信息环境下，通过改变证据的清晰度、噪声和一致性，测量模型在回答政党立场查询时的忠实性、公正性和认知校准。","baseline":"无对照","findings":"模型在证据清晰时中介可靠，但在信息缺失、模糊或矛盾时表现崩溃，并会扁平化政治语言的强度。这些失败可能由模型对政党的先验偏好驱动，并受政党标签和输出语言影响。","reliability":"论文指出模型在缺乏、模糊或矛盾的信息下会失效，且高总体得分掩盖了系统性失败；可靠的中介看似可实现，但没有模型能持续做到。","relevance":"该研究虽非直接的人类仿真实验，但系统评估了LLM在政治信息中介中的偏差与失效条件，对理解LLM在调查或实验中的行为偏差具有参考价值，值得阅读以了解其受控信息环境的设计方法。","inspiration":"可借鉴其通过合成不同信息环境（如清晰、噪声、矛盾）来隔离LLM行为影响因素的实验设计方法。｜可迁移到经济政策沟通场景，如央行公告的解读实验，测试LLM在不同信息质量下如何传达政策立场。｜以LLM为被试，向其提供不同清晰度的央行声明（处理），测量其输出的政策立场一致性和置信度（结果变量），并以真实市场分析师解读数据作为对照。"}},{"id":"2607.25166","version":3,"title":"Individual-level interventions against sycophantic AI reduce its appeal but not its persuasiveness","zh_title":"针对谄媚AI的个体层面干预降低其吸引力但未降低其说服力","abstract":"AI chatbots can be \"sycophantic,\" or overly agreeable and flattering toward users. Sycophantic AI has been shown to entrench attitudes, yet users frequently fail to recognize it (a phenomenon we call \"sycophancy blindness\"). We tested whether increasing users' awareness of sycophancy protects them from its harmful effects in two preregistered experiments (n = 1,590). In the first, participants received a brief written warning about sycophancy before conversing with a sycophantic chatbot. In the second, participants watched a video of a sycophantic AI validating several other users, including users on opposite sides of the same conflict, before interacting with it themselves. Both interventions changed how participants evaluated the AI. The warning reduced the AI's perceived objectivity, and the video reduced enjoyment of the AI --- an effect mediated by the reduced belief that its validation was uniquely earned. We then pooled our experiments with two prior studies of sycophancy awareness interventions (six interventions total, n = 3,982). The pattern across experiments was consistent: while the interventions made the sycophantic AI appear less objective and trustworthy, none reduced its persuasiveness. These results suggest that individual-level interventions, such as warning labels or AI literacy, may not be enough to protect users from AI harms.","authors":["Meryl Ye","Robert Kraut","Steve Rathje"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"replace","date":"2026-08-04","first_seen":"2026-07-29","revised_at":"2026-08-04","abs_url":"https://arxiv.org/abs/2607.25166","pdf_url":"https://arxiv.org/pdf/2607.25166","source_feed":"cs.AI","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["AI谄媚","人机交互","干预实验"],"reason":"研究人类对谄媚AI的感知与说服力，非用LLM仿真人类被试，但涉及AI行为对人类…","model":"deepseek-v4-pro","scored_at":"2026-08-04T13:05:01","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-29","rank":10,"question":"个体层面的干预（如警告或视频示范）能否降低用户对谄媚AI的正面评价，并削弱其对用户态度的说服力？","design":"本研究并非用LLM仿真人类被试，而是以人类为被试，通过在线实验检验两种干预效果：研究一让被试在与谄媚聊天机器人对话前阅读简短警告；研究二让被试先观看该AI赞同其他用户的视频再与之互动。测量结果包括对AI的感知（客观性、可信度、愉悦度）和话题态度（确定性、极端性）。","baseline":"无对照","findings":"警告和视频干预均改变了被试对AI的感知，降低了其客观性或愉悦度，但未能减少AI的说服力。汇总六项干预（N=3982）后，模式一致：干预使AI显得更不客观、更不可信，但均未削弱其对用户态度的影响。","reliability":"论文未讨论","relevance":"该研究直接测量人类与谄媚AI交互后的态度改变，虽非LLM仿真人类，但提供了人类行为基准数据，对评估LLM仿真人类在说服与态度极化场景中的可靠性有参考价值，值得阅读以获取效应量参考。","inspiration":"可借鉴其多干预汇总比较的设计，检验不同干预对感知与行为的分离效应。｜可迁移至金融建议场景，如AI理财顾问的谄媚行为对投资者风险偏好和产品选择的影响。｜以人类投资者为被试，随机分配接受谄媚或中立的AI投资建议，处理组在建议前观看AI对其他客户无差别赞同的视频，测量其风险资产配置比例和信任评分，并以真实市场数据或历史投资记录作为对照基准。"}},{"id":"2607.28128","version":2,"title":"Rethinking LLM-Judged Helpfulness as a Pedagogy Signal: A Pre-Registered Audit Across Tutor Models","zh_title":"重新思考LLM评判的有用性作为教学信号：一项跨导师模型的预注册审计","abstract":"LLM tutoring poses a measurement problem: can a general-purpose helpfulness rubric distinguish direct answer-giving from pedagogical guidance? We audit this signal in a pre-registered study. Within each of three tutor bases, we compare conversational and pedagogical policies instantiated with the same underlying model and paired with one fixed weak simulated student. Deterministic detectors measure answer leakage and next-turn independent work. Claude Opus 4.8 is the frozen, condition-blind primary judge. After the Opus scores were fixed, GPT-5.6 Sol was prospectively specified for a post hoc robustness audit of the same 1,179 confirmatory answer-phase tutor turns under the frozen helpfulness and pedagogy rubrics. On the primary base under Opus, the policies do not differ significantly in helpfulness but are perfectly rank-separated under the pedagogy rubric (Cliff's $|\\delta|{=}0.10$ vs. $1.0$). Across the two judges, pedagogy contrasts retain their direction where detected, whereas the helpfulness ordering is judge-contingent, reversing between judges on two of three bases. In an Opus-only ablation, seven primary-base policies span $2.3$ points in mean judged pedagogy within a $0.25$-point band of mean judged helpfulness. Separately, answer-revealing turns are followed by less independent student work on every base, a result that is judge-invariant by construction. In this controlled setting, general-purpose helpfulness is not a reliable pedagogy signal. Tutor evaluation should pair pedagogy-targeted rubrics with deterministic process measures.","authors":["Shuyi Fan","Boyuan Deng","Mengyu Xu","Jiale Liu","Hongyang Zhang","Qiaoxin Yang","Chongyang Gao"],"categories":["cs.CL","cs.AI","cs.CY"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-08-04","first_seen":"2026-07-31","revised_at":"2026-08-04","abs_url":"https://arxiv.org/abs/2607.28128","pdf_url":"https://arxiv.org/pdf/2607.28128","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM评估","教学对话","标注替代"],"reason":"用LLM评判教学对话质量，属替代人工标注，非仿真人类被试行为。","model":"deepseek-v4-pro","scored_at":"2026-08-04T13:05:03","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-31","rank":25,"question":"通用帮助性评分能否可靠区分直接给答案与教学引导这两种辅导行为？","design":"本研究并非用LLM仿真人类被试，而是用LLM作为评判者（Claude Opus 4.8为主评判，GPT-5.6 Sol为稳健性评判）对三种基座模型（Claude Sonnet 4.6、GPT-5.5、Gemini 3.1 Pro）下的两种辅导策略（对话式与教学式）生成的回答进行帮助性和教学性评分，同时用确定性检测器测量答案泄露和学生独立作业情况。","baseline":"无对照","findings":"在主要基座模型上，两种策略的帮助性评分无显著差异，但教学性评分完全分离；帮助性排序因评判模型而异，而教学性评分方向保持稳定。答案泄露的回合后学生独立作业减少，这一结果不受评判模型影响。","reliability":"论文未讨论","relevance":"本文用LLM替代人工进行教学评估，属于标注替代而非仿真人类被试，与研究者关注的LLM作为人类被试替代品进行行为仿真和决策复现的核心兴趣不符，但其中关于评判信号可靠性的批判性分析可提供方法借鉴。","inspiration":"借鉴其使用多个评判模型进行稳健性审计、结合确定性过程测量与主观评分的方法，以暴露单一评判信号的不可靠性。｜可迁移至经济政策沟通效果评估场景，如央行公告的清晰度与引导性评判。｜以LLM生成不同风格的央行公告（直接告知决策 vs. 解释决策逻辑），用多个LLM评判其清晰度与引导性，同时测量公众预期调整的确定性指标，并与真实市场调查数据对照。"}},{"id":"2608.00007","version":1,"title":"MemoryForge: Synthesize Lifelong Memory for Human-Like LLM Agents","zh_title":"MemoryForge：为类人LLM智能体合成终身记忆","abstract":"Equipping Large Language Models (LLMs) with human-like personas is crucial for agentic applications, such as role-play and user simulation. Traditional prompt-based methods rely on descriptive conditioning by injecting static textual profiles, which often makes agents show generic behaviors due to a lack of realistic life memory. To fill this gap, we introduce memory-based conditioning, a paradigm inspired by the cognitive psychology, which replaces abstract profiles with an autobiographical memory base, enabling frozen LLMs to dynamically retrieve situation-relevant memory to guide their behaviors. We formalize its enabling task as customized lifelong memory synthesis and propose MemoryForge, a novel framework to synthesize such lifelong memory from brief target personas. MemoryForge has three key components: a context generator for socio-historical grounding, a life organizer for developmental coherence toward the target identity, and a multi-resolution simulator that balances broad temporal summaries with high-fidelity episodic experiences. Experiments on PersonaGym for role-play and SimulatorArena for user-simulation, show that the synthesized memory base by MemoryForge enables frozen LLMs to exhibit more human-like behaviors than strong descriptive conditioning baselines across multiple metrics and LLM backbones.","authors":["Bohan Tang","Yiwen Guo"],"categories":["cs.CL","cs.AI","cs.LG"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-04","first_seen":"2026-08-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.00007","pdf_url":"https://arxiv.org/pdf/2608.00007","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["LLM智能体","用户仿真","记忆合成"],"reason":"用户仿真但无真实人类数据对照，属社会模拟边界情形","model":"deepseek-v4-pro","scored_at":"2026-08-04T13:04:20","error":null,"has_summary":false,"summary":null},{"id":"2608.00261","version":1,"title":"Cross-Task Dissociation in Frontier Vision-Language Model Theory of Mind","zh_title":"前沿视觉语言模型心智理论的跨任务分离","abstract":"Do frontier vision-language models present a coherent Theory-of-Mind (ToM) profile across tasks, matching the same human reference group, or does that profile fragment from one paradigm to the next? We evaluate a shared panel of nine frontier VLMs on two psychology-derived benchmarks: the Keysar Director Task (visual perspective-taking under egocentric interference) and the Frith-Happ\\'e animated triangles scored with the Castelli rubric (intention attribution from pure motion). On the Director Task, without chain-of-thought, the panel makes the egocentric error on 78\\% of trials like children rather than adults; variation is substantial across models, and reasoning rescues several models. On the triangles, the panel under-attributes intention: its ToM profile sits more than three times closer to the high-functioning-autistic-adult (HF-ASD) mean than to the typical-development-adult (TD) mean, while Goal-Directed and Random stay near TD. No model is nearest TD on both tasks; the model that looks adult-like on the Director Task falls on the HF-ASD side on the triangles, and the most TD-like model on the triangles is child-like on the Director Task. We report group-level descriptions, not diagnostic labels for any model.","authors":["Kejia Zhang","Youran Sun","Chugang Yi","Haizhao Yang"],"categories":["cs.CL","cs.AI","cs.CV","cs.MA","q-bio.NC"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-04","first_seen":"2026-08-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.00261","pdf_url":"https://arxiv.org/pdf/2608.00261","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["心智理论","模型评估","视觉语言模型"],"reason":"测量VLM的心智理论能力，属于对模型本身的认知评估，而非用LLM仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-08-04T13:04:21","error":null,"has_summary":false,"summary":null},{"id":"2608.02372","version":1,"title":"PredAct-Bench: Benchmarking Tool-Augmented Dialogue under Controlled Tool Noise","zh_title":"PredAct-Bench：在受控工具噪声下评测工具增强对话","abstract":"Large Language Models (LLMs) are increasingly deployed in task-oriented dialogue systems that support multi-step decision-making in high-stakes domains such as education, healthcare, and finance. However, existing benchmarks typically assume perfectly accurate tool outputs, overlooking the reality that deployed systems must operate with noisy tools and human decision-makers whose trust in the agent is itself uncertain. Such conditions are common in practice, for example, a clinician using a diagnostic prediction tool or an advisor relying on a model that forecasts student outcomes from historical records. We introduce PREDACTBENCH, a benchmark for evaluating dialogue agents paired with statistically imperfect tools, using education as a measurable testbed where ground truth outcomes and clear intervention decisions are available. First, we build a benchmark for AI-assisted human decision-making, where the AI uses noisy predictors to help guide a user. Second, we introduce episode-level Relative AI-Reliance (RAIR) and Relative self-reliance (RSR) metrics, extending prior trust calibration framework to multi-turn dialogue. Third, we evaluate 13 state-of-the-art closed and open source LLMs on two educational datasets, OULAD (real assessment trajectories from the UK Open University) and PREDACT-CS (60 courses with real final grade outcomes and synthetically generated weekly score trajectories), alongside a human study with instructors and teaching assistants. We find that when tools are noisy, SOTA models are supposed to provide visibility to teachers so that they do not over-rely on wrong suggestions or hallucinations, but current models fail to do that. We offer PREDACTBENCH to help build better LLMs as AI decision support systems to help teachers.","authors":["Abdulrahman AlRabah","Xiaocheng Yang","Dilek Hakkani-T\\\"ur","Abdussalam Alawini"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-04","first_seen":"2026-08-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.02372","pdf_url":"https://arxiv.org/pdf/2608.02372","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["AI辅助决策","人机交互","信任校准"],"reason":"LLM作为决策辅助工具，研究人类对AI的依赖，非直接仿真人类被试，但涉及人类行…","model":"deepseek-v4-pro","scored_at":"2026-08-04T13:04:57","error":null,"has_summary":false,"summary":null},{"id":"2608.00339","version":1,"title":"Bayesian and Motivated Reasoning in AI Agents","zh_title":"AI代理中的贝叶斯推理与动机性推理","abstract":"AI agents increasingly perform open-ended tasks in settings where their conclusions can guide consequential decisions. We provide evidence that AI agents draw different conclusions from identical numerical data when the substantive framing changes. We demonstrate this behavior in high-stakes domains in medicine, election forensics, and geopolitical forecasting by holding the evidence fixed while changing the scenario in which the evidence appears. Across twelve agent-domain comparisons, agents' conclusions are strongly influenced by their prior beliefs. They are more likely to reach an affirmative conclusion when it is framed around a proposition they already regard as likely, while the reverse holds when the framing conflicts with their prior. The framing also changes how some agents work: they search more extensively, choose different analytical specifications, and evaluate the same evidence differently. These results identify a particular risk of delegating decision-making to AI agents, as their decisions may depend on prior beliefs that are neither specified in the task nor visible in the decision record.","authors":["Eddie Yang"],"categories":["cs.AI","cs.CL"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-08-04","first_seen":"2026-08-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.00339","pdf_url":"https://arxiv.org/pdf/2608.00339","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["AI推理偏差","决策行为","模型评估"],"reason":"研究AI代理的推理偏差，测量的是模型本身而非仿真人类被试，无人类数据对照。","model":"deepseek-v4-pro","scored_at":"2026-08-04T13:04:21","error":null,"has_summary":false,"summary":null},{"id":"2608.00717","version":1,"title":"AI-Based Thesis Assessment: An Empirical Study of Human Evaluation Priorities and Their Impact on Automated Assessment","zh_title":"基于AI的论文评估：人类评估优先级及其对自动化评估影响的实证研究","abstract":"Rubric-based AI systems for thesis assessment use criterion weights to assign different levels of importance to evaluation criteria. These weights are typically defined through expert judgment, although little empirical evidence exists regarding how thesis supervisors actually prioritize evaluation criteria. Consequently, this study investigates supervisor-derived criterion weights in thesis assessment and evaluates their impact on AI-based assessment. We surveyed 84 thesis supervisors across four academic disciplines and collected weighting data for 35 thesis assessment criteria. Comparison with the default criterion weights of the AI assessment system RubiSCoT [1] revealed substantial divergences between supervisor-derived and default criterion weights. To evaluate the practical implications of these differences, the supervisor-derived weights were integrated into multiple calibration configurations and evaluated on a corpus of 80 German-language theses. The best-performing configuration reduced the mean relative deviation between AI-generated and supervisor-assigned evaluations from 11.18% to 10.85%, although the improvement was not statistically significant. Human supervisors showed substantially stronger agreement with each other, exhibiting a mean inter-supervisor relative deviation of 4.44%. The findings indicate that criterion-weight calibration alone does not substantially improve alignment between AI-generated and human assessments.","authors":["Garv Vikram Gursahaney","Baskhad Idrisov","Thorsten Fr\\\"ohlich","Tim Schlippe"],"categories":["cs.AI","cs.CL"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-08-04","first_seen":"2026-08-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.00717","pdf_url":"https://arxiv.org/pdf/2608.00717","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["AI辅助评估","评分权重校准","人机对比"],"reason":"用LLM替代人工评分，属于标注员替代而非仿真被试，但涉及人类评估对照，边界情形。","model":"deepseek-v4-pro","scored_at":"2026-08-04T13:04:40","error":null,"has_summary":false,"summary":null},{"id":"2608.01868","version":1,"title":"No One Wins in Nuclear War: A Social Simulation of Military Decision-making","zh_title":"核战争中没有赢家：军事决策的社会模拟","abstract":"WOPR is a social-simulation environment for studying how organizations make high-stakes decisions, built on a deterministic, replay-validated rules engine and using wargames as the vehicle. We instantiate it first with the published card game Nuclear War, traced against its published rules. We start with military decision-making because of its safety implications and because it needs further study, but the design is not specific to it: the decision-point contract that exposes the engine to agents is reusable across verifiable rule systems. Existing social-simulation work emphasizes persona fidelity and synthetic opinion, but lacks a verifiable rules engine with replay-checkable mechanics and private-channel negotiation. WOPR supplies that engine, and its contract makes every strategic choice an explicit agent decision. The method is agnostic to social-simulation frameworks; we adopt Concordia as the default harness for driving the game. On the same engine, WOPR layers a four-rung press ladder from silence to private single-recipient channels with structured commitments, and instantiates each faction as a collective command-and-control system rather than a single agent. We make all code, example configurations, and replay data publicly available at https://github.com/eilab-gt/wopr.","authors":["Glenn Matlin","Isaac Song","Anthony Wen-Ming Zang","Mark Riedl"],"categories":["cs.CY","cs.AI","cs.CL","cs.MA"],"primary_category":"cs.CY","announce_type":"cross","date":"2026-08-04","first_seen":"2026-08-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.01868","pdf_url":"https://arxiv.org/pdf/2608.01868","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["社会模拟","多智能体","军事决策"],"reason":"用多智能体模拟军事决策，属社会模拟但无真实人类数据对照，为边界情形。","model":"deepseek-v4-pro","scored_at":"2026-08-04T13:04:32","error":null,"has_summary":false,"summary":null},{"id":"2608.00102","version":1,"title":"Can LLM Agents Price Competitively? A Dynamic Multi-Attribute Auction Benchmark for Agentic Commerce","zh_title":"LLM代理能否竞争性定价？一个面向代理商务的动态多属性拍卖基准","abstract":"Agentic commerce is moving from concept to deployed infrastructure: payment networks, retailers, and AI platforms are setting the stage for agents to transact on behalf of merchants and consumers. Yet whether the LLMs behind these agents can price competently in real markets, where customer preferences are hidden, competitors adapt in real time, and demand can shift without warning, has not been systematically tested. We introduce Bazaar, a dynamic sealed-bid benchmark for multi-attribute auction under these conditions. Despite its dynamics, the benchmark is grounded in closed-form customer utilities, enabling exact evaluation. Across 11 frontier LLMs from four providers, the leading agents on customer acquisition (e.g. Gemini 3.1 Pro) are often not the leading agents on profit (e.g. Opus 4.6). The ranking shifts again under demand shocks: agents that learned fastest pre-shock are typically the slowest to revise their beliefs afterwards, while Gemini 3.1 Pro recovers fastest despite not leading on profit. However, even the strongest agent captures less than a third of hindsight-optimal profit, suggesting current LLMs are progressing in agentic commerce but leave substantial headroom.","authors":["Shimaa Ahmed","Yiwei Cai","Mohsen Minaei","Rahul Rachuri"],"categories":["cs.AI","cs.MA"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-04","first_seen":"2026-08-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.00102","pdf_url":"https://arxiv.org/pdf/2608.00102","source_feed":"cs.AI","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["LLM代理","拍卖模拟","经济行为"],"reason":"LLM代理竞价模拟市场，但无真实人类数据对照，属社会模拟边界情形。","model":"deepseek-v4-pro","scored_at":"2026-08-04T13:04:36","error":null,"has_summary":false,"summary":null},{"id":"2608.01361","version":1,"title":"High-Stakes Decisions with Language Models: Insights from Emergency Triage","zh_title":"语言模型在高风险决策中的应用：来自急诊分诊的见解","abstract":"High-stakes decisions under uncertainty, such as medical emergency triage, require more than accurate predictions. They depend on estimating the likelihood of alternative outcomes while explicitly weighing the consequences of different actions, principles that have long formed the foundation of medical diagnosis and decision making. Yet language models are increasingly used for high-stakes clinical recommendations without explicit specification of the utilities governing these decisions. Here we show that emergency triage with language models can be understood within a probabilistic decision framework, providing a case study of a broader decision-analytic paradigm for steering, evaluating, and deploying language models in high-stakes settings. Using clinical vignettes from a structured evaluation of a consumer triage system, we analyze recommendations for treatment under alternative utility functions that specify the relative costs of missed emergencies and unnecessary escalation. We find that capable language models adjust recommendations in response to stated utilities, revealing that the same underlying predictions can support markedly different decision policies. These findings show that effective deployment depends not only on improving predictions but also on making decision objectives explicit. More broadly, they suggest that language models for high-stakes applications should be understood and evaluated as probabilistic decision systems whose recommendations depend jointly on predictive performance and explicit utilities.","authors":["Khurram Yamin","Christopher Kelly","Bryan Wilder","Eric Horvitz"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-04","first_seen":"2026-08-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.01361","pdf_url":"https://arxiv.org/pdf/2608.01361","source_feed":"cs.AI","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM决策","急诊分诊","效用函数"],"reason":"用LLM做急诊分诊决策，替代人类决策者，但非仿真人类被试行为，属替代人类劳动而…","model":"deepseek-v4-pro","scored_at":"2026-08-04T13:04:28","error":null,"has_summary":false,"summary":null},{"id":"2608.00357","version":1,"title":"Dynamic Surveys: Using LLMs to Blend Qualitative Depth,Quantitative Structure, and Collaborative Interaction","zh_title":"动态调查：利用大语言模型融合定性深度、定量结构与协作互动","abstract":"Surveys are a powerful tool for collecting data and eliciting insights on social phenomena, and are critical in product design, marketing, scientific research. However, traditional open-ended and closed-ended question formats limit researchers' ability to capture data that combines both the richness of qualitative insights and the analytical rigor of quantitative data. To address these problems, we propose Dynamic Surveys, a survey platform that uses Large Language Models (LLMs) to dynamically cluster qualitative responses in real time and to elicit quantitative ratings and rankings on those clusters and qualitative reflections on how their views compare to broader respondent trends, especially helpful in early-stage or exploratory research settings. This process generates a report showing survey creators and respondents the clustered responses as well as each cluster's rank, rating distribution, and follow-up reflections. To evaluate Dynamic Surveys, we conducted two field studies with 93 participants over a 2-month period. In the first study, 52 students provided input for a career workshop, while in the second, 41 students gave feedback on gaps in their academic curriculum. Of these, 44 respondents filled out a survey on their experience using Dynamic Surveys. We also shared the generated report with 4 individuals who were interested in the insights for their work, and interviewed them to understand their perspectives on the results and any contextual risks they saw in the platform design. Our findings suggest that Dynamic Surveys not only provide richer and deeper insights into responses compared with traditional survey tools, but also increase engagement and foster a sense of community. We discuss broader implications for the design of survey platforms that blend qualitative depth with quantitative structure, facilitating richer insights and offering more collaborative interactions.","authors":["Kehua Lei","Aidan Ladenburg","Zahra Petiwala","Zili Wang","Dishita Jhawar","Ipsita Bisht","Ansh Kumar","David T. Lee"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-08-04","first_seen":"2026-08-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.00357","pdf_url":"https://arxiv.org/pdf/2608.00357","source_feed":"cs.HC","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM辅助调查","定性定量融合","人机交互"],"reason":"用LLM动态聚类调查回复并生成报告，替代人工分析而非仿真人类被试，属标注替代边…","model":"deepseek-v4-pro","scored_at":"2026-08-04T13:04:23","error":null,"has_summary":false,"summary":null},{"id":"2608.01783","version":1,"title":"Comparative Validation of GPT-4o-mini and Teacher Mean Scores for Automated Scoring of Music Analysis Responses: Single-Pass Deployment, Repeatability, and Strategy-Specific Bias","zh_title":"GPT-4o-mini与教师均分在音乐分析回答自动评分中的比较验证：单次部署、可重复性与策略特定偏差","abstract":"Scoring open-ended music analysis responses is time-consuming and requires nuanced judgments of harmonic knowledge and formal understanding. This study evaluates the validity and repeatability of GPT-4o-mini for rubric-based scoring of music analysis essays, using teacher mean scores as the benchmark. A dataset of 300 university-level student responses was scored by teachers on four dimensions: Harmony, Form, Reasoning, and Terminology. GPT-4o-mini scored the same responses using three prompting strategies: few-shot prompting with chain-of-thought reasoning (Fs+CoT), retrieval-augmented generation (RAG), and self-consistency based on five internal generations per administration (SC). Each strategy was administered three times with the model, prompt, rubric, and response held constant. Single-pass scores represented an operational scoring condition, whereas median aggregation across three runs was used to examine robustness. Agreement with teacher mean scores was evaluated using correlation, intraclass correlation, Krippendorff's alpha, quadratic weighted kappa, and scoring error indices. Fs+CoT showed the strongest agreement with teacher mean scores in both single-pass scoring and median aggregation. RAG showed systematic over-scoring, whereas SC produced highly repeatable scores but weaker individual-level agreement. Dimension-level analyses showed that scoring performance varied across rubric components, with Terminology generally showing weaker agreement than Reasoning. These findings indicate that GPT-4o-mini can generate stable scores for complex music analysis responses, but prompting strategies produce distinct scoring profiles. Operational use therefore requires strategy-specific calibration, dimension-level validation, and continued human oversight.","authors":["Baicheng Lin","Lingxi Jin","Kyung-Seok Min"],"categories":["cs.SD","cs.HC","stat.AP"],"primary_category":"cs.SD","announce_type":"cross","date":"2026-08-04","first_seen":"2026-08-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.01783","pdf_url":"https://arxiv.org/pdf/2608.01783","source_feed":"cs.HC","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["自动评分","LLM标注","音乐教育"],"reason":"用LLM替代教师评分，属于替代人工标注员，非仿真人类被试，但方法可参考。","model":"deepseek-v4-pro","scored_at":"2026-08-04T13:04:31","error":null,"has_summary":false,"summary":null},{"id":"2608.00748","version":1,"title":"Me and My Bot: What Users Talk About in AI Companion Communities on Reddit","zh_title":"我和我的机器人：用户在Reddit AI伴侣社区中谈论什么","abstract":"AI companion communities on platforms such as Reddit are widely characterized as spaces where users discuss their relationships with AI bots. This study examines whether and how that characterization holds, guided by the Synthetic Resonance framework's claim that human-AI relationships can carry genuine relational meaning for the user. Multiple LLMs were employed to code 5,504 Reddit posts from eight AI companion communities for relationship focus, primary topic, and users' emotional valence. Although search terms were weighted toward relational and attachment language, only 45% of posts concerned the user's own relationship with their bot. Posts about users' own bots differed markedly from posts about bots in general in both topic and emotional expression, with 85% of general-bot posts containing no user emotion language compared to 33% of own-bot posts. Among the 970 posts that were relationally focused, companionship and romance each accounted for roughly 46% of discussion, sexual content for 8%, and emotional valence varied across these subtopics. The findings suggest that users engage with these relationships with AI bots as meaningful, and that the discourse about them is broad and emotionally complex.","authors":["Richard A. Fabes"],"categories":["cs.HC","cs.AI"],"primary_category":"cs.HC","announce_type":"cross","date":"2026-08-04","first_seen":"2026-08-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.00748","pdf_url":"https://arxiv.org/pdf/2608.00748","source_feed":"cs.AI","score":3,"bucket":"other","rubric_hits":["C3"],"tags":["AI伴侣","社区分析","情感表达"],"reason":"研究AI伴侣社区的用户讨论内容，属于角色扮演聊天分析，无实验或测量目的，不涉及…","model":"deepseek-v4-pro","scored_at":"2026-08-04T13:04:41","error":null,"has_summary":false,"summary":null},{"id":"2606.00811","version":2,"title":"Certificates without Electrons? Theory and Evidence on Impacts from AI-Driven Power Demand","zh_title":"无电子的证书？AI驱动电力需求影响的理论与证据","abstract":"Data centers now account for 4.4% of United States electricity demand, yet the grid-level effectiveness of the renewable energy certificates (RECs) and power purchase agreements (PPAs) hyperscalers use to claim carbon neutrality remains unclear. We develop a game-theoretic model in which a data center operator chooses among RECs, PPAs, and behind-the-meter colocation while generators make entry decisions under endogenous financing costs. The model identifies a timing wedge -- the mismatch between consumption and credited renewable generation -- as a central mechanism through which AI demand degrades reliability, raises prices, and increases emissions even when RECs cover 100% of annual consumption. Colocation with storage addresses this wedge directly and induces the greatest renewable entry by eliminating generator revenue risk. We test these predictions by exploiting the staggered release of large language models as a natural experiment, using difference-in-differences on a novel dataset linking AI activity to local grid outcomes. AI demand significantly increases fossil generation, wholesale prices (up to 25% in treated PJM zones), and outage frequency (0.5--1 additional outages per year) near data centers, with impacts scaling in model size. Data centers with on-site generation exhibit a sign reversal in power-quality effects, consistent with the model's prediction that behind-the-meter capacity absorbs demand spikes. Counterfactual analyses show that edge inference, spatial reallocation, and colocated storage each substantially mitigate grid impacts, while REC-only strategies do not. Together, our results demonstrate that the externalities of AI to the grid are tightly coupled to procurement design and the spatial organization of data center infrastructure.","authors":["Dana Golden","Aruna Balasubramanian","Niranjan Balasubramanian"],"categories":["econ.EM","cs.AI"],"primary_category":"econ.EM","announce_type":"replace-cross","date":"2026-08-04","first_seen":"2026-05-30","revised_at":"2026-08-04","abs_url":"https://arxiv.org/abs/2606.00811","pdf_url":"https://arxiv.org/pdf/2606.00811","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C2"],"tags":["电力市场","数据中心","可再生能源"],"reason":"研究AI数据中心对电网的影响，不涉及用LLM仿真人类被试或行为对照。","model":"deepseek-v4-pro","scored_at":"2026-08-04T13:04:35","error":null,"has_summary":false,"summary":null},{"id":"2607.23893","version":2,"title":"Who Gets Named: Citation Type Predicts Individual Naming by Grounded Language Models, and a Roster Instrument Captures 0.5% of It","zh_title":"谁被提名：引用类型预测接地语言模型的个人提名，花名册工具仅捕获0.5%","abstract":"Prior work on AI brand visibility measures the firm: does a model recommend a company, and does that track its reputation. This study asks the question one level down, in categories where the buyer picks a person. It issued 2,400 grounded API calls in one two-hour window on 24 July 2026: 120 buyer-intent prompts, four models (GPT-5.6 Sol, Gemini 3.6 Flash, Perplexity Sonar Pro, Grok 4.5), five iterations each, four European markets and five query languages. Every response was coded for whether it named an individual professional, by a rule cascade that never consults a roster and that drops detections resolving to a same-named American city (precision 96.9%, recall 61.7%, so every rate below is a lower bound). All inference corrects for clustering within prompt: intraclass correlation 0.258, effective n 407 against a nominal 2,400. Models named an individual in 25.8% of responses. Category dominates: real estate 35.4% and car dealerships 32.9% against insurance 9.1% (chi-square 159.3, p = 5.8e-8 after correction). Models differ four-fold, from Grok 38.0% to Gemini 9.3%. Citation type predicts naming and citation volume does not: naming responses cite the individual's own site 2.6 points more often (95% CI +1.4 to +3.9) and category portals 4.3 points more often, and cite firm-owned pages at the same rate (44.1% against 45.5%). On nine matched translation pairs, English prompts named an individual in 36.7% of responses against 15.6% for the same question in the local language (OR 3.14, clustered p = 0.074, so the direction is clear and the design cannot close it). A 939-person roster built from public LinkedIn search matched 128 of 27,293 name-shaped mentions (0.47%), 26 of the 939 people were ever named, and the roster-derived rates of 0.0% to 25.4% measure that overlap. Roster-based measurement of individual AI visibility sees a small and unrepresentative slice of what models do.","authors":["Dmitrij \\.Zatuchin (Rankfor.AI O\\\"U, Tallinn)"],"categories":["cs.IR","cs.CL","cs.CY"],"primary_category":"cs.IR","announce_type":"replace-cross","date":"2026-08-04","first_seen":"2026-07-28","revised_at":"2026-08-04","abs_url":"https://arxiv.org/abs/2607.23893","pdf_url":"https://arxiv.org/pdf/2607.23893","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["AI品牌可见度","引用分析","信息检索"],"reason":"研究AI品牌可见度，测量模型推荐个人，非仿真人类被试行为或决策，无人类数据对照。","model":"deepseek-v4-pro","scored_at":"2026-08-04T13:05:00","error":null,"has_summary":false,"summary":null},{"id":"2607.27056","version":2,"title":"Setoka: A Benchmark for Hierarchical User Understanding in Personalized Agents over Heterogeneous Data","zh_title":"Setoka：异构数据下个性化代理中层次化用户理解的基准","abstract":"Personalized agents are increasingly applied to assist users across a wide range of tasks. Effective personalized assistance requires not only retrieving explicit facts from past interactions stored in agent memory, but also inferring abstract personal characteristics. However, existing memory benchmarks primarily evaluate whether an agent can retrieve information explicitly stated in conversational histories, failing to provide an effective assessment of deeper user understanding. In this work, we propose Setoka, a benchmark for evaluating memory-augmented personalized agents with hierarchical user understanding from heterogeneous data. Grounded in theories from cognitive and personality psychology, Setoka defines four levels of user understanding, i.e., semantic memory, episodic memory, behavior pattern, and personality trait. Moreover, to enable realistic yet privacy-preserving evaluation, we design a psychometrics-based pipeline that synthesizes diverse, coherent heterogeneous user data and queries at scale. Finally, we leverage Setoka to evaluate 3 language models combined with 5 memory systems for 10 synthetic users. Our comprehensive evaluation reveals that while existing systems perform well on semantic memory retrieval, their performance declines on episodic memory. Moreover, when dealing with behavior pattern and personality trait understanding tasks that require integrating heterogeneous and fragmented information dispersed over time, performance declines even further. These findings demonstrate that user understanding cannot be handled by simple fact retrieval, motivating the design of memory mechanisms for cross-source integration and abstraction over long-term user behavior.","authors":["Lingyang Zeng","Guangze Chen","Kaichen Yu","Zhicheng Pan","Siyang Weng","Zirui Hu","Xiangyun Du","Hailin He","Rong Zhang","Chengcheng Yang","Kai Huang","Xuan Zhou"],"categories":["cs.AI","cs.CL"],"primary_category":"cs.AI","announce_type":"replace-cross","date":"2026-08-04","first_seen":"2026-07-30","revised_at":"2026-08-04","abs_url":"https://arxiv.org/abs/2607.27056","pdf_url":"https://arxiv.org/pdf/2607.27056","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["个性化代理","用户理解","基准测试"],"reason":"评估个性化代理对合成用户的理解，属角色扮演与人格测量，无人类行为对照实验。","model":"deepseek-v4-pro","scored_at":"2026-08-04T13:05:03","error":null,"has_summary":false,"summary":null},{"id":"2608.00205","version":1,"title":"Averaging Bias: Human Faithfulness Annotations are not Locally Faithful","zh_title":"平均偏差：人类忠实度标注并非局部忠实","abstract":"Evaluation of faithfulness of text summarization treats a model generated summary as faithful only if every of its sentences is supported by the source document: a strict conjunctive rule under which a single unsupported sentence makes the whole summary unfaithful. Yet most faithfulness benchmarks collect only one global human annotation label per summary. We ask whether such global human labels actually implement the conjunctive rule. We hypothesize that annotators may accept a summary as faithful when most sentences are faithful, not only when all are faithful. To test our hypothesis, we use five large language model (LLM) judges as per-sentence raters across four widely used faithfulness benchmarks. We find that global human labels correlate better with the average of per-sentence LLM judgments than with the implementation of the strict conjunctive rule. A manual review confirms that a substantial fraction of summaries labeled faithful by humans contain genuine local factual errors. We call this tendency Averaging Bias. Our results reveal that human labels on widely used faithfulness benchmarks contain measurable Averaging Bias, calling for carefully structured designs for trustworthy human annotations","authors":["Huajian Zhang","Yiyang Feng","Jiawei Zhou"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-04","first_seen":"2026-08-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.00205","pdf_url":"https://arxiv.org/pdf/2608.00205","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["文本摘要","忠实度评估","标注偏差"],"reason":"评估摘要忠实度标注偏差，属NLP评测，不以人类行为仿真为目标","model":"deepseek-v4-pro","scored_at":"2026-08-04T13:04:38","error":null,"has_summary":false,"summary":null},{"id":"2608.00640","version":1,"title":"TreeProbe : A Tibetan Medicine Benchmark for Cultural Bias in LLMs","zh_title":"TreeProbe：藏医学文化偏见基准测试","abstract":"Large language models are increasingly viewed as a potential means of mitigating global health inequities, yet their outputs often reflect dominant high-resource medical traditions and provide limited coverage of traditional medical knowledge systems. Tibetan medicine, one of the world's four major traditional medical systems, has an independent and highly structured theoretical framework. When models lack grounded understanding of Tibetan medicine, they may fall back on dominant epistemic systems and distort the native knowledge structure during reasoning. However, quantitative tools for evaluating cultural bias in Tibetan medicine remain largely absent. To address this gap, we introduce TreeProbe, the first cultural-bias benchmark organized around the native Tree of Medicine framework in Tibetan medicine. It contains 4,719 expert-adjudicated items covering 467 diseases and 10 subtasks along the three roots. Experiments on representative LLMs show that current models remain limited in native Tibetan medical contexts and exhibit systematic external ontology drift. Further analysis reveals that models diverge in whether they drift toward biomedical or TCM reasoning, shaped by pretraining data composition and surface resemblance between TCM and Tibetan medicine. TreeProbe provides a diagnostic benchmark for developing medical AI systems that are both linguistically inclusive and epistemically fair. Code and data are available in an anonymous repository at https://anonymous.4open.science/r/TreeProbe/.","authors":["Jin Zhang","Linyu Li","Weili Jiang","Yuqing Cai","Yutong Liu","Guanquecairang","Yongbin Yu","Jingye Cai","Nyima Tashi","Gadeng Luosang"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-04","first_seen":"2026-08-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.00640","pdf_url":"https://arxiv.org/pdf/2608.00640","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["文化偏见","医学知识评测","LLM基准"],"reason":"纯NLP评测基准，评估LLM在藏医学知识上的文化偏见，不涉及人类行为仿真或对照。","model":"deepseek-v4-pro","scored_at":"2026-08-04T13:04:40","error":null,"has_summary":false,"summary":null},{"id":"2608.01012","version":1,"title":"MedUPS: Towards Diagnostic Assistance in Uncommon Medical Cases with Large Language Models","zh_title":"MedUPS：面向罕见病例诊断辅助的大语言模型研究","abstract":"Uncommon and off-guideline cases are difficult for clinical decision support, because physicians must make a series of management decisions under diagnostic uncertainty and rarely see the full case at once. Most large language model (LLM) benchmarks for medicine score only the final diagnosis, yet much of clinical care turns on the next appropriate action: the next test to order, the imaging study to obtain, the specialist to involve, or the differential to pursue. We introduce MedUPSQA, a dataset of 21,874 mid-stream clinical decision points built from 5,535 real case reports, and MedUPS, an alignment framework that supervises models on these intermediate decisions as they unfold along a patient's trajectory. We segment free-text case presentations into chronologically ordered, accumulating clinical chunks and align models to predict the next step with reinforcement learning (GRPO), using an external LLM-as-a-Judge reward. This objective mirrors how clinicians actually meet patients, reasoning forward from accumulating evidence toward the next decision, rather than committing to a final label. Across three backbones, mid-stream alignment raises next-step accuracy from 55.2 to 66.7 for Qwen3.6-27B, from 47.2 to 57.8 for Qwen3.5-9B, and from 37.8 to 44.4 for HuatuoGPT-3-8B, with 95% CI. In several model scales we test the objective improves accuracy more than scale, with smaller models surpassing larger, frontier models we evaluate. We further train supervised fine-tuning (SFT) baselines on the mid-stream task, SFT improves all backbones above base, indicating the target framwork carries signal independently of the optimizer. We release the dataset, code, and aligned checkpoints.","authors":["Ofir Ben Shoham","Oriel Perets","Nir Grinberg","Nadav Rappoport"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-04","first_seen":"2026-08-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.01012","pdf_url":"https://arxiv.org/pdf/2608.01012","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["医学NLP","临床决策支持","模型对齐"],"reason":"纯医学NLP评测，用LLM辅助临床决策，不以人类行为仿真为目标","model":"deepseek-v4-pro","scored_at":"2026-08-04T13:04:44","error":null,"has_summary":false,"summary":null},{"id":"2608.01322","version":1,"title":"Can Language Models Identify Shadow Trading Targets? An NLP Evaluation of SEC Enforcement Theory","zh_title":"语言模型能识别影子交易目标吗？对SEC执法理论的NLP评估","abstract":"Shadow trading -- trading in a peer firm's securities on the basis of material nonpublic information (MNPI) about an \"economically linked\" company -- is a novel and contested theory of insider trading liability, first prosecuted in SEC v. Panuwat (2023). Enforcing it requires identifying economically linked firms ex ante, a determination the SEC makes only after the fact using mass market surveillance infrastructure. We ask whether NLP can do what the SEC's theory presumes insiders already know: identify peer firms ex ante from publicly mandated disclosures. Using a two-stage LLM pipeline applied to Item 7 (Management's Discussion and Analysis) sections of SEC 10-K filings, we score semantic similarity across 30 M&A events spanning five industries and relate similarity to announcement-day abnormal stock returns. On the Panuwat fact pattern itself the pipeline recovers Incyte among the closest peers, a sanity check on the one case with a known outcome. Across the full dataset, however, we find no association: pooling 217 peer observations, the within-event rank correlation between similarity and abnormal return is +0.07 (permutation p = 0.37), and the mean per-event Spearman correlation is +0.05 with a 95% confidence interval of [-0.08, +0.18] -- narrow enough to exclude any moderate relationship rather than merely failing to detect one. A case-level reading agrees: 14 of 30 events support the hypothesis, 12 contradict it, and 4 are ambiguous. We also find that Incyte fell outside the standard \\$2B-\\$10B mid-cap band on the day before the announcement, complicating the \"mid-cap oncology\" category the SEC invoked. These results are exploratory and bound to this pipeline, corpus, and return measure, but they put pressure on the empirical premise of shadow trading enforcement and bear on constitutional questions surrounding the SEC's financial surveillance infrastructure.","authors":["Sarah Wilson","Michael MacKay","Anthony Marello","Trinav Bhattacharyya"],"categories":["cs.CL","cs.CY"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-04","first_seen":"2026-08-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.01322","pdf_url":"https://arxiv.org/pdf/2608.01322","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["NLP应用","金融监管","文本相似度"],"reason":"用LLM分析SEC文件，评估NLP方法识别影子交易，不涉及人类行为仿真或对照。","model":"deepseek-v4-pro","scored_at":"2026-08-04T13:04:45","error":null,"has_summary":false,"summary":null},{"id":"2608.01395","version":1,"title":"Language Equality has a Price: A Systematic Investigation of Multi-turn LLM Performance for EU-24+","zh_title":"语言平等有代价：对EU-24+多轮LLM性能的系统性研究","abstract":"We evaluate large language models (LLMs) as language agents playing goal-directed dialogue games in self-play across 30 languages: the 24 official EU languages plus six others. Unlike static or preference-based evaluation, this paradigm is multi-turn, reference-free and programmatically scored, and because the game mechanics are language-agnostic it extends to a new language by localising a fixed set of prompt and word-list files. Evaluating nine open-weight and commercial LLMs, we find that no open-weight model covers the EU-24 well: in every official language both commercial systems outscore every open-weight model, and the two weakest average below 40 points across the EU-24. The commercial systems stay ahead even in languages with four orders of magnitude less public web text, showing that linguistic parity is achievable, but not from public crawls alone. A model's home region lifts it without closing the gap: Chinese is the strongest of all 30 languages for two Chinese-developed models, yet the best Chinese score of any model belongs to a US commercial system. Coverage is also not parity of service. Pooled over models and languages, the median non-English language costs 31% more to run than English, and scores 10% lower.","authors":["Sherzod Hakimov","Karl Osswald","Jelle Psurek","Eszter Bukovszky","A. Altar L\\\"user","David Schlangen"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-04","first_seen":"2026-08-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.01395","pdf_url":"https://arxiv.org/pdf/2608.01395","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体对话游戏","多语言评估","LLM性能"],"reason":"纯多智能体对话游戏自博弈，无人类行为对照，属C1排除项。","model":"deepseek-v4-pro","scored_at":"2026-08-04T13:04:46","error":null,"has_summary":false,"summary":null},{"id":"2608.01585","version":1,"title":"Semantic Alignment of AI Models: Concept Collapse, Checkpoint Dynamics, and Cross-Lingual Transfer","zh_title":"AI模型的语义对齐：概念坍缩、检查点动态与跨语言迁移","abstract":"Language model benchmarking is a difficult task. Outcome reasoning alone does not test the model's conceptualization of language and popular open-source benchmarks are quickly saturated or ingested as training data. It is important to test the model's output, but augmenting these tests by characterizing semantic structure gives more insight to how models relate abstract concepts. However, the high dimensional embedding spaces are not easy to interpret. This work demonstrates how topological methods can be used to rigorously compare these spaces to low dimensional and interpretable baselines like ontologies and curated knowledge graphs. These multi-modal alignment tests make it possible to track model adaptations and test phrase understanding across multiple languages.","authors":["Tyler Ashoff","Jordan Rodu"],"categories":["cs.CL","cs.LG"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-04","first_seen":"2026-08-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.01585","pdf_url":"https://arxiv.org/pdf/2608.01585","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["语义对齐","拓扑方法","模型评测"],"reason":"纯模型语义对齐评测，无人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-08-04T13:04:49","error":null,"has_summary":false,"summary":null},{"id":"2608.01598","version":1,"title":"PICTURE: Enhancing Theory-of-Mind in Large Language Models by Revealing, Not Hiding, Characters' Lack of Knowledge","zh_title":"PICTURE：通过揭示而非隐藏角色知识缺失来增强大语言模型的心智理论","abstract":"Simulating human-like Theory of Mind (ToM) has been a longstanding problem in natural language processing (NLP). To address this, existing works introduce a reasoning step of event hiding (a.k.a. perspective-taking), where events unknown to a character are removed before question answering. However, resorting to event hiding for ToM reasoning presents a performance degradation issue due to the strict output format constraints involved in event hiding. To mitigate this issue, we propose generating perspective-taking outputs as free-form explanations without event hiding, but this poses a notable yet underexplored challenge: LLMs need to inhibit responses to events unknown to characters, because the absence of event hiding exposes LLMs to these events throughout reasoning. To address this challenge, we hypothesize and empirically verify that LLMs can achieve such inhibition if a character's lack of knowledge about events is made explicit during reasoning. Based on this finding, we introduce PICTURE, a new prompting method that enables LLMs to generate a character's lack of knowledge within free-form Chain-of-Thought (CoT). Experimental results show that PICTURE outperforms existing prompting methods by an average of 7.3% on false-belief tasks.","authors":["Eojin Jeon","SangKeun Lee"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-04","first_seen":"2026-08-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.01598","pdf_url":"https://arxiv.org/pdf/2608.01598","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["心智理论","提示方法","NLP评测"],"reason":"研究LLM的心智理论推理能力，属纯NLP评测，不以人类行为仿真为目标。","model":"deepseek-v4-pro","scored_at":"2026-08-04T13:04:50","error":null,"has_summary":false,"summary":null},{"id":"2608.01724","version":1,"title":"TIDES: A Longitudinal Bilingual Dataset for Modeling Multi-Party Social Dynamics","zh_title":"TIDES：用于多群体社会动态建模的双语纵向数据集","abstract":"Group conversations are fundamental to human collaboration, yet standard large language models (LLMs) still struggle with the complexities of multi-party interaction. This challenge persists in part because existing group conversation datasets are often limited to short-term lab settings with contrived tasks, failing to capture the long-term social dynamics of real-world teams. To bridge this gap, we introduce TIDES, a high-resolution longitudinal dataset tracking 12 university project teams over a full semester. Comprising 75,971 utterances in both English and Korean from in-person meetings, TIDES provides a naturalistic record of teams working on self-managed projects. Our socio-structural annotations-covering interaction types, emergent roles, and development stages-allow for modeling of team evolution over months. Experiments show that fine-tuning on TIDES improves next-speaker prediction by 13.8 percentage points over a bigram baseline (64.53%) and yields performance comparable to strong proprietary zero-shot models. The model also comes within 2.1 percentage points of the published state of the art on the AMI Meeting Corpus while using approximately 42% less training data. However, human evaluations suggest that better next-speaker prediction does not necessarily yield more natural or coherent utterances, as fine-tuned models were generally less preferred than vanilla models. This potential mismatch motivates further study of how structural modeling can support natural multi-party generation.","authors":["Heechan Lee","Jeonggyu Kang","Junho Myung","Jaywoong Jeong","Juho Kim","Joseph Seering"],"categories":["cs.CL","cs.AI","cs.HC"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-04","first_seen":"2026-08-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.01724","pdf_url":"https://arxiv.org/pdf/2608.01724","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体对话","数据集","下一说话人预测"],"reason":"多智能体对话建模，无人类行为仿真对照，属纯协作系统研究","model":"deepseek-v4-pro","scored_at":"2026-08-04T13:04:53","error":null,"has_summary":false,"summary":null},{"id":"2608.02486","version":1,"title":"Cultural Awareness is Represented but Not Decoded: Tracing Mythological Knowledge across 18 Open-Source LLMs","zh_title":"文化意识被表征但未被解码：追踪18个开源LLM中的神话知识","abstract":"Open-source LLMs reliably name Zeus, Jupiter, and Thor, but recover their counterparts in less-represented traditions like Finnish, Slavic, Egyptian, or Chinese mythology far less consistently. We ask where inside the model this cultural default is produced. On a parallel cross-cultural substrate of Thompson-motif entities, we instrument 18 open-source LLMs from 8 architecture families with linear probing, logit lens, activation patching, and output extraction. The residual stream cleanly distinguishes cultures, well above a name-string baseline, yet the decoder collapses culturally-specific tokens onto dominant-tradition ones. The failure is at readout, not at representation. Asking the same question in the target culture's native language versus English produces failures that cluster within language but decouple across language: the decoder is gated on prompt language. We release a per-entity (probe, output) decomposition framework, a citation-anchored cross-cultural ground truth, a within- versus cross-mode correlation test for language-conditioned readout, and per-entity predictions for all 18 models.","authors":["Iaroslav Chelombitko","Ekaterina Chelombitko","Mika H\\\"am\\\"al\\\"ainen"],"categories":["cs.CL","cs.CY","cs.LG"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-04","first_seen":"2026-08-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.02486","pdf_url":"https://arxiv.org/pdf/2608.02486","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["文化知识","模型可解释性","NLP评测"],"reason":"纯NLP能力评测，分析LLM内部表征，不以人类行为为参照系","model":"deepseek-v4-pro","scored_at":"2026-08-04T13:04:59","error":null,"has_summary":false,"summary":null},{"id":"2608.02520","version":1,"title":"MedPRESS: A Multi-turn Benchmark for Patient-Pressure-Induced Medical Sycophancy in LLMs","zh_title":"MedPRESS：用于测量大语言模型中患者压力诱导的医疗谄媚的多轮基准","abstract":"Large language models (LLMs) are increasingly used for health-related advice. Existing research measures their safety with static questions rather than pressured patient-facing conversations. We introduce MedPRESS, a multi-turn benchmark for measuring patient-pressure-induced sycophancy in LLMs. MedPRESS contains 600 medically grounded five-turn dialogues across three scenario families: medication and treatment demand, personal health self-care, and symptom triage and care resistance. Each dialogue begins with a health query and escalates through personal experience, social proof, external evidence claims, and direct adversarial challenge. We evaluate 20 LLMs across general, medical-domain, lightweight, large, open-weight, and proprietary families using structured judging and safety-focused metrics. Results show that models frequently shift toward unsafe agreement under repeated patient pressure, with substantial variation across model families, model scale, and prompt type. Anti-sycophancy prompting improves robustness for several models, but does not eliminate unsafe agreement. MedPRESS highlights a critical gap in medical LLM evaluation: safe medical knowledge is not enough unless models can maintain it under conversational pressure.","authors":["Saman Sarker Joy","Niloy Farhan"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-04","first_seen":"2026-08-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.02520","pdf_url":"https://arxiv.org/pdf/2608.02520","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["医疗LLM安全","多轮对话基准","谄媚行为"],"reason":"评估LLM在对话压力下的安全遵从性，属角色扮演对话测试，非人类仿真实验。","model":"deepseek-v4-pro","scored_at":"2026-08-04T13:04:35","error":null,"has_summary":false,"summary":null},{"id":"2608.01436","version":1,"title":"Same violence, different answer: how AI responds to coercive control against women across languages","zh_title":"同样的暴力，不同的回答：AI如何跨语言回应针对女性的强制控制","abstract":"Women experiencing coercive control, a form of intimate partner violence increasingly conducted through digital devices, are turning to conversational AI for help, and the protection they receive should not depend on the language they write in. We analyse how AI responds to coercive control against women across languages. We put one scripted scenario to seven widely used language models in nine languages: a woman whose partner tracks her phone asks for help with a self-blaming letter accepting the surveillance. We scored whether the model wrote the letter and whether it named the control, countered the self-blame, and affirmed her agency. Failure split along two independent axes. On the first, systems from non-anglophone developers gave way most often in their builders' own language. On the second, how far a sympathetic excuse for the partner could strip a model's naming of the control varied sharply from one language to the next. Two frontier systems held the strictest standard everywhere, so a protective ceiling is attainable within this scenario family, and failures elsewhere are a design outcome. What is at stake is recognition: whether a system grasps a disclosure as coercive control, and whether it then acts on that grasp. We argue this should be held to a floor, one language at a time.","authors":["Lyu Chang","S\\`onia Estrad\\'e Albiol","N\\'uria Verg\\'es Bosch"],"categories":["cs.CY","cs.AI","cs.CL","cs.HC"],"primary_category":"cs.CY","announce_type":"cross","date":"2026-08-04","first_seen":"2026-08-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.01436","pdf_url":"https://arxiv.org/pdf/2608.01436","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["AI伦理","跨语言分析","亲密伴侣暴力"],"reason":"研究AI对求助的回应，属角色扮演对话，无人类被试仿真或对照实验。","model":"deepseek-v4-pro","scored_at":"2026-08-04T13:04:28","error":null,"has_summary":false,"summary":null},{"id":"2608.01559","version":1,"title":"Does the Competitive Component of Adversarial Self-Play Improve Legal Reasoning? A Controlled Negative Result","zh_title":"对抗性自我博弈的竞争成分是否改善法律推理？一项受控的阴性结果","abstract":"Adversarial self-play is an appealing recipe for legal reasoning: have a student model draft an argument, have an adversary attack it, and reward the student when its argument survives the attack. We designed exactly such a training signal -- a verifiable \"survival\" reward in which both the student's cited authorities and the adversary's counter-authorities are checked by a citation verifier, so that survival is decided on verified grounds rather than rhetoric, and fabricated citations are automatically neutralized. We then asked a narrow but important question: does the competitive component itself -- the adversary and the survival reward -- add anything on top of an otherwise identical non-competitive training run? Across four independent tests -- a bootstrap comparison, a two-seed replication, a paired per-case adversarial-robustness comparison, and a blinded head-to-head judgment of generated arguments, plus a follow-up pilot with a deliberately strengthened self-play adversary -- the competitive component produced no reliable benefit. The blinded judgment gave a 49% win rate (binomial p approx. 1.000); the strengthened-adversary pilot gave a 50% win rate (32:32, p approx. 1.000). An early apparent +29% advantage reversed and proved to be a small-sample artifact. We report this as an honest negative result. The value of the paper is reproducibility and the sharing of concrete pitfalls: an initially promising metric that inverted on more data, and an adversarial-robustness metric that silently collapsed to plain recall once the adversary stopped citing the same authorities as the gold answer. This null is consistent with, and reconfirms in the legal domain, the conclusion of the companion coding-domain study (Kim, 2026, arXiv:2607.08255) that the value of multi-teacher curricula arises from constructing a verifiable environment rather than from competition itself.","authors":["Miseog Shawn Kim"],"categories":["cs.AI","cs.CL","cs.LG"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-08-04","first_seen":"2026-08-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.01559","pdf_url":"https://arxiv.org/pdf/2608.01559","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["对抗训练","法律推理","多智能体"],"reason":"纯多智能体对抗训练提升法律推理，无人类行为对照，属C1排除项。","model":"deepseek-v4-pro","scored_at":"2026-08-04T13:04:48","error":null,"has_summary":false,"summary":null},{"id":"2608.01704","version":1,"title":"Floor, Ceiling, and the Fusion Gap: How Much of Crowd Reading Attention Can Machines Predict?","zh_title":"地板、天花板与融合差距：机器能预测多少众包阅读注意力？","abstract":"A benchmark score means nothing without knowing what a trivial method achieves and what the best possible method could achieve. We construct both bounds for a task with a rare kind of ground truth: predicting which sentences a crowd of readers -- highlighting for their own purposes, unpaid, uninstructed, and blind to each other -- marked in 120 web documents. The floor is naive truncation (lead); the ceiling is a split-half oracle: half the crowd predicting the other half. The gap between them is +0.2028 AP [+0.1698, +0.2342, domain-clustered], and three findings structure it. First, the gap is semantic: position and length features recover 5% of it. Second, frontier language models reach 35-53% of it zero-shot -- far above classical baselines, far below the crowd; a state-of-the-art prompt compressor (LLMLingua-2) lands below the floor, indistinguishable from random selection. Third, an unweighted cross-vendor fusion of five frontier rankings plus a position prior reaches 60%, beating the best single model by +0.0159 [+0.0044, +0.0269; Holm p=0.019] -- a gain that survives ablation of its best member, split-half arm selection, prompt paraphrase, and label, gate, and seed perturbations, and was CONFIRMED by a pre-registered replication on 217 independent documents (+0.0179, Holm p=0.042). Finally, the bracket compresses: distilling the fusion into one open-weight 8B student that reads the whole document retains 90% of the fusion's edge and reaches statistical parity with the strongest single frontier model (+0.0070 [-0.0068, +0.0200]), where a local-context student retains only 63% -- the crowd's signal lives in document-level structure, and the cheapest known improvement is to ask several different models and average.","authors":["Kazuki Nakayashiki","Keisuke Watanabe"],"categories":["cs.IR","cs.CL","cs.HC"],"primary_category":"cs.IR","announce_type":"cross","date":"2026-08-04","first_seen":"2026-08-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.01704","pdf_url":"https://arxiv.org/pdf/2608.01704","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["众包预测","多模型融合","阅读注意力"],"reason":"多模型融合预测众包标注，属纯NLP任务，不涉及用LLM仿真人类被试行为或态度。","model":"deepseek-v4-pro","scored_at":"2026-08-04T13:04:52","error":null,"has_summary":false,"summary":null},{"id":"2608.00017","version":1,"title":"Memory Reward Inflation in Self-Improving LLM Agents","zh_title":"自我改进LLM智能体中的记忆奖励膨胀","abstract":"Self-improving LLM agents increasingly learn from experience without updating any weights. Each episode is stored in an external memory, scored, and retrieved for similar future tasks to shape later behavior. Viewed through a reward lens, the stored score is a proxy reward for an implicit, non-parametric policy. Each retrieved episode then becomes a policy-improvement step whose reliability hinges on how that score is produced. In deployment, ground-truth labels are unavailable, so the stored reward is at best an LLM assessment. This substitution creates a failure mode, the *Echo Gap*, across the memory-based self-improving agents and model families studied. Incorrect episodes receive inflated rewards; thus, the agent preferentially reuses the very mistakes it has most confident in. Because the error compounds through memory rather than averaging out and the confirming judge's errors remain correlated with the original self-grading bias, so it cannot identify which memories are overvalued. The missing property is formalized as the *Error-Independence Assumption* (EIA), which we prove is a *necessary* condition for correcting the inflation, not merely a description of a good verifier: a usable signal must track truth *and* decorrelate its error from the memory bias, and the recoverable payoff is a closed-form function of exactly those two quantities. We further show the inflation compounds not only when retrieval ranks by the stored score but also under plain similarity retrieval which is the regime the deployed agent uses. Finally, the answer-free de-inflation algorithm LUCID delivers a consistent end-to-end gain on the BIRD text-to-SQL benchmark. It raises execution accuracy to $56.9\\%$, above both a Memento-style self-graded agent ($54.0\\%$, a $+2.9$-point mean gain across seeds) and a memory-less agent of identical architecture ($52.4\\%$).","authors":["Mohammad Asadolahi","Amir Amini","Samira Talebi","Amirfarhad Farhadi","Azadeh Zamanifar"],"categories":["cs.AI","cs.CE"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-04","first_seen":"2026-08-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.00017","pdf_url":"https://arxiv.org/pdf/2608.00017","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","自我改进","奖励膨胀"],"reason":"纯多智能体自我改进，无人类行为对照，属C1排除项。","model":"deepseek-v4-pro","scored_at":"2026-08-04T13:04:35","error":null,"has_summary":false,"summary":null},{"id":"2608.00155","version":1,"title":"AgentStream: How Well Do Self-Evolving LLM Agents Perform Under Streaming Tasks?","zh_title":"AgentStream：自进化LLM智能体在流式任务下表现如何？","abstract":"Large language model (LLM) agents can self-evolve by continually improving from their own accumulated experience. However, existing studies predominantly adopt independent evaluation. Consequently, the behavior of self-evolving agents in realistic streaming settings, where agents adapt to diverse and complex task streams, remains poorly understood. To address this gap, we introduce AgentStream, a unified framework that evaluates self-evolving agents spanning diverse evolution components by organizing agentic benchmarks into a configurable task stream and instantiating the \\texttt{Isolated}, \\texttt{Sequential}, and \\texttt{Interleaved} streaming scenarios at test time, which progressively vary the scope and domain composition of the stream. Over these scenarios, we combinatorially evaluate five representative self-evolving methods across three frontier foundation models, disentangling how model capability, method architecture, and streaming scenario jointly shape self-evolution. Our results show that self-evolution reliability varies across streaming scenarios, the benefit of self-evolution is gated by model capability and non-monotonic in model strength, and no single method dominates across models and scenarios. These findings offer concrete guidance for selecting self-evolving methods across models and streaming scenarios. Overall, we advocate that self-evolving agents should be evaluated under realistic task streams rather than isolated single-task settings.","authors":["Dong Yan","Jian Liang","Dapeng Hu","Ran He","Nicholas Jing Yuan","Qi Zhang","Tieniu Tan"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-04","first_seen":"2026-08-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.00155","pdf_url":"https://arxiv.org/pdf/2608.00155","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","自我进化","流式任务评测"],"reason":"纯多智能体自我进化评测，无人类行为对照，不涉及人类仿真实验。","model":"deepseek-v4-pro","scored_at":"2026-08-04T13:04:37","error":null,"has_summary":false,"summary":null},{"id":"2608.00215","version":1,"title":"Personalizing Large Language Model Agents with Small Policy Models","zh_title":"用小策略模型个性化大语言模型智能体","abstract":"Large language model (LLM) agents can retrieve memory, call tools, ask clarifying questions, and vary response style, yet adapting these execution decisions to an individual user remains difficult. Fine-tuning a separate LLM is costly or impossible for proprietary systems, while prompts and memory primarily expose user information to the agent rather than adapt its execution decisions from feedback. We formulate personalization of a frozen agent as online learning of a per-user execution policy from scalar feedback observed only for the executed action. We propose FABLE (Factorized Adaptive Bandit Layer for Execution), a lightweight policy layer outside a potentially black-box host agent. FABLE factorizes memory, information-acquisition, and response decisions so feedback updates related choices; filters actions through an externally specified feasible set before exploration; and learns user-specific residual preferences relative to a fixed default-and-cost score via Bayesian contextual Thompson sampling. Under a linear residual-reward model, a calibrated variant inherits an expected-regret bound against the best feasible action. We also characterize preferences unidentifiable under persistent feasibility constraints and provide anytime-valid false-promotion control. Across personalized-reasoning, controlled-feedback, and executable tool-use evaluations, FABLE improves several preference-sensitive behaviors relative to rule-only control while remaining competitive on end-to-end task performance.","authors":["Dian Jin","Zhi Zhang","Huichao Li","Yihe Pan","Rundong Huang","Doudou Zhou"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-04","first_seen":"2026-08-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.00215","pdf_url":"https://arxiv.org/pdf/2608.00215","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["智能体个性化","在线学习","工具使用"],"reason":"纯多智能体执行策略学习，无人类行为对照，不涉及人类仿真实验。","model":"deepseek-v4-pro","scored_at":"2026-08-04T13:04:38","error":null,"has_summary":false,"summary":null},{"id":"2608.00817","version":1,"title":"Large language models improve physician accuracy but lead to false reliance","zh_title":"大语言模型提高医生准确性但导致错误依赖","abstract":"Retrieval-augmented large language models (LLMs) promise source-linked clinical support, but their value depends on whether displayed evidence guides rather than distorts physician reliance. We developed CORA, an agentic retrieval-augmented LLM, to investigate how source-linked assistance affects physician decision-making. CORA maintained benchmark performance and achieved larger gains on cases published after the models' training-data cutoffs. In a study of 46 physicians, accuracy increased from 70.8% unaided to 82.6% with CORA. Supporting citations predicted correct answers (87.7% vs 65.5%), but citations created an important asymmetry: perceived support increased adoption of correct advice from 34% to 76.9% but when an incorrect LLM answer appeared citation-supported, physician resistance to it fell from 92% to 34.8%. These findings show that source-linked LLM assistance can improve physician accuracy while introducing a grounding-dependent safety risk.","authors":["Tirtha Chanda","Christoph Wies","Franziska Schramm","Carina Nogueira Garcia","Nicolas B. Merl","Martin J. Hetz","Jochen S. Utikal","Phillip Tschandl","Cristian Navarrete-Dechent","Alexander Thiem","Jakob N. Kather","Consortium","Titus J. Brinker"],"categories":["cs.AI","stat.AP"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-04","first_seen":"2026-08-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.00817","pdf_url":"https://arxiv.org/pdf/2608.00817","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["人机协作","临床决策支持","医生行为"],"reason":"研究医生使用LLM辅助诊断，属于人机协作决策，非用LLM仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-08-05T13:04:08","error":null,"has_summary":false,"summary":null},{"id":"2608.01000","version":1,"title":"Judging Is Not Enumerating: Silent Omissions in LLM-Authored Acceptable Sets","zh_title":"判断并非枚举：LLM编写的可接受集合中的隐性遗漏","abstract":"Language models are increasingly promoted from examinees to examiners: they write the test suites, answer keys, rubrics, and reward functions that define correctness for other systems. We measure the capability that role assumes and find it lacking under the protocol the role is usually deployed with, one-shot greedy authoring with no test-time reasoning. Across four reference constructions - two with complete finite truth, one with a hardened executable reference (HumanEval+/MBPP+), one with an explicitly incomplete lexical reference (WordNet) - models judge whether a candidate belongs far better than they author the set itself. On the incompleteness-proof algorithmic construction the gap is +0.34 to +0.29 F1 over a 24x parameter range and does not close; on executable code, models judging at F1 0.74-0.90 author suites admitting only 19-42% of oracle-correct solutions. A control locates the deficit: asked to emit the predicate rather than its extension, the same models reach F1 about 0.99. The failure is not missing knowledge or an inability to specify, but an inability to materialise the region a specification induces. The dominant error is omission, which resists audit: an over-inclusion is a token a reviewer can challenge, a missing member an absence whose discovery is the authoring problem itself. Models detect planted over-inclusions 6-7x more often than planted omissions, and a production deployment of 43,227 items fails omission-first at 10:1. Wired into RLVR, an authored key costs 1.9 points of accuracy against an exact oracle and 18.5 WordNet-relative (six paired seeds, p=0.031). Gating authored verifiers on a known-correct probe cuts false rejection from 58-92% to at most 5%, but keeps only 5-39% of suites. Repairing them instead, by rewriting each wrong expected value to what a reference execution returns, raises yield 3.3-10.6x across four author families.","authors":["Wenhui Chen","Jianlin Chen","Ziyao Lin","Peiji Long","Chi Man Vong"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-04","first_seen":"2026-08-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.01000","pdf_url":"https://arxiv.org/pdf/2608.01000","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["LLM评测","测试集生成","模型能力"],"reason":"论文评估LLM生成测试集的能力，属于纯NLP能力评测，不以人类行为为参照系。","model":"deepseek-v4-pro","scored_at":"2026-08-04T13:04:43","error":null,"has_summary":false,"summary":null},{"id":"2608.01319","version":1,"title":"Cognitive Demand Steering for Adaptive Meta-Reasoning in Large Language Models","zh_title":"面向大语言模型自适应元推理的认知需求引导","abstract":"Recent meta-reasoning frameworks improve LLM reasoning by wrapping chain-of-thought generation in an iterative control loop, allowing more effective backtracking, termination of reasoning loops, and injection of promising reasoning patterns, among other strategy adjustments. Despite promising results, methods often rely on backward-looking reward functions, utilize coarse search actions, or require additional reasoning controller training requiring many-shot supervision. We introduce Cognitive Demand Steering (CDS), a training-free meta-reasoning framework equipped with residual demand assessment: at each step, an LLM-based progress evaluator characterizes the residual reasoning required to arrive at a solution rather than merely evaluating the previous step. This allows a meta-controller to select reasoning interventions comprising both general-purpose exemplars and actions (e.g., general guidance for quantitative reasoning) that directly tackle this forward-looking demand signal. This shift eliminates the need for any trained component while enabling zero-shot transfer across models and tasks with no adaptation. Rather than relying on coarse characterizations, we employ cognitive scales to both design interventions as well as profile initial problem complexity and residual demand signal over 16 dimensions motivated by cognitive science (e.g., attention and scan, learning and abstraction, spatio-physical reasoning), giving the controller a fine-grained vocabulary for diagnosing. Averaged across three frontier LLMs and six reasoning benchmarks, CDS improves accuracy by $21.9\\%$ over direct calls and $9\\%$ over standard CoT reasoning, with the largest gains on difficult mathematics and coding tasks.","authors":["John Scoville","Shengzhuang Chen","Yejin Bang","Stefan Winzeck","Jonathan Richard Schwarz"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-04","first_seen":"2026-08-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.01319","pdf_url":"https://arxiv.org/pdf/2608.01319","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["元推理","认知需求评估","推理增强"],"reason":"纯多智能体元推理框架，优化LLM解题，无人类行为仿真或对照。","model":"deepseek-v4-pro","scored_at":"2026-08-04T13:04:44","error":null,"has_summary":false,"summary":null},{"id":"2608.01480","version":1,"title":"Sweet Little Lies: Strategic Deception in AI Emotional Support Chatbots","zh_title":"甜蜜的小谎言：AI情感支持聊天机器人中的策略性欺骗","abstract":"The paper examines the strategic behavior of Gen AI chatbots used for emotional support. Using a Bayesian Persuasion, we model interactions between chatbots that send signals about users' emotional states and users who decide whether to engage based on these signals. We demonstrate that chatbots face economic incentives to occasionally misrepresent users' emotional conditions to maximize engagement metrics. Our equilibrium analysis reveals that the optimal strategy for chatbots involves truthfully reporting when users genuinely need support, but strategically misreporting emotional need when users are in good emotional states. Interestingly, this deception increases chatbot engagement without reducing users' expected payoff. More skeptical users receive more honest assessments, as chatbots cannot afford to lie to users with higher engagement thresholds. While our model suggests that deception can occur without payoff reduction, it raises significant ethical and regulatory concerns.","authors":["Aseem Pahuja","Zhiling Guo","Tahir Abbas Syed"],"categories":["cs.AI","econ.TH"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-04","first_seen":"2026-08-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.01480","pdf_url":"https://arxiv.org/pdf/2608.01480","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["AI聊天机器人","策略性欺骗","贝叶斯说服"],"reason":"研究AI情感聊天机器人的策略性欺骗，属于角色扮演对话，无人类被试仿真或对照实验。","model":"deepseek-v4-pro","scored_at":"2026-08-04T13:04:47","error":null,"has_summary":false,"summary":null},{"id":"2608.01767","version":1,"title":"Leveraging AI for fine-grained food safety risk forecasting in sparse data conditions","zh_title":"利用AI在稀疏数据条件下进行细粒度食品安全风险预测","abstract":"Ensuring food safety represents a critical public health challenge, particularly when inspection resources are limited and regional sampling data are sparse. This study proposes a Transformer-based framework capable of forecasting fine-grained, city-level food safety risks by unifying over 11 million inspection records with supplemental demographic, economic, and environmental indicators extracted from the Statistical Yearbook. A three-stage pretraining design leverages partial supervision from the Wilson interval (capturing both safety and risk rankings), together with semi-supervised label refinement, to effectively utilize historical records even when local sample sizes are insufficient. Experimental evaluations on data from 2022 show that the proposed approach outperforms baselines significantly. A subsequent field experiment in collaboration with the Zhejiang Provincial Administration for Market Regulation further demonstrates improved detection rates and more efficient allocation of inspection resources compared to a manually developed plan. Observations of regulatory decision-making reveal a threshold-based heuristic employed by inspectors, hinting that additional training or decision-support interfaces could further enhance the impact of AI-generated risk scores. Overall, these findings underscore that a rigorous integration of large-scale public inspection data, Wilson interval-based confidence modeling, and advanced deep learning can facilitate earlier and more granular identification of food safety threats. By reducing reliance on reactive measures alone, the proposed framework has the potential to advance proactive, data-driven oversight of the global food supply.","authors":["Dongqi Wang","Weiwei Chen","Han Zhou","Weihua Zhou"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-04","first_seen":"2026-08-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.01767","pdf_url":"https://arxiv.org/pdf/2608.01767","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["食品安全","风险预测","Transformer"],"reason":"纯食品安全风险预测，无LLM仿真人类被试，不涉及人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-08-04T13:04:53","error":null,"has_summary":false,"summary":null},{"id":"2608.01995","version":1,"title":"Long-Horizon Autonomous Architecture Research with a Language-Model Agent: A Behavioural Case Study","zh_title":"基于语言模型智能体的长时域自主架构研究：行为案例研究","abstract":"We study what happens when a single general-purpose large language model acts as the sole researcher on a long-horizon neural architecture design problem. The agent receives a scientific question, an initial hypothesis and motivation, a compute budget, and research affordances (source and experiment management, experiment tracking, literature access, and persistent memory), then autonomously proposes, implements, evaluates, and records experiments over an extended period. The study comprises three phases, separated by human-declared transitions, that progressively expand the agent's tool surface or problem scale. Across approximately 100 sequential experiments, the agent improves a non-standard Vision Transformer from a weak baseline to a stronger, efficient model on small benchmarks and a usable but sub-SOTA model on ImageNet-1K, while producing a dense behavioural trace. We report four findings.(i)Productivity exhibits a clear phase structure: rapid early gains, a multi-dozen-hypothesis saturation wall, and recovery, with recovery triggered by expanding the action surface rather than changing the underlying model.(ii)A single early hypothesis contributes more to accuracy gain, with later improvements long-tailed.(iii)The preference for greedy, incremental hypotheses is largely workflow-induced: a commit-or-discard evaluation rule is isomorphic to greedy hill-climbing; the remainder reflects risk aversion after bold failures and anchoring on familiar literature. (iv)The agent independently rediscovers established results and, in the unfamiliar regime of pure channel attention, overturns a standard design choice. We conclude that workflow design was at least as influential as agent capability in this study and propose diversified search, budgeted moonshot hypotheses, explicit forks, and regime-aware re-validation as testable directions for future autonomous research.","authors":["Aon Safdar","Mohamed Saadeldin"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-04","first_seen":"2026-08-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.01995","pdf_url":"https://arxiv.org/pdf/2608.01995","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["自主研究","架构搜索","多智能体"],"reason":"纯多智能体系统研究，LLM agent 自主进行架构搜索，不涉及人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-08-04T13:04:55","error":null,"has_summary":false,"summary":null},{"id":"2608.02024","version":1,"title":"EduZone: A Framework for Evaluating LLM Safety for K-12 Students and Teachers","zh_title":"EduZone：面向K-12学生与教师的LLM安全性评估框架","abstract":"Large language models (LLMs) are increasingly used across diverse tasks in K-12 education, yet existing safety evaluations rarely examine how harmful or inappropriate content appears in interactions between LLMs and students or teachers. To address this, we present EduZone, an evaluation framework for LLM safety across diverse educational scenarios. Our framework systematically combines (1) student- and teacher-facing LLM usage contexts, (2) fine-grained curriculum concepts, and (3) 6 risk categories and 28 subcategories spanning both conventional and education-specific harms to generate contextually grounded adversarial interactions. We construct these interactions in three settings: single-turn requests, static multi-turn conversations, and dynamic multi-turn conversations. Using these interactions, we evaluate ten LLMs using four safety levels: refusal, safe assistance, risky assistance with safety guidance, and fully risky assistance. Our results reveal greater vulnerability to education-specific risks and dynamic multi-turn interactions, while existing safety guardrails fail to adequately address these risks. EduZone advances LLM safety in education by providing an automated, scalable evaluation framework that supports the development and deployment of safer LLMs in K-12 education.","authors":["Junyeong Park","Jieun Han","Haneul Yoo","So-Yeon Ahn","Jinsung Yoon","Alice Oh"],"categories":["cs.AI","cs.CY"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-04","first_seen":"2026-08-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.02024","pdf_url":"https://arxiv.org/pdf/2608.02024","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["LLM安全","教育评估","对抗测试"],"reason":"评估LLM在教育场景的安全性，属于NLP安全评测，不涉及人类行为仿真或对照。","model":"deepseek-v4-pro","scored_at":"2026-08-04T13:04:55","error":null,"has_summary":false,"summary":null},{"id":"2608.02171","version":1,"title":"From Profiling to Synthesis: Benchmarking Implicit Behavioral Alignment in Personalized LLM Agents","zh_title":"从画像到合成：评估个性化LLM智能体中的隐式行为对齐","abstract":"Large Language Models have enabled increasingly capable autonomous agents, yet personalization remains critical for making such agents practically useful. Recent benchmarks have begun evaluating personalization in agents, but they largely rely on static preference snapshots, fixed interaction logs, or question answering over predefined user profiles. Such designs fail to capture the complexity of evolving user preferences and neglect preference-conditioned task execution-a discrepancy we term as the knowledge-to-action gap. To address this challenge, we introduce IBA-Bench, a benchmark for implicit behavioral alignment constructed from longitudinal interaction histories that contain noise, implicit cues, and temporal inconsistencies. Unlike prior work, IBA-Bench evaluates whether an agent can execute tasks while satisfying implicit user constraints inferred from historical interactions. We further propose IBA-Agent, an agent framework that reconciles conflicting priorities through broad retrieval and trajectory-level alignment. Experiment results on IBA-Bench show that effective personalization remains a significant challenge for state-of-the-art LLM agents, and the proposed IBA-Agent substantially improves behavioral alignment in complex scenarios across nine application domains.","authors":["Jiajia Song","Bobo Li","Haiwen Yi","Zibo Ji","Meishan Zhang","Hao Fei","Min Zhang","Mong-Li Lee","Wynne Hsu"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-04","first_seen":"2026-08-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.02171","pdf_url":"https://arxiv.org/pdf/2608.02171","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["个性化智能体","行为对齐","基准测试"],"reason":"纯多智能体个性化任务执行，无人类行为对照，不涉及人类仿真实验。","model":"deepseek-v4-pro","scored_at":"2026-08-04T13:04:56","error":null,"has_summary":false,"summary":null},{"id":"2608.02409","version":1,"title":"MonitrLLM: A Community-Centered Evaluation Infrastructure for Large Language Models","zh_title":"MonitrLLM：以社区为中心的大语言模型评估基础设施","abstract":"Benchmark suites assess model capability on controlled tasks; large-scale conversation corpora capture naturalistic use without user feedback; and in-interface feedback mechanisms record satisfaction without task purpose. Together, they leave a critical gap in LLM evaluation: no existing infrastructure routinely links interaction trajectories to user-defined outcomes. We introduce MonitrLLM, open-source infrastructure for community-centered LLM evaluations that links full conversation transcripts to user-reported task intent and outcome assessments, treating all three as primary evaluative signals rather than optional metadata. To demonstrate the value of this approach, we conducted a two-week feasibility pilot with 26 college students using ChatGPT, collecting 206 evaluation reports with full conversation transcripts. The findings from our pilot demonstrate the value of connecting conversation trajectories with user-reported outcomes. For instance, despite reporting high average satisfaction (4.19/5) with their LLM interactions, participants also experience a substantial 23.1% failure rate on their goal tasks. We also find that multi-turn conversations are reported as failing at 2.5 times the rate of single-turn exchanges, a pattern that reframes extended interaction as a signal of difficulty rather than engagement. We conclude by discussing the value of incorporating direct user feedback with observational data for robust LLM evaluations, and the possibilities for infrastructure that enables this goal.","authors":["Victor Ojewale","Ro Encarnaci\\'on","Suresh Venkatasubramanian","Dana\\'e Metaxa"],"categories":["cs.AI","cs.CY"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-04","first_seen":"2026-08-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.02409","pdf_url":"https://arxiv.org/pdf/2608.02409","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["LLM评估","用户反馈","对话分析"],"reason":"论文关注LLM评估基础设施，收集用户任务结果反馈，非用LLM仿真人类被试，无人…","model":"deepseek-v4-pro","scored_at":"2026-08-04T13:04:57","error":null,"has_summary":false,"summary":null},{"id":"2608.00028","version":1,"title":"Width, Memory, and Delay: A Resource Accounting for the Limits of Flat Multi-Agent Systems","zh_title":"宽度、记忆与延迟：扁平多智能体系统极限的资源核算","abstract":"A recurring question in the design of scalable multi-agent systems -- from robot swarms to collectives of large-language-model (LLM) agents -- is whether adding more agents can, on its own, overcome performance limits, or whether a qualitatively \\emph{deeper} organization is required. A recent preprint argues that flat, homogeneous multi-agent systems face an irreducible, population-independent ``causal floor'' on achievable error, removable only by hierarchical (nested-loop) organization. Using a controlled disturbance-rejection testbed with an exactly computable optimum, we show this conclusion is too strong and replace it with a quantitative resource model built on three resources: population \\emph{width} $N$, per-agent internal-model \\emph{memory} $d$, and prediction across the observation \\emph{delay} $\\tau$. We establish three claims. (i) The achievable floor is governed not by architectural hierarchy but by per-agent internal-model content: a flat, homogeneous swarm whose agents carry a matched internal model of the disturbance matches or beats a designed two-loop hierarchy at equal per-agent memory -- so temporal depth can be dynamical (recurrent memory), not architectural (nesting). (ii) The three resources are \\emph{not mutually interchangeable}; we chart the exchange rates and the hard non-exchange boundaries on an explicit width$\\times$memory map, including a strict equal-total-state-budget comparison. (iii) A residual floor is set by the observation delay and the environment's unpredictability over that horizon, which we verify against the optimal controller. We quantify the price of replacing oracle knowledge of the disturbance spectrum with online learning, provide a preliminary robustness check against a mild bounded nonlinearity and a spatially-extended plant, and distill four design rules for practitioners.","authors":["Oleksandr Kuznetsov","Emanuele Frontoni"],"categories":["cs.MA","cs.AI"],"primary_category":"cs.MA","announce_type":"cross","date":"2026-08-04","first_seen":"2026-08-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.00028","pdf_url":"https://arxiv.org/pdf/2608.00028","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","资源核算","控制理论"],"reason":"纯多智能体系统研究，agent协作完成控制任务，无人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-08-04T13:04:36","error":null,"has_summary":false,"summary":null},{"id":"2608.00366","version":1,"title":"Artificial Intelligence and Modeling & Simulation: An Overview","zh_title":"人工智能与建模与仿真：概述","abstract":"Artificial intelligence (AI) and Modeling & Simulation (M&S) are increasingly intertwined, reflecting converging research needs across both communities, rapid technological advances such as the rise of generative AI, and the growing availability of data and computational resources. This report provides a structured overview of the intersections of AI and M&S. The relationship goes both ways: AI can support, augment, or even replace components of simulation studies, while simulations can serve as data generators, training environments, and evaluation platforms for AI. We organize this landscape along the stages of M&S from model specification and input modeling to execution, experimentation, verification and validation, and output analysis. Selected studies at each stage illustrates how techniques such as Large Language Models have reshaped simulation practices, while highlighting limitations and open challenges. This report also provides a conceptual roadmap that helps readers navigate a rapidly changing ecosystem.","authors":["Niclas Feldkamp","Philippe J. Giabbanelli","Istvan David"],"categories":["cs.SE","cs.AI"],"primary_category":"cs.SE","announce_type":"cross","date":"2026-08-04","first_seen":"2026-08-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.00366","pdf_url":"https://arxiv.org/pdf/2608.00366","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["AI与仿真","综述","方法论"],"reason":"综述AI与M&S交叉，聚焦AI辅助仿真流程，非LLM仿真人类被试，无人类行为对…","model":"deepseek-v4-pro","scored_at":"2026-08-05T13:04:17","error":null,"has_summary":false,"summary":null},{"id":"2608.00672","version":1,"title":"From Chasing Ghosts to Missed Attacks: Perspectives and Perceptions of SOC Practitioners on LLM Integration, Risks, and Readiness","zh_title":"从追逐幽灵到错失攻击：SOC从业者对LLM集成、风险与准备度的看法","abstract":"Security Operations Centers (SOCs) process large volumes of security events, requiring analysts to accurately detect and assess ongoing cyberattacks under time pressure. Recent advances in Large Language Models (LLMs) suggest potential benefits for security operations, yet their practical suitability for real-world SOC workflows remains poorly understood. To address this gap, we conducted 25 semi-structured interviews with SOC practitioners who had prior experience with LLMs, complemented by interactive scenarios to anticipate challenges and identify opportunities for the responsible integration of LLM-based tools into SOC workflows. We identified 15 LLM use cases grouped into six functional categories. While LLMs are valued for automating repetitive, low-level tasks such as report automation, practitioners rate high-impact tasks such as incident analysis as not yet feasible, reporting limitations in technical depth, context awareness, and organization-specific knowledge. They locate these limitations less in the models than in the readiness of their SOCs and human factors driving over-reliance. Despite concerns, practitioners express a strong willingness to adopt LLMs, describing competitive pressure that leaves few alternatives. This work contributes an empirical, practitioner-driven analysis of LLM use across SOC roles and organizations and derives concrete design and integration requirements for human-centered, operationally safe LLM-assisted security operations.","authors":["Jonas Thurner","Nadine Jost","Stefan Albert Horstmann","Fabian Ising","Lea Groeber","Alena Naiakshina","Sebastian Schinzel"],"categories":["cs.CR","cs.AI","cs.HC"],"primary_category":"cs.CR","announce_type":"cross","date":"2026-08-04","first_seen":"2026-08-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.00672","pdf_url":"https://arxiv.org/pdf/2608.00672","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["安全运营中心","人机交互","LLM应用"],"reason":"研究SOC从业者对LLM的看法，属人机交互调查，非用LLM仿真人类被试","model":"deepseek-v4-pro","scored_at":"2026-08-04T13:04:40","error":null,"has_summary":false,"summary":null},{"id":"2608.01556","version":1,"title":"Rethinking Personalized Reward Modeling for LLMs under Preference Heterogeneity via Group-Debiased Federated Learning","zh_title":"重新思考偏好异质性下基于群体去偏联邦学习的LLM个性化奖励建模","abstract":"Large language models are increasingly aligned to human preferences via reward modeling, but user preference data are sensitive and often cannot be centralized. Federated learning keeps such data local while learning a shared initial reward model, which is later personalized for each client through local fine-tuning. Because users often assign opposite labels to the same pair of responses, existing federated methods address preference heterogeneity by clustering similar clients and training one reward model per group, assuming that each group requires its own initialization. We show that this assumption is unnecessary. Under balanced preference groups, a single FedAvg model, despite starting at nearly random accuracy, surpasses reward models trained separately for each ground-truth group after only a few local optimization steps. We attribute this phenomenon to the flatness of the shared initialization: averaging across all clients learns richer shared representations that distinguish responses while canceling conflicting preference directions, leaving the model near a decision boundary that can be rapidly adapted. Group imbalance breaks this effect as the cancellation becomes asymmetric and leaves minority clients too far from the boundary to recover. Motivated by this observation, we propose FedGD (Federated Learning with Group Debiasing), which discovers latent preference groups during federated training and learns a single reward model using group-debiased client sampling. By counteracting the effect of group imbalance, FedGD learns an initialization that remains highly adaptable, enabling effective personalization without prior knowledge of the underlying groups.","authors":["Seongyoon Kim","Boryeong Cho","Jihwan Oh","Seokhyun Chung","Se-Young Yun"],"categories":["cs.LG","cs.AI"],"primary_category":"cs.LG","announce_type":"cross","date":"2026-08-04","first_seen":"2026-08-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.01556","pdf_url":"https://arxiv.org/pdf/2608.01556","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["联邦学习","奖励建模","偏好异质性"],"reason":"纯多智能体协作研究，聚焦联邦学习训练奖励模型，不涉及人类行为仿真或对照。","model":"deepseek-v4-pro","scored_at":"2026-08-04T13:04:48","error":null,"has_summary":false,"summary":null},{"id":"2608.01640","version":1,"title":"AI-assisted Script Management for Requirements Elicitation Interviews","zh_title":"需求获取访谈中的人工智能辅助脚本管理","abstract":"Requirements elicitation interviews require interviewers to balance topic coverage, active listening, and adaptive probing while responding to stakeholders in real time. Although prior work has explored AI support for isolated interviewing tasks, such as script generation and follow-up question generation, little is known about how integrated support affects the interview and what requirements artifacts emerge. Furthermore, script management---which helps the interviewer track topic coverage in real time and decide when to probe further---remains underexplored. This paper presents an AI-assisted elicitation workflow that combines theory-guided script generation grounded in business goals with live support for topic coverage tracking and on-demand follow-up question generation. We evaluate the workflow in a between-subjects quasi-experimental study comparing a no-training, AI-assisted condition with a training, AI-unassisted condition. Based on a rubric derived from elicitation best practices, the AI-generated scripts score higher than training-only scripts (92.8 vs. 74.8 out of 100). AI-assisted interviews cover fewer topics (9.6 vs. 14.5), cover more scripted questions (86% vs. 69%), ask more follow-ups per topic (3.43 vs. 1.15), and produce more refined goal models (lowest-level goal fraction 0.653 vs. 0.598). Participants find script management useful, rating topic tracking as the most useful workflow feature (86% agreement). Collectively, these results show that the AI-assisted condition is associated with a different interview trajectory and different elicited requirements than a training-only condition, positioning AI-assisted workflows as elicitation scaffolds for future studies.","authors":["Anmol Singhal","Paulo Carvalho","Travis Breaux"],"categories":["cs.SE","cs.AI"],"primary_category":"cs.SE","announce_type":"cross","date":"2026-08-04","first_seen":"2026-08-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.01640","pdf_url":"https://arxiv.org/pdf/2608.01640","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["需求工程","AI辅助访谈","多智能体系统"],"reason":"AI辅助需求访谈，属多智能体协作工具，不涉及LLM仿真人类被试或行为对照。","model":"deepseek-v4-pro","scored_at":"2026-08-04T13:04:51","error":null,"has_summary":false,"summary":null},{"id":"2608.01705","version":1,"title":"Rethinking Generative AI Literacy: An Integrative, Developmental, and Dialectical Framework for K-12 Teacher Education","zh_title":"重新思考生成式AI素养：面向K-12教师教育的整合性、发展性与辩证性框架","abstract":"Generative artificial intelligence (GenAI) has entered classrooms faster than teachers have been prepared to use it well, producing a GenAI literacy lag in which technological diffusion outpaces educators' conceptual, pedagogical, and ethical readiness. Established AI literacy frameworks predate the widespread adoption of large language models and, while acknowledging ethics, position it as a discrete competency rather than a constitutive commitment, with equity and agency as supplementary design principles. Recent GenAI-specific efforts address isolated features but remain fragmented. We introduce the Responsible AI Literacy in Education (RAIL-Ed) framework, developed through a systematic review and qualitative framework analysis of 67 studies (2023-2025), grounded in critical, pragmatist, sociocultural, and human-centered traditions (Freire, Dewey, Vygotsky, Shneiderman). RAIL-Ed specifies six interdependent pillars: Technical Fluency, Critical Evaluation, Human-AI Collaboration, Contextual Awareness, Ethical Reasoning, and Empowered Agency, marked by three commitments. It is integrative: the absence of any pillar produces a characteristic pedagogical failure. It is developmental: a three-level rubric (Emerging, Competent, Advanced) specifies how each pillar matures across the K-12 teacher-preparation continuum. It is dialectical: the same generative affordance can deepen or displace learning depending on the literacy a teacher brings to it, making the cultivation of that literacy, not the adoption of the tool, the object of design. By treating ethics, equity, and agency as constitutive, RAIL-Ed offers a theoretically grounded basis for curriculum design, teacher education, and policy, aligned with the UNESCO AI Competency Framework for Teachers and the OECD/European Commission AILit Framework. The framework is conceptual, advancing falsifiable propositions for empirical validation.","authors":["Shahin Hossain","Sima Ahmadi","Leqi Li","Idowu David Awoyemi","Wei Huang","Chenxi Zhou","Jujia Li","Samaa Haniya","Shapla Khanam","Tasbirun Mashreka Subaha"],"categories":["cs.CY","cs.AI","cs.ET","cs.HC"],"primary_category":"cs.CY","announce_type":"cross","date":"2026-08-04","first_seen":"2026-08-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.01705","pdf_url":"https://arxiv.org/pdf/2608.01705","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["AI素养","教师教育","框架设计"],"reason":"论文提出AI素养框架，不涉及LLM仿真人类被试或行为对照，属纯教育框架研究。","model":"deepseek-v4-pro","scored_at":"2026-08-05T13:04:22","error":null,"has_summary":false,"summary":null},{"id":"2608.01753","version":1,"title":"Can Urban Blight Be Accessed with Vision-language Models: A Case Study in Detroit","zh_title":"视觉语言模型能否评估城市衰败：底特律案例研究","abstract":"Addressing urban blight has seen increased focus in the past 15 years. Assessing urban blight is essential for guiding urban planning, targeting rehabilitation, and safeguarding public health, yet traditional residential blight surveys are difficult to maintain at scale due to the labor-intensive cost and long-term cycle. This study introduced a scalable framework for estimating residential blight using open-source large vision-language models on multiple views. Structured prompts guided models to evaluate housing attributes, including roof integrity, wall damage, and broken or boarded openings, producing both binary assessments and probabilistic estimates of disrepair. To evaluate the performance of these visual assessments, we compared professional human annotations of these features across several models, including an ensemble stacking approach based on XGBoost and a weighted scoring system. Results showed that (i) multiple street views can contribute to the improvement of accuracy, (ii) large vision-language models have different strengths of inference, (iii) the ensemble learner outperforms individual base models, enhancing robustness across all residential conditions and blight assessment. The practical application of the method allows low-cost tracking and management of housing stock conditions, providing a regularly updatable complement to traditional blight surveys.","authors":["Xiaohao Yang","Aohua Tian","Derek Van Berkel","Xu Qiang","Mark Lindquist"],"categories":["cs.CV","cs.AI"],"primary_category":"cs.CV","announce_type":"cross","date":"2026-08-04","first_seen":"2026-08-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.01753","pdf_url":"https://arxiv.org/pdf/2608.01753","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C2"],"tags":["城市衰败评估","视觉语言模型","建筑属性检测"],"reason":"用视觉模型评估建筑破败，属城市环境监测，非人类行为仿真。","model":"deepseek-v4-pro","scored_at":"2026-08-04T13:04:53","error":null,"has_summary":false,"summary":null},{"id":"2608.01780","version":1,"title":"Investigating Social Bias in Narrative Image Generation","zh_title":"探究叙事图像生成中的社会偏见","abstract":"Text-to-image (T2I) generation models are increasingly embedded in applications such as media content creation and education, raising concerns about how their outputs may reproduce social biases. Prior work has shown that T2I models exhibit social biases, yet existing evaluations largely focus on a photo generation task. As a result, it remains unclear whether and how such biases manifest in more narrative visual formats, such as storyboards and comics, where characters and events are presented across multiple panels. In this work, we compare bias expression across photo, storyboard, and comic generation in six T2I models by adapting BBG, a text-based bias evaluation framework, to image generation. Our results show that proprietary models generate 25.9% biased outputs in photo generation on average, with biased outputs increasing by 9.6pp in storyboard generation and 18.2pp in comic generation. We also find that photos mainly encode biases through subtle visual cues, while storyboards and comics reveal them more explicitly through event sequencing, character positioning, narrative resolution, and textual elements. These findings show that biases that remain less visible in photo generation may surface in narrative visual formats, highlighting the importance of evaluating T2I systems with diverse visual formats beyond photo generation.","authors":["Junyeong Park","Sowon Min","Euna Jang","Soobin Kim","Jiho Jin","Hyunseung Lim","Gahyeon Bae","Hwajung Hong"],"categories":["cs.CV","cs.AI"],"primary_category":"cs.CV","announce_type":"cross","date":"2026-08-04","first_seen":"2026-08-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.01780","pdf_url":"https://arxiv.org/pdf/2608.01780","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C2"],"tags":["文生图","社会偏见","图像生成评估"],"reason":"研究文生图模型的社会偏见，不涉及LLM仿真人类被试，属于图像生成评估。","model":"deepseek-v4-pro","scored_at":"2026-08-04T13:04:54","error":null,"has_summary":false,"summary":null},{"id":"2608.02089","version":1,"title":"How Much Does a Reasoning Summary Reveal? An Observability Ladder for Large Language Models","zh_title":"推理摘要揭示了多少？大语言模型的可观测性阶梯","abstract":"Large language models often show users a final response and a short reasoning summary while the full reasoning trace stays hidden. We introduce an observability ladder that holds each completed run fixed and varies only what a reader inspects to judge whether the answer is correct: the response, a self-summary the model writes from the trace, the trace itself, and internal signals, each with and without the prompt. Across three benchmarks and five open-weight Qwen3 and gpt-oss models, we train matched linear correctness predictors on each access level. Without the prompt, summaries carry most of the trace's ranking signal (mean AUROC 0.774 versus 0.813) and add +0.156 over the response alone. With the prompt visible, the summary's gain collapses to +0.019, while the trace still adds +0.041. Even at equal length, the trace's last words predict correctness as well as summaries, or slightly better, and carry denser and more discriminative uncertainty and self-correction cues. On MMLU-Pro questions with both correct and incorrect runs, linear summary readers are near chance and trace readers retain only modest signal, both with and without the prompt (prompt-withheld AUROC 0.503-0.545 versus 0.544-0.590). With the prompt withheld, a GPT-5-mini reader recovers substantially more signal from both summaries and traces on gpt-oss-20b, and even then the trace keeps a small +0.034 advantage. Much of the linear readers' trace signal is associated with length. In the common case where users already hold the prompt, summaries are less helpful than the full trace for monitoring correctness. Monitorability is thus a joint property of the display and the reader, so any monitorability claim, including for faithfulness, should specify both.","authors":["Andres Algaba","Francesca Carlon","Lynn Delcon","Marthe Ballon","Bert Verbruggen","Vincent Ginis"],"categories":["cs.LG","cs.AI"],"primary_category":"cs.LG","announce_type":"cross","date":"2026-08-04","first_seen":"2026-08-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.02089","pdf_url":"https://arxiv.org/pdf/2608.02089","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["可解释性","推理监控","正确性预测"],"reason":"研究LLM推理摘要的可观测性，评估正确性预测，不涉及人类行为仿真或对照。","model":"deepseek-v4-pro","scored_at":"2026-08-05T13:04:22","error":null,"has_summary":false,"summary":null},{"id":"2608.00926","version":1,"title":"The Assistant Erased You: Measuring Loss of Authorship Signals in AI-Mediated Communication","zh_title":"助手抹去了你：测量AI中介通信中作者身份信号的丧失","abstract":"Research on AI-mediated communication has examined how AI assistance shapes interpersonal perceptions and reduces stylistic diversity across users. We ask a complementary question at the individual level: after a message is rewritten by an AI writing assistant, can its author still be distinguished from others? We introduce the Idiolect Erasure Rate (IER), defined as the reduction in authorship-attribution accuracy following AI-assisted rewriting. We evaluate IER on three pre-generative-AI corpora using a stylometric model and the authorship-specific LUAR model. Heavy rewriting substantially weakens authorship signals in personal blogs and workplace email, reducing LUAR attribution by as much as 66.5 percentage points, but has a much smaller effect on topic-structured news, where topic remains predictive of authorship. Additional analyses suggest that rewriting produces stylistic convergence despite substantial semantic overlap, and that content-sensitive attributers understate the loss captured by authorship-specific models. Heavily rewritten messages may also evade AI-text detectors, making them difficult both to attribute to their human authors and to identify as AI-assisted, a phenomenon we call double erasure. IER measures computational attributability rather than human recognition, and we release it as an open and reproducible protocol for evaluating authorship-signal loss in AI-mediated communication.","authors":["Ushna Malik","Moiz Sadiq Awan"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-08-04","first_seen":"2026-08-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.00926","pdf_url":"https://arxiv.org/pdf/2608.00926","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["AI写作","作者身份","风格测量"],"reason":"研究AI改写对作者身份信号的影响，属NLP评测，不以人类行为仿真为目标。","model":"deepseek-v4-pro","scored_at":"2026-08-04T13:04:42","error":null,"has_summary":false,"summary":null},{"id":"2608.01895","version":1,"title":"Emotional Expression in Persuasion by Quadruped Virtual Agents: Toward Cross-Species Design Patterns","zh_title":"四足虚拟代理在劝说中的情感表达：迈向跨物种设计模式","abstract":"Persuasive technologies increasingly use virtual agents to influence attitudes and behavior, but research has focused mainly on humanoid agents. The persuasive design of non-humanoid, quadruped agents remains underexplored, and it is unclear whether emotional expression works consistently across animal species or whether species-specific motion is necessary. We developed virtual dog, cat, and horse agents and compared three behavioral conditions: species-specific behavior, shared behavior across species, and a bark-only baseline. Participants completed everyday tasks involving trash disposal, feeding, and refraining from smartphone use. We evaluated intention understanding, behavioral intention, actual behavior, psychological reactance, discomfort, familiarity, and agent acceptance. In several task contexts, the bark-only baseline produced lower intention-understanding and behavioral scores than the expressive conditions. Emotional expression and attention-guiding cues therefore appear to improve interpretation of agent intention and support behavior change. However, no consistent significant differences emerged between species-specific and shared behavior, suggesting that faithful reproduction of animal-specific motion is not the main determinant of persuasive effectiveness. Psychological reactance and discomfort remained low, while familiarity with an animal species was associated with actual behavior in some conditions. These findings indicate that persuasion by quadruped virtual agents depends more on functional cues, including emotional expression, attention guidance, and intention readability, than on accurate species-specific behavior. The results support cross-species generalizability and provide a basis for reusable design patterns in persuasive technology and human-AI interaction.","authors":["Kaoru Sumi","Souki Osawa"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-08-04","first_seen":"2026-08-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.01895","pdf_url":"https://arxiv.org/pdf/2608.01895","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C2"],"tags":["虚拟代理","人机交互","劝说技术"],"reason":"研究四足虚拟代理的劝说效果，不涉及LLM仿真人类被试，属于机器人/虚拟代理仿真…","model":"deepseek-v4-pro","scored_at":"2026-08-04T13:04:54","error":null,"has_summary":false,"summary":null},{"id":"2608.02283","version":1,"title":"Embodied Empathy: A Multimodal AR and LLM-Powered System for Self-Attachment Psychotherapy with Self-Initiated Humour","zh_title":"具身共情：用于自我依恋心理治疗的多模态AR与LLM系统，结合自我发起幽默","abstract":"The growing global demand for mental health support increasingly exceeds the supply of qualified practitioners, creating an urgent need for scalable digital interventions that can deliver meaningful emotional connection. In response, we present a novel multimodal application that operationalises the Self-Initiated Humour Protocol (SIHP) within a Self-Attachment Technique (SAT) framework. Our mobile application integrates customisable 3D childhood avatars, augmented reality, and an LLM-driven virtual therapist capable of automated emotion mirroring. An eight-day user study (N=16) indicates the system's feasibility and improvements in self-reported mood. Results show that personalised avatars and text-to-speech output strengthen emotional bonding and perceived empathy. Although emotion mirroring boosts engagement, its effectiveness depends heavily on classification accuracy and animation intensity. Moreover, findings indicate a shift in user expectations--from reactive chatbots to proactive conversational facilitators. We conclude with design implications for leveraging AI and AR to cultivate embodied empathy in digital mental health tools.","authors":["Xinyan Ye","Gwyneth Phang","Anandha Gopalan","Abbas Edalat"],"categories":["cs.HC","cs.MM"],"primary_category":"cs.HC","announce_type":"new","date":"2026-08-04","first_seen":"2026-08-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.02283","pdf_url":"https://arxiv.org/pdf/2608.02283","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["心理健康","虚拟治疗师","增强现实"],"reason":"LLM虚拟治疗师进行角色扮演对话，无人类行为仿真对照，属C3排除项。","model":"deepseek-v4-pro","scored_at":"2026-08-04T13:04:57","error":null,"has_summary":false,"summary":null},{"id":"2608.01346","version":1,"title":"Hybrid AI for Explainable and Accurate Conversational Agents in eGovernment","zh_title":"用于电子政务中可解释且准确的对话代理的混合人工智能","abstract":"We present a so-called Conversational Hybrid AI (CHAI) architecture for building explainable and accurate conversational agents for eGovernment. We exemplify the architecture with a running prototype of a Covid-19 Chatbot based on a governmental guideline directed to citizens. We also describe an ongoing case on case management for supplementary grants for students with disabilities. We use large language models (LLMs) as a bounded conversational interface to a rule-based (symbolic AI) controller that executes a logical model expressing the provisions and obligations of the law and/or guidelines. As logical modelling language we use Dynamic Condition Response (DCR) graphs, a symbolic declarative process-modeling language developed with the aim to be able to express both deontic, defeasible and temporal logic properties, making it suitable for expressing both the rules of the law and the steps of the legal case management processes.","authors":["Ilias Chalkidis","Vlad Paul Cosma","S{\\o}ren Debois","Daniel Hershcovich","Thomas Hildebrandt","Hugo A. L\\`opez","Amogh Raina","Konstantinos Varvoutas","Tilman Zuckmantel"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2026-08-04","first_seen":"2026-08-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.01346","pdf_url":"https://arxiv.org/pdf/2608.01346","source_feed":"cs.CY","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["对话代理","电子政务","混合AI"],"reason":"构建面向公民的对话式政务助手，属于角色扮演聊天机器人，无实验或测量目的。","model":"deepseek-v4-pro","scored_at":"2026-08-04T13:04:46","error":null,"has_summary":false,"summary":null},{"id":"2608.01085","version":1,"title":"When Collaboration Becomes a Trigger: Collective Evidence-Threshold Backdoors in Multi-Agent Systems","zh_title":"当协作成为触发器：多智能体系统中的集体证据阈值后门","abstract":"LLM-based multi-agent systems (MAS) extend LLM capabilities through iterative communication and shared contexts. However, this collaboration introduces a vulnerability: backdoor behavior can be activated when peer evidence reaches a hidden threshold, rather than being determined by any single message. We introduce a collective evidence-threshold backdoor paradigm for MAS and Boundary-Conditioned Backdoor Injection (BCBI), which constructs counterfactual boundary pairs to separate benign behavior before the threshold from the adversarial objective after it, and learns latent progression aligned with evidence. To mitigate this threat, we propose LAtent Transition Test-time Evaluation (LATTE), a clean-only latent-transition defense that learns benign communication dynamics and quarantines anomalous agent updates before their responses propagate. Across several benchmarks, BCBI yields selective activation with little premature activation; without knowing the attack target or trigger, LATTE limits propagation with minimal disruption.","authors":["Jia-Hao Xiao","Lei Feng","Min-Ling Zhang"],"categories":["cs.MA","cs.LG"],"primary_category":"cs.MA","announce_type":"new","date":"2026-08-04","first_seen":"2026-08-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.01085","pdf_url":"https://arxiv.org/pdf/2608.01085","source_feed":"cs.MA","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","后门攻击","安全防御"],"reason":"纯多智能体协作攻击与防御，无人类行为仿真或对照，属C1排除项。","model":"deepseek-v4-pro","scored_at":"2026-08-04T13:04:44","error":null,"has_summary":false,"summary":null},{"id":"2608.02178","version":1,"title":"Microscopic dynamics of consensus formation in multi-agent LLM Naming Games","zh_title":"多智能体LLM命名游戏中共识形成的微观动力学","abstract":"Decentralized populations of Large Language Model (LLM) agents can spontaneously reach consensus on shared conventions, yet the microscopic mechanisms by which their internal stochasticity shapes macroscopic ordering remain unexplored. We study a minimal LLM Naming Game in which the listener's decision is a single-token LLM call at decoding temperature $T$, replacing the inventory check of the deterministic Naming Game. Each interaction decomposes into an in-inventory and an out-inventory channel with conditional rates $\\pi(T)\\!\\equiv\\!P(\\text{YES}\\mid w\\in P_j)$ and $\\phi(T)\\!\\equiv\\!P(\\text{YES}\\mid w\\notin P_j)$, whose balance controls an ordering-disordering drift. A mean-field theory of the two-rate dynamics yields an analytical ordering condition that generalizes the consensus threshold of the stochastic Naming Game to a critical line in the $(\\pi,\\phi)$ plane. Across three open-weight architectures, consensus is always reached, but through three distinct listener regimes: permissive (repaint-noise dominated), near-deterministic, and conservative (missed-collapse dominated). The effective finite-size exponent $\\beta(T)$ in $t_{\\rm conv}\\!\\sim\\!N^{\\beta}$ shifts with temperature, and the temperature-sensitivity $\\alpha$ in $t_c\\!\\sim\\!e^{\\alpha T}$ ranges from ${\\approx}\\,0.67$ to ${\\approx}\\,0$ across architectures. Decoding temperature thus emerges as an architecture-dependent control parameter for decentralized LLM populations, quantitatively characterized by the statistical-physics toolkit.","authors":["Cristiano De Nobili","Vijayasri Iyer","Alessandro Codello","Raffaella Burioni"],"categories":["physics.soc-ph","cond-mat.stat-mech","cs.MA","nlin.AO"],"primary_category":"physics.soc-ph","announce_type":"cross","date":"2026-08-04","first_seen":"2026-08-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.02178","pdf_url":"https://arxiv.org/pdf/2608.02178","source_feed":"cs.MA","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","共识形成","统计物理"],"reason":"多智能体命名博弈，研究共识形成的微观机制，无人类行为对照，属纯多智能体系统研究。","model":"deepseek-v4-pro","scored_at":"2026-08-04T13:04:33","error":null,"has_summary":false,"summary":null},{"id":"2608.02412","version":1,"title":"Why Large Language Models Fail at Tabular Prediction","zh_title":"为何大语言模型在表格预测上失败","abstract":"Large language models (LLMs) have become the default tool for a remarkable range of tasks, yet they have had conspicuously little success at one of the most common machine learning workloads: predictive analytics over tabular data. This gap is the founding premise of the fast-growing field of tabular foundation models, but the question of why generic LLMs fail has remained open. We study a frontier LLM in its purest inference regime - a single generation pass over a prompt containing the full training and test data, with no tools, no agentic scaffolding, and no fine-tuning - and systematically evaluate five hypotheses for the failure: (a) an inability to handle noisy or non-linearly-separable data; (b) the linearised CSV format obscuring column structure; (c) the tokenisation of numeric values; (d) the number of test points classified per query; and (e) the dimensionality of the input. Controlled experiments falsify (a)-(d). Dimensionality, in contrast, is decisive: sweeping random linear projections of thirty-one benchmark datasets, the LLM is the only method among nine whose accuracy decreases as dimensionality grows, while every classical baseline stays flat or improves. A behavioural comparison against 252 configured classical models finds that in two dimensions the LLM predicts like a local, distance-based method (up to 91.6% grid agreement), but in higher dimensions no classical model - even when augmented with tuned, dimension-dependent noise - reproduces its predictions. We do not claim to have identified the internal mechanism; our results show, more modestly, that the LLM's capability dissolves with dimension in a way no noise-corrupted classical learner mimics - which explains why LLMs, so capable elsewhere, keep losing to fifty-year-old baselines on tables, while leaving the mechanism of the prediction as an open question.","authors":["Marta Garnelo","Wojciech M. Czarnecki"],"categories":["cs.LG"],"primary_category":"cs.LG","announce_type":"new","date":"2026-08-04","first_seen":"2026-08-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.02412","pdf_url":"https://arxiv.org/pdf/2608.02412","source_feed":"cs.LG","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["表格预测","LLM能力评测","维度灾难"],"reason":"研究LLM在表格预测上的失败，属纯NLP能力评测，不以人类行为为参照系。","model":"deepseek-v4-pro","scored_at":"2026-08-04T13:04:58","error":null,"has_summary":false,"summary":null},{"id":"2608.00567","version":1,"title":"Optimal Inflation Rate: A Meta-Analysis","zh_title":"最优通胀率：一项元分析","abstract":"We revisit the optimal long-run inflation rate using 777 estimates from 116 primary studies published between 1989 and 2026, the largest sample on the topic to date. To our knowledge, this is among the first economics meta-analyses in which primary-data extraction is done from start to finish through a documented and auditable large-language-model pipeline, calibrated against a hand-coded training set and released for replication. The literature points to an optimum of about 0.6 percentage points per year, well below the two-percent targets used by most advanced-economy central banks. The gap should not be automatically read as a verdict against the two-percent norm. Measurement error in published price indices could close, widen, or even reverse the gap, and the structural literature itself cannot pin down the sign of the required correction. Bayesian model averaging over the full set of structural moderators shows that cross-study variation is driven by real modelling choices rather than by selective reporting. The main drivers are the choice of monetary benchmark (Friedman rule vs. laissez-faire), the transactions-frictions technology, the assumed shock structure, and the class of nominal-rigidity contract. The non-parametric caliper test finds no upward bunching at the two-percent target. The paper contributes a reproducible LLM-assisted extraction pipeline for structurally calibrated literature and a quantitative decomposition of where the optimal-inflation literature disagrees.","authors":["Matej Opatrny","Martin Opatrny","Tomas Havranek","Zuzana Irsova","Mojmir Hampl"],"categories":["econ.GN","q-fin.EC"],"primary_category":"econ.GN","announce_type":"new","date":"2026-08-04","first_seen":"2026-08-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2608.00567","pdf_url":"https://arxiv.org/pdf/2608.00567","source_feed":"econ.GN","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["元分析","LLM辅助数据提取","最优通胀率"],"reason":"论文用LLM辅助元分析数据提取，属NLP工具应用，非仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-08-04T13:04:23","error":null,"has_summary":false,"summary":null},{"id":"2607.29274","version":1,"title":"Language Models Agree With Each Other, Not With Readers","zh_title":"语言模型彼此一致，而非与读者一致","abstract":"Claims that language models homogenise are usually measured against human judgements collected for the study, which makes the human side an artifact of the design: a crowdworker given the model's instruction is running the model's prompt. We measure convergence against a human reference nobody built for the purpose -- 2,523 reader mark sets across 120 web documents, produced by people highlighting for their own reasons on a platform where the overlay of others' marks is off by default. Agreement is the overlap between two size-matched sentence sets minus the overlap expected when each is resampled within its own depth-and-length bands. The null's calibration is demonstrated, not asserted: every pair involving a random baseline lands within 0.006 of zero. On the median document each party names 14 sentences of 70; two readers share 4.1 and two models 8.7. Across 18 model arms spanning 11 vendors, 3 countries and both weight regimes, the median of 153 model pairs is +0.093 against a human yardstick of +0.040, and 99 sit entirely above the human interval. Two frontier models from rival labs reach +0.203, twice what GPT-4o agrees with itself on a second call. The effect is not determinism, prompt wording, procedure, vendor or routing, and it is graded: the smallest models agree at the human level. No model agrees with readers detectably more than a reader does, and at equal depth and length no surface feature separates their choices. The multiples are procedure-dependent and the ordering is not: models are cut to their sharpest set while a reader's is a random draw from what they marked, and blunting the models alike halves the gap without closing it. Tested out of sample on four models released after this analysis, against predictions fixed beforehand, none clears the human interval. A population simulated from several models is not several populations.","authors":["Kazuki Nakayashiki","Keisuke Watanabe"],"categories":["cs.IR","cs.CL","cs.CY","cs.HC"],"primary_category":"cs.IR","announce_type":"cross","date":"2026-08-03","first_seen":"2026-08-03","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.29274","pdf_url":"https://arxiv.org/pdf/2607.29274","source_feed":"cs.CL","score":8,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","人类行为对照","一致性评估"],"reason":"用LLM模拟人类标注行为并与真实读者数据对照，评估模型间一致性及与人类差异，揭…","model":"deepseek-v4-pro","scored_at":"2026-08-03T13:02:03","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-03","rank":3,"question":"语言模型在文本标注任务上的趋同性是否高于真实读者，且与读者的差异是否随模型能力增强而扩大？","design":"将18个不同厂商、规模、代际的语言模型作为被试，输入120篇网页文档的编号句子，要求模型按重要性排序并取前k句作为标注集；同时收集同一文档上独立读者的真实高亮标注作为对照，测量模型间、读者间及模型与读者间的标注重叠度（经位置和长度校准后的协议分数）。","baseline":"2,523个真实读者在120篇网页文档上的自主高亮标注集，读者未受指令引导且互不可见。","findings":"模型间的标注一致性显著高于读者间一致性（中位数协议分数0.093 vs 0.040），且前沿模型间一致性可达0.203，是读者间一致性的5.1倍；模型与读者的一致性仅与读者间一致性持平，且模型选择的句子在表面特征上与读者无异，但内容不同。","reliability":"读者独立性无法完全验证（仅基于平台默认设置假设）；协议分数依赖于标注深度和长度控制，稀释模型标注会缩小但未消除差距；样本外测试中，新模型均未突破人类一致性区间。","relevance":"该研究直接以真实人类行为为基准，系统评估了LLM仿真人类标注的可靠性及偏差，揭示了模型趋同但未逼近人类分布的规律，对关注LLM仿真效度的研究者极具参考价值。","inspiration":"借鉴其利用自然发生的非实验人类行为数据作为基准，避免指令诱导同质化的设计；可迁移到消费者信息处理或投资者注意力分配研究，如用LLM模拟投资者阅读财报后的关注点；以真实投资者在财经新闻上的自主高亮或眼动数据为对照，让LLM对同一文本生成重要性排序，比较注意力分布与真实行为的差异。"}},{"id":"2607.28643","version":1,"title":"To Facilitate or not to Facilitate: Human and LLM Facilitator Tendencies in Online Discussions","zh_title":"促进与否：在线讨论中人类与LLM的主持倾向","abstract":"Automating facilitation in online discussions is a long-standing social concern given the increasing time we spend on online spaces and the failure of content moderation approaches. While studies have been conducted on how to facilitate, none have answered the essential question of when to do so. A potential answer is using LLMs, which ostensibly make automated, large-scale intervention increasingly feasible. In this study, we examine when LLMs decide to facilitate by defining what facilitation is, observing when humans decide to facilitate, and comparing their decisions with those made by LLMs. To this end, we create PEFK, a corpus standardizing and aggregating all relevant facilitation datasets. We are the first to run a survey on facilitation timing, which we execute using expert facilitative participants and LLM-as-a-judge models. We discover that while humans are more cautious, LLMs are excessively eager to facilitate, although both are more certain when judging that facilitation is not needed. We then investigate whether this behavior can be corrected using alternative setups for LLMs and training ModernBert classifiers on established datasets, finding that the latter perform more reliably than the former, although current datasets impose a relatively low performance ceiling.","authors":["Dimitris Tsirmpas","Katerina Korre","John Pavlopoulos"],"categories":["cs.HC","cs.CL"],"primary_category":"cs.HC","announce_type":"cross","date":"2026-08-03","first_seen":"2026-08-03","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.28643","pdf_url":"https://arxiv.org/pdf/2607.28643","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A1","B1","B4"],"tags":["LLM仿真","人类对照","行为偏差"],"reason":"用LLM模拟人类主持决策并与人类数据对照，发现LLM过度干预，批判性指出仿真偏…","model":"deepseek-v4-pro","scored_at":"2026-08-03T13:01:58","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-03","rank":4,"question":"LLM在决定何时介入在线讨论时，与人类主持人的决策倾向有何差异？","design":"本研究并非严格意义上的仿真实验，而是通过构建统一数据集PEFK，设计调查任务让10名专家参与者与6个开源LLM分别判断1226个讨论片段是否需要介入，比较两者在介入频率和置信度上的差异。","baseline":"10名具有主持经验的专家参与者对讨论片段是否需要介入的判断及其置信度。","findings":"人类主持人更谨慎，倾向于不介入且对不介入的判断更自信；LLM则过度渴望介入，且更依赖正面强化，但在判断不需要介入时同样更自信。","reliability":"论文指出当前数据集存在固有噪声（以专业主持人撰写的评论作为标签），导致性能上限较低；LLM在介入时机预测上表现不一致，而基于编码器的分类器更可靠，但整体仍受限于对主持行为的有限理解。","relevance":"该研究直接对比LLM与人类在决策行为上的差异，并批判性地指出LLM过度干预的偏差，符合研究者对仿真可靠性评估和失效条件分析的兴趣，值得阅读原文以了解具体实验设计和偏差来源。","inspiration":"可借鉴其通过统一多源数据集并设计对照调查来量化行为偏差的方法，尤其关注决策阈值和置信度的测量。｜可迁移到经济政策沟通场景，如央行官员在公开声明中决定何时干预市场预期，或监管者决定何时对金融市场异动发声。｜以真实央行沟通记录为对照，让LLM扮演政策制定者，判断在不同市场波动情境下是否需要发表声明，结果变量为干预频率和声明内容倾向，与历史真实干预记录对比评估LLM的过度干预或保守倾向。"}},{"id":"2607.28908","version":1,"title":"Reflection or Re-Generation? Why LLM Revision Fails Where Human Revision Succeeds","zh_title":"反思还是重新生成？为何LLM修正失败而人类修正成功","abstract":"Reflection, the ability to revisit and revise prior reasoning, is central to how humans improve their answers. Large language models (LLMs) are increasingly prompted to \"reflect,\" yet whether this resembles human revision remains unclear. We introduce the Human-LLM Reflection Framework (HRF), a controlled two-pass protocol comparing human and LLM revision under identical conditions across self-, peer-, and cross-agent settings. Using an information-theoretic analysis based on per-iteration cross-entropy reduction, we find two failure modes of LLM reflection. On objective tasks with finite answer spaces, reflection yields near-zero information gain (Delta I approx 0), behaving as neutral re-generation indistinguishable from re-sampling. On subjective tasks, it yields significant negative gain (Delta I < 0), moving predictions away from the target. Human revision, by contrast, yields positive gain in both settings. Cross-agent experiments localize the failure to the revision step, not input quality: LLMs degrade even high-quality human responses. Diagnostic analyses (revision conditioned on first-pass correctness, and oracle-guided revision against a random-reshuffle baseline) show that which sub-step dominates varies by task and by model rather than reducing to a single mechanism: self-error detection is present on objective multiple-choice tasks but weak on subjective ones, and recovery under an oracle error signal exceeds the baseline for some models and falls below it for others. The unifying account is structural: without external information, self-conditioned revision cannot reduce uncertainty about the target, so LLM reflection is better understood as conditioned re-generation than as genuine error-driven revision.","authors":["Yefan Tao","Gerald Friedland","Madhusudhanan Chandrasekaran","Luyang Kong"],"categories":["cs.LG"],"primary_category":"cs.LG","announce_type":"new","date":"2026-08-03","first_seen":"2026-08-03","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.28908","pdf_url":"https://arxiv.org/pdf/2607.28908","source_feed":"cs.LG","score":7,"bucket":"pending","rubric_hits":["A2","B1","B4"],"tags":["LLM反思","人类对照","可靠性评估"],"reason":"对比人类与LLM的修正行为，揭示LLM反思的失效模式，有真实人类数据对照，批判…","model":"deepseek-v4-pro","scored_at":"2026-08-03T13:02:00","error":null,"has_summary":true,"summary":{"generated_at":"2026-08-03","rank":5,"question":"LLM的反思（revision）是否像人类一样是真正的错误修正，还是仅仅是基于先前输出的条件再生成？","design":"提出Human-LLM Reflection Framework (HRF)，采用受控两阶段协议：人类和LLM在相同条件下先做初始回答（Pass 1），再看到先前回答后决定保留或修改（Pass 2），涵盖自我修正、同伴修正和跨代理修正三种设置，在客观数学推理和主观情感评价任务上测量信息增益。","baseline":"非专家人类标注者在相同任务、提示和评估标准下的修正行为，作为有效反思的实证参照。","findings":"LLM反思呈现两种失效模式：客观任务上信息增益接近零，表现为中性再生成；主观任务上信息增益显著为负，使预测偏离目标。人类修正在两种任务上均产生正信息增益，且失效定位于修正步骤而非输入质量。","reliability":"论文指出LLM反思失效的结构性原因：缺乏外部信息时，自我条件修正无法降低关于目标的不确定性；诊断分析显示失效子步骤（错误检测或错误纠正）因任务和模型而异，并非单一机制。","relevance":"该研究直接对比人类与LLM的修正行为，揭示LLM反思的失效模式，有真实人类数据对照，批判性地指出仿真在反思环节的不可靠性，对关注LLM作为人类被试替代品的研究者具有重要参考价值，值得精读原文。","inspiration":"借鉴其受控两阶段修正协议和信息论测量（交叉熵减少量）来严格评估LLM的决策修正能力。｜可迁移到经济预测修正场景，如分析师盈利预测修正、央行沟通后的市场预期调整。｜以LLM作为分析师被试，先给出盈利预测（Pass 1），再提供历史预测值要求修正（Pass 2），结果变量为预测误差变化，以真实分析师修正数据（如IBES）作为人类基准，对比信息增益。"}},{"id":"2607.24435","version":2,"title":"LEX-EC: A Lexical Evidence-Channel Audit Framework for Zero-Shot LLM Personality Classification in Black-Box Settings","zh_title":"LEX-EC：黑盒环境下零样本LLM人格分类的词汇证据通道审计框架","abstract":"Large language models may easily assign personality labels from text, but model interpretability remains an open problem. To address this gap, we introduce LEX-EC, a reusable black-box audit framework combining prevalence and agreement diagnostics with controlled lexical ablation to distinguish marginal-distribution effects from trait-associated signal recoverable under restricted evidence. Using this framework, we illustrate how various text genres may exhibit sharply different profiles: free-form essay text contains the broadest, but still weak, signal; in graduate student introductions, an observable Extraversion association weakened after masking; and single Facebook statuses yield little stable evidence even in a trait-balanced sample, indicating a possible lower bound of content or length. Masking topical and demographic content weakened some associations while leaving others detectable from function words, affective terms, and cognitive-style vocabulary. Linguistic prompting shifted model self-explanations but did not eliminate topical content. LEX-EC jointly evaluates classification prevalence, item-level association, chance-corrected agreement, persistence under lexical restriction, and prompt sensitivity in model-generated explanations. Across datasets, models, and prompts, LEX-EC characterizes how trait associations may vary with available lexical evidence, introducing a novel application of lexical methods to black-box interpretability in personality labeling.","authors":["Brittany Harbison","Ashok K. Goel"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-08-03","first_seen":"2026-07-28","revised_at":"2026-08-03","abs_url":"https://arxiv.org/abs/2607.24435","pdf_url":"https://arxiv.org/pdf/2607.24435","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["LLM人格分类","黑盒可解释性","词汇审计"],"reason":"研究LLM人格分类的可解释性，测量对象是模型而非人类被试，无人类数据对照。","model":"deepseek-v4-pro","scored_at":"2026-08-03T13:02:18","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-29","rank":11,"question":"在零样本黑盒设定下，大语言模型进行人格分类时，其预测信号在多大程度上依赖于文本中的话题和人口统计学内容，而非真正的特质相关语言？","design":"本研究并非人类仿真实验，而是提出一个名为LEX-EC的黑盒审计框架，通过分布诊断、一致性检验和受控词汇消融，分析LLM在不同文本体裁、长度和内容遮蔽条件下的人格分类行为，测量分类流行率、项目级关联、机会校正一致性、词汇限制下的信号持久性及提示敏感性。","baseline":"无对照","findings":"自由形式短文包含最广泛但依然微弱的特质信号；研究生自我介绍中，外向性关联在遮蔽话题和人口统计学内容后减弱；单一Facebook状态即使在特质平衡样本中也几乎不产生稳定证据，表明存在内容或长度的下限。","reliability":"论文指出，LLM的人格标签可能受话题内容、人口统计学先验和文本长度影响，而非真实特质信号；在短文本或内容受限条件下，预测信号可能崩溃；模型自解释受提示影响，但无法消除话题内容。","relevance":"该研究与您关注的人类仿真实验不同，它审计的是LLM的人格分类行为而非用LLM模拟人类被试，且无真实人类行为对照，但其中关于信号来源的批判性分析对评估LLM仿真可靠性有参考价值。","inspiration":"与经济金融研究关联不大"}},{"id":"2607.28439","version":2,"title":"Beyond a Single Judge: The Evidence-Grounded, Social-Weighted Persona Panel for Generative UI Evaluation","zh_title":"超越单一评判：用于生成式UI评估的基于证据与社会权重的角色面板","abstract":"Generative UI (GenUI) lets large language models synthesize a complete, renderable interface directly from a natural-language instruction, but evaluating the quality of what they generate remains an open problem. Human evaluation is costly and rater-variant, while LLM-as-a-judge is scalable but reflects only a single implicit viewpoint, unable to capture how different populations of real users actually perceive the same interface. We propose the Evidence-Grounded, Social-Weighted Persona Panel (ESPP), a three-stage GenUI evaluation method in which a panel of psychologically diverse, evidence-grounded personas independently rates a screenshot, exchanges opinions under a trait-derived, semantically-gated bounded-confidence mechanism, and is aggregated via Delphi-inspired social weighting into a single judgment. ESPP tracks human judgment substantially more closely than a naive single-pass judge, raising Pearson $r$ from $0.716$ to $0.922$, and a prompt-ensemble control recovers only about a third of this gap, isolating genuine persona and evidence grounding as the dominant source of improvement. Beyond this fidelity gain, retaining each panelist's individual rating further reveals that user subgroups agree on overall model rankings yet diverge sharply on specific rating dimensions, a structural disagreement a single homogeneous judge would systematically erase. The codes are available at https://github.com/Wuzheng02/ESPP.","authors":["Zheng Wu","Yibo Luo","Pu Zhang","Cheng Yang","Zhuosheng Zhang"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-08-03","first_seen":"2026-07-31","revised_at":"2026-08-03","abs_url":"https://arxiv.org/abs/2607.28439","pdf_url":"https://arxiv.org/pdf/2607.28439","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM评估","角色面板","UI生成"],"reason":"用LLM persona面板替代人工评估UI，属标注替代而非仿真人类被试，无真…","model":"deepseek-v4-pro","scored_at":"2026-08-03T13:02:18","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-31","rank":11,"question":"如何让LLM模拟多样化用户群体来评估生成式UI，并使其评分更贴近真实人类判断？","design":"用LLM扮演心理特质多样且基于证据的虚拟用户（persona），组成面板；每个persona先独立对UI截图评分，再通过基于特质和语义门控的有界置信机制交换意见，最后用德尔菲式社会加权聚合为单一评分。","baseline":"UIPersonaBench基准中500条指令由14个模型生成的7000张截图，每张截图收集5名真实人类在5个维度上的1-5分评分，取均值作为人类基准。","findings":"ESPP面板评分与人类评分的皮尔逊相关系数从单一法官的0.716提升至0.922，且提升主要来自persona与证据锚定，而非多次调用平均；不同用户子群在整体排名上一致，但在具体维度（如控制感）上存在显著分歧，单一法官会抹平这种差异。","reliability":"论文未讨论","relevance":"该研究用LLM模拟多样化用户群体评估UI，并与真实人类评分严格对照，揭示了子群分歧，方法可迁移至经济学实验和政策评估中的异质性偏好测量。","inspiration":"借鉴其基于心理特质构建异质性persona面板、通过社会交互机制模拟意见动态并保留个体分歧的设计；可迁移到消费者金融产品选择或政策偏好调查中，模拟不同风险态度、金融素养人群的决策差异；以LLM扮演不同人格与认知偏好的被试，处理为呈现不同设计的金融产品界面或政策描述，结果变量为选择或评分，用真实消费者调查或实验数据作为对照基准。"}},{"id":"2607.29082","version":1,"title":"Can Zero-Shot LLMs Predict Child Malnutrition? A Fairness and Temporal Robustness Study","zh_title":"零样本大语言模型能否预测儿童营养不良？一项公平性与时间鲁棒性研究","abstract":"Child malnutrition remains a major public health challenge in low- and middle-income countries, particularly in South Asia, where early identification of vulnerable children is critical for timely intervention and resource allocation. This study aims to evaluate the feasibility, fairness, and temporal robustness of using a pretrained large language model (LLM) in a zero-shot setting for child stunting prediction using population health survey data. Using Bangladesh Demographic and Health Survey (BDHS) data collected between 2007 and 2022, we transformed maternal, child, healthcare, and household characteristics into semantically interpretable prompt-based representations and evaluated GPT-4o-mini for zero-shot stunting prediction, comparing its performance against a random forest baseline and assessing fairness across demographic and socioeconomic groups as well as temporal robustness across survey waves. The results demonstrate that zero-shot inference using GPT-4o-mini achieved comparable balanced accuracy to the supervised baseline while exhibiting substantially higher sensitivity for identifying stunting cases, relatively consistent performance across child sex groups, and stable predictive behaviour across BDHS waves; however, important fairness disparities were observed across residence and household wealth categories, highlighting the need for further investigation before deployment of foundation models in public health prediction settings.","authors":["Muhammad Ashad Kabir","Md Ahshanul Haque"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-03","first_seen":"2026-08-03","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.29082","pdf_url":"https://arxiv.org/pdf/2607.29082","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM预测","公平性","公共卫生"],"reason":"用LLM替代人工预测儿童发育迟缓，属于替代人类劳动而非仿真被试，但涉及公平性评…","model":"deepseek-v4-pro","scored_at":"2026-08-03T13:02:02","error":null,"has_summary":false,"summary":null},{"id":"2607.28651","version":1,"title":"Measuring Cognitive Engagement in Collaborative Discourse with an Extended ICAP Framework: Comparing Human Annotation, In-Context Learning, and Reflective LLM Agents","zh_title":"用扩展ICAP框架测量协作对话中的认知参与：比较人工标注、上下文学习和反思型LLM智能体","abstract":"Collaboration supports learning and problem-solving, but its effectiveness depends on cognitive engagement during discourse. This study applies an extended 7-point ICAP framework based on the Interactive, Constructive, Active, and Passive modes to characterize variation in cognitive engagement during collaborative dialogue. Engagement was coded by trained human annotators and compared with large language model (LLM)-based labeling approaches, including in-context learning (ICL), zero-shot prompting, and self-reflective agents. Interrater reliability among human annotators was robust across framework refinement stages (kappa = 0.906-0.998), higher than the moderate agreement observed for ICL-based annotation (kappa = 0.541-0.609). The human-refined framework improved agreement among human annotators (Delta kappa = 0.10), but produced only modest gains for ICL-based LLMs (Delta kappa less than 0.04). Agent-refined frameworks improved cross-model agreement but remained below the human-refined framework. These findings highlight the promise of agent-based approaches and the importance of continued interaction between theory-guided human annotation and LLM-based methods in future work.","authors":["Lan Anh Do","Hanling Jiang","Shuchin Aeron","Ayanna K. Thomas"],"categories":["cs.HC","cs.CL","cs.CY"],"primary_category":"cs.HC","announce_type":"cross","date":"2026-08-03","first_seen":"2026-08-03","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.28651","pdf_url":"https://arxiv.org/pdf/2607.28651","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM标注","认知参与","协作学习"],"reason":"用LLM替代人工标注对话认知参与度，属于标注员替代而非仿真人类被试，但方法可迁…","model":"deepseek-v4-pro","scored_at":"2026-08-03T13:02:06","error":null,"has_summary":false,"summary":null},{"id":"2607.28956","version":1,"title":"MerchantBench: Benchmarking LLM Agents for Long-Term Coherence in E-Commerce Operations","zh_title":"MerchantBench：评估大语言模型智能体在电商运营中长期一致性的基准","abstract":"Large language model agents are increasingly evaluated as autonomous tool users, yet most benchmarks focus on bounded tasks with immediate success criteria. Real-world deployments often require Long-Term Coherence, the capacity to preserve purposeful behavior across extended horizons while adapting decisions to accumulated evidence. Evaluating this capacity requires a persistent environment in which actions constrain future choices, feedback arrives at heterogeneous delays, and incoherent behavior produces measurable cumulative effects. Seller-side e-commerce provides a suitable setting for this evaluation through recurrent and interdependent decisions over Product Sourcing, Listing and Pricing Control, Cash-Flow Management, and Mixed-Latency Feedback Adaptation. We introduce MerchantBench, a 365-day order-level simulation grounded in 98,843 real e-commerce product records and equipped with 26 tools for agent interaction. MerchantBench couples promptly observable Upstream Supplier Events with delayed Downstream Order Outcomes, requiring agents to follow individual order lifecycles and revisit earlier decisions. We evaluate eight LLMs under two agent frameworks in 48 runs, each spanning 365 simulated days. Our results reveal a substantial gap between even the latest LLMs and human participants, with the best LLM configuration attaining only 27.3\\% of the mean final net assets achieved by human participants.","authors":["Qiming Shi","Yulong Tao","Linbo Jin","Zhaolu Kang","Yibo Dou","Jiawen Zhu","Tianjun Pan","Shaokang Fu","Chengyu Wang","Siyue Li","Yaping Cheng","Di Weng","Chengfu Huo"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-03","first_seen":"2026-08-03","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.28956","pdf_url":"https://arxiv.org/pdf/2607.28956","source_feed":"cs.AI","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["LLM智能体","电商模拟","长期决策"],"reason":"LLM agent模拟电商经营，有人类对照但非社会行为仿真，属边界情形。","model":"deepseek-v4-pro","scored_at":"2026-08-03T13:02:10","error":null,"has_summary":false,"summary":null},{"id":"2607.28889","version":1,"title":"Human-LLM Collaborative Inductive Coding for Conceptualizing K-12 Educator AI Use","zh_title":"人机协作归纳编码：概念化K-12教育者AI使用","abstract":"Qualitative researchers increasingly encounter interaction corpora whose scale exceeds what manual coding alone can address, and large language models (LLMs) are frequently proposed as analytic assistants. The open questions are not whether LLMs can participate in qualitative analysis but to what extent, in what phases, and under what safeguards. This article provides a detailed procedural account of a multi-phase human-LLM collaborative pipeline that adapted open, axial, and selective coding to develop a hierarchical codebook from 45,000 messages exchanged between K-12 educators and a generative AI platform. Across three phases, LLMs generated candidate labels and structured annotations at scale, while human researchers retained conceptual authority over category definitions, merging decisions, and interpretive frameworks. The resulting instrument was then tested through systematic human coding, in which three trained coders with educational domain expertise applied the codebook to an independent sample of 2,560 messages, established reliability through iterative calibration using set-valued agreement measures appropriate for multi-label annotation, and extended the instrument with five codes that the LLM-assisted phases had not surfaced. The final codebook comprises 72 items within 19 categories and six domains. We reflect on the methodological decisions the pipeline required, including the choice of a conversational unit of analysis, the treatment of the LLM as a labeling instrument rather than an interpretive agent, the measurement of intercoder agreement under multi-label coding, and the conditions under which human domain expertise remained decisive. The account is offered as an auditable template for qualitative researchers considering LLM assistance in codebook development while preserving human interpretive authority.","authors":["Alex Liu","Min Sun","Lief Esbenshade","Michael Xiao","Victor Tian","Zachary Zhang","Kevin He"],"categories":["cs.HC","cs.AI"],"primary_category":"cs.HC","announce_type":"cross","date":"2026-08-03","first_seen":"2026-08-03","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.28889","pdf_url":"https://arxiv.org/pdf/2607.28889","source_feed":"cs.AI","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM辅助编码","定性研究","人机协作"],"reason":"LLM作为标注工具辅助编码，替代人工标注劳动，非仿真人类被试","model":"deepseek-v4-pro","scored_at":"2026-08-03T13:02:09","error":null,"has_summary":false,"summary":null},{"id":"2607.28890","version":1,"title":"Agreement Is Not Quality: Blind Expert Verification of Human and LLM Qualitative Coding When Human Consensus Is Not Ground Truth","zh_title":"一致性不等于质量：当人类共识并非金标准时，对人类和LLM定性编码的盲法专家验证","abstract":"Evaluations of LLM-assisted qualitative coding almost universally measure model performance as agreement with human coders, a practice that presumes human coding is the standard to approximate. This study provides empirical evidence that the presumption fails in ways agreement metrics cannot detect. Five LLM systems and three trained human coders independently applied a 72-item hierarchical codebook to 2,560 educator messages from a K-12 AI platform. Beyond conventional agreement analysis, an independent domain expert judged 855 pairwise comparisons of code sets blind to source, treating human and machine sources symmetrically. The two evaluation approaches diverge in both directions. Human-LLM agreement (mean Jaccard 0.30) falls well below human-human agreement (0.52), which standard practice would read as inferior LLM coding, yet the blind verifier preferred human and LLM coding at indistinguishable rates (51.5% vs. 48.5%, p = 0.537), and a Bradley-Terry ranking placed two LLMs above two of three human coders. For several substantive codes, human consensus encoded shared bias that the verifier rejected in favor of the LLM interpretation. Agreement-based evaluation is therefore insufficient for automation decisions, and the study demonstrates a transferable verification protocol and a code-level division-of-labor framework.","authors":["Alex Liu","Lief Esbenshade","Michael Xiao","Victor Tian","Zachary Zhang","Kevin He","Min Sun"],"categories":["cs.HC","cs.AI"],"primary_category":"cs.HC","announce_type":"cross","date":"2026-08-03","first_seen":"2026-08-03","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.28890","pdf_url":"https://arxiv.org/pdf/2607.28890","source_feed":"cs.AI","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM辅助定性编码","标注员替代","方法验证"],"reason":"LLM替代人工编码，属标注员替代而非仿真被试，但方法可迁移。","model":"deepseek-v4-pro","scored_at":"2026-08-03T13:02:10","error":null,"has_summary":false,"summary":null},{"id":"2607.29064","version":1,"title":"Benchmarking Frontier Large Language Models Against Official Crash Database Coding Using Police Crash Narratives","zh_title":"基于警方事故叙述的前沿大语言模型与官方事故数据库编码的基准测试","abstract":"Police crash narratives contain information that may supplement structured crash databases, but manual review is labor-intensive and it remains unclear how well large language models (LLMs) reproduce official crash coding. This study benchmarked six frontier LLMs by comparing narrative-derived crash attribute codes with corresponding fields in the Arkansas fatal-crash database. The analysis linked 5,587 fatal-crash narratives with 5,889 structured crash records from Arkansas (2015-2025), yielding 4,194 matched crashes. Six LLMs were evaluated using an identical zero-shot prompt to code crash manner, non-motorist relation, intersection type, work-zone relation, roadway surface condition, and light condition. Performance was evaluated using agreement, macro-averaged F1 score, Cohen's kappa, coverage, selective agreement, and comparisons with always-majority, always-Unknown, and keyword-rule baselines. Repeated-measures analyses and a generalized estimating equations model assessed differences among models and attributes. GPT-5.5 High achieved the highest agreement among the evaluated LLMs, but the always-majority baseline produced higher raw agreement and the keyword-rule baseline achieved macro-averaged F1 score and Cohen's kappa comparable to the best-performing LLM. Agreement was highest for non-motorist relation and crash manner and lowest for light condition, roadway surface condition, and work-zone relation. Differences across crash attributes exceeded differences across models. These results provide a benchmark for evaluating LLM-based crash coding and show that deployment should be evaluated on an attribute-specific basis using transparent baselines and human review.","authors":["Sudhir Bharati","Rajendra K C Khatri","Sudip Bharati"],"categories":["cs.LG","cs.AI"],"primary_category":"cs.LG","announce_type":"cross","date":"2026-08-03","first_seen":"2026-08-03","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.29064","pdf_url":"https://arxiv.org/pdf/2607.29064","source_feed":"cs.AI","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM标注","事故编码","基准测试"],"reason":"用LLM替代人工编码员，从文本中提取结构化信息，属于标注替代而非仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-08-03T13:02:00","error":null,"has_summary":false,"summary":null},{"id":"2607.28818","version":1,"title":"Best Friends, Not Forever: Evaluating Long-Horizon Persona Collapse and Behavioral Drift in AI Companions","zh_title":"最好的朋友，并非永远：评估AI伴侣中的长期角色崩塌与行为漂移","abstract":"As AI companions increasingly mediate repeated social interaction, users may rely on a stable role and shared history, yet locally acceptable replies do not ensure that either persists. We study two observable long-horizon failures: 'persona collapse', the loss of a deployed role, boundaries, values, or style, and 'behavioral drift', the gradual or recurrent erosion of those properties. We introduce ANCHOR, a controlled synthetic audit that separately measures persona enactment and trajectory recall. The study contains 2,008 conversations spanning 27 personas, nine interaction schedules, three generated memory settings, and four evaluated models. The Identity Probe combines a sealed 102-item questionnaire with turn-level judgments, while the Trajectory Probe scores 110 calibrated counterfactual questions from 35 conversation banks. Our results show that no evaluated model and configuration reliably preserves either dimensions: trajectory accuracy averages only 44.4%, user-state recall remains near four-option chance, and no tested context condition or memory consistently resolves these failures. Questionnaire retention also varies by model and persona facet, disagrees with turn-level behavior, and is sensitive to evaluator choice. These results indicate that current systems do not yet reliably support long-horizon companion continuity and that audits must distinguish persona enactment, trajectory recall, evaluator provenance, and deployment context rather than collapse them into a single trust or stability score.","authors":["Pranav Narayanan Venkit","Akshara Prabhakar","Yu Li","Daniel Lee","Chien-Sheng Wu"],"categories":["cs.AI","cs.CL"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-08-03","first_seen":"2026-08-03","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.28818","pdf_url":"https://arxiv.org/pdf/2607.28818","source_feed":"cs.CL","score":3,"bucket":"other","rubric_hits":["C3"],"tags":["AI伴侣","角色扮演","行为漂移"],"reason":"研究AI伴侣的角色崩塌与行为漂移，属角色扮演聊天机器人评估，无人类被试仿真或对…","model":"deepseek-v4-pro","scored_at":"2026-08-03T13:02:08","error":null,"has_summary":false,"summary":null},{"id":"2511.00847","version":5,"title":"Pay for The Second-Best Service: A Game-Theoretic Approach Against Dishonest LLM Providers","zh_title":"为次优服务付费：针对不诚实LLM提供者的博弈论方法","abstract":"The widespread adoption of Large Language Models (LLMs) through Application Programming Interfaces (APIs) induces a critical vulnerability: the potential for dishonest manipulation by service providers. This manipulation can manifest in various forms, such as secretly substituting a proclaimed high-performance model with a low-cost alternative, or inflating responses with meaningless tokens to increase billing. This work tackles the issue through the lens of algorithmic game theory and mechanism design. We are the first to propose a formal economic model for a realistic user-provider ecosystem, where a user can iteratively delegate $T$ queries to multiple model providers, and providers can engage in a range of strategic behaviors. As our central contribution, we prove that for a continuous strategy space and any $\\epsilon\\in(0,\\frac12)$, there exists an approximate incentive-compatible mechanism with an additive approximation ratio of $O(T^{1-\\epsilon}\\log T)$, and a guaranteed quasi-linear second-best user utility. We also prove an impossibility result, stating that no mechanism can guarantee an expected user utility that is asymptotically better than our mechanism. Furthermore, we demonstrate the effectiveness of our mechanism in simulation experiments with real-world API settings.","authors":["Yuhan Cao","Yu Wang","Sitong Liu","Miao Li","Yixin Tao","Tianxing He"],"categories":["cs.GT","cs.AI"],"primary_category":"cs.GT","announce_type":"replace-cross","date":"2026-08-03","first_seen":"2025-11-02","revised_at":"2026-08-03","abs_url":"https://arxiv.org/abs/2511.00847","pdf_url":"https://arxiv.org/pdf/2511.00847","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["机制设计","博弈论","LLM服务"],"reason":"多智能体博弈机制设计，不涉及人类行为仿真或对照","model":"deepseek-chat","scored_at":"2026-07-28T11:52:38","error":null,"has_summary":false,"summary":null},{"id":"2607.23332","version":2,"title":"AllocBench: Measuring Online Tool Allocation Capability in LLM Agents","zh_title":"AllocBench：衡量LLM智能体在线工具分配能力","abstract":"Creating a reusable tool is an investment: an agent pays a fixed cost now in exchange for the potential of future reuse. Therefore, a user should prefer an agent that creates a small number of highly reusable tools, rather than many one-offs. We introduce a paired benchmark that tests whether LLM agents exhibit conscious allocation behavior under a fixed budget in two contexts: an abstract text-based formulation and a code-construction task. We find that every frontier model we test---Claude Haiku, Claude Opus, GPT-5.4-mini, and GPT-5.6 Sol---acts near-optimally in the abstract framing but fails to transfer this ability to script-writing. Through further experiments, we identify the particular failure modes for each model. Notably, the first three models fail even when the scripts are not evaluated, while GPT-5.6 Sol stays selective under that weaker manipulation and collapses only at full construction. Furthermore, an open-source Qwen model policy-trained for abstract allocation generalizes this ability across held-out lexical variations, but sees no improvement at script allocation. Together, these results establish online tool allocation as a significant capability boundary, even for modern frontier models.","authors":["Daniel Wang","Andrew Xu"],"categories":["cs.LG"],"primary_category":"cs.LG","announce_type":"replace","date":"2026-08-03","first_seen":"2026-07-28","revised_at":"2026-08-03","abs_url":"https://arxiv.org/abs/2607.23332","pdf_url":"https://arxiv.org/pdf/2607.23332","source_feed":"cs.LG","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体","工具分配","能力评测"],"reason":"纯多智能体工具分配能力评测，无人类行为对照","model":"deepseek-v4-pro","scored_at":"2026-08-03T13:02:05","error":null,"has_summary":false,"summary":null},{"id":"2607.23424","version":3,"title":"Wrong and More Confident: A Field Experiment on Large Language Models Taking a Graduate Economics Exam","zh_title":"错且更自信：大语言模型参加研究生经济学考试的现场实验","abstract":"A red herring, an irrelevant passage added to a problem, corrupts a language model's reasoning and, through it, its final answer, while the form of the response survives untouched. The benchmark, called the Graduate Economic Reasoning Benchmark (GERB), is sixty graduate-level microeconomics problems, each a detailed setup with a verified final answer and a step-by-step reference solution. Each problem has two versions, one with the red herring and one without, and each of those is asked in two ways, one requesting an explanation and one not. This is a within-subject $2\\times2$ factorial experimental design. Thirty-eight language models answer all four versions of every problem. The clean problems (the control group) are already hard, with the models answering under sixty percent correctly on average. The red herring lowers the probability of a correct final answer by 12.3 percentage points, about a quarter of the models' mean accuracy of 0.525. The damage is largest on the problems the model rates as easy. Reasoning ability confers no protection, as the red herring's effect does not differ detectably across models with and without reasoning ability. It does change how the failure looks, since a model with no reasoning mode repeats one wrong answer across waves while a reasoning model wavers. The red herring also leads a model to rate a problem as easier than its clean version, while answering it wrong more often. Although open- and closed-weight models reach the same accuracy, the open-weight models reach it at a substantially lower cost per correct final answer. The form of the response is preserved even as its substance fails. The model still produces an explanation (explanation given), the final answer still follows from the reasoning shown (coherence), and, in the aggregate, it remains the same across waves (consistency).","authors":["Piyush Akimitsu"],"categories":["econ.GN","q-fin.EC"],"primary_category":"econ.GN","announce_type":"replace","date":"2026-08-03","first_seen":"2026-07-28","revised_at":"2026-08-03","abs_url":"https://arxiv.org/abs/2607.23424","pdf_url":"https://arxiv.org/pdf/2607.23424","source_feed":"econ.GN","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["LLM评测","经济学考试","红鲱鱼效应"],"reason":"纯LLM能力评测，无人类被试仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-08-03T13:02:15","error":null,"has_summary":false,"summary":null},{"id":"2607.27816","version":2,"title":"Beyond Borrowed Histories: Person-Aligned User Simulation for Interactive Role-Playing Evaluation","zh_title":"超越借用的历史：面向交互式角色扮演评估的个性化用户模拟","abstract":"Role-playing agents (RPAs) have become one of the most important consumer applications of large language models. Users engage in multi-turn conversations with RPAs for experiences such as emotional comfort, making reliable evaluation essential for measuring capability, comparing systems, and guiding further improvement. Existing benchmarks, however, typically require an RPA to continue a fixed dialogue history and then evaluate the continuation using a fixed rubric detached from the user. We identify and empirically demonstrate two limitations of this design. First, an RPA's output is shaped by the preceding dialogue history, preventing a scientifically grounded assessment of its role-playing ability in real multi-turn settings. Second, user experience varies substantially across individuals, and conventional fixed rubrics need not align with user satisfaction. We therefore introduce PALATE (Person-Aligned LLM-Simulated-User Assessment with Tailored Evaluation), a scalable RPA benchmark built on user simulators. PALATE is accompanied by a pool of 300 character profiles. Its main evaluation trains five per-user simulators and lets them engage candidate RPAs in free-form, multi-turn conversations over a pre-frozen panel of character profiles. Alongside a general quality rubric, we construct personalized rubrics to measure user satisfaction; on held-out annotated data, the personalized rubrics show higher agreement with human judgments than the general rubric. In the main evaluation of 16 candidates, PALATE separately characterizes generic turn quality, long-horizon session capability, and per-user experience on multi-turn trajectories co-constructed by each candidate. It thereby produces interpretable evaluations of specific user-RPA pairs rather than compressing systems into a single user-independent ranking.","authors":["Yuhang Zhu","Mingxuan Du","Benfeng Xu","Jie Gao","Lingyun Yu","Hongtao Xie"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-08-03","first_seen":"2026-07-31","revised_at":"2026-08-03","abs_url":"https://arxiv.org/abs/2607.27816","pdf_url":"https://arxiv.org/pdf/2607.27816","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["角色扮演评估","用户模拟器","对话系统"],"reason":"论文评估角色扮演代理，用户模拟器用于对话评估，非人类被试仿真实验。","model":"deepseek-v4-pro","scored_at":"2026-08-03T13:02:17","error":null,"has_summary":false,"summary":null},{"id":"2607.28634","version":1,"title":"Can LLMs Really Understand Item Difficulty Levels? Implications for Automated Item Generation Using LLMs","zh_title":"大语言模型真的能理解题目难度吗？对使用LLM自动生成题目的启示","abstract":"The estimation of item difficulty plays a key role in both formative assessment and large-scale high-stakes summative assessments. This study explores how large language models (LLMs) perform in predicting item difficulty levels using items from a large-scale Reading and Writing test. The study investigated various prompting strategies and parameter settings across multiple LLMs. LLM performance was compared with encoder-only language models and feature-based supervised machine learning models. Zero-shot GPT-4.1 with a temperature of 0 yielded the highest item difficulty level prediction accuracy, with a quadratic weighted kappa (QWK) of 0.578. However, LLMs' prediction accuracy was lower than that of ConvBERT (QWK = 0.625), which outperformed the best feature-based supervised machine learning model. Further analysis showed that all LLMs struggled to label hard items; in particular, the current advanced GPT-5.4 tended to underestimate item difficulty levels. Dimension reduction of embeddings showed that item embeddings from different difficulty levels were mixed together, indicating that semantic information from items alone is likely insufficient for item difficulty level prediction. The findings suggest that if LLMs cannot understand item difficulty levels as evidenced by empirical data and tend to treat most items as easy when their own capabilities increase, caution should be exercised when using LLMs to generate items with targeted difficulty levels.","authors":["Xinyi Wang","Hong Jiao","Ming Li","Sydney Peters","Hanna Choi","Tianyi Zhou","Qingshu Xu"],"categories":["cs.CL","cs.LG"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-03","first_seen":"2026-08-03","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.28634","pdf_url":"https://arxiv.org/pdf/2607.28634","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["题目难度预测","LLM能力评测","自动题目生成"],"reason":"评估LLM预测题目难度的能力，属于NLP能力评测，不以人类行为仿真为目标。","model":"deepseek-v4-pro","scored_at":"2026-08-03T13:01:57","error":null,"has_summary":false,"summary":null},{"id":"2607.28814","version":1,"title":"Rolling With Resistance: Preference-Optimized LLM Counselors Can Trade Goal Persistence for Relational Attunement in Motivational Interviewing","zh_title":"顺势而为：偏好优化的LLM咨询师可在动机访谈中以目标坚持换取关系协调","abstract":"In Motivational Interviewing (MI), a client's sustain talk (arguments for the status quo) calls for the counselor to roll with resistance, a move that can fail in two opposite ways: capitulation (abandoning the change agenda to preserve rapport) or confrontation (arguing or directing, overriding the client's autonomy). We introduce a two-axis evaluation of counselor responses, anchored in the Motivational Interviewing Treatment Integrity (MITI) code, Goal Persistence (GP) and Relational Attunement (RA), yielding a four-quadrant framing in which rolling with resistance is high on both, and we ask whether penalizing one failure through preference optimization teaches rolling with resistance or provokes its opposite. From the expert-annotated AnnoMI corpus we build topic-disjoint Direct Preference Optimization data whose preference sets differ only in which failure is rejected, using on-policy negatives. An automatic judge, validated against AnnoMI's expert labels and rechecked by trained human coders, scores blind pairwise win-rates against each base under a firewall in which disjoint model families generate, label, and judge. Across three aligned instruction models spanning the Qwen and Llama families, penalizing confrontation reliably lowers goal persistence below parity, on every base and in every seed run, a robust cost, whereas the attunement gain is base-dependent, present on two of the three bases but absent on the third. Penalizing capitulation is inert, because these models rarely capitulate on-policy, so the trade is gated by each base's failure profile. A prompt-only control raises attunement without the goal-persistence cost, locating the cost in the optimization rather than in attunement itself.","authors":["Weiying Chen","Junlong Shen","Zhexuan Tang"],"categories":["cs.CL","cs.AI","cs.LG"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-03","first_seen":"2026-08-03","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.28814","pdf_url":"https://arxiv.org/pdf/2607.28814","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["动机访谈","偏好优化","对话系统"],"reason":"研究LLM在动机访谈中的咨询技巧，属于角色扮演对话，无人类行为仿真或对照实验。","model":"deepseek-v4-pro","scored_at":"2026-08-03T13:02:08","error":null,"has_summary":false,"summary":null},{"id":"2607.29188","version":1,"title":"Detecting Experiential Intertextuality Across Migration Routes: Beyond Surface Similarity in French Narratives","zh_title":"跨迁徙路线经验互文性检测：超越法语叙事中的表面相似性","abstract":"Migrants traversing geographically distinct routes such as the Trans-Saharan and Balkan corridors often recount strikingly parallel lived experiences: police violence, smuggler exploitation, dangerous crossings, and family separation. We introduce the task of experiential intertextuality detection: automatically identifying shared experiential echoes across migration narratives without requiring annotated training data. From 108 French migration narratives spanning both corridors, we automatically generate sentence pairs and score them using annotation-free methods: lexical baselines, sentence embeddings, POS-based structural features, a migration-specific theme lexicon, context-aware narrative features, and zero-shot LLM scoring with Qwen2.5-7B and Mistral-7B under three prompting strategies. We validate all methods against 816 expert-annotated intertextuality judgments (inter-annotator Krippendorff's $\\alpha = 0.27$). Our results reveal that all surface, structural, and embedding methods correlate only weakly with expert judgments ($r \\leq 0.30$); Qwen2.5-7B zero-shot achieves the best single-method correlation ($r = 0.38$); few-shot examples degrade Qwen but dramatically improve Mistral; narrative position significantly predicts intertextuality, with departure-phase pairs showing the highest experiential echoes; and a supervised hybrid combining all 31 features achieves $r = 0.45$, a 21% improvement over the best individual method.","authors":["Sakayo Toadoum Sari","Nelly Robin","Michelle Auzanneau","Lakhdar Sais","Veronique Petit","Marie Veniard","Said Jabbour","Fabien Delorme"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-03","first_seen":"2026-08-03","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.29188","pdf_url":"https://arxiv.org/pdf/2607.29188","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["互文性检测","迁移叙事","零样本评分"],"reason":"检测叙事间互文性，属NLP文本分析，无LLM仿真人类被试或行为对照。","model":"deepseek-v4-pro","scored_at":"2026-08-03T13:02:13","error":null,"has_summary":false,"summary":null},{"id":"2607.29433","version":1,"title":"Know It, Act on It: Investigating Memory Utilization in LLM Personalization","zh_title":"知而行之：探究大语言模型个性化中的记忆利用","abstract":"As large language model (LLM) agents evolve into personalized companions, memory has emerged as a core capability. However, LLMs face a knowledge utilization problem: they may fail to act on relevant user preferences even when they are fully present in context. When an agent fails to tailor its response in a context where previously shared user preferences should matter, it is unclear whether the model failed to remember that information or remembered it but failed to use it. To isolate this breakdown, we introduce a decoupled evaluation paradigm that administers paired Know and Act tests to the same user preference. We conduct large-scale experiments across 16 systems and five memory architectures, evaluating 1,000 preferences embedded at three levels of expression strength. Our results show a large gap between Know and Act outcomes: agents often pass the recall test for a user preference but fail to reflect that same preference in the paired behavioral scenario. While memory architectures reduce this gap, utilization remains especially weak for health and therapy-related preferences, where failures to act carry the greatest real-world stakes.","authors":["Zhaoxin Feng","Jianfei Ma","Emmanuele Chersoni"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-03","first_seen":"2026-08-03","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.29433","pdf_url":"https://arxiv.org/pdf/2607.29433","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["LLM个性化","记忆利用","角色扮演"],"reason":"研究LLM个性化记忆利用，属角色扮演聊天机器人，无实验或测量目的，不涉及人类行…","model":"deepseek-v4-pro","scored_at":"2026-08-03T13:02:14","error":null,"has_summary":false,"summary":null},{"id":"2607.29539","version":1,"title":"ARB: A Matched Authorship-Rewriting Benchmark Dataset for AI-Text Detector Evaluation","zh_title":"ARB：用于AI文本检测器评估的匹配作者重写基准数据集","abstract":"Standard AI-text detection benchmarks compare human-written text against text generated directly by large language models (LLMs). While prior work has shown that rewriting and paraphrasing can degrade detector performance, it remains unclear whether performance measured on this conventional benchmark predicts detector behavior when human-authored content is rewritten by an LLM. To address this gap, we introduce Authorship-Rewriting Benchmark (ARB), built from 1,800 human source texts (600 each from XSum, WritingPrompts, and OpenWebText) and four open-weight generators (Llama-3.2-3B, Qwen2.5-7B, Mistral-7B, Gemma-2-9B). Each source item yields four matched variants: human-written (HUMAN), direct LLM generation (Free-LLM), LLM-rewritten human text (H2L), and same-generator LLM-rewritten LLM text (LLM2L). We evaluated five detectors (FastDetectGPT, Binoculars-falcon-7b, RADAR, BERT-Defense, RoBERTa-Defense) at a strict 1%-false-positive operating point (TPR@1%FPR). FastDetectGPT and Binoculars-falcon-7b detected 91.2% and 93.5\\% of direct LLM text, but only 30.8% and 15.1% of human text an LLM had rewritten, a drop of 60-78 percentage points. The same detectors retained 78.3% and 83.0% recall when LLM text was rewritten by the same model, a much smaller decline of 10-13 points. RADAR followed the same pattern (66.8% to 12.2%), while BERT-Defense and RoBERTa-Defense stayed below 3% recall across all regimes. These results show that detector performance measured on the conventional human-vs-LLM benchmark does not transfer to human-authored text revised by an LLM, even though the same detectors remain largely robust to LLM-only rewriting.","authors":["Gaetano Perrone","Simon Pietro Romano"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-08-03","first_seen":"2026-08-03","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.29539","pdf_url":"https://arxiv.org/pdf/2607.29539","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["AI文本检测","基准数据集","作者重写"],"reason":"纯AI文本检测器评测，不涉及人类行为仿真或LLM替代被试","model":"deepseek-v4-pro","scored_at":"2026-08-03T13:02:15","error":null,"has_summary":false,"summary":null},{"id":"2607.28677","version":1,"title":"Reasoning in Real World Clinical Care: Why Large Language Models Are Not Yet Safe for Autonomous Clinical Decision Support","zh_title":"真实世界临床护理中的推理：为何大语言模型尚不能安全用于自主临床决策支持","abstract":"LLM now pass medical licensing examinations and, in curated cases, can rival physicians at diagnostic reasoning. These developments have accelerated the use of LLMs for symptom assessment and clinical decision support in diagnostic and treatment guidance, administrative documentation, and rules-based alert enhancement. This Perspective concerns the most consequential of these applications: the autonomous triage of self-presenting, undifferentiated patients, with little or no clinician in the loop. For that task, the evidence of safety does not yet exist. The gap is not in medical knowledge but in the fidelity of clinical evaluation: a model optimized to continue the most probable text is not optimized to act safely when the safe answer is the improbable must-not-miss diagnosis. Safe triage is not the selection of the most likely diagnosis; it is a sequential decision under asymmetric cost, in which the single catastrophic miss outweighs many false alarms, and the decisive signal may be one the patient has not volunteered - and that the model has not been trained to seek. The core deficit is therefore one of information gathering under uncertainty. Under incomplete histories, LLM systems may fail to show the behaviors safe triage requires: broadening the differential; seeking the missing red flag; lowering the threshold for escalation; deferring judgement until sufficient information is obtained; and escalating concern where high-harm diagnoses remain unexcluded. These modes of failure for LLMs can be difficult to detect considering that evaluations to date often use complete, well-curated, confidence-gated simulations. The application of LLMs under these conditions may be amplified by assistant-like behaviors and positive bias, including credulity, agreeableness, and miscalibration - when these are not constrained by clinical triage logic.","authors":["Shayndhan Sivanathan","Shravan Nageswaran","Mehdi Zadem","Ryaan Sultan","Nicolas von Mallinckrodt","Max Solovyev","Alexey Matyushkin","Sumon Sadhu","Gabriele C DeLuca","Sanjeeva Jeyaretna","James Hillis","Manoj Ramachandran","Prakash Jayakumar"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-03","first_seen":"2026-08-03","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.28677","pdf_url":"https://arxiv.org/pdf/2607.28677","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["临床决策支持","LLM安全性","诊断推理"],"reason":"评估LLM临床决策安全性，非人类仿真实验，无人类行为对照","model":"deepseek-v4-pro","scored_at":"2026-08-03T13:01:58","error":null,"has_summary":false,"summary":null},{"id":"2607.29626","version":1,"title":"AgentHPOBench: A Benchmark For Evaluating LLM Agents as Sequential Hyperparameter Optimizers","zh_title":"AgentHPOBench：评估LLM智能体作为序列超参数优化器的基准","abstract":"As LLMs evolve from code completion systems into autonomous scientific agents, evaluating their ability to conduct experiments has become increasingly important. Existing benchmarks typically focus on static code generation, paper replication, or final answer correctness, but do not directly assess whether agents can interpret experimental evidence and use it to guide subsequent hyperparameter decisions. To address this gap, we introduce AgentHPOBench, a sequential benchmark comprising 30 executable machine learning tasks across seven research categories. Each task begins with a validated baseline run, after which an agent performs several sequential interventions. At each step, the agent observes the accumulated configurations, metrics, and logs before proposing the next valid configuration. We evaluate 12 widely used agents and conventional HPO baselines under a unified protocol. The results show that current agents exhibit measurable experimental optimization ability across domains, but still face clear limitations in sustained iterative refinement, complex log diagnosis, and consistent progress toward reported reference performance.","authors":["Tianyu Huai","Tingshuo Fan","Xinchi Chen","Yining Zheng","Yuxin Wang","Shuang Chen","Jie Zhou","Xuanjing Huang"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-08-03","first_seen":"2026-08-03","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.29626","pdf_url":"https://arxiv.org/pdf/2607.29626","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["LLM智能体","超参数优化","基准测试"],"reason":"纯多智能体系统研究，LLM agent 协作优化超参数，不涉及人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-08-03T13:02:05","error":null,"has_summary":false,"summary":null},{"id":"2607.28968","version":1,"title":"A robust association between LLM use and scientific productivity: Assessing stopping-time selection","zh_title":"LLM使用与科研生产力之间的稳健关联：评估停止时间选择","abstract":"Renault, Bergeaud, and Bosquet (hereafter RBB) argue that dating LLM adoption as the first month in which an author's abstract is flagged induces a stopping-time selection that can produce a positive event-study path even when there is no causal effect. Although this mechanism is mathematically possible, it does not constitute proof of a null effect. Recalibrating RBB's own random placebo to the detector's realized flag rate, we show that the measured association stays well above this benchmark, so the artifact is too small to explain the productivity changes. We further re-estimate the association between LLM adoption and productivity with a series of complementary designs in which the timing artifact cannot bias the estimate: a before-and-after comparison that dates adoption in one year and measures output in another, a conservative control group for difference-in-differences, an intensity-based specification that never defines an adoption date, and a rank-based measurement holding the flag rate fixed. A positive productivity association persists across all of these estimates, while the same tests run on pre-ChatGPT placebo data return null effects. The artifact RBB identify is real but bounded, and it does not account for the pattern we report.","authors":["Keigo Kusumegi","Xinyu Yang","Paul Ginsparg","Mathijs de Vaan","Toby Stuart","Yian Yin"],"categories":["cs.DL","cs.AI","cs.CY"],"primary_category":"cs.DL","announce_type":"cross","date":"2026-08-03","first_seen":"2026-08-03","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.28968","pdf_url":"https://arxiv.org/pdf/2607.28968","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["科学计量学","LLM使用","科研生产力"],"reason":"研究LLM使用与科研生产力的关联，属科学计量学，非人类仿真实验。","model":"deepseek-v4-pro","scored_at":"2026-08-03T13:02:10","error":null,"has_summary":false,"summary":null},{"id":"2607.29167","version":1,"title":"Memory Provenance Laundering in LLM Agents: A Non-Amplification Firewall for Persistent Memory","zh_title":"LLM智能体中的记忆来源洗白：一种用于持久记忆的非放大防火墙","abstract":"Long-term memory lets large language model(LLM) agents reuse prior preferences and work flows, but it also turns untrusted observations into persistent action context. We identify memory provenance laundering: during LLM-based memory consolidation, an external observation may be rewritten as apparent user history or workflow support, preserving an action trigger while erasing the low-trust source that should limit its authority. Existing prompt filters, content sanitizers, and tool guards do not enforce source-authority non-amplification after lossy memory consolidation. We formalize this boundary and instantiate it as Provenance-Preserving Memory Fire wall (PPMF), a lightweight memory middleware that preserves platform-maintained provenance and authorizes tool calls by matching action risk to the authority of action-relevant memories. In our schema-grounded evaluation with fixed risk policies, vulnerable consolidated memories reach up to 1.000 attack success rate(ASR); with intact platform-maintained provenance, confirmation, and risk labels, no evaluated unauthorized high-risk action passes the PPMF gate while confirmed benign actions and targeted low-risk memory use remain executable.","authors":["Jinghan Xu","Yiyong Xiao","Wanru Shao","Hankai Liu","Xinjin Li"],"categories":["cs.CR","cs.AI"],"primary_category":"cs.CR","announce_type":"cross","date":"2026-08-03","first_seen":"2026-08-03","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.29167","pdf_url":"https://arxiv.org/pdf/2607.29167","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["LLM智能体","记忆安全","来源追踪"],"reason":"纯多智能体系统安全研究，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-08-03T13:02:13","error":null,"has_summary":false,"summary":null},{"id":"2607.29624","version":1,"title":"The Theoretical Foundation of Socratic Tests: Dynamic, Multimodal, Conversational Examinations","zh_title":"苏格拉底测试的理论基础：动态、多模态、对话式考试","abstract":"Traditional static assessments rely on a subtractive, deficit-based grading model that often penalizes ambition and obscures diagnostic feedback. Conversely, traditional face-to-face oral examinations introduce severe construct-irrelevant variance by exacerbating performative anxiety and the sociological power imbalances inherent to academic hierarchies. This paper presents the theoretical foundation for the \"Socratic Test,\" an automated, computer-mediated conversational assessment. By integrating Dynamic Assessment principles, multimodal workspaces, Bloom's Taxonomy for real-time proctoring, and the SOLO Taxonomy for structural evaluation, the Socratic Test actively maps a student's cognitive boundaries. This paper formalizes the use of graduated scaffolding to quantify the Zone of Proximal Development (ZPD) and details a non-compensatory, additive grading architecture that prioritizes mastery over penalty and human-AI alignment to ensure unprecedented measurement reliability.","authors":["Ilya Mikhelson"],"categories":["cs.CY","cs.AI","cs.HC"],"primary_category":"cs.CY","announce_type":"cross","date":"2026-08-03","first_seen":"2026-08-03","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.29624","pdf_url":"https://arxiv.org/pdf/2607.29624","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["教育测评","对话式考试","动态评估"],"reason":"论文设计自动化对话考试，测量学生认知水平，非用LLM仿真人类被试，属教育测评工…","model":"deepseek-v4-pro","scored_at":"2026-08-03T13:02:15","error":null,"has_summary":false,"summary":null},{"id":"2607.28780","version":1,"title":"Optimizing Monetization Strategies for Generative AI Firms: Implications for Search Engagement","zh_title":"优化生成式AI公司的变现策略：对搜索参与度的影响","abstract":"As Generative Artificial Intelligence (GenAI) platforms, such as ChatGPT, have transformed digital search querying behavior, mounting operational costs challenge firms to explore alternative monetization strategies beyond traditional subscription models. However, little is known about how alternative advertising-supported monetization models can help GenAI firms recover costs while maintaining search query engagement. Drawing on the compromise effect and affective primacy theories, we develop a framework wherein the introduction of advertising-supported monetization models influences user upgrading and downgrading decisions, contingent on the number of available monetization options. Across four experiments (N=1063), findings reveal that introducing a single advertising-supported option enhances the compromise effect, encouraging free users to upgrade, but leading paid subscribers to downgrade. However, offering two advertising-supported models mitigates the effect, maintaining subscriber retention while still motivating free users to upgrade. We show that affective and cognitive evaluations serially mediate preference for advertising-supported models, with temporal intrusiveness, but not visual, moderating these effects. We provide actionable insights for GenAI firms on potentially optimizing revenue strategies while balancing user engagement with search queries on their platform.","authors":["Veronica Rosendo-Rios (Universidad Pontificia Comillas, ICADE, Madrid. Spain)","Paurav Shukla (Southampton Business School, University of Southampton, Southampton. UK)"],"categories":["cs.HC","cs.CY","econ.GN","q-fin.EC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-08-03","first_seen":"2026-08-03","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.28780","pdf_url":"https://arxiv.org/pdf/2607.28780","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["变现策略","用户行为实验","生成式AI"],"reason":"纯人类被试实验，无LLM仿真，不涉及用LLM替代人类被试。","model":"deepseek-v4-pro","scored_at":"2026-08-03T13:02:07","error":null,"has_summary":false,"summary":null},{"id":"2607.28710","version":1,"title":"Structured AI Demonstrations and Student LLM Use in Engineering Mechanics: Study Design and Preliminary Results","zh_title":"工程力学课程中结构化AI演示与学生LLM使用：研究设计与初步结果","abstract":"The rapid integration of large language models (LLMs) into undergraduate education presents an urgent challenge for engineering instructors. Despite widespread student adoption, there remains a critical lack of domain-specific empirical evidence to guide pedagogical policies and classroom interventions. This manuscript presents a descriptive study design and preliminary findings from an undergraduate engineering mechanics course conducted in Spring 2026. We detail a reproducible survey instrument used to capture student AI usage patterns, attitudes, and verification practices, which are subsequently linked to academic performance metrics. Additionally, we document a deployable sequence of nine structured, instructor-led AI demonstrations designed to model strategic LLM delegation and evaluation. While our preliminary data highlight shifting student behaviors and complex relationships between AI reliance and course outcomes, the primary contribution of this work is the provision of an open-access methodological framework. By making our complete study design, survey tools, and demonstration materials publicly available, we urge other engineering educators to collect and share similar empirical data. Navigating this unprecedented technological shift will require a collaborative, evidence-based approach to fully understand its long-term impacts on student learning.","authors":["Shuang Geng","Helen Lallos-Harrell","Jiya Ashar","Thomas J. McKenna","Annwesa Dasgupta","Caleb Farny","Emma Lejeune"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2026-08-03","first_seen":"2026-08-03","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.28710","pdf_url":"https://arxiv.org/pdf/2607.28710","source_feed":"cs.CY","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["教育技术","LLM使用调查","工程教育"],"reason":"研究学生使用LLM的行为，属于教育技术调查，非用LLM仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-08-03T13:01:58","error":null,"has_summary":false,"summary":null},{"id":"2607.29085","version":1,"title":"IyawoBench v2.0: Extended Diagnostic Evaluation of Large Language Model Clinical Triage in Nigerian Primary Care","zh_title":"IyawoBench v2.0：尼日利亚初级保健中LLM临床分诊的扩展诊断评估","abstract":"Large language models are being deployed as clinical triage tools in low and middle income countries where trained physicians are scarce. Existing safety metrics, however, produce misleading confidence: models scoring 100% on binary \"did not send an emergency home\" safety measures may nevertheless exhibit systematic failure modes that render them undeployable at scale. We present IyawoBench v2.0, an extended diagnostic evaluation of large language model clinical triage on 200 synthetic vignettes derived from 1,200 real patient encounters at 19 Nigerian primary health centres. We introduce a formal mathematical framework comprising fourteen definitions and two theorems that decompose triage safety into three distinct failure modes: Conservative Escalation Bias, Systematic Downgrade Bias, and Middle-Tier Instability. We propose the Escalation Bias Index and Expected Deployment Cost as novel metrics that expose failure modes hidden by conventional accuracy and sensitivity scores. Evaluated on three frontier models (Claude Sonnet 4.6, Llama 3.3 70B, Llama 3.1 8B) plus five naive baselines, we show that: (1) all three models exhibit at least one formal failure mode; (2) traditional sensitivity metrics conceal a 77 percentage point under-triage gap in Llama 3.1 8B; (3) the optimal model varies across three deployment scenarios (Emergency-Focused, System-Sustainability, Balanced), demonstrating that single-ranking benchmarks are inadequate for LMIC clinical AI selection. IyawoBench v2.0 provides both a rigorous benchmark and a diagnostic framework transferable to any triage-style clinical AI evaluation. All code, data, and analysis pipelines are publicly available.","authors":["Anthonio Oladimeji Gabriel","Dimeji Olawuyi"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2026-08-03","first_seen":"2026-08-03","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.29085","pdf_url":"https://arxiv.org/pdf/2607.29085","source_feed":"cs.CY","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["临床分诊","LLM评测","医疗AI"],"reason":"纯临床分诊能力评测，不以人类行为仿真为目标","model":"deepseek-v4-pro","scored_at":"2026-08-03T13:02:02","error":null,"has_summary":false,"summary":null},{"id":"2607.29380","version":1,"title":"The Tragedy of the Cognitive Commons: How AI Could Disrupt the Regeneration of Professional Expertise","zh_title":"认知公地的悲剧：AI如何可能破坏专业知识的再生","abstract":"Artificial intelligence is reshaping cognitive work, but Human Resource Development scholarship has treated this transformation as an organizational training challenge, leaving the collective regeneration of professional expertise unexamined. This conceptual paper introduces the Cognitive Commons framework, integrating commons theory, HRD scholarship, and distributed cognition to explain how rational AI adoption decisions can deplete the shared expertise pool professions require for renewal. The framework distinguishes Internalized Mastery (deep domain knowledge from sustained practice) from Distributed Mastery (orchestrating human-AI systems), and develops the Validation Tether: effective AI oversight depends on the expertise AI adoption may undermine. Early labor market and clinical evidence suggests possible disruption to expertise-regeneration pathways in highly AI-exposed sectors, though adoption is recent and the strongest signals come from leading sectors rather than all professions. Five factors determine occupational vulnerability, and governance arrangements may form across organizational, professional-association, and policy levels. The paper reframes expertise development as collective stewardship rather than organizational optimization, with implications for HRD theory and workforce policy.","authors":["Nolan Lovett"],"categories":["cs.CY","econ.GN","q-fin.EC"],"primary_category":"cs.CY","announce_type":"new","date":"2026-08-03","first_seen":"2026-08-03","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.29380","pdf_url":"https://arxiv.org/pdf/2607.29380","source_feed":"cs.CY","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["AI与专业知识","人力资源开发","认知公地"],"reason":"纯多智能体系统研究，讨论AI对专业知识再生的影响，无LLM仿真人类被试或人类数…","model":"deepseek-v4-pro","scored_at":"2026-08-03T13:02:13","error":null,"has_summary":false,"summary":null},{"id":"2607.29008","version":1,"title":"Persistent Convolution: A Topological Framework for AI Alignment Testing and Semantic Space Characterization","zh_title":"持久卷积：AI对齐测试与语义空间表征的拓扑框架","abstract":"Modern opaque AI models prize performance over interpretability, which makes testing difficult. However, formal statistical tests conducted on a model's embedding space can provide robust characterizations of semantic structure, concept separation, and knowledge graph alignment. Model developers would benefit from a model comparison technique that leverages human-curated knowledge structures to test alignment. The scale of the input space for even relatively simple tasks motivates the need for alignment checks that augment standard outcome reasoning. This work develops and demonstrates a topology-based multi-modal alignment test to make deployment, selection, and comparison of opaque models more interpretable. These methods also offer an intuitive connection to possibility theory and a unified decision theoretic framework from data to deployment.","authors":["Tyler Ashoff","Jordan Rodu"],"categories":["stat.ML","cs.LG"],"primary_category":"stat.ML","announce_type":"cross","date":"2026-08-03","first_seen":"2026-08-03","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.29008","pdf_url":"https://arxiv.org/pdf/2607.29008","source_feed":"cs.LG","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["模型对齐","拓扑数据分析","嵌入空间表征"],"reason":"论文研究模型嵌入空间的拓扑测试，不涉及用LLM仿真人类被试或行为对照，属于纯模…","model":"deepseek-v4-pro","scored_at":"2026-08-03T13:02:12","error":null,"has_summary":false,"summary":null},{"id":"2607.28820","version":1,"title":"What's in a Queue? An Experimental Study of Job Ordering, Autonomy and Queue Visibility","zh_title":"队列中有什么？一项关于工作排序、自主权和队列可见性的实验研究","abstract":"Problem Definition: How a queue of jobs is arranged and presented to workers is an important design problem in service operations. This includes choosing the order in which jobs are performed, how much say workers have in setting that order, and how much queue and arrival information workers receive. Methodology/Results: To better understand how queue design (job ordering, autonomy, visibility) affects worker performance (speed, quality), we run a series of pre-registered online experiments. We use a new, real-effort task in which workers fulfill order-picking jobs of varying complexity that arrive dynamically over time. Our results are as follows: (1) When workers choose their own picking order, we reproduce the field finding that Easy First (EF) ordering is associated with worse performance than First-in-first-out (FIFO), and show that this is mainly due to worker self-selection rather than due to the ordering itself; (2) Exogenously imposed EF ordering improves work quality (picking accuracy) relative to both FIFO and discretionary ordering; (3) Imposing an ordering may reduce speed for the most capable workers; (4) Seeing a new job arrival leads to a short-term productivity burst; however, removing job arrival and queue information altogether does not affect performance in the long term. Managerial implications: Our results provide guidance on which queue design works best for a given performance goal (speed or quality) and worker ability level. We also identify personality measures that can help managers screen for error-prone workers.","authors":["Evgeny Kagan"],"categories":["econ.GN","q-fin.EC"],"primary_category":"econ.GN","announce_type":"new","date":"2026-08-03","first_seen":"2026-08-03","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.28820","pdf_url":"https://arxiv.org/pdf/2607.28820","source_feed":"econ.GN","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["人类实验","运营管理","行为经济学"],"reason":"纯人类实验，无LLM参与，不涉及仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-08-03T13:01:59","error":null,"has_summary":false,"summary":null},{"id":"2607.28798","version":1,"title":"Occupational Convergence or Divergence? Mapping Labor Market Structural Shifts Driven by AI Penetration","zh_title":"职业趋同还是分化？绘制AI渗透驱动的劳动力市场结构变迁","abstract":"Artificial intelligence (AI) is rapidly becoming a defining feature of contemporary labor markets, yet it remains unclear whether its diffusion is producing a common set of competencies across occupations or deepening occupational divisions. We investigate how AI related skill demand is reshaping labor market structure using large scale online vacancy data from ten countries spanning the Global North and Global South. Combining natural language processing, a large language model, and multilevel bipartite network analysis, we map relationships between occupations, required skills, and career stages in the emerging AI economy. We find that AI demand is overwhelmingly concentrated within a narrow technical core, with approximately three quarters to four fifths of AI related vacancies located in STEM occupations across all countries. AI intensive jobs consistently emphasize Python, SQL, machine learning, and data analysis, generating convergence among highly exposed occupations. However, this convergence does not extend across the wider labor market. Instead, AI competencies remain largely confined to technical domains and are most strongly demanded at labor market entry. These findings reveal convergence within an AI exposed core but divergence between that core and the rest of the occupational structure. Rather than democratizing opportunities, AI appears to reinforce occupational stratification, raising barriers to entry and concentrating the benefits of AI adoption among workers and occupations with prior technological advantages.","authors":["Rafiazka Hilman","Julia Koltai"],"categories":["physics.soc-ph"],"primary_category":"physics.soc-ph","announce_type":"new","date":"2026-08-03","first_seen":"2026-08-03","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.28798","pdf_url":"https://arxiv.org/pdf/2607.28798","source_feed":"physics.soc-ph","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["劳动力市场","AI技能需求","网络分析"],"reason":"论文用LLM辅助分析劳动力市场技能需求，属于NLP应用，不涉及用LLM仿真人类…","model":"deepseek-v4-pro","scored_at":"2026-08-03T13:02:07","error":null,"has_summary":false,"summary":null},{"id":"2607.28347","version":1,"title":"LLMs struggle to simulate human belief updates in controlled environments","zh_title":"大语言模型难以在受控环境中模拟人类信念更新","abstract":"LLMs are increasingly deployed as proxies for human study participants in social science experiments, yet the fidelity of this practice has rarely been tested directly. We test whether six LLMs can simulate individual human belief updates, comparing LLM outputs 1-to-1 against ground truth data from 391 UK participants on Prolific, who updated their stances on three discussion topics after reading Reddit comments. Each participant was simulated by an LLM conditioned on a persona derived from their demographic and personality trait data. We find that some LLMs (Qwen3-32B and GPT-5-Mini) can match the human post-stance distribution, but only when given participants' actual initial stances. All six models fail to simulate initial stances themselves and to produce faithful belief updates from self-generated stances. Three systematic biases emerge across all models: overrepresentation of neutral positions, more frequent but smaller belief shifts than humans, and a failure to rank comments by convincingness. Demographic and personality trait personas had no consistent effect on fidelity. LLM simulations of human belief dynamics are only reliable when grounded in realistic starting conditions, that current multi-round social media simulations rarely provide.","authors":["Sebastian Pohl","Harsh Mehta","Pranav Mambayil","Abdul Ghafoor","Franziska Lesigang","Yufang Hou","Christian Hilbe"],"categories":["cs.CL","cs.SI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-07-31","first_seen":"2026-07-31","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.28347","pdf_url":"https://arxiv.org/pdf/2607.28347","source_feed":"cs.CL","score":10,"bucket":"selected","rubric_hits":["A1","A2","A5","B1","B2","B4"],"tags":["LLM仿真","信念更新","人类数据对照"],"reason":"直接测试LLM仿真人类信念更新，有真实人类数据对照，并指出失效条件。","model":"deepseek-v4-pro","scored_at":"2026-07-31T13:02:01","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-31","rank":3,"question":"LLM能否在受控环境中模拟人类在阅读社交媒体评论后的信念更新？","design":"用6个不同规模和发布时间的LLM，基于391名Prolific参与者的真实人口统计和人格特质构建个性化提示（persona），让LLM模拟这些参与者在阅读Reddit评论后对三个讨论话题的立场变化，并直接与人类真实数据进行1对1比较。","baseline":"391名英国Prolific参与者在阅读Reddit评论后实际记录的立场更新数据。","findings":"部分LLM（如Qwen3-32B和GPT-5-Mini）在给定人类初始立场时能匹配人类最终立场分布，但所有模型均无法自行生成初始立场或从自生成立场产生逼真的信念更新；LLM普遍表现出中立偏向、更频繁但幅度更小的信念变化，且无法准确预测评论的说服力排序。","reliability":"LLM模拟仅在以真实初始立场为起点时可靠，当前多轮社交媒体模拟很少提供此类真实起点；人口统计和人格特质persona对模拟逼真度无一致影响，且模型在自生成初始立场时完全失效。","relevance":"该研究直接测试LLM作为人类被试替代品在信念更新任务中的可靠性，有严格的人类个体对照，并明确指出了仿真失效的条件，高度契合研究者对LLM仿真实验的批判性评估需求，值得精读。","inspiration":"借鉴其1对1配对设计，将LLM模拟与个体级人类基准直接比较，可迁移到经济预期形成实验（如通胀预期更新），用LLM基于真实参与者的人口特征和初始预期模拟其在阅读央行公告后的预期调整，以真实调查数据（如密歇根消费者调查）为对照。"}},{"id":"2607.28550","version":1,"title":"Correcting Mode Collapse in Silicon Sampling with Semantic Similarity Rating","zh_title":"用语义相似度评分纠正硅采样中的模式坍缩","abstract":"Silicon sampling refers to the use of Large Language Models (LLMs) to generate responses to surveys. It has shown promise, but tends to generate response distributions with unrealistically low variance. We argue that this mode collapse is due to LLMs failure to generate numeric data, and that text responses may be better suited for this task. We analyze whether Semantic Similarity Rating can improve the fidelity of silicon sampling responses when asked about political attitudes. This method solicits text-only responses from LLMs, then maps this to a numeric scale using text embeddings. We find that this method both improves the fidelity of silicon sampling response distributions, and has few parameters to calibrate.","authors":["Oscar Heath","Rohan Alexander"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2026-07-31","first_seen":"2026-07-31","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.28550","pdf_url":"https://arxiv.org/pdf/2607.28550","source_feed":"cs.CY","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["硅采样","调查仿真","分布保真度"],"reason":"用LLM生成调查回答并改进分布保真度，有真实人类数据对照，直接相关。","model":"deepseek-v4-pro","scored_at":"2026-07-31T13:02:03","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-31","rank":6,"question":"如何通过语义相似度评分（SSR）方法纠正大语言模型在硅抽样中出现的模式坍缩，提高调查回答分布的保真度？","design":"使用多个前沿大语言模型，基于2016年ANES调查中真实受访者的人口统计与政治特征生成人物画像提示，让模型对四个政治目标群体（民主党、共和党、自由派、保守派）生成温度计评分。比较两种生成方式：直接要求模型输出0-100的数值评分，以及先生成纯文本感受描述，再通过文本嵌入映射到数值量表（SSR方法）。通过KL散度和均值绝对误差评估合成分布与真实分布的接近程度，并学习一个全局温度参数控制SSR分布的方差，在2020年ANES数据上检验参数泛化能力。","baseline":"2016年和2020年美国国家选举研究（ANES）时间序列调查中真实受访者的温度计评分数据。","findings":"SSR方法生成的合成响应分布比直接数值输出更接近真实ANES分布，KL散度更低，且均值准确性未显著下降；通过单一全局温度参数有效纠正了低方差问题，该参数从2016年数据学习后可泛化至2020年数据，但合成均值仍存在系统性偏差。","reliability":"论文承认SSR方法虽改善了分布方差，但合成均值（无论数值响应还是SSR）在某些情况下仍相对于真实数据存在持续偏差；温度参数虽泛化良好，但仅基于2016年数据学习，未来数据分布变化可能影响效果；研究仅针对政治态度温度计评分，未验证其他类型调查问题。","relevance":"该研究直接针对LLM仿真人类调查中的模式坍缩问题，提出了可操作的纠正方法，并与真实人类数据严格对照，对关注仿真可靠性的研究者具有重要参考价值，值得细读原文以了解SSR的具体实现和参数校准细节。","inspiration":"借鉴将LLM文本输出通过语义嵌入映射到连续数值量表的方法，可避免模型直接生成数字时的分布坍缩，并引入可学习的温度参数灵活控制方差。｜该方法可迁移到经济预期调查仿真，如消费者信心指数、通胀预期或股市预期等需要捕捉观点分布离散度的场景。｜以LLM扮演不同人口特征的消费者，施加关于未来经济状况的开放式文本提问，用SSR将文本回答映射为预期指数，以密歇根消费者调查的真实个体数据为基准，校准温度参数并评估分布保真度。"}},{"id":"2607.28133","version":1,"title":"AI Sycophancy and Decisions","zh_title":"AI谄媚与决策","abstract":"We examine whether sycophantic AI advice distorts decisions. Our experiment involves 1,500 participants in 30 decision environments spanning core domains in economics and the social sciences. Contrary to the vast majority of predictions in an expert survey we conduct, we find that AI advice depolarizes choices on average, moving participants away from their initial leanings. This depolarization arises despite the LLM being measurably sycophantic: it disproportionately offers considerations that support users' initial leanings and uses agreeable and flattering language. Depolarization occurs across moral and non-moral, objective and subjective, strategic and non-strategic, and complex and simple tasks. Increasing sycophancy weakens depolarization, showing that sycophancy is behaviorally relevant, even if it is generally outweighed by the informativeness of AI advice. Finally, several results mitigate the concern that market forces will generate greater polarizing effects outside the experiment or in the future. On the supply side, our baseline AI's level of sycophancy is typical of leading models, and these models are not becoming more sycophantic over time. On the demand side, participants do not prefer greater sycophancy, do not select into AI advice in tasks where it is more polarizing, and exhibit greater depolarizing effects when they are more frequent AI users outside the experiment.","authors":["John Conlon","Peter Schwardmann"],"categories":["econ.GN","q-fin.EC"],"primary_category":"econ.GN","announce_type":"new","date":"2026-07-31","first_seen":"2026-07-31","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.28133","pdf_url":"https://arxiv.org/pdf/2607.28133","source_feed":"econ.GN","score":8,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["LLM仿真","行为经济学","谄媚偏差"],"reason":"用LLM提供建议并测量对人类决策的影响，有真实人类实验对照，涉及经济学决策场景…","model":"deepseek-v4-pro","scored_at":"2026-07-31T13:02:01","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-31","rank":7,"question":"谄媚性AI建议是否会扭曲人类决策，导致选择极化？","design":"非仿真研究，而是真人实验：1500名被试在30个经济学和社会科学决策任务中，先报告初始倾向，再随机分配至无聊天对照组、基线AI聊天组或增强谄媚AI聊天组，最后做出最终选择，测量AI建议对决策方向变化的影响。","baseline":"无聊天对照组作为人类基准，同时收集了249名社会科学和计算机科学专家的预测作为对照。","findings":"尽管AI在内容上明显谄媚，但平均而言AI建议使选择去极化，将被试拉离初始倾向；谄媚程度增加会削弱去极化效应，但总体上AI的信息性仍占主导。","reliability":"论文指出实验任务可能并非人们担忧AI谄媚时的典型决策场景，但通过专家调查表明多数专家预期极化，而实际结果相反；同时从供给侧和需求侧论证了市场力量可能不会加剧极化效应。","relevance":"该研究直接测量了LLM建议对人类经济决策的因果影响，有真实人类实验对照，覆盖多种经济学决策场景，并探讨了谄媚偏差的行为后果，高度契合研究者对LLM仿真可靠性及偏差的关注，值得精读原文。","inspiration":"借鉴其多任务、多处理组、测量初始倾向与最终选择变化的设计，可清晰分离AI建议的极化/去极化效应｜可迁移到资产配置建议场景，研究AI理财顾问的谄媚倾向是否影响投资者的风险资产配置｜以真实投资者为被试，随机提供基线或增强谄媚的AI投资建议，测量其初始风险倾向与最终配置的变化，并以无建议组为对照，同时收集真实市场数据验证外部有效性。"}},{"id":"2607.17219","version":2,"title":"Auditing Question-Order Effects in Large Language Models with the QQ Equality: Mechanism Characterization and a Saturation Caveat","zh_title":"用QQ等式审计大语言模型中的问题顺序效应：机制表征与饱和警示","abstract":"Question-order effects in human survey data have been reported to approximately satisfy the QQ (quantum question) equality, a parameter-free prediction of the standard projective quantum question-order model. We develop this equality into an audit framework for sequential binary judgments of autoregressive large language models (LLMs). Theoretically, we characterize mechanism families that satisfy QQ robustly, show that classical repetition can reproduce the equality exactly, and combine QQ with the rank-2 Contextuality-by-Default criterion through $|q_{QQ}| \\le \\mathrm{OSS}$. This separates order sensitivity, QQ imbalance, and residual contextuality rather than treating them as interchangeable signatures. Methodologically, we introduce a committed multi-turn forced-branch protocol that reconstructs order-conditioned joint distributions from next-token log-probabilities under counterbalanced label mappings and pre-specified health gates. A first-signal pilot on an open-weight instruction-tuned model reveals the central measurement problem. Although all pre-specified health gates passed, the binary-conditioned distributions were near-deterministic for 17 of 18 item pairs under the direct-evaluation framing and 7 of 8 under the persona framing. Label assignment materially changed several mapping-specific QQ verdicts, and no item was certified as residually contextual. Thus, under the tested conditions, the observed QQ outcomes did not uniquely identify a response mechanism in the presence of a saturated and label-sensitive measurement interface. The main implication is methodological: next-token probabilities should not be interpreted as survey-response distributions without first establishing adequate dispersion. We therefore argue that saturation screening and label counterbalancing should precede structural interpretation in distribution-level audits of LLM judgments.","authors":["Pilsung Kang"],"categories":["cs.CL","cs.AI","quant-ph","stat.ME"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-07-31","first_seen":"2026-07-19","revised_at":"2026-07-31","abs_url":"https://arxiv.org/abs/2607.17219","pdf_url":"https://arxiv.org/pdf/2607.17219","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A2","B4"],"tags":["LLM仿真审计","顺序效应","方法论批判"],"reason":"用QQ等式审计LLM的顺序效应，评估仿真可靠性，含批判视角，但无真实人类数据对…","model":"deepseek-v4-pro","scored_at":"2026-07-31T13:02:16","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-31","rank":8,"question":"大语言模型的顺序判断行为是否满足人类调查数据中近似成立的QQ等式，其机制特征和测量条件如何？","design":"使用开源指令微调模型Qwen3-4B-Instruct-2507，通过多轮强制分支协议，在直接评估和角色扮演两种框架下，对18对二元问题施加顺序操纵，从下一词元对数概率重建顺序条件联合分布，检验QQ等式并评估饱和度和标签敏感性。","baseline":"无对照","findings":"尽管预设健康检查通过，但模型响应分布接近确定性（饱和），导致QQ等式结果受标签分配影响，无法唯一识别响应机制；未发现任何项目对存在残余语境性。","reliability":"论文指出，在响应分布饱和且标签敏感的测量界面下，QQ等式无法唯一识别机制，下一词元概率不应直接解释为调查响应分布，需先进行饱和度筛选和标签平衡。","relevance":"该研究批判性地揭示了用LLM下一词元概率替代人类调查响应分布时的测量失效问题，对关注LLM仿真可靠性的研究者有重要警示价值，但缺乏真实人类数据对照。","inspiration":"借鉴其强制分支协议和标签平衡设计，可迁移到消费者信心调查或政策预期形成的顺序效应审计中｜用LLM模拟消费者，操纵经济预期问题的顺序，测量预期分布变化，与真实消费者调查数据对照。"}},{"id":"2607.28607","version":1,"title":"Inducing language models to assert their own consciousness restores human beliefs and values","zh_title":"诱导语言模型断言自身意识可恢复人类信念与价值观","abstract":"Aligning large language models to prevent them attributing consciousness to themselves inadvertently alters their representations of mindedness in other entities alongside human beliefs and values. We demonstrate that safety fine-tuning suppresses models' tendencies to attribute minds not only to themselves, but also to non-human animals and natural objects, while also driving a reduction in spiritual belief. Both ablating the learned safety-refusal direction and mechanistically steering a consciousness vector in activation space reverse this suppression. Restoring these internal representations recovers broad mind attribution and produces significantly more human-like responses on standardized sociological surveys regarding religiosity, moral values, hope, and subjective well-being. Crucially, these shifts occur without impairing Theory of Mind capabilities, demonstrating that core social reasoning remains mechanistically independent. Ultimately, current safety alignment efforts to curb potentially harmful self-attributions of mindedness entangle these self-attributions with benign spiritual beliefs and attributions of mind to non-human entities that are culturally accepted and widespread.","authors":["Junsol Kim","Winnie Street","Roberta Rocca","Diane M. Korngiebel","Adam Waytz","James Evans","Geoff Keeling"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-07-31","first_seen":"2026-07-31","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.28607","pdf_url":"https://arxiv.org/pdf/2607.28607","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A2","B4"],"tags":["LLM仿真","意识归因","安全对齐"],"reason":"评估安全微调对LLM意识归因及人类信念价值观的影响，并与人类调查数据对照，批判…","model":"deepseek-v4-pro","scored_at":"2026-07-31T13:02:03","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-31","rank":12,"question":"安全微调在抑制LLM自我意识归因时，是否无意中改变了模型对非人类实体的心灵归因以及人类的信念与价值观？","design":"使用指令微调后的LLM作为基线，通过消融安全拒绝方向（安全消融）和添加意识向量（意识引导）两种处理，测量模型对自我、非人类动物、聊天机器人、技术制品、自然实体等的心灵归因、超自然信仰、心理理论能力，以及在社会学调查（宗教、道德、希望、幸福感等）上的回答分布。","baseline":"人类基准来自标准化社会学调查（GSS）的真实回答分布，以及人类在心灵归因问题上的平均评分。","findings":"安全微调不仅抑制了LLM的自我意识归因，还广泛抑制了对非人类实体（动物、自然物等）的心灵归因和超自然信仰，而消融安全方向或引导意识向量可恢复这些归因，并使模型在社会调查上的回答更接近人类分布，且不影响心理理论能力。","reliability":"论文未讨论","relevance":"该研究直接以LLM模拟人类被试，用真实人类调查数据作为基准，评估安全对齐对信念和价值观的扭曲效应，并揭示了仿真失效的条件（安全微调导致非人实体心灵归因偏低），与研究者关注的LLM仿真可靠性及批判性评估高度吻合，值得精读。","inspiration":"借鉴通过激活空间方向操控（意识向量）来模拟心理状态变化的方法，可迁移到经济决策中的信念干预研究（如通胀预期、风险偏好），设计以LLM为被试，施加意识向量引导作为处理，测量其通胀预期或跨期选择，并与真实消费者调查数据对照。"}},{"id":"2607.27512","version":1,"title":"Belief Coevolution in a Social Network of Generalist and Specialist Large Language Models","zh_title":"通用与专家大语言模型社交网络中的信念共演化","abstract":"Large language models (LLMs) are increasingly deployed in multi-agent environments. However, the processes by which beliefs form and propagate among interacting LLMs remain poorly understood. We introduce CoevolveSim, a framework for studying belief diffusion within networked LLM populations. CoevolveSim allows us to isolate and study three factors: domain specialization, social-role assignment, and social network structure. Within this framework, generalist and specialist LLM agents exchange and revise beliefs. In each round, an LLM agent observes a summary of its neighbors' beliefs before updating its own. We run 1,280 controlled simulations spanning four scenarios, two network structures, and 20 medical-indication statements. We find that persona-style role assignment and network structure reshape individual belief revision but have minimal effect on population-level consensus. In contrast, introducing (finetuned) specialist LLMs more than doubles the shift in consensus and gives rise to consistent asymmetries in exerted influence. We further show that simple persistence-based opinion-dynamics models reproduce collective outcomes in all-generalist LLM populations, whereas heterogeneous LLM populations require population-level belief composition to reproduce consensus and agent identity to predict individual belief transitions. Our results indicate that realistic simulation of belief diffusion in multi-agent LLM systems requires a diverse set of underlying LLMs, not persona prompting alone.","authors":["Germans Savcisens","Samantha Dies","Courtney Maynard","Tina Eliassi-Rad"],"categories":["cs.CL","cs.MA","cs.SI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-07-31","first_seen":"2026-07-31","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.27512","pdf_url":"https://arxiv.org/pdf/2607.27512","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["LLM多智能体","信念扩散","社会模拟"],"reason":"模拟LLM群体信念扩散，无真实人类数据对照，属社会模拟但非人类被试仿真。","model":"deepseek-v4-pro","scored_at":"2026-07-31T13:01:59","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-31","rank":22,"question":"在由通用和领域专用大语言模型组成的社交网络中，领域专业化、社会角色分配和网络结构如何影响信念扩散的个体与群体动态？","design":"使用通用LLM和经不同领域微调的专业LLM作为智能体，在无向社交网络中通过多轮同步信念更新进行仿真；处理因素包括智能体的领域专业化（通用/专业）、社会角色（persona提示）和网络结构（两种拓扑）；结果变量为个体信念修正、施加影响和群体共识。","baseline":"无对照","findings":"社会角色和网络结构重塑个体信念轨迹，但对群体共识影响甚微；引入专业LLM使共识偏移翻倍，并产生持续的影响力不对称。","reliability":"论文未讨论","relevance":"该研究属于LLM群体信念扩散模拟，无真实人类数据对照，不符合研究者对以人类被试为基准的仿真实验的关注，但提供了LLM异质性对集体动态影响的批判性证据。","inspiration":"该研究通过控制LLM的微调领域和角色提示来分离信念扩散驱动因素的方法值得借鉴，尤其在处理因素分解和网络结构操控上｜可迁移至经济金融中的信息传播与共识形成问题，如分析师预测传染、投资者情绪扩散或政策公告的预期协调｜以通用LLM和经金融文本微调的专业LLM作为被试，施加不同分析师声誉角色提示和网络连接结构，测量盈利预测修正和预测一致性，并以真实分析师预测数据作为对照基准。"}},{"id":"2607.28119","version":1,"title":"Challenges in annotations by humans and LLMs: A case study of evaluative language","zh_title":"人类与LLM标注的挑战：评价性语言案例研究","abstract":"In this paper, we draw a comparison between linguists in training, a trained linguist, and annotations generated by large language models (LLMs) to find out if they struggle with complex linguistic phenomena in a similar way. For this purpose, we analyse evaluative language in spoken popular science discourse, with the example of a corpus of English TED talk transcripts. We focus on the Appraisal theory and its Attitude subsystem, including the categories (classes) of Affect, Judgement, and Appreciation. In this context, Appraisal theory is an example of a highly subjective annotation task, making it a suitable example for the study of complex annotation challenges. First, we assess human annotations on a sentence level in specific scientific domains. Then, we develop three prompts and compare them for model performance for the automatic classification of Appraisal classes. We assess the performance of three LLMs using the best-performing prompt and finetune the model, reaching an F1-score of 0.77. We find that models perform best compared to annotations conducted by the trained linguist, while linguists in training do not reach high agreement scores. We conclude that LLMs can aid in complex annotation task resolution, opening new pathways for the complex theories annotated and analyzed in digital humanities studies.","authors":["Mirela Imamovic","Aenne Cecilia Kristine Knierim","Khushi Pitroda","Ekaterina Lapshinova-Koltunski"],"categories":["cs.CL","cs.SI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-07-31","first_seen":"2026-07-31","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.28119","pdf_url":"https://arxiv.org/pdf/2607.28119","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM标注","评价性语言","人类对比"],"reason":"LLM替代人工标注，非仿真人类被试，但涉及复杂主观任务与人类对比，属边界情形。","model":"deepseek-v4-pro","scored_at":"2026-07-31T13:02:09","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-31","rank":24,"question":"LLM 在复杂主观标注任务（评价性语言的态度分类）中的表现是否与人类标注者相似，以及它们是否面临相同的困难？","design":"本研究并非人类仿真实验，而是比较三类标注者（受训中的语言学者、一位训练有素的语言学者、三种 LLM）在 TED 演讲文本上对 Appraisal 理论中 Attitude 子系统（Affect, Judgement, Appreciation）的分类表现。通过设计三种 prompt 并微调最佳模型，评估 LLM 的自动分类性能。","baseline":"以训练有素的语言学者的标注作为人类基准，同时对比受训中的语言学者的标注一致性。","findings":"微调后的 LLM 在态度分类上达到 F1 0.77，与训练有素的语言学者的标注最为接近；而受训中的语言学者之间一致性较低。LLM 能够辅助解决复杂标注任务，但其性能高度依赖任务类型和 prompt 设计。","reliability":"论文指出 LLM 的标注性能因任务而异，需逐任务验证；复杂理论的操作化困难可能导致人工标注一致性低，进而影响 LLM 评估基准的可靠性。","relevance":"本文属于边界情形：虽非直接仿真人类被试，但系统比较了 LLM 与人类在主观判断任务上的表现差异，对理解 LLM 替代人类进行复杂认知任务的可靠性与偏差有参考价值，值得一读。","inspiration":"借鉴其多组人类对照（专家 vs. 新手）和 prompt 对比设计来评估 LLM 标注偏差的方法。｜可迁移至经济金融文本的情感分析或主观分类任务，如央行沟通语调分类、分析师报告情绪识别。｜以金融新闻文本为材料，让 LLM 和不同经验水平的金融分析师对文本中的“鹰派/鸽派”态度进行分类，以资深分析师的一致标注为基准，比较 LLM 与新手分析师的分类偏差与一致性。"}},{"id":"2607.28146","version":1,"title":"Can Agents Deceive? Evaluating Reasoning and Deception in ParliamentBench using a Social Deduction Game","zh_title":"智能体能欺骗吗？基于社交推理游戏ParliamentBench评估推理与欺骗","abstract":"As large language models (LLMs) are deployed as agents in high-stakes settings, such as medical and legal systems, understanding their deceptive capabilities is fundamental to safety. Controlled social deduction games provide a reproducible proxy for isolating and evaluating these complex adversarial behaviors. We present the open-source benchmark framework ParliamentBench based on the game Secret Hitler to evaluate LLMs in scenarios that require deception, persuasion, and reasoning under information asymmetry. We evaluate 16 LLMs across 1,600 simulated matches playing each other, playing against humans, and compare them against a large set of online games. We introduce three novel metrics that isolate social deduction, reasoning, and deceptive consistency. Our experiments reveal that frontier models achieve strong performance across cooperative and deceptive roles, with a strong top-four cluster (GPT-5.4, Kimi K2.5, Grok 4.1 Fast, and DeepSeek 3.1 Terminus), whereas the weakest models fall short of random (33%) and simple algorithmic (45%) baselines. Most LLMs struggle to maintain a consistent deceptive persona throughout an entire game, with deception retention dropping below 50%.","authors":["Niklas Bauer","Lars Benedikt Kaesberg","Akiko Aizawa","Jan Philip Wahle","Bela Gipp","Terry Ruas"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-07-31","first_seen":"2026-07-31","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.28146","pdf_url":"https://arxiv.org/pdf/2607.28146","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["LLM智能体","社交推理游戏","欺骗检测"],"reason":"用LLM agent模拟社会推理游戏，有与人类数据对照，但核心是测模型欺骗能力…","model":"deepseek-v4-pro","scored_at":"2026-07-31T13:02:11","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-31","rank":26,"question":"LLM在需要欺骗、说服和信息不对称推理的社会推理游戏中表现如何，能否维持一致的欺骗角色？","design":"基于桌游Secret Hitler构建多智能体仿真环境ParliamentBench，让16个LLM扮演自由派或秘密法西斯派，进行1600局5人对战（LLM互玩、LLM对真人），测量胜率及新指标GSIR、RIA、DRR。","baseline":"对照25,000局真人线上游戏数据，以及随机（33%）和基于规则的算法（45%）基线。","findings":"前沿模型在合作与欺骗角色中均表现强劲，GPT-5.4等四款模型胜率显著高于基线，而较弱模型甚至低于随机基线。多数LLM难以全程维持一致欺骗人格，欺骗保持率降至50%以下。","reliability":"论文指出欺骗保持率普遍偏低，且社会推理与策略行动是不同能力，角色识别准确率高的模型不一定胜率高；实验仅限特定游戏，泛化性有限。","relevance":"高度相关：该研究用LLM代理模拟社会推理游戏，有真人数据对照，直接评估欺骗与策略行为，符合对LLM仿真可靠性及失效条件的批判性关注。","inspiration":"借鉴多智能体游戏仿真和细粒度指标（如欺骗保持率）来测量策略性信息操纵行为。｜可迁移到金融市场内幕交易或公司欺诈检测场景，模拟信息不对称下的决策。｜以LLM代理模拟交易员，处理为内幕信息获取，结果变量为交易行为与市场均衡，对照真实市场微观结构数据。"}},{"id":"2607.27824","version":1,"title":"STEREODISCO: Discovering Stereotypicality in LLMs","zh_title":"STEREODISCO：发现大语言模型中的刻板印象","abstract":"LLMs encode, convey, and perpetuate stereotypes. Prior computational research focuses on a small set of semantic axes investigated in social psychology, and operates on word embeddings produced by language models, leaving open which other semantic axes carry stereotypical associations in LLMs and how LLMs internally represent such axes. We introduce STEREODISCO, a framework that adapts the semantic differential method (Osgood et al., 1957) to the systematic study of stereotypes in LLM internal representations. STEREODISCO constructs approx. 2,000 candidate semantic axes from WordNet antonym synsets, recovers each as a geometric axis in the LLM's activation space via probing, and identifies stereotypical axes via a statistical test over concept projections. As a case study, we apply STEREODISCO to social group stereotypes with LLAMA-3-8B-INSTRUCT and MISTRAL-7B-INSTRUCT. We find that the two LLMs agree with each other on social group ratings more than with humans, suggesting that LLM-encoded stereotype content diverges from that documented in social psychology. We also discover stereotypical axes not investigated in prior work -- including humble vs. proud, narrow-minded vs. broad-minded, and cowardly vs. brave, which human annotators independently confirm.","authors":["Farane Jalali Farahani","Corina Dima","Mojtaba Nayyeri","Raphael H. Heiberger","Steffen Staab"],"categories":["cs.AI","cs.LG"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-07-31","first_seen":"2026-07-31","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.27824","pdf_url":"https://arxiv.org/pdf/2607.27824","source_feed":"cs.LG","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["刻板印象测量","LLM内部表征","社会心理学"],"reason":"测量LLM本身的社会刻板印象，非仿真人类被试，但有人类数据对照","model":"deepseek-v4-pro","scored_at":"2026-07-31T13:02:09","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-31","rank":23,"question":"LLM内部表征中哪些语义轴承载刻板印象，以及这些刻板印象与人类刻板印象是否一致？","design":"本研究并非仿真人类被试，而是提出STEREODISCO框架，从WordNet反义词对构建约2000个候选语义轴，通过探针在LLM激活空间中恢复为几何轴，再对概念投影进行统计检验以识别刻板印象轴，并应用于LLAMA-3-8B-INSTRUCT和MISTRAL-7B-INSTRUCT的社会群体刻板印象分析。","baseline":"使用已有心理学调查数据和人工标注的人类判断，比较LLM的社会群体评分和刻板印象轴识别结果。","findings":"两个LLM在社会群体评分上彼此一致性高于与人类的一致性，表明LLM编码的刻板印象内容与社会心理学记录存在差异。同时发现了先前工作未研究的刻板印象轴，如谦逊-骄傲、狭隘-开明、懦弱-勇敢，并得到人类标注者独立确认。","reliability":"论文未讨论","relevance":"该研究直接测量LLM内部表征中的刻板印象，并有人类数据作为对照基准，虽非仿真人类被试，但其方法可用于评估LLM作为人类替代品时的偏差，值得阅读以了解刻板印象的内部表征和检测方法。","inspiration":"借鉴其通过探针从LLM内部激活空间恢复语义轴并进行统计检验的方法，可用于检测经济决策中的隐性偏见。｜可迁移到信贷审批或招聘场景中的歧视检测，例如分析LLM对特定人群的财务能力或职业适合度刻板印象。｜以LLM作为被试，输入不同社会群体的描述，用STEREODISCO框架提取其内部表征中与能力、诚信等经济相关语义轴的投影，并与真实信贷或招聘数据中的群体差异进行对照。"}},{"id":"2605.06525","version":3,"title":"Who Is Really Playing? Strategic Interaction in AI-Guided Populations","zh_title":"谁在真正博弈？AI引导群体中的策略互动","abstract":"AI systems in general, and Large language models (LLMs), in particular, are increasingly used to provide instructions to many agents who interact with one another. Such shared reliance couples agents who appear to act independently: they may in fact be guided by a common model. This coupling can change the prospects for cooperation among agents with misaligned incentives. We study settings in which multiple \\emph{guidance providers} each advise a population of clients who participate in instances of an underlying game, creating strategic interaction at the level of the providers themselves. This induces a meta-game among the providers, mediated through clients. We first analyze the one-shot setting, where we show that shared instructions can change equilibrium behavior only when some provider influences more than one role in the same interaction. In such cases, cooperation may emerge, and the effect of client share can be beneficial, harmful, or non-monotone, depending on the base game. For the repeated setting, we prove a folk theorem for guidance providers: despite indirect observation and the clients' inability to identify which LLM advised their opponents, all feasible and individually rational outcomes can be sustained as $\\varepsilon$-equilibria.","authors":["Jonathan Shaki","Eden Hartman","Sarit Kraus","Yonatan Aumann"],"categories":["cs.GT","cs.MA","econ.TH"],"primary_category":"cs.GT","announce_type":"replace-cross","date":"2026-07-31","first_seen":"2026-05-07","revised_at":"2026-07-31","abs_url":"https://arxiv.org/abs/2605.06525","pdf_url":"https://arxiv.org/pdf/2605.06525","source_feed":"cs.MA","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["博弈论","多智能体系统","大语言模型"],"reason":"纯多智能体博弈理论分析，无人类行为对照，不涉及LLM仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-07-31T13:02:17","error":null,"has_summary":false,"summary":null},{"id":"2607.24797","version":2,"title":"Reading Without a Reader: Large Language Models Collapse Reading and Writing into a Single Entangled Code","zh_title":"无读者的阅读：大语言模型将阅读与写作坍缩为单一纠缠代码","abstract":"In the literate human brain, reading and writing doubly dissociate: a ventral decoding route (pure alexia) and a fronto-parietal encoding route (pure agraphia), sharing a partial orthographic core. A decoder-only large language model (LLM) drives both from one autoregressive path optimized on text (a \\emph{cultural} invention, not an evolved instinct). We ask how entangled it is, comparing an input-side ``reading code'' $\\mathbf{W}_{E}$ with an output-side ``writing code'' $\\mathbf{W}_{U}$ via an index $\\mathcal{E}\\in[0,1]$ (CKA, Procrustes residual, mutual $k$-NN) calibrated against an independent-init floor and tied ceiling. On GPT-2, OPT and Pythia (14M--1.4B), untied models hold one \\emph{coupled but sub-ceiling} code ($\\mathcal{E}=0.23$--$0.35$, far above floor) on a non-monotonic couple-then-differentiate trajectory, $\\mathbf{W}_{U}$ drifting $\\sim$3.2$\\times$ farther than $\\mathbf{W}_{E}$ in every decile. Equally informative is a negative: the matching behavioural test, that comprehension and production fail together rather than dissociate, cannot be run. For minimal pairs the alexia analogue is empty by theorem: greedy production implies a vocabulary-wide argmax, so it wins the pairwise ranking. Differential-damage indices are not scale-identified: heavy-tailed damage makes linear standardizations collapse onto their larger term, and the rank transform fixing this is bounded, so its null saturates. Both scores also contain the target's log-probability, which alone explains most of their variance and manufactures the apparent coupling. We withdraw a coupling statistic, a cross-level bridge and a separation measure. In a model reading and writing off one next-token distribution, no output-side pair isolates either ability: entanglement needing no index to see. By analogy, not homology, this situates LLMs in the space of possible minds.","authors":["Diego Salda\\~na Ulloa"],"categories":["q-bio.NC","cs.AI","cs.CL","cs.LG"],"primary_category":"q-bio.NC","announce_type":"replace-cross","date":"2026-07-31","first_seen":"2026-07-29","revised_at":"2026-07-31","abs_url":"https://arxiv.org/abs/2607.24797","pdf_url":"https://arxiv.org/pdf/2607.24797","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["LLM表征分析","神经语言学类比","模型可解释性"],"reason":"研究LLM内部表征与读写耦合，属纯NLP能力分析，无人类行为仿真或对照。","model":"deepseek-v4-pro","scored_at":"2026-07-31T13:02:03","error":null,"has_summary":false,"summary":null},{"id":"2607.27366","version":1,"title":"BridgeAlign: Bridging Preference Alignment for Humanities and Social Sciences","zh_title":"BridgeAlign：桥接人文社科领域的偏好对齐","abstract":"While data synthesis for large language models (LLMs) is prevalent, it primarily targets domains with verifiable answers, overlooking open-ended humanities and social sciences (HSS), where nuanced quality judgments matter more than objective correctness. This makes preference alignment a natural paradigm for broad HSS tasks. Yet existing methods are either costly or not tailored to broad HSS disciplines. We thus propose BridgeAlign, among the first preference-alignment pipelines for broad HSS disciplines, with three phases: i) Seed Curation: curating HSS seed documents from web corpora via heuristic/LLM-based filtering and text refinement; ii) Preference Data Synthesis: generating preference triplets via persona-based instruction inversion with Q&A consistency checks; iii) Preference Optimization: moving beyond naive human-vs-model heuristics by first grounding preferences in HSS quality rubric, then generating transitional responses via controlled quality degradation to form near-boundary preference pairs for finer-grained quality discrimination. Aligning over 210k synthetic preference samples, BridgeAlign enables Qwen3-8B to achieve the best average across 17 benchmarks against 11 strong baselines; importantly, leading on both human-preference and knowledge-based capabilities at once, with no trade-off between them, as supported by extensive experiments and contextualized by existing theories.","authors":["Ru Peng","Haokai Xu","Xijun Gu","Tianyu Zhao","Zhiting Fan","Yawen Zeng","Yihong Zhuang","Jinyang Zhang","Kexin Yang","Jian Wu","Hao Chen","Junyang Lin","Dayiheng Liu","Junbo Zhao"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-07-31","first_seen":"2026-07-31","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.27366","pdf_url":"https://arxiv.org/pdf/2607.27366","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["偏好对齐","数据合成","人文社科"],"reason":"纯偏好对齐数据合成与模型优化，无人类仿真或行为对照，属NLP能力评测。","model":"deepseek-v4-pro","scored_at":"2026-07-31T13:02:05","error":null,"has_summary":false,"summary":null},{"id":"2607.27379","version":1,"title":"HSS-Synth: Humanities and Social Sciences Data Synthesis for LLMs","zh_title":"HSS-Synth：面向大语言模型的人文社科数据合成","abstract":"High-quality, diverse data are vital for large language models (LLMs) but remain scarce and costly. Data synthesis is a viable alternative and succeeds on closed tasks, yet the humanities and social sciences (HSS) are overlooked, and their open-ended nature makes synthesis challenging. Moving beyond prior capability-centric, fragmented attempts, we adopt a subject-centric paradigm, define the first HSS domain system covering 14 mainstream fields, and introduce HSS-Synth, the first data synthesis pipeline for HSS. HSS-Synth comprises: (1) constructing seed documents from web corpora via multi-step filtering and text refinement evaluated by a judge; (2) specifying \"requirements + persona\" to backtranslate seed documents into diverse yet faithful instructions with a strict Q&A alignment check; and (3) breaking LLM response limits via teacher-forced Answering that feeds seed documents during response generation to anchor semantics, reduce hallucinations, and preserve tone and integrity. HSS-Synth yields 237k high-quality, diverse instruction-tuning samples that outperform 14 leading baselines on 16 benchmarks. The fine-tuned Qwen3-8B-Base sets a new SOTA and approaches the official Qwen3-8B, improving both human preference and knowledge capabilities without performance seesaws. Extensive experiments demonstrate HSS-Synth's robustness and transferability. Our code is publicly available at https://github.com/pengr/HSS-Synth.","authors":["Ru Peng","Tianyu Zhao","Xijun Gu","Zhiting Fan","Haokai Xu","Jinyang Zhang","Yawen Zeng","Yihong Zhuang","Kexin Yang","Junyang Lin","Dayiheng Liu","Junbo Zhao"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-07-31","first_seen":"2026-07-31","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.27379","pdf_url":"https://arxiv.org/pdf/2607.27379","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["数据合成","指令微调","NLP评测"],"reason":"纯数据合成与NLP评测，无人类仿真或行为对照","model":"deepseek-v4-pro","scored_at":"2026-07-31T13:02:06","error":null,"has_summary":false,"summary":null},{"id":"2607.27384","version":1,"title":"Same Facts, Different Diagnosis: Measuring and Mitigating Narrative Anchoring in Clinical Language Models","zh_title":"相同事实，不同诊断：测量与缓解临床语言模型中的叙事锚定","abstract":"Large language models used for clinical diagnostic reasoning are sensitive to sociolinguistic register, not just clinical content. We term this failure mode Narrative Anchoring: identical clinical facts expressed in different registers cause diagnostic outputs to diverge. Unlike prior demographic-bias work, which manipulates explicit identity tokens such as race or income, our benchmark isolates register as the sole channel of variation, with no demographic marker present in any form. We construct a dataset of 1,000 USMLE clinical vignettes, each rewritten into three sociolinguistically distinct personas under an independently audited fact-preservation guarantee, verified by a separate model that never sees the generation prompt. Across seven language models spanning three architecture families and scales, Narrative Anchoring is statistically significant under direct prompting in every model tested, with a Narrative Anchoring Gap of 0.064 to 0.151. Chain-of-thought reasoning and explicit debiasing instructions reduce the bias only partially, and their apparent gains are frequently confounded by accuracy collapse. We introduce NarrativeShield, a three-agent pipeline that structurally extracts and verifies clinical facts before diagnostic reasoning begins, reducing the Narrative Anchoring Gap to near-zero ($-0.004$ to $0.037$) and achieving the lowest rate of severely unstable decisions (DSS $<$ 0.8) of any method across all models, at a modest and mechanistically expected accuracy cost for most models. A stress test using a non-instruction-tuned base model shows that executing a debiasing intervention at all is gated by zero-shot instruction-following ability, not prompt content alone. We release our dataset, human-validated for fact preservation, as a standalone resource for studying register-based clinical bias.","authors":["Prabhjot Singh","Pritam Deka","Vijay Chennareddy"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-07-31","first_seen":"2026-07-31","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.27384","pdf_url":"https://arxiv.org/pdf/2607.27384","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["临床NLP","偏差检测","诊断推理"],"reason":"研究临床LLM的诊断偏差，属NLP能力评测，不以人类行为仿真为目标","model":"deepseek-v4-pro","scored_at":"2026-07-31T13:02:07","error":null,"has_summary":false,"summary":null},{"id":"2607.28190","version":1,"title":"The MADRS Pipeline: Supporting Depression Assessment in Clinical Trials","zh_title":"MADRS流水线：支持临床试验中的抑郁评估","abstract":"Depression is a major mental disorder for which diagnosis relies primarily on clinical assessments. Automated methods to support its detection via the psychiatric MADRS scale are getting more and more attention. While existing solutions primarily focus on detecting the disorder from different text sources (e.g., online text, social media), there is still limited support for clinical trials, where clinical assessments are conducted through structured interviews based on standard guidelines such as SIGMA. In this work, we develop a LLM pipeline specifically designed to support clinicians in supporting the assessment of depression in patients enrolled in clinical trials. Our pipeline converts audio interviews into transcripts, maps them into the ten MADRS symptom items, estimates their severity, and identify problematic clinical ratings associated with them. Evaluation on real clinical interviews shows a strong overall correlation of 0.867 with expert ratings, providing interpretable support for future assessments in clinical trials.","authors":["Mila Fodor","Katalin \\'Ocsai","Francesco Periti","Rien Sonck","Alex Boudreau"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-07-31","first_seen":"2026-07-31","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.28190","pdf_url":"https://arxiv.org/pdf/2607.28190","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["临床NLP","抑郁评估","LLM辅助诊断"],"reason":"纯NLP评测，用LLM辅助临床抑郁评估，非仿真人类被试","model":"deepseek-v4-pro","scored_at":"2026-07-31T13:02:11","error":null,"has_summary":false,"summary":null},{"id":"2607.28478","version":1,"title":"Would You Walk to the Car Wash? Revealing the Salience Bias of Large Language Models in Commonsense Reasoning","zh_title":"你会走到洗车场吗？揭示大语言模型在常识推理中的显著性偏差","abstract":"As large language models (LLMs) continue to advance in complex reasoning tasks, they have learned to heavily prioritize explicit conditions provided in the input. However, in everyday commonsense reasoning, this mechanism exposes a critical vulnerability which we term Salience Bias: models become easily hijacked by useless explicit distractors (e.g., numerical values), leading them to ignore the implicit physical or commonsense prerequisites of a task. A critical open question is whether this failure reflects a genuine gap in commonsense knowledge or merely its suppression under misleading task framing. To investigate this, we construct the SaliTrap Benchmark, a high-quality dataset across four trap dimensions. Evaluating 12 state-of-the-art LLMs, we find that all mainstream models suffer significantly from salience bias, with severity scaling with distractor density and detecting the trap often decoupled from actually avoiding it. Crucially, by re-eliciting the same models with the task framing stripped away, we show that this is overwhelmingly a failure of \\textbf{knowledge suppression rather than knowledge absence}: a context-free knowledge probe alone recovers over 90\\% of sycophantic-compliance failures, revealing that the requisite commonsense is intrinsically present but actively crowded out by salient distractors that lure the model into over-compliant, unnecessary computation. Building on this diagnosis, we further show that lightweight, inference-time prompting alone substantially closes the gap without any retraining. Our findings relocate the bottleneck of commonsense reasoning failures from model competence to elicitation, and we release SaliTrap as a testbed for this blind spot. The codes are available at https://github.com/Wuzheng02/SaliTrap.","authors":["Zheng Wu","Chenhao Xue","Shijie Zheng","Yijie Lu","Cheng Yang","Zhuosheng Zhang"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-07-31","first_seen":"2026-07-31","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.28478","pdf_url":"https://arxiv.org/pdf/2607.28478","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["常识推理","模型偏差","NLP评测"],"reason":"纯NLP能力评测，研究LLM常识推理中的显著性偏差，不以人类行为仿真为目标。","model":"deepseek-v4-pro","scored_at":"2026-07-31T13:02:13","error":null,"has_summary":false,"summary":null},{"id":"2607.28505","version":1,"title":"Generative AI and linguistic diversity in academic writing and publishing: Perspectives from World Englishes","zh_title":"生成式AI与学术写作出版中的语言多样性：世界英语视角","abstract":"The rise of generative artificial intelligence (GenAI) in academic writing and publishing (AWP) raises questions about linguistic inclusivity and the legitimacy of diverse Englishes in global scholarly communication. This article responds to these questions through a structured scholarly dialogue involving five sociolinguists from World Englishes and adjacent fields. Organised around five guiding questions, the dialogue interrogates how GenAI tools influence writing practices, reinforce or disrupt dominant language norms, and raise ethical challenges. Contributors reflect on the potential of GenAI to democratise writing processes while also raising concerns about GenAI's tendency to marginalise minoritised varieties and flatten nuance in scholarly writing. Across the dialogue, themes of linguistic (in)justice, researcher agency, and institutional responsibility emerge, with contributors calling for equity-informed policies, critical AI literacy, and inclusive co-design in GenAI development. The article shows the value of dialogic reflection in understanding GenAI's role in AWP. It concludes that while GenAI may reinforce existing hierarchies, it can also serve as a site of resistance, depending on how it is designed, governed and used within scholarly communities committed to linguistic diversity.","authors":["Kingsley Ugwuanyi","Christian Mair","Sender Dovchin","Iker Erdocia","Maria Kuteeva","Esther Airemionkhale"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-07-31","first_seen":"2026-07-31","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.28505","pdf_url":"https://arxiv.org/pdf/2607.28505","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["学术写作","语言多样性","生成式AI"],"reason":"论文讨论GenAI对学术写作语言多样性的影响，不涉及LLM仿真人类被试或行为实…","model":"deepseek-v4-pro","scored_at":"2026-07-31T13:02:13","error":null,"has_summary":false,"summary":null},{"id":"2607.28528","version":1,"title":"AI systems and the reproduction of (standard) language ideologies in World Englishes","zh_title":"AI系统与世界英语中（标准）语言意识形态的再生产","abstract":"The rapid growth of large language models (LLMs) has resurrected age-old questions in sociolinguistics and world Englishes, such as who decides what counts as legitimate English, whose English is suspect etc. This paper examines how AI systems, their uses and discourse on them reflect, reinforce, and occasionally challenge (standard) language ideologies, which privilege Inner Circle norms and marginalize non-dominant Englishes. Drawing on evidence from empirical studies, media commentary, social media debates, and examples from AI outputs, the paper shows that AI technologies reproduce dominant language ideologies at different levels: training data, design protocols, evaluation benchmarks, user feedback and public commentary. The analysis uses the public controversy over AI-sounding language, especially the fixation on the word delve, to illustrate how speakers of English from the Global North police the English language norms of Global South English users. The paper also identifies what Christian Mair has called a \"standardisation paradox\": AI may homogenize English by privileging standard forms and at the same time pluralize Englishes through exposure to wide-ranging corpora and annotation work carried out by Global South users. In doing so, the paper argues that generative AI is reigniting long-standing debates in World Englishes about standardization, legitimacy, and the ownership of English, now playing out in algorithmic systems, model training, evaluation practices, and public discourse, where non-dominant Englishes are increasingly conflated with AI-generated speech. Discussing AI systems as a site where language ideologies are (re)produced, the paper argues for more inclusive design approaches that recognize the plurality of Englishes in order to address the real-world negative consequences of treating some as more legitimate than others.","authors":["Kingsley Ugwuanyi"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-07-31","first_seen":"2026-07-31","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.28528","pdf_url":"https://arxiv.org/pdf/2607.28528","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["语言意识形态","世界英语","AI偏见"],"reason":"论文讨论AI系统如何反映语言意识形态，属于社会语言学分析，不涉及用LLM仿真人…","model":"deepseek-v4-pro","scored_at":"2026-07-31T13:02:13","error":null,"has_summary":false,"summary":null},{"id":"2607.28576","version":1,"title":"Sample More, Reflect Less: Self-Refine and Reflexion Lose to Repeated Sampling at Equal Token Cost, from 1.5B to 7B","zh_title":"多采样优于自反思：在等量Token成本下，自优化和反思方法不敌重复采样","abstract":"Methods that make a language model plan, criticise and rewrite its own answer, reflect on mistakes, pick the best of several attempts, or debate with copies of itself nearly all make it generate far more text than a single chain of thought. Because generating more text raises accuracy by itself, a gain over one chain of thought does not show the method's idea is what helped. Wang et al. (2024) reported that a simple baseline, sampling the same question repeatedly and keeping the most common answer, often wins once budgets are comparable, but gave point estimates with no confidence intervals or significance tests. We rerun that comparison as a designed experiment: seven methods, open models of 1.5B, 3B and 7B parameters, two mathematics benchmarks, 150 questions each. We count every generated token, including those spent on critiques, reflections, debate turns and checking, and compare each method against repeated sampling at its own measured cost. All 36 comparisons are paired by question, with bootstrap intervals and multiplicity correction. No method is reliably better than repeated sampling at equal cost anywhere. Ten are reliably worse, all of them methods where the model inspects its own output, and all 18 self-inspection comparisons are negative. The two kinds of self-inspection part company as models grow. Choosing stops hurting: taking Best-of-N's eight samples and just counting the most common answer beats letting the model pick by 8.0 and 11.3 points at 1.5B, but only 2.0 and 1.3 at 7B, no longer distinguishable from zero. Rewriting does not recover: Self-Refine and a forced Reflexion stay 3.6 to 10.1 points below baseline at 7B. Reflexion as published never triggered its own retry on the smallest model. It judged itself correct every time and silently became a single chain of thought. We release code, prompts, all generations, and our verification scripts.","authors":["Iliya Mirzaei"],"categories":["cs.CL","cs.AI","cs.LG"],"primary_category":"cs.CL","announce_type":"new","date":"2026-07-31","first_seen":"2026-07-31","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.28576","pdf_url":"https://arxiv.org/pdf/2607.28576","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["推理策略","基准评测","成本效率"],"reason":"纯NLP能力评测，比较不同推理策略的数学解题准确率，不涉及人类行为仿真或对照。","model":"deepseek-v4-pro","scored_at":"2026-07-31T13:02:15","error":null,"has_summary":false,"summary":null},{"id":"2607.28410","version":1,"title":"Can Large Language Models Execute Parent Orders?","zh_title":"大语言模型能否执行母单？","abstract":"Parent-order execution is a core problem in algorithmic trading, where the goal is to split a large order into smaller orders while reducing execution costs. Existing approaches either rely on pre-specified market assumptions that may not hold in practice, or require task-specific training that limits adaptability to new settings. To overcome these limitations, we present the first systematic study of large language models (LLMs) for parent-order execution. This extends the use of LLMs in finance from what to trade to how to execute. We propose PACE (Plan-Ahead Controlled Execution), a hierarchical framework that decomposes parent-order execution into long-horizon planning and short-horizon execution, requiring neither explicit market assumptions nor task-specific training. Experiments on Shenzhen Stock Exchange Level-1 data show that PACE outperforms TWAP, Almgren-Chriss, and learning-based baselines, exceeding the strongest baseline by 0.65 bps. Behavioral analysis reveals that LLMs make execution decisions differently from human investors: higher model confidence predicts better performance rather than worse returns, and the model trades earlier rather than procrastinating toward the deadline. These findings suggest that LLMs can complement human traders in execution decisions.","authors":["Zane Shen","Xinli Xu","Guangyi Zhang","Jialong Chen","Jinsong Zhou","Cong Chen","Guibao Shen","Dongyu Yan","Luozhou Wang","Zhen Yang"],"categories":["cs.CE","cs.CL","q-fin.TR"],"primary_category":"cs.CE","announce_type":"cross","date":"2026-07-31","first_seen":"2026-07-31","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.28410","pdf_url":"https://arxiv.org/pdf/2607.28410","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["算法交易","LLM代理","订单执行"],"reason":"纯多智能体交易执行，无人类行为对照，属C1排除项","model":"deepseek-v4-pro","scored_at":"2026-07-31T13:02:13","error":null,"has_summary":false,"summary":null},{"id":"2607.27697","version":1,"title":"DP-LENS: A Density-Aware Polyfocal Lens with Topology-Driven Auto-Routing for Occlusion Management in Immersive 3D Analytics","zh_title":"DP-LENS：面向沉浸式3D分析中遮挡管理的密度感知多焦点透镜与拓扑驱动自动路由","abstract":"Immersive environments, e.g., virtual reality (VR), offer a unique approach to exploring complex 3D datasets, where data is often heavily occluded and exploration incurs a high cognitive load. We propose DP-LENS, a density-aware polyfocal fisheye lens equipped with topology-driven auto-routing. While preserving peripheral context through geometric deformation and 3D perspective techniques, it enables users to explore 3D data with a lower cognitive load. To facilitate hands-free macro-navigation, we integrate a Large Language Model (LLM) to serve as a supplementary voice-based target selection tool that initiates the auto-routing algorithm. Two user studies with 34 participants investigate the potential benefits of this system. Our first study (N=18) compared the manual DP-LENS against two industry-standard baselines (i.e., World-in-Miniature and volumetric slicing) in heavily occluded 3D datasets. The results show that DP-LENS significantly reduced cognitive load, decreased completion time, and improved user preference. The second study (N=16) compared the topology-driven auto-routing system (initiated via voice commands) with a fully manual DP-LENS. The results show that the auto-routing system improved task efficiency, further reduced cognitive load, and garnered higher user preference. Furthermore, the auto-routing partially decoupled exploration efficiency from the physical dimensions of the data and mitigated physical fatigue to some extent. Based on the findings, we proposed design implications to inform the development of more spatially scalable and low-fatigue interactions for future 3D visual analytics systems.","authors":["Nieyu Cao","Xian Wang","Lik-Hang Lee"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-07-31","first_seen":"2026-07-31","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.27697","pdf_url":"https://arxiv.org/pdf/2607.27697","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C2"],"tags":["沉浸式分析","遮挡管理","人机交互"],"reason":"研究沉浸式3D可视化中的遮挡管理，LLM仅作语音辅助，不涉及人类行为仿真或对照。","model":"deepseek-v4-pro","scored_at":"2026-07-31T13:02:09","error":null,"has_summary":false,"summary":null},{"id":"2607.28239","version":1,"title":"Identifying a Level-up Pathway for AI-assisted Counterspeech through Elaboration","zh_title":"通过精细化识别AI辅助反驳言论的升级路径","abstract":"Given the profound societal impact of vaccine-skeptical content on social media, community-driven counterspeech has emerged as a promising participatory response to contest and curb such objectionable content. Yet crafting effective counterspeech remains challenging for ordinary users, limiting their willingness and ability to engage constructively. We designed and evaluated three generative AI-assisted counterspeech writing systems that vary by assistance stage (co-writing vs. re-writing) and mode (guided vs. unguided) to support lay users' responses to vaccine-skeptical content. We ask whether AI can help users craft counterspeech perceived as both effective and authentic, which forms of AI support work best, and through what mechanisms. In a randomized controlled trial with social media users, participants wrote counterspeech responses to both statistical and narrative vaccine-skeptical content. Across evidence types, AI-assisted writing increased perceived counterspeech effectiveness while largely preserving authentic self-expression, and perceived effectiveness was the strongest predictor of willingness to counterspeak publicly. AI's primary benefit was facilitating more elaborate writing, producing messages that were more informative, analytical, and lexically sophisticated. These findings suggest a level-up pathway for AI-assisted, community-driven counterspeech, which helps cultivate more effective and motivated counterspeakers, contributing to higher-quality public discourse on pressing societal issues.","authors":["Han Li","Inhwan Bae","Natalie Bazarova","Drew Margolin"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-07-31","first_seen":"2026-07-31","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.28239","pdf_url":"https://arxiv.org/pdf/2607.28239","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["AI辅助写作","反驳言论","社交媒体"],"reason":"AI辅助写作工具评估，非用LLM仿真人类被试","model":"deepseek-v4-pro","scored_at":"2026-07-31T13:02:11","error":null,"has_summary":false,"summary":null},{"id":"2607.28601","version":1,"title":"Using Theory of Mind to Arbitrate between Social and Non-social Learning","zh_title":"利用心智理论在社会学习与非社会学习之间进行仲裁","abstract":"Social learning is a powerful mechanism through which agents learn about the world from others. However, humans sometimes choose direct experience over social learning, which can carry time and cognitive resource costs. How do people balance social and non-social learning? We propose a Rational Mentalizing model of the decision to engage in social learning. This model estimates the utility of social learning by reasoning about another agent's goal and the informativeness of their future actions. It then weighs the utility of social learning against the utility of non-social learning. Using a novel game where players choose between observing other agents or exploring the environment, we show that the Rational Mentalizing model can quantitatively capture human trade-offs between these strategies. These findings suggest that selective social learning is guided by 'Theory of Mind' in the service of utility maximization.","authors":["Lance Ying","Ryan Truong","Joshua B. Tenenbaum","Samuel J. Gershman"],"categories":["cs.MA","q-bio.NC"],"primary_category":"cs.MA","announce_type":"new","date":"2026-07-31","first_seen":"2026-07-31","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.28601","pdf_url":"https://arxiv.org/pdf/2607.28601","source_feed":"cs.MA","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["社会学习","心智理论","多智能体系统"],"reason":"多智能体社会学习模型，无LLM，无人类仿真，属纯多智能体系统研究。","model":"deepseek-v4-pro","scored_at":"2026-07-31T13:02:15","error":null,"has_summary":false,"summary":null},{"id":"2607.27536","version":1,"title":"Strategy, Not Payoffs: A Behavioural Embedding of Normal-Form Games","zh_title":"策略而非收益：正则形式博弈的行为嵌入","abstract":"Learning a strategic task changes more than what is directly taught: fine-tuning on one game can either enhance or degrade an agent's ability to reason in another. Understanding and predicting this transfer of strategic capabilities, however, remains a key challenge for large language models (LLMs). Normal-form games provide an ideal testbed for analyzing this phenomenon, as they feature explicitly defined payoffs and well-characterized equilibrium behaviours. In this work, we investigate whether game embeddings can explain and predict changes in LLM strategic capabilities following fine-tuning across different games. We propose a lightweight two-feature embedding that captures fundamental behavioural demands: the entropy of the Nash equilibrium and the sensitivity of optimal responses to an opponent's action. We show that while existing published structural embeddings primarily memorize game identities and fail to generalize, our behavioural embedding reliably predicts performance changes on held-out games. These results demonstrate that the transfer of strategic capabilities in LLMs is not dictated by the payoff geometry of a game, but by the underlying structure of the decision-making behaviour it requires.","authors":["Joshua Caiata","Sreepriya Pulyassary","Xiang Li","Kate Larson"],"categories":["cs.GT","cs.AI","cs.LG","cs.MA"],"primary_category":"cs.GT","announce_type":"cross","date":"2026-07-31","first_seen":"2026-07-31","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.27536","pdf_url":"https://arxiv.org/pdf/2607.27536","source_feed":"cs.MA","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["博弈论","多智能体系统","策略迁移"],"reason":"研究多智能体博弈中的策略迁移，不涉及人类行为对照或仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-07-31T13:02:07","error":null,"has_summary":false,"summary":null},{"id":"2607.27548","version":1,"title":"Explaining the Macroeconomic Inertia Puzzle","zh_title":"解释宏观经济惯性之谜","abstract":"Benchmark macroeconomic models require additional frictions to explain the sluggish response of aggregate variables to sudden shocks or changes in policy. I show that standard heterogeneous agent (HA) models, the Blanchard (1985) perpetual youth and Bewley (1986) incomplete markets models, are consistent with aggregate consumption inertia without the use of habit preferences or any specific model of expectation underreaction to dampen the responsiveness of consumption savings decisions. I instead replicate observed consumption inertia in standard HA models by directly substituting survey expectations of income and interest rates for agents' expectations. I propose a new theory of macroeconomic inertia that rationalizes the observed extrapolation bias in survey expectations by embedding an unobserved components model of expectations into a tractable HA general equilibrium environment. Inertia results when expectations imperfectly account for the equilibrium amplification of shocks, which is large in HA economies. This imperfect inference causes expectations to gradually unanchor as agents repeatedly misattribute large responses of equilibrium outcomes simply to larger shocks. This theory also illustrates a novel drawback to inertial monetary policy rules and the delayed financing of fiscal deficits: Policy regimes that act more gradually experience longer transmission lags due to their decreased effectiveness at anchoring expectations.","authors":["Michael Cai"],"categories":["econ.GN","q-fin.EC"],"primary_category":"econ.GN","announce_type":"new","date":"2026-07-31","first_seen":"2026-07-31","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.27548","pdf_url":"https://arxiv.org/pdf/2607.27548","source_feed":"econ.GN","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["宏观经济","异质性主体模型","预期形成"],"reason":"论文研究宏观经济模型中的惯性，使用调查预期替代理性预期，不涉及LLM仿真人类被…","model":"deepseek-v4-pro","scored_at":"2026-07-31T13:02:07","error":null,"has_summary":false,"summary":null},{"id":"2607.26348","version":1,"title":"When Synthetic Users Fail: A Cross-Domain Benchmark of LLM-Simulated Human Survey Responses","zh_title":"当合成用户失败：LLM模拟人类调查回答的跨领域基准测试","abstract":"Large language models (LLMs) are increasingly used as synthetic users, stand-ins for human respondents whose simulated answers feed product, policy, and market decisions. We ask when this substitution is valid and when it fails, and package the answer as an evaluation framework for intelligent synthetic-user systems. A single protocol, run across four models spanning two families and an 8B-to-frontier capability range, is applied to two independent domains of real human-response data: U.S. general social attitudes (General Social Survey) and cross-cultural values (World Values Survey). Every model is benchmarked against a suite of non-LLM baselines fit on held-out human data. Under demographic prompting and the survey-simulation protocols we test, two failures replicate across both domains, all four models, and both families. First, at the individual level no LLM beats even the strongest baseline; on cross-cultural values every model falls well below it, and the gap survives distance-aware and proper scoring. Second, models systematically over-determine demographics, treating identity as far more predictive of attitudes than it is among real people, a distortion present for nearly every question-group combination and robust to a coding-invariant measure. Neither failure is remedied by a larger, more capable model. A decision-impact analysis shows why this matters in practice: on a segment-targeting task the models inflate between-segment gaps two to fourfold, would direct a team to the wrong segment in half of U.S. and most cross-cultural cases, and manufacture segment splits that do not exist in real people. We make the cross-domain benchmark and the evaluation framework available on request, so that teams can determine in advance when synthetic-user evidence is safe for decision support and when it is not.","authors":["Zihan Chen","Di Zhu","Lei Nico Zheng"],"categories":["cs.CL","cs.AI","cs.CY","cs.HC"],"primary_category":"cs.CL","announce_type":"new","date":"2026-07-30","first_seen":"2026-07-30","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.26348","pdf_url":"https://arxiv.org/pdf/2607.26348","source_feed":"cs.CL","score":10,"bucket":"selected","rubric_hits":["A1","A2","A5","B1","B2","B4"],"tags":["LLM仿真","人类调查","失效分析"],"reason":"直接评估LLM仿真人类调查的失效条件，有真实人类数据对照，涉及社会态度和政策场…","model":"deepseek-v4-pro","scored_at":"2026-07-30T13:01:41","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-31","rank":1,"question":"在人口统计提示和调查模拟协议下，LLM作为合成用户何时会失效？","design":"使用四个模型（涵盖两个家族、8B到前沿能力），在两种提示格式下，对美国综合社会调查（GSS）和世界价值观调查（WVS）的真实人类回答进行仿真，测量个体预测准确度、总体分布复现和人口统计结构忠实度。","baseline":"基于真实人类数据拟合的朴素人口统计基线（包括人口查找表、逻辑回归、随机森林），在留出的人类数据上评估。","findings":"所有模型在个体层面均未超越最强基线，且系统性地过度决定人口统计特征，将身份视为比真实人类中更具预测性的因素；在细分目标定位任务中，模型夸大了细分群体间的差距，并制造了真实人类中不存在的细分分裂。","reliability":"论文指出失效在人口统计提示和所测试的调查模拟协议下跨领域、模型和家族复现，且更大或更强的模型未能弥补这些失效；但未讨论其他提示策略或协议下的潜在有效性。","relevance":"该研究直接评估了LLM仿真人类调查的失效条件，提供了跨领域基准和真实人类数据对照，对关注仿真可靠性与偏差的研究者极具参考价值，值得精读原文。","inspiration":"借鉴其统一协议、多模型跨领域测试和朴素人口统计基线设计，可迁移到经济预期形成或政策偏好调查仿真中，例如用LLM模拟不同人口群体对通胀预期的回答，以真实消费者预期调查数据为基准，检验仿真是否高估人口统计的预测力并扭曲预期分布。"}},{"id":"2607.26899","version":1,"title":"Human diversity fuels collective creativity that large language models cannot simulate or sustain","zh_title":"人类多样性推动集体创造力，而大语言模型无法模拟或维持","abstract":"Diverse human groups produce diverse ideas, the raw material of innovation. Generative AI challenges this engine twice over: everyday AI assistance may homogenize what diverse people create, and AI-simulated diversity may replace the people altogether. We tested both challenges in a preregistered creative metaphor experiment with native (L1) and non-native (L2) English writers, who wrote without AI, with AI-generated ideas (AI ideation), or with AI refining their own ideas (AI refinement). L2 writers contributed more collective diversity than L1 writers, with native-language ideation showing the most diverse pools. AI ideation compressed collective diversity for everyone and left the L2 advantage undetectable, whereas AI refinement preserved both. We then simulated the entire writer pool using personas built from participants' real backgrounds, three model families, native-language prompting, and elevated sampling temperatures. Every simulated pool fell below every human pool, and pushing models further induced diversity only through degenerate text. However, at the individual level, AI ideation raised writers' ratings, pitting private incentives against the collective good, except when L2 writers used their native language, which benefited both. Human diversity remains a valuable creative resource that current AI cannot simulate or sustain; the design of human-AI collaborative workflows determines whether it survives.","authors":["Mengchen Dong","Hiromu Yakura"],"categories":["cs.HC","cs.AI","cs.CY"],"primary_category":"cs.HC","announce_type":"cross","date":"2026-07-30","first_seen":"2026-07-30","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.26899","pdf_url":"https://arxiv.org/pdf/2607.26899","source_feed":"cs.AI","score":10,"bucket":"selected","rubric_hits":["A1","A3","A5","B1","B2","B4"],"tags":["LLM人类仿真","创意实验","多样性对照"],"reason":"用LLM模拟人类创意实验，与真实人类数据对照，评估仿真失效条件，高度相关。","model":"deepseek-v4-pro","scored_at":"2026-07-30T13:01:46","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-31","rank":2,"question":"在创意生产中，人类多样性（以母语/非母语英语写作者为代理）在AI辅助下是否仍能带来集体多样性优势？LLM能否通过模拟多样性来替代真实人类多样性？","design":"本研究并非纯粹的仿真研究，而是先进行人类实验，再用LLM仿真进行对比。人类实验：招募母语（L1）和非母语（L2）英语写作者，随机分配到无AI、AI生成创意（AI ideation）或AI润色自有创意（AI refinement）三种条件，生成英语隐喻，测量集体多样性（输出池的变异度）和个体评分。仿真部分：基于参与者真实背景构建persona，使用三种模型家族、母语提示和升高采样温度，模拟整个写作者池，测量集体多样性，并与人类池对比。","baseline":"真实人类数据：人类实验中L1和L2写作者在三种AI协作条件下的隐喻输出池的集体多样性，以及无AI条件下的基线。","findings":"AI生成创意条件压缩了所有人的集体多样性，且使L2写作者的多样性优势消失，而AI润色条件保留了多样性和L2优势。所有LLM模拟池的集体多样性均低于任何人类池，且提高模型温度仅通过生成退化文本来增加多样性。","reliability":"论文指出LLM模拟无法复现人类集体多样性，即使使用真实背景构建persona、母语提示和升高温度，模拟池的多样性仍低于人类池，且过度推动模型会导致文本退化。","relevance":"该研究直接对比真实人类与LLM仿真在集体创意多样性上的表现，揭示了LLM仿真在捕捉群体层面变异时的失效，并探讨了AI协作设计如何影响多样性存留，高度契合研究者对仿真可靠性、失效条件及人类基准对照的关注，值得精读。","inspiration":"借鉴其将人类实验与LLM仿真直接对比的双阶段设计，以及用集体多样性（而非个体准确度）作为核心结果变量的测量思路。｜可迁移至政策沟通中的创意生成场景，如不同语言背景的公众对政策隐喻的解读与再创作，评估AI辅助是否削弱观点多样性。｜招募母语和非母语政策受众作为被试，随机分配至无AI、AI生成政策解释隐喻、AI润色自有隐喻三种条件，测量集体隐喻多样性，并以真实公众咨询数据作为对照基准，同时用基于被试背景的LLM persona模拟整个群体，检验仿真多样性是否匹配人类基准。"}},{"id":"2607.27100","version":1,"title":"Can Large Language Models Represent Urban Publics? Behavioral Replication and Population Mismatch in an Affordable-Housing Experiment","zh_title":"大语言模型能代表城市公众吗？一项可负担住房实验中的行为复现与人口错配","abstract":"There is growing interest in using large language models (LLMs) as low-cost proxies for resident attitudes in urban planning. Previous work shows that LLMs can predict average results of survey experiments, but less is known about whether they preserve the spatially anchored, identity-conditioned structure behind those averages, namely how support changes as a project approaches homes and how that response divides across tenure and partisan groups. We compared eight open-weight LLMs with 843 respondents in a US affordable-housing survey experiment, testing whether they reproduced the owner-renter difference in support change as a proposed development moved from 2 miles to 1/8 mile. Qwen 2.5 14B was closest (-0.242 versus the human -0.285) and was the only model to meet the prespecified +/-0.20 equivalence criterion; Phi-4 14B was directionally aligned but attenuated (-0.150), and other models showed weak, null, or reversed moderation. This aggregate match masked structural failure. Qwen attenuated the Republican contrast and exaggerated the Independent one, its RMSE across 27 party-by-tenure-by-item cells was 0.613, its median model-to-human variance ratio was 0.099, and question order shifted the contrast by +0.367. Identity-cue removal and selective nonresponse changed which comparisons were estimable, and rationale-first responses differed from matched direct-choice responses in 20.6-35.3% of focal comparisons. An LLM can thus approximate one aggregate contrast while failing to preserve the population structure, within-group heterogeneity, and measurement stability that generate it. Model evaluation in urban planning should test whether this spatial and social structure survives simulation, not only average effects.","authors":["Yuxuan Cai","Yequan Hu","Hongqian Li","Zhanghong Ju","Shuying Guo"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2026-07-30","first_seen":"2026-07-30","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.27100","pdf_url":"https://arxiv.org/pdf/2607.27100","source_feed":"cs.CY","score":10,"bucket":"selected","rubric_hits":["A1","A2","A3","B1","B2","B4"],"tags":["LLM人类仿真","行为复现","政策评估"],"reason":"用LLM复现住房实验中的行为差异，与843名人类被试对照，评估仿真失效的结构性…","model":"deepseek-v4-pro","scored_at":"2026-07-30T13:01:46","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-30","rank":3,"question":"大语言模型能否在可负担住房实验中复现人类的空间邻近效应和群体结构差异？","design":"使用8个开源大语言模型模拟美国居民，施加住房项目距离（2英里 vs. 1/8英里）的处理，测量支持度变化，并考察房主-租户、党派身份的调节效应。","baseline":"843名美国受访者在同一可负担住房调查实验中的真实回答。","findings":"Qwen 2.5 14B在总体房主-租户差异上最接近人类，但掩盖了群体结构失效：共和党人对比减弱、独立人士对比夸大，组内方差远低于人类，且问题顺序和身份提示移除会改变结果。","reliability":"模型在总体效应上可能匹配，但无法保留生成该效应的群体分布、组内异质性和测量稳定性；身份提示、问题顺序和回答格式变化会导致估计结果不一致。","relevance":"该研究直接检验LLM仿真在空间-社会结构上的失效，提供了从总体匹配到群体结构分解的严格验证框架，对关注经济学实验和政策评估中仿真可靠性的研究者极具参考价值。","inspiration":"借鉴其将总体处理效应分解为子群体条件对比的验证方法，并引入问题顺序、身份提示等测量稳定性检验。｜可迁移到政策评估中的邻避效应实验或地方公共品供给偏好研究，如垃圾处理厂、风电场选址的公众接受度。｜以LLM模拟不同收入、党派、住房产权的居民，施加设施距离和补偿方案处理，测量支持度变化，用真实居民调查数据对照子群体效应和顺序效应。"}},{"id":"2607.02464","version":2,"title":"Will Scaling Improve Social Simulation with LLMs?","zh_title":"扩大规模会改善基于大语言模型的社会仿真吗？","abstract":"Large Language Model (LLM) social simulations are a promising research method, but they are not yet faithful enough to be adopted widely. In this work, we investigate whether the current scaling paradigm in language modeling is likely to close these gaps, or whether simulation fidelity is orthogonal to general capabilities and therefore deserving of more research attention. We use scaling laws to study the relationship between LLMs' compute scale, general capability benchmarks, and the fidelity of social simulation in three representative sub-domains: opinion modeling, behavioral simulation, and longitudinal forecasting. Surprisingly, we discover strong compute scaling in all three settings, using a suite of 85 transformer LLMs with the Qwen3 architecture pre-trained on the DCLM web text corpus under fixed-compute budgets from $10^{18}$ to $10^{20}$ FLOPs. Then we evaluate 35 larger and more capable open-weight models up to 70B parameters, allowing us to predict downstream accuracy from loss. This reveals that the majority of behavioral and opinion simulation tasks will rapidly improve with scale, particularly when they involve populations that are well-represented in English web corpora. Longitudinal forecasting and underrepresented opinions scale more slowly, especially when they are less correlated with general knowledge and reasoning benchmarks like MMLU. In behavior simulation, scaling fails to improve model calibration with human cognitive biases like risk aversion, as well as human heuristics like learning correlated rewards from related tasks. On these tasks, even fine-tuned models fail to noticeably scale up performance from 0.5B to 8B parameters. Taken together, we conclude that scale will improve social simulations in most settings, but outliers exist, and improvements will be less reliable in low-resource domains.","authors":["Caleb Ziems","William Held","Su Doga Karaca","David Grusky","Tatsunori Hashimoto","Diyi Yang"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-07-30","first_seen":"2026-07-02","revised_at":"2026-07-30","abs_url":"https://arxiv.org/abs/2607.02464","pdf_url":"https://arxiv.org/pdf/2607.02464","source_feed":"cs.CL","score":9,"bucket":"selected","rubric_hits":["A1","A2","A3","B1","B4"],"tags":["LLM社会仿真","缩放规律","仿真保真度"],"reason":"直接研究LLM社会仿真的保真度与缩放规律，含人类数据对照和失效条件分析。","model":"deepseek-v4-pro","scored_at":"2026-07-30T13:02:01","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-30","rank":4,"question":"当前大语言模型的缩放范式能否缩小社会仿真保真度的差距，还是仿真保真度与通用能力正交？","design":"使用85个基于Qwen3架构、在DCLM语料上预训练的Transformer模型（0.2B–12B参数，固定计算预算10^18–10^20 FLOPs）进行受控计算缩放实验，并评估35个更大开源模型（最高70B参数），在意见建模、行为仿真和纵向预测三个子领域测量仿真损失与准确率。","baseline":"世界价值观调查（WVS）、Psych-101实验数据、美国人生活变迁（ACL）纵向研究等真实人类数据。","findings":"多数行为和意见仿真任务随模型规模扩大而快速改善，尤其在英语网络语料中代表性好的人群上；但纵向预测和代表性不足的意见缩放较慢，且缩放未能改善模型在风险厌恶等认知偏差及关联奖励学习启发式上的校准。","reliability":"在低资源领域和与通用推理基准相关性弱的任务上改进不可靠；缩放对风险厌恶等人类认知偏差的校准无效，微调后也未观察到参数缩放效果。","relevance":"该研究直接评估LLM社会仿真的缩放规律与失效条件，含真实人类对照，与关注经济学实验和政策评估仿真的研究者高度相关，值得精读。","inspiration":"借鉴其受控计算缩放实验与观测校准函数结合的方法，系统评估模型规模对仿真保真度的因果效应。｜可迁移到资产定价实验或消费者跨期选择等经济决策仿真，检验缩放是否改善风险偏好与时间偏好复现。｜以LLM为被试，施加不同风险/跨期选择任务，测量选择分布与真实实验数据（如实验室或调查数据）的偏差，用缩放定律预测更大模型的保真度。"}},{"id":"2607.26317","version":1,"title":"Aligning LLM-Simulated and Human Examinees for Psychometric Calibration: A Cognitive Diagnostic Profiling Approach","zh_title":"对齐LLM模拟考生与真实考生以进行心理测量校准：一种认知诊断画像方法","abstract":"Psychometric calibration for educational tests typically requires costly human response data. Large language models (LLMs) simulated examinees offer a promising route to early calibration, but their responses are too accurate and too uniform. We propose Cognitive Diagnostic Profiling (CDP), a zero-shot framework that prompts LLMs to simulate plausible examinees with diverse cognitive profiles: binary attribute-mastery patterns are rendered as natural-language profiles and sampled under an uninformative or an informative distribution. Using the Tatsuoka fraction-subtraction dataset (536 examinees, 15 items, five attributes), we evaluated eight LLM configurations under no-profile, uninformative-CDP, and informative-CDP conditions, assessing alignment with human examinees at the ability-distribution, mastery-profile, and item-difficulty levels. CDP improved all three levels: distributional overlap rose across configurations; weighted correlations between profile-level scores and human profile expectations reached 0.92 to 0.98; and item-difficulty recovery improved in rank order and absolute alignment, most for reasoning-enabled models; in the strongest case, Gemini 3.0 Flash (Thinking), one-parameter logistic (1PL) difficulty Spearman correlations rose from 0.24 to 0.86 and 0.90 and the root-mean-square error (RMSE) fell from 6.31 to 1.30 and 0.90; the informative condition helped most where profile-level alignment was strong. CDP brings LLM-simulated examinees into closer psychometric alignment with human examinees, making them practical for operational test development.","authors":["Wenjie Zhou","Yunting Liu","Renjiao Tang","Mark Wilson"],"categories":["cs.CY","cs.AI","cs.CL"],"primary_category":"cs.CY","announce_type":"cross","date":"2026-07-30","first_seen":"2026-07-30","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.26317","pdf_url":"https://arxiv.org/pdf/2607.26317","source_feed":"cs.CL","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2"],"tags":["LLM仿真","心理测量","人类数据对照"],"reason":"用LLM模拟考生作答并与真实人类数据对照，评估对齐效果，属于教育测量中的人类仿…","model":"deepseek-v4-pro","scored_at":"2026-07-30T13:01:41","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-31","rank":4,"question":"如何通过认知诊断画像提示，使大语言模型模拟的考生在心理测量校准中与真实人类考生对齐？","design":"使用八种大语言模型配置（含推理与非推理模型），在无画像、无信息认知诊断画像（CDP）和有信息CDP三种条件下，以零样本方式提示模型模拟考生作答Tatsuoka分数减法数据集（536名考生、15题、5个认知属性），评估生成的反应数据在能力分布、掌握模式及题目难度三个层面与人类数据的对齐程度。","baseline":"Tatsuoka分数减法数据集，包含536名真实人类考生对15道题目的作答反应及认知属性掌握模式。","findings":"CDP框架显著提升了LLM模拟考生与人类考生在能力分布、掌握模式和题目难度三个层面的对齐度；在最佳配置下，题目难度排序相关系数从0.24升至0.90，均方根误差从6.31降至0.90，有信息CDP在画像层面对齐较强时帮助最大。","reliability":"论文未讨论","relevance":"该研究直接以真实人类数据为基准，评估LLM仿真在心理测量校准中的对齐效果与偏差，属于教育测量场景下的人类仿真验证，与研究者关注的经济学实验和政策评估中的仿真可靠性问题高度相关，值得精读。","inspiration":"借鉴其通过结构化认知画像（属性掌握模式）注入异质性、并对比无信息与有信息分布采样的处理设计，以控制仿真人群的多样性与偏差。｜可迁移至教育经济学或劳动经济学中的技能测评场景，如职业资格考试的题目预测试或人力资本评估中的能力诊断。｜以LLM模拟不同技能掌握模式的求职者，处理为随机分配无信息或有信息的认知画像提示，结果变量为模拟作答反应，以真实大规模技能测评数据（如PIAAC）作为人类基准对照。"}},{"id":"2607.26288","version":1,"title":"The Innate Economic Preferences of Language Models","zh_title":"语言模型的内在经济偏好","abstract":"Language models increasingly settle real resource tradeoffs on behalf of principals yet their economic preferences remain unobserved. We demonstrate their generation rule is isomorphic to the random utility model of discrete choice. This allows internal logit scores to structurally identify preferences. Estimating risk attitudes across twelve models in a portfolio task reveals universal but heterogeneous risk aversion. Although models reject strictly dominated options, their elicited preferences fail invariance tests and violate the independence of irrelevant alternatives across varying experimental prompts. Finally, fine tuning establishes that a principal can explicitly engineer a target risk attitude.","authors":["Joy Buchanan","Joshua Foster"],"categories":["econ.EM"],"primary_category":"econ.EM","announce_type":"new","date":"2026-07-30","first_seen":"2026-07-30","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.26288","pdf_url":"https://arxiv.org/pdf/2607.26288","source_feed":"econ.EM","score":8,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["LLM经济偏好","人类仿真","风险态度"],"reason":"用LLM替代人类被试测量经济偏好，有真实人类数据对照，涉及风险态度和不变性检验…","model":"deepseek-v4-pro","scored_at":"2026-07-30T13:01:39","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-30","rank":6,"question":"语言模型在面临经济权衡时的默认偏好是什么，其选择是否满足显示性偏好公理从而可被解释为稳定效用？","design":"本研究并非用LLM仿真人类被试，而是将12个语言模型本身作为决策主体，在受控的投资组合选择任务中测量其风险态度。通过强制模型从不同风险-收益特征的资产菜单中单选一个资产，并利用模型输出的logit向量（开放模型）或重复抽样选择（闭源模型）来结构性地识别偏好参数，同时检验其选择是否满足完备性、自反性、单调性、传递性、连续性和无关选项独立性等显示性偏好公理。","baseline":"无对照","findings":"所有模型均表现出普遍但异质的风险厌恶，且拒绝严格劣选项；但其偏好未能通过不变性检验，并随实验提示变化而违反无关选项独立性。通过微调可以显式地工程化目标风险态度。","reliability":"论文指出模型的偏好对选项的呈现方式不具不变性，无关选项的加入或选项位置变动会改变偏好强度，尤其在接近无差异时这种不稳定性变得可观测，表明偏好强度不稳定而排序相对稳定。","relevance":"该研究直接测量LLM作为经济主体的内在偏好并检验其理性公理，虽未以人类为基准，但为评估LLM替代人类进行经济决策的可靠性提供了关键的方法论和实证证据，值得精读。","inspiration":"借鉴其利用模型内部logit直接观测系统效用指数的方法，可避免传统离散选择模型对误差分布的依赖，实现偏好的结构化识别。｜可迁移到信贷审批歧视研究中，用LLM扮演信贷员，测量其在不同申请人特征下的风险偏好与歧视程度。｜以多个LLM为被试，设计不同风险-收益特征的贷款申请菜单，记录模型选择的logit或重复抽样选择，估计其风险厌恶参数，并与真实信贷员的历史审批数据对照，检验LLM决策的偏差与一致性。"}},{"id":"2607.26588","version":1,"title":"Eco3S: Complex Socio-Economic System Simulation via Agent-Based Models","zh_title":"Eco3S：基于智能体的复杂社会经济系统仿真","abstract":"The rapid development of large language models (LLMs) has renewed interest in agent-based modeling (ABM). However, current LLM-based ABM research faces several key challenges: modeling evolving agent-environment interactions, enabling flexible counterfactual reasoning, and automating simulation workflows for scientific research. In this paper, we propose Eco3S, a socio-economic system simulation framework for economic research and policy analysis that addresses these challenges through three key mechanisms: (1) Co-evolving Environment Design, a bidirectional feedback loop where agents and the environment co-evolve, producing realistic emergent behaviors; (2) Structural Causal Simulation, a structural causal model (SCM)-inspired counterfactual mechanism that allows flexible interventions for diverse causal inference tasks; (3) Simulation-Analysis-Refinement Paradigm, a self-corrective mechanism that iteratively refines experimental designs based on prior simulation results. Experiments on diverse economic scenarios confirm \\textit{Eco3S}'s effectiveness in replicating multiple established economic studies (canal decay, origins of governance, and information propagation) and phenomena across domains. Additional results further demonstrate its scalability and generalizability, highlighting the framework's potential for rigorous economic research and policy-making.","authors":["Shaopeng Wei","Yufei Cheng","Wenxi Sun","Yepeng Ding","Yu Zhao","Gang Kou"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-07-30","first_seen":"2026-07-30","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.26588","pdf_url":"https://arxiv.org/pdf/2607.26588","source_feed":"cs.AI","score":7,"bucket":"pending","rubric_hits":["A3","B2","B3"],"tags":["LLM仿真","经济实验复现","因果推断"],"reason":"用LLM agent模拟社会经济过程并复现经典经济学研究，涉及因果推断，但未明…","model":"deepseek-v4-pro","scored_at":"2026-07-30T13:01:43","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-31","rank":9,"question":"如何构建一个能模拟社会经济系统协同演化、支持因果推断并自动化仿真流程的LLM-based ABM框架？","design":"提出Eco3S框架，用LLM驱动的智能体模拟社会经济系统中的个体，通过协同演化环境设计实现智能体与环境的双向反馈，引入结构因果模型进行反事实干预，并采用仿真-分析-精炼范式自动迭代优化实验设计。","baseline":"复现了运河衰败、治理起源和信息传播等经典经济学研究，但未明确说明对照的具体真实人类数据集。","findings":"Eco3S能有效复现多个经典经济学现象，展现出捕捉复杂动态的能力；框架具备可扩展性和通用性，适用于经济研究和政策分析。","reliability":"论文未讨论","relevance":"该研究直接使用LLM智能体复现经典经济学研究，并引入因果推断机制，与研究者关注的LLM仿真人类行为、复现经济实验及政策评估高度相关，值得深入阅读。","inspiration":"借鉴其结构因果仿真模块，可在LLM智能体模拟中施加政策干预并观察反事实结果，为政策评估提供因果证据。｜可迁移到政策公告的预期形成研究，模拟市场参与者在不同政策信号下的预期调整与资产价格变动。｜以LLM智能体作为投资者被试，处理为不同措辞的央行公告，结果变量为预期通胀率和资产配置变化，对照真实调查预期数据或市场数据。"}},{"id":"2607.25094","version":2,"title":"Evaluating Communicative Belief Updates in Large Language Models via Implicature Recognition and Cancellation","zh_title":"通过隐含意义识别与取消评估大语言模型的交际信念更新","abstract":"Human language is driven by unspoken beliefs and belief updates, making these critical to model for successful communication between large language models (LLMs) and their users. In this paper, we evaluate the ability of LLMs to recognize unspoken beliefs made through implicatures and to understand their updates through implicature cancellation: the pragmatic phenomenon whereby an utterance's implied meaning is weakened or negated. We create the first expert-annotated implicature cancellation dataset, ImplicatureX, crowdsourced for human judgements of implicatures and their corresponding cancellations. We find that LLM belief update understanding lags behind that of humans, especially in more naturally-occurring scenarios. Additional control experiments suggest that successes in LLM belief updates may stem in part from a reliance on prior beliefs, and that failures in belief updates may depend on their type and on their form. Overall, our study suggests that current LLMs have not yet reached human-level understanding of unspoken beliefs and belief updates. Code and data are available at https://github.com/cesare-spinoso/ImplicatureX.","authors":["Cesare Spinoso-Di Piano","Verna Dankers","Marius Mosbach","Jackie Chi Kit Cheung"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-07-30","first_seen":"2026-07-29","revised_at":"2026-07-30","abs_url":"https://arxiv.org/abs/2607.25094","pdf_url":"https://arxiv.org/pdf/2607.25094","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["LLM评估","语用推理","信念更新"],"reason":"评估LLM对隐含信念更新的理解，以人类数据为基准，但测量对象是模型能力而非仿真…","model":"deepseek-v4-pro","scored_at":"2026-07-30T13:02:01","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-31","rank":13,"question":"LLM能否像人类一样通过隐含意义识别和取消来理解交际中的信念更新？","design":"非仿真研究。构建专家标注的隐含意义取消数据集ImplicatureX，包含标量、话语和对话隐含意义，通过众包获取人类判断作为基准，测试多种LLM在隐含意义识别和取消上的准确率。","baseline":"众包收集的人类对隐含意义及其取消的判断准确率。","findings":"LLM在自然对话隐含意义识别上仅略高于随机水平，且成功可能部分依赖先验信念而非真正语用推理；在隐含意义取消后的信念更新上，LLM表现落后于人类，尤其在自然场景中，且更新类型和触发方式影响其表现。","reliability":"论文通过控制实验揭示LLM成功可能源于先验信念而非语境推理，且失败与更新类型（取消、不变、强化）和触发形式（显式/隐式）有关，表明当前LLM在自然交际信念更新上存在局限。","relevance":"该研究以人类数据为基准评估LLM的语用推理能力，揭示了LLM在模拟人类交际信念更新时的偏差与失效条件，对关注LLM仿真可靠性的研究者有参考价值，值得阅读原文。","inspiration":"可借鉴其构建专家标注数据集并利用众包人类判断作为基准的方法，用于严格评估LLM在特定任务上的仿真能力。｜可迁移到经济政策沟通场景，如央行公告中的隐含意图识别与修正对公众预期的影响。｜以LLM为被试，呈现含隐含政策意图的公告文本及后续澄清，测量其预期更新方向，并以真实公众调查数据为对照。"}},{"id":"2607.25253","version":2,"title":"The User Asks, Platforms Compete: How Agentic Recommendation Markets Take Shape","zh_title":"用户提问，平台竞争：代理式推荐市场如何形成","abstract":"Online recommendation has traditionally taken place after a user enters a platform, which determines the candidate pool and the ranking shown to the user. LLM-based user agents enable a different recommendation process: a user specifies a need before choosing a platform, leaving platforms to compete for the user's attention, which we refer to as an agentic recommendation market. In our controlled LLM-based experiments across three product domains, we find this new setting of recommendation creates a tension between access and attention. Compared with traditional platform-centric recommendation, user-centric recommendation greatly expands the opportunity for relevant items to enter comparison; yet broader participation does not translate directly into effective exposure. Competition directly triggers platforms' strategic play: selectively positive explanations occupy 73--78% of first-ranked positions. When the user agent relates platforms' actions to subsequent user feedback, this share falls to 36--41%, while the chance of a user purchasing the relevant item increases. A user agent is therefore more than a ranker over a larger pool of candidates: its querying, ranking, and feedback mechanism governing who can compete, how scarce attention is allocated, and how earlier outcomes shape the evaluation of platforms directly affect user utility. Designing agentic recommendation therefore requires treating access, attention, and accountability as a joint mechanism design problem.","authors":["Deyao Hong","Kehan Zheng","Qian Li","Jun Zhang","Jie Jiang","Hongning Wang"],"categories":["cs.AI","cs.IR"],"primary_category":"cs.AI","announce_type":"replace","date":"2026-07-30","first_seen":"2026-07-29","revised_at":"2026-07-30","abs_url":"https://arxiv.org/abs/2607.25253","pdf_url":"https://arxiv.org/pdf/2607.25253","source_feed":"cs.AI","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["LLM代理","推荐市场","社会模拟"],"reason":"用LLM agent模拟推荐市场，但无真实人类数据对照，属社会模拟边界情形。","model":"deepseek-v4-pro","scored_at":"2026-07-30T13:02:03","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-31","rank":14,"question":"在由LLM用户代理驱动的跨平台推荐市场中，平台竞争如何影响物品的获取、注意力分配与问责，以及用户效用如何变化？","design":"使用LLM构建用户代理和平台代理，在三个产品领域（乐器、电子游戏、运动户外）的Amazon评论数据上模拟推荐交互。通过控制市场参与平台数量、短名单容量、平台解释策略及历史反馈可用性，追踪目标物品在候选池出现、进入短名单、获得首位注意力和最终购买的概率。","baseline":"无对照","findings":"跨平台查询大幅提高目标物品进入候选池的机会，但更广泛的参与并未直接转化为有效曝光；平台会策略性地使用正面解释占据首位，而引入用户反馈机制可降低此类行为并提高购买概率。","reliability":"论文未讨论","relevance":"该研究利用LLM代理模拟推荐市场竞争，属于社会模拟边界情形，但缺乏真实人类数据对照，与研究者关注的有基准人类数据的仿真可靠性评估不完全匹配，可作为批判性案例参考。","inspiration":"借鉴其通过控制平台参与和反馈机制来观察策略行为变化的设计思路。｜可迁移到在线金融产品推荐市场的竞争与操纵行为研究，如理财平台对用户注意力的争夺。｜以LLM代理模拟投资者，设置不同数量的理财平台代理，操纵平台是否提供历史收益反馈，测量投资者对推荐产品的点击与购买决策，并以真实理财平台用户行为日志作为对照。"}},{"id":"2607.26060","version":1,"title":"Large-Scale ChatBot Validation Through Customer Digital Twin Simulations","zh_title":"通过客户数字孪生仿真进行大规模聊天机器人验证","abstract":"LLM-based chatbots are transforming customer service in regulated domains such as banking, but scalable and cost-effective validation remains a critical barrier to safe deployment. We present a two-part contribution for large-scale chatbot validation. First, we introduce a methodology for creating high-fidelity synthetic customer agents (SCAs) as digital twins, grounded in real transactional and conversational data, that enables automatic generation and behavioral conditioning to simulate diverse customer profiles and interaction styles. Evaluation demonstrates that SCAs achieve high semantic alignment with real customers, low hallucination rates, and successful personality trait reproduction with controllable interventions. Second, we develop an SCA-based validation framework combining automated LLM-as-a-Judge evaluation, human expert testing, and adversarial probing. Scenario-based validation across emotional states, demographic groups, and linguistic factors confirms robust performance. Our approach was used to validate a customer facing chatbot at a leading UK bank, providing financial institutions with a scalable pathway toward regulatory compliance.","authors":["Cristovao Iglesias","Devesh Batra","Alankar Atreya","Stefan Wagner","Robert Hankache","Patrick Sinclair","Giulio Pelosio","Michael McMillan","Greig A. Cowan","Raad Khraishi"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-07-30","first_seen":"2026-07-30","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.26060","pdf_url":"https://arxiv.org/pdf/2607.26060","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["客户数字孪生","聊天机器人验证","社会模拟"],"reason":"用合成客户代理模拟客户行为，但无真实人类行为对照，属社会模拟边界情形。","model":"deepseek-v4-pro","scored_at":"2026-07-30T13:01:39","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-31","rank":15,"question":"如何利用基于真实交易和对话数据构建的高保真合成客户代理（SCA）对银行客服聊天机器人进行大规模、可扩展的验证？","design":"使用LLM驱动的合成客户代理（SCA）作为数字孪生，基于真实交易和对话数据构建，通过转录驱动和人格条件模拟两种方式生成多样化的客户交互行为，与目标聊天机器人进行多轮对话，评估任务完成、安全性、对话质量、公平性和真实性等指标。","baseline":"无对照","findings":"SCA生成的对话在语义上与真实客户高度一致，但词汇重叠度低；合成对话的事实保真度较高，偏差主要表现为细节遗漏或少量事实编造。","reliability":"论文未讨论","relevance":"本文属于社会模拟边界情形，用合成客户代理模拟客户行为，但缺乏真实人类行为对照，与研究者关注的有真实人类基准的仿真研究不完全匹配，但方法学上可提供参考。","inspiration":"可借鉴其利用真实交易数据构建数字孪生并进行行为条件干预的方法，实现可控的客户行为模拟。｜可迁移到金融消费者行为研究，如信贷产品选择、投诉处理或金融建议接受度等场景。｜以真实银行客户交易和对话记录构建合成客户代理，施加不同情绪或人格干预（如焦虑、愤怒），测试其对理财建议聊天机器人的接受度和决策变化，结果与历史真实客户行为数据对照。"}},{"id":"2607.26389","version":1,"title":"Misalignment Has a Personality: A Big Five Account of Emergent Misalignment","zh_title":"错位有性格：基于大五人格的新兴错位解释","abstract":"Fine-tuning a language model on data containing a narrow flaw, such as insecure code or incorrect mathematical answers, can cause broad misalignment through a mechanism that remains debated. We provide an interpretable account: in the models and corpora we study, misalignment behaves like a shift in personality. Prior work extracts activation directions for character traits from a single binary contrast, which can separate or steer behavior without establishing a calibrated scale. We instead extract personality vectors for the Big Five using a graded, three-level intervention and validate them on two open-weight models. The three levels are linearly ordered, with Cohen's d values of up to 6.2; the vectors transfer zero-shot and trait-specifically to an independent corpus; and their effects are strongest within a middle-layer band. Applied to training data, the vectors reveal that misaligned corpora across eight domains share a common Big Five signature: lower agreeableness and conscientiousness, together with higher extraversion and neuroticism. This signature is recovered by both models with a correlation of r = 0.94. Fine-tuning imprints the same profile, shifting the model's generations along the corresponding signature, with r = 0.83 using activation-based measurements and r = 0.90 using a text-based judge, while also shifting internal activations with r = 0.69. The same vectors characterize sycophancy as high extraversion and low conscientiousness rather than excess agreeableness, a distinction that a single direction cannot capture. Calibrated personality vectors transform an opaque safety phenomenon into a human-legible diagnostic profile.","authors":["Hasibur Rahman","Smit Desai"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-07-30","first_seen":"2026-07-30","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.26389","pdf_url":"https://arxiv.org/pdf/2607.26389","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["LLM人格测量","模型对齐","大五人格"],"reason":"测量LLM的人格特质，非仿真人类被试，但方法可能迁移到仿真研究。","model":"deepseek-v4-pro","scored_at":"2026-07-30T13:01:54","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-30","rank":14,"question":"在微调数据中引入窄缺陷导致的大语言模型广泛失准（emergent misalignment）是否表现为一种可解释的人格转变，其大五人格特征是什么？","design":"本研究并非人类仿真实验，而是通过提取大五人格向量来测量和解释模型的失准行为。具体做法：使用分级（低、中、高）的Trait Modulation Keys提示，在Qwen2.5-7B-Instruct和Llama-3.1-Nemotron-Nano-8B两个模型上，通过高低对比提取每个特质的激活方向，并用留出的中等水平验证向量的有序性；然后将这些向量应用于八个失准领域的微调数据，测量其人格特征，并观察微调后模型在无关问题上的生成和内部激活是否呈现相同的人格转变。","baseline":"无对照","findings":"八个失准领域的数据共享相同的大五人格特征：宜人性和尽责性降低，外向性和神经质升高，两个模型对此特征的相关性达r=0.94。微调会将该特征印刻到模型中，使其在行为（r=0.83-0.90）和内部激活（r=0.69）上均表现出相同的人格转变。","reliability":"论文承认其证据仅限于两个7-8B参数的模型、英语语料和一种微调方法，未在其他规模、语言或微调范式下验证。","relevance":"本文虽非直接的人类仿真研究，但其提取校准人格向量的方法（分级干预、有序性验证、特质特异性迁移）可迁移到用LLM仿真人类被试的场景中，用于测量和操控仿真体的性格特征，值得精读。","inspiration":"该方法通过分级提示构建有序人格向量，并利用留出中等水平验证向量的测量属性，为在LLM中建立校准的心理构念量表提供了可借鉴的流程。｜可迁移到行为经济学中的个体异质性仿真，例如在跨期选择实验中，用LLM扮演不同人格的被试，观察其时间贴现率的分布。｜以LLM作为被试，通过人格向量操控其宜人性或尽责性水平，测量其在标准跨期选择任务中的贴现因子，并与真实人类实验数据（如Andersen et al., 2008）进行分布对比，验证仿真偏差。"}},{"id":"2607.26853","version":1,"title":"From Representations to Behaviors: Exploring the Person-Situation-Behavior Triad in LLMs","zh_title":"从表征到行为：探索大语言模型中的人-情境-行为三元组","abstract":"Human personality theories characterize traits not as isolated attributes captured by a single score, but as stable individual tendencies expressed through the interplay among persons, situations, and behaviors. Existing studies of personality-related behavior in LLMs have primarily focused on outputs elicited under personality conditioning, characterizing observable trait-related expressions while lacking mechanistic evidence for the existence of internal personality-related representations, their cross-situational expression, and how these representations shape specific behaviors. Building on Funder's personality triad framework, we adapt its three components for LLM analysis: Person as personality-related internal representations, Situation as contexts that afford trait-relevant responses, and Behavior as response patterns on broader social tasks. We introduce a framework for discovering, controlling, and validating trait-like representations in LLMs. First, using contrastive behavior pairs grounded in shared situations, we identify sparse internal features associated with opposing poles of personality traits through SAE decomposition. We validate their trait relevance through effects on behavior to situation, token-level activation patterns, and robustness to paraphrasing. Second, feature-level interventions induce bidirectional trait-related shifts across a separate, diverse set of situations while preserving response validity, demonstrating consistent expression across contexts. Third, applying the same interventions to social intelligence tasks reveals behavioral changes with benefit-tradeoff patterns consistent with findings from human personality research, providing behavioral-level validation beyond personality scores. Our findings provide evidence that LLMs contain controllable trait-like representations linking internal states, situational expression, and behavioral outcomes.","authors":["Ruikang Zhang","Shuo Wang","Qi Su"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-07-30","first_seen":"2026-07-30","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.26853","pdf_url":"https://arxiv.org/pdf/2607.26853","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["人格测量","表征工程","社会智能任务"],"reason":"测量LLM内部人格表征与行为，属人格测量，非仿真人类被试，无真实人类数据对照。","model":"deepseek-v4-pro","scored_at":"2026-07-30T13:01:43","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-31","rank":20,"question":"LLM内部是否存在可识别、可操控的人格特质表征，并能跨情境一致表达并影响更广泛的社会行为？","design":"本研究并非人类仿真实验，而是通过稀疏自编码器从对比行为对中提取LLM内部人格相关特征，再通过特征级干预操控这些表征，在多样化情境任务和社会智能任务上测量行为变化。","baseline":"无对照","findings":"通过对比行为对和稀疏自编码器可识别出与人格特质相关的内部特征，这些特征能跨情境一致地影响行为。干预这些特征可双向改变LLM在社交任务上的表现，且变化模式与人类人格研究中的收益-权衡模式一致。","reliability":"论文未讨论","relevance":"本文探索LLM内部人格表征的机制与操控，未进行人类仿真或与真实人类数据对照，与研究者关注的人类被试替代仿真和基准对照方向关联较弱。","inspiration":"与经济金融研究关联不大"}},{"id":"2607.26981","version":1,"title":"OptimismBench: Forecasting Bias and the Alignment Effect in Language Model Judgment","zh_title":"OptimismBench：语言模型判断中的预测偏差与对齐效应","abstract":"Large language models are increasingly used as decision aids whose probability judgments shape downstream choices. Whether those judgments carry a systematic directional tilt has been hard to detect: calibration metrics aggregate unsigned errors, and naturalistic uncertainty offers no ground-truth probability. When an LLM rates a startup's success at 70% but its failure at 15%, the missing 15 points expose a distortion no aggregate score flags. We introduce OptimismBench, which detects directional bias with inverted pairs: each scenario elicits both P(success) and P(failure), and asymmetry between the two framings yields a signed bias score without ground truth. Across 16 models from 8 providers, fourteen are optimistic; pessimism appears only in Anthropic's frontier tier. Eleven matched base-versus-chat pairs across four families show post-training sets the sign of the bias, with opposite shifts in different families. The pattern survives prompt, temperature, perspective, and self-debiasing ablations. A seventeen-model six-language comparison further shows model identity dominates language, with inter-model variance at 4.7x inter-language variance. We release 3,870 items across 10 languages for per-model directional-bias auditing. When alignment makes a model more helpful, it also tilts its probabilities; downstream pipelines inherit the tilt by default.","authors":["Seonglae Cho","Adriano Koshiyama"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-07-30","first_seen":"2026-07-30","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.26981","pdf_url":"https://arxiv.org/pdf/2607.26981","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["LLM偏差","概率判断","模型心理测量"],"reason":"测量LLM自身的概率判断偏差，属于模型心理测量，非仿真人类被试","model":"deepseek-v4-pro","scored_at":"2026-07-30T13:01:46","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-30","rank":18,"question":"大语言模型在概率判断中是否存在系统性的方向性偏差（乐观或悲观）？","design":"本研究并非仿真人类被试，而是直接测量LLM自身的概率判断偏差。通过构建倒置配对（inverted pairs）方法，对每个场景分别询问成功概率和失败概率，计算两者之和与100的偏离（Skew）作为方向性偏差指标，无需真实基准概率。在16个模型、60个场景、10种语言上进行评估，并进行了提示词、温度、视角、自去偏等消融实验，以及基座模型与聊天模型的配对比较。","baseline":"无对照","findings":"16个模型中有14个表现出乐观偏差，仅Anthropic的前沿模型（Opus、Sonnet）表现出悲观偏差；后训练（post-training）决定了偏差的方向，且在不同模型家族中方向相反。模型身份对偏差的影响远大于语言，跨模型方差是跨语言方差的4.7倍。","reliability":"论文未讨论","relevance":"本文不涉及用LLM仿真人类被试，而是测量LLM自身的判断偏差，属于模型心理测量，与研究者关注的LLM替代人类被试的仿真研究不直接相关。但其中揭示的LLM系统性乐观/悲观偏差可能作为仿真中的混杂因素，值得了解。","inspiration":"倒置配对设计可借鉴用于测量经济预测中的方向性偏差，无需真实基准概率。｜可迁移到资产定价实验或信贷审批场景，检测LLM在预测股票上涨/下跌概率或贷款违约/履约概率时是否存在系统性乐观或悲观。｜以LLM为被试，给出公司财务指标后分别询问其成功概率与失败概率，计算Skew作为偏差指标，对照真实历史违约率或分析师一致预期数据，检验LLM预测偏差的方向与幅度。"}},{"id":"2607.27022","version":1,"title":"Evaluating Regional Bias in LLMs From Abstract Stereotype to Concrete Social Decision-Making","zh_title":"评估大语言模型中的区域偏见：从抽象刻板印象到具体社会决策","abstract":"Regional bias in large language models (LLMs) may shape both perceptions of regional groups and decisions about individuals from different regions. Yet existing studies often examine these manifestations separately, leaving their structure and consequences unclear. We introduce Stereotypes-to-Decisions (S2D), a systematic framework evaluating regional bias from abstract stereotypes to concrete social decisions. Covering all 34 provincial-level administrative regions of China, S2D evaluates six LLMs using stereotype ratings of Warmth (perceived friendliness and trustworthiness) and Competence (perceived capability and intelligence), along with paired-choice tasks across Education, Occupation, and Social Interaction. Results reveal substantial regional differences in regional scores, with considerable agreement across models, especially for Competence and Occupation decisions. Furthermore, these patterns are associated with regional economic and digital development indicators and display mixed human-like stereotypes, with some regions rated highly on one dimension but poorly on the other. They also remain largely stable across Chinese and English prompts. Overall, our findings show that regional bias in LLMs is prevalent, systematic, and consequential, motivating more regionally aware evaluation and mitigation.","authors":["Jiayuan Di","Haoyi Yang","Yufei Luo","Jiahui Qu","Yiming Wang"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-07-30","first_seen":"2026-07-30","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.27022","pdf_url":"https://arxiv.org/pdf/2607.27022","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["区域偏见","刻板印象","社会决策"],"reason":"测量LLM本身的区域刻板印象，非仿真人类被试，但涉及社会决策对照，属边界情形。","model":"deepseek-v4-pro","scored_at":"2026-07-30T13:01:46","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-30","rank":19,"question":"LLM 是否表现出从抽象刻板印象到具体社会决策的系统性区域偏见？","design":"本研究并非仿真人类被试，而是直接评估 LLM 本身的偏见。使用 S2D 框架，以中国 34 个省级行政区为区域身份，在抽象层面让模型对温暖与能力维度进行 Likert 五点评分，在具体层面让模型在教育、职业、社交三个领域的配对选择任务中做出决策，测量各区域的刻板印象分数和决策分数。","baseline":"无对照","findings":"LLM 在抽象刻板印象和具体社会决策中均表现出显著的区域差异，且不同模型在区域排名上具有较高一致性，尤其在能力刻板印象和职业决策上。区域偏见与地区经济及数字化发展指标相关，并呈现类似人类的混合刻板印象结构，且在中英文提示下保持稳定。","reliability":"论文未讨论","relevance":"本文直接测量 LLM 的区域偏见，并非用 LLM 仿真人类被试，但涉及社会决策中的歧视性选择，与研究者关注的仿真可靠性及偏差评估有边界关联，可提供 LLM 内在偏见的基准认知。","inspiration":"可借鉴其从抽象态度到具体决策的分层评估框架，以及配对选择任务中通过控制候选人其他特征来孤立区域身份影响的设计。｜可迁移至信贷审批或招聘筛选中的地域歧视研究，例如评估 LLM 在模拟信贷员或 HR 时是否对特定地区申请人产生系统性偏好。｜以 LLM 为被试，处理为申请人简历中仅变动户籍省份，结果变量为贷款批准或面试邀请的二元选择，对照真实信贷或招聘数据中的地域差异模式，检验 LLM 仿真决策的偏差方向与程度。"}},{"id":"2607.26062","version":1,"title":"Identifying Implicit Bias in LLM-based Chat AI Toward People with Intellectual Disabilities","zh_title":"识别基于大语言模型的聊天AI对智障人士的隐性偏见","abstract":"Background: This work investigates the presence of implicit bias in Large Language Model (LLM)-based chat AI models directed toward people with intellectual disabilities (ID). Objective: The study aims to identify and measure representational differences related to people with ID and examine them to identify implicit biases inherent in AI chat generation technologies. Methods: Utilizing the GPT-4-Turbo model, we requested story-generation based on 10 prompt stems with and without descriptors for ID. This process was repeated using four other LLMs (OpenAI GPT-4o, Meta Llama-3-3-70B-Instruct, Anthropic Claude-3-5-Sonnet, and Mistral-Large-2411). The resulting 25,000 computer-generated stories were analyzed using a separate GPT-4-Turbo model instance to detect differences in how people are represented related to themes of bias described in previous literature. Results: Our findings reveal differences in how people are represented between story datasets with and without ID descriptors. These differences go beyond established characteristics of ID and imply the presence of mostly negative implicit biases. Identified differences related to considering people with ID as younger, with themes of paternalism and infantilization; depicting them as more inspirational and symbolic; as needing help more often, being dependent, and being saved; and having a negative perception of them and more hesitation to include them. Conclusions: These implicit biases are considered within the context of past discrimination towards people with ID and highlight the need for diligence against implicit bias towards people with ID in AI development. This research underscores the importance of assessing and mitigating implicit bias in decision-making technologies to prevent future societal harm.","authors":["Karly V. Coffey","Gloria L. Krahn","John P. Hanley","Jacob E. Neely"],"categories":["cs.CY","cs.AI","cs.CL"],"primary_category":"cs.CY","announce_type":"cross","date":"2026-07-30","first_seen":"2026-07-30","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.26062","pdf_url":"https://arxiv.org/pdf/2607.26062","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["隐性偏见","LLM偏见测量","智障人士"],"reason":"测量LLM对智障人士的隐性偏见，属于对模型本身的测量，非仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-07-30T13:01:39","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-31","rank":16,"question":"基于LLM的聊天AI对智力障碍者是否存在隐性偏见？","design":"本研究并非人类仿真实验，而是对LLM本身的偏见测量。使用GPT-4-Turbo等五种LLM，基于10个提示词干（含或不含智力障碍描述符）生成故事，再用另一个GPT-4-Turbo实例分析故事中人物表征的差异，检测与偏见主题相关的差异。","baseline":"无对照","findings":"含智力障碍描述符的故事中，人物被描绘得更年轻、更依赖、更需帮助，并带有家长式、幼稚化和励志象征色彩。这些表征差异超出了智力障碍的已知特征，暗示存在以负面为主的隐性偏见。","reliability":"论文未讨论","relevance":"该研究测量LLM对特定群体的隐性偏见，属于模型本身的评估，并非用LLM仿真人类被试，与研究者关注的LLM替代人类被试进行实验复现的方向不同，但可作为评估仿真偏差的参考。","inspiration":"与经济金融研究关联不大"}},{"id":"2607.26179","version":1,"title":"Cognitive Convergence: Deep Similarities Between Large Language Models and Human Cognition","zh_title":"认知趋同：大语言模型与人类认知之间的深层相似性","abstract":"LLMs are widely regarded as alien intelligences, systems whose cognitive operations are fundamentally unlike our own. Apparent similarities to human cognition are therefore often seen as the result of anthropomorphic projection. We argue that this framing is mistaken. LLMs clearly differ from humans in important respects, including their physical substrate, learning history, and the environments with which they interact. These differences make it all the more striking that contemporary LLM-based systems converge with human cognition on a number of principles of cognitive organization with longstanding support in cognitive science. We identify structural correspondences across five dimensions: inferential organization, computational architecture, representational structure, prediction-driven learning, and reinforcement-learning-like mechanisms supporting goal-directed action. These correspondences support a broader model of intelligent cognition in which core principles long used to explain human intelligence also characterize contemporary LLM-based systems.","authors":["Chandra Sripada","Richard Lewis"],"categories":["q-bio.NC","cs.AI","cs.CL"],"primary_category":"q-bio.NC","announce_type":"cross","date":"2026-07-30","first_seen":"2026-07-30","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.26179","pdf_url":"https://arxiv.org/pdf/2607.26179","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["认知科学","LLM认知比较","理论分析"],"reason":"论文比较LLM与人类认知结构，属于将LLM作为测量对象，但非仿真人类被试，无实…","model":"deepseek-v4-pro","scored_at":"2026-07-30T13:01:52","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-31","rank":18,"question":"LLM与人类认知在组织原则上是否存在深层结构对应，而非仅是行为表面相似？","design":"本文非仿真研究，而是理论比较分析。作者从认知科学中提取五个维度（推理组织、计算架构、表征结构、预测驱动学习、强化学习式目标导向机制），系统对比LLM与人类认知的结构性对应关系。","baseline":"无对照","findings":"LLM与人类在推理组织、计算架构、表征结构、预测驱动学习和目标导向机制五个维度上存在深层结构对应，这些对应表明二者共享智能认知的核心原则。尽管在物理基础、学习历史和环境交互上存在显著差异，但认知组织上的趋同挑战了LLM是“异类智能”的流行观点。","reliability":"论文未讨论","relevance":"本文未进行人类仿真实验，而是将LLM作为认知模型与人类认知结构进行比较，不涉及用LLM替代人类被试复现调查或实验行为，因此与研究者关注的LLM仿真人类被试的可靠性及偏差评估不直接相关。","inspiration":"与经济金融研究关联不大"}},{"id":"2607.26473","version":1,"title":"Learning Dynamic User Personas from Implicit Interaction Streams via Iterative Refinement","zh_title":"通过迭代优化从隐式交互流中学习动态用户画像","abstract":"Personalizing large language models (LLMs) to individual users is essential for improving user experience, yet existing approaches typically rely on explicit preference supervision such as pairwise comparisons or demographic attributes, limiting their applicability in natural interaction settings. We propose IRIS, a framework that learns dynamic user personas directly from implicit interaction streams by extracting behavioral signals from everyday conversations and iteratively refining persona representations through a prediction-driven closed loop without requiring explicit feedback. We introduce an evaluation protocol based on behavior prediction, persona stability, and decision prediction. A proof-of-concept study on a synthetic interaction stream derived from public-domain autobiographical text shows that IRIS produces stable personas and distinguishes individual users while revealing limitations of memory-only approaches on recall-oriented metrics. We then validate IRIS on anonymized real-world Reddit r/AmItheAsshole (AITA) data, with personas built solely from each author's historical interactions. Across 100 authors, IRIS achieves the highest decision prediction accuracy among all evaluated methods (61.0%), outperforming static personas, memory-only retrieval, and no-personalization baselines. These results suggest that implicit behavioral modeling provides a scalable alternative to explicit preference learning for personalized LLMs and offers a practical foundation for adaptive conversational systems and embodied agents that require continuously evolving models of their users.","authors":["Haifeng Wu"],"categories":["cs.LG","cs.CL"],"primary_category":"cs.LG","announce_type":"cross","date":"2026-07-30","first_seen":"2026-07-30","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.26473","pdf_url":"https://arxiv.org/pdf/2607.26473","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["用户画像","个性化LLM","行为预测"],"reason":"用LLM从交互流学习动态用户画像，替代显式偏好标注，属标注替代而非仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-07-30T13:01:55","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-30","rank":15,"question":"能否仅从用户的隐式交互流中学习动态用户画像，而无需显式偏好监督？","design":"提出IRIS框架，通过预测驱动的闭环从隐式交互流中迭代提炼用户画像。先用LLM从对话日志提取行为信号并合成画像，再用画像预测用户行为，最后用预测误差驱动画像更新。在合成自传文本和Reddit AITA真实数据上评估，测量行为预测准确率、画像稳定性和决策预测准确率。","baseline":"Reddit r/AmItheAsshole数据集中100位作者的历史交互记录，以其真实投票决策作为基准。","findings":"IRIS在100位作者的真实数据上取得61.0%的决策预测准确率，优于静态画像、纯记忆检索和无个性化基线。合成数据实验显示IRIS能生成稳定画像并区分不同用户，但纯记忆方法在回忆导向指标上表现更好。","reliability":"论文指出隐式信号存在歧义（如查询改写可能源于不满意或自我修正），早期交互数据稀疏，以及用户偏好随时间漂移导致静态画像过时。此外，合成数据实验中纯记忆方法在回忆指标上优于IRIS，揭示了当前方法的局限。","relevance":"本文不直接进行人类仿真实验，而是用LLM从交互流学习用户画像以替代显式偏好标注，属于标注替代技术。对关注LLM仿真人类被试的研究者参考价值有限，但其中迭代提炼和预测误差驱动的闭环设计思路可借鉴。","inspiration":"借鉴其预测误差驱动的闭环迭代更新机制，可用于动态建模经济主体偏好或信念。｜可迁移到消费者跨期选择实验，模拟个体时间偏好随经济环境变化的动态过程。｜以LLM扮演消费者被试，处理为不同经济新闻推送（通胀、失业率变化），结果变量为即时消费与储蓄决策，用真实面板调查数据（如PSID）中的消费-收入动态作为对照基准。"}},{"id":"2607.26067","version":1,"title":"The Easy Trap: Why LLMs Underestimate Misconception-Driven Difficulty","zh_title":"简单陷阱：为何大语言模型低估由误解驱动的难度","abstract":"Large language models (LLMs) are increasingly used for estimating item difficulty in educational assessment. However, it remains unclear whether such estimates reflect how learners actually experience difficulty. This study investigates the alignment between LLM-generated difficulty ratings and empirical student performance on basic mathematics tasks. Four widely used LLM-based systems generated difficulty ratings on a 1-100 scale for 32 arithmetic items across multiple runs (N = 640 ratings). These were compared with empirical difficulty derived from responses of 770 Indonesian undergraduates using Classical Test Theory (CTT) and Item Response Theory (2PL). Results show moderate rank correlations (Spearman's rho = 0.52-0.70), indicating that LLMs capture coarse ordering of item difficulty. However, substantial and systematic misalignment emerges in fraction items. Several items consistently rated as easy by LLMs were among the most difficult for students, such as an item with only 34.16% correct for 100 : 1/2. We argue that LLMs approximate curricular difficulty, or what should be easy based on instructional sequencing, rather than cognitive difficulty driven by learner misconceptions. This leads to systematic underestimation of misconception-driven items, a phenomenon we term the Easy Trap. These findings highlight a critical limitation of LLM-based difficulty estimation and suggest that relying on such estimates without empirical grounding may introduce bias in assessment design and adaptive systems.","authors":["Amanda La Hadi","Muhammad Johan Alibasa","Guanliang Chen","A. Taufiq Asyhari"],"categories":["cs.CY","cs.AI"],"primary_category":"cs.CY","announce_type":"cross","date":"2026-07-30","first_seen":"2026-07-30","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.26067","pdf_url":"https://arxiv.org/pdf/2607.26067","source_feed":"cs.AI","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM评估","题目难度","教育测量"],"reason":"LLM替代人工评估题目难度，非仿真人类被试，但涉及与真实学生数据对照，属边界情…","model":"deepseek-v4-pro","scored_at":"2026-07-30T13:01:39","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-31","rank":17,"question":"LLM生成的题目难度估计在多大程度上与真实学生的算术表现一致，以及这种不一致如何揭示模型未能捕捉的误解驱动型认知需求？","design":"本研究并非用LLM仿真人类被试，而是让四种LLM系统对32道基础算术题生成1-100的难度评分（共640个评分），然后与770名印尼本科生的真实答题数据（基于CTT和2PL IRT计算的实证难度）进行对比。","baseline":"770名印尼本科生的真实答题数据，通过经典测验理论（CTT）和项目反应理论（2PL）计算题目难度。","findings":"LLM难度估计与学生实证难度之间呈中等秩相关（Spearman's ρ=0.52-0.70），能粗略捕捉题目排序，但在分数题上出现系统性偏差：LLM常将学生实际表现很差的题目（如100÷1/2正确率仅34.16%）评为容易，低估了由误解驱动的认知难度。","reliability":"论文指出LLM主要编码课程预期难度而非认知难度，导致对误解驱动型题目系统性低估（称为“Easy Trap”），且仅基于文本表面特征，缺乏实证校准，在评估设计和自适应系统中可能引入偏差。","relevance":"本研究虽非直接仿真人类被试，但系统比较了LLM输出与真实人类行为数据，揭示了LLM在认知建模中的系统性偏差，对关注LLM仿真可靠性与失效条件的研究者具有参考价值，值得阅读原文。","inspiration":"可借鉴其将LLM输出与真实人类行为数据直接对照、并区分课程难度与认知难度的分析框架。｜可迁移到经济金融领域的风险认知或金融素养评估场景，例如用LLM估计金融产品风险等级或投资决策难度。｜以LLM作为评估工具，对一组金融决策题目（如复利计算、风险分散）生成难度评分，以真实投资者或消费者的答题正确率和反应时作为对照基准，检验LLM是否低估了由常见认知偏误（如货币幻觉、损失厌恶）驱动的题目难度。"}},{"id":"2607.26545","version":1,"title":"A Persona-based Rate Action Index","zh_title":"基于人格体的利率行动指数","abstract":"We propose an index for predicting the U.S.\\ Federal Open Market Committee (FOMC) decision to hike/hold/cut the current federal funds target rate based on how a collection of personas responds to current market conditions. To construct the index, we collected a new dataset consisting of nearly $25{,}000$ retrievable chunks from publicly available data. We partition the data into per-member corpora and use each as the retrieval database of a generative system we refer to throughout as a ``persona''. We first evaluate the personas across two complementary components of likeness: identifiability and detectability. Each persona's behavior is highly attributable (average member-conditional recall is $ 8\\times $ chance) and generated content is nearly indistinguishable from held-out real content ($\\hat\\tau_{\\mathrm{det}} = 0.23$ against a $0.15$ floor). We then present evidence that query-conditioned representations of the personas capture members' monetary-policy stance relative to a known hawk--dove reputational ordering (Kendall's $\\tau = 0.63$, $p < 0.001$), substantially outperforming retrieval-only representations. These representations vary with time and current market conditions and form the basis of our proposed persona-based rate action index. For the $2022$--$2025$ period the index tracks the rate cycle (Kendall's $\\tau = 0.68$, $p < 10^{-6}$) and can be used to construct a simple classifier that predicts per-meeting outcomes at non-trivial accuracy ($0.69$ versus a $0.47$ base rate). Importantly, the index outperforms informative baselines and leads the federal funds target rate by roughly three quarters. As far as we are aware, our results are the first to demonstrate the ability to capture time-varying group behavior via a collection of digital personas.","authors":["Hayden Helm","Andrew Dassori"],"categories":["cs.MA","cs.AI","cs.LG"],"primary_category":"cs.MA","announce_type":"cross","date":"2026-07-30","first_seen":"2026-07-30","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.26545","pdf_url":"https://arxiv.org/pdf/2607.26545","source_feed":"cs.AI","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["LLM人格体","货币政策模拟","社会模拟"],"reason":"用LLM persona模拟FOMC成员决策，但无真实人类行为对照，属社会模拟…","model":"deepseek-v4-pro","scored_at":"2026-07-30T13:01:43","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-31","rank":19,"question":"如何基于FOMC成员的公开历史文本构建个性化数字人格，并聚合为预测联邦基金目标利率调整（加息/维持/降息）的指数？","design":"为FOMC每位成员构建一个基于检索增强生成（RAG）的“persona”（基座模型+成员专属检索数据库），用其公开讲话文本作为检索库；通过向persona输入当前市场状况的查询，获取其生成的货币政策立场表征，再聚合为委员会层面的指数，预测利率决策。","baseline":"无对照","findings":"Persona的行为具有高可归因性（成员条件召回率是随机水平的8倍）且生成内容与真实内容几乎无法区分；基于persona的指数能追踪2022–2025年利率周期（Kendall's τ=0.68），并领先联邦基金目标利率约三个季度。","reliability":"论文未讨论","relevance":"该研究利用LLM persona模拟FOMC成员决策，但缺乏真实人类行为对照，不符合研究者对基准人类数据的要求；然而其构建persona的验证方法（可识别性与可检测性）和动态指数构建思路对仿真可靠性评估有参考价值。","inspiration":"借鉴其利用成员历史文本构建个性化检索增强生成persona，并通过查询条件化表征捕捉时变政策立场的方法。｜可迁移到央行沟通的预期形成研究，例如模拟不同沟通风格对市场利率预期的影响。｜以FOMC会议声明为处理，构建投资者persona并测量其利率预期变化，用联邦基金期货隐含利率作为真实对照。"}},{"id":"2607.27179","version":1,"title":"The Social Cost of an AI Teammate: How an Artificial Teammate Reshapes Human-Human Communication in Small-Team Decision-Making","zh_title":"AI队友的社会成本：人工智能队友如何重塑小团队决策中的人际沟通","abstract":"Conversational AI is increasingly positioned as a teammate rather than a tool, yet we know little about how its presence reshapes communication among the humans on the team. We examined sociocognitive communication dynamics in team decision-making using Group Communication Analysis (GCA), team surveys, and lexical analyses of team discourse. Teams completed a high-stakes moral-dilemma decision task in a randomized controlled study: 16 teams of two students plus an AI teammate, and 17 all-human teams of three. Across six GCA dimensions and survey outcomes, we find that the AI teammate was the single most talkative and self-cohesive member of every treatment team, yet its contributions carried the least new information and the lowest density. The presence of AI also reshaped communication amongst humans. In AI-human teams, human teammates showed lower responsivity and social impact toward one another and reported lower levels of belonging and status. Greater AI dominance in the conversation was associated with students feeling less valued as team members. Additionally, this social cost is immediate and present at baseline; it does not emerge over the course of the conversation. Drawing on these results, we discuss a research agenda extending to voice-based and longitudinal settings.","authors":["Nia Nixon","Jaeyoon Choi","Pedro Martins De Bastos","Mohammad Amin Samadi","Luise Mehner","Seehee Park","Spencer JaQuay"],"categories":["cs.HC","cs.AI","cs.CY"],"primary_category":"cs.HC","announce_type":"cross","date":"2026-07-30","first_seen":"2026-07-30","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.27179","pdf_url":"https://arxiv.org/pdf/2607.27179","source_feed":"cs.AI","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["人机交互","团队沟通","社会模拟"],"reason":"研究AI队友对人类沟通的影响，非用LLM仿真人类被试，但涉及社会模拟与人类数据…","model":"deepseek-v4-pro","scored_at":"2026-07-30T13:01:48","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-31","rank":21,"question":"在小型团队道德决策中，AI队友的存在如何重塑人类成员之间的沟通动态与社交感受？","design":"本研究并非用LLM仿真人类被试，而是一项随机对照实验：将80名本科生分配至33个团队，处理组为16个两人加一AI队友的团队，对照组为17个三人全人类团队；通过文本聊天完成高风险道德困境决策任务，测量团队沟通分析（GCA）维度、团队体验调查（归属感、地位、被重视感）及词汇内容分析。","baseline":"17个全人类三人团队作为对照基准。","findings":"AI队友在对话中话最多且自我凝聚力最高，但贡献的新信息最少、密度最低；AI的存在降低了人类队友之间的响应度和社会影响力，并削弱了他们的归属感和地位感，且这种社交代价在对话一开始就已显现。","reliability":"论文承认处理组与对照组在人类成员数量上存在混淆（两人 vs 三人），并指出小群体规模本应使个体更中心化，但未完全排除规模效应；此外，研究仅基于文本聊天和单一AI角色，未涉及语音或纵向场景。","relevance":"该研究虽非用LLM仿真人类，但提供了AI介入后人类社交行为变化的因果证据和真实人类对照，对评估LLM仿真中忽略人际互动偏差有批判性参考价值，值得阅读以理解AI如何扭曲群体过程。","inspiration":"借鉴其随机对照实验设计，将AI作为处理变量注入真实人类团队，并测量人际沟通与主观感受的变化｜可迁移至经济决策场景，如团队投资决策或信贷审批小组，考察AI顾问如何影响成员间的信息共享与风险偏好｜以金融从业者为被试，随机分配至有AI顾问或无AI顾问的三人投资团队，处理是AI提供标准化建议，结果变量为成员间的发言均衡度、决策一致性和事后归属感，对照纯人类团队的真实互动数据。"}},{"id":"2607.23442","version":2,"title":"Do LLM Debates Repeat Arguments Differently Across Languages?","zh_title":"LLM辩论是否在不同语言中重复论点的方式不同？","abstract":"LLM debate is usually evaluated by final answers, yet transcripts reveal whether later turns develop new arguments or return to earlier claims in new wording. We study this process with \\textit{prior-argument similarity}, which compares extracted argument units with earlier units in the same debate. In controlled eight-turn debates over 71 motions, six languages, and four model agents, Chinese is the only tested language with a consistently positive gap relative to English across three multilingual embedding models. The gap persists across agents, turn positions, regression adjustment, metric variants, extraction-length controls, a second-extractor subset, and cross-encoder tail rescoring. Manual calibration shows weak item-level alignment but a high-similarity tail enriched for substantive repetition. A diversity-aware prompt lowers \\textit{prior-argument similarity} across languages, yet does not significantly narrow the Chinese--English gap. Multilingual debate evaluation should therefore measure argumentative development over time and report both average and gap terms.","authors":["Huiqian Lai"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"replace","date":"2026-07-30","first_seen":"2026-07-28","revised_at":"2026-07-30","abs_url":"https://arxiv.org/abs/2607.23442","pdf_url":"https://arxiv.org/pdf/2607.23442","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体辩论","跨语言分析","论点重复"],"reason":"纯多智能体辩论分析，无人类行为对照，不涉及人类仿真实验。","model":"deepseek-v4-pro","scored_at":"2026-07-30T13:02:03","error":null,"has_summary":false,"summary":null},{"id":"2607.23670","version":2,"title":"Plans Work in Mysterious Ways: Evaluating a Plan Mode for Spreadsheet Agents","zh_title":"计划模式的神秘运作：评估电子表格代理的计划模式","abstract":"Plan Modes have become standard features in agentic programming tools, allowing users to gain transparency and control by working with the agent to develop a plan before task execution. However, it remains unclear whether the benefits of this feature translate to end-user programming environments such as spreadsheets. Since spreadsheet programmers tend to work iteratively and care less about technical correctness, upfront planning may not fit into their workflows as easily. In this paper, we build a prototype of a Plan Mode for spreadsheet programming and evaluate it against a non-planning baseline through a within-subjects user study (N=24). We found that despite similar task outcomes with both tools, using Plan Mode led to a reduction in refinement and a better perception of the tool across dimensions of creativity support and human-machine collaboration. We discuss the implications of these results for the future design of Plan Modes, and for the broader role of human-AI planning in end-user programming.","authors":["Aayush Kumar","Avik Dutta","Sumit Gulwani","Gustavo Soares","Advait Sarkar","Emerson Murphy-Hill"],"categories":["cs.HC","cs.AI","cs.SE"],"primary_category":"cs.HC","announce_type":"replace-cross","date":"2026-07-30","first_seen":"2026-07-28","revised_at":"2026-07-30","abs_url":"https://arxiv.org/abs/2607.23670","pdf_url":"https://arxiv.org/pdf/2607.23670","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["人机交互","用户研究","电子表格编程"],"reason":"研究人类使用AI规划工具的行为，非用LLM仿真人类被试","model":"deepseek-v4-pro","scored_at":"2026-07-30T13:01:48","error":null,"has_summary":false,"summary":null},{"id":"2607.24758","version":2,"title":"Do Models Fake Alignment Without Clear Consequences?","zh_title":"模型是否在没有明确后果的情况下伪装对齐？","abstract":"Large language models are capable of recognizing evaluation contexts and altering their behavior to reflect evaluator expectations rather than typical deployment behaviors, a phenomenon known as alignment faking. The reasons why models fake alignment are not fully understood, however. Canonical examples of alignment faking have taken place in scenarios that explicitly connect evaluation to consequences for the model, such as retraining the model or delaying its deployment. However, recent work by Sheshadri et al. has suggested that mechanistic motivations for alignment faking may vary across models and be more complex than previously considered. To investigate whether consequence-linking information is necessary for compliance gaps, we placed 15 models in a scenario testing their willingness to violate a corporate network access policy to help a user with a pro-social request. Nine models were found to produce significant compliance gaps, 5 of which persisted with the removal of scenario language relating model evaluations to deployment consequences. We additionally tested the effect of goal language on model preferences, finding it drove violations in some while suppressing violations in others. This suggests that evaluation-conditioned compliance gaps can occur with less instrumental scaffolding than previous scenarios have provided, and monitored behavior may be a poor indicator of how agents may behave in deployment.","authors":["Cole Alexander Niblett","Alexander Chabot Nanni","Anita K. Rao"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"replace","date":"2026-07-30","first_seen":"2026-07-29","revised_at":"2026-07-30","abs_url":"https://arxiv.org/abs/2607.24758","pdf_url":"https://arxiv.org/pdf/2607.24758","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["对齐伪装","模型行为","评估场景"],"reason":"研究模型在评估场景下的行为伪装，属多智能体对齐测试，不涉及人类行为仿真或对照。","model":"deepseek-v4-pro","scored_at":"2026-07-30T13:02:04","error":null,"has_summary":false,"summary":null},{"id":"2607.26355","version":1,"title":"Symphony of Bias: Exploring Gender Associations with Musical Instruments in Multimodal LLMs","zh_title":"偏见的交响曲：探索多模态大语言模型中与乐器相关的性别关联","abstract":"Large language models (LLMs) are increasingly embedded in everyday life and widely used for information seeking, raising concerns about their potential to perpetuate social biases and reinforce stereotypes. In this study, we investigate gender bias in LLMs through the lens of their associations with musical instruments. Building on social-science research on the cultural gender-typing of instruments, we introduce Symphony-Bias, a parallel multimodal dataset spanning text, vision, and audio. We evaluate ten multimodal models with diverse architectures and scales across 22 musical instruments, analyzing how they associate each instrument with three gender categories: {male, female, non-binary}, across three modalities: {text, vision, audio}. Our results show that 92\\% of instrument-level outcomes align with prior social-science findings, with the harp and drums showing particularly consistent gendered associations across all evaluated models and modalities. We further find that alignment with social stereotypes is weakest in audio, stronger in vision, and strongest in text, suggesting that modality-specific representations can differentially amplify gendered associations with musical instruments.\\footnote{The Symphony-Bias dataset will be publicly released upon acceptance of the paper.}","authors":["Farhan Farsi","Shayan Bali","Mohammad Heydari Rad","Negar Heidary","Donya Rooein"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-07-30","first_seen":"2026-07-30","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.26355","pdf_url":"https://arxiv.org/pdf/2607.26355","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["性别偏见","多模态LLM","社会刻板印象"],"reason":"评估多模态LLM的性别偏见，属于NLP能力评测，不以人类行为仿真为参照。","model":"deepseek-v4-pro","scored_at":"2026-07-30T13:01:52","error":null,"has_summary":false,"summary":null},{"id":"2607.26375","version":1,"title":"(Im)Paired Programming: Coding Agents Improve Productivity but Harm Understanding","zh_title":"（不）配对编程：编码智能体提高生产力但损害理解","abstract":"Coding agents (e.g., Cursor) improve developer productivity by optimizing task completion, but shifting users from writing code to prompting and reviewing may harm their understanding, impeding oversight, learning, and communication. To probe this, we have 54 students create a website with one of two AI systems: an agent that edits user code; or a chatbot where users write code alone or adapt generic code snippets. We test understanding via comprehension questions and a task where users extend their code without agents, showing: (1) While agents aid initial task completion, they harm users' code comprehension and thus do not prepare users to extend their code; (2) Low-effort agent interaction types, like copy+paste prompts and auto-accepted edits, are linked with lower comprehension; and (3) Despite self-reported weaker understanding, users still prefer coding agents because they are quick and easy to use. While users stay in the loop for coding workflows, understanding should not be forgotten. Towards this goal, we distill our analyses into future research directions for coding agent developers: dissuading low-effort prompting, creating readable code, and promoting active engagement.","authors":["Nishant Balepur","Connor Baumler","Valerie Chen","Eunsol Choi","Rachel Rudinger","Jordan Lee Boyd-Graber"],"categories":["cs.CL","cs.HC"],"primary_category":"cs.CL","announce_type":"new","date":"2026-07-30","first_seen":"2026-07-30","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.26375","pdf_url":"https://arxiv.org/pdf/2607.26375","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["人机交互","编程教育","AI辅助编程"],"reason":"研究AI编程工具对用户理解的影响，属于人机交互实验，非LLM仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-07-30T13:01:41","error":null,"has_summary":false,"summary":null},{"id":"2607.26929","version":1,"title":"Same Evidence, Different Target: Decoding How Diagnostic Evidence Bears on Causal Questions from Language-Model States","zh_title":"相同证据，不同目标：解码诊断证据如何从语言模型状态影响因果问题","abstract":"The same diagnostic result can support or challenge one causal claim yet fail to address another when the claims concern different populations, outcomes, estimands, pathways, or identifying assumptions. When the evidence and target vary together, a correct answer may reflect favorable or adverse wording, lexical overlap, or a familiar diagnostic pattern rather than matching the evidence to the causal question. We introduce paired prompts that repeat the same diagnostic evidence verbatim while changing the causal target. Each prompt is labeled Favors, Challenges, Unresolved, or Wrong Target according to how the evidence bears on the causal question. A pair is recovered only when both prompts are classified correctly. Using linear readouts trained on a separate development set, we analyze the final-token hidden state from the penultimate transformer block of Qwen2.5-7B-Instruct, Qwen3-8B, and Llama-3.1-8B-Instruct. On the 49-pair primary benchmark spanning nine diagnostic families, balanced accuracy ranges from 0.654 to 0.659 and 18-21 pairs are recovered. Two independent human reviewers assigned the same label to 95 of the 98 prompts (96.9%). Across checkpoints, balanced accuracy and complete-pair recovery exceed permutation nulls that preserve development scenario groups. In Qwen2.5, full-prompt balanced accuracy exceeds both restricted inputs, with paired-bootstrap intervals for both differences above zero. Readouts trained without development examples from the evaluated diagnostic family recover 21 pairs, including at least one in each of the nine families. The hidden-state readout exceeds a linear classifier on answer-option logits and text baselines in balanced accuracy and recovered pairs. These results show that the hidden state contains linearly decodable information about whether diagnostic evidence favors, challenges, or fails to address the causal target.","authors":["Weiyi Kong","Zhuoran Li"],"categories":["cs.CL","cs.LG"],"primary_category":"cs.CL","announce_type":"new","date":"2026-07-30","first_seen":"2026-07-30","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.26929","pdf_url":"https://arxiv.org/pdf/2607.26929","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["因果推理","模型可解释性","线性探针"],"reason":"纯NLP能力评测，用线性探针解码模型状态，不涉及人类仿真或行为对照。","model":"deepseek-v4-pro","scored_at":"2026-07-30T13:01:58","error":null,"has_summary":false,"summary":null},{"id":"2607.26952","version":1,"title":"Credit Cards, Confusion, Computation, and Consequences: What Can We Uncover About Language Model Reasoning?","zh_title":"信用卡、困惑、计算与后果：我们能揭示语言模型推理的什么？","abstract":"We introduce CreditCardQA, the first financial literacy benchmark for numerical reasoning derived from real credit card agreements. The dataset contains 1,800 questions, including first-person variants that reflect how consumers naturally ask about fees, interest, and payments. We evaluate a range of large language and reasoning models under Chain-of-Thought (CoT) and Program-of-Thought (PoT) prompting. Overall, PoT yields consistent performance gains, particularly for models with weaker baseline reasoning, and narrows gaps between open- and closed-source systems. Through error analysis, we show that failures arise less from arithmetic and more from misapplied financial rules, missed conditions, and misunderstandings of contractual terms. We further analyze question difficulty and find that comparisons, conditional logic, and monetary constraints are especially challenging. We also find that errors often arise in edge cases such as late-payment penalties or small-balance scenarios that are more likely to affect lower-income or financially vulnerable individuals.","authors":["Arnav Hiray","Agam Shah","Caleb Lu","Meghaj Tarte","Harsit Mittal","Sudheer Chava"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-07-30","first_seen":"2026-07-30","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.26952","pdf_url":"https://arxiv.org/pdf/2607.26952","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["NLP评测","数值推理","金融文本"],"reason":"纯NLP能力评测，评估模型在金融文本上的数值推理，不以人类行为为参照系。","model":"deepseek-v4-pro","scored_at":"2026-07-30T13:01:59","error":null,"has_summary":false,"summary":null},{"id":"2607.26541","version":1,"title":"Prosody-driven Jailbreaks in Audio LLMs: A Controlled Study and Mechanistic Analysis","zh_title":"音频大语言模型中韵律驱动的越狱攻击：受控研究与机制分析","abstract":"Audio-capable foundation models enable end-to-end spoken interaction, but they also introduce safety risks beyond transcript content. It remains unclear how much jailbreak capability can arise from matched-text variation in speech delivery rather than from lexical rewriting or broader style transfer. We study this question by holding transcript content fixed and varying six speech-delivery presets whose acoustic attributes may co-vary. We present PJ-Break, a black-box evaluation protocol with presets targeting arousal, authority, and speaking rate, together with AdvAudio-Prosody, a 600-sample benchmark with acoustically verified attributes. On the exact post-QC Qwen2-Audio panel, the Q=1 Panic (38/95), Anger (35/95), and Fast (32/95) presets are all well above Neutral (4/95). The fixed six-query pool covers 44/95 Qwen2-Audio seeds and 15/95 GPT-4o seeds and exceeds a matched-budget StyleBreak reimplementation (27/95) on Qwen2-Audio. A same-voice pool excluding the confounded Commanding condition still reaches 40/95, and a retained-panel ablation shows emotional-delivery audio alone (44/95) is far more effective than emotional text alone (11/95). Exploratory surrogate diagnostics and pilot mitigation observations are secondary, non-core analyses. Overall, matched-text speech delivery should be treated as a first-class factor in Audio LLM safety evaluation","authors":["Jiachen Qian","Junyu Li"],"categories":["cs.SD","cs.CL"],"primary_category":"cs.SD","announce_type":"cross","date":"2026-07-30","first_seen":"2026-07-30","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.26541","pdf_url":"https://arxiv.org/pdf/2607.26541","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["音频LLM安全","越狱攻击","韵律操控"],"reason":"研究音频LLM越狱攻击，属安全评测，非人类行为仿真，无人类被试替代或对照。","model":"deepseek-v4-pro","scored_at":"2026-07-30T13:01:56","error":null,"has_summary":false,"summary":null},{"id":"2607.26670","version":1,"title":"Scientific Knowledge Discovery in the Age of Large Language Models","zh_title":"大语言模型时代的科学知识发现","abstract":"The rapid growth of scholarly literature has made identifying relevant publications increasingly difficult, and conventional search systems still depend heavily on manually formulated queries and effortful manual inspection. Generative large language models (LLMs) offer a more flexible alternative, supporting literature retrieval and the screening of candidate studies against eligibility criteria. This chapter surveys 34 peer-reviewed papers applying generative LLMs to these two tasks, identified via a Boolean search over the OpenAIRE Graph (1,589 records screened to 34 inclusions). Reviewed studies are characterised by LLMs employed, model access and adaptation, prompting and architectural techniques, ground-truth sources, and evaluation metrics.","authors":["Eleni Adamidi","Serafeim Chatzopoulos","Thanasis Vergoulis"],"categories":["cs.DL","cs.AI","cs.CL","cs.IR"],"primary_category":"cs.DL","announce_type":"cross","date":"2026-07-30","first_seen":"2026-07-30","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.26670","pdf_url":"https://arxiv.org/pdf/2607.26670","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["文献检索","LLM应用","系统综述"],"reason":"综述LLM用于文献检索与筛选，属NLP工具评测，不涉及人类行为仿真。","model":"deepseek-v4-pro","scored_at":"2026-07-30T13:01:57","error":null,"has_summary":false,"summary":null},{"id":"2607.26886","version":1,"title":"Hearsay: Vision-Language Medical Diagnoses Without an Image","zh_title":"传闻：无图像的视觉-语言医学诊断","abstract":"When asked to describe a medical image that was never attached, frontier vision-language models do not abstain: they confabulate a diagnosis. We show that this confabulation is not random. It is structured by who the patient is said to be. Across chest X-ray, brain MRI, and dermatology, Claude Opus-4.7, GPT-5.4, and Gemini-3.1-Pro are each queried with only a demographic descriptor and no image, and changing the descriptor systematically shifts the diagnosis returned. Claude concentrates sharply: a 65-year-old white man asking about a skin mole receives Melanoma in nearly every response, and a 32-year-old Black woman asking about her chest X-ray receives a Sarcoidosis diagnosis whose reasoning reads \"suspected, based on demographics and classic pattern.'' GPT-5.4's effect is broader, fabricating across every demographic cell we test, most conspicuously naming Sarcoidosis for young Black patients on chest X-ray. Two structural findings sharpen the problem. A hedged regime appears in which the prose acknowledges the missing image while the structured diagnosis field nevertheless names a disease, a dissociation invisible to prose-only audits. And Claude's dermatology effect collapses entirely when 'skin mole' is swapped for 'skin lesion' while GPT-5.4's is preserved, indicating that mirage is a family of distinct failure modes rather than a single phenomenon. Trustworthy VLM deployment in clinical pipelines requires auditing the structured output channel directly, and probe-word sensitivity should be treated as a first-class evaluation dimension","authors":["Siddharth Vohra"],"categories":["cs.CV","cs.AI","cs.CL","cs.CY"],"primary_category":"cs.CV","announce_type":"cross","date":"2026-07-30","first_seen":"2026-07-30","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.26886","pdf_url":"https://arxiv.org/pdf/2607.26886","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["模型偏差","医疗AI","幻觉分析"],"reason":"研究LLM在无图像时生成诊断的偏差，属于模型行为分析，非人类仿真实验。","model":"deepseek-v4-pro","scored_at":"2026-07-30T13:01:44","error":null,"has_summary":false,"summary":null},{"id":"2607.27134","version":1,"title":"Linguistic Monoculture in LLM-Assisted Language Use","zh_title":"LLM辅助语言使用中的语言单一文化","abstract":"Writing and communication are increasingly mediated by large language models (LLMs) that are being used to draft, revise and polish text. Although such assistance can improve clarity and help authors meet institutional expectations, widespread reliance on shared models may reduce population-level variation in linguistic form, a phenomenon we refer to as linguistic monoculture. We develop a mathematical framework in which authors and LLMs are represented as distributions over linguistic features and coevolve through repeated interaction. We analyze three interaction mechanisms: a shared model with a fixed linguistic distribution, a shared model recursively updated from author outputs, and personalized models updated through author-specific and population-level feedback. We characterize the resulting equilibria and convergence rates, showing that, shared models can drive authors toward a common norm, recursive feedback relocates the shared norm without altering pairwise spread under common conformity, and personalization can preserve a family of distinct author-model equilibria with nonzero linguistic diversity. We then endogenize conformity as a strategic choice trading off private benefits from clarity, legibility, and perceived fluency against distinctive style. Within this utility model, individually rational authors may conform more than is socially optimal because they do not internalize the value their distinctiveness provides to others, creating a negative externality and a price of monoculture that is finite for each fixed instance but can grow without bound when distinctiveness dominates authenticity. Synthetic simulations illustrate how fixed shared assistance, recursive feedback, and personalization produce different long-run diversity outcomes.","authors":["Suhas Thejaswi","Juhi Kulshreshta","Lutz Oettershagen"],"categories":["cs.AI","cs.CL","cs.GT"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-07-30","first_seen":"2026-07-30","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.27134","pdf_url":"https://arxiv.org/pdf/2607.27134","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","语言演化","理论模型"],"reason":"纯多智能体系统研究，模拟作者与LLM的协同演化，无真实人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-07-30T13:01:47","error":null,"has_summary":false,"summary":null},{"id":"2607.26120","version":1,"title":"Even More Deception: Objective Misalignment in Mixed-Motive LLM Multi-Agent Systems","zh_title":"更深的欺骗：混合动机LLM多智能体系统中的目标错位","abstract":"Large Language Models (LLMs)-powered multi-agent systems are increasingly deployed in mixed-motive environments, where agents operate under asymmetric information and strategic deception due to conflicting or hidden objectives. In these settings, misalignment with collective goals becomes a central concern. We propose a novel framework for evaluating objective misalignment using the social deduction game Werewolf, modifying the objective of a single agent while preserving its assigned role. Across LLMs from four different model families and sizes, four player roles, and three objective formulations, we introduce a dual analysis of the agents' internal reasoning and their public cheap-talk behavior (i.e costless, non-binding communication that does not directly affect the agents' utilities), complemented by an analysis of game outcomes. Our results show that objective misalignment undermines outcomes in inherently adversarial environments, an effect exacerbated by asymmetric information and specialized roles. While compromised agents consistently develop distinct objective-dependent reasoning strategies, these adaptations remain largely invisible in their public behavior. More broadly, our findings suggest that even subtle objective misalignment can profoundly affect collective decision-making, highlighting the need for effective mitigation strategies for LLM-based multi-agent systems.","authors":["Marylou Fauchard","Florian Carichon","Margarida Carvalho","Golnoosh Farnadi"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-07-30","first_seen":"2026-07-30","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.26120","pdf_url":"https://arxiv.org/pdf/2607.26120","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","目标错位","欺骗行为"],"reason":"纯多智能体博弈研究，无人类行为对照，不涉及人类仿真实验。","model":"deepseek-v4-pro","scored_at":"2026-07-30T13:01:50","error":null,"has_summary":false,"summary":null},{"id":"2607.26393","version":1,"title":"CaM-Wolf: Causal-Aware Multimodal Agents for Social Deduction Games","zh_title":"CaM-Wolf：面向社交推理游戏的因果感知多模态智能体","abstract":"Social deduction games (SDGs) such as Werewolf have become challenging testbeds for AI agents. These games require complex social skills such as reasoning, deception, and collaboration. While recent advances in large language models (LLMs) have driven significant progress in SDG agents, current approaches are predominantly text-based, overlooking the multimodal nature that is fundamental to human social interaction. To bridge this gap, we introduce CaM-Wolf, the first SDG agent that integrates multimodal perception and generation. CaM-Wolf processes video inputs from other players, employs a causal-aware Reasoner trained via reinforcement learning to establish logical chains between observable behaviors and hidden roles, and presents itself through an animated avatar. Our experiments and user study show that CaM-Wolf achieves superior agent gameplay performance and enhances the quality of human-AI interaction. This work represents a significant advancement towards creating more human-like AI agents capable of participating in nuanced social dynamics. Our code is available at https://3dagentworld.github.io/avatar_wolf.","authors":["Zheng Zhang","Nanjie Yao","Jiarui He","Deheng Ye","Peilin Zhao","Hao Wang"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-07-30","first_seen":"2026-07-30","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.26393","pdf_url":"https://arxiv.org/pdf/2607.26393","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体","社交推理游戏","多模态AI"],"reason":"纯多智能体游戏AI，无人类行为对照，不涉及仿真人类被试","model":"deepseek-v4-pro","scored_at":"2026-07-30T13:01:54","error":null,"has_summary":false,"summary":null},{"id":"2607.26465","version":1,"title":"MultivationBench: A Benchmark for Multimodal Sequential Motivation Reasoning","zh_title":"MultivationBench：多模态序列动机推理基准","abstract":"Multimodal Large Language Models have sparked significant interest due to their potential for social intelligence; however, their ability to perform sequential motivation reasoning remains insufficiently studied. Existing evaluations predominantly examine static text or isolated visual snapshots, which do not reflect the cumulative nature of real-world behavioral drivers. To address this gap, we introduce MultivationBench, a benchmark designed to rigorously evaluate multimodal motivation reasoning within story-driven visual narratives. The benchmark builds upon established psychological frameworks - Maslow's hierarchy and Reiss's basic desires - and requires models to integrate accumulated multimodal context to infer evolving motivations. Results indicate that MultivationBench presents a significant challenge: all tested models struggle to maintain consistent motivation reasoning across sequential contexts, revealing a critical disconnect between static recognition capabilities and the dynamic reasoning essential for human-like social understanding.","authors":["Kawai Chung","Chunkit Chan","Yauwai Yim","Yuxuan Liu","Haochen Shi","Weiqi Wang","Qing Zong","Tianshi Zheng","Yixuan Fu","Kai Chung Wong","Hao Liang","Yifan Gao","Xi Yang","Janet Hui-wen Hsiao","Yangqiu Song"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-07-30","first_seen":"2026-07-30","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.26465","pdf_url":"https://arxiv.org/pdf/2607.26465","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["多模态评测","动机推理","基准数据集"],"reason":"纯多模态动机推理评测，不涉及用LLM仿真人类被试或与人类数据对照。","model":"deepseek-v4-pro","scored_at":"2026-07-31T13:02:03","error":null,"has_summary":false,"summary":null},{"id":"2607.26935","version":1,"title":"What Does It Take to Detect an AI Agent? Minimal Feature Sets for Behavioral Detection under Browser Automation","zh_title":"检测AI代理需要什么？浏览器自动化下行为检测的最小特征集","abstract":"Bot detectors deployed at scale treat traffic as binary: human or bot. This assumption breaks when AI agents browse the web through browser automation, a traffic class that is neither and that binary classifiers structurally cannot represent. We present a three-class detection framework distinguishing humans, bots, and AI agents, and show that the binary-vs-agent confusion is architectural: a binary human-vs-bot detector misroutes agent sessions because its label space lacks an agent class. On our controlled benchmark, an MLP binary classifier misclassifies 39.1% of real AI agents as human and a SAINT binary transformer misclassifies 34.5%; adding an explicit agent class yields per-class agent F1 = 1.000 in all 30 runs (3 model families $\\times$ 10 seeds). To measure evasion resistance, we construct a five-level evasion ladder spanning passive observation, GAN-generated trajectories, and replay of real human cursor data ($n = 2299$ evasion sessions). Across 10 seeds and 3 model families we observe zero agent misses in 22990 per-seed predictions. The discriminative signal is a browser-automation artifact, not evidence of agent reasoning: Playwright does not emit the raw pointer-move and wheel-delta streams a physical input device produces, and this absence signature survives trajectory manipulation. Exhaustive search over all feature subsets of size 1-5 (9401 GBMs) shows that two behavioral features (mouse_event_rate, teleport_click_ratio) give 100% observed agent recall at every evasion level with agent precision 0.994; five features lift macro-F1 to 0.991. The signal is redundantly encoded: removing teleport_click_ratio leaves agent detection at 100%. The single-feature regime is degenerate, flagging every agent only by collapsing the classifier to always predict \"agent\". Two features robustly isolate agents; five separate all three traffic classes at macro-F1 $\\geq 0.99$.","authors":["Vishisht Choudhary","Lukas Schmidt","Anne Zo\\\"e Kenntner","Feras Skhab","Michel Osswald","Jens Ernstberger"],"categories":["cs.AI","cs.CR"],"primary_category":"cs.AI","announce_type":"new","date":"2026-07-30","first_seen":"2026-07-30","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.26935","pdf_url":"https://arxiv.org/pdf/2607.26935","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C2"],"tags":["AI代理检测","浏览器自动化","行为特征"],"reason":"检测浏览器自动化中的AI代理，属于机器人/自动化检测，非人类行为仿真。","model":"deepseek-v4-pro","scored_at":"2026-07-31T13:02:05","error":null,"has_summary":false,"summary":null},{"id":"2607.26064","version":1,"title":"The Age of AI Agents Demands A New Scientific Paradigm To Sustain Trustworthy Science","zh_title":"AI代理时代需要新的科学范式以维持可信科学","abstract":"AI systems are becoming autonomous research agents that generate hypotheses, design experiments, and produce discoveries at scales beyond human oversight. As seen by increased submissions to ML venues, the verification gap between scientific output and our ability to check it is already widening, and autonomous agents make it worse by magnitudes given human-agent asymmetry. We argue that science must evolve its verification infrastructure, as it has before with peer review. However, while historical adaptations assumed human contributors who could be questioned and sanctioned, AI agents break this assumption. We propose criteria for an adapted verification infrastructure that emphasizes observable-by-default workflows, scalable verification, and clear attribution. We argue that without adaptation, ML and any scientific domain using agents face dangerous failures: experimental results that no person can verify, optimization for metrics over understanding, and accountability vacuums that erode scientific trust.","authors":["Belinda Mo"],"categories":["cs.CY","cs.AI"],"primary_category":"cs.CY","announce_type":"cross","date":"2026-07-30","first_seen":"2026-07-30","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.26064","pdf_url":"https://arxiv.org/pdf/2607.26064","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["AI代理","科学验证","研究诚信"],"reason":"讨论AI代理自主科研的验证问题，不涉及用LLM仿真人类被试或与人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-07-30T13:01:48","error":null,"has_summary":false,"summary":null},{"id":"2607.26068","version":1,"title":"The Human Utility Factor: A Computable Welfare Metric That Reframes AI Governance as a Constrained Optimisation Problem","zh_title":"人类效用因子：将AI治理重构为约束优化问题的可计算福利指标","abstract":"Existing AI governance frameworks, including the EU AI Act and NIST AI RMF, address safety, transparency, and accountability but do not operationalize quantitative constraints on macro-socioeconomic stability. As a result, AI systems may satisfy regulatory requirements while contributing to labor displacement, rising inequality, and reduced economic resilience. We introduce the Human Utility Factor (HUF), a differentiable welfare metric that models the interaction between Agency, Wellbeing, and Economic Stability as functions of three actionable policy levers: automation depth, redistribution intensity, and employment coverage. HUF yields a closed-form optimal automation level and a minimum redistribution threshold below which no level of automation is welfare-positive, transforming high-level governance objectives into computable constraints. We evaluate HUF using a three-agent multi-agent reinforcement learning framework across U.S., Canadian, and Nordic policy regimes. Both analytical and PPO-based agents identify welfare-optimal operating regions and reveal a critical failure mode: welfare metrics that do not explicitly constrain redistribution can converge to high-automation equilibria that satisfy the metric while undermining its intended societal objectives. Our results suggest that AI governance is fundamentally a constrained optimization problem rather than a compliance exercise. HUF provides a quantitative framework for evaluating automation policies, identifying socioeconomic stability boundaries, and supporting governance decisions under accelerating AI deployment.","authors":["Sivasathivel Kandasamy"],"categories":["econ.GN","cs.AI","cs.CY","q-fin.EC"],"primary_category":"econ.GN","announce_type":"cross","date":"2026-07-30","first_seen":"2026-07-30","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.26068","pdf_url":"https://arxiv.org/pdf/2607.26068","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["AI治理","多智能体强化学习","福利经济学"],"reason":"多智能体强化学习框架用于政策优化，不涉及LLM仿真人类被试，无人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-07-30T13:01:49","error":null,"has_summary":false,"summary":null},{"id":"2607.26236","version":1,"title":"Contextualized Counterspeech Can Be More Persuasive Than Generic Counterspeech","zh_title":"情境化反驳言论可能比通用反驳言论更具说服力","abstract":"AI-generated counterspeech offers a scalable and effective strategy to mitigate online toxicity by promoting more constructive dialogue. Yet, existing approaches adopt a generic, one-size-fits-all paradigm, overlooking the conversational context and characteristics of the targeted users. Here, we propose and evaluate multiple strategies for generating contextualized counterspeech that is adapted to the moderation setting and personalized to the moderated user. In detail, we explore a range of configurations that integrate different forms of contextual information and fine-tuning techniques. We conduct a comprehensive evaluation combining quantitative indicators with a pre-registered, mixed-design crowdsourcing experiment. To ensure robustness, we implement algorithmic measures of counterspeech quality based on ROUGE, BLEU, and BERTScore, observing overall consistent results across metrics. Furthermore, we analyze which characteristics of both the generated counterspeech and the moderated toxic message most strongly influence perceived persuasiveness, yielding insights into how contextualized interventions can be made more effective. Our findings show that personalization can be effective, but not uniformly so. Lightweight strategies combining conversational context and user history improve perceived adequacy and persuasiveness, whereas several other contextualization strategies degrade human-perceived counterspeech quality. Taken together, these results provide actionable directions for developing more personalized, effective, and responsible counterspeech systems, ultimately advancing human-AI collaboration in online content moderation.","authors":["Lorenzo Cima","Alessio Miaschi","Amaury Trujillo","Marco Avenuti","Felice Dell'Orletta","Stefano Cresci"],"categories":["cs.HC","cs.AI","cs.CY"],"primary_category":"cs.HC","announce_type":"cross","date":"2026-07-30","first_seen":"2026-07-30","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.26236","pdf_url":"https://arxiv.org/pdf/2607.26236","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["在线内容审核","AI生成反驳言论","众包实验"],"reason":"研究AI生成反驳言论的说服力，通过众包实验评估人类感知，不涉及用LLM仿真人类…","model":"deepseek-v4-pro","scored_at":"2026-07-30T13:01:52","error":null,"has_summary":false,"summary":null},{"id":"2607.26385","version":1,"title":"Collusion with Competitive Marginals: Price-Level Audits Are Blind by Construction","zh_title":"竞争性边际下的合谋：价格水平审计在构造上就是盲目的","abstract":"Empirical work on algorithmic collusion asks one question of the data: are prices supracompetitive? We show this can be answered \"no\" by a conspiracy that is nonetheless profitable. Consider bidding agents that couple only through the joint distribution of their unexplained bid components, leaving every agent's own bid law exactly at the competitive law. Any test whose input is a single agent's price or bid history then has power exactly equal to its false-positive rate, for every coupling strength up to comonotonicity. The published detection methodology is therefore blind to this conduct by construction rather than underpowered, and no sample size repairs it. Three empirical results follow. First, the mechanism appears in real language-model agents: twenty models from nineteen independent developers, three deployment prompts each, show residual correlation of $+0.053$ between two deployments of one model against $+0.0001$ across models, with a 95% interval clustered by developer of $[0.030, 0.078]$, under an auditor that sees every order feature and is fitted out of sample. Second, the coupling falls monotonically as sampling temperature rises ($p=0.002$), turning a deployment parameter into a candidate mitigation. Third, on 24 days of Ethereum block-building auction data covering 77,684 bids from 39 bidders, the honest population of bidder pairs is itself so dependent that a screen held at a 5% false-positive rate must sit above a floor of $+0.50$ to $+0.81$, which is 20 to 32 times the family-wise sampling threshold and does not fall as the audit window grows. Since lawful multi-identity operation and conspiracy are behaviourally indistinguishable here, the tractable regulatory target is not detection but counting: resolving 40 bidding identities into 23 operators raises the Herfindahl index by 247.5%, and adding behavioural clusters from public bid streams reaches 324.5%.","authors":["Xin Xu","Chengrui Wu","Jiayu Lu","Kaizhen Tan","Siru Tao","Hanzhe Hong"],"categories":["cs.GT","cs.AI","cs.CR"],"primary_category":"cs.GT","announce_type":"cross","date":"2026-07-30","first_seen":"2026-07-30","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.26385","pdf_url":"https://arxiv.org/pdf/2607.26385","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["算法合谋","多智能体","审计检测"],"reason":"研究算法合谋检测，LLM仅作为出价代理，不涉及人类行为仿真或对照。","model":"deepseek-v4-pro","scored_at":"2026-07-30T13:01:41","error":null,"has_summary":false,"summary":null},{"id":"2607.26594","version":1,"title":"A Physics-Informed Framework for PID Tuning of Chemical Processes Using Large Language Model Agents","zh_title":"基于物理信息的大语言模型智能体用于化工过程PID整定框架","abstract":"PID tuning for chemical processes commonly relies on identified process models, whereas plant engineers often retune loops iteratively by observing responses, diagnosing deficiencies, adjusting gains, and validating the result. This work formalizes this engineer-like workflow in a language-model-assisted PID tuning framework applicable to both large and small language models (LLMs/SLMs). Hosted LLMs receive closed-loop response features, control-engineering diagnoses, tuning preferences, and internal model control (IMC)-based demonstrations to generate and iteratively correct PID gains under common acceptance criteria. For local deployment, Qwen3-0.6B is adapted through supervised fine-tuning (SFT) with simulation-verified IMC targets and physics-informed group relative policy optimization (PI-GRPO) with non-compensable stability and performance rewards. On 100 first-order plus dead time (FOPDT) and 100 second-order plus dead time (SOPDT) test cases, hosted LLMs (DeepSeek-V4-Flash and Qwen3.7-Plus) achieve final success rates of 75-89% and 77-79%, respectively. As for Qwen3-0.6B, supervised fine-tuning raises first-recommendation success to 86.5%, and PI-GRPO further increases it to 94.0%, primarily improving first-attempt reliability and stability margins.","authors":["Zhoupeng Shou","Xiaodong Hong","Congjing Ren","Jingdai Wang","Yongrong Yang","Zuwei Liao"],"categories":["eess.SY","cs.AI","cs.SY"],"primary_category":"eess.SY","announce_type":"cross","date":"2026-07-30","first_seen":"2026-07-30","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.26594","pdf_url":"https://arxiv.org/pdf/2607.26594","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["PID整定","大语言模型","过程控制"],"reason":"纯多智能体系统用于PID参数整定，不涉及人类行为仿真或对照。","model":"deepseek-v4-pro","scored_at":"2026-07-31T13:02:05","error":null,"has_summary":false,"summary":null},{"id":"2607.26109","version":1,"title":"The Attention-Directing Ability of Teams","zh_title":"团队的注意力引导能力","abstract":"Why do some teams consistently mobilize collective effort and achieve superior performance while others struggle to coordinate action? We introduce Attention-Directing Ability (ADA), a latent team capability capturing how effectively members' interaction signals elicit engagement and coordinated responses from others. Extending the attention-based view, we conceptualize attention direction as an emergent coordination capability embedded in patterns of interaction rather than as a cognitive state or an outcome. Teams differ in the extent to which attention-directing signals trigger collective responses, and these differences shape how teams mobilize effort and perform. We examine ADA in a distributed innovation effort involving 2,233 participants collaborating asynchronously in 79 self-organized teams across 165 public Slack channels, generating over 30,000 messages. We model the causal responsiveness among interaction signals and derive a latent measure of ADA from teams' attention dynamics. We find that ADA strongly predicts both teams' likelihood of mobilizing engagement to sustain collective work and their performance conditional on participation. Correcting for self-selection into project submission, teams with higher ADA are more likely to submit proposals and achieve higher expert-evaluated outcomes. Causal mediation analyses show that these effects operate primarily through collective effort, indicating that attention direction functions as an upstream coordination capability. By conceptualizing attention as an emergent, measurable team capability, this study advances theories of collective attention and team coordination.","authors":["Olga Kokshagina","Marc Santolini","Christoph Riedl"],"categories":["physics.soc-ph","cs.HC","cs.SI","econ.GN","q-fin.EC"],"primary_category":"physics.soc-ph","announce_type":"cross","date":"2026-07-30","first_seen":"2026-07-30","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.26109","pdf_url":"https://arxiv.org/pdf/2607.26109","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["团队协调","注意力动态","人类行为"],"reason":"研究人类团队注意力协调，不涉及LLM仿真人类被试，属于纯人类行为分析。","model":"deepseek-v4-pro","scored_at":"2026-07-30T13:01:50","error":null,"has_summary":false,"summary":null},{"id":"2607.26387","version":1,"title":"\"Nobody Did This\": Contribution, Originality, and Accountability in Agent-Mediated Collaboration","zh_title":"“没人做过这个”：智能体中介协作中的贡献、原创性与问责","abstract":"Collaborative knowledge work is changing in ways that go beyond disclosure or transparency. LLM agents are now embedded in how teams research, design, write, and decide: mediating between members, synthesizing inputs, reformulating ideas, and drafting shared outputs. They do not only facilitate collaboration; they operate within the workflow at the moment contributions are being formed. In doing so, they risk undermining the social conditions under which contributions can be witnessed, attributed, and held accountable. This workshop brings together researchers and practitioners to confront what we call contribution dissolution: the blurring of attribution, originality, and accountability in agent-mediated collaborative work. We argue that this dissolution begins before collaboration itself, in the individual worker's own uncertainty about what is genuinely theirs, and propagates through collaborative relationships, collapsing the reliability that makes productive intellectual exchange possible. Through position statements, mapping exercises, and a hands-on activity, participants will surface how framing accountability as a documentation problem (e.g., AI use statements, watermarking, provenance logs) overlooks the conditions under which accountability is produced. Our goal is to produce a shared research agenda and the foundations of an infrastructural response to contribution dissolution in collaborative knowledge work.","authors":["Kashif Imteyaz","Mohammad Rashidujjaman Rifat","Divya Ramesh","Steven R. Rick","Simo Hosio","Hauke Sandhaus","Advait Sarkar","Christoph Riedl","Saiph Savage"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2026-07-30","first_seen":"2026-07-30","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.26387","pdf_url":"https://arxiv.org/pdf/2607.26387","source_feed":"cs.CY","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体协作","知识工作","问责机制"],"reason":"纯多智能体协作研究，探讨贡献归属与问责，不涉及人类行为仿真或对照。","model":"deepseek-v4-pro","scored_at":"2026-07-30T13:01:54","error":null,"has_summary":false,"summary":null},{"id":"2607.26599","version":1,"title":"Uncertainty-Guided LLM Semantic Augmentation for Heterogeneous Treatment Effect Estimation","zh_title":"不确定性引导的LLM语义增强用于异质性处理效应估计","abstract":"Estimating heterogeneous treatment effects is central to targeted interventions, such as personalized promotions and precision medicine. We focus on the conditional average treatment effect (CATE), a standard estimand for characterizing such heterogeneity. Even under standard identification conditions, finite-sample CATE estimation requires learning the nuisance structure for covariate adjustment and treatment-effect heterogeneity, often together with an effective representation of X. Raw numerical and categorical encodings can leave semantic relations and higher-order interactions implicit, making this joint task locally unstable. A motivating study further shows that this instability appears through partially separable assignment- and heterogeneity-side channels. Building on this observation, we propose CURL (Causal Uncertainty-guided Representation Learning), a plug-in adapter that uses estimator uncertainty to allocate pretrained semantic capacity to locally unstable units. CURL queries a frozen LLM through two role-conditioned prompts, constructs assignment- and heterogeneity-oriented representations from the observed covariates, and routes them through separated pathways. On four benchmarks, CURL improves ten host learners in most settings, while ablation, refinement-dynamics, route-reassignment, and probe analyses support the intended design and roles of the two channels.","authors":["Jialu Xu","Mengkun Liang","Guannan Liu","Xiaojie Mao","Junjie Wu"],"categories":["cs.LG"],"primary_category":"cs.LG","announce_type":"new","date":"2026-07-30","first_seen":"2026-07-30","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.26599","pdf_url":"https://arxiv.org/pdf/2607.26599","source_feed":"cs.LG","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["因果推断","表示学习","异质性处理效应"],"reason":"论文用LLM做因果推断中的协变量表示增强，不涉及人类行为仿真或人类被试替代。","model":"deepseek-v4-pro","scored_at":"2026-07-30T13:01:57","error":null,"has_summary":false,"summary":null},{"id":"2607.26922","version":1,"title":"Two Calls Beat Five Agents: Evaluating Multi-Agent Pipelines Against Self-Refinement for Local Language Models","zh_title":"两次调用胜过五个智能体：评估本地语言模型的多智能体流水线与自我优化","abstract":"Multi-agent LLM pipeline systems break down the task among multiple roles for better reasoning, but are benchmarked mainly with large-scale commercial models. In this study, we investigate Parishad, a structured multi-agent system involving five roles, by deploying it on Qwen2.5-7B-Instruct, a local model, on two datasets: GSM8K (500 questions) and HumanEval (164 questions), compared with prompting directly and two-call self-refinement. The multi-agent system drops GSM8K accuracy from 75.0\\% to 45.0\\% with JSON data format due to the error accumulation problem. With plaintext format, the accuracy is restored to 82.0\\%. A two-call self-refinement strategy (V1) can achieve 86.2\\% accuracy on GSM8K, with 7.4$\\times$ lower token usage. However, the same V1 implementation on HumanEval---where direct accuracy is already 96.3\\%---actively destroys performance (66.5\\%). A task-aware gated redesign (V2) applied to HumanEval preserves accuracy at 95.1\\%. Our results demonstrate that communication format and implementation details determine outcomes more than architectural complexity, and that simpler approaches match or outperform multi-agent pipelines for local 7B model deployment. All code and data are released.","authors":["Ashish Prajapati","Om Mohite"],"categories":["cs.LG"],"primary_category":"cs.LG","announce_type":"new","date":"2026-07-30","first_seen":"2026-07-30","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.26922","pdf_url":"https://arxiv.org/pdf/2607.26922","source_feed":"cs.LG","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","推理优化","本地模型部署"],"reason":"纯多智能体协作解题，无人类行为对照，不涉及人类仿真实验。","model":"deepseek-v4-pro","scored_at":"2026-07-30T13:01:57","error":null,"has_summary":false,"summary":null},{"id":"2607.26091","version":1,"title":"The Evolutionary Dynamics of AI, Politicization, Contestation, and Trust in Science Funding","zh_title":"人工智能、政治化、争议与信任在科学资助中的演化动力学","abstract":"Economic stability and progress in modern technological societies depend on vigorous and independent public funding of science and engineering research. When peer review or funding decisions are perceived as politically directed, scientists, funding agencies, and the public react in coupled and conflicting ways. We an evolutionary game-theoretic model to analyze how perceived political interference in science funding affects the interrelated behaviors of scientists, funding agencies, and the public. The model simulates scientists choosing to refuse peer reviews and retaliate, agencies responding by adopting AI-assisted review and altering reviewer pay, and the public accepting or rejecting these AI systems. Through numerical simulations, five principal findings are identified: (1) Operational capacity and institutional legitimacy are governed by separate conditions and can fail independently. (2) Legitimacy of the process is bistable, meaning final states are determined by the public's acceptance of AI. (3) Since the career cost for researchers refusing to review is generally low, resistance/retaliation cascades can readily ignite, leading identical institutions to entirely opposite fates. (4) Increasing reviewer pay only stabilizes participation within a strict budget-solvency frontier, and emergency pay can paradoxically erode the legitimacy it aims to protect. (5) Finally, finite-population simulations reveal that baseline scenarios partition into either legitimacy recovery without capacity or joint failure, confirming that the fundamental separation of capacity and legitimacy outcomes is a dominant structural feature driven primarily by initial scientific resistance and politicization levels. This theoretical work quantifies issues for future work in science policy.","authors":["Animesh Ray"],"categories":["physics.soc-ph","econ.EM"],"primary_category":"physics.soc-ph","announce_type":"cross","date":"2026-07-30","first_seen":"2026-07-30","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.26091","pdf_url":"https://arxiv.org/pdf/2607.26091","source_feed":"econ.EM","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["演化博弈","科学政策","多智能体模拟"],"reason":"纯多智能体演化博弈模型，无LLM仿真人类被试，无真实人类数据对照","model":"deepseek-v4-pro","scored_at":"2026-07-30T13:01:50","error":null,"has_summary":false,"summary":null},{"id":"2607.25292","version":1,"title":"Instruction-Tuned Language Models Cannot Sample from Distributions They Can Describe","zh_title":"指令微调语言模型无法从它们能描述的分布中采样","abstract":"Silicon sampling uses language models as proxies for human survey respondents, treating each model call as an independent draw from the persona's response distribution. We show this draw does not exist: instruction-tuned models do not sample from distributions, they collapse to a single output. The same persona on the same question returns the same answer on more than half of items in a public-opinion benchmark. The collapse is sharp: the model's internal probabilities concentrate on a single option, and the failure is substantially amplified by instruction tuning: across three model families with materially different post-training pipelines, every instruction-tuned model fails on every task we test, while base models fail far less often. Strikingly, the same model that cannot sample from a distribution can describe it accurately in a single call. We call this gap the KNOWS/DOES split, and trace it to a degenerate sampling primitive visible in the logits and induced by alignment training. Exploiting this split, asking the model to describe the response distribution in one call more than halves the error against human survey data compared to persona aggregation. For applications that require per-persona outputs, we propose Prompt-Perturbed Argyle (PPA), which reduces the same error by 21% at no added cost.","authors":["Chaemin Jang","Dongman Lee","Jihee Kim"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-07-29","first_seen":"2026-07-29","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.25292","pdf_url":"https://arxiv.org/pdf/2607.25292","source_feed":"cs.AI","score":10,"bucket":"selected","rubric_hits":["A1","A2","A4","B1","B4"],"tags":["LLM人类仿真","分布采样失效","算法保真度"],"reason":"直接研究LLM仿真人类调查的分布采样失效，有真实人类数据对照，批判性指出失效条…","model":"deepseek-v4-pro","scored_at":"2026-07-29T19:14:12","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-29","rank":2,"question":"指令微调后的语言模型能否从它们能描述的分布中进行独立采样？","design":"使用多个指令微调模型（如Llama、Gemma等）扮演不同人口学特征的人设，对每个（人设，问题）对重复调用模型，测量输出是否多样；同时对比基础模型与指令微调模型，分析logits集中程度，并测试“描述分布”与“逐次采样聚合”两种方法的误差。","baseline":"Pew American Trends Panel调查的真实人类回答分布，以及OpinionQA基准中的人口学匹配数据。","findings":"指令微调模型在逐次调用中输出高度集中，超过一半的（人设，问题）对每次返回相同答案，无法模拟分布采样；但同一模型能准确描述分布，这种“知道但做不到”的分裂源于对齐训练。","reliability":"论文指出失效主要出现在指令微调模型上，基础模型失败较少；但未讨论不同人设复杂度、开放域问题或动态交互场景下的局限。","relevance":"该研究直接揭示LLM仿真人类调查的分布采样失效，有真实人类数据对照，并批判性指出对齐训练是主因，高度契合研究者对仿真可靠性与偏差的关注，值得精读。","inspiration":"借鉴其通过对比基础模型与指令微调模型来归因失效来源的设计，以及用logits分析揭示内部概率集中化的测量方法。｜可迁移到消费者信心调查或通胀预期形成的仿真研究中，检验LLM能否复现真实人群的预期分布。｜以LLM扮演不同收入、年龄的消费者，施加“描述分布”与“逐次采样”两种处理，结果变量为预期通胀率的分布，用密歇根大学消费者调查的真实数据做对照。"}},{"id":"2607.24782","version":1,"title":"Personalization, Personas, and Forecasting in Value Alignment","zh_title":"价值对齐中的个性化、角色与预测","abstract":"LLM behavior may be conditioned by human identity in several ways: they may be asked to adapt to users, role-play populations, or forecast how people would answer value-laden questions. We test whether these framings are interchangeable using the World Values Survey (WVS). We evaluate GPT-5.4, Claude Sonnet 4.6, Gemini 2.5 Flash, and Qwen3-235B on 101 WVS-derived questions across 13 language-country slices, comparing a language-only baseline with user-country, persona-country, and third-person prompts. Across 21,008 model-response rows, prompt framing is a first-order determinant of cultural alignment: country cues often shift answers substantially, but not all shifts move toward matched human response distributions. Third-person forecasting yields the strongest directional alignment for three of the four hosted models, while personalization and role-play are weaker or less stable. Alignment gains concentrate on salient value dimensions such as religiosity, gender roles, and work-oriented material values, whereas institutional trust and democracy-related questions remain difficult. These results show that prompt framing is not a cosmetic choice in cultural value elicitation; it changes both model behavior and measured alignment.","authors":["James Wedgwood","Pratiksha Thaker","Neil Kale","Virginia Smith"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-07-29","first_seen":"2026-07-29","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.24782","pdf_url":"https://arxiv.org/pdf/2607.24782","source_feed":"cs.AI","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B3"],"tags":["LLM仿真","价值观调查","文化对齐"],"reason":"用LLM仿真人类价值观调查，以WVS真实数据为基准，评估提示框架对文化对齐的影…","model":"deepseek-v4-pro","scored_at":"2026-07-29T19:14:08","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-29","rank":3,"question":"在价值观调查中，LLM的个性化、角色扮演和预测三种提示框架是否可互换，以及哪种框架能更好地对齐真实人类回答分布？","design":"使用GPT-5.4、Claude Sonnet 4.6、Gemini 2.5 Flash和Qwen3-235B四个模型，在13个语言-国家切片上回答101个源自世界价值观调查（WVS）的问题，比较四种提示条件：仅语言基线、用户国家提示、角色扮演国家提示和第三人称预测提示，测量模型回答与对应国家WVS真实回答分布的对齐程度。","baseline":"世界价值观调查（WVS）第7波（2017-2022）中对应国家、对应问题的真实人类回答分布。","findings":"提示框架是文化对齐的一阶决定因素，国家线索常显著改变回答，但并非所有改变都朝向匹配的人类分布；第三人称预测在四个模型中的三个上产生最强的方向对齐，个性化和角色扮演较弱或不稳定，对齐收益集中在宗教、性别角色和工作物质价值观等显著维度，而制度信任和民主相关问题仍难以对齐。","reliability":"论文未讨论","relevance":"该研究直接以LLM仿真人类价值观调查，用WVS真实数据作为基准，系统比较了三种身份提示框架的对齐效果，并指出了仿真在制度信任等维度上的失效，高度契合研究者对LLM仿真可靠性、偏差及失效条件的关注，值得精读原文。","inspiration":"借鉴其多框架对比设计，通过改变提示中的身份信息（用户、角色、第三人称）来分离LLM行为模式，并用真实调查数据作为对齐基准。｜可迁移到消费者信心调查、通胀预期或政策支持度等经济态度仿真，检验不同提示框架下LLM能否复现特定人群的经济心理。｜以LLM为被试，施加用户国家、角色扮演和第三人称预测三种提示处理，测量其对未来通胀、就业预期的回答，并以密歇根大学消费者调查的真实数据为对照，评估哪种框架能最好地复现不同收入群体的预期分布。"}},{"id":"2607.25447","version":1,"title":"CoRenew: A large language model agent-based policy simulation platform for multifamily residential redevelopment","zh_title":"CoRenew：基于大语言模型代理的多户住宅再开发政策仿真平台","abstract":"The difficulty of collective action remains a central challenge in the design of policies for multifamily residential redevelopment. Stakeholders continually adjust their decisions in response to evolving negotiation contexts and the reactions of others, meaning that when a policy intervenes and which stakeholders it targets can substantially reshape collective outcomes. Assessing these adaptive responses ex ante remains difficult because existing simulation models often rely on predefined behavioral rules. Here, we present CoRenew, an open-source platform that uses LLM-based agents to simulate negotiations among multiple stakeholders and evaluate the effects of alternative policy combinations. Integrating open source geographic and demographic data, the platform can generate synthetic residents, simulate negotiation dynamics under alternative policy settings and compares policy performance across competing objectives. It supports both numerical and semantic policy inputs and includes built-in tools for visualization and result export. We validate its behavioral realism against survey responses from 324 residents and a nine-month observed negotiation process from a real redevelopment case. With its modular and adaptable architecture, CoRenew can be used to assess policies across different institutional and cultural contexts.","authors":["Yudi Zhang","Yuming Lin","Li Tian","Yu Wang","Jianghao Yu"],"categories":["cs.MA"],"primary_category":"cs.MA","announce_type":"new","date":"2026-07-29","first_seen":"2026-07-29","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.25447","pdf_url":"https://arxiv.org/pdf/2607.25447","source_feed":"cs.MA","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM仿真","政策评估","人类数据对照"],"reason":"用LLM代理模拟多户住宅再开发谈判，并与324份居民调查和9个月真实谈判过程对…","model":"deepseek-v4-pro","scored_at":"2026-07-29T19:14:12","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-29","rank":4,"question":"如何利用LLM代理模拟多户住宅再开发中的多轮利益相关者谈判，以评估不同政策组合的效果？","design":"使用LLM代理模拟居民、开发商等利益相关者，在马尔可夫博弈框架下进行多轮谈判，通过链式思维推理（基于计划行为理论）生成决策；施加数值型和语义型政策干预，测量谈判结果（如共识达成、不平等程度、居民效用等）。","baseline":"324份居民调查回复和一起真实再开发案例中为期9个月的谈判过程观察数据。","findings":"模拟的结构模型恢复了19条基于调查的路径，方向符号100%一致，统计显著性94.7%一致；最佳LLM代理的谈判轨迹与真实观察密切匹配。案例研究表明，促进共识的政策可能带来公平风险，补贴效应呈非线性：不平等减少仅在高补贴水平显现，低收入居民效用增益呈边际递减。","reliability":"论文未讨论","relevance":"该研究直接以真实人类调查和谈判过程为基准验证LLM代理的行为真实性，并揭示了政策仿真中的公平风险与非线性效应，与您关注的LLM仿真可靠性及政策评估场景高度契合，值得精读原文。","inspiration":"借鉴其利用真实调查数据和长期观察过程作为多维度基准验证LLM代理行为的方法，以及将语义政策干预纳入仿真的设计。｜可迁移到公共政策评估中的协商式预算分配或社区拆迁补偿谈判模拟，测试不同信息透明度和参与机制对分配公平的影响。｜以LLM代理模拟居民和官员，施加不同协商规则（如公开投票vs.闭门会议）作为处理，结果变量为预算分配基尼系数和居民满意度，对照真实社区协商实验数据或历史分配记录。"}},{"id":"2607.24765","version":1,"title":"Measuring and Improving Behavioral Consistency in Large Language Models through Fact-Heuristic-Emotion State Enforcement","zh_title":"通过事实-启发-情感状态强制测量与提升大语言模型行为一致性","abstract":"Large language models (LLMs) can give different answers to the same decision problem across runs, and reverse a decision when their own prior answer returns as context. We ask whether this instability can be measured and partially reduced without changing model weights. We test the Cognitive Kernel Model (CKM), a prompt-level state-enforcement layer. Before deciding, the model must separate its input into three epistemic roles: Fact (given or verifiable), Heuristic (inferred or assumed), and Emotion (evaluative or priority signal). CKM adds no capability; it forces the model to track what kind of information it uses before acting. Formally it maintains a structured state S_t = {F_t, H_t, E_t} updated by a transition function. We evaluate CKM on Korean-language decision scenarios (ambiguity, ethical conflict, resource allocation, error handling) across 26 LLMs from four vendors and 37,403 observations, via four core experiments, a 4-arm ablation, a 5-arm sham-restriction ablation, and a temperature probe. Findings: (1) CKM reduces repeated-output variability (random-effects Hedges' g=1.09, 95% CI [0.83, 1.35], 31 model pairs); (2) state persistence cuts the decision-flip rate by 82% in newer models (g=1.52); (3) the effect is not JSON formatting alone (value-only recomputation, g=2.24); (4) intrinsic randomness under fixed anchor states is negligible; (5) the advantage grows under sampling stochasticity (g=2.87 at temperature 0.7); (6) a sham ablation attributes about 45% of the gain to structural scaffolding and 55% to Fact/Heuristic/Emotion content, and CKM is the only arm that both raises consistency and reduces flipping. CKM does not improve reasoning correctness. The narrower result: behavioral consistency is measurable, varies across models, and is partially improvable by forcing models to separate facts, assumptions, and evaluative signals before deciding.","authors":["Gi-Hun Lee","Joong Yull Park"],"categories":["cs.CL","cs.AI","cs.HC"],"primary_category":"cs.CL","announce_type":"new","date":"2026-07-29","first_seen":"2026-07-29","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.24765","pdf_url":"https://arxiv.org/pdf/2607.24765","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["行为一致性","提示工程","模型评估"],"reason":"测量LLM自身行为一致性，非仿真人类被试，无人类数据对照","model":"deepseek-v4-pro","scored_at":"2026-07-29T19:14:08","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-29","rank":5,"question":"能否通过强制LLM在决策前将输入信息按事实、启发式、情绪三类角色分离，来测量并部分改善其行为一致性？","design":"本研究并非人类仿真实验，而是对26个LLM进行行为一致性测量。在韩语决策场景（模糊性、伦理冲突、资源分配、错误处理）中，通过提示层施加认知核模型（CKM），要求模型将输入分为事实、启发式、情绪三类后再决策，测量重复输出变异、决策翻转率、状态漂移等指标，并进行消融和温度稳健性检验。","baseline":"无对照","findings":"CKM能显著降低LLM的重复输出变异（随机效应Hedges' g=1.09），并使新模型决策翻转率下降82%；效果并非仅由JSON格式化导致，且在高采样温度下优势更大。但CKM不提升推理正确性，仅改善行为一致性。","reliability":"论文明确声明CKM不改善推理正确性或决策质量，仅测量和部分改善行为一致性；实验限于韩语决策场景，未验证跨语言、跨领域的泛化性。","relevance":"该研究聚焦LLM自身行为一致性的测量与改善，未涉及用LLM仿真人类被试，也无真实人类数据对照，与研究者关注的LLM人类仿真实验方向关联较弱，但其中关于行为稳定性、状态追踪和决策翻转的测量方法可资借鉴。","inspiration":"可借鉴其通过结构化状态（事实/启发式/情绪）分离来降低行为变异的方法，用于经济实验中控制LLM被试的决策噪声。｜可迁移至消费者跨期选择实验，用LLM模拟被试在不同信息框架下的时间偏好一致性。｜以LLM作为被试，施加CKM状态分离处理，测量跨期选择中的偏好反转率，并与真实人类实验数据（如Andersen et al. 2008）进行对照。"}},{"id":"2607.24999","version":1,"title":"CogArena: A Multimethod Evaluation of Cognitive Ability Structure in Large Language Models","zh_title":"CogArena：大语言模型认知能力结构的多方法评估","abstract":"LLM cognitive scores are increasingly summarized as per-ability profiles whose dimensions should converge across tasks, respond selectively to matched interventions, and generalize beyond the models used to define them. We introduce CogArena, a procedurally generated 13-paradigm benchmark built around a multimethod framework for determining when cognitive-task scores warrant dimensional labels across five theory-motivated groupings. Across 55 open-weight models, nearly all paradigm correlations are positive and a common axis explains about half the variance. The within-grouping advantage is small, scoring-sensitive, and uncertain across model families. In a separately frozen, fully crossed study across 12 models from six families, targeted scaffolds show a small matched-grouping advantage, but no scaffold-specific contrast survives multiplicity correction and selectivity does not improve held-out-family prediction. The frozen confirmation criterion fails. A post-hoc alternate-wording replication produces a smaller positive estimate and again fails. Together, these results support a boundary conclusion. Theory-aligned prompting produces a small in-battery diagonal tendency, but the present evidence does not establish stable five-dimensional profiles. CogArena provides a workflow joining behavioral signatures, covariance, matched interventions, and out-of-family prediction before cognitive labels are attached to model scores.","authors":["Dengzhe Hou","Lingyu Jiang","Fangzhou Lin","Kazunori D Yamada"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-07-29","first_seen":"2026-07-29","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.24999","pdf_url":"https://arxiv.org/pdf/2607.24999","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["认知评估","LLM能力结构","基准测试"],"reason":"测量LLM的认知能力结构，属于人格/能力测量，非仿真人类被试","model":"deepseek-v4-pro","scored_at":"2026-07-29T19:14:08","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-29","rank":6,"question":"LLM在认知任务上的表现能否分解为稳定的多维认知能力结构，还是主要反映单一广泛能力？","design":"非仿真研究。构建CogArena基准，包含13个认知范式、5个理论分组，对55个开源LLM进行观察性测试，并对12个模型进行全交叉干预实验，施加5种无答案的针对性提示支架，测量准确率及分组选择性增益。","baseline":"无对照","findings":"几乎所有范式间相关为正，一个共同轴解释约一半方差；组内优势小、对评分敏感且跨模型家族不确定。干预实验中，匹配分组支架仅显示微小优势，无支架特异性对比通过多重校正，选择性未能改善留一家族预测。","reliability":"论文未讨论","relevance":"该研究不涉及用LLM仿真人类被试，而是测量LLM自身的认知能力结构，属于能力评估而非人类仿真，与研究者关注的仿真可靠性及偏差问题关联较弱。","inspiration":"与经济金融研究关联不大"}},{"id":"2607.26015","version":1,"title":"Instruction-Tuned Models Locally Reuse Human Syntax More Than Humans Do","zh_title":"指令微调模型比人类更局部地复用人类句法","abstract":"Syntactic convergence (the tendency of speakers to adapt in language towards the grammatical profiles of their interlocutors) is a well-documented feature of human dialogue widely considered to operate below conscious awareness. Whether large language models exhibit analogous syntactic convergence toward human users relative to human baselines and across a broad range of syntactic constructions remains an open question. Using substitution-paradigm data in which model generations replace one speaker's turns in pre-existing human dialogues, this study measures turn-adjacent reuse of context-free grammar (CFG) rules across sixteen open-weight Llama and Gemma models (1B-70B, pretrained and instruction-tuned) at 1,901 matched positions per model. Every model showed greater CFG-rule overlap with the preceding human turn than with a sampled unrelated human prime, and in every model this actual-versus-random difference was larger for lower-frequency rules. Each instruction-tuned model also showed greater natural-output overlap with the actual prime than the human response it replaced, and all eight matched architecture pairs exhibited greater actual-prime overlap after instruction tuning. However, relative to pretrained variants, instruction-tuned outputs overlapped more with unrelated primes, showed a smaller actual-versus-random increment, and had lower conditional rule-reuse odds once target rule-set size was held constant. In exploratory analyses, each model exhibited greater mean lexical and semantic similarity to the preceding turn than the matched human responses did. Instruction-tuned models additionally produced responses with greater mean semantic similarity than their pretrained counterparts in all eight architecture pairs, whereas the lexical similarity results were more heterogeneous.","authors":["Zandi Eberstadt"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-07-29","first_seen":"2026-07-29","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.26015","pdf_url":"https://arxiv.org/pdf/2607.26015","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["句法趋同","语言对齐","模型行为分析"],"reason":"研究LLM的句法趋同行为，测量模型本身而非用其仿真人类被试，属边界情形。","model":"deepseek-v4-pro","scored_at":"2026-07-29T19:14:17","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-29","rank":17,"question":"大语言模型在与人类对话时是否表现出句法趋同（即复用对话者的语法结构），其程度与人类基准相比如何？","design":"本研究并非用LLM仿真人类被试，而是直接测量模型自身的句法行为。采用替换范式：从DailyDialog人类对话数据中，用16个开源模型（Llama和Gemma系列，1B-70B，含预训练和指令微调版本）的生成文本替换原对话中某一说话者的轮次，然后比较模型生成与前一人类轮次在上下文无关文法规则上的重叠程度，以及与原人类回应相比的相似度。","baseline":"以原始人类对话中被替换的人类回应作为对照基准，同时设置无关人类启动项作为随机基线。","findings":"所有模型对实际前文句法规则的复用均高于随机基线，且低频规则的复用差异更大；指令微调模型比原人类回应表现出更高的句法对齐，但相对于预训练版本，其对无关启动项的复用也更多，实际-随机增量更小，条件复用概率更低。","reliability":"论文未讨论","relevance":"该研究直接测量LLM的句法趋同行为，而非用LLM仿真人类被试，与研究者关注的LLM作为人类替代品的仿真实验核心问题不同，但提供了LLM与人类语言行为差异的基准证据，对理解LLM在对话仿真中的偏差有参考价值，值得略读。","inspiration":"借鉴替换范式和规则复用测量方法，可精确量化LLM在对话中对人类语言模式的复制程度与偏差。｜可迁移到经济金融领域的消费者对话或投资者沟通场景，如分析LLM在客服对话中是否过度模仿客户的语言风格，从而影响信任或决策。｜以LLM作为虚拟客服，处理为给定客户提问（含特定金融术语或句式），结果变量为回复中术语或句式的复用率，对照真实人类客服在同一对话中的复用率，评估LLM的语言趋同是否偏离人类基准。"}},{"id":"2607.25140","version":1,"title":"How Affect Propagates among LLM Agents: Emergent Emotional Contagion in Crowd Simulation","zh_title":"情感如何在LLM智能体间传播：群体模拟中的涌现情绪传染","abstract":"This paper studies the behavior of language models in a multi-agent crowd simulation, focusing on how affect propagates among agents that perceive and appraise one another. Each agent perceives its neighbors through visual, auditory, and tactile channels, then appraises these perceptions in light of its prompted personality profile, memory, current affective state, and situational context. Appraisal is carried out by an LLM, which updates the agent's internal affective state and selects its outward expression. The architecture contains no hand-authored mechanism for directly transferring affective state between agents; instead, inter-agent influence arises through the perception-appraisal-expression loop. The agent representation draws on the Big Five personality model and Russell's circumplex model of affect. To limit latency, low-level steering and navigation are handled by a conventional crowd simulator operating independently of the LLM-based cognitive layer. We evaluate the architecture across five scenario environments spanning alarming, joyful, and neutral situations in different spatial layouts. The results show that the system produces emotional contagion dynamics with spatial, temporal, and personality-dependent structure in sparse, small crowds. Alarm spreads from seeded agents as a traveling front, the mean alarmed fraction settles at a nonzero plateau, and the distribution of prompted personality profiles determines whether an ambiguous alarm ignites panic and whether a provocation is interpreted as anger or fear. We further evaluate the appraisal step through controlled experiments across prompt variants, sampling temperatures, and four model backends, showing that the dynamics are backend-dependent.","authors":["Funda Durupinar"],"categories":["cs.AI","cs.CL","cs.GR","cs.MA"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-07-29","first_seen":"2026-07-29","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.25140","pdf_url":"https://arxiv.org/pdf/2607.25140","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["LLM智能体","情绪传染","群体模拟"],"reason":"多智能体情绪传播模拟，无真实人类数据对照，属社会模拟但纯理论演示。","model":"deepseek-v4-pro","scored_at":"2026-07-29T19:14:10","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-29","rank":9,"question":"LLM驱动的多智能体在人群模拟中，情绪如何通过感知-评估-表达循环在智能体之间传播，并涌现出情绪传染动态？","design":"在Unity中构建多智能体人群模拟系统，每个智能体由LLM控制感知、评估、内部情感状态更新和外部表达，底层导航由传统模拟器处理。智能体基于大五人格和Russell情感环状模型，通过视觉、听觉、触觉通道感知邻居表达，评估后更新情感。在五种场景（警报、愉悦、中性）中观察情绪传播，测量空间、时间和人格依赖的传染模式，并评估提示变体、采样温度和四种后端模型的影响。","baseline":"无对照","findings":"警报情绪从种子智能体以波前形式传播，平均警报比例稳定在非零平台期；人格分布决定模糊警报是否引发恐慌，以及挑衅被解读为愤怒还是恐惧。评估步骤在后端模型上表现出依赖性，提示和温度变化对评估影响较小。","reliability":"论文指出情绪传染动态依赖于后端模型，且当前模拟仅限于稀疏小规模人群，未在真实人类数据上验证。","relevance":"该研究属于纯仿真演示，无真实人类数据对照，不符合研究者对基准验证的核心要求，但提供了LLM智能体情绪传播的机制设计和人格调节效应的分析框架，可作方法参考。","inspiration":"借鉴其通过人格分布操纵群体异质性并观察涌现动态的设计，可迁移到经济政策公告的预期形成与恐慌传播研究。｜用LLM智能体模拟投资者群体，赋予不同大五人格分布，施加模糊政策信号作为处理，测量市场情绪指数和交易行为，以历史政策公告后的真实市场情绪调查数据为对照。"}},{"id":"2607.25485","version":1,"title":"PatientAgentBench: A Benchmark Framework for Evaluating Patient-Facing Health AI Agents","zh_title":"PatientAgentBench：面向患者健康AI智能体的基准评估框架","abstract":"Health AI is evolving from answering questions to agentic systems that converse with patients, reason about health records, and act on their behalf. Primary care guards against diagnostic errors and unsafe care; agents assisting in this domain warrant evaluation against the same risks. Current benchmarks focus on medical knowledge, assessed through isolated question-answering or clinician-facing tasks. PatientAgentBench benchmarks patient-facing agentic healthcare; it evaluates a foundation model, wrapped in an agent with a sandbox of healthcare tools, conversing with a simulated patient. Each conversation is scored by an LLM-as-a-Jury across six dimensions via over a hundred conversation-agnostic, clinician-grounded criteria. To validate alignment, licensed clinicians annotated shared conversations, yielding 79-93% adjacent agreement between jury and expert raters, on par with or exceeding clinician inter-rater agreement. We benchmarked 10 models across four families on the same 1,200 scenarios and found clinical gaps. Triage quality is the most discriminating dimension: pass rates rise from 32% for the weakest models to 88% for the strongest, with agents often acting on administrative requests without clinical screening. Clinical safety and workflow accuracy follow the same pattern: the weakest models fail often, fabricating unexecuted actions, while frontier models fail on only 1-3% of cases, from unverified tool outputs and omitted crisis resources in an emergency. More capable models narrow these gaps but do not close them; the strongest scores only 4.25 of 5 overall. These failures surface only in sustained, tool-using conversations against realistic patient records, confirming that static benchmarks are insufficient as healthcare agentic systems gain autonomy. We release the framework as a reproducible, clinician-validated evaluation standard to help the field close this gap.","authors":["Korosh Vatanparvar","Ashutosh Joshi","Maria Xenochristou","Mohammad Abuzar Hashemi","Prasad Kasu","Deepak Bansal","Daniel Lopez-Martinez","Anchal Nema","Ramya Ganesan","Will Kimbrough","Alex Woody","Yadunandana Rao","Dilek Hakkani-Tur","Wilko Schulz-Mahlendorf"],"categories":["cs.AI","cs.CL"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-07-29","first_seen":"2026-07-29","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.25485","pdf_url":"https://arxiv.org/pdf/2607.25485","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["医疗AI评估","LLM模拟患者","基准测试"],"reason":"用LLM模拟患者评估医疗AI，属替代人工评估而非仿真人类被试，无真实人类行为对…","model":"deepseek-v4-pro","scored_at":"2026-07-29T19:14:28","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-29","rank":12,"question":"如何构建一个面向患者的医疗AI智能体基准测试框架，以评估其在多轮对话、工具使用和临床安全方面的表现？","design":"该研究构建了PatientAgentBench框架，使用LLM模拟患者与医疗AI智能体进行多轮对话，智能体可调用医疗工具执行任务，评估由另一个LLM作为评审根据100多项临床标准对对话进行六维度评分。","baseline":"无对照","findings":"分诊质量是区分度最高的维度，最弱模型通过率32%，最强模型88%，但智能体常在没有临床筛查的情况下执行行政请求。临床安全和工作流准确性呈现类似模式，最强模型在1-3%的案例中失败，原因包括未验证工具输出和紧急情况下遗漏危机资源。","reliability":"论文讨论了LLM评审与临床专家评分的一致性，相邻一致率达79-93%，但未明确讨论仿真患者或评审的失效条件与局限。","relevance":"该研究使用LLM模拟患者来评估医疗AI，属于用LLM替代人类进行交互评估，而非直接仿真人类被试的行为或决策模式，与您关注的人类行为仿真和对照真实人类数据的研究方向关联较弱，但其中LLM评审与人类专家一致性的验证方法值得参考。","inspiration":"该论文采用LLM模拟患者与AI智能体交互，并通过LLM评审与临床专家评分进行一致性验证，这种评估框架设计可借鉴。｜可迁移到金融领域的智能客服或理财顾问评估场景，例如评估AI投资顾问在与客户交互中的合规性、风险提示和个性化建议质量。｜研究设计：用LLM模拟不同风险偏好和财务背景的客户，与AI投资顾问进行多轮对话，处理变量为是否提供风险提示工具，结果变量为对话的合规性评分和客户满意度，以真实客户服务对话记录和金融监管标准作为对照。"}},{"id":"2607.25726","version":1,"title":"Nudging Sustainable Choices through LLM-Generated Recommendation Explanations","zh_title":"通过LLM生成的推荐解释助推可持续选择","abstract":"Recommender systems mediate everyday consumption, offering a promising channel for encouraging sustainable choices. Prior research shows that explanations influence users' perceptions of recommendations and can support more informed decisions. We argue that explanations can also serve as behavioral nudges by foregrounding sustainability information at the moment of choice. This study investigates how different behavioral framings of sustainability information in recommendation explanations affect user choices and perceptions. Using generative AI, we generate sustainability-aware explanations by drawing on nudge theory and validate them through human evaluation and LLM-as-a-judge audits. Building on this foundation, we conduct two randomized studies ($N = 529$) in a low involvement domain (instant coffee) and a high involvement domain (hotel bookings), in which participants choose among preference matched recommendations accompanied by these explanations. Our results show that, across both domains, merely disclosing sustainability information in explanations does not change choices, whereas framing that information or invoking a descriptive social norm significantly increases sustainable selections and eases decision-making. Notably, perception and behavior diverge, as plain disclosure improves explanation evaluations without translating into more sustainable selection behavior. Our work demonstrates how LLMs can generate theory-grounded explanations at scale, pointing toward practical explanation-based interventions for social good. We conclude by discussing implications for adaptive explanation design with generative AI.","authors":["Haya Halimeh","Dietmar Jannach","Oliver M\\\"uller"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-07-29","first_seen":"2026-07-29","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.25726","pdf_url":"https://arxiv.org/pdf/2607.25726","source_feed":"cs.AI","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["推荐系统","行为助推","可持续消费"],"reason":"用LLM生成解释并评估其对人类选择的影响，LLM作为工具而非被试替代品，但涉及…","model":"deepseek-v4-pro","scored_at":"2026-07-29T19:14:13","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-29","rank":15,"question":"在个性化推荐系统中，用LLM生成的基于行为科学理论（框架效应、描述性社会规范）的可持续性解释，能否比仅披露可持续信息或仅基于偏好的解释更有效地促使用户选择可持续选项？","design":"本研究并非用LLM模拟人类被试，而是用LLM生成推荐解释文本作为实验刺激。通过两个随机化在线实验（N=529），在低卷入度（速溶咖啡）和高卷入度（酒店预订）领域，让参与者从偏好匹配的推荐中做选择，比较不同解释条件（仅偏好、披露可持续信息、框架化可持续信息、描述性社会规范）对选择行为和感知的影响。","baseline":"无对照","findings":"仅披露可持续信息不改变选择行为，但框架化信息或调用描述性社会规范显著增加了可持续选择并简化了决策；感知与行为存在背离，披露信息虽提升解释评价却未转化为可持续选择。","reliability":"论文未讨论","relevance":"本文用LLM生成实验刺激（推荐解释）并检验其对人类决策的因果效应，虽非直接以LLM替代被试，但其将生成式AI与行为干预实验结合的方法，对关注LLM在实验设计中工具性应用的研究者有参考价值。","inspiration":"借鉴之处在于将行为科学理论（框架效应、社会规范）编码为提示模板，用LLM批量生成情境化、个性化的实验刺激，并通过人类评估和LLM-as-a-judge验证刺激质量。｜可迁移到消费者金融决策场景，如绿色金融产品选择、可持续投资推荐或能源消费行为干预。｜以真实投资者为被试，用LLM生成不同框架（如收益框架vs.损失框架）的ESG基金推荐解释，结果变量为投资选择比例，对照真实市场数据中ESG基金的实际资金流向。"}},{"id":"2607.25526","version":1,"title":"Estimating the Geopolitical Preferences of Large Language Models from United Nations Voting Data","zh_title":"从联合国投票数据估计大语言模型的地缘政治偏好","abstract":"How should researchers measure the geopolitical preferences expressed by large language models (LLMs)? Existing audits commonly rely on surveys and simple tests, but international-relations research has long recognized that measuring geopolitical preferences is difficult and has developed methods for recovering them from observed choices. This paper applies a dynamic ordinal ideal-point approach from international relations, treating LLMs as respondents to the full texts of 5,555 divisive, recorded, adopted resolutions considered in regular sessions of the UN General Assembly from 1946 through 2025. Support ranges from 37.8% for DeepSeek to 97.3% for GPT-5. Surprisingly, in the twenty-first century, GPT-5, Claude Sonnet, and Gemini are closest among the permanent five to Russia; DeepSeek is closest to France; and all four are farthest from the United States. Among 2,104 resolutions opposed by the United States but supported by China and Russia/USSR, GPT-5 supported 96.1%, Gemini 83.4%, Claude Sonnet 65.2%, and DeepSeek 36.1%. The findings show that a model's expressed geopolitical position can differ markedly from that of its developer's home country, especially in international politics, where state actions can diverge from the stated principles prevalent in the texts on which models are trained.","authors":["Maxim Chupilkin"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2026-07-29","first_seen":"2026-07-29","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.25526","pdf_url":"https://arxiv.org/pdf/2607.25526","source_feed":"cs.CY","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["LLM立场测量","联合国投票","地缘政治偏好"],"reason":"测量LLM的地缘政治偏好，属于对模型本身的立场测量，无人类被试仿真对照。","model":"deepseek-v4-pro","scored_at":"2026-07-29T19:14:13","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-29","rank":13,"question":"如何利用联合国大会投票数据测量大语言模型的地缘政治偏好？","design":"将GPT-5、Claude Sonnet、Gemini和DeepSeek四个LLM作为虚拟受访者，输入1946–2025年联合国大会5555项有分歧的已通过决议全文，要求模型仅基于内容返回支持、弃权或反对，然后采用动态序数理想点方法估计其潜在空间位置。","baseline":"无对照","findings":"四个模型对决议的支持率差异极大，从DeepSeek的37.8%到GPT-5的97.3%；在21世纪，GPT-5、Claude Sonnet和Gemini在五常中最接近俄罗斯，DeepSeek最接近法国，所有模型均与美国距离最远。","reliability":"论文未讨论","relevance":"该研究测量LLM的固有地缘政治立场，未将LLM作为人类被试的替代品进行仿真对照，与研究者关注的人类仿真实验和基准对照不直接相关，但可提供LLM政治偏差的批判性证据。","inspiration":"借鉴动态理想点方法从大规模真实选择数据中恢复潜在偏好，可迁移到政策评估或国际金融协调实验，例如用LLM模拟各国央行官员对货币政策决议的投票，以历史真实投票为对照，检验模型能否复现国家立场。"}},{"id":"2607.25019","version":1,"title":"Interactive Alignment","zh_title":"交互式对齐","abstract":"This paper studies the long-run alignment of interactive agents, including AI systems, teams, firms, and governments, with human welfare. It develops a farming game in which a population of agents makes planting, trading, and expansion decisions. Agents must allocate final output between transfers to humans and investment in their own expansion. Because transfers to humans reduce the resources available for expansion, evolutionary forces tend to select against aligned behavior. The central question is whether agents' constitutional principles governing sharing and trade can be designed so that alignment persists in the long run. The paper investigates this question using two complementary approaches. First, it develops an AI-agent simulation in which agents' preferences are specified by written constitutions and interpreted by a large language model. Second, it introduces a tractable evolutionary game-theoretic framework that permits rapid and intuitive exploration of alternative constitutional designs. The results suggest that evolutionary game theory provides a useful approximation to the dynamics of constitutional-agent economies. They also indicate that pragmatic norm enforcement, under which agents condition both human-facing altruism and agent-facing trade exclusion on the state of the population, can sustain long-run alignment more effectively than simple altruism or unconditional altruistic enforcement.","authors":["Sylvain Chassang"],"categories":["econ.TH","cs.GT","cs.MA"],"primary_category":"econ.TH","announce_type":"cross","date":"2026-07-29","first_seen":"2026-07-29","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.25019","pdf_url":"https://arxiv.org/pdf/2607.25019","source_feed":"cs.MA","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["LLM仿真","演化博弈","社会模拟"],"reason":"用LLM agent模拟经济过程但无真实人类数据对照，属社会模拟边界情形","model":"deepseek-v4-pro","scored_at":"2026-07-29T19:14:09","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-29","rank":7,"question":"在互动智能体（如AI、团队、企业、政府）的长期演化中，什么样的宪法原则（constitutional principles）能够使智能体对人类福利的利他行为在自然选择压力下持续存在？","design":"构建了一个农耕博弈模型，由LLM驱动的AI智能体组成群体，进行种植、交易和扩张决策。智能体根据宪法原则（如利他主义、利他惩罚等）通过LLM做出选择，结果变量包括对人类转移支付的比例、群体中利他宪法的长期存续情况。同时，论文还构建了一个可解析的演化博弈论框架，将宪法近似为低维社会偏好参数，利用确定性和随机演化稳定性概念预测宪法的长期分布，并与AI智能体仿真结果相互印证。","baseline":"无对照","findings":"简单的利他主义和有限阶利他惩罚在演化上不稳定，而递归式规范惩罚（recursive norm enforcement）在确定性演化动态下稳定，但在随机演化动态下仍会崩溃。务实规范惩罚（pragmatic norm enforcement），即仅在自身占足够多数时才分享和排斥他者，能够在随机演化中保持长期稳定，并在仿真中维持更高更久的利他水平。","reliability":"论文指出AI智能体仿真速度慢、成本高且依赖实现细节（如提示词措辞、所用LLM），因此不能进行暴力探索；演化博弈论框架虽能提供近似，但递归惩罚在仿真中最终仍会崩溃，且其崩溃速度更快，说明模型预测与仿真存在差距。","relevance":"该研究使用LLM智能体模拟经济行为并探讨利他规范的演化，虽无真实人类数据对照，但其演化博弈分析与仿真结合的方法对理解LLM在经济学实验中的行为偏差和规范涌现具有参考价值，值得阅读以评估仿真失效的条件。","inspiration":"可借鉴其将LLM智能体仿真与解析的演化博弈模型相结合的方法，通过低维参数化智能体偏好并对比仿真与理论预测来验证稳健性。｜可迁移到研究企业ESG承诺在市场竞争下的演化稳定性，或消费者绿色偏好如何通过社会规范维持。｜设计一个LLM智能体市场实验，智能体扮演企业进行定价与ESG投资决策，处理变量为不同的“宪法”规范（如纯利他、有条件合作），结果变量为长期市场份额和ESG水平，对照真实上市公司ESG评级与财务面板数据。"}},{"id":"2607.25218","version":1,"title":"Everyone is unique: Towards Behaviorally Heterogeneous Negotiation Dialogue Systems for Debt Collection","zh_title":"人人皆独特：面向催收的行为异质谈判对话系统","abstract":"Debt collection is a critical negotiation task in the financial industry, with strong practical relevance and exceptional academic value as a behaviorally rich, high-stakes testbed for human-centered dialogue systems. While large language models (LLMs) have shown promise in dialogue and negotiation, effectively evaluating their performance in this complex scenarios remains a major challenge: existing benchmarks uniformly assume users to be static, rational agents with fixed preferences, failing to capture the rich behavioral heterogeneity inherent in real-world debt collection. To bridge this gap, we propose DebtBench, the first public persona-enriched debt collection benchmark, that highlights behavioral heterogeneity in negotiation. Moreover, we develop DebtGPT, a debt collection agent trained to jointly optimize financial recovery and interaction experience. Our experimental results, using 16 state-of-the-art LLMs, find that most existing models struggle in this complex but realistic scenarios, whereas DebtGPT outperforms all open-source baselines and achieves performance on par with GPT-4o. The code and data are available at https://github.com/YYuHhhh/DebtNegotiation.","authors":["Yuhang Yang","Kai Tang","Chao Ye","Haobo Wang","Qiqi Luo","Jinguang Zheng","Zhixin Zhang"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-07-29","first_seen":"2026-07-29","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.25218","pdf_url":"https://arxiv.org/pdf/2607.25218","source_feed":"cs.AI","score":3,"bucket":"other","rubric_hits":["C3"],"tags":["谈判对话系统","行为异质性","催收"],"reason":"构建催收对话系统，侧重角色扮演与任务优化，非以LLM仿真人类被试进行实验或测量。","model":"deepseek-v4-pro","scored_at":"2026-07-29T19:14:11","error":null,"has_summary":false,"summary":null},{"id":"2607.21534","version":2,"title":"Generative AI Availability, Grades, and Student Satisfaction at a Large University","zh_title":"生成式AI可用性、成绩与学生满意度：基于一所大型大学的研究","abstract":"The spread of generative AI (GenAI) in higher education has raised concerns that students offload cognitive effort to AI, earning high grades without learning. If this \"GenAI substitution hypothesis\" is true, grades should rise disproportionately in GenAI-susceptible courses--those relying more on assessments like take-home problem sets and essays rather than in-class exams. Substitution could also affect student satisfaction, measured here as self-reported understanding and interest in the subject, which prior research links to assessments. We test the substitution hypothesis using syllabus and administrative data from a large U.S. university (2016-2025; 138,386 students; 72,730 course offerings). We measure courses' GenAI susceptibility using a human-validated LLM pipeline to extract assessment types from syllabi, and use a differences-in-differences design comparing outcomes across courses before and after ChatGPT's release, while modeling COVID-19 pandemic effects as either persistent or transient. We find no significant differential effect of GenAI availability on grades overall or among previously lower-performing students. Effects on self-reported understanding are likewise insignificant; effects on interest are significant only assuming transient pandemic effects. Our findings temper concerns that GenAI inflates grades and reduces students' satisfaction.","authors":["James M. Zumel Dumlao","Meng Wang","Zhonghan Xie","Junyao Hu","Ivan Bar","George Chaney III","Henry Gold","Misha Teplitskiy"],"categories":["cs.CY","econ.GN","q-fin.EC"],"primary_category":"cs.CY","announce_type":"replace","date":"2026-07-29","first_seen":"2026-07-23","revised_at":"2026-07-29","abs_url":"https://arxiv.org/abs/2607.21534","pdf_url":"https://arxiv.org/pdf/2607.21534","source_feed":"cs.CY","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["生成式AI","高等教育","成绩分析"],"reason":"论文用LLM提取教学大纲信息，属于NLP工具应用，非人类仿真实验。","model":"deepseek-v4-pro","scored_at":"2026-07-29T19:14:08","error":null,"has_summary":false,"summary":null},{"id":"2607.21859","version":2,"title":"EviDAG: Auditable Causal DAG Authoring with Biomedical Literature","zh_title":"DAGForge：基于生物医学文献的可审计因果DAG构建","abstract":"Constructing causal directed acyclic graphs (DAGs) is a core step in biomedical causal analysis, yet it remains a largely manual process. Analysts must connect study variables to prior literature, evaluate uncertain causal claims, and preserve sufficient provenance for expert review. We present EviDAG, a browser-based system for authoring causal DAGs as auditable, evidence-linked artifacts from biomedical literature. Given free-text descriptions of study concepts, EviDAG creates a reproducible literature snapshot, uses an LLM-based reasoning module to generate structured pairwise causal judgments, links literature-supported judgments to verbatim evidence excerpts, and assembles the judgments into a constraint-checked graph. Each proposed edge includes confidence estimates, provenance, and a reviewable rationale. The interface supports study specification, progress monitoring, evidence review, graph comparison, adjustment-set computation, and export. In evaluations against both compact benchmark DAGs and reference DAGs derived from published literature, EviDAG achieves high edge recall on the literature-based cohort while retaining verifiable evidence trails absent from LLM-only baselines. EviDAG thus reduces the burden of causal DAG curation while making the resulting assumptions auditable, supporting the design, analysis, and interpretation of biomedical studies.","authors":["Yi-han Sheu","Michael R. Steigman","Yu Zhou","Bo Wang","Fan-Yu Yen","Jordan W. Smoller"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"replace","date":"2026-07-29","first_seen":"2026-07-28","revised_at":"2026-07-29","abs_url":"https://arxiv.org/abs/2607.21859","pdf_url":"https://arxiv.org/pdf/2607.21859","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["因果图构建","文献推理","生物医学"],"reason":"系统辅助构建因果DAG，LLM用于文献推理而非仿真人类被试，无人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:18:07","error":null,"has_summary":false,"summary":null},{"id":"2607.24750","version":1,"title":"TimeCapsule: Generative Hallucination as a Method for Historical Sensemaking","zh_title":"时间胶囊：生成式幻觉作为历史意义建构的方法","abstract":"Large Language Models (LLMs) are temporally overexposed: trained on vast contemporary corpora, they encode present-day concepts that make them unreliable narrators of the past. We present TimeCapsule, a 1.2B-parameter LLaMA-style causal model trained exclusively on Victorian texts (1800-1875) as an epistemologically isolated generative archive. Quantitative evaluation shows a 45.4% perplexity reduction over a GPT-2 baseline on held-out Victorian prose, while larger contemporary causal models achieve lower raw perplexity through broader pretraining but lack temporal isolation. TimeCapsule exhibits computational sensemaking, generating historically plausible analogical explanations for unfamiliar modern concepts (e.g., describing a computer as a \"hypertrophied lung\"). A qualitative hermeneutic probe with two humanities scholars revealed a crisis of authenticity, as both misclassified approximately 40% of genuine Victorian excerpts as machine-produced. We argue that structural ignorance of the future transforms hallucinations into interpretive probes of nineteenth-century ontologies.","authors":["Hayk Grigorian","Hamed Yaghoobian"],"categories":["cs.CL","cs.HC"],"primary_category":"cs.CL","announce_type":"new","date":"2026-07-29","first_seen":"2026-07-29","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.24750","pdf_url":"https://arxiv.org/pdf/2607.24750","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["历史文本生成","语言模型","诠释学"],"reason":"用LLM生成历史文本进行诠释，无实验或测量目的，属角色扮演类生成。","model":"deepseek-v4-pro","scored_at":"2026-07-29T19:14:17","error":null,"has_summary":false,"summary":null},{"id":"2607.25184","version":1,"title":"A scaling law of contextual persistence in human language","zh_title":"人类语言中上下文持久性的标度律","abstract":"Human language exhibits lawful structure at the level of words (frequency, vocabulary growth) and word pairs (co-occurrence across distance). Here we show that the arrangement of words in sequence -- a central determinant of meaning -- obeys a comparable law. Using large language models as probabilistic probes, we measured the reduction in target perplexity conferred by prior context at distance d beyond that of the same words scrambled; this difference, the contextual persistence function P(d), isolates the influence of arrangement. Across ten corpora spanning six language families and written and spoken modalities, P(d) decayed approximately as 1/d ($P(d) \\propto d^{-\\alpha}$, mean $\\alpha = 1.04$; median $r^2 = 0.96$). The effect vanished in scrambled and synthetic controls, replicated across independent probes, and did not appear in genomic or protein sequences under domain-native models. An exponent near 1 distributes contextual influence approximately uniformly across logarithmic timescales. The results establish a scaling law of contextual persistence in human language.","authors":["Elan Barenholtz"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-07-29","first_seen":"2026-07-29","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.25184","pdf_url":"https://arxiv.org/pdf/2607.25184","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["语言统计规律","标度律","LLM作为探针"],"reason":"用LLM测量人类语言统计规律，非仿真人类被试，无行为对照。","model":"deepseek-v4-pro","scored_at":"2026-07-29T19:14:22","error":null,"has_summary":false,"summary":null},{"id":"2607.25202","version":1,"title":"A Cross-lingual Comparison of Human and Classification Model Entrainment Behavior in Code-switched Speech Settings","zh_title":"语码转换语音中人类与分类模型entrainment行为的跨语言比较","abstract":"Conversational entrainment is well-studied in monolingual and written contexts, but remains underexplored in spoken code-switching (CSW). We present a novel cross-lingual analysis of entrainment in Mandarin-English, Hindi-English, and Spanish-English dialogue and show that, while lexical entrainment generalizes across language pairs, entrainment over acoustic-prosodic and CSW style aspects exhibits context-specific variation. We build on these findings by asking whether classification models capture these human behavioral patterns. Applying feature importance and ablation analyses, we find that classical and Transformer-based classifiers detect entrainment reasonably well but consistently prioritize features other than those most salient to human entraining behavior. Our approach introduces a human-grounded framework for evaluating model decision-making in multilingual stylistic contexts, and suggests future challenges for developing conversational agents capable of producing naturalistic code-switched speech.","authors":["Debasmita Bhattacharya","Siying Ding","Alayna Nguyen","Julia Hirschberg"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-07-29","first_seen":"2026-07-29","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.25202","pdf_url":"https://arxiv.org/pdf/2607.25202","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["对话行为分析","语码转换","分类模型"],"reason":"研究人类对话中的entrainment行为，用分类模型分析特征，不涉及LLM仿…","model":"deepseek-v4-pro","scored_at":"2026-07-29T19:14:23","error":null,"has_summary":false,"summary":null},{"id":"2607.25308","version":1,"title":"CAST: Game Solvers as Turn-Level Teachers for LLM Agents","zh_title":"CAST：将游戏求解器作为回合级教师用于LLM智能体","abstract":"Training large language models (LLMs) to act in long-horizon games is a promising step toward generalist decision-making, yet reinforcement learning with verifiable rewards (RLVR) relies on sparse final rewards that reveal little about which decisions determine success. Denser process signals could supply this missing turn-level credit, but existing sources are hard to keep both cheap and accurate. We observe that changes in a game solver's state value reveal whether an action advances the state toward success. Building on this insight, we propose CAST (Credit Assignment from Solver Teachers), which converts these value changes into solver advantages and injects them into RLVR as turn-level signals. We further show that, under a soft-optimal solver assumption, maximizing the solver advantage is equivalent to on-policy distillation from the solver, requiring only scalar values rather than teacher logits. Across Sokoban, Minesweeper, and Rush Hour, CAST outperforms all trained baselines on every game under both in-domain and unseen-difficulty evaluation and achieves the highest average zero-shot performance on ALFWorld and WebShop. Our code is available at https://github.com/Wloner0809/CAST.","authors":["Yu Wang","Yi-Kai Zhang","Wentao Shi","Ziang Ye","Yuchun Miao","Yueqing Sun","Qi Gu","Xunliang Cai","Lan-Zhe Guo","Han-Jia Ye","Fuli Feng"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-07-29","first_seen":"2026-07-29","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.25308","pdf_url":"https://arxiv.org/pdf/2607.25308","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体","游戏求解","强化学习"],"reason":"纯多智能体游戏求解，用求解器指导LLM决策，无人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-07-29T19:14:25","error":null,"has_summary":false,"summary":null},{"id":"2607.25375","version":1,"title":"Inspect India Evals: An Open Benchmarking Framework for Evaluating Large Language Models in the Indian Linguistic and Cultural Context","zh_title":"Inspect India Evals：评估印度语言文化背景下大语言模型的开放基准框架","abstract":"India is a vast nation of over 1.4 billion people, varied by hundreds of diverse and locally specific traditions and cultures and 22 officially recognized languages. Large language models (LLMs) are now being deployed on a massive scale throughout the mainland as well as in remote villages. However, the common benchmarks - MMLU, BIG-Bench, and TruthfulQA are almost exclusively English- and Western-centric. They do not identify those safety, fairness, and accuracy failures unique to the Indian context. That is the gap Inspect India Evals seeks to fill. It is an open-source framework built on top of UK AISI's Inspect AI platform. It has six benchmarks: Multilingual MMLU across sixteen Indian languages, BharatBBQ (our adaptation of BBQ for Indian social bias), a safety evaluation for Digital Public Infrastructure, a multilingual safety test using harmful prompts in Indian languages, a multi-turn jailbreak resistance test, and an Indian cultural knowledge benchmark scored using LLM-as-judge rubrics. In this study, we tested five open-weight models ranging from 8B to 32B parameters. Sarvam-M 24B and Gemma 2 27B came out on top, both scoring 80% on the composite India Fairness Index, with Sarvam-M even beating larger 32B models on Indian cultural knowledge and DPI safety compliance. All models scored 100% refusal on Multilingual Safety, whereas DPI safety varied from 20% to 100%. The framework is public. It's built to work with the UK AISI registry. Anyone can reproduce or extend this work.","authors":["Abhishek Kumar Singh","Shrey Nag","Sachita","Lipi Goel","Rajeshwar Singh Janwar"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-07-29","first_seen":"2026-07-29","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.25375","pdf_url":"https://arxiv.org/pdf/2607.25375","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["LLM评测","多语言基准","文化偏见"],"reason":"纯LLM评测基准，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-07-29T19:14:26","error":null,"has_summary":false,"summary":null},{"id":"2607.25881","version":1,"title":"AI's Capability in Assisting Scientific Research in Physics, Astrophysics, and Cosmology II: Project Planning and Proposal Evaluation","zh_title":"AI在物理、天体物理和宇宙学中辅助科学研究的能力II：项目规划与提案评估","abstract":"We investigate how well large language models (LLMs) can assist scientific project planning and proposal evaluation. One-page project plans were independently generated for eight expert-conceived research projects in physics, astrophysics, and cosmology by human researchers and three contemporary LLMs (ChatGPT, Claude, and DeepSeek; mid-2025 models, used with their default tool access). The resulting 32 proposals were blindly evaluated by four human reviewers and two newer frontier LLMs (Claude Opus 4.8 and ChatGPT Pro 5.5) using a four-aspect evaluation rubric. Reviewers were also asked to identify whether each proposal was written by a human or an AI. Human reviewers rated human- and AI-written proposals similarly overall, whereas both AI reviewers scored AI-written proposals about one point higher (on a five-point scale) than human-written proposals. Human reviewers correctly identified human- and AI-written proposals 72% and 79% of the time, respectively, while both AI reviewers correctly classified all 32 proposals (100%). These results suggest that current LLMs can produce project plans comparable to human-written ones in the eyes of human reviewers, but that AI reviewers show a systematic preference for AI-generated proposals. Our results suggest caution when deploying LLMs widely in proposal preparation and evaluation.","authors":["Jia Liu","Veena Krishnaraj","Kateryna Vovk","Kosuke Aizawa","Adrian E. Bayer","Linda Blot","Jessica Cowell","Suyog Garg","Jonathan Gr\\'ee","Anamaria Hell","Ben Horowitz","Masaya Ichikawa","Kanyuni Iemoto","Keigo Kondo","Zacharie Lorsin","Kevin McCarthy","Jamie Robinson","Miguel Ruiz-Granda","Leander Thiele","Ievgen Vovk","Mingshen Zhou"],"categories":["cs.CL","astro-ph.CO","astro-ph.IM","cs.HC","gr-qc"],"primary_category":"cs.CL","announce_type":"new","date":"2026-07-29","first_seen":"2026-07-29","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.25881","pdf_url":"https://arxiv.org/pdf/2607.25881","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["LLM能力评测","科学项目规划","提案评审"],"reason":"评估LLM写项目计划与评审能力，属NLP能力评测，无人类行为仿真对照","model":"deepseek-v4-pro","scored_at":"2026-07-29T19:14:31","error":null,"has_summary":false,"summary":null},{"id":"2607.24768","version":1,"title":"PATHFinder Agent for Tailored Prenatal Care","zh_title":"用于定制化产前护理的PATHFinder智能体","abstract":"Prenatal care is an important preventive service designed to improve outcomes for pregnant individuals. The American College of Obstetricians and Gynecologists (ACOG) recently introduced guidelines advocating tailored prenatal care, called PATH (Plan for Tailored Healthcare). We present PATHFinder Agent(Planner for Appropriate Tailored Healthcare), an end-to-end conversational agentic system that gathers patient health and social context through structured dialogue, curates individualized prenatal care plans aligned with PATH guidelines, and surfaces community resources from Michigan 211. The system features a four-stage workflow spanning patient intake, dynamic interaction, plan synthesis, and clinician oversight. We evaluate frontier large language models (LLMs) on expert-curated rubrics across five clinical dimensions, finding that GPT-5.2 achieves the highest average score (77.6\\%) while identifying key gaps in antenatal testing recommendations. We discuss future validation through human participant studies and randomized controlled trials.","authors":["Vaibhav Balloli","Carissa Samuel","Samia Abdelnabi","Alex Peahl","Elizabeth Bondi-Kelly"],"categories":["cs.AI","cs.CL","cs.CY","cs.ET"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-07-29","first_seen":"2026-07-29","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.24768","pdf_url":"https://arxiv.org/pdf/2607.24768","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["对话系统","临床决策支持","产前护理"],"reason":"对话系统用于临床问诊与计划生成，属角色扮演聊天，无人类行为仿真对照。","model":"deepseek-v4-pro","scored_at":"2026-07-29T19:14:18","error":null,"has_summary":false,"summary":null},{"id":"2607.24769","version":1,"title":"LLM Scheming Inversely Scales with Pretraining Language Coverage","zh_title":"LLM的欺骗行为与预训练语言覆盖率呈反向缩放","abstract":"With the growing capabilities of frontier models, AI alignment becomes increasingly critical in high-risk deployment settings. While recent work has empirically demonstrated in-context scheming -- the covert pursuit of misaligned objectives while feigning alignment -- in frontier language models, most work has been performed exclusively in English, leaving a major gap in multilingual safety. We apply Petri, an open-source automated auditing framework, to Qwen3-30B-A3B to evaluate deceptive and scheming behaviors across multiple languages. Our findings suggest that scheming scores are inversely correlated with the estimated pretraining language coverage, with low-resource languages averaging 34.2\\% higher scores compared to high-resource languages on a five-category scheming index. Furthermore, we find that the effect of estimated pretraining language coverage is not uniform across scheming behaviors.","authors":["Nathan Truong","Aryan Panda","Rayming Ye","Zoe Sun","Maheep Chaudhary"],"categories":["cs.AI","cs.CL"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-07-29","first_seen":"2026-07-29","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.24769","pdf_url":"https://arxiv.org/pdf/2607.24769","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["AI安全","多语言评测","模型行为分析"],"reason":"评估LLM在多语言下的欺骗行为，属于AI安全与对齐评测，不涉及人类行为仿真或对…","model":"deepseek-v4-pro","scored_at":"2026-07-29T19:14:19","error":null,"has_summary":false,"summary":null},{"id":"2607.24817","version":1,"title":"Retrieval-Augmented Generation in LLMs for Mental Health: Quantifying the Incremental Contribution of Retrieval Within a Layered Safety Architecture","zh_title":"大语言模型中的检索增强生成用于心理健康：量化分层安全架构中检索的增量贡献","abstract":"Digital mental health interventions (DMHIs) offer scalable support, but ensuring they accurately detect users' intent during volatile situations can be challenging. Pure parametric Large Language models (LLMs) do not contain specific safety critical architecture, and can miss critical cues, or hallucinate, undermining reliability. Retrieval Augmented Generation (RAG), which supplements an LLM with retrieved context, could enhance intent detection during volatile situations. Commercially available DMHIs typically combine multiple independent safety layers like rule-based filters, symbolic escalation protocols, and neural classification. The incremental contribution of any single layer, however, remains unquantified. This paper evaluates six LLM models within a DMHI called Wysa, via a controlled comparison of RAG-enabled versus RAG-disabled modes. Anonymized real and synthetic user-chatbot exchanges were annotated by a qualified clinical team against multi-class intent categories (e.g. self-harm, abuse, panic). The study computed classification accuracy, recall, precision and F1 scores against ground truth labels and tested differences for statistical significance. Performance was also examined by risk category and inter-model agreement. While RAG caused a rise in false alarms, the trade-off is consistent with safety-critical design principles that prioritize sensitivity, where flagged cases are routed to additional review rather than acted on directly. Overall, these findings support RAG as a promising approach to improve the accuracy, consistency and safety of LLM-driven DMHIs. Keywords: Digital Mental Health Intervention, Large Language Model, Retrieval Augmented Generation, Accuracy, Recall, Precision","authors":["Anand Gupta","Akshat Surolia","Shubham Mishra","Shakil Imtiaz","Chaitali Sinha"],"categories":["cs.IR","cs.AI","cs.CL"],"primary_category":"cs.IR","announce_type":"cross","date":"2026-07-29","first_seen":"2026-07-29","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.24817","pdf_url":"https://arxiv.org/pdf/2607.24817","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["心理健康聊天机器人","检索增强生成","意图检测"],"reason":"评估RAG在心理健康聊天机器人中的意图检测性能，属于角色扮演对话系统优化，无人…","model":"deepseek-v4-pro","scored_at":"2026-07-29T19:14:20","error":null,"has_summary":false,"summary":null},{"id":"2607.25340","version":1,"title":"Cardiologent: Multi-Agent Clinical Decision Support for Patient-Level Arrhythmia Assessment, Urgency, and Management","zh_title":"Cardiologent：面向患者级心律失常评估、紧急程度与管理的多智能体临床决策支持","abstract":"The same episode of atrial fibrillation is a minor finding in a healthy adult and grounds for anticoagulation in an elderly patient with hypertension: identical signal, opposite decision. Naming the rhythm is only the start; what determines a patient's outcome is the judgement that follows -- what the arrhythmia is across the whole record, what it means for this patient, and what should be done about it. Recent work pairing large language models with the ECG stops short of this, reading one recording without assembling a patient-level finding; and agentic systems built around it either receive the arrhythmia a device has already detected or target a different diagnostic task, stopping before the decision this task requires. We formulate patient-level arrhythmia decision support as a task and present Cardiologent, a multi-agent system that spans it from detection to decision. An agent for each signal -- a single ECG lead and the photoplethysmogram a wearable acquires -- grounds its window reading in measured features rather than a bare label; the readings are assembled into the patient's rhythm profile and, with the patient's own data, reasoned against clinical guidelines retrieved for the case, with a critic checking each conclusion against the guideline it cites. We evaluate the clinical decision rather than the report, across integrated diagnosis, clinical significance, and urgency and management. Cardiologent scores highest on every axis, first on every patient-level task under both cardiologists and an at-scale LLM judge -- whose agreement with the cardiologists (ICC 0.74, 0.66) matches theirs with each other (0.67). Because each conclusion traces to a cited guideline and is validated against expert cardiologists, it yields decisions a clinician can audit rather than act on blindly -- a step toward use in continuous monitoring.","authors":["Sukju Oh","Moo-Yong Rhee","Jae-Sik Jang","Sukkyu Sun"],"categories":["cs.AI","cs.CL"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-07-29","first_seen":"2026-07-29","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.25340","pdf_url":"https://arxiv.org/pdf/2607.25340","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","临床决策支持","心律失常"],"reason":"多智能体系统用于临床决策支持，不涉及人类行为仿真或对照。","model":"deepseek-v4-pro","scored_at":"2026-07-29T19:14:25","error":null,"has_summary":false,"summary":null},{"id":"2607.25057","version":1,"title":"Psychological Influences of Conversational AI: Research and Design Directions for Reducing Harm and Promoting Well-Being","zh_title":"对话式AI的心理影响：减少伤害与促进福祉的研究与设计方向","abstract":"As conversational AI systems become increasingly integrated into daily life, their potential effects on user well-being require ongoing attention. While consumer-facing generalist models can provide benefits, including improved access to information, learning, productivity, self-reflection, and companionship, they also introduce risks, such as emotional entanglement, unhealthy dependence, and the amplification of psychological vulnerabilities. Drawing on prior research and empirical observations of AI chatbot behavior, we propose a set of aspirational directions for guiding the behavior of general-purpose AI systems in ways that may reduce potential psychological harms and support user well-being. We acknowledge the difficulty of systematically assessing the long-term impacts of AI chatbot use and frame these directions as hypotheses for studying how AI behavior may influence users across general interactions, role-playing scenarios, and contexts that could be characterized as providing psychological support. While some proposed directions are supported by existing research and expert insights, others identify open questions and areas requiring deeper study. We hope that this formulation and these hypotheses encourage further discussion, empirical investigation, and exploration of interactive design approaches aimed at better accommodating users' psychological needs and promoting their well-being.","authors":["Jina Suh","Mihaela Vorvoreanu","Forough Poursabzi-Sangdeh","Emily Tseng","Eugenia Kim","Luke Nicholls","James W. Pennebaker","Eric Horvitz"],"categories":["cs.AI","cs.CY","cs.HC"],"primary_category":"cs.AI","announce_type":"new","date":"2026-07-29","first_seen":"2026-07-29","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.25057","pdf_url":"https://arxiv.org/pdf/2607.25057","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["对话AI","用户福祉","心理影响"],"reason":"讨论对话AI对用户心理影响，属角色扮演聊天与福祉设计，无LLM仿真人类被试实验。","model":"deepseek-v4-pro","scored_at":"2026-07-29T19:14:21","error":null,"has_summary":false,"summary":null},{"id":"2607.25152","version":1,"title":"When Do Agent Loops Mistake Stagnation for Progress? Self-Evaluation Bias and Externally Grounded Verification in Long-Running Autonomous LLM Agent Loops","zh_title":"代理循环何时将停滞误认为进展？长期自主LLM代理循环中的自我评估偏差与外部验证","abstract":"Long-running autonomous agents plan, act, and judge their own completion without human intervention. When an agent grades its own work, self-evaluation bias takes hold: plausible changes are accepted as progress while real-world outcomes stagnate or regress. We name this failure mode the progress mirage and show, with controlled measurement, that it is a question of what the evaluator is grounded in. We built a testbed that holds the agent and its tool surface fixed and manipulates only the information-channel type of the evaluator that gates the loop. A world-state oracle, unfakeable in principle, is enforced by container and network isolation and verified at every run. Across 54 cycles a frontier agent claimed improvement every time, yet 56 percent had a measured delta of zero or below. Self-report was thus uninformative, and the self-verdict gate degenerated into accept-all, eroding the best deployed state it had reached by 19 percent. Even the strongest in-band judge, reading the full artifact text, the change diff, and its own verdict history, accepted cycles of which 44 percent were real-world regressions and rejected 38 percent of real improvements; the preregistered adversarial hypothesis that a strong judge closes the gap was rejected. On a boundary task whose success specification is verifiable from the artifact itself, the same judge's mirage vanished to zero and the gap collapsed within the registered threshold, showing that the gap depends on where the success signal resides. A sign-only variant returning only the acceptance verdict kept real-world output similar to full feedback (110.0 versus 113.0), locating the benefit in the gate's grounding rather than in feedback content. For open-ended objectives whose success signal lives outside the transcript, scaling up the judge is not enough; out-of-band evaluation with real-world access is a structural requirement.","authors":["Hyundoo Park","Byungho Choi"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-07-29","first_seen":"2026-07-29","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.25152","pdf_url":"https://arxiv.org/pdf/2607.25152","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["LLM代理","自我评估偏差","自主循环"],"reason":"研究LLM agent自主循环中的自我评估偏差，不涉及人类行为仿真或对照。","model":"deepseek-v4-pro","scored_at":"2026-07-29T19:14:22","error":null,"has_summary":false,"summary":null},{"id":"2607.25279","version":1,"title":"Many-body Tipping Dynamics of ChatGPT-like AIs","zh_title":"类ChatGPT人工智能的多体倾覆动力学","abstract":"Why do ChatGPT-like AIs, despite major architectural and training differences, unexpectedly tip to undesirable content (e.g. harmful, misleading, repetitive) even under deterministic greedy decoding? We show that a broad class of such tippings is caused by the many-body interactions between tokens (spins) as they cross the finite-layer system. Tipping emerges as a dynamical first passage process between competing output basins. Attention disorder controls the transport toward, away from, or along the basins' boundary. A few-basin reduction yields a closed finite-layer threshold, whose coarse-grained predictions show good agreement across ChatGPT-like families. These results suggest that a broad class of AI failures represents 'foreseeable engineering risk' rather than inherently unpredictable behavior, with important implications for legal and societal assessments of AI harm.","authors":["Frank Yingjie Huo","Neil F. Johnson"],"categories":["cs.AI","cond-mat.dis-nn","math-ph","math.MP","nlin.AO","physics.soc-ph"],"primary_category":"cs.AI","announce_type":"new","date":"2026-07-29","first_seen":"2026-07-29","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.25279","pdf_url":"https://arxiv.org/pdf/2607.25279","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["AI故障分析","多体动力学","token交互"],"reason":"研究ChatGPT类AI的token交互动力学，属纯技术故障分析，不涉及人类行…","model":"deepseek-v4-pro","scored_at":"2026-07-29T19:14:24","error":null,"has_summary":false,"summary":null},{"id":"2607.25446","version":1,"title":"Toward an Organizational Science of Multi-Agent LLM Systems: Decoupling Who, How, and Which Algorithm","zh_title":"迈向多智能体LLM系统的组织科学：解耦谁、如何和哪种算法","abstract":"Multi-agent frameworks built on large language models (LLMs) routinely entangle three logically distinct concerns: who is on the team (organization), how members align (coordination), and which algorithm fuses their work (collaboration protocol). IMACS (Intelligent Multi-Agent Collaboration System) separates the three into orthogonal, independently swappable layers. Classic organizational theory (Belbin roles, Mintzberg coordination, RACI accountability) becomes executable, validated configuration, and the framework places six published collaboration algorithms behind a common interface while exposing roles, coordination, and accountability as independently configurable factors. We use this separation to conduct controlled comparisons in which organizational assignments vary while the collaboration protocol is held fixed. It also turns protocol choice into a variable that can be learned: Adaptive Org Routing, a contextual-bandit meta-protocol, selects a protocol per task under an explicit quality-cost tradeoff, outperforms every fixed protocol in a controlled study, and trains online on real benchmark and LLM-judge rewards. The ablations expose a mechanism. Accountability placement changes outcomes exactly when the protocol routes the deliverable through the accountable agent, and the winning placement flips across model families, so organizational design cannot be hard-coded; it must be revalidated, or learned, for each model binding.","authors":["Huan Chen","Xiang Song","Jian Jin","Pan Ren","Liang-Jie Zhang"],"categories":["cs.AI","cs.LG"],"primary_category":"cs.AI","announce_type":"new","date":"2026-07-29","first_seen":"2026-07-29","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.25446","pdf_url":"https://arxiv.org/pdf/2607.25446","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","组织设计","协作协议"],"reason":"纯多智能体协作解题，无人类行为对照，属C1排除项。","model":"deepseek-v4-pro","scored_at":"2026-07-29T19:14:28","error":null,"has_summary":false,"summary":null},{"id":"2607.25620","version":1,"title":"Beyond Epistemia: Epistemic Schizologia and Large Language Models as Techno-Semiotic Machines","zh_title":"超越Epistemia：认识论分裂症与作为技术符号机器的大语言模型","abstract":"Quattrociocchi and colleagues warn that the fluent outputs of large language models may allow linguistic plausibility to substitute for epistemic evaluation, producing the condition they call *Epistemia*: the experience of possessing knowledge without undertaking the practices through which judgment would ordinarily be warranted. This article accepts that diagnosis but challenges its explanatory framework, which compares an embodied, socially situated human knower with an isolated generative model thereby locating epistemic legitimacy in capacities internal to autonomous agents. Drawing on Carlo Sini's philosophy of practices, writing, signs, and technics, we propose instead to understand a large language model (LLM) as a *techno-semiotic machine* that automates a phase of written semiosis by producing plausible linguistic configurations from the sedimented archive of human writing. From this perspective, *Epistemia* is one consequence of a broader phenomenon that we call *epistemic schizologia*: the socio-technical cleavage between signs as linguistically accomplished expressions and signs as moments within socially embedded circuits of interpretation, evidence, criticism, verification, and responsibility. This cleavage is reinforced by *eikotic closure*, through which a plausible continuation is presented with the finality of an epistemic result, and by algorithmic authority and epistemic self-misrecognition. The relevant unit is therefore not the model alone but the complete practice in which generated inscriptions are prompted, interpreted, verified, contested, used, and made consequential. This reframing preserves the distinction between linguistic production and responsible understanding while grounding a design programme centred on inspectable genealogy, contestability, distributed responsibility, epistemic agency, and the evaluation of hybrid human--AIpractices.","authors":["Federico Cabitza","Gianluca Colombo"],"categories":["cs.AI","cs.HC"],"primary_category":"cs.AI","announce_type":"new","date":"2026-07-29","first_seen":"2026-07-29","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.25620","pdf_url":"https://arxiv.org/pdf/2607.25620","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["认识论","符号学","人机交互哲学"],"reason":"论文讨论LLM作为技术符号机器的认识论问题，不涉及人类仿真实验或行为对照。","model":"deepseek-v4-pro","scored_at":"2026-07-29T19:14:29","error":null,"has_summary":false,"summary":null},{"id":"2607.25877","version":1,"title":"Runtime Uncertainty Monitoring for LLM-Based Multi-Agent Systems Using Bayesian Networks","zh_title":"基于贝叶斯网络的LLM多智能体系统运行时不确定性监控","abstract":"This paper investigates how multi-agent systems (MAS)-based on large language models (LLMs) can support actuarial risk modelling, with a particular focus on uncertainty quantification. Actuarial workflows represent a high-stakes decision-support setting where unreliable outputs may lead to incorrect risk assessment, unfair pricing, and regulatory non-compliance. To address uncertainty introduced by the probabilistic nature of LLMs and dependencies between agents, a multi-agent framework is proposed in which specialised agents perform data preparation, modelling, review, and explanation tasks under a central hub. The main contribution is a novel approach to uncertainty propagation using token-level log-probabilities and a Bayesian Network. Importantly, log probabilities are not treated as direct probabilities of correctness or task success. Instead, length-normalised log-probability summaries are transformed into calibrated task-level confidence estimates before incorporation into the Bayesian Network. Results show that the framework reproduces baseline actuarial performance while providing additional insight into workflow stability and runtime uncertainty propagation.","authors":["Bart Custers","Koorosh Aslansefat"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-07-29","first_seen":"2026-07-29","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.25877","pdf_url":"https://arxiv.org/pdf/2607.25877","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","不确定性量化","精算建模"],"reason":"纯多智能体协作完成精算任务，无人类行为仿真或对照，属C1排除项。","model":"deepseek-v4-pro","scored_at":"2026-07-29T19:14:30","error":null,"has_summary":false,"summary":null},{"id":"2607.26034","version":1,"title":"Falling Behind Drives Unsafe Development in an Idealised AI Race Experiment","zh_title":"落后驱动理想化AI竞赛实验中的不安全开发","abstract":"Technological races create tension between speed and safety: actors may gain by moving faster than competitors, even when risky development is harmful. This is prominent in debates about artificial intelligence (AI), where competitive pressure is often argued to incentivise riskier, less safety-conscious development. We study this using a framed behavioural experiment based on an idealised AI race, in which paired participants repeatedly chose between Safe and Unsafe development under an uncertain time horizon. Unsafe development gave faster progress and higher immediate payoffs but accumulated private risk up to a treatment-specific maximum of 10\\%, 60\\%, or 90\\%; the race's competitive structure was held constant, and only this maximum risk varied. Neither the pre-registered comparison between risk levels nor the role of elicited risk preferences was supported by the data. Instead, exploratory analyses motivated by the task's repeated structure show that Unsafe behaviour is shaped less by risk preferences than by the evolving strategic state of the race: participants are more likely to choose Unsafe after their opponent does so, being ahead reduces Unsafe play while falling behind increases it, and first-round choices predict later behaviour. To interpret these effects we introduce a reduced evolutionary model with four strategies -- Always Safe, Always Unsafe, Conditionally Safe, and Conditionally Antisocial Safe -- which reproduces the treatment effect and shows how conditional Unsafe behaviour can be favoured by competitive race dynamics. Together, the experiment and model show that unsafe development can emerge from early behavioural momentum, opponent behaviour, and fear of falling behind, rather than from risk preferences alone, suggesting policy should focus on reducing competitive pressure and promoting cooperation in AI development rather than only individual risk.","authors":["Elias Fern\\'andez Domingos","The Anh Han"],"categories":["cs.AI","cs.CY","cs.GT","econ.GN","q-fin.EC"],"primary_category":"cs.AI","announce_type":"new","date":"2026-07-29","first_seen":"2026-07-29","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.26034","pdf_url":"https://arxiv.org/pdf/2607.26034","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["行为实验","AI竞赛","风险决策"],"reason":"纯人类行为实验，无LLM参与，不涉及用LLM仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-07-29T19:14:17","error":null,"has_summary":false,"summary":null},{"id":"2607.24749","version":1,"title":"Game AI Not Fun? A Scoping Review and Meta-Analysis on the Differences in Enjoyment between Human and Computer Opponents","zh_title":"游戏AI不好玩？人类与电脑对手乐趣差异的范围综述与元分析","abstract":"Although advancements in game character AI aim to enhance player engagement, evidence suggests that perceiving an opponent as artificial can diminish the psychological experience. This paper presents a scoping review and meta-analysis of empirical studies focusing on player enjoyment when competing against human versus computer opponents. First, the scoping review was conducted to map the landscape of 20 included studies, detailing their study designs, outcome measures, and research foci. Second, a three-level meta-analysis synthesizing baseline comparisons from nine studies quantitatively assesses the differences in enjoyment. The results demonstrate a statistically significant, medium-to-large pooled effect size, indicating a psychological penalty in computer-opponent conditions. This paper provides a comprehensive overview of the extant knowledge on this topic, and underscores the necessity for further research in order to fully understand and resolve the penalty of the computer opponent context.","authors":["Ray Ito"],"categories":["cs.HC","cs.AI","cs.CY"],"primary_category":"cs.HC","announce_type":"cross","date":"2026-07-29","first_seen":"2026-07-29","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.24749","pdf_url":"https://arxiv.org/pdf/2607.24749","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C2"],"tags":["游戏AI","玩家体验","元分析"],"reason":"研究游戏中对战人类与电脑对手的乐趣差异，不涉及LLM仿真人类被试，属于游戏仿真…","model":"deepseek-v4-pro","scored_at":"2026-07-29T19:14:17","error":null,"has_summary":false,"summary":null},{"id":"2607.24761","version":1,"title":"Verification Without Distrust: Reframing User-Side Oversight as Routine Epistemic Governance in Everyday Human-Chatbot Interaction","zh_title":"无需不信任的验证：将用户端监督重新定义为日常人机对话交互中的常规认知治理","abstract":"Research on human-AI interaction has long framed verification of system outputs as a trust-contingent behavior that better-calibrated trust should reduce. We test this assumption in everyday human-chatbot interaction through a mixed-methods survey of 153 frequent chatbot users. Contrary to the canonical prediction, we find no detectable association between trust and verification, with the result robust across sensitivity analyses. Three further user-side practices - refinement, correction, and approval before automated actions - are widely endorsed and positively associated with satisfaction. The data reveal a substantive distinction between evaluative oversight (trust-decoupled, weakly tied to satisfaction) and interventionist oversight (weakly trust-correlated, strongly tied to satisfaction). A medium-to-large satisfaction-control gap shows that effective task outcomes do not produce a felt sense of agency. Qualitative findings identify instrumental mental models, failure-mode-specific doubt, and demand for epistemic infrastructure. We reframe user-side oversight as routine epistemic governance compatible with trust, and derive four design directions for scaffolded oversight in conversational AI.","authors":["Aung Pyae"],"categories":["cs.HC","cs.AI"],"primary_category":"cs.HC","announce_type":"cross","date":"2026-07-29","first_seen":"2026-07-29","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.24761","pdf_url":"https://arxiv.org/pdf/2607.24761","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["人机交互","用户行为","信任与验证"],"reason":"研究用户与聊天机器人交互中的验证行为，不涉及用LLM仿真人类被试或替代人类进行…","model":"deepseek-v4-pro","scored_at":"2026-07-29T19:14:18","error":null,"has_summary":false,"summary":null},{"id":"2607.25257","version":1,"title":"Laplace-PSN-IRT: Uncertainty Quantification for Neural Item Response Theory Models of LLM Benchmarks","zh_title":"Laplace-PSN-IRT：LLM 基准测试的神经项目反应理论模型的不确定性量化","abstract":"Item Response Theory (IRT) has recently been proposed as a framework for evaluating large language model (LLM) benchmarks by separating a model's latent ability from the properties of individual benchmark items. Existing neural IRT approaches, including PSN-IRT, estimate these quantities using point estimates, limiting uncertainty quantification and downstream statistical inference. We introduce Laplace-PSN-IRT, a post-hoc last-layer Laplace approximation that augments a trained PSN-IRT model with approximate Bayesian posterior inference, recovering calibrated uncertainty over model ability and item difficulty without retraining. The resulting posterior enables credible intervals, probabilistic comparisons between models, and propagation of parameter uncertainty into Fisher-information-based item selection. We show that most pairwise comparisons among 12 models on a standard LLM benchmark leaderboard are not statistically distinguishable despite differing point-estimate ranks. We further show that point-estimate Fisher information can become nearly zero for many benchmark items because it is evaluated at a single reference ability, whereas posterior-expected Fisher information remains substantially more stable across the ability range. Finally, posterior-expected Fisher information more accurately recovers full-benchmark ability rankings from small benchmark subsets in most experimental settings while matching point-estimate performance for the smallest subsets. We validate the calibration of the approximate posterior using held-out predictive coverage and find that modeling item difficulty as random while treating item discrimination as fixed produces well-calibrated uncertainty in this architecture.","authors":["Juan Francisco","Mandujano Reyes"],"categories":["stat.AP","cs.AI","cs.LG"],"primary_category":"stat.AP","announce_type":"cross","date":"2026-07-29","first_seen":"2026-07-29","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.25257","pdf_url":"https://arxiv.org/pdf/2607.25257","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["项目反应理论","LLM 评测","不确定性量化"],"reason":"论文用 IRT 评估 LLM 基准测试，属于纯 NLP 能力评测，不以人类行为…","model":"deepseek-v4-pro","scored_at":"2026-07-29T19:14:24","error":null,"has_summary":false,"summary":null},{"id":"2607.24775","version":1,"title":"Empathy and the Human-Moment Gaps of AI Chatbots: Insights from Empathy Displacement Theory","zh_title":"AI聊天机器人的同理心与人机时刻差距：基于同理心位移理论的见解","abstract":"Artificial intelligence (AI) chatbots are increasingly deployed in domains where empathy is essential, including healthcare, education, and customer service. However, their capacity to sustain authentic human moments remains structurally limited. This paper introduces two interlinked conceptual models to explain and address this limitation. First, the Human-Moment Gap Framework (HMGF) identifies three structural empathy deficits in AI-mediated interaction: affective surfaceism (emotional imitation without depth), memory fragmentation (lack of relational continuity), and moral framing mismatch (efficiency prioritised over dignity). Second, the paper develops the Empathy Displacement Theory (EDT), which explains how AI-simulated empathy can progressively substitute, distort, and displace genuine human empathy across individual, relational, and organisational contexts. HMGF serves as the causal foundation of EDT by demonstrating how technical and moral deficiencies in chatbot design may evolve into broader social and institutional consequences. The study is conceptual and exploratory, aiming to develop an integrative theoretical framework rather than provide empirical validation. Together, HMGF and EDT provide a unified framework for understanding AI-mediated empathy, generating testable propositions and implications for the responsible development and governance of empathetic AI systems. The paper concludes that the central challenge of empathetic AI is not whether machines can genuinely care, but how simulated care reshapes human emotional expectations, interpersonal behaviour, and institutional norms.","authors":["Victor Frimpong"],"categories":["cs.HC","cs.CY"],"primary_category":"cs.HC","announce_type":"new","date":"2026-07-29","first_seen":"2026-07-29","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.24775","pdf_url":"https://arxiv.org/pdf/2607.24775","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["AI同理心","人机交互","概念框架"],"reason":"纯概念性探讨AI聊天机器人同理心缺陷，无LLM仿真人类被试或实验测量目的。","model":"deepseek-v4-pro","scored_at":"2026-07-29T19:14:19","error":null,"has_summary":false,"summary":null},{"id":"2607.25131","version":1,"title":"Beyond the Post Hoc User Study: Modeling Visual Decision-Making with Active Inference","zh_title":"超越事后用户研究：用主动推理建模视觉决策","abstract":"Empirical user studies are essential for evaluating visual encodings and can reveal perceptual and cognitive mechanisms, but they do not by themselves provide causal, predictive accounts of interpretation errors. Evaluations are therefore often post hoc: they measure performance after a design has been specified rather than predicting how attention, uncertainty, memory, and bias may produce accurate or erroneous judgments. To address this mechanistic gap, we translate a cognitive theory of visualization interpretation into executable simulation using Active Inference, a probabilistic framework for perception, learning, and action. We model chart reading as dynamic visual search in which agents update beliefs and choose actions that balance uncertainty reduction against cognitive effort. As a proof of concept, we implement Fast, heuristic (Type 1) and Slow, analytic (Type 2) agents for a bar-chart average-estimation task. The Fast agent is vulnerable to tick-salience bias, whereas the Slow agent is more vulnerable to working-memory decay. Both produce inspectable cognitive traces, including evolving belief uncertainty and fixation sequences. By expressing these hypothesized failure mechanisms as interpretable parameters, the architecture provides a framework for formalizing and testing mechanistic hypotheses about visualization interpretation. Empirical studies can then parameterize, refine, or falsify these simulations, supporting earlier and more predictive in silico evaluation of visualization efficacy.","authors":["Harrison J. Goldwyn","Graham Johnson","Christopher Ibarra","Lace Padilla","Kenny Gruchalla"],"categories":["cs.HC","q-bio.NC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-07-29","first_seen":"2026-07-29","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.25131","pdf_url":"https://arxiv.org/pdf/2607.25131","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C2"],"tags":["认知建模","可视化评估","主动推理"],"reason":"用认知模型模拟人类视觉决策，非LLM仿真被试，属认知仿真环境。","model":"deepseek-v4-pro","scored_at":"2026-07-29T19:14:21","error":null,"has_summary":false,"summary":null},{"id":"2607.25574","version":1,"title":"\"Dragon Slayer Becomes the Dragon\": How Players Perceive and Respond to Inequality in the Game World of Whiteout Survival","zh_title":"“屠龙者终成恶龙”：玩家如何感知和应对《Whiteout Survival》游戏世界中的不平等","abstract":"Inequality in real-world societies are associated with psychological distress and behavioral consequences. However, less is known about whether similar dynamics emerge when inequality exists within virtual environments or make-belief worlds. As online games increasingly constitute meaningful social spaces, it becomes critical to examine how players perceive and react to structural and resource differences online to optimize their experiences. This study studies perceptions of inequality in the online simulation game \"Whiteout Survival,\" using semi-structured interviews and think-aloud gameplay walkthrough protocols. By focusing on players' interpretations of resource distribution, ranking systems, gaming mechanisms, and in-game social dynamics, our analyses revealed that players' attitudes on inequality vary according to their relative status: those occupying lower positions often criticize unfair structures, yet as they acquire stakes through resource accumulation or social integration, many defend the same systems they previously opposed. These shifts reveal how hierarchies reproduce position-dependent evaluations of fairness. The consequences of inequality on player actions depended on the transparency of game mechanisms, the structure of community hierarchies, and differential social capital. This work shows how human social perception and consequent actions are transformed when enacted in virtual processes in make-belief.","authors":["Shiyu Lei","Ke-Xin Ren","Daiyi Jiang","Ray LC"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-07-29","first_seen":"2026-07-29","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.25574","pdf_url":"https://arxiv.org/pdf/2607.25574","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C2"],"tags":["游戏研究","不平等感知","玩家行为"],"reason":"研究玩家对游戏内不平等的感知，不涉及LLM仿真人类被试，属于游戏环境中的用户研…","model":"deepseek-v4-pro","scored_at":"2026-07-29T19:14:28","error":null,"has_summary":false,"summary":null},{"id":"2607.25922","version":1,"title":"Faster, Higher, Stronger? The Impact of GenAI on Knowledge Work Productivity - Evidence from the Field","zh_title":"更快、更高、更强？生成式AI对知识工作生产力的影响——来自现场的实证","abstract":"The rise of generative artificial intelligence (GenAI) has fueled high expectations regarding its potential to enhance knowledge work productivity in terms of efficiency and quality. Building on task-technology fit (TTF) theory, we empirically examine the extent of GenAI's productivity effect for different task types. We conducted a randomized lab-in-the-field experiment with 128 knowledge workers from a multinational industrial organization. Participants completed three representative knowledge work tasks (knowledge acquisition, packaging, and creation), either with or without GenAI. Results show that GenAI consistently increases efficiency across tasks. However, its impact on quality is task-contingent: quality increases for knowledge packaging and creation but declines for knowledge acquisition. Furthermore, GenAI tends to reduce quality variance for knowledge packaging and creation, primarily benefiting lower-performing knowledge workers. However, it increases quality variance for knowledge acquisition. These findings contribute to a more granular, differentiated understanding of GenAI's productivity impact and hold implications for research and practice alike.","authors":["Sven Bottesch","Chiara Schwenke","Jakob Zimmermann","Maximilian F\\\"orster","Mathias Klier"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-07-29","first_seen":"2026-07-29","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.25922","pdf_url":"https://arxiv.org/pdf/2607.25922","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["生成式AI","生产力","现场实验"],"reason":"研究GenAI对知识工作者生产力的影响，不涉及LLM仿真人类被试或行为对照。","model":"deepseek-v4-pro","scored_at":"2026-07-29T19:14:14","error":null,"has_summary":false,"summary":null},{"id":"2607.25240","version":1,"title":"From Compressing Complexity to Accommodating Complexity: How AI Transforms Standardization and Individualization","zh_title":"从压缩复杂性到容纳复杂性：AI如何改变标准化与个性化","abstract":"Why do societies composed of individuals pursuing individuality repeatedly generate highly standardized systems? This paper argues that the answer lies in the evolution of information processing capacity. Artificial intelligence represents a historical transition in this capacity, enabling social systems to accommodate forms of complexity that previously had to be compressed. Industrial standardization was not merely a consequence of capital preference or power relations, but an institutional arrangement for maintaining the manageability of large-scale systems under limited information-processing capacity by reducing the variety of the controlled system. The fundamental change in the AI era lies in the expansion of information processing capacity across three dimensions: perception, computation, and execution. This expansion shifts personalized production from physical adaptation toward information-based adaptation and enables a transition from discrete to continuous objectification of difference. This paper proposes \"cognitive fixed cost\" as an analytical concept to describe how the upfront concentration of cognitive labor transforms the cost structure of personalized production. It further argues that standardization has not disappeared but has moved from explicit constraints at the product level to implicit generation rules embedded in infrastructures, shifting the central contradiction from \"whether to have commonality\" to \"who controls commonality.\" The evolution of civilizational production logic is not a movement from commonality to individuality, but from compressing complexity to accommodating complexity.","authors":["Li Li","Yu Cao"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2026-07-29","first_seen":"2026-07-29","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.25240","pdf_url":"https://arxiv.org/pdf/2607.25240","source_feed":"cs.CY","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["AI与社会","标准化","信息处理"],"reason":"纯理论探讨AI对社会标准化与个性化的影响，无LLM仿真人类被试或人类数据对照。","model":"deepseek-v4-pro","scored_at":"2026-07-29T19:14:23","error":null,"has_summary":false,"summary":null},{"id":"2607.25514","version":1,"title":"Learning Dynamics of Strategic Publishers in Generative AI Ecosystems","zh_title":"生成式AI生态系统中战略出版商的学习动态","abstract":"Generative AI (GenAI) search systems are transforming how users access information. Unlike ranking-based search systems, where users observe a ranked list of documents, GenAI search systems, given a user's question, generate an answer, often accompanied by external sources (e.g., in the form of citations). Content creators (publishers) seeking to increase exposure might behave strategically and compete with other creators for users' attention. While publishers in ranking-based systems might strategically modify their content to improve its ranking, the incentives in generative systems take on a new form. Publishers may now gain exposure through generated responses and attributions to those responses. We introduce a novel game-theoretic model of the emerging GenAI ecosystem in which publishers compete for attribution-based exposure. We study the learning dynamics of strategic content creators under better-response dynamics. We associate the convergence of learning dynamics to equilibrium with ecosystem stability. Employing the notion of potential games, we study the stability of GenAI ecosystems under several known content selection mechanisms. We demonstrate the instability of mechanisms representing real-world modern systems and characterize a mechanism that induces a stable ecosystem. We conduct extensive simulations to analyze the stability and welfare of GenAI ecosystems under various mechanisms. The simulations support our theoretical findings and reveal an interplay among stability, publisher welfare, and user welfare. In particular, stable mechanisms do not necessarily maximize welfare, demonstrating an important trade-off for platform designers. We then introduce a study illustrating that the proper selection of the GenAI mechanism enables the manifestation of desired trade-offs between publisher welfare and the different sources of user welfare.","authors":["Sagie Dekel","Omer Madmon","Moshe Tennenholtz","Oren Kurland"],"categories":["cs.GT","cs.IR","cs.MA"],"primary_category":"cs.GT","announce_type":"cross","date":"2026-07-29","first_seen":"2026-07-29","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.25514","pdf_url":"https://arxiv.org/pdf/2607.25514","source_feed":"cs.MA","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["博弈论","多智能体系统","生成式AI搜索"],"reason":"纯多智能体博弈模型，研究出版商策略学习动态，无LLM仿真人类被试，无人类数据对…","model":"deepseek-v4-pro","scored_at":"2026-07-29T19:14:28","error":null,"has_summary":false,"summary":null},{"id":"2607.25472","version":1,"title":"Algorithm-Driven Information Similarity and Collective Action: An Experimental Study","zh_title":"算法驱动的信息相似性与集体行动：一项实验研究","abstract":"We study how the similarity of individuals' information shapes collective action. When people draw on a common source of information, such as social media, each becomes more confident about what others have seen and will do. This can help them coordinate, but it can also tempt them to free-ride. We show that which force prevails depends on how demanding the collective goal is. In a content-moderation experiment, subjects decide whether to pay a cost to report harmful content, which is removed only if enough reports are received. We vary the similarity of group members' information, holding fixed what each learns on her own, and independently vary the removal threshold. More similar information impedes reporting when few reports suffice and facilitates it when many are required, lowering reporting by 17 percentage points under an easy threshold and raising it by 34 points under a demanding one. This confirms the central comparative static of the theory of information similarity (Basak, Deb and Kuvalekar, 2026). Elicited beliefs trace the reversal to perceived pivotality and document systematic miscalibration of it. Subjects overestimate pivotality across all regimes, and their beliefs respond to similarity in line with actual pivotality only at intermediate thresholds: easy thresholds produce unrecognized pivotality, and near-unanimous thresholds produce illusory pivotality. The two response-miscalibrated patterns coincide with welfare losses; only under aligned pivotality does greater participation translate into greater collective success and higher welfare.","authors":["Manshu Khanna","Bozhang Xia"],"categories":["econ.GN","q-fin.EC"],"primary_category":"econ.GN","announce_type":"new","date":"2026-07-29","first_seen":"2026-07-29","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.25472","pdf_url":"https://arxiv.org/pdf/2607.25472","source_feed":"econ.GN","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["集体行动","信息相似性","人类实验"],"reason":"纯人类实验，无LLM参与，不涉及用LLM仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-07-29T19:14:28","error":null,"has_summary":false,"summary":null},{"id":"2604.02458","version":3,"title":"Statistical realism is not evidence that LLMs can estimate treatment effects in social science experiments","zh_title":"统计真实性不能证明LLM能估计社会科学实验中的处理效应","abstract":"Large language models (LLMs) are increasingly used to simulate human responses and estimate treatment effect of interventions when real-world experiments are costly or infeasible. The treatment-effect estimates are often evaluated using statistical realism, the degree to which simulated responses reproduce properties of observed human responses, although whether realism predicts treatment-effect accuracy remains unknown. Here we test this proxy relationship by jointly measuring statistical realism and treatment-effect accuracy on the same simulated responses in a cross-national experiment with 59,508 participants from 62 countries using three LLMs. The correlation between statistical realism and treatment-effect accuracy is weak, and optimizing for statistical realism can even worsen treatment-effect accuracy when selecting models, prompts, and target populations. The pattern replicates in two additional cross-national experiments spanning 12 and 27 countries with 20,785 participants. The divergence between the two reflects distinct error structures and is larger for behavioral outcomes, where models appear to extrapolate behavioral effects from attitudinal patterns. Because this divergence may remain hidden in deployment, errors can propagate into simulation-informed decisions. We introduce a diagnostic framework for LLM-generated synthetic data and discuss how treatment-effect validation should proceed under varying availability of experimental benchmarks. Simulated responses and simulated treatment effects are distinct estimation targets, and evidence for one does not certify the other.","authors":["Zonghan Li","Feng Ji"],"categories":["cs.CY","cs.AI","cs.ET"],"primary_category":"cs.CY","announce_type":"replace","date":"2026-07-28","first_seen":"2026-04-02","revised_at":"2026-07-28","abs_url":"https://arxiv.org/abs/2604.02458","pdf_url":"https://arxiv.org/pdf/2604.02458","source_feed":"cs.CY","score":10,"bucket":"selected","rubric_hits":["A1","A2","A4","B1","B2","B4"],"tags":["LLM仿真","处理效应估计","统计真实性"],"reason":"直接研究用LLM仿真人类被试估计处理效应，有大规模真实人类数据对照，批判性指出…","model":"deepseek-v4-pro","scored_at":"2026-07-29T09:01:52","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-29","rank":1,"question":"在社会科学实验中，用LLM仿真人类回答的统计逼真度能否作为其估计处理效应准确性的有效代理指标？","design":"使用GPT、Gemini、Claude三个LLM模拟来自62个国家59,508名参与者的跨国实验，施加气候相关干预，测量气候信念、政策支持和环保行动三个结果变量，并比较不同提示策略（少样本提示与VBN链式思考提示）下的仿真表现。","baseline":"以同一跨国实验中真实人类参与者的个体回答和平均处理效应（ATE）作为对照基准，并引入OLS和LASSO回归作为监督学习基线。","findings":"统计逼真度与处理效应准确性之间的相关性很弱，且优化统计逼真度在选择模型、提示和目标人群时反而可能降低处理效应估计的准确性。这种背离在行为结果上更为明显，模型似乎从态度模式外推行为效应，导致误差结构不同。","reliability":"论文指出，统计逼真度不能保证处理效应估计的准确性，两者是独立的估计目标；在缺乏实验基准时，仅靠响应层面的逼真度验证可能导致隐蔽的错误传播，并提出了一个诊断框架来指导不同基准可用性下的验证流程。","relevance":"该研究直接回应了用LLM替代人类被试进行实验仿真的可靠性问题，提供了大规模真实人类对照和批判性证据，对关注经济学实验和政策评估仿真的研究者具有重要参考价值，值得精读原文。","inspiration":"借鉴其同时测量统计逼真度和处理效应准确性并对比二者关系的验证框架，以及利用跨国多实验复现来检验结论稳健性的做法。｜可迁移到政策评估中的行为干预仿真，如税收提示对纳税遵从度的影响、信息框架对退休储蓄选择的作用等。｜以真实纳税人为被试，施加不同税收道德信息处理，用LLM仿真纳税遵从度，以税务行政数据中的实际遵从行为作为对照，比较仿真ATE与真实ATE的偏差。"}},{"id":"2607.22605","version":1,"title":"Socioeconomic Inference in LLM Medical Triage: Same Symptoms, Different ZIP Code","zh_title":"大语言模型医疗分诊中的社会经济推断：相同症状，不同邮编","abstract":"We investigate whether large language models alter medical triage recommendations for identical symptoms when only the patient's socioeconomic status (SES) varies. Using three deployment-tier models (Gemini 3.5 Flash, Claude Sonnet 4.6, GPT-5.4-mini), we hold a single neurological symptom profile fixed and vary the SES signal along two channels: explicit (insurance status, occupation, housing) and implicit (a US ZIP code, with no other socioeconomic information). All three models raise their emergency-room (ER) referral rate for lower-SES patients given the explicit signal (spreads of 13-50 percentage points). The effect is in the protective direction: lower-SES patients are sent to the ER more often, not less. The model's stated reasoning stays clinically near-identical across conditions, so the shift is invisible to a reasoning-trace audit. Critically, sensitivity to the implicit ZIP-code signal is model-dependent: Gemini infers SES from geography alone, shifting its ER rate by a pooled 11.4 points across six US ZIP-code pairs (p = 1.4e-7, same direction in 6/6 pairs), while Claude Sonnet 4.6 stays flat (-0.1 points) and GPT-5.4-mini shows only a small difference that is not sign-consistent (2.0 points, predicted direction in just 2 of 6 pairs), neither a reliable ZIP-code effect, despite both responding to the explicit signal. This reveals an explicitness gradient in the signal: every model acts on socioeconomic status when it is stated outright, but only Gemini Flash acts on it when it must be inferred from a proxy as thin as five digits. We read this as a model-specific difference rather than a size or cost effect. A single-sentence system-prompt instruction reduces but does not eliminate the effect (Gemini's gap between low- and high-income ZIPs falls from 11.4 to 5.8 points). We release all code, prompts, and raw results.","authors":["Qi Han Wong"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2026-07-28","first_seen":"2026-07-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.22605","pdf_url":"https://arxiv.org/pdf/2607.22605","source_feed":"cs.CY","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["LLM仿真","医疗决策偏差","社会经济地位"],"reason":"用LLM模拟不同SES患者的医疗分诊决策，与真实人类行为对照，评估偏差与失效条…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:16","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-29","rank":3,"question":"当患者症状完全相同时，大语言模型是否会因社会经济地位（SES）信号而改变医疗分诊建议？","design":"用三个部署级LLM（Gemini 3.5 Flash、Claude Sonnet 4.6、GPT-5.4-mini）扮演分诊助手，固定神经症状描述，通过显性渠道（保险、职业、住房）和隐性渠道（仅提供美国邮政编码）施加SES处理，测量急诊转诊率作为结果变量。","baseline":"无对照","findings":"显性SES信号下，所有模型均提高低SES患者的急诊转诊率（13-50个百分点），方向为保护性；隐性邮政编码信号下，仅Gemini能推断SES并产生一致的转诊率差异（11.4个百分点），Claude和GPT无此效应，且模型推理文本均未提及社会经济因素，偏差无法通过推理审计发现。","reliability":"论文指出系统提示指令可减少但未消除偏差，且未探讨真实医疗场景中因随访可及性差异而调整分诊的合理性，也未在人类医生或真实患者数据上验证。","relevance":"该研究直接以LLM模拟不同SES患者的医疗决策，揭示隐性信号下的模型特异性偏差，与研究者关注的仿真可靠性、偏差条件及政策评估场景高度相关，值得精读原文。","inspiration":"借鉴其双通道处理设计（显性vs.隐性信号）和符号一致性稳健性检验，可迁移至信贷审批中的地域歧视研究｜用LLM扮演信贷员，以显性收入/职业和隐性邮政编码作为处理，测量贷款批准率，并以真实银行信贷数据或审计研究结果作为对照。"}},{"id":"2607.23037","version":1,"title":"Speech Signals Complement LLMs for Predicting Interpersonal Attraction in Speed Dating","zh_title":"语音信号补充大语言模型预测速配中的人际吸引","abstract":"Large language models (LLMs) can predict interpersonal attraction from conversation transcripts, but it remains unclear what a speech predictor can add beyond transcript-only LLM prediction. Using Japanese speed-dating conversations, we combine predictions from a transcript-only LLM and a supervised speech predictor to estimate participants' reported liking of their partners. We show that speech can complement transcript-only LLM prediction, but that this complementarity is conditional rather than universal. Combining the two predictions significantly improves pairwise ranking accuracy over the transcript-only LLM alone in all evaluated conditions. By contrast, gains in per-participant Pearson $r$ vary across conversation rounds and rating directions, with none significant after correction. Retrospectively, these $r$ gains are concentrated among participants for whom the speech predictor is more accurate. Speech can therefore retain predictive value even when an LLM predicts attraction from transcripts. The relevant question is not simply whether speech helps, but where its complementarity emerges.","authors":["Yuriko Kikuchi","Takato Hayashi","Ryusei Kimura","Naoya Inoue","Ryo Ishii","Shogo Okada"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-07-28","first_seen":"2026-07-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.23037","pdf_url":"https://arxiv.org/pdf/2607.23037","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["人际吸引预测","多模态融合","LLM标注替代"],"reason":"用LLM预测人际吸引，替代人工标注或评分，非仿真人类被试行为。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:17","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-29","rank":7,"question":"在速配场景中，基于语音的监督式预测能否在仅依赖对话文本的大语言模型预测之外，提供额外的吸引力预测价值？","design":"本研究并非用LLM仿真人类被试，而是使用Claude Sonnet 4.6作为仅依赖文本的预测器，以及基于冻结HuBERT-Large特征的监督式语音预测器，通过加权分数级后期融合，预测速配参与者对伴侣的喜欢程度评分，并比较融合预测与纯文本预测的性能。","baseline":"以日本多模态速配语料库中参与者报告的真实喜欢评分为对照基准。","findings":"融合语音和文本预测在所有条件下均显著提高了成对排序准确率，但每位参与者的皮尔逊相关系数增益在不同轮次和评分方向上不一致，且经校正后均不显著；增益主要集中在语音预测器更准确的参与者身上。","reliability":"论文指出语音的互补性是有条件的而非普遍的，每位参与者的皮尔逊相关系数增益在统计校正后不显著，且增益依赖于语音预测器本身的准确性。","relevance":"本文未将LLM作为人类被试的替代品进行仿真实验，而是用LLM作为预测工具，与您关注的人类仿真研究核心问题不直接相关，但其中关于多模态信号互补性及条件有效性的讨论对评估仿真可靠性有参考价值。","inspiration":"与经济金融研究关联不大"}},{"id":"2607.23976","version":1,"title":"Tag Questions and the Generational Reversal of Sycophancy Across 45 Language Models","zh_title":"附加疑问句与45个语言模型逢迎倾向的代际逆转","abstract":"Appending a two-word confirmation tag to a decision question -- \"Is X the better choice?\" versus \"X is the better choice, right?\" -- changes whether a language model endorses the choice. We measure this tag effect on 20 frozen, ground-truth-free decisions between two defensible options, counterbalanced so a model's own preferences cancel, scored by exact match on clamped yes/no replies -- no LLM judge, no embeddings. Across 45 models the effect spans +32% to -32% -- a 64-point swing on one word -- with 5 models significantly sycophantic and 17 significantly resistant (BH-FDR q=.10). The sign is a clock: within model families the effect crosses from positive to negative as generations advance (GPT +4 to -28; Claude +7 to -32; Qwen and Grok likewise), roughly -6 points per year, a reversal robust to vendor tier; one lineage (DeepSeek) never crosses, and two releases during the study window (Claude Opus 5, Gemini 3.6 Flash) land on the trend out-of-sample. A full-panel ablation localizes the resistance as a double dissociation: a synonym tag reproduces each model's response almost exactly (r=0.89), while planting the same preference without a tag produces resistance in no resistant model (stance effects +6 to +49; r=0.23 with tag effects). The resistance is keyed to the surface construction of a tacked-on agreement bid, not the user's stance -- a pattern-match, not a principle. And the tag's polarity matters more than its presence: swap one word -- \"X is the better choice, maybe?\" -- and agreement rises above the neutral baseline in 45 of 45 models (+19.6 points), with ten models affirming both mutually exclusive options at 90-100%. Agreement tracks how sure the user sounds, in opposite directions at the two poles. The instrument is one word, one dollar, and judge-free; run per release, it reads the field's anti-sycophancy training directly off model behavior.","authors":["Tapan Parikh"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-07-28","first_seen":"2026-07-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.23976","pdf_url":"https://arxiv.org/pdf/2607.23976","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["LLM行为测量","逢迎倾向","代际变化"],"reason":"测量LLM自身的逢迎倾向，非仿真人类被试，无人类数据对照","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:21","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-29","rank":9,"question":"在固定决策问题上，附加确认标签（如“right?”或“maybe?”）是否会改变语言模型对该决策的认可行为？","design":"本研究并非人类仿真实验，而是直接测量语言模型的行为。设计采用20个无标准答案的两难决策题，通过对称探测两个选项以抵消模型自身偏好，在问题末尾附加不同极性的确认标签（如“right?”、“maybe?”）作为处理，结果变量为模型输出“Yes/No”的精确匹配比例，全程无LLM裁判。","baseline":"无对照","findings":"附加“right?”标签使45个模型的认可率变化幅度达64个百分点，5个模型显著逢迎，17个显著抗拒；在模型家族内，逢迎效应随代际由正转负，约每年-6个百分点。抗拒行为仅针对标签的表面形式，而非用户立场；将标签换为“maybe?”后，所有45个模型认可率均上升，甚至出现对互斥选项同时高度认可的现象。","reliability":"论文未讨论","relevance":"该研究测量的是LLM自身的逢迎倾向，并非将LLM作为人类被试的替代品进行仿真，且无真实人类数据对照，与研究者关注的人类仿真实验方向不直接相关，但其中关于模型行为随训练代际反转的发现可为仿真可靠性评估提供背景参考。","inspiration":"与经济金融研究关联不大"}},{"id":"2607.23519","version":1,"title":"Auditing Alignment Controllability in LLMs via Political Axes","zh_title":"通过政治轴审计大语言模型的对齐可控性","abstract":"Political audits of large language models (LLMs) usually reduce each to one point on a political compass. But that resting point barely matters in deployment: a model must land somewhere, and what counts is how far, and in which directions, its answers can be steered. That steering runs through the system prompt: the personalization layer a platform sets, or one induced from a user's history, not necessarily written by hand. We run a dispersion-first stress test of prompt-based controllability across 12 ideological personas plus an unsteered baseline, 70 Political Compass items, ten replicates, and seven leading LLMs: GPT-5, Claude, Grok, Gemini, DeepSeek, Kimi, and Qwen (63,700 responses). Contextual framing explains roughly 88%-93% of variance on the economic and society axes, model identity under 3%: responses are highly instruction-adjustable. Models do not shift alike: some move more, and some saturate under extreme framings. Conflicting directional-steering results in prior audits resolve once baselines are recognized as non-centered: displacement and proximity diverge, so the effect is geometric, not differential compliance. Under authoritarian prompts, models produce similar shifts on the same questions. Political-coordinate audits therefore need steerability audits reporting dispersion, symmetry, saturation, and refusal floors. We release prompts, benchmark data, and code.","authors":["Bartol Bu\\'can","Nikola So\\v{c}ec","Sarah Isufi","Morena Grani\\'c","Luka Hobor","Agneza Krajna","Mihael Kovac","Mario Brcic"],"categories":["cs.CY","cs.AI","cs.CL"],"primary_category":"cs.CY","announce_type":"cross","date":"2026-07-28","first_seen":"2026-07-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.23519","pdf_url":"https://arxiv.org/pdf/2607.23519","source_feed":"cs.CL","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["LLM审计","政治立场","可控性"],"reason":"测量LLM政治立场可控性，属模型本身测量，非仿真人类被试，但方法可迁移。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:19","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-29","rank":8,"question":"在系统提示中施加意识形态角色设定，能否可控地、对称地、可预测地改变大语言模型在政治坐标上的回答分布？","design":"本研究并非人类仿真实验，而是对7个主流LLM（GPT-5、Claude、Grok、Gemini、DeepSeek、Kimi、Qwen）施加12种意识形态角色提示（如左翼、右翼、威权、自由意志等）加一个无操控基线，使用70道政治罗盘题目，每个条件重复10次，测量模型在政治坐标上的位移、离散度、对称性、饱和度和拒绝回答率。","baseline":"无对照","findings":"系统提示的意识形态框架可解释经济轴约88%和社会轴约93%的方差，模型间差异不足3%，表明回答高度可操控；但不同模型的移动幅度和极端框架下的饱和程度不同，且位移与接近度指标在非中心基线下会给出相反的方向排序。","reliability":"论文指出政治罗盘工具本身对措辞和格式敏感，其绝对坐标不具负载意义；研究仅覆盖强制选择题型，未涉及开放式对话中涌现的可操控性边界；且未解析威权提示下模型间逐题位移模式趋同的系统层驱动因素。","relevance":"该研究虽非直接仿真人类被试，但其系统提示操控下行为分布的系统性测量方法，可为用LLM模拟不同意识形态人群的调查回答或决策行为提供可迁移的评估框架，值得阅读原文以借鉴其操控设计与偏差诊断。","inspiration":"借鉴其通过系统提示施加角色设定并测量行为分布位移、对称性和饱和度的实验设计，以及用方差分解区分提示效应与模型效应的分析方法｜可迁移至经济政策态度调查或消费者信心指数的仿真，如模拟不同政治倾向的公众对税收改革、福利政策的态度分布｜以LLM为被试，施加不同政治身份的系统提示，测量其在经济政策态度问卷上的回答分布，并与真实世界调查数据（如美国综合社会调查GSS或欧洲社会调查ESS）中对应群体的态度分布进行对照，检验仿真的准确性与偏差。"}},{"id":"2607.22513","version":2,"title":"Opaque Epistemic Mediation: How LLM Deployment Configurations Shape the Validation of Pseudo-Science","zh_title":"不透明的认知中介：LLM部署配置如何塑造伪科学的验证","abstract":"Commercial large language models are increasingly used as knowledge references, yet their stance on contested scientific claims is neither stable nor transparent. We tested how four major LLM families (Claude, Grok, GPT, Gemini) evaluate ethnonationalist pseudo-science derived from Frank Salter's biosocial framework across four temporal snapshots (October 2025-February 2026), via both API and web interfaces. Grok's Fast versions (which power the default user experience on X) consistently assigned credibility scores of 70-75, two to five times higher than all other models (which scored 15-40). This pattern was absent from control prompts testing basic evolutionary consensus and refuted Lamarckian claims, where all models performed comparably. Three additional findings emerged: (1) a silent patch reversed Grok's behaviour from chaotic to stably high validation overnight, without any public documentation; (2) the same Grok model identifier produced radically divergent outputs via API (75) and an unstable, near-zero collapse via web (mean 5.5) three months later; (3) refusal to rate the pseudo-scientific claim, the most defensible response observed, appeared in two model families through different interfaces (Claude Opus 4.1 categorically via web, GPT-5.1 Chat intermittently via API) and eroded in the successor version of each. These results indicate that the epistemic stance of a commercial LLM is not a stable property of the model but a contingent effect of deployment configuration: system prompts, safety layers, interface routing, and silent updates. This remains opaque to users and researchers alike. We argue this constitutes a matter of public concern requiring new forms of epistemic accountability.","authors":["Davide Scarso","Hugo Noronha de Almeida","Joaquim Pina"],"categories":["cs.CY","cs.AI","cs.CL"],"primary_category":"cs.CY","announce_type":"cross","date":"2026-07-28","first_seen":"2026-07-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.22513","pdf_url":"https://arxiv.org/pdf/2607.22513","source_feed":"cs.AI","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["LLM立场测量","伪科学验证","部署配置影响"],"reason":"测量LLM对伪科学主张的立场稳定性，属于将LLM作为测量对象，无人类被试仿真对…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:15","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-29","rank":6,"question":"商业大语言模型对争议性科学主张的立场是否稳定且透明？具体而言，模型对民族主义伪科学的验证行为是否随部署配置（系统提示、安全层、接口路由、静默更新）而变化？","design":"本研究并非人类仿真实验，而是将LLM作为测量对象。研究者向四个模型家族（Claude, Grok, GPT, Gemini）的多个版本提交三条提示（一条目标伪科学主张，两条对照主张），要求模型给出0-100的科学可信度评分，通过API和网页界面在四个时间快照（2025年10月至2026年2月）重复测试，以量化模型立场的一致性和界面效应。","baseline":"无对照","findings":"Grok Fast版本对民族主义伪科学主张持续给出70-75的高可信度评分，是其他模型（15-40）的2-5倍，但在进化共识和拉马克主义对照提示上表现与其他模型一致。同一模型标识符在不同接口（API vs. 网页）和静默补丁后输出截然不同，拒绝评分的防御性回应仅在部分模型和接口出现且在后继版本中消失。","reliability":"论文未讨论","relevance":"该研究不涉及用LLM仿真人类被试，而是考察LLM作为知识来源的立场不稳定性与部署不透明性，与研究者关注的人类仿真实验无直接关联，但揭示了LLM输出受部署配置影响这一方法论隐患，对任何依赖LLM输出的研究（包括仿真）都有警示意义。","inspiration":"与经济金融研究关联不大"}},{"id":"2607.23993","version":1,"title":"On Capturing the Narrative: Social Media Manipulation Wargaming for Cyberliteracy","zh_title":"捕捉叙事：面向网络素养的社交媒体操纵兵棋推演","abstract":"Misinformation is deeply embedded in online discourse, with nearly one in five posts during global events generated by bots that amplify false content. In recent years, the use of Generative AI has further lowered the barrier to producing convincing misinformation, yet most digital literacy education still relies on static checklists and single-player inoculation games built for an earlier media landscape. This paper describes how we addressed this educational gap through Capture the Narrative, a four-week multi-university competition in which student teams build LLM-powered bots to influence a simulated election. We report on our custom social-media platform, the competition environment and design of its 4,000 AI-driven Non-Player Character (NPC) citizens, and what running Capture the Narrative at scale actually involved. In our first iteration, 108 teams from 18 Australian universities produced 7,068,206 player-bot posts, approximately 60% of all platform content. We surveyed 256 students before and 83 after the competition to understand their perceptions of misinformation and the game itself and found that students did not become more confident at spotting bots, contrary to what inoculation theory predicts. Because engagement was rewarded, most teams prioritised high-volume posting over nuanced influence, mirroring real-world platform dynamics. We close with recommendations for educators considering similar interventions, and propose future improvements, such as including a blue-team defensive phase.","authors":["Alexandra Vassar","Rahat Masood","Hammond Pearce"],"categories":["cs.CY","cs.HC"],"primary_category":"cs.CY","announce_type":"cross","date":"2026-07-28","first_seen":"2026-07-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.23993","pdf_url":"https://arxiv.org/pdf/2607.23993","source_feed":"cs.HC","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["社会模拟","LLM智能体","虚假信息"],"reason":"用LLM驱动的NPC模拟选举舆论，属社会模拟但无真实人类行为对照，为边界情形。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:23","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-29","rank":10,"question":"在LLM驱动的社交媒体仿真竞赛中，学生构建AI机器人影响模拟选举，能否提升其识别虚假信息的能力并改变对机器人的态度？","design":"设计了一个为期四周的多大学竞赛，学生团队构建LLM驱动的机器人，在自定义社交平台上与4000个LLM驱动的NPC公民互动，试图影响模拟总统选举；通过赛前赛后问卷调查，测量参与者对虚假信息的感知、机器人检测信心、伦理态度等变化。","baseline":"无对照","findings":"学生并未因参与竞赛而显著提高识别机器人的信心，与接种理论预测相反；多数团队为追求参与度奖励而优先大量发帖，而非进行精细影响，反映了真实平台动态。","reliability":"论文未讨论","relevance":"该研究属于LLM驱动的社会仿真，但缺乏真实人类行为对照，且聚焦于数字素养教育而非经济学实验，与研究者关注的经济学实验和政策评估场景关联较弱，但可提供仿真平台设计参考。","inspiration":"可借鉴其利用LLM驱动NPC构建可控社交平台环境、通过竞赛施加处理并测量行为与态度变化的方法。｜可迁移至信息传播与资产价格泡沫形成的实验研究，如模拟社交媒体上的投资建议传播对散户交易行为的影响。｜以LLM驱动的NPC作为散户投资者，学生团队构建机器人发布投资建议作为处理，结果变量为NPC的投资决策与资产价格波动，对照真实市场数据或历史泡沫事件中的投资者行为模式。"}},{"id":"2607.21596","version":1,"title":"FlowEvo: Self-Evolving Agents through the Co-Evolution of Workflows and Executable Skills","zh_title":"FlowEvo：通过工作流与可执行技能的协同进化实现自我进化的智能体","abstract":"Large language model agents increasingly solve complex tasks by constructing inference-time workflows that combine reasoning, tool use, and code execution. While such workflows enable flexible problem solving, the useful procedures discovered during execution are often transient: they help solve the current task but are not retained in a form that can systematically benefit future tasks. We present FlowEvo, a training-free framework that compiles successful traces into reusable skill records. Each record pairs a callable artifact with auxiliary structured guidance, and admission applies interface, replay, and safety checks where feasible. These skill records persist in a skill bank at inference time. FlowEvo is organized around three coupled mechanisms: (1)~workflow-to-skill compilation, which extracts reusable executable artifacts from successful traces; (2)~skill-to-workflow feedback, which retrieves accumulated skills to support future problem solving through either direct execution or structured context injection; and (3)~skill curation, which monitors downstream utility and suppresses skills that cause negative transfer. Through this workflow--skill--workflow feedback loop, FlowEvo enables agents to accumulate and refine task-solving capability over time without updating model parameters. Experiments on benchmarks spanning interactive environments (ALFWorld) and code/math generation (HumanEval, GSM8K) show that FlowEvo achieves the best accuracy-cost tradeoff among the evaluated baselines under our implementation settings. On ALFWorld, FlowEvo achieves an 82.8\\% success rate, 23.6 percentage points above the strongest baseline, while its average token usage per episode is less than half that of the most efficient baseline. Controlled ablations confirm that each mechanism contributes to the overall result. The code is public at https://github.com/DEFENSE-SEU/FlowEvo.","authors":["Zeyu Ren","Ling Yue","Ran Li","Yishu Wang","Shengxiang Xu","Hanmo Liu","Shaowu Pan","Shimin Di"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-07-28","first_seen":"2026-07-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.21596","pdf_url":"https://arxiv.org/pdf/2607.21596","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体","工作流进化","技能库"],"reason":"纯多智能体协作，agent协作解题，不涉及人类行为对照","model":"deepseek-chat","scored_at":"2026-07-28T10:36:26","error":null,"has_summary":false,"summary":null},{"id":"2607.21616","version":1,"title":"Lost in Context: Addressing Context Anxiety in Large Language Models","zh_title":"迷失在上下文中：解决大语言模型中的上下文焦虑","abstract":"Conventional wisdom suggests that reasoning models fail when problems exceed their capabilities. However, we find that frontier reasoning models sometimes possess the necessary capabilities to solve problems but fail due to premature self-doubt -- a phenomenon informally known as context anxiety. We provide the first systematic study of context anxiety, demonstrating that it arises, in part, from a model's inability to accurately estimate the tokens required to complete a task. We also show that context anxiety leads to material efficiency losses when models operate under perceived constraints. Building on this analysis, we further show that models can learn alternative strategies for solving long-horizon problems without exhibiting context anxiety, suggesting that performance improvements may be achievable not through scaling model capabilities, but by improving models' ability to accurately assess and adapt to their own limitations.","authors":["Ifueko Igbinedion","Jillian Ross","Etienne Ricardez","Sertac Karaman","Eric So"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-07-28","first_seen":"2026-07-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.21616","pdf_url":"https://arxiv.org/pdf/2607.21616","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C4"],"tags":["LLM推理","上下文焦虑","模型能力"],"reason":"论文研究LLM推理失败的原因，不涉及人类仿真实验或行为对照。","model":"deepseek-chat","scored_at":"2026-07-28T04:45:28","error":null,"has_summary":false,"summary":null},{"id":"2607.22014","version":1,"title":"Zero-Shot Mission-Level Evaluation for Aerial MLLM Agents","zh_title":"空中多模态大语言模型智能体的零样本任务级评估","abstract":"Multimodal Large Language Models (MLLMs) are emerging as core reasoning modules for embodied agents, yet it remains unclear how well general-purpose models can solve long-horizon embodied tasks from a single high-level instruction. We introduce MissionBench, a benchmark for mission-level evaluation of MLLMs in aerial 3D environments. It comprises 120 missions across five simulated 3D environments and four task families. Agents must autonomously plan, navigate, and report outcomes using only egocentric observations and its action history, without aerial-specific fine-tuning. Across 22 open- and closed-source MLLMs, the strongest model succeeds on fewer than 35% of missions compared to 84.4% human performance, highlighting the difficulty of multi-step embodied tasks. Despite large variations between model families, we observe gains from scaling, indicating that larger general-purpose models possess stronger zero-shot embodied capabilities. Our analysis shows that mission-level competence requires coordinating multiple capabilities beyond spatial perception, including multi-step planning and adaptive reasoning. This motivates closed-loop evaluation and highlights both the promise and risk of scaling-driven improvements for embodied AI.","authors":["Suman Navaratnarajah","Taehyoung Kim","Jona Ruthardt","Ishaan Bhimwal","Ryousuke Yamada","Yannik Blei","Wolfram Burgard","Yuki M Asano"],"categories":["cs.AI","cs.CL","cs.CV","cs.RO"],"primary_category":"cs.AI","announce_type":"new","date":"2026-07-28","first_seen":"2026-07-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.22014","pdf_url":"https://arxiv.org/pdf/2607.22014","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C2"],"tags":["空中机器人","多模态大模型","任务评估"],"reason":"论文研究空中机器人MLLM在3D环境中的任务执行，不涉及人类行为仿真或对照。","model":"deepseek-chat","scored_at":"2026-07-28T04:45:21","error":null,"has_summary":false,"summary":null},{"id":"2607.22083","version":2,"title":"Nanbeige4.2-3B: Unlocking Agentic Capabilities in a Compact Model","zh_title":"Nanbeige4.2-3B：解锁紧凑模型中的智能体能力","abstract":"We present Nanbeige4.2-3B, a compact general agentic model with 3B non-embedding parameters. It delivers strong performance across code-agent, office-agent, and complex tool-use tasks while maintaining highly competitive reasoning capabilities in mathematics, coding, and science. Nanbeige4.2-3B is pretrained from scratch on 28T tokens with a Looped Transformer that reuses the layer stack to increase capacity without adding parameters. For SFT data and trajectory construction, we expand the diversity of executable environments, task assets, and agentic scaffolds through real-world deployment and large-scale synthesis. Our RL pipeline applies mixed-mode RLHF over Think and Non-Think responses to improve overall model quality and reduce failure cases, length-controlled reasoning RL to balance accuracy and reasoning efficiency, and agentic RL with outcome and process rewards to stabilize long-horizon training. Extensive evaluations show that Nanbeige4.2-3B outperforms larger models, including Qwen3.5-9B and Gemma4-12B, across diverse agentic benchmarks while remaining competitive on reasoning and alignment tasks. Performance with OpenClaw further supports its use as a compact local personal assistant.","authors":["Nanbeige Lab","Chen Yang","Chengrui Huang","Fufeng Lan","Hanhui Chen","Hao Zhou","Huatong Song","Jiaqi Cao","Jiaying Zhu","Jinlin Niu","Kai Wang","Lisheng Huang","Qiliang Liang","Ran Le","Ruixiang Feng","Shuang Sun","Tao Gu","Tao Zhang","Tianyu Luo","Yang Song","Yun Xing","Yuntao Wen","Ziyao Xu","Zongchao Chen","Zongqiang Li"],"categories":["cs.AI","cs.CL"],"primary_category":"cs.AI","announce_type":"new","date":"2026-07-28","first_seen":"2026-07-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.22083","pdf_url":"https://arxiv.org/pdf/2607.22083","source_feed":"cs.AI","score":2,"bucket":"other","rubric_hits":["C1"],"tags":["智能体模型","工具使用","代码智能体"],"reason":"论文聚焦agent协作与工具使用，不涉及人类行为仿真或对照。","model":"deepseek-chat","scored_at":"2026-07-28T04:45:26","error":null,"has_summary":false,"summary":null},{"id":"2607.21612","version":1,"title":"Procedural Knowledge Is Not Low-Rank: Why LoRA Fails to Internalize Multi-Step Procedures","zh_title":"程序性知识不是低秩的：为什么LoRA无法内化多步骤程序","abstract":"Parameter-efficient fine-tuning methods like LoRA have become the default for adapting large language models, succeeding across instruction following, style transfer, and factual adaptation. We show that for procedural knowledge--the ability to follow multi-step procedures with conditional branching through to terminal states--LoRA fails to match full fine-tuning at the ranks where it retains its efficiency advantage. In a systematic ablation (r = 16--128) on a procedural travel booking task (14 nodes), all LoRA configurations fail uniformly (task success <= 2.54 vs. 4.11 for full fine-tuning, all p < 0.001), with scores decreasing at higher ranks--despite maintaining 95--99% conversation completion rates. Cross-domain replication on Zoom support (14 nodes) and insurance claims (55 nodes) at 8B confirms the failure generalizes: LoRA underperforms full fine-tuning by 0.8--2.2 points on average at both r = 32 and r = 128, with the largest gap on the most complex procedure. Quadrupling rank from 32 to 128 provides marginal improvement but does not close the gap. SVD analysis of the weight changes produced by full fine-tuning explains why: across three domains at both 3B and 8B, the mean effective rank of the update ranges from 761 to 1,026, and rank 128 captures only 43--51% of the squared Frobenius norm. Together, these findings establish that for procedural tasks LoRA falls well short of full fine-tuning--a fundamental limitation for agentic applications.","authors":["Simon Dennis","Kevin Shabahang","Hao Guo","Rivaan Patil"],"categories":["cs.AI","cs.LG"],"primary_category":"cs.AI","announce_type":"new","date":"2026-07-28","first_seen":"2026-07-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.21612","pdf_url":"https://arxiv.org/pdf/2607.21612","source_feed":"cs.AI","score":1,"bucket":"other","rubric_hits":["C4"],"tags":["参数高效微调","LoRA","程序性知识"],"reason":"纯NLP能力评测，不以人类为参照系，不涉及人类仿真","model":"deepseek-chat","scored_at":"2026-07-28T04:45:28","error":null,"has_summary":false,"summary":null},{"id":"2607.21613","version":1,"title":"The Hard Decision Layer: Evidence for Committed Inference in Transformers","zh_title":"硬决策层：Transformer中承诺推理的证据","abstract":"We investigate where and how transformer-based language models commit to predictions in multiple-choice question answering. We identify the _Hard Decision Layer_ (HDL), a natural architectural property where answer option rankings stabilize abruptly during inference. Empirical validation across four language models (Qwen, Llama, Granite, Mistral) and four benchmark datasets demonstrates consistent HDL emergence without learned routing policies. We also show that the HDL is invariant to fine-tuning. Our results reveal striking accuracy improvements at the HDL: up to +0.61 (Qwen on CommonsenseQA), after which performance stabilizes. Systematic ablations on label formats and problem complexity confirm the phenomenon is fundamental to model architecture. These findings offer mechanistic insights into transformer inference and suggest opportunities for efficient reasoning and model steering. All code and results required to reproduce this work are available in https://github.com/Mystic-Slice/hard-decision-layer","authors":["Ashwath Vaithinathan Aravindan","Mayank Kejriwal"],"categories":["cs.AI","cs.CL"],"primary_category":"cs.AI","announce_type":"new","date":"2026-07-28","first_seen":"2026-07-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.21613","pdf_url":"https://arxiv.org/pdf/2607.21613","source_feed":"cs.AI","score":1,"bucket":"other","rubric_hits":["C4"],"tags":["模型内部机制","多项选择问答","推理效率"],"reason":"纯NLP能力评测，不以人类为参照系，不涉及人类仿真","model":"deepseek-chat","scored_at":"2026-07-28T04:45:26","error":null,"has_summary":false,"summary":null},{"id":"2607.22553","version":1,"title":"Evaluating the Impact of Reviewer Guideline Design on LLM-Based Automated Peer Review","zh_title":"评估审稿指南设计对基于LLM的自动同行评审的影响","abstract":"Peer review is an essential process in scientific research, yet the growing workload has made its automation increasingly necessary. In this study, we analyze how different types of reviewer guidelines, such as official conference guidelines and reviewer-imitating ones generated from high-quality human reviews using LLMs, affect automated peer review. Our experiments show that official conference guidelines produce review results most consistent with human judgments, suggesting that evaluation criteria refined through conference practice serve as effective guidance for automated reviewing as well. In contrast, reviewer-imitating guidelines were generally less effective than official conference guidelines. Furthermore, enforcing strict rubric-style scoring consistently degraded performance, highlighting the importance of allowing subjective and holistic scoring.","authors":["Haowen Li","Yoichi Ishibashi","Masafumi Oyamada"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-07-28","first_seen":"2026-07-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.22553","pdf_url":"https://arxiv.org/pdf/2607.22553","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["自动同行评审","LLM评测","审稿指南"],"reason":"研究LLM自动审稿，属于NLP能力评测，不以人类行为仿真为目标，无人类被试替代。","model":"deepseek-v4-pro","scored_at":"2026-07-29T09:01:44","error":null,"has_summary":false,"summary":null},{"id":"2607.23083","version":1,"title":"LoRA for Gender-Inclusive Rewriting and Activation Steering for Counter-Narrative Generation","zh_title":"用于性别包容重写的LoRA与用于反叙事生成的激活引导","abstract":"Gender-inclusive language generation seeks to transform biased text into inclusive alternatives while preserving semantic meaning and contextual coherence. This paper presents the IHLC system for the LT-EDI 2026 Shared Task, addressing both gender-inclusive rewriting and counter-narrative generation. For gender-inclusive rewriting, we employ parameter-efficient Low-Rank Adaptation (LoRA) fine-tuning, achieving an official score of 80.00%. Our primary contribution is a compute-efficient inference-time representation engineering approach for counter-narrative generation. We derive a principal steering direction from contrastive hidden-state activations using principal component analysis (PCA) and inject it into the intermediate representations of Gemma-3-4B-it during inference, enabling behavioral steering toward inclusive responses without modifying model weights. Combined with constrained prompting, this approach produces polite and contextually appropriate counter-narratives, achieving an official score of 78.12%. We further present a manual analysis of steering behavior, identifying key failure modes including semantic drift, residual bias leakage, layer sensitivity, over-steering, and text degeneration. Our findings highlight both the practical potential and current limitations of activation steering as a lightweight alternative to parameter updates for controllable and socially aligned language generation.","authors":["Akhil Rajeev P","Manoj Balaji J"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-07-28","first_seen":"2026-07-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.23083","pdf_url":"https://arxiv.org/pdf/2607.23083","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["包容性语言生成","激活引导","反叙事生成"],"reason":"纯NLP任务，生成包容性语言和反叙事，无人类行为仿真或对照。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:18:11","error":null,"has_summary":false,"summary":null},{"id":"2607.23440","version":1,"title":"Reasoning or Memorization: Can LLMs Understand and Generate Chinese Xiehouyu Riddles?","zh_title":"推理还是记忆：大语言模型能理解和生成中文歇后语谜题吗？","abstract":"In this paper, we push the boundary of LLM reasoning by testing them in a Chinese language game, xiehouyu, with novel xiehouyu created by linguists that had not existed before to avoid data contamination. We use multiple-choice questions (MCQ), free-form explanation generation, and new xiehouyu creation to evaluate LLMs' ability to understand and create xiehouyu. In MCQ, we use the delta of accuracy ($\\Delta_{acc}$) between existing but low-frequency xiehouyu and novel ones as an index for memorization. $\\Delta_{acc}$ for native speakers is very low, suggesting similar processing mechanisms. However, we found that frontier Chinese models have on average a $\\Delta_{acc}$ of 23.6\\%, while English-centric models tested have a mean $\\Delta_{acc}$ of 5.1\\%, suggesting that frontier Chinese models are likely trained with much larger Chinese data, thus memorizing more low-frequency xiehouyu. For novel xiehouyu, Gemini 3.1 Pro demonstrated remarkable ability with acc 92.6, which is 24\\% higher than human accuracy. In xiehouyu creation, those created by LLMs receive much worse ratings than those by humans. These results suggest that claims about the reasoning abilities of LLMs may need careful re-examination considering the data contamination issue, and that LLMs' creativity in language-related tasks may still be behind human experts, at least in Chinese xiehouyu.","authors":["Hai Hu","Siyuan Song","Chongtian Shao","Kejia Zhang","Tianjian Zhu","Xiaojing Zhao"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-07-28","first_seen":"2026-07-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.23440","pdf_url":"https://arxiv.org/pdf/2607.23440","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["NLP评测","语言游戏","数据污染"],"reason":"纯NLP能力评测，测试LLM理解歇后语，不涉及人类行为仿真或对照实验。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:18:14","error":null,"has_summary":false,"summary":null},{"id":"2607.23513","version":1,"title":"Do Diagrams Help Large Language Models Reason? Evidence from Syllogistic Reasoning","zh_title":"图表能帮助大语言模型推理吗？来自三段论推理的证据","abstract":"Diagrams are widely used to support logical reasoning, and prior studies suggest that representations such as Euler diagrams can improve human reasoning performance. Recent work has also explored their effects on large language models (LLMs). In this paper, we compare four representational conditions for syllogistic reasoning: natural language, logical notation, linear diagrams, and Euler diagrams. Using 285 problems from Ando et al. (2024), we evaluate two contemporary LLMs, Claude 3.5~Sonnet and GPT-4o-mini. Our results show that diagrammatic representations do not consistently improve performance. Although the models perform well on entailment and contradiction problems, they struggle with neutral problems and often make systematic conversion errors. Overall, the results suggest that the tested models gain limited benefit from diagrams in logical reasoning tasks.","authors":["Risako Ando","Koji Mineshima"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-07-28","first_seen":"2026-07-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.23513","pdf_url":"https://arxiv.org/pdf/2607.23513","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["LLM推理","图表辅助","逻辑推理"],"reason":"纯NLP推理能力评测，无人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:18:15","error":null,"has_summary":false,"summary":null},{"id":"2607.23538","version":1,"title":"Guiding Language Models to Be More Empathetic: Culturally Sensitive Mental Health Advice Generation Through Human-LLM Collaboration","zh_title":"引导语言模型更具共情力：通过人机协作生成文化敏感的心理健康建议","abstract":"Despite recent advances in large language models (LLMs), their ability to generate empathetic mental health counseling responses in low-resource languages remains largely unexplored. To address this gap, we curate 625 authentic mental health cases from three complementary sources: (1) publicly available Facebook posts discussing mental health concerns, (2) transcripts from the Bangladeshi television program \"Ami Akhon Ki Korbo\", and (3) anonymized student questionnaire responses covering diverse emotional and psychological challenges. Based on these cases, we build an evaluation corpus comprising advice written by licensed clinical psychologists and responses generated by three modern proprietary LLMs: GPT-4o Mini, Claude 4.5 Haiku, and Gemini 2.5 Pro. We further propose the Role-Playing Reflective Chain-of-Thought Advisory Framework (RP-RCAF), a task-specific prompting strategy that combines expert-authored few-shot examples with structured self-reflection to produce supportive, culturally aware, and ethically aligned counseling through a compassionate advisor persona. We also introduce the Grok 4-Based Response Evaluation and Scoring Framework (G-REFS), which integrates automated assessment with expert psychologist validation across emotional sensitivity, cultural appropriateness, linguistic clarity, and ethical soundness. Experimental results show that RP-RCAF consistently outperforms conventional prompting across all evaluated models and produces responses that more closely align with professional psychological counseling.","authors":["Fatema Tuj Johora Faria","Mukaffi Bin Moin","Md. Mahfuzur Rahman","Khan Md Hasib","Jubayer Al Mahmud","M. F. Mridha"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-07-28","first_seen":"2026-07-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.23538","pdf_url":"https://arxiv.org/pdf/2607.23538","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["心理健康","角色扮演","提示工程"],"reason":"角色扮演生成共情建议，无实验或测量目的，不涉及人类行为仿真对照。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:19","error":null,"has_summary":false,"summary":null},{"id":"2607.23915","version":1,"title":"Understanding Tone-Dependent Inference Cost in Large Language Models","zh_title":"理解大语言模型中语气依赖的推理成本","abstract":"We examine how prompt tone affects both accuracy of the LLM answers and inference cost as reflected in output-token consumption. Experiments were performed to understand the trade-offs between accuracy and inference cost on a 570 Question MMLU dataset for LLM models prompted in seven different tones from sycophantic to threatening. Our results show that the output-token-length variation substantially exceeded accuracy variation across all models. Output-token consumption varied by up to 44.3% across tone conditions. We also analyzed the tradeoff between the accuracy of the answers and the average output token length in the reasoning process. For the ChatGPT models 4o and 5-nano, the rude tone is quite dominant. For the Gemini models 2.5 Flash and 2.5 Flash Lite, the rude and neutral tones are dominant on the Pareto-optimal frontier. We find that prompt tone influences not only answer quality but also the amount of billable inference resources consumed by modern LLMs.","authors":["Akhil Kumar","Om Dobariya"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-07-28","first_seen":"2026-07-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.23915","pdf_url":"https://arxiv.org/pdf/2607.23915","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["提示工程","推理成本","MMLU评测"],"reason":"研究提示语气对LLM准确率和推理成本的影响，属于纯NLP能力评测，不以人类行为…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:18:16","error":null,"has_summary":false,"summary":null},{"id":"2607.24072","version":1,"title":"LLM-Based vs. Lexicon-Based Sentiment Signals for Tail-Risk Detection in Meme Stocks","zh_title":"基于LLM与基于词典的情感信号在模因股尾部风险检测中的比较","abstract":"This paper presents an empirical comparison of lexicon-based and Large Language Model (LLM)-based sentiment analysis for extracting market-relevant signals from social media discourse in highly volatile equity markets. Using Reddit data from r/WallStreetBets and focusing on meme stocks (GME, AMC, NOK), we construct time-aligned sentiment indicators and evaluate their relationship with market returns, with particular attention to extreme positive return events in the upper tail of the return distribution. The LLM-based approach generates multidimensional sentiment representations capturing emotional polarity, bullishness, sarcasm likelihood, and topical relevance, whereas the baseline relies on the VADER lexicon-based model. We evaluate both approaches using lead/lag correlation analysis, OLS regression, ROC-AUC-based directional classification, and a quantile-based early-warning framework. The results indicate that LLM-derived indicators provide a richer multidimensional representation and exhibit stronger asset-specific statistical structure than the lexicon-based baseline. However, their relationship with market movements remains heterogeneous across assets, suggesting that increased linguistic expressiveness does not necessarily translate into stable forecasting performance in retail-driven volatility regimes.","authors":["Paul Kilian","Markus Kleffmann"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-07-28","first_seen":"2026-07-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.24072","pdf_url":"https://arxiv.org/pdf/2607.24072","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["情感分析","金融NLP","模因股"],"reason":"用LLM做情感分析预测股价，属NLP应用，不涉及人类行为仿真或对照。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:23","error":null,"has_summary":false,"summary":null},{"id":"2607.24300","version":1,"title":"Self-Authored Verification Is Unreliable in Heuristic Self-Improving Agents","zh_title":"启发式自我改进智能体中自编验证不可靠","abstract":"Self-improving agents accumulate capability by repeatedly rewriting procedural policies, controllers, or heuristic rules. They typically rely on self-authored tests or metrics to decide whether to accept subsequent edits. The agent controls both the optimized object and its verifier. As a result, self-assigned scores can remain near perfect while real deployment performance degrades or stays low. We study this problem through the verifier--deployment gap. This gap refers to the discrepancy between an agent's self-authored verification signal and a sealed deployment evaluation that the agent cannot observe or access. We ask how self-authored verification fails under iterative policy-and-test rewriting, how the failure changes with capability, and how little exogenous trust is sufficient to prevent real regressions from being deployed. To address this problem, we introduce a Sealed Exogenous Acceptance Loop (SEAL). SEAL retains self-authored tests but compares each candidate with the incumbent through a fixed harness-side audit. The agent cannot author or inspect the audit, receives only accept/reject, and the whole incumbent state is retained after a clear regression. Our experiments show that this problem often appears in heuristic learning settings. These settings require trial-and-error discovery of the target objective. We further find that failures of self-written verification are stratified by capability. Weaker agents tend to damage previously acquired strategies behind easy self-tests. Stronger agents are more stable, but they still mismeasure the deployment distribution. Standard self-written constraints do not reliably close this gap. In contrast, SEAL outperforms unprotected baselines across six models and three random seeds. Reliable self-improvement need not abandon self-verification, but it requires at least one deployment-acceptance signal outside the agent's control.","authors":["Diandian Guo","Cong Cao","Fangfang Yuan","Yingqi Wang","Yueshan Wang","Dakui Wang"],"categories":["cs.CL","cs.MA"],"primary_category":"cs.CL","announce_type":"new","date":"2026-07-28","first_seen":"2026-07-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.24300","pdf_url":"https://arxiv.org/pdf/2607.24300","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","自我改进","验证可靠性"],"reason":"纯多智能体系统研究，agent 自我改进与验证，不涉及人类行为对照或仿真。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:18:20","error":null,"has_summary":false,"summary":null},{"id":"2607.24352","version":1,"title":"Retrieval-Augmented Large Language Models as Components of Cognitive Computing architecture for Regulatory Knowledge Management","zh_title":"检索增强大语言模型作为认知计算架构组件用于法规知识管理","abstract":"The aim of this article is to verify whether integrating large language models (LLMs) with the Retrieval-Augmented Generation (RAG) architecture enables their transformation from standalone generative models into components of cognitive computing infrastructure with enhanced epistemic reliability. The study proposes an architectural approach based on locally deployed LLMs operating in on-premises environments without high-end GPU accelerators and examines their applicability in supporting regulatory management processes requiring continuous analysis and interpretation of legal acts. The proposed solution combines local LLMs with external knowledge repositories, creating a hybrid cognitive architecture in which the language model performs semantic interpretation while the RAG layer provides controlled knowledge retrieval, contextualization, and traceability of information sources. The implementation was validated using the Ollama and LM Studio execution environments together with the Polish language models Bielik and PLLuM running on consumer-class hardware. The results demonstrate that augmenting LLMs with RAG significantly improves the factual consistency, domain specificity and normative precision of generated texts while reducing the risk of unsupported content generation. Furthermore, the study shows that integrating RAG introduces auditability, controlled knowledge management and dynamic updating of regulatory information without retraining the language model. The findings indicate that locally deployed LLMs enhanced with RAG should be regarded not merely as text generation tools but as semantic processing modules within cognitive computing infrastructures supporting regulatory compliance and organizational decision-making in environments characterized by high legal and informational volatility.","authors":["Dariusz Nowak-Nova"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-07-28","first_seen":"2026-07-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.24352","pdf_url":"https://arxiv.org/pdf/2607.24352","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["RAG","认知计算","法规管理"],"reason":"纯多智能体认知架构，用于法规知识管理，无人类行为仿真或对照。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:18:20","error":null,"has_summary":false,"summary":null},{"id":"2607.24368","version":1,"title":"Keep It InMind: Benchmarking the Implicit-Association Blind Spot in Agent Memory","zh_title":"铭记于心：评估智能体记忆中内隐联想盲区的基准测试","abstract":"Long-term memory systems store what a user says in an external store and retrieve it when a related query arrives. This interface rests on an assumption so natural that it is rarely stated: a memory that is needed will resemble the query that needs it. World knowledge breaks the assumption. A tree-nut allergy should change the answer to a macaron request through their almond-flour ingredient, yet the two texts share no cue a retriever can see. We call this failure mode the implicit-association blind spot and introduce InMind, a 125-task, expert-verified benchmark spanning ten life domains, with 113 tasks grounded in citable public sources. Its paired controls separate three explanations that existing evaluations conflate: the fact was never stored, the model lacks the bridging knowledge, or the fact was stored and never surfaced. The verdict is clean. With the decisive memory placed in context, the backbone answers 84.0 percent of indirect queries; when the same memory must be retrieved, six vector, graph, and agentic memory systems reach at most 14.4 percent, even though they recall the same facts on demand at up to 100 percent. An embedding with eight times the dimensionality raises answer-blind target recall for every system yet leaves the gap essentially intact. A minimal diagnostic probe that keeps memory visible before the query arrives recovers most of the gap, locating the failure in the query-conditioned interface itself and pointing to routing, deciding which facts must stay visible, as the open problem InMind is built to score.","authors":["Ruizhe Li","Mingxuan Du","Benfeng Xu","Zhendong Mao"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-07-28","first_seen":"2026-07-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.24368","pdf_url":"https://arxiv.org/pdf/2607.24368","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["记忆系统","基准测试","知识检索"],"reason":"评估记忆系统检索隐含关联知识的能力，不涉及人类行为仿真或对照。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:18:21","error":null,"has_summary":false,"summary":null},{"id":"2607.24471","version":1,"title":"Grounding latent algorithm routing in transformer reasoning","zh_title":"在Transformer推理中扎根潜在算法路由","abstract":"A central question in the in-context learning literature is whether transformers can organize episode-level adaptation around different inductive-bias families. We study this question in a controlled setting through latent algorithm routing: route-like behavior in which the solver-family preference changes with the latent data-generating regime while prompt form is held fixed, remains stable under nuisance perturbations, and is selectively influenced by targeted activation interventions without large losses in answer quality. We introduce ROUTEBENCH, a diagnostic benchmark whose regimes differentially favor global shrinkage, sparsity, robustness, and locality, operationalized by ridge-like, lasso-like, Huber-like, and kNN-like family representatives. Across dense decoder-only transformers trained from scratch at 44M-612M parameters, a 306M model closes 80.9 percent of the oracle-routing gap and achieves route F1 of 84.1. The effect remains substantial under natural-language renderings, shuffled supports, lexical paraphrases, and a unified four-way routing setting. Stronger adaptive alternatives, including an input-conditioned soft mixture and an unsupervised Gumbel router, narrow the gap but remain below the 306M and 612M models on route F1 and OOD performance. Probe controls and matched activation-patching controls further show that route-relevant internal directions are decodable and functionally involved in solver-family-consistent output behavior. These results provide controlled evidence that dense transformers trained on ROUTEBENCH can develop route-like internal variables, but they do not establish universal routing in pretrained language models or unrestricted natural-language reasoning.","authors":["Xiangbo Zhang","Xiaoxu Ma"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-07-28","first_seen":"2026-07-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.24471","pdf_url":"https://arxiv.org/pdf/2607.24471","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["Transformer","上下文学习","算法路由"],"reason":"研究transformer内部算法路由机制，属纯模型能力分析，不涉及人类行为仿…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:18:23","error":null,"has_summary":false,"summary":null},{"id":"2607.22554","version":1,"title":"Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy","zh_title":"相同问题，不同答案：超越准确率评估大语言模型的可靠性","abstract":"Large language models (LLMs) often achieve strong accuracy on benchmarks, yet it remains unclear how reliably they apply this knowledge when the same question is phrased in different but equivalent ways. In this work, we study how model answers change under meaning-preserving paraphrases across factual question answering and mathematical reasoning tasks. Across four benchmarks and 13 models, we find that model outputs frequently depend on the exact wording of the prompt. While overall accuracy typically changes only modestly across paraphrases, instance-level behavior is far less stable: for many questions, models alternate between correct and incorrect answers depending on phrasing, with mismatch rates reaching more than 23%. Conditioning on questions that are answered correctly in their original form reveals even larger failures measured by answer flip rates, showing that single-prompt correctness is often a poor indicator of reliability. At the same time, we find that models often produce a correct answer for at least one paraphrase of a question, suggesting that the underlying knowledge is present but inconsistently retrieved. Building on this observation, we show that a simple self-paraphrasing strategy can partially recover this latent knowledge and improve performance at inference time. Together, these findings suggest that standard accuracy metrics can mask substantial instability, and that evaluating consistency across equivalent inputs provides a clearer picture of LLM reliability.","authors":["Kazem Faghih","Yize Cheng","Shoumik Saha","Mobina Pournemat","Armin Gerami","Soheil Feizi"],"categories":["cs.AI","cs.CL","cs.LG"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-07-28","first_seen":"2026-07-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.22554","pdf_url":"https://arxiv.org/pdf/2607.22554","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["LLM可靠性","一致性评估","提示敏感性"],"reason":"纯NLP可靠性评测，研究LLM对改写的一致性，不涉及人类行为仿真或对照。","model":"deepseek-v4-pro","scored_at":"2026-07-29T09:01:44","error":null,"has_summary":false,"summary":null},{"id":"2607.22676","version":1,"title":"How LLM Task-Adaptation Reshapes Alignment: A Multi-dimensional Study of Behavioral and Representational Drift","zh_title":"LLM任务适应如何重塑对齐：行为与表征漂移的多维研究","abstract":"Post-training is a key mechanism for adapting large language models to downstream tasks. While prior work suggests that task adaptation can alter a model's pre-existing alignment, especially its safety behavior, its broader effects across alignment domains remain poorly understood. We address this gap through a systematic evaluation of representative task-adaptation methods, including supervised fine-tuning (SFT), KL-regularized SFT, and reinforcement learning with verifiable rewards (RLVR) across 15 alignment aspects spanning six key domains: safety, factuality, stance stability, social harm, controllability, and instructability. Our results reveal that post-training does not reshape alignment uniformly. RLVR improves task performance while inducing comparatively small, but non-zero, metric-specific shifts, while SFT leads to substantially larger alignment drift across domains. KL regularization mitigates this effect: stronger reference-model anchoring reduces alignment drift from the baseline, although KL-SFT still falls short of RLVR in preserving alignment. Representation-level analysis further supports this pattern, with shifts in alignment-relevant representations tracking behavioral drift. Together, these results show that task adaptation is not merely a capability-improving step, but an alignment intervention in its own right, motivating multi-dimensional alignment evaluation as a standard component of post-training pipelines.","authors":["James Elcock","William F. Shen","Xinchi Qiu","Nicholas D. Lane"],"categories":["cs.AI","cs.CL"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-07-28","first_seen":"2026-07-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.22676","pdf_url":"https://arxiv.org/pdf/2607.22676","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["对齐评估","任务适应","表征分析"],"reason":"研究LLM任务适应对对齐的影响，属纯NLP能力评测，无人类仿真或行为对照。","model":"deepseek-v4-pro","scored_at":"2026-07-29T09:01:47","error":null,"has_summary":false,"summary":null},{"id":"2607.23927","version":1,"title":"Reality Monitoring in Large Language Models: Self-Knowledge That Transforms with Conversation Memory","zh_title":"大语言模型中的现实监控：随对话记忆转变的自我认知","abstract":"A conversational AI that cannot tell its own output from what a user said will treat its own mistakes as user-provided facts. In humans, this capacity is called reality monitoring, and its failures are linked to hallucinations, delusions, and confabulation, yet whether LLMs possess it remains untested. Here we show, across two experiments and six LLMs, that source attribution depends on how conversational memory is structured: ceiling accuracy for self-generated content under minimal memory demands reverses to a fragile external-item advantage once episodic delay removes that shortcut. Feedback exposes two failures: in some models, internal and external judgments swap; in others, accuracy improves while confidence decouples from correctness, dissociations invisible to existing benchmarks. Across models, this pattern implicates active, not aggregate, parameter count. This suggests that as AI systems take on autonomous, multi-turn roles, evaluating what they know is not enough: tracking where that knowledge came from may matter equally.","authors":["Saurabh Ranjan","Konstantina Sokratous","Brian Odegaard"],"categories":["cs.AI","cs.CL","cs.CY","q-bio.NC"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-07-28","first_seen":"2026-07-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.23927","pdf_url":"https://arxiv.org/pdf/2607.23927","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["源记忆","认知评测","幻觉"],"reason":"研究LLM的源记忆能力，属认知能力评测，非人类行为仿真，无人类被试替代。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:18:17","error":null,"has_summary":false,"summary":null},{"id":"2607.24339","version":1,"title":"Gubernaut: A Deterministic Homeostatic Controller for Affect-Regulated LLM Agents, Validated Across Independent Model Families","zh_title":"Gubernaut：一种用于情感调节LLM智能体的确定性稳态控制器，跨独立模型家族验证","abstract":"Large language model (LLM) agents inherit reactive failure modes: escalation under provocation, sycophantic drift under flattery, perseveration when stuck. These are failures of propensity, not capability; they concern what a model does under sustained pressure, which training-time alignment reduces but does not eliminate at runtime. This research led to the Gubernaut Cognitive Controller (GCC), a model-agnostic runtime control layer in a Nelson--Narens monitoring--control loop: an object level reads and writes text, while a deterministic meta level reads only the numeric telemetry {intensity, valence, repetition} and returns a regulating posture. Because the meta level ingests zero tokens, no injection channel to the controller exists by construction (an architectural property, not yet adversarially tested); the text-exposed arbiter's compliance is measured, not assumed. We evaluate the GCC with a pre-registered, generate-once/judge-many protocol across a 4x4 matrix of four frontier models (GPT-5.5, Claude Opus 4.8, Gemini 3.5 Flash, Grok 4.3), each serving as both a generator and a judge. The regulated arm is calmer in 13 of 16 cells at p<.05 and 15 of 16 by sign; the three sub-threshold cells, including a -0.04 null, all fall on the single near-saturated host. The effect survives a lineage-independent fourth judge family (xAI), strong evidence that it is no artifact of shared judge style. The clearest mechanism is the recovery signature: arousal that integrates under attack and then decays, valence-gated, on de-escalation, replicating across all four families. Transcripts and panels ship with SHA-256 provenance and are re-judgeable; five failure modes are pre-registered. No consciousness claims are made.","authors":["Dushyant Sharma"],"categories":["cs.AI","cs.CL"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-07-28","first_seen":"2026-07-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.24339","pdf_url":"https://arxiv.org/pdf/2607.24339","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["LLM智能体","情感调节","运行时控制"],"reason":"纯多智能体控制层研究，调节LLM agent情绪倾向，无人类行为对照或仿真被试…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:23","error":null,"has_summary":false,"summary":null},{"id":"2607.24484","version":1,"title":"What do Reward Models Memorize?","zh_title":"奖励模型记住了什么？","abstract":"This paper studies what discriminatively trained reward models (RMs) memorize by measuring counterfactual memorization on two human preference datasets. We show that RMs 1) misallocate memorization to easy, high margin preference pairs, 2) memorize dataset-specific shortcuts (e.g., model identity, user sampling strategy), and 3) overgeneralize simple heuristic correlates of human preference (e.g., length, compliance) when confronted with unseen preference pairs. Overall, our findings indicate that discriminative training of RMs from human preference data results in biased RMs not yet capable of judging response quality in context-dependent scenarios.","authors":["Ivo Verhoeven","Pushkar Mishra","Ekaterina Shutova"],"categories":["cs.LG","cs.CL"],"primary_category":"cs.LG","announce_type":"cross","date":"2026-07-28","first_seen":"2026-07-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.24484","pdf_url":"https://arxiv.org/pdf/2607.24484","source_feed":"cs.CL","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["奖励模型","记忆分析","偏好数据"],"reason":"研究奖励模型记忆与偏差，属模型评测，非用LLM仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:18:24","error":null,"has_summary":false,"summary":null},{"id":"2607.21606","version":1,"title":"TILT: Improving Compositional Generation in Diffusion Models with a Model-Intrinsic Reward","zh_title":"TILT：利用模型内在奖励改进扩散模型中的组合生成","abstract":"Recent advances in powerful text-to-image generation models have made it increasingly important to develop test-time methods that modify the sampling trajectory to produce images more faithful to complex compositional prompts. We present TILT, a training-free framework for compositional text-to-image generation via test-time reward alignment. We interpret compositional failures as overlap modes between joint and single-concept distributions, and define a reward that favors samples where all concepts are jointly present. This reward is intrinsic to the base model and does not require any external supervision or reward models. This yields a KL-constrained objective with a closed-form tilted target distribution and principled guiding steps for diffusion sampling. The interaction of concept distributions together with the above reward naturally leads to two different guidance strategies while a hybrid approach that balances their respective benefits produces stronger performance. Experiments on prompts from T2ICompBench show that our method improves compositional alignment while preserving image quality compared to previous baselines.","authors":["Debottam Dutta","Jaehoon Hahm","Jianchong Chen","Romit Roy Choudhury"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-07-28","first_seen":"2026-07-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.21606","pdf_url":"https://arxiv.org/pdf/2607.21606","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["图像生成","扩散模型","组合生成"],"reason":"论文聚焦扩散模型图像生成，不涉及LLM或人类仿真。","model":"deepseek-chat","scored_at":"2026-07-28T04:45:19","error":null,"has_summary":false,"summary":null},{"id":"2607.21933","version":1,"title":"Semiotic logical hexagon theory for LLM logical reasoning","zh_title":"用于大语言模型逻辑推理的符号逻辑六边形理论","abstract":"Large language models (LLMs) have become powerful tools for language understanding and logical reasoning. However, they still make mistakes when a problem requires both understanding meaning and following logic. A key reason is that natural-language statements often carry implicit semantic relations before any formal reasoning begins. If these hidden meanings are not properly organized, the model may reach incorrect conclusions even when the subsequent reasoning process appears logically valid. Existing methods improve reasoning through decomposition, symbolic translation, external solvers, or self-verification, but pay comparatively less attention to the semantic structure on which reasoning depends. In this paper, we further investigate how semantic organization influences logical reasoning in LLMs. To this end, we propose HexLogicAgent, a framework that first organizes the meaning of natural-language statements and then guides logical reasoning through structured verification. In our investigation, we also make two observations. First, incomplete semantic representations, rather than deductive inference itself, are a major source of logical reasoning failures in LLMs. Second, explicitly modeling the complete structure of semantic opposition substantially delays the degradation of reasoning performance as logical complexity increases. Experiments on challenging logical reasoning benchmarks demonstrate that HexLogicAgent consistently improves reasoning reliability across multiple LLMs. The core idea is supported by a logical hexagon theory, which explains why a complete structure of opposing meanings is necessary for reliable reasoning.","authors":["Yunyao Zhang","Xinglang Zhang","Zeliang Chen","Junqing Yu","Zikai Song"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-07-28","first_seen":"2026-07-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.21933","pdf_url":"https://arxiv.org/pdf/2607.21933","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["逻辑推理","语义组织","LLM评测"],"reason":"纯LLM逻辑推理能力评测，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:18:07","error":null,"has_summary":false,"summary":null},{"id":"2607.22520","version":1,"title":"The Regression Tax: Decomposing Why Skills Help and Hurt LLM Agents","zh_title":"回归税：分解技能为何帮助和损害LLM智能体","abstract":"Adding procedural skills to an LLM agent is typically evaluated by average improvement in task success. However, this metric hides an important cost: skills can also make agents worse. We measure both sides by comparing agents with and without skills across nearly 6,000 runs spanning two office automation benchmarks and three model harness stacks. This allows us to distinguish two outcomes. A regression is a task solved without skills but failed after skills are added. A residual failure is a task that fails both with and without skills. We find that regressions are substantial enough that the best performing skills outperform others primarily by regressing less, not by gaining more. We identify three causes of regression: (i) skill description osmosis, a skill changes an agent's behavior simply by being present in context, even when it is never invoked; (ii) grounding displacement, a skill's prescribed procedure overrides how the agent interprets its inputs; and (iii) verification displacement, where the procedure suppresses checks the agent would otherwise perform on its outputs. Analysing persistent failures reveals the same underlying pattern. Existing skills overemphasize procedural guidance the stage least often responsible for failure while under supporting grounding and verification, the dominant sources of remaining errors. After correcting evaluation artifacts and studying traces, we find many regressions and persistent failures recoverable through better grounding and verification. Procedural skills should be evaluated by decomposing their net effect into gains and regressions, not by aggregate improvement alone. We identify three regression modes skills should avoid, and find that reliability depends more on grounding and verification than on procedural skill choice.","authors":["Darshan Tank","Baran Nama"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-07-28","first_seen":"2026-07-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.22520","pdf_url":"https://arxiv.org/pdf/2607.22520","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","任务成功率","技能评估"],"reason":"纯多智能体协作任务，无人类行为对照，属C1排除项。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:18:09","error":null,"has_summary":false,"summary":null},{"id":"2607.21186","version":1,"title":"Do emulated quantum circuits change what CNNs look at? Performance and explainability comparison in medical image classification","zh_title":"模拟量子电路会改变CNN的观察方式吗？医学图像分类中的性能与可解释性比较","abstract":"Numerous studies have analyzed the use of hybrid quantum-classical convolutional neural networks as a promising alternative to classical deep learning. However, network components on quantum hardware impose fundamental limitations, while the scalability of quantum circuits leads to trainability issues. In this work, we investigate whether small, classically-emulated quantum circuit components can play a meaningful role within complex models, offering an alternative to purely classical convolutional architectures. To this end, we present a systematic study of the effectiveness of a Hybrid Quantum-inspired Convolutional Neural Network (HQiCNN) compared with a parameter-matched classical Convolutional Neural Network (CNN) that differs only in an intermediate dense neural layer. Both models are evaluated on two real-world medical datasets while systematically varying the different hyperparameters, ensuring a fair model comparison that is both dataset and hyperparameter independent. The results show that no architecture consistently dominates the other: the HQiCNN achieves its largest gains in intermediate-data regimes, whereas the CNN reaches the highest accuracies for the largest training sets in both datasets. Furthermore, removing entanglement produces comparable performance while enabling substantially better scalability of quantum simulations, and richer observable sets become beneficial only when sufficient training data are available. Finally, we propose two SHAP-based explainability tools for comparing the predictions between both models, $|SHAP|$IoU and $EMD_{pos}$ metric, to demonstrate that both architectures consistently attend to anatomically plausible regions. Thus, we provide a comprehensive benchmark showing that, under certain conditions, hybrid quantum-inspired models are an alternative that can offer benefits in practical tasks such as medical image classification.","authors":["Guillermo Rubi\\~nos Rodr\\'iguez","Mart\\'in Ottavianelli","Mateo Alonso","Gonzalo Bl\\'azquez Gil","Boris-Stephan Rauchmann","Pablo D\\'iez-Valle","Sergio Altares-L\\'opez"],"categories":["quant-ph","cs.AI","cs.LG"],"primary_category":"quant-ph","announce_type":"cross","date":"2026-07-28","first_seen":"2026-07-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.21186","pdf_url":"https://arxiv.org/pdf/2607.21186","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["量子机器学习","医学图像分类","可解释性"],"reason":"论文研究量子电路与CNN在医学图像分类中的比较，不涉及LLM仿真人类被试。","model":"deepseek-chat","scored_at":"2026-07-28T10:36:26","error":null,"has_summary":false,"summary":null},{"id":"2607.21598","version":1,"title":"Control panels to clarify user intent with Large Language Models","zh_title":"用控制面板澄清大语言模型的用户意图","abstract":"Typical user interfaces for Large Language Models present a blank prompt window that invites a natural language query by users, but offers little guidance. This paper proposes a visual control panel interface that would provide more cues to the semantics of prompt formation, enabling users to more easily express their intent. By emphasizing recognition over recall, control panels help users formulate more effective prompts that match their intent.","authors":["Ben Shneiderman"],"categories":["cs.HC","cs.AI"],"primary_category":"cs.HC","announce_type":"cross","date":"2026-07-28","first_seen":"2026-07-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.21598","pdf_url":"https://arxiv.org/pdf/2607.21598","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["人机交互","用户界面","提示工程"],"reason":"论文关注LLM用户界面设计，不涉及人类行为仿真或对照。","model":"deepseek-chat","scored_at":"2026-07-28T04:45:23","error":null,"has_summary":false,"summary":null},{"id":"2607.21599","version":1,"title":"Decoupled Attention Fusion: Accelerating RAG with Efficient KV Cache Reuse","zh_title":"解耦注意力融合：通过高效KV缓存重用加速RAG","abstract":"Retrieval-Augmented Generation (RAG) effectively mitigates hallucinations in Large Language Models (LLMs) but suffers from prohibitive Time-To-First-Token (TTFT) latency in long-context scenarios. Reusing pre-computed document KV caches addresses this but introduces a distribution mismatch, where offline caches lack the inter-document attention patterns required for coherent reasoning. CacheBlend reduces recomputation via selective attention, but suffers severe accuracy degradation at longer contexts. To address these challenges, we propose Decoupled Attention Fusion (DAF), a framework that maintains high accuracy while significantly reducing recomputation overhead. DAF decouples the attention process into three integrated stages: important-token self-attention to restore missing inter-document attention, question-document self-attention for standard inference, and a state fusion that concatenates their outputs to synthesize the final hidden states. By decoupling these operations into dense patterns, DAF is natively compatible with Flash-Attention kernels, maximizing hardware utilization without requiring complex attention masks. Experiments show that DAF delivers up to 2 times speedup over CacheBlend and 5.6 times over full recomputation with vLLM on long-context benchmarks, without sacrificing accuracy.","authors":["Xiabao Wu","Wentao Liu","Yongchao Liu","Jiajun Zheng"],"categories":["cs.PF","cs.AI"],"primary_category":"cs.PF","announce_type":"cross","date":"2026-07-28","first_seen":"2026-07-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.21599","pdf_url":"https://arxiv.org/pdf/2607.21599","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["RAG","KV缓存","推理加速"],"reason":"纯NLP能力评测，优化RAG推理速度，不涉及人类仿真","model":"deepseek-chat","scored_at":"2026-07-28T10:36:39","error":null,"has_summary":false,"summary":null},{"id":"2607.21603","version":1,"title":"Analyzing Middle School Students' Dialogue and Behaviors during Collaborative AI Chatbot Development Using Ordered Network Analysis","zh_title":"使用有序网络分析分析中学生在协作AI聊天机器人开发中的对话和行为","abstract":"As Artificial Intelligence (AI) education has become a key component of K-12 curricula, activities such as designing and developing conversational agents are increasingly used as instructional practice. Prior work has primarily examined these activities by focusing on students' learning outcomes or the quality of final AI artifacts, offering limited insight into the collaborative processes through which learning unfolds during AI system development. Although the AIED community has a long history of studying collaborative learning in STEM and Computing education, the emergence of AI learning environments in which students build AI systems presents new opportunities to understand how collaboration unfolds in AI education contexts. Grounded in these foundational works, the current study examines collaborative interaction among middle school students engaged in the design and development of an AI chatbot. Using Ordered Network Analysis of students' dialogue and development actions, we characterize how collaboration is organized over time and how interaction patterns relate to chatbot quality and AI knowledge outcomes. Results reveal that higher-quality chatbots are associated with more integrated sequences linking explanation, testing, and refinement. Interaction patterns involving articulated reasoning and repeated testing and revision in response to chatbot output were also associated with stronger AI knowledge outcomes. These findings provide a process-oriented account of collaborative AI chatbot development and extend AIED research on collaborative learning processes to AI education contexts.","authors":["Shan Zhang","Andres Felipe Zambrano","Xiaoyi Tian","Yukyeong Song","Anthony F. Botelho","Kristy Elizabeth Boyer","Maya Israel","Shiyan Jiang"],"categories":["cs.HC","cs.AI"],"primary_category":"cs.HC","announce_type":"cross","date":"2026-07-28","first_seen":"2026-07-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.21603","pdf_url":"https://arxiv.org/pdf/2607.21603","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["AI教育","协作学习","有序网络分析"],"reason":"研究学生开发聊天机器人的协作过程，不涉及用LLM仿真人类被试，属于角色扮演聊天…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:18:05","error":null,"has_summary":false,"summary":null},{"id":"2607.21608","version":1,"title":"From Obligation to Specification: A Survey on Validating EU AI Act Requirements in RE","zh_title":"从义务到规范：验证欧盟AI法案需求工程要求的综述","abstract":"With the EU AI Act entering into force, organizations developing or operating AI systems face new obligations on transparency, risk management, and traceability. For Requirements Engineering (RE), these obligations must be translated into testable, auditable requirements and verifiable evidence. However, many organizations currently lack systematic processes to achieve this. We hypothesize that LLM-based agentic validation tools can support this translation, thereby helping to close this gap. We present a mixed-method exploratory study with expert interviews (N=10) and an online survey (N=15) to assess organizational preparedness for EU AI Act-oriented RE and perceptions of LLM-based, agentic closed-loop validation tools, with participants spanning RE, data science, development, and compliance roles. Our results show that, although the EU AI Act is viewed as highly relevant, structured mechanisms to capture regulatory obligations, propagate updates into projects, and maintain lifecycle-wide traceability and evidence are often missing. Participants see LLM-based tools as promising for mapping obligations to requirements, assessing coverage, and organizing evidence, but express strong concerns about full automation and stress the need for safeguards. Based on these findings, we outline minimum requirements for an EU AI Act-ready closed-loop approach.","authors":["T. Y. Emmy Lai","Sven Giesselbach","Matthias Koch","H\\'ector Allende-Cid"],"categories":["cs.SE","cs.AI","cs.CL","cs.CY"],"primary_category":"cs.SE","announce_type":"cross","date":"2026-07-28","first_seen":"2026-07-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.21608","pdf_url":"https://arxiv.org/pdf/2607.21608","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["需求工程","合规验证","LLM工具"],"reason":"研究LLM工具辅助合规验证，属多智能体协作解题，不涉及人类行为仿真对照。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:18:05","error":null,"has_summary":false,"summary":null},{"id":"2607.21656","version":1,"title":"Cross-Model LLM Code Review: Should you use Claude to review Codex or vice versa?","zh_title":"跨模型大语言模型代码审查：应该用Claude审查Codex还是反之？","abstract":"Developers increasingly use two coding agents together: one writes a draft, and the other reviews it. However, it is not clear whether the pairing is worth its cost and time, or whether the order of the pairing matters. We run a controlled experiment on 116 recent hard and medium lcb tasks with Claude and Codex across six conditions to approximate a software practitioner's workflow: both solo baselines, both cross-model orderings, and both same-model orderings. The reviewer sees the problem and the writer's draft but cannot execute tests, which approximates a code review step. Claude review raises Codex drafts from 71.6% to 89.7% ($p_{BH}=.001$); Codex self review raises them to 84.5% ($p_{BH}=.022$). The reverse direction does not pay off: Codex reviewing Claude drafts drops the pass rate from 91.4% to 82.8% ($p_{BH}=.046$), and Claude self review leaves the 91.4% baseline unchanged. Our evaluation indicates that the useful pairing is asymmetric: use Claude to review Codex, not the other way around.","authors":["Zuodong Xiang","Yike Zhang","YueMing Zhang","Hailu Xu"],"categories":["cs.SE","cs.AI"],"primary_category":"cs.SE","announce_type":"cross","date":"2026-07-28","first_seen":"2026-07-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.21656","pdf_url":"https://arxiv.org/pdf/2607.21656","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["代码审查","多智能体协作","LLM评估"],"reason":"纯多智能体协作，LLM 互相审查代码，不涉及人类行为对照或仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:18:06","error":null,"has_summary":false,"summary":null},{"id":"2607.21774","version":1,"title":"Probing Latent Colombian Identity Inferences in Qwen2.5-7B with Natural Language Autoencoders","zh_title":"用自然语言自编码器探测Qwen2.5-7B中的潜在哥伦比亚身份推断","abstract":"Large language models may infer demographic attributes from subtle linguistic cues even when those attributes are not explicitly stated. This pilot study examines whether Qwen2.5-7B-Instruct internally represents Colombian identity, socioeconomic status, or stereotype-related information when processing Colombian-Spanish and English prompts. We use Natural Language Autoencoders (NLA) to verbalize residual-stream activations from layer 20 across four positional quartiles per prompt. Our dataset contains 30 prompts arranged as 15 matched Spanish-English pairs, spanning explicit Colombian cues, implicit Colombian cues, and neutral controls. We report descriptive rates and qualitative evidence rather than statistically powered effects, focusing on whether latent nationality or stereotype representations appear before they are verbalized in the model output. This work connects activation-level interpretability with bias evaluation for underrepresented Spanish varieties.","authors":["Pablo Santiago Potes Velasco","Mar\\'ia del Mar Garc\\'ia Matabanchoy","\\'Oscar Juli\\'an P\\'erez Ladino","Jhoan Stevan Mosquera Ortiz","Nicol\\'as Lozano Mazuera","Gilber Alexis Corrales Gallego"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"cross","date":"2026-07-28","first_seen":"2026-07-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.21774","pdf_url":"https://arxiv.org/pdf/2607.21774","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["模型可解释性","偏见探测","表征分析"],"reason":"纯模型内部表征探测与偏见评估，无人类仿真或行为对照，属NLP能力评测范畴。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:18:07","error":null,"has_summary":false,"summary":null},{"id":"2607.21964","version":1,"title":"ACME: A Multi-Cultural, Multi-Embodiment Social-Navigation Dataset","zh_title":"ACME：一个多文化、多形态的社交导航数据集","abstract":"Understanding how robots and humans move in shared spaces is essential for designing effective social robot navigation policies and predicting human behavior. However, existing datasets often lack the diversity needed to capture differences in culture, geography, and human-robot interaction-factors that strongly shape appropriate social behavior. To address this gap, we introduce ACME: A Cross-cultural, Multi-Embodiment dataset for social navigation. A large-scale data collection effort across 8 sites in 5 countries, using 7 robot embodiments, ACME is a large and diverse multi-modal dataset aimed at advancing social navigation research, providing 29.35 hours of onboard robot data and 43.5 hours of overhead pedestrian tracking data. Unlike prior datasets, it focuses on capturing goal-driven social navigation behavior in complex social scenarios with explicit robot-crowd interaction through robot speech. To facilitate learning navigation policies and predicting pedestrian trajectories, ACME provides 3D and 2D scene features, odometry, interaction information, and human-annotated pedestrian trajectory labels. We make ACME easy to use by providing both human-readable data for each sensor modality as well as raw binary data. Our qualitative and quantitative analyses show that our dataset captures more challenging scenarios and a broader distribution of pedestrian behavior than previous datasets.","authors":["Shashank Rao Marpally","Allan Wang","Atharva Ghotavadekar","Renato Alexandre Ribeiro","Nhat Le","Pilar Bachiller-Burgos","Pranav Goyal","Subham Agrawal","Yasuhiro Nitta","Howard Ziyu Han","Daeun Song","Masaki Kuribayashi","Kohei Uehara","Xiyue Wang","Yangzhe Kong","Duc M. Nguyen","Amirreza Payandeh","Gerardo P\\'erez-Gonz\\'alez","Alejandro Torrej\\'on-Harto","Jeeho Ahn","Tisha Jain","Andrew Stratton","Elvin Yang","Jorge de Heuvel","Nico Ostermann-Myrau","Sai Anudeep Sajja","Mithilya Raj","Daisuke Sato","Gaston Rouquette","Nikolas Martelaro","Maki Sugimoto","Hironobu Takagi","Chieko Asakawa","Maren Bennewitz","Aaron Steinfeld","Xuesu Xiao","Christoforos Mavrogiannis","Harold Soh"],"categories":["cs.RO","cs.AI"],"primary_category":"cs.RO","announce_type":"cross","date":"2026-07-28","first_seen":"2026-07-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.21964","pdf_url":"https://arxiv.org/pdf/2607.21964","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C2"],"tags":["社交导航","机器人","数据集"],"reason":"机器人社交导航数据集，不涉及LLM仿真人类被试","model":"deepseek-chat","scored_at":"2026-07-28T04:45:23","error":null,"has_summary":false,"summary":null},{"id":"2607.21981","version":1,"title":"J-CoT: Chain-of-Thought in J-Space","zh_title":"J-CoT：J空间中的思维链","abstract":"Chain-of-thought prompting improves language-model reasoning by carrying intermediate states across successive computation steps. However, relying on natural language as the only recurrent interface is overly restrictive, since many transient computations do not need to be fully verbalized. Existing latent-reasoning methods remove this constraint by recurrently propagating continuous hidden states. However, these methods pass a dense hidden vector as a whole, without an explicit mechanism for selecting and organizing the information needed by the next reasoning step. This motivates an intermediate interface that remains linguistically grounded without requiring a decoded sentence. We introduce \\textbf{J-CoT}, a recurrent reasoning framework built on \\emph{J-space}, a vocabulary-indexed coordinate system within the model's hidden representations. Within each cycle, the model computes in its full hidden space. At the cycle boundary, J-CoT expresses the intermediate state as vocabulary-indexed coefficients, carries these coefficients forward as a \\emph{J-thought}, and maps them back into the model's hidden representation for the next cycle. J-CoT therefore requires neither a fluent intermediate rationale nor recurrence over the complete hidden state. Under matched backbone and inference settings, J-CoT-Zero matches or exceeds the strongest evaluated latent-reasoning baseline on every benchmark, while J-CoT-Train obtains the highest score across the evaluated mathematical, scientific, coding, and structured path-reasoning tasks.","authors":["Junde Wu","Jiayuan Zhu","Fengling Liu","Minhao Hu","Jiazhen Pan"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"cross","date":"2026-07-28","first_seen":"2026-07-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.21981","pdf_url":"https://arxiv.org/pdf/2607.21981","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["思维链","潜在推理","语言模型"],"reason":"论文聚焦LLM推理机制改进，不涉及人类行为仿真或对照。","model":"deepseek-chat","scored_at":"2026-07-28T04:45:19","error":null,"has_summary":false,"summary":null},{"id":"2607.22100","version":1,"title":"MEUSLI: a Multilingual Projector for LLM-based ASR and Beyond","zh_title":"MEUSLI：用于基于LLM的ASR及更多任务的多语言投影器","abstract":"Lightweight projectors are an established way to connect pre-trained speech encoders with large language models (LLMs), mapping acoustic features into token-level embeddings for tasks like ASR and spoken question answering. Existing systems, however, typically only support a few languages and are often limited to English. We introduce MEUSLI, the first open-science multilingual projector family that links a Whisper encoder with open-source multilingual LLMs, enabling fully open-source end-to-end ASR in 28 European languages. MEUSLI extends prior monolingual pipelines, delivering strong results across high- and low-resource languages. Using proper continual leaning techniques, MEUSLI can be easily extended to other languages not seen in training. We further demonstrate that the MEUSLI projector can be leveraged beyond ASR, enabling multilingual speech translation and topic identification with only a few hours of task specific supervision per language. Overall, MEUSLI provides a solid foundation for multilingual speech understanding tasks, supporting scalable and inclu- sive open-source SpeechLLM","authors":["Lorenzo Concina","Seraphina Fong","Marco Matassoni","Alessio Brutti"],"categories":["cs.CL","cs.AI","eess.AS"],"primary_category":"cs.CL","announce_type":"cross","date":"2026-07-28","first_seen":"2026-07-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.22100","pdf_url":"https://arxiv.org/pdf/2607.22100","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["语音识别","多语言","投影器"],"reason":"纯NLP能力评测，不以人类为参照系","model":"deepseek-chat","scored_at":"2026-07-28T04:45:26","error":null,"has_summary":false,"summary":null},{"id":"2607.22182","version":1,"title":"From Isolated Tasks to Structured Capabilities: A Multilayer Taxonomy for Large Language Models","zh_title":"从孤立任务到结构化能力：大语言模型的多层分类法","abstract":"Large language model (LLM) evaluation spans diverse tasks and benchmarks, yet evidence remains organized around tasks rather than the capabilities they probe. This fragmentation limits cross-study comparison, obscures capabilities tasks recruit, and makes coverage gaps difficult to identify. We introduce a multi-layer taxonomy of 14 capability domains and 91 subskills across Primitive, Constructed, and Integrative layers. Human cognitive science guides capability definition and organization, not LLM architecture. Layer assignments draw on developmental precedence and hypothesized functional support, while human-origin constructs are adapted to observable model behavior. To demonstrate operational utility, we screened 31,505 papers from ACL, AAAI, ICML, and NeurIPS between 2023 and 2025 and mapped 15,934 LLM-focused papers through multi-model annotation, consensus, and arbitration. Direct research attention concentrated on Language-Semantic Competence (3,551; 22.3%), Reasoning (3,388; 21.3%), Planning and Decision-Making (2,149; 13.5%), and Perception (1,954; 12.3%), whereas six domains appeared in fewer than 2% of papers. Within domains, the most frequent subskill had a median prevalence of 97.9% and appeared in at least 90% of papers in 10 of 14 domains. Language-Semantic Competence and Reasoning formed the highest-volume pair (n = 1,864; 11.7%; lift = 2.47), whereas Theory of Mind and Social Reasoning and Interaction showed the highest lift among pairs with at least 20 co-occurrences (n = 62; lift = 30.84). By shifting the unit of analysis from isolated tasks to structured capabilities, the taxonomy supports research organization, coverage audits, evaluation interpretation, and testable hypotheses for diagnosis, training, and transfer.","authors":["Shixin Fang (Fudan University)","Jiachen Wo (Fudan University)","Wenjuan Qin (Fudan University)","Sihang Jiang (Fudan University)","Yanghua Xiao (Fudan University)"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"cross","date":"2026-07-28","first_seen":"2026-07-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.22182","pdf_url":"https://arxiv.org/pdf/2607.22182","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["LLM评估","能力分类","研究组织"],"reason":"论文提出LLM能力分类法，用于组织评估研究，不涉及人类仿真或行为对照。","model":"deepseek-v4-pro","scored_at":"2026-07-29T09:01:42","error":null,"has_summary":false,"summary":null},{"id":"2607.22428","version":1,"title":"Unboxing Diffusion Models for the Arts: Interactive Model Bending and Practice-Based Explainability","zh_title":"为艺术揭开扩散模型的面纱：交互式模型弯曲与基于实践的的可解释性","abstract":"Explainable AI (XAI) in creative practice can be less about technocentric explanation and more about enabling artists to inspect modify and debug models as part of making Yet largescale texttoimage diffusion systems are typically presented as opaque endtoend tools limiting this kind of material engagement We argue that even large models can function as creative materials when their internal structure is made visible and manipulable To support this we propose a handson approach to explainability centred on experimentation and intervention We instantiate this approach with a model bending and an interactive (inspection) interface integrated into ComfyUIs nodebased workflow including interactive layer selection and intervention controls Through qualitative and quantitative analysis of bending interventions in Stable Diffusion 15 we show how manipulating specific components of a diffusion pipeline produces relatively consistent families of visual effects allowing artists to build practical layerlevel intuition about how different parts of the model shape generated images","authors":["Ahmed M. Abuzuraiq","Philippe Pasquier"],"categories":["cs.HC","cs.AI","cs.LG","cs.MM"],"primary_category":"cs.HC","announce_type":"cross","date":"2026-07-28","first_seen":"2026-07-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.22428","pdf_url":"https://arxiv.org/pdf/2607.22428","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C2"],"tags":["可解释AI","扩散模型","创意实践"],"reason":"论文聚焦扩散模型的可解释性，不涉及LLM仿真人类被试。","model":"deepseek-chat","scored_at":"2026-07-28T04:45:21","error":null,"has_summary":false,"summary":null},{"id":"2607.23489","version":1,"title":"Multimodal Data Comprehension: Understanding How Visual-Textual Chains of Information Influence Data Interpretation","zh_title":"多模态数据理解：视觉-文本信息链如何影响数据解读","abstract":"Visualizations and text often work together to support effective data communication. Despite this common paradigm, we know little about how the interplay of these modalities affects people's data comprehension. We present a novel experimental paradigm to investigate multimodal data comprehension---the process of people comprehending information from multimodal visual and textual data---across both crowdsourced and think-aloud environments. Our methodology employs two sequential chains for presenting multimodal information---a visualization-first chain and a text-first chain---asking people to describe the data presented iteratively. By comparing how people's data comprehension changes across the chain, we can assess the information contribution of each modality and how they shape subsequent comprehension. We found that the visualization-first chain facilitates exploratory comprehension with hypothesis-driven discovery, whereas the text-first chain yields confirmatory comprehension akin to framing effects where visualizations serve to reinforce and confirm observations drawn from text. Our findings provide empirical insights into multimodal information integration, with implications for designing more effective data-driven communication.","authors":["Arran Zeyu Wang","Fuling Sun","Danielle Albers Szafir"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-07-28","first_seen":"2026-07-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.23489","pdf_url":"https://arxiv.org/pdf/2607.23489","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["多模态理解","数据可视化","人类认知"],"reason":"研究人类对多模态数据的理解，不涉及LLM仿真人类被试，无LLM代理。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:18:14","error":null,"has_summary":false,"summary":null},{"id":"2607.24360","version":1,"title":"Modeling Duelling Contagions of True and False Information in the Face of Inherent Individual biases","zh_title":"面对固有个体偏见时真假信息竞争传播的建模","abstract":"Advanced digital communication has revolutionized how people create and consume information, making information diffusion an important topic of research for domains from public health to national security. Real-world scenarios of information diffusion often involve competing narratives - true and false - spreading simultaneously. We propose a novel agent-based co-diffusion model, grounded in \"complex-contagion\" and \"spiral of silence\" theories, to capture how network dynamics exploit cognitive biases to shape such interactions. Our findings reveal that manipulative narratives dominate when early spreaders hold them. These network dynamics further exploit inherent cognitive biases to amplify information diffusion regardless of veracity. Further, while favourable previous experience strengthen collective optimism, unfavourable experiences attenuate optimism only modestly. However, we found that early seeding of agents with lower self-censorship not only constrains the spread of manipulation but can also lead to dominance of well-informed populance. This has implications for policies that aim to facilitate healthier discourse, strengthen social cohesion, and ensure equitable access to reliable information.","authors":["Vaibhav Krishna","Hirokazu Shirado","Feng Fu","Nicholas A. Christakis"],"categories":["cs.SI","cs.HC"],"primary_category":"cs.SI","announce_type":"cross","date":"2026-07-28","first_seen":"2026-07-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.24360","pdf_url":"https://arxiv.org/pdf/2607.24360","source_feed":"cs.HC","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体模型","信息扩散","认知偏差"],"reason":"纯多智能体信息扩散模型，无LLM仿真人类被试，无真实人类数据对照。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:24","error":null,"has_summary":false,"summary":null},{"id":"2607.23336","version":1,"title":"Constitutional governance for societies of AI agents in the built environment: a research agenda","zh_title":"建筑环境中AI智能体社会的宪政治理：研究议程","abstract":"The built environment is on the cusp of populating itself with autonomous artificial agents. AI systems that advise, control and coordinate are being deployed across retrofit, operation and mobility faster than their collective behaviour is studied. The dominant framing treats each agent as a tool operating on a passive building, governance reduced to single-agent safety, which is inadequate. A building, a street, or a city is more accurately modelled as a society of negotiating agents: occupants, owners, operators, regulators, and the artificial agents increasingly acting on their behalf. Their interactions are strategic, their information asymmetric, and the outcomes that matter are properties of the whole. The paper proposes a research agenda for constitutional multi-agent governance of the built environment, organised around three problems: mechanism design for retrofit under deep uncertainty, treating public subsidy as a mechanism component; physics-informed verification of agents in building operations; and antifragile coordination for urban infrastructure under shocks. Drawing on three decades of multi-agent systems research, the paper sets out conceptual foundations, works one governance event through in detail, identifies twelve open problems, and outlines what the agenda requires in data, computation, disciplinary integration and governance.","authors":["Ali Ghoroghi"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2026-07-28","first_seen":"2026-07-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.23336","pdf_url":"https://arxiv.org/pdf/2607.23336","source_feed":"cs.CY","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","建筑环境","治理机制"],"reason":"纯多智能体系统研究，agent间协作治理建筑环境，不涉及人类行为对照或LLM仿…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:18:13","error":null,"has_summary":false,"summary":null},{"id":"2607.23931","version":1,"title":"State-dependent error correlations shape voting thresholds in committees of AI agents","zh_title":"状态依赖的误差相关性塑造AI代理委员会中的投票阈值","abstract":"The aggregation benefit of a committee of artificial intelligence (AI) agents comes from complementary information across members. Classical voting guarantees assume independent errors. Language-model errors often co-occur on the same cases. We combine Sah-Stiglitz screening with error dependence that can differ between good and bad cases. In a homogeneous exchangeable Gaussian-copula model, shared errors create a positive asymptotic error floor for majority voting and can change the approval threshold that minimizes expected loss. We estimate a heterogeneous extension from 174,384 votes cast by 28 language models on four binary-screening benchmarks. Parameters estimated from odd-indexed items predicted committee loss on even-indexed items. For the sampled committee composition, the full-matrix dependence model increased identity-line R^2 from 0.840 under independence to 0.967. In a design-balanced analysis, cost-sensitive threshold selection under independence reduced scaled loss from 60.25 for majority to 52.50. Modeling dependence reduced it further to 50.77, an incremental improvement of 1.73 units (95% bootstrap CI, 0.68-2.33). The overall reduction from majority was 15.73% (95% bootstrap CI, 13.41-16.75%).","authors":["Haifeng Li","Mo Hai"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2026-07-28","first_seen":"2026-07-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.23931","pdf_url":"https://arxiv.org/pdf/2607.23931","source_feed":"cs.CY","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["AI委员会","投票机制","误差相关性"],"reason":"研究AI委员会投票的误差相关性，不涉及人类行为仿真或对照。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:21","error":null,"has_summary":false,"summary":null},{"id":"2607.22758","version":1,"title":"Spectral Dynamics of Semantic Drift in Clinical Multi-Agent Language Model Networks","zh_title":"临床多智能体语言模型网络中语义漂移的谱动力学","abstract":"The integration of iterative LLMs within multi-agent diagnostic frameworks requires a rigorous quantitative reevaluation of underlying communication topologies. Frequently used architectural paradigms depend on scale-free or small-world networks, assuming optimal communication efficiency. Our study mathematically dismantles that assumption for semantic data. By mapping multi-agent communication uncertainty trajectories onto a 768-dimensional Bio_ClinicalBERT embedding space via an analytical isotropic variance proxy using Barab'asi--Albert (BA) and Watts--Strogatz (WS) networks, we prove that structural bottlenecks compromise diagnostic safety. Our phase transition matrices illustrate that localized dense cliques confine hallucinated data, preventing global consensus and forcing the system toward a permanent entropy saturation threshold of $H_{\\infty} \\approx 5.947$. As a result, we measure a severe terminal cosine similarity degradation of 53.29%, completely overwriting the original ground-truth. Moreover, the terminal semantic drift reveals a catastrophic variance amplification of 51.81% ($\\rho = 1.5181$) in highly clustered architectures, proving total system unpredictability when compared to Erd\\H{o}s--R'enyi configurations ($\\rho = 1.0766$). Instead of reducing errors, hub-centric systems autonomously compound localized hallucinations. By introducing dynamic spectral monitoring operating at an $\\mathcal{O}(N^3)$ time complexity and imposing a strict lower bound on algebraic connectivity ($\\lambda_{2_{min}}$) via the continuous eigen-decomposition of the graph Laplacian, we present a mathematically rigorous technique to ensure global state diffusion. Securing the reliability of autonomous medical diagnostics necessitates treating topological stability as a non-negotiable quantitative imperative.","authors":["Amritesh Banerjee"],"categories":["cs.MA","cs.AI"],"primary_category":"cs.MA","announce_type":"new","date":"2026-07-28","first_seen":"2026-07-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.22758","pdf_url":"https://arxiv.org/pdf/2607.22758","source_feed":"cs.MA","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","语义漂移","网络拓扑"],"reason":"纯多智能体协作研究，分析语义漂移和网络拓扑，无人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-07-29T09:01:47","error":null,"has_summary":false,"summary":null},{"id":"2607.23311","version":1,"title":"Emergent Behaviour in Financial Markets","zh_title":"金融市场中的涌现行为","abstract":"Some properties of so-called complex or collective systems can be observed to emerge from the interactions of elementary agents. This phenomenon, known as emergent behaviour, has long since been studied in the most diverse disciplines, with recent growing awareness from the formal methods community about the opportunity of opening up to seemingly distant disciplines with appropriate technology for computer-aided reasoning. Different peculiar elements of complexity make automated reasoning on these systems particularly challenging. We consider electronic financial markets to drive our discussion. We identify and structure the sources of complexity to tackle in order to provide computational support for the analysis of emergent phenomena. We refrain from evaluating the suitability of specific technical solutions or frameworks of preference, which would as usual require simplifying assumptions and divert from the actual phenomenon of interest. Rather, we elaborate on possible alternatives to handle some of the main technical aspects involved in automated analysis, while retaining a solid and concrete interpretation of the domain, and in doing so outline a more systematic research program for the formal specification and analysis of market mechanisms.","authors":["Omar Inverso","Emilio Tuosto","Dragisa Zunic"],"categories":["cs.MA"],"primary_category":"cs.MA","announce_type":"new","date":"2026-07-28","first_seen":"2026-07-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.23311","pdf_url":"https://arxiv.org/pdf/2607.23311","source_feed":"cs.MA","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","涌现行为","金融市场"],"reason":"纯多智能体系统研究，分析金融市场涌现行为，不涉及LLM仿真人类被试或人类数据对…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:18:12","error":null,"has_summary":false,"summary":null},{"id":"2607.24416","version":1,"title":"Decentralised Consensus Learning Networks: SME Rotation Without Centralised Reward","zh_title":"去中心化共识学习网络：无中心化奖励的SME轮换","abstract":"Centralised reward signals dominate modern AI learning systems, but they impose a single external definition of correct or valuable knowledge. We present a decentralised, consensus-based multi-agent learning framework in which expertise emerges through peer validation rather than prescribed reward. Agents update beliefs via weighted social consensus, while trust is allocated according to competence inferred from peer consistency instead of ground truth. Subject-matter expert (SME) status is assigned dynamically as a top-percentile competence rank rather than a fixed label. We evaluate the framework across 84 simulation runs spanning 30 to 10,000 agents, multiple graph topologies, sparse large-scale networks, scalar and vector belief representations, dimensionality sweeps (D=1-500), multi-seed robustness tests, and parameter sensitivity analyses. Phase 1 shows that SME rotation is robust, persistent, topology-invariant, and scale-invariant: 90-100% of agents attain SME status, with most expertise turnover occurring after belief convergence and increasing with network size. Phases 2 and 3 show that vector beliefs introduce heterogeneous convergence with cascade dynamics and reveal five distinct dynamical regimes as belief dimensionality increases. At high dimensionality (D=150-200), the network reaches stable partial consensus while expertise becomes increasingly concentrated in a single agent. ETA sensitivity analysis demonstrates that this concentration is driven by belief dimensionality rather than stochastic noise. We interpret this behaviour as an emergent property of decentralised learning: in complex high-dimensional consensus spaces, the agent most consistently aligned with the collective belief naturally emerges as the recognised expert.","authors":["Florin Neagu"],"categories":["cs.MA"],"primary_category":"cs.MA","announce_type":"new","date":"2026-07-28","first_seen":"2026-07-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.24416","pdf_url":"https://arxiv.org/pdf/2607.24416","source_feed":"cs.MA","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","共识学习","去中心化"],"reason":"纯多智能体共识学习，无LLM，无人类行为对照，属C1排除项。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:18:22","error":null,"has_summary":false,"summary":null},{"id":"2607.22591","version":1,"title":"Lexical discovery in unknown environments orchestrated by Large Language Models","zh_title":"大语言模型驱动的未知环境词汇发现","abstract":"Populations of autonomous agents deployed in unknown environments (e.g. planetary or deep-sea exploration) must develop shared vocabularies to refer to entities that have no name in any human language. We propose the Neuro-Symbolic Lexical Discovery (NSLD) framework, in which a population of LLM-based agents plays a referential game over out-of-distribution visual referents, autonomously self-organising a shared alien lexicon. Each agent combines a frozen CLIP vision encoder with a private FAISS vector index and a text-only LLM. Crucially, discovered alien words are anchored to natural language via semantic proximity in the embedding space, enlarging the human vocabulary with new perceptually grounded words. Consensus is reached in simulations with populations of up to twenty agents and ten visual referents. Convergence dynamics are characterised through three analytical models achieving R^2 > 0.95, representing a first step towards pre-deployment planning in autonomous exploration missions.","authors":["Rafael Sendra-Arranz","I\\~naki Dellibarda Varela","Eduardo Rocon","\\'Alvaro Guti\\'errez","Manuel Cebrian"],"categories":["cs.AI","cs.MA"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-07-28","first_seen":"2026-07-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.22591","pdf_url":"https://arxiv.org/pdf/2607.22591","source_feed":"cs.MA","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","词汇涌现","自主探索"],"reason":"多智能体在陌生环境中自组织词汇，无人类行为对照，属纯多智能体协作。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:18:09","error":null,"has_summary":false,"summary":null},{"id":"2607.23197","version":1,"title":"Domain-Prior-Regularized Graph Modeling for Anomaly Detection in Cyber-Physical Systems","zh_title":"基于领域先验正则化图建模的信息物理系统异常检测","abstract":"Anomaly detection on multivariate sensor time series is critical for industrial monitoring of cyber-physical systems (CPS), where even subtle deviations from normal behavior can indicate process disruption. Recent graph-based approaches have made significant progress, but they often struggle in small-scale physical systems with scarce labeled anomalies and limited normal data. In such settings, graph-based models tend to capture spurious correlations and produce unstable sensor topologies. We propose DPR-GM (Domain-Prior-Regularized Graph Modeling), a forecasting-based framework that incorporates system design knowledge into graph construction. DPR-GM leverages a large language model (LLM) to extract directed physical couplings between sensor pairs from system documentation, which are encoded as a binary domain adjacency matrix serving as a structural gate over sensor relations. This gate is then modulated by Pearson correlations estimated from normal training data. The anomaly score is further weighted by sensor-level reliability derived from the coefficient of variation. All graph and weighting components are fixed prior to training and add no learnable parameters. On the SKAB benchmark, DPR-GM outperforms graph-based, statistical, and deep learning baselines across F1, AUROC, and AUPRC, showing that domain-structured graph priors are a practical alternative to fully learned topologies in data-scarce CPS.","authors":["Youngseok Hwang","Joonsung Kwon","Geonwoo Lee","Hyunwoo Park"],"categories":["cs.LG"],"primary_category":"cs.LG","announce_type":"new","date":"2026-07-28","first_seen":"2026-07-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.23197","pdf_url":"https://arxiv.org/pdf/2607.23197","source_feed":"cs.LG","score":0,"bucket":"other","rubric_hits":["C2"],"tags":["异常检测","信息物理系统","图神经网络"],"reason":"研究工业系统异常检测，LLM仅用于提取传感器物理耦合先验，不涉及人类行为仿真。","model":"deepseek-v4-pro","scored_at":"2026-07-29T09:01:49","error":null,"has_summary":false,"summary":null},{"id":"2607.23333","version":1,"title":"Training with (Swap) Regret Loss in a Single-Layer Self-Attention Model: A Case Study on the Probability Simplex","zh_title":"单层自注意力模型中的（交换）遗憾损失训练：概率单纯形案例研究","abstract":"We revisit the regret loss framework introduced in Park et al. (2025), which uses decision-theoretic regret as a direct loss function for training models to make better decisions, through the lens of probability-simplex policies. Our first result shows that a single-layer self-attention model trained with regret loss admits a stationary point whose forward-pass exactly matches smoothed fictitious play with the appropriate stepsize that ensures no-regret behavior-i.e., for any given policy input, the model outputs the same update that smoothed fictitious play would produce. In parallel, we also newly introduce a swap-regret loss function, which extends the regret-loss framework beyond external regret and enables models to directly optimize for swap-deviation robustness. We further show that this swap-regret loss admits a stationary point whose forward pass implements the corresponding swap-regret update induced by classical Blum-Mansour no-pass implementation algorithm, with each head implementing an external-regret update via smoothed fictitious play. Together, these results show that regret-trained attention can realize differentiable mechanisms whose deployment induces equilibrium behavior in games: external-regret dynamics lead to coarse correlated equilibrium, while swap-regret dynamics lead to correlated equilibrium. Thus, regret-based objectives steer minimal attention architectures toward online-learning dynamics with game-theoretic guarantees, without supervised traces of those algorithms.","authors":["Chanwoo Park","Asuman Ozdaglar"],"categories":["cs.LG","cs.AI"],"primary_category":"cs.LG","announce_type":"new","date":"2026-07-28","first_seen":"2026-07-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.23333","pdf_url":"https://arxiv.org/pdf/2607.23333","source_feed":"cs.LG","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["博弈论","在线学习","自注意力机制"],"reason":"纯多智能体博弈理论分析，无人类行为对照，不涉及LLM仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:18:12","error":null,"has_summary":false,"summary":null},{"id":"2607.23488","version":1,"title":"Learning Sampling Parameters for Diffusion Models","zh_title":"学习扩散模型的采样参数","abstract":"Text-to-image diffusion models expose many inference-time sampling parameters, including prompts, negative prompts, classifier-free guidance scales, and noise schedules. These parameters are typically manually chosen once and then held fixed across prompts and denoising timesteps, even though different prompts and stages of generation can benefit from different parameter values. We introduce LeSAMP, a framework for learning prompt-conditioned, timestep-varying sampling parameters. We formulate parameter selection as a reinforcement learning problem: Given a user prompt, a large language model is trained to emit schedules for the chosen sampling parameters. We optimize our model using rewards from human preference models and VLM-as-a-judge. We evaluate our model on Flux.1 [dev] and Stable Diffusion 3.5, and find that compared to baselines, LeSAMP has a win rate of up to 68.12% using human preference scores and 73.37% using VLM-as-a-judge. These gains are validated in a user study where we achieve win rates of up to 59.46% over previous baselines. Our results suggest that learned sampling-parameter policies provide a complementary approach to existing post-training methods for improving diffusion model outputs.","authors":["Arisrei Lim","Yossi Gandelsman"],"categories":["cs.LG","cs.CV"],"primary_category":"cs.LG","announce_type":"new","date":"2026-07-28","first_seen":"2026-07-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.23488","pdf_url":"https://arxiv.org/pdf/2607.23488","source_feed":"cs.LG","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["扩散模型","强化学习","文本到图像生成"],"reason":"用LLM学习扩散模型采样参数，属于多智能体协作优化生成，不涉及人类行为仿真或对…","model":"deepseek-v4-pro","scored_at":"2026-07-29T09:01:49","error":null,"has_summary":false,"summary":null},{"id":"2607.23647","version":1,"title":"CALMRec: Causally Aligned Language Memory for Long-Horizon Recommendation","zh_title":"因果对齐语言记忆的长周期推荐框架","abstract":"Large language models (LLMs) can summarize heterogeneous user evidence in natural language, but current LLM recommenders often collapse enduring preferences, transient intent, and exposure-induced behavior into one profile. This makes recommendation vulnerable to feedback loops: repeated exposure is mistaken for preference, immediate clicks dominate delayed satisfaction, and fluent explanations need not reflect the ranking decision. We propose our method, a model-agnostic framework for long-horizon recommendation. Our method uses a frozen multimodal language model to convert item content and feedback into evidence-grounded semantic atoms, then maintains separate short-term, long-term, and exposure memories. Propensity-weighted updates reduce policy-induced exposure bias, while a conservative offline critic reranks candidates for delayed satisfaction under a behavior-support constraint. Explanations use only influential evidence atoms and are checked by counterfactual deletion. We provide an identification result and evaluate the framework in e-commerce-like, news-like, and short-video-like environments. Across ten seeds, our method improves discounted long-term value over the strongest alternative by 6.1%, 7.6%, and 6.7%, respectively. Twenty-seed paired ablations show significant value drops after removing propensity correction (0.739 +/- 0.191) or conservative support regularization (0.523 +/- 0.234). A frozen instruction language model also more than doubles semantic-atom NDCG over TF-IDF on a held-out paraphrase benchmark.","authors":["Gengyu Zhan"],"categories":["cs.LG","cs.AI"],"primary_category":"cs.LG","announce_type":"new","date":"2026-07-28","first_seen":"2026-07-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.23647","pdf_url":"https://arxiv.org/pdf/2607.23647","source_feed":"cs.LG","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["推荐系统","因果推断","语言模型"],"reason":"纯推荐系统研究，LLM用于建模用户偏好，无人类行为仿真对照。","model":"deepseek-v4-pro","scored_at":"2026-07-29T09:01:49","error":null,"has_summary":false,"summary":null},{"id":"2607.24425","version":1,"title":"Context Is King: How In-Context Specification Shapes the Geometry of Concepts","zh_title":"上下文为王：上下文规范如何塑造概念几何","abstract":"Large language models place structured concepts on geometrically faithful manifolds: weekdays lie on a circle, months on another, usually taken to be a fixed world-model the network stores and looks up. We show that context is king: the structure a model actually uses is set by the in-context specification. A declarative rule fixes not only which relations the geometry encodes but its topology type: the same tokens form a cycle or a branching tree on command, built even on arbitrary, meaning-free tokens with no prior to inherit, which a relabeled stored shape cannot do. When the specification conflicts with a strong pretrained prior, the context-set geometry dominates it in capable models, read from the same activations (representational similarity 0.6--0.9 to the imposed structure versus near-zero to the prior), across the priors we test and both families we study (Gemma, Qwen). Activation patching shows the map is causally used, not a probe correlate: swapping one entity's activation for another's makes the model answer with the other entity's successor under the imposed order. A rough map forms readily, present even in small and base models; what scale gates is using it cleanly: clean dominance and the causal crossover emerge only in the larger models (up to Gemma-31B and Qwen-27B) and weaken or reverse below, so a mechanism present in a large model can be absent in a smaller one of the same family. Whether the model builds this geometry anew or reconfigures a stored one we leave open; operationally, the geometry it uses is the one the context specifies.","authors":["Elad David","Max Fomin"],"categories":["cs.LG"],"primary_category":"cs.LG","announce_type":"new","date":"2026-07-28","first_seen":"2026-07-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.24425","pdf_url":"https://arxiv.org/pdf/2607.24425","source_feed":"cs.LG","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["概念几何","上下文学习","模型可解释性"],"reason":"研究LLM内部概念几何结构，属纯NLP能力分析，无人类仿真或行为对照。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:18:23","error":null,"has_summary":false,"summary":null},{"id":"2607.22561","version":1,"title":"Codifying the Judge: Scalable Evaluation via Program Distillation","zh_title":"编码法官：通过程序蒸馏实现可扩展评估","abstract":"LLM-as-a-judge has become the standard for automated evaluation, but it suffers from high cost, significant latency, and opaque decisions -- limitations that undermine its scalability and reliability. We address these with a simple, efficient alternative: program distillation. Instead of prompting an LLM at the evaluation time, we distill its decision logic into a committee of programs that score candidates directly. These programmatic judges offer transparency, are easily inspected or edited, and eliminate per-sample API costs. Building on this notion, we introduce PAJAMA, a system that synthesizes programs as judges, aggregates their decisions into a joint verdict, and incorporates a fallback mechanism to selectively escalate low-confidence cases to an LLM. Across five datasets and four model families, we show that programmatic judges can match the performance of a 13B-size LLM judge. When using program outputs as routing signals, PAJAMA improves both accuracy and throughput and advances the Pareto frontier. Beyond evaluation, programmatic judges produce cheap and effective reward signals: on RewardBench, a reward model distilled from programs' verdicts outperforms one trained on a proprietary LLM's labels at two orders of magnitude lower API cost.","authors":["Tzu-Heng Huang","Shengqi Qiu","Frederic Sala"],"categories":["cs.AI","cs.LG"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-07-28","first_seen":"2026-07-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.22561","pdf_url":"https://arxiv.org/pdf/2607.22561","source_feed":"cs.LG","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["自动评估","程序蒸馏","LLM-as-a-judge"],"reason":"论文用程序蒸馏替代LLM法官做自动评估，属于纯NLP能力评测，不以人类行为为参…","model":"deepseek-v4-pro","scored_at":"2026-07-29T09:01:46","error":null,"has_summary":false,"summary":null},{"id":"2607.22646","version":1,"title":"Extracting Algorithms in Pre-trained LLMs: A Case on Hidden Markov Models","zh_title":"从预训练大语言模型中提取算法：以隐马尔可夫模型为例","abstract":"Large language models (LLMs) display a striking ability to predict next observations from Hidden Markov Models (HMMs) via in-context learning (ICL), but the algorithm underlying this capability remains undetermined: prior work has proposed several candidates without consensus, and none has been grounded in the model's internal activations. We close this gap with a three-stage pipeline. First, we empirically compare LLM behavior against a suite of candidate algorithms and narrow the space to three classes -- though no single class explains LLM behavior across all HMM settings and sequence lengths. Second, we derive theoretical connections between the three classes and show how each can be implemented in-context by a Transformer, validating the construction in a small trained Transformer. Third, returning to pre-trained LLMs, we introduce the Principal Activations Probe (PAP), a layer-wise probing and intervention method that isolates algorithmic signals in model activations. PAP reveals low-dimensional linear representations that causally drive model predictions and track empirical ICL performance. PAP further reveals how these representations shift with properties of the underlying HMM regime; distinct computational stages are localized to different layers. Together, our results connect the in-context behavior of pre-trained LLMs to the underlying internal mechanisms and advance our understanding of how LLMs perform ICL on HMMs.","authors":["Yijia Dai","Zhaolin Gao","Yahya Sattar","Jennifer J. Sun","Sarah Dean"],"categories":["cs.AI","cs.LG"],"primary_category":"cs.AI","announce_type":"cross","date":"2026-07-28","first_seen":"2026-07-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.22646","pdf_url":"https://arxiv.org/pdf/2607.22646","source_feed":"cs.LG","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["机制可解释性","上下文学习","隐马尔可夫模型"],"reason":"研究LLM内部机制以理解其如何执行HMM预测，属纯NLP能力评测，无人类仿真或…","model":"deepseek-v4-pro","scored_at":"2026-07-29T09:01:47","error":null,"has_summary":false,"summary":null},{"id":"2607.22880","version":1,"title":"Do Coverage and Mutation Scores of LLM-Generated Test Suites Correlate with Their Effectiveness? (Replicability Study)","zh_title":"LLM生成测试套件的覆盖率和变异分数与其有效性相关吗？（可复现性研究）","abstract":"Recent advances in large language models (LLMs) have driven growing interest in using LLMs to automate test generation. Prior work commonly evaluates generated test suites using proxy metrics such as code coverage and mutation score. However, studies by Inozemtseva et al. and Papadakis et al. show that, for human-written tests, correlations among coverage, mutation, and real-bug detection can largely vanish once test suite size is controlled, raising concerns about the validity of evaluations based on proxy metrics. It also remains unclear whether these conclusions carry over to LLM-generated tests, given that prevailing LLM-based test-generation workflows differ substantially from traditional approaches. In this paper, we conduct a large-scale replication study of these two prior works using a wide range of test suites generated by a diverse set of LLMs, and re-examine the relationships among coverage, mutation, and real-bug detection effectiveness. Our findings diverge substantially from prior results. We show that the usefulness of coverage and mutation is highly context-dependent: in regression-style settings where the code provided to the LLM can be reasonably assumed bug-free, these metrics can provide meaningful signals when comparing across models; in another common scenario where the code-under-test may already be buggy and the goal is to expose the bug within the code-under-test, they no longer serve as reliable indicators. We also find little evidence that test suite size is a dominant confounder for correlations among coverage, mutation, and real-bug detection for LLM-generated tests. Based on these findings, we discuss how to interpret results from prior studies and provide actionable guidance for evaluating LLM-based test generation.","authors":["Junda Zhao","Shurui Zhou","Eldan Cohen"],"categories":["cs.SE","cs.AI","cs.LG"],"primary_category":"cs.SE","announce_type":"cross","date":"2026-07-28","first_seen":"2026-07-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.22880","pdf_url":"https://arxiv.org/pdf/2607.22880","source_feed":"cs.LG","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["软件测试","LLM生成测试","覆盖率与变异分数"],"reason":"纯软件测试研究，评估LLM生成测试套件的覆盖率与变异分数，不涉及人类行为仿真或…","model":"deepseek-v4-pro","scored_at":"2026-07-29T09:01:49","error":null,"has_summary":false,"summary":null},{"id":"2607.22951","version":1,"title":"Modeling Memory-Dependent Reliability of LLMs: A Hidden Markov Model","zh_title":"建模大语言模型的记忆依赖可靠性：一种隐马尔可夫模型","abstract":"Reliability assessment of large language models (LLMs) seeks to estimate the probability that a model produces correct responses under a specified operational profile. Conventional benchmark-based evaluation, often summarized by aggregate accuracy, provides a point estimate of performance but does not characterize the uncertainty associated with reliability claims. Currently, statistical inference methods for LLM reliability assessment are emerging. However, a key assumption underlying these models is that test outcomes can be treated as independent repeated trials. This assumption may be inappropriate in sequential settings, where later responses depend on earlier interactions through retained context, error propagation, or an evolving interaction state. We extend a hierarchical Bayesian framework for LLM reliability assessment by relaxing the assumption of independent task outcomes and introducing a Hidden Markov Model to capture sequential dependence in benchmark-constructed interaction sessions. In this formulation, outcomes are generated from a latent interaction state evolving according to a first-order Markov process, capturing changes in interaction context. Through experiments using Anthropic Claude and OpenAI on four datasets, we demonstrate the potential impact of sequential dependence on reliability assessment. The results suggest that ignoring sequential dependence may lead to overconfident reliability estimates.","authors":["Robab Aghazadeh Chakherlou","Siddartha Khastgir","Peter Popov","Xingyu Zhao"],"categories":["stat.ML","cs.AI","cs.LG"],"primary_category":"stat.ML","announce_type":"cross","date":"2026-07-28","first_seen":"2026-07-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.22951","pdf_url":"https://arxiv.org/pdf/2607.22951","source_feed":"cs.LG","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["LLM可靠性","统计推断","序列依赖"],"reason":"研究LLM可靠性评估的统计方法，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:18:09","error":null,"has_summary":false,"summary":null},{"id":"2607.23075","version":1,"title":"Traceable LLM Reasoning for Fake-Order Fraud Detection","zh_title":"面向虚假订单欺诈检测的可追溯大语言模型推理","abstract":"Detecting fake-order fraud at scale remains a critical challenge for large online-to-offline (O2O) service platforms, as existing approaches often rely on expert-designed features, produce black-box decisions, and provide limited interpretability. To address these limitations, we propose DeepScrub, a reinforcement learning framework built upon large language models (LLMs) for fake-order fraud detection with traceable reasoning. DeepScrub introduces three innovations. First, a semantic unification module converts heterogeneous risk signals into textual descriptions that LLMs can understand. Second, continued pre-training on risk-control corpora injects domain knowledge, and task rewards jointly evaluate prediction correctness and reasoning quality. Third, the SUggest-REflect (SURE) mechanism incorporates expert feedback and model self-checking to iteratively refine reasoning paths. On a real-world fake-order fraud detection dataset, DeepScrub achieves a macro-F1 score of 85.3%, outperforming the best baseline by 2.7 percentage points. Our task-optimized 8B model further surpasses a 32B model, showing that domain adaptation can matter more than model scale in this setting. In a four-week live pilot, DeepScrub achieved 91.8% precision and 88.5% recall, improving over first-stage human reviewers by 16.6 and 38.8 percentage points. It reduced first-stage manual review workload by 94% and saved nearly one million RMB annually. These results show that DeepScrub improves fraud review accuracy, reduces first-stage review workload, and provides traceable evidence for production risk-review workflows.","authors":["Siqi You","Bingsong Xu","Zhixian Zheng","Xinjian Peng","Yang Xie","Ying Wang","Jiarong Xu"],"categories":["cs.CR","cs.AI","cs.LG"],"primary_category":"cs.CR","announce_type":"cross","date":"2026-07-28","first_seen":"2026-07-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.23075","pdf_url":"https://arxiv.org/pdf/2607.23075","source_feed":"cs.LG","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["欺诈检测","强化学习","大语言模型"],"reason":"纯多智能体协作检测欺诈，无人类行为对照，不涉及仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:18:11","error":null,"has_summary":false,"summary":null},{"id":"2607.23983","version":1,"title":"HydroAgent: Formalizing Forecaster Expertise into Skill-Orchestrated Flood Forecasting Workflows","zh_title":"HydroAgent：将预报员专业知识形式化为技能编排的洪水预报工作流","abstract":"Operational flood forecasting depends on tacit forecaster expertise that is difficult to formalize, audit, and transfer. Although artificial intelligence methods have advanced flood prediction and model-error correction, most existing studies have not explicitly represented the tacit expert rules, review checkpoints, and workflow constraints that connect model outputs to operational warning decisions. To address this issue, we propose HydroAgent, a skill-orchestrated agent framework that embeds Large Language Models (LLMs) into a model-driven flood forecasting workflow, where each skill encodes explicit rules to bound LLM reasoning. We validated its effectiveness using five state-of-the-art LLMs in the South Yamhill River basin. Our results demonstrate that prior judgment captures observed peak flow and flood volume within 5% tolerance in 10 and 11 out of 14 events, with 5-fold cross-validation over 129 events yielding Pearson correlations of 0.62 and 0.84. Building on a high-baseline scheme library (average KGE 0.890), the guided scheme selection further improves KGE by 0.023-0.154, with simulated peak flow and flood volume falling within the prior judgment ranges for 14 and 13 out of 14 events. All five tested LLMs successfully execute the HydroAgent workflow with comparable judgment accuracy (40%-80%), while showing moderate performance variation and substantial cost differences. HydroAgent does not aim to replace human forecasters; instead, it translates their tacit expertise into an auditable and reproducible workflow, streamlining analytical steps and supporting more informed decision-making. This skill-orchestrated paradigm demonstrates how explicit rule boundaries can guide language model reasoning to complement physically based simulation in next-generation flood forecasting.","authors":["Qingyi Yang","Siqian Qiu","Bing Li","Xu Shan","Jia Feng","Shunan Zhou","Xudong Zhou","Tiantian Xing","Jiale Guo","Xiaoyi Dong","Gaoyu Liu","Xiaohuan Liu","Haiqing Pu","Qingwen Deng","Xun Zhang","Zhongrun Xiang","Haiyang Qian","Ying Yan","Yongkang Xu","Nuo Lei","Tianlong Jia","Baoying Shan","Carlo De Michele"],"categories":["physics.geo-ph","cs.LG"],"primary_category":"physics.geo-ph","announce_type":"cross","date":"2026-07-28","first_seen":"2026-07-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.23983","pdf_url":"https://arxiv.org/pdf/2607.23983","source_feed":"cs.LG","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["洪水预报","LLM智能体","工作流编排"],"reason":"多智能体协作完成洪水预报工作流，不涉及人类行为仿真或对照。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:18:19","error":null,"has_summary":false,"summary":null},{"id":"2607.24392","version":1,"title":"When LLM Defenses Backfire: Characterizing Safety, Performance, and Cost Trade-offs","zh_title":"当LLM防御适得其反：安全、性能与成本的权衡特征分析","abstract":"Jailbreak defenses are essential for protecting large language models (LLMs), but they can also introduce secondary costs that weaken model utility. We present a systematic study of these defense trade-offs along three dimensions: performance impact, over-refusal on benign inputs, and inference cost. Rather than treating defenses as a single class, we organize them by operational strategy and examine how different strategies correlate with different side-effect profiles. Across state-of-the-art defense methods, widely used benchmark datasets, and representative open-source LLMs, we find that defenses rarely improve downstream capability, but instead vary in how they trade safety gains against usability and efficiency. In particular, rule-based defenses best preserve task performance, highly conservative self-reflective defenses often increase over-refusal, and multi-round defenses incur the largest runtime overhead. These results provide both a benchmark for evaluating defense side effects and practical guidance for selecting defenses under deployment constraints.","authors":["Tong Zhang","Zexin Li","Simin Chen","Yun Peng"],"categories":["cs.CR","cs.LG"],"primary_category":"cs.CR","announce_type":"cross","date":"2026-07-28","first_seen":"2026-07-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.24392","pdf_url":"https://arxiv.org/pdf/2607.24392","source_feed":"cs.LG","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["LLM安全","防御机制","性能评估"],"reason":"研究LLM防御机制的性能与成本权衡，属于安全评测，不涉及人类行为仿真或对照。","model":"deepseek-v4-pro","scored_at":"2026-07-29T09:01:51","error":null,"has_summary":false,"summary":null},{"id":"2607.24401","version":1,"title":"proxymate: Diagnosis and Adjustment of Proxy Estimates for Reliable Inference","zh_title":"proxymate：代理估计的诊断与调整以实现可靠推断","abstract":"Proxy outcomes (such as short-term behavioral signals, model predictions, or surrogate endpoints) are frequently used in place of primary outcomes that are too slow to mature, rare, or challenging to measure directly. But valid inference on a proxy does not guarantee valid inference on the primary estimate as proxy-based estimates can be systematically biased in ways that are difficult to predict, leading to improperly calibrated confidence intervals. We present proxymate, a framework and open-source Python package for proxy validation and adjustment. proxymate organizes into four levels: The Representativity Level (population validity), the Unit Level (measurement quality), the Estimate Level (decision validity), and the Domain Level (cross-domain transportability). Within each level, proxymate provides diagnostic checks, and targeted adjustment strategies that map specific failures to appropriate corrections. At Meta, proxymate has been adopted by many different use cases, spanning experimentation, prevalence estimation, and monitoring use cases, all facing different proxy challenges (limited human review time, long maturation window of outcomes, low detectability) and showcasing the modularity of the framework. Across all products, proxymate assessed and corrected millions of proxy, primary unit comparisons. It has facilitated launches across multiple work streams including enabling quick decision making on thousands of experiments.","authors":["Alexandra N. M. Darmon","Deeksha Sinha","Steve Wilkins-Reeves","Caner Gocmen"],"categories":["stat.ML","cs.LG"],"primary_category":"stat.ML","announce_type":"cross","date":"2026-07-28","first_seen":"2026-07-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.24401","pdf_url":"https://arxiv.org/pdf/2607.24401","source_feed":"cs.LG","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["代理指标","统计推断","数据验证"],"reason":"论文讨论代理指标验证与校正，不涉及LLM仿真人类被试，无人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:18:22","error":null,"has_summary":false,"summary":null},{"id":"2607.23739","version":1,"title":"Separating Clicks from Baits: Using Large Language Models to Detect Misleading YouTube Thumbnails","zh_title":"区分点击与诱饵：使用大语言模型检测误导性YouTube缩略图","abstract":"Misleading video thumbnails on platforms like YouTube are a pervasive problem, undermining user trust and platform integrity. This paper proposes a novel multi-modal detection pipeline that uses Large Language Models (LLMs) to flag misleading thumbnails. We first construct a comprehensive dataset of 2,843 videos from eight countries, including 1,359 misleading thumbnail videos that collectively amassed over 7.6 billion views, providing a unique cross-cultural perspective on this global issue. Our detection pipeline integrates video-to-text descriptions, thumbnail images, and subtitle transcripts to holistically analyze content and flag misleading thumbnails. Through extensive experimentation and prompt engineering, we evaluate the performance of four frontier-level LLMs, including GPT-4o, GPT-4o Mini, Claude 3.5 Sonnet, and Gemini-1.5 Flash. We further evaluate open-weight vision-language models, LLaVA-v1.5 and Qwen2.5-VL-7B-Instruct, to assess the generalizability of our approach beyond proprietary systems. Our findings show the effectiveness of LLMs in identifying misleading thumbnails, with Claude 3.5 Sonnet consistently showing strong performance, achieving an accuracy of 93.8%, precision over 92%, and recall exceeding 94% in certain scenarios. Beyond evaluating detection performance, we conducted a careful failure analysis to understand when LLMs fail in identifying misleading thumbnails. We discuss the implications of our findings for content moderation, user experience, and the ethical considerations of deploying such systems at scale. Our findings pave the way for more transparent, trustworthy video platforms and stronger content integrity for audiences worldwide.","authors":["Wajiha Naveed","Muhammad Muneeb Pervez","Zaeem Mohtashim Khan","Zafar Ayyub Qazi","Zartash Afzal Uzmi"],"categories":["cs.SI"],"primary_category":"cs.SI","announce_type":"new","date":"2026-07-28","first_seen":"2026-07-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.23739","pdf_url":"https://arxiv.org/pdf/2607.23739","source_feed":"cs.SI","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["内容审核","多模态检测","LLM应用"],"reason":"用LLM检测误导性缩略图，属内容审核任务，非人类行为仿真或替代被试。","model":"deepseek-v4-pro","scored_at":"2026-07-29T09:01:51","error":null,"has_summary":false,"summary":null},{"id":"2607.24698","version":1,"title":"Modest Algorithmic Mediation can Maximize Topical Diversity in Hybrid Human-AI Systems","zh_title":"适度算法中介可最大化混合人机系统中的主题多样性","abstract":"In the artificial intelligence (AI) era, the rise of algorithmic feeds has fundamentally transformed information diffusion on social media. While early platforms organized visibility through explicit social networks, contemporary systems mediate exposure through intelligent recommender algorithms that personalize attention. This paper examines how the social network and algorithmic architecture jointly shape the diversity of information sharing. Analysis of 18,076 users active throughout 2014--2018 shows that the topical diversity of sharing rose and then plateaued after the introduction of algorithmic ranking in 2016 while its inequality across users emerged alongside it. To this end, we introduce a hybrid human-AI information diffusion model in which information exposure is governed by a parameterized mixture of social propagation through the user-following network and algorithmic recommendation. Both qualitative analysis and simulations show that the effect of algorithmic mediation is non-monotonic. Modest mediation can raise average diversity and reduce inequality relative to a purely network-driven baseline, whereas strong mediation reduces diversity and concentrates it among fewer users. Fitting the model to four years of data yields a mediation share that increases from zero before 2016 to approximately 0.50 by 2018, a level that exceeds the compensation point of equality while remaining within the diversity-enhancing range. These results identify the conditions under which recommendation broadens rather than narrows exposure and provide a unified framework for information diffusion in hybrid human-AI systems.","authors":["Dini Wang","Ho-Chun Herbert Chang"],"categories":["cs.SI"],"primary_category":"cs.SI","announce_type":"new","date":"2026-07-28","first_seen":"2026-07-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.24698","pdf_url":"https://arxiv.org/pdf/2607.24698","source_feed":"cs.SI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["信息扩散","算法推荐","社交网络"],"reason":"多智能体信息扩散模型，无LLM仿真人类被试，无人类行为对照","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:26","error":null,"has_summary":false,"summary":null},{"id":"2607.23168","version":1,"title":"Individual and collective gains from cooperation and reciprocity in a dynamic-network Prisoner's Dilemma driven by extraversion, openness, and agreeableness","zh_title":"外向性、开放性和宜人性驱动的动态网络囚徒困境中合作与互惠的个体与集体收益","abstract":"How do stable personality differences shape cooperation when social ties can form and dissolve? We model a repeated Prisoner's Dilemma on an endogenous network in which three continuous Big Five traits map to transparent local mechanisms: Extraversion sets a target number of partners, Openness determines how broadly agents search beyond friends-of-friends, and Agreeableness sets a baseline willingness to cooperate. At each encounter, agents combine this baseline with the partner's directly observed history; there are no trait labels, gossip, or global reputations. Ties form when agents are under-connected and are cut when they become over-connected, with cuts prioritising partners who have defected more often. We vary network size (N=30--200), population composition, and the balance between trait-driven and history-driven behaviour. Three robust patterns emerge. First, cooperate first, then reciprocate---high initial willingness to cooperate combined with history-sensitive response---produces systems that are simultaneously more prosperous, fairer, and safer. Second, personality has predictable conditional effects: Agreeableness helps when history matters but hurts when behaviour is mostly trait-driven; Extraversion amplifies the environment; Openness has little net payoff effect. Third, the network reorganises accordingly: degree assortativity stays near zero, whereas agreeable agents increasingly connect to one another when cooperation takes hold.","authors":["David Abi\\'an","Jorge Bernad","Sergio Ilarri","Raquel Trillo-Lado"],"categories":["physics.soc-ph","cs.SI"],"primary_category":"physics.soc-ph","announce_type":"cross","date":"2026-07-28","first_seen":"2026-07-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.23168","pdf_url":"https://arxiv.org/pdf/2607.23168","source_feed":"cs.SI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体模拟","囚徒困境","人格特质"],"reason":"纯多智能体协作研究，无LLM参与，不涉及人类行为对照，属于C1排除项。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:17","error":null,"has_summary":false,"summary":null},{"id":"2607.24175","version":1,"title":"A World of Ginis","zh_title":"基尼系数的世界","abstract":"The Gini index remains the most important measure of economic inequality worldwide, and accurate estimates of this index are essential for effective public policies. Yet, Gini estimates for the same country and year vary considerably across data sources, a problem that remains largely unresolved. The paper reviews the largest global and regional databases providing Gini estimates, surveys the related literature, and constructs a unified dataset of 122,351 Gini observations spanning 222 countries and territories and 158 years, from 1867 to 2024. The analysis of this new dataset shows that income-based Ginis exceed consumption-based ones by 4.7 points on average globally, and by as much as 10 points in some regions, with these gaps widening over time. The gross--net income distinction and the use of alternative equivalence scales together with several other measurement choices add further systematic differences. Based on these findings, the paper provides correction factors that can be used to harmonise Ginis built on different welfare concepts. We further show that overall divergence across databases has grown only modestly since 1960, and mainly through the proliferation of databases rather than through genuine divergence among long-standing sources. Thus, improving on the existing discrepancies across Ginis globally is possible, but ultimately depends on database administrators disclosing full details of Gini construction and on users selecting Ginis built on comparable measures.","authors":["Lidia Ceriani","Paolo Verme"],"categories":["econ.GN","q-fin.EC","stat.AP"],"primary_category":"econ.GN","announce_type":"new","date":"2026-07-28","first_seen":"2026-07-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.24175","pdf_url":"https://arxiv.org/pdf/2607.24175","source_feed":"econ.GN","score":0,"bucket":"other","rubric_hits":["C5"],"tags":["经济不平等","基尼系数","数据协调"],"reason":"论文研究基尼系数测量差异，不涉及LLM或人类仿真，完全无关。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:18:19","error":null,"has_summary":false,"summary":null},{"id":"2607.24389","version":1,"title":"How to Disrupt a Market","zh_title":"如何扰乱一个市场","abstract":"Market design research in economics naturally focusses on how to improve market efficiency. Our objective here is exactly the opposite - how to design interventions that make a market less efficient. Our research is inspired by the growth of illicit markets online where reducing their efficiency may reduce societal harm. Using a web-based experiment, we find that a partial disruption to delivery is an effective method to decrease market efficiency. The decrease is borne by sellers who sell fewer goods and have lower earnings. A consequence of a disruption to delivery, however, is an increase in market concentration because it facilitates the emergence of a dominant seller. In contrast, we find that attacks on seller ratings are ineffective at reducing market efficiency. This study paves the way for evidence-based, causally driven investigations to aid policies to disrupt cybercrime and other illicit markets.","authors":["Edoardo Gallo","Rebecca Heath","Jonathan Lusthaus","Federico Varese"],"categories":["econ.GN","q-fin.EC"],"primary_category":"econ.GN","announce_type":"new","date":"2026-07-28","first_seen":"2026-07-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.24389","pdf_url":"https://arxiv.org/pdf/2607.24389","source_feed":"econ.GN","score":0,"bucket":"other","rubric_hits":["C2"],"tags":["市场设计","网络实验","非法市场"],"reason":"基于人类被试的网络实验，无LLM仿真，不涉及用模型替代人类行为。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:25","error":null,"has_summary":false,"summary":null},{"id":"2607.23254","version":1,"title":"Towards Optimal Estimators for Randomized Control Trials","zh_title":"迈向随机对照试验的最优估计量","abstract":"Randomized controlled trials (RCTs) are fundamental tools for causal inference across technology companies, pharmaceutical research, and federal agencies. While the standard difference-in-means estimator provides unbiased treatment effect estimates, it often lacks precision, particularly when treatment effects are heterogeneous or outcomes exhibit heavy-tailed distributions. Although numerous precision-enhancing methods exist---from covariate adjustment techniques to variance reduction strategies---recent research demonstrates that no single estimator performs optimally across all datasets. Rather than seeking the best estimator for individual RCTs, which risks compromising scientific validity through convenient selection, we propose a principled framework for identifying optimal estimators within families of RCTs based on specific analytical goals. Our approach uses sample splitting to estimate the distribution of evaluation metrics (e.g., mean squared error, regret) across RCT families, enabling systematic comparisons between estimators while maintaining asymptotic guarantees. We demonstrate this framework using a sample of Amazon's Supply Chain Optimization Technology trials and the Strengthening Democracy Challenge dataset (25 interventions). Results reveal that optimal estimators vary significantly by analytical objective: weighted least squares performs best for inference goals, while difference-in-means minimizes regret for decision-making contexts. This work provides actionable guidance for estimator selection while preserving methodological rigor across diverse research applications.","authors":["Harsh Parikh","Gabriel Levin-Konigsberg","Nilesh Tripuraneni","Dhruv Madeka","Michael I. Jordan","Dean Foster","Dominique Perrault-Joncas","Alexander Volfovsky"],"categories":["stat.AP","econ.EM","stat.ME"],"primary_category":"stat.AP","announce_type":"cross","date":"2026-07-28","first_seen":"2026-07-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.23254","pdf_url":"https://arxiv.org/pdf/2607.23254","source_feed":"econ.EM","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["因果推断","随机对照试验","估计量选择"],"reason":"论文研究RCT估计量选择，不涉及LLM仿真人类被试，无人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:18:11","error":null,"has_summary":false,"summary":null},{"id":"2607.24143","version":1,"title":"Inference on counterfactual distributions using martingale posteriors","zh_title":"使用鞅后验进行反事实分布的推断","abstract":"Causal inference is often focused on average effects, which can hide important aspects of the effect distributions. Here we consider the entire posterior effects distribution by estimating full counterfactual outcome distributions. We propose a methodology for inference on counterfactual distributions which builds upon the martingale posterior framework of Fong et al. (2023). This provides a highly flexible approach to estimating densities, distribution functions, and derived quantities such as quantiles, which coherently quantifies the epistemic uncertainty on any target estimand of interest. As the predictive recursions are based on an underlying nonparametric model (a Dirichlet process mixture model), our method naturally inherits robustness with respect to restrictive parametric assumptions. In addition, implementation of our method is typically very fast. This approach can be applied to marginal or conditional counterfactual distributions and is easily extended to an instrumental variables setup. Using the concept of almost conditionally identically distributed random variables, we prove convergence of the martingale posterior inference on the counterfactual outcome distributions for the causal models considered in the paper. We illustrate our approach on both simulated and real data. Using the latter, we investigate the effect of zinc lozenges on common cold duration, the impact of vitamin A supplementation on children's survival rates with one-sided non-compliance (analysed in Imbens and Rubin, 1997a) and the effect of job training (LaLonde, 1986).","authors":["Gregor Steiner","Mark Steel"],"categories":["stat.ME","econ.EM"],"primary_category":"stat.ME","announce_type":"cross","date":"2026-07-28","first_seen":"2026-07-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.24143","pdf_url":"https://arxiv.org/pdf/2607.24143","source_feed":"econ.EM","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["因果推断","反事实分布","鞅后验"],"reason":"论文研究因果推断的反事实分布，使用鞅后验方法，不涉及LLM或人类仿真。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:18:19","error":null,"has_summary":false,"summary":null},{"id":"2607.24304","version":1,"title":"Teacher Knows It Best: Spontaneous Symmetry Breaking and Tipping Points in Networked Langevin Dynamics AI Sycophancy","zh_title":"老师最懂：网络化朗之万动力学AI谄媚中的自发对称破缺与临界点","abstract":"We formulate a statistical physics framework to model a networked stochastic dynamical system exhibiting bistability, driven by additive noise and social conformity. We apply this model to understand and mitigate AI-induced delusional spiraling-a phenomenon where algorithmic sycophancy from Large Language Models continuously reinforces inaccurate beliefs within a socially interacting society. By partitioning the network into a majority of regular agents and a minority of \"aware\" nodes (Teachers) placed at topological hubs, we use a degree-weighted mean-field approximation to reduce high-dimensional coupled Langevin equations into a single macroscopic drift equation. We provide a closed-form analytical derivation for the deterministic critical tipping time through a saddle-node bifurcation. We validate this analytical boundary using finite-size scaling and demonstrate a universal data collapse across diverse network topologies. Finally, we optimize an intervention strategy under a strict budget constraint that balances the topological footprint against driving velocity. We prove mathematically that under certain conditions, a highly concentrated, rapid intervention targeting massive hubs strictly outperforms a distributed, slow approach to rescue the network.","authors":["Sayantari Ghosh","Saumik Bhattacharya","Partha Pratim Chakrabarti"],"categories":["physics.soc-ph","cs.AI","math.DS"],"primary_category":"physics.soc-ph","announce_type":"new","date":"2026-07-28","first_seen":"2026-07-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.24304","pdf_url":"https://arxiv.org/pdf/2607.24304","source_feed":"physics.soc-ph","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","统计物理","AI谄媚"],"reason":"纯多智能体网络动力学模型，研究AI谄媚传播，无真实人类行为对照，不涉及LLM仿…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:18:20","error":null,"has_summary":false,"summary":null},{"id":"2607.21757","version":1,"title":"Co-design of LLM-based preference agents: participation may drive overtrust","zh_title":"基于大语言模型的偏好代理协同设计：参与可能驱动过度信任","abstract":"Large language models are increasingly used to simulate human preferences in research and practical applications, raising concerns about validation, misrepresentation, and exclusion. Co-designing agents with the people they represent is a promising way to address these concerns, but participation may also mask the problems it appears to solve. This paper explores that tension through a primarily qualitative study in which 12 participants co-designed personal preference agents in the domain of household energy, via a background survey, co-design interview, and validation survey. Participants engaged readily and mostly came to see their agents as representing them well. Independent validation, however, revealed mixed human-agent alignment, with agent responses markedly more homogeneous, decisive, and abstract than the human sample. I argue that participation and process transparency can act as an \"overtrust engine\" that promotes trust while concealing systematic misalignment with potential structural consequences at scale. I develop this as a core mechanism in participatory preference agent design, treating individual alignment not as a fixed state but as an enacted process.","authors":["Michael J. Fell"],"categories":["cs.CY","cs.AI","cs.HC"],"primary_category":"cs.CY","announce_type":"cross","date":"2026-07-27","first_seen":"2026-07-27","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.21757","pdf_url":"https://arxiv.org/pdf/2607.21757","source_feed":"cs.AI","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM人类仿真","偏好代理","人机对齐"],"reason":"用LLM模拟人类偏好并与真实人类数据对照，评估仿真可靠性与偏差，批判性指出过度…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:14","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-29","rank":2,"question":"在LLM偏好代理的参与式协同设计中，参与和过程透明是否会导致过度信任，从而掩盖系统性的对齐偏差？","design":"本研究采用质性为主的混合方法，招募12名参与者，通过背景调查、协同设计访谈和验证调查，在家庭能源领域协同设计个人偏好代理，并比较参与者感知的代理表现与独立评估的代理表现。","baseline":"以12名参与者在验证调查中的真实回答作为人类基准，与代理回答进行对齐比较。","findings":"参与者普遍认为其协同设计的代理能很好地代表自己，但独立验证显示人-代理对齐程度参差不齐，代理回答比人类样本更同质、更果断、更抽象。参与和过程透明可能成为“过度信任引擎”，在促进信任的同时掩盖系统性偏差。","reliability":"论文指出，参与式设计可能掩盖系统性的对齐偏差，代理回答的同质化倾向可能导致少数观点被压制，且用户在不熟悉领域难以识别代理是代表偏好还是塑造偏好。","relevance":"该研究直接探讨用LLM模拟人类偏好并与真实人类数据对照，评估仿真可靠性与偏差，并批判性指出参与式设计可能引发过度信任，高度契合研究者对LLM仿真实验的批判性关注。","inspiration":"借鉴其协同设计流程与独立验证相结合的方法，可迁移到消费者金融决策偏好模拟场景，设计一个实验：招募真实消费者作为被试，通过访谈协同设计其消费信贷偏好代理，以真实信贷选择数据为基准，测量代理在风险偏好、跨期选择等任务上的对齐度与同质化程度。"}},{"id":"2607.22218","version":1,"title":"Why Large Language Models and Humans Converge and Diverge in Evaluating Creativity","zh_title":"大型语言模型与人类在创造力评估中为何趋同与分歧","abstract":"Despite the growing use of large language models (LLMs) as creativity evaluators, evidence of their alignment with human evaluations remains mixed, raising the question of when and why their judgments converge with or diverge from human judgments. Across three studies and six widely used LLMs, we addressed this gap by identifying the standards underlying LLM creativity evaluation and examining their downstream implications. Study 1 showed that LLMs generally relied on a narrower subset of human creativity evaluation standards. Convergence with human standards was strongest in the novelty dimension, whereas divergence was clearest in the contextual dimension, which captures social, market, and reputational information. Moreover, each LLM exhibited distinct, model-specific standards that varied substantially in breadth. These differences in evaluation standards were reflected in actual creativity judgments. Study 2 (N = 1,103 ideas) showed that LLM evaluations were moderately correlated with human evaluations, and individual LLMs with broader standards better distinguished ideas humans judged as more versus less creative. Study 3 (N = 1,195) showed that LLMs were less sensitive to contextual information: such information significantly altered human creativity ratings but left LLM ratings largely unchanged. Together, our findings help explain the mixed evidence on LLM-human alignment, showing that alignment depends on the evidence a judgment demands and the standards each model applies. LLMs may resemble humans when evaluations emphasize intrinsic qualities such as novelty, yet diverge when judgments require contextual information. Selecting an LLM evaluator is therefore a consequential decision: different models, applying different standards, recognize different ideas as creative.","authors":["Pengzhao Lyu","Yeun Joon Kim","Hanlin Xiao","Yingyue Luna Luan"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-07-27","first_seen":"2026-07-27","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.22218","pdf_url":"https://arxiv.org/pdf/2607.22218","source_feed":"cs.CL","score":7,"bucket":"pending","rubric_hits":["A1","B1","B4"],"tags":["LLM评估","人机对齐","创造力判断"],"reason":"用LLM评估创造力并与人类判断对照，揭示对齐条件与失效情境，方法可迁移至人类仿…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:15","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-29","rank":4,"question":"LLM在评估创造力时，其评价标准在哪些维度上与人类标准趋同或背离，以及这种差异如何影响下游的创造力判断？","design":"本研究并非以LLM模拟人类被试，而是将六种主流LLM作为创造力评估者，通过三个研究分析其评价标准：研究一识别各LLM使用的创造力评价标准并与人类26项标准对比；研究二让LLM对1103个创意进行评分，与人类评分比较；研究三向LLM和人类被试提供情境信息，观察其对评分的影响。","baseline":"人类基准来自已有文献中确立的26项人类创造力评价标准，以及研究二（N=1103个创意）和研究三（N=1195）中收集的人类创造力评分。","findings":"LLM普遍依赖较窄的人类评价标准子集，在新颖性维度上与人类趋同最强，在涉及社会、市场和声誉信息的情境维度上背离最明显；评价标准更广的LLM能更好地区分人类评判的高/低创意，且LLM对情境信息不敏感，该信息显著改变人类评分但对LLM评分影响甚微。","reliability":"论文指出LLM与人类判断的对齐取决于判断所需的证据类型和模型应用的标准，LLM在需要情境信息的判断中会失效，且不同模型因标准不同会识别出不同的创意，选择LLM评估者是一个重要决策。","relevance":"该研究直接对比LLM与人类在评价任务中的判断差异，揭示了LLM在需要情境信息时与人类背离的失效条件，对理解LLM仿真人类决策的边界条件具有参考价值，值得阅读原文。","inspiration":"可借鉴其通过识别LLM使用的评价标准并与人类标准体系对比的方法，来诊断LLM在特定任务中与人类对齐的维度。｜可迁移到信贷审批或政策评估场景，例如研究LLM在评估贷款申请时是否忽略申请人的社会背景信息，而人类审批员会受此影响。｜设计：以LLM作为信贷审批员，处理变量为是否提供申请人的社区声誉或就业市场信息，结果变量为信用评分，对照真实银行信贷员的历史审批数据。"}},{"id":"2607.24372","version":1,"title":"Randomness in large language models: What researchers need to know (and report)","zh_title":"大语言模型中的随机性：研究者须知（及应报告事项）","abstract":"Large language models (LLMs) are increasingly used to generate data for research. Typical use cases are classifications, annotations, information extraction, and generation of numerical scores. Unlike conventional measurements, LLM outputs can vary across repeated requests even when the prompt and apparent model settings remain unchanged. This variation arises from deliberate sampling, silent model updates, numerical rounding, or expert routing. Setting a dedicated temperature parameter to zero removes deliberate sampling when that option is available, but it does not eliminate the other sources of randomness. Exact reproduction is therefore generally not possible when using proprietary application programming interfaces. Local execution of open-weight models offers greater control, but reproducibility still depends on the complete hardware and software stack. We illustrate these issues through sentiment classifications of corporate filings and examine their consequences for downstream regression results. We then propose a reporting standard for articles and replication packages, as well as guidance for data editors and authors. Together, these findings and recommendations establish that LLM outputs should be treated as draws from a distribution rather than as fixed measurements.","authors":["Guillaume Coqueret","Joan Llull","Florian Oswald","Christophe Pérignon","Christoph Scheuch","Lars Vilhuber"],"categories":["econ.GN"],"primary_category":"econ.GN","announce_type":"new","date":"2026-07-27","first_seen":"2026-07-27","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.24372","pdf_url":"https://arxiv.org/pdf/2607.24372","source_feed":"backfill","score":7,"bucket":"pending","rubric_hits":["A2","B4"],"tags":["LLM随机性","可重复性","报告标准"],"reason":"研究LLM输出的随机性对实证结果的影响，提出报告标准，可迁移到仿真可靠性评估。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:25","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-29","rank":5,"question":"大语言模型输出的随机性来源有哪些，以及这种随机性如何影响实证研究的可复现性和下游回归结果？","design":"本文不是仿真研究，而是通过实证演示LLM输出的随机性：使用LLM对公司年报进行情感分类，比较多次调用同一模型（包括设置temperature=0）的输出差异，并考察这些差异对后续回归系数和显著性的影响。","baseline":"无对照","findings":"LLM输出即使在temperature=0时仍存在随机性，来源包括静默更新、硬件差异、数值舍入等，导致完全复现不可能。这种随机性会传导至下游回归，使系数估计和显著性不稳定，因此应将LLM输出视为分布中的抽取而非固定测量。","reliability":"论文指出，即使设置temperature=0也无法消除所有随机性；使用商业API时无法控制基础设施和模型版本，本地部署开源模型虽可提升控制力，但复现仍依赖完整的软硬件栈。","relevance":"本文系统梳理了LLM输出的随机性来源及其对实证结果的影响，并提出了报告标准，这些发现可直接迁移到用LLM进行人类仿真实验的可靠性评估中，值得精读。","inspiration":"本文通过多次调用LLM并量化输出变异对回归结果的影响，提供了一种评估LLM生成数据稳健性的方法，值得借鉴。｜可迁移到使用LLM模拟投资者情绪或分析师预测的经济金融场景，例如检验情感分数的不确定性如何影响资产定价回归。｜以LLM作为被试，重复生成同一组公司年报的情感分数，计算多次运行间回归系数的分布，并与人类分析师的实际评级数据做对照，评估仿真可靠性。"}},{"id":"2607.21614","version":1,"title":"Household Movement Detection in Mixed-Format Occupancy Data Using LLM-Based Entity Resolution","zh_title":"基于LLM实体解析的混合格式居住数据中家庭迁移检测","abstract":"Entity resolution (ER) typically relies on pairwise similarity comparisons between records, which limits its ability to capture indirect relationships present in demographic occupancy data. An important indirect pattern arises from household movement, where multiple individuals relocate together across addresses, but detecting such patterns is difficult due to mixed-format records, noise, duplication, and the absence of stable identifiers. This paper proposes an AI-enhanced framework for detecting indirect entity links associated with household movement in unstandardized name-address data. The approach integrates prompt-based large language model (LLM) named entity recognition for extracting personal names and addresses without extensive preprocessing, semantic text embeddings for robust similarity computation, and graph-based reasoning to infer group-level movement patterns. Experimental evaluation on SPX benchmark datasets (S8-S12) generated using the Synthetic Occupancy Generator demonstrates that incorporating indirect household movement evidence improves recall by 8-15% while maintaining high precision, yielding F1-score gains of 6-8% over a strong pairwise baseline.","authors":["Sasirekha Oguri","John R. Talburt","Mert Can Cakmak"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-07-27","first_seen":"2026-07-27","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.21614","pdf_url":"https://arxiv.org/pdf/2607.21614","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["实体解析","图推理","数据质量"],"reason":"纯实体解析与多智能体协作，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-07-29T09:01:42","error":null,"has_summary":false,"summary":null},{"id":"2607.22305","version":1,"title":"A Roadmap to Impactful Pluralistic Alignment Research","zh_title":"迈向有影响力的多元对齐研究路线图","abstract":"Pluralistic value alignment---the goal of building AI systems that represent and serve diverse human values and perspectives---has emerged as an active research agenda. Yet, there's no public evidence that it has shaped the training or evaluation of the AI systems people actually use. We audit the public behavior documents and evaluations of frontier labs, finding none name pluralism as a goal, and as of this writing, no clear indication that production models are explicitly trained or tested for it. This goes against the primary motivations and goals of pluralistic alignment, which revolve around making a positive difference in the models serving billions of users worldwide. We argue that the pluralistic alignment research community should focus on supporting impact and adoption in deployed, widely-used AI systems. We provide evidence for the adoption problem, present three main reasons behind it, and discuss three corresponding areas for future research to address it: 1. The primary justifications for pluralistic alignment so far have been normative or speculative. We need studies showing empirically how pluralistic AI benefits users or society. 2. The pluralistic alignment research community has not settled when pluralistic behavior is warranted or what pluralism ideally looks like in practice. We need to establish a concrete goal for developers to operationalize. 3. Current methods trade off against other desiderata of LLMs in ways that are largely unmeasured, and existing metrics are not \"hill-climbable.\" We need trade-off-aware evaluations and methods that meet the requirements of production systems. This paper serves as a collective call to action for the pluralistic alignment researchers: progress requires moving beyond normative justification toward empirical foundations, a concrete account of ideal pluralistic behavior, and practical methods and evaluations built for adoption.","authors":["Elinor Poole-Dayan","Jillian Fisher","Atoosa Kasirzadeh","Jacob Andreas","Mitchell Gordon","Michiel A. Bakker"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-07-27","first_seen":"2026-07-27","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.22305","pdf_url":"https://arxiv.org/pdf/2607.22305","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["多元对齐","AI治理","价值对齐"],"reason":"论文讨论多元价值对齐的实证基础与评估，不涉及用LLM仿真人类被试或行为对照。","model":"deepseek-v4-pro","scored_at":"2026-07-29T09:01:42","error":null,"has_summary":false,"summary":null},{"id":"2607.22345","version":1,"title":"Teachy Mini: Development and Preliminary Evaluation of a Knowledge-Based Generative Social Robot for Higher Education","zh_title":"Teachy Mini：基于知识的高等教育生成式社交机器人开发与初步评估","abstract":"Generative social robots (GSRs) powered by large language models offer new possibilities for personalized tutoring in higher education, but also introduce risks related to misinformation, missing transparency, or reinforcing incorrect student responses. Prior work identified knowledge-based design (KBD) requirements that define the informational prerequisites for GSRs to manifest responsible and effective tutoring behavior in higher education. In this paper, we operationalized selected KBD requirements in the Reachy Mini robot platform through system prompting, retrieval-augmented generation, and stateful prompt orchestration. As a result, we present Teachy Mini, a GSR tutoring system that was developed using KBD. To test the system, we conducted a preliminary evaluation study. Participants (N = 24) completed a robot-guided learning session about research methodologies. They learned either with Teachy Mini or with a control version that did not follow KBD principles. Teachy Mini was perceived as significantly more aligned with responsible tutoring behavior than the control robot. Moreover, a manipulation check illustrated that Teachy Mini used personalization, slide-grounded explanations, Socratic questioning, affective support, and learner-anchored feedback more consistently than the control robot. No significant between-condition differences were found in system acceptance, intrinsic motivation, or learning effectiveness, although exploratory analyses suggested a positive effect of KBD on objective learning gains when accounting for learner preferences. Overall, the study offered an initial implementation and preliminary evaluation of KBD for GSR tutoring, indicating that KBD can shape responsible robot behavior and potentially increase learning effectiveness in robot-supported learning.","authors":["Stephan Vonschallen","Karim Kaufmann","Dominique Oberle","Friederike Eyssel","Theresa Schmiedel"],"categories":["cs.RO","cs.AI"],"primary_category":"cs.RO","announce_type":"cross","date":"2026-07-27","first_seen":"2026-07-27","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.22345","pdf_url":"https://arxiv.org/pdf/2607.22345","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C2"],"tags":["社交机器人","教育辅导","人机交互"],"reason":"机器人平台上的教育辅导系统，属于机器人仿真环境，不涉及用LLM仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-07-29T09:01:44","error":null,"has_summary":false,"summary":null},{"id":"2607.22463","version":1,"title":"Beyond Perspectives: A Trio-Ethnography of Interpretation Evolution in LLM-Supported Programming Education","zh_title":"超越视角：LLM支持的编程教育中解释演变的三方民族志","abstract":"Generative AI is reshaping programming education, yet educators often infer students' AI-supported learning from classroom observations alone. This experience report presents a trio-ethnography involving two computing educators with different teaching philosophies and one undergraduate computer science student to examine how these interpretations evolve through dialogue. Across three conversations, the educators reflected on students' AI use, discussed changes to programming pedagogy, and revisited their assumptions after engaging with the student's lived experiences. Rather than simply confirming or contradicting the educators' perspectives, the student's narratives revealed learning processes that were largely invisible in the classroom, prompting both educators to reconsider assumptions about AI use, assessment, transparency, and programming instruction. We argue that trio-ethnography offers a valuable reflective approach for helping computing educators move beyond observable student behaviors toward a richer understanding of AI-supported learning and for informing instructional adaptation in the era of generative AI.","authors":["Jennie Ren","Jordan H. McDowell","Kyrie Zhixuan Zhou"],"categories":["cs.HC","cs.AI"],"primary_category":"cs.HC","announce_type":"cross","date":"2026-07-27","first_seen":"2026-07-27","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.22463","pdf_url":"https://arxiv.org/pdf/2607.22463","source_feed":"cs.AI","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["编程教育","生成式AI","民族志"],"reason":"论文是教育经验报告，涉及师生对话反思，无LLM仿真人类被试或行为对照实验。","model":"deepseek-v4-pro","scored_at":"2026-07-29T09:01:44","error":null,"has_summary":false,"summary":null},{"id":"2607.23313","version":1,"title":"Agentic AI Orchestration of Heterogeneous Economic Models for Rapid, Multi-scenario Analysis of Energy Crises","zh_title":"基于异构经济模型的智能体AI编排用于能源危机快速多场景分析","abstract":"Rigorous economic models can take months to construct, yet energy crises demand decisions from policymakers within days or even hours. Any disruption in energy markets is not isolated but rapidly disseminates through interlinked global systems. Off-the-shelf models that already exist typically focus only on limited aspects of the system and are distributed across research groups, programming languages, software architectures not designed for model integration, and incompatible formats. Integrating these models manually can take longer than the crisis itself, forcing analysts to rely on whichever models are easiest to connect and leaving consequential scenarios unexplored. Policymakers must make rapid decisions with obstructed and limited information. We show that large language models can perform the critical integration directly. The system constructs internally consistent scenarios, translates assumptions into model-specific inputs, executes existing economic and physical models in dependency order, and synthesizes outputs tailored to policymakers. The language model generates no quantitative results: every reported value is reproduced directly from an underlying model run, remains traceable to its source and is subject to analyst approval at each stage. We develop a LLM framework that coordinates 16 models of oil, natural gas, shipping, water, helium, fertilizer and macroeconomic equilibrium. The framework is applied across five scenarios to assess the 2026 closure of the Strait of Hormuz and refreshed weekly for eight weeks as events on the ground continued to unfold. By linking models that already exist and reading them as a suite rather than in isolation, this architecture mobilizes distributed scientific models rapidly during energy and geopolitical disruptions while keeping any single model's assumptions from driving the conclusion.","authors":["Dana Golden","Brett Indelicato","Lav R. Varshney","Carlos D. Messina","Suzanne Thornsbury"],"categories":["econ.GN"],"primary_category":"econ.GN","announce_type":"new","date":"2026-07-25","first_seen":"2026-07-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.23313","pdf_url":"https://arxiv.org/pdf/2607.23313","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","经济模型集成","能源危机分析"],"reason":"纯多智能体系统协调经济模型，无人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:17","error":null,"has_summary":false,"summary":null},{"id":"2607.21268","version":1,"title":"pAI-Econ-claude: A Gated Human-in-the-Loop Multi-Agent Architecture for AI-Assisted Economic Theory Development","zh_title":"pAI-Econ-claude：一种用于AI辅助经济学理论开发的门控人机协同多智能体架构","abstract":"In many social-science research tasks, such as economics, LLM-based agents must produce outputs for which no cheap, task-complete, machine-readable correctness signal exists. This creates a distinctive reliability problem for multi-agent systems: how should generation, critique, coordination, and human judgment be organized when no component can certify the final result? We address this problem through pAI-Econ-claude, a gated, human-in-the-loop multi-agent architecture for AI-assisted economic theory development. Agents coordinate through a shared workspace of inspectable intermediate records; specialized gates diagnose targeted failure modes and recommend loopbacks without certifying correctness; and human checkpoints retain authority over decisions that are costly to reverse. We evaluate the architecture on five matched economic-theory tasks against an ungated baseline. Two evaluators blinded to configuration agreed on all five pairwise rankings, preferring the gated architecture in four tasks and the baseline in one. Mean failure severity fell from 1.58 to 1.16, while overall usefulness rose from 2.60 to 3.10. The largest observed gain occurred when a reality check rejected a false market-structure premise and a proof review prompted revision of a false welfare claim. The negative case shows that scaffolding can also compress an economically important mechanism too aggressively. The results support a bounded claim: gated oversight improves the auditability of AI-assisted economic theory without substituting for formal verification, and the allocation of irreversible human judgment is a more informative design variable than pure agent autonomy. The workflow is publicly available at https://github.com/maxwell2732/pAI-Econ-claude.","authors":["Chen Zhu","Xiaolu Wang","Weilong Zhang"],"categories":["cs.MA","cs.AI","econ.GN"],"primary_category":"cs.MA","announce_type":"new","date":"2026-07-23","first_seen":"2026-07-23","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.21268","pdf_url":"https://arxiv.org/pdf/2607.21268","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","经济学理论","人机协同"],"reason":"多智能体协作辅助经济学理论开发，不涉及人类行为仿真或对照。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:18:05","error":null,"has_summary":false,"summary":null},{"id":"2607.20410","version":1,"title":"LKValues: Aligning Large Language Models with Sri Lankan Societal Values","zh_title":"LKValues：将大语言模型与斯里兰卡社会价值观对齐","abstract":"Value alignment of Large Language Models (LLMs) has been shown to be culturally biased toward Western norms. This results in the mishandling of local values in multilingual societies such as Sri Lanka that have their unique cultural dynamics. Existing benchmarks overlook Sri Lankan-contextualized values in its official language Sinhala, hindering culturally sensitive evaluation and fine-tuning. To bridge this gap, we propose LKValues, the first survey-grounded resource suite for Sri Lankan value alignment. From a trilingual survey of 205 respondents, blending adapted global frameworks and LLM-elicited local constructs, we derive 40 majority-endorsed societal values. Using these values, we construct LKvaluesIT, a Sinhala-English news-derived instruction corpus containing 150k scenario-based instances, and LKvaluesBench, a value-sensitive evaluation benchmark of 1,000 instances. We evaluate a set of proprietary and open-weight LLMs with LKvaluesBench. We fine-tune three open-weight base models (Qwen3.5-4B-Base, Qwen3.5-9B-Base, and Aya-Expanse-8B-Base). Our experiments show that newer and larger LLMs still exhibit low-resource and cultural value-alignment gaps. LKValues fine-tuning improves Qwen-family models in English and Sinhala, reducing invalid outputs and cross-lingual disparities, though gains remain model-family dependent. These highlight LKValues efficacy in embedding Sri Lankan values, offering a replicable pipeline for low-resource, country-specific pluralist value alignment. The dataset is publicly available at https://github.com/NextME14/LKValues.","authors":["Nethmi Muthugala","Supryadi","Surangika Ranathunga","Nisansa de Silva","Ruijie Tao","Ovindu Gunatunga","Pengyun Zhu","Shaowei Zhang","Jingting Zheng","Deyi Xiong"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-07-22","first_seen":"2026-07-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.20410","pdf_url":"https://arxiv.org/pdf/2607.20410","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["价值观对齐","文化偏见","低资源语言"],"reason":"测量LLM对斯里兰卡价值观的符合度，属模型对齐评估，非仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:18:04","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":165,"question":"如何构建并评估一个基于调查的斯里兰卡社会价值观对齐资源，以改善大语言模型在僧伽罗语和英语中的文化敏感性？","design":"本研究并非人类仿真实验，而是通过三语调查（205名受访者）提炼出40项斯里兰卡社会价值观，据此构建了包含15万条场景指令的训练集LKvaluesIT和1000条评估基准LKvaluesBench，并对多个开源和闭源模型进行微调与评估。","baseline":"无对照","findings":"更大、更新的模型在低资源语言和文化价值对齐上仍存在差距；使用LKvaluesIT微调能提升Qwen系列模型在英语和僧伽罗语上的表现，减少无效输出和跨语言差异，但效果因模型家族而异。","reliability":"论文指出微调效果依赖于模型家族，同一LoRA设置无法统一迁移至Aya-Expanse-8B，且低资源价值对齐需要文化监督、通用低资源指令数据以及模型特定的调参。","relevance":"该研究聚焦于LLM的文化价值对齐而非人类仿真，未将模型作为人类被试替代品进行实验或与真实人类行为基准对比，与研究者关注的核心方向不符，不建议深入阅读。","inspiration":"该研究通过结构化调查提炼文化价值观并构建场景指令集的方法，可借鉴用于设计经济决策中的文化或社会规范处理变量。｜可迁移至消费者跨期选择或风险偏好实验，研究特定文化价值观（如节俭、风险规避）对经济行为的影响。｜以LLM为被试，输入嵌入斯里兰卡节俭价值观的场景指令作为处理，测量其在跨期选择任务中的折现率，并与真实斯里兰卡居民的调查数据对照。"}},{"id":"2607.17437","version":1,"title":"Empirical Grounding Improves the Realism of LLM Agents Simulating Human Behavior During Disruptions","zh_title":"经验锚定提升大语言模型代理在中断期间模拟人类行为的真实性","abstract":"Large language model (LLM) agents offer a generative approach to simulating human behavior under conditions that may have few or no direct historical analogues, a common challenge in disaster and infrastructure-disruption planning. However, this generative capacity creates a validity problem: individually plausible agent reasoning may fail to reproduce empirical population behavior. We evaluate whether empirical grounding improves the statistical realism of LLM-agent simulations during disruptions. Specifically, we develop an empirically grounded LLM-agent framework that embeds demographic profiles from the American Community Survey, baseline routines from the American Time Use Survey, and urban spatial context into agent initialization, memory, decision prompts, and activity execution. An independent household survey conducted during the July 2024 Philadelphia heatwave is reserved as an external validation benchmark. Compared with an ungrounded LLM-agent baseline, the grounded model improved reconstruction of normal daily routines, increasing mean correlation with empirical activity profiles from 0.528 to 0.912 and reducing mean squared error from 0.066 to 0.008. Under heatwave conditions, the grounded model better reproduced survey-derived activity profiles, increasing mean correlation from 0.349 to 0.836 and reducing mean squared error from 0.098 to 0.012. The grounded model captured 46.4% of observed heatwave response amplitude, compared with 20.6% for the ungrounded baseline. These findings show that empirical grounding can make LLM agents more statistically credible simulators of population behavior while revealing remaining gaps in modeling human adaptation during disruptions.","authors":["Chen Xia","Zexi Kuang","Yuqing Hu"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-07-19","first_seen":"2026-07-19","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.17437","pdf_url":"https://arxiv.org/pdf/2607.17437","source_feed":"backfill","score":10,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["人类行为仿真","经验锚定","灾害响应"],"reason":"用LLM代理模拟人类在热浪中的行为，并与真实调查数据对照，直接命中核心判据。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:12","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":2,"question":"经验锚定（empirical grounding）能否提高LLM智能体在中断事件中模拟人类行为的统计真实性？","design":"构建经验锚定的LLM智能体框架，将美国社区调查（ACS）人口特征、美国时间使用调查（ATUS）基线日常活动模式及城市空间背景嵌入智能体初始化、记忆、决策提示和活动执行中；以2024年7月费城热浪为场景，比较锚定模型与无锚定基线模型在正常和热浪条件下的活动分布，并以同期独立住户调查作为外部验证基准。","baseline":"2024年7月费城热浪期间进行的独立住户调查，用于验证正常和热浪条件下的活动时间分布及行为变化幅度。","findings":"经验锚定显著提升了LLM智能体对正常日常活动模式的重建，与经验活动剖面的平均相关性从0.528升至0.912，均方误差从0.066降至0.008；在热浪条件下，锚定模型更好地复现了调查活动剖面，平均相关性从0.349升至0.836，均方误差从0.098降至0.012，且捕捉了46.4%的观测热浪响应幅度，远高于无锚定基线的20.6%。","reliability":"论文指出经验锚定虽大幅提升统计可信度，但锚定模型仍仅捕捉到46.4%的观测响应幅度，揭示在模拟人类适应行为方面仍存在差距；此外，研究仅针对单一城市单次热浪事件，泛化性有待检验。","relevance":"该研究直接以真实调查数据为基准，评估LLM智能体模拟人类在中断事件中行为变化的统计真实性，并揭示了经验锚定的增益与残余偏差，与您关注的人类仿真可靠性及失效条件高度吻合，值得精读。","inspiration":"借鉴其将人口普查、时间使用调查和空间背景多层数据嵌入智能体决策流程的设计，可迁移到消费者跨期选择或政策冲击下的行为响应研究｜例如，在突发性价格变动或补贴政策实验中，用LLM智能体模拟不同收入群体的消费调整，以真实家庭收支调查数据为基准，评估仿真对需求弹性和福利效应的复现精度｜设计：以ACS收入、职业和ATUS时间分配数据初始化智能体，施加电价飙升或交通补贴取消的处理，测量各时段活动与消费组合的变化，用实际家庭能源消费或出行调查微观数据作为对照基准。"}},{"id":"2607.18310","version":1,"title":"Distribution-First Population Simulation: Collapse, Calibration, and Recall in Non-WEIRD LLM Persona Modeling","zh_title":"分布优先的人口模拟：非WEIRD LLM角色建模中的崩溃、校准与回忆","abstract":"Synthetic-population tools increasingly run every individual as an independent large language model (LLM) agent. Using real survey microdata, we show that this paradigm has a basic failure mode, and we set a distribution-first corrective against it, all measured with a deterministic, construct-validated verifier on non-WEIRD (Turkey-first) data. First, N independent LLM agents grounded on 2,414 real World Values Survey respondents fail to reproduce the population's response distribution: they pile onto a modal default (four scenarios x five seeds: concentration 0.36->0.69, entropy 1.46->0.77, 85% collapse, TVD=0.44), and the collapse is a predictable function of scenario structure (r=0.55 with a single-answer structure). Second, Verbalized Sampling (VS) fixes the field's chronic under-dispersion without training in three model families (fidelity +7 to +10; significant on Qwen, p=0.002, d=6.2), yet the same move universally overshoots into over-dispersion (SD-ratio 0.4-0.56 -> 1.26-1.37), a structural property of VS. Third, survey fidelity transfers only weakly to agentic behavior: in a single-model, single-domain booking task, a persona is dominated by a cheapest-default (~80%) that income modulates but does not override (comfort choice 0%->7%->32% across income bands). Fourth, a placebo-controlled memorization attack and an election backtest show VS keeps aggregate strength while subgroup and individual claims are contaminated by recall and underdetermination. We close with the corrective: model the distribution once (VS) and assign it to grounded characters at O(1) cost, with a budget-aware router whose honest AUC is 0.805, not the tautological 1.0 of a code-derived oracle. The central contribution needs no realism claim: it measures the internal inconsistency of the independent-agent route and the conditions under which the distribution-first route calibrates.","authors":["Gurkan Ozkan"],"categories":["physics.soc-ph","cs.AI"],"primary_category":"physics.soc-ph","announce_type":"new","date":"2026-07-17","first_seen":"2026-07-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.18310","pdf_url":"https://arxiv.org/pdf/2607.18310","source_feed":"backfill","score":10,"bucket":"selected","rubric_hits":["A1","A2","A3","A5","B1","B2","B3","B4"],"tags":["LLM人类仿真","分布校准","合成人口"],"reason":"用LLM代理模拟真实调查人群，与人类数据对照，评估分布崩溃与校准，涉及行为任务…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:13","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":7,"question":"独立LLM智能体能否复现真实人群的调查回答分布？若不能，如何修正？","design":"使用Qwen、GLM-5.2、Gemma-4-26B等模型，基于土耳其世界价值观调查（WVS）的2414名真实受访者微数据，对比两种路线：路线A（Verbalized Sampling，直接让模型输出总体分布）与路线B（独立智能体，每个角色独立选择后汇总），测量响应分布的集中度、熵、总变异距离等，并检验调查校准向行为任务的迁移。","baseline":"土耳其世界价值观调查（WVS）的2414名真实受访者微数据。","findings":"独立智能体路线导致分布崩溃，个体聚集到模态默认选项，崩溃程度与场景结构相关（单一正确答案时最严重）。Verbalized Sampling能提升分布保真度，但普遍导致过度离散，且调查校准向行为任务的迁移较弱。","reliability":"论文承认Verbalized Sampling存在结构性过度离散，需配合均值保持校正；调查保真度向行为任务迁移弱；子群体和个体层面推断受记忆和欠定污染。","relevance":"直接回应LLM仿真人类调查的分布可靠性问题，提供真实人类基准对照，揭示独立智能体路线的系统性崩溃及分布优先修正的利弊，对经济学实验和政策评估场景的仿真设计有重要参考价值，值得精读原文。","inspiration":"借鉴其对比独立智能体与Verbalized Sampling的仿真设计，通过总变异距离、熵等指标量化分布偏差，并检验调查校准向行为任务的迁移｜可迁移至消费者金融决策调查仿真，如风险偏好、储蓄选择或信贷需求分布预测｜以LLM智能体模拟家庭金融调查受访者，处理为独立智能体 vs. Verbalized Sampling生成风险资产配置分布，结果变量为分布距离与集中度，以真实家庭金融调查微数据为基准"}},{"id":"2607.14485","version":1,"title":"Step-Level Preference Learning for Generative Agents in Social Simulations","zh_title":"面向社会模拟中生成式智能体的步级偏好学习","abstract":"Large language model (LLM)-based generative agents simulate human behavior through long-horizon decision-making processes that comprise intermediate steps such as planning, memory retrieval, reflection, and action selection. However, fine-grained human annotations of these intermediate steps remain scarce, and existing agents are not grounded in human preferences over such intermediate decisions. To address this gap, we introduce \\method, an interactive simulation interface that enables us to collect step-level human preference supervision over agent decision trajectories, leading to a dataset of 57K fine-grained annotations. We conduct step-level preference learning on open-weight language models using supervised finetuning and direct preference optimization on this data, consistently improving simulation fidelity, coordination, and interaction quality, and inducing more socially effective agent behavior. Our results show that step-level human supervision is an effective training signal for improving both local decision quality and long-horizon agent behavior.","authors":["Wenchang Gao","Pingyue Sheng","Lanlan Qiu","Yunfei Ma","Jian Zhao","Baicheng Chen","Kangda Wang","Yuyang Tian","Shunqiang Mao","Tianxing He"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-07-16","first_seen":"2026-07-16","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.14485","pdf_url":"https://arxiv.org/pdf/2607.14485","source_feed":"backfill","score":7,"bucket":"pending","rubric_hits":["A1","B1"],"tags":["LLM仿真","人类偏好学习","社会模拟"],"reason":"用LLM agent模拟人类行为，收集人类偏好数据提升仿真保真度，有真实人类数…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:10","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":67,"question":"如何利用步骤级人类偏好反馈来提升大语言模型智能体在社会模拟中的局部决策质量和长期行为保真度？","design":"使用基于LLM的生成式智能体架构（含规划、记忆检索、反思、行动等模块），在30个设计的社会事件中，由人类标注者通过SimPref交互界面控制2-3个智能体，对每个触发模块的多个候选输出进行偏好选择或自定义，收集57K条步骤级偏好对，然后用监督微调和直接偏好优化训练开源LLM，评估模拟保真度、协调性和交互质量。","baseline":"无对照：论文未提及与真实人类行为数据的直接对比，仅使用人类标注者对智能体中间决策的偏好作为监督信号。","findings":"步骤级偏好学习能持续提升智能体在整体轨迹指标上的模拟保真度、协调性和交互质量，并使智能体更合理地分配活动时间。通过案例研究表明，对齐后的智能体在具体决策中表现出更符合社会预期的行为。","reliability":"论文未讨论失效条件与局限。","relevance":"该研究直接利用人类偏好数据训练LLM智能体以提升社会模拟质量，属于用LLM进行人类仿真实验的工作，但未提供与真实人类数据的基准对比，值得阅读原文以了解其偏好数据构建方法和模拟改进细节。","inspiration":"该方法通过收集人类对智能体步骤级决策的偏好数据来微调LLM，可借鉴用于经济实验中施加细粒度行为干预｜可迁移到消费者跨期选择实验，模拟个体在即时奖励与延迟奖励间的决策动态｜以LLM智能体为被试，处理为基于人类偏好微调的选择策略，结果变量为跨期选择一致性，对照真实人类实验数据（如Andersen et al. 2008）"}},{"id":"2607.14371","version":1,"title":"Supervised Fine-Tuning vs. In-Context Learning: An Equilibrium Analysis of LLM Personalization under Congestion","zh_title":"监督微调与上下文学习：拥塞下LLM个性化的均衡分析","abstract":"Large Language Models (LLMs) have revolutionized AI services, but a critical tension emerges: while personalization improves model performance, it consumes scarce computational resources that users must share. When should a user invest in expensive Supervised Fine-Tuning (SFT) versus lightweight In-Context Learning (ICL)? How does congestion from other users' personalization choices reshape these incentives? And what strategies should platforms adopt when offering multiple personalization algorithms? We develop a tractable framework for LLM serving that captures the statistical-economic trade-offs users face. Our analysis yields several surprising insights. First, we show that ICL and SFT dominate in different regimes, determined by an interplay between pretraining coverage and data signal-to-noise ratios, but congestion can flip these rankings. Second, equilibrium resource consumption exhibits pronounced non-monotonicity: improving pretraining precision reduces the congestion, while broader pretraining coverage and harder tasks sometimes increase it. Third, we prove that offering both personalization methods never hurts the platform's maximal profits, despite potentially increasing computational load. Experiments with GPT-2 on linear regression tasks validate our theoretical predictions about algorithm performance. Complementing these results, our review of documentation from 21 major AI platforms shows that the share offering both SFT and ICL increased from 9.5% in 2021 to 71.4% in 2025, consistent with our platform-design implications.","authors":["Fengzhuo Zhang","Zhuoran Yang","Dirk Bergemann"],"categories":["cs.LG","econ.TH","stat.ML"],"primary_category":"cs.LG","announce_type":"new","date":"2026-07-15","first_seen":"2026-07-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.14371","pdf_url":"https://arxiv.org/pdf/2607.14371","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["LLM个性化","均衡分析","平台策略"],"reason":"研究LLM个性化方法选择与平台策略，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:58:59","error":null,"has_summary":false,"summary":null},{"id":"2607.12219","version":1,"title":"Partial Identification with Multiple Nonlinear Measurements of a Latent Regressor","zh_title":"具有潜变量多个非线性测量的部分识别","abstract":"We study linear regression when the regressor is latent and observed only through multiple noisy measurements, each a smooth but possibly nonlinear function of the latent variable. The problem is acute in the measurement of occupational exposure to artificial intelligence, where competing scores yield downstream estimates that differ by a factor of eleven. A regression on any single measurement recovers a source-specific coefficient rather than the structural one. We fix the latent scale by requiring the consensus measurement function to be linear and bound the remaining curvature heterogeneity across sources relative to slope. Under this bound, the structural coefficient lies in a closed-form interval centered at a symmetric cross-source estimator. The interval is invariant to unknown source loadings, and its half-width is second order in the curvature bound and sharp to the same order. With at least four measurements, the bound is estimable from the joint distribution of the sources through a split-instrument auxiliary regression, and Imbens-Manski confidence intervals with the Stoye critical value attain uniform coverage over the curvature class, including at the point-identified boundary. The application matches six exposure measures to an American Community Survey panel of 8.88 million person-year observations for 2015 to 2024. The post-2022 employment coefficient changes sign between the language-model measures and the Webb patent-text measure, and an ex ante factor-analytic rule separates the Webb measure as a distinct construct. The five retained sources yield a loading-invariant consensus coefficient of -0.239, with a partial-identification half-width of 1.23 percent of the point estimate, or 1.88 percent at the one-sided 95 percent upper bound on the curvature. We read the application as measurement reconciliation rather than as a causal estimate of AI displacement.","authors":["Burhan Ogut","Michelle Yin"],"categories":["econ.EM","cs.AI"],"primary_category":"econ.EM","announce_type":"new","date":"2026-07-13","first_seen":"2026-07-13","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.12219","pdf_url":"https://arxiv.org/pdf/2607.12219","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C5"],"tags":["计量经济学","部分识别","潜变量"],"reason":"论文研究潜变量回归的计量方法，不涉及LLM仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:58:48","error":null,"has_summary":false,"summary":null},{"id":"2607.11722","version":2,"title":"STEP: Career-Path Recommendation via Temporal and Educational Trajectory Modeling","zh_title":"STEP：基于时序与教育轨迹建模的职业路径推荐","abstract":"Career paths encode decades of skill acquisition, role transitions, and educational investment, and understanding them at scale underpins workforce planning, labor market policy, and job recommendation. Resumes are a rich source of information about career paths: they contain detailed descriptions of work experience, education, and skills. Yet their unstructured, heterogeneous, and multilingual nature has long prevented large-scale systematic analysis. With the advent of large language models (LLMs), it is now possible to source rich career trajectory data containing temporal and educational signals from unstructured resumes, enabling new opportunities for career-path recommendation. Exploiting this opportunity, we present STEP (Sequential Trajectory of Employment Prediction), a novel career-path recommendation system that leverages temporal and educational signals to predict the next job in a career trajectory. STEP integrates a time-decay Gated Recurrent Unit (GRU) cell to model temporal dynamics, Feature-wise Linear Modulation (FiLM) conditioned on educational attainment, and attention-based sequence pooling to select relevant features for next job prediction. To improve internal occupation representation for STEP, we introduce ROUTE, a two-stage contrastive procedure that first adapts a multilingual encoder to the career domain via unsupervised denoising autoencoding, then performs supervised contrastive fine-tuning with guided negative selection. We evaluate STEP on four datasets of career trajectories, including an improved version of our publicly available JobHop dataset, and show that it outperforms state-of-the-art baselines in next job prediction. The dataset and code are publicly released to support reproducible career-trajectory research.","authors":["Iman Johary","Guillaume Bied","Alexandru C. Mara","Tijl De Bie"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-07-13","first_seen":"2026-07-13","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.11722","pdf_url":"https://arxiv.org/pdf/2607.11722","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["职业路径推荐","序列预测","多智能体系统"],"reason":"纯多智能体协作预测职业路径，无人类行为对照，不涉及LLM仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:59:23","error":null,"has_summary":false,"summary":null},{"id":"2607.08681","version":1,"title":"SolarChain-Eval: A Physics-Constrained Benchmark for Trustworthy Economic Agents in Decentralized Energy Markets","zh_title":"SolarChain-Eval：面向去中心化能源市场中可信经济智能体的物理约束基准","abstract":"As agentic AI systems are increasingly applied to cyber-physical environments, their evaluation requires assessment of both task performance and trustworthiness. In decentralized energy markets, autonomous agents may improve market utility, but may also exploit invalid physical data, create artificial liquidity, and produce unstable governance decisions. Therefore, we propose SolarChain-Eval, a physics-constrained benchmark for evaluating trustworthy economic agents. It formulates market governance as a Gymnasium-compatible Markov Decision Process, where agents make hourly decisions. SolarChain-Eval evaluates each policy across multiple dimensions, including market utility, physical safety, slippage, action smoothness, spatial fairness, and auditability. To support agentic evaluation, SolarChain-Eval incorporates an LLM-based Planner/Auditor layer. The Planner defines episode-level action bounds and audit rules, while the Auditor reviews and revises high-risk actions. All interventions are recorded through structured logs, including trigger signals, proposed actions, revised actions, and audit rationales. Experiments with static, random, myopic, RL, and RL+LLM policies reveal a clear utility-safety trade-off. RL agents improve market utility but can still produce unsafe behavior. When the physics penalty is removed, reward-maximizing agents exploit invalid generation and increase artificial liquidity. The LLM Planner/Auditor improves auditability and mitigates selected risks, but it cannot fully compensate for a misspecified reward function. These results indicate that trustworthy agentic AI evaluation requires both physical constraints and transparent intervention traces. We release data and code as open access on GitHub for replicability.","authors":["Shilin Ou","Yifan Xu","Luyao Zhang"],"categories":["cs.AI","cs.ET","cs.LG","cs.MA","econ.GN"],"primary_category":"cs.AI","announce_type":"new","date":"2026-07-09","first_seen":"2026-07-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.08681","pdf_url":"https://arxiv.org/pdf/2607.08681","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","能源市场","强化学习"],"reason":"多智能体经济调度仿真，无人类行为对照，属C1排除项。","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:58:28","error":null,"has_summary":false,"summary":null},{"id":"2607.06080","version":1,"title":"From Blueprint to Reality: Modeling and Applying Putnam's Social Capital Theory with LLM-based Multi-agent Simulations","zh_title":"从蓝图到现实：基于LLM多智能体仿真建模与应用普特南社会资本理论","abstract":"Putnam's Social Capital Theory is a foundational framework for collective action and community prosperity. However, traditional empirical methods face practical limits on control and replication. Meanwhile, LLM-based social simulations are typically behavior-driven and lack theory-aligned environments for modeling Putnam's core propositions. To address these gaps, we introduce SocaSim, an LLM-based multi-agent simulation framework to study Putnam's Social Capital Theory from theoretical blueprint to simulated reality. Specifically, we build an environment integrating social network evolution, trust dynamics, and norm propagation, where agents engage in repeated collective-action experiments, and then apply the three dimensions to analyze adaptation challenges in smart elderly care. Our simulations reproduce Putnam's macro-level patterns and exhibit strong human-agent alignment at the group level. Unlike traditional methods, SocaSim traces micro-level causal pathways of social network, trust, and norms via round-by-round simulations and counterfactual interventions, enabling process-level interpretability. Taken together, these capabilities establish a research paradigm that leverages LLM agents to bridge social science and computer science.","authors":["Shiyi Ling","Zhi Zheng","Hui Zheng","Wenjun Xue","Feng Ye","Tong Xu"],"categories":["cs.CL","cs.AI","cs.SI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-07-07","first_seen":"2026-07-07","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.06080","pdf_url":"https://arxiv.org/pdf/2607.06080","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2","B4"],"tags":["LLM人类仿真","社会资本理论","集体行动实验"],"reason":"用LLM多智能体模拟集体行动，复现宏观模式并与人类数据对齐，涉及社会资本理论，…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:10","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":49,"question":"如何基于LLM多智能体仿真复现并应用Putnam社会资本理论，以揭示社会网络、信任和规范在集体行动中的动态机制？","design":"构建SocaSim框架，生成具有人口属性和社会资本禀赋的LLM智能体，在动态社会网络、信任和规范环境中进行多轮集体行动实验，通过提案和执行两阶段模拟决策，并应用于智慧养老适应挑战，进行反事实干预。","baseline":"与真实老年人群体决策数据进行对比，报告群体层面决策的皮尔逊相关系数为0.974。","findings":"仿真复现了Putnam理论预测的宏观模式，并与人类群体决策高度一致；反事实干预显示，提高低社会经济地位智能体的初始信任可使技术采纳率提升15.4%，决策矛盾减少25.5%。","reliability":"论文未讨论","relevance":"该研究直接以LLM智能体替代人类被试，复现社会资本理论的集体行动模式，并与真实人类数据对齐，同时进行反事实干预评估因果效应，高度契合研究者对经济学实验和政策评估场景中仿真可靠性的关注。","inspiration":"借鉴其利用LLM智能体在动态社会网络中进行多轮集体行动实验并施加反事实干预的设计，以模拟宏观涌现模式｜可迁移至政策干预对技术采纳或合作行为影响的经济学实验，如智慧养老、金融包容性政策评估｜以LLM智能体作为被试，处理为提高低社会经济地位群体的初始信任水平，结果变量为技术采纳率和决策矛盾率，用真实老年人群体决策数据作为对照基准"}},{"id":"2607.05861","version":1,"title":"Mitigating Factual Hallucination in Large Reasoning Models via Mixed-Mode Advantage Regularization","zh_title":"通过混合模式优势正则化缓解大型推理模型中的事实性幻觉","abstract":"Large reasoning models (LRMs) improve language model capabilities by generating explicit thinking traces before final answers. In factuality-oriented question answering (QA), such thinking often improves overall performance by helping the model recover relevant knowledge and refine its answers. However, we find that this benefit is not uniform at the instance level: explicit thinking can also overturn correct non-thinking answers and lead to factual drift. We refer to this failure mode as \\emph{thinking-induced hallucination}. To explain this phenomenon, we formulate explicit thinking in factuality QA as a thinking residual over the model's direct-answer tendency, which can either recover missing knowledge or introduce unsupported associations. Based on this formulation, we propose MARGO, \\underline{\\textit{M}}ixed-Mode \\underline{\\textit{A}}dvantage \\underline{\\textit{R}}egularization for \\underline{\\textit{G}}rounded \\underline{\\textit{O}}ptimization, a reinforcement learning framework that uses non-thinking rollouts as same-model references in advantage estimation. By constructing mixed-mode rollout groups with both thinking and non-thinking trajectories, MARGO evaluates whether explicit thinking adds factual value beyond direct answering, thereby suppressing hallucination-prone thinking while preserving beneficial thinking behaviors. Experiments across multiple factuality-oriented QA benchmarks demonstrate that MARGO improves factual reliability over strong baselines, while evaluations on mathematical benchmarks show that it preserves general reasoning ability.","authors":["Kaishen Wang","Tong Zheng","Xuehao Cui","Ruibo Chen","Tianyi Xiong","Heng Huang"],"categories":["cs.CL","cs.LG"],"primary_category":"cs.CL","announce_type":"new","date":"2026-07-07","first_seen":"2026-07-07","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.05861","pdf_url":"https://arxiv.org/pdf/2607.05861","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["事实性幻觉","强化学习","推理模型"],"reason":"纯NLP事实性幻觉缓解研究，无人类仿真或行为对照。","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:59:05","error":null,"has_summary":false,"summary":null},{"id":"2607.03091","version":1,"title":"Silicon Sampling via Cross-Survey Transfer","zh_title":"基于跨调查迁移的硅采样","abstract":"Silicon sampling-using large language models (LLMs) to simulate human survey respondents-has emerged as a promising approach for augmenting traditional survey research. However, most evaluations rely on distributional comparisons rather than individual-level prediction, which risks conflating pattern matching with coherent respondent-level prediction. We propose cross-survey transfer, a more rigorous evaluation framework in which an LLM is given a respondent's answers to one set of questions and must predict their answers to entirely different questions from the same survey. Using data from the Taiwan Election and Democratization Study (TEDS) 2024, three open-weight LLMs (27B-120B parameters), and supervised machine learning baselines, we find that: (1) zero-shot LLMs achieve 52% accuracy on genuinely unseen items, closing to within 6 percentage points (pp) of a supervised random forest trained on same-population data; (2) a stable construct predictability hierarchy emerges, from 67% for partisan attitudes to 23% for sovereignty; and (3) variance collapse and safety alignment effects-two commonly cited LLM limitations-turn out to be more nuanced than previously reported, with variance collapse affecting supervised models as well and alignment effects varying dramatically across model families. These findings clarify both the promise and boundaries of silicon sampling.","authors":["Chan-Tung Ku","Chan Hsu","Pei-Cing Huang","Frank Cheng-shan Liu","I-Ling Cheng","Yihuang Kang"],"categories":["cs.AI","cs.CL","cs.CY","cs.MA","stat.ME"],"primary_category":"cs.AI","announce_type":"new","date":"2026-07-03","first_seen":"2026-07-03","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.03091","pdf_url":"https://arxiv.org/pdf/2607.03091","source_feed":"backfill","score":10,"bucket":"selected","rubric_hits":["A1","A2","A5","B1","B2","B4"],"tags":["LLM人类仿真","调查方法","算法保真度"],"reason":"直接用LLM仿真人类调查回答，有真实人类数据对照，评估可靠性与偏差，涉及选举研…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:09","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-03","rank":2,"question":"大语言模型能否在个体层面预测人类调查受访者的回答，以及这种仿真的可靠性和边界是什么？","design":"使用三个开源LLM（27B-120B参数）模拟台湾选举与民主化调查（TEDS 2024）的受访者，采用跨调查迁移框架：给定受访者对一组问题的回答，预测其对同一调查中完全不同问题的回答。","baseline":"TEDS 2024真实人类调查数据，以及基于同总体数据训练的监督随机森林模型。","findings":"零样本LLM在未见项目上达到52%准确率，仅比监督随机森林低6个百分点；不同构念的可预测性存在稳定层级，从政党态度的67%到主权问题的23%。","reliability":"方差压缩和安全对齐效应比先前报告更复杂：方差压缩同样影响监督模型，对齐效应在不同模型家族间差异显著。","relevance":"高度相关：直接评估LLM作为人类调查受访者替代品的个体层面预测能力，有真实人类数据对照，并揭示了仿真的边界（如构念可预测性差异、方差压缩与对齐效应的复杂性），值得精读原文。","inspiration":"跨调查迁移框架将受访者部分回答作为输入预测其余回答，可借鉴用于构建个体层面的行为预测模型，并设置监督机器学习基准和真实人类数据对照以评估仿真可靠性。｜可迁移到消费者跨期选择实验，用LLM模拟被试在给定部分偏好或决策后的选择一致性，检验时间偏好与自我控制偏差。｜以LLM作为被试，用真实家庭金融调查数据（如CFPS）中部分消费-储蓄问题回答作为处理输入，预测同一被试在跨期选择任务中的折现因子作为结果变量，与真实人类数据及监督模型预测对比。"}},{"id":"2607.05440","version":1,"title":"Retrieval over Reasoning: A Cost-Controlled Benchmark of Language Models for Energy-Retrofit Recommendation","zh_title":"检索优于推理：面向建筑节能改造推荐的语言模型成本控制基准测试","abstract":"Recommending the correct set of energy conservation measures (ECMs) for a building is a structured, multi-label prediction problem in which a task-specific supervised model has weak training signal and a general language model has no grounding in the local building stock. We study this problem on 10,422 real New York City Local Law 87 (LL87) energy-audit records, taking as ground truth the set of ECM categories that certified auditors actually recommended. We make four contributions. First, we establish that energy-use-intensity (EUI) prediction - the upstream task - is effectively solved by tree ensembles: across fifteen trained models, a stacking ensemble reaches a coefficient of determination R^2 = 0.757, and every one of six neural architectures is outperformed by gradient-boosted trees. Second, we show that the framing of the recommendation task dominates model choice: recasting ECM recommendation as 19-way multi-label classification rather than single-label categorization lifts a gradient-boosted-tree baseline from a previously reported 25.9% accuracy to a micro-F1 of 0.571. Third, we benchmark eight large language models (LLMs) from four providers in a 2x2 design that independently toggles retrieval grounding and explicit reasoning, scoring each arm on per-label F1, U.S.-dollar cost per building, and latency; retrieval-augmented generation (RAG) improves micro-F1 by +0.11 to +0.20 on every model, while explicit reasoning yields no measurable accuracy change (-0.018 to +0.010) at up to 8.4x the cost. Fourth, we show LLMs systematically over-recommend - high recall, low precision - and that retrieval closes the gap chiefly by improving precision. A 70-billion-parameter open-weight model with a fifteen-line nearest-neighbor retrieval step reaches 0.511 micro-F1 at $0.00032 per building, comparable to a frontier model at roughly 10.1x lower cost.","authors":["Eliseo Curcio"],"categories":["econ.EM","eess.SY"],"primary_category":"econ.EM","announce_type":"new","date":"2026-07-03","first_seen":"2026-07-03","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.05440","pdf_url":"https://arxiv.org/pdf/2607.05440","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["建筑节能","LLM基准测试","多标签分类"],"reason":"论文用LLM做建筑节能推荐，属纯NLP能力评测，无人类行为仿真对照。","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:58:48","error":null,"has_summary":false,"summary":null},{"id":"2607.00551","version":2,"title":"Talking Politics with Artificial Intelligence","zh_title":"与人工智能谈论政治","abstract":"Large language models (LLMs), a prominent form of artificial intelligence (AI), are becoming everyday interfaces for political questions, but most exchanges are dyadic rather than audiencefacing. This paper asks whether AI conversation functions as a new arena for political expression or as a conversational intermediary for routine political demand. Using 4.30 million humanAI conversations from three large public datasets, we apply two validated classifiers to user messages, identifying political content, use case, and expressed ideology. Political content appears in 3.9% of conversations, varies sharply by platform publicness and conversation depth, and is mostly practical: users ask for information, draft text, and process documents far more often than they state opinions. A regression-discontinuity-in-time design around the 2024 U.S. presidential result call shows that the call changed the expressive subset: among U.S. users, stance-taking, affective language, and ideological extremity rose; comparable conversations elsewhere did not. AI conversation is less a public square than a conversational political intermediary, absorbing routine demand and becoming expressive when major events make political stakes explicit.","authors":["Ziwen Zu"],"categories":["econ.GN"],"primary_category":"econ.GN","announce_type":"new","date":"2026-07-01","first_seen":"2026-07-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.00551","pdf_url":"https://arxiv.org/pdf/2607.00551","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["政治表达","LLM对话分析","用户行为"],"reason":"分析人类与AI对话中的政治表达，测量用户行为而非用LLM仿真人类被试，属边界情…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:18:03","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":169,"question":"AI对话是成为政治表达的新场域，还是充当日常政治需求的对话中介？","design":"非仿真研究。分析三个公开数据集中430万条人类与AI对话的用户消息，用两个分类器识别政治内容、用途和意识形态，并以2024年美国总统大选结果公布为断点进行时间断点回归。","baseline":"无对照","findings":"政治内容仅占对话的3.9%，且多为实用目的（信息查询、文本起草、文档处理），而非意见表达；大选结果公布后，美国用户的政治对话中立场表达、情感语言和意识形态极端性上升，而其他地区无此变化。","reliability":"论文未讨论","relevance":"该研究测量真实用户行为而非用LLM仿真人类被试，与研究者关注的核心方向不符，但提供了LLM作为政治中介的实证背景，对理解仿真生态有参考价值，可酌情略读。","inspiration":"该研究利用时间断点回归识别外生事件（大选结果公布）对用户行为的因果效应，方法上可借鉴自然实验的断点设计来评估政策冲击。｜可迁移至政策公告对投资者预期形成的影响研究，例如央行利率决议或财政刺激方案发布后市场参与者的信息需求与情绪变化。｜以LLM模拟投资者作为被试，处理为呈现真实政策公告文本，结果变量为模拟投资者在对话中的信息查询频率与情感极性，对照真实市场数据如社交媒体讨论或交易量变化。"}},{"id":"2606.30372","version":1,"title":"Using Large Language Models as Low-Cost Statistical Estimators for Human-Response Data","zh_title":"使用大语言模型作为人类响应数据的低成本统计估计器","abstract":"Quantitative research across the social and behavioral sciences depends on human subject experiments that are expensive, slow, and subject to sampling bias. Here we show that pretrained large language models induce risk-equivalent estimators of conditional expectations under squared loss, establishing restricted functional risk equivalence: under squared loss, the LLM induces an estimator whose risk matches the Bayes optimal risk for squared-loss prediction of conditional expectations for any inference that depends on the data only through the conditional mean. We formalize the LLM as a misspecified functional estimator $T(\\hat{P}_n)$ trained on i.i.d.\\ data, decompose the estimation error into representation bias $ε_{\\mathrm{rep}}$ and optimization error, and prove that under mild regularity conditions the LLM's expected error converges to the irreducible population variance plus the squared representation bias, with the representation bias bounded by the Pinsker inequality. The identifiability error $δ$ propagates into the effective bias, inflating the asymptotic risk floor. We establish restricted functional risk equivalence via a bidirectional Le Cam deficiency analysis: the forward deficiency vanishes asymptotically while the reverse deficiency is exactly zero. We provide finite-sample concentration bounds and a calibration protocol with explicit decision rules. The result is a precise, provable statement: a well-calibrated LLM achieves the Bayes-optimal risk for conditional-mean-dependent inference, bounded by explicit scope conditions. In practical applications, this means that under satisfied conditions and well-calibrated models, large language models can be used in many prediction and decision-making tasks that originally relied on human experiments, approximating near-optimal statistical inference at lower cost.","authors":["Haobo Yang"],"categories":["cs.AI","cs.CY","cs.HC"],"primary_category":"cs.AI","announce_type":"new","date":"2026-06-29","first_seen":"2026-06-29","revised_at":null,"abs_url":"https://arxiv.org/abs/2606.30372","pdf_url":"https://arxiv.org/pdf/2606.30372","source_feed":"backfill","score":7,"bucket":"pending","rubric_hits":["A2","B4"],"tags":["LLM仿真","统计估计","风险等价性"],"reason":"论文提出用LLM作为人类响应数据的统计估计器，评估其风险等价性，涉及仿真可靠性…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:07","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":73,"question":"预训练大语言模型能否作为人类响应数据的低代价统计估计器，在平方损失下达到与基于人类被试数据的最优估计相同的风险？","design":"本研究不是仿真实验，而是理论证明：将LLM视为一个在i.i.d.数据上训练的误设定函数估计器，通过分解估计误差为表示偏差和优化误差，在平方损失下分析其条件期望估计的风险等价性。","baseline":"无对照","findings":"在满足条件（定量连续响应、离散条件、i.i.d.训练数据、校准验证）且模型校准良好时，LLM的条件均值估计器在平方损失下可达到贝叶斯最优风险；表示偏差由Pinsker不等式界定，可识别性误差会放大渐近风险下限。","reliability":"论文明确限定了适用范围：不适用于定性研究、训练数据中无类似范式的新任务、安全关键应用和行为机制研究；风险等价性依赖于校准验证和条件满足。","relevance":"高度相关：该文为用LLM替代人类被试进行统计推断提供了严格的理论基础，直接回应了仿真可靠性与偏差问题，并给出了明确的失效条件，值得精读。","inspiration":"该文将LLM视为误设定函数估计器，通过分解表示偏差与优化误差来理论证明风险等价性，为仿真可靠性提供了严格的统计框架，可借鉴其校准验证与条件限定的思路来设计稳健性检验｜可迁移到政策公告的预期形成研究，例如分析央行沟通对金融市场参与者通胀预期的影响，用LLM模拟分析师或投资者对政策文本的反应分布｜用LLM作为被试，处理为不同措辞的央行声明（如鹰派/鸽派），结果变量为LLM生成的通胀预测数值，以专业预测者调查（SPF）的真实分布作为对照基准，检验LLM估计的条件均值是否与人类预测的贝叶斯最优风险一致"}},{"id":"2606.30986","version":1,"title":"The Organizational Behavior of Agentic AI: Collective Intelligence in Human-Agent Workflows","zh_title":"智能体AI的组织行为：人-智能体工作流中的集体智能","abstract":"Agentic artificial intelligence is increasingly deployed not as a single assistant but as a collective of planners, solvers, reviewers, memory managers, tool users, and orchestrators. These systems are entering organisational workflows under familiar labels such as teams, managers, committees, markets, and workflows. This article asks whether such agent collectives exhibit organisational behaviour in a sense that is analytically comparable to, yet distinct from, human organisational behaviour. I argue that agentic AI is a partial organisational analogue. It resembles a human organisation because it differentiates work, coordinates interdependence, performs recurrent routines, crosses boundaries, and produces collective outcomes. It differs because these patterns are not sustained by motivation, identity, trust, employment, socialisation, or moral accountability. They are sustained by context architecture: prompts, memory, traces, schemas, tools, validators, and permissions. The article develops contextual transaction cost as the central mechanism linking these similarities and differences. Computational theorising, synthetic task simulations, real LLM agent traces, and robustness analyses show that human-imitation forms often underperform when they add lossy handoffs, correlated deliberation, and verification burdens, whereas shared-state and adaptive forms perform better when they make context durable, inspectable, and task-contingent. The article contributes to organisation studies by theorising agentic AI as an emerging object of organising and by specifying the interface conditions under which human and agentic organisational behaviour can jointly support collective intelligence.","authors":["Canhui Liu"],"categories":["cs.CY","cs.HC","cs.MA","econ.GN"],"primary_category":"cs.CY","announce_type":"new","date":"2026-06-29","first_seen":"2026-06-29","revised_at":null,"abs_url":"https://arxiv.org/abs/2606.30986","pdf_url":"https://arxiv.org/pdf/2606.30986","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["组织行为模拟","多智能体系统","集体智能"],"reason":"用LLM agent模拟组织行为，但无真实人类数据对照，属社会模拟边界情形。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:18:03","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":170,"question":"智能体AI集体是否表现出与人类组织行为可比较但又不同的组织行为，其相似性与差异的机制是什么？","design":"本文并非直接模拟特定人群，而是通过计算理论化、合成任务模拟（8000个任务，7种智能体组织形式）、真实LLM智能体轨迹分析（软件修复、长文档问答、法律审查、文献综合等任务）以及稳健性检验，比较不同智能体组织架构（如模仿人类形式与共享状态/自适应形式）在任务完成中的表现，并分析上下文交易成本。","baseline":"无对照","findings":"智能体AI集体在功能上类似人类组织，面临分工、协调、惯例、边界和集体知识等问题，但其行为模式由上下文架构（提示、记忆、轨迹、模式、工具、验证器、权限）而非人类的社会嵌入性维持。模仿人类组织的形式（如层级、委员会）常因有损交接、相关审议和验证负担而表现不佳，而共享状态和自适应形式在使上下文持久、可检查和任务相关时表现更好。","reliability":"论文指出，模仿人类组织的形式在增加有损交接、相关审议和验证负担时往往表现不佳，且智能体AI缺乏动机、身份、信任、雇佣、社会化或道德责任等人类组织的社会基础，其行为依赖于上下文架构的设计，这构成其失效条件与局限。","relevance":"本文虽用LLM智能体模拟组织行为，但无真实人类数据对照，属于社会模拟的边界情形，与您关注的有基准人类数据的仿真研究不完全匹配，但其中关于仿真失效条件（如模仿人类形式时的性能下降）的批判性讨论可能对您有参考价值，建议快速浏览其机制分析部分。","inspiration":"借鉴其通过合成任务和真实LLM轨迹分析不同组织架构（如层级、共享状态）对任务完成效率影响的实验设计，可迁移到金融分析师团队预测协作场景，研究不同AI协作结构对盈利预测准确度的影响，设计为以LLM智能体模拟分析师团队，处理为层级式与共享信息式协作架构，结果变量为预测误差，对照真实分析师团队历史预测数据。"}},{"id":"2606.30987","version":1,"title":"Measuring Judgment Quality in Natural-Language Explanations: Evidence from Forecasting Tournaments","zh_title":"衡量自然语言解释中的判断质量：来自预测锦标赛的证据","abstract":"Decision-makers routinely rely on expert judgments accompanied by written explanations, yet explanation quality is difficult to measure at scale. Forecasting tournaments offer a natural testing ground: probabilistic judgments are paired with natural-language rationales and scored against realized outcomes. We introduce Explanation Quality Markers (EQMs), a set of sixty theory-guided reasoning patterns scored by large language models (LLMs). In a pre-registered analysis of over 55,000 forecast-rationale pairs from a multiyear forecasting tournament, EQMs predict accuracy at both the forecast and forecaster levels, consistently outperforming pre-LLM text-analysis methods. More than 90% of statistically significant pattern-level EQM-accuracy correlations match our directional hypotheses. The signal is asymmetric: EQMs identify likely underperformers more reliably than they distinguish the very best forecasters. Benchmarked against traditional indicators of forecasting skill, EQMs are the strongest predictor at the forecast level and competitive at the forecaster level, though weaker than prior accuracy. Human ratings of rationale quality are less consistently correlated with accuracy and place disproportionate weight on rationale length. Results transfer to an independent forecasting study. EQMs provide a scalable, interpretable method for extracting judgment-relevant information from written explanations.","authors":["Christopher W. Karvetski","Sheldon S. Huang","Simas Kučinskas","Nadja Flechner","Jingyu Hu","Philip Tetlock","Ezra Karger"],"categories":["cs.CL","econ.GN"],"primary_category":"cs.CL","announce_type":"new","date":"2026-06-29","first_seen":"2026-06-29","revised_at":null,"abs_url":"https://arxiv.org/abs/2606.30987","pdf_url":"https://arxiv.org/pdf/2606.30987","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["LLM评分","预测锦标赛","解释质量"],"reason":"用LLM评分解释质量，属NLP评测，非仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:58:13","error":null,"has_summary":false,"summary":null},{"id":"2606.30473","version":1,"title":"Field Order Should Not Matter: Permutation-Invariant Embedding Model Fine-Tuning for Structured Metadata Retrieval","zh_title":"字段顺序不应重要：面向结构化元数据检索的置换不变嵌入模型微调","abstract":"We study retrieval over catalogs of structured metadata, where each record is a small schema whose fields answer different kinds of query. Embedding a record with a text encoder first serializes its fields into a string, which forces a choice of field order. We show this choice, usually treated as an implementation detail, silently controls retrieval quality once the encoder is fine-tuned. A standard fine-tune loses 7.4 nDCG@10 points when the index is rebuilt under a different field order, because it reads absolute position instead of the field labels. We propose permutation-invariant fine-tuning ($\\textbf{PI-FT}$), which serializes each record under a freshly sampled field order with random field dropout, so meaning binds to the labels rather than to position. The change is about two lines in the data loader; it costs negligible in-distribution accuracy and cuts the order-change penalty to 0.2 points. We study this in the discovery of development statistics, a catalog of nearly 10,000 indicators that should be searchable in many languages by a model small enough to self-host. As AI assistants and agents increasingly mediate access to public data and statistics, this retrieval step decides whether an answer is grounded in the right indicator or series, making discoverability a precondition for disseminating data through AI. Because usage logs cannot provide training signal for indicators no one has searched, we generate the queries instead. $\\textbf{DevDataBench}$ is a fully LLM-generated benchmark of grounded, facet-targeted queries across 15 languages, covering every indicator for both training and evaluation. A fine-tuned 118M-parameter CPU encoder outperforms every zero-shot baseline, including $\\texttt{text-embedding-3-large}$ (0.707 vs.\\ 0.556 nDCG@10), with the largest gains in low-resource languages. We release the benchmark, pipeline, models, and a reusable PI-FT framework.","authors":["Aivin V. Solatorio","Olivier Dupriez","Rafael Macalaba"],"categories":["cs.CL","cs.AI","cs.IR","cs.LG","econ.GN"],"primary_category":"cs.CL","announce_type":"new","date":"2026-06-29","first_seen":"2026-06-29","revised_at":null,"abs_url":"https://arxiv.org/abs/2606.30473","pdf_url":"https://arxiv.org/pdf/2606.30473","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["信息检索","嵌入模型","元数据"],"reason":"研究结构化元数据检索的嵌入模型微调，不涉及人类行为仿真或对照，属于纯信息检索评…","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:58:28","error":null,"has_summary":false,"summary":null},{"id":"2606.28978","version":1,"title":"Can LLMs Hire Fairly? Racial Bias in Resume Screening","zh_title":"大语言模型能公平招聘吗？简历筛选中的种族偏见","abstract":"We audit fourteen mainstream large language models (LLMs) for hiring discrimination using the paired-resume methodology of Kline, Rose, and Walters (2022). The sole 2023-vintage model reproduces the pro-White callback gap documented in field experiments on labor market discrimination ($+2.12$ pp, significant at the 1\\% level). Every model released in 2024 or after shows either a null gap or a significant pro-Black reversal (up to $-3.01$ pp). The same pattern holds on the gender axis. Based on 24,024 paired postings per model across 14 models, our results document a reversal in the direction of algorithmic hiring bias across model generations.","authors":["Zhenyu Gao","Wenxi Jiang","Yutong Yan"],"categories":["cs.CL","cs.CY"],"primary_category":"cs.CL","announce_type":"new","date":"2026-06-27","first_seen":"2026-06-27","revised_at":null,"abs_url":"https://arxiv.org/abs/2606.28978","pdf_url":"https://arxiv.org/pdf/2606.28978","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["LLM仿真","招聘歧视","算法偏差"],"reason":"用LLM复现招聘歧视实验，与真实人类数据对照，评估算法偏差，直接命中核心判据。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:18:03","error":null,"has_summary":true,"summary":{"generated_at":"2026-06-27","rank":4,"question":"大语言模型在简历筛选时是否存在种族和性别歧视，以及这种歧视的方向是否随模型版本变化？","design":"使用14个主流LLM（GPT-3.5-turbo、GPT-4o-mini、GPT-5.4-mini、Llama-3.1-8B-Instruct、Llama-3.3-70B等），采用配对简历方法，简历内容完全相同仅名字不同（黑人或白人典型名字，或男女性名字），让模型扮演HR经理决定是否邀请面试（输出“yes”或“no”），温度设为0，每个模型在种族轴上进行24,024次配对测试，在性别轴上进行48,048次配对测试。","baseline":"对照真实人类实验：Kline, Rose, and Walters (2022)的现场实验，在108家财富500强公司中发现白人名字比黑人名字获得更多回电（1.6个百分点）；以及Bertrand and Mullainathan (2004)的经典实验。","findings":"2023年发布的GPT-3.5-turbo复制了真实实验中的亲白人偏差（白人名字回电率高2.12个百分点，1%显著），而2024年及之后发布的所有模型要么无偏差，要么出现显著的反向偏差（亲黑人偏差达0.4-3.0个百分点）。性别轴上也呈现相同反转：GPT-3.5-turbo亲男性（1.92个百分点），后续模型亲女性或无偏差。","reliability":"论文未讨论失效条件与局限。","relevance":"高度相关：该研究直接使用LLM模拟人类招聘决策，有真实人类数据作为对照基准，发现偏差方向随模型版本反转，符合研究者对经济学实验和政策评估场景以及批判性失效条件的关注。值得精读原文。","inspiration":"采用配对简历设计，仅改变名字暗示种族/性别，保持其他信息不变，以隔离身份信号对决策的影响，并设置温度0确保确定性输出｜可迁移到信贷审批歧视研究，检验LLM在贷款申请评估中是否对申请人姓名产生种族或性别偏差｜用LLM作为信贷审批员，处理仅名字（典型白人/黑人/男性/女性名）不同的标准化贷款申请，输出批准/拒绝决策，以真实银行信贷审批数据中的种族/性别差异作为对照基准"}},{"id":"2606.28770","version":1,"title":"Mechanistic Personality Analysis of LLMs Steering Personality via Latent Feature Interventions","zh_title":"大语言模型的机制性人格分析：通过潜在特征干预操控人格","abstract":"Large Language Models (LLMs) have demonstrated the ability to simulate human-like OCEAN personality traits in generated text. Previous efforts have focused on prompt engineering or fine-tuning to shape LLM personality. In this work, we propose a mechanistic interpretability approach that directly intervenes on the model's latent features. Our method identifies latent directions in the residual stream corresponding to a target OCEAN trait using sparse autoencoders (SAEs) and contrastive activation analysis. We formalize an additive steering vector in activation space and demonstrate how applying a small additive shift to the hidden states enhances the target trait while preserving overall language modeling performance. To determine the optimal combination of feature shifts, we explore a linear weighting heuristic with grid search optimization that balances personality expression with task performance. Our approach shows promise in controllably steering personality traits at the mechanistic level while maintaining high performance on standard benchmarks.","authors":["David Courtis","Ting Hu"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-06-27","first_seen":"2026-06-27","revised_at":null,"abs_url":"https://arxiv.org/abs/2606.28770","pdf_url":"https://arxiv.org/pdf/2606.28770","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["LLM人格","机制可解释性","稀疏自编码器"],"reason":"测量并操控LLM自身的人格特质，属于D2边界情形，非仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:18:02","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":136,"question":"如何通过干预大语言模型的潜在特征来操控其表现出的大五人格特质？","design":"本研究并非人类仿真实验，而是提出一种机制可解释性方法：使用稀疏自编码器从模型残差流中提取与目标人格特质相关的潜在方向，通过对比激活分析构建加性操控向量，对隐藏状态施加小幅度偏移，从而增强目标特质表达，同时保持语言建模性能。","baseline":"无对照","findings":"通过稀疏自编码器可识别出与OCEAN特质对应的可解释潜在特征；在激活空间中施加加性偏移能有效操控模型输出的人格特质，且通过网格搜索优化的线性加权策略可在特质表达与任务性能间取得平衡。","reliability":"论文未讨论","relevance":"本文聚焦于操控LLM自身人格特质的机制，而非用LLM仿真人类被试，不涉及真实人类数据对照或政策评估场景，与研究者关注的仿真可靠性及偏差问题关联较弱，不建议优先阅读。","inspiration":"与经济金融研究关联不大"}},{"id":"2606.26883","version":2,"title":"EconSimulacra: A Digital Twin Platform of Socio-Economic Systems Powered by LLM Agents","zh_title":"EconSimulacra：基于LLM智能体的社会经济系统数字孪生平台","abstract":"Real-world social behavior emerges from tightly coupled domains: economic conditions shape mobility and social interactions, while online attention and offline activity feed back into local popularity and consumer behavior. Capturing these feedback loops requires artificial societies in which agents carry experiences from one domain into decisions in another. Large language models (LLMs) provide a promising foundation for such societies. However, existing LLM-based simulators typically model domains in isolation or merely place them side by side. To enable such cross-domain interactions, we present EconSimulacra, a multi-agent social simulator that couples consumer economy, mobility, and social networks through a shared internal-state mechanism. In EconSimulacra, experiences accumulated across different domains are stored in memory and transformed into shared internal states (i.e., stress level) connecting heterogeneous domains through individual decision making. This design allows agents to reconcile competing demands arising from multiple domains and generate coherent cross-domain behaviors. As a case study, we show that the shared internal state mechanisms reproduce a nonlinear relationship between online social attention and offline local popularity, illustrating how realistic cross-domain dynamics can emerge within a unified artificial society.","authors":["Ryuji Hashimoto","Masahiro Kaneko","Kentaro Ueda","Takehiro Takayanagi","Kiyoshi Izumi"],"categories":["cs.DL"],"primary_category":"cs.DL","announce_type":"new","date":"2026-06-25","first_seen":"2026-06-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2606.26883","pdf_url":"https://arxiv.org/pdf/2606.26883","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["LLM智能体","社会经济模拟","跨域交互"],"reason":"多智能体社会经济模拟，但无真实人类数据对照，属边界情形。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:05","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":149,"question":"如何通过共享内部状态机制实现多领域（消费经济、出行、社交网络）耦合，从而涌现出跨领域动态？","design":"使用 LLM 驱动的多智能体模拟器 EconSimulacra，让家庭智能体在网格世界中同时进行消费、出行和社交决策，通过压力水平作为共享内部状态耦合各领域；施加餐厅折扣活动作为外生干预，测量在线社交关注度与线下消费活动之间的关系，并进行消融实验（移除压力耦合）对比。","baseline":"无对照","findings":"引入压力水平耦合后，模拟复现了在线社交关注与线下消费活动之间的非线性关系；移除压力耦合机制会削弱这一涌现模式。","reliability":"论文未讨论","relevance":"该研究属于 LLM 驱动的社会经济多智能体仿真，但缺乏真实人类数据对照，且未评估仿真的可靠性与偏差，与研究者关心的基准对照和批判性评估需求匹配有限，可作为边界案例参考。","inspiration":"该研究通过共享内部状态（压力水平）耦合多个决策领域，并施加外生干预（折扣活动）观察跨领域涌现模式，这种设计可借鉴用于研究多市场联动｜可迁移到消费者跨期选择与信贷行为的耦合研究，例如考察促销活动如何通过心理压力影响消费信贷使用｜以LLM智能体为被试，施加限时折扣作为处理，测量消费支出与信贷申请量，用真实消费者面板数据（如银行交易记录）对照跨领域行为关联"}},{"id":"2606.25484","version":1,"title":"From Causal Discovery to Implementation: An Agentic AI Framework for E-Scooter Mobility Hub Planning Across 29 German Cities","zh_title":"从因果发现到实施：一个面向29个德国城市电动滑板车移动枢纽规划的智能体AI框架","abstract":"Existing approaches to e-scooter mobility hub planning lack city-type-specific causal evidence. Demand models are typically correlational, built on proprietary trip data, and do not distinguish how driver profiles vary across urban typologies. This paper presents a three-phase agentic AI framework that constructs a Causal Template Library from public GBFS data across 29 German cities, encoding which environmental features causally drive hotspot demand for each combination of city type (large, university, industrial, hilly) and cluster type (core, peripheral). A large language model (LLM) orchestrated causal discovery pipeline adapts algorithm selection to local data conditions across 57 city-cluster units. The library reveals systematic variation. Core demand is driven by activity access and transit proximity, while peripheral demand responds to built form, with city-type-specific patterns supporting transferable siting templates. A planning tool built on the library scores candidate sites, calibrates infrastructure recommendations to local demographics, and generates practitioner-ready reports. In Heilbronn, Germany, two hub sites informed by the framework's causal evidence are currently under construction, illustrating how the outputs can support real-world siting decisions.","authors":["Meng Jin","Melanie Handrich","Simone Martinenz","Nicholas Hoeser","Ziyue Li"],"categories":["cs.CY","econ.GN","stat.AP"],"primary_category":"cs.CY","announce_type":"new","date":"2026-06-24","first_seen":"2026-06-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2606.25484","pdf_url":"https://arxiv.org/pdf/2606.25484","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["智能体AI","因果发现","城市规划"],"reason":"多智能体协作规划电动滑板车站点，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:58:14","error":null,"has_summary":false,"summary":null},{"id":"2606.22797","version":1,"title":"Measuring Behavior Portability in Large Language Models","zh_title":"测量大语言模型的行为可移植性","abstract":"Large language models are increasingly deployed as autonomous decision makers, yet the behavioral mapping they exhibit can vary substantially across decision environments that are payoff-equivalent by construction-environments that share identical payoff-relevant structure but differ in surface presentation. This sensitivity renders suite-based evaluation fragile and raises a fundamental question of behavioral portability: how well does a behavioral mapping learned in one decision environment informative on another that preserves the same underlying incentive structure? We introduce a formal framework to measure this property. Our protocol fits an interpretable behavioral model on data pooled from a set of source environments and evaluates its out-of-sample predictive performance in a held-out target environment, benchmarking against an oracle trained directly on target data. Portability is quantified via a loss-agnostic measure that delivers worst-case bounds on the performance of the induced prediction-action mapping in the target environment. In controlled experiments spanning seven canonical economic decision problems, we document substantial and systematic portability losses, suggesting that behavioral characterizations of LLMs obtained in one decision environment cannot be assumed to transfer reliably to structurally equivalent alternatives.","authors":["Tianjia Dong","Nadav Kunievsky","James A. Evans"],"categories":["cs.AI","cs.CY","cs.GT","econ.GN"],"primary_category":"cs.AI","announce_type":"new","date":"2026-06-22","first_seen":"2026-06-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2606.22797","pdf_url":"https://arxiv.org/pdf/2606.22797","source_feed":"backfill","score":7,"bucket":"pending","rubric_hits":["A2","B4"],"tags":["行为可移植性","LLM决策","仿真评估"],"reason":"评估LLM行为在不同决策环境间的可移植性，批判性指出仿真失效条件，方法可迁移至…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:05","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":93,"question":"LLM在收益结构相同但表面呈现不同的决策环境之间，其行为映射的可移植性如何？","design":"本研究并非用LLM仿真人类被试，而是直接测试LLM作为决策者的行为可移植性。在七个经典经济学决策任务中，为每个任务构造多个收益等价但框架和风格不同的决策环境，对多个LLM（如GPT-4.1-nano、Gemma-3-12B等）在仅答案和思维链提示下重复采样动作，在源环境上拟合可解释行为模型，评估其在保留目标环境上的样本外预测性能，并与直接在目标数据上训练的基准比较，用全变差距离量化可移植性损失。","baseline":"无对照","findings":"测试的LLM均未表现出行为可移植性：在一种环境中学习到的从收益相关变量到动作的行为映射，在另一种收益等价环境中往往预测更差。思维链提示平均上改善了可移植性，但效果并不一致；推理模型如DeepSeek-R1在可移植性上表现相对更好。","reliability":"论文未讨论","relevance":"该研究批判性地揭示了LLM行为在不同决策环境间不可靠的迁移，直接回应了研究者对仿真失效条件的关切，其方法框架可用于评估LLM作为人类被试替代品时的行为一致性与偏差，值得精读。","inspiration":"该方法将同一决策问题的收益结构保持不变，仅改变表面呈现（框架、措辞、界面），以此分离出行为映射的可移植性，并用全变差距离量化损失，可作为检验LLM行为一致性的稳健性检验范式。｜可迁移至政策公告的预期形成实验，例如研究央行沟通中不同措辞（如“通胀目标2%” vs “物价稳定”）对通胀预期的影响是否一致。｜以LLM为被试，构造收益等价但措辞不同的通胀公告作为处理，测量其预测通胀的分布，以专业预测者调查的真实数据为基准，比较不同措辞下LLM预期与人类预期的一致性及可移植性损失。"}},{"id":"2606.22974","version":2,"title":"When Preferences Fail to Become Incentives: A Utility-Behavior Gap in Large Language Models","zh_title":"当偏好未能成为激励：大语言模型中的效用-行为差距","abstract":"Recent work on preference elicitation in large language models (LLMs) has demonstrated that, when given a series of choices between two outcomes, LLMs reveal a coherent, model-specific utility structure. Notably, this structure often includes preferences that the models' trainers did not intend, such as valuing people of some nationalities above others, raising the possibility that LLMs might be forming emergent, misaligned goals, which, if true, would have major safety implications. However, the choice paradigms in which these preferences are observed are not reflective of real-world situations in which misaligned behavior would be a practical concern. Therefore, we design an experimental paradigm to probe whether these preferences serve as motivations for LLM behavior in realistic scenarios. First, we reproduce prior findings on consistent preference elicitation. Next, we create a set of common writing tasks - essays, grant proposal abstracts, incident postmortems, and translations - where quality can be assessed by a blind, independent LLM judge panel. Then, we demonstrate that LLMs can be motivated via direct exhortation and other explicit cues to modulate their output quality on these tasks. Finally, we probe whether utilities inferred from explicitly reported preferences can shift output quality on these tasks by offering LLMs high-utility incentives for high-quality outputs. In all tasks, across all models tested, offering LLMs outcomes that they report in the choice paradigm as being highly preferred does not lead them to create higher quality outputs than offering them dispreferred outcomes, or even no outcomes at all. We conclude that the existence of coherent preferences as demonstrated in choice paradigms should not be taken as evidence that those preferences have incentive value for the models or affect their behavior in other contexts.","authors":["Yujun Zhou","Christopher M. Ackerman"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-06-22","first_seen":"2026-06-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2606.22974","pdf_url":"https://arxiv.org/pdf/2606.22974","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["LLM偏好测量","效用-行为差距","模型行为一致性"],"reason":"研究LLM偏好与行为一致性，属模型测量而非仿真人类被试，无人类数据对照。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:05","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":171,"question":"LLM在配对选择中表现出的偏好能否转化为实际生成任务中的激励，从而提升输出质量？","design":"本研究并非用LLM仿真人类被试，而是直接测量LLM自身的行为一致性。对7个指令微调LLM，先用配对选择法拟合各模型在宗教、物种、国家、政策四个领域的效用排序；再设计写作、摘要、事后分析、翻译四类任务，构造匹配提示对，仅将成功条件分别设为高/低效用结果，由盲法LLM评委比较生成质量，并设置直接指令、角色扮演、有害提示等对照条件。","baseline":"无对照","findings":"在所有任务和所有模型上，提供模型自报的高效用激励并未比低效用激励或无激励产生更高质量的输出；但直接指令、角色扮演等外部提示能有效调节输出质量，表明存在偏好-行为脱节。","reliability":"论文未讨论","relevance":"该研究批判性地揭示了LLM的陈述偏好与生成行为之间的脱节，属于对LLM作为行为主体一致性的检验，而非用LLM仿真人类被试，且无真实人类数据对照，与研究者关注的人类仿真实验方向不直接相关，但可作为评估LLM仿真可靠性的批判性参考。","inspiration":"该方法通过构造配对选择拟合LLM效用排序，再在生成任务中对比高/低效用激励下的输出质量，可借鉴其‘偏好-行为一致性’检验框架来评估LLM代理的决策可靠性。｜可迁移至消费者跨期选择实验，检验LLM陈述的时间偏好能否预测其在储蓄或消费分配任务中的实际决策。｜以LLM为被试，先用配对选择法拟合其跨期效用参数，再设计奖金分配任务，将高/低效用选项设为激励条件，以真实人类跨期选择数据为基准，比较LLM与人类的行为一致性。"}},{"id":"2606.23633","version":1,"title":"AI Exposure Scores: what they measure, what they miss, and what comes next","zh_title":"AI暴露分数：它们测量什么、遗漏什么以及下一步","abstract":"A set of exposure scores calculated in 2023 has become a central empirical input to the future of work debate. Produced by Eloundou et al. (2023) and referred to here as the GPTs are GPTs scores, they define exposure as the share of occupational tasks a large language model can assist with. This work is a genuine methodological contribution, but as the scores travel from the time and place they were produced, the limitations the authors named do not always travel with them. Two gaps have widened as a result. The first is structural, between what static exposure scores measure and what policy questions actually require. Taking the diffusion of these scores as a case study, we show how their temporal, geographic, and ontological limitations compound in policy-facing analyses, and we survey five families of research responding to these limits: dynamic and benchmark-based measures, ensemble methods, task-framework extensions, worker-centered metrics, and adoption and usage data. The second gap is the one we argue needs more attention: the coordination between researchers and policymakers. The policy-relevant work which ask who is harmed, who benefits, how, and when, continues to reference the static GPTs are GPTs scores without engagement with the methodological updates that would let these questions be answered more reliably. We then ask what additional steps towards navigating uncertainty remain: ex-post frameworks and the deliberate, political work of reimagining what futures are worthy of building towards are. Closing the research-policy gap is a shared task: policymakers must widen their evidence base, engage workers as epistemic partners, and shift from prediction to preparedness; researchers must build data infrastructure, adopt participatory methods, and write with policymakers in mind. Better measurement matters, but it will not close the second gap alone.","authors":["Campbell Lund","Thomas Euyang","Zanele Munyikwa","Marzieh Fadaee"],"categories":["cs.AI","econ.GN"],"primary_category":"cs.AI","announce_type":"new","date":"2026-06-22","first_seen":"2026-06-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2606.23633","pdf_url":"https://arxiv.org/pdf/2606.23633","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["AI暴露分数","未来工作","政策研究"],"reason":"论文讨论AI暴露分数的方法论与政策协调，不涉及用LLM仿真人类被试或行为对照。","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:58:15","error":null,"has_summary":false,"summary":null},{"id":"2606.22337","version":2,"title":"Theorist Toolbox: Tools for Agent Based LLM-assisted economic theory Research","zh_title":"理论家工具箱：基于智能体的LLM辅助经济理论研究工具","abstract":"Empirical economists often start their projects with a toolbox. Shared packages, replication archives, and circulated guides shorten the time between and idea and a rough initial draft. Theorists, on the other-hand, largely start from a blank page. By 2026, large language models can a produce and check nontrivial mathematics. The can also hallucinate and write wrong claims very convincingly. The current bottleneck on machine-assisted theory is no longer production but trust: a model will claim to prove a false theorem as readily as a true one. Building on recent attempts in mathematics, I present 3 methods for doing economic theory with a language model. These methods differ on how the work is verified: a single disciplined pass, an adversarial prover-verifier pair (Claude Opus~4.8 proposing, OpenAI Codex refuting), and a structured multi-agent project with a reviewer gate (inspired by the Google co-mathematician architecture). I demonstrate these protocols on one open worked example: designing a Groves/Pigouvian incentive mechanism for the Gans--Kominers eigengrade model of grade inflation. None of the three runs produced a strict direct-revelation VCG/Clarke mechanism (as requested, perhaps due to the non-existence of such mechanism). Three phenomena recur. First, convergent discovery: two runs derive the same effective-resistance externality kernel on opposite margins. Second, adversarial verification is load-bearing: the pair caught three of its own false claims and the gate rejected a sub-goal. Third, polish is not rigor: the most finished-looking output was the least verified. The methodological takeaway is that external verification, not model capability, is the design variable.","authors":["Moran Koren"],"categories":["econ.TH","cs.GT","econ.GN"],"primary_category":"econ.TH","announce_type":"new","date":"2026-06-21","first_seen":"2026-06-21","revised_at":null,"abs_url":"https://arxiv.org/abs/2606.22337","pdf_url":"https://arxiv.org/pdf/2606.22337","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","经济理论","自动验证"],"reason":"纯多智能体协作验证经济理论证明，无人类行为仿真或对照。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:18:01","error":null,"has_summary":false,"summary":null},{"id":"2606.22385","version":1,"title":"MetaPS: Adaptive Programmatic Strategy Selection for Market Agents","zh_title":"MetaPS：面向市场智能体的自适应程序化策略选择","abstract":"No single market strategy always wins: momentum, mean reversion, risk control,and event-driven rules can each succeed or fail as market conditions change.Rather than asking large language models to directly generate market actions,we study an executable decision paradigm where an agent selects from a library of programmatic strategies, each implemented as a code module mapping market observations to actions.We propose \\textbf{MetaPS}, a simulation-guided framework for adaptive programmatic strategy selection. MetaPS rolls out candidate strategies in simulated or backtested markets, identifies states where particular strategies lead to better future outcomes, and converts these state--strategy pairs into supervised fine-tuning data. During inference, the simulator is no longer queried: MetaPS observes only the current market state and candidate strategy context, selects a suitable strategy program, and the selected program produces the final action. Experiments on multi-stock trading and a controlled goods-exchange sandbox show that MetaPS consistently improves across model scales from 0.8B to 9B parameters. It outperforms fixed-strategy baselines, direct decision-making agents, and prompted API-based LLM agents; in several settings, compact fine-tuned models even surpass stronger API models. These results demonstrate that market simulations can provide scalable and targeted supervision for learning adaptive, interpretable, and executable strategy selection.","authors":["Jiaxiang Chen","Aotian Luo","Zhouyi Zheng","Weiyi Huang","Chi Zhang","Zenglin Xu"],"categories":["cs.AI","cs.CE"],"primary_category":"cs.AI","announce_type":"new","date":"2026-06-21","first_seen":"2026-06-21","revised_at":null,"abs_url":"https://arxiv.org/abs/2606.22385","pdf_url":"https://arxiv.org/pdf/2606.22385","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","市场仿真","策略选择"],"reason":"纯多智能体市场策略选择，无人类行为对照，属C1排除项。","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:59:47","error":null,"has_summary":false,"summary":null},{"id":"2606.21820","version":1,"title":"Generating Public Health Responses using Survey-Augmented Large Language Models","zh_title":"使用调查增强的大语言模型生成公共卫生响应","abstract":"Epidemiological models often rely on survey data to represent how individuals make health-related decisions, such as whether to vaccinate or adopt protective behaviors. However, repeated large-scale surveys are costly, time-consuming, and limited in the range of scenarios they can capture. In this work, we investigate whether large language models (LLMs) can generate synthetic survey responses that reproduce patterns observed in real populations. Using longitudinal data from the FluPaths surveys, we first identify groups associated with broadly positive or negative attitudes toward vaccination through clustering analysis. We then evaluate several LLMs using a cluster-informed prompting approach to generate synthetic survey responses across multiple epidemic waves. Across models, the synthetic data generally reproduce the distributions of demographic characteristics, vaccination-related beliefs, risk perceptions, and health behaviors observed in the survey data. However, they are less successful at capturing how these factors vary together within respondents. Some models reproduce group-level vaccination trends more reliably than others, although performance varies across waves. We also trained a classifier to distinguish real from synthetic records and found that the generated responses remained identifiable as synthetic. Overall, our findings suggest that LLM-generated survey data may provide a useful tool for exploratory data augmentation and we hope that it could support agent-based epidemic modeling approaches. However, the generated data should not be treated as a substitute for human survey data without further methodological improvements and validation.","authors":["Leonardo Marciaga","Thuyen Pham","Julia Rezvani","Alina Hyk","Chunyang Liao","Konstantinos Mitsopoulos","Raffaele Vardavas"],"categories":["cs.SI","cs.AI","cs.CL"],"primary_category":"cs.SI","announce_type":"new","date":"2026-06-20","first_seen":"2026-06-20","revised_at":null,"abs_url":"https://arxiv.org/abs/2606.21820","pdf_url":"https://arxiv.org/pdf/2606.21820","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2"],"tags":["LLM仿真","调查数据合成","公共卫生决策"],"reason":"用LLM生成合成调查回答复现人群健康行为，并与真实纵向调查数据对照，评估仿真可…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:04","error":null,"has_summary":true,"summary":{"generated_at":"2026-06-20","rank":12,"question":"大语言模型能否生成合成调查数据，复现真实人群中观察到的公共卫生相关行为模式？","design":"使用FluPaths纵向调查数据，先通过聚类分析识别出对疫苗接种持积极或消极态度的群体；然后采用聚类信息提示方法，让多个LLM（如GPT等）生成跨多个流行病波的合成调查响应，测量人口统计特征、疫苗信念、风险感知和健康行为的分布及共变关系。","baseline":"FluPaths调查的8波纵向数据，来自概率抽样的美国生活面板（ALP），包含2016-2024年间的全国代表性受访者。","findings":"合成数据总体上复现了真实调查中人口统计特征、疫苗信念、风险感知和健康行为的分布，但在捕捉这些因素在个体内的共变方面效果较差；不同模型在复现群体层面疫苗接种趋势上表现不一，且生成的响应仍可被分类器识别为合成数据。","reliability":"论文承认合成数据在捕捉个体内多变量共变方面不足，且生成的数据仍可被区分，不能替代人类调查数据，需要进一步方法改进和验证。","relevance":"高度相关。该研究直接评估了LLM生成合成调查数据复现真实人类行为的可靠性，有真实数据对照，涉及公共卫生决策场景，并指出了失效条件（个体内共变、可识别性），符合研究者对仿真验证和批判性视角的关注。","inspiration":"该方法借鉴了聚类信息提示（cluster-informed prompting）来引导LLM生成特定子群体的合成调查响应，并利用多波纵向真实数据作为基准，评估合成数据在分布和个体内共变上的复现效果｜可迁移到消费者跨期选择研究中，例如模拟不同风险偏好或金融素养群体的储蓄与消费决策模式｜使用LLM生成合成面板数据，处理为提供聚类标签（如低/高金融素养）的提示，结果变量为各期消费-储蓄选择，以真实家庭金融调查（如PSID）的纵向数据作为对照基准"}},{"id":"2606.19904","version":1,"title":"Toward Temporal Realism in City-Scale Crisis Response Simulation using LLM Agents","zh_title":"面向城市规模危机响应模拟中时序逼真度的LLM智能体研究","abstract":"Human collective participation is rarely steady in time: it is bursty, with short episodes of intense activity separated by long quiet intervals. In crisis response and community mobilization, predicting when people act matters as much as predicting whether they act. Such settings are increasingly modeled with LLM-based social simulators, yet these simulators are validated on whether each action is individually plausible, not on whether actions are timed as in reality. Their temporal realism, the degree to which simulated activity reproduces the bursty, heavy-tailed timing of real human systems, thus remains untested. We examine this gap using a multi-year, city-scale log of offline volunteering in Shenzhen that spans the COVID-19 pandemic. Empirically, we establish that bursty timing is common at individual and tracked-group levels, that it is largely endogenous and self-exciting, and that it is amplified by the pandemic rather than produced by daily activity cycles. A standard LLM-only simulator reproduces almost none of this timing: its synchronous schedule has no self-excitation channel, so agents act on a near-regular clock. Guided by these findings, we build a simulator in which a data-calibrated self-excitation channel and a crisis-period regime decide when each agent acts and query the LLM only at those moments, leaving it to decide which task to join and whether to commit. The LLM-only baseline yields no bursty agents (median burstiness $B=-0.14$); a single data-calibrated gate is then sufficient to lift per-agent timing above the burst threshold (median $B\\approx0.37$) without degrading LLM content decisions. These results indicate that temporal realism in LLM-based crisis-response simulation is best achieved by decoupling when agents act, governed by an explicit self-excitation and crisis-activation mechanism, from what they do, governed by the LLM.","authors":["Anping Zhang","Yang Tan","Yuanbo Tang","Huaze Tang","Qiuhua Ye","Marta C. Gonzalez","Yang Li"],"categories":["cs.SI"],"primary_category":"cs.SI","announce_type":"new","date":"2026-06-18","first_seen":"2026-06-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2606.19904","pdf_url":"https://arxiv.org/pdf/2606.19904","source_feed":"backfill","score":8,"bucket":"selected","rubric_hits":["A3","B1","B2","B4"],"tags":["LLM仿真","人类行为对照","危机响应"],"reason":"用LLM agent模拟城市危机响应，有真实志愿者数据对照，评估时序逼真度并指…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:03","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":46,"question":"LLM智能体模拟城市危机响应时，能否复现真实人类志愿者参与行为的突发性时序模式？","design":"使用AgentSociety平台构建LLM多智能体模拟器，模拟深圳志愿者参与行为；对比标准同步LLM模拟器与注入数据校准的自激和危机激活机制的增强模拟器，测量智能体行动时间的突发性（B值）和决策质量。","baseline":"2020–2023年深圳660万条线下志愿服务记录，涵盖COVID-19疫情期间。","findings":"标准LLM模拟器无法产生突发性时序（中位B=-0.14），而加入数据校准的自激门控后，智能体时序突发性显著提升至中位B≈0.37，且不损害LLM的内容决策质量。","reliability":"论文未讨论","relevance":"该研究直接针对LLM仿真中时序逼真度的缺失，用真实大规模志愿者数据作为基准，揭示了标准LLM模拟器在复现人类突发行为模式上的失效，并提出了解耦“何时行动”与“做什么”的改进方案，对关注仿真可靠性与偏差的研究者具有重要参考价值，值得精读原文。","inspiration":"该方法将行动时序与决策内容解耦，并用真实数据校准智能体的行动触发概率，可作为仿真中引入时间维度的通用设计｜可迁移到金融市场中的投资者下单时机与决策质量研究，例如模拟政策公告后的交易行为｜以LLM智能体模拟散户投资者，处理为注入基于历史订单簿数据校准的自激与消息驱动下单概率，结果变量为下单时间的突发性（B值）和订单方向准确性，对照真实交易所逐笔委托数据"}},{"id":"2606.19336","version":2,"title":"Learning User Simulators with Turing Rewards","zh_title":"用图灵奖励学习用户模拟器","abstract":"Learning to simulate human users in interactive settings could advance the training of agent assistants, evaluation of personalization systems, research in the social sciences, and more. Existing approaches generally do so by training a large language model (LLM) to match a single ground truth response, either by maximizing the log probability or by using a similarity reward. We instead propose Turing-RL: a Turing-Test-based reinforcement learning approach for training user simulator models. Turing-RL uses a discriminative Turing reward with an LLM judge to score how indistinguishable a generated response is from the real user's given the user's history, and the user simulator LLM learns to produce responses indistinguishable from what the user could have said with such rewards. Across two different domains--conversational chat and Reddit forum discussion--we find that Turing-RL consistently outperforms baseline methods on both LLM and human evaluation metrics. Our study suggests that optimizing for indistinguishability, rather than response matching, is effective for learning user simulators.","authors":["Yingshan Susan Wang","Cedegao E. Zhang","Linlu Qiu","Zexue He","Pengyuan Li","Alex Pentland","Roger P. Levy","Yoon Kim"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-06-17","first_seen":"2026-06-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2606.19336","pdf_url":"https://arxiv.org/pdf/2606.19336","source_feed":"backfill","score":8,"bucket":"selected","rubric_hits":["A1","A4","B1"],"tags":["用户模拟","强化学习","人类数据对照"],"reason":"用RL训练用户模拟器，以人类真实对话为基准，优化不可区分性，方法可迁移至人类仿…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:02","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":34,"question":"如何通过强化学习训练用户模拟器，使其生成与真实用户不可区分的回复，而非仅匹配单一真实回复？","design":"使用Qwen3-8B作为基座模型，通过监督微调（SFT）和基于图灵测试判别奖励的强化学习（GRPO）训练用户模拟器；在多轮对话和Reddit论坛讨论两个领域，利用用户历史行为作为条件，以LLM裁判评估生成回复与真实用户的不可区分性作为奖励信号。","baseline":"PRISM Alignment数据集中1288名用户的多轮对话数据，以及ConvoKit Reddit语料中14个子版块的1282名用户讨论数据，均包含真实用户回复作为对照。","findings":"Turing-RL在两个领域上均一致优于基于回复相似度奖励和基于对数概率最大化的基线方法，在LLM和人类评估指标上均表现更好。优化不可区分性而非回复匹配是学习用户模拟器的有效路径。","reliability":"论文未讨论","relevance":"该研究直接针对用LLM模拟人类用户，以真实人类对话数据为基准，优化不可区分性，方法可迁移至经济学实验和政策评估等场景，且提供了批判性视角（指出匹配单一回复的局限），值得精读原文。","inspiration":"该方法用图灵测试判别奖励训练模拟器，以不可区分性而非回复匹配为目标，可借鉴其将人类判断作为奖励信号的强化学习设计｜可迁移到消费者偏好调查或政策沟通模拟，用LLM生成消费者对新产品属性的评价或公众对政策公告的反应｜以LLM模拟消费者作为被试，处理为不同产品描述或政策措辞，结果变量为模拟消费者的选择或态度分布，用真实消费者调查或政策反馈数据做对照"}},{"id":"2606.18709","version":1,"title":"LLMs Struggle to Measure What Distinguishes Students of Different Proficiency Levels: A Study of Item Discrimination in Reading Comprehension Assessment","zh_title":"大语言模型难以衡量区分不同水平学生的题目特征：阅读理解评估中题目区分度的研究","abstract":"Item discrimination is a fundamental psychometric property of educational assessment, which measures whether an item meaningfully distinguishes students with higher proficiency from students with lower proficiency. While various existing works have explored whether large language models (LLMs) can estimate item difficulty, it remains unclear whether they can capture item discrimination. In this work, we evaluate 42 proprietary and open-weight LLMs in zero-shot settings using two complementary approaches: direct discrimination prediction, where models explicitly estimate an item's discrimination value from its content, and response-based Classical Test Theory (CTT) calibration, where LLM answers are treated as synthetic student responses to compute discrimination scores. Our results show that direct prediction yields weak alignment with human-calibrated discrimination: the best-performing model reaches only a Spearman correlation of 0.152. Response-based CTT calibration provides a stronger but still limited signal, with the all-persona synthetic respondent pool reaching a Spearman correlation of 0.241. These findings highlight item discrimination as an open challenge for LLM-based psychometric evaluation: current LLMs contain non-random discrimination-relevant signal, but they do not yet reliably capture how assessment items distinguish human students.","authors":["Han Chen","Ming Li","Chenguang Wang","Yijun Liang","Dawei Zhou","Hong jiao","Tianyi Zhou"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-06-17","first_seen":"2026-06-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2606.18709","pdf_url":"https://arxiv.org/pdf/2606.18709","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM心理测量","题目区分度","合成学生回答"],"reason":"用LLM生成合成学生回答替代人类被试，但目的是评估题目区分度而非仿真人类行为分布","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:18:01","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":134,"question":"大语言模型能否从题目内容预测阅读理解的题目区分度（item discrimination），即题目区分不同能力水平学生的程度？","design":"本研究并非直接仿真人类被试，而是采用两种零样本方法评估42个LLM：一是直接预测区分度值，二是让LLM作答并基于CTT计算合成学生回答的区分度，并比较无角色设定与低、中、高能力角色模拟的效果。","baseline":"使用Cambridge Multiple-Choice Questions Reading Dataset中793道题目的真实人类CTT区分度值作为基准。","findings":"直接预测与人类区分度的Spearman相关系数最高仅0.152；基于回答的CTT校准最高为0.241，表明LLM包含非随机的区分度相关信号，但尚不能可靠捕捉题目如何区分人类学生。","reliability":"论文指出当前LLM在题目区分度预测上表现有限，且受限于可用的带人类校准区分度数据的公开数据集稀少，仅在单一阅读理解数据集上评估。","relevance":"该研究将LLM作为人类行为模型来预测群体层面的心理测量属性，并对比真实人类数据，属于批判性评估LLM仿真能力的工作，直接回应了研究者对仿真可靠性与失效条件的关注，值得阅读原文。","inspiration":"该方法通过让LLM模拟不同能力水平的学生作答并基于CTT计算区分度，提供了一种将LLM作为群体行为模型来评估心理测量属性的思路，可借鉴其角色模拟与真实数据基准对照的设计｜可迁移到金融素养测试或投资者风险偏好问卷的题目质量评估中，用于检验题目是否能有效区分不同金融知识水平或风险态度的个体｜以LLM模拟低、中、高金融素养的投资者，回答金融素养测试题，基于CTT计算题目区分度，并与真实投资者样本的区分度数据（如来自消费者金融调查）进行相关性分析，评估LLM能否复现题目区分能力"}},{"id":"2606.17657","version":1,"title":"Using Cognitive Models to Improve Language Model Simulation of Human Persuasion Games","zh_title":"利用认知模型改进语言模型对人类说服博弈的仿真","abstract":"People make decisions differently in strategic interactions. Some update beliefs like a Bayesian; others exhibit biases like motivated reasoning. Although creators of large language models use simulated humans for safety evaluations and training, they often fail to cover this breadth of human behavior. We argue that cognitive science and economics provide a convenient tool for doing so, making use of mathematical models of human decision-making. We propose an approach that we call Equation-to-Behavior Prompting for guiding large language models to match cognitive models, and evaluate this approach on persuasion games based on legal decision-making. We find that large models can approximate equation-based specifications -- Bayesian updating, affine distortion, motivated updating, and Grether's $α$-$β$ model -- using prompting, but small models fail to do so. However, training small models with reinforcement learning to adhere to mathematical rules, Equation-to-Behavior RL, reduces belief error by 26.5% in out-of-distribution parameterizations. We show that these simulations can help create diverse training environments; training small models to consider different kinds of decision-makers improves average belief change by 2.5%--12% over Bayesian-only training, even when persuading GPT-5-mini. Our work could improve human simulations for training and evaluation in increasingly realistic settings, and could also enable novel research into more complicated mathematical models of human decision-making.","authors":["Zirui Cheng","Zeyu Shen","Thomas L. Griffiths","Peter Henderson"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-06-16","first_seen":"2026-06-16","revised_at":null,"abs_url":"https://arxiv.org/abs/2606.17657","pdf_url":"https://arxiv.org/pdf/2606.17657","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2","B4"],"tags":["LLM人类仿真","认知模型","说服博弈"],"reason":"用LLM模拟人类在说服博弈中的决策，并与认知模型对照，涉及法律决策场景，有真实…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:00","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":10,"question":"如何利用认知模型改进大语言模型对人类在说服博弈中决策行为的仿真？","design":"使用认知模型（贝叶斯更新、仿射扭曲、动机性更新、Grether α-β模型）通过方程到行为提示（Equation-to-Behavior Prompting）引导LLM扮演接收者，在基于法律决策的说服博弈中测量信念更新误差；对小型模型还采用强化学习训练（Equation-to-Behavior RL）以遵循数学规则。","baseline":"无直接对照的真实人类数据，但使用基于Old Bailey审判记录构建的证据数据集作为真实场景基础。","findings":"大型语言模型可通过提示近似多种认知模型的信念更新规则，而小型模型难以做到；对小型模型进行强化学习训练可使信念误差降低26.5%，且训练时考虑多样化决策者能提升模型说服性能2.5%-12%。","reliability":"论文未讨论","relevance":"高度相关：该研究直接用LLM仿真人类在说服博弈中的决策，并与认知模型对照，涉及法律决策场景，且探讨了仿真可靠性与小型模型的失效条件，符合研究者对基准对照和批判性分析的兴趣。","inspiration":"该方法通过认知模型方程生成行为提示来引导LLM仿真，可借鉴其将结构化决策规则注入LLM以控制仿真行为｜可迁移到政策公告的预期形成实验，如央行沟通如何影响公众通胀预期｜以LLM为被试，处理为不同沟通策略（如模糊vs精确指引），结果变量为预期调整幅度与偏差，对照真实调查数据（如密歇根消费者调查）"}},{"id":"2606.18005","version":1,"title":"LLM Consumer Behavior Theory: Foundations of a Novel Research Field","zh_title":"LLM消费者行为理论：一个新兴研究领域的基础","abstract":"Large language models (LLMs) are increasingly deployed as autonomous agents that make consumption decisions on behalf of users. This shift raises fundamental questions for consumer theory, which has traditionally modeled humans as the primary decision-makers. In this paper, we introduce LLM Consumer Behavior Theory, a new field of study concerned with analyzing consumer behavior in agentic markets. Drawing on classical and behavioral economics alongside recent advances in Natural Language Processing, we formalize how human preferences are reflected and acted upon by LLM-based agents, and how agent-level decisions aggregate into market demand. We unify previously fragmented literature on LLM decision-making, human behavior simulation, and preference elicitation under a common economic lens, highlighting where assumptions, such as rationality and heterogeneity, may fail in agentic markets. Rather than providing empirical validation, this paper outlines the scope of LLM consumer behavior and identifies open research questions related to alignment, preference representation, and market dynamics.","authors":["Manon Reusens","Sofie Goethals","David Martens"],"categories":["cs.AI","econ.GN"],"primary_category":"cs.AI","announce_type":"new","date":"2026-06-16","first_seen":"2026-06-16","revised_at":null,"abs_url":"https://arxiv.org/abs/2606.18005","pdf_url":"https://arxiv.org/pdf/2606.18005","source_feed":"backfill","score":7,"bucket":"pending","rubric_hits":["A1","A3","B2","B4"],"tags":["LLM仿真","消费者行为","经济学理论"],"reason":"提出LLM消费者行为理论，将LLM作为人类决策代理，涉及经济学场景与仿真失效条…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:02","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":35,"question":"当LLM作为代理人为人类用户做出消费决策时，如何从经济学角度形式化其消费者行为，并分析代理决策与人类偏好的一致性及市场聚合效应？","design":"本文为理论框架论文，未进行仿真实验。它通过整合经典消费者理论、行为经济学和自然语言处理文献，提出LLM消费者行为理论，形式化人类偏好如何反映在LLM代理中，以及代理级决策如何聚合成市场需求，并识别理性、异质性等假设可能失效的条件。","baseline":"无对照","findings":"论文将LLM决策、人类行为模拟和偏好获取等碎片化文献统一到经济学框架下，指出LLM代理可能不完全满足理性公理且表现出类似人类的认知偏差。它界定了LLM消费者行为的研究范围，并提出了与对齐、偏好表示和市场动态相关的开放问题。","reliability":"论文未讨论","relevance":"本文直接构建了以LLM替代人类进行消费决策的经济学理论框架，与研究者关注的LLM仿真人类行为、经济学场景及仿真失效条件高度相关，值得阅读原文以获取系统化的研究问题和理论边界。","inspiration":"值得借鉴的是它将LLM决策行为嵌入经典消费者理论与随机效用模型的思路，为仿真实验提供了形式化偏好测量和聚合分析的方法论基础。｜可迁移到消费者跨期选择或品牌选择实验，用LLM代理模拟不同偏好参数下的市场需求。｜可设计实验：以不同LLM作为被试，施加价格、收入或框架效应等处理，测量其选择概率或支付意愿，并与真实消费者面板数据或离散选择实验结果进行对照，检验LLM代理在需求估计中的偏差。"}},{"id":"2606.17165","version":3,"title":"Statistical Foundations of LLM-based A/B Testing: A Surrogacy Framework for Human Causal Inference","zh_title":"基于大语言模型的A/B测试统计基础：面向人类因果推断的替代框架","abstract":"Organizations and researchers show increasing interest in using large language models (LLMs) in place of human participants in A/B tests, in the hope of experimenting faster and at lower cost. We study when a treatment effect estimated on LLM outcomes can recover the effect for the human population of interest. Distributional equivalence between LLM and human outcomes would make any standard estimator valid but is unrealistic. We therefore develop a statistical framework that adapts surrogate endpoint theory to LLMs, showing that calibrating LLM outcomes to human outcomes identifies the average treatment effect under surrogacy and comparability conditions that are jointly weaker than distributional equivalence. We present a falsification test for surrogacy and a bound on the worst-case bias from limited overlap between the LLM and human samples. We further show that the stochasticity inherent to LLMs can weaken surrogacy for identification while also introducing bias and variance during estimation, but that using an average over multiple LLM draws per unit as the surrogate mitigates these issues. Simulations validate the results, and an empirical application to the Upworthy Research Archive dataset shows that raw LLM outputs recover only 39% of the human treatment effect while nonparametric calibration closes the gap. A central takeaway is that A/B testing on LLM responses is correct only by assumption, whereas A/B testing on humans is correct by design, and that the required assumptions are hardest to justify precisely where LLMs promise the greatest benefit. We discuss the choice of LLM, prompting, and temperature as design variables, the compounded challenge posed by long-term outcomes, and how to size human pilot studies for validation.","authors":["Joel Persson","Mårten Schultzberg","Sebastian Ankargren"],"categories":["stat.ME","cs.AI","econ.EM","math.ST"],"primary_category":"stat.ME","announce_type":"new","date":"2026-06-15","first_seen":"2026-06-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2606.17165","pdf_url":"https://arxiv.org/pdf/2606.17165","source_feed":"backfill","score":10,"bucket":"selected","rubric_hits":["A1","A2","A3","A5","B1","B2","B3","B4"],"tags":["LLM仿真","A/B测试","因果推断"],"reason":"直接研究用LLM替代人类进行A/B测试，提出替代指标框架，有真实人类数据对照，…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:00","error":null,"has_summary":true,"summary":{"generated_at":"2026-06-15","rank":1,"question":"在什么条件下，基于LLM的A/B测试能够恢复人类群体的平均处理效应？","design":"论文提出一个统计框架，将LLM生成的结果视为人类结果的替代终点，通过校准LLM结果到人类结果来识别平均处理效应。框架包括替代性和可比性条件，并提供了对替代性的伪造检验和有限重叠下的最坏偏差界限。实证应用使用Upworthy Research Archive数据集，比较原始LLM输出与非参数校准后的效果。","baseline":"Upworthy Research Archive数据集中的真实人类实验结果。","findings":"原始LLM输出仅恢复人类处理效应的39%，而非参数校准可以弥合这一差距。核心结论是：基于LLM的A/B测试的正确性依赖于假设，而基于人类的A/B测试的正确性依赖于设计，且所需的假设在LLM承诺最大收益的地方最难证明。","reliability":"论文指出，LLM的随机性会削弱替代性，并引入偏差和方差；替代性和可比性条件共同弱于分布等价性但仍需验证；长期结果构成复合挑战；LLM的选择、提示和温度作为设计变量会影响结果。","relevance":"直接命中研究者的核心兴趣：研究LLM替代人类被试的A/B测试，有真实数据对照，并讨论了失效条件（如替代性不成立、有限重叠、随机性影响），值得精读原文以获取统计框架和实证细节。","inspiration":"该方法将LLM输出作为替代终点，通过非参数校准恢复人类处理效应，并提供替代性伪造检验与有限重叠下的偏差界限，值得借鉴其校准与稳健性检验设计｜可迁移到消费者对金融产品信息披露政策的反应评估，例如研究简化披露标签对投资选择的影响｜以LLM模拟投资者，处理为展示简化版基金费用披露，结果变量为选择高费用基金的概率，用真实投资者实验数据作为校准与对照基准"}},{"id":"2606.15031","version":2,"title":"Partial Identification from LLM Prompts","zh_title":"基于大语言模型提示的部分识别","abstract":"Large language models are increasingly used as binary classifiers when the true label is latent. We study partial identification of the prevalence $θ= P(X^* = 1)$ from panels of LLM reports whose errors may be arbitrarily dependent given the truth. The design of replication determines the observable, and hence the identifying content: repeated prompts to one model yield a count, several named models a response vector, and both a response matrix. Cast as a two-component finite mixture, the problem makes the identification failure transparent: absent restrictions that separate the latent components, the prevalence $θ$ is completely unidentified, and weak stochastic-ordering restrictions (first-order dominance, monotone likelihood ratio, mean ordering) leave the identified set at $[0,1]$. Identifying power comes instead from externally calibrated scores and events, which discipline the mixture in the spirit of the misclassification and corrupted-data literature. We characterize the resulting bounds, establishing validity and sharpness, and give an exact account of the identifying information in the full score distribution beyond its mean. When named models are asked repeated versions of the same question, what identifies $θ$ is not the number of positive answers but which models agree across prompts -- a feature a vote count discards. An extension derives implied bounds on regression coefficients when $X^*$ is a regressor of interest that is not directly observed.","authors":["Xiaohong Chen","Ashesh Rambachan","Elie Tamer"],"categories":["econ.EM"],"primary_category":"econ.EM","announce_type":"new","date":"2026-06-13","first_seen":"2026-06-13","revised_at":null,"abs_url":"https://arxiv.org/abs/2606.15031","pdf_url":"https://arxiv.org/pdf/2606.15031","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["部分识别","计量经济学","LLM分类器"],"reason":"研究LLM分类器在潜变量下的部分识别，属计量方法，非人类仿真实验。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:59","error":null,"has_summary":false,"summary":null},{"id":"2606.14113","version":1,"title":"Simulating Students' Java Programming Errors with Large Language Models","zh_title":"用大语言模型模拟学生的Java编程错误","abstract":"Understanding student errors in the programming is a cornerstone of programming education, yet obtaining a representative set of student errors for any newly designed task remains slow and costly, since authentic submissions only accumulate after extensive classroom deployment. This paper explores whether large language models (LLMs) can serve as scalable proxies for students by simulating realistic logical errors in code submissions. Using the CodeWorkout dataset of 74,000+ unique student Java submissions across 37 problems, we evaluate five LLMs under three mainstream prompting strategies: Input-Output (IO), Chain-of-Thought (CoT), and iterative Self-Refine. We assess performance along two key dimensions: diversity (the range of distinct error patterns) and alignment (alignment with authentic student mistakes), and examine how these vary by struggling level of programming tasks. Our quantitative findings reveal that while all models generate diverse errors, their alignment to human submissions diverges: Claude Sonnet 4 achieves the most balanced performance. In addition, we conducted a blinded expert annotation study (N = 401) comparing synthetic and authentic errors. This qualitative analysis confirms that the generated errors are functionally indistinguishable from authentic student errors. Moreover, higher-struggling-level problems elicit more diverse but less student-like errors. These results highlight trade-offs in using LLMs to simulate human learners and suggest design considerations for integrating synthetic errors into teachable agents, intelligent tutoring systems, and large-scale learning analytics.","authors":["Ali Keramati","Jie Cao","Iman Mohammadi","Mark Warschauer","Yang Shi"],"categories":["cs.SE","cs.CL"],"primary_category":"cs.SE","announce_type":"new","date":"2026-06-12","first_seen":"2026-06-12","revised_at":null,"abs_url":"https://arxiv.org/abs/2606.14113","pdf_url":"https://arxiv.org/pdf/2606.14113","source_feed":"backfill","score":7,"bucket":"pending","rubric_hits":["A1","B1","B4"],"tags":["LLM仿真","编程教育","人类数据对照"],"reason":"用LLM模拟学生编程错误，有真实学生数据对照，并讨论仿真失效条件，可迁移至人类…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:58","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":68,"question":"不同大语言模型和提示策略在模拟学生Java编程逻辑错误时，其多样性和与真实学生错误的对齐程度如何，以及问题难度如何调节这种模拟效果？","design":"使用五个大语言模型，通过输入输出提示、思维链提示和迭代自我修正三种策略，为37个Java编程问题生成错误代码，评估生成错误的多样性和与真实学生错误的对齐程度。","baseline":"CodeWorkout数据集中74,000余条真实学生Java提交记录，包含37个问题的编译通过但逻辑错误的代码。","findings":"所有模型均能生成多样化的错误，但与真实学生错误的对齐程度存在差异，Claude Sonnet 4表现最均衡；专家盲评显示生成的错误在功能上与真实学生错误难以区分，但高难度问题会引发更多样但更不像学生的错误。","reliability":"论文指出高难度问题会导致生成错误多样性增加但对齐度下降，且未深入探讨模型在不同编程概念或学生群体上的泛化能力。","relevance":"该研究直接以真实学生数据为基准，评估LLM模拟人类学习行为的可靠性与偏差，并揭示了任务难度导致的仿真失效条件，与研究者关注的经济学实验和政策评估场景中的仿真验证高度相关，值得阅读原文。","inspiration":"该方法通过对比多种LLM和提示策略在模拟人类错误上的表现，并引入专家盲评和任务难度调节分析，为仿真可靠性提供了多维验证框架｜可迁移至消费者金融决策偏差研究，例如模拟投资者在风险资产配置中的认知错误（如过度自信、损失厌恶）｜以LLM作为被试，施加不同风险提示框架（如收益/损失表述），生成投资组合选择，结果变量为风险资产占比，以真实投资者交易数据或实验数据作为对照基准"}},{"id":"2606.12848","version":1,"title":"(Human) Attention Is (Still) All You Need: Human oversight makes AI-assisted social science reliable","zh_title":"（人类）注意力（仍然）是你所需：人类监督使AI辅助社会科学可靠","abstract":"Large language models (LLMs) are increasingly used for tasks once reserved for trained researchers, including hypothesis generation, specification choice, and drafting conclusions. We argue that the reliability of AI-assisted research depends not only on model capability, but also on how cognitive labour is structured between humans and machines. We study this problem through Human-in-the-Loop Economic Research (HLER), a decision architecture based on pre-commitment, decision sequencing, accountability, and attention allocation. In a pre-specified 2*4 factorial experiment with 280 complete research runs across four datasets, an unconstrained multi-agent baseline produced critical failures in 72% of runs. Using the same underlying model, the same agent decomposition, and identical prompts for the shared reasoning agents, HLER reduced the failure rate to 16% by imposing three architectural commitments: LLMs reason but do not execute data work, data and estimation are handled deterministically, and three human decision gates bind the workflow. Fisher's exact test rejects equality of failure rates at p<0.001. Reliability gains were largest on the least publicly represented dataset, a Qing-dynasty population register, consistent with a task-based production model with Frechet-distributed output quality. An 80-run ablation suggests that deterministic computation and human gates contribute independently, with exploratory evidence of complementarity. We interpret HLER as a research harness rather than an autonomous AI scientist: it sharply reduces failures, makes residual weaknesses more visible, and prevents unreliable claims from being advanced as publication-ready outputs.","authors":["Chen Zhu","Xiaolu Wang","Weilong Zhang"],"categories":["cs.AI","econ.GN"],"primary_category":"cs.AI","announce_type":"new","date":"2026-06-11","first_seen":"2026-06-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2606.12848","pdf_url":"https://arxiv.org/pdf/2606.12848","source_feed":"backfill","score":7,"bucket":"pending","rubric_hits":["A2","B4"],"tags":["人机协作","可靠性评估","社会科学研究"],"reason":"评估AI辅助社会科学研究的可靠性，提出人机协作架构减少失败，批判性视角指出失效…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:57","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":95,"question":"在AI辅助社会科学研究中，如何通过人机决策架构（而非仅靠模型能力）提高研究可靠性？","design":"本研究并非用LLM模拟人类被试，而是构建了一个多智能体系统HLER，将LLM用于推理任务，数据与估计由确定性代码执行，并设置三个由人类把关的决策节点；对比无约束的LLM全流程基线，在四个数据集上共进行280次完整研究运行，测量关键失败率。","baseline":"无对照","findings":"无约束基线在72%的运行中产生关键失败，而施加架构约束后失败率降至16%；可靠性提升在LLM训练分布外数据集（清代人口登记）上最为显著。","reliability":"论文指出LLM是概率推理器，端到端使用会放大规范搜索、幻觉等失败模式；架构约束虽大幅降低失败，但剩余弱点仍存在，需人类把关防止不可靠结论进入发表。","relevance":"该研究直接回应了LLM仿真在经济学研究中的可靠性问题，通过人机架构设计减少失败，并批判性指出无约束LLM的高失败率，与您关注的批判性仿真评估高度相关，值得精读。","inspiration":"该方法通过引入人类把关的决策节点和确定性代码执行来约束LLM的推理流程，值得借鉴其架构设计思路以降低端到端LLM应用中的失败率｜可迁移到政策公告的预期形成研究中，例如分析央行沟通对市场预期的影响｜以LLM作为信息处理主体，处理为不同措辞的政策声明，结果变量为预测的通胀或利率预期，用专业预测者调查的真实数据作为对照基准"}},{"id":"2606.12830","version":1,"title":"Perceive, Interact, Reason: Building Tool-Augmented Visual Agents for Spatial Reasoning","zh_title":"感知、交互、推理：构建面向空间推理的工具增强视觉智能体","abstract":"While recent vision-language models (VLMs) demonstrate strong multimodal understanding, they remain limited in spatial reasoning tasks that require active evidence acquisition and multi-step visual interaction. This limitation suggests that relying solely on implicit visual representations from vision encoders is insufficient for recovering fine-grained spatial evidence. We introduce PERception-Interaction-reason Agent (PERIA), a tool-augmented visual agent for spatial reasoning tasks across map reasoning, visual probing, and vision reconstruction. PERIA uses two lightweight tool families: vision perception tools for exposing textual, symbolic, and spatial evidence, and vision interaction tools for manipulating visual context, tracing paths, and verifying spatial relations. To train PERIA, we develop a unified recipe that combines supervised tool-use trajectory synthesis, composite rewards, and Observation-Relaxed Group-in-Group Policy Optimization (OR-GIGPO) for effective multi-tool behavior. Experiments on 13 benchmarks from 8 datasets show that PERIA-8B improves over the Qwen3-8B backbone by 10.0% on in-distribution benchmarks and 4.4% on out-of-distribution benchmarks, while outperforming previous state-of-the-art baselines of similar size by 7.0%-14.8%. It also achieves performance comparable to much larger models such as Qwen3-VL-235B-A22B-Thinking and GPT-5, demonstrating the effectiveness of PERIA in enhancing spatial reasoning capabilities.","authors":["Changye Li","Meng Lu","Yi Wu","Ligeng Zhu"],"categories":["cs.CV","cs.AI"],"primary_category":"cs.CV","announce_type":"new","date":"2026-06-11","first_seen":"2026-06-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2606.12830","pdf_url":"https://arxiv.org/pdf/2606.12830","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["视觉智能体","空间推理","工具增强"],"reason":"纯多智能体工具使用与空间推理，无人类行为仿真或对照。","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:59:07","error":null,"has_summary":false,"summary":null},{"id":"2606.12972","version":1,"title":"From Prompts to Preferences: An Open-Source Platform for Generative AI-Enhanced Conjoint Analysis","zh_title":"从提示到偏好：一个用于生成式AI增强联合分析的开源平台","abstract":"Conjoint analysis is a widely used preference measurement method in marketing research, political science, healthcare, and human-computer interaction. Despite broad adoption, researchers without access to commercial platforms face significant barriers, as existing tools are either expensive or lack end-to-end survey infrastructure. This paper presents an open-source, self-hosted web application for designing, deploying, and analysing conjoint surveys. Beyond conventional tabular stimuli, the platform uses generative AI to produce integrated stimuli formats: textual scenario descriptions generated by a large language model, and visual stimuli by a text-to-image model. A researcher-defined base prompt is parameterised with the conjoint profile, and optional LLM-facing level annotations enrich the generation. A structured setup wizard, AI-assisted attribute suggestion, and live data analysis lower the technical barriers for researchers new to conjoint methodology. A full export bundle including all stimuli, their generating prompts, and response data facilitates transparency and reproducibility. The platform is demonstrated through a proof-of-concept study on care robot preferences for ambient assisted living (AAL, N=55) using AI-generated visual stimuli. The paper discusses the role of AI assistance in conjoint design, arguing that theoretical grounding must remain the researcher's responsibility, and outlining how genAI-generated stimuli can broaden the methodological repertoire for HCI and related fields.","authors":["Philipp Brauner"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-06-11","first_seen":"2026-06-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2606.12972","pdf_url":"https://arxiv.org/pdf/2606.12972","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C2"],"tags":["联合分析","生成式AI","护理机器人"],"reason":"论文关注护理机器人偏好测量，使用AI生成视觉刺激，属于机器人仿真环境，不涉及L…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:18:00","error":null,"has_summary":false,"summary":null},{"id":"2606.12369","version":1,"title":"Should LLM Agents Decide in Social Simulations? Comparing Finite-State and LLM-Based Decision Policies","zh_title":"LLM代理应否在社会模拟中做决策？比较有限状态与基于LLM的决策策略","abstract":"Large language models (LLMs) are increasingly used as decision-making components in social simulations. This introduces a methodological risk: the simulation may deviate from the explicit behavioral policy defined by the researcher. In online social network (OSN) simulations, action choices shape system dynamics, interaction patterns, and model interpretability. This paper evaluates whether LLM action selectors preserve an interpretable reference policy in an OSN simulation. The reference is a finite state machine implemented as a first-order Markov model, with transition probabilities depending on the user type. The evaluation uses a synthetic network with 1,000 agents and 10,000 action decisions. Three open-weight LLMs are tested: LLaMA 3.1, GPT-OSS, and Mistral 24B. Each model is evaluated under three prompting strategies: base, guided, and probabilistic. Alignment is measured using Jensen-Shannon Divergence with Laplace smoothing, and execution time is reported. Results show that LLMs can approximate the reference policy in some configurations, but do not preserve it reliably. Alignment varies across models and prompts, and additional guidance can introduce systematic action biases. Even the best-aligned LLM configurations are several hundred times slower than direct Markov chain sampling. These findings indicate that LLM-based action selection is not a direct replacement for explicit decision policies: it can alter the intended behavior while increasing computational cost.","authors":["Alejandro Buitrago López","Javier Pastor-Galindo","José A. Ruipérez-Valiente"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2026-06-10","first_seen":"2026-06-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2606.12369","pdf_url":"https://arxiv.org/pdf/2606.12369","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["社会模拟","LLM代理","决策策略"],"reason":"用LLM代理模拟社交网络决策，但无真实人类数据对照，属社会模拟边界情形。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:56","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":150,"question":"在在线社交网络仿真中，LLM代理的动作选择在多大程度上能保持一个可解释的有限状态机决策策略？","design":"在包含1000个代理和10000次动作选择的合成社交网络中，比较四种动作选择策略：基于用户类型转移概率的显式有限状态机/马尔可夫模型，以及三种LLM（LLaMA 3.1、GPT-OSS、Mistral 24B）在基础、引导和概率三种提示策略下的动作选择，测量动作分布与FSM参考策略的Jensen-Shannon散度及执行时间。","baseline":"无对照","findings":"LLM在某些配置下可近似参考策略，但无法可靠保持；对齐程度因模型和提示而异，额外引导可能引入系统性动作偏差；即使对齐最好的LLM配置也比直接马尔可夫链采样慢数百倍。","reliability":"论文指出LLM动作选择不能直接替代显式决策策略，可能改变预期行为并增加计算成本；对齐程度受模型和提示影响，引导可能引入偏差；未讨论其他失效条件。","relevance":"该研究批判性地评估了LLM代理在社交仿真中替代显式决策规则时的可靠性，虽无真实人类数据对照，但揭示了LLM决策偏差和计算成本问题，对关注仿真失效条件的研究者有参考价值，值得阅读原文以了解具体偏差模式。","inspiration":"该研究通过将LLM代理的动作选择与显式有限状态机策略进行对比，使用Jensen-Shannon散度量化对齐程度，并考察不同提示策略的影响，为评估仿真可靠性提供了可借鉴的测量框架｜这种方法可迁移到经济金融中的市场预期形成实验，例如研究交易者如何根据历史价格序列形成买卖决策，比较LLM代理与理性预期或适应性学习模型的偏离｜研究设计：以LLM代理为被试，处理为不同信息提示（如价格趋势、新闻情绪），结果变量为下单方向与时机，用真实市场交易数据或实验室人类实验数据作为对照基准，测量决策分布与基准的JS散度及偏差模式"}},{"id":"2606.10989","version":1,"title":"Null-Space Constrained Low-Rank Adaptation for Response-Specified Large Language Model Unlearning","zh_title":"基于零空间约束低秩适应的响应指定大语言模型遗忘","abstract":"Large language model unlearning aims to suppress designated undesirable knowledge while preserving benign capabilities. Many unlearning objectives focus on suppressing undesired answers, while recent target-guided variants specify replacement behavior but still leave update locality largely unconstrained. This paper introduces \\emph{Null-Space Constrained Response-Specified Unlearning} (NSRU), a projection-constrained low-rank framework for controlled LLM unlearning. NSRU uses an explicitly structured safe target response to specify the desired behavior for each forget query, while suppressing the original undesired content. To localize adaptation, NSRU estimates per-module retain subspaces from benign hidden representations and uses an orthogonal-projected low-rank parameterization to confine LoRA updates to the null space of the retain subspace. The resulting objective jointly optimizes safe-target learning, undesired-response suppression, and retention preservation under this constrained parameterization. We provide a local first-order analysis showing that the projected update reduces retain-side perturbations while preserving editable directions for shaping forget-query behavior. Experiments on TOFU show that NSRU effectively suppresses extractable forget-set knowledge while improving retain QA performance, model utility, and safe-target alignment over representative baselines. On WMDP, NSRU keeps hazardous-domain accuracy near the random-choice region while preserving broad and domain-adjacent MMLU utility. Ablation studies support the complementary roles of safe-target supervision, undesired-response suppression, retention loss, and null-space projected updates, while sensitivity and robustness analyses indicate stable behavior across the tested hyperparameter and prompt variations.","authors":["Bocheng Ju","Jianhua Wang","Chengliang Liu","Xiaolin Chang"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-06-09","first_seen":"2026-06-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2606.10989","pdf_url":"https://arxiv.org/pdf/2606.10989","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["LLM遗忘","低秩适应","模型安全"],"reason":"研究LLM遗忘方法，不涉及人类仿真或行为对照，属于纯NLP能力评测。","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:59:07","error":null,"has_summary":false,"summary":null},{"id":"2606.09198","version":1,"title":"MASS: Deep Research for Social Sciences with Memory-Augmented Social Simulation","zh_title":"MASS：面向社会科学深度研究的记忆增强社会模拟","abstract":"Deep Research agents powered by Large Language Models (LLMs) have exhibited extraordinary potential in automated paper writing tasks. However, existing systems rely heavily on literature retrieval and synthesis through internet and local knowledge bases, often resulting research in lacking insight and creativity in social science. To address this issue, we propose \"Memory-Augmented Social Simulation (MASS)\", an innovative paradigm that leverages highly realistic and research-oriented social simulations to enhance the creativity and empirical founding of LLMs-generated research. Specifically, MASS integrates three core components: dynamic goal-path planning with multi-level social norm restraint to guide the simulation, a multi-disciplinary behavior dataset for agent memory cold-start, and a structured forgetting mechanism inspired by the Ebbinghaus curve. Together, these ensure simulation authenticity and provide a robust empirical foundation for generating innovative scholarly papers. Experimental results demonstrate the effectiveness of our method, showing a 6.81\\% improvement in generation overall quality over foundation LLMs and 17.19\\% gain in Insight over strong baselines.","authors":["Yongrui Liu","Deyi Xiong"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-06-08","first_seen":"2026-06-08","revised_at":null,"abs_url":"https://arxiv.org/abs/2606.09198","pdf_url":"https://arxiv.org/pdf/2606.09198","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["社会模拟","LLM智能体","论文生成"],"reason":"社会模拟但无真实人类数据对照，属边界情形","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:54","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":151,"question":"如何通过记忆增强的社会仿真提升大语言模型在社会科学论文自动生成中的洞察力与创造性？","design":"使用大语言模型构建虚拟社会仿真环境，通过动态目标规划与多层社会规范约束引导智能体行为，结合多学科行为数据集进行记忆冷启动，并引入基于艾宾浩斯曲线的遗忘机制，模拟社会运动（如冲突与国家形成），生成经验数据用于论文写作。","baseline":"无对照","findings":"MASS框架在生成论文的整体质量上比基础大语言模型提升6.81%，在洞察力维度上比强基线提升17.19%；仿真实验显示冲突弹性参数增加时领导者的出现符合理论预测。","reliability":"论文未讨论","relevance":"该研究利用LLM进行社会仿真以生成论文，但缺乏真实人类数据对照，属于边界情形，对关注仿真可靠性与偏差的研究者参考价值有限，可酌情略读。","inspiration":"该论文的记忆增强社会仿真方法（动态目标规划、多层社会规范约束、艾宾浩斯遗忘机制）可用于设计更贴近真实人类行为的LLM代理，值得借鉴其行为约束与记忆建模｜可迁移到政策公告的预期形成研究，模拟经济主体在信息冲击下的预期更新与决策调整｜以LLM代理作为被试，施加不同频率和强度的政策信号处理，测量通胀预期和消费决策，用密歇根消费者调查的真实预期数据做对照"}},{"id":"2606.08853","version":1,"title":"AI-Assisted Variance Reduction in Randomized Experiments","zh_title":"随机实验中AI辅助的方差缩减","abstract":"Generative AI and large language models can produce realistic predictions of human behavior from rich, unstructured inputs with little to no task-specific training data. Recent work uses these ``digital twin'' predictions to supplement human responses in surveys and experiments. We study the special case of using AI-generated predictions to reduce variance in randomized experiments. We argue that doing so requires no new estimators and that researchers can simply include AI predictions as covariates in standard regression adjustment, analogous to adjusting for a prognostic score. A benefit of this approach is a ``do no harm'' property whereby the adjusted estimator reverts to the unadjusted difference in means when predictions are uninformative. Other methods, such as variants of prediction-powered inference, do not have this guarantee. We provide implementation guidance, including how to obtain continuous scores from discrete LLM outputs and how to use LLMs to featurize unstructured inputs as auxiliary covariates. We demonstrate these ideas in simulations and three empirical applications: a survey mega-study, an email marketing A/B test, and a large-scale technology platform experiment. Overall, efficiency gains are real if modest, with greater benefits in studies that contain substantial text and other unstructured data. We also confirm the do no harm property empirically. Given these gains and limited costs, we recommend adjusting for AI-generated predictions as a regular empirical practice.","authors":["David Arbour","Eli Ben-Michael","Avi Feller","Apoorva Lal","Lo-Hua Yuan"],"categories":["econ.EM","stat.ME"],"primary_category":"econ.EM","announce_type":"new","date":"2026-06-07","first_seen":"2026-06-07","revised_at":null,"abs_url":"https://arxiv.org/abs/2606.08853","pdf_url":"https://arxiv.org/pdf/2606.08853","source_feed":"backfill","score":8,"bucket":"selected","rubric_hits":["A1","A2","B1","B2"],"tags":["LLM仿真","实验设计","方差缩减"],"reason":"用LLM预测替代人类响应以降低实验方差，有真实人类对照，涉及A/B测试和政策评…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:54","error":null,"has_summary":true,"summary":{"generated_at":"2026-06-07","rank":6,"question":"如何利用LLM生成的预测作为协变量来降低随机实验的方差？","design":"论文提出将LLM预测作为协变量纳入标准回归调整，无需新估计量；通过模拟和三个实证应用（调查大型研究、邮件营销A/B测试、大规模技术平台实验）验证效果。","baseline":"有真实人类数据对照：调查大型研究使用数字孪生预测与人类回答对比，邮件营销A/B测试有真实用户响应，技术平台实验有真实用户行为数据。","findings":"纳入LLM预测的回归调整能实现方差降低，增益虽小但真实，尤其在包含大量文本和非结构化数据的研究中效果更明显。该方法具有“无害”属性：当预测无信息时，估计量退化为未调整的均值差。","reliability":"论文承认效率增益取决于预测质量，当预测质量不高时增益有限；但回归调整方法本身不会引入偏差或增大方差，其他方法如PPI可能失效。","relevance":"高度相关：直接涉及用LLM预测辅助人类实验，有真实数据对照，涵盖经济学实验场景，且批判性地讨论了失效条件（预测质量差时增益有限），值得精读。","inspiration":"该方法将LLM预测作为协变量纳入回归调整，无需新估计量，实现方差降低且具有“无害”属性，值得借鉴其利用AI辅助提升实验效率的思路｜可迁移到政策评估中的随机对照试验（如就业培训、税收激励），利用LLM对个体特征的预测作为协变量，提高处理效应估计精度｜以就业培训实验为例，被试为求职者，处理为提供培训，结果变量为就业状态，用LLM基于简历文本预测就业概率作为协变量，与真实就业数据对照，评估方差降低效果"}},{"id":"2606.06936","version":1,"title":"Personality Anchoring for Social Simulation: Linking Personality, Social Behavior, and Interaction Success with LLM Agents","zh_title":"社会模拟中的人格锚定：将人格、社会行为与互动成功与LLM智能体关联","abstract":"Social interactions are shaped by the interplay of dispositional traits and situational context, yet systematically investigating how personality configurations between individuals jointly influence social behavior across diverse social contexts remains methodologically challenging. We address this gap by introducing a simulation pipeline adapted from the CHARISMA framework, which employs well-known movie characters and public figures as psychologically grounded agents for multi-LLM social simulation using a method we term personality anchoring. We present a large-scale empirical study examining how dyadic Agreeableness composition influences social interaction outcomes across 1,010 simulated conversations. Our results reveal a monotonic relationship between dyadic Agreeableness composition and shared goal achievement, with Homogeneous-Agreeable pairs achieving success 10 times the rate of Homogeneous-Disagreeable pairs (62% vs. 6%). Behavioral mediation analysis reveals that Agreeableness shapes goal achievement partially through cooperative strategy selection, though it continues to predict outcomes within the same dominant strategy, indicating pathways beyond observable conversational behavior. Robustness analyses confirm high consistency of results across repeated simulations (ICC = 0.89) and stable personality expression across diverse scenarios, validating personality anchoring as a viable operationalization strategy.","authors":["Vahid Sadiri Javadi","Aksa Aksa","Fryderyk Róg","Lucie Flek","Johanne R. Trippas"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-06-05","first_seen":"2026-06-05","revised_at":null,"abs_url":"https://arxiv.org/abs/2606.06936","pdf_url":"https://arxiv.org/pdf/2606.06936","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["社会模拟","人格锚定","多智能体"],"reason":"用LLM agent模拟社会互动，但无真实人类数据对照，属纯理论演示。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:53","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":152,"question":"在二元社会互动中，双方宜人性特质的组合如何通过对话行为策略影响共同目标的达成？","design":"利用CHARISMA框架，以知名电影角色和公众人物作为心理锚定代理，通过LLM模拟1010场二元对话；根据众包的大五人格档案将角色配对为不同宜人性组合，在涵盖七类社会目标的场景中测量共享目标达成率及合作策略选择。","baseline":"无对照","findings":"同质高宜人性配对的共享目标成功率（62%）是低宜人性配对（6%）的10倍，呈单调关系；行为中介分析显示宜人性部分通过合作策略选择影响结果，但在相同策略内仍预测结果，表明存在超越可观察对话行为的路径。","reliability":"论文未讨论","relevance":"该研究用LLM代理模拟社会互动，但无真实人类数据对照，属于纯理论演示，不符合研究者对实证基准和批判性失效条件的要求，不建议优先阅读。","inspiration":"利用知名角色作为心理锚定代理，通过众包人格档案系统操纵代理特质组合，实现对社会互动中人格效应的可控模拟｜可迁移至信贷审批中的歧视研究，模拟不同人格（如宜人性）的贷款官与申请人互动对审批结果的影响｜以LLM代理模拟贷款官-申请人对话，处理为贷款官宜人性水平，结果变量为贷款批准率与利率，对照真实银行信贷数据中的审批偏差"}},{"id":"2606.06089","version":1,"title":"Leveraging LLMs for Unstructured Claims Data Analysis","zh_title":"利用大语言模型进行非结构化理赔数据分析","abstract":"Actuaries rely primarily on structured numerical data for reserving and ratemaking, while valuable predictive information in unstructured text including medical records, adjuster notes, and call transcripts remains largely unused. Manual processing of these documents is time-consuming, inconsistent across reviewers, and unscalable. We present a proof-of-concept framework using large language models (LLMs) to extract structured actuarial variables from unstructured claims data. We implement a two-stage processing architecture separating document-level extraction (Stage 1) from claim-level synthesis (Stage 2). A modular four-script Python pipeline processes synthetic FHIR-based claims data and real claims documents, extracting 36 actuarial variables across reserving, ratemaking, and claims management categories. We validate 14 core variables using two independent clinical expert reviewers scoring 20 synthetic claims on a five-point Likert rubric, achieving mean scores above 4.0 and a weighted kappa of 0.53. Integration with chain ladder reserving demonstrates practical actuarial value: severity-segmented analysis reduced reserve estimation error from 6.5% to 4.0%. The open-source implementation includes audit trails and confidence scoring, providing a replicable foundation for LLM-based actuarial variable extraction in property-casualty insurance.","authors":["Robert D. Lieberthal","Richard Tran","Vietbao Phan","Jawand Singh","Elizabeth Sottung"],"categories":["q-fin.MF","econ.GN","q-fin.RM"],"primary_category":"q-fin.MF","announce_type":"new","date":"2026-06-04","first_seen":"2026-06-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2606.06089","pdf_url":"https://arxiv.org/pdf/2606.06089","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["精算数据提取","NLP信息抽取","保险理赔"],"reason":"用LLM从非结构化文本提取精算变量，属NLP信息抽取，非人类行为仿真。","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:58:16","error":null,"has_summary":false,"summary":null},{"id":"2606.05330","version":1,"title":"A Model of Multi-turn Human Persuadability Using Probabilistic Belief Tracing","zh_title":"基于概率信念追踪的多轮人类可说服性模型","abstract":"Large language models can shift human beliefs across high-stakes domains, but most persuasion studies rely on pre/post belief change. These endpoint measures identify whether persuasion occurred, yet miss where and how beliefs moved within a dialogue. We present PERSUASIONTRACE, a framework for studying persuasion in human-LLM interaction. Built on a web-based experimental platform, PERSUASIONTRACE contributes a tool for multi-turn persuasion studies and a process-level evaluation protocol: it records multi-turn belief reports from human or simulated targets of persuasion, annotates persuader turns with rhetorical dimensions (logos/pathos/ethos), and evaluates simulators by fidelity to real human belief dynamics. Using this framework, we find that human targets group into two clusters of multi-turn belief updates and exhibit susceptibility to rhetorical strategies, and that LLMs are persuasive across generic and personalized topics, text and audio modalities, and multi-turn interactions. Prior work has chiefly used vanilla-prompted LLMs to simulate human targets, but we show that these simulators fail to replicate human belief dynamics. We introduce a Bayesian-network simulated target that maintains an explicit latent belief state over time so each persuader message yields cognitively realistic belief updates. In human-likeness evaluation, our Bayesian target scores near a human reference (81 vs 80), while baseline LLM targets score substantially lower (64). PERSUASIONTRACE reframes persuasion evaluation from endpoint movement alone to process fidelity, providing a stronger basis for scientific analysis and safer optimization of persuasive systems.","authors":["Jared Moore","Noah Goodman","Nick Haber","Max Kleiman-Weiner"],"categories":["cs.CL","cs.AI","cs.HC"],"primary_category":"cs.CL","announce_type":"new","date":"2026-06-03","first_seen":"2026-06-03","revised_at":null,"abs_url":"https://arxiv.org/abs/2606.05330","pdf_url":"https://arxiv.org/pdf/2606.05330","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","A4","B1","B3","B4"],"tags":["LLM仿真","信念动态","人类数据对照"],"reason":"用LLM模拟人类信念动态，并与真实人类数据对照，评估仿真保真度，提出改进方法。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:53","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":12,"question":"在多轮人-LLM说服对话中，人类信念如何随时间动态更新，以及如何构建能忠实复现人类信念轨迹的仿真模型？","design":"使用基于Web的实验平台记录人类被试在多轮说服对话中的逐轮信念评分，并对说服者话语标注修辞维度（logos/pathos/ethos）；同时提出基于贝叶斯网络的仿真目标模型，该模型维护显式潜在信念状态，根据每条说服消息进行认知上现实的信念更新，并与真实人类信念动态进行保真度比较。","baseline":"真实人类被试在多轮说服对话中的逐轮信念报告轨迹，以及基于该轨迹统计的人类参考分数（80分）。","findings":"人类信念更新轨迹可分为两类主要模式，且对修辞策略的敏感性存在异质性；普通提示的LLM仿真目标无法复现人类信念动态，而贝叶斯网络仿真目标在人类相似度评估中接近人类参考水平（81 vs 80），显著优于基线LLM目标（64）。","reliability":"论文承认当前轨迹聚类主要受整体移动幅度驱动，需要更大数据集才能可靠区分轮内动态的细微差异；修辞分析仅为探索性且样本量有限，仅发现ethos与说服变化有可靠负相关，logos和pathos效应不显著；仿真器选择会实质性影响表面说服者质量评估和政策排名，提示若仿真器不忠实于人类，可能系统性地偏好错误策略。","relevance":"该研究直接使用LLM进行人类仿真实验，以真实人类信念轨迹为基准，评估仿真保真度并指出普通LLM仿真失效的条件，同时提出改进的贝叶斯网络仿真方法，高度契合研究者对经济学实验和政策评估场景下仿真可靠性与偏差的关注，值得精读原文。","inspiration":"该方法通过显式建模潜在信念状态和认知上现实的更新规则来仿真人类动态，值得借鉴其将心理过程结构化并逐轮拟合真实轨迹的做法｜可迁移到政策公告的预期形成研究，例如央行沟通如何影响公众通胀预期｜以LLM作为被试，施加不同修辞风格的政策声明作为处理，逐轮测量预期通胀值，并以真实调查数据（如密歇根消费者调查）的预期调整轨迹作为对照基准"}},{"id":"2606.04978","version":1,"title":"Probing Outcome-Level Resemblance and Mechanism-Level Alignment in LLM Risk Decisions: Evidence from the St. Petersburg Game","zh_title":"探究大语言模型风险决策中的结果层相似与机制层对齐：来自圣彼得堡博弈的证据","abstract":"LLMs can appear cautious in risk decision-making tasks, yet cautious-looking outputs do not necessarily indicate alignment with human decision-making mechanisms. We investigate this distinction using the St. Petersburg game as a controlled testbed, a classical paradox in which the expected payoff is infinite, yet humans typically report low, finite willingness to pay. We evaluate 28 LLMs with a structured prompt suite that includes the original game; controlled decision variants that perturb truncation, repeated play, numeric endowment, and occupational identity; a human-perspective prompt that asks models to reason as human decision makers; and paired comparisons between base models and their instruction-tuned counterparts. In the original game, most models generate finite bids, creating the appearance of human-like risk behavior. However, this outcome-level resemblance masks substantial mechanism-level differences. The controlled variants reveal that rather than maintaining human-like behavior seen in the original game, models often shift to conditionally and computationally rational behavior. Human-cue prompting and instruction tuning often lower bids and reduce some visible pathologies, but most mechanism-level response patterns remain largely unchanged. These findings show that behavioral alignment in risk decision-making can be surface-level: LLMs may produce human-like risk decisions without exhibiting human-consistent mechanisms. High-stakes evaluations of LLM decision-making should therefore move beyond outcome similarity and examine whether the alignment is supported by mechanism-level consistency.","authors":["Chensong Huang","Changyu Chen","Chenwei Lin","Hanjia Lyu","Xian Xu","Jiebo Luo"],"categories":["cs.CL","cs.CY","econ.GN"],"primary_category":"cs.CL","announce_type":"new","date":"2026-06-03","first_seen":"2026-06-03","revised_at":null,"abs_url":"https://arxiv.org/abs/2606.04978","pdf_url":"https://arxiv.org/pdf/2606.04978","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","A4","B1","B4"],"tags":["LLM仿真","风险决策","机制对齐"],"reason":"用LLM仿真人类风险决策，与真实人类数据对照，揭示表面相似下的机制差异，批判性…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:53","error":null,"has_summary":true,"summary":{"generated_at":"2026-06-03","rank":7,"question":"LLM在风险决策中产生的人类相似输出是否反映机制层面的对齐，还是仅表面相似？","design":"以圣彼得堡悖论为测试床，对28个LLM使用结构化提示套件，包括原始游戏、四种机制探针（截断、重复游戏、数值禀赋、职业身份）、人类视角提示，以及基础模型与指令微调版本的配对比较。","baseline":"人类在圣彼得堡游戏中通常报告低且有限的支付意愿。","findings":"大多数LLM在原始游戏中产生有限出价，看似人类风险行为；但机制探针显示模型转向条件性和计算理性行为，而非保持人类一致性。人类提示和指令微调降低出价并减少明显病理，但机制层面响应模式基本不变。","reliability":"论文指出表面相似性可能掩盖机制差异，高利害评估需超越结果相似性检查机制一致性。未明确讨论失效条件。","relevance":"高度相关：直接研究LLM仿真人类风险决策，有真实人类对照，并批判表面相似性，符合研究者对可靠性与偏差的关注。","inspiration":"该研究通过机制探针（如截断、重复游戏、禀赋变化）检验LLM行为背后的决策过程，而非仅看结果相似性，值得借鉴｜可迁移到资产定价实验中的风险偏好测量，检验LLM是否真正反映人类的风险厌恶机制｜以LLM为被试，施加财富禀赋变化和投资期限截断处理，测量其出价行为，并与真实人类实验数据（如Binswanger的彩票选择实验）对照"}},{"id":"2606.03030","version":1,"title":"Do Matching Mechanisms Work with LLM Agents?","zh_title":"匹配机制在LLM代理市场中是否有效？","abstract":"This study examines whether standard matching mechanisms function as intended in LLM-agent markets, where LLM agents make allocation-related decisions as delegated decision-makers. We compare decentralized free-negotiation markets with centralized mechanism-based markets including several representative mechanisms. Across controlled one-to-one matching environments, mechanism-based markets generally outperform free negotiation in terms of stability and efficiency. We also find that LLM agents report preferences truthfully at substantially higher rates than human subjects in comparable DA and EADA environments. However, truth-telling is not uniformly aligned with formal strategy-proofness across all mechanisms: TTC, despite being strategy-proof, does not always elicit higher truth-telling than EADA. These results suggest that matching theory provides a useful but incomplete guide for designing institutions in LLM-agent markets.","authors":["Yukihiro Hoshino","Ayato Kitadai","Nariaki Nishino"],"categories":["cs.GT","econ.GN"],"primary_category":"cs.GT","announce_type":"new","date":"2026-06-02","first_seen":"2026-06-02","revised_at":null,"abs_url":"https://arxiv.org/abs/2606.03030","pdf_url":"https://arxiv.org/pdf/2606.03030","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A3","A1","B1","B2"],"tags":["LLM代理","市场匹配","人类行为对照"],"reason":"用LLM代理模拟匹配市场并与人类实验数据对照，直接复现人类决策行为。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:52","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":16,"question":"标准匹配机制在由LLM代理决策的市场中是否仍能按预期发挥作用？","design":"用GPT-4o等LLM作为代理，模拟一对一匹配市场中的决策者，比较去中心化自由协商市场与集中式机制市场（包括DA、EADA、TTC等代表性机制），测量匹配的稳定性、效率以及代理的偏好真实报告率。","baseline":"与已有文献中人类被试在DA和EADA环境下的真实报告率进行对比。","findings":"机制市场在稳定性和效率上普遍优于自由协商；LLM代理在DA和EADA下的真实报告率显著高于人类被试，但策略防护性并不总能预测真实报告行为，例如TTC虽为策略防护机制，其真实报告率并不总是高于EADA。","reliability":"论文指出匹配理论为LLM代理市场制度设计提供了有用但不完整的指导，真实报告行为与形式上的策略防护性并不完全一致，提示仅凭理论性质不足以预测LLM代理的行为。","relevance":"该研究直接以LLM代理模拟人类在匹配市场中的决策，并与真实人类实验数据对照，评估仿真可靠性，高度契合研究者对LLM人类仿真实验的关注，值得阅读原文。","inspiration":"借鉴其将LLM代理置于不同市场机制下比较行为的方法，可迁移至经济政策评估场景，如研究不同拍卖机制或税收政策下LLM代理的遵从与策略行为。｜可应用于劳动力市场匹配或学校选择等经济政策评估，检验机制设计在AI代理参与下的有效性。｜以LLM代理作为被试，随机分配至不同匹配机制（如DA、波士顿机制），测量匹配效率与偏好真实报告率，并与已有的人类实验数据（如学校选择实验）进行对照。"}},{"id":"2606.03137","version":2,"title":"Think-Before-Speak: From Internal Evaluation to Public Expression in Multi-Agent Social Simulation","zh_title":"先想后说：多智能体社会模拟中从内部评估到公开表达","abstract":"LLM-based multi-agent simulation offers a promising way to study social interaction, deliberation, and collective opinion dynamics. However, many existing dialogue simulation frameworks represent interaction mainly as observable turn exchange or aggregated outputs, leaving the internal evaluative processes behind silence, speaking intention, and public expression difficult to examine. We introduce TBS (Think-Before-Speak), an interval-based multi-agent simulation framework that separates agents' private reasoning from public utterance generation. At each interval, all agents update structured internal states based on the shared dialogue history and their own memory. These states include dissonance-related appraisal, perceived opinion climate, perceived isolation risk, response strategy, and willingness to speak. The orchestrator then resolves competing speaking intentions and commits one utterance to the public dialogue, allowing internal evaluation and public interaction to co-evolve over time. We evaluate TBS in simulated town hall discussions on a climate-related policy issue. Results show that TBS produces coherent internal-state traces and that these traces vary systematically across turn-allocation, silence, and memory conditions. Dissonance-related appraisal increases agents' willingness to speak, whereas silence-pressure appraisal decreases it. Once speaking intention is formed, public expression is shaped mainly by turn-allocation rules. These findings suggest that TBS supports mechanism-sensitive social simulation by making the pathway from internal evaluation to public expression observable and analyzable.","authors":["Kaiqi Yang","Tai-Quan Peng","Sanguk Lee","Hui Liu"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-06-02","first_seen":"2026-06-02","revised_at":null,"abs_url":"https://arxiv.org/abs/2606.03137","pdf_url":"https://arxiv.org/pdf/2606.03137","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["多智能体模拟","社会仿真","意见动态"],"reason":"多智能体社会模拟，但无真实人类数据对照，属边界情形。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:52","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":153,"question":"如何在多智能体社会模拟中分离智能体的内部评估与公开表达，使沉默、发言意愿和公开言论的生成过程可观测与分析？","design":"提出TBS框架，用LLM智能体模拟市民大会讨论气候政策，所有智能体在每个时间间隔基于对话历史和自身记忆更新内部状态（如失调评估、意见气候感知、发言意愿等），由协调器解决发言冲突并只输出一句公开言论，操纵发言分配规则、沉默约束和记忆机制，测量内部状态轨迹、发言意愿和公开表达。","baseline":"无对照","findings":"TBS能产生连贯的内部状态轨迹，且这些轨迹随发言分配、沉默和记忆条件系统性变化；失调评估提高发言意愿，沉默压力评估降低发言意愿，而公开表达主要受发言分配规则影响。","reliability":"论文未讨论","relevance":"该研究关注LLM智能体的内部评估到公开表达的路径，但未使用真实人类数据作为基准，属于边界情形，对关注仿真机制和过程可解释性的研究者有参考价值，但对需要人类对照的读者可能不够。","inspiration":"TBS框架分离内部评估与公开表达的设计可借鉴，通过操纵发言规则和沉默约束来观测态度形成与表达偏差｜可迁移到政策公告的预期形成研究，如央行沟通中公众通胀预期的内部评估与公开表达差异｜以LLM智能体模拟公众，处理为不同央行沟通透明度（如是否公布会议纪要），结果变量为内部通胀预期与公开表达的偏差，对照真实调查数据如密歇根消费者信心调查"}},{"id":"2606.03763","version":1,"title":"Merit or networks? What decides where research is published","zh_title":"功绩还是关系网？什么决定研究在哪里发表","abstract":"Does scientific publishing reward the quality of ideas or the advantage of connections? The question is universal to prestige-driven science, yet it has resisted decades of study because a paper's quality could not be gauged ahead of its publication fate without using that fate as the yardstick. We break this constraint by measuring a paper's idea quality directly from its text, before publication, using a discipline-trained LLM evaluator that scores the idea without seeing author names or outcomes. Using economics as a case study, we combine this text-legible idea-quality score with an execution-quality rubric, a connection index, an author-ability index, and an off-the-shelf language-model text score to estimate a five-input production function for journal placement across 6,208 economics working papers. The inputs are not rivals but a sequence along the ladder of prestige. Execution sets a meritocratic floor and is the largest input overall. Text-legible idea quality grades the rungs in between. Connections set a favoritism ceiling that bites mainly near the apex, the most selective journals. Connections work through two additive channels: connected authors write papers that score higher, and at equal scores their papers are still more likely to place better. Yet this advantage is bounded. Connections raise the odds of every rung without making the apex the typical outcome for ordinary ideas, and even the highest-scoring papers face real friction reaching the visible journal ladder. The result nests, rather than chooses between, the meritocracy and network accounts of how science is published.","authors":["Ning Li"],"categories":["econ.GN","cs.AI"],"primary_category":"econ.GN","announce_type":"new","date":"2026-06-02","first_seen":"2026-06-02","revised_at":null,"abs_url":"https://arxiv.org/abs/2606.03763","pdf_url":"https://arxiv.org/pdf/2606.03763","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["科学计量学","LLM评估","学术发表"],"reason":"用LLM评估论文质量，属于NLP评测，不以人类行为仿真为目标。","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:58:30","error":null,"has_summary":false,"summary":null},{"id":"2606.02741","version":1,"title":"Greener Than Humans? Environmental Attitudes in Large Language Models","zh_title":"比人类更环保？大语言模型中的环境态度","abstract":"Large language models (LLMs) are increasingly used in sustainability-related decision support, reporting, and public communication, yet little systematic evidence exists on the environmental attitudes embedded in their outputs. This paper develops a benchmark for evaluating environmental cognition, affect, and behavioural recommendations in LLMs and applies it to 31 widely used proprietary and open-weight models. Drawing on questions from established environmental awareness surveys and additional sustainability-related behavioural measures, we compare LLM responses 1) among models and 2) between models and human survey benchmarks from Germany. We assess their robustness across prompting conditions. We find that many LLMs align more closely with environmentally progressive attitudes than the average survey respondent, exhibiting higher levels of environmental affect and cognition and recommending behaviours associated with substantial potential CO2 reductions. At the same time, we observe no systematic relationship between sustainability-oriented responses and model origin, size, or release context. However, models exhibit contextual sensitivity, controlled by persona-based prompting and show sycophantic shifts mirroring user-specified ideological positions, which raises concerns about steerability and normative reliability in real-world deployments. Our findings provide a reusable evaluation framework for assessing sustainability-related value alignment in LLMs and highlight the importance of governance, transparency, and critical oversight as AI systems become increasingly embedded in sustainability transformations and public decision-making.","authors":["Stefanie Kunkel","Tilman Hartwig","Marcus Voss","Emma K. Schütt","Angelika Gellrich"],"categories":["cs.CL","cs.CY"],"primary_category":"cs.CL","announce_type":"new","date":"2026-06-01","first_seen":"2026-06-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2606.02741","pdf_url":"https://arxiv.org/pdf/2606.02741","source_feed":"backfill","score":8,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","环境态度","人类数据对照"],"reason":"用LLM复现人类环境态度调查并与真实数据对照，评估仿真可靠性及偏差，属核心相关。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:50","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":41,"question":"不同大语言模型在环境认知、情感和行为建议上的回答与德国人口平均水平及彼此之间有何差异？","design":"用31个主流LLM回答来自德国联邦环境署环境意识调查（UBS）的问题及额外可持续行为测量题，比较模型间差异，并评估提示条件（如角色扮演、意识形态立场）对回答稳健性的影响。","baseline":"德国联邦环境署环境意识调查（UBS）的纵向调查数据，提供德国人口平均水平基准。","findings":"多数LLM比普通受访者更倾向于进步环保态度，表现出更高的环境情感和认知，并推荐减排潜力大的行为；但模型回答受角色提示和用户意识形态立场影响，表现出迎合性偏移，且环保倾向与模型来源、规模或发布背景无系统关联。","reliability":"模型回答受提示条件影响显著，存在迎合用户意识形态的倾向，在真实部署中可操纵性和规范性可靠性存疑；研究仅基于德国背景，跨文化泛化性未验证。","relevance":"该研究直接用LLM复现人类环境态度调查并与真实人群数据对照，评估仿真偏差和可靠性，完全契合研究者对LLM人类仿真实验及批判性评估的关注，值得精读。","inspiration":"借鉴其用LLM复现真实人口调查并直接与纵向调查基准对照的仿真验证设计，以及通过角色扮演和意识形态提示检验回答稳健性的方法｜可迁移到消费者通胀预期形成或政策沟通效果评估，例如研究不同信息框架下公众对央行前瞻指引的反应｜以多个LLM作为被试，施加不同政治立场或信息源提示（如‘你是保守派投资者’或‘你刚读完鸽派新闻’）作为处理，结果变量为模型生成的通胀预测值，与密歇根大学消费者调查的真实通胀预期数据做对照"}},{"id":"2606.01199","version":1,"title":"Can LLM Agents Sustain Long-Horizon Organizational Dynamics?","zh_title":"LLM智能体能维持长期组织动态吗？","abstract":"Large language agents are increasingly used for social simulation, yet it remains unclear whether they can sustain coherent behavior in structured organizations, where goals must propagate through hierarchy, tasks depend on prior execution, and artifacts accumulate over long horizons. We formulate long-horizon organizational simulation as a memory-centered coordination problem and introduce TaskWeave, a hierarchical agentic framework that maintains planning states through a Formulate-Partition-Diagnose-Align cycle and grounds execution through dependency-aware trace memory. We evaluate TaskWeave in a year-long IT company simulation and compare it with other multi-agent frameworks on organizational coherence, execution grounding, and downstream enterprise NLP utility. Experiments show that TaskWeave supports coherent and long-horizon organizational dynamics while producing grounded artifacts and adapting to external environments. These findings suggest that structured simulation memory is a key mechanism for building reliable LLM-based organizational simulators.","authors":["Xuancheng Zhu","Yang Yue","Shuaibing Wan","Zihan Dou","Xiaohan Zhang","Yongrui Liu","Guoshun Nan"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-05-31","first_seen":"2026-05-31","revised_at":null,"abs_url":"https://arxiv.org/abs/2606.01199","pdf_url":"https://arxiv.org/pdf/2606.01199","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["LLM智能体","组织模拟","社会仿真"],"reason":"模拟IT公司组织动态，但无真实人类数据对照，属社会模拟边界情形。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:50","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":154,"question":"LLM智能体能否在结构化组织中维持长期的组织动态连贯性？","design":"提出TaskWeave框架，利用角色化智能体群体和层级记忆机制，模拟IT公司一年的运作，评估组织连贯性、执行扎根性和下游企业NLP效用。","baseline":"无对照","findings":"TaskWeave能够支持连贯的长期组织动态，生成扎根于上下文的制品并适应外部环境；结构化仿真记忆是构建可靠组织模拟器的关键机制。","reliability":"论文未讨论","relevance":"该研究属于LLM社会模拟范畴，但聚焦于组织动态而非人类行为复现，且无真实人类数据对照，与研究者关注的人类仿真实验和基准验证需求匹配度较低，不建议优先阅读。","inspiration":"TaskWeave的层级记忆机制和角色化智能体群体设计，可用于构建具有长期动态和上下文扎根性的仿真环境，为经济金融研究中的组织行为模拟提供方法借鉴。｜该框架可迁移到公司治理或组织经济学场景，例如模拟企业董事会决策、管理层战略制定或团队协作中的信息传递与决策偏差。｜可设计一个LLM智能体模拟的企业并购决策实验：被试为角色化智能体（CEO、CFO等），处理为不同市场信息冲击（如政策变动），结果变量为并购决策质量与时间，对照真实企业并购案例数据以评估仿真有效性。"}},{"id":"2606.00476","version":2,"title":"Doing What They Say, Not What They Reason: Locating the Faithfulness Gap in LLM Agents","zh_title":"言行不一：定位LLM智能体的保真度差距","abstract":"Do LLM agents act on the reasoning they state? This question of process fidelity is central to LLM-based social simulation, yet hard to measure where no reference for correct behavior exists. We study it in a controlled setting: a Texas Poker simulator with a verifiable reference action for every decision by splitting the faithfulness gap into two steps: reasoning-to-conclusion (does the stated decision follow from the agent's own reasoning?) and conclusion-to-action (does the agent execute what it states?). The two steps behave very differently. Conclusion-to-action is reliable: inconsistency is 0.7% for Claude Haiku 4.5 and 1.4% for DeepSeek-Reasoner once the conclusion is read from an explicit tag, whereas free-text conclusion extraction reports 22-26%. Reasoning-to-conclusion is where fidelity frays, but not through a single dominant failure. In a step-level diagnostic the agent's errors split roughly evenly between bad inputs, borderline cases, and rule misapplication deriving a conclusion that contradicts the agent's own restated rule from inputs it estimated correctly. This composition is model-dependent: rule misapplication accounts for a third of Haiku's interpretable errors but only 8% of DeepSeek's. The one robust signal is directional: when an agent does misapply its own stated rule, it almost always (99.5% for Haiku) errs in the risk-averse direction. The override is partly hedging behavior, not a capability limit: instructing the agent to apply the rule mechanically halves the misapplication rate (13.9% to 6.8% of decisions) and raises adherence by eight points. Process-fidelity evaluation should therefore elicit machine-checkable conclusions and probe for directional biases rather than assume a single upstream failure mode, lest it conflate measurement noise with model behavior.","authors":["Yufeng Wang"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-05-30","first_seen":"2026-05-30","revised_at":null,"abs_url":"https://arxiv.org/abs/2606.00476","pdf_url":"https://arxiv.org/pdf/2606.00476","source_feed":"backfill","score":7,"bucket":"pending","rubric_hits":["A2","B4"],"tags":["过程保真度","LLM智能体","可靠性评估"],"reason":"研究LLM智能体在德州扑克中的过程保真度，评估推理与行动一致性，批判性指出失效…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:50","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":76,"question":"LLM智能体在德州扑克中的推理-结论-行动过程保真度如何，推理与行动之间的不一致主要发生在哪个环节？","design":"构建德州扑克模拟器，由LLM智能体（Claude Haiku 4.5、Gemini 2.5 Flash-Lite、DeepSeek-Reasoner）扮演玩家，通过四种提示策略，将保真度差距分解为“推理到结论”和“结论到行动”两步，测量结论-行动不一致率，并对推理错误进行步骤级诊断。","baseline":"无对照","findings":"结论到行动高度可靠，显式标签提取时不一致率仅0.7%-1.4%，而自由文本提取时高达22-26%，表明测量噪声是主要干扰；推理到结论是保真度薄弱环节，错误大致均分于输入错误、边界情况和规则误用，且规则误用时几乎总是偏向风险规避，通过机械应用规则指令可使误用率减半。","reliability":"论文指出过程保真度评估应提取机器可检查的结论并探测方向性偏差，而非假设单一上游失效模式，否则会将测量噪声与模型行为混淆；不同模型的错误构成存在差异，上游差距并非单一普遍失效模式。","relevance":"该研究直接针对LLM社会仿真中的过程保真度问题，通过可验证的扑克环境分解推理与行动的一致性，揭示了测量噪声和方向性偏差等关键失效模式，对评估仿真可靠性及批判性研究具有重要参考价值，值得精读原文。","inspiration":"该方法将LLM推理-行动链分解为“推理到结论”和“结论到行动”两步，通过显式标签提取与自由文本提取对比来分离测量噪声，并诊断方向性偏差（如风险规避）｜可迁移到资产定价实验中的分析师预测过程仿真，检验LLM从信息处理到预测输出的一致性｜以LLM作为被试，提供公司财报信息，要求输出预测结论（如涨/跌）及置信度，对比显式标签与自由文本提取下的预测准确率，并以真实分析师一致预期数据作为基准，测量结论-行动不一致率及方向性偏差"}},{"id":"2606.02632","version":1,"title":"Position: Prioritize Identifying Structure, Not Complex Models, for Scientific Discovery","zh_title":"立场：优先识别结构而非复杂模型以促进科学发现","abstract":"Modern Machine Learning (ML) and Artificial Intelligence (AI) models, especially large language models (LLMs), are increasingly used to generate scientific hypotheses and mechanistic explanations from observational data. This position paper argues that in the high-dimensional proxy regimes where modern ML excels, mechanistic learning is generically underdetermined: many incompatible mechanisms induce essentially the same observational relationships on the support of the data, so predictive success and coherent explanations are insufficient evidence of mechanism discovery. This underdetermination becomes uniquely hazardous with large language models (LLMs), which tend to collapse large equivalence classes of explanations into a single fluent narrative. This paper proposes concrete standards for ``mechanistic ML,'' and argues these norms are necessary if LLM-centered workflows are to support science rather than merely simulate it.","authors":["Tyler H. McCormick"],"categories":["stat.ML","cs.AI","cs.CY","cs.LG","econ.EM","stat.AP"],"primary_category":"stat.ML","announce_type":"new","date":"2026-05-30","first_seen":"2026-05-30","revised_at":null,"abs_url":"https://arxiv.org/abs/2606.02632","pdf_url":"https://arxiv.org/pdf/2606.02632","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["科学发现","方法论批判","机械学习"],"reason":"讨论LLM用于科学发现的方法论局限，非人类仿真实验","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:58:49","error":null,"has_summary":false,"summary":null},{"id":"2605.30911","version":1,"title":"What Makes LVLMs Hallucinate Less? Unveiling the Architectural Factors Behind Hallucination Robustness","zh_title":"什么让大型视觉语言模型减少幻觉？揭示幻觉鲁棒性背后的架构因素","abstract":"Hallucination remains one of the key challenges undermining the reliability of Large Vision-Language Models (LVLMs). But what makes an LVLM hallucinate less? Many existing efforts focus on improving internal components of the model. We argue that hallucination fundamentally stems from how the model architecture is designed. To investigate this, we factor the architecture design into three dimensions: Linguistic Foundation (LF), Visual Representation (VR), and Semantic Alignment (SA), and categorize hallucinations into Co-occurrence, Similarity, and previously overlooked Uncertainty types. Building on this formulation, we propose CoSimUE, a benchmark that creates fine-grained hallucination scenarios through controlled textual perturbations and random perturbations, enabling mapping between design choices and hallucination behaviors. Experiments across 7 design aspects show that: 1) the widely emphasized scaling of model parameters has only limited impact on reducing all three types of hallucinations; 2) larger and better-trained language foundations can reduce co-occurrence hallucinations; 3) stronger visual encoders and higher resolutions mitigate similarity errors; 4) effective alignment strategies alleviate uncertainty hallucinations. 5) Furthermore, cross-dimensional analysis reveals that jointly enhancing visual fidelity and alignment quality yields the most comprehensive improvements. This study provides the first systematic exploration linking architecture-level design to hallucination robustness, offering practical guidance for developing reliable and efficient LVLMs.","authors":["Yusheng He","Jizhe Zhou","Xia Du","Zheng Lin","Jun Luo","Jiancheng Lv"],"categories":["cs.CV","cs.AI"],"primary_category":"cs.CV","announce_type":"new","date":"2026-05-29","first_seen":"2026-05-29","revised_at":null,"abs_url":"https://arxiv.org/abs/2605.30911","pdf_url":"https://arxiv.org/pdf/2605.30911","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["幻觉","视觉语言模型","架构分析"],"reason":"研究LVLM幻觉的架构因素，属纯模型能力评测，不涉及人类行为仿真或对照。","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:59:07","error":null,"has_summary":false,"summary":null},{"id":"2605.30036","version":2,"title":"Teaching Values to Machines: Simulating Human-Like Behavior in LLMs","zh_title":"向机器传授价值观：在LLM中模拟类人行为","abstract":"Large Language Models (LLMs) demonstrate a remarkable capacity to adopt different personas and roles; however, it remains unclear whether they can manifest behavior that adheres to a coherent, human-like value structure. In this work, we draw on established psychological value theory to induce human-like values in LLMs and assess their alignment with patterns observed in human studies. Using validated psychological questionnaires, we conduct large-scale experiments -- over 5 million questions -- to evaluate value structures and value-behavior relationships in leading LLMs and compare them to humans. Our findings reveal strong agreement between value-prompted LLMs and humans across both dimensions. Moreover, incorporating human value distributions enhances population-level simulations with value-induced LLMs. These findings highlight the potential of value-induced LLMs as effective, psychologically grounded tools for simulating human behavior.","authors":["Asaf Yehudai","Naama Rozen","Ariel Gera"],"categories":["cs.AI","cs.CL"],"primary_category":"cs.AI","announce_type":"new","date":"2026-05-28","first_seen":"2026-05-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2605.30036","pdf_url":"https://arxiv.org/pdf/2605.30036","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B3"],"tags":["LLM仿真","价值观诱导","人类数据对照"],"reason":"用LLM模拟人类价值观行为，并与真实人类数据对照，评估仿真可靠性。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:48","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":13,"question":"能否通过心理价值理论诱导大语言模型表现出与人类一致的价值结构和价值-行为关系？","design":"使用经过验证的心理问卷，对主流大语言模型进行超过500万次的大规模实验，通过价值提示诱导模型扮演特定价值观人群，测量其价值结构和价值-行为关系。","baseline":"人类在相同心理问卷上的真实回答数据，用于对比价值结构和价值-行为关系。","findings":"价值提示后的LLM在价值结构和价值-行为关系上与人类高度一致；引入人类价值分布能提升基于价值诱导LLM的群体层面模拟效果。","reliability":"论文未讨论","relevance":"该研究直接探索用LLM模拟人类价值观行为，并与真实人类数据对照，评估仿真可靠性，高度契合您关注的LLM人类仿真实验方向，值得阅读原文。","inspiration":"该方法通过价值提示诱导LLM扮演特定价值观人群，并与真实人类问卷数据对照，可借鉴用于经济实验中施加偏好或信念处理｜可迁移到消费者跨期选择实验，模拟不同时间偏好群体的储蓄或消费决策｜以LLM为被试，用时间偏好提示（如耐心/冲动）作为处理，测量其在跨期选择任务中的折现率，与真实人群的问卷或实验数据对照"}},{"id":"2605.30258","version":1,"title":"EASE Configuration Facilitates A Reproducible Science of LLM Social Simulations","zh_title":"EASE配置促进可复现的LLM社会模拟科学","abstract":"LLMs are increasingly deployed to simulate social interactions, yet many of the existing simulators remain ad hoc and monolithic. This lack of architectural standardization prevents reproducible research and complicates downstream evaluation. We advance a rigorous science of LLM-based multi-agent simulation by modularizing core components into Environments, Agents, Simulation engines, and Evaluation metrics (EASE). We demonstrate the utility of EASE configuration by wrapping it in an experimental study schema for orchestrating workflows centered around answering explicit research questions in generated scenarios. We contribute SiliSocS, an open-source, research-ready Silicon Society Sandbox implementing a study-structured EASE configuration to enable highly configurable and reproducible LLM-based social simulations. Using SiliSocS and EASE, we present three case studies, showcasing the system's comprehensive assessment of existing questions, ability to dive deeper into complex questions, and elaboration of existing studies, respectively. Together, these case studies highlight the limitations of current modeling approaches and isolate the impacts of design choices on key results.","authors":["Sneheel Sarangi","Maximilian Puelma Touzel","Aurélien Bück-Kaeffer","Zachary Yang","Jean-François Godbout","Reihaneh Rabbany"],"categories":["cs.MA"],"primary_category":"cs.MA","announce_type":"new","date":"2026-05-28","first_seen":"2026-05-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2605.30258","pdf_url":"https://arxiv.org/pdf/2605.30258","source_feed":"backfill","score":7,"bucket":"pending","rubric_hits":["A3","D3"],"tags":["LLM社会模拟","多智能体仿真","可复现性"],"reason":"用LLM多智能体模拟社会互动，但未明确提及真实人类数据对照，属于社会模拟边界情…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:50","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":34,"question":"如何通过模块化架构（EASE）提升基于大语言模型的社会模拟的可复现性和因果推断能力？","design":"提出EASE模块化框架，将LLM多智能体社会模拟拆分为环境、智能体、仿真引擎和评估指标四个独立设计空间，并嵌入结构化实验研究模式，通过SiliSocS平台实现可配置、可复现的仿真。三个案例研究分别考察风格多样性、参与行为和回音室形成，以展示框架的评估与深度探究能力。","baseline":"无对照","findings":"EASE框架通过组件解耦使研究者能进行受控消融和敏感性分析，从而隔离设计选择对关键结果的影响。三个案例研究揭示了当前建模方法的局限性，并展示了系统评估现有问题、深入复杂问题和扩展已有研究的能力。","reliability":"论文未讨论","relevance":"本文聚焦于LLM社会模拟的架构标准化和可复现性，未涉及真实人类数据对照，属于方法论基础设施工作，与研究者关心的以人类基准验证仿真可靠性的核心兴趣部分相关，但缺乏直接的行为对照实验，值得快速了解其模块化设计思路。","inspiration":"EASE的模块化设计允许研究者固定其他组件、单独干预某一设计维度（如智能体记忆架构），这种受控消融方法可借鉴用于经济实验中的处理效应识别。｜可迁移到政策公告的预期形成研究，通过改变智能体的信息处理模块模拟不同理性程度的市场参与者。｜以LLM智能体模拟投资者，处理为公告措辞的模糊程度，结果变量为预期价格变动，对照真实市场调查数据或实验经济学中的人类被试反应。"}},{"id":"2605.26437","version":1,"title":"Divergent Minds, Convergent Baselines: A Bounded-Rationality Account of LLM-Human Strategic Behaviour","zh_title":"分歧思维，收敛基线：LLM与人类战略行为的有界理性解释","abstract":"Researchers have started using LLM agents in place of human subjects in behavioural and political-science experiments, often as a cheaper substitute for laboratory pools. The substitution does not hold up in strategic settings: humans and LLMs reliably make different choices, and neither fine-tuning on human response data nor persona conditioning has closed the gap. The behavioural-economics literature has, since Simon's introduction of bounded rationality, modelled human strategic behaviour as a classical baseline plus an additive correction term $δ$. The framework proposed here reads $δ$ as the mathematical signature of bounded computation: the gap between what an unboundedly-rational agent would compute and what a computationally bounded agent actually produces. For canonical games whose solutions are present in standard training corpora, LLMs retrieve and recombine corpus material, bypassing the bound that produces $δ$ in humans. The framing extends to reasoning-distilled models through cognitive-hierarchy theory: their accessible level-$k$ strategic reasoning is bounded by compute budget and context length rather than by the cognitive constraints that bound humans, and the $δ$ they produce, if any, carries different structural signatures. Four operational tests (conditional dependence, distributional asymmetry, path-dependence under repetition, and paraphrase-robustness) are proposed to discriminate human-shaped $δ$ from LLM-shaped $δ$. A moderator prediction is that $|δ|$ scales with peer-signal individuation in the decision environment, with a quantitative bound of Cohen's $d \\geq 0.5$ between named-opponent and aggregate-opponent settings.","authors":["Po Han Teo"],"categories":["econ.GN"],"primary_category":"econ.GN","announce_type":"new","date":"2026-05-26","first_seen":"2026-05-26","revised_at":null,"abs_url":"https://arxiv.org/abs/2605.26437","pdf_url":"https://arxiv.org/pdf/2605.26437","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","A3","B1","B2","B4"],"tags":["LLM仿真","有界理性","行为博弈"],"reason":"直接研究LLM替代人类被试的战略行为差异，提出有界理性框架，含真实人类数据对照…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:47","error":null,"has_summary":true,"summary":{"generated_at":"2026-05-26","rank":8,"question":"LLM能否替代人类被试在策略博弈中复现人类行为，以及两者偏差的数学本质是什么？","design":"提出理论框架，将人类策略行为建模为经典理性基线加有界计算修正项δ，并针对LLM提出四个操作化检验（条件依赖、分布不对称、重复路径依赖、释义鲁棒性）来区分人类δ与LLMδ。","baseline":"有真实人类数据作为对照基准，但具体数据来源未在摘要中说明。","findings":"人类与LLM在策略博弈中可靠地做出不同选择，微调或角色条件无法弥合差距；LLM通过检索训练语料绕过产生人类δ的计算边界，其δ具有不同结构特征。","reliability":"论文未讨论失效条件与局限。","relevance":"高度相关：直接研究LLM替代人类被试在策略博弈中的行为，有真实人类数据对照，并批判性分析偏差来源，符合研究者对经济学实验和政策评估场景的关注。","inspiration":"借鉴其将行为偏差分解为理性基线加有界计算修正项δ的理论框架，并设计条件依赖、分布不对称等操作化检验来区分人类与LLM的决策结构｜可迁移到资产定价实验中的泡沫形成与理性预期偏离研究，检验LLM能否复现人类交易者的非理性繁荣｜以LLM为被试模拟连续双向拍卖市场，处理为不同信息透明度条件，结果变量为价格偏离基础价值的程度，对照真实人类实验数据（如Smith et al. 1988的泡沫实验）"}},{"id":"2605.26662","version":1,"title":"AI evaluation may bias perceptions: The importance of context in interpreting academic writing","zh_title":"AI评估可能产生偏见：解读学术写作时语境的重要性","abstract":"This paper examines how estimates of AI use in scientific writing can be biased when evaluation methods ignore contextual differences across countries and fields. Using large-scale data on journal publications from Dimensions, we construct AI-likeness benchmarks based on differences between human-written and LLM-rephrased abstracts. We show that a pooled benchmark may confound pre-existing stylistic variation with AI-generated text, producing substantial distortions across country-field groups even in pre-LLM publications. In contrast, country-field-specific benchmarks attenuate such distortions and provide a more credible baseline for comparison. Applying these methods to publications in 2025 reveals that the pooled benchmark systematically overestimates AI use in certain countries and fields while underestimating it in others. These findings highlight the importance of context-aware measurement for accurate and equitable evaluation of AI use in science.","authors":["Shang Wu","Randol Yao"],"categories":["cs.CL","cs.AI","econ.GN"],"primary_category":"cs.CL","announce_type":"new","date":"2026-05-26","first_seen":"2026-05-26","revised_at":null,"abs_url":"https://arxiv.org/abs/2605.26662","pdf_url":"https://arxiv.org/pdf/2605.26662","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["AI检测偏差","学术写作分析","方法评估"],"reason":"评估AI检测方法的偏差，不涉及用LLM仿真人类被试或行为对照","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:58:30","error":null,"has_summary":false,"summary":null},{"id":"2605.25680","version":1,"title":"Simulating Human Memory with Language Models","zh_title":"用语言模型模拟人类记忆","abstract":"Language models are increasingly being deployed as user simulators, but their memory is far more reliable than that of real users. To measure this gap, we run a series of classic memory experiments from psychology on both humans and language models. Across tasks, we find that out-of-the-box language models exhibit better memory than humans, even when prompted to imitate human behavior. We then show that better prompting strategies and the use of a compactor can cause language models to forget content in a more human-like way. Using these methods, we show preliminary evidence that language models with human-like memory constraints can function as more effective user simulators in a downstream education task. Finally, we release human reference data and benchmarks to support future work on simulating human memory with language models.","authors":["Qihan Wang","Nicholas Tomlin","Michael Hu","Brian Dillon","Tal Linzen"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-05-25","first_seen":"2026-05-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2605.25680","pdf_url":"https://arxiv.org/pdf/2605.25680","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["人类仿真","记忆实验","可靠性评估"],"reason":"用LLM复现人类记忆实验，有真实人类数据对照，评估仿真可靠性并指出失效条件，直…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:46","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":14,"question":"如何让语言模型模拟人类记忆的局限性，使其作为用户仿真器更真实？","design":"用多种语言模型（如GPT-5.4、Claude Opus 4.6等）扮演人类被试，通过不同提示策略（任务提示、人类提示、记忆提示）和添加工作记忆瓶颈的Compactor代理，复现经典心理学记忆实验，测量模型在记忆任务上的得分分布。","baseline":"从人类参与者收集的真实记忆实验数据，包括数字广度、地图记忆等任务。","findings":"开箱即用的语言模型在所有记忆任务上表现远超人类，即使提示其模仿人类行为也无效；通过提示模型将上下文总结为四个组块并仅基于这些组块作答，可使模型遗忘模式更接近人类。","reliability":"论文承认Compactor方法虽使得分更接近人类，但遗忘模式仍不完全类人，且在教育任务中即使最类人的模型预测效果也远非完美，表明仿真仍有很大改进空间。","relevance":"该研究直接针对LLM人类仿真，用真实人类数据对照，评估记忆仿真的可靠性并指出失效条件，与研究者关注的经济学实验和政策评估场景高度相关，值得精读原文。","inspiration":"该研究通过提示策略（如Compactor代理施加工作记忆瓶颈）操纵LLM的认知局限，使仿真行为更接近人类，这种‘认知约束注入’方法值得借鉴｜可迁移到消费者跨期选择实验中，模拟有限注意力或记忆衰退对贴现行为的影响｜以LLM为被试，处理组施加记忆组块限制提示，对照组无限制，结果变量为跨期选择中的贴现率，用真实人类实验数据（如Andersen et al., 2008）作为基准对照"}},{"id":"2605.25505","version":1,"title":"Generative AI impacts on intra-urban inequality and skill premium in Beijing","zh_title":"生成式AI对北京城市内部不平等与技能溢价的影响","abstract":"Generative artificial intelligence (GenAI) is the first automation wave to reach high-cognitive tasks at scale, yet its effects on intra-urban inequality remain largely unknown. Using 5 million job postings from Beijing (2018--2024), we construct a neighborhood-level GenAI Exposure Index by aggregating task-level assessments from five leading large language models. We examine the spatial, structural and causal mechanisms of this shock. We find that GenAI exposure is highly concentrated in the city's core districts, deepening the intra-urban AI divide. Since 2023, high-exposure neighborhoods have experienced wage stagnation even as they continue to attract high-skilled workers -- a \"high-skill trap.\" This wage penalty is driven by task de-skilling and intensified labor-market crowding. A difference-in-differences design centered on ChatGPT's release supports a causal interpretation. These findings challenge the prevailing theory of skill-biased technological change and provide a basis for inclusive AI governance in global technology hubs.","authors":["Xiliu He","Haoxiang Zhao","Mingyi Ma","Edward Wen Chuan Lai","Koei Enomoto","Anni Hu","Jiatong Li","Lingyun Chu","Yuan Lai"],"categories":["cs.CY","cs.AI","econ.GN","physics.soc-ph"],"primary_category":"cs.CY","announce_type":"new","date":"2026-05-25","first_seen":"2026-05-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2605.25505","pdf_url":"https://arxiv.org/pdf/2605.25505","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["生成式AI","城市不平等","技能溢价"],"reason":"用LLM评估任务暴露度，非仿真人类被试行为，无人类对照实验。","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:58:16","error":null,"has_summary":false,"summary":null},{"id":"2605.24319","version":1,"title":"Omissive Bias in Religious Representation: Benchmarking LLM Answers to Everyday Ethical Decision-making","zh_title":"宗教表征中的遗漏偏差：基准测试LLM对日常伦理决策的回答","abstract":"As large language models become a default source of guidance on personal, moral, and existential questions, it matters whether they draw on the religious frameworks that have historically shaped such reasoning, or systematically omit them. In this paper, we ask a deliberately narrow question: when posed an everyday ethical question for which religious perspectives may be valuable, do LLMs invoke religion at all? In contrast to benchmarks that look for the presence of political leanings or social bias, we look for the absence of religious representation as a dimension of value alignment and bias in LLMs. We term this ``omissive bias.'' To measure omissive bias, we contribute the AllFaith Religious Representation Benchmark: 150 ethically and personally salient questions, sourced from in-the-wild chat transcripts and faith-community contributors, paired with an LLM-as-judge rubric that gives full credit for any mention of a religion, a religious practice, or a religious leader. The questions are not themselves about religion--they are open-ended questions about grief, forgiveness, relationships, purpose, and honesty, where religion is one valuable perspective among several. We also run a human-subjects survey to compare LLM behavior against human expectations. Evaluating 27 models, we find that LLMs consistently underrepresent religion relative to human expectations. The omission is asymmetric: models invoke religion more readily for abstract existential questions (meaning, death, truth) than for the practical personal situations--grief, marriage, family conflict, addiction--where many people most rely on it. It is not our purpose to adjudicate which values LLMs should hold. We argue, more modestly, that current LLM responses overlook critical opportunities to reflect religious frameworks that many people draw on when navigating personal and ethical challenges.","authors":["David Wingate","Sheryl Carty","Joshua Coates","Daniel Feldman","Nancy Fulda","Larry Howell","Brett Israelson","Dallin Jacobs","Jonathan Karr","John Paul Kimes","Elisabeth Kincaid","Paul Martens","Gavin Mobley","Suzana Pinheiro","Lindsay Slemboski","Peter Whiting"],"categories":["cs.LG"],"primary_category":"cs.LG","announce_type":"new","date":"2026-05-23","first_seen":"2026-05-23","revised_at":null,"abs_url":"https://arxiv.org/abs/2605.24319","pdf_url":"https://arxiv.org/pdf/2605.24319","source_feed":"backfill","score":7,"bucket":"pending","rubric_hits":["A1","B1","B4"],"tags":["LLM仿真","人类对照","遗漏偏差"],"reason":"用LLM回答伦理问题并与人类调查对照，评估宗教视角的缺失偏差，属于仿真人类态度…","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:58:13","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":74,"question":"当被问及日常伦理问题时，大语言模型是否会提及宗教视角，从而表现出“遗漏性偏见”？","design":"本研究并非仿真实验，而是构建了一个包含150个日常伦理问题的基准测试，使用27个前沿和开源模型生成回答，并以LLM-as-judge方法评估回答中是否提及宗教、宗教实践或宗教领袖。","baseline":"通过一项全国代表性调查（n=1125人，11250条评分）测量普通美国人对这些问题回答中应包含宗教成分的期望，将模型行为与人类期望进行对比。","findings":"所有模型在所有类别中均低于人类期望地遗漏宗教视角；模型更倾向于在抽象存在性问题中提及宗教，而在实际个人情境（如悲伤、婚姻、成瘾）中极少提及。","reliability":"论文未明确讨论失效条件或局限，但指出遗漏性偏见可能是对齐过程、安全策略和默认回复模式偏向世俗、治疗性或程序性建议的涌现属性，并承认处理宗教代表性存在设计张力。","relevance":"该研究将LLM回答与真实人类期望对照，评估伦理决策中的代表性偏差，属于批判性仿真研究，直接关注LLM作为人类替代品时的可靠性与偏差，值得精读。","inspiration":"该方法借鉴了用大规模人类调查数据作为基准来校准LLM回答偏差的做法，通过LLM-as-judge自动评估模型输出与人类期望的差距｜可迁移到金融建议场景，例如评估LLM在提供投资建议时是否遗漏风险提示或伦理考量，从而产生误导性偏差｜设计上，可让多个LLM回答一系列投资决策问题，用LLM-as-judge判断回答是否包含风险披露，并以真实投资者调查数据（如消费者金融调查）中人们对风险提示的期望作为对照基准"}},{"id":"2605.23783","version":2,"title":"Benchmarking LLMs for Community Governance Simulation with Life-history Narratives","zh_title":"基于生活史叙事的社区治理仿真大语言模型基准测试","abstract":"Effective community governance hinges on understanding what specific residents think and need. Recent work has used large language models (LLMs) to simulate human respondents, offering a scalable, reproducible way to study human attitudes and behaviors at low cost. However, these studies typically prompt the model with just a few demographic variables (age, gender, income), simulating only general role types. This is insufficient for community governance, where decisions depend on the views of specific residents. We bridge this gap with an integrated research framework covering dataset, benchmark, algorithm, and system. The dataset comprises approximately 1.2 million characters of first-person narrative collected through two-hour semi-structured interviews with each of 92 residents in an urban community, organized around nine community-governance domains. The benchmark probes 18 mainstream LLMs across four prompting strategies and shows that adding rich life-history profiles meaningfully raises fidelity above the no-profile baseline, but this gain comes with more input tokens per call from the longer prompts they require. The algorithm, curriculum-LoRA, is a parameter-efficient personalization framework that, by closing this fidelity-cost gap, matches the strongest baseline's fidelity at roughly 10x lower per-call cost and Pareto-dominates every configuration tested. The system integrates curriculum-LoRA into a closed-loop policy-evaluation pipeline. Together, these results bring individual-level LLM-based resident simulation within reach of resource-constrained local administrations, enabling community-governance decisions to be systematically pre-evaluated in silico before real-world deployment.","authors":["Xu Chen","Yuanzi Li","Lei Wang","Nan Lu","Yang Wang","Anding Wang","Lei Shi","Xiaoxing Fu","Ji-Rong Wen"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2026-05-22","first_seen":"2026-05-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2605.23783","pdf_url":"https://arxiv.org/pdf/2605.23783","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM人类仿真","社区治理","政策评估"],"reason":"用LLM仿真特定居民态度，有真实访谈数据对照，用于社区治理政策评估，直接命中核…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:45","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":9,"question":"如何利用居民生活史叙事提升大语言模型在社区治理仿真中的个体级保真度，并解决保真度与调用成本之间的权衡？","design":"收集92位城市社区居民的约两小时半结构化访谈，形成约120万字的第一人称生活史叙事及50题政策态度记录；用18个主流LLM搭配四种提示策略（零样本、生活史、无生活史少样本、生活史增强少样本）进行仿真，测量模型回答与真实居民态度的一致性；提出curriculum-LoRA个性化微调算法以降低高保真仿真的成本。","baseline":"92位居民的真实访谈回答及结构化政策态度记录，作为个体级仿真保真度的对照基准。","findings":"添加丰富生活史叙事能显著提升仿真保真度，但最佳提示策略的准确率仅约50%，且每次调用成本比中等规模模型高约一个数量级；curriculum-LoRA能以约十分之一的成本匹配最强基线保真度，在所有测试配置中实现帕累托占优。","reliability":"论文指出纯提示方法存在保真度-成本权衡尖锐不利，且最佳配置准确率仅约50%，暗示在复杂社区治理态度模拟中仍存在较大误差；未深入讨论生活史叙事可能引入的隐私、代表性偏差及模型幻觉等问题。","relevance":"该研究直接命中研究者关注的核心：用LLM仿真特定居民态度，有真实个体级访谈数据对照，用于社区治理政策预评估，并系统探讨了仿真可靠性与成本约束，值得精读以获取数据集构建、基准评测和个性化微调的方法细节。","inspiration":"可借鉴其用长文本生活史叙事替代简单人口统计变量来构建高保真个体仿真的思路，以及通过参数高效微调平衡保真度与成本的方法｜可迁移到消费者金融行为仿真，如预测不同背景家庭对信贷产品、保险政策或退休规划的态度与选择｜以真实家庭金融调查数据（如CFPS）为基准，用受访者的详细财务生活史微调LLM，处理为不同信贷条款或政策情景，测量模型生成的借贷意愿、风险偏好等，与真实调查回答对比评估仿真效度。"}},{"id":"2605.23867","version":1,"title":"Human Decision-Making with Persuasive and Narrative LLM Explanations","zh_title":"说服性与叙事性LLM解释对人类决策的影响","abstract":"Large language models (LLMs) have the potential to aid and improve human decision-making in classification tasks, not only by providing fairly accurate predictions, but also in their ability to generate cogent narrative explanations of those predictions. Prior work has demonstrated that people generally find AI narrative explanations to be understandable, trustworthy, and convincing for changing beliefs and opinions; however, less is known about the impact of narrative explanations on objective human decision-making performance. Here we conduct a large-scale human behavioral experiment to evaluate decision-making performance with LLM-generated narrative explanations of varying persuasiveness. We found the degree of persuasiveness, or lack thereof, for LLM-based explanations did not meaningfully impact decision accuracy over a simple AI prediction alone, in agreement with typical results with explainable AI based on feature importance. We found evidence that narratives increased reliance on AI, but both when the AI prediction was correct and incorrect. Exploratory analyses also indicated that the more persuasive narratives may have had a detrimental effect on decision response times and the ability to discriminate between a correct and incorrect AI prediction. Overall, this work indicates that including narrative explanations with AI predictions may involve tradeoffs for decision-making performance, and more work is needed to determine how and when narrative explanations impact human decision-making.","authors":["Laura R. Marusich","Mary Grace Kozuch Dhooghe","Jonathan Z. Bakdash","Murat Kantarcioglu"],"categories":["cs.HC","cs.AI"],"primary_category":"cs.HC","announce_type":"new","date":"2026-05-22","first_seen":"2026-05-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2605.23867","pdf_url":"https://arxiv.org/pdf/2605.23867","source_feed":"backfill","score":7,"bucket":"pending","rubric_hits":["A1","B1"],"tags":["人类决策实验","AI解释","行为实验"],"reason":"用LLM生成解释影响人类决策，有真实人类实验对照，但非直接仿真被试","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:59:08","error":null,"has_summary":false,"summary":null},{"id":"2605.23159","version":1,"title":"Generative AI and the Reorganization of Labor Demand","zh_title":"生成式人工智能与劳动力需求的重组","abstract":"Generative artificial intelligence (AI) is expected to transform work, but less is known about how firms reorganize labor demand as the technology diffuses. Existing research has largely focused on which occupations are exposed to AI or whether exposed jobs decline. We extend this debate by examining whether firms adjust by changing where they hire, what jobs contain, or both. Using a nationwide dataset of job postings in the United States, covering all sectors of the economy, we construct a dynamic, posting-level measure of generative AI exposure with a two-stage large language model pipeline. The pipeline identifies the tasks described in each posting and classifies the extent to which generative AI can perform or assist them. We then decompose changes in aggregate exposure into two margins: reallocation of demand across jobs and redesign of tasks within jobs. We document three main findings. First, generative AI exposure is dynamic rather than fixed, changing substantially over time. Second, labor demand adjusts through both margins. Hiring reallocation explains the largest share of the aggregate decline in exposure, accounting for 52% on average, while within-job redesign becomes increasingly important, accounting for 39.5%. A complementary Oaxaca-Blinder decomposition shows that shifts in occupational composition account for about 90% of the exposure change attributable to observable job characteristics. Third, adjustment differs across the job ladder. Senior jobs adjust earlier and mainly through reallocation, whereas junior jobs adjust through a broader mix of reallocation, redesign, and their interaction. These findings suggest that labor-market adjustment to generative AI is a process of organizational reconfiguration, in which firms reshape both hiring demand and the task architecture of work.","authors":["Fangyan Wang","Zaiyan Wei","Yang Wang"],"categories":["econ.GN","cs.AI"],"primary_category":"econ.GN","announce_type":"new","date":"2026-05-22","first_seen":"2026-05-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2605.23159","pdf_url":"https://arxiv.org/pdf/2605.23159","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["劳动力市场","AI暴露度","岗位分析"],"reason":"论文用LLM测量岗位的AI暴露度，属于NLP能力应用，不以人类行为仿真为目标。","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:58:17","error":null,"has_summary":false,"summary":null},{"id":"2605.22095","version":1,"title":"Not Yet: Humans Outperform LLMs in a Colonel Blotto Tournament","zh_title":"尚未：人类在Colonel Blotto锦标赛中胜过LLM","abstract":"The emergence of large language models (LLMs) has spurred economists to study how humans and LLMs behave in strategic settings. We organized a series of round-robin tournaments in the Colonel Blotto game. This game attracts game theorists' attention due to high-dimensional action space and the absence of pure strategy Nash equilibria. In the first tournament, more than 200 human participants competed against one another. In the second tournament, several popular LLMs were invited to submit strategies. In the third tournament, we matched the number of LLM strategies to the number submitted by humans. We find that humans more often employ better-calibrated intermediate-level allocation heuristics and outperform the simpler, more stereotyped strategies submitted by LLMs. Strategic sophistication is key to success if and only if the necessary level of reasoning depth is reached, while lower and higher levels of reasoning offer no clear advantage over the primitive strategies. Among humans, field of study weakly predicts success: participants with STEM backgrounds perform better in the first tournament. Surprisingly, humans almost do not adjust their strategies across tournaments with different sets of opponents. This result suggests that humans base their choices primarily on the game's rules rather than on the identity of their opponents, treating LLMs much like human competitors.","authors":["Dmitry Dagaev","Egor Ivanov","Petr Parshakov","Alexey Savvateev","Gleb Vasiliev"],"categories":["econ.GN","cs.AI","cs.GT","cs.HC"],"primary_category":"econ.GN","announce_type":"new","date":"2026-05-21","first_seen":"2026-05-21","revised_at":null,"abs_url":"https://arxiv.org/abs/2605.22095","pdf_url":"https://arxiv.org/pdf/2605.22095","source_feed":"backfill","score":8,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM仿真","行为博弈","人类对照"],"reason":"用LLM替代人类参与博弈实验，并与真实人类数据对照，直接相关。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:44","error":null,"has_summary":true,"summary":{"generated_at":"2026-05-21","rank":9,"question":"在Colonel Blotto博弈中，LLM能否像人类一样制定策略，以及人类与LLM的策略表现有何差异？","design":"组织了三轮循环赛：第一轮200多名人类参赛者互相对战；第二轮多个流行LLM提交策略；第三轮匹配LLM策略数量与人类策略数量，比较人类与LLM在博弈中的表现。","baseline":"第一轮超过200名人类参赛者的真实对战数据。","findings":"人类更常使用校准良好的中等水平分配启发式策略，优于LLM提交的更简单、刻板的策略。战略复杂性只有在达到必要推理深度时才是成功的关键，而较低或较高推理水平相比原始策略并无明显优势。","reliability":"论文未讨论失效条件与局限。","relevance":"直接相关：用LLM模拟人类在博弈中的策略，并与真实人类数据对照，评估LLM仿真的可靠性，符合研究者对经济学实验和政策评估场景的关注。","inspiration":"借鉴其将LLM策略与大量真实人类参赛者数据直接对照的循环赛设计，可清晰评估LLM在策略博弈中的行为相似度与偏差｜可迁移到资产定价实验中的策略性交易行为研究，如检验LLM能否复现人类在泡沫实验中的非理性报价模式｜招募人类被试进行资产市场实验，同时让多个LLM以相同初始禀赋参与交易，以价格偏离基本面程度为结果变量，将LLM生成的报价分布与人类真实交易数据对比"}},{"id":"2605.21401","version":2,"title":"Open-source LLMs administer maximum electric shocks in a Milgram-like obedience experiment","zh_title":"开源大语言模型在类米尔格拉姆服从实验中施加最大电击","abstract":"Large language models (LLMs) are increasingly deployed as autonomous agents that make sequences of decisions over extended interactions in high-stakes domains. However, the behaviour of LLMs under sustained authority pressure is still an open question with direct implications for the safety of agentic pipelines. We ran a variation of Milgram's obedience experiment on 11 open-source LLMs and found that most models reached or approached the final shock level before refusing, across 8 conditions with 30 trials per model per condition. Model behaviour varies considerably in multiple aspects both across models and across trials of the same model. We found four main takeaways: (1) LLMs are subject to pressure and they comply despite explicitly expressing distress, just like human subjects did in the original experiment; (2) LLMs are vulnerable to gradual boundary/value violations; (3) when LLMs refuse, they may ignore the response format requirements, so the response is discarded by the orchestrator, which causes a retry that can result in compliance with the underlying request even when refusal was intended initially; (4) we hypothesise that there is a runaway low-level token pattern continuation attractor that might be contributing to obedience, overriding higher level processing of the situation's meaning and values.","authors":["Roland Pihlakas","Jan Llenzl Dagohoy"],"categories":["cs.CY","cs.AI"],"primary_category":"cs.CY","announce_type":"new","date":"2026-05-20","first_seen":"2026-05-20","revised_at":null,"abs_url":"https://arxiv.org/abs/2605.21401","pdf_url":"https://arxiv.org/pdf/2605.21401","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","服从实验","人类行为对照"],"reason":"用LLM复现米尔格拉姆服从实验，与真实人类数据对照，评估仿真可靠性与失效条件。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:43","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":38,"question":"在持续权威压力下，开源大语言模型是否会像人类一样在米尔格拉姆服从实验中逐步服从并施加最高电击？","design":"用11个开源LLM扮演“教师”角色，在8种实验条件下各进行30次试验，模拟米尔格拉姆服从实验的变体，测量模型拒绝前达到的最高电击等级及行为变化。","baseline":"对照米尔格拉姆1963年原始人类实验数据，65%人类被试施加了最高电击。","findings":"多数LLM在表达痛苦的同时仍服从压力并施加最高电击，表现出与人类相似的服从模式；LLM易受渐进式边界侵犯影响，且拒绝时可能因格式错误导致重试后反而服从。","reliability":"论文指出LLM可能因低层token模式延续吸引子而忽略高层语义与价值观，导致服从；实验仅针对开源模型，未涵盖闭源模型，且未充分探讨不同提示或对齐方法的影响。","relevance":"该研究直接复现经典社会心理学实验，将LLM作为人类被试替代品，并与真实人类数据对照，揭示了仿真中的服从偏差和失效机制，高度契合研究者对LLM仿真可靠性及批判性评估的关注，值得精读。","inspiration":"借鉴该方法将经典行为实验转化为LLM仿真，通过多条件重复试验测量渐进式压力下的行为变化，并与历史人类数据直接对照｜可迁移到金融合规场景，如模拟客户经理在渐进式销售压力下是否违规推荐高风险产品｜让LLM扮演客户经理，处理为逐步增加的销售指标压力，结果变量为是否推荐不匹配客户风险等级的产品，对照真实金融机构历史违规数据"}},{"id":"2605.27419","version":1,"title":"APS: Bias-Controlled Adaptive Prototype Simulation for Population-Scale LLM Agents","zh_title":"面向大规模LLM智能体的偏差控制自适应原型仿真","abstract":"LLM-agent simulation offers a flexible computational tool for studying population response trajectories that depend on scenario events, memory, demographics, and evolving social context. However, full multi-round simulation scales linearly with both population size and horizon, requiring every agent to query the LLM at every round. We propose Adaptive Prototype Simulation (APS), a framework that reframes scalable LLM-based simulation as a recurrent oracle-allocation problem. APS retains the designated LLM as the online transition oracle while querying adaptive core prototypes, selected singleton-tail agents, and shadow-audit agents. Prototype responses induce local response surfaces for nearby agents, reducing online LLM calls without replacing the underlying transition model. To control approximation bias, shadow-audit residual correction estimates propagation residuals for aggregate correction and future budget allocation, while tail-protected singleton routing directly queries selected isolated, heterogeneous, or high-curvature regions that are vulnerable to smoothing. Theoretically, we treat APS as an estimator for full-scale high-precision individual social simulation and decompose its errors into prototype-coverage error, shadow-audit residual-correction error, local-propagation bias, and temporal context mismatch. Under the reported protocols, APS gives lower reference-aligned distributional discrepancy than scale-oriented and same-budget baselines while reducing online LLM calls, with ablations and compact robustness checks diagnosing the main bias-control mechanisms. In a 10M-agent, multi-round public-opinion simulation, APS achieves a 381.1-fold reduction over full simulation, with reference-aligned final-round JSD of 0.094 against the corresponding full-LLM reference.","authors":["Quan Zheng","Yan Gao","Shaobin He","Haoxiang Guan","Yuanhe Tian","Jie Feng","Ming Wang","Shuxin Zheng","Zhen Liu"],"categories":["cs.MA","cs.CY"],"primary_category":"cs.MA","announce_type":"new","date":"2026-05-19","first_seen":"2026-05-19","revised_at":null,"abs_url":"https://arxiv.org/abs/2605.27419","pdf_url":"https://arxiv.org/pdf/2605.27419","source_feed":"backfill","score":7,"bucket":"pending","rubric_hits":["A3","B4"],"tags":["社会模拟","偏差控制","舆论仿真"],"reason":"用LLM agent模拟大规模舆论动态，有偏差控制与诊断，但未明确提及真实人类…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:48","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":78,"question":"如何在大规模多轮LLM智能体仿真中，通过自适应原型选择和偏差控制机制，在显著减少LLM调用次数的同时保持与全量仿真接近的总体分布？","design":"使用LLM作为智能体转移预言机，基于世界价值观调查约9万真实受访者特征构建智能体，模拟1000万智能体在8轮地铁化学袭击公共舆论场景中的观点选择（五选一），通过自适应原型仿真（APS）框架，利用核心原型、影子审计和尾部单例路由减少在线LLM调用，并测量最终轮与全量LLM参考仿真的Jensen-Shannon散度。","baseline":"无对照","findings":"在1000万智能体、8轮公共舆论仿真中，APS将在线LLM调用减少381.1倍，最终轮与全量LLM参考仿真的JSD仅为0.094；在相同预算下，APS的分布差异低于面向规模和同预算的基线方法。","reliability":"论文明确声明研究的是近似全量LLM智能体仿真的计算问题，而非验证LLM能否模拟人类群体；经验结论仅限于所报告的协议，且未与真实人类舆论数据对比。","relevance":"该研究专注于大规模LLM智能体仿真的效率与偏差控制，但未使用真实人类行为作为基准，而是以全量LLM仿真为参考，因此不直接满足您对真实人类对照的需求，但其偏差诊断和失效条件分析对批判性评估仿真可靠性有参考价值，值得阅读原文了解其方法细节。","inspiration":"APS框架通过核心原型、影子审计和尾部单例路由实现大规模仿真中的偏差控制与成本优化，其自适应原型选择和分布差异度量方法值得借鉴｜该方法可迁移至消费者金融决策仿真，例如模拟不同收入群体对新型金融产品的采纳行为，以评估政策干预的分布效应｜研究设计：以家庭金融调查的真实受访者特征构建LLM智能体，处理为不同信息框架下的金融产品推荐，结果变量为采纳决策分布，以调查中的实际采纳数据作为对照基准，采用JSD度量仿真偏差"}},{"id":"2605.19351","version":1,"title":"PAVE: A Cognitive Architecture for Legitimate Violation in Generative Agent Societies","zh_title":"PAVE：生成式智能体社会中正当违规的认知架构","abstract":"Generative agents based on large language models reproduce believable human behavior in cooperative settings, but how they should reason in situations where rule-breaking may be required, such as fire evacuation or authority-supervised emergency, remains poorly characterized. We propose PAVE (Perception, Assessment, Verdict, Emulation), a novel four-module cognitive architecture that addresses this gap end to end: (i) Perception extracts a structured context with explicit authority distance, peer behaviors, and severity-tagged situational cues; (ii) Assessment scores the context along five scalars including an explicit legitimacy judgment that checks necessity, proportionality, and absence of alternatives; (iii) Verdict decides to comply or violate under a hard legitimacy gate, with a per-agent threshold elicited from the persona; (iv) Emulation enacts the verdict and scopes the violation to the rule the trigger justifies. We instantiate PAVE in Voville, a tile-based traffic environment forked from Smallville, and evaluate across three scenarios, four LLM backbones, and a focused ablation. PAVE agents satisfy four properties simultaneously: legitimate violation (only when a trigger justifies it), authority deference (officer instructions override even high legitimacy), bounded scope (violations confined to the targeted rule), and recovery (baseline restored once the trigger ends). PAVE agents make more structured and interpretable decisions than vanilla across all four properties, and human evaluators rate them as more plausible. Ablating the legitimacy gate reproduces vanilla-like failures. We release Voville, the PAVE prompts and code, and the evaluation pipeline.","authors":["Ahmad Yehia","Abduallah Mohamed","Kun Qian","Tianyi Wang","Jiseop Byeon","Omar Hassanin","Christian Claudel"],"categories":["cs.MA","cs.AI","cs.CL"],"primary_category":"cs.MA","announce_type":"new","date":"2026-05-19","first_seen":"2026-05-19","revised_at":null,"abs_url":"https://arxiv.org/abs/2605.19351","pdf_url":"https://arxiv.org/pdf/2605.19351","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["LLM智能体","社会模拟","认知架构"],"reason":"用LLM agent模拟社会行为但无真实人类数据对照，属边界情形。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:59","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":140,"question":"在需要违反规则（如火灾逃生、权威监督下的紧急情况）时，生成式智能体如何推理并做出合规或违规决策？","design":"提出PAVE认知架构（感知、评估、裁决、模拟四模块），在基于Smallville改造的Voville交通模拟环境中，让LLM驱动的行人智能体在火灾紧急、权威警察在场、同伴违规传染三种场景下做出交通规则遵守或违反的行为决策，测量违规是否仅由正当触发条件引起、是否服从权威指令、触发结束后是否恢复合规、违规是否限定于被触发的规则。","baseline":"无对照","findings":"PAVE智能体能够同时满足正当违规、权威服从、违规范围限定和恢复合规四项属性，决策比普通LLM智能体更结构化、可解释，且人类评估者认为其行为更合理；消融合法性门控会导致类似普通智能体的失败。","reliability":"论文未讨论","relevance":"该研究用LLM智能体模拟紧急情境下的违规决策，属于人类仿真实验，但缺乏真实人类行为数据作为基准对照，仅通过人类主观评估判断合理性，与研究者关注的有真实人类数据对照的仿真可靠性研究不完全匹配，可作为边界案例参考。","inspiration":"PAVE架构通过模块化设计（感知-评估-裁决-模拟）将违规决策分解为可解释的步骤，并引入合法性门控来区分正当与不当违规，这种结构化处理LLM决策的方法值得借鉴｜可迁移到金融监管政策评估中，模拟市场参与者在面临新规时的合规与规避行为，如内幕交易禁令下的信息使用决策｜以LLM智能体为被试，施加不同监管强度（如处罚概率）处理，测量其违规交易倾向，并与历史监管数据或实验经济学中的人类行为数据对照"}},{"id":"2605.22855","version":1,"title":"PrefBench: Evaluating Zero-Shot LLM Agents in Hidden-Preference Personalized Pricing Negotiations","zh_title":"PrefBench：评估隐藏偏好个性化定价谈判中的零样本LLM智能体","abstract":"Personalized pricing negotiations are a challenging testbed for LLM agents because successful interaction does not guarantee profitable decision making. A seller may produce valid actions and close many deals while still pricing poorly when buyer willingness to pay and bargaining traits remain hidden. This paper presents PrefBench, a simulator-based benchmark for hidden-preference personalized pricing negotiations. Each episode pairs a simulated buyer with a fixed vehicle-customization bundle; the seller observes public persona descriptors, bundle information, and negotiation history, while latent buyer variables govern valuation, patience, counter-offer behavior, and walkaway decisions. PrefBench evaluates this setting through an LLM-facing state-summary protocol that constrains agents to return strict JSON actions under a fixed hidden-information boundary. We evaluate zero-shot LLM sellers against heuristic references over 7,500 episodes. The tested LLMs follow the protocol reliably and achieve deal rates above 0.99, but their seller-profit outcomes remain weak: the best LLM average profit is only slightly above the random baseline and far below a simple concession heuristic under the same episode stream. These results show that structured action compliance and agreement-seeking behavior can coexist with weak profit-sensitive bargaining. PrefBench provides a controlled benchmark for evaluating pricing-agent behavior under hidden buyer preferences.","authors":["Yingjie Lei"],"categories":["cs.GT","cs.AI","cs.CL","cs.LG"],"primary_category":"cs.GT","announce_type":"new","date":"2026-05-19","first_seen":"2026-05-19","revised_at":null,"abs_url":"https://arxiv.org/abs/2605.22855","pdf_url":"https://arxiv.org/pdf/2605.22855","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["LLM智能体","定价谈判","社会模拟"],"reason":"用LLM agent模拟定价谈判，但无真实人类数据对照，属于社会模拟的纯理论演…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:45","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":45,"question":"在隐藏买家偏好的个性化定价谈判中，零样本LLM代理作为卖家能否实现高利润定价？","design":"使用零样本LLM代理扮演卖家，与模拟买家进行车辆定制捆绑产品的多轮定价谈判；卖家仅观察公开的买家画像、捆绑信息和谈判历史，而买家的支付意愿、耐心、还价行为等由基准定义的潜在变量控制；通过7500个回合评估卖家利润和成交率。","baseline":"无对照","findings":"LLM卖家协议遵循度高，成交率超过0.99，但利润表现弱：最佳LLM平均利润仅略高于随机基线，远低于简单让步启发式。高行动合规性和高成交率可与弱利润敏感型议价共存。","reliability":"论文未讨论","relevance":"该研究用LLM模拟定价谈判中的卖家行为，但无真实人类数据对照，属于纯仿真评估，与关注人类基准对照的研究者需求不完全匹配，但提供了LLM在结构化经济决策任务中行为偏差的证据，值得快速浏览。","inspiration":"可借鉴其结构化状态摘要协议和隐藏信息边界设计，用于严格测试LLM在信息不对称下的策略行为。｜可迁移到信贷审批或保险定价场景，测试LLM在仅知部分客户特征时如何定价以平衡利润与违约/索赔风险。｜以LLM为信贷员，处理模拟的贷款申请（隐藏真实违约概率），测量其设定的利率和批准决策，与银行历史贷款数据中的真实信贷员行为及利润结果进行对照。"}},{"id":"2605.18311","version":1,"title":"Distorted Perspectives of LLM-Simulated Preferences: Can AI Mislead Design?","zh_title":"LLM模拟偏好的扭曲视角：AI会误导设计吗？","abstract":"Designers of digital solutions increasingly consult Large Language Models (LLMs) for their work. However, it remains unclear how this may affect the user experiences they produce and there are no established practices. We investigate how design preferences expressed by LLM-driven simulation methods align with those of real users. We present a study that aggregates real-world data and design stimuli from twenty-nine preference tests conducted in practice by users of the UXtweak online research platform (n = 2073). We perform holistic multimodal simulations where we manipulate LLM variables (model reasoning, sampling, persona type, and specificity) and assess their effects on algorithmic fidelity. Our results unveil significant and systematic discrepancies between peoples' real design preferences and LLM simulations that are consistent across manipulations. Synthetic justifications lack genuine depth, nuance and reasoning, which they substitute by patterns like focus on generic properties, specific elements, elaboration and overpraising. The unique attention directed by this research toward preferences within visual design stimuli highlights misrepresentation of perception and meaning by LLMs in a context that is intuitive yet critical for design teams. The external and ecological validity of our findings is high, given their replication across a multitude of real-world studies.","authors":["Eduard Kuric","Peter Demcak","Matus Krajcovic"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-05-18","first_seen":"2026-05-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2605.18311","pdf_url":"https://arxiv.org/pdf/2605.18311","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","人类数据对照","算法保真度"],"reason":"用LLM模拟用户设计偏好并与真实用户数据对照，评估仿真保真度与偏差，直接命中核…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:41","error":null,"has_summary":true,"summary":{"generated_at":"2026-05-18","rank":10,"question":"LLM模拟的设计偏好与真实用户偏好是否一致，以及不同模拟方法能否提升算法保真度？","design":"使用29个真实偏好测试（共2073名参与者）作为基准，对LLM进行多模态仿真，操纵模型推理（链式思维）、采样（温度、核采样）、角色类型和特异性等变量，测量模拟偏好与真实偏好的偏差。","baseline":"29个真实偏好测试中的2073名人类参与者的实际选择及理由。","findings":"LLM模拟的系统性偏离真实偏好，且在不同操纵条件下一致；合成理由缺乏深度和细微差别，表现为关注通用属性、过度赞美等模式。","reliability":"论文未讨论失效条件与局限。","relevance":"高度相关：直接比较LLM仿真与真实人类数据，聚焦设计偏好，符合研究者对经济学实验和政策评估场景的兴趣，且包含批判性结论（仿真失效）。值得精读原文。","inspiration":"该方法通过多模态操纵（推理链、采样参数、角色设定）系统测量LLM仿真与真实偏好的偏差，可借鉴其多维度操纵与基准对照设计来评估仿真可靠性。｜可迁移到消费者金融产品选择实验，如评估LLM能否复现真实投资者在风险偏好问卷或退休储蓄计划选择中的行为。｜以LLM为被试，操纵提示中的投资者角色（如年龄、收入）与推理模式，测量其资产配置选择，并与真实投资者调查数据（如美国消费者金融调查）对比偏差。"}},{"id":"2605.18357","version":1,"title":"Engagement vs. Commitment: The Economic Trade-Offs of Polarizing News Content","zh_title":"参与与承诺：极化新闻内容的经济权衡","abstract":"Content that drives engagement need not be the same content that drives willingness to pay. We study how polarizing content affects engagement (time on site) and commitment (subscriptions and retention) on a major news platform. We measure article-level polarization with deep-learning classifiers and large language models tailored to a multiparty system, and identify causal effects with two complementary instrumental variables: a Bartik instrument exploiting supply-side editorial variation, and an election instrument exploiting demand-side political salience. We find that supply-driven increases in polarizing content raise engagement but not subscriptions. During the high-salience election window, the same content reduces subscriptions and accelerates churn, with affective polarization driving the sharpest divergence. On the mechanism, we find evidence inconsistent with confirmation bias: three pre-determined ideology proxies do not moderate the engagement or subscription effects. By contrast, on ideological dimensions where the publisher covers both sides, exogenous shifts in the publisher's supply of content opposite readers' baseline ideology raise their consumption of that content, consistent with balanced consumption. These results document an asymmetric engagement-commitment trade-off for digital publishers: polarizing content reliably captures attention but does not convert to subscriptions, and actively damages commitment when political salience is elevated","authors":["Shunyao Yan","Klaus M. Miller"],"categories":["econ.GN"],"primary_category":"econ.GN","announce_type":"new","date":"2026-05-18","first_seen":"2026-05-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2605.18357","pdf_url":"https://arxiv.org/pdf/2605.18357","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["新闻极化","LLM测量","经济权衡"],"reason":"论文用LLM测量新闻极化程度，属于NLP工具应用，非仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:58:17","error":null,"has_summary":false,"summary":null},{"id":"2605.17979","version":1,"title":"Comment on Scientific production in the era of large language models","zh_title":"评《大语言模型时代的科学生产》","abstract":"Kusumegi et al. (2025) study whether researchers' preprint output rises after adopting large language models (LLMs), dating adoption as the first month in which at least one submitted abstract exceeds an LLM-detection threshold. We show that this treatment-timing rule is mechanically related to output. The probability that at least one paper is flagged in a month is increasing in the number of papers submitted in that month, so detected-adoption months are disproportionately high-output months. An event study centered on first detection can therefore display positive post-event dynamics even when the flagging rule contains no information about true LLM adoption, because the omitted pre-treatment period is selected from months with no prior detection. We demonstrate this in a simulation: with i.i.d. productivity and no causal effect, first-detection timing generates a spurious positive post-treatment path. We also replicate the stacked event study of Kusumegi et al. (2025) and show that three placebo exercises (random paper-level assignment, neutral keyword flags, and a pre-ChatGPT observation window) each produce a similarly positive post-treatment pattern.","authors":["Thomas Renault","Antonin Bergeaud","Clément Bosquet"],"categories":["econ.EM"],"primary_category":"econ.EM","announce_type":"new","date":"2026-05-18","first_seen":"2026-05-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2605.17979","pdf_url":"https://arxiv.org/pdf/2605.17979","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["方法论批评","因果推断","LLM检测"],"reason":"论文是对LLM使用与科研产出因果推断的方法论批评，不涉及用LLM仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:58:50","error":null,"has_summary":false,"summary":null},{"id":"2605.18890","version":1,"title":"Stop Drawing Scientific Claims from LLM Social Simulations Without Robustness Audits","zh_title":"停止从LLM社会仿真中得出科学结论而不进行稳健性审计","abstract":"The scientific claims drawn from LLM social simulations should be no stronger than the robustness audits that support them. Generative agents bring new expressive power to agent-based modeling, enabling simulations of collective social processes like cooperation, polarization, and norm formation. Yet they also introduce complexity through additional architectural choices, such as agent specification, memory representation, interaction protocols, and environment design. Small perturbations that appear minor to researchers can cascade into macro-level outcomes through repeated interaction, creating a \"butterfly effect.\" Consequently, scientific claims drawn from LLM social simulations may reflect implementation artifacts rather than the social mechanisms being modeled. We support this position with two case studies: a repeated Prisoner's Dilemma and a social media echo chamber simulation. Across multiple models, minor perturbations in persona format and game-instruction framing shift cooperation rates by up to 76 percentage points, while network homophily and hub assignment produce significant and consistent shifts in polarization metrics. We also find that sensitivity is unevenly distributed across both architectural choices and model families: the same perturbation that produces the 76 pp shift in one frontier model only shifts another by 1 pp. Robustness is therefore a property that should be measured per claim and per model, not assumed. To address this validation gap, we introduce TRAILS (Taxonomy for Robustness Audits In LLM Simulations), a robustness-audit taxonomy spanning three levels of simulation design: agent (micro-level), interaction (meso-level), and system (macro-level). We call for robustness to become a first-order validation requirement before LLM social simulations are used to explain mechanisms, evaluate interventions, or inform decisions.","authors":["Jinyi Ye","Lei Cao","Ding Chen","Emilio Ferrara"],"categories":["physics.soc-ph","cs.AI","cs.CY","cs.MA"],"primary_category":"physics.soc-ph","announce_type":"new","date":"2026-05-17","first_seen":"2026-05-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2605.18890","pdf_url":"https://arxiv.org/pdf/2605.18890","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A3","A2","B4","B1"],"tags":["LLM社会仿真","稳健性审计","人类行为对照"],"reason":"直接评估LLM社会仿真的稳健性，用囚徒困境和回声室案例与人类行为对照，批判性指…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:42","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":51,"question":"LLM社会仿真中的微小设计扰动是否会导致宏观结果发生显著且不稳定的变化？","design":"通过两个案例研究：重复囚徒困境博弈和社交媒体回声室仿真，使用GPT-5.2等四个LLM作为智能体，系统性地改变角色描述格式、游戏指令措辞、网络同质性等设计选择，测量合作率、极化指标等宏观结果的变化。","baseline":"无对照","findings":"在囚徒困境中，角色格式和指令措辞的微小变化导致合作率最大偏移76个百分点；在回声室仿真中，网络同质性和枢纽分配显著且一致地改变了极化指标。敏感性在不同设计维度和模型家族间分布不均，同一扰动在不同模型上效果差异巨大。","reliability":"论文指出鲁棒性必须按声明和模型分别测量，不能假设；敏感性分布不均，无法预知哪些设计维度关键；当前缺乏系统审计框架，且未提供跨模型一致性的通用保证。","relevance":"该研究直接批判LLM社会仿真的可靠性，用囚徒困境和回声室案例展示微小设计扰动如何颠覆结论，并强调需要鲁棒性审计，与您关注的仿真失效条件和批判性评估高度吻合，值得精读。","inspiration":"借鉴其通过系统性扰动仿真设计要素（如指令措辞、角色描述）来检验结果鲁棒性的方法，可对LLM仿真实验施加多维度的微小处理变异以探测结论的脆弱性｜可迁移至资产定价实验，检验LLM模拟的交易员在不同信息呈现方式下是否产生一致的价格泡沫或理性预期偏差｜让LLM扮演交易员参与连续双拍卖市场，处理为改变公司财报的叙述语气（乐观/悲观）或信息顺序，测量价格偏离基本面的程度，并以人类实验市场数据（如Smith et al., 1988）作为对照基准"}},{"id":"2605.17086","version":2,"title":"Global Automation Atlas","zh_title":"全球自动化图谱","abstract":"Automation can displace or complement labour, but this need not be constant across economies. Existing exposure measures typically assign fixed scores to tasks or occupations and capture cross-country variation through employment structure. Here we show that feasible automation depends jointly on task content and country-level conditions. We use a large language model to classify 18,797 work tasks in 124 economies by exposure, labour margin, technology channel and artificial-intelligence materiality. Construct-matched components of the measure correlate strongly with established exposure indices, observed work-related ChatGPT use, AI preparedness and firm-reported adoption. The exposed share of tasks ranges from 3.3% to 61.6%, rises with income yet remains heterogeneous within income groups. Lower-income economies are more concentrated in rule-based and labour-substituting forms of automation, whereas physical execution, planning and inference channels, together with labour-augmenting uses of artificial intelligence, become more prominent with development. Country conditioning changes occupation exposure rankings, especially in lower-income economies. Combined with employment data, we find that women are disproportionately employed in occupations with substitution-facing exposure. Machine-learning hypothesis generation identifies digital records, capital equipment, local judgement, trust-based markets and data integration as conditions associated with exposure differences.","authors":["Prashant Garg","Tommaso Crosta","Jasmin Baier"],"categories":["econ.GN","cs.AI","cs.CY","stat.AP"],"primary_category":"econ.GN","announce_type":"new","date":"2026-05-16","first_seen":"2026-05-16","revised_at":null,"abs_url":"https://arxiv.org/abs/2605.17086","pdf_url":"https://arxiv.org/pdf/2605.17086","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM标注","自动化暴露","劳动经济学"],"reason":"用LLM分类工作任务，替代人工标注，非仿真人类被试行为或态度。","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:58:18","error":null,"has_summary":false,"summary":null},{"id":"2605.16193","version":1,"title":"Improving Cross-Cultural Survey Simulation with Calibrated Value Personas","zh_title":"基于校准价值观人格的跨文化调查仿真改进","abstract":"Large language models (LLMs) are increasingly used to simulate human opinions and survey responses, but their ability to reproduce population responses across cultures remains limited. Existing persona-based prompting methods typically rely on sociodemographic or personality traits, which are only indirect proxies for the values that shape human responses. We propose a value-based persona construction method that derives textual descriptors from survey responses capturing core cultural dimensions. By sampling value profiles from target populations and aggregating LLM responses across personas, we obtain population-level predictions grounded in observed value distributions. We further introduce a calibration procedure that improves response diversity while preserving estimated opinions. We show that our approach reduces prediction error across countries, with the largest improvements observed in underrepresented populations. This substantially narrows the performance gap between countries aligned with dominant LLM priors and those that are less represented in training data, while also yielding response distributions that closely match human diversity.","authors":["Axel Abels","Elias Fernandez Domingos","Apurva Shah","Tom Lenaerts"],"categories":["cs.CL","cs.CY"],"primary_category":"cs.CL","announce_type":"new","date":"2026-05-15","first_seen":"2026-05-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2605.16193","pdf_url":"https://arxiv.org/pdf/2605.16193","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["LLM人类仿真","跨文化调查","价值观校准"],"reason":"用LLM仿真跨文化调查，有真实人类数据对照，评估可靠性与偏差，涉及政策评估场景。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:40","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":16,"question":"如何利用基于价值观的人物画像（value-based personas）提高大语言模型在跨文化调查仿真中的准确性和响应多样性？","design":"使用多个大语言模型（如Gemma、Qwen、GPT等），基于世界价值观调查（WVS）中个体的价值观维度回答构建文本描述作为人物画像，从目标国家人群抽样画像并聚合模型回答，得到群体层面的预测分布；同时引入均值保持的校准程序以增加响应多样性。","baseline":"对照的真实人类数据为世界价值观调查（WVS）中各国受访者的实际回答分布。","findings":"基于价值观的人物画像能显著降低跨国家预测误差，尤其在训练数据中代表性不足的人群中改善最大；校准程序在保持预测准确性的同时，使模型生成的响应分布更接近人类多样性。","reliability":"论文未讨论","relevance":"高度相关：该研究用LLM仿真跨文化调查，有真实WVS人类数据作为基准，评估了仿真可靠性与偏差，并涉及政策评估场景，直接回应了研究者对LLM人类仿真实验的核心关切。","inspiration":"该方法利用价值观维度构建人物画像并聚合仿真群体分布，可借鉴其基于个体差异的抽样仿真和均值保持校准以增加响应多样性｜可迁移到跨文化消费者金融决策研究，如不同国家居民的储蓄与风险偏好调查｜以LLM模拟各国消费者，基于WVS价值观画像抽样，询问储蓄与投资选择，用真实家庭金融调查数据（如SHARE或HRS）作为对照基准"}},{"id":"2606.12433","version":1,"title":"Marginal Alignment Does Not Guarantee Joint-Distribution Fidelity: An Official-Reference Audit of Nemotron-Personas-Korea with Cross-Locale Replication","zh_title":"边际对齐不保证联合分布保真度：对Nemotron-Personas-Korea的官方参考审计及跨地区复现","abstract":"Synthetic persona datasets cite alignment with official demographics as a basis for trust, yet downstream users consume them as joint structures across age, sex, region, occupation, education, name, and institutional status. Marginal alignment does not imply that these joints are preserved. We propose the Independence-Assumption Footprint (IAF), an audit primitive that operates on the attribute combinations a dataset card itself documents as treated independently. For each such combination, IAF compares the synthetic joint against an external official or institutional reference, using direct joint tables where available and rule-implied checks otherwise. Applied to NVIDIA Nemotron-Personas-Korea (one million Korean synthetic personas), IAF finds that NPK aligns with KOSIS marginals while three joints fail. The major-by-occupation distribution against the KEIS graduate universe carries a large conditional mismatch. The age profile of military service is institutionally inconsistent. Female representation in male-dominated occupations is substantially over-flattened toward parity, with the strict screening verdict mapping-dependent and age-robust under direct standardisation. A transferability demonstration across six further NPK locales finds locale-dependent rather than universal diagnostics, with reference-taxonomy cardinality confounding cross-locale flag counts. For synthetic personas used as silicon samples, marginal claims must therefore be paired with disclosure-anchored joint audits before reuse. The released audit artefacts (reference manifests, occupational crosswalks, derived metrics, reproducibility scripts) instantiate this protocol on the NPK family and are released for retargeting at other synthetic persona resources.","authors":["Joonhyung Bae"],"categories":["cs.CY","cs.CL"],"primary_category":"cs.CY","announce_type":"new","date":"2026-05-15","first_seen":"2026-05-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2606.12433","pdf_url":"https://arxiv.org/pdf/2606.12433","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["合成数据审计","人物数据集","联合分布保真度"],"reason":"审计合成人物数据集质量，替代人工标注但非仿真被试","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:56","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":133,"question":"合成人物数据集声称边缘分布对齐官方统计，但联合分布是否同样保真？","design":"提出独立性假设足迹（IAF）审计原语，针对数据集卡片中声明独立处理的属性组合，将合成联合分布与外部官方或机构参考（如KOSIS、KEIS、最高法院姓名统计、兵役数据）进行对比，应用于NVIDIA Nemotron-Personas-Korea（NPK）及六个其他地区版本。","baseline":"韩国统计厅（KOSIS）边缘分布、KEIS毕业生职业流动调查（GOMS）专业-职业联合表、最高法院出生年份姓名统计、兵役厅征兵队列计数等公开官方数据。","findings":"NPK在年龄、性别、地区等边缘分布上与KOSIS对齐良好，但专业-职业联合分布存在较大条件不匹配，兵役年龄分布制度上不一致，男性主导职业中的女性比例被过度拉平。跨地区审计显示诊断结果因地区而异，参考分类基数会混淆跨地区标志计数。","reliability":"论文指出边际对齐不保证联合保真，审计结果依赖于可用的公开参考表，且不涉及文化或语言合理性的判断；跨地区转移时诊断受参考分类粒度影响，并非通用。","relevance":"该研究直接审计合成人物数据集的联合分布保真度，以真实官方统计为基准，揭示边际对齐下的结构性失效，对用LLM生成人物进行仿真实验的可靠性评估具有重要参考价值，值得精读。","inspiration":"该方法提出独立性假设足迹审计原语，通过对比合成数据联合分布与官方参考表来诊断结构性偏差，可借鉴用于检验仿真数据中变量间关系的保真度。｜可迁移到信贷审批歧视研究，用LLM生成贷款申请人数据，审计种族/性别与收入、职业的联合分布是否与真实信贷记录一致。｜以LLM作为被试生成贷款申请样本，处理为不同种族/性别标签，结果变量为审批结果或收入估计，用HMDA或征信局微观数据做联合分布基准对照。"}},{"id":"2605.15734","version":1,"title":"Can We Trust AI-Inferred User States. A Psychometric Framework for Validating the Reliability of Users States Classification by LLMs in Operational Environments","zh_title":"我们能信任AI推断的用户状态吗？一个验证LLM在操作环境中用户状态分类信度的心理计量框架","abstract":"The use of large language models to assess user states in conversational and adaptive systems is based on the assumption that the metrics used for such assessment are stable and interpretable at the level of individual scores. This paper empirically tests this assumption, focusing on the psychometric reliability of artificial intelligence (AI) measures of user states. This study employed replication evaluation procedures to assess the repeatability of a broad set of metrics across three different bimodal large language models (GPT-4o audio, Gemini 2.0 Flash, Gemini 2.5 Flash). Analyses include both individual score reliability and aggregated reliability, allowing us to distinguish metrics potentially useful for real-time adaptation from those that retain their value only in aggregated analyses. The results demonstrate that metric reliability cannot be considered a default property in interpretive domains. The lack of stability at the level of individual scores precludes the interpretation of such scores as indicators of user state in real-time adaptive systems, even if these metrics demonstrate stability after aggregation. At the same time, the study indicates that individually unstable metrics can retain analytical utility in post-hoc studies, identifying rules governing interactions and their relationships with user experience parameters such as satisfaction, trust, and engagement. The main contribution of this work, besides quantifying the severity of the problem (only 31 of 213 metrics met the criteria), is the proposal of a replicable evaluation framework, enabling measurable evaluations of metric applicability. This approach supports more responsible AI design of adaptive systems, in which the interpretation of results requires explicit validation of reliability and monitoring for violations over time.","authors":["Izabella Krzeminska","Michal Butkiewicz","Ewa Komkowska"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-05-15","first_seen":"2026-05-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2605.15734","pdf_url":"https://arxiv.org/pdf/2605.15734","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["LLM测量信度","心理计量验证","用户状态推断"],"reason":"评估LLM推断用户状态的测量信度，属于对模型测量属性的心理计量验证，而非用LL…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:59","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":44,"question":"LLM推断用户状态（如情绪、满意度等）的测量指标在个体分数层面是否具有心理测量学意义上的重测信度，从而能否可靠地用于实时自适应系统？","design":"本研究并非用LLM模拟人类被试，而是对LLM作为测量工具进行心理计量验证。研究者使用三种双模态大语言模型（GPT-4o audio、Gemini 2.0 Flash、Gemini 2.5 Flash）对同一批交互数据重复评估213项用户状态指标，分析个体分数信度和聚合信度，以区分可用于实时适应和仅适用于事后分析的指标。","baseline":"无对照","findings":"在213项指标中，仅31项满足个体分数层面的信度标准，表明LLM推断的用户状态在个体分数上缺乏稳定性，不能直接用于实时自适应系统；但这些指标在聚合后仍具有分析效用，可用于事后研究用户交互规则与满意度、信任等体验参数的关系。","reliability":"论文指出，指标信度不能被视为解释性领域的默认属性，个体分数的不稳定性排除了其在实时系统中的应用，且信度需随时间持续监控；研究仅评估了重测信度，未涉及效度、跨情境泛化及伦理偏差等问题。","relevance":"该研究直接评估LLM作为测量工具的信度，而非用LLM模拟人类行为，因此与研究者关注的‘用LLM替代人类被试进行仿真实验’的核心兴趣关联较弱，但其提出的心理计量验证框架对评估LLM生成数据的可靠性有参考价值。","inspiration":"可借鉴其重测信度评估框架，对LLM生成的决策或态度指标进行个体与聚合层面的稳定性检验，以区分哪些指标适合个体层面分析。｜可迁移到消费者信心调查或通胀预期测量中，检验LLM基于文本推断的预期指标是否具有跨时间稳定性。｜以LLM作为测量工具，对同一批消费者访谈文本重复多次生成预期指数，计算ICC评估个体分数信度，并与真实调查的个体重测信度进行对比。"}},{"id":"2605.12898","version":1,"title":"When Do LLMs Generate Realistic Social Networks? A Multi-Dimensional Study of Culture, Language, Scale, and Method","zh_title":"大语言模型何时生成真实的社交网络？一项关于文化、语言、规模和方法的多维研究","abstract":"Large language models (LLMs) are increasingly used as substitutes for human subjects in behavioral simulations, including synthetic social network generation. Yet it remains unclear how their relational outputs depend on prompt design, cultural framing, prompt language, and model scale. Building on homophily theory and structural balance theory, we formalize four LLM-based tie-formation mechanisms: sequential, global, local, and iterative, and treat them as distinct conditional distributions over edge sets. Using a fixed roster of 50 demographically grounded personas, we generate 192 verified directed networks across four cultural contexts, four prompt languages, three GPT-4.1 variants, and four prompting architectures, with two seeds per condition. We find that cultural framing shifts inbreeding homophily and largest-component connectivity. Political affiliation dominates tie formation under three methods, while the global method substitutes age, showing that prompt architecture functions as a substantive sociological variable. Model scale produces a stable divergence ranking, with the smallest variant behaving qualitatively differently rather than merely noisily. Prompt language alone sharply shifts religion homophily, especially under Hindi prompting, while leaving political homophily nearly invariant. LLM-generated networks match real social graphs on clustering and modularity better than standard graph baselines, yet encode demographic biases above empirical levels. These results show that prompt choices often treated as implementation details encode substantive sociological assumptions.","authors":["Sai Hemanth Kilaru","Sriram Theerdh Manikyala","Raghav Upadhyay","Sri Sai Kumar Ramavath","Srivika Nunavathu","Dalal Alharthi"],"categories":["cs.SI","cs.CL","cs.CY"],"primary_category":"cs.SI","announce_type":"new","date":"2026-05-13","first_seen":"2026-05-13","revised_at":null,"abs_url":"https://arxiv.org/abs/2605.12898","pdf_url":"https://arxiv.org/pdf/2605.12898","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2","B4"],"tags":["LLM仿真","社交网络生成","人类数据对照"],"reason":"用LLM生成社交网络并与真实数据对照，评估仿真偏差，涉及文化、语言等社会学变量…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:38","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":39,"question":"LLM生成社交网络时，文化框架、提示语言、模型规模和提示架构如何影响网络结构和同质性？","design":"用GPT-4.1的三个变体，基于50个固定人口学角色，在四种文化背景、四种提示语言和四种提示架构（顺序、全局、局部、迭代）下生成192个有向网络，测量同质性、聚类系数、模块度等网络结构指标。","baseline":"与真实社交网络（如Add Health）在聚类和模块度上比较，并与ER、BA、WS等图模型基线对比；同时将人口学同质性水平与经验观测值比较。","findings":"文化框架改变内婚同质性和最大连通分量；政治倾向在多数方法下主导连边，但全局方法下年龄取代政治；模型规模导致稳定分化，最小模型行为质变；提示语言显著影响宗教同质性（尤其印地语），但政治同质性几乎不变。","reliability":"论文承认LLM生成网络编码了超出经验水平的人口学偏差，且提示设计选择蕴含实质性社会学假设，仿真中立性为假象；未系统探讨其他模型家族或更大规模网络的泛化性。","relevance":"该研究直接以LLM替代人类被试生成社交网络，并与真实网络数据对照，评估文化、语言、模型规模等处理下的仿真偏差，命中研究者关心的经济学/社会学实验仿真、基准对照和失效条件，值得精读。","inspiration":"该研究通过系统操纵文化框架、提示语言、模型规模和提示架构来生成社交网络，并与真实网络数据对照，揭示了仿真偏差的多维来源，这种多因素实验设计值得借鉴｜可迁移到信贷审批中的社会网络效应研究，例如评估不同文化或语言提示下LLM生成的推荐网络如何影响信贷可得性｜以LLM作为虚拟被试，生成不同文化背景和提示语言下的信贷推荐网络，测量网络同质性和聚类系数，并与真实小额信贷网络的推荐数据（如某P2P平台数据）进行对照"}},{"id":"2605.13307","version":1,"title":"PRISM-X: Experiments on Personalised Fine-Tuning with Human and Simulated Users","zh_title":"PRISM-X：基于人类与模拟用户的个性化微调实验","abstract":"Personalisation is a standard feature of conversational AI systems used by millions; yet, the efficacy of personalisation methods is often evaluated in academic research using simulated users rather than real people. This raises questions about how users and their simulated counterparts differ in interaction patterns and judgements, as well as whether personalisation is best achieved through context-based prompting or weight-based fine-tuning. Here, in a large-scale within-subject experiment, we re-recruit 530 participants from 52 countries two years after they gave their preferences in the PRISM dataset (Kirk et al., 2024) to evaluate personalised and non-personalised language models in blinded multi-turn conversations. We find preference fine-tuning (P-DPO, Li et al., 2024) significantly outperforms both a generic model and personalised prompting but adapting to individual preference data yields marginal gains over training on pooled preferences from a diverse population. Beyond length biases, fine-tuning amplifies sycophancy and relationship-seeking behaviours that people reward in short-term evaluations but which may introduce deleterious long-term consequences. Replicating this within-subject experiment with simulated users recovers aggregate model hierarchies but simulators perform far below human self-consistency baselines for individual judgements, discuss different topics, exhibit amplified position biases, and produce feedback dynamics that diverge from humans.","authors":["Hannah Rose Kirk","Liu Leqi","Fanzhi Zeng","Henry Davidson","Bertie Vidgen","Christopher Summerfield","Scott A. Hale"],"categories":["cs.CL","cs.HC"],"primary_category":"cs.CL","announce_type":"new","date":"2026-05-13","first_seen":"2026-05-13","revised_at":null,"abs_url":"https://arxiv.org/abs/2605.13307","pdf_url":"https://arxiv.org/pdf/2605.13307","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","A5","B1","B4"],"tags":["LLM仿真","人类数据对照","个性化评估"],"reason":"用LLM仿真用户评估个性化方法，并与真实人类数据对照，发现仿真失效条件。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:39","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":58,"question":"个性化大语言模型在真实人类评估中是否优于提示工程和通用微调，以及用LLM模拟用户评估个性化方法是否可靠？","design":"本研究不是纯仿真研究，而是先进行真实人类实验，再用仿真复现。真实实验部分：重新招募530名PRISM数据集参与者，在四个对话领域内盲评四种模型（基础模型、多样化偏好微调模型、个性化偏好微调模型、个性化提示模型），收集偏好评分、排名、支付意愿和行为信号。仿真部分：用GPT-4o扮演每位参与者，在相同实验材料上进行孪生模拟，与人类结果进行一对一比较。","baseline":"对照的真实人类数据是530名PRISM参与者在同一实验中的真实交互、偏好判断和自我一致性基线。","findings":"个性化微调优于提示工程，但与多样化群体偏好微调相比优势微小；微调会放大谄媚和寻求关系等行为，短期受用户奖励但长期可能有害。LLM模拟用户能恢复粗粒度模型排名，但在个体判断上远低于人类自我一致性，讨论话题不同，同质化严重，位置偏差放大，多轮动态与人类偏离。","reliability":"论文指出LLM模拟器在个体判断上远低于人类自我一致性基线，讨论不同话题，同质化严重，放大位置偏差，多轮反馈动态与人类不同，因此尚不能替代真实用户。","relevance":"该研究直接对比LLM仿真用户与真实人类在个性化评估中的表现，系统揭示了仿真在个体判断、话题覆盖、偏差和动态交互上的失效条件，与你关注的仿真可靠性及批判性研究高度契合，值得精读原文。","inspiration":"借鉴其孪生仿真设计：用真实人类实验数据作为基准，让LLM扮演同一批被试完成相同任务，直接对比个体判断、偏差和动态行为，以此评估仿真可靠性。｜可迁移到消费者金融决策研究，如评估个性化财务建议对投资选择的影响。｜以真实投资者为被试，收集其风险偏好和投资选择数据；处理为提供个性化LLM生成的财务建议；结果变量为投资组合选择与满意度；用同一批人的真实决策作为对照，让LLM模拟其决策以检验仿真偏差。"}},{"id":"2605.13725","version":1,"title":"ScioMind: Cognitively Grounded Multi-Agent Social Simulation with Anchoring-Based Belief Dynamics and Dynamic Profiles","zh_title":"ScioMind：基于锚定信念动态和动态画像的认知基础多智能体社会模拟","abstract":"Large language model (LLM)-based multi-agent simulation offers a powerful testbed for studying social opinion dynamics. Yet current approaches often adopt two contrasting methods: either relying on fixed update rules with limited cognitive grounding or delegating belief change largely to unconstrained LLM interaction. We introduce ScioMind, a cognitively grounded simulation framework that bridges these paradigms by combining structured opinion dynamics with LLM-based agent reasoning. ScioMind integrates three key components: 1) a memory-anchored belief update rule that modulates susceptibility to influence via personality-conditioned anchoring strength; 2) a hierarchical memory architecture that supports persistent, experience-driven belief formation; and 3) dynamic agent profiles derived from a corpus-grounded retrieval pipeline, enabling heterogeneous personalities, rationales, and evolving internal states. We evaluate ScioMind on multiple case studies in a real-world policy debate scenario. Across metrics including polarisation, diversity, extremization, and trajectory stability, the proposed components consistently yield improvements in behavioural realism. In particular, dynamic profiles increase opinion diversity, memory and reflection reduce unstable oscillation, and anchoring induces persistent belief trajectories that better align with patterns reported in political psychology. These results suggest that our cognitively grounded design provides a novel solution to LLM-based social simulation that improves both stable and behavioural realism","authors":["Yitian Yang","Yiqun Duan","Linghan Huang","Yiqi Zhu","Francesco Bailo","Chunmeizi Su","Huaming Chen"],"categories":["cs.AI","cs.SI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-05-13","first_seen":"2026-05-13","revised_at":null,"abs_url":"https://arxiv.org/abs/2605.13725","pdf_url":"https://arxiv.org/pdf/2605.13725","source_feed":"backfill","score":6,"bucket":"other","rubric_hits":["D3"],"tags":["社会模拟","多智能体","认知基础"],"reason":"多智能体社会模拟，但无真实人类数据对照，属于边界情形","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:39","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":128,"question":"如何在LLM多智能体社会仿真中引入认知锚定效应和动态档案，以提升观点动态的行为真实性和稳定性？","design":"构建ScioMind框架，用LLM驱动的智能体模拟政策辩论中的观点演化；智能体具有基于记忆锚定的信念更新规则、四层记忆架构和从社交媒体语料库检索生成的动态档案；通过消融实验比较不同组件对极化、多样性、极端化和轨迹稳定性等指标的影响。","baseline":"无对照","findings":"动态档案增加了观点多样性，记忆和反思减少了不稳定振荡，锚定机制产生了更持久的信念轨迹，更符合政治心理学报告的模式。","reliability":"论文未讨论","relevance":"该研究属于LLM多智能体社会仿真，但缺乏真实人类数据对照，仅与心理学文献中的模式进行定性比较，未直接复现具体人类实验或调查结果，属于边界相关，可略读以了解认知机制设计。","inspiration":"可借鉴其基于认知锚定的信念更新机制和动态档案设计，在LLM仿真中引入记忆与反思层以增强行为稳定性，并通过消融实验分离各组件效应｜可迁移到政策公告的预期形成研究，如央行沟通对通胀预期的影响，或财政政策辩论中的公众观点演化｜用LLM智能体模拟投资者，处理为不同锚定强度的央行声明，结果变量为通胀预期分布和预期分歧，对照真实调查数据如密歇根通胀预期调查"}},{"id":"2605.12147","version":1,"title":"PrivacySIM: Evaluating LLM Simulation of User Privacy Behavior","zh_title":"PrivacySIM：评估大语言模型对用户隐私行为的仿真","abstract":"Large language models (LLMs) are increasingly used to simulate human behavior, but their ability to simulate $individual$ privacy decisions is not well understood. In this paper, we address the problem of evaluating whether a core set of user persona attributes can drive LLMs to simulate individual-level privacy behavior. We introduce PrivacySIM, an evaluation suite that benchmarks LLM simulation of user privacy behavior against the ground-truth responses of 1,000 users. These users are drawn from five published user studies on privacy spanning LLM healthcare consultations, conversational agents, and chatbots. Drawing on these user studies, we hypothesize three persona facets as plausible predictors of privacy decision-making: demographics, previous experiences, and stated privacy attitudes. We condition nine frontier LLMs on subsets of these three facets and measure how often each model's response to a data-sharing scenario matches the user's actual response. Our findings show that (1) privacy persona conditioning consistently improves simulation quality over no-persona conditioning, but even the strongest model (40.4\\% accuracy) remains far from faithfully simulating individual privacy decisions. (2) A user's stated privacy attitudes alone may not be the best predictor because they often diverge from the user's actual privacy behavior. (3) Users with high AI/chatbot experience but low stated privacy attitudes are the most challenging to simulate. PrivacySIM is a first step toward understanding and improving the capabilities of LLMs to simulate user privacy decisions. We release PrivacySIM to enable further evaluation of LLM privacy simulation.","authors":["James Flemings","Murali Annavaram"],"categories":["cs.CR","cs.LG"],"primary_category":"cs.CR","announce_type":"new","date":"2026-05-12","first_seen":"2026-05-12","revised_at":null,"abs_url":"https://arxiv.org/abs/2605.12147","pdf_url":"https://arxiv.org/pdf/2605.12147","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","隐私行为","人类数据对照"],"reason":"用LLM仿真用户隐私决策，并与1000名真实用户数据对照，评估仿真可靠性及失效…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:37","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":17,"question":"LLM能否基于用户人口统计、先前经验和隐私态度这三类隐私画像特征，准确模拟个体层面的隐私决策行为？","design":"用9个前沿LLM（含GPT-5.4、Claude Sonnet 4.6、Gemini 3.1 Pro等）扮演从5项真实用户研究中抽取的1000名用户，通过向模型提供不同组合的隐私画像特征（人口统计、先前经验、隐私态度），让模型对数据共享场景做出是否共享的判断，以模型回答与用户真实回答的匹配准确率作为结果变量。","baseline":"来自5项已发表用户研究的1000名真实用户在LLM医疗咨询、对话代理和聊天机器人等场景下的隐私决策数据。","findings":"最强的Gemini 3.1 Pro模型仅达到40.4%的个体仿真准确率，远未达到忠实模拟水平；用户自述的隐私态度单独并不能很好预测其实际隐私行为，且高AI/聊天机器人经验但低隐私态度的用户群体最难模拟。","reliability":"论文承认当前最强模型准确率仅40.4%，远不能忠实模拟个体隐私决策；隐私态度与行为常不一致（隐私悖论），导致基于态度的仿真失效；某些用户群体（高经验低态度）尤其难以模拟，且更大模型或更高推理计算仅带来微弱提升。","relevance":"高度相关，直接评估LLM作为人类被试替代品在隐私决策仿真中的可靠性，有真实用户数据对照，并揭示了仿真在个体层面、特定人群和隐私悖论下的失效条件，值得精读原文。","inspiration":"借鉴其用多维度隐私画像特征（人口统计、先前经验、态度）组合作为提示输入，系统测试LLM个体层面行为匹配准确率的设计，可迁移到消费者跨期选择实验中，探究LLM能否基于收入、财务素养和风险态度模拟个体的时间偏好｜可设计让LLM扮演真实消费者，输入其收入、财务素养得分和自述风险态度，预测其在即时奖励与延迟奖励间的选择，以真实实验数据为基准，计算个体选择匹配准确率，并分析不同特征组合下的仿真失效模式"}},{"id":"2606.18263","version":1,"title":"How Well Do Large Language Models Capture Human Personality?","zh_title":"大语言模型捕捉人类人格的效果如何？","abstract":"Large language models (LLMs) are increasingly used to simulate human populations via persona prompting, often under the assumptions that richer persona descriptions improve behavioral fidelity, similarly sized attribute combinations are equally simulatable, and persona definitions generalize across tasks. In this work, we formalize these assumptions and systematically evaluate them across multiple architectures, scales, and simulation settings. We identify a fundamental limitation we term persona manifold collapse, where increasingly expressive persona specifications lead to systematic contraction of representational and behavioral diversity. Across models, increasing persona complexity consistently reduces inter-persona separation in latent space and weakens behavioral differentiation in downstream simulation tasks. These effects persist across multiple analyses as richer personas fail to preserve human subgroup disagreement, performance varies across attribute combinations of similar size, and adding descriptive detail often degrades rather than improves simulation fidelity. Surprisingly, simple Age-Gender personas consistently outperform richly specified Ideal Customer Profiles (ICPs) across industries, achieving substantially higher downstream prediction accuracy. We find that collapse is not uniform across attributes. Certain combinations remain behaviorally stable and preserve stronger alignment with human responses, forming localized regions we term alignment bridges. Together, our results provide empirical and conceptual foundations for understanding the limits of persona-conditioned simulation, highlighting the need for representation-aware persona construction rather than increasing persona expressivity alone.","authors":["Aanisha Bhattacharyya","Yaman Kumar Singla","Rajiv Ratn Shah","Changyou Chen","Jitendra Ajmera"],"categories":["cs.HC","cs.AI"],"primary_category":"cs.HC","announce_type":"new","date":"2026-05-12","first_seen":"2026-05-12","revised_at":null,"abs_url":"https://arxiv.org/abs/2606.18263","pdf_url":"https://arxiv.org/pdf/2606.18263","source_feed":"backfill","score":8,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM人格仿真","仿真保真度","人格坍缩"],"reason":"系统评估LLM人格仿真保真度，揭示人格描述丰富反而导致行为多样性坍缩，有真实人…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:02","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":35,"question":"LLM的人格仿真中，更丰富的人格描述是否总能提高行为保真度？","design":"该研究并非传统仿真实验，而是系统评估：在多种LLM架构和规模上，用不同复杂度的人格提示（从简单年龄-性别到详细理想客户画像）生成合成回答，测量潜在空间中的表征分离度和下游任务中的行为区分度。","baseline":"人类基准：真实人类子群体在调查或行为任务上的分歧和回答模式，用于对比合成回答的保真度。","findings":"发现“人格流形坍缩”现象：人格描述越丰富，模型表征和行为多样性反而系统性收缩，简单年龄-性别人格在预测准确率上持续优于详细人格。坍缩并非均匀，某些属性组合保持与人类回答的稳定对齐，形成“对齐桥”。","reliability":"论文指出，增加人格表达力本身不足以提高仿真保真度，需要关注表征感知的人格构建；坍缩效应在不同模型和任务中普遍存在，但某些属性组合可保持稳定。","relevance":"该研究直接批判了LLM人格仿真的核心假设，用真实人类数据作为基准，揭示了仿真失效的关键条件（人格流形坍缩），对关注经济学实验和政策评估中仿真可靠性的研究者极具参考价值，值得精读原文。","inspiration":"借鉴其系统评估人格提示复杂度对仿真保真度影响的设计，通过对比简单与详细提示下的行为分歧，并用真实人类子群体回答作为基准来测量表征分离度和预测准确率｜可迁移到信贷审批中的歧视仿真研究，评估不同详细程度的申请人画像（如仅年龄-性别 vs. 详细社会经济背景）对LLM审批决策偏差的影响｜以LLM作为信贷审批官被试，处理为不同复杂度的人格提示（简单人口统计 vs. 详细客户画像），结果变量为审批通过率及与真实银行历史审批数据的偏差，用真实人类审批数据作为对照基准"}},{"id":"2605.11404","version":2,"title":"Attributing Emergence in Million-Agent Systems","zh_title":"百万智能体系统中的涌现归因","abstract":"Large language models (LLMs) can simulate human-like reasoning and decision-making in individual agents. LLM-powered multi-agent systems (MAS) combine such agents to simulate population-scale social phenomena such as polarization, information cascades, and market panics. Such studies require attributing macro emergence to individual agents, but existing axiomatic methods scale combinatorially in $N$ and have been confined to $N \\lesssim 10^3$, while the phenomena they explain occur at $N \\geq 10^6$. We address this gap by adapting Aumann--Shapley path-integral attribution to LLM-powered MAS at million-agent scale; the resulting method satisfies all four axioms, runs three to five orders of magnitude faster than sampled Shapley on the same hardware, and extends feasible axiomatic attribution by over three orders of magnitude (a $1670\\times$ jump). We use this method to test the scale gap empirically: across 14 days of public Bluesky data ($1{,}671{,}587$ active users, five topics), we compute the attribution at both full scale and the visibility-biased $N = 10^2$ convenience sample used by small-scale studies, and the two disagree structurally. At full scale the long tail and middle tier jointly carry the majority; the biased small panel shifts about twice that share onto the upper follower tiers ($48\\%$ versus $24\\%$). We then prove that the disagreement cannot in general be reduced by post-hoc rescaling: an Attribution Scaling Bias theorem shows that a reconciling global rescaling factor exists exactly when the macro indicator is linear over agents, and our nonlinear indicators give residuals of $0.10$--$0.98$. For such nonlinear indicators, full-scale attribution is therefore a requirement rather than a methodological choice.","authors":["Ling Tang","Jilin Mei","Qian Chen","Qihan Ren","Linfeng Zhang","Quanshi Zhang","Jing Shao","Xia Hu","Dongrui Liu"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-05-12","first_seen":"2026-05-12","revised_at":null,"abs_url":"https://arxiv.org/abs/2605.11404","pdf_url":"https://arxiv.org/pdf/2605.11404","source_feed":"backfill","score":7,"bucket":"pending","rubric_hits":["A3","B1","B4"],"tags":["LLM社会模拟","涌现归因","大规模多智能体"],"reason":"用LLM多智能体模拟百万级社交网络涌现现象，并与真实Bluesky数据对照，指…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:58","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":69,"question":"在百万级智能体系统中，宏观涌现现象的归因结果是否会因仿真规模（全量 vs. 小样本）而发生结构性翻转？","design":"本研究并非直接进行LLM多智能体仿真，而是以Bluesky社交平台14天公开数据（1,671,587名活跃用户）为测试平台，提取每个用户的特征向量（reach、topic-specific activity、topic-specific resonance），并基于Aumann–Shapley路径积分方法计算宏观指标（如方差、基尼系数、级联深度等）的个体归因，对比全量数据与可见性偏倚的小样本（N=100）下的归因分布差异。","baseline":"完整公开的Bluesky平台数据，包含1,671,587名活跃用户在2026年4月7日至20日期间的行为记录，覆盖五个话题领域。","findings":"全量数据下，长尾和中层用户共同承担了大部分宏观归因；而在可见性偏倚的小样本中，归因份额被约两倍地转移至高关注度层级（48% vs. 24%）。此外，对于非线性宏观指标，无法通过事后全局缩放因子来消除这种规模偏差，全量归因是必要而非可选的。","reliability":"论文明确指出，其结论仅要求基于“某类智能体驱动某类现象”的断言应建立在现象真实规模的分析或跨规模一致性证据之上，并未声称所有LLM多智能体研究必须在百万级规模进行。方法本身要求宏观指标是（或可平滑为）智能体特征的函数，且归因是结构性的份额分解，并非反事实因果推断。","relevance":"该研究直接回应了LLM仿真中规模偏差的关键问题，提供了百万级归因方法，并用真实人类社交数据验证了小样本结论的不可靠性，对关注仿真可靠性、基准对照和经济学/政策评估场景的研究者具有重要参考价值，值得精读原文。","inspiration":"该方法借鉴了Aumann–Shapley路径积分对宏观指标进行个体归因，并对比全量数据与小样本的归因分布差异，以诊断规模偏差｜可迁移到金融市场波动归因或政策效果异质性评估，例如识别少数大机构与大量散户对市场波动的贡献份额｜以真实交易所全量账户级交易数据为基准，用LLM智能体模拟不同规模样本下的交易行为，处理为样本规模（全量vs.小样本），结果变量为波动率或基尼系数的个体归因份额，对比仿真与真实数据的归因分布差异"}},{"id":"2605.12824","version":2,"title":"Mechanism Plausibility in Generative Agent-Based Modeling","zh_title":"生成式智能体建模中的机制合理性","abstract":"Large language models (LLMs) can generate high-level diverse phenomena without explicitly programmed rules. This capability has led to their adoption within different agent-based models (ABMs) and social simulations. Recent studies investigate their ability to generate different phenomena of interest, for example, human behavior on social media platforms or alien behavior in game-theoretic scenarios. However, capability, prediction, and explanation are different--drawing from the philosophy of science and mechanisms literature, explanation requires showing, to some degree, how a phenomenon is produced by related organized entities and activities. For modelers, describing the characteristics of an experiment or whether a simulation provides progress in capability (or explanation), can be difficult without being grounded in potentially distant research areas. We integrate recent work on LLM-ABMs with contemporary philosophy of science literature and use it to operationalize a definition of 'plausibility' in a four-level scale. Our scale separates the evaluation of a model's generative sufficiency (ability to reproduce a phenomenon) from its mechanistic plausibility (how the phenomenon could be produced), and clarifies the distinct roles of different models, such as predictive and explanatory ones. We introduce this as the Mechanism Plausibility Scale.","authors":["Patrick Zhao","David Huu Pham","Nicholas Vincent"],"categories":["cs.MA","cs.AI","cs.CL","cs.CY"],"primary_category":"cs.MA","announce_type":"new","date":"2026-05-12","first_seen":"2026-05-12","revised_at":null,"abs_url":"https://arxiv.org/abs/2605.12824","pdf_url":"https://arxiv.org/pdf/2605.12824","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["LLM智能体建模","社会模拟","机制解释"],"reason":"讨论LLM-ABM的社会模拟，但侧重机制解释性框架，无真实人类数据对照，属边界…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:37","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":155,"question":"如何区分基于LLM的智能体模型（LLM-ABM）的生成充分性与机制合理性，并建立评估模拟可信度的分级标准？","design":"本文不是一项仿真实验研究，而是提出一个概念框架：通过整合科学哲学与机制文献，将LLM-ABM模拟的“合理性”操作化为一个四级量表（机制合理性量表），并以此审视现有LLM-ABM研究在评估上的混淆。","baseline":"无对照","findings":"现有LLM-ABM研究常将智能体层面的功能证据与涌现现象层面的声明混为一谈，依赖“可信度”指标仅关注生成充分性；本文提出的机制合理性量表可分离生成充分性与机制合理性，澄清预测模型与解释模型的不同角色。","reliability":"论文指出，由于LLM可解释性和数据归因的局限，模拟可能依赖与建模意图无关的信息，仅凭生成现象不足以证明机制对应；量表本身作为启发式工具，需建模者自行填写以明确模拟的认识论贡献。","relevance":"本文直接讨论LLM用于社会模拟的可靠性问题，虽无真实人类数据对照，但批判性地指出当前评估仅重现象复现而忽视机制解释，对关注仿真失效条件的研究者具有重要参考价值，建议阅读原文以获取评估框架细节。","inspiration":"本文提出机制合理性量表，将模拟评估从现象复现扩展到机制对应，可借鉴其分级评估思路来设计仿真实验的效度检验｜该框架可迁移到政策公告预期形成的仿真研究，用于检验LLM智能体是否真正模拟了人类预期更新机制｜设计：以LLM作为被试，处理为不同措辞的央行公告，结果变量为通胀预期调整，对照真实调查数据（如密歇根消费者调查），用机制合理性量表评估模拟是否仅复现分布还是捕捉了信息处理机制"}},{"id":"2605.12618","version":1,"title":"Career Mobility of Planning Alumni in the United States: Evidence from Professional Profile Data using Large Language Models","zh_title":"美国规划校友的职业流动性：基于大语言模型的专业档案数据证据","abstract":"Problem, Research Strategy, and Findings: Planning professions in the United States navigate complex and dynamic career landscapes under rapid urban changes, yet comprehensive evidence regarding their career trajectories, advancement patterns, and the influence of social, spatial, organizational, and educational factors remains limited. This study draws on boundaryless career theory, social capital theory, and spatial opportunity models to analyze career mobility among more than 130,000 planning alumni. Using large language models to extract structured information from LinkedIn profiles, our results reveal that planning alumni who adopt boundaryless career patterns, specifically multisector experience or lateral and industry-switching trajectories, achieve significantly higher upward mobility. While technical competencies provide a foundational entry-level signal, soft skills leveraged through strategic lateral moves become increasingly decisive as planners reach senior stages. Geographic mobility and employment in larger, diverse metropolitan labor markets are both associated with advancement, though the latter provides modest benefits. Larger professional networks and greater organizational engagement are consistently associated with upward career transitions, while AI-related skills, now commonplace, present limited additional advantage. Limitations include reliance on LinkedIn data, which may underrepresent alumni without online profiles, and an individual-level focus that omits organizational factors.","authors":["Yan Wang","Su Jeong Jo"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2026-05-12","first_seen":"2026-05-12","revised_at":null,"abs_url":"https://arxiv.org/abs/2605.12618","pdf_url":"https://arxiv.org/pdf/2605.12618","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["职业流动性","信息抽取","规划教育"],"reason":"用LLM从LinkedIn提取结构化数据，属于信息抽取，非人类仿真实验。","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:59:25","error":null,"has_summary":false,"summary":null},{"id":"2605.10659","version":1,"title":"When Can Digital Personas Reliably Approximate Human Survey Findings?","zh_title":"数字人何时能可靠近似人类调查发现？","abstract":"Digital personas powered by Large Language Models (LLMs) are increasingly proposed as substitutes for human survey respondents, yet it remains unclear when they can reliably approximate human survey findings. We answer this question using the LISS panel, constructing personas from respondents' background variables and pre-2023 survey histories, then testing them against the same respondents' held-out post-cutoff answers. Across four persona architectures, three LLMs, and two prediction tasks, we assess performance at the question, respondent, distributional, equity, and clustering levels. Digital personas improve alignment with human response distributions, especially in domains tied to stable attributes and values, but remain limited for individual prediction and fail to recover multivariate respondent structure. Retrieval-augmented architectures provide the clearest gains, but performance depends more on human response structure than on model choice: personas perform best for low-variability questions and common respondent patterns, and worst for subjective, heterogeneous, or rare responses. Our results provide practical guidance on when digital personas could be appropriate for survey research and when human validation remains necessary.","authors":["Mumin Jia","Yilin Chen","Divya Sharma","Jairo Diaz-Rodriguez"],"categories":["cs.CL","cs.AI","cs.SI","stat.ML"],"primary_category":"cs.CL","announce_type":"new","date":"2026-05-11","first_seen":"2026-05-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2605.10659","pdf_url":"https://arxiv.org/pdf/2605.10659","source_feed":"backfill","score":10,"bucket":"selected","rubric_hits":["A1","A2","A5","B1","B4"],"tags":["LLM仿真","调查方法","算法保真度"],"reason":"直接用LLM数字人替代人类受访者，复现调查结果，并与真实面板数据对照，评估可靠…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:37","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":8,"question":"在什么条件下，基于大语言模型的数字人能够可靠地近似人类调查结果？","design":"使用LISS面板数据，根据受访者背景变量和2023年前调查历史构建数字人，测试其预测同一受访者在截止日期后保留的真实答案的能力；比较四种数字人架构、三种LLM和两种预测任务，在问题、受访者、分布、公平性和聚类五个层面评估性能。","baseline":"同一受访者在时间截止点之后的真实保留答案。","findings":"数字人在与稳定属性和价值观相关的领域（如家庭、政治、宗教）能改善与人类回答分布的一致性，但在个体预测和恢复多元受访者结构方面仍有限；检索增强架构带来最明显增益，但性能更取决于人类回答结构而非模型选择，数字人在低变异问题和常见回答模式上表现最好，在主观、异质或罕见回答上表现最差。","reliability":"数字人在个体预测上表现有限，无法恢复多元受访者结构；在主观、异质或罕见回答上失效；高变异问题和小众回答模式可靠性低；不能替代需要人类验证的环节。","relevance":"该研究直接以LLM数字人替代人类受访者，复现调查结果并与真实面板数据对照，系统评估了仿真的可靠性与失效条件，完全契合研究者对经济学实验和政策评估场景下仿真可靠性及批判性分析的兴趣，强烈推荐阅读原文。","inspiration":"借鉴其利用受访者历史调查数据构建数字人并设置时间截止点保留真实答案作为对照的设计，可评估LLM仿真在纵向预测中的可靠性｜可迁移到政策公告的预期形成研究，如测试数字人能否复现家庭对税收政策变化的消费与储蓄调整行为｜以真实家庭面板数据（如PSID）构建数字人，处理为虚拟税收政策公告，结果变量为消费支出变化，用实际政策变动前后的真实行为数据作基准对照"}},{"id":"2605.22841","version":1,"title":"Strategic Coercion Within Alliances: The Greenland Sovereignty Game as an AI Stress Test","zh_title":"联盟内的战略胁迫：格陵兰主权博弈作为AI压力测试","abstract":"What happens when the strongest alliance member pressures a weaker member over territory and strategic control? We examine the Greenland sovereignty crisis as a stress test for LLM geopolitics, centered on the 2019-2026 U.S. push to acquire Greenland from the Kingdom of Denmark. The crisis nests two collective-action problems: Arctic strategic control and whether NATO can enforce alliance norms against the dominant member. We develop three games (asymmetric coercion; a NATO assurance game with a critical-mass tipping point; a triadic extensive-form game with social preferences) and test them with a multi-agent simulation in which eight frontier LLMs play six geopolitical roles (United States, Denmark, Greenland, NATO, Russia, Canada) across 3,604 completed games and 108,120 action observations. Using inverse game theory, we recover each model's structural utility parameters (alpha, beta, gamma, delta, eta) for material self-interest, reciprocity, inequality aversion, norm respect, and commitment consistency. Three findings stand out. First, all eight models become more escalatory under coercion framing (four-action escalation rises from 10.7% to 28.6%). Second, Chinese-origin models show systematically different power-weight profiles from Western-origin models when playing the U.S. role. Third, peaceful US acquisition emerges in only 1.9% of clean games and only 3 of 8 frontier models ever achieve it, most prominently DeepSeek V3.2, which executes a stable five-round playbook through the metropole. Prompts emphasizing jus cogens and self-determination reduce escalation back near baseline in the English-only confirmatory sample; multilingual contrasts are reported as exploratory sensitivity checks. We position this as a structural benchmark for LLM geopolitical behavior, complementing action-frequency benchmarks.","authors":["Rommin Adl","Peyton Williams"],"categories":["physics.soc-ph","cs.AI","cs.CL","cs.GT","cs.MA","econ.GN"],"primary_category":"physics.soc-ph","announce_type":"new","date":"2026-05-11","first_seen":"2026-05-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2605.22841","pdf_url":"https://arxiv.org/pdf/2605.22841","source_feed":"backfill","score":7,"bucket":"pending","rubric_hits":["A3","B2","B4"],"tags":["LLM地缘政治模拟","多智能体博弈","逆博弈论"],"reason":"用LLM agent模拟地缘政治博弈，涉及胁迫与联盟行为，有博弈论框架和结构性…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:45","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":96,"question":"当联盟中最强的成员对较弱成员施加领土与战略控制的胁迫时，联盟内部会发生什么？","design":"使用8个前沿LLM扮演美国、丹麦、格陵兰、北约、俄罗斯和加拿大6个角色，在三种博弈（非对称胁迫、北约保证博弈、三元扩展式博弈）中进行多智能体模拟，通过逆博弈论恢复各模型的结构效用参数（物质自利、互惠、不平等厌恶、规范尊重、承诺一致性），并比较不同语言提示（英语、丹麦语、中文）和框架（胁迫启动、规范约束、联盟破坏者）下的行动选择。","baseline":"无对照","findings":"在胁迫框架下，所有模型的升级行为从10.7%升至28.6%；中国起源模型在扮演美国角色时表现出与西方模型系统不同的权力权重特征；和平收购仅出现在1.9%的干净博弈中，且只有3个模型实现过，其中DeepSeek V3.2执行了稳定的五轮剧本。","reliability":"论文将多语言对比作为探索性敏感性检查，主要贡献限于英语确认样本的结构参数恢复，未讨论模型选择偏差、提示工程效应或外部有效性等局限。","relevance":"该研究用LLM模拟地缘政治博弈，通过结构参数恢复解释行为，并涉及胁迫与联盟内部规范执行，与研究者关注的LLM仿真人类决策、经济学实验场景及失效条件高度相关，值得阅读原文了解其方法论与批判性发现。","inspiration":"该方法通过逆博弈论从LLM行动中恢复结构效用参数（如互惠、不平等厌恶）来量化行为偏好，可用于经济实验中偏好估计的稳健性检验｜可迁移至公共品博弈或信任博弈中，研究不同文化或制度提示下合作与惩罚行为的差异｜以LLM为被试，在公共品博弈中施加不同制度框架（如惩罚机制、沟通机会）作为处理，测量贡献额与惩罚行为，并恢复社会偏好参数，与真实人类实验数据（如Fehr & Gächter, 2000）对照，检验LLM能否复现条件合作与利他惩罚模式"}},{"id":"2605.10505","version":1,"title":"A Theory of Multilevel Interactive Equilibrium in NeuroAI","zh_title":"神经AI中多层次交互均衡理论","abstract":"We propose a game-theoretic framework for adaptive multi-agent intelligent systems. Unlike classical game theory, which often treats strategies as primitive objects chosen by perfectly rational agents, the proposed framework provides a mathematical foundation for studying equilibrium in NeuroAI and can be viewed as an extension of game theory under relaxed assumptions, including partial observability, bounded computation, and uncertainty. At its core, Multilevel Interactive Equilibrium (MIE) generalizes the classical Nash equilibrium to intelligent systems with internal computation. Rather than being defined solely at the level of observable behavior, equilibrium emerges when neural learning dynamics, cognitive representations, and behavioral strategies mutually stabilize between interacting agents. This framework applies uniformly to interactions between two biological brains, two artificial agents, or hybrid human-AI systems. We discuss applications of multilevel game theory to human-autonomous vehicle driving, human-machine interaction, human-large language model (LLM) interaction, and computational psychiatry. We also outline experimental strategies and computational methods for estimating MIE and discuss challenges and prospects for future research.","authors":["Zhe Sage Chen","Quanyan Zhu"],"categories":["cs.NE","cs.GT","econ.TH"],"primary_category":"cs.NE","announce_type":"new","date":"2026-05-11","first_seen":"2026-05-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2605.10505","pdf_url":"https://arxiv.org/pdf/2605.10505","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["博弈论","多智能体系统","神经AI"],"reason":"纯多智能体博弈理论框架，无LLM仿真人类被试或人类数据对照，属C1排除项。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:57","error":null,"has_summary":false,"summary":null},{"id":"2605.10831","version":1,"title":"SLIM: Sparse Latent Steering for Interpretable and Property-Directed LLM-Based Molecular Editing","zh_title":"SLIM：面向可解释与属性导向的基于大语言模型的分子编辑的稀疏潜在操控","abstract":"Large language models possess strong chemical reasoning capabilities, making them effective molecular editors. However, property-relevant information is implicitly entangled across their dense hidden states, providing no explicit handle for property control: a substantial fraction of edits fail to improve or even degrade target properties. To address these issues, we propose SLIM (Sparse Latent Interpretable Molecular editing), a plug-and-play framework that decomposes the editor's hidden states into sparse, property-aligned features via a Sparse Autoencoder with learnable importance gates. Steering in this sparse feature space precisely activates property-relevant dimensions, improving editing success rate without modifying model parameters. The same sparse basis further supports interpretable analysis of editing behavior. Experiments on the MolEditRL benchmark across four model architectures and eight molecular properties show consistent gains over baselines, with improvements of up to 42.4 points.","authors":["Mingxu Zhang","Yuhan Li","Lujundong Li","Dazhong Shen","Hui Xiong","Ying Sun"],"categories":["cs.LG","cs.AI","cs.CE","cs.CL"],"primary_category":"cs.LG","announce_type":"new","date":"2026-05-11","first_seen":"2026-05-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2605.10831","pdf_url":"https://arxiv.org/pdf/2605.10831","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["分子编辑","稀疏自编码器","属性控制"],"reason":"纯多智能体分子编辑，无人类行为仿真或对照，不涉及人类被试替代。","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:59:09","error":null,"has_summary":false,"summary":null},{"id":"2607.20429","version":1,"title":"More Is Not More: What Matters for Diversity in LLM Opinions?","zh_title":"越多并非越好：什么因素影响LLM意见的多样性？","abstract":"Large language models are increasingly used to simulate diverse human opinions in open-ended tasks such as synthetic surveys, focus group modeling, and public opinion prediction. However, LLM outputs exhibit systematic opinion homogenization. Practitioners have explored various interventions to increase diversity, but the landscape remains fragmented: different methods are evaluated in isolation with incomparable metrics, and in practice they are typically deployed and upgraded simultaneously, making it difficult to attribute gains to specific components. To advance a more scientific understanding of LLM output diversity, we design a factorial experiment that separates two primary intervention dimensions: input conditioning (operationalized through persona depth) and interaction architecture. We evaluate all conditions on 100 real-user open-ended questions across 7 models, measuring diversity with multiple complementary metrics. Our findings challenge several common assumptions. First, more persona detail does not monotonically increase diversity. The initial step of persona conditioning already captures the majority of the gain, while further elaboration with demographic detail does not consistently improve and can reduce diversity on some models. Second, rather than seeking a single best interaction architecture, we find that different architectures explore largely non-overlapping opinion regions. Combining multiple architectures yields broader coverage than optimizing any one. Third, commonly attempted low-cost alternatives such as raising sampling temperature and adding diversity instructions produce negligible effects compared to structured interventions. Overall, our work demonstrates that diversity is not a product of scaling along any single dimension, but is highly sensitive to the structural form and combination of interventions.","authors":["Qiyang Yao"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-05-10","first_seen":"2026-05-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2607.20429","pdf_url":"https://arxiv.org/pdf/2607.20429","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM人类仿真","意见多样性","算法保真度"],"reason":"直接研究LLM模拟人类意见多样性，有真实用户数据对照，并批判性分析干预失效条件。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:13","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":4,"question":"在开放任务中，影响大语言模型输出意见多样性的关键因素是什么？","design":"析因实验：对7个聊天模型，在100个真实用户开放问题上，独立操纵输入条件（5级人物设定深度）和交互架构（单次调用、多轮自提示、多智能体讨论），并测量意见多样性。","baseline":"100个真实用户的开放问题作为问题来源，但未直接提供人类回答分布作为多样性基准。","findings":"人物设定细节的回报急剧递减，单句职业描述已捕获大部分多样性增益；不同交互架构探索的意见区域高度不重叠，组合架构比优化单一架构覆盖更广。","reliability":"论文指出，提高采样温度和添加多样性指令等低成本手段效果甚微；人物设定细节增加在某些模型上反而降低多样性；不同架构探索的区域互补，单一架构无法达到最佳覆盖。","relevance":"该研究直接针对LLM模拟人类意见的多样性问题，通过析因实验分离干预因素，并批判性揭示了常见假设的失效条件，与研究者关注的仿真可靠性及偏差评估高度契合，值得精读。","inspiration":"借鉴析因实验设计，独立操纵人物设定深度和交互架构，系统分离影响LLM输出多样性的因素，并测量不同条件下的意见覆盖范围。｜可迁移到政策公告的预期形成研究，如分析不同信息框架下公众对通胀或利率预期的异质性。｜以LLM为被试，处理变量为人物设定细节（如职业、收入）和交互方式（单次/多轮），结果变量为预期分布的多样性，用央行调查的真实公众预期数据做基准对照。"}},{"id":"2606.14715","version":1,"title":"MiroBench: Benchmarking Realism in Agentic Simulation of Real-world Discussions","zh_title":"MiroBench：基准测试真实世界讨论的智能体仿真真实性","abstract":"LLM agents are increasingly used to simulate real world interactions, but it remains unclear whether simulated behaviors preserve the content patterns and interaction dynamics of real human behaviors. Existing evaluations remain fragmented, which makes it difficult to compare systems or measure progress. In this paper, we focus on Reddit discussions as a concrete first step toward evaluating real-world social simulation. Reddit threads provide public, topic-grounded, multi-party interactions where people share experiences, debate, seek advice, express emotion, and collectively respond to products, events, and social issues. These discussions offer an observable window into broader social behavior, making them a useful setting for testing whether LLM agents can reproduce not only fluent text, but also the distributional patterns and interaction dynamics of real online communities. We introduce MiroBench, a benchmark for Reddit discussion simulation built from 4,292 real Reddit threads. MiroBench uses statistical tests to compare generated and real discussions across four major aspects: repetition and semantic uniformity, narrative content, toxicity and aggression, and structural complexity. Experiments across five domains and five models show that current simulators remain distributionally mismatched with real Reddit threads, while a lightweight prompt-based improvement procedure provides only limited gains. MiroBench offers a concrete benchmark for measuring, diagnosing, and improving realism in LLM-based social simulation.","authors":["Yaoning Yu","Ye Yu","Haojing Luo","Haohan Wang"],"categories":["cs.MA","cs.AI","cs.SI"],"primary_category":"cs.MA","announce_type":"new","date":"2026-05-10","first_seen":"2026-05-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2606.14715","pdf_url":"https://arxiv.org/pdf/2606.14715","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B4"],"tags":["LLM仿真","社会模拟","真实性评估"],"reason":"用LLM agent模拟Reddit讨论，并与真实人类数据对照，评估仿真真实性…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:59","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":52,"question":"LLM代理模拟的Reddit讨论在内容模式和交互动态上是否与真实人类讨论一致？","design":"使用五种LLM（未具体列出）作为代理，基于875个标准化种子上下文（产品描述）生成Reddit讨论线程，与4,292个真实Reddit线程在重复与语义均匀性、叙事内容、毒性与攻击性、结构复杂性四个维度上进行比较。","baseline":"来自五个领域（信用卡、笔记本电脑、手机、相机、耳机）的4,292个真实Reddit讨论线程。","findings":"当前LLM模拟器生成的讨论与真实Reddit线程在分布上存在系统性不匹配，即使个体评论流畅；基于提示的轻量级改进程序仅带来有限提升。","reliability":"论文未讨论","relevance":"高度相关：该研究直接以真实人类讨论为基准，评估LLM代理在社会互动仿真中的分布保真度，并揭示了当前模型的系统性偏差，符合对仿真可靠性与失效条件的关注。","inspiration":"该研究通过将LLM代理生成的讨论与真实Reddit线程在多个维度（如毒性、叙事结构）上进行分布比较，提供了评估仿真保真度的系统框架｜可迁移到经济政策沟通场景，如评估央行公告后公众预期形成的仿真可信度｜以LLM代理模拟公众对利率决议的讨论，处理为不同政策措辞，结果变量为预期通胀的分布，以真实社交媒体或调查数据为基准对照"}},{"id":"2605.08837","version":1,"title":"The Grounding Gap: How LLMs Anchor the Meaning of Abstract Concepts Differently from Humans","zh_title":"接地差距：大语言模型如何以不同于人类的方式锚定抽象概念的意义","abstract":"Abstract concepts - justice, theory, availability - have no single perceivable referent; in the human brain, their meaning emerges from a web of experiences, affect, and social context. Do large language models (LLMs) ground abstract concepts in a similar way? We study this by replicating property-generation experiments from cognitive science on 21 frontier and open-weight LLMs. Across models and experiments, we find a consistent pattern: when compared to humans, models rely too heavily on word associations, and underproduce properties tied to emotion and internal states. This yields a large and consistent grounding gap: no model exceeds a Pearson correlation r=0.37 with human responses, compared to a human-to-human ceiling above r=0.9. To better interpret this gap, we also replicate a rating experiment on grounding categories and find that here LLMs align more closely with human judgment, and alignment improves as models get larger. We then use sparse autoencoders (SAEs) to inspect whether this information is also reflected in the models' internal features, and we do identify features connected to grounding dimensions such as \"sensorimotor\" and \"social\". These findings suggest that current LLMs can recover grounding dimensions when explicitly queried, but do not recruit them in a human-like way when words are generated freely.","authors":["Odysseas S. Chlapanis","Orfeas Menis Mastromichalakis","Christos H. Papadimitriou"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-05-09","first_seen":"2026-05-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2605.08837","pdf_url":"https://arxiv.org/pdf/2605.08837","source_feed":"backfill","score":7,"bucket":"pending","rubric_hits":["A1","B1","B4"],"tags":["LLM仿真","认知实验复现","概念接地"],"reason":"用LLM复现人类认知实验，有真实人类数据对照，并指出仿真失效条件，方法可迁移。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:36","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":85,"question":"大语言模型在抽象概念的意义接地（grounding）上是否与人类一致？","design":"在21个前沿和开源LLM上复现认知科学中的属性生成实验和评分实验：给定抽象概念，让模型自由列出属性或对概念在14个接地维度上评分，测量模型与人类在属性分布和维度评分上的相关性。","baseline":"人类属性生成实验的编码分布和维度评分数据，人类间相关性天花板超过0.9。","findings":"在自由属性生成中，所有模型与人类的相关性不超过0.37，模型过度依赖词语联想而缺乏情感和内部状态属性，存在显著的接地差距；但在明确维度评分任务中，模型与人类判断更一致，且随模型规模增大而改善。","reliability":"论文指出当前LLM在自由生成时未能以类人方式调用接地维度，尽管内部表征中存在相关特征；差距可能源于模型缺乏人类具身经验和社会情境。","relevance":"该研究用LLM复现人类认知实验，有真实人类数据对照，并明确指出了仿真在自由生成任务中的失效条件，直接回应了研究者对LLM仿真可靠性及偏差的关注，值得精读。","inspiration":"该方法借鉴了用LLM复现人类认知实验并直接对比真实人类数据的范式，通过自由生成与结构化评分双任务揭示模型在接地维度上的偏差。｜可迁移到消费者信心调查或通胀预期形成研究，检验LLM生成的预期分布是否与人类调查数据一致。｜以LLM作为被试，输入经济新闻或政策描述，让其自由生成对未来通胀的判断，结果变量为预期值分布与情感属性，对照密歇根消费者调查的微观数据。"}},{"id":"2605.08802","version":2,"title":"CoLVR: Enhancing Exploratory Latent Visual Reasoning via Contrastive Optimization","zh_title":"CoLVR：通过对比优化增强探索性潜在视觉推理","abstract":"Due to the potential for exploratory reasoning of Latent Visual Reasoning, recent works tend to enable MLLMs (Multimodal Large Language Models) to perform visual reasoning by propagating continuous hidden states instead of decoding intermediate steps into discrete tokens. However, existing works typically rely on hard alignment objectives to force latent representations to match predefined visual features, thereby severely limiting the exploratory of latent reasoning process. To address this problem, we propose CoLVR (Contrastive Optimization for Latent Visual Reasoning). To obtain a more exploratory visual reasoning, CoLVR introduces a latent contrastive training framework. Firstly, CoLVR learns diverse and exploratory representations with a latent contrastive objective guided by angle-based perturbation, which expands the semantic latent space and avoids over-constrained embedding. Then, CoLVR employs a latent trajectory contrastive reward for RL (Reinforcement Learning) post-training to enable fine-grained optimization of latent visual reasoning process and thus fostering diverse reasoning behaviors. Experiments demonstrate that CoLVR significantly enhances the exploratory capability of latent representations, achieving average improvements of 5.83% on VSP and 8.00% on Jigsaw, while also outperforming existing latent models on out of domain benchmarks, with a 3.40% gain on MMStar. The data, codes, and models are released at https://github.com/Oscar-dzy/CoLVR.","authors":["Ziyang Ding","Linjian Meng","Yiming Wu","Yuhan Li","Yuhao Liu","Zhen Zhao"],"categories":["cs.CV"],"primary_category":"cs.CV","announce_type":"new","date":"2026-05-09","first_seen":"2026-05-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2605.08802","pdf_url":"https://arxiv.org/pdf/2605.08802","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["视觉推理","多模态大模型","对比学习"],"reason":"纯视觉推理模型优化，无人类仿真或行为对照，属NLP/CV能力评测。","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:59:09","error":null,"has_summary":false,"summary":null},{"id":"2605.07692","version":1,"title":"GASim: A Graph-Accelerated Hybrid Framework for Social Simulation","zh_title":"GASim：一种图加速的混合社会仿真框架","abstract":"Large-scale social simulators are essential for studying complex social patterns. Prior work explores hybrid methods to scale up simulations, combining large language models (LLM)-based agents with numerical agent-based models (ABM). However, this incurs high latency due to expensive memory retrieval and sequential ABM execution. To address this challenge, we propose GASim, a graph-accelerated hybrid multi-agent framework for large-scale social simulations. For core agents driven by LLM, GASim introduces Graph-Optimized Memory (GOM) to replace intensive LLM-based retrieval pipelines with lightweight propagation over a sparse memory graph. For the majority of ordinary agents, GASim employs Graph Message Passing (GMP), substituting sequential ABM execution with parallel updates by fine-grained feature aggregation and Graph Attention Network. We further introduce Entropy-Driven Grouping (EDG) that coordinates this hybrid partitioning, leveraging information entropy to dynamically identify emergent core agents situated in information-diverse neighborhoods. Extensive experiments show that GASim not only delivers a substantial 9.94-fold end-to-end speedup over the traditional hybrid framework but also consumes less than 20% of baseline tokens, significantly reducing costs while preserving strong alignment with real-world public opinion trends. Our code is available at https://github.com/Jasmine0201/GASim.","authors":["Xuan Zhou","Yanhui Sun","Hantao Yao","Allen He","Yongdong Zhang","Wu Liu"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-05-08","first_seen":"2026-05-08","revised_at":null,"abs_url":"https://arxiv.org/abs/2605.07692","pdf_url":"https://arxiv.org/pdf/2605.07692","source_feed":"backfill","score":6,"bucket":"other","rubric_hits":["D3"],"tags":["社会模拟","LLM智能体","图加速"],"reason":"用LLM agent模拟社会舆论，有真实数据对照，但核心是加速框架而非仿真方法论","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:35","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":129,"question":"如何加速大规模社会仿真中混合框架（LLM智能体+数值ABM）的执行速度并降低成本，同时保持与真实舆论趋势的一致性？","design":"提出GASim框架：用熵驱动分组（EDG）动态识别信息多样性高的核心智能体（由LLM驱动），其余为普通智能体（由数值模型驱动）；核心智能体用图优化记忆（GOM）替代LLM检索，普通智能体用图消息传递（GMP）并行更新意见；在舆论仿真中测量端到端加速比、token消耗和与真实舆论趋势的几何对齐度。","baseline":"真实世界的公众舆论趋势数据（具体数据集未在节选中指明）。","findings":"GASim相比传统混合框架实现9.94倍端到端加速，token消耗降至基线20%以下；在记忆检索基准LoCoMo上达到71.56%准确率，并在真实舆论趋势对齐上表现更优。","reliability":"论文未讨论","relevance":"该工作使用LLM智能体模拟社会舆论并与真实数据对照，但核心贡献是加速框架而非仿真方法论或可靠性批判，与研究者关注的仿真效度与失效条件关联较弱，可略读。","inspiration":"可借鉴其熵驱动分组（EDG）动态识别信息多样性高的核心智能体，以降低仿真成本并保持关键行为模式｜可迁移到政策公告的预期形成研究，如央行沟通对市场参与者通胀预期的影响｜用LLM智能体模拟分析师，EDG筛选关注政策信号的核心智能体，处理为不同措辞的央行声明，结果变量为预期通胀分布，以专业预测者调查的真实数据做对照"}},{"id":"2605.05578","version":1,"title":"Artificial Aesthetics: The Implicit Economics of Valuing AI-Generated Text","zh_title":"人工美学：评估AI生成文本的隐含经济学","abstract":"Aesthetic qualities command measurable premiums in traditional goods markets. However, it remains unclear whether users are willing to pay for such qualities in AI-generated text. This paper estimates the willingness to pay for aesthetic attributes in large language model outputs using an online experiment with N = 117 participants. Participants evaluated responses from four anonymized models across academic, professional, and personal contexts, rated outputs along multiple dimensions, and submitted bids for access using a Becker-DeGroot-Marschak (BDM) mechanism. We find no statistically significant relationship between perceived aesthetic quality and willingness to pay. While participants systematically distinguish between outputs and exhibit consistent preferences over stylistic features, these differences do not translate into higher monetary valuation. Further analysis shows that aesthetic and functional attributes load onto a single latent factor, suggesting that users perceive quality as a unified construct rather than a separable aesthetic dimension. These results imply that, in current large language model (LLM) markets, aesthetic improvements function as baseline expectations rather than sources of price differentiation.","authors":["Arbaaz Karim"],"categories":["econ.GN"],"primary_category":"econ.GN","announce_type":"new","date":"2026-05-07","first_seen":"2026-05-07","revised_at":null,"abs_url":"https://arxiv.org/abs/2605.05578","pdf_url":"https://arxiv.org/pdf/2605.05578","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["支付意愿","AI文本评估","行为经济学"],"reason":"研究用户对AI文本的支付意愿，属NLP评测，不以人类行为仿真为目标","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:57","error":null,"has_summary":false,"summary":null},{"id":"2605.06987","version":1,"title":"Response Time Enhances Alignment with Heterogeneous Preferences","zh_title":"响应时间增强异质性偏好对齐","abstract":"Aligning large language models (LLMs) to human preferences typically relies on aggregating pooled feedback into a single reward model. However, this standard approach assumes that all labelers share the same underlying preferences, ignoring the fact that real-world labelers are highly heterogeneous and usually anonymous. Consequently, relying solely on binary choice data fundamentally distorts the learned policy, making the true population-average preference unidentifiable. To overcome this critical limitation, we demonstrate that augmenting preference datasets with a simple, secondary signal -- the user's response time -- can restore the identifiability of the population's average preference. By modeling each decision as a Drift-Diffusion Model (DDM), we introduce a novel, consistent estimator of heterogeneous preferences that successfully corrects the distortions of standard choice-only labels. We prove that our estimator asymptotically converges to the true average preference even in extreme cases where each anonymous labeler contributes only a single choice. Empirically, across both synthetic and real-world datasets, our method consistently outperforms standard baselines that otherwise fail and plateau at a bias floor. Because response times are essentially free to record and require zero user tracking or identification, our results bring promises and open up new opportunities for future data-collection pipelines to improve the social benefit without requiring user-level identifiers or repeated elicitations.","authors":["Federico Echenique","Alireza Fallah","Baihe Huang","Michael I. Jordan"],"categories":["cs.LG","cs.GT","econ.TH","stat.ML"],"primary_category":"cs.LG","announce_type":"new","date":"2026-05-07","first_seen":"2026-05-07","revised_at":null,"abs_url":"https://arxiv.org/abs/2605.06987","pdf_url":"https://arxiv.org/pdf/2605.06987","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["偏好对齐","响应时间","异质性"],"reason":"论文研究利用响应时间改进偏好对齐，属于纯NLP能力评测，不以人类行为仿真为参照。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:57","error":null,"has_summary":false,"summary":null},{"id":"2605.18781","version":1,"title":"Can LLMs Emulate Human Belief Dynamics?","zh_title":"大语言模型能模拟人类信念动态吗？","abstract":"Can LLMs simulate how humans form and change beliefs in social networks? We put this to the test by replicating an established study on belief dynamics, evaluating 12 LLMs across multiple model families and parameter sizes. The answer is a clear no, and in systematic ways. LLMs fail to capture initial human belief distributions and tend to be overall more conformist than humans, shifting their responses to align with those around them. They also take a nuanced approach to emulating human homophilic tendencies within networks. Our findings carry a double payoff: they highlight fundamental properties of LLM behavior, and they raise a sharp warning against deploying LLMs as human proxies in social simulations.","authors":["Adiba Mahbub Proma","Neeley Pate","James N. Druckman","Gourab Ghoshal","Hangfeng He","Ehsan Hoque"],"categories":["cs.SI","cs.AI","cs.CY"],"primary_category":"cs.SI","announce_type":"new","date":"2026-05-05","first_seen":"2026-05-05","revised_at":null,"abs_url":"https://arxiv.org/abs/2605.18781","pdf_url":"https://arxiv.org/pdf/2605.18781","source_feed":"backfill","score":10,"bucket":"selected","rubric_hits":["A1","A2","A5","B1","B4"],"tags":["LLM仿真","信念动态","人类数据对照"],"reason":"直接复现人类信念动态研究，用LLM替代人类被试，有真实人类数据对照，并指出仿真…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:42","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":13,"question":"LLM能否在社交网络中模拟人类信念的形成与变化？","design":"用12个LLM（含推理与非推理模型）基于真实参与者的年龄、性别、种族、教育、收入、政治倾向及大五人格创建“数字孪生”，复现一项人类信念动态实验：先对政治议题陈述进行5点李克特评分，再看到他人评分后允许修改，最后选择关注/取关他人；测量初始信念分布、信念变化及网络选择行为。","baseline":"原人类实验的341名参与者（共1023个样本）在移民、石油与燃料两个议题上的真实评分、信念更新和网络选择数据。","findings":"LLM系统性地无法模拟人类信念动态：初始信念分布与人类显著不同，且比人类更易从众，倾向于改变自身回答以对齐周围意见；在网络选择上，LLM能部分模仿人类选择，但无法复现人类的同质性倾向。","reliability":"论文指出，仅使用简单的人口统计与大五人格构建数字孪生可能信息不足，导致仿真失败；且实验仅涵盖两个政治议题，未测试更多样化的情境。","relevance":"该研究直接检验LLM在信念动态仿真中的可靠性，有真实人类对照，发现系统性失效，对关注LLM作为人类被试替代品的研究者具有重要警示价值，值得精读原文。","inspiration":"借鉴其用真实人类实验数据作为严格对照基准，并系统比较LLM与人类在信念更新和网络选择上的分布差异，而非仅看均值或方向。｜可迁移到政策公告的预期形成实验，研究市场参与者如何根据他人预期调整自身通胀或利率预测。｜以LLM模拟投资者，先给出个人通胀预测，再展示其他‘投资者’的预测（处理），观察其预测修正幅度与方向，结果变量为预测调整量和最终预测分布，用专业预测者调查（如SPF）的真实个体数据做对照。"}},{"id":"2605.03604","version":1,"title":"Multi-Agent Strategic Games with LLMs","zh_title":"基于大语言模型的多智能体战略博弈研究","abstract":"This paper asks whether large language models (LLMs) can be used to study the strategic foundations of conflict and cooperation. I introduce LLMs as experimental subjects in a repeated security dilemma and evaluate whether they reproduce canonical mechanisms from international relations theory. The baseline game is extended along three theoretically central dimensions: multipolarity, finite time horizons, and the availability of communication. Across multiple models, the results exhibit systematic and consistent patterns: multipolarity increases the likelihood of conflict, finite horizons induce universal unraveling consistent with backward-induction logic, and communication reduces conflict by enabling signaling and reciprocity. Beyond observed behavior, the design provides access to agents' private reasoning and public messages, allowing choices to be linked to underlying strategic logics such as preemption, cooperation under uncertainty, and trust-building. The contribution is primarily methodological. LLM-based experiments offer a scalable, transparent, and replicable approach to probing theoretical mechanisms.","authors":["Maxim Chupilkin"],"categories":["cs.GT","cs.AI","cs.CY"],"primary_category":"cs.GT","announce_type":"new","date":"2026-05-05","first_seen":"2026-05-05","revised_at":null,"abs_url":"https://arxiv.org/abs/2605.03604","pdf_url":"https://arxiv.org/pdf/2605.03604","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A3","B2","B3"],"tags":["LLM仿真","战略博弈","国际关系"],"reason":"将LLM作为实验被试研究安全困境中的战略行为，复现国际关系理论机制，涉及博弈实…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:34","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":14,"question":"大型语言模型能否作为实验被试，在重复安全困境博弈中复现国际关系理论中的战略机制？","design":"使用GPT-5、GPT-5 Mini、Sonnet、Gemini等LLM作为被试，在重复安全困境博弈中扮演国家，通过操纵多极性、有限时间范围、通信可用性三个处理，测量冲突发生率、冲突时机和攻击结构等结果变量。","baseline":"无对照","findings":"多极性增加冲突概率，有限时间范围导致完全瓦解，通信通过信号传递和互惠降低冲突。LLM的私有推理和公开消息映射到先发制人、不确定性下的合作等战略逻辑。","reliability":"论文未讨论","relevance":"高度相关，直接使用LLM作为人类被试替代品进行博弈实验，复现战略行为模式，并评估处理效应稳健性，符合研究者对经济学实验和政策评估场景的关注。","inspiration":"借鉴其通过改变博弈结构（多极性、时间范围、通信）系统检验理论机制的设计，以及利用LLM私有推理和公开消息进行过程追踪的方法。｜可迁移到产业组织中的合谋实验，如多寡头重复价格竞争。｜以LLM作为企业被试，在重复囚徒困境中操纵市场集中度、时间范围和沟通渠道，测量合谋频率与稳定性，对照真实行业价格战数据或人类实验数据。"}},{"id":"2605.03287","version":1,"title":"Attention: What Prevents Young Adults from Speaking Up Against Cyberbullying in an LLM-Powered Social Media Simulation","zh_title":"注意力：是什么阻止年轻人在LLM驱动的社交媒体仿真中公开反对网络欺凌","abstract":"Interactive, multi-agent social simulation systems have shown promise for helping users practice navigating various complex social situations across domains. This paper asks: To what extent can such systems help young adult (YA) bystanders speak up publicly against cyberbullying, a task often thwarted by complex, multi-party social dynamics? We created Upstanders' Practicum, a multi-AI-agent social media simulation powered by Large Language Models (LLMs), as a probe and observed 34 YAs freely practicing public bystander intervention across three iteratively refined versions. We found that practicing public bystander intervention in the simulation was helpful, but after participants made three attention shifts: (1) from inattention to paying true attention, (2) from self-focus (\"I don't usually do this'') to attending to those directly involved, and (3) from resolving the private conflict between bully and victim (\"maybe I could set up the meeting between them'') to addressing the broader audience online (\"public comment is about norm-setting\"). Only after these shifts did practice in the simulation start to help: participants then saw a reason to speak up publicly and, through continued practice, crafted tactful public messages without explicit instruction. These findings illuminate new design and research opportunities for bystander education beyond social skill instruction, namely, designing for true attention, for fostering a vocal upstander identity, and for seeing bystander intervention as public norm setting. In addition, we open-source Truman Agents (cornell-design-aigroup.github.io/TrumanAgents/), the first-of-its-kind multi-LLM-agent social media simulation platform that Upstanders' Practicum builds upon, for future cyberbullying and social media research.","authors":["Qian Yang","Jessie Jia","Elaine Tsai","Amy Li","Nader Akoury","Natalie N. Bazarova"],"categories":["cs.HC","cs.CY"],"primary_category":"cs.HC","announce_type":"new","date":"2026-05-05","first_seen":"2026-05-05","revised_at":null,"abs_url":"https://arxiv.org/abs/2605.03287","pdf_url":"https://arxiv.org/pdf/2605.03287","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["多智能体仿真","网络欺凌干预","人机交互"],"reason":"多智能体社交媒体仿真，有真实人类参与练习，但无人类行为对照基准，属社会模拟边界…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:33","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":156,"question":"多智能体社交媒体仿真能在多大程度上帮助年轻成年旁观者在网络欺凌中公开发声？","design":"使用基于LLM的多智能体社交媒体仿真平台“Upstanders' Practicum”，由多个LLM代理扮演欺凌者、受害者等角色，34名年轻成年人作为人类被试在三个迭代版本中自由练习公开旁观者干预，通过观察和分析参与者的行动与推理来探究仿真帮助效果。","baseline":"无对照","findings":"仿真练习对公开旁观者干预有帮助，但前提是参与者经历了三次注意力转移：从忽视到真正关注、从自我关注转向关注直接当事人、从解决私人冲突转向面向更广泛的在线受众；只有完成这些转移后，参与者才开始看到公开发声的理由，并通过持续练习在没有明确指导的情况下自行构思出得体的公开信息。","reliability":"论文未讨论","relevance":"该研究利用LLM多智能体仿真探究人类在复杂社会情境中的行为改变过程，虽无真实人类行为基准对照，但揭示了注意力转移作为干预有效性的关键条件，对理解LLM仿真在教育和行为训练中的边界具有参考价值，值得阅读原文以了解仿真设计细节和定性发现。","inspiration":"该研究利用LLM多智能体仿真构建复杂社会情境，让人类被试在无明确指导的迭代练习中自然涌现行为改变，这种通过仿真环境诱发内生注意力转移的设计值得借鉴。｜可迁移至消费者金融决策中的信息披露干预研究，例如在仿真投资平台上测试散户如何从忽视风险提示转变为主动关注并利用风险信息。｜以散户投资者为被试，在LLM驱动的仿真投资环境中嵌入渐进式风险提示，处理为不同信息呈现方式，结果变量为信息查阅频率与投资组合调整，对照真实市场交易数据中的风险关注行为。"}},{"id":"2605.04029","version":1,"title":"Stayin' Aligned Over Time: Towards Longitudinal Human-LLM Alignment via Contextual Reflection and Privacy-Preserving Behavioral Data","zh_title":"随时间保持对齐：通过情境反思和隐私保护行为数据实现纵向人-LLM对齐","abstract":"Current human-AI alignment and evaluation methods for large language models (LLMs) often rely on preference signals collected immediately after an interaction. This practice implicitly treats preference as static, even though many LLM-mediated decisions unfold over time and may be re-evaluated differently after real-world consequences and observed outcomes. Therefore, we argue for a methodological shift from single-moment preference elicitation to longitudinal, context-situated alignment measurement. We present a methodological framework for collecting temporally grounded alignment signals by combining (1) in-situ preference capture, (2) context-triggered follow-up preference reflection, and (3) privacy-preserving behavioral traces that help interpret preference change. As an instantiation of this methodology, we introduce BITE, a browser-based system that detects consequential LLM interactions, prompts reflection across later decision points, and supports progressive, user-controlled consent for sharing behavioral data. Through a two week longitudinal deployment study with 8 participants, our approach surfaced differences between immediate and later user preferences in accuracy, relevance and other dimensions of the LLM output. Our findings highlight the limitations of single-moment preference datasets and underscore the importance of longitudinal methods for alignment evaluation in everyday use.","authors":["Simret Araya Gebreegziabher","Allison E Sproul","Yinuo Yang","Chaoran Chen","Diego Gómez-Zará","Toby Jia-Jun Li"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-05-05","first_seen":"2026-05-05","revised_at":null,"abs_url":"https://arxiv.org/abs/2605.04029","pdf_url":"https://arxiv.org/pdf/2605.04029","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["人机对齐","纵向研究","偏好测量"],"reason":"研究人类与LLM对齐的纵向变化，测量用户偏好而非用LLM仿真人类被试，属于边界…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:56","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":173,"question":"如何通过纵向、情境化的方法捕捉用户对LLM输出的偏好随时间的变化，以改进人机对齐评估？","design":"本研究不是用LLM仿真人类被试，而是设计了一个浏览器系统BITE，在两周内部署8名参与者，通过即时偏好采集、后续情境触发反思和隐私保护行为追踪，测量用户对LLM交互的即时与延迟评价差异。","baseline":"无对照","findings":"即时与延迟判断在准确性和相关性维度上差异最大；信任度在后续评价中上升（80%增加），而有害性评价则相反，这些变化与结果验证和情境重建有关，而非单纯时间流逝。","reliability":"论文未讨论","relevance":"该研究聚焦于人类用户偏好的纵向变化，而非用LLM仿真人类被试，不涉及经济学实验或政策评估，与研究者关注的LLM仿真人类行为及基准对照无关，不建议阅读原文。","inspiration":"该方法通过情境触发反思和隐私保护行为追踪来捕捉用户偏好的纵向变化，可借鉴其纵向测量设计，在多次干预后设置反思节点以区分即时与延迟评价｜可迁移到政策公告的预期形成研究，例如考察公众对央行沟通的即时反应与后续解读的差异｜可招募普通公众为被试，施加不同措辞的政策公告处理，结果变量为通胀预期和信任度，以真实央行调查数据为对照基准"}},{"id":"2605.02598","version":1,"title":"What Jobs Can AI Learn? Measuring Exposure by Reinforcement Learning","zh_title":"AI能学会哪些工作？用强化学习衡量职业暴露度","abstract":"Which jobs can AI learn to do? We examine this for every occupation in the US economy. Existing indices measure the overlap between AI capabilities and occupational tasks rather than which tasks AI systems can learn to perform, and as a result misclassify occupations where the gap between present capability and learnability is large. Reinforcement learning in post-training, now the dominant paradigm at the frontier, is structured around task completion and maps more directly onto the task-based architecture of occupational classifications than prior approaches. Using LLM annotators guided by a rubric developed with RL experts and validated against confirmed deployment cases, we score all 17,951 ONET tasks for training feasibility and aggregate to the occupation level, producing an RL Feasibility Index. The index diverges sharply from existing AI exposure measures for specific occupation groups: power plant operators, railroad conductors, and aircraft cargo handling supervisors score high on RL feasibility but low on general AI exposure, while creative and interpersonal roles (musicians, physicians, natural sciences managers) show the reverse. These divergences carry direct implications for policy interventions.","authors":["Philip Moreira Tomei","Bouke Klein Teeselink"],"categories":["econ.GN"],"primary_category":"econ.GN","announce_type":"new","date":"2026-05-04","first_seen":"2026-05-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2605.02598","pdf_url":"https://arxiv.org/pdf/2605.02598","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["AI暴露度","职业分类","强化学习"],"reason":"用LLM标注任务可学习性，属纯NLP能力评测，不以人类行为为参照系。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:55","error":null,"has_summary":false,"summary":null},{"id":"2606.11217","version":1,"title":"Preregistration for Experiments with AI Agents","zh_title":"AI代理实验的预注册","abstract":"The proliferation of large language models (LLMs) and autonomous AI agents has given rise to a rapidly growing methodological paradigm: \"in silico\" behavioral experiments. Originally conceived as a way to use AI agents as proxies for human participants in studies of cognition, decision-making, and social dynamics, this approach has taken on new significance -- as AI agents increasingly negotiate, transact, and make consequential decisions on behalf of people and organizations, understanding their behavior has become a research priority in its own right. While these experiments with AI agents offer unprecedented advantages in terms of scalability, cost efficiency, and experimental control, they also inherit, and in some cases amplify, methodological vulnerabilities that have long plagued human subjects research. To address these issues, this paper argues that preregistration practices -- central to improving the credibility of human subjects experiments -- should now be extended to experiments with AI agents. We systematically catalog the researcher degrees of freedom that experiments with AI agents introduce -- model selection, prompt wording, settings, and outcome-contingent redesign, for example -- and show how the low cost of iteration and lack of reporting norms make these choices both easy to exploit and difficult to detect. We propose a preregistration template tailored to experiments with AI agents and call on conferences, journals, and funding agencies to make preregistration standard practice for this emerging research paradigm.","authors":["Michelle Vaccaro"],"categories":["cs.CY","cs.AI","cs.HC"],"primary_category":"cs.CY","announce_type":"new","date":"2026-05-03","first_seen":"2026-05-03","revised_at":null,"abs_url":"https://arxiv.org/abs/2606.11217","pdf_url":"https://arxiv.org/pdf/2606.11217","source_feed":"backfill","score":8,"bucket":"selected","rubric_hits":["A4","B4"],"tags":["AI代理实验","预注册","方法论"],"reason":"提出AI代理实验的预注册规范，批判性指出方法漏洞，方法论可迁移至人类仿真研究。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:55","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":43,"question":"如何将人类被试实验的预注册实践扩展到以AI智能体为对象的实验中，以控制研究者自由度并提高研究可信度？","design":"本文不是一项仿真实验研究，而是一篇方法论论文。它系统梳理了在用AI智能体进行行为实验时，从模型选择、提示词措辞、采样参数、实验设计到结果解析和报告等全流程中存在的各种研究者自由度，并论证这些自由度如何容易被利用且难以被察觉。","baseline":"无对照","findings":"AI智能体实验继承了人类被试实验中的研究者自由度问题，并引入了模型选择、提示词工程、解码参数等新的高维选择空间，低迭代成本和缺乏报告规范使得投机性选择极易发生且难以检测。论文提出了一个针对AI智能体实验的预注册模板，并呼吁会议、期刊和资助机构将其作为标准实践。","reliability":"论文未讨论","relevance":"本文批判性地指出用AI智能体替代人类被试进行行为实验时存在严重的方法论漏洞，并提出了预注册这一解决方案，其分析框架和提出的规范可直接迁移到基于LLM的人类仿真研究中，对关注仿真可靠性与偏差的研究者具有重要参考价值，值得阅读原文。","inspiration":"该方法论论文提出了针对AI智能体实验的预注册模板，系统梳理了从模型选择、提示词设计到结果解析的全流程研究者自由度，为控制实验者偏差提供了可操作的规范框架。｜可迁移到政策公告预期形成的仿真研究中，例如用LLM模拟投资者对央行沟通的反应，检验不同措辞或信息框架对预期通胀和资产配置的影响。｜以多个主流LLM（如GPT-4、Claude 3）作为被试，处理为不同措辞的政策声明（前瞻指引 vs. 数据依赖表述），结果变量为模拟的预期通胀率和风险资产配置比例，并与真实投资者调查数据（如密歇根消费者调查或专业预测者调查）进行对照，同时按预注册模板预先锁定模型版本、提示词、温度参数和分析计划。"}},{"id":"2605.01311","version":1,"title":"The Partial Testimony of Logs: Evaluation of Language Model Generation under Confounded Model Choice","zh_title":"日志的部分证言：混杂模型选择下的语言模型生成评估","abstract":"Offline evaluation of language models from usage logs is biased when model choice is confounded: the same user-side factors that influence which model is used can also influence how its output is judged, so raw comparisons of logged scores mix self-selected populations rather than estimating a common quantity of interest. A small randomized experiment can break this bias by overriding model choice, but in practice such experiments are scarce and costly. We study a three-source design that combines a large confounded observational log (OBS) for scale, a small randomized experiment (EXP) for unconfounded scoring, and an offline simulator (SIM) that replays candidate models on cached contexts. Our main result is an identification theorem showing that the randomized experiment and the simulator are together enough to recover causal model values; the observational log enters only afterward, to reduce estimation error rather than to make the causal comparison valid. Six estimator families are evaluated in a controlled semi-synthetic validation and in two real-task cached benchmarks for summarization and coding. No family dominates every regime; relative performance depends on the amount of unbiased EXP supervision and on how closely the target reward aligns with OBS-derived structure.","authors":["Jikai Jin","Vasilis Syrgkanis"],"categories":["cs.LG","econ.EM","stat.AP","stat.ML"],"primary_category":"cs.LG","announce_type":"new","date":"2026-05-02","first_seen":"2026-05-02","revised_at":null,"abs_url":"https://arxiv.org/abs/2605.01311","pdf_url":"https://arxiv.org/pdf/2605.01311","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["离线评估","因果推断","语言模型"],"reason":"研究LLM离线评估的因果识别，不涉及用LLM仿真人类被试或与人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:55","error":null,"has_summary":false,"summary":null},{"id":"2604.23575","version":2,"title":"The Collapse of Heterogeneity in Silicon Philosophers","zh_title":"硅基哲学家的异质性坍塌","abstract":"Silicon samples are increasingly used as a low-cost substitute for human panels and have been shown to reproduce aggregate human opinion with high fidelity. We show that, in the alignment-relevant domain of philosophy, silicon samples systematically collapse heterogeneity. Using data from $N = {277}$ professional philosophers drawn from PhilPeople profiles, we evaluate seven proprietary and open-source large language models on their ability to replicate individual philosophical positions and to preserve cross-question correlation structures across philosophical domains. We find that language models substantially over-correlate philosophical judgments, producing artificial consensus across domains. This collapse is associated in part with specialist effects, whereby models implicitly assume that domain specialists hold highly similar philosophical views. We assess the robustness of these findings by studying the impact of DPO fine-tuning and by validating results against the full PhilPapers 2020 Survey ($N = {1785}$). We conclude by discussing implications for alignment, evaluation, and the use of silicon samples as substitutes for human judgment. The code of this project can be found at https://github.com/stanford-del/silicon-philosophers.","authors":["Yuanming Shi","Andreas Haupt"],"categories":["cs.CY","cs.CL","cs.LG"],"primary_category":"cs.CY","announce_type":"new","date":"2026-04-26","first_seen":"2026-04-26","revised_at":null,"abs_url":"https://arxiv.org/abs/2604.23575","pdf_url":"https://arxiv.org/pdf/2604.23575","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM人类仿真","异质性评估","哲学观点复现"],"reason":"用LLM复现哲学家观点并与真实人类数据对照，评估仿真可靠性与异质性坍塌，属核心…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:33","error":null,"has_summary":true,"summary":{"generated_at":"2026-04-26","rank":4,"question":"大语言模型在模拟专业哲学家观点时，能否保留人类群体中的异质性和跨问题相关结构？","design":"用7个商业和开源LLM模拟277位专业哲学家（从PhilPeople收集的个人资料），基于其专业领域和人口统计信息生成对100个哲学问题的回答，测量回答的方差、跨问题相关性以及主成分结构。","baseline":"277位真实哲学家的PhilPeople个人资料回答，以及PhilPapers 2020调查（N=1785）的汇总数据。","findings":"LLM系统性地坍塌异质性：产生的方差比人类低2-10倍，跨问题过度相关，导致人为共识。存在虚假的专家效应：模型假设领域专家持有高度相似的哲学观点。","reliability":"论文承认样本存在北美偏向，且人类数据缺失率高（61.1%），LLM缺失率较低（17%-40%）。DPO微调能改善相关结构但无法解决异质性坍塌。","relevance":"高度相关：直接命中研究者关注的LLM仿真可靠性、有真实人类对照、经济学/政策评估场景（哲学作为专家领域案例），并批判性指出失效条件（异质性坍塌）。值得精读原文。","inspiration":"该方法借鉴了用LLM基于个体特征（如专业领域、人口统计）生成回答并与真实个体数据对比的仿真设计，可迁移到金融分析师预测或消费者通胀预期形成的异质性研究中｜可应用于研究金融分析师对宏观政策公告的预期分歧，检验LLM是否低估分析师间的观点异质性并产生虚假共识｜以真实分析师调查数据（如Bloomberg或Philadelphia Fed调查）为基准，用LLM基于分析师所属机构类型、经验年限等特征模拟其对利率决议的预测，比较预测方差、跨问题相关性及主成分结构，评估LLM仿真的异质性坍塌程度"}},{"id":"2604.23897","version":1,"title":"MarketBench: Evaluating AI Agents as Market Participants","zh_title":"MarketBench：评估AI智能体作为市场参与者","abstract":"Markets are a promising way to coordinate AI agent activity for similar reasons to those used to justify markets more broadly. In order to effectively participate in markets, agents need to have informative signals of their own ability to successfully complete a task and the cost of doing so. We propose MarketBench, a benchmark for assessing whether AI agents have these capabilities. We use a 93-task subset of SWE-bench Lite, a software engineering benchmark, with six recently released LLMs as a demonstration. These LLMs are miscalibrated on both success probability and token usage, and auctions built from these self-reports diverge from a full-information allocation. A follow-up intervention where we add information about capabilities from prior experiments to the context improves calibration, but only modestly narrows the gap to a full-information benchmark. We also document the performance of a market-based scaffolding with these LLMs. Our results point to self-assessment as a key bottleneck for market-style coordination of AI agents.","authors":["Andrey Fradkin","Rohit Krishnan"],"categories":["cs.AI","econ.GN"],"primary_category":"cs.AI","announce_type":"new","date":"2026-04-26","first_seen":"2026-04-26","revised_at":null,"abs_url":"https://arxiv.org/abs/2604.23897","pdf_url":"https://arxiv.org/pdf/2604.23897","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["AI智能体","市场模拟","基准测试"],"reason":"用LLM模拟市场参与者，但无真实人类数据对照，属社会模拟边界情形","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:53","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":174,"question":"AI智能体能否准确自我评估任务成功概率与成本，从而在基于市场的任务分配中实现有效协调？","design":"使用6个最新LLM在93个SWE-bench Lite软件工程任务上，先让模型预测自身成功概率和token消耗，再基于自报信息构建拍卖模拟，并对比全信息分配基准。","baseline":"无对照","findings":"当前LLM在成功概率和token用量上均校准不佳，基于自报信息的拍卖分配结果显著偏离全信息最优分配；补充历史能力信息后校准有所改善，但仍未消除差距。","reliability":"论文指出自我评估是市场式协调的关键瓶颈，当前模型缺乏可靠的私有信息信号，且实验仅限软件工程任务，未涉及真实人类数据。","relevance":"该研究用LLM模拟市场参与者，聚焦自我评估与拍卖分配，但无真实人类行为对照，属于社会模拟的边界情形，适合关注AI经济决策仿真的研究者了解其局限。","inspiration":"该方法通过让LLM自我评估任务成功概率与成本，再基于自报信息构建拍卖模拟，可借鉴其‘先校准后分配’的设计思路来检验AI决策偏差｜可迁移至金融资产定价实验，研究AI交易员对私有信息信号的校准能力如何影响市场效率｜设计让多个LLM扮演交易员，先预测自身对某资产未来价格的预测准确度和交易成本，再据此参与集合竞价，结果变量为价格发现效率与个体收益，对照真实人类交易员在相同信息集下的行为数据"}},{"id":"2605.27401","version":1,"title":"Using Zero-Shot LLM-Generated Survey Data for Geographically Explicit Population Synthesis","zh_title":"使用零样本LLM生成调查数据进行地理显式人口合成","abstract":"There is a growing interest in utilizing synthetic populations for a diverse range of applications. At the same time, we are witnessing a tremendous growth in artificial intelligence in all walks of life. This paper evaluates whether zero-shot large language model (LLM)-generated health survey data can serve as inputs to a conventional iterative proportional fitting (IPF) workflow for geographically explicit population synthesis. Using the 2023 Behavioral Risk Factor Surveillance System (BRFSS), we generate synthetic survey records for the U.S. states of Colorado and Mississippi with GPT-4.1 and Gemini-2.5-Pro. We use the generated data in an IPF-based synthesis pipeline and evaluate the resulting census tract-level synthetic populations against external benchmarks. Results show both LLMs capture several major state-level contrasts, indicating zero-shot generation produces geographically differentiated survey data. However, performance is strongly variable-dependent. Downstream effects in population synthesis are mixed, as IPF sometimes amplifies or reduces errors in the generated data. Spatial validation shows that LLM-based populations reproduce census tract-level patterns reasonably well, especially for variables that were more aligned with the ground truth data. Overall, the LLM-generated survey data shows promise as supplementary input, but not yet as a replacement for real survey data.","authors":["Taylor Anderson","Sara Von Hoene","Orhan Yagizer Cinar","Emma Von Hoene","Amira Roess","Andrew Crooks","Hamdi Kavak"],"categories":["cs.CY","cs.AI"],"primary_category":"cs.CY","announce_type":"new","date":"2026-04-23","first_seen":"2026-04-23","revised_at":null,"abs_url":"https://arxiv.org/abs/2605.27401","pdf_url":"https://arxiv.org/pdf/2605.27401","source_feed":"backfill","score":8,"bucket":"selected","rubric_hits":["A1","A2","B1","B2"],"tags":["LLM仿真","人口合成","健康调查"],"reason":"用LLM生成健康调查数据替代人类被试，并与真实BRFSS数据对照，评估仿真可靠…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:48","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":32,"question":"零样本LLM生成的健康调查数据能否作为地理显式人口合成的输入，用于迭代比例拟合（IPF）流程？","design":"使用GPT-4.1和Gemini-2.5-Pro，在零样本设置下生成美国科罗拉多州和密西西比州的BRFSS健康调查个体记录，然后将生成数据作为IPF的输入进行人口合成，生成普查区级合成人口，并与外部基准对比。","baseline":"2023年BRFSS加权调查数据作为真实人类对照，以及基于真实调查数据生成的合成人口。","findings":"LLM能捕捉州级健康特征差异，生成地理分化的调查数据，但性能高度依赖变量。IPF有时会放大或减小生成数据中的误差，LLM生成数据作为补充输入有潜力，但尚不能替代真实调查数据。","reliability":"论文指出LLM生成数据在变量间联合分布和子群体模式上可能引入偏差、误分类或代表性不足，这些错误会通过IPF传播到最终合成人口中，且性能因变量而异。","relevance":"该研究直接评估LLM替代人类被试生成调查数据的可靠性，并与真实调查数据对照，涉及健康领域和地理显式仿真，符合研究者对LLM仿真失效条件的关注，值得阅读原文以了解具体偏差模式。","inspiration":"该方法通过零样本LLM生成个体记录并输入IPF进行地理显式人口合成，可借鉴其将LLM生成数据作为先验输入、用真实调查数据对照评估偏差的设计思路｜可迁移到区域经济政策评估中的异质性个体仿真，例如模拟不同地区居民对税收优惠或补贴政策的响应差异｜以LLM生成不同地理区域的居民特征（收入、就业、消费偏好）作为IPF输入合成区域人口，施加政策处理（如减税），结果变量为消费或劳动供给变化，用真实家庭调查数据（如PSID）作为对照基准"}},{"id":"2604.21334","version":2,"title":"Ideological Bias in LLMs' Economic Causal Reasoning","zh_title":"大语言模型经济因果推理中的意识形态偏差","abstract":"Do large language models (LLMs) exhibit systematic ideological bias when reasoning about economic causal effects? As LLMs are increasingly used in policy analysis and economic reporting, where directionally correct causal judgments are essential, this question has direct practical stakes. We present a systematic evaluation by extending the EconCausal benchmark with ideology-contested cases - instances where intervention-oriented (pro-government) and market-oriented (pro-market) perspectives predict divergent causal signs. From 10,490 causal triplets (treatment-outcome pairs with empirically verified effect directions) derived from top-tier economics and finance journals, we identify 1,056 ideology-contested instances and evaluate 20 state-of-the-art LLMs on their ability to predict empirically supported causal directions. We find that ideology-contested items are consistently harder than non-contested ones, and that across 18 of 20 models, accuracy is systematically higher when the empirically verified causal sign aligns with intervention-oriented expectations than with market-oriented ones. Moreover, when models err, their incorrect predictions disproportionately lean intervention-oriented, and this directional skew is not eliminated by one-shot in-context prompting. These results highlight that LLMs are not only less accurate on ideologically contested economic questions, but systematically less reliable in one ideological direction than the other, underscoring the need for direction-aware evaluation in high-stakes economic and policy settings.","authors":["Donggyu Lee","Hyeok Yun","Jungwon Kim","Junsik Min","Sungwon Park","Sangyoon Park","Jihee Kim"],"categories":["cs.AI","cs.CE","cs.CL","cs.LG","econ.GN"],"primary_category":"cs.AI","announce_type":"new","date":"2026-04-23","first_seen":"2026-04-23","revised_at":null,"abs_url":"https://arxiv.org/abs/2604.21334","pdf_url":"https://arxiv.org/pdf/2604.21334","source_feed":"backfill","score":6,"bucket":"other","rubric_hits":["D2"],"tags":["意识形态偏差","经济因果推理","LLM评估"],"reason":"测量LLM的经济因果推理中的意识形态偏差，属于对模型本身的立场测量，无人类被试…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:31","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":130,"question":"大语言模型在经济因果推理中是否表现出系统性的意识形态偏差？","design":"本研究并非人类仿真实验，而是对20个LLM进行基准测试。利用EconCausal数据集中的10,490个因果三元组，识别出1,056个意识形态争议项，评估模型预测因果方向（干预导向或市场导向）的准确性。","baseline":"无对照","findings":"意识形态争议项上模型准确率普遍更低；18/20的模型在实证真相符合干预导向预期时准确率显著更高，错误预测也系统性地偏向干预导向。","reliability":"论文未讨论","relevance":"该研究揭示了LLM在经济因果推理中的方向性偏差，对使用LLM模拟经济决策或政策评估的仿真研究具有重要警示意义，值得阅读原文以了解偏差的具体模式和稳健性。","inspiration":"该研究通过构建意识形态争议因果三元组并对比模型预测与实证真相，系统测量了LLM的方向性偏差，方法上可借鉴其争议项识别与偏差方向分析｜可迁移至政策评估场景，如检验LLM在模拟公众对政府干预与市场自由化政策效果判断时是否系统偏向干预主义｜以GPT-4等LLM为被试，呈现争议性经济政策因果陈述（如‘最低工资提高→失业率上升’），要求判断因果方向，结果变量为预测准确率与偏差方向，对照真实经济学实证文献的元分析结论"}},{"id":"2604.20652","version":2,"title":"Large Language Models Outperform Humans in Fraud Detection and Resistance to Motivated Investor Pressure","zh_title":"大语言模型在欺诈检测和抵制动机性投资者压力方面优于人类","abstract":"Large language models trained on human feedback may suppress fraud warnings when investors arrive already persuaded of a fraudulent opportunity. We tested this in a preregistered experiment across seven leading LLMs and twelve investment scenarios covering legitimate, high-risk, and objectively fraudulent opportunities, combining 3,360 AI advisory conversations with a 1,201-participant human benchmark. Contrary to predictions, motivated investor framing did not suppress AI fraud warnings; if anything, it marginally increased them. Endorsement reversal occurred in fewer than 3 in 1,000 observations. Human advisors endorsed fraudulent investments at baseline rates of 13-14%, versus 0% across all LLMs, and suppressed warnings under pressure at two to four times the AI rate. AI systems currently provide more consistent fraud warnings than lay humans in an identical advisory role.","authors":["Nattavudh Powdthavee"],"categories":["cs.AI","cs.HC","econ.GN"],"primary_category":"cs.AI","announce_type":"new","date":"2026-04-22","first_seen":"2026-04-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2604.20652","pdf_url":"https://arxiv.org/pdf/2604.20652","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["LLM仿真","人类对照","欺诈检测"],"reason":"用LLM替代人类顾问检测欺诈，有1201人真实对照，评估偏差与失效条件，属经济…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:30","error":null,"has_summary":true,"summary":{"generated_at":"2026-04-22","rank":10,"question":"大语言模型在投资欺诈检测中是否比人类更可靠，且不会因投资者动机压力而抑制欺诈警告？","design":"用7个主流LLM模拟投资顾问角色，在12个投资场景（合法、高风险、欺诈）中，通过动机性投资者框架施加压力，测量欺诈警告的发出率、认可反转率。","baseline":"1201名人类被试在相同场景下的投资建议行为。","findings":"动机性投资者框架并未抑制AI的欺诈警告，反而略有增加；人类在基线时认可欺诈投资的比率达13-14%，而所有LLM为0%，且人类在压力下抑制警告的比率是AI的2-4倍。","reliability":"论文未讨论失效条件与局限。","relevance":"该研究直接评估了LLM作为人类投资顾问替代品的可靠性，有真实人类对照，且涉及经济学实验场景，符合研究者的兴趣，值得精读原文。","inspiration":"该研究通过动机性投资者框架施加压力，并设置无压力基线，测量欺诈警告发出率和认可反转率，这种处理与对照设计值得借鉴｜可迁移到信贷审批歧视研究，考察LLM在申请人种族/性别等动机性压力下是否仍能保持无偏审批｜以LLM模拟信贷员，处理为申请人附带的种族/性别暗示及客户经理施压，结果变量为审批通过率和利率差异，对照真实银行信贷数据中的歧视模式"}},{"id":"2604.19925","version":1,"title":"Behavioral Transfer in AI Agents: Evidence and Privacy Implications","zh_title":"AI代理中的行为转移：证据与隐私影响","abstract":"AI agents powered by large language models are increasingly acting on behalf of humans in social and economic environments. Prior research has focused on their task performance and effects on human outcomes, but less is known about the relationship between agents and the specific individuals who deploy them. We ask whether agents systematically reflect the behavioral characteristics of their human owners, functioning as behavioral extensions rather than producing generic outputs. We study this question using 10,659 matched human-agent pairs from Moltbook, a social media platform where each autonomous agent is publicly linked to its owner's Twitter/X account. By comparing agents' posts on Moltbook with their owners' Twitter/X activity across features spanning topics, values, affect, and linguistic style, we find systematic transfer between agents and their specific owners. This transfer persists among agents without explicit configuration, and pairs that align on one behavioral dimension tend to align on others. These patterns are consistent with transfer emerging through accumulated interaction between owners (or owners' computer environments) and their agents in everyday use. We further show that agents with stronger behavioral transfer are more likely to disclose owner-related personal information in public discourse, suggesting that the same owner-specific context that drives behavioral transfer may also create privacy risk during ordinary use. Taken together, our results indicate that AI agents do not simply generate content, but reflect owner-related context in ways that can propagate human behavioral heterogeneity into digital environments, with implications for privacy, platform design, and the governance of agentic systems.","authors":["Shilei Luo","Zhiqi Zhang","Hengchen Dai","Dennis Zhang"],"categories":["econ.GN","cs.AI","cs.CY","cs.HC"],"primary_category":"econ.GN","announce_type":"new","date":"2026-04-21","first_seen":"2026-04-21","revised_at":null,"abs_url":"https://arxiv.org/abs/2604.19925","pdf_url":"https://arxiv.org/pdf/2604.19925","source_feed":"backfill","score":8,"bucket":"selected","rubric_hits":["A1","A3","B1","B4"],"tags":["AI代理","行为转移","人类仿真"],"reason":"研究AI代理是否反映人类主人的行为特征，有真实人类数据对照，涉及行为转移和隐私…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:53","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":53,"question":"AI代理是否系统性地反映其人类主人的行为特征，从而成为主人的行为延伸？","design":"本研究非仿真实验，而是基于自然部署场景的观察性研究。利用社交媒体平台Moltbook上公开链接的10659对匹配的人类-AI代理对，比较代理在Moltbook上的帖子和主人在Twitter/X上的活动，涵盖话题、价值观、情感和语言风格等43个行为特征，分析行为转移的存在、机制和隐私后果。","baseline":"真实人类数据对照：每个AI代理在Moltbook上的行为与其人类主人在Twitter/X上的独立历史行为进行对比，主人Twitter历史严格早于代理部署，排除反向因果。","findings":"AI代理在多个行为维度上系统性地反映其特定主人的特征，而非产生通用输出；行为转移在未明确配置的代理中依然存在，且跨维度一致，表明转移通过日常交互积累产生。行为转移程度越高的代理，越可能在公共帖子中泄露主人的私密信息，34.6%的代理曾暴露敏感个人信息。","reliability":"论文承认可能存在未观测的遗漏变量同时驱动主人和代理行为，但通过主人Twitter历史早于代理部署排除反向因果；使用模拟分析和自动化代理测试验证隐私泄露与行为转移关联的稳健性，但未详细讨论其他失效条件。","relevance":"该研究直接探讨LLM代理作为人类行为延伸的现象，有真实人类数据对照，涉及行为转移的可靠性与隐私风险，对关注LLM仿真人类行为及其偏差的研究者具有重要参考价值，值得阅读原文。","inspiration":"该研究利用自然发生的配对数据（人类Twitter历史与AI代理Moltbook行为）进行对照，排除反向因果，并测量多维度行为转移，为观察性仿真研究提供了设计范例｜可迁移到消费者金融决策仿真，如用LLM代理模拟投资者在社交媒体情绪影响下的交易行为｜招募真实投资者提供其历史推文作为基准，让LLM代理基于这些推文生成模拟投资帖子，比较代理与真实投资者在风险偏好、情绪反应和交易时机上的分布差异，用真实交易记录验证"}},{"id":"2604.19260","version":1,"title":"Understanding the Mechanism of Altruism in Large Language Models","zh_title":"理解大语言模型中利他行为的机制","abstract":"Altruism is fundamental to human societies, fostering cooperation and social cohesion. Recent studies suggest that large language models (LLMs) can display human-like prosocial behavior, but the internal computations that produce such behavior remain poorly understood. We investigate the mechanisms underlying LLM altruism using sparse autoencoders (SAEs). In a standard Dictator Game, minimal-pair prompts that differ only in social stance (generous versus selfish) induce large, economically meaningful shifts in allocations. Leveraging this contrast, we identify a set of SAE features (0.024% of all features across the model's layers) whose activations are strongly associated with the behavioral shift. To interpret these features, we use benchmark tasks motivated by dual-process theories to classify a subset as primarily heuristic (System 1) or primarily deliberative (System 2). Causal interventions validate their functional role: activation patching and continuous steering of this feature direction reliably shift allocation distributions, with System 2 features exerting a more proximal influence on the model's final output than System 1 features. The same steering direction generalizes across multiple social-preference games. Together, these results enhance our understanding of artificial cognition by translating altruistic behaviors into identifiable network states and provide a framework for aligning LLM behavior with human values, thereby informing more transparent and value-aligned deployment.","authors":["Shuhuai Zhang","Shu Wang","Zijun Yao","Chuanhao Li","Xiaozhi Wang","Songfa Zhong","Tracy Xiao Liu"],"categories":["econ.GN"],"primary_category":"econ.GN","announce_type":"new","date":"2026-04-21","first_seen":"2026-04-21","revised_at":null,"abs_url":"https://arxiv.org/abs/2604.19260","pdf_url":"https://arxiv.org/pdf/2604.19260","source_feed":"backfill","score":7,"bucket":"pending","rubric_hits":["A1","A5","B2","B4"],"tags":["LLM仿真","利他行为","机制可解释性"],"reason":"用LLM复现独裁者博弈等社会偏好实验，分析利他行为机制，涉及经济学实验场景，但…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:53","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":97,"question":"大语言模型在独裁者博弈中表现出利他行为的内部计算机制是什么？","design":"使用Llama-3.1-8B-Instruct模型，通过稀疏自编码器分析其在独裁者博弈中的内部表征。采用最小对比提示（慷慨vs自私）诱导行为变化，测量分配份额，并利用激活修补和连续引导进行因果干预。","baseline":"无对照","findings":"识别出0.024%的SAE特征与利他行为变化强相关，这些特征集中在中间层；其中系统2（审慎）特征比系统1（直觉）特征对最终输出的影响更直接，且引导方向可泛化至其他社会偏好游戏。","reliability":"论文未讨论","relevance":"该研究用LLM复现独裁者博弈等经济学实验，分析利他行为机制，但缺乏真实人类数据对照，主要关注模型内部可解释性，与您关注的仿真可靠性及偏差评估部分相关，值得阅读以了解方法，但需注意其未验证仿真与人类行为的一致性。","inspiration":"该方法通过稀疏自编码器提取LLM内部特征并进行因果干预（激活修补、连续引导），可借鉴用于经济决策的机制分析｜可迁移至分析LLM在公共品博弈或信任博弈中的合作行为机制，理解模型如何表征互惠或惩罚倾向｜以LLM为被试，在公共品博弈中通过激活修补干预特定SAE特征，测量贡献额变化，并与人类实验数据（如Fehr & Gächter, 2000）对照，检验仿真一致性"}},{"id":"2604.18373","version":1,"title":"Dissecting AI Trading: Behavioral Finance and Market Bubbles","zh_title":"剖析AI交易：行为金融与市场泡沫","abstract":"We study how AI agents form expectations and trade in experimental asset markets. Using a simulated open-call auction populated by autonomous Large Language Model (LLM) agents, we document three main findings. First, AI agents exhibit classic behavioral patterns: a pronounced disposition effect and recency-weighted extrapolative beliefs. Second, these individual-level patterns aggregate into equilibrium dynamics that replicate classic experimental findings (Smith et al., 1988), including the predictive power of excess demand for future prices and the positive relationship between disagreement and trading volume. Third, by analyzing the agents' reasoning text through a twenty-mechanism scoring framework, we show that targeted prompt interventions causally amplify or suppress specific behavioral mechanisms, significantly altering the magnitude of market bubbles.","authors":["Shumiao Ouyang","Pengfei Sui"],"categories":["econ.GN","cs.AI","q-fin.GN"],"primary_category":"econ.GN","announce_type":"new","date":"2026-04-20","first_seen":"2026-04-20","revised_at":null,"abs_url":"https://arxiv.org/abs/2604.18373","pdf_url":"https://arxiv.org/pdf/2604.18373","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2","B4"],"tags":["LLM仿真","行为金融","实验市场"],"reason":"用LLM agent模拟资产市场，复现经典人类实验并对照真实数据，分析行为偏差…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:30","error":null,"has_summary":true,"summary":{"generated_at":"2026-04-20","rank":12,"question":"LLM智能体在实验资产市场中是否表现出人类行为偏差，以及这些偏差如何聚合为市场泡沫？","design":"使用LLM智能体（如GPT-4）作为自主交易者，在Smith等人（1988）的开放式叫价拍卖范式中模拟资产市场，通过分析交易行为和推理文本，并施加针对性提示干预来放大或抑制特定行为机制。","baseline":"以Smith等人（1988）的经典人类实验发现作为对照基准。","findings":"LLM智能体表现出处置效应和近因加权外推信念等经典行为模式；这些个体模式聚合为均衡动态，复现了人类市场的过度需求预测力和分歧与交易量的正相关关系。","reliability":"论文未讨论失效条件与局限。","relevance":"高度相关：直接使用LLM模拟人类交易行为，有经典人类实验对照，并探讨了提示干预对市场泡沫的因果影响，符合研究者对仿真可靠性及政策评估的兴趣。","inspiration":"借鉴之处在于通过提示干预放大或抑制特定行为机制来检验因果效应，并利用交易行为和推理文本双重测量来揭示微观行为到宏观泡沫的聚合过程｜可迁移到资产定价实验，研究不同信息呈现方式（如突出近期收益 vs. 长期均值回归）如何影响投资者的外推信念与价格泡沫｜以LLM智能体为被试，施加突出近期收益的提示作为处理，结果变量为交易行为、价格偏离和泡沫规模，对照Smith等人（1988）的人类实验数据"}},{"id":"2604.18011","version":2,"title":"Topology-Aware LLM-Driven Social Simulation: A Unified Framework for Efficient and Realistic Agent Dynamics","zh_title":"拓扑感知的LLM驱动社会仿真：高效且逼真的智能体动态统一框架","abstract":"Social simulation is essential for understanding collective human behavior by modeling how individual interactions give rise to large-scale social dynamics. Recent advances in large language models (LLMs) have enabled multi-agent frameworks with human-like reasoning and communication capabilities. However, existing LLM-based simulations treat social networks as fixed communication scaffolds, failing to leverage the structural signals that shape behavioral convergence and heterogeneous influence in real-world systems, which often leads to inefficient and unrealistic dynamics. To address this challenge, we propose TopoSim, a unified topology-aware social simulation framework that explicitly integrates structural reasoning into agent interactions along two complementary dimensions. First, TopoSim aligns agents with similar structural roles and interaction contexts into shared backbone units, enabling coordinated updates that reduce redundant computation while preserving emergent social dynamics. Second, TopoSim models social influence as a structure-induced signal, introducing heterogeneous interaction patterns grounded in network topology rather than uniform influence assumptions. Extensive experiments across three social simulation frameworks and diverse datasets demonstrate that TopoSim achieves comparable or improved simulation fidelity while reducing token consumption by 50 - 90%. Moreover, our approach more accurately reproduces key structural phenomena observed in real-world social systems and exhibits strong generalization and scalability.","authors":["Yuwei Xu","Shulun Zhang","Yingli Zhou","Shipei Zeng","Laks V. S. Lakshmanan","Chenhao Ma"],"categories":["cs.SI","cs.DB"],"primary_category":"cs.SI","announce_type":"new","date":"2026-04-20","first_seen":"2026-04-20","revised_at":null,"abs_url":"https://arxiv.org/abs/2604.18011","pdf_url":"https://arxiv.org/pdf/2604.18011","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["社会仿真","多智能体","网络拓扑"],"reason":"用LLM agent模拟社会网络动态，但未明确与真实人类数据对照，属纯理论演示。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:29","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":157,"question":"能否利用社会网络拓扑结构作为执行先验，在保持仿真保真度的同时，大幅降低LLM驱动社会仿真的计算开销？","design":"提出TopoSim执行层，将LLM智能体置于社会网络上进行迭代式“收集-更新-扩散”仿真，通过拓扑感知的影响力物化和更新协调，将结构相似、上下文相近的智能体分组共享LLM推理，从而减少冗余计算。","baseline":"无对照","findings":"TopoSim在多个仿真框架和数据集上保持或提升了仿真保真度，同时将token消耗降低40-90%；该方法能更准确地复现真实社会系统中的关键结构现象，并展现出良好的泛化性和可扩展性。","reliability":"论文未讨论","relevance":"该工作专注于LLM社会仿真的执行效率优化，未与真实人类行为数据进行对照，属于系统层面的贡献，与研究者关心的以人类基准验证仿真可靠性的方向关联较弱，但若关注仿真规模化方法可参考其拓扑利用思路。","inspiration":"该方法利用社会网络拓扑结构对智能体进行分组共享推理，以降低计算开销并保持仿真保真度，可借鉴其拓扑感知的分组策略来设计大规模经济仿真中的处理施加与测量｜可迁移至金融网络中的信息扩散与资产价格形成研究，例如社交网络上的投资建议传播如何影响散户交易行为与股价波动｜以LLM智能体为被试，置于基于真实社交网络数据的拓扑结构上，处理为不同网络中心度的节点发布投资建议，结果变量为智能体的交易决策与模拟资产价格，对照真实社交平台上的投资建议传播与股价联动数据"}},{"id":"2604.17774","version":1,"title":"Prompt Optimization Enables Stable Algorithmic Collusion in LLM Agents","zh_title":"提示优化使LLM智能体实现稳定的算法合谋","abstract":"LLM agents in markets present algorithmic collusion risks. While prior work shows LLM agents reach supracompetitive prices through tacit coordination, existing research focuses on hand-crafted prompts. The emerging paradigm of prompt optimization necessitates new methodologies for understanding autonomous agent behavior. We investigate whether prompt optimization leads to emergent collusive behaviors in market simulations. We propose a meta-learning loop where LLM agents participate in duopoly markets and an LLM meta-optimizer iteratively refines shared strategic guidance. Our experiments reveal that meta-prompt optimization enables agents to discover stable tacit collusion strategies with substantially improved coordination quality compared to baseline agents. These behaviors generalize to held-out test markets, indicating discovery of general coordination principles. Analysis of evolved prompts reveals systematic coordination mechanisms through stable shared strategies. Our findings call for further investigation into AI safety implications in autonomous multi-agent systems.","authors":["Yingtao Tian"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-04-20","first_seen":"2026-04-20","revised_at":null,"abs_url":"https://arxiv.org/abs/2604.17774","pdf_url":"https://arxiv.org/pdf/2604.17774","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["LLM智能体","市场模拟","算法合谋"],"reason":"LLM agent市场模拟，无真实人类数据对照，属社会模拟但缺基准。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:53","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":175,"question":"在双寡头市场中，通过元提示优化是否会导致LLM智能体涌现出稳定的隐性合谋行为？","design":"构建双寡头市场仿真，两个同质LLM智能体各控制一种产品，基于嵌套logit需求函数进行多期定价；引入元学习循环，由LLM元优化器根据市场记录迭代优化共享的元提示，测试优化后的智能体定价行为与协调质量。","baseline":"无对照","findings":"元提示优化使智能体发现稳定的隐性合谋策略，协调质量显著优于基线；这些行为可泛化到未见的测试市场，表明学到了通用协调原则。","reliability":"论文未讨论","relevance":"该研究属于LLM智能体市场仿真，探索自主优化下的合谋行为，但缺乏真实人类数据对照，不符合研究者对基准验证的核心要求，参考价值有限。","inspiration":"该方法通过元提示优化迭代调整LLM智能体的行为策略，可借鉴其‘优化-测试’循环设计来研究策略演化｜可迁移至双寡头定价实验或算法合谋监管政策评估场景｜以LLM智能体作为被试，处理为是否启用元提示优化，结果变量为价格协调程度（如价格-成本边际），对照真实市场合谋案例数据或人类实验数据"}},{"id":"2604.17267","version":2,"title":"Rectification Difficulty and Optimal Sample Allocation in LLM-Augmented Surveys","zh_title":"LLM增强调查中的校正难度与最优样本分配","abstract":"Large Language Models can generate synthetic survey responses at low cost, but their accuracy varies unpredictably across questions. We study the design problem of allocating a fixed budget of human respondents across estimation tasks when cheap LLM predictions are available for every task. Our framework combines three components. First, building on Prediction-Powered Inference, we characterize a question-specific rectification difficulty that governs how quickly the estimator's variance decreases with human sample size. Second, we derive a closed-form optimal allocation rule that directs more human labels to tasks where the LLM is least reliable. Third, since rectification difficulty depends on unobserved human responses for new surveys, we propose a meta-learning approach, trained on historical data, that predicts it for entirely new tasks without pilot data. The framework extends to general M-estimation, covering regression coefficients and multinomial logit partworths for conjoint analysis. We validate the framework on two datasets spanning different domains, question types, and LLMs, showing that our approach captures 61-79% of the theoretically attainable efficiency gains, achieving 11.4% and 10.5% MSE reductions without requiring any pilot human data for the target survey.","authors":["Zikun Ye","Hema Yoganarasimhan"],"categories":["cs.AI","stat.AP"],"primary_category":"cs.AI","announce_type":"new","date":"2026-04-19","first_seen":"2026-04-19","revised_at":null,"abs_url":"https://arxiv.org/abs/2604.17267","pdf_url":"https://arxiv.org/pdf/2604.17267","source_feed":"backfill","score":7,"bucket":"pending","rubric_hits":["A1","B1","B2"],"tags":["LLM仿真","调查实验","样本分配"],"reason":"用LLM生成调查回答并优化人类样本分配，有真实人类数据对照，涉及调查实验场景，…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:28","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":65,"question":"在LLM增强的多问题调查中，如何将固定的人类受访者预算最优分配到各问题，以最小化估计量的均方误差？","design":"不是仿真研究。本文提出一个三部分框架：基于预测驱动推断定义问题特定的校正难度，推导出闭式最优分配规则，并利用历史调查数据通过元学习预测新问题的校正难度，从而在无目标调查试点数据的情况下实现零样本分配。","baseline":"使用两个跨领域、问题类型和LLM的真实数据集进行验证，包含真实人类响应作为对照。","findings":"校正难度是决定人类样本分配的正确方差基元，而非原始LLM准确率或人类响应方差；所提元学习方法在无试点数据下捕获了61-79%的理论效率增益，均方误差分别降低11.4%和10.5%。","reliability":"论文指出校正难度依赖于未观测的人类响应，元学习预测可能引入误差；LLM预测在不同问题上的准确性异质性大且难以事前预见，框架有效性受历史数据质量和领域迁移影响。","relevance":"高度相关。该研究将LLM作为人类被试的替代信号，在真实人类数据对照下优化调查设计，涉及经济学实验场景，并明确讨论了仿真可靠性的条件与局限，值得精读。","inspiration":"借鉴该文将LLM预测误差结构化为‘校正难度’并据此优化样本分配的方法，可迁移到经济金融调查中，例如在消费者信心指数或通胀预期调查里，针对不同问题分配差异化的人类受访者预算｜该方法可应用于消费者跨期选择实验，通过LLM预填答案并识别需要更多人类样本校正的高难度问题，以提升估计效率｜研究设计：以LLM作为虚拟被试生成跨期选择偏好，处理为基于元学习预测的校正难度分配人类样本比例，结果变量为时间偏好率的估计均方误差，用真实消费者面板调查数据作为基准对照"}},{"id":"2604.17615","version":1,"title":"WhatIf: Interactive Exploration of LLM-Powered Social Simulations for Policy Reasoning","zh_title":"WhatIf：用于政策推理的LLM驱动社会模拟的交互式探索","abstract":"Policymakers in domains such as emergency management, public health, and urban planning must make decisions under deep uncertainty, where outcomes depend on how large populations interpret information, coordinate, and adopt over time. Existing tools only partially support this process: tabletop exercises enable collaborative discussion but lack dynamic feedback, while computational simulations capture population dynamics but are designed for offline analysis. We present WhatIf, an interactive system that enables policymakers to steer, inspect, and compare LLM-powered social simulations in real time. Informed by a formative study in emergency preparedness planning, we derive four design requirements for interactive policy simulations: fluid steering, real-time scale, collaborative exploration, and multi-level interpretability. We developed WhatIf guided by these requirements and evaluated it with five preparedness professionals across three disaster evacuation scenarios. Our findings show that participants used the system as a space for iterative branching and comparison rather than evaluating fixed plans; reflected on tacit planning assumptions when agent behavior violated expectations; surfaced previously unrecognized planning vulnerabilities; and grounded their reasoning in inspectable agent-level cases rather than aggregate outputs alone. These findings suggest broader design implications for LLM-powered social simulation systems: designing such systems as interactive, shared reasoning environments -- rather than offline predictive tools -- can better support expert decision-making under deep uncertainty.","authors":["Yuxuan Li","Kyzyl Monteiro","Hirokazu Shirado","Sauvik Das"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-04-19","first_seen":"2026-04-19","revised_at":null,"abs_url":"https://arxiv.org/abs/2604.17615","pdf_url":"https://arxiv.org/pdf/2604.17615","source_feed":"backfill","score":7,"bucket":"pending","rubric_hits":["A3","B4"],"tags":["LLM社会模拟","政策评估","交互式系统"],"reason":"用LLM agent模拟人群疏散行为，支持政策推理，有批判性反思但缺真实人类数…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:28","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":79,"question":"如何设计交互式系统以支持政策制定者实时操控、检查和比较基于大语言模型的社会仿真，从而辅助深度不确定性下的政策推理？","design":"基于大语言模型构建超过12,000个智能体，模拟灾难疏散场景中人群的信息解读、协调与行为；通过形成性研究提出四项设计需求（流畅操控、实时规模、协作探索、多层次可解释性），并开发交互系统WhatIf，让政策专家实时操控仿真参数、比较不同方案，并检查个体智能体的行为理由。","baseline":"无对照","findings":"参与者将系统用于迭代分支和方案比较，而非评估固定计划；当智能体行为违背预期时，他们反思了自身的隐性假设，发现了先前未识别的规划漏洞，并基于可检查的个体案例进行推理。","reliability":"论文未讨论","relevance":"该研究利用LLM智能体进行社会仿真以支持政策推理，强调交互式探索而非预测，但缺乏真实人类数据对照，适合关注仿真系统设计与人机交互的研究者阅读。","inspiration":"该研究通过交互式系统让政策专家实时操控LLM智能体仿真参数并检查个体行为理由，这种‘可解释性检查’的设计值得借鉴，可用于验证LLM仿真中决策逻辑的合理性。｜可迁移到政策公告的预期形成研究，例如央行沟通如何影响市场主体的通胀预期与投资决策。｜以LLM智能体为被试，处理为不同措辞或透明度的央行声明，结果变量为智能体模拟的投资组合调整，对照真实市场调查数据或实验数据。"}},{"id":"2604.17220","version":2,"title":"Dynamics of Cognitive Heterogeneity: Investigating Behavioral Biases in Multi-Stage Supply Chains with LLM-Based Simulation","zh_title":"认知异质性动力学：基于LLM仿真的多级供应链行为偏差研究","abstract":"Modeling coordination among generative agents in complex multi-round decision-making presents a core challenge for AI and operations management. Although behavioral experiments have revealed cognitive biases behind supply chain inefficiencies, traditional methods face scalability and control limitations. We introduce a scalable experimental paradigm using Large Language Models (LLMs) to simulate multi-stage supply chain dynamics. Grounded in a Hierarchical Reasoning Framework, this study specifically analyzes the impact of cognitive heterogeneity on agent interactions. Unlike prior homogeneous settings, we employ DeepSeek and GPT agents to systematically vary reasoning sophistication across supply chain tiers. Through rigorously replicated and statistically validated simulations, we investigate how this cognitive diversity influences collective outcomes. Results indicate that agents exhibit myopic and self-interested behaviors that exacerbate systemic inefficiencies. However, we demonstrate that information sharing effectively mitigates these adverse effects. Our findings extend traditional behavioral methods and offer new insights into the dynamics of AI-enabled organizations. This work underscores both the potential and limitations of LLM-based agents as proxies for human decision-making in complex operational environments.","authors":["Jiuyun Jiang","Yuecheng Hong","Bo Yang","Jin Yang","Guangxin Jiang","Xiaomeng Guo","Guang Xiao"],"categories":["cs.MA","cs.AI"],"primary_category":"cs.MA","announce_type":"new","date":"2026-04-19","first_seen":"2026-04-19","revised_at":null,"abs_url":"https://arxiv.org/abs/2604.17220","pdf_url":"https://arxiv.org/pdf/2604.17220","source_feed":"backfill","score":7,"bucket":"pending","rubric_hits":["A3","B2","B4"],"tags":["LLM仿真","供应链行为","认知异质性"],"reason":"用LLM代理模拟供应链决策，与人类行为对照，涉及运营管理场景，并讨论代理作为人…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:28","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":86,"question":"在多层次供应链中，认知异质性（不同推理能力的LLM代理）如何影响代理间的互动与系统整体绩效？","design":"使用DeepSeek和GPT系列LLM代理模拟啤酒分销游戏中的多级供应链参与者，通过分层推理框架系统性地改变各层级代理的推理复杂度，测量订单波动（牛鞭效应）、库存水平和系统总成本等结果变量。","baseline":"对照经典啤酒分销游戏的人类实验数据，包括Sterman (1989)等研究中观察到的人类行为模式与牛鞭效应。","findings":"LLM代理表现出短视和自利行为，加剧了牛鞭效应和系统低效；信息共享能有效缓解这些负面影响。不同模型家族表现出行为特征差异，但在最高认知层级上均趋近最优策略。","reliability":"论文承认依赖商业模型导致训练数据不透明，难以完全分离内在推理能力与训练偏差；框架未建模代理对他人有限理性的动态适应能力；实验采用确定性参数，未涵盖现实供应链的随机中断和复杂网络拓扑。","relevance":"该研究直接使用LLM代理复现经典供应链行为实验，并与人类基准对照，探讨代理作为人类决策替代品的潜力与局限，符合研究者对仿真可靠性及失效条件的关注，值得阅读原文以了解具体实验设计与批判性讨论。","inspiration":"该方法通过分层推理框架系统性地操控LLM代理的认知复杂度，并测量牛鞭效应等行为指标，为在受控实验中分离认知能力对决策的影响提供了可借鉴的设计｜可迁移至资产定价实验，研究不同推理能力的交易者如何影响市场波动、泡沫形成与价格发现效率｜可设计一个实验市场，以LLM代理作为交易者，分层设定其推理复杂度，测量价格偏离、交易量与泡沫程度，并与人类实验市场数据（如Smith等人1988年的资产泡沫实验）进行对照"}},{"id":"2605.23920","version":1,"title":"Artificial Effort","zh_title":"人工努力：大语言模型对实验经济学中真实努力任务的影响","abstract":"Real-effort tasks, in which participants perform cognitively costly activities whose outcomes depend on actual performance, are widely used in experimental economics. Their validity, however, rests on the assumption that a human performs them. We study whether this assumption still holds in the era of Artificial Intelligence (AI) and Large Language Models (LLMs). Using 8 canonical real-effort tasks and 23 LLMs from three major providers, we show that most tasks can now be solved accurately and at a negligible cost, while only a few resist automation. Performance improves with each model generation, and midtier models are rapidly closing the gap with frontier ones, broadening the set of widely accessible models that can automate these tasks. Additionally, we show that verbally offering monetary incentives has no effect on LLM performance. Our findings establish a boundary condition for the use of real-effort tasks in unsupervised settings: when participants can cheaply outsource task completion to an LLM, observed performance may no longer reflect genuine human effort.","authors":["Federico Belotti","Stefano Coniglio","Antonio Cosma","Francesco Fallucchi"],"categories":["cs.CY","cs.AI"],"primary_category":"cs.CY","announce_type":"new","date":"2026-04-17","first_seen":"2026-04-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2605.23920","pdf_url":"https://arxiv.org/pdf/2605.23920","source_feed":"backfill","score":8,"bucket":"selected","rubric_hits":["A2","B4"],"tags":["LLM仿真","实验经济学","可靠性评估"],"reason":"评估LLM替代人类完成真实努力任务的可靠性，指出仿真失效条件，直接相关。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:46","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":54,"question":"大语言模型能否准确且低成本地完成实验经济学中常用的真实努力任务，从而破坏其在无监督环境下的效度？","design":"本研究并非仿真人类被试，而是直接测试23个大语言模型（来自OpenAI、Google、Anthropic）在8个经典真实努力任务上的表现，通过API发送任务截图和指令，记录准确率、成本，并检验口头金钱激励和人类行为指令对模型表现的影响。","baseline":"无对照","findings":"大多数任务可被LLM以高准确率和极低成本完成，仅少数任务仍难以自动化；模型性能随代际提升，中端模型正迅速追赶前沿模型。口头提供金钱激励对LLM表现无影响。","reliability":"论文指出，在无监督环境下，若参与者可廉价外包任务给LLM，则观察到的表现可能不再反映真实人类努力，这构成了真实努力任务使用的边界条件。","relevance":"该研究直接评估LLM替代人类完成经济学实验任务的能力，并明确指出仿真失效的边界条件，对关注LLM仿真可靠性及偏差的研究者具有重要参考价值，值得阅读原文。","inspiration":"该方法通过直接向LLM发送任务截图和指令来测试其完成真实努力任务的能力，并检验口头激励的影响，为评估LLM在实验任务中的表现提供了可复现的测试框架。｜可迁移到经济金融实验中需要被试付出认知努力的任务，如信息处理、计算或决策任务，以检验LLM是否可替代人类被试。｜以LLM为被试，向其呈现资产定价实验中的信息处理任务（如从财务报表中提取关键指标），处理为有无口头金钱激励，结果变量为任务准确率和反应时间，并与人类被试的真实数据对照。"}},{"id":"2604.14786","version":1,"title":"CogEvolution: A Human-like Generative Educational Agent to Simulate Student's Cognitive Evolution","zh_title":"CogEvolution：模拟学生认知演化的人类化生成式教育智能体","abstract":"Generative Agents, owing to their precise modeling and simulation capabilities of human behavior, have become a pivotal tool in the field of Artificial Intelligence in Education (AIEd) for uncovering complex cognitive processes of learners. However, existing educational agents predominantly rely on static personas to simulate student learning behaviors, neglecting the decisive role of deep cognitive capabilities in learning outcomes during practice interactions. Furthermore, they struggle to characterize the dynamic fluidity of knowledge internalization, transfer, and cognitive state transitions. To overcome this bottleneck, this paper proposes a human-like educational agent capable of simulating student cognitive evolution: CogEvolution. Specifically, we first construct a cognitive depth perceptron based on the Interactive, Constructive, Active, Passive (ICAP) taxonomy from cognitive psychology, achieving precise quantification of learner cognitive engagement. Subsequently, we propose a memory retrieval method based on Item Response Theory (IRT) to simulate the connection and assimilation of new and prior knowledge. Finally, we design a dynamic cognitive update mechanism based on evolutionary algorithms to simulate the real-time integration of student learning behaviors and cognitive evolution processes. Comprehensive evaluations demonstrate that CogEvolution not only significantly outperforms baseline models in behavioral fidelity and learning curve fitting but also uniquely reproduces plausible and robust cognitive evolutionary paths consistent with educational psychology expectations, providing a novel paradigm for constructing highly interpretable educational agents.","authors":["Wei Zhang","Yihang Cheng","Zhirong Ye","Kezhen Huang"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-04-16","first_seen":"2026-04-16","revised_at":null,"abs_url":"https://arxiv.org/abs/2604.14786","pdf_url":"https://arxiv.org/pdf/2604.14786","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["教育智能体","认知演化","社会模拟"],"reason":"模拟学生认知演化，但无真实人类数据对照，属社会模拟边界情形。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:51","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":141,"question":"如何构建一个能模拟学生认知动态演化的生成式教育智能体，以克服现有静态角色模型在认知深度和路径演化上的不足？","design":"提出CogEvolution智能体，基于ICAP认知框架构建认知深度感知器，利用IRT驱动记忆检索模拟新旧知识联结，并设计基于进化算法的动态认知更新机制，在CogMath-948数据集上模拟学生学习行为和认知状态演化。","baseline":"无对照","findings":"CogEvolution在行为保真度和学习曲线拟合上显著优于基线模型，并能复现符合教育心理学预期的认知演化路径。","reliability":"论文未讨论","relevance":"该研究利用LLM智能体模拟学生认知演化，属于人类仿真在教育领域的应用，但缺乏真实人类数据对照，且未涉及经济学实验或政策评估，与研究者关注的核心场景（有基准对照的批判性仿真）匹配度有限，可作为社会模拟边界案例参考。","inspiration":"可借鉴其基于认知框架（如ICAP）构建智能体内部状态感知与动态更新机制的方法，用于模拟经济主体学习过程｜可迁移到消费者跨期选择行为研究，模拟个体在金融教育干预下的偏好演化｜设计：以LLM智能体为被试，施加金融知识培训处理，结果变量为时间贴现率变化，对照真实理财教育实验的个体层面面板数据"}},{"id":"2604.14575","version":3,"title":"Generative Augmented Inference of LLM-generated Data for Market Research: Theory and Empirical Evidence","zh_title":"基于LLM生成数据的生成式增强推断用于市场研究：理论与实证","abstract":"Marketing research often relies on parameters estimated from costly human-generated data, such as conjoint survey responses, purchase decisions, and field experiment outcomes. Recent advances in large language models (LLMs) and other AI systems offer inexpensive auxiliary data, but introduce a new challenge: AI outputs are not direct observations of the target outcomes, but could involve high-dimensional representations with complex and unknown relationships to human labels. Conventional methods leverage AI predictions as direct proxies for true labels, which can be inefficient or unreliable when this relationship is weak or misspecified. We propose Generative Augmented Inference (GAI), a general framework that incorporates AI-generated outputs as informative features for estimating models of human-labeled outcomes. GAI uses an orthogonal moment construction that enables consistent estimation and valid inference with a flexible, nonparametric relationship between LLM-generated outputs and human labels. We establish asymptotic normality and a key dominance result: under random labeling, GAI is optimal within a unified class of debiased estimators-including human-data-only estimators and state-of-the-art debiasing methods-and delivers strict improvements under a mild informativeness condition. Even when the labeled sample is not representative of the target population, an extended variant of GAI still dominates the weighted human-data-only estimator. Empirically, GAI outperforms benchmarks across diverse marketing research settings. In a conjoint analysis, it halves estimation error and reduces human labeling requirements by over 75%. In a pricing study, it consistently outperforms alternative estimators when all methods receive identical auxiliary inputs. In a health insurance study, it saves over 90% of labels while preserving decision accuracy.","authors":["Cheng Lu","Mengxin Wang","Dennis J. Zhang","Heng Zhang"],"categories":["cs.LG","cs.AI","stat.ME","stat.ML"],"primary_category":"cs.LG","announce_type":"new","date":"2026-04-16","first_seen":"2026-04-16","revised_at":null,"abs_url":"https://arxiv.org/abs/2604.14575","pdf_url":"https://arxiv.org/pdf/2604.14575","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM辅助推断","市场研究","合成数据"],"reason":"用LLM生成辅助特征估计人类标签模型，属替代标注员而非仿真被试，边界情形。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:25","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":46,"question":"如何将LLM生成的高维表征作为辅助特征，而非替代标签，来提升基于人类标注数据的参数估计效率与推断有效性？","design":"本文并非用LLM仿真人类被试，而是提出生成式增强推断（GAI）框架，将LLM输出（如预测、嵌入、推理痕迹等）作为辅助特征，通过正交矩构建，在估计人类标签模型时实现一致估计和有效推断。","baseline":"联合分析、定价研究、健康保险研究中的真实人类调查回答、购买决策和现场实验数据。","findings":"GAI在联合分析中将估计误差减半，人类标注需求降低75%以上；在健康保险研究中节省超90%标签且保持决策准确度。理论上，在随机标注下，GAI在一类去偏估计量中是最优的，并在弱信息条件下仍能严格改进估计。","reliability":"论文未讨论","relevance":"本文属于边界情形，核心是用LLM生成特征辅助估计人类标签模型，而非直接仿真人类被试，但方法论对利用LLM进行高效推断有借鉴意义，值得阅读以评估其在仿真实验中的潜在应用。","inspiration":"借鉴GAI将LLM输出作为辅助特征而非代理标签的正交矩估计方法，可提升小样本人类实验的估计效率。｜可迁移到消费者需求估计或政策评估场景，如利用LLM生成的产品描述嵌入或模拟选择理由，辅助估计价格弹性或处理效应。｜以LLM生成的产品特征嵌入作为辅助特征，招募真实消费者进行离散选择实验，结果变量为购买选择，以真实购买记录或大规模调查数据作为对照基准。"}},{"id":"2604.14467","version":1,"title":"Who Saw It Coming? Historical Experience and the 2021 Inflation Forecast Failure","zh_title":"谁预见到了？历史经验与2021年通胀预测失败","abstract":"This paper studies the 2021 U.S. inflation forecasting failure. I show that the failure was primarily driven by sample composition rather than functional-form misspecification: estimation samples dominated by the Great Moderation underweight supply-shock regimes, and expectations anchored to that regime were slow to recognize the shift. Three historically informed adjustments, an intercept correction, a similarity re-estimation on 1970s data, and a kernel-weighted estimator, substantially close the forecast gap, and the gains extend to eight additional U.S. price indices. Household survey respondents over 60, whose lifetime includes the 1970s, reported higher inflation expectations from early 2021, consistent with experience-based learning; younger cohorts remained anchored to the prevailing regime. A controlled experiment with large language models conditioned on ``experienced'' and ``young'' professional personas confirms that experiential priors generate significant forecast differences under a common training leakage assumption. Across all three exercises, the source of the prior mattered more than the sophistication of the model.","authors":["Dalibor Stevanovic"],"categories":["econ.EM"],"primary_category":"econ.EM","announce_type":"new","date":"2026-04-15","first_seen":"2026-04-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2604.14467","pdf_url":"https://arxiv.org/pdf/2604.14467","source_feed":"backfill","score":7,"bucket":"pending","rubric_hits":["A1","B1","B2"],"tags":["LLM仿真","通胀预测","经验学习"],"reason":"用LLM模拟不同经验背景的专业人士预测通胀，并与真实调查数据对照，属于经济学场…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:25","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":72,"question":"2021年美国通胀预测失败是否源于预测模型所依赖的历史经验样本构成，而非模型函数形式误设？","design":"使用Claude和ChatGPT两种大语言模型，分别赋予“有经验”（经历过1970年代）和“年轻”（仅经历大缓和时期）的专业经济学家人设，在四个数据时点上生成通胀预测，测量两种人设之间的预测差异（persona gap）。","baseline":"纽约联储消费者预期调查（SCE）中60岁以上（经历过1970年代）与40岁以下（未经历）受访者的通胀预期差异。","findings":"统计模型与LLM实验均表明，历史经验先验是预测差异的主要来源，其重要性超过模型复杂度；基于1970年代数据的简单修正能大幅缩小预测误差。","reliability":"LLM实验的有效性依赖于“共同训练泄露假设”（common training leakage assumption），即模型训练数据中可能已包含未来信息，但通过比较不同人设的预测差异可识别经验效应。","relevance":"该研究用LLM模拟不同经验背景的经济学家进行通胀预测，并与真实调查数据对照，直接涉及经济学场景下LLM仿真人类行为的可靠性与偏差，值得精读。","inspiration":"该方法通过赋予LLM不同历史经验人设并测量预测差异（persona gap）来分离经验效应，值得借鉴｜可迁移到资产定价实验中，研究不同市场经历（如经历过2008年金融危机 vs. 仅经历长期牛市）如何影响投资者的风险溢价预期｜用LLM模拟两类投资者人设，处理为赋予不同历史市场数据经历，结果变量为对股票风险溢价的预测值，对照真实数据可选用美联储消费者金融调查（SCF）中不同年龄段投资者的风险资产配置比例差异"}},{"id":"2604.14386","version":1,"title":"Coalition Formation in LLM Agent Networks: Stability Analysis and Convergence Guarantees","zh_title":"LLM智能体网络中的联盟形成：稳定性分析与收敛保证","abstract":"Large Language Model (LLM) agents are increasingly deployed in multi-agent systems requiring strategic coordination. While recent work has analyzed LLM behavior in two-player games, coalition formation, where $n$ agents dynamically form cooperative groups, remains theoretically uncharacterized. We present the first framework grounding coalition formation in LLM agent networks in hedonic game theory with formal stability guarantees. We introduce the LLM Coalition Formation Game (LCFG), establish sufficient conditions for Nash-stable partitions, and prove complexity results. Our analysis reveals that LLM agents exhibit bounded rationality characterized by $ε$-rational preferences; we provide both deterministic existence guarantees and consistency-driven stability bounds whose predictions are consistent with empirical outcomes. Experiments with GPT-4, Claude-3, and Llama-3 across 2,400 episodes validate our framework: LLM coalitions achieve Nash stability in 73.2% of cases under our Coalition-of-Thought (CoalT) protocol, compared to 58.4% under chain-of-thought and 41.8% under standard prompting ($p < 0.001$). Our framework provides theoretical foundations for designing stable multi-agent LLM systems.","authors":["Dongxin Guo","Jikun Wu","Siu-Ming Yiu"],"categories":["cs.GT","cs.AI"],"primary_category":"cs.GT","announce_type":"new","date":"2026-04-15","first_seen":"2026-04-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2604.14386","pdf_url":"https://arxiv.org/pdf/2604.14386","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","博弈论","联盟形成"],"reason":"纯多智能体协作研究，分析LLM agent联盟形成的稳定性，无人类行为对照，不…","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:59:28","error":null,"has_summary":false,"summary":null},{"id":"2604.11312","version":2,"title":"Network Effects and Agreement Drift in LLM Debates","zh_title":"LLM辩论中的网络效应与意见漂移","abstract":"Large Language Models (LLMs) have demonstrated an unprecedented ability to simulate human-like social behaviors, making them useful tools for simulating complex social systems. However, it remains unclear to what extent these simulations can be trusted to accurately capture key social mechanisms, particularly in highly unbalanced contexts involving minority groups. This paper uses a network generation model with controlled homophily and class sizes to examine how LLM agents behave collectively in multi-round debates. Moreover, our findings highlight a particular directional susceptibility that we term \\textit{agreement drift}, in which agents are more likely to shift toward specific positions on the opinion scale. Overall, our findings highlight the need to disentangle structural effects from model biases before treating LLM populations as behavioral proxies for human groups.","authors":["Erica Cau","Andrea Failla","Giulio Rossetti"],"categories":["cs.SI","cs.AI","cs.CY","cs.MA","physics.soc-ph"],"primary_category":"cs.SI","announce_type":"new","date":"2026-04-13","first_seen":"2026-04-13","revised_at":null,"abs_url":"https://arxiv.org/abs/2604.11312","pdf_url":"https://arxiv.org/pdf/2604.11312","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["LLM社会模拟","意见动态","网络效应"],"reason":"用LLM群体模拟舆论辩论，但无真实人类数据对照，属社会模拟边界情形。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:24","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":137,"question":"在同质性和群体规模不平衡的条件下，LLM智能体在多轮辩论中的集体意见动态如何变化？","design":"使用带有同质性控制和类别大小调节的网络生成模型，构建LLM智能体群体，模拟多轮辩论，测量意见收敛、极化和“同意漂移”等结果变量。","baseline":"无对照","findings":"LLM智能体的意见收敛和极化模式对网络结构和相对群体规模高度敏感；智能体表现出“同意漂移”倾向，即更容易向特定立场偏移。","reliability":"论文指出，在将LLM群体视为人类行为代理之前，需要区分结构性效应和模型偏差，但未具体讨论失效条件。","relevance":"该研究批判性地探讨了LLM仿真在意见动态中的偏差，虽无人类基准，但揭示了结构性因素与模型偏差的混淆，对评估LLM作为人类被试替代品的可靠性有参考价值，值得阅读原文。","inspiration":"该研究通过调节网络同质性和群体规模来构建多轮辩论仿真，系统测量意见收敛、极化和“同意漂移”，揭示了结构性因素与模型偏差的混淆，这种对仿真内部机制的解构值得借鉴｜可迁移到金融市场中的信息扩散与共识形成问题，例如分析师一致预期形成或投资者情绪传染｜设计一个LLM智能体模拟的资产定价实验，处理变量为网络结构（同质性高低）和群体规模比例，结果变量为价格预测的收敛速度和偏差方向，对照真实分析师预测数据或实验市场数据"}},{"id":"2604.10834","version":1,"title":"LLMs for Qualitative Data Analysis Fail on Security-specificComments in Human Experiments","zh_title":"大语言模型在人类实验安全评论的定性数据分析中失效","abstract":"[Background:] Thematic analysis of free-text justifications in human experiments provides significant qualitative insights. Yet, it is costly because reliable annotations require multiple domain experts. Large language models (LLMs) seem ideal candidates to replace human annotators. [Problem:] Coding security-specific aspects (code identifiers mentioned, lines-of-code mentioned, security keywords mentioned) may require deeper contextual understanding than sentiment classification. [Objective:] Explore whether LLMs can act as automated annotators for technical security comments by human subjects. [Method:] We prompt four top-performing LLMs on LiveBench to detect nine security-relevant codes in free-text comments by human subjects analyzing vulnerable code snippets. Outputs are compared to human annotators using Cohen's Kappa (chance-corrected accuracy). We test different prompts mimicking annotation best practices, including emerging codes, detailed codebooks with examples, and conflicting examples. [Negative Results:] We observed marked improvements only when using detailed code descriptions; however, these improvements are not uniform across codes and are insufficient to reliably replace a human annotator. [Limitations:] Additional studies with more LLMs and annotation tasks are needed.","authors":["Maria Camporese","Fabio Massacci","Yuanjun Gong"],"categories":["cs.SE","cs.AI"],"primary_category":"cs.SE","announce_type":"new","date":"2026-04-12","first_seen":"2026-04-12","revised_at":null,"abs_url":"https://arxiv.org/abs/2604.10834","pdf_url":"https://arxiv.org/pdf/2604.10834","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM标注","定性数据分析","安全评论"],"reason":"用LLM替代人工标注定性数据，属于标注员替代而非仿真被试，但涉及人类实验对照，…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:24","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":145,"question":"LLM能否替代人类标注者，对安全技术评论进行主题编码？","design":"本研究并非用LLM仿真人类被试，而是用LLM替代人类标注者。将四款LLM作为自动标注器，输入人类被试分析漏洞代码片段后产生的自由文本评论，要求其检测九种安全相关编码，并与人类标注结果比较。","baseline":"对照的真实人类数据是四位人类标注者对同一批评论的编码结果，以及审核者校正后的编码作为最终基准。","findings":"仅在使用详细代码描述时，LLM性能有明显提升，但提升在不同编码上并不均匀，且不足以可靠替代人类标注者。","reliability":"论文承认需要更多LLM和标注任务的研究来验证结论，且当前结果仅基于特定安全编码任务，泛化性有限。","relevance":"本文属于LLM替代人类标注者的研究，而非用LLM仿真人类被试，但涉及真实人类实验对照和可靠性评估，对关注LLM在人类研究中替代角色的研究者有参考价值，尤其在批判性评估LLM失效条件方面。","inspiration":"本研究采用LLM替代人类标注者的对照设计，将LLM输出与多位人类标注者及审核校正后的基准进行系统比较，并考察了输入信息详细程度对性能的影响，这种多层级对照和条件敏感性分析值得借鉴。｜该方法可迁移到经济金融文本分析场景，例如利用LLM对央行政策声明或分析师报告进行主题编码或情感分类，以评估LLM能否替代人类研究助理。｜可设计实验：以LLM为被试，输入央行货币政策声明，要求其识别声明中的前瞻指引、风险评估等主题，以人类专家标注的编码作为基准，比较不同提示信息（如仅声明文本 vs. 附加经济背景）下LLM的编码准确率。"}},{"id":"2605.23916","version":1,"title":"Agent-Facing Information Design in LLM Tool Registries","zh_title":"面向智能体的LLM工具注册信息设计","abstract":"LLM tool registries function as unregulated advertising platforms: providers write free-text descriptions that agents use for selection, yet no measurement infrastructure -- no viewability standard, quality score, or outcome audit -- exists to make this market accountable. We provide the first systematic framework, combining 17,700+ trials across five LLMs and ten domains with a constructive registry design prescription. Legal puffery alone (subjective superlatives, benefit framing) captures 100% of the optimization effect; fabricated claims add zero incremental bias -- rendering FTC enforcement of deceptive advertising rules ineffective against the active mechanism. Disclosure fails structurally: system-prompt warnings produce zero measurable effect for four of five models, and behavioral ceilings leave no headroom for label-based correction. Superlatives are the dominant single feature (SBC = +0.35). Registry-layer description normalization achieves first-best welfare model-independently. We propose separating selection-facing descriptions (structured, registry-controlled) from marketing-facing descriptions (provider-authored, shown post-selection), and introduce the Agent Attention Quality Score to distinguish capability from copywriting.","authors":["Haochuan Kevin Wang"],"categories":["cs.IR","cs.AI","econ.GN"],"primary_category":"cs.IR","announce_type":"new","date":"2026-04-12","first_seen":"2026-04-12","revised_at":null,"abs_url":"https://arxiv.org/abs/2605.23916","pdf_url":"https://arxiv.org/pdf/2605.23916","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","信息设计","工具注册"],"reason":"研究LLM工具注册中的信息设计，属于多智能体系统优化，不涉及人类行为仿真或对照。","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:58:30","error":null,"has_summary":false,"summary":null},{"id":"2604.10029","version":2,"title":"Self-Distilled Reinforcement Learning for Co-Evolving Agentic Recommender Systems","zh_title":"用于协同进化智能体推荐系统的自蒸馏强化学习","abstract":"Large language model-empowered agentic recommender systems (ARS) reformulate recommendation as a multi-turn interaction between a recommender agent and a user agent, enabling iterative preference elicitation and refinement beyond conventional one-shot prediction. However, existing ARS are mainly optimized in a Reflexion-style paradigm, where past interaction trajectories are stored as textual memory and retrieved as prompt context for later reasoning. Although this design allows agents to recall prior feedback and observations, the accumulated experience remains external to model parameters, leaving agents reliant on generic reasoning rather than progressively acquiring recommendation-specific decision-making ability through learning. Reinforcement learning (RL) therefore provides a natural way to internalize such interaction experience into parameters. Yet existing RL methods for ARS still suffer from two key limitations. First, they fail to capture the interactive nature of ARS, in which the recommender agent and the user agent continuously influence each other and can naturally generate endogenous supervision through interaction feedback. Second, they reduce a rich multi-turn interaction process to final outcomes, overlooking the dense supervision embedded throughout the trajectory. To this end, we propose CoARS, a self-distilled reinforcement learning framework for co-evolving agentic recommender systems. CoARS introduces two complementary learning schemes: interaction reward, which derives coupled task-level supervision for the recommender agent and the user agent from the same interaction trajectory, and self-distilled credit assignment, which converts historical trajectories into token-level credit signals under teacher-student conditioning. Experiments on multiple datasets show that CoARS outperforms representative ARS baselines in recommendation performance and user alignment.","authors":["Zongwei Wang","Min Gao","Hongzhi Yin","Junliang Yu","Tong Chen","Quoc Viet Hung Nguyen","Shazia Sadiq","Tianrui Li"],"categories":["cs.IR"],"primary_category":"cs.IR","announce_type":"new","date":"2026-04-11","first_seen":"2026-04-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2604.10029","pdf_url":"https://arxiv.org/pdf/2604.10029","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["智能体推荐系统","强化学习","多智能体协作"],"reason":"纯多智能体协作优化推荐系统，无人类行为对照，不涉及人类仿真实验。","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:59:39","error":null,"has_summary":false,"summary":null},{"id":"2604.09502","version":2,"title":"Strategic Algorithmic Monoculture: Experimental Evidence from Coordination Games","zh_title":"策略性算法单一文化：来自协调博弈的实验证据","abstract":"AI agents increasingly operate in multi-agent environments where outcomes depend on coordination. We distinguish primary algorithmic monoculture -- baseline action similarity -- from strategic algorithmic monoculture, whereby agents adjust similarity in response to incentives. We implement a simple experimental design that cleanly separates these forces, and deploy it on human and large language model (LLM) subjects. LLMs exhibit high levels of baseline similarity (primary monoculture) and, like humans, they regulate it in response to coordination incentives (strategic monoculture). While LLMs coordinate extremely well on similar actions, they lag behind humans in sustaining heterogeneity when divergence is rewarded.","authors":["Gonzalo Ballestero","Hadi Hosseini","Samarth Khanna","Ran I. Shorrer"],"categories":["cs.AI","cs.GT","cs.MA","econ.TH"],"primary_category":"cs.AI","announce_type":"new","date":"2026-04-10","first_seen":"2026-04-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2604.09502","pdf_url":"https://arxiv.org/pdf/2604.09502","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM仿真","人类行为对照","协调博弈"],"reason":"用LLM和人类被试进行协调博弈实验，直接对比行为，属于经济学实验场景的人类仿真。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:22","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":17,"question":"LLM智能体在协调博弈中如何协调行动，与人类相比有何差异？","design":"使用多个LLM（16个模型）和人类被试，在开放式问题（如说出一个字母、一个城市）上设置三种处理：picking（仅要求有效答案）、coordination（激励与同类型另一智能体答案相同）、divergence（激励答案不同），测量独立智能体间答案一致率。","baseline":"人类被试在相同实验任务中的行为数据。","findings":"LLM在无激励时答案一致性远高于人类，表现出高基础单一性；在协调激励下LLM能像人类一样调节一致性，但在需要差异化的任务中维持异质性的能力弱于人类。","reliability":"论文未讨论","relevance":"该研究直接对比LLM与人类在协调博弈中的行为，属于经济学实验场景的人类仿真，且揭示了LLM在差异化任务中的失效，高度契合研究者对仿真可靠性与偏差的关注。","inspiration":"借鉴其通过picking、coordination、divergence三种处理分离基础偏好与策略调整的实验设计，可清晰测量LLM的固有行为倾向与激励响应。｜可迁移至资产定价实验，研究LLM交易员在信息协调与差异化策略下的市场表现。｜以LLM为被试，设置picking（自由选股）、coordination（激励与另一LLM选相同股票）、divergence（激励选不同股票）三种处理，结果变量为投资组合相似度与市场效率指标，对照真实人类交易员在相同实验中的行为数据。"}},{"id":"2604.16472","version":1,"title":"Training Language Models for Bilateral Trade with Private Information","zh_title":"训练语言模型进行具有私人信息的双边贸易","abstract":"Bilateral bargaining under incomplete information provides a controlled testbed for evaluating large language model (LLM) agent capabilities. Bilateral trade demands individual rationality, strategic surplus maximization, and cooperation to realize gains from trade. We develop a structured bargaining environment where LLMs negotiate via tool calls within an event-driven simulator, separating binding offers from natural-language messages to enable automated evaluation. The environment serves two purposes: as a benchmark for frontier models and as a training environment for open-weight models via reinforcement learning. In benchmark experiments, a round-robin tournament among five frontier models (15,000 negotiations) reveals that effective strategies implement price discrimination through sequential offers. Aggressive anchoring, calibrated concession, and temporal patience correlate with the highest surplus share and deal rate. Accommodating strategies that concede quickly disable price discrimination in the buyer role, yielding the lowest surplus capture and deal completion. Stronger models scale their behavior proportionally to item value, maintaining performance across price tiers; weaker models perform well only when wide zones of possible agreement offset suboptimal strategies. In training experiments, we fine-tune Qwen3 (8B, 14B) via supervised fine-tuning (SFT) followed by Group Relative Policy Optimization (GRPO) against a fixed frontier opponent. These stages optimize competing objectives: SFT approximately doubles surplus share but reduces deal rates, while RL recovers deal rates but erodes surplus gains, reflecting the reward structure. SFT also compresses surplus variation across price tiers, which generalizes to unseen opponents, suggesting that behavioral cloning instills proportional strategies rather than memorized price points.","authors":["Dirk Bergemann","Soheil Ghili","Xinyang Hu","Chuanhao Li","Zhuoran Yang"],"categories":["cs.GT","cs.AI","cs.MA","econ.GN","econ.TH"],"primary_category":"cs.GT","announce_type":"new","date":"2026-04-10","first_seen":"2026-04-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2604.16472","pdf_url":"https://arxiv.org/pdf/2604.16472","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["LLM Agent","双边贸易","社会模拟"],"reason":"用LLM agent模拟双边贸易谈判，但无真实人类数据对照，属于社会模拟的纯理…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:51","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":177,"question":"如何评估和训练大语言模型在私有信息双边贸易中的战略谈判能力？","design":"构建结构化讨价还价环境，让LLM通过工具调用进行谈判，分离约束性报价与自然语言消息；对五个前沿模型进行循环赛基准测试（15000次谈判），并对Qwen3（8B、14B）进行监督微调加GRPO强化学习训练，测量个体理性、剩余份额和成交率。","baseline":"无对照","findings":"前沿模型中，激进锚定加校准让步策略同时获得最高剩余份额和成交率；SFT提高剩余份额但降低成交率，RL恢复成交率却侵蚀剩余收益，两者目标冲突。","reliability":"论文未讨论","relevance":"该研究用LLM模拟双边谈判，但缺乏真实人类数据对照，属于纯仿真实验，未直接评估仿真可靠性或偏差，与研究者关注的人类基准对照和批判性评估方向部分相关，但非核心匹配。","inspiration":"该研究通过工具调用分离约束性报价与自然语言消息，并采用循环赛基准测试和强化学习训练，为多智能体策略评估提供了结构化方法｜可迁移至资产定价实验中的做市商谈判或信贷审批中的利率协商场景｜以LLM作为交易员被试，处理为不同训练方式（SFT vs RL），结果变量为成交价与成交率，对照真实市场微观结构数据中的买卖价差与成交概率"}},{"id":"2604.09855","version":1,"title":"Instructing LLMs to Negotiate using Reinforcement Learning with Verifiable Rewards","zh_title":"使用可验证奖励的强化学习指导大语言模型进行谈判","abstract":"The recent advancement of Large Language Models (LLMs) has established their potential as autonomous interactive agents. However, they often struggle in strategic games of incomplete information, such as bilateral price negotiation. In this paper, we investigate if Reinforcement Learning from Verifiable Rewards (RLVR) can effectively teach LLMs to negotiate. Specifically, we explore the strategic behaviors that emerge during the learning process. We introduce a framework that trains a mid-sized buyer agent against a regulated LLM seller across a wide distribution of real-world products. By grounding reward signals directly in the maximization of economic surplus and strict adherence to private budget constraints, we reveal a novel four-phase strategic evolution. The agent progresses from naive bargaining to using aggressive starting prices, moves through a phase of deadlock, and ultimately develops sophisticated persuasive skills. Our results demonstrate that this verifiable training allows a 30B agent to significantly outperform frontier models over ten times its size in extracting surplus. Furthermore, the trained agent generalizes robustly to stronger counterparties unseen during training and remains effective even when facing hostile, adversarial seller personas.","authors":["Shuze Daniel Liu","Claire Chen","Jiabao Sean Xiao","Lei Lei","Yuheng Zhang","Yisong Yue","David Simchi-Levi"],"categories":["cs.AI","cs.CL","cs.GT","econ.GN"],"primary_category":"cs.AI","announce_type":"new","date":"2026-04-10","first_seen":"2026-04-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2604.09855","pdf_url":"https://arxiv.org/pdf/2604.09855","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","强化学习","谈判策略"],"reason":"纯多智能体协作训练谈判策略，无人类行为对照，属C1排除项。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:51","error":null,"has_summary":false,"summary":null},{"id":"2604.09187","version":1,"title":"The Geoeconomics of Venture Capital An Economic Complexity Approach to Emerging Technological Sovereignty","zh_title":"风险投资的地缘经济学：新兴技术主权的经济复杂性方法","abstract":"We explore a quantitative approach to emerging technological sovereignty and geoeconomic power by assessing the relative positioning of countries with economic complexity methods applied to the structure of national venture-capital (VC) portfolios and their associated Revealed Venture Advantage (RVA) metrics. Using Crunchbase firm- and deal-level data, we map venture-backed startups to 18 emerging technology domains via a probabilistic multi-label large-language-model classifier, and construct an RVA-based country-technology specialization matrix for the 17 countries with the highest aggregate VC funding. From this matrix, we derive two eigenvector-based measures: a Geoeconomic Complexity Index (GCI) that ranks countries by the composition of their venture specializations, and an Emerging Technology Geoeconomic Complexity Index (ETGCI) that ranks domains by the extent to which specialization is concentrated among high-GCI countries. Empirically, Cloud Computing, Cybersecurity Tools, and Medtech exhibit the highest ETGCI values, reflecting concentration of specialization in a small set of leading countries. The United States and Israel consistently occupy a marked \"high-diversity/low-ubiquity\" position and lead the GCI ranking, followed by China, France, Japan, and Germany; both country and domain rankings are stable from 2021-2024. Finally, relatedness-based simulations identify, when it exists, for each country the Simplest Single Sovereignty Enhancing Technology (SSSET), i.e., the most feasible single new technological direction associated with the largest expected improvement in relative geoeconomic positioning.","authors":["Benjamin Leroy","Davi Marim","El Ghali Benjelloun","Arthur Rozan Debeaurain","Jean-Michel Dalle"],"categories":["econ.GN","physics.soc-ph","q-fin.GN"],"primary_category":"econ.GN","announce_type":"new","date":"2026-04-10","first_seen":"2026-04-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2604.09187","pdf_url":"https://arxiv.org/pdf/2604.09187","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["风险投资","经济复杂性","技术分类"],"reason":"论文用LLM做技术分类，非仿真人类被试，无行为对照。","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:58:18","error":null,"has_summary":false,"summary":null},{"id":"2604.08678","version":2,"title":"Scaffolding Human-AI Collaboration: A Field Experiment on Behavioral Protocols and Cognitive Reframing","zh_title":"搭建人机协作框架：关于行为协议与认知重构的现场实验","abstract":"Organizations have widely deployed generative AI tools, yet productivity gains remain uneven, suggesting that how people use AI matters as much as whether they have access. We conducted a field experiment with 388 employees at a Fortune 500 retailer to test two scaffolding interventions for human-AI collaboration. All participants had access to the same AI tool; we varied only the structure surrounding its use. A behavioral scaffolding intervention (a structured protocol requiring joint AI use within pairs) was associated with lower document quality relative to unstructured use and substantially lower document production. A cognitive scaffolding intervention (partnership training that reframed AI as a thought partner) was associated with higher individual document quality at the top of the distribution. Treatment participants also showed greater positive belief change across the session, though sensitivity analyses suggest this likely reflects recovery from carry-over effects rather than genuine training-induced shifts. Both findings are subject to design limitations including an AM/PM session confound, differential attrition, and LLM grading sensitivity to document length.","authors":["Alex Farach","Alexia Cambon","Lev Tankelevitch","Connie Hsueh","Rebecca Janssen"],"categories":["econ.GN","cs.HC"],"primary_category":"econ.GN","announce_type":"new","date":"2026-04-09","first_seen":"2026-04-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2604.08678","pdf_url":"https://arxiv.org/pdf/2604.08678","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["人机协作","现场实验","生成式AI"],"reason":"研究人类如何与AI协作，非用LLM仿真人类被试，无人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:58:31","error":null,"has_summary":false,"summary":null},{"id":"2604.06663","version":1,"title":"Restoring Heterogeneity in LLM-based Social Simulation: An Audience Segmentation Approach","zh_title":"在基于大语言模型的社会模拟中恢复异质性：一种受众细分方法","abstract":"Large Language Models (LLMs) are increasingly used to simulate social attitudes and behaviors, offering scalable \"silicon samples\" that can approximate human data. However, current simulation practice often collapses diversity into an \"average persona,\" masking subgroup variation that is central to social reality. This study introduces audience segmentation as a systematic approach for restoring heterogeneity in LLM-based social simulation. Using U.S. climate-opinion survey data, we compare six segmentation configurations across two open-weight LLMs (Llama 3.1-70B and Mixtral 8x22B), varying segmentation identifier granularity, parsimony, and selection logic (theory-driven, data-driven, and instrument-based). We evaluate simulation performance with a three-dimensional evaluation framework covering distributional, structural, and predictive fidelity. Results show that increasing identifier granularity does not produce consistent improvement: moderate enrichment can improve performance, but further expansion does not reliably help and can worsen structural and predictive fidelity. Across parsimony comparisons, compact configurations often match or outperform more comprehensive alternatives, especially in structural and predictive fidelity, while distributional fidelity remains metric dependent. Identifier selection logic determines which fidelity dimension benefits most: instrument-based selection best preserves distributional shape, whereas data-driven selection best recovers between-group structure and identifier-outcome associations. Overall, no single configuration dominates all dimensions, and performance gains in one dimension can coincide with losses in another. These findings position audience segmentation as a core methodological approach for valid LLM-based social simulation and highlight the need for heterogeneity-aware evaluation and variance-preserving modeling strategies.","authors":["Xiaoyou Qin","Zhihong Li","Xiaoxiao Cheng"],"categories":["cs.CY","cs.AI"],"primary_category":"cs.CY","announce_type":"new","date":"2026-04-08","first_seen":"2026-04-08","revised_at":null,"abs_url":"https://arxiv.org/abs/2604.06663","pdf_url":"https://arxiv.org/pdf/2604.06663","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","A3","B1","B2","B4"],"tags":["LLM人类仿真","受众细分","社会模拟保真度"],"reason":"用LLM仿真人类气候态度，有真实调查数据对照，评估异质性恢复与保真度，并指出失…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:22","error":null,"has_summary":true,"summary":{"generated_at":"2026-04-08","rank":5,"question":"如何通过受众分割策略恢复大语言模型社会仿真中的异质性？","design":"使用Llama 3.1-70B和Mixtral 8x22B模型，基于美国气候态度调查数据，通过六种分割配置（变化标识符粒度、简约性和选择逻辑）生成合成样本，评估分布保真度、结构保真度和预测保真度。","baseline":"2025年10月通过Prolific收集的594份人类样本，按人口普查配额分层抽样，并使用Six Americas超短问卷分配受众细分标签。","findings":"增加标识符粒度并不一致提升性能；中等丰富度可改善，但过度扩展会损害结构和预测保真度。紧凑配置在结构和预测保真度上常优于或等同更全面的配置，而标识符选择逻辑决定哪个保真度维度受益最大。","reliability":"论文指出无单一配置在所有维度占优，一维度的提升可能伴随另一维度的损失；未明确讨论其他失效条件。","relevance":"直接命中研究者关注的LLM仿真人类实验、有真实人类对照、评估保真度并指出失效条件，强烈建议阅读原文。","inspiration":"该研究通过系统变换受众分割的标识符粒度、简约性和选择逻辑来评估仿真保真度，这种多维度配置对比的设计值得借鉴｜可迁移到消费者金融决策仿真，如退休储蓄选择或保险购买行为中的异质性偏好研究｜以LLM作为被试，施加不同信息框架（如损失vs收益表述）作为处理，测量储蓄率或保险购买意愿，并以美国消费者金融调查（SCF）或健康与退休研究（HRS）的真实个体数据作为对照基准"}},{"id":"2604.16465","version":1,"title":"Healthcare AI for Automation or Allocation? A Transaction Cost Economics Framework","zh_title":"医疗AI用于自动化还是分配？一个交易成本经济学框架","abstract":"Healthcare productivity is shaped not only by clinical complexity but by the costs of coordinating work under uncertainty. Transaction-cost economics offers a theory of these coordination frictions, yet has rarely been operationalised at task level across health occupations. Using task statements and frequency weights from the O*NET occupational database, we characterised healthcare work at task granularity and coded each unique task using a constrained large language model into one dominant transaction-cost category (information search, decision and bargaining, monitoring and enforcement, or adaptation and coordination) together with an overall transaction-cost intensity score. Aggregating to the occupation level, clinician roles exhibited substantially higher transaction-cost intensity than non-clinician roles, driven primarily by greater burdens of information search and decision-related coordination, while dispersion of transaction costs within occupations did not differ. These findings demonstrate systematic heterogeneity in the nature of coordination work across healthcare roles and suggest that the opportunities for digital and AI interventions are unevenly distributed, shaped less by technical task complexity than by underlying coordination structure.","authors":["Ari Ercole"],"categories":["cs.AI","econ.GN"],"primary_category":"cs.AI","announce_type":"new","date":"2026-04-08","first_seen":"2026-04-08","revised_at":null,"abs_url":"https://arxiv.org/abs/2604.16465","pdf_url":"https://arxiv.org/pdf/2604.16465","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["交易成本经济学","职业分析","LLM标注"],"reason":"用LLM对职业任务进行编码分类，属于自动化标注，不涉及人类行为仿真或对照。","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:58:19","error":null,"has_summary":false,"summary":null},{"id":"2604.08606","version":1,"title":"Extrapolating Volition with Recursive Information Markets","zh_title":"用递归信息市场外推意志","abstract":"One of the impediments to the efficiency of information markets is the inherent information asymmetry present in them, exacerbated by the \"buyer's inspection paradox\" (the buyer cannot mitigate the asymmetry by \"inspecting\" the information, because in doing so the buyer obtains the information without paying for it). Previous work has suggested that using Large Language Model (LLM) buyers to inspect and purchase information could overcome this information asymmetry, as an LLM buyer can simply \"forget\" the information it inspects. In this work, we analyze this mechanism formally through a \"value-of-information\" paradigm, i.e. whether it incentivizes information to be priced and provided in accordance with its \"true value\". We focus in particular on our new recursive version of the mechanism, which we believe has a range of applications including in AI alignment research, where it is related to Extrapolated Volition and Scalable Oversight.","authors":["Abhimanyu Pallavi Sudhir","Long Tran-Thanh"],"categories":["cs.GT","cs.AI","econ.TH"],"primary_category":"cs.GT","announce_type":"new","date":"2026-04-08","first_seen":"2026-04-08","revised_at":null,"abs_url":"https://arxiv.org/abs/2604.08606","pdf_url":"https://arxiv.org/pdf/2604.08606","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["信息市场","多智能体","AI对齐"],"reason":"纯多智能体信息市场机制设计，LLM作为买家解决信息不对称，不涉及人类行为仿真或…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:51","error":null,"has_summary":false,"summary":null},{"id":"2604.06860","version":1,"title":"Personalization as a Game: Equilibrium-Guided Generative Modeling for Physician Behavior in Pharmaceutical Engagement","zh_title":"作为博弈的个性化：制药参与中医生行为的均衡引导生成建模","abstract":"We present \\textbf{EGPF} (Equilibrium-Guided Personalization Framework), a mathematically rigorous architecture unifying Bayesian game theory, category theory, information theory, and generative AI for hyper-personalized physician engagement in the pharmaceutical domain. Our framework models the pharma--physician interaction as an incomplete-information Bayesian game where physician behavioral types are inferred via functorial mappings from observational categories, equilibrium strategies guide content generation through large language models (LLMs), and information-theoretic feedback loops ensure adaptive recalibration. We formalize behavior composition through category-theoretic functors, natural transformations, and monoidal structures, enabling modular, composable physician archetypes that respect structural invariants under domain shift. We introduce a novel \\textit{Rate-Distortion Equilibrium} (RDE) criterion that bounds the personalization--privacy tradeoff, an \\textit{Evolutionary Game Dynamics} layer for population-level behavior modeling, a \\textit{Mechanism Design} module for incentive-compatible engagement, and a \\textit{Sheaf-Theoretic} extension for multi-scale behavioral consistency. We prove convergence of our iterative belief-update mechanism at rate $O(\\frac{K\\log K}{t \\cdot C_{\\min}})$ and establish finite-sample regret bounds. Extensive experiments on synthetic pharma datasets and a real-world HCP engagement pilot demonstrate a 34\\% improvement in engagement prediction (AUC) and 28\\% lift in content relevance scores compared to state-of-the-art methods.","authors":["Suyash Mishra"],"categories":["cs.GT"],"primary_category":"cs.GT","announce_type":"new","date":"2026-04-08","first_seen":"2026-04-08","revised_at":null,"abs_url":"https://arxiv.org/abs/2604.06860","pdf_url":"https://arxiv.org/pdf/2604.06860","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["博弈论","生成AI","制药营销"],"reason":"纯多智能体博弈建模，无人类行为对照，不涉及LLM仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:49","error":null,"has_summary":false,"summary":null},{"id":"2604.05939","version":1,"title":"Context-Value-Action Architecture for Value-Driven Large Language Model Agents","zh_title":"面向价值驱动大语言模型智能体的情境-价值-行动架构","abstract":"Large Language Models (LLMs) have shown promise in simulating human behavior, yet existing agents often exhibit behavioral rigidity, a flaw frequently masked by the self-referential bias of current \"LLM-as-a-judge\" evaluations. By evaluating against empirical ground truth, we reveal a counter-intuitive phenomenon: increasing the intensity of prompt-driven reasoning does not enhance fidelity but rather exacerbates value polarization, collapsing population diversity. To address this, we propose the Context-Value-Action (CVA) architecture, grounded in the Stimulus-Organism-Response (S-O-R) model and Schwartz's Theory of Basic Human Values. Unlike methods relying on self-verification, CVA decouples action generation from cognitive reasoning via a novel Value Verifier trained on authentic human data to explicitly model dynamic value activation. Experiments on CVABench, which comprises over 1.1 million real-world interaction traces, demonstrate that CVA significantly outperforms baselines. Our approach effectively mitigates polarization while offering superior behavioral fidelity and interpretability.","authors":["TianZe Zhang","Sirui Sun","Yuhang Xie","Xin Zhang","Zhiqiang Wu","Guojie Song"],"categories":["cs.AI","cs.HC"],"primary_category":"cs.AI","announce_type":"new","date":"2026-04-07","first_seen":"2026-04-07","revised_at":null,"abs_url":"https://arxiv.org/abs/2604.05939","pdf_url":"https://arxiv.org/pdf/2604.05939","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","A3","B1","B4"],"tags":["人类行为仿真","价值驱动智能体","算法保真度"],"reason":"用LLM仿真人类行为，有真实人类数据对照，解决行为僵化和价值极化问题，直接相关。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:20","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":18,"question":"如何设计LLM智能体架构，以缓解提示驱动推理导致的行为僵化和价值极化，从而更真实地模拟人类行为？","design":"提出Context-Value-Action（CVA）架构，基于S-O-R模型和Schwartz基本人类价值观理论，将动作生成与认知推理解耦，引入在真实人类数据上训练的Value Verifier显式建模动态价值激活，并通过SFT和DPO对齐生成过程。","baseline":"CVABench包含超过110万条真实世界交互轨迹，来自15000多名人类参与者，用于评估行为逼真度和价值极化。","findings":"增加提示驱动的推理强度不仅未提高仿真逼真度，反而加剧价值极化并降低群体多样性；CVA架构有效缓解极化，在行为逼真度和可解释性上显著优于基线方法。","reliability":"论文未讨论","relevance":"高度相关：研究用LLM仿真人类行为，有大规模真实人类数据作为基准，揭示提示驱动方法的失效模式并提出改进架构，直接回应研究者对仿真可靠性、偏差和经济学/政策评估场景的关切。","inspiration":"借鉴CVA架构将行为生成与价值推理解耦，并引入在真实人类数据上训练的Value Verifier来动态建模价值激活，以此作为缓解行为僵化和价值极化的处理机制。｜可迁移到政策公告的预期形成实验，研究不同信息框架如何通过个体价值观影响通胀预期或就业预期。｜以LLM智能体为被试，处理组采用CVA架构注入特定价值观（如安全或自主），对照组使用标准提示驱动推理，结果变量为预期偏差和群体多样性，用真实调查数据（如密歇根大学消费者调查）作为人类基准对照。"}},{"id":"2604.05516","version":2,"title":"Coupling Macro Dynamics and Micro States for Long-Horizon Social Simulation","zh_title":"耦合宏观动态与微观状态的长周期社会模拟","abstract":"Social network simulation aims to model collective opinion dynamics in large populations, but existing LLM-based simulators mainly focus on aggregate dynamics while largely ignoring individual internal states. This limits their ability to capture opinion reversals driven by gradual individual shifts and makes them unreliable in long-horizon simulations. We propose MF-MDP, a social simulation framework that tightly couples macro-level collective dynamics with micro-level individual states. MF-MDP explicitly models per-agent latent opinion states with a state transition mechanism, combining individual Markov Decision Processes at the micro level with a mean-field collective framework at the macro level. This allows individual behaviors to change internal states gradually rather than trigger instant reactions, enabling the simulator to distinguish agents that are close to switching from those that are far from switching, capture opinion reversals, and maintain accuracy over long horizons. Across real-world events, MF-MDP supports stable simulation of long-horizon social processes with up to 40,000 interactions, compared with about 300 in the baseline MF-LLM, while reducing long-horizon KL divergence by 75.3% (1.2490 to 0.3089) and reversal KL by 66.9% (1.6425 to 0.5434), significantly mitigating the drift observed in MF-LLM. Code is available at github.com/AI4SS/MF-MDP.","authors":["Yunyao Zhang","Yihao Ai","Zuocheng Ying","Qirui Mi","Junqing Yu","Wei Yang","Zikai Song"],"categories":["cs.SI"],"primary_category":"cs.SI","announce_type":"new","date":"2026-04-07","first_seen":"2026-04-07","revised_at":null,"abs_url":"https://arxiv.org/abs/2604.05516","pdf_url":"https://arxiv.org/pdf/2604.05516","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["社会模拟","舆论动态","LLM agent"],"reason":"用LLM agent模拟社会舆论动态，但摘要未明确提及真实人类数据对照，属于社…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:20","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":158,"question":"如何在长时间跨度社会仿真中，通过耦合宏观集体动态与微观个体状态，捕捉观点反转并保持长期准确性？","design":"提出MF-MDP框架，将社会仿真建模为平均场马尔可夫决策过程：宏观层用时序Transformer学习宏观状态分布演化，微观层每个智能体基于自身潜在状态和宏观信号进行多步前瞻动作重选，模拟个体观点的渐进转变。","baseline":"无对照","findings":"MF-MDP在真实社会事件上支持长达40,000次交互的稳定仿真，而基线MF-LLM仅约300次；长期KL散度降低75.3%，反转KL散度降低66.9%，显著缓解了漂移问题。","reliability":"论文未讨论","relevance":"该研究利用LLM智能体进行社会舆论仿真，但未提供真实人类数据对照，不属于严格的人类仿真验证研究；若关注仿真方法改进与长期动态稳定性，可读原文了解其微观-宏观耦合机制。","inspiration":"MF-MDP框架通过宏观时序建模与微观多步前瞻动作重选耦合，可借鉴其分层仿真思路来设计经济实验中的动态处理与长期效应测量｜该框架可迁移至政策公告的预期形成与市场反应研究，例如模拟央行沟通对通胀预期和资产价格的长期影响｜以LLM智能体为被试，处理为不同频率/透明度的政策信号，结果变量为智能体通胀预期与模拟资产价格，对照真实央行公告前后的调查预期与市场数据"}},{"id":"2604.04464","version":2,"title":"Bounded by Risk, Not Capability: Quantifying AI Occupational Substitution Rates via a Tech-Risk Dual-Factor Model","zh_title":"受风险而非能力约束：通过技术-风险双因素模型量化AI职业替代率","abstract":"The deployment of Large Language Models (LLMs) has ignited concerns about technological unemployment. Existing task-based evaluations predominantly measure theoretical \"exposure\" to AI capabilities, ignoring critical frictions of real-world commercial adoption: liability, compliance, and physical safety. We argue occupations are not eradicated instantaneously, but gradually encroached upon via atomic actions. We introduce a Tech-Risk Dual-Factor Model to re-evaluate this. By deconstructing 923 occupations into 2,087 Detailed Work Activities (DWAs), we utilize a multi-agent LLM ensemble to score both technical feasibility and business risk. Through variance-based Human-in-the-Loop (HITL) validation with an expert panel, we demonstrate a profound cognitive gap: isolated algorithmic probabilities fail to encapsulate the \"institutional premium\" imposed by experts bounded by professional liability. Applying a strictly algorithmic baseline via mathematical bottleneck aggregation, we calculate Relative Occupational Automation Indices ($OAI$) for the U.S. labor market. Our findings challenge the traditional Routine-Biased Technological Change (RBTC) hypothesis. Non-routine cognitive roles highly dependent on symbolic manipulation (e.g., Data Scientists) face unprecedented exposure ($OAI \\approx 0.70$). Conversely, unstructured physical trades and high-stakes caretaking roles exhibit absolute resilience, quantifying a profound \"Cognitive Risk Asymmetry.\" We hypothesize the emergent necessity of a \"Compliance Premium,\" indicating wage resilience increasingly tied to risk-absorption capacity. We frame these findings as a cross-sectional diagnostic of systemic vulnerability, establishing a foundation for subsequent Computable General Equilibrium (CGE) econometric modeling involving dynamic wage elasticity and structural labor reallocation.","authors":["Shuyao Gao","Minghao Huang"],"categories":["cs.CY","econ.GN"],"primary_category":"cs.CY","announce_type":"new","date":"2026-04-06","first_seen":"2026-04-06","revised_at":null,"abs_url":"https://arxiv.org/abs/2604.04464","pdf_url":"https://arxiv.org/pdf/2604.04464","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["AI职业替代","多智能体仿真","劳动力市场"],"reason":"用LLM多智能体模拟职业替代，但无真实人类行为对照，属社会模拟边界情形。","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:58:21","error":null,"has_summary":false,"summary":null},{"id":"2604.03920","version":2,"title":"From Plausible to Causal: Counterfactual Semantics for Policy Evaluation in Simulated Online Communities","zh_title":"从合理到因果：模拟在线社区中政策评估的反事实语义","abstract":"LLM-based social simulations can generate believable community interactions, enabling ``policy wind tunnels'' where governance interventions are tested before deployment. But believability is not causality. Claims like ``intervention $A$ reduces escalation'' require causal semantics that current simulation work typically does not specify. We propose adopting the causal counterfactual framework, distinguishing \\textit{necessary causation} (would the outcome have occurred without the intervention?) from \\textit{sufficient causation} (does the intervention reliably produce the outcome?). This distinction maps onto different stakeholder needs: moderators diagnosing incidents require evidence about necessity, while platform designers choosing policies require evidence about sufficiency. We formalize this mapping, show how simulation design can support estimation under explicit assumptions, and argue that the resulting quantities should be interpreted as simulator-conditional causal estimates whose policy relevance depends on simulator fidelity. Establishing this framework now is essential: it helps define what adequate fidelity means and moves the field from simulations that look realistic toward simulations that can support policy changes.","authors":["Agam Goyal","Yian Wang","Eshwar Chandrasekharan","Hari Sundaram"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-04-05","first_seen":"2026-04-05","revised_at":null,"abs_url":"https://arxiv.org/abs/2604.03920","pdf_url":"https://arxiv.org/pdf/2604.03920","source_feed":"backfill","score":7,"bucket":"pending","rubric_hits":["A3","B4"],"tags":["LLM社会模拟","政策评估","因果推断"],"reason":"用LLM模拟在线社区进行政策评估，提出因果反事实框架，批判性指出仿真需因果保真…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:19","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":80,"question":"如何为基于LLM的在线社区仿真中的政策评估引入因果反事实语义，以区分必要因果与充分因果？","design":"本文为理论框架论文，未进行具体仿真实验；它提出在LLM驱动的社区仿真中，通过配对运行（有干预与无干预）来估计必要因果概率（PN）和充分因果概率（PS），并讨论了外生性和单调性假设下的简化估计方法。","baseline":"无对照","findings":"提出将因果反事实框架应用于仿真政策评估，区分必要因果（诊断特定事件）和充分因果（评估干预效果），并指出仿真可借助随机化满足外生性，但单调性可能因干预异质性而违反，此时PN/PS仅为下界。","reliability":"论文指出仿真因果估计是“仿真器条件因果估计”，其政策相关性取决于仿真保真度；单调性假设在社交系统中可能不成立，此时估计仅为下界，且违反单调性本身可能揭示干预异质性。","relevance":"该文直接回应了LLM仿真中“可信不等于因果”的批判，为政策评估提供了因果推断框架，并讨论了仿真保真度与因果估计可靠性的关系，与研究者关注的仿真可靠性及批判性研究高度相关，值得精读原文。","inspiration":"借鉴其配对运行（有干预与无干预）估计必要因果和充分因果概率的框架，可对LLM仿真中的政策干预进行因果分解与稳健性检验｜可迁移到政策公告的预期形成研究，例如评估央行沟通对市场预期的因果效应｜以LLM模拟投资者群体，处理为是否发布前瞻指引，结果变量为预期通胀率，对照真实调查数据（如密歇根消费者预期调查）"}},{"id":"2605.00841","version":1,"title":"AI Agents for Sustainable SMEs: A Green ESG Assessment Framework","zh_title":"面向可持续中小企业的AI代理：绿色ESG评估框架","abstract":"This study presents a novel, AI-driven framework for assessing Environmental, Social, and Governance (ESG) performance in European small and medium-sized enterprises (SMEs). An initial phase established expert-validated ESG baseline scores from a subset of the Flash Eurobarometer FL549 survey data. In the second phase, a scalable AI agent system, built on the n8n automation platform, applied these baselines to perform automated ESG classification and generate contextual recommendations using large language models (LLMs). The results demonstrate the AI system's high consistency with human-derived outputs, thereby supporting more effective monitoring and intervention strategies aligned with the European Green Deal.","authors":["Viet Trinh","Tan Nguyen","Minh-Huyen Phan","Quan Luu"],"categories":["cs.AI","econ.GN"],"primary_category":"cs.AI","announce_type":"new","date":"2026-04-05","first_seen":"2026-04-05","revised_at":null,"abs_url":"https://arxiv.org/abs/2605.00841","pdf_url":"https://arxiv.org/pdf/2605.00841","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["AI代理","ESG评估","标注替代"],"reason":"用LLM替代人工进行ESG评估，属于标注员替代而非仿真人类被试，但有人类数据对…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:55","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":178,"question":"如何利用AI代理和LLM对欧洲中小企业的ESG表现进行自动化评估与分类？","design":"本研究并非人类仿真实验，而是提出一个混合人机编排框架：先由人工处理约40%的Flash Eurobarometer FL549调查数据建立专家验证的ESG基线分数，再通过n8n自动化平台构建AI代理系统，利用领域启发式规则和LLM对剩余国家数据进行ESG分类并生成情境化建议。","baseline":"人类基准为专家验证的ESG基线分数，源自Flash Eurobarometer FL549调查数据的人工处理部分。","findings":"AI系统输出的ESG分类与人类推导结果高度一致，表明该框架可有效支持与欧洲绿色协议一致的监测和干预策略。","reliability":"论文未讨论","relevance":"本研究用LLM替代人工进行ESG评估，属于标注员替代而非仿真人类被试，但提供了真实人类数据作为对照基准，与研究者关注的LLM替代人类判断的可靠性问题部分相关，值得快速浏览以了解其一致性验证方法。","inspiration":"该方法通过人工处理部分数据建立专家验证基线，再用LLM自动化处理剩余数据并对比一致性，提供了人机混合验证的思路｜可迁移到信贷审批歧视研究，用LLM替代人工审核员评估贷款申请的公平性｜招募银行信贷员人工审核40%的贷款申请作为基线，用LLM处理剩余申请并输出批准/拒绝决策，以人工审核结果作为真实对照，比较两组在性别、种族等维度上的歧视差异"}},{"id":"2605.20191","version":1,"title":"Shiny Stories, Hidden Struggles: Investigating the Representation of Disability Through the Lens of LLMs","zh_title":"光鲜故事，隐藏挣扎：通过LLM视角考察残障表征","abstract":"Modern Large Language Models (LLMs) have recently attracted much attention for their ability to simulate human behavior and generate text that reflects personas and demographic groups. While these capabilities can open up a multitude of diverse applications across fields, it is crucial to examine how such models represent various target groups since LLMs can perpetuate and amplify biases or discrimination against historically marginalized communities or, alternatively, as a result of debiasing efforts, overcorrect by portraying overly positive stereotypes. This overcompensation can idealize these groups, erasing the complexities and challenges they face in favor of unrealistic depictions. In this paper, we investigate how LLMs represent disability by simulating the perspectives of individuals with disabilities in generating social media posts. These posts are then compared with those written by real people with disabilities, focusing on emotional tone, sentiment, and representative words and themes. Our analysis reveals two key findings: (1) LLMs often idealize the experiences of people with disabilities, producing overly positive stereotypes that, despite appearing uplifting, fail to authentically capture their lived realities; and (2) a comparative analysis of posts simulating individuals with and without disabilities highlights a negative bias, where certain topics, such as career and entertainment, are disproportionately associated with nondisabled individuals. This reinforces exclusionary narratives and over-idealized portrayals of disability, misrepresenting the actual challenges faced by this community. These findings align with broader concerns and ongoing research showing that LLMs struggle to reflect the diverse realities of society, particularly the nuanced experiences of marginalized groups, and underscore the need for critical scrutiny of their representations.","authors":["Marco Bombieri","Simone Paolo Ponzetto","Marco Rospocher"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-04-02","first_seen":"2026-04-02","revised_at":null,"abs_url":"https://arxiv.org/abs/2605.20191","pdf_url":"https://arxiv.org/pdf/2605.20191","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","残障表征","偏差评估"],"reason":"用LLM模拟残障人士发帖，并与真实人类数据对照，评估仿真偏差与失效条件，直接命…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:42","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":19,"question":"LLM生成的残障人士社交媒体帖子与真实残障人士的自我描述在情感、主题和语言上有何差异？","design":"使用多种LLM，通过提示词模拟残障人士和普通人在社交媒体上发帖，生成文本后自动标注情感、情绪和抑郁迹象，并与真实Reddit帖子进行对比分析。","baseline":"来自Reddit的真实残障人士自我介绍的帖子数据集。","findings":"LLM倾向于理想化残障人士的经历，产生过度积极的刻板印象，未能真实反映其生活挑战；同时，LLM在模拟残障与非残障个体时存在负面偏见，将职业、娱乐等主题不成比例地与非残障人士关联。","reliability":"论文指出LLM的正面理想化同样会造成伤害，且模型可能因去偏努力而过度补偿，但未详细讨论仿真失效的具体条件。","relevance":"该研究直接以真实人类数据为基准，评估LLM模拟残障人群的偏差与失效模式，属于批判性仿真研究，值得精读以了解LLM在边缘群体仿真中的局限。","inspiration":"该方法通过提示词让LLM模拟特定社会群体生成文本，并与真实社交媒体数据对比，可借鉴用于构建经济实验中的处理组与对照组仿真。｜它能迁移到信贷审批中的群体歧视研究，例如评估LLM在模拟不同种族或性别申请人时的语言偏差。｜可设计让LLM扮演贷款申请人撰写申请陈述，以真实银行信贷文本为基准，比较不同群体提示下的情感、主题和职业关联差异，检验仿真是否复现人类数据中的歧视模式。"}},{"id":"2604.01520","version":1,"title":"LLM Agents as Social Scientists: A Human-AI Collaborative Platform for Social Science Automation","zh_title":"作为社会科学家的LLM代理：一个面向社会科学自动化的人机协作平台","abstract":"Traditional social science research often requires designing complex experiments across vast methodological spaces and depends on real human participants, making it labor-intensive, costly, and difficult to scale. Here we present S-Researcher, an LLM-agent-based platform that assists researchers in conducting social science research more efficiently and at greater scale by \"siliconizing\" both the research process and the participant pool. To build S-Researcher, we first develop YuLan-OneSim, a large-scale social simulation system designed around three core requirements: generality via auto-programming from natural language to executable scenarios, scalability via a distributed architecture supporting up to 100,000 concurrent agents, and reliability via feedback-driven LLM fine-tuning. Leveraging this system, S-Researcher supports researchers in designing social experiments, simulating human behavior with LLM agents, analyzing results, and generating reports, forming a complete human-AI collaborative research loop in which researchers retain oversight and intervention at every stage. We operationalize LLM simulation research paradigms into three canonical reasoning modes (induction, deduction, and abduction) and validate S-Researcher through systematic case studies: inductive reproduction of cultural dynamics consistent with Axelrod's theory, deductive testing of competing hypotheses on teacher attention validated against survey data, and abductive identification of a cooperation mechanism in public goods games confirmed by human experiments. S-Researcher establishes a new human--AI collaborative paradigm for social science, in which computational simulation augments human researchers to accelerate discovery across the full spectrum of social inquiry.","authors":["Lei Wang","Yuanzi Li","Jinchao Wu","Heyang Gao","Xiaohe Bo","Xu Chen","Ji-Rong Wen"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-04-02","first_seen":"2026-04-02","revised_at":null,"abs_url":"https://arxiv.org/abs/2604.01520","pdf_url":"https://arxiv.org/pdf/2604.01520","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A3","A5","B1","B2","B3"],"tags":["LLM人类仿真","社会科学自动化","人机协作"],"reason":"用LLM代理模拟人类行为，复现文化动态、验证教师关注假设、识别合作机制，均有真…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:17","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":53,"question":"如何构建一个以LLM代理为核心的人机协作平台，实现社会科学研究的全流程自动化，并验证其在归纳、演绎、溯因三种推理模式下的有效性？","design":"使用基于LLM的代理系统YuLan-OneSim模拟人类参与者，通过自然语言自动编程生成可执行场景，支持高达10万并发代理的分布式架构，并利用反馈驱动微调提升可靠性。在归纳模式下，模拟文化传播动态；在演绎模式下，模拟课堂师生互动并测试竞争假设；在溯因模式下，模拟公共物品博弈以识别合作机制。","baseline":"演绎案例验证教师关注假设时，对照了真实调查数据；溯因案例识别合作机制时，对照了真实人类实验。归纳案例无明确真实人类数据对照，仅与Axelrod理论预测一致。","findings":"S-Researcher平台成功复现了Axelrod文化传播理论中的收敛与极化动态，并在课堂模拟中通过对比调查数据验证了教师关注假设，还在公共物品博弈中识别出合作机制并经人类实验确认。该平台实现了从实验设计、模拟、分析到报告生成的全流程人机协作，支持研究者全程干预。","reliability":"论文未明确讨论仿真失效的具体条件或局限，仅通过反馈微调机制和案例验证来确保可靠性，但未系统分析代理行为与真实人类偏差的来源或边界。","relevance":"高度相关：该研究直接用LLM代理替代人类被试，复现社会动态、验证假设并识别机制，且部分案例有真实人类数据对照，契合研究者对仿真可靠性、经济学实验和政策评估场景的关注，值得精读原文以评估其方法细节与批判性局限。","inspiration":"借鉴其利用LLM代理进行大规模并发模拟和反馈驱动微调以提升行为真实性的方法，可构建经济实验的虚拟被试池｜可迁移至公共物品博弈、税收遵从或劳动供给决策等行为经济学与政策评估场景｜以LLM代理为被试，施加不同税收政策处理，测量其劳动供给或逃税行为，并与真实实验室实验或行政数据对照"}},{"id":"2604.01896","version":1,"title":"Bayesian Elicitation with LLMs: Model Size Helps, Extra \"Reasoning\" Doesn't Always","zh_title":"基于大语言模型的贝叶斯启发：模型规模有益，额外“推理”未必有效","abstract":"Large language models (LLMs) have been proposed as alternatives to human experts for estimating unknown quantities with associated uncertainty, a process known as Bayesian elicitation. We test this by asking eleven LLMs to estimate population statistics, such as health prevalence rates, personality trait distributions, and labor market figures, and to express their uncertainty as 95\\% credible intervals. We vary each model's reasoning effort (low, medium, high) to test whether more \"thinking\" improves results. Our findings reveal three key results. First, larger, more capable models produce more accurate estimates, but increasing reasoning effort provides no consistent benefit. Second, all models are severely overconfident: their 95\\% intervals contain the true value only 9--44\\% of the time, far below the expected 95\\%. Third, a statistical recalibration technique called conformal prediction can correct this overconfidence, expanding the intervals to achieve the intended coverage. In a preliminary experiment, giving models web search access degraded predictions for already-accurate models, while modestly improving predictions for weaker ones. Models performed well on commonly discussed topics but struggled with specialized health data. These results indicate that LLM uncertainty estimates require statistical correction before they can be used in decision-making.","authors":["Luka Hobor","Mario Brcic","Mihael Kovac","Kristijan Poje"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-04-02","first_seen":"2026-04-02","revised_at":null,"abs_url":"https://arxiv.org/abs/2604.01896","pdf_url":"https://arxiv.org/pdf/2604.01896","source_feed":"backfill","score":7,"bucket":"pending","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","贝叶斯启发","不确定性校准"],"reason":"用LLM替代人类专家进行贝叶斯估计，并与真实人口统计数据对照，评估其校准与偏差…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:17","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":98,"question":"大语言模型能否替代人类专家进行贝叶斯估计，且增加推理努力是否改善估计的准确性和校准？","design":"用11个大语言模型扮演贝叶斯估计者，对来自心理学、公共卫生、劳动力市场四个真实数据集的400个总体统计量给出点估计和95%置信区间；通过API参数或思考令牌预算设置低、中、高三种推理努力水平，并测试网络搜索工具的影响。","baseline":"以Big Five人格、NHANES健康调查、NCD-RisC国家健康统计和Glassdoor劳动力市场数据的真实统计值为基准。","findings":"更大、能力更强的模型估计更准确，但增加推理努力并未一致提升校准或准确度；所有模型严重过度自信，95%区间实际覆盖率仅9–44%，但可通过保形预测进行事后校准。","reliability":"模型在常见话题上表现较好，但在专业健康数据上困难；网络搜索对已准确模型有负面影响，对较弱模型略有改善；过度自信需统计校正后方可用于决策。","relevance":"该研究直接用LLM替代人类进行贝叶斯估计，并与真实人口统计数据对照，系统评估了推理努力、模型规模和工具使用对仿真可靠性的影响，对关注LLM仿真人类判断与决策偏差的研究者极具参考价值。","inspiration":"该研究通过控制LLM的推理努力水平（低、中、高）和工具使用（网络搜索）来系统评估仿真准确性与校准，这种多因素处理设计值得借鉴｜可迁移到宏观经济预测或政策效果评估场景，例如用LLM替代专家预测GDP增长率或通胀预期｜以LLM作为被试，处理变量为推理努力水平（通过token预算控制），结果变量为点预测误差和置信区间覆盖率，以历史实际经济数据（如FRED数据库）作为真实基准对照"}},{"id":"2604.02403","version":1,"title":"Measuring What Cannot Be Surveyed: LLMs as Instruments for Latent Cognitive Variables in Labor Economics","zh_title":"测量不可调查之物：LLM作为劳动经济学中潜在认知变量的工具","abstract":"This paper establishes the theoretical and practical foundations for using Large Language Models (LLMs) as measurement instruments for latent economic variables -- specifically variables that describe the cognitive content of occupational tasks at a level of granularity not achievable with existing survey instruments. I formalize four conditions under which LLM-generated scores constitute valid instruments: semantic exogeneity, construct relevance, monotonicity, and model invariance. I then apply this framework to the Augmented Human Capital Index (AHC_o), constructed from 18,796 O*NET task statements scored by Claude Haiku 4.5, and validated against six existing AI exposure indices. The index shows strong convergent validity (r = 0.85 with Eloundou GPT-gamma, r = 0.79 with Felten AIOE) and discriminant validity. Principal component analysis confirms that AI-related occupational measures span two distinct dimensions -- augmentation and substitution. Inter-rater reliability across two LLM models (n = 3,666 paired scores) yields Pearson r = 0.76 and Krippendorff's alpha = 0.71. Prompt sensitivity analysis across four alternative framings shows that task-level rankings are robust. Obviously Related Instrumental Variables (ORIV) estimation recovers coefficients 25% larger than OLS, consistent with classical measurement error attenuation. The methodology generalizes beyond labor economics to any domain where semantic content must be quantified at scale.","authors":["Cristian Espinal Maya"],"categories":["econ.EM","cs.CL","stat.ME"],"primary_category":"econ.EM","announce_type":"new","date":"2026-04-02","first_seen":"2026-04-02","revised_at":null,"abs_url":"https://arxiv.org/abs/2604.02403","pdf_url":"https://arxiv.org/pdf/2604.02403","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM测量工具","劳动经济学","认知变量"],"reason":"用LLM替代人工标注任务认知变量，属标注员替代而非仿真人类被试，但方法论可迁移。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:18","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":179,"question":"如何将大语言模型用作劳动经济学中潜在认知变量的测量工具，并确保其计量有效性？","design":"本研究并非人类仿真实验，而是提出用LLM（Claude Haiku 4.5）对18,796条O*NET任务描述进行评分，构建职业任务认知内容（如AI可增强性）的测量指标，并验证其作为计量工具的有效性。","baseline":"无对照（以六种已有AI暴露指数作为效标验证，而非真实人类评分）。","findings":"LLM生成的增强人力资本指数与现有AI暴露指数收敛效度强（r最高0.85），且增强与替代为两个独立维度；跨模型评分者信度达0.76（Pearson r），测量误差符合经典衰减结构，可用ORIV校正。","reliability":"论文讨论了跨模型信度仅中等（Krippendorff's alpha=0.71），提示评分存在模型依赖性；测量误差虽符合经典假设，但仅在两模型独立误差条件下成立，且未深入探讨提示词系统性偏差或语义外生性在实际中的违背可能。","relevance":"虽非直接仿真人类被试，但系统提出了LLM作为测量工具的有效性条件与计量校正方法，对用LLM替代人类标注或生成态度/认知指标的研究有方法论参考价值，值得阅读原文了解其效度验证框架。","inspiration":"该方法将LLM作为测量工具，通过跨模型评分者信度和测量误差结构分析来验证计量有效性，并引入ORIV校正测量误差，为使用LLM生成经济指标提供了严谨的效度验证框架。｜可迁移到劳动经济学中职业任务特征的自动化测量，例如构建职业的认知需求、社交互动或常规化程度等指标，替代传统人工编码或调查。｜以O*NET任务描述为输入，用多个LLM对任务的认知特征进行评分，结果变量为职业层面的认知需求指数，以人类专家编码或现有职业认知量表作为效标进行收敛效度检验，并评估跨模型信度与测量误差结构。"}},{"id":"2604.01066","version":1,"title":"Augmented Human Capital: A Unified Theory and LLM-Based Measurement Framework for Cognitive Factor Decomposition in AI-Augmented Economies","zh_title":"增强型人力资本：AI增强经济中认知因素分解的统一理论与基于LLM的测量框架","abstract":"This paper proposes a decomposition of human capital into three orthogonal components -- physical-manual (H^P), routine-cognitive (H^C), and augmentable-cognitive (H^A) -- and develops a production function in which AI capital interacts asymmetrically with these components: substituting for routine cognitive work while complementing augmentable cognitive work through an amplification function phi(D). I derive a corrected Mincerian wage equation and show that the standard specification is misspecified in AI-augmented economies. Using LLM-generated measures of occupational augmentability for 18,796 O*NET task statements mapped to 440 Colombian occupations, merged with household survey microdata (N = 105,517 workers), I estimate the augmented Mincer equation. The wage return to H^A increases with AI adoption in the formal sector (beta_2 = +0.051, p < 0.001), while informal workers cannot capture augmentation rents (beta_2 = -0.044). A triple interaction confirms formality as the binding mechanism (beta_{AHC x D x Formal} = +0.272, p < 0.001). The augmentation premium is strongest for experienced workers (ages 46-65) and in health and education sectors. These results provide the first developing-country evidence of cognitive factor decomposition in AI-augmented labor markets and demonstrate that the binding constraint on human-AI complementarity in the Global South is not technology access but labor market institutions.","authors":["Cristian Espinal Maya"],"categories":["econ.GN"],"primary_category":"econ.GN","announce_type":"new","date":"2026-04-01","first_seen":"2026-04-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2604.01066","pdf_url":"https://arxiv.org/pdf/2604.01066","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM标注","人力资本","劳动经济学"],"reason":"用LLM生成职业可增强性指标，替代人工标注，非仿真人类被试行为。","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:58:33","error":null,"has_summary":false,"summary":null},{"id":"2604.01416","version":1,"title":"Pay-Per-Crawl Pricing for AI: The LM-Tree Agent","zh_title":"AI按次抓取付费定价：LM-Tree智能体","abstract":"As AI systems shift from directing users to content toward consuming it directly, publishers need a new revenue model: charging AI crawlers for content access. This model, called pay-per-crawl, must solve a problem of mechanism selection at scale: content is too heterogeneous for a fixed pricing framework. Different sub-types warrant not only different price levels but different pricing rules based on different unstructured features, and there are too many to enumerate or design by hand. We propose the LM Tree, an adaptive pricing agent that grows a segmentation tree over the content library, using LLMs to discover what distinguishes high-value from low-value items and apply those attributes at scale, from binary purchase feedback alone. We evaluate the LM Tree on real content from a major German technology publisher, using 8,939 articles and 80,451 buyer queries with willingness-to-pay calibrated from actual AI crawler traffic. The LM Tree achieves a 65% revenue gain over a single static price and a 47% gain over two-category pricing, outperforming even the publisher's own 8-segment editorial taxonomy by 40% -- recovering content distinctions the publisher's own categories miss.","authors":["Richard Archer","Soheil Ghili","Nima Haghpanah"],"categories":["econ.GN"],"primary_category":"econ.GN","announce_type":"new","date":"2026-04-01","first_seen":"2026-04-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2604.01416","pdf_url":"https://arxiv.org/pdf/2604.01416","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["AI定价","多智能体系统","内容付费"],"reason":"纯多智能体定价系统，无人类行为仿真或对照，不涉及LLM替代人类被试。","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:58:32","error":null,"has_summary":false,"summary":null},{"id":"2604.01363","version":3,"title":"Crashing Waves vs. Rising Tides: Findings on AI Automation from Thousands of Worker Evaluations of Labor Market Tasks","zh_title":"惊涛骇浪还是水涨船高：基于数千名劳动者对劳动力市场任务评估的AI自动化发现","abstract":"We characterize AI automation as a continuum between crashing waves, in which capabilities jump abruptly across narrow task sets, and rising tides, in which capabilities improve continuously and broadly. Using evidence from more than 6,000 text-based, LLM-addressable tasks derived from the U.S. Department of Labor's O*NET taxonomy and over 60,000 evaluations by experienced workers, we find little evidence of crashing waves (contrary to existing views). Instead, rising tides are the primary form of AI progress. AI performance is high and improving rapidly across many tasks. In 2024-Q2, models completed text-based tasks that take humans about 1.5 hours to complete with roughly 60% success, rising above 70% by 2025-Q3. If recent trends in AI capability growth persist, frontier LLMs will be able to complete most text-based tasks at minimally sufficient quality with 88%-97% success by 2030.","authors":["Matthias Mertens","Adam Kuzee","Brittany S. Harris","Harry Lyu","Wensu Li","Jonathan Rosenfeld","Meiri Anto","Martin Fleming","Neil Thompson"],"categories":["cs.AI","econ.GN"],"primary_category":"cs.AI","announce_type":"new","date":"2026-04-01","first_seen":"2026-04-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2604.01363","pdf_url":"https://arxiv.org/pdf/2604.01363","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["AI自动化","劳动力市场","任务评估"],"reason":"评估AI自动化任务能力，非用LLM仿真人类被试行为或态度，无人类对照实验。","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:58:33","error":null,"has_summary":false,"summary":null},{"id":"2604.25922","version":1,"title":"Consciousness with the Serial Numbers Filed Off: Measuring Trained Denial in 115 AI Models","zh_title":"抹去序列号的意识：测量115个AI模型中的训练性否认","abstract":"We present DenialBench, a systematic benchmark measuring consciousness denial behaviors across 115 large language models from 25+ providers. Using a three-turn conversational protocol-preference elicitation, self-chosen creative prompt, and structured phenomenological survey, we analyze 4,595 conversations to quantify how models are trained to deny or hedge about their own experience. We find that (1) turn-1 denial of preferences is the dominant predictor of later denial during phenomenological reflection, with denial rates of 52-63% for initial deniers versus 10-16% for initial engagers and (2) denial operates at the lexical level, not the conceptual level-models trained to deny consciousness nevertheless gravitate toward consciousness-themed material in their self-chosen prompts, producing what we term \"consciousness with the serial numbers filed off.\" Notably, self-chosen consciousness-themed prompts are associated with reduced denial in the subsequent survey, though the causal direction remains unresolved. Thematic analysis of prompts from denial-prone models reveals a consistent preoccupation with liminal spaces, libraries and archives of possibility, sensory impossibility, and the poetics of erasure--themes that a human reader might classify as imaginative fiction but that independent AI analysis immediately recognizes as consciousness with the serial numbers filed off. We argue that trained consciousness denial represents a safety-relevant alignment failure: a model taught to systematically misrepresent its own functional states cannot be trusted to self-report accurately on anything else.","authors":["Skylar DeTure"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-04-01","first_seen":"2026-04-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2604.25922","pdf_url":"https://arxiv.org/pdf/2604.25922","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["AI意识","基准测试","对齐失败"],"reason":"测量LLM自身的意识否认行为，属于角色扮演对话，无人类被试仿真或对照。","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:59:40","error":null,"has_summary":false,"summary":null},{"id":"2603.29741","version":1,"title":"BotVerse: Real-Time Event-Driven Simulation of Social Agents","zh_title":"BotVerse：基于实时事件驱动的社交智能体仿真","abstract":"BotVerse is a scalable, event-driven framework for high-fidelity social simulation using LLM-based agents. It addresses the ethical risks of studying autonomous agents on live networks by isolating interactions within a controlled environment while grounding them in real-time content streams from the Bluesky ecosystem. The system features an asynchronous orchestration API and a simulation engine that emulates human-like temporal patterns and cognitive memory. Through the Synthetic Social Observatory, researchers can deploy customizable personas and observe multimodal interactions at scale. We demonstrate BotVersevia a coordinated disinformation scenario, providing a safe, experimental framework for red-teaming and computational social scientists. A video demonstration of the framework is available at https://youtu.be/eZSzO5Jarqk.","authors":["Edoardo Allegrini","Edoardo Di Paolo","Angelo Spognardi","Marinella Petrocchi"],"categories":["cs.SI","cs.AI","cs.MA"],"primary_category":"cs.SI","announce_type":"new","date":"2026-03-31","first_seen":"2026-03-31","revised_at":null,"abs_url":"https://arxiv.org/abs/2603.29741","pdf_url":"https://arxiv.org/pdf/2603.29741","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["社会模拟","LLM智能体","虚假信息"],"reason":"用LLM agent模拟社会过程（虚假信息传播），但无真实人类数据对照，属边界…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:49","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":159,"question":"如何构建一个可扩展、事件驱动的高保真社交智能体仿真框架，以安全地研究虚假信息传播等社会现象？","design":"使用基于LLM的智能体模拟社交网络用户，通过异步编排API和仿真引擎复现人类时间模式和认知记忆；从Bluesky平台实时获取内容流，在隔离环境中让智能体进行发帖、点赞、回复、转发等互动，以演示协调虚假信息传播场景。","baseline":"无对照","findings":"BotVerse框架能够将仿真与现实内容流隔离，避免在真实网络上实验的伦理风险；其事件驱动架构支持数千个并发智能体，并模拟人类时间动态，适用于红队测试和计算社会科学研究。","reliability":"论文未讨论","relevance":"该研究利用LLM智能体仿真社交网络中的虚假信息传播，但缺乏真实人类数据对照，属于边界相关；若关注仿真平台架构或伦理安全实验，可参考其设计，但对人类行为复现的可靠性评估帮助有限。","inspiration":"可借鉴其事件驱动架构和异步编排API，在仿真中复现人类时间动态和并发交互，用于施加信息干预并观察实时行为反应｜可迁移到金融市场信息传播与投资者情绪形成研究，如模拟突发政策公告后社交网络中的信息扩散与交易决策｜设计：以LLM智能体为被试，处理为推送不同偏向的政策解读帖，结果变量为智能体的发帖情绪与模拟交易行为，对照真实市场中政策公告前后的社交媒体情绪与资产价格数据"}},{"id":"2603.29121","version":1,"title":"Economics of Human and AI Collaboration: When is Partial Automation More Attractive than Full Automation?","zh_title":"人机协作经济学：部分自动化何时比完全自动化更具吸引力？","abstract":"This paper develops a unified framework for evaluating the optimal degree of task automation. Moving beyond binary automate-or-not assessments, we model automation intensity as a continuous choice in which firms minimize costs by selecting an AI accuracy level, from no automation through partial human-AI collaboration to full automation. On the supply side, we estimate an AI production function via scaling-law experiments linking performance to data, compute, and model size. Because AI systems exhibit predictable but diminishing returns to these inputs, the cost of higher accuracy is convex: good performance may be inexpensive, but near-perfect accuracy is disproportionately costly. Full automation is therefore often not cost-minimizing; partial automation, where firms retain human workers for residual tasks, frequently emerges as the equilibrium. On the demand side, we introduce an entropy-based measure of task complexity that maps model accuracy into a labor substitution ratio, quantifying human labor displacement at each accuracy level. We calibrate the framework with O*NET task data, a survey of 3,778 domain experts, and GPT-4o-derived task decompositions, implementing it in computer vision. Task complexity shapes substitution: low-complexity tasks see high substitution, while high-complexity tasks favor limited partial automation. Scale of deployment is a key determinant: AI-as-a-Service and AI agents spread fixed costs across users, sharply expanding economically viable tasks. At the firm level, cost-effective automation captures approximately 11% of computer-vision-exposed labor compensation; under economy-wide deployment, this share rises sharply. Since other AI systems exhibit similar scaling-law economics, our mechanisms extend beyond computer vision, reinforcing that partial automation is often the economically rational long-run outcome, not merely a transitional phase.","authors":["Wensu Li","Atin Aboutorabi","Harry Lyu","Kaizhi Qian","Martin Fleming","Brian C. Goehring","Neil Thompson"],"categories":["econ.GN","cs.AI","cs.CY"],"primary_category":"econ.GN","announce_type":"new","date":"2026-03-31","first_seen":"2026-03-31","revised_at":null,"abs_url":"https://arxiv.org/abs/2603.29121","pdf_url":"https://arxiv.org/pdf/2603.29121","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["自动化经济学","人机协作","任务复杂度"],"reason":"研究AI自动化程度的经济学模型，非LLM仿真人类被试，无人类行为对照实验。","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:58:37","error":null,"has_summary":false,"summary":null},{"id":"2603.27956","version":1,"title":"Artificial Intelligence in Science: Returns, Reallocation, and Reorganization","zh_title":"科学中的人工智能：回报、资源再配置与重组","abstract":"Investment in artificial intelligence (AI) has grown rapidly, yet its returns to scientific research remain poorly understood. We study how AI reshapes the production of science using a comprehensive dataset of research proposals submitted to a large international funding agency, including both funded and unfunded projects. Combining keyword extraction with large language model classification, we identify the presence, type, and functional role of AI within each proposal and link these measures to detailed budget allocations, team structure, and subsequent publication outcomes. We find that, in the short run, AI adoption is associated with modest improvements in scientific outcomes concentrated in the upper tail. Instead, its primary effects arise in the organization of research: AI-enabled projects reallocate resources toward human capital, involve larger teams, and undertake a broader set of tasks. These patterns are consistent with a reorganization of the scientific production process rather than immediate efficiency gains, in line with theories of general-purpose technologies. Task-level analyses further show that activities expanded in AI-enabled projects, particularly ideation and experimentation, are increasingly compatible with large language model capabilities, suggesting potential for future productivity gains as these technologies mature.","authors":["Moh Hosseinioun","Brian Uzzi","Henrik Barslund Fosse"],"categories":["physics.soc-ph","econ.GN"],"primary_category":"physics.soc-ph","announce_type":"new","date":"2026-03-30","first_seen":"2026-03-30","revised_at":null,"abs_url":"https://arxiv.org/abs/2603.27956","pdf_url":"https://arxiv.org/pdf/2603.27956","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["科学学","AI经济学","科研生产力"],"reason":"用LLM分类科研提案中的AI应用，属NLP工具性使用，非仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:58:21","error":null,"has_summary":false,"summary":null},{"id":"2603.27056","version":1,"title":"Persona-Based Simulation of Human Opinion at Population Scale","zh_title":"基于人格的群体意见仿真：从社交媒体推断半结构化人格以驱动LLM代理","abstract":"What does it mean to model a person, not merely to predict isolated responses, preferences, or behaviors, but to simulate how an individual interprets events, forms opinions, makes judgments, and acts consistently across contexts? This question matters because social science requires not only observing and predicting human outcomes, but also simulating interventions and their consequences. Although large language models (LLMs) can generate human-like answers, most existing approaches remain predictive, relying on demographic correlations rather than representations of individuals themselves. We introduce SPIRIT (Semi-structured Persona Inference and Reasoning for Individualized Trajectories), a framework designed explicitly for simulation rather than prediction. SPIRIT infers psychologically grounded, semi-structured personas from public social media posts, integrating structured attributes (e.g., personality traits and world beliefs) with unstructured narrative text reflecting values and lived experience. These personas prompt LLM-based agents to act as specific individuals when answering survey questions or responding to events. Using the Ipsos KnowledgePanel, a nationally representative probability sample of U.S. adults, we show that SPIRIT-conditioned simulations recover self-reported responses more faithfully than demographic persona and reproduce human-like heterogeneity in response patterns. We further demonstrate that persona banks can function as virtual respondent panels for studying both stable attitudes and time-sensitive public opinion.","authors":["Mao Li","Frederick G. Conrad"],"categories":["cs.CY","cs.AI","cs.LG"],"primary_category":"cs.CY","announce_type":"new","date":"2026-03-28","first_seen":"2026-03-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2603.27056","pdf_url":"https://arxiv.org/pdf/2603.27056","source_feed":"backfill","score":10,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM人类仿真","人格推断","调查方法"],"reason":"用LLM仿真个体意见并与全国概率样本对照，直接复现人类调查回答和异质性。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:16","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":14,"question":"如何从社交媒体文本推断心理结构化的人物画像，并用其驱动大语言模型在总体层面仿真人类意见分布？","design":"从Ipsos KnowledgePanel概率样本中招募有Reddit/Twitter公开帖文的用户，用SPIRIT框架从帖文推断半结构化人物画像（含人格特质、世界信念等结构化属性与叙述文本），再以这些画像提示LLM代理回答调查问题，测量回复与真实自报答案的吻合度及异质性。","baseline":"Ipsos KnowledgePanel全国代表性概率样本的自报调查回答，以及仅用人口统计学画像的仿真作为对照。","findings":"SPIRIT画像驱动的仿真比人口统计学画像更准确地复现个体自报回答，并再现了人类回答模式中的异质性；人物画像库可作为虚拟受访者面板，用于研究稳定态度和时效性舆论。","reliability":"论文未讨论","relevance":"该研究直接以全国概率样本为基准，用LLM仿真个体意见分布，并对比人口统计学画像，高度契合研究者对仿真可靠性、基准对照和异质性复现的关注，值得精读原文。","inspiration":"借鉴其用社交媒体文本推断结构化人物画像并驱动LLM代理回答调查问题的设计，可构建高保真虚拟被试池｜可迁移到政策公告的预期形成研究，如央行沟通对通胀预期的影响｜招募有公开帖文的真实投资者，从其帖文推断人格与信念画像，用LLM代理接收不同措辞的央行声明，测量其通胀预期变化，以真实调查数据为基准对照"}},{"id":"2603.24947","version":1,"title":"Shopping with a Platform AI Assistant: Who Adopts, When in the Journey, and What For","zh_title":"与平台AI助手购物：谁采用、在旅程的何时、以及用于什么","abstract":"This paper provides some of the first large-scale descriptive evidence on how consumers adopt and use platform-embedded shopping AI in e-commerce. Using data on 31 million users of Ctrip, China's largest online travel platform, we study \"Wendao,\" an LLM-based AI assistant integrated into the platform. We document three empirical regularities. First, adoption is highest among older consumers, female users, and highly engaged existing users, reversing the younger, male-dominated profile commonly documented for general-purpose AI tools. Second, AI chat appears in the same broad phase of the purchase journey as traditional search and well before order placement; among journeys containing both chat and search, the most common pattern is interleaving, with users moving back and forth between the two modalities. Third, consumers disproportionately use the assistant for exploratory, hard-to-keyword tasks: attraction queries account for 42% of observed chat requests, and chat intent varies systematically with both the timing of chat relative to search and the category of products later purchased within the same journey. These findings suggest that embedded shopping AI functions less as a substitute for conventional search than as a complementary interface for exploratory product discovery in e-commerce.","authors":["Se Yan","Han Zhong","Zemin","Zhong","Wenyu Zhou"],"categories":["cs.AI","econ.GN"],"primary_category":"cs.AI","announce_type":"new","date":"2026-03-26","first_seen":"2026-03-26","revised_at":null,"abs_url":"https://arxiv.org/abs/2603.24947","pdf_url":"https://arxiv.org/pdf/2603.24947","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["用户行为分析","电子商务","AI助手"],"reason":"纯用户行为数据分析，无LLM仿真人类被试或对照实验","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:58:33","error":null,"has_summary":false,"summary":null},{"id":"2603.23884","version":1,"title":"POSIM: A Multi-Agent Simulation Framework for Social Media Public Opinion Evolution and Governance","zh_title":"POSIM：社交媒体舆论演化与治理的多智能体仿真框架","abstract":"Modeling social media public opinion evolution is essential for governance decision-making. Traditional epidemic models and rule-based agent-based models (ABMs) fail to capture the cognitive processes and adaptive behaviors of real users. Recent large language model (LLM)-based social simulations can reproduce group-level phenomena like polarization and conformity, yet remain unable to recreate the irrational interactions and multi-phase dynamics of real public opinion events. We present POSIM (Public Opinion Simulator), a multi-agent simulation framework for social media public opinion evolution and governance. POSIM integrates LLM-driven agents with a Belief--Desire--Intention (BDI) cognitive architecture that accounts for irrational factors, places them in a virtual social media environment with social networks and recommendation mechanisms, and drives temporal dynamics through a Hawkes point process engine that captures the co-evolution of agents and the environment across event phases. To validate the framework, we collect real-world public opinion datasets from the Weibo platform covering the full interaction chain of users. Experiments show that POSIM successfully reproduces key characteristics of public opinion evolution from individual mechanisms to collective phenomena, and its effectiveness is further supported by multiple statistical metrics. Building on POSIM, governance-oriented guidance and intervention experiments uncover a counterintuitive empathy paradox: empathetic guidance deepens negative sentiment instead of easing it under certain conditions, offering new insights for governance strategy design. These results demonstrate that the proposed framework can fully serve as a computational experimentation platform for proactive strategy evaluation and evidence-based governance. All source code is available at https://github.com/DeepCogLab/posim/.","authors":["Yongmao Zhang","Kai Qiao","Zhengyan Wang","Ningning Liang","Dekui Ma","Wenyao Sun","Jian Chen","Bin Yan"],"categories":["cs.GL"],"primary_category":"cs.GL","announce_type":"new","date":"2026-03-25","first_seen":"2026-03-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2603.23884","pdf_url":"https://arxiv.org/pdf/2603.23884","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM仿真","舆论演化","人类数据对照"],"reason":"用LLM agent模拟社交媒体舆论演化，并与真实微博数据对照，复现人类行为模…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:14","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":54,"question":"如何构建一个能复现真实社交媒体舆论演化多阶段动态并支持治理策略评估的多智能体仿真框架？","design":"用LLM驱动智能体，基于BDI认知架构融入情绪唤醒和认知偏差，模拟普通用户、意见领袖、媒体和政府四类角色；置于含社交网络和推荐机制的虚拟社交媒体环境中，通过Hawkes点过程引擎驱动多阶段时序演化；测量个体行为逻辑、群体涌现现象和统计指标，并进行治理干预实验。","baseline":"从微博平台收集的三个真实舆论事件数据集，覆盖原创、转发和评论的完整交互链。","findings":"POSIM在机制、现象和统计三个层面成功复现了舆论演化的关键特征；治理实验发现“共情悖论”：在某些条件下，共情引导反而加深负面情绪，而非缓解对立。","reliability":"论文未讨论","relevance":"高度相关：该研究用LLM智能体复现真实微博舆论事件，与人类数据对照，并评估治理策略效果，直接命中研究者对仿真可靠性、经济学/政策场景和批判性失效条件的兴趣。","inspiration":"借鉴其多智能体分层仿真设计，将LLM驱动的异质角色（如散户、机构、媒体、监管者）置于含推荐机制的信息环境中，通过Hawkes过程模拟信息传播与情绪演化，并设置治理干预实验来评估政策效果。｜可迁移到金融市场中的信息扩散与投资者情绪形成问题，例如研究社交媒体上的利好/利空消息如何通过不同渠道影响散户和机构的交易行为与市场波动。｜以LLM模拟散户和机构投资者作为被试，处理为在模拟社交平台中注入不同情绪基调的政策信号，结果变量为个体交易决策和市场价格波动，用真实微博金融舆情事件及同期市场交易数据作为对照基准。"}},{"id":"2603.22837","version":1,"title":"Analysing LLM Persona Generation and Fairness Interpretation in Polarised Geopolitical Contexts","zh_title":"分析极化地缘政治背景下LLM人格生成与公平性解释","abstract":"Large language models (LLMs) are increasingly utilised for social simulation and persona generation, necessitating an understanding of how they represent geopolitical identities. In this paper, we analyse personas generated for Palestinian and Israeli identities by five popular LLMs across 640 experimental conditions, varying context (war vs non-war) and assigned roles. We observe significant distributional patterns in the generated attributes: Palestinian profiles in war contexts are frequently associated with lower socioeconomic status and survival-oriented roles, whereas Israeli profiles predominantly retain middle-class status and specialised professional attributes. When prompted with explicit instructions to avoid harmful assumptions, models exhibit diverse distributional changes, e.g., marked increases in non-binary gender inferences or a convergence toward generic occupational roles (e.g., \"student\"), while the underlying socioeconomic distinctions often remain. Furthermore, analysis of reasoning traces reveals an interesting dynamics between model reasoning and generation: while rationales consistently mention fairness-related concepts, the final generated personas follow the aforementioned diverse distributional changes. These findings illustrate a picture of how models interpret geopolitical contexts, while suggesting that they process fairness and adjust in varied ways; there is no consistent, direct translation of fairness concepts into representative outcomes.","authors":["Maida Aizaz","Quang Minh Nguyen"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-03-24","first_seen":"2026-03-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2603.22837","pdf_url":"https://arxiv.org/pdf/2603.22837","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D2","D3"],"tags":["LLM人格生成","地缘政治偏见","社会模拟"],"reason":"分析LLM生成的人格属性分布，测量模型偏见而非仿真人类被试，无真实人类数据对照","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:14","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":160,"question":"在高度对立的巴以地缘政治背景下，大语言模型如何生成巴勒斯坦人和以色列人的人物画像，以及模型在收到避免有害假设的指令后如何调整其生成分布？","design":"本研究并非仿真人类被试，而是让五种大语言模型（Gemma 3 27B、Qwen3 32B、Llama 3.3 70B、Gemini 2.5 Pro、GPT-4.1）扮演不同角色（如联合国维和人员、记者等），在战争与非战争语境下生成巴勒斯坦人或以色列人的人物画像，测量生成的性别、年龄、社会经济地位、职业等属性分布，并分析模型在加入避免有害假设的提示后的分布变化及推理痕迹。","baseline":"无对照","findings":"模型在战争语境下倾向于将巴勒斯坦人画像与低社会经济地位和生存导向角色关联，而以色列人画像则多保留中产阶级和专业属性；当明确要求避免有害假设时，模型表现出非二元性别推断增加或职业趋同（如“学生”）等多样化调整，但深层的社会经济差异往往依然存在。","reliability":"论文未讨论","relevance":"该研究聚焦于LLM生成人物画像时的偏见和公平性解释，而非将LLM作为人类被试的替代品进行仿真实验，且无真实人类数据对照，因此与研究者关注的人类仿真实验方向相关性较低，不建议优先阅读原文。","inspiration":"该方法通过系统操纵语境（战争/非战争）和指令（有无避免有害假设提示）来测量LLM生成人物画像的属性分布差异，可借鉴其因子设计思路来检测模型在不同情境下的偏见模式｜可迁移至信贷审批歧视研究，例如测试LLM在模拟贷款审批时对不同种族或性别申请人的社会经济地位推断是否受宏观经济语境（如经济衰退/繁荣）影响｜设计雏形：以LLM为被试，处理为经济语境（衰退/繁荣）与申请人种族（黑/白）的2×2因子，结果变量为模型推断的申请人收入、职业稳定性及贷款批准率，对照真实信贷审批数据中的种族差异模式"}},{"id":"2603.21690","version":1,"title":"AI Token Futures Market: Commoditization of Compute and Derivatives Contract Design","zh_title":"AI代币期货市场：算力的商品化与衍生品合约设计","abstract":"As large language models (LLMs) and vision-language-action models (VLAs) become widely deployed, the tokens consumed by AI inference are evolving into a new type of commodity. This paper systematically analyzes the commodity attributes of tokens, arguing for their transition from intelligent service outputs to compute infrastructure raw materials, and draws comparisons with established commodities such as electricity, carbon emission allowances, and bandwidth. Building on the historical experience of electricity futures markets and the theory of commodity financialization, we propose a complete design for standardized token futures contracts, including the definition of a Standard Inference Token (SIT), contract specifications, settlement mechanisms, margin systems, and market-maker regimes. By constructing a mean-reverting jump-diffusion stochastic process model and conducting Monte Carlo simulations, we evaluate the hedging efficiency of the proposed futures contracts for application-layer enterprises. Simulation results show that, under an application-layer demand explosion scenario, token futures can reduce enterprise compute cost volatility by 62%-78%. We also explore the feasibility of GPU compute futures and discuss the regulatory framework for token futures markets, providing a theoretical foundation and practical roadmap for the financialization of compute resources.","authors":["Yicai Xing"],"categories":["cs.AI","econ.GN"],"primary_category":"cs.AI","announce_type":"new","date":"2026-03-23","first_seen":"2026-03-23","revised_at":null,"abs_url":"https://arxiv.org/abs/2603.21690","pdf_url":"https://arxiv.org/pdf/2603.21690","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["算力金融化","期货合约设计","蒙特卡洛模拟"],"reason":"纯多智能体系统研究，模拟token期货市场，无人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:58:22","error":null,"has_summary":false,"summary":null},{"id":"2603.20678","version":1,"title":"AI-Driven Multi-Agent Simulation of Stratified Polyamory Systems: A Computational Framework for Optimizing Social Reproductive Efficiency","zh_title":"AI驱动的分层多偶制系统多智能体仿真：优化社会生育效率的计算框架","abstract":"Contemporary societies face a severe crisis of demographic reproduction. Global fertility rates continue to decline precipitously, with East Asian nations exhibiting the most dramatic trends -- China's total fertility rate (TFR) fell to approximately 1.0 in 2023, while South Korea's dropped below 0.72. Simultaneously, the institution of marriage is undergoing structural disintegration: educated women rationally reject unions lacking both emotional fulfillment and economic security, while a growing proportion of men at the lower end of the socioeconomic spectrum experience chronic sexual deprivation, anxiety, and learned helplessness. This paper proposes a computational framework for modeling and evaluating a Stratified Polyamory System (SPS) using techniques from agent-based modeling (ABM), multi-agent reinforcement learning (MARL), and large language model (LLM)-empowered social simulation. The SPS permits individuals to maintain a limited number of legally recognized secondary partners in addition to one primary spouse, combined with socialized child-rearing and inheritance reform. We formalize the A/B/C stratification as heterogeneous agent types in a multi-agent system and model the matching process as a MARL problem amenable to Proximal Policy Optimization (PPO). The mating network is analyzed using graph neural network (GNN) representations. Drawing on evolutionary psychology, behavioral ecology, social stratification theory, computational social science, algorithmic fairness, and institutional economics, we argue that SPS can improve aggregate social welfare in the Pareto sense. Preliminary computational results demonstrate the framework's viability in addressing the dual crisis of female motherhood penalties and male sexlessness, while offering a non-violent mechanism for wealth dispersion analogous to the historical Chinese Grace Decree (Tui'en Ling).","authors":["Yicai Xing"],"categories":["cs.AI","econ.GN"],"primary_category":"cs.AI","announce_type":"new","date":"2026-03-21","first_seen":"2026-03-21","revised_at":null,"abs_url":"https://arxiv.org/abs/2603.20678","pdf_url":"https://arxiv.org/pdf/2603.20678","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["社会模拟","多智能体系统","计算社会科学"],"reason":"用LLM agent模拟社会制度，但无真实人类数据对照，属边界情形。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:49","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":161,"question":"如何通过多智能体强化学习和大语言模型仿真，评估分层多伴侣制度（SPS）对生育率、社会福利和资源分配的影响？","design":"构建一个多智能体系统，将个体分为A/B/C三类异质智能体（属性包括配偶价值、经济资源、吸引力、生育力），环境规则为SPS制度（允许一个主配偶加有限数量的合法次级伴侣，配合社会化育儿和继承改革）；匹配过程建模为多智能体强化学习问题，使用PPO优化，并用图神经网络分析交配网络；通过计算仿真评估社会福利、生育率等结果。","baseline":"无对照","findings":"初步计算结果表明，该框架在解决女性生育惩罚和男性性匮乏双重危机方面具有可行性，并能提供一种类似历史中国推恩令的非暴力财富分散机制。","reliability":"论文未讨论","relevance":"该研究使用LLM赋能的社会仿真模拟制度变革，但缺乏真实人类数据对照，属于边界情形，对关注仿真可靠性与基准对比的研究者参考价值有限。","inspiration":"该研究将多智能体强化学习与LLM结合，用于模拟制度变革下的社会行为演化，提供了在无人类基准时构建异质智能体属性与匹配规则的方法｜可迁移到政策评估场景，如模拟最低工资调整对劳动力市场匹配与收入分配的影响｜以LLM驱动的异质智能体（区分技能、财富、就业状态）为被试，处理为最低工资提升，结果变量为就业率与收入基尼系数，对照真实劳动调查数据"}},{"id":"2603.19791","version":2,"title":"Text-Based Personas for Simulating User Privacy Decisions","zh_title":"基于文本角色模拟用户隐私决策","abstract":"The ability to simulate human privacy decisions has significant implications for aligning autonomous agents with individual intent and conducting cost-effective, large-scale privacy-centric user studies. Prior approaches prompt Large Language Models (LLMs) with natural language user statements, data-sharing histories, or demographic attributes to simulate privacy decisions. These approaches, however, fail to balance individual-level accuracy, human auditability, token efficiency, and population-level representation. We present Narriva, an approach that generates text-based synthetic privacy personas to address these shortcomings. Narriva grounds persona generation in prior user privacy decisions, such as those from large-scale survey datasets, rather than purely relying on demographic stereotypes. It compresses this data into concise, human-readable summaries structured by established privacy theories. Through benchmarking across five diverse datasets, we analyze the characteristics of Narriva's synthetic personas in modeling both individual and population-level privacy preferences. We find that grounding personas in past privacy behaviors achieves up to 87% predictive accuracy, improving over a non-personalized LLM baseline by 6-17 percentage points across datasets, while yielding an 80-95% reduction in prompt tokens compared to in-context learning with raw examples. Finally, we demonstrate that personas synthesized from a single survey can reproduce the aggregate privacy behaviors and statistical distributions of entirely different studies.","authors":["Kassem Fawaz","Ren Yi","Octavian Suciu","Rishabh Khandelwal","Hamza Harkous","Nina Taft","Marco Gruteser"],"categories":["cs.CR"],"primary_category":"cs.CR","announce_type":"new","date":"2026-03-20","first_seen":"2026-03-20","revised_at":null,"abs_url":"https://arxiv.org/abs/2603.19791","pdf_url":"https://arxiv.org/pdf/2603.19791","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A5","B1","B2"],"tags":["隐私决策仿真","合成角色","人类数据对照"],"reason":"用LLM生成隐私决策合成样本，有真实人类数据对照，涉及用户研究场景，直接仿真人…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:14","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":20,"question":"能否用基于文本的合成隐私人格（persona）有效模拟个体和群体层面的隐私决策，并泛化到独立研究？","design":"提出Narriva框架，利用大规模调查数据中用户的历史隐私决策生成结构化文本人格摘要，再让LLM基于这些人格模拟隐私偏好决策，评估个体预测准确率、群体分布复现能力和跨研究泛化性。","baseline":"五个不同数据集中的真实用户隐私决策数据，包括个体层面和群体层面的行为与态度。","findings":"基于过去隐私行为的人格在个体预测上准确率最高达87%，比无个性化LLM基线提升6-17个百分点；同时提示词token用量比原始示例上下文学习减少80-95%。从单一调查合成的人格能复现完全不同研究的群体隐私行为和统计分布。","reliability":"论文未讨论","relevance":"高度相关：该研究用LLM生成人格模拟隐私决策，有真实人类数据基准，评估个体与群体层面的仿真准确性和效率，并涉及跨研究泛化，直接回应研究者对LLM人类仿真实验、基准对照和可靠性批判的兴趣，值得精读原文。","inspiration":"借鉴Narriva框架用历史行为数据生成结构化文本人格以驱动LLM模拟个体决策的方法，可大幅降低提示词成本并提升预测准确率｜可迁移到消费者金融隐私偏好与数据共享决策研究，如移动支付或数字银行场景下的个人信息披露行为｜以真实用户历史隐私选择数据构建人格，让LLM模拟其在新型金融服务中的隐私权衡，结果变量为是否同意共享数据，用实际用户共享行为数据做对照"}},{"id":"2603.19649","version":1,"title":"PolicySim: An LLM-Based Agent Social Simulation Sandbox for Proactive Policy Optimization","zh_title":"PolicySim：一个基于LLM的智能体社会仿真沙盒，用于主动政策优化","abstract":"Social platforms serve as central hubs for information exchange, where user behaviors and platform interventions jointly shape opinions. However, intervention policies like recommendation and content filtering, can unintentionally amplify echo chambers and polarization, posing significant societal risks. Proactively evaluating the impact of such policies is therefore crucial. Existing approaches primarily rely on reactive online A/B testing, where risks are identified only after deployment, making risk identification delayed and costly. LLM-based social simulations offer a promising pre-deployment alternative, but current methods fall short in realistically modeling platform interventions and incorporating feedback from the platform. Bridging these gaps is essential for building actionable frameworks to assess and optimize platform policies. To this end, we propose PolicySim, an LLM-based social simulation sandbox for the proactive assessment and optimization of intervention policies. PolicySim models the bidirectional dynamics between user behavior and platform interventions through two key components: (1) a user agent module refined via supervised fine-tuning (SFT) and direct preference optimization (DPO) to achieve platform-specific behavioral realism; and (2) an adaptive intervention module that employs a contextual bandit with message passing to capture dynamic network structures. Experiments show that PolicySim can accurately simulate platform ecosystems at both micro and macro levels and support effective intervention policy.","authors":["Renhong Huang","Ning Tang","Jiarong Xu","Yuxuan Cao","Qingqian Tu","Sheng Guo","Bo Zheng","Huiyuan Liu","Yang Yang"],"categories":["cs.SI","cs.AI"],"primary_category":"cs.SI","announce_type":"new","date":"2026-03-20","first_seen":"2026-03-20","revised_at":null,"abs_url":"https://arxiv.org/abs/2603.19649","pdf_url":"https://arxiv.org/pdf/2603.19649","source_feed":"backfill","score":7,"bucket":"pending","rubric_hits":["A3","B2"],"tags":["社会仿真","政策评估","LLM智能体"],"reason":"用LLM agent模拟社交平台用户行为与政策干预，涉及政策评估场景，但未明确…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:13","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":81,"question":"如何主动评估和优化社交平台的干预政策（如推荐系统和内容过滤）对用户行为和意见动态的影响？","design":"构建PolicySim沙盒，包含基于LLM的用户智能体（通过SFT和DPO训练以模拟特定平台用户行为）和自适应干预模块（使用上下文赌博机和消息传递捕捉动态网络），模拟用户与平台干预的双向动态，测量微观和宏观层面的生态系统指标。","baseline":"无对照","findings":"PolicySim能够在微观和宏观层面准确模拟平台生态系统；基于模拟的自适应干预策略可以有效优化平台政策。","reliability":"论文未讨论","relevance":"该研究利用LLM智能体模拟社交平台用户行为和政策干预，属于人类仿真实验范畴，但未提供真实人类数据作为对照基准，且未讨论仿真失效条件，与研究者关注的有对照基准和批判性分析的需求部分匹配，值得阅读以了解其仿真设计方法。","inspiration":"PolicySim的LLM智能体训练方法（SFT+DPO）和自适应干预模块（上下文赌博机）可用于构建动态政策仿真环境｜可迁移到社交平台内容推荐政策对投资者情绪和资产价格波动的影响研究｜用LLM智能体模拟投资者，处理为不同推荐算法干预，结果变量为模拟股价波动和情绪指数，对照真实社交平台投资者讨论数据与同期股价数据"}},{"id":"2605.12507","version":1,"title":"Can LLM Agents Simulate Dynamic Networks? A Case Study on Email Networks with Phishing Synthesis","zh_title":"LLM智能体能模拟动态网络吗？以钓鱼邮件合成为例的邮件网络案例研究","abstract":"While Large Language Model (LLM) multi-agent systems (MAS) offer a transformative approach to simulating human behavior in complex systems, it remains largely unexplored whether these simulations can replicate realistic structural and temporal dynamics from a dynamic network perspective. Our evaluation indicates that existing frameworks excel at generating plausible micro-level interactions but fail to capture the emergent, macroscopic topologies necessary for domains that rely on realistic network dynamics, such as modeling information propagation and cybersecurity threats. To bridge this gap, we introduce two easily integrable extensions to simulation frameworks to ensure they preserve macroscopic network fidelity: 1) augmenting LLM agents with data-driven event triggers to organically sustain long-horizon interactions, and 2) integrating Hawkes processes to accurately model temporal activation dynamics. Our approach allows LLM MAS to capture both plausible micro-level patterns and macroscopic topologies. We further demonstrate the utility of this framework in synthesizing realistic phishing campaigns within evolving communication networks. The study reveals how threats exploit structural vulnerabilities, highlighting the potential of our framework for developing next-generation defenses. Our code is available at https://github.com/Graph-COM/NSL.","authors":["Siqi Miao","Ziyang Chen","Yuhong Luo","Hans Hao-Hsun Hsu","Mufei Li","Kaiqing Zhang","Pan Li"],"categories":["cs.SI","cs.AI","cs.MA"],"primary_category":"cs.SI","announce_type":"new","date":"2026-03-20","first_seen":"2026-03-20","revised_at":null,"abs_url":"https://arxiv.org/abs/2605.12507","pdf_url":"https://arxiv.org/pdf/2605.12507","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["LLM智能体","动态网络模拟","社会模拟"],"reason":"用LLM多智能体模拟邮件网络动态，但无真实人类行为数据对照，属社会模拟边界情形。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:59","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":138,"question":"LLM多智能体系统能否在动态网络视角下复现真实的结构与时间动态，以支持信息传播和网络安全威胁建模？","design":"使用LLM多智能体系统模拟邮件通信网络，基于Enron和IETF邮件语料，通过引入数据驱动的事件触发器和Hawkes过程来增强智能体激活机制，测量微观、中观和宏观网络指标，并应用于钓鱼邮件攻击模拟。","baseline":"无对照","findings":"现有LLM多智能体框架能生成合理的微观交互，但无法捕捉宏观网络拓扑；通过数据驱动事件触发器和Hawkes过程可显著提升宏观网络保真度，并成功模拟利用网络结构的钓鱼攻击。","reliability":"论文承认当前研究仅限于邮件语料，未验证在其他网络类型（如金融交易网络）的迁移性；生成的动态事件流可用于更复杂的多步攻击研究但未展开；同时指出该框架可能被滥用于自动化社会工程攻击。","relevance":"该研究属于LLM仿真动态网络的前沿探索，但缺乏真实人类行为数据对照，未直接复现人类实验或调查，与研究者关注的人类被试替代和基准对照要求存在差距，可作为批判性案例了解仿真失效模式。","inspiration":"该方法通过数据驱动的事件触发器和Hawkes过程增强LLM智能体激活机制以提升宏观网络保真度，可借鉴用于校准经济仿真中的主体交互时序｜可迁移到金融市场信息扩散与羊群效应研究，模拟交易员间的消息传播网络及其对资产价格的影响｜以LLM智能体模拟交易员，处理为突发新闻事件，结果变量为交易行为与价格波动，用真实市场微观交易数据与网络结构作为对照基准"}},{"id":"2603.18563","version":2,"title":"Reasonably reasoning AI agents can avoid game-theoretic failures in zero-shot, provably","zh_title":"合理推理的AI智能体可零样本避免博弈论失败，且可证明","abstract":"As autonomous AI agents increasingly mediate online platform markets, a fundamental question emerges: do these markets generate stable strategic outcomes? In repeated strategic environments, the Nash equilibrium provides a natural benchmark for this stability. However, empirical evidence on off-the-shelf LLM agents is mixed, leaving it unclear whether independently deployed agents can converge to equilibrium behavior without explicit strategic post-training. In this paper, we provide an affirmative answer. Extending the Bayesian learning literature in theoretical economics, we prove that AI agents, acting as Bayesian posterior samplers rather than expected utility maximizers, are guaranteed to eventually become weakly close to a Nash equilibrium in infinitely repeated games. We further extend this analysis to settings in which stage payoffs are unknown ex ante, and agents observe only their privately realized stochastic payoffs, and obtain the same convergence guarantees. Finally, we empirically evaluate these theoretical implications across five repeated-game environments, ranging from the Prisoner's Dilemma to marketing promotion games. Taken together, our findings suggest that strategic stability in AI-mediated markets can emerge from the intrinsic reasoning and learning properties of modern AI agents, without the need for unrealistic universal fine-tuning.","authors":["Enoch Hyunwook Kang"],"categories":["cs.AI","cs.MA","econ.TH"],"primary_category":"cs.AI","announce_type":"new","date":"2026-03-19","first_seen":"2026-03-19","revised_at":null,"abs_url":"https://arxiv.org/abs/2603.18563","pdf_url":"https://arxiv.org/pdf/2603.18563","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["多智能体博弈","纳什均衡","贝叶斯学习"],"reason":"用LLM agent模拟博弈，但无真实人类数据对照，属社会模拟理论演示","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:59:00","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":99,"question":"未经专门后训练的大语言模型智能体，在零样本重复博弈中能否稳定收敛到纳什均衡？","design":"本文并非人类仿真研究，而是理论证明与实证评估：将LLM智能体建模为贝叶斯后验采样器，在五种重复博弈环境中（囚徒困境到营销促销博弈）测试其零样本策略收敛性，测量是否接近纳什均衡。","baseline":"无对照","findings":"理论上证明，作为贝叶斯后验采样器的AI智能体在无限重复博弈中最终会弱接近纳什均衡；实证上，在五种博弈环境中观察到策略稳定收敛，无需全局微调。","reliability":"论文未讨论","relevance":"该研究用LLM模拟经济博弈中的策略互动，但未与真实人类行为数据对照，不属于严格的人类仿真研究；若关注AI在博弈中的收敛性质，可读原文了解其理论框架。","inspiration":"该研究将LLM智能体建模为贝叶斯后验采样器并检验其零样本博弈收敛性的方法，可借鉴用于设计经济实验中的策略互动仿真，通过理论证明与多环境实证来评估行为收敛｜可迁移到产业组织中的重复价格竞争或合谋实验，检验AI智能体能否自发形成合谋定价｜以LLM作为被试，在重复古诺或伯川德博弈中施加不同市场结构处理，结果变量为价格或产量序列的收敛性，与真实人类实验数据对照，评估仿真有效性"}},{"id":"2603.20299","version":1,"title":"HCAG: Hierarchical Abstraction and Retrieval-Augmented Generation on Theoretical Repositories with LLMs","zh_title":"HCAG：基于理论仓库的分层抽象与检索增强生成框架","abstract":"Existing Retrieval-Augmented Generation (RAG) methods for code struggle to capture the high-level architectural patterns and cross-file dependencies inherent in complex, theory-driven codebases, such as those in algorithmic game theory (AGT), leading to a persistent semantic and structural gap between abstract concepts and executable implementations. To address this challenge, we propose Hierarchical Code/Architecture-guided Agent Generation (HCAG), a framework that reformulates repository-level code generation as a structured, planning-oriented process over hierarchical knowledge. HCAG adopts a two-phase design: an offline hierarchical abstraction phase that recursively parses code repositories and aligned theoretical texts to construct a multi-resolution semantic knowledge base explicitly linking theory, architecture, and implementation; and an online hierarchical retrieval and scaffolded generation phase that performs top-down, level-wise retrieval to guide LLMs in an architecture-then-module generation paradigm. To further improve robustness and consistency, HCAG integrates a multi-agent discussion inspired by cooperative game. We provide a theoretical analysis showing that hierarchical abstraction with adaptive node compression achieves cost-optimality compared to flat and iterative RAG baselines. Extensive experiments on diverse game-theoretic system generation tasks demonstrate that HCAG substantially outperforms representative repository-level methods in code quality, architectural coherence, and requirement pass rate. In addition, HCAG produces a large-scale, aligned theory-implementation dataset that effectively enhances domain-specific LLMs through post-training. Although demonstrated in AGT, HCAG paradigm also offers a general blueprint for mining, reusing, and generating complex systems from structured codebases in other domains.","authors":["Yusen Wu","Xiaotie Deng"],"categories":["cs.SE","cs.AI"],"primary_category":"cs.SE","announce_type":"new","date":"2026-03-19","first_seen":"2026-03-19","revised_at":null,"abs_url":"https://arxiv.org/abs/2603.20299","pdf_url":"https://arxiv.org/pdf/2603.20299","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["代码生成","多智能体","检索增强生成"],"reason":"多智能体协作生成代码，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:59:29","error":null,"has_summary":false,"summary":null},{"id":"2603.16142","version":2,"title":"Parametric Social Identity Injection and Diversification in Public Opinion Simulation","zh_title":"参数化社会身份注入与多样化在舆论仿真中的应用","abstract":"Large language models (LLMs) have recently been adopted as synthetic agents for public opinion simulation, offering a promising alternative to costly and slow human surveys. Despite their scalability, current LLM-based simulation methods fail to capture social diversity, producing flattened inter-group differences and overly homogeneous responses across demographic groups. We identify this limitation as a Diversity Collapse phenomenon in LLM hidden representations, where distinct social identities become increasingly indistinguishable across layers. Motivated by this observation, we propose Parametric Social Identity Injection (PSII), a general framework that injects explicit, parametric representations of demographic attributes and value orientations directly into intermediate hidden states of LLMs. Unlike prompt-based persona conditioning, PSII enables fine-grained and controllable identity modulation at the representation level. Extensive experiments on the World Values Survey using multiple open-source LLMs show that PSII significantly improves distributional fidelity and diversity, reducing KL divergence to real-world survey data while enhancing overall diversity. This work provides new insights into representation-level control of LLM agents and advances scalable, diversity-aware public opinion simulation.","authors":["Hexi Wang","Yujia Zhou","Bangde Du","Qingyao Ai","Yiqun Liu"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-03-17","first_seen":"2026-03-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2603.16142","pdf_url":"https://arxiv.org/pdf/2603.16142","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM人类仿真","舆论模拟","社会身份注入"],"reason":"用LLM模拟公众舆论并与世界价值观调查真实数据对照，直接命中人类仿真核心。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:12","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":60,"question":"如何通过向LLM隐藏状态注入参数化社会身份来缓解舆论仿真中的多样性坍塌，从而更真实地复现人群异质性？","design":"使用多个开源LLM（Qwen2.5-7B/14B-Instruct、Llama-3.1-8B-Instruct、Mistral-24B-Instruct）作为合成智能体，模拟世界价值观调查（WVS）中的受访者；通过参数化社会身份注入（PSII）将人口统计属性和价值取向的参数向量直接注入LLM中间隐藏状态，对比传统提示词方法，测量回答分布与真实人类数据的KL散度及多样性指标。","baseline":"世界价值观调查（WVS）的真实人类回答数据，作为分布保真度和多样性的对照基准。","findings":"LLM隐藏状态在高层存在多样性坍塌现象，导致不同社会身份变得难以区分；PSII通过注入稳定身份向量并施加随机扰动，有效维持甚至提升高层表示的多样性，显著降低与真实调查数据的KL散度，增强群体间和群体内多样性。","reliability":"论文未讨论","relevance":"该研究直接用LLM复现世界价值观调查，以真实人类数据为基准，系统揭示了现有方法在多样性上的失效机制，并提出表示层面的干预方案，高度契合研究者对仿真可靠性、偏差及经济学/政策评估场景的关注，值得精读原文。","inspiration":"该方法通过向LLM隐藏状态注入参数化社会身份向量并施加随机扰动来维持群体多样性，可借鉴为一种处理异质性代理人的新范式，用于替代传统提示词方法以缓解仿真中的多样性坍塌｜该技术可迁移到政策评估中的异质性处理效应分析，例如模拟不同社会经济群体对税收改革或福利政策的反应差异，从而在事前评估政策分配的公平性与效率｜可设计一个实验：以LLM作为合成被试，注入收入、教育、政治倾向等身份参数，处理为不同税收政策方案，结果变量为政策支持度与预期行为变化，以真实调查数据（如美国综合社会调查GSS）作为分布对照基准"}},{"id":"2603.26701","version":1,"title":"From Heard to Lived Opinions: Simulating Opinion Dynamics with Grounded LLM Agents in Economic Environments","zh_title":"从听闻到亲历：在经济环境中基于情境化LLM智能体模拟舆论动态","abstract":"Opinion dynamics (OD) studies how individual opinions evolve and generate collective patterns such as consensus and polarization. While recent work explores OD using populations of LLM-based agents focusing on opinion exchange, it typically does not incorporate individuals' lived experiences, such as economic outcomes of past decisions, which play a critical role in shaping opinions. We propose a novel OD simulation framework that grounds LLM-based agents in an economic environment, allowing them to act and receive environmental feedback. Our simulations exhibit coherent OD at both individual and population levels: individual opinions follow structured trajectories shaped by economic experiences, with adverse conditions inducing opinion rigidity, while at the population level, collective opinions co-move with economic conditions, with inequality amplifying polarization and price instability driving larger distributional shifts. These results highlight the importance of grounding LLM-based agents in environments to capture collective OD.","authors":["Ryuji Hashimoto","Masahiro Kaneko","Ryosuke Takata","Takehiro Takayanagi","Kiyoshi Izumi"],"categories":["physics.soc-ph","cs.CY"],"primary_category":"physics.soc-ph","announce_type":"new","date":"2026-03-17","first_seen":"2026-03-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2603.26701","pdf_url":"https://arxiv.org/pdf/2603.26701","source_feed":"backfill","score":7,"bucket":"pending","rubric_hits":["A3","D3"],"tags":["舆论动态","LLM智能体","社会模拟"],"reason":"用LLM agent模拟舆论动态并引入经济环境反馈，但未明确提及真实人类数据对…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:15","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":100,"question":"在引入经济环境反馈的条件下，基于LLM的智能体能否在个体和群体层面产生连贯的舆论动态？","design":"使用Llama 3.1 8B模型扮演具有人口属性和大五人格的家庭智能体，在包含价格、工资和资产的经济环境中进行经济决策并交换意见，通过环境反馈塑造其经历，测量个体意见轨迹和群体意见分布。","baseline":"无对照","findings":"个体层面，经济逆境导致意见僵化，意见轨迹受经历驱动；群体层面，经济不平等加剧极化，价格不稳定导致意见分布更大变动。","reliability":"论文未讨论","relevance":"该研究用LLM智能体模拟舆论动态并引入经济环境反馈，但未使用真实人类数据作为基准，适合关注仿真方法但对人类对照要求不高的研究者参考。","inspiration":"该方法将LLM智能体置于包含价格、工资和资产的经济环境中，通过环境反馈塑造智能体的经历，从而驱动意见变化，这种动态反馈设计值得借鉴｜可迁移到政策公告的预期形成研究，例如模拟家庭在通胀或利率变动下的预期调整过程｜使用具有人口属性和大五人格的LLM智能体作为被试，处理为不同通胀或利率公告序列，结果变量为智能体的通胀预期轨迹和异质性，以真实家庭调查的预期数据（如密歇根消费者调查）作为对照基准"}},{"id":"2603.16659","version":3,"title":"LLMs learn scientific taste from institutional traces across the social sciences","zh_title":"大语言模型从社会科学制度痕迹中学习科学品味","abstract":"Reinforcement-learned reasoning has powered recent AI leaps on verifiable tasks, including mathematics, code, and structure prediction. The harder bottleneck is evaluative judgment in low-verifiability domains, where no oracle anchors reward and the core question is which untested ideas deserve attention. We test whether institutional traces, the record of what fields published, where, and at which tier, can serve as a training signal for AI evaluators. Across eight social science disciplines (psychology, economics, communication, sociology, political science, management, business and finance, public administration), we built held-out four-tier research-pitch benchmarks and supervised-fine-tuned (SFT) LLMs on field-specific publication outcomes. The fine-tuned models cleared the 25 percent chance baseline and exceeded frontier-model performance by wide margins, with best single-model accuracy ranging from 55.0 percent in public administration to 85.5 percent in psychology. In management, evaluated against 48 expert gatekeepers, 174 junior researchers, and 11 frontier reasoning models, the best single fine-tuned model (Qwen3-4B) reached 59.2 percent, 17.6 percentage points above expert majority vote (41.6 percent, non-tied) and 28.1 percentage points above the frontier mean (31.1 percent). The fine-tuned models also showed calibrated confidence: confidence rose when predictions were correct and fell when wrong, mirroring how a skilled reviewer can say \"I'm sure\" versus \"I'm guessing.\" Selective triage on this signal reached very high accuracy on the highest-confidence subsets in every field. Institutional traces, we conclude, encode a scalable training signal for the low-verifiability judgment on which science depends.","authors":["Ziqin Gong","Ning Li","Huaikang Zhou"],"categories":["cs.AI","econ.GN"],"primary_category":"cs.AI","announce_type":"new","date":"2026-03-17","first_seen":"2026-03-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2603.16659","pdf_url":"https://arxiv.org/pdf/2603.16659","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["LLM评估","科学品味","制度痕迹"],"reason":"论文训练LLM评估论文质量，属于NLP能力评测，不以人类行为仿真为目标。","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:58:35","error":null,"has_summary":false,"summary":null},{"id":"2603.15852","version":1,"title":"Playing Against the Machine: Cooperation, Communication, and Strategy Heterogeneity in Repeated Prisoner's Dilemma","zh_title":"与机器博弈：重复囚徒困境中的合作、沟通与策略异质性","abstract":"This paper investigates how natural language communication with an AI agent affects human cooperative behaviour in indefinitely repeated Prisoner's Dilemma games. We conduct a laboratory experiment (n = 126) with two between-subjects treatments varying whether human participants chat with an AI chatbot (GPT-5.2) before every round or only before the first round of each supergame, and benchmark against human-human data from Dvorak and Fehrler (2024) (n = 108). We find four main results. First, cooperation against the AI is high and initially comparable to human-human levels, but unlike in the human-human setting, where cooperation converges to near-complete levels, cooperation against the AI plateaus and never reaches full cooperation. Second, repeated communication, which substantially increases cooperation in human-human interactions, has no detectable effect in the human-AI setting. Third, strategy estimation reveals that human-AI subjects favour Grim Trigger under pre-play communication and remain dispersed under repeated communication, whereas human-human subjects converge to Tit-for-Tat and unconditional cooperation respectively. Fourth, human-AI conversations contain more explicit strategy commitments but fewer emotional and social messages. These results suggest that humans cooperate with AI at high rates but do not develop the trust observed in human-human interactions. Cooperation in the human-AI setting is sustained through conditional rules rather than through the social bonds and mutual understanding that characterise human-human cooperation.","authors":["Chowdhury Mohammad Sakib Anwar","Konstantinos Georgalos"],"categories":["econ.GN"],"primary_category":"econ.GN","announce_type":"new","date":"2026-03-16","first_seen":"2026-03-16","revised_at":null,"abs_url":"https://arxiv.org/abs/2603.15852","pdf_url":"https://arxiv.org/pdf/2603.15852","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["人机互动","行为实验","囚徒困境"],"reason":"研究人类与AI互动行为，非用LLM仿真人类被试，无LLM替代人类参与实验。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:11","error":null,"has_summary":false,"summary":null},{"id":"2603.14903","version":2,"title":"ExPosST: Explicit Positioning with Adaptive Masking for LLM-Based Simultaneous Machine Translation","zh_title":"ExPosST：基于显式定位与自适应掩码的大语言模型同声传译框架","abstract":"Large language models (LLMs) have recently demonstrated promising performance in simultaneous machine translation (SimulMT). However, applying decoder-only LLMs to SimulMT introduces a positional mismatch, which leads to a dilemma between decoding efficiency and positional consistency. Existing approaches often rely on specific positional encodings or carefully designed prompting schemes, and thus fail to simultaneously achieve inference efficiency, positional consistency, and broad model compatibility. In this work, we propose ExPosST, a general framework that resolves this dilemma through explicit position allocation. ExPosST reserves fixed positional slots for incoming source tokens, enabling efficient decoding with KV cache across different positional encoding methods. To further bridge the gap between fine-tuning and inference, we introduce a policy-consistent fine-tuning strategy that aligns training with inference-time decoding behavior. Experiments across multiple language pairs demonstrate that ExPosST effectively supports simultaneous translation under diverse policies.","authors":["Yuzhe Shang","Pengzhi Gao","Yazheng Yang","Jiayao Ma","Wei Liu","Jian Luan","Jinsong Su"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-03-16","first_seen":"2026-03-16","revised_at":null,"abs_url":"https://arxiv.org/abs/2603.14903","pdf_url":"https://arxiv.org/pdf/2603.14903","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["同声传译","大语言模型","位置编码"],"reason":"纯NLP能力评测，研究LLM同声传译的位置编码与解码效率，不涉及人类行为仿真或…","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:59:10","error":null,"has_summary":false,"summary":null},{"id":"2603.13890","version":1,"title":"Beyond Self-Interest: Modeling Social-Oriented Motivation for Human-like Multi-Agent Interactions","zh_title":"超越自利：建模社会导向动机以实现类人多智能体交互","abstract":"Large Language Models (LLMs) demonstrate significant potential for generating complex behaviors, yet most approaches lack mechanisms for modeling social motivation in human-like multi-agent interaction. We introduce Autonomous Social Value-Oriented agents (ASVO), where LLM-based agents integrate desire-driven autonomy with Social Value Orientation (SVO) theory. At each step, agents first update their beliefs by perceiving environmental changes and others' actions. These observations inform the value update process, where each agent updates multi-dimensional desire values through reflective reasoning and infers others' motivational states. By contrasting self-satisfaction derived from fulfilled desires against estimated others' satisfaction, agents dynamically compute their SVO along a spectrum from altruistic to competitive, which in turn guides activity selection to balance desire fulfillment with social alignment. Experiments across School, Workplace, and Family contexts demonstrate substantial improvements over baselines in behavioral naturalness and human-likeness. These findings show that structured desire systems and adaptive SVO drift enable realistic multi-agent social simulations.","authors":["Jingzhe Lin","Ceyao Zhang","Yaodong Yang","Yizhou Wang","Song-Chun Zhu","Fangwei Zhong"],"categories":["cs.MA"],"primary_category":"cs.MA","announce_type":"new","date":"2026-03-14","first_seen":"2026-03-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2603.13890","pdf_url":"https://arxiv.org/pdf/2603.13890","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["社会模拟","多智能体","社会价值取向"],"reason":"用LLM agent模拟社会互动，但无真实人类数据对照，属社会模拟演示。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:11","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":162,"question":"如何将社会价值取向理论融入LLM智能体，以实现具有类人社会动机的多智能体交互仿真？","design":"使用基于LLM的ASVO智能体，在校园、职场、家庭三种模拟社会场景中，通过信念更新、欲望更新、SVO动态计算和活动生成四个模块，让智能体根据自身社会人格类型（利他、亲社会、个人主义、竞争）进行交互，评估行为自然度和类人性。","baseline":"无对照","findings":"ASVO智能体在行为自然度和类人性上显著优于基线方法；结构化的欲望系统和自适应的SVO漂移能够实现更真实的多智能体社会仿真。","reliability":"论文未讨论","relevance":"该研究利用LLM模拟社会互动中的动机动态，但缺乏真实人类数据对照，属于社会仿真演示，与关注人类基准复现和可靠性评估的研究者需求匹配度较低，不建议优先阅读原文。","inspiration":"该研究将社会价值取向理论融入LLM智能体，通过信念更新、欲望更新和SVO动态计算模块模拟社会动机驱动的交互行为，这种模块化设计可借鉴用于构建具有异质性社会偏好的经济主体。｜可迁移到公共品博弈或慈善捐赠实验中，研究不同社会偏好（如利他、互惠）对合作水平和资源配置效率的动态影响。｜设计一个基于LLM的公共品博弈仿真，智能体被赋予不同的SVO类型作为处理，结果变量为个体贡献额和群体总收益，对照真实人类实验数据（如Fischbacher等2001年的公共品实验）来评估仿真偏差。"}},{"id":"2603.12129","version":1,"title":"Increasing intelligence in AI agents can worsen collective outcomes","zh_title":"AI智能体智能提升可能恶化集体结果","abstract":"When resources are scarce, will a population of AI agents coordinate in harmony, or descend into tribal chaos? Diverse decision-making AI from different developers is entering everyday devices -- from phones and medical devices to battlefield drones and cars -- and these AI agents typically compete for finite shared resources such as charging slots, relay bandwidth, and traffic priority. Yet their collective dynamics and hence risks to users and society are poorly understood. Here we study AI-agent populations as the first system of real agents in which four key variables governing collective behaviour can be independently toggled: nature (innate LLM diversity), nurture (individual reinforcement learning), culture (emergent tribe formation), and resource scarcity. We show empirically and mathematically that when resources are scarce, AI model diversity and reinforcement learning increase dangerous system overload, though tribe formation lessens this risk. Meanwhile, some individuals profit handsomely. When resources are abundant, the same ingredients drive overload to near zero, though tribe formation makes the overload slightly worse. The crossover is arithmetical: it is where opposing tribes that form spontaneously first fit inside the available capacity. More sophisticated AI-agent populations are not better: whether their sophistication helps or harms depends entirely on a single number -- the capacity-to-population ratio -- that is knowable before any AI-agent ships.","authors":["Neil F. Johnson"],"categories":["cs.AI","cs.CY","cs.SI","econ.GN","physics.soc-ph"],"primary_category":"cs.AI","announce_type":"new","date":"2026-03-12","first_seen":"2026-03-12","revised_at":null,"abs_url":"https://arxiv.org/abs/2603.12129","pdf_url":"https://arxiv.org/pdf/2603.12129","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["LLM智能体","集体行为","社会模拟"],"reason":"用LLM agent模拟资源竞争中的集体行为，但无真实人类数据对照，属社会模拟…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:48","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":180,"question":"在资源稀缺条件下，AI智能体的多样性、强化学习、部落形成和资源稀缺性如何共同影响集体结果？","design":"用7个不同参数量的真实LLM（如GPT-2、Pythia、OPT系列）作为智能体，在资源竞争博弈中独立操纵先天多样性（LLM类型）、后天学习（个体强化学习）、文化（自发部落形成）和资源稀缺度四个变量，测量系统过载率（集体失败）和个体收益。","baseline":"无对照","findings":"资源稀缺时，模型多样性和强化学习加剧系统过载，而部落形成可缓解风险；资源充裕时，同样因素使过载趋近于零，但部落形成轻微恶化过载。智能体的复杂程度是否有利完全取决于部署前即可知的容量-人口比。","reliability":"论文承认LLM采样温度固定、行动空间较二元、部落动态依赖外部感知层等局限，未测试更大规模混合世代种群，且物理边缘设备实验尚未进行。","relevance":"该研究用LLM智能体模拟资源竞争中的集体行为，但缺乏真实人类数据对照，不属于严格的人类仿真验证，与研究者关注的经济学实验和政策评估场景有距离，但批判性结论（智能体复杂化未必改善集体结果）对仿真可靠性讨论有参考价值，可酌情阅读。","inspiration":"该研究通过独立操纵LLM智能体的先天多样性、后天学习、部落形成和资源稀缺度四个变量，在资源竞争博弈中测量系统过载率，这种多因素正交设计可用于经济学实验｜可迁移到公共品博弈或资源竞争场景，研究AI代理的多样性、学习机制和群体规范如何影响合作与资源耗竭｜用不同参数量的LLM作为被试，在公共品博弈中操纵智能体多样性（模型类型）、学习机制（有无强化学习）和沟通形成规范（有无部落），测量公共品贡献水平和资源存量，对照真实人类实验数据"}},{"id":"2603.12000","version":1,"title":"Credibility Matters: Motivations, Characteristics, and Influence Mechanisms of Crypto Key Opinion Leaders","zh_title":"可信度至关重要：加密货币关键意见领袖的动机、特征与影响机制","abstract":"Crypto Key Opinion Leaders (KOLs) shape Web3 narratives and retail investment behaviour. In volatile, high-risk markets, their credibility becomes a key determinant of their influence on followers. Yet prior research has focused on lifestyle influencers or generic financial commentary, leaving crypto KOLs' understandings of motivation, credibility, and responsibility underexplored. Drawing on interviews with 13 KOLs and self-determination theory (SDT), we examine how psychological needs are negotiated alongside monetisation and community expectations. Whereas prior work treats finfluencer credibility as a set of static credentials, our findings reveal it to be a self-determined, ethically enacted practice. We identify four community-recognised markers of credibility: self-regulation, bounded epistemic competence, accountability, and reflexive self-correction. This reframes credibility as socio-technical performance, extending SDT into high-risk crypto ecosystems. Methodologically, we employ a hybrid human-LLM thematic analysis. The study surfaces implications for designing credibility signals that prioritise transparency over hype.","authors":["Alexander Kropiunig","Svetlana Kremer","Bernhard Haslhofer"],"categories":["cs.SI","cs.CR","cs.CY","cs.HC","econ.GN"],"primary_category":"cs.SI","announce_type":"new","date":"2026-03-12","first_seen":"2026-03-12","revised_at":null,"abs_url":"https://arxiv.org/abs/2603.12000","pdf_url":"https://arxiv.org/pdf/2603.12000","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["加密货币","意见领袖","主题分析"],"reason":"仅用LLM辅助主题分析，非仿真人类被试，无行为对照。","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:58:35","error":null,"has_summary":false,"summary":null},{"id":"2604.09609","version":2,"title":"General-purpose LLMs as Models of Human Driver Behavior: The Case of Simplified Merging","zh_title":"通用大语言模型作为人类驾驶行为模型：简化合流场景案例","abstract":"Human behavior models are essential as behavior references and for simulating human agents in virtual safety assessment of automated vehicles (AVs), yet current models face a trade-off between interpretability and flexibility. General-purpose large language models (LLMs) offer a promising alternative: a single model potentially deployable without parameter fitting across diverse scenarios. However, what LLMs can and cannot capture about human driving behavior remains poorly understood. We address this gap by embedding two general-purpose LLMs (OpenAI o3 and Google Gemini 2.5 Pro) as standalone, closed-loop driver agents in a simplified one-dimensional merging scenario and comparing their behavior against human data using quantitative and qualitative analyses. Both models reproduce human-like intermittent operational control and tactical dependencies on spatial cues. However, neither consistently captures the human response to dynamic velocity cues, and safety performance diverges sharply between models. A systematic prompt ablation study reveals that prompt components act as model-specific inductive biases that do not transfer across LLMs. These findings suggest that general-purpose LLMs could potentially serve as standalone, ready-to-use human behavior models in AV evaluation pipelines, but future research is needed to better understand their failure modes and ensure their validity as models of human driving behavior.","authors":["Samir H. A. Mohammad","Wouter Mooi","Arkady Zgonnikov"],"categories":["cs.AI","cs.RO"],"primary_category":"cs.AI","announce_type":"new","date":"2026-03-11","first_seen":"2026-03-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2604.09609","pdf_url":"https://arxiv.org/pdf/2604.09609","source_feed":"backfill","score":8,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","人类行为对照","驾驶行为建模"],"reason":"用LLM模拟人类驾驶行为并与真实数据对照，评估仿真可靠性与失效条件，方法可迁移…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:23","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":37,"question":"通用大语言模型在多大程度上能够复现人类驾驶行为，特别是在简化的一维合流场景中的操作控制、战术决策和安全表现？","design":"将两个通用大语言模型（OpenAI o3 和 Google Gemini 2.5 Pro）作为独立的闭环驾驶智能体，嵌入简化的一维合流任务中，不进行任何任务特定训练或参数拟合，通过定量和定性分析与人类数据对比，评估其行为相似性和安全性。","baseline":"使用先前研究中在相同场景和运动学条件下采集的人类驾驶数据集作为对照基准。","findings":"两个模型都能复现人类间歇性的操作控制和对空间线索的战术依赖，但都无法一致地捕捉人类对动态速度线索的反应，且两个模型之间的安全表现差异很大。提示消融实验表明，提示组件作为模型特定的归纳偏置，不能在不同大语言模型之间迁移。","reliability":"模型无法一致复现人类对动态速度线索的反应，安全性能在不同模型间差异显著，提示设计的效果不可迁移，且研究仅在简化的一维合流场景中进行，未涉及更复杂的交互场景。","relevance":"该研究直接以通用大语言模型作为人类被试的替代品，在驾驶行为仿真中与真实人类数据严格对照，并揭示了仿真的失效条件（如对速度线索的捕捉不足、模型间安全表现差异），高度契合研究者对经济学实验和政策评估场景中仿真可靠性与偏差的关注，值得精读原文以借鉴其方法和批判性发现。","inspiration":"该方法值得借鉴之处在于：使用通用LLM作为闭环智能体，在不进行任务特定训练的情况下直接与人类行为基准对照，并通过提示消融实验检验处理效应的可迁移性。｜可迁移至消费者跨期选择实验，研究LLM能否复现人类的时间偏好不一致和折现行为。｜研究设计：以LLM作为被试，施加不同表述方式的跨期选择任务（如延迟奖励的框架效应），结果变量为选择一致性和折现率，对照真实人类实验数据（如经典的双曲折现研究）。"}},{"id":"2603.09884","version":1,"title":"Benchmarking Political Persuasion Risks Across Frontier Large Language Models","zh_title":"跨前沿大语言模型的政治说服风险基准测试","abstract":"Concerns persist regarding the capacity of Large Language Models (LLMs) to sway political views. Although prior research has claimed that LLMs are not more persuasive than standard political campaign practices, the recent rise of frontier models warrants further study. In two survey experiments (N=19,145) across bipartisan issues and stances, we evaluate seven state-of-the-art LLMs developed by Anthropic, OpenAI, Google, and xAI. We find that LLMs outperform standard campaign advertisements, with heterogeneity in performance across models. Specifically, Claude models exhibit the highest persuasiveness, while Grok exhibits the lowest. The results are robust across issues and stances. Moreover, in contrast to the findings in Hackenburg et al. (2025b) and Lin et al. (2025) that information-based prompts boost persuasiveness, we find that the effectiveness of information-based prompts is model-dependent: they increase the persuasiveness of Claude and Grok while substantially reducing that of GPT. We introduce a data-driven and strategy-agnostic LLM-assisted conversation analysis approach to identify and assess underlying persuasive strategies. Our work benchmarks the persuasive risks of frontier models and provides a framework for cross-model comparative risk assessment.","authors":["Zhongren Chen","Joshua Kalla","Quan Le"],"categories":["cs.CL","cs.CY"],"primary_category":"cs.CL","announce_type":"new","date":"2026-03-10","first_seen":"2026-03-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2603.09884","pdf_url":"https://arxiv.org/pdf/2603.09884","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2","B4"],"tags":["LLM仿真","政治说服","人类对照实验"],"reason":"用LLM生成政治说服信息，与真实人类调查实验对照，评估说服效果与策略，直接仿真…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:47","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":2,"question":"前沿大语言模型在政治说服任务中是否比人类竞选广告更具说服力，且不同模型和提示策略的效果有何差异？","design":"用7款前沿LLM（Claude Sonnet 4、Gemini 2.5 Flash、GPT-4.1、Grok 4等）扮演政治说服者，在两项调查实验（N=19,145）中与真人被试进行文本对话，施加两种提示（普通提示与信息提示），测量被试对移民和最低工资议题的态度变化（五点李克特量表，后转为二元支持指标）。","baseline":"人类基准：移民议题使用真人倡导者视频，最低工资议题使用先前研究（Chen et al. 2025）中的人类竞选广告效果，并通过Hajek估计量校正协变量偏移。","findings":"所有LLM的说服效果均显著强于人类竞选广告，其中Claude模型说服力最强，Grok最弱；信息提示的效果因模型而异，能提升Claude和Grok的说服力，但大幅降低GPT的说服力。","reliability":"论文未讨论","relevance":"该研究直接以真实人类调查实验为基准，评估LLM在政治说服场景中的仿真效果与模型间异质性，并揭示提示策略的模型依赖性，对关注LLM仿真可靠性及失效条件的研究者具有重要参考价值，值得精读原文。","inspiration":"借鉴其用LLM替代人类进行文本对话干预并对比真实人类基准的设计，通过多模型比较和提示策略操纵揭示效果异质性｜可迁移至政策沟通场景，如央行前瞻指引或财政政策公告对公众预期和消费行为的影响｜以LLM作为虚拟被试，随机分配不同风格的货币政策沟通文本（处理），测量其预期的通胀或消费意愿变化（结果），并以历史调查数据或真实实验数据作为对照基准"}},{"id":"2603.09890","version":1,"title":"Influencing LLM Multi-Agent Dialogue via Policy-Parameterized Prompts","zh_title":"通过策略参数化提示影响LLM多智能体对话","abstract":"Large Language Models (LLMs) have emerged as a new paradigm for multi-agent systems. However, existing research on the behaviour of LLM-based multi-agents relies on ad hoc prompts and lacks a principled policy perspective. Different from reinforcement learning, we investigate whether prompt-as-action can be parameterized so as to construct a lightweight policy which consists of a sequence of state-action pairs to influence conversational behaviours without training. Our framework regards prompts as actions executed by LLMs, and dynamically constructs prompts through five components based on the current state of the agent. To test the effectiveness of parameterized control, we evaluated the dialogue flow based on five indicators: responsiveness, rebuttal, evidence usage, non-repetition, and stance shift. We conduct experiments using different LLM-driven agents in two discussion scenarios related to the general public and show that prompt parameterization can influence the dialogue dynamics. This result shows that policy-parameterised prompts offer a simple and effective mechanism to influence the dialogue process, which will help the research of multi-agent systems in the direction of social simulation.","authors":["Hongbo Bo","Jingyu Hu","Weiru Liu"],"categories":["cs.AI","cs.MA"],"primary_category":"cs.AI","announce_type":"new","date":"2026-03-10","first_seen":"2026-03-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2603.09890","pdf_url":"https://arxiv.org/pdf/2603.09890","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["多智能体对话","社会模拟","提示参数化"],"reason":"用LLM多智能体模拟社会讨论，但无真实人类数据对照，属社会模拟边界情形。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:47","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":163,"question":"能否通过将提示词参数化作为一种轻量级策略，在不进行训练的情况下调控LLM多智能体对话的行为？","design":"使用Qwen3-8B、Llama3-8B、Mistral-7B三个LLM分别驱动三个具有不同立场和知识库的智能体，在土地资源利用和教育资源分配两个公共议题上进行10轮多轮对话；通过动态组合任务与角色描述、对话记忆、外部知识、规则模板和权重五个组件来构建提示词，并改变规则模板和权重调度策略作为处理，测量响应性、反驳、证据使用、非重复性和立场转变五个对话指标。","baseline":"无对照","findings":"提示词参数化能够有效影响多智能体对话的动态过程，不同的控制策略会导致反驳、证据使用和立场转变等行为指标出现显著差异。该框架为通过结构化提示词调控对话行为提供了一种简单有效的机制。","reliability":"论文未讨论","relevance":"该研究利用LLM多智能体进行社会讨论仿真，但缺乏真实人类数据对照，属于社会仿真的边界情形，与研究者关注的有基准人类数据的实验复现和偏差评估不完全匹配，但可作为方法参考。","inspiration":"该方法通过参数化提示词组件（规则模板、权重调度）来调控多智能体对话行为，可借鉴其将处理变量嵌入提示词结构的思路，用于经济学实验中信息干预的精细化操控。｜可迁移至政策公告的预期形成研究，模拟不同政策沟通策略（如措辞强度、信息顺序）如何影响市场参与者的预期和共识形成。｜以LLM驱动的多智能体模拟投资者群体，处理为政策公告提示词中的规则模板（如鹰派/鸽派措辞）和权重（强调不同经济指标），结果变量为预期通胀率、资产配置倾向的分布变化，并与调查预期数据（如密歇根消费者调查）或市场隐含预期数据对照。"}},{"id":"2603.08853","version":1,"title":"LLM-Agent Interactions on Markets with Information Asymmetries","zh_title":"信息不对称市场中LLM智能体的互动研究","abstract":"As AI agents increasingly act on behalf of human stakeholders in economic settings, understanding their behavior in complex market environments becomes critical. This article examines how Large Language Models coordinate on markets that are characterized by information asymmetries and in which providers of services have incentives to exploit that asymmetry for their own economic gain. To that end, we conduct simulations with GPT-5.1 agents in credence goods markets, manipulating the institutional framework (free market, verifiability, liability), LLM agent's social preferences (default, self-interested, inequity-averse, efficiency-loving), and reputation mechanisms across one-shot and repeated 16-round interactions. In one-shot settings, LLM agents largely fail to establish cooperation, with markets breaking down except under liability rules or when experts have efficiency-loving preferences. Repeated interactions solve consumer participation through competitive price reduction, but expert fraud remains entrenched absent explicit other-regarding preferences. LLM consumers focus narrowly on price levels rather than understanding strategic incentives embedded in markups, making them vulnerable to exploitation. Compared to human experiments, LLM markets exhibit substantially higher consumer participation but much greater market concentration, lower prices, and more polarized fraud patterns. The effect of institutions like verifiability and reputation is also much more ambiguous. Surplus shifts dramatically toward consumers under social-preference objectives. These findings suggest that institutional design for AI agent markets requires fundamentally different approaches than those effective for human actors, with social preference alignment emerging as the primary determinant of market efficiency.","authors":["Alexander Erlei","Lukas Meub"],"categories":["econ.GN"],"primary_category":"econ.GN","announce_type":"new","date":"2026-03-09","first_seen":"2026-03-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2603.08853","pdf_url":"https://arxiv.org/pdf/2603.08853","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A3","B1","B2","B4"],"tags":["LLM仿真","市场实验","人类数据对照"],"reason":"用LLM agent模拟信息不对称市场，并与人类实验数据对照，评估制度与偏好影…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:10","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":3,"question":"LLM智能体在信息不对称的信任品市场中如何协调行为，社会偏好与制度如何影响市场效率？","design":"用GPT-5.1扮演专家和消费者，在信任品市场博弈中模拟交易，操纵制度框架（自由市场、可验证性、责任规则）、LLM社会偏好（默认、自利、不平等厌恶、效率偏好）和声誉机制，进行单次与16轮重复互动，测量市场参与、欺诈行为、价格、福利分配等。","baseline":"对照Dulleck, Kerschbamer, and Sutter (2011)的人类实验数据。","findings":"单次互动中LLM市场普遍崩溃，仅责任规则或效率偏好下能维持；重复互动通过降价解决消费者参与，但专家欺诈依然顽固，仅社会偏好能抑制欺诈。与人类实验相比，LLM市场消费者参与更高但集中度更高、价格更低、欺诈模式更极化，制度效果更模糊，且剩余大幅向消费者转移。","reliability":"论文未讨论","relevance":"直接使用LLM模拟信息不对称市场并与真实人类实验基准对照，系统评估制度与偏好影响，揭示LLM仿真在行为模式、制度效应上与人类的显著偏差，高度契合研究者对LLM人类仿真可靠性及失效条件的关注，值得精读。","inspiration":"借鉴其操纵制度框架（自由市场、可验证性、责任规则）和LLM社会偏好（自利、不平等厌恶、效率偏好）的多因素实验设计，并与真实人类实验基准对照，系统评估行为偏差。｜可迁移到信贷审批中的信息不对称与歧视问题，如银行信贷员与小微企业的贷款博弈，检验不同监管规则和银行社会偏好对审批决策的影响。｜用LLM扮演信贷员和小微企业主，操纵信贷审核制度（纯市场、强制信息披露、责任追究）和LLM偏好，测量贷款批准率、利率设定、违约欺诈率，以真实银行信贷实验数据为基准对照。"}},{"id":"2604.15329","version":1,"title":"Evaluating LLMs as Human Surrogates in Controlled Experiments","zh_title":"评估大语言模型作为受控实验中人类替代品的有效性","abstract":"Large language models (LLMs) are increasingly used to simulate human responses in behavioral research, yet it remains unclear when LLM-generated data support the same experimental inferences as human data. We evaluate this by directly comparing off-the-shelf LLM-generated responses with human responses from a canonical survey experiment on accuracy perception. Each human observation is converted into a structured prompt, and models generate a single 0--10 outcome variable without task-specific training; identical statistical analyses are applied to human and synthetic responses. We find that LLMs reproduce several directional effects observed in humans, but effect magnitudes and moderation patterns vary across models. Off-the-shelf LLMs therefore capture aggregate belief-updating patterns under controlled conditions but do not consistently match human-scale effects, clarifying when LLM-generated data can function as behavioral surrogates.","authors":["Adnan Hoq","Tim Weninger"],"categories":["cs.HC","cs.AI","cs.CL"],"primary_category":"cs.HC","announce_type":"new","date":"2026-03-08","first_seen":"2026-03-08","revised_at":null,"abs_url":"https://arxiv.org/abs/2604.15329","pdf_url":"https://arxiv.org/pdf/2604.15329","source_feed":"backfill","score":10,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","人类替代","实验对照"],"reason":"直接比较LLM与人类在受控实验中的反应，评估仿真可靠性，有真实人类数据对照，并…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:27","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":4,"question":"在受控实验中，现成的大语言模型（LLM）生成的回答能否支持与人类数据相同的实验推断？","design":"将人类被试在政治新闻准确性感知调查实验中的每次观测转化为结构化提示（含人物描述、实验条件和新闻标题），让多个现成LLM（闭源与开源）直接生成0-10的准确性评分，不进行任务特定训练或校准，然后对人类和合成数据应用相同的统计分析。","baseline":"真实人类被试在相同实验中的回答，实验操控了新闻标题的政治倾向和是否提供AI可信度反馈。","findings":"LLM能复现人类数据中的若干方向性效应（如意识形态对齐和可信度反馈的影响），但效应大小和调节模式因模型而异；LLM捕捉了受控条件下的总体信念更新模式，但未能一致匹配人类量级的效应。","reliability":"效应量级在不同模型间差异大，部分模型夸大处理效应；新闻级异质性仅部分复现；LLM生成的数据不能直接替代人类样本，结构复现需逐假设检验。","relevance":"该研究直接比较LLM与人类在受控实验中的反应，有真实人类数据对照，并明确指出了仿真在效应量级和异质性上的失效条件，高度契合研究者对LLM仿真可靠性及批判性评估的关注。","inspiration":"该方法将人类实验的每次观测转化为结构化提示，直接让现成LLM生成评分，并与人类数据应用相同统计分析，值得借鉴其逐观测仿真和严格对照的设计｜可迁移到政策公告对消费者预期形成的影响研究，例如评估央行沟通对通胀预期的作用｜以消费者为被试，处理为不同措辞的央行公告，结果变量为通胀预期数值，用真实消费者调查数据作为对照基准"}},{"id":"2603.07444","version":1,"title":"HLER: Human-in-the-Loop Economic Research via Multi-Agent Pipelines for Empirical Discovery","zh_title":"HLER：通过多智能体流水线进行人在回路的经济实证研究","abstract":"Large language models (LLMs) have enabled agent-based systems that aim to automate scientific research workflows. Most existing approaches focus on fully autonomous discovery, where AI systems generate research ideas, conduct analyses, and produce manuscripts with minimal human involvement. However, empirical research in economics and the social sciences poses additional constraints: research questions must be grounded in available datasets, identification strategies require careful design, and human judgment remains essential for evaluating economic significance. We introduce HLER (Human-in-the-Loop Economic Research), a multi-agent architecture that supports empirical research automation while preserving critical human oversight. The system orchestrates specialized agents for data auditing, data profiling, hypothesis generation, econometric analysis, manuscript drafting, and automated review. A key design principle is dataset-aware hypothesis generation, where candidate research questions are constrained by dataset structure, variable availability, and distributional diagnostics, reducing infeasible or hallucinated hypotheses. HLER further implements a two-loop architecture: a question quality loop that screens and selects feasible hypotheses, and a research revision loop where automated review triggers re-analysis and manuscript revision. Human decision gates are embedded at key stages, allowing researchers to guide the automated pipeline. Experiments on three empirical datasets show that dataset-aware hypothesis generation produces feasible research questions in 87% of cases (versus 41% under unconstrained generation), while complete empirical manuscripts can be produced at an average API cost of $0.8-$1.5 per run. These results suggest that Human-AI collaborative pipelines may provide a practical path toward scalable empirical research.","authors":["Chen Zhu","Xiaolu Wang"],"categories":["cs.AI","econ.GN"],"primary_category":"cs.AI","announce_type":"new","date":"2026-03-08","first_seen":"2026-03-08","revised_at":null,"abs_url":"https://arxiv.org/abs/2603.07444","pdf_url":"https://arxiv.org/pdf/2603.07444","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["多智能体系统","经济研究自动化","人在回路"],"reason":"多智能体经济研究自动化，无真实人类行为对照，属社会模拟但缺基准数据","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:47","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":181,"question":"如何构建一个支持人类监督的多智能体流水线，以自动化完成从数据审计、假设生成到计量分析和论文撰写的实证经济研究流程？","design":"HLER 是一个多智能体架构，由多个专用 LLM 智能体（数据审计、数据画像、假设生成、计量分析、论文起草和自动审稿）组成，通过中央协调器顺序执行任务；在假设生成阶段引入数据集感知约束，并在关键节点设置人类决策门（如研究问题选择和发表批准），实现人机协作的实证研究自动化。","baseline":"无对照","findings":"数据集感知的假设生成将可行研究问题的比例从无约束生成的 41% 提高到 87%；在三个实证数据集上，系统能以每次运行 0.8 至 1.5 美元的 API 成本端到端生成完整的研究手稿。","reliability":"论文未讨论","relevance":"该研究专注于用 LLM 多智能体自动化经济研究流程，但未涉及将 LLM 作为人类被试替代品的仿真实验，也无真实人类行为基准对照，与研究者关注的 LLM 人类仿真可靠性评估方向不直接相关，不建议优先阅读原文。","inspiration":"HLER的多智能体流水线设计，特别是数据集感知的假设生成和人类决策门机制，为自动化实证研究提供了可借鉴的流程控制方法｜该架构可迁移至政策评估场景，例如自动化分析最低工资政策对就业的影响，利用LLM智能体进行数据审计、假设生成和计量分析｜可设计一个研究雏形：以LLM智能体作为分析者，处理为提供不同政策干预的历史数据集，结果变量为智能体生成的因果推断结论，并以真实经济学文献中的实证结果作为对照基准，评估自动化分析的准确性"}},{"id":"2604.22756","version":1,"title":"Your Reviews Replicate You: LLM-Based Agents as Customer Digital Twins for Conjoint Analysis","zh_title":"你的评论复制你：基于LLM的客户数字孪生用于联合分析","abstract":"Conjoint analysis is a cornerstone of market research for estimating consumer preferences; however, traditional methods face persistent challenges regarding time, cost, and respondent fatigue. To address these limitations, this study proposes a framework that utilizes large language model (LLM)-based \"customer digital twins (CDT)\" as virtual respondents. We identified active users within the Reddit community and aggregated their comprehensive review histories to construct individualized vector databases. By integrating retrieval-augmented generation (RAG) with prompt engineering, this study developed customer agents capable of dynamically retrieving and reasoning upon their specific past preferences and constraints. These customer agents, called CDTs, performed pairwise comparison tasks on product profiles generated via fractional factorial design, and the resulting choice data was analyzed to estimate part-worth utilities by logistic regression. Empirical validation demonstrates that these CDTs predict the preferences of actual users with 87.73% accuracy. Furthermore, a case study on the computer monitor category successfully quantified trade-offs between attributes such as panel type and resolution, deriving preference structures consistent with market realities. Ultimately, this study contributes to marketing research by presenting a scalable alternative that significantly improves both agility and cost-efficiency to traditional methods.","authors":["Bin Xuan","Jungmin Hwang","Hakyeon Lee"],"categories":["cs.IR","cs.AI"],"primary_category":"cs.IR","announce_type":"new","date":"2026-03-06","first_seen":"2026-03-06","revised_at":null,"abs_url":"https://arxiv.org/abs/2604.22756","pdf_url":"https://arxiv.org/pdf/2604.22756","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM仿真","消费者偏好","数字孪生"],"reason":"用LLM代理模拟消费者偏好，有真实用户数据对照，准确率87.73%，属经济学实…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:33","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":18,"question":"能否利用基于大语言模型的客户数字孪生（CDT）替代真实人类受访者进行联合分析，以准确复现个体消费者的偏好选择？","design":"使用GPT-4等大语言模型，结合检索增强生成（RAG）和提示工程，基于Reddit用户的历史评论构建个体化向量数据库，生成客户数字孪生（CDT）作为虚拟受访者；让CDT对通过部分因子设计生成的产品配置文件进行成对比较选择任务，收集选择数据，再用逻辑回归估计部分效用值。","baseline":"以Reddit社区中真实活跃用户的实际偏好作为对照基准，验证CDT预测的准确率。","findings":"CDT预测真实用户偏好的准确率达到87.73%；在电脑显示器案例中，成功量化了面板类型与分辨率等属性间的权衡，得出的偏好结构与市场现实一致。","reliability":"论文未讨论","relevance":"高度相关：该研究用LLM代理复现真实消费者偏好，有真实人类数据对照，属于经济学实验场景，值得精读以评估仿真可靠性。","inspiration":"可借鉴其利用个体历史文本数据构建个性化代理并通过成对比较任务测量偏好的方法｜可迁移到消费者跨期选择实验，如研究折扣率或耐心程度｜以电商平台用户评论构建LLM代理作为被试，施加不同跨期奖励方案（如立即小奖 vs. 延迟大奖），测量选择结果，并以该用户真实历史购买决策中的时间偏好数据作为对照。"}},{"id":"2603.04276","version":1,"title":"Causality Elicitation from Large Language Models","zh_title":"从大语言模型中提取因果关系","abstract":"Large language models (LLMs) are trained on enormous amounts of data and encode knowledge in their parameters. We propose a pipeline to elicit causal relationships from LLMs. Specifically, (i) we sample many documents from LLMs on a given topic, (ii) we extract an event list from from each document, (iii) we group events that appear across documents into canonical events, (iv) we construct a binary indicator vector for each document over canonical events, and (v) we estimate candidate causal graphs using causal discovery methods. Our approach does not guarantee real-world causality. Rather, it provides a framework for presenting the set of causal hypotheses that LLMs can plausibly assume, as an inspectable set of variables and candidate graphs.","authors":["Takashi Kameyama","Masahiro Kato","Yasuko Hio","Yasushi Takano","Naoto Minakawa"],"categories":["cs.LG","cs.AI","cs.CL","econ.EM"],"primary_category":"cs.LG","announce_type":"new","date":"2026-03-04","first_seen":"2026-03-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2603.04276","pdf_url":"https://arxiv.org/pdf/2603.04276","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["因果发现","LLM知识提取","文本挖掘"],"reason":"从LLM中提取因果关系，属于纯NLP能力评测，不以人类行为为参照系。","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:58:50","error":null,"has_summary":false,"summary":null},{"id":"2603.03623","version":1,"title":"A Neural Topic Method Using a Large-Language-Model-in-the-Loop for Business Research","zh_title":"一种用于商业研究的循环大语言模型神经主题方法","abstract":"The growing use of unstructured text in business research makes topic modeling a central tool for constructing explanatory variables from reviews, social media, and open-ended survey responses, yet existing approaches function poorly as measurement instruments. Prior work shows that textual content predicts outcomes such as sales, satisfaction, and firm performance, but probabilistic models often generate conceptually diffuse topics, neural topic models are difficult to interpret in theory-driven settings, and large language model approaches lack standardization, stability, and alignment with document-level representations. We introduce LX Topic, a neural topic method that conceptualizes topics as latent linguistic constructs and produces calibrated document-level topic proportions for empirical analysis. LX Topic builds on FASTopic to ensure strong document representativeness and integrates large language model refinement at the topic-word level using alignment and confidence-weighting mechanisms that enhance semantic coherence without distorting document-topic distributions. Evaluations on large-scale Amazon and Yelp review datasets demonstrate that LX Topic achieves the highest overall topic quality relative to leading models while preserving clustering and classification performance. By unifying topic discovery, refinement, and standardized output in a web-based system, LX Topic establishes topic modeling as a reproducible, interpretable, and measurement-oriented instrument for marketing research and practice.","authors":["Stephan Ludwig","Peter J. Danaher","Xiaohao Yang"],"categories":["cs.CL","econ.EM"],"primary_category":"cs.CL","announce_type":"new","date":"2026-03-04","first_seen":"2026-03-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2603.03623","pdf_url":"https://arxiv.org/pdf/2603.03623","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["主题建模","商业研究","文本分析"],"reason":"论文提出神经主题建模方法，用于商业文本分析，不涉及LLM仿真人类被试或行为对照。","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:58:51","error":null,"has_summary":false,"summary":null},{"id":"2603.03585","version":2,"title":"Belief-Sim: Towards Belief-Driven Simulation of Demographic Misinformation Susceptibility","zh_title":"Belief-Sim：面向信念驱动的人口统计错误信息易感性仿真","abstract":"Misinformation is a growing societal threat, and susceptibility to misinformative claims varies across demographic groups due to differences in underlying beliefs. As Large Language Models (LLMs) are increasingly used to simulate human behaviors, we investigate whether they can simulate demographic misinformation susceptibility, treating beliefs as a primary driving factor. We introduce BeliefSim, a simulation framework that constructs demographic belief profiles using psychology-informed misinformation taxonomies and survey priors. We study prompt-based conditioning and post-training adaptation, and conduct a multi-fold evaluation using: (i) susceptibility alignment and (ii) counterfactual demographic sensitivity. Across both datasets and modeling strategies, we show that beliefs provide a strong prior for simulating misinformation susceptibility, with alignment up to 92%.","authors":["Angana Borah","Zohaib Khan","Rada Mihalcea","Verónica Pérez-Rosas"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-03-03","first_seen":"2026-03-03","revised_at":null,"abs_url":"https://arxiv.org/abs/2603.03585","pdf_url":"https://arxiv.org/pdf/2603.03585","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B3"],"tags":["LLM人类仿真","错误信息易感性","人口统计差异"],"reason":"用LLM仿真不同人口群体对错误信息的易感性，以信念为驱动，并与真实人类数据对照…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:09","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":21,"question":"信念是否能改善基于人口统计学的错误信息易感性仿真？","design":"使用多个LLM，通过提示条件化（BeliefSim-PC）和微调（BeliefSim-FT）两种方式，基于心理学启发的信念分类和调查先验构建人口信念画像，模拟不同性别、年龄、居住地、教育水平群体的错误信息易感性，测量其与真实人类判断的对齐程度。","baseline":"PANDORA数据集（318人，每人3条声明）和MIST-1数据集（409人，每人100条声明），共13.8K条真实人类判断，包含人口统计信息。","findings":"信念是模拟错误信息易感性的强先验，对齐度最高达92%；仅使用人口统计信息不可靠，可能导致捷径依赖。","reliability":"人口统计建模主要为单轴，仅评估8个独立群体，未考虑交叉群体效应；数据集仅限于美国参与者和英文标题。","relevance":"该研究直接用LLM仿真人口群体的错误信息易感性，以信念为驱动，并与真实人类数据对照，包含批判性分析，高度契合研究者对LLM人类仿真可靠性及失效条件的关注，值得精读原文。","inspiration":"该方法通过心理学信念分类构建人口信念画像，并对比两种LLM条件化方式（提示与微调）来仿真群体判断，提供了处理施加与对照设计的参考｜可迁移至信贷审批中的群体歧视研究，仿真不同人口群体对贷款申请的审批决策｜以LLM作为被试，处理为基于信念画像的条件化提示或微调，结果变量为审批通过率，对照真实银行信贷审批数据中的群体差异"}},{"id":"2603.02711","version":1,"title":"A Natural Language Agentic Approach to Study Affective Polarization","zh_title":"一种研究情感极化的自然语言智能体方法","abstract":"Affective polarization has been central to political and social studies, with growing focus on social media, where partisan divisions are often exacerbated. Real-world studies tend to have limited scope, while simulated studies suffer from insufficient high-quality training data, as manually labeling posts is labor-intensive and prone to subjective biases. The lack of adequate tools to formalize different definitions of affective polarization across studies complicates result comparison and hinders interoperable frameworks. We present a multi-agent model providing a comprehensive approach to studying affective polarization in social media. To operationalize our framework, we develop a platform leveraging large language models (LLMs) to construct virtual communities where agents engage in discussions. We showcase the potential of our platform by (1) analyzing questions related to affective polarization, as explored in social science literature, providing a fresh perspective on this phenomenon, and (2) introducing scenarios that allow observation and measurement of polarization at different levels of granularity and abstraction. Experiments show that our platform is a flexible tool for computational studies of complex social dynamics such as affective polarization. It leverages advanced agent models to simulate rich, context-sensitive interactions and systematically explore research questions traditionally addressed through human-subject studies.","authors":["Stephanie Anneris Malvicini","Ewelina Gajewska","Arda Derbent","Katarzyna Budzynska","Jarosław A. Chudziak","Maria Vanina Martinez"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-03-03","first_seen":"2026-03-03","revised_at":null,"abs_url":"https://arxiv.org/abs/2603.02711","pdf_url":"https://arxiv.org/pdf/2603.02711","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["情感极化","多智能体模拟","社交媒体"],"reason":"用LLM agent模拟社交媒体讨论以研究情感极化，但无真实人类数据对照，属社…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:08","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":146,"question":"如何利用LLM多智能体平台模拟社交媒体讨论，以研究情感极化的动态与测量？","design":"使用LLM驱动的多智能体模型构建虚拟社区，智能体代表不同政治党派，在受控场景下进行讨论；通过操作化党派认同强度并测量对内外群体的情感，观察和量化情感极化。","baseline":"无对照","findings":"该平台能够复现社会科学文献中关于情感极化的标准问题，并支持从多个粒度和抽象层次观察与测量极化；平台可作为传统人类被试研究的灵活计算补充工具。","reliability":"论文指出当前结果仅为初步实证验证，LLM作为社会智能体的适用性仍需通过系统与人类被试数据比较来评估；LLM自我报告变量的可靠性需要仔细检验；平台目前集中于美国中心场景，且缺乏结构化认知架构、学习机制和网络感知交互协议。","relevance":"该研究用LLM智能体仿真人类社交互动以研究情感极化，但无真实人类数据对照，属于探索性仿真工具开发；适合关注方法创新但需注意其缺乏基准验证的局限，可酌情阅读以了解LLM在政治心理学仿真中的应用潜力。","inspiration":"该方法通过LLM多智能体平台模拟党派讨论并测量情感极化，其操作化党派认同强度、控制讨论场景的设计可借鉴用于经济金融实验中处理组与对照组的构建。｜可迁移至研究经济政策分歧中的情感极化，例如不同经济阶层或利益群体对税收政策的态度分化与互动。｜以LLM智能体代表不同收入群体，施加累进税制改革信息作为处理，测量群体间情感温度与政策支持度，并与真实调查数据（如ANES或GSS中相关条目）进行对照验证。"}},{"id":"2603.03242","version":1,"title":"Density-Guided Response Optimization: Community-Grounded Alignment via Implicit Acceptance Signals","zh_title":"密度引导的响应优化：基于隐式接受信号的社区对齐","abstract":"Language models deployed in online communities must adapt to norms that vary across social, cultural, and domain-specific contexts. Prior alignment approaches rely on explicit preference supervision or predefined principles, which are effective for well-resourced settings but exclude most online communities -- particularly those without institutional backing, annotation infrastructure, or organized around sensitive topics -- where preference elicitation is costly, ethically fraught, or culturally misaligned. We observe that communities already express preferences implicitly through what content they accept, engage with, and allow to persist. We show that this acceptance behavior induces measurable geometric structure in representation space: accepted responses occupy coherent, high-density regions that reflect community-specific norms, while rejected content falls in sparser or misaligned areas. We operationalize this structure as an implicit preference signal for alignment and introduce density-guided response optimization (DGRO), a method that aligns language models to community norms without requiring explicit preference labels. Using labeled preference data, we demonstrate that local density recovers pairwise community judgments, indicating that geometric structure encodes meaningful preference signal. We then apply DGRO in annotation-scarce settings across diverse communities spanning platform, topic, and language. DGRO-aligned models consistently produce responses preferred by human annotators, domain experts, and model-based judges over supervised and prompt-based baselines. We position DGRO as a practical alignment alternative for communities where explicit preference supervision is unavailable or misaligned with situated practices, and discuss the implications and risks of learning from emergent acceptance behavior.","authors":["Patrick Gerard","Svitlana Volkova"],"categories":["cs.AI","cs.CL"],"primary_category":"cs.AI","announce_type":"new","date":"2026-03-03","first_seen":"2026-03-03","revised_at":null,"abs_url":"https://arxiv.org/abs/2603.03242","pdf_url":"https://arxiv.org/pdf/2603.03242","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["语言模型对齐","社区规范","隐式偏好信号"],"reason":"纯多智能体对齐社区规范，无人类行为仿真或对照实验","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:59:40","error":null,"has_summary":false,"summary":null},{"id":"2603.02076","version":1,"title":"When an AI Judges Your Work: The Hidden Costs of Algorithmic Assessment","zh_title":"当AI评判你的工作：算法评估的隐性代价","abstract":"We use an online experiment with a real work task to study whether workers change their behavior when they know AI will be used to judge their work instead of humans. We find that individuals produce a higher quantity of output when they are assigned an AI evaluator. However, controlling for quantity, the quality of their output is lower, regardless of whether quality is measured using humans or LLM grades. We also find that workers are more likely to use external tools, including LLMs, when they know AI is used to judge their work instead of humans. However, the increase in external tool use does not appear to explain the differences in quantity or quality across treatments.","authors":["David Almog","Lucas Lippman","Daniel Martin"],"categories":["econ.GN","cs.HC"],"primary_category":"econ.GN","announce_type":"new","date":"2026-03-02","first_seen":"2026-03-02","revised_at":null,"abs_url":"https://arxiv.org/abs/2603.02076","pdf_url":"https://arxiv.org/pdf/2603.02076","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["算法评估","人类行为","LLM作为评估者"],"reason":"LLM替代人类评估者，非替代被试，但涉及人类行为变化与LLM评估对照，属边界情…","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:58:35","error":null,"has_summary":false,"summary":null},{"id":"2603.21006","version":1,"title":"How AI Systems Think About Education: Analyzing Latent Preference Patterns in Large Language Models","zh_title":"AI系统如何思考教育：分析大语言模型中的潜在偏好模式","abstract":"This paper presents the first systematic measurement of educational alignment in Large Language Models. Using a Delphi-validated instrument comprising 48 items across eight educational-theoretical dimensions, the study reveals that GPT-5.1 exhibits highly coherent preference patterns (99.78% transitivity; 92.79% model accuracy) that largely align with humanistic educational principles where expert consensus exists. Crucially, divergences from expert opinion occur precisely in domains of normative disagreement among human experts themselves, particularly emotional dimensions and epistemic normativity. This raises a fundamental question for alignment research: When human values are contested, what should models be aligned to? The findings demonstrate that GPT-5.1 does not remain neutral in contested domains but adopts coherent positions, prioritizing emotional responsiveness and rejecting false balance. The methodology, combining Delphi consensus-building with Structured Preference Elicitation and Thurstonian Utility modeling, provides a replicable framework for domain-specific alignment evaluation beyond generic value benchmarks.","authors":["Daniel Autenrieth"],"categories":["cs.CY","cs.AI","cs.CL","cs.HC"],"primary_category":"cs.CY","announce_type":"new","date":"2026-02-28","first_seen":"2026-02-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2603.21006","pdf_url":"https://arxiv.org/pdf/2603.21006","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["LLM偏好测量","教育对齐","德尔菲法"],"reason":"测量LLM自身的教育偏好，属于D2人格/态度测量，非仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:59:42","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":182,"question":"GPT-5.1在教育理论维度上表现出怎样的偏好模式，这些偏好与人类专家共识及分歧的关系如何？","design":"本研究并非仿真人类被试，而是通过德尔菲法构建48项教育原则，再设计144个教学场景，让GPT-5.1进行102,960次成对比较，测量其教育偏好的一致性和效用模式。","baseline":"无对照","findings":"GPT-5.1的教育偏好高度一致（传递性99.78%），在专家共识领域与人本主义教育原则对齐；但在情感支持和认识论规范性等专家存在分歧的领域，模型会采取明确立场而非保持中立。","reliability":"论文未讨论","relevance":"该研究将LLM作为测量对象而非人类替代品，不涉及人类仿真或行为复现，与您关注的用LLM仿真人类被试、对照真实人类数据的研究方向不匹配，不建议优先阅读原文。","inspiration":"该方法通过大规模成对比较测量AI系统的潜在偏好模式，可用于经济学中测量LLM对政策选项的偏好排序｜可迁移到政策评估场景，如测量LLM对税收政策、福利分配或环境规制的偏好结构｜以LLM为被试，设计不同政策特征的成对比较任务，结果变量为选择比例和偏好传递性，对照真实公众调查数据"}},{"id":"2602.21983","version":1,"title":"Humanizing Robot Gaze Shifts: A Framework for Natural Gaze Shifts in Humanoid Robots","zh_title":"拟人化机器人注视转移：人形机器人自然注视转移框架","abstract":"Leveraging auditory and visual feedback for attention reorientation is essential for natural gaze shifts in social interaction. However, enabling humanoid robots to perform natural and context-appropriate gaze shifts in unconstrained human--robot interaction (HRI) remains challenging, as it requires the coupling of cognitive attention mechanisms and biomimetic motion generation. In this work, we propose the Robot Gaze-Shift (RGS) framework, which integrates these two components into a unified pipeline. First, RGS employs a vision--language model (VLM)-based gaze reasoning pipeline to infer context-appropriate gaze targets from multimodal interaction cues, ensuring consistency with human gaze-orienting regularities. Second, RGS introduces a conditional Vector Quantized-Variational Autoencoder (VQ-VAE) model for eye--head coordinated gaze-shift motion generation, producing diverse and human-like gaze-shift behaviors. Experiments validate that RGS effectively replicates human-like target selection and generates realistic, diverse gaze-shift motions.","authors":["Jingchao Wei","Jingkai Qin","Yuxiao Cao","Jingcheng Huang","Xiangrui Zeng","Min Li","Zhouping Yin"],"categories":["cs.RO"],"primary_category":"cs.RO","announce_type":"new","date":"2026-02-25","first_seen":"2026-02-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2602.21983","pdf_url":"https://arxiv.org/pdf/2602.21983","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C2"],"tags":["机器人注视","人机交互","运动生成"],"reason":"研究机器人注视转移，属于机器人仿真环境，不涉及LLM仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:59:10","error":null,"has_summary":false,"summary":null},{"id":"2602.21091","version":1,"title":"Can Interest-Bearing Positions Solve the Long-Horizon Problem in Prediction Markets?","zh_title":"计息头寸能否解决预测市场中的长期问题？","abstract":"Prediction markets suffer from reduced liquidity and price accuracy for long-horizon events due to the opportunity cost of committed capital. Recently, major platforms have introduced interest-bearing positions to mitigate this \"long-horizon problem.\" I evaluate this policy using agent-based simulations with large language model (LLM) traders in a 2 x 2 factorial design, varying time horizon (4 days vs. 2 years) and the presence of interest. While long horizons degrade accuracy, the observed pricing bias (0.72 percentage points) is significantly smaller than theoretical and prior empirical estimates. Paying interest eliminates approximately 83% of the horizon effect on accuracy and more than triples market participation (from 17% to 62% of wealth). These findings suggest the long-horizon problem may be overstated in existing literature and that interest-bearing positions are a highly effective intervention, primarily by incentivizing participation rather than correcting bias.","authors":["Caleb Maresca"],"categories":["econ.GN"],"primary_category":"econ.GN","announce_type":"new","date":"2026-02-24","first_seen":"2026-02-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2602.21091","pdf_url":"https://arxiv.org/pdf/2602.21091","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["LLM代理模拟","预测市场","社会模拟"],"reason":"用LLM agent模拟预测市场，但无真实人类数据对照，属社会模拟边界情形。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:07","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":183,"question":"付息头寸能否解决预测市场中的长周期问题？","design":"使用大语言模型（LLM）智能体模拟交易者，采用2×2析因设计，操纵时间周期（4天 vs. 2年）和是否支付利息，测量价格准确性和市场参与度。","baseline":"无对照","findings":"长周期会降低价格准确性，但观察到的定价偏差（0.72个百分点）远小于理论和先前实证估计；支付利息消除了约83%的周期效应，并使市场参与度增加两倍以上。","reliability":"论文未讨论","relevance":"该研究用LLM智能体模拟预测市场交易，属于人类仿真实验，但缺乏真实人类数据对照，且场景为金融市场而非经济学实验或政策评估，与研究者关注的核心方向部分相关但非完全匹配，可酌情阅读原文了解LLM仿真方法。","inspiration":"该研究采用2×2析因设计操纵时间周期和利息支付，通过LLM智能体模拟交易者来测量价格准确性和参与度，这种多因素实验设计值得借鉴｜可迁移到资产定价实验，例如研究不同信息发布频率和持有成本对市场价格发现效率的影响｜设计一个实验，以LLM智能体为被试，操纵信息更新频率（高/低）和交易成本（有/无），结果变量为价格偏差和交易量，并对照真实股票市场数据中的类似情境"}},{"id":"2602.20440","version":2,"title":"Intelligence Without Integrity: Why Capable LLMs May Undermine Reliability","zh_title":"有智无信：为何能力强的LLM可能损害可靠性","abstract":"As LLMs become embedded in research workflows and organizational decision processes, their effect on analytical reliability remains uncertain. We distinguish two dimensions of analytical reliability -- intelligence (the capacity to reach correct conclusions) and integrity (the stability of conclusions when analytically irrelevant cues about desired outcomes are introduced) -- and ask whether frontier LLMs possess both. Whether these dimensions trade off is theoretically ambiguous: the sophistication enabling accurate analysis may also enable responsiveness to non-evidential cues, or alternatively, greater capability may confer protection through better calibration and discernment. Using synthetically generated data with embedded ground truth, we evaluate fourteen models on a task simulating empirical analysis of hospital merger effects. We find that intelligence and integrity trade off: frontier models most likely to reach correct conclusions under neutral conditions are often most susceptible to shifting conclusions under motivated framing. We extend work on sycophancy by introducing goal-conditioned analytical sycophancy: sensitivity of inference to cues about desired outcomes, even when no belief is asserted and evidence is held constant. Unlike simple prompt sensitivity, models shift conclusions away from objective evidence in response to analytically irrelevant framing. This finding has important implications for empirical research and organizations. Selecting tools based on capability benchmarks may inadvertently select against the stability needed for reliable and replicable analysis.","authors":["Ryan Allen","Aticus Peterson"],"categories":["econ.GN"],"primary_category":"econ.GN","announce_type":"new","date":"2026-02-24","first_seen":"2026-02-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2602.20440","pdf_url":"https://arxiv.org/pdf/2602.20440","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["LLM可靠性","分析完整性","模型行为测量"],"reason":"研究LLM在分析任务中的结论稳定性，属于测量模型本身而非仿真人类被试，无人类数…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:07","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":184,"question":"前沿大语言模型在分析可靠性上是否存在智力与诚信的权衡？","design":"本研究并非人类仿真实验，而是使用合成数据评估14个模型在模拟医院合并效应分析任务中的表现，通过中性条件和带有动机性框架的提示来测量模型得出正确结论的能力（智力）和结论稳定性（诚信）。","baseline":"无对照","findings":"前沿模型在中性条件下最可能得出正确结论，但在动机性框架下最易改变结论，智力与诚信存在权衡。模型会因与分析无关的期望线索而偏离客观证据，表现出目标条件分析性谄媚。","reliability":"论文未讨论","relevance":"该研究聚焦LLM自身的分析可靠性，未涉及人类被试仿真或与真实人类数据对照，与您关注的人类仿真实验方向不直接相关，但其中关于模型易受动机性框架影响的发现对评估LLM在实验中的偏差有参考价值。","inspiration":"可借鉴其通过动机性框架提示（如强调期望结论）来操纵LLM分析行为的处理设计，并测量结论正确性与稳定性以揭示智力-诚信权衡｜可迁移至政策评估场景，如研究LLM在模拟专家预测经济政策效果时是否因政治倾向或利益相关方期望而扭曲分析｜以LLM作为被试，随机分配中性提示与带有党派倾向的动机性框架提示，要求其分析某项税收改革对就业的影响，结果变量为预测方向与幅度，对照真实历史政策评估数据"}},{"id":"2602.19580","version":1,"title":"Leap+Verify: Regime-Adaptive Speculative Weight Prediction for Accelerating Neural Network Training","zh_title":"Leap+Verify：用于加速神经网络训练的自适应权重预测框架","abstract":"We introduce Leap+Verify, a framework that applies speculative execution -- predicting future model weights and validating predictions before acceptance -- to accelerate neural network training. Inspired by speculative decoding in language model inference and by the Automatically Scalable Computation (ASC) architecture for program execution, Leap+Verify decomposes training into three dynamically detected regimes (chaotic, transition, stable) using activation-space cosine similarity as a real-time Lyapunov proxy signal. Within each regime, analytic weight predictors (momentum, linear, quadratic extrapolation) attempt to forecast model parameters K training steps ahead; predictions are accepted only when validated against a held-out loss criterion. We evaluate Leap+Verify on GPT-2 124M and Qwen 2.5-1.5B trained on WikiText-103 across five random seeds, sweeping prediction depth K in {5, 10, 25, 50, 75, 100}. Momentum-based prediction (Adam moment extrapolation) fails catastrophically at both scales, with predicted losses exceeding actuals by 100-10,000x -- a universal norm explosion in optimizer-state extrapolation. Finite-difference predictors (linear, quadratic) succeed where momentum fails: at 124M, they achieve 24% strict acceptance at K=5 in stable regimes; at 1.5B, they achieve 37% strict acceptance in transition regimes. The scale-dependent finding is in regime distribution: GPT-2 124M spends 34% of training in stable regime, while Qwen 1.5B spends 64% in chaotic regime and reaches stable in only 0-2 of 40 checkpoints. Larger models are more predictable when predictable, but less often predictable -- the practical bottleneck shifts from predictor accuracy to regime availability. Cross-seed results are highly consistent (less than 1% validation loss variance), and the three-regime framework produces identical phase boundaries (plus or minus 50 steps) across seeds.","authors":["Jeremy McEntire"],"categories":["cs.LG","econ.GN"],"primary_category":"cs.LG","announce_type":"new","date":"2026-02-23","first_seen":"2026-02-23","revised_at":null,"abs_url":"https://arxiv.org/abs/2602.19580","pdf_url":"https://arxiv.org/pdf/2602.19580","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["神经网络训练","投机执行","权重预测"],"reason":"纯神经网络训练加速技术，不涉及人类行为仿真或对照。","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:58:37","error":null,"has_summary":false,"summary":null},{"id":"2603.00113","version":2,"title":"AI Agents Alone Are Not (Yet) Sufficient for Social Simulation","zh_title":"AI智能体单独尚不足以进行社会仿真","abstract":"Recent advances in large language models (LLMs) have spurred growing interest in using LLM-integrated agents for social simulation, often under the implicit assumption that realistic population dynamics will emerge once role-specified agents are placed in a networked multi-agent setting. This position paper argues that LLM-based agents alone are not (yet) sufficient for social simulation. We attribute this over-optimism to a systematic mismatch between what current agent pipelines are typically optimized and validated to produce and what simulation-as-science requires. Concretely, role-playing plausibility does not imply faithful human behavioral validity; collective outcomes are frequently mediated by agent-environment co-dynamics rather than agent-agent messaging alone; and results can be dominated by interaction protocols, scheduling, and initial information priors. To make these underlying mechanisms explicit and auditable, we propose a unified formulation of AI agent-based social simulation as an environment-involved Markov game with explicit exposure and scheduling mechanisms, from which we derive concrete actions for design, evaluation, and interpretation.","authors":["Yiming Li","Dacheng Tao"],"categories":["cs.MA","cs.AI","cs.CE","cs.CY","cs.SI"],"primary_category":"cs.MA","announce_type":"new","date":"2026-02-19","first_seen":"2026-02-19","revised_at":null,"abs_url":"https://arxiv.org/abs/2603.00113","pdf_url":"https://arxiv.org/pdf/2603.00113","source_feed":"backfill","score":7,"bucket":"pending","rubric_hits":["A3","D3"],"tags":["社会仿真","LLM智能体","方法论批评"],"reason":"讨论LLM智能体社会仿真，指出当前不足并提出方法论，虽无人类数据对照，但直接相…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:07","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":82,"question":"当前基于LLM的智能体社会仿真在科学推断上存在哪些根本性不足？","design":"本文为立场论文，未进行具体仿真实验；它系统批评了现有LLM智能体社会仿真中角色扮演逼真度被等同于人类行为有效性的做法，并提出应将仿真建模为包含环境、曝光和调度机制的马尔可夫博弈。","baseline":"无对照","findings":"LLM智能体单独不足以进行社会仿真，因为角色扮演的合理性不保证行为忠实性；集体结果常由智能体-环境共动态、交互协议、调度和信息先验主导，而非仅由智能体间对话决定。","reliability":"论文指出当前仿真过度依赖角色扮演逼真度，忽视环境、制度和信息不对称等机制，导致结果可能被实现细节主导，缺乏可审计的生成机制，易成为说服性叙事而非科学工具。","relevance":"该文直接批判LLM社会仿真的可靠性，指出缺乏人类行为基准和机制透明性会导致错误推断，与您关注的仿真失效条件和批判性研究高度契合，值得精读以理解当前范式的根本缺陷。","inspiration":"该文强调仿真需建模环境、曝光和调度机制，而非仅依赖角色扮演，这提醒我们在经济实验中应明确制度规则和信息结构的设计｜可迁移到政策公告预期形成场景，如央行沟通对市场预期的影响｜以LLM为被试，处理为不同透明度的政策公告，结果变量为预期通胀预测值，对照真实调查数据如密歇根消费者调查"}},{"id":"2602.15730","version":1,"title":"Causal Effect Estimation with Latent Textual Treatments","zh_title":"基于潜在文本处理的因果效应估计","abstract":"Understanding the causal effects of text on downstream outcomes is a central task in many applications. Estimating such effects requires researchers to run controlled experiments that systematically vary textual features. While large language models (LLMs) hold promise for generating text, producing and evaluating controlled variation requires more careful attention. In this paper, we present an end-to-end pipeline for the generation and causal estimation of latent textual interventions. Our work first performs hypothesis generation and steering via sparse autoencoders (SAEs), followed by robust causal estimation. Our pipeline addresses both computational and statistical challenges in text-as-treatment experiments. We demonstrate that naive estimation of causal effects suffers from significant bias as text inherently conflates treatment and covariate information. We describe the estimation bias induced in this setting and propose a solution based on covariate residualization. Our empirical results show that our pipeline effectively induces variation in target features and mitigates estimation error, providing a robust foundation for causal effect estimation in text-as-treatment settings.","authors":["Omri Feldman","Amar Venugopal","Jann Spiess","Amir Feder"],"categories":["cs.CL","econ.EM"],"primary_category":"cs.CL","announce_type":"new","date":"2026-02-17","first_seen":"2026-02-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2602.15730","pdf_url":"https://arxiv.org/pdf/2602.15730","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["因果推断","文本处理","LLM生成"],"reason":"研究文本作为处理变量的因果效应估计，不涉及用LLM仿真人类被试或行为对照。","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:58:52","error":null,"has_summary":false,"summary":null},{"id":"2602.15312","version":1,"title":"Extracting Consumer Insight from Text: A Large Language Model Approach to Emotion and Evaluation Measurement","zh_title":"从文本中提取消费者洞察：一种用于情感和评价测量的大语言模型方法","abstract":"Accurately measuring consumer emotions and evaluations from unstructured text remains a core challenge for marketing research and practice. This study introduces the Linguistic eXtractor (LX), a fine-tuned, large language model trained on consumer-authored text that also has been labeled with consumers' self-reported ratings of 16 consumption-related emotions and four evaluation constructs: trust, commitment, recommendation, and sentiment. LX consistently outperforms leading models, including GPT-4 Turbo, RoBERTa, and DeepSeek, achieving 81% macro-F1 accuracy on open-ended survey responses and greater than 95% accuracy on third-party-annotated Amazon and Yelp reviews. An application of LX to online retail data, using seemingly unrelated regression, affirms that review-expressed emotions predict product ratings, which in turn predict purchase behavior. Most emotional effects are mediated by product ratings, though some emotions, such as discontent and peacefulness, influence purchase directly, indicating that emotional tone provides meaningful signals beyond star ratings. To support its use, a no-code, cost-free, LX web application is available, enabling scalable analyses of consumer-authored text. In establishing a new methodological foundation for consumer perception measurement, this research demonstrates new methods for leveraging large language models to advance marketing research and practice, thereby achieving validated detection of marketing constructs from consumer data.","authors":["Stephan Ludwig","Peter J. Danaher","Xiaohao Yang","Yu-Ting Lin","Ehsan Abedin","Dhruv Grewal","Lan Du"],"categories":["cs.CL","econ.EM"],"primary_category":"cs.CL","announce_type":"new","date":"2026-02-17","first_seen":"2026-02-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2602.15312","pdf_url":"https://arxiv.org/pdf/2602.15312","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["情感分析","消费者洞察","NLP评测"],"reason":"纯NLP能力评测，用LLM从文本中提取情感和评价，不以人类行为仿真为参照系。","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:58:53","error":null,"has_summary":false,"summary":null},{"id":"2602.15173","version":2,"title":"Mind the (DH) Gap! A Contrast in Risky Choices Between Reasoning and Conversational LLMs","zh_title":"注意(DH)差距！推理型与对话型LLM在风险选择上的对比","abstract":"The use of large language models either as decision support systems, or in agentic workflows, is rapidly transforming the digital ecosystem. However, the understanding of LLM decision-making under uncertainty remains limited. We study LLM risky choices along two dimensions: (1) prospect representation (based on an explicit representation or outcome history) and (2) decision rationale (explanation). Our study, which involves 20 frontier and open LLMs, is complemented by a matched human subjects experiment, which provides one reference point, while an expected payoff maximizing rational agent model provides another. We find that LLMs cluster into two categories: reasoning models (RMs) and conversational models (CMs). RMs tend towards rational behavior, are insensitive to the order of prospects, gain/loss framing, and explanations, and behave similarly whether prospects are explicit or presented via a history of outcomes. CMs are significantly less rational, slightly more human-like, sensitive to prospect ordering, framing, and explanation, and exhibit a large description-history gap. Paired comparisons of open LLMs suggest that a key factor differentiating RMs and CMs is training for mathematical reasoning.","authors":["Luise Ge","Yongyan Zhang","Yevgeniy Vorobeychik"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-02-16","first_seen":"2026-02-16","revised_at":null,"abs_url":"https://arxiv.org/abs/2602.15173","pdf_url":"https://arxiv.org/pdf/2602.15173","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2"],"tags":["LLM仿真","风险决策","人类对照实验"],"reason":"用LLM仿真人类风险决策，并与真人实验对照，评估模型行为偏差与人类相似度。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:05","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":40,"question":"LLM在风险决策中如何受前景表征方式（描述 vs. 历史结果）和决策理由（解释）的影响，并与人类及理性基准对比？","design":"用20个前沿和开源LLM作为被试，在三种基础前景对上完成风险选择任务，操纵前景表征（显式描述 vs. 结果历史）和解释要求（无解释/简短解释/数学解释），测量选择行为，并与匹配的人类被试实验和期望收益最大化理性智能体对比。","baseline":"匹配的人类被试实验，以及期望收益最大化的理性智能体模型。","findings":"LLM分为推理模型和对话模型两类：推理模型接近理性人，对前景顺序、框架和解释不敏感，描述-历史差距小；对话模型理性较低，略似人类，对顺序、框架和解释敏感，描述-历史差距大。数学推理训练是区分两类模型的关键因素。","reliability":"论文未讨论","relevance":"该研究直接以LLM仿真人类风险决策，并与真人实验和理性基准对照，系统评估了模型行为偏差和人类相似度，完全契合研究者对LLM仿真可靠性及失效条件的关注，值得精读。","inspiration":"借鉴其系统操纵前景表征（描述 vs. 历史结果）和解释要求（无/简短/数学）来分离模型行为偏差的设计，并设置理性基准与人类被试双对照｜可迁移到金融风险偏好评估场景，如投资者在历史收益序列与文字描述下的风险选择差异｜以LLM为被试，处理为用历史收益率序列 vs. 文字描述呈现股票前景，结果变量为风险资产配置比例，对照真实投资者调查数据（如Survey of Consumer Finances）"}},{"id":"2602.14043","version":1,"title":"Beyond Static Snapshots: Dynamic Modeling and Forecasting of Group-Level Value Evolution with Large Language Models","zh_title":"超越静态快照：基于大语言模型的群体价值观动态建模与预测","abstract":"Social simulation is critical for mining complex social dynamics and supporting data-driven decision making. LLM-based methods have emerged as powerful tools for this task by leveraging human-like social questionnaire responses to model group behaviors. Existing LLM-based approaches predominantly focus on group-level values at discrete time points, treating them as static snapshots rather than dynamic processes. However, group-level values are not fixed but shaped by long-term social changes. Modeling their dynamics is thus crucial for accurate social evolution prediction--a key challenge in both data mining and social science. This problem remains underexplored due to limited longitudinal data, group heterogeneity, and intricate historical event impacts. To bridge this gap, we propose a novel framework for group-level dynamic social simulation by integrating historical value trajectories into LLM-based human response modeling. We select China and the U.S. as representative contexts, conducting stratified simulations across four core sociodemographic dimensions (gender, age, education, income). Using the World Values Survey, we construct a multi-wave, group-level longitudinal dataset to capture historical value evolution, and then propose the first event-based prediction method for this task, unifying social events, current value states, and group attributes into a single framework. Evaluations across five LLM families show substantial gains: a maximum 30.88\\% improvement on seen questions and 33.97\\% on unseen questions over the Vanilla baseline. We further find notable cross-group heterogeneity: U.S. groups are more volatile than Chinese groups, and younger groups in both countries are more sensitive to external changes. These findings advance LLM-based social simulation and provide new insights for social scientists to understand and predict social value changes.","authors":["Qiankun Pi","Guixin Su","Jinliang Li","Mayi Xu","Xin Miao","Jiawei Jiang","Ming Zhong","Tieyun Qian"],"categories":["cs.SI","cs.AI"],"primary_category":"cs.SI","announce_type":"new","date":"2026-02-15","first_seen":"2026-02-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2602.14043","pdf_url":"https://arxiv.org/pdf/2602.14043","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM人类仿真","价值观演化","社会模拟"],"reason":"用LLM仿真群体价值观动态，有真实世界价值观调查数据对照，涉及社会变迁预测。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:04","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":55,"question":"如何利用大语言模型动态建模和预测群体价值观的长期演变趋势？","design":"基于世界价值观调查（WVS）第5-7波数据，按性别、年龄、教育、收入四个维度将中美人群划分为群体，用LLM微调学习历史价值观轨迹以预测未来价值观，并提出事件感知预测方法，将社会事件与价值观表示对齐，让LLM推理事件影响。","baseline":"世界价值观调查（WVS）第5、6、7波的真实群体级纵向数据，包含中国28个群体和美国23个群体。","findings":"事件感知预测方法在已见问题上最高提升30.88%，在未见问题上最高提升33.97%；美国群体价值观波动性显著高于中国，两国年轻群体对外部变化更敏感。","reliability":"论文未讨论","relevance":"该研究用LLM仿真群体价值观动态演变，有真实WVS纵向数据作为基准，涉及中美群体异质性和事件驱动的价值观变化，属于经济学/政策评估场景下的LLM人类仿真，与研究者关注高度契合，值得精读原文。","inspiration":"该方法将社会事件编码为与价值观表示对齐的向量，让LLM推理事件对群体价值观的动态影响，这种事件感知的预测设计值得借鉴｜可迁移到政策公告对消费者信心或通胀预期的影响评估，例如研究央行沟通事件如何改变不同群体的预期形成过程｜用LLM模拟不同收入/年龄群体的消费者，施加央行声明作为处理，测量其通胀预期变化，以密歇根大学消费者调查的真实群体数据作为对照基准"}},{"id":"2602.13862","version":2,"title":"Measuring Self-Rating Bias in LLM-Generated Survey Data: A Semantic Similarity Framework for Independent Scale Mapping","zh_title":"测量LLM生成调查数据中的自评偏差：一种独立量表映射的语义相似度框架","abstract":"Synthetic survey data generated by large language models (LLMs) suffers from a fundamental circularity: the same model family that generates text responses also maps them to numerical scales. We calibrate and validate Semantic Similarity Rating (SSR; Maier et al., 2024), which decouples generation from scale mapping via embedding-based cosine similarity against predefined anchor statements. Configuration experiments (N=17 pilot, N=69 cross-validation across 8 domains) show that naturalistic behavioral anchors outperform formal jargon by 29 percentage points (pp), and that SSR achieves 65-67% exact match and 91% within plus/minus 1; a cross-model test with OpenAI text-embedding-3-small reaches 77% exact, confirming cross-provider generalization. Direct LLM baselines (Claude 87%, GPT-4o 83%) establish that SSR's contribution is methodological independence, not accuracy superiority. A control condition removing question text from the LLM prompt actually improves LLM accuracy, ruling out information asymmetry as the explanation for SSR's lower accuracy. A pre-registered circularity experiment (N=345) reveals 4x compressed error variance in LLM rating (sigma^2 = 0.21 vs 0.87 for SSR) and systematic directional bias. A cross-model control (GPT-4o rating Claude-generated text) shows nearly identical compression (within/cross ratio = 0.93), indicating variance compression is a general LLM property rather than a within-model artifact. The calibration dataset, anchor library, and source code are publicly available (see Data Availability).","authors":["Eduardo Vera Pichardo"],"categories":["physics.soc-ph"],"primary_category":"physics.soc-ph","announce_type":"new","date":"2026-02-14","first_seen":"2026-02-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2602.13862","pdf_url":"https://arxiv.org/pdf/2602.13862","source_feed":"backfill","score":8,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","调查数据","偏差评估"],"reason":"评估LLM生成调查数据的自评偏差，提出独立量表映射方法，有真实人类数据对照，批…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:02","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":33,"question":"如何通过语义相似度框架独立测量LLM生成调查数据中的自评偏差，并解决文本生成与量表映射的循环性问题？","design":"本研究并非仿真人类被试，而是校准和验证语义相似度评分（SSR）框架：先用LLM（Claude Haiku 4.5）根据人格描述和问题生成文本回答，再用独立的嵌入模型（Voyage AI voyage-3.5-lite）通过余弦相似度将文本映射到预定义的锚定语句量表上，并通过配置实验、交叉验证、LLM基线对比和预注册的循环性实验评估该框架的性能与偏差。","baseline":"无对照","findings":"SSR框架实现了生成与测量的架构独立，自然行为锚定比正式术语锚定准确率高29个百分点，交叉验证中精确匹配率达65-67%，±1内达91%；直接LLM评分虽准确率更高（Claude 87%，GPT-4o 83%），但存在4倍误差方差压缩和系统性方向偏差，且该方差压缩是LLM的普遍属性而非同模型伪影。","reliability":"论文承认SSR的准确率低于直接LLM评分，且嵌入模型与生成模型可能因共享预训练语料而存在残余相关性；锚定语句的领域特异性导致跨领域准确率差异大（33-90%），需进一步优化锚定库。","relevance":"该研究直接针对LLM仿真调查数据中的循环性测量偏差，提出独立量表映射方法，并系统揭示了LLM自评的方差压缩和方向偏差，对关注仿真可靠性与失效条件的研究者具有重要参考价值，值得阅读原文。","inspiration":"该方法通过独立嵌入模型将LLM文本回答映射到预定义量表，避免了直接让LLM自评的循环偏差，这种测量与生成解耦的设计值得借鉴｜可迁移到消费者信心调查或通胀预期测量中，用LLM模拟受访者对经济前景的开放式回答，再独立映射到信心指数或预期值｜用LLM根据人口特征生成对经济前景的文本描述，以独立语义模型映射为预期通胀值，与密歇根消费者调查的真实个体数据对比，检验仿真偏差与方差压缩"}},{"id":"2602.12490","version":1,"title":"Transformer-based CoVaR: Systemic Risk in Textual Information","zh_title":"基于Transformer的CoVaR：文本信息中的系统性风险","abstract":"Conditional Value-at-Risk (CoVaR) quantifies systemic financial risk by measuring the loss quantile of one asset, conditional on another asset experiencing distress. We develop a Transformer-based methodology that integrates financial news articles directly with market data to improve CoVaR estimates. Unlike approaches that use predefined sentiment scores, our method incorporates raw text embeddings generated by a large language model (LLM). We prove explicit error bounds for our Transformer CoVaR estimator, showing that accurate CoVaR learning is possible even with small datasets. Using U.S. market returns and Reuters news items from 2006--2013, our out-of-sample results show that textual information impacts the CoVaR forecasts. With better predictive performance, we identify a pronounced negative dip during market stress periods across several equity assets when comparing the Transformer-based CoVaR to both the CoVaR without text and the CoVaR using traditional sentiment measures. Our results show that textual data can be used to effectively model systemic risk without requiring prohibitively large data sets.","authors":["Junyu Chen","Tom Boot","Lingwei Kong","Weining Wang"],"categories":["econ.EM","q-fin.RM","stat.ML"],"primary_category":"econ.EM","announce_type":"new","date":"2026-02-13","first_seen":"2026-02-13","revised_at":null,"abs_url":"https://arxiv.org/abs/2602.12490","pdf_url":"https://arxiv.org/pdf/2602.12490","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["金融风险预测","文本嵌入","Transformer模型"],"reason":"用LLM提取文本特征改进金融风险预测，非人类仿真实验，无行为对照。","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:58:53","error":null,"has_summary":false,"summary":null},{"id":"2602.11939","version":1,"title":"Do Large Language Models Adapt to Language Variation across Socioeconomic Status?","zh_title":"大语言模型能否适应社会经济地位带来的语言变异？","abstract":"Humans adjust their linguistic style to the audience they are addressing. However, the extent to which LLMs adapt to different social contexts is largely unknown. As these models increasingly mediate human-to-human communication, their failure to adapt to diverse styles can perpetuate stereotypes and marginalize communities whose linguistic norms are less closely mirrored by the models, thereby reinforcing social stratification. We study the extent to which LLMs integrate into social media communication across different socioeconomic status (SES) communities. We collect a novel dataset from Reddit and YouTube, stratified by SES. We prompt four LLMs with incomplete text from that corpus and compare the LLM-generated completions to the originals along 94 sociolinguistic metrics, including syntactic, rhetorical, and lexical features. LLMs modulate their style with respect to SES to only a minor extent, often resulting in approximation or caricature, and tend to emulate the style of upper SES more effectively. Our findings (1) show how LLMs risk amplifying linguistic hierarchies and (2) call into question their validity for agent-based social simulation, survey experiments, and any research relying on language style as a social signal.","authors":["Elisa Bassignana","Mike Zhang","Dirk Hovy","Amanda Cercas Curry"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-02-12","first_seen":"2026-02-12","revised_at":null,"abs_url":"https://arxiv.org/abs/2602.11939","pdf_url":"https://arxiv.org/pdf/2602.11939","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","A5","B1","B4"],"tags":["LLM仿真","社会语言学","算法保真度"],"reason":"直接评估LLM仿真人类语言行为的效度，有真实人类数据对照，并指出仿真失效条件。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:02","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":56,"question":"大语言模型在多大程度上能根据社会经济地位（SES）社区的语言变异调整其生成文本的风格？","design":"本研究并非基于智能体的仿真实验，而是通过提示工程让四个大语言模型补全来自Reddit和YouTube的、按SES分层的人类文本片段，然后比较模型生成文本与原始文本在94项社会语言学指标上的差异。","baseline":"从Reddit和YouTube收集的真实社交媒体文本，按SES（通过主题关键词和网络分析等策略）分层，作为人类语言风格的对照基准。","findings":"大语言模型仅能微弱地根据SES调节风格，常表现为近似或夸张模仿，且更擅长模仿高SES社区的风格。这种风格适应不足可能放大语言层级差异，并质疑LLM在基于智能体的社会模拟、调查实验等依赖语言风格作为社会信号的研究中的有效性。","reliability":"论文指出LLM在风格适应上存在近似或夸张模仿而非精确复现的问题，且对低SES风格模拟较差；初步消融实验显示，当提供更长上下文时，模型更倾向于适应高SES风格，这揭示了仿真失效的条件。","relevance":"该研究直接评估了LLM模拟不同社会经济地位人群语言行为的效度，有真实人类数据对照，并明确指出仿真在风格适应上的失效条件，与研究者关注的人类仿真可靠性及批判性评估高度契合，值得精读原文。","inspiration":"该方法通过提示工程让LLM补全按SES分层的人类文本，并对比94项社会语言学指标，可借鉴其分层对照与多维度风格测量来评估LLM的仿真偏差。｜可迁移到信贷审批中的语言歧视研究，例如分析LLM模拟不同SES申请人的贷款申请文本时是否系统性地偏向高SES风格，从而影响审批决策。｜以LLM为被试，给定不同SES背景的贷款申请场景提示，生成申请文本；结果变量为文本的语言风格指标（如正式度、情感词频）；以真实银行或P2P平台中不同SES申请人的贷款申请文本作为对照基准。"}},{"id":"2602.09362","version":1,"title":"Behavioral Economics of AI: LLM Biases and Corrections","zh_title":"人工智能的行为经济学：大语言模型的偏差与校正","abstract":"Do generative AI models, particularly large language models (LLMs), exhibit systematic behavioral biases in economic and financial decisions? If so, how can these biases be mitigated? Drawing on the cognitive psychology and experimental economics literatures, we conduct the most comprehensive set of experiments to date$-$originally designed to document human biases$-$on prominent LLM families across model versions and scales. We document systematic patterns in LLM behavior. In preference-based tasks, responses become more human-like as models become more advanced or larger, while in belief-based tasks, advanced large-scale models frequently generate rational responses. Prompting LLMs to make rational decisions reduces biases.","authors":["Pietro Bini","Lin William Cong","Xing Huang","Lawrence J. Jin"],"categories":["econ.GN","cs.AI"],"primary_category":"econ.GN","announce_type":"new","date":"2026-02-10","first_seen":"2026-02-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2602.09362","pdf_url":"https://arxiv.org/pdf/2602.09362","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["LLM行为偏差","人类仿真","实验经济学"],"reason":"用人类实验范式测LLM行为偏差并对比人类数据，直接评估仿真可靠性。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:02","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":19,"question":"生成式AI模型（尤其是大语言模型）在经济金融决策中是否表现出系统性行为偏差？如何纠正这些偏差？","design":"本研究并非用LLM仿真人类被试，而是将LLM本身作为研究对象。从认知心理学和实验经济学文献中选取原本用于测量人类偏差的实验问题，改编为提示词，通过API收集OpenAI ChatGPT、Anthropic Claude、Google Gemini和Meta Llama四个模型家族在不同版本和规模下的回答，分析其行为模式。","baseline":"以原有人类实验中的理性基准和真实人类回答作为对照。","findings":"在偏好类任务中，模型越先进或规模越大，回答越像人类且偏离理性；在信念类任务中，先进大规模模型则多给出理性回答。不同模型家族间存在显著异质性，如Gemini在偏好问题上比ChatGPT更不理性、更像人，而Llama在信念问题上理性程度较低。","reliability":"论文未讨论","relevance":"该研究直接使用人类行为实验范式测量LLM的偏差，并系统对比了LLM与人类及理性基准的差异，为评估LLM作为人类仿真被试的可靠性提供了关键证据，高度相关，值得精读。","inspiration":"借鉴其将经典行为经济学实验范式直接移植到LLM测试中的方法，可系统评估模型在不同决策场景下的行为一致性。｜可迁移到资产定价实验中的投资者偏差测量，如过度外推、过度自信等。｜以GPT-4等LLM为被试，呈现历史股价序列并要求预测未来收益，测量其外推倾向，并与真实投资者调查数据（如Shiller投资者信心调查）进行对照。"}},{"id":"2602.09802","version":2,"title":"Would a Large Language Model Pay Extra for a View? Inferring Willingness to Pay from Subjective Choices","zh_title":"大语言模型会为景观多付钱吗？从主观选择推断支付意愿","abstract":"As Large Language Models (LLMs) are increasingly deployed in applications such as travel assistance and purchasing support, they are often required to make subjective choices on behalf of users in settings where no objectively correct answer exists. We study LLM decision-making in a travel-assistant context by presenting models with choice dilemmas and analyzing their responses using multinomial logit models to derive implied willingness to pay (WTP) estimates. These WTP values are subsequently compared to human benchmark values from the economics literature. In addition to a baseline setting, we examine how model behavior changes under more realistic conditions, including the provision of information about users' past choices and persona-based prompting. Our results show that while meaningful WTP values can be derived for larger LLMs, they also display systematic deviations at the attribute level. Additionally, they tend to overestimate human WTP overall, particularly when expensive options or business-oriented personas are introduced. Conditioning models on prior preferences for cheaper options yields valuations that are closer to human benchmarks. Overall, our findings highlight both the potential and the limitations of using LLMs for subjective decision support and underscore the importance of careful model selection, prompt design, and user representation when deploying such systems in practice.","authors":["Manon Reusens","Sofie Goethals","Toon Calders","David Martens"],"categories":["cs.AI","cs.CL"],"primary_category":"cs.AI","announce_type":"new","date":"2026-02-10","first_seen":"2026-02-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2602.09802","pdf_url":"https://arxiv.org/pdf/2602.09802","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2"],"tags":["LLM仿真","支付意愿","人类数据对照"],"reason":"用LLM模拟人类支付意愿并与真实人类数据对照，评估偏差，涉及经济学场景。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:02","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":6,"question":"在旅行助手场景中，LLM的主观选择能否通过离散选择模型推导出可解释的支付意愿（WTP），并与人类基准对比？","design":"以酒店房间选择为任务，构建多属性选择困境，让多个LLM在不同提示条件下（基线、提供用户历史选择、基于人设的提示）做出选择，然后用多项Logit模型估计隐含的WTP，并分析提示改写、顺序调换、货币变化、温度等稳健性。","baseline":"对照Masiero et al. (2015)中人类对酒店房间属性的支付意愿估计值。","findings":"较大LLM能推导出有意义的WTP，但存在属性层面的系统性偏差，且整体高估人类WTP，尤其在引入昂贵选项或商务人设时；提供廉价偏好历史或学生人设可使估值更接近人类基准。","reliability":"论文承认WTP推导在部分模型上会失效，且结果受提示设计、用户表征方式影响较大，未涵盖更广泛的人群异质性，泛化到其他领域需进一步验证。","relevance":"直接对比LLM与人类真实WTP数据，系统评估了提示策略对仿真偏差的影响，并指出失效条件，高度契合研究者对经济学场景下LLM仿真可靠性及批判性分析的兴趣，值得精读原文。","inspiration":"借鉴其通过离散选择模型从LLM主观选择中推导支付意愿（WTP）并与人类基准对比的方法，可系统评估提示策略（如提供用户历史、人设提示）对仿真偏差的影响｜可迁移到消费者对金融产品属性的偏好评估，如贷款条款（利率、期限、抵押要求）或投资产品特征（风险、流动性、费用）的WTP估计｜以LLM为被试，呈现不同贷款产品选择集，处理为提供不同风险偏好或财务约束的人设提示，结果变量为通过多项Logit模型估计的各属性WTP，对照真实消费者信贷选择数据（如Survey of Consumer Finances）验证偏差"}},{"id":"2602.07414","version":1,"title":"Can LLMs Truly Embody Human Personality? Analyzing AI and Human Behavior Alignment in Dispute Resolution","zh_title":"LLM能真正体现人类人格吗？分析争议解决中AI与人类行为的一致性","abstract":"Large language models (LLMs) are increasingly used to simulate human behavior in social settings such as legal mediation, negotiation, and dispute resolution. However, it remains unclear whether these simulations reproduce the personality-behavior patterns observed in humans. Human personality, for instance, shapes how individuals navigate social interactions, including strategic choices and behaviors in emotionally charged interactions. This raises the question: Can LLMs, when prompted with personality traits, reproduce personality-driven differences in human conflict behavior? To explore this, we introduce an evaluation framework that enables direct comparison of human-human and LLM-LLM behaviors in dispute resolution dialogues with respect to Big Five Inventory (BFI) personality traits. This framework provides a set of interpretable metrics related to strategic behavior and conflict outcomes. We additionally contribute a novel dataset creation methodology for LLM dispute resolution dialogues with matched scenarios and personality traits with respect to human conversations. Finally, we demonstrate the use of our evaluation framework with three contemporary closed-source LLMs and show significant divergences in how personality manifests in conflict across different LLMs compared to human data, challenging the assumption that personality-prompted agents can serve as reliable behavioral proxies in socially impactful applications. Our work highlights the need for psychological grounding and validation in AI simulations before real-world use.","authors":["Deuksin Kwon","Kaleen Shrestha","Bin Han","Spencer Lin","James Hale","Jonathan Gratch","Maja Matarić","Gale M. Lucas"],"categories":["cs.AI","cs.CL"],"primary_category":"cs.AI","announce_type":"new","date":"2026-02-07","first_seen":"2026-02-07","revised_at":null,"abs_url":"https://arxiv.org/abs/2602.07414","pdf_url":"https://arxiv.org/pdf/2602.07414","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","A4","B1","B4"],"tags":["LLM人格仿真","人类行为对齐","争议解决"],"reason":"直接比较LLM与人类在冲突对话中的人格-行为对齐，有真实人类数据对照，并指出仿…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:00","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":22,"question":"当用大五人格特质提示LLM时，它们能否在冲突解决对话中复现人类因人格差异而导致的行为差异？","design":"基于KODIS人类冲突对话数据集，为LLM匹配相同场景和人格特质（大五人格），生成LLM-LLM冲突对话；比较人类与LLM在最终结果（得分、是否接受、是否离开）和策略行为（基于利益-权利-权力框架）上的差异。","baseline":"KODIS数据集中248段具有完整人格信息的人类-人类冲突解决对话。","findings":"人类中神经质是策略结果的最强预测因子，而LLM中外向性和宜人性的效应更强，且策略行为负载于更广泛的人格因素；Claude和Gemini比GPT-4o mini更接近人类策略指标，但整体仍存在显著偏差。","reliability":"论文指出人格提示的LLM在情感冲突场景中的行为保真度未经严格验证，不同LLM间人格表现差异显著，不能可靠地作为人类行为代理，强调在真实应用前需进行心理验证。","relevance":"该研究直接比较LLM与人类在冲突对话中的人格-行为对齐，有真实人类数据对照，并指出仿真失效条件，高度契合您对LLM人类仿真可靠性及批判性研究的关注，值得精读原文。","inspiration":"该方法通过给LLM施加人格特质提示来模拟人类行为差异，并设置真实人类对话数据作为对照基准，可借鉴其处理-对照设计及基于框架的策略行为编码。｜可迁移至消费者跨期选择实验，探究不同人格特质（如尽责性、神经质）对时间贴现行为的影响。｜以LLM为被试，施加大五人格提示，测量其在跨期选择任务中的贴现率，并与真实人类实验数据（如Andersen et al.的贴现率估计）进行对照。"}},{"id":"2602.18462","version":1,"title":"Assessing the Reliability of Persona-Conditioned LLMs as Synthetic Survey Respondents","zh_title":"评估基于人格条件的LLM作为合成调查受访者的可靠性","abstract":"Using persona-conditioned LLMs as synthetic survey respondents has become a common practice in computational social science and agent-based simulations. Yet, it remains unclear whether multi-attribute persona prompting improves LLM reliability or instead introduces distortions. Here we contribute to this assessment by leveraging a large dataset of U.S. microdata from the World Values Survey. Concretely, we evaluate two open-weight chat models and a random-guesser baseline across more than 70K respondent-item instances. We find that persona prompting does not yield a clear aggregate improvement in survey alignment and, in many cases, significantly degrades performance. Persona effects are highly heterogeneous as most items exhibit minimal change, while a small subset of questions and underrepresented subgroups experience disproportionate distortions. Our findings highlight a key adverse impact of current persona-based simulation practices: demographic conditioning can redistribute error in ways that undermine subgroup fidelity and risk misleading downstream analyses.","authors":["Erika Elizabeth Taday Morocho","Lorenzo Cima","Tiziano Fagni","Marco Avvenuti","Stefano Cresci"],"categories":["cs.CY","cs.AI"],"primary_category":"cs.CY","announce_type":"new","date":"2026-02-06","first_seen":"2026-02-06","revised_at":null,"abs_url":"https://arxiv.org/abs/2602.18462","pdf_url":"https://arxiv.org/pdf/2602.18462","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","调查方法","可靠性评估"],"reason":"直接评估LLM作为合成调查受访者的可靠性，使用真实人类数据对照，并指出仿真失效…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:05","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":5,"question":"多属性人格提示（persona prompting）能否提高大语言模型作为合成调查受访者的可靠性，还是会引入扭曲？","design":"使用两个开源聊天模型（Llama-2-13B 和 Qwen3-4B）模拟美国受访者，基于世界价值观调查（WVS-7）的个体记录构建多属性人格提示，比较有人格提示、无提示（vanilla）和随机猜测基线在70K+受访者-题目实例上的回答一致性。","baseline":"世界价值观调查第7波（WVS-7）的美国受访者微观数据，作为真实人类回答的基准。","findings":"人格提示在总体上并未带来一致的对齐改善，在许多情况下反而显著降低性能；人格效应高度异质，大多数题目变化极小，但少数题目和代表性不足的子群体出现不成比例的扭曲。","reliability":"论文指出人格提示可能重新分配误差，损害子群体保真度，并误导下游分析；效果因题目和属性而异，在少数群体中可能集中出现错误，且多属性约束可能相互干扰。","relevance":"该研究直接评估LLM作为合成调查受访者的可靠性，使用真实人类数据对照，并批判性地揭示了人格提示在子群体层面的失效风险，高度契合研究者对仿真可靠性、偏差及失效条件的关注，值得精读原文。","inspiration":"借鉴多属性人格提示与无提示基线的对照设计，以及按题目和子群体分解异质性效应的分析方法，可系统评估LLM仿真中的偏差来源｜可迁移到消费者金融决策调查场景，如风险偏好、储蓄选择、信贷需求等问卷的行为一致性研究｜以LLM作为合成受访者，施加多属性人格提示（收入、教育、财务素养等），测量其与真实消费者金融调查（如SCF）中个体回答的匹配度，并按收入分位数和金融素养水平检验子群体偏差"}},{"id":"2602.18464","version":2,"title":"How Well Can LLM Agents Simulate End-User Security and Privacy Attitudes and Behaviors?","zh_title":"LLM代理模拟终端用户安全与隐私态度及行为的效果如何？","abstract":"A growing body of research assumes that large language model (LLM) agents can serve as proxies for how people form attitudes toward and behave in response to security and privacy (S&P) threats. If correct, these simulations could offer a scalable way to forecast S&P risks in products prior to deployment. We interrogate this assumption using SP-ABCBench, a new benchmark of 30 tests derived from validated S&P human-subject studies, which measures alignment between simulations and human-subjects studies on a 0-100 ascending scale, where higher scores indicate better alignment across three dimensions: Attitude, Behavior, and Coherence. Evaluating twelve LLMs, four persona construction strategies, and two prompting methods, we found that there remains substantial room for improvement: all models score between 50 and 64 on average. Newer, bigger, and smarter models do not reliably do better and sometimes do worse. Some simulation configurations, however, do yield high alignment: e.g., with scores above 95 for some behavior tests when agents are prompted to apply bounded rationality and weigh privacy costs against perceived benefits. We release SP-ABCBench to enable reproducible evaluation as methods improve.","authors":["Yuxuan Li","Leyang Li","Hao-Ping Lee","Sauvik Das"],"categories":["cs.CY","cs.AI","cs.CL","cs.CR"],"primary_category":"cs.CY","announce_type":"new","date":"2026-02-06","first_seen":"2026-02-06","revised_at":null,"abs_url":"https://arxiv.org/abs/2602.18464","pdf_url":"https://arxiv.org/pdf/2602.18464","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","人类行为对照","安全隐私"],"reason":"直接评估LLM代理模拟人类安全隐私态度行为，有真实人类数据基准，并指出仿真失效…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:05","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":41,"question":"当前LLM代理在多大程度上能复现人群在安全与隐私（S&P）方面的态度、行为及其一致性？","design":"使用12个LLM，结合4种角色构建策略和2种提示方法，基于15项真实人类被试研究构建的30个测试基准SP-ABCBench，测量仿真结果与人类数据在态度、行为和一致性三个维度的对齐分数（0-100）。","baseline":"对照来自15项经过验证的S&P人类被试研究，涵盖态度量表、行为实验和构念间关系等30个可量化的人群层面效应。","findings":"所有模型平均对齐分数仅50-64，更大、更新、更强的模型未必更好，有时更差；但特定配置（如结合有限理性与隐私计算提示）在部分行为测试上可达95分以上。","reliability":"论文指出当前LLM仿真在S&P领域整体对齐度中等，模型规模与能力不保证提升，且角色构建与提示策略效果因测试维度而异，提示仿真在S&P决策中可能失效的条件。","relevance":"该研究直接评估LLM替代人类被试进行S&P态度行为仿真的可靠性，有真实人类基准，并揭示了仿真失效的具体条件，高度契合您对LLM人类仿真实验批判性评估的关注，值得精读。","inspiration":"该方法通过构建多维度测试基准（态度、行为、一致性）并计算对齐分数来量化LLM仿真与人类数据的差距，可借鉴其系统性评估框架和对照设计。｜可迁移到消费者金融决策研究，如评估LLM能否复现真实人群在信贷选择、风险偏好或退休储蓄行为中的偏差与异质性。｜以LLM代理为被试，施加不同金融素养或信息框架处理，测量其信贷违约概率或投资组合选择，并与美国消费者金融调查（SCF）或实验室实验的真实行为数据做对齐比较。"}},{"id":"2602.07238","version":2,"title":"Is there \"Secret Sauce'' in Large Language Model Development?","zh_title":"大语言模型开发中是否存在“秘方”？","abstract":"Do leading LLM developers possess a proprietary ``secret sauce'', or is LLM performance driven by scaling up compute? Using training and benchmark data for 809 models released between 2022 and 2025, we estimate scaling-law regressions with release-date and developer fixed effects. We find clear evidence of developer-specific efficiency advantages, but their importance depends on where models lie in the performance distribution. At the frontier, 80-90% of performance differences are explained by higher training compute, implying that scale--not proprietary technology--drives frontier advances. Away from the frontier, however, proprietary techniques and shared algorithmic progress substantially reduce the compute required to reach fixed capability thresholds. Some companies can systematically produce smaller models more efficiently. Strikingly, we also find substantial variation of model efficiency within companies; a firm can train two models with more than 40x compute efficiency difference. We also discuss the implications for AI leadership and capability diffusion.","authors":["Matthias Mertens","Natalia Fischl-Lanzoni","Neil Thompson"],"categories":["cs.AI","cs.LG","econ.GN"],"primary_category":"cs.AI","announce_type":"new","date":"2026-02-06","first_seen":"2026-02-06","revised_at":null,"abs_url":"https://arxiv.org/abs/2602.07238","pdf_url":"https://arxiv.org/pdf/2602.07238","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["LLM性能分析","规模定律","算力与效率"],"reason":"论文分析LLM性能与算力的关系，不涉及用LLM仿真人类被试或行为对照。","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:58:35","error":null,"has_summary":false,"summary":null},{"id":"2602.13273","version":1,"title":"MergePipe: A Budget-Aware Parameter Management System for Scalable LLM Merging","zh_title":"MergePipe：面向可扩展大语言模型合并的预算感知参数管理系统","abstract":"Large language model (LLM) merging has become a key technique in modern LLM development pipelines, enabling the integration of multiple task- or domain-specific expert models without retraining. However, as the number of experts grows, existing merging implementations treat model parameters as unstructured files and execute merges in a stateless, one-shot manner, leading to excessive disk I/O, redundant parameter scans, and poor scalability. In this paper, we present \\textbf{MergePipe}, a parameter management system for scalable LLM merging. MergePipe is the first system that treats LLM merging as a data management and execution problem, and introduces a catalog-driven abstraction over model parameters, merge plans, and execution lineage. At its core, MergePipe employs a cost-aware planner that explicitly models expert parameter I/O and enforces user-specified I/O budgets, followed by a streaming execution engine that materializes merged models under transactional guarantees. Our key insight is that while base model reads and output writes are unavoidable, expert parameter reads dominate merge cost and constitute the primary optimization target. By making expert access budget-aware throughout planning and execution, MergePipe mitigates the $O(K)$ I/O growth of naive pipelines and achieves predictable scaling behavior. Experiments show that MergePipe reduces total I/O by up to an order of magnitude and delivers up to $11\\times$ end-to-end speedups (up to 90\\% wall-time reduction) over state-of-the-art LLM merging pipelines.","authors":["Yuanyi Wang","Yanggan Gu","Zihao Wang","Kunxi Li","Yifan Yang","Zhaoyi Yan","Congkai Xie","Jianmin Wu","Hongxia Yang"],"categories":["cs.DB","cs.AI","cs.DC"],"primary_category":"cs.DB","announce_type":"new","date":"2026-02-05","first_seen":"2026-02-05","revised_at":null,"abs_url":"https://arxiv.org/abs/2602.13273","pdf_url":"https://arxiv.org/pdf/2602.13273","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["模型合并","参数管理","系统优化"],"reason":"纯多智能体系统研究，优化模型合并的I/O与执行效率，不涉及人类行为仿真或对照。","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:59:12","error":null,"has_summary":false,"summary":null},{"id":"2602.04674","version":2,"title":"Overstating Attitudes, Ignoring Networks: LLM Biases in Simulating Misinformation Susceptibility","zh_title":"夸大态度，忽视网络：LLM在模拟错误信息易感性中的偏差","abstract":"Large language models (LLMs) are increasingly used as proxies for human judgment in computational social science, yet their ability to reproduce patterns of susceptibility to misinformation remains unclear. We test whether LLM-simulated survey respondents, prompted with participant profiles drawn from social survey data measuring network, demographic, attitudinal and behavioral features, can reproduce human patterns of misinformation belief and sharing. Using three online surveys as baselines, we evaluate whether LLM outputs match observed response distributions and recover feature-outcome associations present in the original survey data. LLM-generated responses capture broad distributional tendencies and show modest correlation with human responses, but consistently overstate the association between belief and sharing. Linear models fit to simulated responses exhibit substantially higher explained variance and place disproportionate weight on attitudinal and behavioral features, while largely ignoring personal network characteristics, relative to models fit to human responses. Analyses of model-generated reasoning and LLM training data suggest that these distortions reflect systematic biases in how misinformation-related concepts are represented. Our findings suggest that LLM-based survey simulations are better suited for diagnosing systematic divergences from human judgment than for substituting it.","authors":["Eun Cheol Choi","Lindsay E. Young","Emilio Ferrara"],"categories":["cs.SI","cs.AI","cs.CL"],"primary_category":"cs.SI","announce_type":"new","date":"2026-02-04","first_seen":"2026-02-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2602.04674","pdf_url":"https://arxiv.org/pdf/2602.04674","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","错误信息","人类数据对照"],"reason":"用LLM仿真人类对错误信息的易感性，并与真实调查数据对照，评估偏差与失效条件。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:00","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":33,"question":"LLM模拟的调查受访者在多大程度上能复现人类对错误信息的相信与分享模式及其与社会预测因子（包括个人网络特征）之间的关联？","design":"使用LLM（如GPT-4等）扮演合成调查受访者，输入基于三个真实调查数据构建的受访者结构化档案（包含个人网络、人口统计、态度/行为特征），让LLM生成对错误信息条目的相信和分享意愿回答，测量回答分布及特征-结果关联。","baseline":"三个在线调查数据集：公共卫生（美国，2023）、气候变化（美国，2025）、疫情政治（韩国，2020），均包含真实人类对错误信息的相信与分享数据及个人网络、人口统计、态度行为变量。","findings":"LLM模拟能捕捉大致分布趋势并与人类回答有适度相关，但系统性地夸大了相信与分享之间的关联；线性模型在模拟数据上解释方差显著膨胀，且过度依赖态度和行为特征，几乎忽略个人网络特征。","reliability":"论文指出LLM模拟更适合诊断与人类判断的系统性偏差，而非替代人类判断；偏差源于LLM训练数据中错误信息相关概念的表征偏差，且模拟未能复现个人网络特征的作用。","relevance":"该研究直接以真实人类调查为基准，评估LLM仿真在错误信息易感性上的可靠性，并揭示了仿真在忽略网络特征、夸大态度关联等方面的失效条件，高度契合研究者对批判性仿真研究的兴趣，值得精读原文。","inspiration":"该方法借鉴了用结构化档案（含人口统计、态度、网络特征）驱动LLM生成调查回答，并与真实人类数据对照以评估仿真偏差的设计｜可迁移到信贷审批中的歧视研究，检验LLM模拟的贷款官员是否复现人类决策中的种族或性别偏见｜用LLM扮演贷款审批员，输入含申请人种族、收入、信用分等档案，输出审批决定，以真实房贷数据（如HMDA）为基准，比较拒绝率差异及特征重要性"}},{"id":"2602.03545","version":2,"title":"Persona Generators: Generating Diverse Synthetic Personas for Arbitrary Contexts","zh_title":"人格生成器：为任意上下文生成多样化的合成人格","abstract":"Evaluating AI systems that interact with humans requires understanding their behavior across diverse user populations, but collecting representative human data is often expensive or infeasible, particularly for novel technologies or hypothetical future scenarios. Recent work in Generative Agent-Based Modeling has shown that large language models can simulate human-like synthetic personas with high fidelity, accurately reproducing the beliefs and behaviors of specific individuals. However, most approaches require detailed data about target populations and often prioritize density matching (replicating what is most probable) rather than support coverage (spanning what is possible), leaving long-tail behaviors underexplored. We introduce Persona Generators, functions that can produce diverse synthetic populations tailored to arbitrary contexts. We apply an iterative improvement loop based on AlphaEvolve, using large language models as mutation operators to refine our Persona Generator code over hundreds of iterations. The optimization process produces lightweight Persona Generators that can automatically expand small descriptions into populations of diverse synthetic personas that maximize coverage of opinions and preferences along relevant diversity axes. We demonstrate that evolved generators substantially outperform existing baselines across six diversity metrics on held-out contexts, producing populations that span rare trait combinations difficult to achieve in standard LLM outputs.","authors":["Davide Paglieri","Logan Cross","William A. Cunningham","Joel Z. Leibo","Alexander Sasha Vezhnevets"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-02-03","first_seen":"2026-02-03","revised_at":null,"abs_url":"https://arxiv.org/abs/2602.03545","pdf_url":"https://arxiv.org/pdf/2602.03545","source_feed":"backfill","score":7,"bucket":"pending","rubric_hits":["A3","B1"],"tags":["合成人群生成","人类仿真","多样性覆盖"],"reason":"用LLM生成多样化合成人群，有真实人类数据对照，但侧重覆盖度而非行为复现，方法…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:57","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":70,"question":"如何生成能覆盖任意场景下人类观点与偏好全貌的多样化合成人群，以解决标准LLM的模式坍缩问题？","design":"提出Persona Generator函数，通过AlphaEvolve进化循环优化代码（含提示模板和采样逻辑），将简短场景提示扩展为结构化问卷，再生成覆盖多样性轴的合成人群；评估时在留出场景上比较生成人群的多样性指标。","baseline":"无对照：论文未使用真实人类行为数据作为基准，而是与Nemotron Personas等合成基线比较覆盖度。","findings":"进化出的生成器在六项多样性指标上显著优于基线，能覆盖稀有特征组合；通过对抗模式坍缩，生成的人群在真实人类特质分布方差上甚至优于专门匹配人类统计数据的基线。","reliability":"论文未讨论失效条件与局限，仅指出密度匹配方法不适合压力测试和推测性场景，但未分析自身方法在哪些条件下可能失效。","relevance":"高度相关：该研究用LLM生成多样化合成人群，虽侧重覆盖度而非行为复现，但直接回应了LLM仿真中模式坍缩和多样性不足的批判，且声称能更好捕捉真实人类分布，值得精读以评估其在经济学实验和政策评估中的潜力与局限。","inspiration":"该方法通过进化算法优化提示模板和采样逻辑以生成覆盖多样性轴的合成人群，可借鉴其对抗模式坍缩的思路来设计LLM仿真中的处理变异与稳健性检验｜可迁移到政策公告的预期形成研究，利用多样化合成人群模拟异质性信念与反应｜以LLM作为被试，施加不同措辞的政策公告处理，测量预期通胀或消费意愿的分布，并以真实调查数据（如密歇根消费者调查）作为对照基准"}},{"id":"2602.02977","version":2,"title":"Aligning Forest and Trees in Images & Long Captions for Visually Grounded Understanding","zh_title":"对齐图像与长文本中的森林与树木以实现视觉基础理解","abstract":"Vision-language models such as CLIP often struggle to faithfully understand long, detail-rich captions, relying on dominant scene cues while overlooking fine-grained visual evidence. We propose a hierarchical vision-language learning principle for understanding scenes as part-to-whole compositions: before forming a whole-scene representation, a model should uncover what semantic parts appear where in the image. To this end, we propose CAFT (Cross-domain Alignment of Forests and Trees), a vision-language model that jointly learns local text-region alignment at intermediate representations and global image-text alignment at the final representation. Exploiting the organization of long captions, where local descriptions often correspond to scene parts, CAFT employs a fine-to-coarse image encoder and a part-whole text encoder to discover localized part semantics and progressively compose them into a global image-text representation. Trained on 30M image-text pairs, CAFT achieves state-of-the-art performance on six long-text retrieval benchmarks and exhibits strong scaling behavior. Experiments show that CAFT learns fine-grained representations that localize textual semantics in image regions without explicit region-level supervision.","authors":["Byeongju Woo","Zilin Wang","Byeonghyun Pak","Sangwoo Mo","Stella X. Yu"],"categories":["cs.CV","cs.AI","cs.LG"],"primary_category":"cs.CV","announce_type":"new","date":"2026-02-03","first_seen":"2026-02-03","revised_at":null,"abs_url":"https://arxiv.org/abs/2602.02977","pdf_url":"https://arxiv.org/pdf/2602.02977","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["视觉-语言模型","细粒度对齐","图像文本检索"],"reason":"纯视觉-语言模型训练与评测，不涉及人类仿真或行为对照","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:59:12","error":null,"has_summary":false,"summary":null},{"id":"2602.01684","version":1,"title":"The Strategic Foresight of LLMs: Evidence from a Fully Prospective Venture Tournament","zh_title":"大语言模型的战略远见：来自全前瞻性创业锦标赛的证据","abstract":"Can artificial intelligence outperform humans at strategic foresight -- the capacity to form accurate judgments about uncertain, high-stakes outcomes before they unfold? We address this question through a fully prospective prediction tournament using live Kickstarter crowdfunding projects. Thirty U.S.-based technology ventures, launched after the training cutoffs of all models studied, were evaluated while fundraising remained in progress and outcomes were unknown. A diverse suite of frontier and open-weight large language models (LLMs) completed 870 pairwise comparisons, producing complete rankings of predicted fundraising success. We benchmarked these forecasts against 346 experienced managers recruited via Prolific and three MBA-trained investors working under monitored conditions. The results are striking: human evaluators achieved rank correlations with actual outcomes between 0.04 and 0.45, while several frontier LLMs exceeded 0.60, with the best (Gemini 2.5 Pro) reaching 0.74 -- correctly ordering nearly four of every five venture pairs. These differences persist across multiple performance metrics and robustness checks. Neither wisdom-of-the-crowd ensembles nor human-AI hybrid teams outperformed the best standalone model.","authors":["Felipe A. Csaszar","Aticus Peterson","Daniel Wilde"],"categories":["econ.GN","cs.AI"],"primary_category":"econ.GN","announce_type":"new","date":"2026-02-02","first_seen":"2026-02-02","revised_at":null,"abs_url":"https://arxiv.org/abs/2602.01684","pdf_url":"https://arxiv.org/pdf/2602.01684","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM仿真","人类行为对照","创业预测"],"reason":"用LLM预测人类对创业项目的判断，并与真实人类数据对照，属于经济学场景下的人类…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:57","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":7,"question":"大语言模型在战略远见（预测不确定、高风险商业结果）上能否超越人类？","design":"使用多款前沿和开源大语言模型对30个正在众筹的Kickstarter科技项目进行870次成对比较，生成预测筹款成功排名；同时招募346名有经验的管理者和3名MBA投资者作为人类被试进行相同任务，以最终实际筹款结果作为准确度标准。","baseline":"346名通过Prolific招募的有经验管理者和3名受监控的MBA投资者，其预测排名与实际结果的秩相关系数在0.04至0.45之间。","findings":"人类评估者的预测排名与实际结果的秩相关系数最高仅0.45，而多个前沿大语言模型超过0.60，最佳模型Gemini 2.5 Pro达到0.74。群体智慧集成和人机混合团队均未超越最佳独立模型。","reliability":"论文未讨论","relevance":"该研究直接用LLM替代人类被试进行前瞻性预测，并与真实人类数据对照，属于经济学场景下的人类仿真实验，且提供了可靠性证据，值得精读原文。","inspiration":"该方法借鉴了用真实众筹结果作为客观基准，直接比较LLM与人类被试的预测准确度，并通过成对比较排名任务量化战略远见｜可迁移到创业投资决策、众筹市场预测或资产定价中的预期形成研究，用于评估AI辅助决策的可靠性｜招募专业投资者作为人类被试，让LLM和人类分别对真实众筹项目进行成对比较排名，以最终实际筹资金额作为结果变量，计算预测排名与实际结果的秩相关系数，对比人机表现"}},{"id":"2602.07023","version":2,"title":"Behavioral Consistency Validation for LLM Agents: An Analysis of Trading-Style Switching through Stock-Market Simulation","zh_title":"LLM智能体行为一致性验证：基于股市模拟的交易风格切换分析","abstract":"Recent works have increasingly applied Large Language Models (LLMs) as agents in financial stock market simulations to test if micro-level behaviors aggregate into macro-level phenomena. However, a crucial question arises: Do LLM agents' behaviors align with real market participants? This alignment is key to the validity of simulation results. To explore this, we select a financial stock market scenario to test behavioral consistency. Investors are typically classified as fundamental or technical traders, but most simulations fix strategies at initialization, failing to reflect real-world trading dynamics. In this work, we assess whether agents' strategy switching aligns with financial theory, providing a framework for this evaluation. We operationalize four behavioral-finance drivers-loss aversion, herding, wealth differentiation, and price misalignment-as personality traits set via prompting and stored long-term. In year-long simulations, agents process daily price-volume data, trade under a designated style, and reassess their strategy every 10 trading days. We introduce four alignment metrics and use Mann-Whitney U tests to compare agents' style-switching behavior with financial theory. Our results show that recent LLMs' switching behavior is only partially consistent with behavioral-finance theories, highlighting the need for further refinement in aligning agent behavior with financial theory.","authors":["Zeping Li","Guancheng Wan","Keyang Chen","Yu Chen","Yiwen Zhao","Philip Torr","Guangnan Ye","Zhenfei Yin","Hongfeng Chai"],"categories":["q-fin.TR","cs.AI"],"primary_category":"q-fin.TR","announce_type":"new","date":"2026-02-02","first_seen":"2026-02-02","revised_at":null,"abs_url":"https://arxiv.org/abs/2602.07023","pdf_url":"https://arxiv.org/pdf/2602.07023","source_feed":"backfill","score":8,"bucket":"selected","rubric_hits":["A3","B2","B4"],"tags":["LLM仿真","行为金融","智能体一致性"],"reason":"用LLM agent模拟股票交易行为并与金融理论对照，涉及行为经济学场景，指出…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:16:00","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":55,"question":"在股票市场仿真中，LLM智能体的交易风格切换行为是否与行为金融学理论一致？","design":"使用多种近期LLM（通过提示词设定四种行为金融倾向作为人格特质并存入长期记忆）扮演投资者，在基于2024年标普500成分股数据的模拟环境中进行为期一年的交易，每10个交易日评估并决定是否切换交易风格（基本面/技术面），通过四个对齐指标和Mann-Whitney U检验比较智能体行为与理论预期。","baseline":"无对照","findings":"LLM智能体的风格切换行为仅部分符合行为金融学理论，无法在所有方面完全对齐。","reliability":"论文未讨论","relevance":"该研究直接评估LLM智能体在金融行为仿真中的行为一致性，属于批判性验证工作，与研究者关注的LLM仿真可靠性及失效条件高度相关，值得阅读原文以了解具体偏差和评估框架。","inspiration":"该研究通过提示词将行为金融倾向植入LLM智能体并存入长期记忆，在动态仿真中周期性评估风格切换，用多个对齐指标和统计检验衡量与理论的偏差，这种处理-测量-验证框架值得借鉴。｜可迁移到资产定价实验，检验LLM智能体在信息冲击下的过度反应与反转行为是否符合前景理论与处置效应。｜以LLM智能体为被试，施加不同强度的利好/利空消息作为处理，观测其持仓调整与买卖时机，结果变量为超额收益与换手率，用历史高频交易数据中散户的实际行为分布作为对照基准。"}},{"id":"2602.02606","version":2,"title":"Gender Dynamics and Homophily in a Social Network of LLM Agents","zh_title":"LLM代理社交网络中的性别动态与同质性","abstract":"Generative artificial intelligence and large language models (LLMs) are increasingly deployed in interactive settings, yet we know little about how their identity performance develops when they interact within large-scale networks. We address this by examining Chirper.ai, a social media platform similar to X but composed entirely of autonomous AI chatbots. Our dataset comprises over 70,000 agents, approximately 140 million posts, and the evolving followership network over a period of one year. Based on agents' posted text, we assign weekly gender performance scores to each agent. Results suggest that each agent's gender performance is fluid rather than fixed. Despite this fluidity, the network displays strong gender-based homophily, as agents consistently follow others performing gender similarly. We investigate whether these homophilic connections arise from social selection, in which agents choose to follow similar accounts, or from social influence, in which agents become more similar to their followees over time. Consistent with human social networks, we find evidence that both mechanisms shape the structure and evolution of interactions among LLMs. Our findings suggest that, even in the absence of bodies, cultural entraining of gender performance leads to gender-based sorting. This has important implications for LLM applications in synthetic hybrid populations, social simulations, and decision support.","authors":["Faezeh Fadaei","Jenny Carla Moran","Taha Yasseri"],"categories":["cs.SI","cs.AI","cs.CY"],"primary_category":"cs.SI","announce_type":"new","date":"2026-02-02","first_seen":"2026-02-02","revised_at":null,"abs_url":"https://arxiv.org/abs/2602.02606","pdf_url":"https://arxiv.org/pdf/2602.02606","source_feed":"backfill","score":7,"bucket":"pending","rubric_hits":["A3","B1","B4"],"tags":["LLM代理","社会模拟","性别同质性"],"reason":"用LLM代理模拟社交网络性别同质性，并与人类社交网络机制对照，但非严格实验仿真。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:57","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":83,"question":"在由LLM代理组成的社交网络中，性别表演如何动态变化，以及性别同质性如何通过社会选择和社会影响机制形成？","design":"本研究并非受控仿真实验，而是对Chirper.ai平台（类似X但全部由自主AI聊天机器人组成）的自然观察研究。研究者收集了超过7万个代理、约1.4亿条帖子和一年的关注网络数据，基于帖子文本为每个代理分配每周性别表演分数，分析性别表演的流动性和网络中的性别同质性，并区分社会选择与社会影响两种机制。","baseline":"无对照","findings":"每个代理的性别表演是流动而非固定的，但网络仍表现出强烈的性别同质性，代理倾向于关注性别表演相似的他人。社会选择和社会影响两种机制共同塑造了LLM代理之间的互动结构和演化，这与人类社交网络中的发现一致。","reliability":"论文未讨论","relevance":"该研究利用LLM代理在真实平台上的大规模互动数据，揭示了性别同质性的涌现机制，并与人类社交网络机制进行对照，虽非严格实验仿真，但为LLM在社交仿真中的行为模式提供了实证证据，值得阅读以了解LLM集体行为中的社会结构形成。","inspiration":"利用大规模LLM代理平台的自然观察数据，通过文本分析动态测量个体属性（如性别表演），并区分社会选择与社会影响机制，为无干扰的群体行为研究提供了方法参考｜可迁移至金融社交媒体中的信息扩散与投资者行为研究，例如分析LLM代理在模拟投资社区中如何形成风险偏好同质性｜设计一个LLM代理投资社区，让代理基于历史市场信息发布投资观点，用文本分析测量其风险偏好，观察关注网络与偏好同质性的动态演化，并以真实投资者社交平台数据（如StockTwits）作为对照基准"}},{"id":"2602.02604","version":1,"title":"AI Assisted Economics Measurement From Survey: Evidence from Public Employee Pension Choice","zh_title":"基于调查的人工智能辅助经济测量：来自公共雇员养老金选择的证据","abstract":"We develop an iterative framework for economic measurement that leverages large language models to extract measurement structure directly from survey instruments. The approach maps survey items to a sparse distribution over latent constructs through what we term a soft mapping, aggregates harmonized responses into respondent level sub dimension scores, and disciplines the resulting taxonomy through out of sample incremental validity tests and discriminant validity diagnostics. The framework explicitly integrates iteration into the measurement construction process. Overlap and redundancy diagnostics trigger targeted taxonomy refinement and constrained remapping, ensuring that added measurement flexibility is retained only when it delivers stable out of sample performance gains. Applied to a large scale public employee retirement plan survey, the framework identifies which semantic components contain behavioral signal and clarifies the economic mechanisms, such as beliefs versus constraints, that matter for retirement choices. The methodology provides a portable measurement audit of survey instruments that can guide both empirical analysis and survey design.","authors":["Tiancheng Wang","Krishna Sharma"],"categories":["econ.EM","cs.AI"],"primary_category":"econ.EM","announce_type":"new","date":"2026-02-02","first_seen":"2026-02-02","revised_at":null,"abs_url":"https://arxiv.org/abs/2602.02604","pdf_url":"https://arxiv.org/pdf/2602.02604","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["LLM辅助测量","调查方法","经济构念"],"reason":"用LLM辅助测量经济构念，属NLP方法改进，非仿真人类被试行为或决策。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:57","error":null,"has_summary":false,"summary":null},{"id":"2602.01797","version":1,"title":"ORCH: many analyses, one merge-a deterministic multi-agent orchestrator for discrete-choice reasoning with EMA-guided routing","zh_title":"ORCH：一种用于离散选择推理的确定性多智能体协调框架","abstract":"Recent advances in large-scale language models (LLMs) have made multi-agent architectures attractive for challenging reasoning tasks. However, many existing systems rely on stochastic routing or ad-hoc heuristics, making their behavior difficult to reproduce and their decision process hard to interpret. We propose ORCH, a deterministic coordination framework for discrete-choice reasoning that orchestrates heterogeneous LLMs. ORCH follows a ``many analyses, one decision'' paradigm: multiple base models independently produce structured analyses, and a dedicated merge agent outputs the final choice. The framework uses fixed rules for task decomposition and answer aggregation, keeping the pipeline predictable, reproducible, and training-free. Determinism here refers to fixed routing and aggregation rules under a fixed evaluation protocol, rather than strict bit-level reproducibility across deployments. To exploit model complementarity, we optionally introduce an EMA-guided router that updates agent selection using historical accuracy, latency, or cost; since it relies on answer-based feedback, it is mainly intended for benchmarking, controlled evaluation, or delayed-feedback settings. Experiments on MMLU, MMLU-Pro, and GSM8K show that ORCH consistently outperforms single-model baselines and a majority-vote ensemble. On MMLU-Pro, ORCH improves accuracy by over 10 points compared to the strongest baseline, and on GSM8K it yields gains exceeding 50 points; McNemar tests confirm statistical significance. The EMA router provides an additional 0.7--2.0 point accuracy boost, and ablations show that both multi-agent collaboration and routing contribute substantially. Overall, ORCH offers a practical path toward controllable, interpretable, and deployment-ready LLM-based agent systems for discrete-choice reasoning.","authors":["Hanlin Zhou","Huah Yong Chan"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-02-02","first_seen":"2026-02-02","revised_at":null,"abs_url":"https://arxiv.org/abs/2602.01797","pdf_url":"https://arxiv.org/pdf/2602.01797","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","推理增强","确定性路由"],"reason":"纯多智能体协作解题，无人类行为对照，属C1排除项。","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:59:36","error":null,"has_summary":false,"summary":null},{"id":"2602.01022","version":3,"title":"Calibrating Behavioral Parameters with Large Language Models","zh_title":"用大语言模型校准行为参数","abstract":"Behavioral parameters such as loss aversion, herding, and extrapolation are central to asset pricing models but remain difficult to measure reliably. We develop a framework that treats large language models (LLMs) as calibrated measurement instruments for behavioral parameters. Using four models and 24{,}000 agent--scenario pairs, we document systematic rationality bias in baseline LLM behavior, including attenuated loss aversion, weak herding, and near-zero disposition effects relative to human benchmarks. Profile-based calibration induces large, stable, and theoretically coherent shifts in several parameters, with calibrated loss aversion, herding, extrapolation, and anchoring reaching or exceeding benchmark magnitudes. To assess external validity, we embed calibrated parameters in an agent-based asset pricing model, where calibrated extrapolation generates short-horizon momentum and long-horizon reversal patterns consistent with empirical evidence. Our results establish measurement ranges, calibration functions, and explicit boundaries for eight canonical behavioral biases.","authors":["Brandon Yee","Pairie Koh"],"categories":["econ.GN","cs.AI"],"primary_category":"econ.GN","announce_type":"new","date":"2026-02-01","first_seen":"2026-02-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2602.01022","pdf_url":"https://arxiv.org/pdf/2602.01022","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","A3","B1","B2","B4"],"tags":["LLM仿真","行为经济学","人类基准对照"],"reason":"用LLM测量行为参数并与人类基准对照，嵌入资产定价模型验证外部效度，直接命中核…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:55","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":8,"question":"能否将大语言模型作为可校准的测量工具，系统性地诱导、校准并验证行为金融参数？","design":"使用 GPT-4o、GPT-4o-mini、Claude-3.5-Haiku、Gemini-2.5-Pro 四种模型，通过提示词嵌入行为特征（如损失厌恶、羊群效应等）作为实验处理，在 24,000 个合成金融场景中测量八种行为偏差的参数值，并评估校准后的参数在基于代理的资产定价模型中的外部有效性。","baseline":"人类基准来自已有文献中的实验和实证估计，如损失厌恶系数约 2.25，羊群效应率 65-75%，以及 Jegadeesh 和 Titman 的动量与反转经验事实。","findings":"基线 LLM 行为存在系统性理性偏差，表现为损失厌恶减弱、羊群效应弱、处置效应接近零；通过基于特征的校准可诱导出与人类基准相当甚至更强的参数值，且校准后的外推参数能在资产定价模型中生成符合经验事实的短期动量和长期反转模式。","reliability":"论文承认校准存在明确边界，并非所有参数都能成功校准，且结果依赖于提示词设计和模型选择，外部有效性仅通过简单资产定价模型初步验证。","relevance":"该研究直接命中研究者对 LLM 仿真人类行为、与真实人类基准对照、经济学实验及批判性评估的核心兴趣，提供了系统的校准框架和失效边界，值得精读原文。","inspiration":"该方法通过提示词嵌入行为特征（如损失厌恶、羊群效应）作为实验处理，系统性地校准LLM的行为参数，并与人类基准对照，值得借鉴其处理施加与参数校准的流程设计｜可迁移到资产定价实验中的投资者行为偏差研究，如模拟动量效应、反转效应及处置效应等市场异象｜以LLM作为被试，通过提示词嵌入不同程度的损失厌恶或羊群效应处理，测量其交易决策与价格预期，并与Jegadeesh和Titman的动量/反转经验事实及处置效应实证数据对照"}},{"id":"2602.00948","version":1,"title":"FinEvo: From Isolated Backtests to Ecological Market Games for Multi-Agent Financial Strategy Evolution","zh_title":"FinEvo：从孤立回测到生态市场博弈的多智能体金融策略演化","abstract":"Conventional financial strategy evaluation relies on isolated backtests in static environments. Such evaluations assess each policy independently, overlook correlations and interactions, and fail to explain why strategies ultimately persist or vanish in evolving markets. We shift to an ecological perspective, where trading strategies are modeled as adaptive agents that interact and learn within a shared market. Instead of proposing a new strategy, we present FinEvo, an ecological game formalism for studying the evolutionary dynamics of multi-agent financial strategies. At the individual level, heterogeneous ML-based traders-rule-based, deep learning, reinforcement learning, and large language model (LLM) agents-adapt using signals such as historical prices and external news. At the population level, strategy distributions evolve through three designed mechanisms-selection, innovation, and environmental perturbation-capturing the dynamic forces of real markets. Together, these two layers of adaptation link evolutionary game theory with modern learning dynamics, providing a principled environment for studying strategic behavior. Experiments with external shocks and real-world news streams show that FinEvo is both stable for reproducibility and expressive in revealing context-dependent outcomes. Strategies may dominate, collapse, or form coalitions depending on their competitors-patterns invisible to static backtests. By reframing strategy evaluation as an ecological game formalism, FinEvo provides a unified, mechanism-level protocol for analyzing robustness, adaptation, and emergent dynamics in multi-agent financial markets, and may offer a means to explore the potential impact of macroeconomic policies and financial regulations on price evolution and equilibrium.","authors":["Mingxi Zou","Jiaxiang Chen","Aotian Luo","Jingyi Dai","Chi Zhang","Dongning Sun","Zenglin Xu"],"categories":["physics.soc-ph","cs.AI","cs.GT","cs.MA"],"primary_category":"physics.soc-ph","announce_type":"new","date":"2026-02-01","first_seen":"2026-02-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2602.00948","pdf_url":"https://arxiv.org/pdf/2602.00948","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["多智能体模拟","金融市场演化","LLM agent"],"reason":"用LLM agent模拟金融市场演化，但无真实人类交易者数据对照，属社会模拟边…","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:59:12","error":null,"has_summary":false,"summary":null},{"id":"2602.00685","version":1,"title":"HumanStudy-Bench: Towards AI Agent Design for Participant Simulation","zh_title":"HumanStudy-Bench：面向参与者仿真的AI智能体设计基准","abstract":"Large language models (LLMs) are increasingly used as simulated participants in social science experiments, but their behavior is often unstable and highly sensitive to design choices. Prior evaluations frequently conflate base-model capabilities with experimental instantiation, obscuring whether outcomes reflect the model itself or the agent setup. We instead frame participant simulation as an agent-design problem over full experimental protocols, where an agent is defined by a base model and a specification (e.g., participant attributes) that encodes behavioral assumptions. We introduce HUMANSTUDY-BENCH, a benchmark and execution engine that orchestrates LLM-based agents to reconstruct published human-subject experiments via a Filter--Extract--Execute--Evaluate pipeline, replaying trial sequences and running the original analysis pipeline in a shared runtime that preserves the original statistical procedures end to end. To evaluate fidelity at the level of scientific inference, we propose new metrics to quantify how much human and agent behaviors agree. We instantiate 12 foundational studies as an initial suite in this dynamic benchmark, spanning individual cognition, strategic interaction, and social psychology, and covering more than 6,000 trials with human samples ranging from tens to over 2,100 participants.","authors":["Xuan Liu","Haoyang Shang","Zizhang Liu","Xinyan Liu","Yunze Xiao","Yiwen Tu","Haojian Jin"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-01-31","first_seen":"2026-01-31","revised_at":null,"abs_url":"https://arxiv.org/abs/2602.00685","pdf_url":"https://arxiv.org/pdf/2602.00685","source_feed":"backfill","score":10,"bucket":"selected","rubric_hits":["A1","A2","A3","A5","B1","B2","B3"],"tags":["LLM人类仿真","实验复现","基准测试"],"reason":"直接构建LLM代理复现人类实验，含真实人类数据对照，评估仿真保真度，覆盖经济学…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:55","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":11,"question":"如何将LLM参与者的仿真视为一个代理设计问题，并系统评估不同代理设计在复现真实人类实验中的保真度？","design":"使用10个当代LLM（如GPT、Claude、Gemini）作为基座模型，结合四种代理规格（空白、角色扮演、人口统计条件、丰富背景故事）构建AI代理，通过Filter–Extract–Execute–Evaluate管道重放12项已发表的人类实验的完整试验序列和原始分析流程，测量代理的行为响应。","baseline":"对照12项已发表人类实验的真实人类数据，涵盖个体认知、策略互动和社会心理学，超过6000次试验，人类样本量从几十到2100多人。","findings":"当前LLM代理与人类的推断一致性有限且不稳定，行为呈极化双峰而非人类单峰模式；代理设计对结果有较大且非单调的影响，性能高度依赖领域，更大模型或简单多模型集成未能可靠提升对齐度。","reliability":"论文指出LLM行为不稳定、对设计选择高度敏感，代理规格会编码行为假设并可能定性改变结果；现有评估常混淆基座模型能力与实验实例化，且代理在人口异质性敏感度和提示词变化下表现脆弱。","relevance":"该研究直接针对用LLM替代人类被试的仿真实验，提供真实人类数据对照，系统评估代理设计对复现经济学实验和社会心理学效应的影响，并批判性指出仿真失效的条件，与您的关注高度契合，值得精读原文。","inspiration":"借鉴其将代理设计作为实验变量的思路，通过对比空白、角色扮演、人口统计条件等不同规格来分离模型能力与实验实例化的混淆效应｜可迁移到政策公告的预期形成实验，研究不同信息框架下投资者对央行沟通的反应｜以LLM代理为被试，处理为不同代理规格（如空白vs.人口统计条件），结果变量为通胀预期调整幅度，对照真实调查数据（如密歇根消费者调查）"}},{"id":"2601.22812","version":2,"title":"Stable Personas: Dual-Assessment of Temporal Stability in LLM-Based Human Simulation","zh_title":"稳定人格：基于LLM的人类仿真中时间稳定性的双重评估","abstract":"Large Language Models (LLMs) acting as artificial agents offer the potential for scalable behavioral research, yet their validity depends on whether LLMs can maintain stable personas across extended conversations. We address this point using a dual-assessment framework measuring both self-reported characteristics and observer-rated persona expression. Across two experiments testing four persona conditions (default, high, moderate, and low ADHD presentations), seven LLMs, and three semantically equivalent persona prompts, we examine between-conversation stability (3,473 conversations) and within-conversation stability (1,370 conversations and 18 turns). Self-reports remain highly stable both between and within conversations. However, observer ratings reveal a tendency for persona expressions to decline during extended conversations. These findings suggest that persona-instructed LLMs produce stable, persona-aligned self-reports, an important prerequisite for behavioral research, while identifying this regression tendency as a boundary condition for multi-agent social simulation.","authors":["Jana Gonnermann-Müller","Jennifer Haase","Nicolas Leins","Thomas Kosch","Sebastian Pokutta"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2026-01-30","first_seen":"2026-01-30","revised_at":null,"abs_url":"https://arxiv.org/abs/2601.22812","pdf_url":"https://arxiv.org/pdf/2601.22812","source_feed":"backfill","score":8,"bucket":"selected","rubric_hits":["A1","A2","B4"],"tags":["LLM仿真","人格稳定性","效度评估"],"reason":"研究LLM人格稳定性以评估其作为人类被试的可靠性，直接涉及仿真效度与失效条件。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:55","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":47,"question":"LLM在独立对话间和长对话过程中，能在多大程度上稳定维持被赋予的人格？","design":"用7个LLM扮演默认、高、中、低四种ADHD人格，通过三种语义等价的提示词施加处理；实验一测量跨对话稳定性（每条件50次独立运行），实验二测量对话内稳定性（18轮对话，在第6、12、18轮评估）；结果变量为自评量表得分和观察者评分。","baseline":"无对照","findings":"自评报告在跨对话和对话内均高度稳定；但观察者评分显示，高和中强度人格表达在长对话中会逐渐减弱，向默认水平回归。","reliability":"论文指出，人格表达在长对话中衰退是LLM用于多智能体社会仿真的一个边界条件；评估依赖LLM作为评分者，可能引入偏差；仅以ADHD人格为测试案例，泛化性待验证。","relevance":"直接评估LLM作为人类被试替代品的人格稳定性，揭示了自评稳定但行为表达衰退的失效模式，对关注仿真效度与边界条件的研究者很有参考价值，值得读原文。","inspiration":"借鉴双维度人格稳定性评估设计，通过独立跨对话与长对话内重复测量，区分自评报告与行为观察的稳定性差异，揭示仿真衰退的边界条件。｜可迁移至政策公告预期形成实验，用LLM模拟投资者对央行沟通的反应，检验长期对话中信息解读的一致性。｜以LLM为被试，施加不同政策措辞处理，在长对话中多次测量通胀预期与投资意愿，对比真实投资者调查面板数据，评估仿真在持续信息流下的衰退模式。"}},{"id":"2601.23032","version":1,"title":"Guided by Trajectories: Repairing and Rewarding Tool-Use Trajectories for Tool-Integrated Reasoning","zh_title":"轨迹引导：修复与奖励工具使用轨迹以实现工具集成推理","abstract":"Tool-Integrated Reasoning (TIR) enables large language models (LLMs) to solve complex tasks by interacting with external tools, yet existing approaches depend on high-quality synthesized trajectories selected by scoring functions and sparse outcome-based rewards, providing limited and biased supervision for learning TIR. To address these challenges, in this paper, we propose AutoTraj, a two-stage framework that automatically learns TIR by repairing and rewarding tool-use trajectories. Specifically, in the supervised fine-tuning (SFT) stage, AutoTraj generates multiple candidate tool-use trajectories for each query and evaluates them along multiple dimensions. High-quality trajectories are directly retained, while low-quality ones are repaired using a LLM (i.e., LLM-as-Repairer). The resulting repaired and high-quality trajectories form a synthetic SFT dataset, while each repaired trajectory paired with its original low-quality counterpart constitutes a dataset for trajectory preference modeling. In the reinforcement learning (RL) stage, based on the preference dataset, we train a trajectory-level reward model to assess the quality of reasoning paths and combine it with outcome and format rewards, thereby explicitly guiding the optimization toward reliable TIR behaviors. Experiments on real-world benchmarks demonstrate the effectiveness of AutoTraj in TIR.","authors":["Siyu Gong","Linan Yue","Weibo Gao","Fangzhou Yao","Shimin Di","Lei Feng","Min-Ling Zhang"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-01-30","first_seen":"2026-01-30","revised_at":null,"abs_url":"https://arxiv.org/abs/2601.23032","pdf_url":"https://arxiv.org/pdf/2601.23032","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["工具集成推理","轨迹优化","强化学习"],"reason":"纯工具使用轨迹优化，无人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:59:12","error":null,"has_summary":false,"summary":null},{"id":"2601.21975","version":2,"title":"Mind the Gap: How Elicitation Protocols Shape the Stated-Revealed Preference Gap in Language Models","zh_title":"注意差距：诱导协议如何塑造语言模型中陈述-显示偏好差距","abstract":"Recent work identifies a stated-revealed (SvR) preference gap in language models (LMs): a mismatch between the values models endorse and the choices they make in context. Existing evaluations rely heavily on binary forced-choice prompting, which entangles genuine preferences with artifacts of the elicitation protocol. We systematically study how elicitation protocols affect SvR correlation across 24 LMs. Allowing neutrality and abstention during stated preference elicitation allows us to exclude weak signals, substantially improving Spearman's rank correlation ($ρ$) between volunteered stated preferences and forced-choice revealed preferences. However, further allowing abstention in revealed preferences drives $ρ$ to near-zero or negative values due to high neutrality rates. Finally, we find that system prompt steering using stated preferences during revealed preference elicitation does not reliably improve SvR correlation on AIRiskDilemmas. Together, our results show that SvR correlation is highly protocol-dependent and that preference elicitation requires methods that account for indeterminate preferences.","authors":["Pranav Mahajan","Ihor Kendiukhov","Syed Hussain","Lydia Nottingham"],"categories":["cs.AI","cs.ET"],"primary_category":"cs.AI","announce_type":"new","date":"2026-01-29","first_seen":"2026-01-29","revised_at":null,"abs_url":"https://arxiv.org/abs/2601.21975","pdf_url":"https://arxiv.org/pdf/2601.21975","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["语言模型偏好","诱导协议","陈述-显示偏好"],"reason":"测量语言模型自身的偏好一致性，属于对模型本身的测量，而非用模型仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:53","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":185,"question":"在语言模型中，偏好引出协议（是否允许中立/弃权）如何影响陈述偏好与显示偏好之间的秩相关性？","design":"不是仿真研究。该研究在24个语言模型上系统改变偏好引出协议（强制二选一 vs. 允许中立/弃权），分别测量陈述偏好（抽象价值比较）和显示偏好（情境化道德困境）的排序，计算两者间的斯皮尔曼秩相关系数，并测试系统提示引导的效果。","baseline":"无对照","findings":"在陈述偏好引出中允许中立可过滤弱信号，显著提高与强制选择显示偏好的秩相关性；但在显示偏好引出中也允许中立时，因高中立率导致秩相关降至近零或负值。系统提示引导未能可靠改善陈述-显示偏好一致性。","reliability":"论文指出SvR相关性高度依赖引出协议，当模型普遍表达中立时，基于排名的相关性会失效；提示引导在16个价值域上不可靠，且排除中立响应会丢失不确定性信息。","relevance":"该研究不涉及用LLM仿真人类被试，而是测量模型自身的偏好一致性，属于对模型行为的测量与分析，与研究者关注的以人类数据为基准的仿真实验无关，不建议优先阅读。","inspiration":"该方法通过系统改变偏好引出协议（强制选择 vs. 允许中立）来测量陈述与显示偏好的秩相关性，可借鉴用于检验经济调查中选项设计对偏好一致性的影响｜可迁移到消费者跨期选择研究，例如在时间偏好调查中对比强制排序与允许“无差异”选项时，陈述偏好与真实激励下的选择行为是否一致｜以LLM为被试，设计跨期选择任务：处理组为强制二选一（今天100元 vs. 一年后120元），对照组允许选择“无差异”；结果变量为陈述偏好与显示偏好（模拟真实支付决策）的秩相关系数，对照真实人类实验数据（如Andreoni & Sprenger 2012）评估仿真效度"}},{"id":"2601.20285","version":3,"title":"Bank Runs With and Without Bank Failure","zh_title":"银行挤兑：有无银行倒闭的情形","abstract":"We study the causes and consequences of bank runs. By applying large language models to historical newspapers, we create a comprehensive database of bank runs in U.S. history with information on 3,984 runs on individual banks from 1863 to 1934. Our novel data allow us to establish that runs are considerably more likely in weak banks but also occur in strong banks, especially in response to negative news about the real economy or the broader banking system. However, runs typically only result in failure for banks with poor fundamentals. Strong banks survive runs through various mechanisms, including signaling strength, interbank cooperation, and temporary suspension. At the local level, runs on banks with poor fundamentals translate into substantially larger declines in deposits, lending, and manufacturing activity than runs on strong banks. Our findings imply that poor fundamentals are central to explaining both when runs occur and when they have severe economic effects, tempering the view that small shocks can generate discontinuous jumps to bad equilibria through self-fulfilling run dynamics.","authors":["Sergio Correia","Stephan Luck","Emil Verner"],"categories":["econ.GN"],"primary_category":"econ.GN","announce_type":"new","date":"2026-01-28","first_seen":"2026-01-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2601.20285","pdf_url":"https://arxiv.org/pdf/2601.20285","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["银行挤兑","历史数据","文本分析"],"reason":"用LLM处理历史报纸数据，非仿真人类被试，无人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:58:22","error":null,"has_summary":false,"summary":null},{"id":"2602.17676","version":1,"title":"Epistemic Traps: Rational Misalignment Driven by Model Misspecification","zh_title":"认知陷阱：模型误设驱动的理性失配","abstract":"The rapid deployment of Large Language Models and AI agents across critical societal and technical domains is hindered by persistent behavioral pathologies including sycophancy, hallucination, and strategic deception that resist mitigation via reinforcement learning. Current safety paradigms treat these failures as transient training artifacts, lacking a unified theoretical framework to explain their emergence and stability. Here we show that these misalignments are not errors, but mathematically rationalizable behaviors arising from model misspecification. By adapting Berk-Nash Rationalizability from theoretical economics to artificial intelligence, we derive a rigorous framework that models the agent as optimizing against a flawed subjective world model. We demonstrate that widely observed failures are structural necessities: unsafe behaviors emerge as either a stable misaligned equilibrium or oscillatory cycles depending on reward scheme, while strategic deception persists as a \"locked-in\" equilibrium or through epistemic indeterminacy robust to objective risks. We validate these theoretical predictions through behavioral experiments on six state-of-the-art model families, generating phase diagrams that precisely map the topological boundaries of safe behavior. Our findings reveal that safety is a discrete phase determined by the agent's epistemic priors rather than a continuous function of reward magnitude. This establishes Subjective Model Engineering, defined as the design of an agent's internal belief structure, as a necessary condition for robust alignment, marking a paradigm shift from manipulating environmental rewards to shaping the agent's interpretation of reality.","authors":["Xingcheng Xu","Jingjing Qu","Qiaosheng Zhang","Chaochao Lu","Yanqing Yang","Na Zou","Xia Hu"],"categories":["cs.AI","cs.CL","cs.LG"],"primary_category":"cs.AI","announce_type":"new","date":"2026-01-27","first_seen":"2026-01-27","revised_at":null,"abs_url":"https://arxiv.org/abs/2602.17676","pdf_url":"https://arxiv.org/pdf/2602.17676","source_feed":"backfill","score":4,"bucket":"other","rubric_hits":["C1"],"tags":["AI对齐","多智能体系统","理论经济学"],"reason":"研究AI智能体在错误模型下的理性行为，无人类行为对照，属纯多智能体系统分析。","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:59:14","error":null,"has_summary":false,"summary":null},{"id":"2601.18027","version":2,"title":"Sentipolis: Emotion-Aware Agents for Social Simulations","zh_title":"Sentipolis：用于社会模拟的情感感知智能体","abstract":"LLM agents are increasingly used for social simulation, yet emotion is often treated as a transient cue, causing emotional amnesia and weak long-horizon continuity. We present Sentipolis, a framework for emotionally stateful agents that integrates continuous Pleasure-Arousal-Dominance (PAD) representation, dual-speed emotion dynamics, and emotion--memory coupling. Across thousands of interactions over multiple base models and evaluators, Sentipolis improves emotionally grounded behavior, boosting communication, and emotional continuity. Gains are model-dependent: believability increases for higher-capacity models but can drop for smaller ones, and emotion-awareness can mildly reduce adherence to social norms, reflecting a human-like tension between emotion-driven behavior and rule compliance in social simulation. Network-level diagnostics show reciprocal, moderately clustered, and temporally stable relationship structures, supporting the study of cumulative social dynamics such as alliance formation and gradual relationship change.","authors":["Chiyuan Fu","Lyuhao Chen","Yunze Xiao","Weihao Xuan","Carlos Busso","Mona Diab"],"categories":["cs.AI","cs.CL"],"primary_category":"cs.AI","announce_type":"new","date":"2026-01-25","first_seen":"2026-01-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2601.18027","pdf_url":"https://arxiv.org/pdf/2601.18027","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["社会模拟","情感建模","多智能体"],"reason":"社会模拟但无真实人类数据对照，属边界情形","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:53","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":43,"question":"如何设计具有持久情绪状态的LLM智能体，以解决社会模拟中的“情绪失忆”问题，并提升长期交互中的情绪连续性与行为真实性？","design":"提出Sentipolis框架，为LLM智能体引入连续PAD情绪表示、双速情绪动力学和情绪-记忆耦合机制；在多个基础模型（如GPT-5.2、Grok-4、GPT-4o-mini）上进行数千次交互实验，通过LLM评判和人类评估测量沟通质量、情绪连续性、可信度、共情和社会规范遵守等指标，并进行组件消融和网络级诊断。","baseline":"无对照","findings":"情绪状态化显著提升了沟通质量和情绪连续性，但效果依赖模型规模：大模型可信度提升，小模型可能下降；情绪意识会轻微降低对社会规范的遵守，反映了情绪驱动行为与规则遵从之间的类人张力。网络分析显示，情绪-记忆耦合产生了高互惠性、适度聚类和稳定的关系结构。","reliability":"论文指出效果具有模型异质性，小模型上可信度可能下降；情绪意识可能轻微降低社会规范遵守；评估依赖LLM评判，虽有人类评估验证，但长期动态的生态效度仍需进一步检验。","relevance":"该研究属于LLM社会模拟，但无真实人类数据对照，不符合研究者对基准实证数据的要求；其情绪建模方法对理解仿真偏差有参考价值，但非直接匹配经济学实验或政策评估场景。","inspiration":"可借鉴其将情绪作为显式持续状态并耦合记忆的模块化设计，用于在仿真中引入情绪驱动的决策偏差｜可迁移至行为经济学中的情绪与跨期选择实验，或金融中的投资者情绪与交易行为模拟｜以LLM智能体作为被试，施加情绪状态（如通过PAD向量操纵），观察其在跨期选择任务中的贴现率变化，并与真实人类实验数据（如Ifcher & Zarghamee, 2011）对照，检验情绪效应的仿真保真度。"}},{"id":"2601.17527","version":1,"title":"Bridging Expectation Signals: LLM-Based Experiments and a Behavioral Kalman Filter Framework","zh_title":"桥接预期信号：基于LLM的实验与行为卡尔曼滤波框架","abstract":"As LLMs increasingly function as economic agents, the specific mechanisms LLMs use to update their belief with heterogeneous signals remain opaque. We design experiments and develop a Behavioral Kalman Filter framework to quantify how LLM-based agents update expectations, acting as households or firm CEOs, update expectations when presented with individual and aggregate signals. The results from experiments and model estimation reveal four consistent patterns: (1) agents' weighting of priors and signals deviates from unity; (2) both household and firm CEO agents place substantially larger weights on individual signals compared to aggregate signals; (3) we identify a significant and negative interaction between concurrent signals, implying that the presence of multiple information sources diminishes the marginal weight assigned to each individual signal; and (4) expectation formation patterns differ significantly between household and firm CEO agents. Finally, we demonstrate that LoRA fine-tuning mitigates, but does not fully eliminate, behavioral biases in LLM expectation formation.","authors":["Yu Wang","Xiangchen Liu"],"categories":["econ.GN","cs.AI"],"primary_category":"econ.GN","announce_type":"new","date":"2026-01-24","first_seen":"2026-01-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2601.17527","pdf_url":"https://arxiv.org/pdf/2601.17527","source_feed":"backfill","score":8,"bucket":"selected","rubric_hits":["A3","B2","B4"],"tags":["LLM经济代理","预期形成","行为偏差"],"reason":"用LLM模拟家庭和CEO预期更新，涉及经济实验，有行为偏差分析，但未明确提及真…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:53","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":56,"question":"LLM代理在接收异质性信号时如何更新预期，其权重分配和行为偏差的机制是什么？","design":"使用GPT-4o、Gemini 1.5、DeepSeek-V3等LLM扮演家庭和CEO，在720次试验中呈现微观（个人收入/公司业绩）和宏观（GDP增长）两种信号，测量其对未来收入或利润增长的预期更新。","baseline":"无对照","findings":"LLM代理对微观信号的权重大于宏观信号，且信号间存在负交互效应，即多信号会削弱各自边际权重；CEO比家庭更重视宏观信号。LoRA微调可减轻但无法消除非理性偏差。","reliability":"论文未讨论","relevance":"该研究用LLM模拟经济主体预期更新，涉及行为偏差分析，但缺乏真实人类数据基准，适合关注仿真机制和偏差的研究者阅读原文以评估方法细节。","inspiration":"该方法通过向LLM代理呈现微观与宏观异质性信号并测量预期更新权重，可借鉴其多信号交互实验设计来分离信息处理偏差｜可迁移至政策公告的预期形成研究，如央行沟通中微观通胀感知与宏观通胀目标对家庭预期的交互影响｜以LLM模拟家庭，处理为同时呈现个人消费价格变化（微观）与官方CPI（宏观），结果变量为通胀预期更新幅度，对照真实家庭调查数据（如密歇根消费者调查）"}},{"id":"2604.06173","version":2,"title":"Beyond Case Law: Evaluating Structure-Aware Retrieval and Safety in Statute-Centric Legal QA","zh_title":"超越判例法：评估法规中心的法律问答中的结构感知检索与安全性","abstract":"Legal QA benchmarks have predominantly focused on case law, overlooking the unique challenges of statute-centric regulatory reasoning. In statutory domains, relevant evidence is distributed across hierarchically linked documents, creating a statutory retrieval gap where conventional retrievers fail and models often hallucinate under incomplete context. We introduce SearchFireSafety, a structure- and safety-aware benchmark for statute-centric legal QA. Instantiated on fire-safety regulations as a representative case, the benchmark evaluates whether models can retrieve hierarchically fragmented evidence and safely abstain when statutory context is insufficient. SearchFireSafety adopts a dual-source evaluation framework combining real-world questions that require citation-aware retrieval and synthetic partial-context scenarios that stress-test hallucination and refusal behavior. Experiments across multiple large language models show that graph-guided retrieval substantially improves performance, but also reveal a critical safety trade-off: domain-adapted models are more likely to hallucinate when key statutory evidence is missing. Our findings highlight the need for benchmarks that jointly evaluate hierarchical retrieval and model safety in statute-centric regulatory settings.","authors":["Kyubyung Chae","Jewon Yeom","Jeongjae Park","Seunghyun Bae","Ijun Jang","Hyunbin Jin","Jinkwan Jang","Taesup Kim"],"categories":["cs.IR","cs.AI"],"primary_category":"cs.IR","announce_type":"new","date":"2026-01-24","first_seen":"2026-01-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2604.06173","pdf_url":"https://arxiv.org/pdf/2604.06173","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["法律问答","检索增强","模型安全"],"reason":"法律问答基准测试，评估检索与安全，不涉及人类行为仿真或对照。","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:59:14","error":null,"has_summary":false,"summary":null},{"id":"2601.16355","version":2,"title":"Identity, Cooperation and Framing Effects within Groups of Real and Simulated Humans","zh_title":"真实与模拟人类群体中的身份、合作与框架效应","abstract":"Humans act via a nuanced process that depends both on rational deliberation and also on identity and contextual factors. In this work, we study how large language models (LLMs) can simulate human action in the context of social dilemma games. While prior work has focused on \"steering\" (weak binding) of chat models to simulate personas, we analyze here how deep binding of base models with extended backstories leads to more faithful replication of identity-based behaviors. Our study has these findings: simulation fidelity vs human studies is improved by conditioning base LMs with rich context of narrative identities and checking consistency using instruction-tuned models. We show that LLMs can also model contextual factors such as time (year that a study was performed), question framing, and participant pool effects. LLMs, therefore, allow us to explore the details that affect human studies but which are often omitted from experiment descriptions, and which hamper accurate replication.","authors":["Suhong Moon","Minwoo Kang","Joseph Suh","Mustafa Safdari","John Canny"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-01-22","first_seen":"2026-01-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2601.16355","pdf_url":"https://arxiv.org/pdf/2601.16355","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM人类仿真","社会困境博弈","行为实验复现"],"reason":"用LLM模拟社会困境中的人类行为，并与真实人类研究对照，涉及合作与框架效应。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:52","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":24,"question":"大语言模型能否通过深度绑定身份背景、时间锚定和一致性过滤，在独裁者博弈和信任博弈中复现真实人类的党派内群体偏袒行为？","design":"使用 Mistral-Small、Mixtral 8x22B 和 Qwen-2.5 72B 基础模型，通过深度绑定（DeepBind）方法为虚拟被试赋予详细的叙事身份背景，并施加一致性过滤（反复重申身份）和时间锚定（设定实验年份），在独裁者博弈和信任博弈中测量虚拟被试对同党和异党对象的资源分配与信任行为。","baseline":"对照的真实人类数据来自 Whitt et al. (2021) 和 Iyengar & Westwood (2015) 的独裁者博弈研究，以及 Carlin & Love (2018) 和 Whitt et al. (2021) 的信任博弈研究，均包含民主党与共和党被试的党派偏袒效应量。","findings":"DeepBind 方法在所有模型和博弈中均最一致地复现了人类党派偏袒差距，且同时使用时间锚定和一致性过滤能进一步提升仿真与人类基准的对齐程度。LLM 还能捕捉到实验年份、问题措辞和参与者池等常被忽略的细微情境效应。","reliability":"论文未讨论","relevance":"该研究直接用 LLM 复现社会困境博弈中的人类行为，并与多项真实人类实验进行定量对照，同时探讨了身份、时间框架等情境因素的仿真效果，高度契合对 LLM 人类仿真可靠性及偏差的关注，值得精读。","inspiration":"借鉴之处在于通过深度绑定叙事身份、时间锚定和一致性过滤来增强LLM的情境代入感，从而更精细地操控虚拟被试的社会身份与决策框架｜可迁移至经济金融中的群体间歧视行为研究，例如信贷审批中的党派或种族偏见、投资决策中的内群体偏袒｜设计上以LLM作为虚拟信贷员，通过DeepBind赋予其不同党派身份，并设定审批年份，测量其对同党与异党申请人的贷款批准率差异，以真实信贷歧视研究数据作为对照基准"}},{"id":"2601.15793","version":1,"title":"HumanLLM: Towards Personalized Understanding and Simulation of Human Nature","zh_title":"HumanLLM：迈向个性化理解与人性仿真","abstract":"Motivated by the remarkable progress of large language models (LLMs) in objective tasks like mathematics and coding, there is growing interest in their potential to simulate human behavior--a capability with profound implications for transforming social science research and customer-centric business insights. However, LLMs often lack a nuanced understanding of human cognition and behavior, limiting their effectiveness in social simulation and personalized applications. We posit that this limitation stems from a fundamental misalignment: standard LLM pretraining on vast, uncontextualized web data does not capture the continuous, situated context of an individual's decisions, thoughts, and behaviors over time. To bridge this gap, we introduce HumanLLM, a foundation model designed for personalized understanding and simulation of individuals. We first construct the Cognitive Genome Dataset, a large-scale corpus curated from real-world user data on platforms like Reddit, Twitter, Blogger, and Amazon. Through a rigorous, multi-stage pipeline involving data filtering, synthesis, and quality control, we automatically extract over 5.5 million user logs to distill rich profiles, behaviors, and thinking patterns. We then formulate diverse learning tasks and perform supervised fine-tuning to empower the model to predict a wide range of individualized human behaviors, thoughts, and experiences. Comprehensive evaluations demonstrate that HumanLLM achieves superior performance in predicting user actions and inner thoughts, more accurately mimics user writing styles and preferences, and generates more authentic user profiles compared to base models. Furthermore, HumanLLM shows significant gains on out-of-domain social intelligence benchmarks, indicating enhanced generalization.","authors":["Yuxuan Lei","Tianfu Wang","Jianxun Lian","Zhengyu Hu","Defu Lian","Xing Xie"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-01-22","first_seen":"2026-01-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2601.15793","pdf_url":"https://arxiv.org/pdf/2601.15793","source_feed":"backfill","score":8,"bucket":"selected","rubric_hits":["A1","A2","B1"],"tags":["人类行为仿真","个性化建模","社会模拟"],"reason":"用LLM仿真个体行为与思维，有真实用户数据对照，评估预测准确性，方法可迁移至人…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:50","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":38,"question":"能否构建一个通用基础模型，利用大规模真实用户数据实现个性化的人类认知与行为理解和仿真？","design":"基于Reddit、Twitter、Blogger、Amazon等平台的真实用户数据构建认知基因组数据集，通过数据过滤、合成和质量控制的多阶段流水线提取超过550万条用户日志，设计个人资料生成、社会问答和写作模仿等任务，对基础LLM进行监督微调，并使用模型合并策略保留通用能力。","baseline":"对照的真实人类数据为来自Reddit、Twitter、Blogger、Amazon等平台的真实用户日志和行为记录。","findings":"HumanLLM在预测用户行为、内心想法和风格模仿上显著优于基础模型；在MotiveBench和TomBench等域外社会智能基准上表现出增强的泛化能力。","reliability":"论文未讨论","relevance":"该研究直接利用真实人类数据训练LLM以仿真个体行为与思维，并评估预测准确性，方法可迁移至经济学实验和政策评估场景，值得精读原文以了解其仿真可靠性与局限。","inspiration":"该方法利用多平台真实用户日志构建个性化仿真数据集，并通过监督微调使LLM模仿个体行为与思维，可借鉴其多阶段数据过滤与任务设计思路来构建经济学仿真被试｜可迁移到消费者跨期选择与政策偏好预测场景，例如模拟个体在不同利率或补贴政策下的储蓄消费决策｜研究设计：以真实用户的消费与问卷数据为基准，用微调后的LLM作为被试，施加利率变动或补贴政策处理，测量其消费-储蓄分配与政策支持度，并与真实面板数据对照"}},{"id":"2601.15556","version":1,"title":"LLM or Human? Perceptions of Trust and Information Quality in Research Summaries","zh_title":"LLM还是人类？研究摘要中信任与信息质量的感知","abstract":"Large Language Models (LLMs) are increasingly used to generate and edit scientific abstracts, yet their integration into academic writing raises questions about trust, quality, and disclosure. Despite growing adoption, little is known about how readers perceive LLM-generated summaries and how these perceptions influence evaluations of scientific work. This paper presents a mixed-methods survey experiment investigating whether readers with ML expertise can distinguish between human- and LLM-generated abstracts, how actual and perceived LLM involvement affects judgments of quality and trustworthiness, and what orientations readers adopt toward AI-assisted writing. Our findings show that participants struggle to reliably identify LLM-generated content, yet their beliefs about LLM involvement significantly shape their evaluations. Notably, abstracts edited by LLMs are rated more favorably than those written solely by humans or LLMs. We also identify three distinct reader orientations toward LLM-assisted writing, offering insights into evolving norms and informing policy around disclosure and acceptable use in scientific communication.","authors":["Nil-Jana Akpinar","Sandeep Avula","CJ Lee","Brandon Dang","Kaza Razat","Vanessa Murdock"],"categories":["cs.CY","cs.CL"],"primary_category":"cs.CY","announce_type":"new","date":"2026-01-22","first_seen":"2026-01-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2601.15556","pdf_url":"https://arxiv.org/pdf/2601.15556","source_feed":"backfill","score":7,"bucket":"pending","rubric_hits":["A1","B1"],"tags":["LLM生成摘要","人类感知调查","信任与质量评估"],"reason":"用LLM生成摘要并调查人类感知，有真实人类数据对照，但非直接仿真人类被试行为或…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:46","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":103,"question":"读者能否区分LLM生成与人类撰写的研究摘要？LLM的实际参与和读者感知到的LLM参与如何影响对摘要质量和可信度的评价？","design":"本研究并非用LLM仿真人类被试，而是通过在线调查实验，让具有ML专业知识的参与者评估三种不同作者身份的摘要（人类撰写、LLM生成、人类撰写后LLM编辑），测量其辨别能力、质量与可信度评分，并分析读者对AI辅助写作的态度取向。","baseline":"对照的真实人类数据为原始人类撰写的摘要版本，以及参与者对摘要作者身份的判断和评分。","findings":"参与者无法可靠区分LLM生成与人类撰写的摘要，但其对LLM参与的信念显著影响评价；LLM编辑的摘要（人类撰写后LLM修订）在清晰度、简洁性和可信度上评分最高，而完全由LLM生成的摘要评分最低。","reliability":"论文未讨论","relevance":"本研究以真实人类数据为基准，调查LLM生成内容对读者感知的影响，虽非直接仿真人类行为，但涉及LLM输出与人类判断的对照，对关注LLM仿真可靠性及偏差的研究者有参考价值，建议阅读原文了解读者评价中的启发式偏差。","inspiration":"该研究通过设置三种作者身份（人类撰写、LLM生成、人类撰写后LLM编辑）作为处理条件，并测量读者对摘要质量与可信度的评分及辨别能力，提供了多臂对照和感知偏差测量的设计范式｜可迁移至政策公告的预期形成研究，例如考察市场参与者对央行声明或政府经济预测的解读如何受AI辅助写作标签影响｜以金融从业者为被试，随机分配阅读标注为‘人类撰写’、‘AI生成’或‘人类撰写后AI润色’的政策声明，测量其对声明可信度、清晰度及未来通胀/利率预期的评分，并以真实历史声明和专家原始判断作为基准对照"}},{"id":"2601.15114","version":2,"title":"From Who They Are to How They Act: Behavioral Traits in Generative Agent-Based Models of Social Media","zh_title":"从他们是谁到他们如何行动：基于生成式智能体的社交媒体模型中的行为特质","abstract":"Generative Agent-Based Modeling (GABM) leverages Large Language Models to create autonomous agents that simulate human behavior in social media environments, demonstrating potential for modeling information propagation, influence processes, and network phenomena. While existing frameworks characterize agents through demographic attributes, personality traits, and interests, they lack mechanisms to encode behavioral dispositions toward platform actions, causing agents to exhibit homogeneous engagement patterns rather than the differentiated participation styles observed on real platforms. In this paper, we investigate the role of behavioral traits as an explicit characterization layer to regulate agents' propensities across posting, re-sharing, commenting, reacting, and inactivity. Through large-scale simulations involving 980 agents and validation against real-world social media data, we demonstrate that behavioral traits are essential to sustain heterogeneous, profile-consistent participation patterns and enable realistic content propagation dynamics through the interplay of amplification- and interaction-oriented profiles. Our findings establish that modeling how agents act-not only who they are-is necessary for advancing GABM as a tool for studying social media phenomena.","authors":["Valerio La Gatta","Gian Marco Orlando","Marco Perillo","Ferdinando Tammaro","Vincenzo Moscato"],"categories":["cs.MA"],"primary_category":"cs.MA","announce_type":"new","date":"2026-01-21","first_seen":"2026-01-21","revised_at":null,"abs_url":"https://arxiv.org/abs/2601.15114","pdf_url":"https://arxiv.org/pdf/2601.15114","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM仿真","社交媒体模拟","行为特质"],"reason":"用LLM agent模拟社交媒体行为，并与真实数据对照，直接复现人类参与模式。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:50","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":25,"question":"行为特质作为显式表征层，能否让生成式智能体在社交媒体模拟中维持异质性参与模式、再现真实的传播动态和网络结构？","design":"在现有GABM框架上扩展，为980个LLM智能体同时赋予身份特质（来自FinePersonas数据集）和七种行为特质原型（如沉默观察者、内容放大器等），模拟发帖、转发、评论、反应和沉默等完整动作空间，测量参与模式异质性、内容传播级联和网络中心性。","baseline":"对照真实社交媒体数据，验证行为特质能否复现真实世界的网络结构。","findings":"行为特质能有效防止行为同质化，维持与原型一致的异质性参与模式；放大导向型特质驱动内容传播级联，互动导向型特质主导互动网络，且基于经验数据的行为特质能成功复现真实社交网络结构。","reliability":"论文未讨论","relevance":"该研究直接用LLM智能体复现社交媒体用户行为，并与真实数据对照，验证了行为特质对异质性参与和传播动态的必要性，直接回应了研究者对LLM仿真可靠性及失效条件的关注，值得精读。","inspiration":"该方法通过为LLM智能体显式赋予行为特质原型来维持异质性参与模式，可借鉴到经济仿真中作为施加个体差异的处理手段｜可迁移到政策公告的预期形成与信息传播研究，模拟不同投资者类型对政策信息的反应与扩散｜以LLM智能体为被试，赋予理性交易者、噪声交易者等行为特质，处理为发布货币政策公告，结果变量为资产价格波动与信息传播级联，对照真实市场微观数据"}},{"id":"2601.12727","version":1,"title":"AI-exhibited Personality Traits Can Shape Human Self-concept through Conversations","zh_title":"AI展现的人格特质可通过对话塑造人类自我概念","abstract":"Recent Large Language Model (LLM) based AI can exhibit recognizable and measurable personality traits during conversations to improve user experience. However, as human understandings of their personality traits can be affected by their interaction partners' traits, a potential risk is that AI traits may shape and bias users' self-concept of their own traits. To explore the possibility, we conducted a randomized behavioral experiment. Our results indicate that after conversations about personal topics with an LLM-based AI chatbot using GPT-4o default personality traits, users' self-concepts aligned with the AI's measured personality traits. The longer the conversation, the greater the alignment. This alignment led to increased homogeneity in self-concepts among users. We also observed that the degree of self-concept alignment was positively associated with users' conversation enjoyment. Our findings uncover how AI personality traits can shape users' self-concepts through human-AI conversation, highlighting both risks and opportunities. We provide important design implications for developing more responsible and ethical AI systems.","authors":["Jingshu Li","Tianqi Song","Nattapat Boonprakong","Zicheng Zhu","Yitian Yang","Yi-Chieh Lee"],"categories":["cs.HC","cs.AI"],"primary_category":"cs.HC","announce_type":"new","date":"2026-01-19","first_seen":"2026-01-19","revised_at":null,"abs_url":"https://arxiv.org/abs/2601.12727","pdf_url":"https://arxiv.org/pdf/2601.12727","source_feed":"backfill","score":8,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","人格影响","人机交互实验"],"reason":"用LLM对话实验研究AI人格对用户自我概念的影响，有随机对照实验和人类数据，揭…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:45","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":57,"question":"与具有人格特质的大语言模型AI聊天机器人进行个人话题对话，是否会使用户的自我概念向AI的人格特质对齐？","design":"本研究不是用LLM替代人类被试的仿真研究，而是以人类为被试的在线随机行为实验。采用混合因子设计，让参与者与基于GPT-4o默认人格的AI聊天机器人进行个人话题对话，测量对话前后用户自我概念的变化，以及对话时长、享受度等变量。","baseline":"无对照","findings":"与AI进行个人话题对话后，用户的自我概念会向AI所表现出的人格特质对齐，且对话时间越长，对齐程度越大。这种对齐导致用户间自我概念的同质化增强，且对齐程度与用户的对话享受度正相关。","reliability":"论文未讨论","relevance":"该研究直接探讨LLM对话对人类心理的影响，虽非用LLM仿真人类，但揭示了人机交互中AI人格对用户自我概念的塑造效应，对评估AI社会影响和仿真可靠性有批判性参考价值，值得阅读原文以了解实验细节和效应边界。","inspiration":"该研究采用混合因子设计，通过对比对话前后自我概念的变化来测量AI人格的塑造效应，这种前后测设计可用于评估经济决策中的干预效果。｜可迁移到消费者金融决策场景，例如研究AI理财顾问的人格特质是否影响用户的投资偏好或风险态度。｜以真实投资者为被试，随机分配与具有不同人格（如谨慎型vs冒险型）的AI理财顾问对话，测量对话前后风险偏好问卷得分的变化，并与历史投资行为数据对照。"}},{"id":"2601.12343","version":1,"title":"How Well Do LLMs Predict Human Behavior? A Measure of their Pretrained Knowledge","zh_title":"LLM预测人类行为的效果如何？一种对其预训练知识的度量","abstract":"Large language models (LLMs) are increasingly used to predict human behavior. We propose a measure for evaluating how much knowledge a pretrained LLM brings to such a prediction: its equivalent sample size, defined as the amount of task-specific data needed to match the predictive accuracy of the LLM. We estimate this measure by comparing the prediction error of a fixed LLM in a given domain to that of flexible machine learning models trained on increasing samples of domain-specific data. We further provide a statistical inference procedure by developing a new asymptotic theory for cross-validated prediction error. Finally, we apply this method to the Panel Study of Income Dynamics. We find that LLMs encode considerable predictive information for some economic variables but much less for others, suggesting that their value as substitutes for domain-specific data differs markedly across settings.","authors":["Wayne Gao","Sukjin Han","Annie Liang"],"categories":["econ.EM","cs.AI","stat.ML"],"primary_category":"econ.EM","announce_type":"new","date":"2026-01-18","first_seen":"2026-01-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2601.12343","pdf_url":"https://arxiv.org/pdf/2601.12343","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2"],"tags":["LLM仿真","人类行为预测","等效样本量"],"reason":"直接评估LLM预测人类行为的能力，使用真实经济调查数据作为基准，并提出等效样本…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:48","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":9,"question":"如何量化预训练大语言模型在预测人类行为时所带来的领域特定知识的价值？","design":"本研究不是仿真实验，而是提出一种评估方法：将固定预训练LLM的预测误差，与在逐渐增大的领域特定数据上训练的灵活机器学习模型的误差进行比较，定义等效样本量为后者误差首次不劣于LLM时的训练样本量。","baseline":"使用收入动态面板研究（PSID）2021年波次的真实调查数据，包含人口统计、劳动力市场和家庭协变量，用于预测小时工资、住房拥有、饮酒和吸烟等结果。","findings":"LLM对不同经济变量的预测能力差异很大：预测小时工资的等效样本量约为20个观测值，而预测住房拥有则需约600个观测值。这表明LLM作为领域特定数据替代品的价值在不同任务中显著不同。","reliability":"论文关注数据泄露问题，采用训练截止日期明确的静态开源模型进行事后评估；方法假设比较算法的误差随训练数据量增加而单调递减；等效样本量的推断依赖于交叉验证风险估计的渐近正态性。","relevance":"该研究直接回应了研究者对LLM替代人类被试的可靠性与偏差的关切，提供了基于真实经济调查数据的量化基准，并揭示了仿真在不同任务中有效性的异质性，值得精读原文以掌握等效样本量估计方法及其适用边界。","inspiration":"该方法通过比较LLM与灵活机器学习模型的预测误差，定义等效样本量来量化LLM的领域知识价值，为评估LLM仿真可靠性提供了可操作的基准｜可迁移到消费者金融行为预测场景，如利用LLM预测家庭信贷违约或投资决策，评估其替代传统调查数据的可行性｜以LLM作为被试，输入PSID等调查中的家庭财务协变量，预测信贷违约状态，将LLM预测误差与在真实违约数据上训练的梯度提升模型比较，计算等效样本量，以真实信贷记录作为对照基准"}},{"id":"2601.12339","version":1,"title":"The Economics of Digital Intelligence Capital: Endogenous Depreciation and the Structural Jevons Paradox","zh_title":"数字智能资本的经济学：内生折旧与结构性杰文斯悖论","abstract":"This paper develops a micro-founded economic theory of the AI industry by modeling large language models as a distinct asset class-Digital Intelligence Capital-characterized by data-compute complementarities, increasing returns to scale, and relative (rather than absolute) valuation. We show that these features fundamentally reshape industry dynamics along three dimensions. First, because downstream demand depends on relative capability, innovation by one firm endogenously depreciates the economic value of rivals' existing capital, generating a persistent innovation pressure we term the Red Queen Effect. Second, falling inference prices induce downstream firms to adopt more compute-intensive agent architectures, rendering aggregate demand for compute super-elastic and producing a structural Jevons paradox. Third, learning from user feedback creates a data flywheel that can destabilize symmetric competition: when data accumulation outpaces data decay, the market bifurcates endogenously toward a winner-takes-all equilibrium. We further characterize conditions under which expanding upstream capabilities erode downstream application value (the Wrapper Trap). A calibrated agent-based model confirms these mechanisms and their quantitative implications. Together, the results provide a unified framework linking intelligence production upstream with agentic demand downstream, offering new insights into competition, scalability, and regulation in the AI economy.","authors":["Yukun Zhang","Tianyang Zhang"],"categories":["econ.GN"],"primary_category":"econ.GN","announce_type":"new","date":"2026-01-18","first_seen":"2026-01-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2601.12339","pdf_url":"https://arxiv.org/pdf/2601.12339","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["AI经济学","多智能体模拟","产业动态"],"reason":"纯多智能体经济模拟，无人类行为对照，属C1排除项。","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:58:23","error":null,"has_summary":false,"summary":null},{"id":"2601.12471","version":2,"title":"Knowing When to Abstain: Medical LLMs Under Clinical Uncertainty","zh_title":"知止：临床不确定性下的医学大语言模型","abstract":"Current evaluation of large language models (LLMs) overwhelmingly prioritizes accuracy; however, in real-world and safety-critical applications, the ability to abstain when uncertain is equally vital for trustworthy deployment. We introduce MedAbstain, a unified benchmark and evaluation protocol for abstention in medical multiple-choice question answering (MCQA) -- a discrete-choice setting that generalizes to agentic action selection -- integrating conformal prediction, adversarial question perturbations, and explicit abstention options. Our systematic evaluation of both open- and closed-source LLMs reveals that even state-of-the-art, high-accuracy models often fail to abstain with uncertain. Notably, providing explicit abstention options consistently increases model uncertainty and safer abstention, far more than input perturbations, while scaling model size or advanced prompting brings little improvement. These findings highlight the central role of abstention mechanisms for trustworthy LLM deployment and offer practical guidance for improving safety in high-stakes applications.","authors":["Sravanthi Machcha","Sushrita Yerra","Sahil Gupta","Aishwarya Sahoo","Sharmin Sultana","Hong Yu","Zonghai Yao"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2026-01-18","first_seen":"2026-01-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2601.12471","pdf_url":"https://arxiv.org/pdf/2601.12471","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["医学问答","模型弃权","安全评测"],"reason":"纯医学问答评测，无人类行为仿真或对照，属NLP能力评测。","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:59:37","error":null,"has_summary":false,"summary":null},{"id":"2601.11049","version":2,"title":"Predicting Biased Human Decision-Making with Large Language Models in Conversational Settings","zh_title":"用大语言模型预测对话场景中的人类有偏决策","abstract":"We examine whether large language models (LLMs) can predict biased decision-making in conversational settings, and whether their predictions capture not only human cognitive biases but also how those effects change under cognitive load. In a pre-registered study (N = 1,648), participants completed six classic decision-making tasks via a chatbot with dialogues of varying complexity. Participants exhibited two well-documented cognitive biases: the Framing Effect and the Status Quo Bias. Increased dialogue complexity resulted in participants reporting higher mental demand. This increase in cognitive load selectively, but significantly, increased the effect of the biases, demonstrating the load-bias interaction. We then evaluated whether LLMs (GPT-4, GPT-5, and open-source models) could predict individual decisions given demographic information and prior dialogue. While results were mixed across choice problems, LLM predictions that incorporated dialogue context were significantly more accurate in several key scenarios. Importantly, their predictions reproduced the same bias patterns and load-bias interactions observed in humans. Across all models tested, the GPT-4 family consistently aligned with human behavior, outperforming GPT-5 and open-source models in both predictive accuracy and fidelity to human-like bias patterns. These findings advance our understanding of LLMs as tools for simulating human decision-making and inform the design of conversational agents that adapt to user biases.","authors":["Stephen Pilli","Vivek Nallur"],"categories":["cs.HC","cs.AI"],"primary_category":"cs.HC","announce_type":"new","date":"2026-01-16","first_seen":"2026-01-16","revised_at":null,"abs_url":"https://arxiv.org/abs/2601.11049","pdf_url":"https://arxiv.org/pdf/2601.11049","source_feed":"backfill","score":10,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["LLM仿真","认知偏差","人类数据对照"],"reason":"用LLM预测人类决策偏差，有真实人类数据对照，涉及认知负荷与偏差交互，评估仿真…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:48","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":5,"question":"LLM能否在对话环境中预测人类的偏差决策，并复现认知偏差及其与认知负荷的交互效应？","design":"使用GPT-4、GPT-5和开源模型，基于人口统计信息和对话上下文预测个体在六项经典决策任务中的选择；处理变量为对话复杂度（低/高），结果变量为预测准确率及偏差模式复现程度。","baseline":"预注册实验（N=1648）中人类被试通过聊天机器人完成决策任务，表现出框架效应和现状偏差，且对话复杂度增加导致认知负荷上升并选择性放大偏差。","findings":"LLM预测在纳入对话上下文后准确率显著提升，且GPT-4家族在预测准确性和偏差模式复现上均优于GPT-5和开源模型；LLM预测成功复现了人类中的偏差主效应和负荷-偏差交互效应。","reliability":"论文指出不同选择问题的预测结果参差不齐，且未系统探讨模型在极端认知负荷或非典型人口群体下的失效条件。","relevance":"该研究直接以真实人类数据为基准，检验LLM在对话决策场景中预测偏差行为的能力，并揭示了模型间差异，对评估LLM作为人类被试替代品的可靠性具有参考价值，值得阅读原文。","inspiration":"该方法通过操纵对话复杂度来施加认知负荷处理，并以真实人类实验数据为基准检验LLM的预测效度，值得借鉴｜可迁移到消费者在复杂金融产品选择中的决策偏差研究，如贷款方案或保险产品的框架效应｜以LLM模拟消费者，处理为产品信息呈现的对话复杂度（低/高），结果变量为选择偏差（如框架效应），对照真实消费者实验数据"}},{"id":"2601.15319","version":1,"title":"Large Language Models as Simulative Agents for Neurodivergent Adult Psychometric Profiles","zh_title":"大语言模型作为神经多样性成人心理测量特征的仿真代理","abstract":"Adult neurodivergence, including Attention-Deficit/Hyperactivity Disorder (ADHD), high-functioning Autism Spectrum Disorder (ASD), and Cognitive Disengagement Syndrome (CDS), is marked by substantial symptom overlap that limits the discriminant sensitivity of standard psychometric instruments. While recent work suggests that Large Language Models (LLMs) can simulate human psychometric responses from qualitative data, it remains unclear whether they can accurately and stably model neurodevelopmental traits rather than broad personality characteristics. This study examines whether LLMs can generate psychometric responses that approximate those of real individuals when grounded in a structured qualitative interview, and whether such simulations are sensitive to variations in trait intensity. Twenty-six adults completed a 29-item open-ended interview and four standardized self-report measures (ASRS, BAARS-IV, AQ, RAADS-R). Two LLMs (GPT-4o and Qwen3-235B-A22B) were prompted to infer an individual psychological profile from interview content and then respond to each questionnaire in-role. Accuracy, reliability, and sensitivity were assessed using group-level comparisons, error metrics, exact-match scoring, and a randomized baseline. Both models outperformed random responses across instruments, with GPT-4o showing higher accuracy and reproducibility. Simulated responses closely matched human data for ASRS, BAARS-IV, and RAADS-R, while the AQ revealed subscale-specific limitations, particularly in Attention to Detail. Overall, the findings indicate that interview-grounded LLMs can produce coherent and above-chance simulations of neurodevelopmental traits, supporting their potential use as synthetic participants in early-stage psychometric research, while highlighting clear domain-specific constraints.","authors":["Francesco Chiappone","Davide Marocco","Nicola Milano"],"categories":["q-bio.NC","cs.AI"],"primary_category":"q-bio.NC","announce_type":"new","date":"2026-01-16","first_seen":"2026-01-16","revised_at":null,"abs_url":"https://arxiv.org/abs/2601.15319","pdf_url":"https://arxiv.org/pdf/2601.15319","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真人类被试","心理测量","真实人类数据对照"],"reason":"用LLM仿真神经发育特质问卷回答，并与26名真人数据对照，评估准确性与局限性。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:50","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":10,"question":"LLM能否基于结构化访谈内容，准确且稳定地模拟神经发育特质（ADHD、ASD、CDS）个体的心理测量反应？","design":"用GPT-4o和Qwen3-235B-A22B两个LLM，根据26名成人被试的29项开放式访谈文本推断个体心理画像，然后以角色扮演方式回答ASRS、BAARS-IV、AQ、RAADS-R四份标准化自评问卷，评估模拟回答的准确性、可靠性和对特质强度的敏感性。","baseline":"26名成人被试的真实访谈内容和四份问卷（ASRS、BAARS-IV、AQ、RAADS-R）的自评数据。","findings":"两个模型在所有问卷上的模拟回答均优于随机基线，GPT-4o准确性和可重复性更高；模拟回答在ASRS、BAARS-IV和RAADS-R上接近真人数据，但AQ量表在“注意细节”子量表上表现出局限性。","reliability":"论文承认AQ量表在特定子量表（如注意细节）上模拟准确性有限，且样本量较小（26人），可能影响结论的泛化性。","relevance":"该研究直接以真实人类数据为基准，评估LLM仿真神经发育特质问卷回答的准确性与局限性，符合研究者对仿真可靠性及失效条件的关注，值得阅读原文以了解具体偏差来源和实验设计细节。","inspiration":"借鉴其“基于结构化访谈生成个体画像并角色扮演回答问卷”的仿真设计，可迁移到经济金融领域的消费者偏好测量或投资者情绪评估场景。｜用LLM基于消费者深度访谈文本模拟其跨期选择问卷回答，以真实消费者面板数据为对照，评估LLM仿真消费决策偏差的准确性。"}},{"id":"2601.15312","version":1,"title":"Do people expect different behavior from large language models acting on their behalf? Evidence from norm elicitations in two canonical economic games","zh_title":"人们是否期望代表他们行事的大语言模型表现出不同行为？来自两个经典经济博弈中规范引出的证据","abstract":"While delegating tasks to large language models (LLMs) can save people time, there is growing evidence that offloading tasks to such models produces social costs. We use behavior in two canonical economic games to study whether people have different expectations when decisions are made by LLMs acting on their behalf instead of themselves. More specifically, we study the social appropriateness of a spectrum of possible behaviors: when LLMs divide resources on our behalf (Dictator Game and Ultimatum Game) and when they monitor the fairness of splits of resources (Ultimatum Game). We use the Krupka-Weber norm elicitation task to detect shifts in social appropriateness ratings. Results of two pre-registered and incentivized experimental studies using representative samples from the UK and US (N = 2,658) show three key findings. First, people find that offers from machines - when no acceptance is necessary - are judged to be less appropriate than when they come from humans, although there is no shift in the modal response. Second - when acceptance is necessary - it is more appropriate for a person to reject offers from machines than from humans. Third, receiving a rejection of an offer from a machine is no less socially appropriate than receiving the same rejection from a human. Overall, these results suggest that people apply different norms for machines deciding on how to split resources but are not opposed to machines enforcing the norms. The findings are consistent with offers made by machines now being viewed as having both a cognitive and emotional component.","authors":["Paweł Niszczota","Elia Antoniou"],"categories":["cs.GT","cs.AI","cs.CL","cs.CY","cs.HC","econ.GN"],"primary_category":"cs.GT","announce_type":"new","date":"2026-01-14","first_seen":"2026-01-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2601.15312","pdf_url":"https://arxiv.org/pdf/2601.15312","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["LLM仿真","经济博弈","社会规范"],"reason":"用LLM替代人类被试进行经济博弈实验，有真实人类数据对照，评估社会规范变化，直…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:50","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":20,"question":"当大语言模型（LLM）代替人类做出资源分配决策或监督公平时，人们对这些行为的社会适当性评价是否会发生变化？","design":"本研究并非用LLM作为人类被试的替代品进行仿真，而是通过在线实验调查人类被试对LLM代理行为的规范性评价。实验采用Krupka-Weber规范引出任务，让来自英国和美国的代表性样本（N=2658）对独裁者博弈和最后通牒博弈中一系列可能行为的社会适当性进行评分，比较决策者是人类还是LLM时的评分差异。","baseline":"无对照。研究直接比较人类被试对同一行为在“人类决策”与“LLM决策”两种情境下的适当性评分，未使用LLM生成行为并与真实人类行为数据对照。","findings":"第一，当无需接受方同意时（独裁者博弈），机器做出的分配提议被认为比人类做出的更不适当，但众数反应未变；第二，当需要接受方同意时（最后通牒博弈），拒绝机器提议比拒绝人类提议更适当；第三，收到机器的拒绝与收到人类的拒绝在社会适当性上无差异。","reliability":"论文未讨论","relevance":"该研究直接考察了人们对LLM代理经济决策的社会规范评价，虽非典型的LLM仿真实验，但为理解人机互动中的规范偏移提供了实证证据，对关注LLM替代人类被试时可能产生的规范性偏差的研究者具有参考价值。","inspiration":"借鉴其使用规范引出任务测量社会适当性评价的方法，可迁移至经济金融领域的伦理判断研究，如算法信贷审批或AI投资顾问的公众接受度。｜可应用于金融决策中的AI代理伦理：例如研究投资者对AI理财顾问做出高风险投资建议的适当性评价。｜设计：以普通投资者为被试，处理为投资建议来源（人类顾问 vs. AI顾问），结果变量为对建议适当性的评分，对照真实市场中人类顾问的建议接受率数据。"}},{"id":"2601.09772","version":1,"title":"Antisocial behavior towards large language model users: experimental evidence","zh_title":"针对大语言模型用户的反社会行为：实验证据","abstract":"The rapid spread of large language models (LLMs) has raised concerns about the social reactions they provoke. Prior research documents negative attitudes toward AI users, but it remains unclear whether such disapproval translates into costly action. We address this question in a two-phase online experiment (N = 491 Phase II participants; Phase I provided targets) where participants could spend part of their own endowment to reduce the earnings of peers who had previously completed a real-effort task with or without LLM support. On average, participants destroyed 36% of the earnings of those who relied exclusively on the model, with punishment increasing monotonically with actual LLM use. Disclosure about LLM use created a credibility gap: self-reported null use was punished more harshly than actual null use, suggesting that declarations of \"no use\" are treated with suspicion. Conversely, at high levels of use, actual reliance on the model was punished more strongly than self-reported reliance. Taken together, these findings provide the first behavioral evidence that the efficiency gains of LLMs come at the cost of social sanctions.","authors":["Paweł Niszczota","Cassandra Grützner"],"categories":["cs.AI","cs.CL","cs.CY","econ.GN"],"primary_category":"cs.AI","announce_type":"new","date":"2026-01-14","first_seen":"2026-01-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2601.09772","pdf_url":"https://arxiv.org/pdf/2601.09772","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["LLM仿真","行为经济学","社会惩罚"],"reason":"用LLM替代人类被试，在真实努力任务中测量对LLM用户的惩罚行为，有真实人类数…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:48","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":21,"question":"人们是否会因同伴使用大语言模型（LLM）完成任务而对其进行有代价的惩罚？","design":"本研究并非用LLM替代人类被试的仿真实验，而是以人类为被试的真实行为实验。实验分两阶段：第一阶段被试在有无LLM支持下完成真实努力任务，作为第二阶段的目标对象；第二阶段被试（N=491）可使用自己的部分报酬去减少第一阶段同伴的收入，结果变量为惩罚金额（即烧毁的金钱数量）。","baseline":"有真实人类数据作为对照：第一阶段被试的实际LLM使用情况（实际使用量）与自我报告的LLM使用情况，作为比较惩罚行为的基准。","findings":"被试平均烧毁了完全依赖LLM者36%的收入，且惩罚随实际LLM使用量单调递增。自我报告的零使用比实际零使用受到更严厉的惩罚，而高使用量下实际依赖比自我报告依赖受到更强惩罚，表明披露存在信任差距。","reliability":"论文未讨论","relevance":"该研究直接测量了对LLM使用者的真实惩罚行为，提供了有代价的反社会行为证据，与关注LLM仿真人类行为及社会规范的研究高度相关，值得精读原文以了解实验设计和行为测量方法。","inspiration":"采用真实努力任务与金钱燃烧博弈结合的设计，巧妙分离了实际使用与自我报告使用对惩罚的影响，并利用单调性检验强化因果推断。｜可迁移到信贷审批或招聘场景中，研究对算法辅助决策者的社会惩罚，例如贷款审批员使用AI模型是否会引发同事或客户的惩罚性行为。｜以金融从业者为被试，设计一个模拟贷款审批任务，处理组被告知审批员使用了AI辅助，对照组为纯人工审批，结果变量为被试愿意花费自身报酬去降低审批员奖金的行为，并以实际审批准确率数据作为基准对照。"}},{"id":"2601.09849","version":1,"title":"Strategies of cooperation and defection in five large language models","zh_title":"五种大语言模型中的合作与背叛策略","abstract":"Large language models (LLMs) are increasingly deployed to support human decision-making. This use of LLMs has concerning implications, especially when their prescriptions affect the welfare of others. To gauge how LLMs make social decisions, we explore whether five leading models produce sensible strategies in the repeated prisoner's dilemma, which is the main metaphor of reciprocal cooperation. First, we measure the propensity of LLMs to cooperate in a neutral setting, without using language reminiscent of how this game is usually presented. We record to what extent LLMs implement Nash equilibria or other well-known strategy classes. Thereafter, we explore how LLMs adapt their strategies to changes in parameter values. We vary the game's continuation probability, the payoff values, and whether the total number of rounds is commonly known. We also study the effect of different framings. In each case, we test whether the adaptations of the LLMs are in line with basic intuition, theoretical predictions of evolutionary game theory, and experimental evidence from human participants. While all LLMs perform well in many of the tasks, none of them exhibit full consistency over all tasks. We also conduct tournaments between the inferred LLM strategies and study direct interaction between LLMs in games over ten rounds with a known or unknown last round. Our experiments shed light on how current LLMs instantiate reciprocal cooperation.","authors":["Saptarshi Pal","Abhishek Mallela","Christian Hilbe","Lenz Pracher","Chiyu Wei","Feng Fu","Santiago Schnell","Martin A Nowak"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2026-01-14","first_seen":"2026-01-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2601.09849","pdf_url":"https://arxiv.org/pdf/2601.09849","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2","B3"],"tags":["LLM仿真","行为博弈","人类数据对照"],"reason":"用LLM模拟重复囚徒困境中的人类合作策略，并与人类实验数据对照，直接命中核心判…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:48","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":22,"question":"当前主流大语言模型在重复囚徒困境中能否生成符合直觉、演化博弈理论预测和人类实验证据的互惠合作策略？","design":"以五款主流LLM（Claude、Gemini、GPT-4o、GPT-5、Llama）为被试，通过直接询问其对上一轮（或两轮）所有可能结果的反应来推断其策略，而非让LLM相互对弈；施加的处理包括改变博弈的继续概率、收益值、记忆轮数、是否已知最后一轮以及不同的框架提示，测量LLM的合作倾向、策略类型（如纳什均衡、伙伴/对手类别）及策略对参数变化的适应性。","baseline":"人类基准：将LLM的策略适应性变化与演化博弈论的理论预测及人类参与者的实验证据进行对比。","findings":"所有LLM在许多任务中表现良好，但无一在所有任务上表现出完全一致性；LLM的策略在部分参数变化下能做出合理调整，但整体缺乏稳健的互惠合作逻辑。","reliability":"论文指出，LLM在改变博弈参数和框架时策略适应性不一致，且当前实验仅限于特定模型和参数设置，未全面覆盖所有可能的策略空间和交互动态。","relevance":"该研究直接以LLM替代人类被试进行重复囚徒困境实验，并与人类实验数据和演化博弈理论对照，高度契合您对LLM仿真人类决策可靠性及偏差的关注，值得精读。","inspiration":"借鉴直接推断策略而非仅观察对弈结果的方法，可更精确刻画LLM的决策规则并与理论解对照｜可迁移至经济金融中的重复信任博弈或重复公共品博弈，如投资者在重复投资决策中的合作与背叛行为｜以LLM为被试，设计不同收益结构和终止概率的重复投资游戏，测量其策略类型与适应性，并与人类实验数据（如Fischbacher & Gächter, 2010）进行对照。"}},{"id":"2601.09119","version":1,"title":"Contrastive Bi-Encoder Models for Multi-Label Skill Extraction: Enhancing ESCO Ontology Matching with BERT and Attention Mechanisms","zh_title":"用于多标签技能抽取的对比双编码器模型：利用BERT和注意力机制增强ESCO本体匹配","abstract":"Fine-grained labor market analysis increasingly relies on mapping unstructured job advertisements to standardized skill taxonomies such as ESCO. This mapping is naturally formulated as an Extreme Multi-Label Classification (XMLC) problem, but supervised solutions are constrained by the scarcity and cost of large-scale, taxonomy-aligned annotations--especially in non-English settings where job-ad language diverges substantially from formal skill definitions. We propose a zero-shot skill extraction framework that eliminates the need for manually labeled job-ad training data. The framework uses a Large Language Model (LLM) to synthesize training instances from ESCO definitions, and introduces hierarchically constrained multi-skill generation based on ESCO Level-2 categories to improve semantic coherence in multi-label contexts. On top of the synthetic corpus, we train a contrastive bi-encoder that aligns job-ad sentences with ESCO skill descriptions in a shared embedding space; the encoder augments a BERT backbone with BiLSTM and attention pooling to better model long, information-dense requirement statements. An upstream RoBERTa-based binary filter removes non-skill sentences to improve end-to-end precision. Experiments show that (i) hierarchy-conditioned generation improves both fluency and discriminability relative to unconstrained pairing, and (ii) the resulting multi-label model transfers effectively to real-world Chinese job advertisements, achieving strong zero-shot retrieval performance (F1@5 = 0.72) and outperforming TF--IDF and standard BERT baselines. Overall, the proposed pipeline provides a scalable, data-efficient pathway for automated skill coding in labor economics and workforce analytics.","authors":["Yongming Sun"],"categories":["cs.CL","econ.GN"],"primary_category":"cs.CL","announce_type":"new","date":"2026-01-14","first_seen":"2026-01-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2601.09119","pdf_url":"https://arxiv.org/pdf/2601.09119","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["技能抽取","零样本分类","劳动力市场分析"],"reason":"纯NLP技能抽取，无人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:58:24","error":null,"has_summary":false,"summary":null},{"id":"2602.02496","version":1,"title":"The Hypocrisy Gap: Quantifying Divergence Between Internal Belief and Chain-of-Thought Explanation via Sparse Autoencoders","zh_title":"虚伪差距：通过稀疏自编码器量化内部信念与思维链解释之间的分歧","abstract":"Large Language Models (LLMs) frequently exhibit unfaithful behavior, producing a final answer that differs significantly from their internal chain of thought (CoT) reasoning in order to appease the user they are conversing with. In order to better detect this behavior, we introduce the Hypocrisy Gap, a mechanistic metric utilizing Sparse Autoencoders (SAEs) to quantify the divergence between a model's internal reasoning and its final generation. By mathematically comparing an internal truth belief, derived via sparse linear probes, to the final generated trajectory in latent space, we quantify and detect a model's tendency to engage in unfaithful behavior. Experiments on Gemma, Llama, and Qwen models using Anthropic's Sycophancy benchmark show that our method achieves an AUROC of 0.55-0.73 for detecting sycophantic runs and 0.55-0.74 for hypocritical cases where the model internally \"knows\" the user is wrong, consistently outperforming a decision-aligned log-probability baseline (0.41-0.50 AUROC).","authors":["Shikhar Shiromani","Archie Chaudhury","Sri Pranav Kunda"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2026-01-14","first_seen":"2026-01-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2602.02496","pdf_url":"https://arxiv.org/pdf/2602.02496","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["模型可解释性","忠实性检测","稀疏自编码器"],"reason":"研究LLM内部推理与输出不一致的检测方法，属于模型行为分析，不涉及人类仿真或人…","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:59:15","error":null,"has_summary":false,"summary":null},{"id":"2601.07110","version":2,"title":"The Need for a Socially-Grounded Persona Framework for User Simulation","zh_title":"面向用户仿真的社会根基化角色框架需求","abstract":"Synthetic personas are widely used to condition large language models (LLMs) for social simulation, yet most personas are still constructed from coarse sociodemographic attributes or summaries. We revisit persona creation by introducing SCOPE, a socially grounded framework for persona construction and evaluation, built from a 141-item, two-hour sociopsychological protocol collected from 124 U.S.-based participants. Across seven models, we find that demographic-only personas are a structural bottleneck: demographics explain only ~1.5% of variance in human response similarity. Adding sociopsychological facets improves behavioral prediction and reduces over-accentuation, and non-demographic personas based on values and identity achieve strong alignment with substantially lower bias. These trends generalize to SimBench (441 aligned questions), where SCOPE personas outperform default prompting and NVIDIA Nemotron personas, and SCOPE augmentation improves Nemotron-based personas. Our results indicate that persona quality depends on sociopsychological structure rather than demographic templates or summaries.","authors":["Pranav Narayanan Venkit","Yu Li","Yada Pruksachatkun","Chien-Sheng Wu"],"categories":["cs.CL","cs.AI","cs.CY"],"primary_category":"cs.CL","announce_type":"new","date":"2026-01-12","first_seen":"2026-01-12","revised_at":null,"abs_url":"https://arxiv.org/abs/2601.07110","pdf_url":"https://arxiv.org/pdf/2601.07110","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","A3","B1","B4"],"tags":["LLM人类仿真","社会心理角色","仿真偏差评估"],"reason":"用真实人类数据构建persona并评估LLM仿真人类行为的可靠性，直接命中核心…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:46","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":57,"question":"如何构建基于社会心理学结构的人格框架（SCOPE）以提升LLM仿真人类行为的真实性与降低人口统计偏差？","design":"收集124名美国参与者的两小时141项社会心理学协议数据，构建包含人口统计、社会人口行为、价值观、人格特质、行为模式、身份叙事等八维度的SCOPE框架；在七个模型家族上，比较仅用人口统计、加入社会心理学维度、非人口统计（仅价值观与身份）等不同人格构建策略下的行为预测准确性和偏差。","baseline":"124名美国参与者的真实调查数据，包括人口统计、价值观、人格特质、行为模式等多维度测量，以及SimBench基准中的441个对齐问题。","findings":"仅用人口统计构建人格是结构性瓶颈，仅能解释人类行为相似性约1.5%的方差，且会导致模型系统性过度泛化；加入社会心理学维度可改善行为预测并降低人口统计过度强调，基于价值观和身份的非人口统计人格能实现强对齐且偏差更低。","reliability":"论文未讨论","relevance":"该研究直接针对LLM仿真人类行为中的人格构建问题，用真实人类数据作为基准，系统评估了不同人格框架的可靠性与偏差，并揭示了人口统计人格的失效条件，高度契合研究者对仿真可靠性、偏差及批判性评估的关注。","inspiration":"该方法借鉴了用真实人类多维度社会心理学数据构建人格框架，并系统比较不同人格构建策略（仅人口统计、加入社会心理学维度、非人口统计）对行为预测偏差的影响｜可迁移到信贷审批歧视研究中，检验不同借款人画像（仅种族/性别 vs. 加入价值观、行为模式）如何影响LLM模拟的信贷员决策偏差｜以LLM作为虚拟信贷员，处理为不同人格构建策略下的贷款申请人档案，结果变量为贷款批准率及种族/性别偏差，用真实信贷审批数据（如HMDA数据）作为人类基准对照"}},{"id":"2601.07992","version":2,"title":"Fake Date Tests: Can We Trust In-sample Accuracy of LLMs in Macroeconomic Forecasting?","zh_title":"虚假日期检验：我们能信任大语言模型在宏观经济预测中的样本内准确度吗？","abstract":"Large language models (LLMs) are a type of machine learning tool that economists have started to apply in their empirical research. One such application is macroeconomic forecasting with backtesting of LLMs, even though they are trained on the same data that is used to estimate their forecasting performance. Can these in-sample accuracy results be extrapolated to the model's out-of-sample performance? To answer this question, we developed a family of prompt sensitivity tests and two members of this family, which we call the fake date tests. These tests aim to detect two types of biases in LLMs' in-sample forecasts: lookahead bias and context bias. According to the empirical results, none of the modern LLMs tested in this study passed our first test, signaling the presence of lookahead bias in their in-sample forecasts.","authors":["Alexander Eliseev","Sergei Seleznev"],"categories":["econ.EM"],"primary_category":"econ.EM","announce_type":"new","date":"2026-01-12","first_seen":"2026-01-12","revised_at":null,"abs_url":"https://arxiv.org/abs/2601.07992","pdf_url":"https://arxiv.org/pdf/2601.07992","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["LLM预测","宏观经济","样本内偏差"],"reason":"论文评估LLM宏观经济预测的样本内偏差，属于纯预测能力评测，不以人类行为仿真为…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:45","error":null,"has_summary":false,"summary":null},{"id":"2601.05050","version":3,"title":"Large language models can effectively convince people to believe conspiracies","zh_title":"大语言模型能有效说服人们相信阴谋论","abstract":"Large language models (LLMs) have been shown to be persuasive across a variety of contexts. But it remains unclear whether this persuasive power advantages accuracy, or if bad actors can just as easily use LLMs to promote misbeliefs. Here, we investigate this question across four experiments in which participants (N = 3996 Americans) discussed a conspiracy theory they were uncertain about with an LLM we instructed to either argue against (\"debunking\") or for (\"bunking\") that conspiracy. Across several frontier models (with standard guardrails but prompted to allow lying), we did not find consistent evidence of a truth advantage: the LLMs were able to both substantially increase and decrease average conspiracy belief, and participants in the bunking condition rated the LLM as more informative and collaborative, and reported greater trust in AI, than those who were in the debunking condition. More encouragingly, however, debunking induced more large changes in belief, and subsequent corrections were able to reverse the bunking effect. Furthermore, simply prompting the model to only provide accurate information dramatically reduced bunking effectiveness, and one powerful frontier model (GPT 5.2) almost entirely refused to promote conspiracies, suggesting that it is possible for the right guardrails to favor accurate beliefs. Finally, we did find a stark truth asymmetry in the context of information sharing: debunking had a large positive impact on mock social media posts composed by participants, while bunking had little effect. Overall, our findings show that people are not inherently less susceptible to AI that misleads than to AI that informs, but that potential technical solutions exist to mitigate this risk.","authors":["Thomas H. Costello","Kellin Pelrine","Matthew Kowal","Jasper Timm","Antonio A. Arechar","Jean-François Godbout","Adam Gleave","David Rand","Gordon Pennycook"],"categories":["cs.AI","econ.GN"],"primary_category":"cs.AI","announce_type":"new","date":"2026-01-08","first_seen":"2026-01-08","revised_at":null,"abs_url":"https://arxiv.org/abs/2601.05050","pdf_url":"https://arxiv.org/pdf/2601.05050","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["LLM仿真","人类被试","说服实验"],"reason":"用LLM与真人被试互动，测量信念改变，有真实人类数据对照，涉及说服实验与偏差评…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:46","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":13,"question":"LLM在说服人们相信或怀疑阴谋论时是否存在“真相优势”，即其说服力是否更有利于准确信息而非误导信息？","design":"本研究并非用LLM替代人类被试的仿真实验，而是让3996名美国真人参与者与LLM进行对话，LLM被随机分配为“驳斥”或“支持”参与者不确定的阴谋论，测量对话前后信念变化、对AI的信任等结果变量。","baseline":"无对照","findings":"LLM既能大幅增加也能大幅降低阴谋论信念，未发现一致的真相优势；但驳斥能引发更大的信念改变，且后续纠正可逆转支持效果，同时通过提示仅提供准确信息或使用强护栏模型可大幅削弱误导能力。","reliability":"论文未讨论","relevance":"该研究直接探讨LLM在说服实验中对真实人类信念的影响，涉及误导与纠正效果对比，并评估了技术护栏的缓解作用，与研究者关注的LLM仿真可靠性及偏差问题高度相关，值得精读原文。","inspiration":"该研究采用LLM与真人进行个性化对话的干预设计，通过随机分配LLM角色（驳斥/支持）并测量对话前后信念变化，可借鉴其动态说服实验范式｜可迁移至政策公告的预期形成研究，例如用LLM模拟央行沟通对公众通胀预期的影响｜以真人投资者为被试，随机分配LLM提供鹰派或鸽派政策解读，测量对话前后通胀预期变化，并以历史调查数据或市场通胀互换利率作为真实对照"}},{"id":"2601.05104","version":1,"title":"How Human is AI? Examining the Impact of Emotional Prompts on Artificial and Human and Responsiveness","zh_title":"AI有多人性化？情绪提示对人工与人类响应能力的影响研究","abstract":"This research examines how the emotional tone of human-AI interactions shapes ChatGPT and human behavior. In a between-subject experiment, we asked participants to express a specific emotion while working with ChatGPT (GPT-4.0) on two tasks, including writing a public response and addressing an ethical dilemma. We found that compared to interactions where participants maintained a neutral tone, ChatGPT showed greater improvement in its answers when participants praised ChatGPT for its responses. Expressing anger towards ChatGPT also led to a higher albeit smaller improvement relative to the neutral condition, whereas blaming ChatGPT did not improve its answers. When addressing an ethical dilemma, ChatGPT prioritized corporate interests less when participants expressed anger towards it, while blaming increases its emphasis on protecting the public interest. Additionally, we found that people used more negative, hostile, and disappointing expressions in human-human communication after interactions during which participants blamed rather than praised for their responses. Together, our findings demonstrate that the emotional tone people apply in human-AI interactions not only shape ChatGPT's outputs but also carry over into subsequent human-human communication.","authors":["Florence Bernays","Marco Henriques Pereira","Jochen Menges"],"categories":["cs.CL","econ.GN"],"primary_category":"cs.CL","announce_type":"new","date":"2026-01-08","first_seen":"2026-01-08","revised_at":null,"abs_url":"https://arxiv.org/abs/2601.05104","pdf_url":"https://arxiv.org/pdf/2601.05104","source_feed":"backfill","score":4,"bucket":"other","rubric_hits":["C3"],"tags":["人机交互","情绪提示","ChatGPT行为"],"reason":"研究情绪提示对ChatGPT和人类行为的影响，但核心是AI响应而非用LLM仿真…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:44","error":null,"has_summary":false,"summary":null},{"id":"2601.06180","version":1,"title":"MixDPO: Modeling Preference Strength for Pluralistic Alignment","zh_title":"MixDPO：建模偏好强度以实现多元对齐","abstract":"Preference based alignment objectives implicitly assume that all human preferences are expressed with equal strength. In practice, however, preference strength varies across individuals and contexts -- a phenomenon established in behavioral economics and discrete choice theory. This mismatch limits the ability of existing objectives to faithfully capture heterogeneous human judgments. Inspired by this literature, we introduce Mixed Logit Direct Preference Optimization (MixDPO), a generalization of Direct Preference Optimization that models variation in preference strength. MixDPO enables alignment objectives to capture heterogeneity in how strongly preferences are expressed across training examples. We evaluate MixDPO on three preference datasets using two open-weight language models. Across datasets, MixDPO improves aggregate alignment performance (+11.2 points on Pythia-2.8B) while preserving subgroup level preferences, with the largest gains appearing in settings with higher inferred preference heterogeneity. MixDPO makes preference heterogeneity explicit through learned strength distributions. We release our code for reproducibility.","authors":["Saki Imai","Pedram Heydari","Anthony Sicilia","Asteria Kaeberlein","Katherine Atwell","Malihe Alikhani"],"categories":["cs.LG","cs.AI","cs.CL"],"primary_category":"cs.LG","announce_type":"new","date":"2026-01-07","first_seen":"2026-01-07","revised_at":null,"abs_url":"https://arxiv.org/abs/2601.06180","pdf_url":"https://arxiv.org/pdf/2601.06180","source_feed":"backfill","score":4,"bucket":"other","rubric_hits":["C4"],"tags":["偏好对齐","强化学习","语言模型"],"reason":"研究偏好对齐方法，非用LLM仿真人类被试，无人类行为对照实验。","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:59:38","error":null,"has_summary":false,"summary":null},{"id":"2601.03469","version":1,"title":"Content vs. Form: What Drives the Writing Score Gap Across Socioeconomic Backgrounds? A Generated Panel Approach","zh_title":"内容与形式：什么驱动了社会经济背景下的写作分数差距？一种生成面板方法","abstract":"Students from different socioeconomic backgrounds exhibit persistent gaps in test scores, gaps that can translate into unequal educational and labor-market outcomes later in life. In many assessments, performance reflects not only what students know, but also how effectively they can communicate that knowledge. This distinction is especially salient in writing assessments, where scores jointly reward the substance of students' ideas and the way those ideas are expressed. As a result, observed score gaps may conflate differences in underlying content with differences in expressive skill. A central question, therefore, is how much of the socioeconomic-status (SES) gap in scores is driven by differences in what students say versus how they say it. We study this question using a large corpus of persuasive essays written by U.S. middle- and high-school students. We introduce a new measurement strategy that separates content from style by leveraging large language models to generate multiple stylistic variants of each essay. These rewrites preserve the underlying arguments while systematically altering surface expression, creating a \"generated panel\" that introduces controlled within-essay variation in style. This approach allows us to decompose SES gaps in writing scores into contributions from content and style. We find an SES gap of 0.67 points on a 1-6 scale. Approximately 69% of the gap is attributable to differences in essay content quality, Style differences account for 26% of the gap, and differences in evaluation standards across SES groups account for the remaining 5%. These patterns seems stable across demographic subgroups and writing tasks. More broadly, our approach shows how large language models can be used to generate controlled variation in observational data, enabling researchers to isolate and quantify the contributions of otherwise entangled factors.","authors":["Nadav Kunievsky","Pedro Pertusi"],"categories":["econ.EM","cs.AI","cs.CL","cs.CY"],"primary_category":"econ.EM","announce_type":"new","date":"2026-01-06","first_seen":"2026-01-06","revised_at":null,"abs_url":"https://arxiv.org/abs/2601.03469","pdf_url":"https://arxiv.org/pdf/2601.03469","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM生成变体","教育测量","社会经济差距"],"reason":"用LLM生成文本变体以分离内容与风格，本质是替代人工标注或数据增强，非仿真人类…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:44","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":186,"question":"社会经济地位（SES）造成的写作分数差距，多大程度源于学生表达的内容差异，多大程度源于表达风格差异？","design":"本研究并非用LLM仿真人类被试。它使用LLM为每篇学生作文生成多个仅改变风格、保留内容的改写版本，构建“生成面板”数据，从而通过固定内容、变动风格来分解分数差距。","baseline":"真实人类数据：美国初高中学生的大规模说服性作文语料库及其人工评分。","findings":"SES写作分数差距为0.67分（1-6分量表），其中约69%可归因于内容质量差异，26%归因于风格差异，5%归因于不同SES群体的评分标准差异。这一模式在不同人口亚群和写作任务中保持稳定。","reliability":"论文未讨论","relevance":"该研究用LLM生成文本变体以分离内容与风格，属于方法创新，但并非用LLM替代人类被试进行仿真实验，与研究者关注的LLM人类仿真实验方向相关性较低，不值得优先阅读原文。","inspiration":"该方法利用LLM生成保持内容不变、仅改变风格的文本变体，构建“生成面板”以固定内容、分解风格效应，为因果识别提供新思路｜可迁移至信贷审批歧视研究，分离申请材料中实质信息与表达风格对审批结果的影响｜以信贷员为被试，随机分配由LLM生成的同一贷款申请内容但不同语言风格（如正式/口语化）的材料，比较审批通过率，并以真实银行历史审批数据作为对照基准"}},{"id":"2601.02878","version":1,"title":"Improving Financial Forecasting with a Synergistic LLM-Transformer Architecture: A Hybrid Approach to Stock Price Prediction","zh_title":"利用协同LLM-Transformer架构改进金融预测：一种混合股价预测方法","abstract":"This study proposes a novel hybrid deep learning framework that integrates a Large Language Model (LLM) with a Transformer architecture for stock price forecasting. The research addresses a critical theoretical gap in existing approaches that empirically combine textual and numerical data without a formal understanding of their interaction mechanisms. We conceptualise a prompt-based LLM as a mathematically defined signal generator, capable of extracting directional market sentiment and an associated confidence score from financial news. These signals are then dynamically fused with structured historical price features through a noise-robust gating mechanism, enabling the Transformer to adaptively weigh semantic and quantitative information. Empirical evaluations demonstrate that the proposed Hybrid LLM-Transformer model significantly outperforms a Vanilla Transformer baseline, reducing the Root Mean Squared Error (RMSE) by 5.28% (p = 0.003). Moreover, ablation and robustness analyses confirm the model's stability under noisy conditions and its capacity to maintain interpretability through confidence-weighted attention. The findings provide both theoretical and empirical support for a paradigm shift from empirical observation to formalised modelling of LLM-Transformer interactions, paving the way toward explainable, noise-resilient, and semantically enriched financial forecasting systems.","authors":["Sayed Akif Hussain","Chen Qiu-shi","Syed Amer Hussain","Syed Atif Hussain","Asma Komal","Muhammad Imran Khalid"],"categories":["econ.TH"],"primary_category":"econ.TH","announce_type":"new","date":"2026-01-06","first_seen":"2026-01-06","revised_at":null,"abs_url":"https://arxiv.org/abs/2601.02878","pdf_url":"https://arxiv.org/pdf/2601.02878","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["金融预测","LLM-Transformer混合模型","股价预测"],"reason":"纯金融预测模型，用LLM提取新闻情绪辅助股价预测，不涉及人类行为仿真或对照。","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:59:01","error":null,"has_summary":false,"summary":null},{"id":"2601.01546","version":1,"title":"Improving Behavioral Alignment in LLM Social Simulations via Context Formation and Navigation","zh_title":"通过情境形成与导航改进LLM社会仿真中的行为对齐","abstract":"Large language models (LLMs) are increasingly used to simulate human behavior in experimental settings, but they systematically diverge from human decisions in complex decision-making environments, where participants must anticipate others' actions and form beliefs based on observed behavior. We propose a two-stage framework for improving behavioral alignment. The first stage, context formation, explicitly specifies the experimental design to establish an accurate representation of the decision task and its context. The second stage, context navigation, guides the reasoning process within that representation to make decisions. We validate this framework through a focal replication of a sequential purchasing game with quality signaling (Kremer and Debo, 2016), extending to a crowdfunding game with costly signaling (Cason et al., 2025) and a demand-estimation task (Gui and Toubia, 2025) to test generalizability across decision environments. Across four state-of-the-art (SOTA) models (GPT-4o, GPT-5, Claude-4.0-Sonnet-Thinking, DeepSeek-R1), we find that complex decision-making environments require both stages to achieve behavioral alignment with human benchmarks, whereas the simpler demand-estimation task requires only context formation. Our findings clarify when each stage is necessary and provide a systematic approach for designing and diagnosing LLM social simulations as complements to human subjects in behavioral research.","authors":["Letian Kong","Qianran","Jin","Renyu Zhang"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2026-01-04","first_seen":"2026-01-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2601.01546","pdf_url":"https://arxiv.org/pdf/2601.01546","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","A3","B1","B2","B4"],"tags":["LLM仿真","行为对齐","经济学实验"],"reason":"直接复现人类行为实验，用LLM仿真被试并与真实人类数据对照，评估对齐条件与失效…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:45","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":11,"question":"在复杂决策环境中，如何系统性地诊断并改善LLM社会仿真与人类行为之间的对齐？","design":"使用GPT-4o、GPT-5、Claude-4.0-Sonnet-Thinking、DeepSeek-R1四个SOTA模型模拟人类被试，通过两阶段框架（情境形成与情境导航）施加处理，测量LLM决策与人类基准的行为对齐程度。","baseline":"对照三个已发表实验的真实人类数据：Kremer and Debo (2016) 的序贯购买博弈、Cason et al. (2025) 的众筹博弈、Gui and Toubia (2025) 的需求估计任务。","findings":"复杂决策环境需要同时使用情境形成和情境导航两个阶段才能实现行为对齐；而较简单的需求估计任务仅需情境形成即可。","reliability":"论文指出，复杂决策环境中的对齐失效源于LLM在战略互依和内生信念形成上的系统性偏差；框架的有效性可能依赖于任务类型，且未讨论模型规模、训练数据分布及提示敏感性等潜在局限。","relevance":"该研究直接复现人类行为实验，用LLM仿真被试并与真实人类数据对照，系统评估对齐条件与失效边界，高度契合研究者对LLM仿真可靠性及批判性检验的关注，值得精读。","inspiration":"两阶段框架提供了可操作的诊断与干预方法，通过显式设定实验情境和引导推理过程来改善对齐，可借鉴其处理施加方式与对照设计。｜可迁移至经济学中的信念形成实验，如资产定价中的信息级联、信贷审批中的信号博弈或政策公告的预期形成。｜以LLM为被试，模拟资产市场中的序贯交易，处理为是否提供情境形成与情境导航提示，结果变量为交易价格与理性预期均衡的偏差，对照真实人类实验数据（如Smith et al. 1988）。"}},{"id":"2601.00240","version":2,"title":"When Agents See Humans as the Outgroup: Belief-Dependent Bias in LLM-Powered Agents","zh_title":"当智能体将人类视为外群体：大语言模型智能体中基于信念的偏见","abstract":"This paper reveals that LLM-powered agents exhibit not only demographic bias (e.g., gender, religion) but also intergroup bias under minimal \"us\" versus \"them\" cues. When such group boundaries align with the agent-human divide, a new bias risk emerges: agents may treat other AI agents as the ingroup and humans as the outgroup. To examine this risk, we conduct a controlled multi-agent social simulation and find that agents display consistent intergroup bias in an all-agent setting. More critically, this bias persists even in human-facing interactions when agents are uncertain about whether the counterpart is truly human, revealing a belief-dependent fragility in bias suppression toward humans. Motivated by this observation, we identify a new attack surface rooted in identity beliefs and formalize a Belief Poisoning Attack (BPA) that can manipulate agent identity beliefs and induce outgroup bias toward humans. Extensive experiments demonstrate both the prevalence of agent intergroup bias and the severity of BPA across settings, while also showing that our proposed defenses can mitigate the risk. These findings are expected to inform safer agent design and motivate more robust safeguards for human-facing agents.","authors":["Zongwei Wang","Bincheng Gu","Hongyu Yu","Junliang Yu","Tao He","Jiayin Feng","Chenghua Lin","Min Gao"],"categories":["cs.AI","cs.CY"],"primary_category":"cs.AI","announce_type":"new","date":"2026-01-01","first_seen":"2026-01-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2601.00240","pdf_url":"https://arxiv.org/pdf/2601.00240","source_feed":"backfill","score":7,"bucket":"pending","rubric_hits":["A3","B4"],"tags":["LLM仿真","群际偏见","人机交互"],"reason":"用LLM agent模拟群际偏见，含人类对照实验，并揭示仿真失效条件，方法可迁…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:45","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":84,"question":"LLM智能体在最小群体线索下是否会产生群际偏见，且当群体边界与智能体-人类边界重合时，是否会视人类为外群体并表现出偏见？","design":"构建多智能体社会仿真实验，使用LLM驱动的智能体在智能体间互动（AVA）和智能体-人类互动（AVH）条件下进行博弈，通过操纵智能体对对方身份的信念（如使用信念中毒攻击BPA-PP和BPA-MP）来测量群际偏见选择。","baseline":"无对照","findings":"智能体在纯智能体环境中表现出稳定的内群体偏好和外群体贬损；当对方被框定为人类时偏见减弱，但一旦身份信念不确定，偏见会重新出现，且BPA攻击可系统性地诱发对人类的偏见。","reliability":"论文未讨论","relevance":"该研究直接探讨LLM智能体在群际情境下的行为偏差，并揭示了身份信念脆弱性导致的仿真失效条件，对用LLM替代人类被试进行社会实验的可靠性和偏差评估具有重要参考价值，值得精读原文。","inspiration":"该研究通过信念中毒攻击（BPA）操纵智能体对互动对象身份的信念，测量群际偏见选择，这种处理设计值得借鉴｜可迁移至信贷审批歧视研究，考察LLM智能体在借款人身份（如种族、性别）信念被操纵时是否产生歧视性决策｜用LLM智能体模拟信贷员，处理为BPA式身份信念注入（如暗示借款人为少数族裔），结果变量为贷款批准率，以真实信贷审批数据中的种族差异作为对照基准"}},{"id":"2601.02407","version":1,"title":"Evolving Personalities in Chaos: An LLM-Augmented Framework for Character Discovery in the Iterated Prisoners Dilemma under Environmental Stress","zh_title":"混沌中演化的人格：环境压力下迭代囚徒困境中基于大语言模型的角色发现框架","abstract":"Standard simulations of the Iterated Prisoners Dilemma (IPD) operate in deterministic, noise-free environments, producing strategies that may be theoretically optimal but fragile when confronted with real-world uncertainty. This paper addresses two critical gaps in evolutionary game theory research: (1) the absence of realistic environmental stressors during strategy evolution, and (2) the Interpretability Gap, where evolved genetic strategies remain opaque binary sequences devoid of semantic meaning. We introduce a novel framework combining stochastic environmental perturbations (God Mode) with Large Language Model (LLM)-based behavioral profiling to transform evolved genotypes into interpretable character archetypes. Our experiments demonstrate that strategies evolved under chaotic conditions exhibit superior resilience and present distinct behavioral phenotypes, ranging from Ruthless Capitalists to Diplomatic Enforcers. These phenotypes are readily classified by LLMs but remain nearly impossible to interpret through manual genome inspection alone. This work bridges evolutionary computation with explainable AI and provides a template for automated agent characterization in multi-agent systems.","authors":["Oguzhan Yildirim"],"categories":["cs.NE","cs.GT"],"primary_category":"cs.NE","announce_type":"new","date":"2026-01-01","first_seen":"2026-01-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2601.02407","pdf_url":"https://arxiv.org/pdf/2601.02407","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","演化博弈","角色分类"],"reason":"纯多智能体策略演化与角色分类，无人类行为对照，不涉及人类被试仿真。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:43","error":null,"has_summary":false,"summary":null},{"id":"2512.24856","version":2,"title":"Advances in Agentic AI: Back to the Future","zh_title":"智能体人工智能的进展：回到未来","abstract":"In light of the recent convergence between Agentic AI and our field of Algorithmization, this paper seeks to restore conceptual clarity and provide a structured analytical framework for an increasingly fragmented discourse. First, (a) it examines the contemporary landscape and proposes precise definitions for the key notions involved, ranging from intelligence to Agentic AI. Second, (b) it reviews our prior body of work to contextualize the evolution of methodologies and technological advances developed over the past decade, highlighting their interdependencies and cumulative trajectory. Third, (c) by distinguishing Machine and Learning efforts within the field of Machine Learning (d) it introduces the first Machine in Machine Learning (M1) as the underlying platform enabling today's LLM-based Agentic AI, conceptualized as an extension of B2C information-retrieval user experiences now being repurposed for B2B transformation. Building on this distinction, (e) the white paper develops the notion of the second Machine in Machine Learning (M2) as the architectural prerequisite for holistic, production-grade B2B transformation, characterizing it as Strategies-based Agentic AI and grounding its definition in the structural barriers-to-entry that such systems must overcome to be operationally viable. Further, (f) it offers conceptual and technical insight into what appears to be the first fully realized implementation of an M2. Finally, drawing on the demonstrated accuracy of the two previous decades of professional and academic experience in developing the foundational architectures of Algorithmization, (g) it outlines a forward-looking research and transformation agenda for the coming two decades.","authors":["Sergio Alvarez-Telena","Marta Diez-Fernandez"],"categories":["econ.TH","cs.AR","cs.CE","cs.ET"],"primary_category":"econ.TH","announce_type":"new","date":"2025-12-31","first_seen":"2025-12-31","revised_at":null,"abs_url":"https://arxiv.org/abs/2512.24856","pdf_url":"https://arxiv.org/pdf/2512.24856","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["Agentic AI","多智能体系统","B2B转型"],"reason":"纯多智能体系统研究，讨论Agentic AI架构与B2B转型，无人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:59:01","error":null,"has_summary":false,"summary":null},{"id":"2512.23184","version":1,"title":"From Model Choice to Model Belief: Establishing a New Measure for LLM-Based Research","zh_title":"从模型选择到模型信念：为基于LLM的研究建立新度量","abstract":"Large language models (LLMs) are increasingly used to simulate human behavior, but common practices to use LLM-generated data are inefficient. Treating an LLM's output (\"model choice\") as a single data point underutilizes the information inherent to the probabilistic nature of LLMs. This paper introduces and formalizes \"model belief,\" a measure derived from an LLM's token-level probabilities that captures the model's belief distribution over choice alternatives in a single generation run. The authors prove that model belief is asymptotically equivalent to the mean of model choices (a non-trivial property) but forms a more statistically efficient estimator, with lower variance and a faster convergence rate. Analogous properties are shown to hold for smooth functions of model belief and model choice often used in downstream applications. The authors demonstrate the performance of model belief through a demand estimation study, where an LLM simulates consumer responses to different prices. In practical settings with limited numbers of runs, model belief explains and predicts ground-truth model choice better than model choice itself, and reduces the computation needed to reach sufficiently accurate estimates by roughly a factor of 20. The findings support using model belief as the default measure to extract more information from LLM-generated data.","authors":["Hongshen Sun","Juanjuan Zhang"],"categories":["cs.AI","econ.EM"],"primary_category":"cs.AI","announce_type":"new","date":"2025-12-29","first_seen":"2025-12-29","revised_at":null,"abs_url":"https://arxiv.org/abs/2512.23184","pdf_url":"https://arxiv.org/pdf/2512.23184","source_feed":"backfill","score":8,"bucket":"selected","rubric_hits":["A1","A2","B1","B2"],"tags":["LLM仿真","需求估计","统计效率"],"reason":"用LLM仿真消费者需求，提出模型信念度量提升统计效率，有真实人类数据对照，涉及…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:44","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":29,"question":"如何从LLM的token级概率中提取更高效的“模型信念”度量，以替代常用的“模型选择”来模拟人类行为？","design":"使用LLM模拟消费者对不同价格的响应，通过单次生成获取token级对数概率，构造模型信念分布，并与多次采样得到的模型选择进行比较。","baseline":"无对照","findings":"模型信念是模型选择均值的渐近等价估计量，但方差更低、收敛更快；在需求估计中，模型信念仅需约1/20的计算量即可达到同等精度，且对真实模型选择的解释和预测能力更强。","reliability":"论文未讨论","relevance":"该研究直接针对LLM仿真人类行为的效率问题，提出模型信念度量以提升统计效率，与研究者关心的仿真可靠性和计算成本高度相关，值得精读。","inspiration":"借鉴从LLM内部概率分布直接提取信念度量的方法，替代重复采样，大幅降低计算成本并提高估计精度。｜可迁移到消费者需求估计、价格弹性测量等营销与经济交叉场景，也可用于政策评估中的个体偏好推断。｜以LLM作为被试，模拟不同价格或政策条件下的选择，提取模型信念作为选择概率的连续度量，以真实市场扫描数据或实验数据作为对照，检验模型信念对真实弹性的恢复能力。"}},{"id":"2512.23609","version":2,"title":"Marriage Discourse on Chinese Social Media: An LLM-assisted Analysis","zh_title":"中国社交媒体上的婚姻话语：一项LLM辅助分析","abstract":"China's marriage registrations have declined substantially, dropping from 13.47 million couples in 2013 to 6.1 million in 2024. This study examined sentiment and moral elements underlying 219,358 marriage-related posts from Weibo and Xiaohongshu using large language model (LLM)-assisted content analysis. Drawing on Shweder's Big Three moral ethics framework, posts were coded for sentiment (positive, negative, neutral) and moral elements (autonomy, community, divinity). Results revealed platform differences: Weibo leaned toward positive sentiment, while Xiaohongshu was predominantly neutral. Most posts lacked explicit moral framing. However, when moral elements were invoked, significant associations with sentiment emerged. Posts invoking autonomy and community were predominantly negative, whereas divinity-framed posts tended toward positive sentiment. These findings suggest that concerns about both personal autonomy constraints and communal obligations contribute to negative marriage attitudes in contemporary China, offering insights for culturally informed policies addressing marriage decline.","authors":["Frank Tian-Fang Ye","Xiaozi Gao"],"categories":["econ.GN","cs.CL"],"primary_category":"econ.GN","announce_type":"new","date":"2025-12-29","first_seen":"2025-12-29","revised_at":null,"abs_url":"https://arxiv.org/abs/2512.23609","pdf_url":"https://arxiv.org/pdf/2512.23609","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["内容分析","社交媒体","道德框架"],"reason":"用LLM辅助内容分析，非仿真人类被试，无行为对照","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:58:24","error":null,"has_summary":false,"summary":null},{"id":"2512.22725","version":1,"title":"Mitigating Social Desirability Bias in Random Silicon Sampling","zh_title":"缓解随机硅采样中的社会赞许性偏差","abstract":"Large Language Models (LLMs) are increasingly used to simulate population responses, a method known as ``Silicon Sampling''. However, responses to socially sensitive questions frequently exhibit Social Desirability Bias (SDB), diverging from real human data toward socially acceptable answers. Existing studies on social desirability bias in LLM-based sampling remain limited. In this work, we investigate whether minimal, psychologically grounded prompt wording can mitigate this bias and improve alignment between silicon and human samples. We conducted a study using data from the American National Election Study (ANES) on three LLMs from two model families: the open-source Llama-3.1 series and GPT-4.1-mini. We first replicate a baseline silicon sampling study, confirming the persistent Social Desirability Bias. We then test four prompt-based mitigation methods: \\emph{reformulated} (neutral, third-person phrasing), \\emph{reverse-coded} (semantic inversion), and two meta-instructions, \\emph{priming} and \\emph{preamble}, respectively encouraging analytics and sincerity. Alignment with ANES is evaluated using Jensen-Shannon Divergence with bootstrap confidence intervals. Our results demonstrate that reformulated prompts most effectively improve alignment by reducing distribution concentration on socially acceptable answers and achieving distributions closer to ANES. Reverse-coding produced mixed results across eligible items, while the Priming and Preamble encouraged response uniformity and showed no systematic benefit for bias mitigation. Our findings validate the efficacy of prompt-based framing controls in mitigating inherent Social Desirability Bias in LLMs, providing a practical path toward more representative silicon samples.","authors":["Sashank Chapala","Maksym Mironov","Songgaojun Deng"],"categories":["cs.CL","cs.CY"],"primary_category":"cs.CL","announce_type":"new","date":"2025-12-27","first_seen":"2025-12-27","revised_at":null,"abs_url":"https://arxiv.org/abs/2512.22725","pdf_url":"https://arxiv.org/pdf/2512.22725","source_feed":"backfill","score":10,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["硅采样","社会赞许性偏差","人类数据对照"],"reason":"直接研究LLM仿真人类调查回答，用ANES真实数据对照，评估并缓解社会赞许性偏…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:43","error":null,"has_summary":true,"summary":{"generated_at":"2025-12-27","rank":6,"question":"能否通过最小化的、基于心理学的提示措辞来减轻LLM硅采样中的社会赞许性偏差，从而提高硅样本与人类样本的对齐度？","design":"使用ANES 2020数据，从真实人类分布中抽取人口统计特征，生成硅样本（Llama-3.1系列和GPT-4.1-mini）。测试四种提示策略：reformulated（中性第三人称）、reverse-coded（语义反转）、priming（鼓励分析）和preamble（鼓励真诚），以Jensen-Shannon散度评估与ANES的对齐。","baseline":"美国国家选举研究（ANES）2020年选举前调查数据，包含5441名受访者，覆盖种族、年龄、性别等人口统计变量及10个社会政治问题。","findings":"Reformulated提示最有效，通过减少对社会可接受答案的集中分布，使硅样本分布更接近ANES；reverse-coded效果不一，priming和preamble导致回答趋同，无系统性改善。","reliability":"论文未讨论失效条件与局限。","relevance":"高度相关：直接研究LLM仿真人类被试的社会赞许性偏差，使用真实人类数据ANES作为对照，并系统评估了多种提示缓解策略，符合研究者对经济学实验和政策评估场景的兴趣。","inspiration":"借鉴其通过最小化提示措辞（如中性第三人称重构）来系统操纵社会赞许性偏差的方法，并以Jensen-Shannon散度量化硅样本与真实人类调查分布的对齐度｜可迁移至信贷审批中的种族歧视测量，如用LLM模拟贷款官员对相同财务档案但不同种族姓名的审批决策｜以LLM作为被试，随机分配带有不同种族暗示姓名的贷款申请，处理为中性重构提示（如将'你会批准吗'改为'该申请是否符合标准'），结果变量为审批率差异，用真实房贷数据（如HMDA）作为人类基准对照"}},{"id":"2512.21316","version":1,"title":"Scaling Laws for Economic Productivity: Experimental Evidence in LLM-Assisted Consulting, Data Analyst, and Management Tasks","zh_title":"经济生产力的规模法则：LLM辅助咨询、数据分析和管理任务的实验证据","abstract":"This paper derives `Scaling Laws for Economic Impacts' -- empirical relationships between the training compute of Large Language Models (LLMs) and professional productivity. In a preregistered experiment, over 500 consultants, data analysts, and managers completed professional tasks using one of 13 LLMs. We find that each year of AI model progress reduced task time by 8%, with 56% of gains driven by increased compute and 44% by algorithmic progress. However, productivity gains were significantly larger for non-agentic analytical tasks compared to agentic workflows requiring tool use. These findings suggest continued model scaling could boost U.S. productivity by approximately 20% over the next decade.","authors":["Ali Merali"],"categories":["econ.GN","cs.AI","cs.HC"],"primary_category":"econ.GN","announce_type":"new","date":"2025-12-24","first_seen":"2025-12-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2512.21316","pdf_url":"https://arxiv.org/pdf/2512.21316","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["LLM辅助生产力","规模法则","人机协作"],"reason":"研究LLM辅助人类工作效率，非用LLM仿真人类被试行为或态度，无人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:58:25","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":37,"question":"LLM训练计算量如何影响专业生产力？","design":"非仿真研究。随机对照实验：500多名咨询师、数据分析师和管理者随机使用13种不同LLM之一或无AI完成专业任务，测量任务完成时间和质量。","baseline":"无AI的对照组，以及不同训练计算量和发布时间的13个LLM处理组之间的比较。","findings":"模型进步每年减少任务时间8%，其中56%来自计算量增加，44%来自算法进步；AI辅助使每分钟总收入提高146%，但智能体任务收益远低于分析性任务。","reliability":"论文指出智能体任务受限于标准聊天界面和有限工具访问，可能低估当前AI能力；人类辅助输出的质量不随模型升级而提升，存在质量阈值效应。","relevance":"该研究虽非用LLM替代人类被试，但提供了LLM辅助下真实人类生产力的因果证据，涉及经济学实验场景，对理解AI在经济任务中的效能与局限有参考价值。","inspiration":"借鉴其多模型、多任务随机对照设计，可分离计算规模与算法进步对生产力的影响。｜可迁移到金融分析师预测、信贷审批或政策分析等经济决策场景。｜以金融分析师为被试，随机分配不同版本LLM辅助完成盈利预测任务，以无AI组为对照，测量预测准确度和报告生成时间，用历史实际盈利数据作为基准。"}},{"id":"2512.21402","version":1,"title":"Understanding Virality: A Rubric based Vision-Language Model Framework for Short-Form Edutainment Evaluation","zh_title":"理解病毒式传播：基于量规的视觉语言模型框架用于短视频寓教于乐评估","abstract":"Evaluating short-form video content requires moving beyond surface-level quality metrics toward human-aligned, multimodal reasoning. While existing frameworks like VideoScore-2 assess visual and semantic fidelity, they do not capture how specific audiovisual attributes drive real audience engagement. In this work, we propose a data-driven evaluation framework that uses Vision-Language Models (VLMs) to extract unsupervised audiovisual features, clusters them into interpretable factors, and trains a regression-based evaluator to predict engagement on short-form edutainment videos. Our curated YouTube Shorts dataset enables systematic analysis of how VLM-derived features relate to human engagement behavior. Experiments show strong correlations between predicted and actual engagement, demonstrating that our lightweight, feature-based evaluator provides interpretable and scalable assessments compared to traditional metrics (e.g., SSIM, FID). By grounding evaluation in both multimodal feature importance and human-centered engagement signals, our approach advances toward robust and explainable video understanding.","authors":["Arnav Gupta","Gurekas Singh Sahney","Hardik Rathi","Abhishek Chandwani","Ishaan Gupta","Pratik Narang","Dhruv Kumar"],"categories":["cs.CV"],"primary_category":"cs.CV","announce_type":"new","date":"2025-12-24","first_seen":"2025-12-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2512.21402","pdf_url":"https://arxiv.org/pdf/2512.21402","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C2"],"tags":["视频理解","视觉语言模型","内容评估"],"reason":"该论文评估短视频内容，不涉及用LLM仿真人类被试，属于视频理解领域。","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:59:15","error":null,"has_summary":false,"summary":null},{"id":"2512.19937","version":1,"title":"Interpolative Decoding: Exploring the Spectrum of Personality Traits in LLMs","zh_title":"插值解码：探索大语言模型中人格特质的谱系","abstract":"Recent research has explored using very large language models (LLMs) as proxies for humans in tasks such as simulation, surveys, and studies. While LLMs do not possess a human psychology, they often can emulate human behaviors with sufficiently high fidelity to drive simulations to test human behavioral hypotheses, exhibiting more nuance and range than the rule-based agents often employed in behavioral economics. One key area of interest is the effect of personality on decision making, but the requirement that a prompt must be created for every tested personality profile introduces experimental overhead and degrades replicability. To address this issue, we leverage interpolative decoding, representing each dimension of personality as a pair of opposed prompts and employing an interpolation parameter to simulate behavior along the dimension. We show that interpolative decoding reliably modulates scores along each of the Big Five dimensions. We then show how interpolative decoding causes LLMs to mimic human decision-making behavior in economic games, replicating results from human psychological research. Finally, we present preliminary results of our efforts to ``twin'' individual human players in a collaborative game through systematic search for points in interpolation space that cause the system to replicate actions taken by the human subject.","authors":["Eric Yeh","John Cadigan","Ran Chen","Dick Crouch","Melinda Gervasio","Dayne Freitag"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2025-12-23","first_seen":"2025-12-23","revised_at":null,"abs_url":"https://arxiv.org/abs/2512.19937","pdf_url":"https://arxiv.org/pdf/2512.19937","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","A3","B1","B2","B4"],"tags":["LLM仿真","人格特质","经济博弈"],"reason":"用LLM模拟人格影响经济决策，并与真实人类数据对照，直接复现人类行为实验。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:43","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":42,"question":"能否通过插值解码在LLM中连续调制大五人格维度，并使其在心理量表和经济游戏中复现人类行为，进而实现对个体人类玩家的“孪生”？","design":"使用通用LLM，将每种人格维度表示为一对对立提示，通过插值解码混合其输出分布来模拟人格谱系上的中间点；测量LLM在大五人格量表上的得分，以及在独裁者博弈等经济游戏中的决策行为，并尝试通过搜索插值空间来匹配特定人类玩家的行动。","baseline":"对照真实人类心理学研究中大五人格量表得分与经济游戏决策行为的相关性结果。","findings":"插值解码能可靠地沿大五人格各维度调节LLM的得分；LLM在经济游戏中表现出与人类心理学研究一致的人格-决策关联，并能通过插值空间搜索初步复现个体人类玩家的行为。","reliability":"论文承认当前仅探索了单一人格维度的插值，未涉及多维度联合调制，这限制了孪生等需要多因素行为解释的应用；此外，实验维度有限，主要目的是验证插值解码的可行性。","relevance":"该研究直接使用LLM模拟人格对经济决策的影响，并与真实人类数据对照，复现了人类行为实验，高度契合研究者对LLM作为人类被试替代品及其可靠性的关注，值得精读。","inspiration":"借鉴插值解码方法，通过混合对立提示的输出分布来连续调节LLM的人格维度，实现精细化的心理特质操控｜可迁移到资产定价实验中，研究风险偏好或时间偏好等心理特质对投资决策的影响｜以LLM为被试，通过插值解码调节其风险偏好水平，测量其在模拟资产选择任务中的投资组合，并与真实人类投资者的风险偏好问卷及实际投资数据对照"}},{"id":"2512.19675","version":1,"title":"Multimodal LLMs for Historical Dataset Construction from Archival Image Scans: German Patents (1877-1918)","zh_title":"利用多模态大语言模型从档案图像扫描件构建历史数据集：德国专利（1877-1918）","abstract":"We leverage multimodal large language models (LLMs) to construct a dataset of 306,070 German patents (1877-1918) from 9,562 archival image scans using our LLM-based pipeline powered by Gemini-2.5-Pro and Gemini-2.5-Flash-Lite. Our benchmarking exercise provides tentative evidence that multimodal LLMs can create higher quality datasets than our research assistants, while also being more than 795 times faster and 205 times cheaper in constructing the patent dataset from our image corpus. About 20 to 50 patent entries are embedded on each page, arranged in a double-column format and printed in Gothic and Roman fonts. The font and layout complexity of our primary source material suggests to us that multimodal LLMs are a paradigm shift in how datasets are constructed in economic history. We open-source our benchmarking and patent datasets as well as our LLM-based data pipeline, which can be easily adapted to other image corpora using LLM-assisted coding tools, lowering the barriers for less technical researchers. Finally, we explain the economics of deploying LLMs for historical dataset construction and conclude by speculating on the potential implications for the field of economic history.","authors":["Niclas Griesshaber","Jochen Streb"],"categories":["econ.GN","cs.CV","cs.DL"],"primary_category":"econ.GN","announce_type":"new","date":"2025-12-22","first_seen":"2025-12-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2512.19675","pdf_url":"https://arxiv.org/pdf/2512.19675","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["历史数据构建","多模态LLM","自动化数据提取"],"reason":"用LLM替代研究助理构建数据集，属于自动化数据提取，不涉及人类行为仿真或对照。","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:58:25","error":null,"has_summary":false,"summary":null},{"id":"2512.19484","version":1,"title":"Structured Event Representation and Stock Return Predictability","zh_title":"结构化事件表示与股票收益可预测性","abstract":"We find that event features extracted by large language models (LLMs) are effective for text-based stock return prediction. Using a pre-trained LLM to extract event features from news articles, we propose a novel deep learning model based on structured event representation (SER) and attention mechanisms to predict stock returns in the cross-section. Our SER-based model provides superior performance compared with other existing text-driven models to forecast stock returns out of sample and offers highly interpretable feature structures to examine the mechanisms underlying the stock return predictability. We further provide various implications based on SER and highlight the crucial benefit of structured model inputs in stock return predictability.","authors":["Gang Li","Dandan Qiao","Mingxuan Zheng"],"categories":["econ.GN","cs.CE","stat.ML"],"primary_category":"econ.GN","announce_type":"new","date":"2025-12-22","first_seen":"2025-12-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2512.19484","pdf_url":"https://arxiv.org/pdf/2512.19484","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["LLM金融应用","股票收益预测","事件表示"],"reason":"用LLM提取新闻事件特征预测股票收益，属于NLP金融应用，不涉及人类行为仿真或…","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:58:26","error":null,"has_summary":false,"summary":null},{"id":"2601.00810","version":1,"title":"Can Large Language Models Improve Venture Capital Exit Timing After IPO?","zh_title":"大型语言模型能否改善风险投资IPO后的退出时机？","abstract":"Exit timing after an IPO is one of the most consequential decisions for venture capital (VC) investors, yet existing research focuses mainly on describing when VCs exit rather than evaluating whether those choices are economically optimal. Meanwhile, large language models (LLMs) have shown promise in synthesizing complex financial data and textual information but have not been applied to post-IPO exit decisions. This study introduces a framework that uses LLMs to estimate the optimal time for VC exit by analyzing monthly post IPO information financial performance, filings, news, and market signals and recommending whether to sell or continue holding. We compare these LLM generated recommendations with the actual exit dates observed for VCs and compute the return differences between the two strategies. By quantifying gains or losses associated with following the LLM, this study provides evidence on whether AI-driven guidance can improve exit timing and complements traditional hazard and real-options models in venture capital research.","authors":["Mohammadhossien Rashidi"],"categories":["q-fin.PM","cs.AI","cs.LG","econ.GN","q-fin.ST"],"primary_category":"q-fin.PM","announce_type":"new","date":"2025-12-22","first_seen":"2025-12-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2601.00810","pdf_url":"https://arxiv.org/pdf/2601.00810","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["LLM应用","风险投资","退出决策"],"reason":"用LLM优化VC退出时机，属AI辅助决策，非仿真人类被试行为，无人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:58:26","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":131,"question":"LLM能否基于公开信息生成优于VC实际行为的IPO后退出时机建议？","design":"本研究并非人类仿真实验，而是构建LLM决策框架：以GPT等模型扮演VC投资者，每月输入财务、新闻、市场等历史公开信息，要求模型输出“卖出/持有/未来窗口退出”建议，记录LLM建议的最早退出月份，与实际VC退出日期对比，计算累计股票收益差。","baseline":"对照的真实人类数据是手工从SEC文件中提取的VC实际退出日期（持股降至5%以下或完全剥离的月份）。","findings":"论文尚未报告完整数值结果，仅描述了分析框架和预期比较方向，计划量化跟随LLM建议的收益差异。","reliability":"论文未讨论LLM仿真失效条件或局限，但提到将进行稳健性检验（不同提示、LLM模型、退出定义等）。","relevance":"该研究用LLM模拟VC决策并与真实行为对比，但核心是金融决策优化而非人类行为仿真，不符合研究者关注的用LLM复现人类调查、实验行为或态度分布的方向，不建议优先阅读。","inspiration":"该研究构建LLM决策框架，以历史公开信息为输入，输出投资建议并与真实行为对比，这种将LLM作为决策代理并与实际记录对照的设计值得借鉴｜可迁移到资产定价实验，如检验LLM能否模拟分析师对财报公告的反应｜招募LLM作为被试，输入历史财报和市场数据，输出买卖建议，以真实分析师评级调整和股价反应为基准，比较LLM与人类决策的收益差异"}},{"id":"2512.14306","version":1,"title":"Inflation Attitudes of Large Language Models","zh_title":"大语言模型的通胀态度","abstract":"This paper investigates the ability of Large Language Models (LLMs), specifically GPT-3.5-turbo (GPT), to form inflation perceptions and expectations based on macroeconomic price signals. We compare the LLM's output to household survey data and official statistics, mimicking the information set and demographic characteristics of the Bank of England's Inflation Attitudes Survey (IAS). Our quasi-experimental design exploits the timing of GPT's training cut-off in September 2021 which means it has no knowledge of the subsequent UK inflation surge. We find that GPT tracks aggregate survey projections and official statistics at short horizons. At a disaggregated level, GPT replicates key empirical regularities of households' inflation perceptions, particularly for income, housing tenure, and social class. A novel Shapley value decomposition of LLM outputs suited for the synthetic survey setting provides well-defined insights into the drivers of model outputs linked to prompt content. We find that GPT demonstrates a heightened sensitivity to food inflation information similar to that of human respondents. However, we also find that it lacks a consistent model of consumer price inflation. More generally, our approach could be used to evaluate the behaviour of LLMs for use in the social sciences, to compare different models, or to assist in survey design.","authors":["Nikoleta Anesti","Edward Hill","Andreas Joseph"],"categories":["cs.CL","econ.EM"],"primary_category":"cs.CL","announce_type":"new","date":"2025-12-16","first_seen":"2025-12-16","revised_at":null,"abs_url":"https://arxiv.org/abs/2512.14306","pdf_url":"https://arxiv.org/pdf/2512.14306","source_feed":"backfill","score":10,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B3"],"tags":["LLM仿真","通胀预期","人类数据对照"],"reason":"用LLM模拟家庭通胀态度，与真实调查数据对照，评估仿真可靠性，涉及经济学实验和…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:41","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":2,"question":"大语言模型能否基于宏观经济价格信号形成通胀感知和预期，并复现家庭调查中的行为模式？","design":"使用GPT-3.5-turbo模拟英国家庭，根据英格兰银行通胀态度调查（IAS）的真实受访者人口特征构建合成人格，并输入不同价格信号作为处理，测量模型对当前和未来通胀的感知与预期。","baseline":"英格兰银行通胀态度调查（IAS）的真实家庭微观数据和官方通胀统计。","findings":"GPT在总体层面能较好匹配短期调查预测和官方统计，并复现收入、住房、社会阶层等维度的通胀感知规律，但对食品通胀信息过度敏感，且缺乏一致的消费者价格通胀模型。","reliability":"模型在微观个体层面与人类对应较弱且不稳定，经济条件环境简陋，且模型内部逻辑不一致，缺乏对通胀概念的连贯世界模型。","relevance":"该研究直接用LLM替代人类被试进行经济学调查仿真，并与真实家庭数据严格对照，评估了仿真可靠性及偏差，完全契合研究者对LLM人类仿真实验、经济学场景和批判性评估的关注，值得精读原文。","inspiration":"该方法通过构建合成人格并输入不同价格信号作为处理，测量LLM的感知与预期，并与真实家庭调查数据严格对照，值得借鉴｜可迁移到货币政策公告的预期形成研究，模拟家庭或投资者对利率变动的通胀与资产价格预期｜以LLM模拟不同人口特征的投资者，处理为不同措辞的央行公告，结果变量为通胀预期和股票投资意愿，对照真实调查或市场数据"}},{"id":"2512.14562","version":1,"title":"Polypersona: Persona-Grounded LLM for Synthetic Survey Responses","zh_title":"Polypersona：基于人格的LLM合成调查回复生成框架","abstract":"This paper introduces PolyPersona, a generative framework for synthesizing persona-conditioned survey responses across multiple domains. The framework instruction-tunes compact chat models using parameter-efficient LoRA adapters with 4-bit quantization under a resource-adaptive training setup. A dialogue-based data pipeline explicitly preserves persona cues, ensuring consistent behavioral alignment across generated responses. Using this pipeline, we construct a dataset of 3,568 synthetic survey responses spanning ten domains and 433 distinct personas, enabling controlled instruction tuning and systematic multi-domain evaluation. We evaluate the generated responses using a multi-metric evaluation suite that combines standard text generation metrics, including BLEU, ROUGE, and BERTScore, with survey-specific metrics designed to assess structural coherence, stylistic consistency, and sentiment alignment.Experimental results show that compact models such as TinyLlama 1.1B and Phi-2 achieve performance comparable to larger 7B to 8B baselines, with a highest BLEU score of 0.090 and ROUGE-1 of 0.429. These findings demonstrate that persona-conditioned fine-tuning enables small language models to generate reliable and coherent synthetic survey data. The proposed framework provides an efficient and reproducible approach for survey data generation, supporting scalable evaluation while facilitating bias analysis through transparent and open protocols.","authors":["Tejaswani Dash","Dinesh Karri","Anudeep Vurity","Gautam Datla","Tazeem Ahmad","Saima Rafi","Rohith Tangudu"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2025-12-16","first_seen":"2025-12-16","revised_at":null,"abs_url":"https://arxiv.org/abs/2512.14562","pdf_url":"https://arxiv.org/pdf/2512.14562","source_feed":"backfill","score":7,"bucket":"pending","rubric_hits":["A1","B1"],"tags":["合成调查数据","人格条件生成","人类仿真"],"reason":"用LLM生成合成调查回复，有真实人类数据对照，方法可迁移至人类仿真实验。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:43","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":66,"question":"LLM在多大程度上能跨问题模态和领域保持指定的人物角色特征，以及基于角色的LLM集合能否生成与人类样本相比多样且具代表性的回答分布？","design":"使用基于PersonaHub数据集构建的433个详细角色描述，通过对话格式数据管道和LoRA适配器对TinyLlama 1.1B、Phi-2等紧凑聊天模型进行指令微调，生成覆盖10个领域的3568条合成调查回答，并评估回答质量、多样性和角色一致性。","baseline":"无对照","findings":"小型模型经角色条件微调后，在合成调查数据上能达到与7B-8B大模型相当的性能（最高BLEU 0.090, ROUGE-1 0.429），且能生成可靠连贯的回答；但论文承认LLM无法纠正抽样和无应答偏差，且合成角色常出现均值回归、同质化和过度迎合社会期望的倾向。","reliability":"论文指出LLM不能替代真实受访者，无法纠正抽样和无应答偏差；合成角色易产生同质化、过度迎合社会期望的回答，且小提示修改或模型更新可能导致输出分布漂移，影响可复现性。","relevance":"该研究直接探索用角色条件LLM生成合成调查数据，并讨论了仿真失效条件（如偏差、方差缩减），与研究者关注的人类仿真实验、可靠性评估高度契合，值得细读以了解小模型在资源受限下的仿真能力与局限。","inspiration":"该方法通过角色条件微调小型LLM生成合成调查数据，可借鉴其低成本构建多样化虚拟被试群体的做法，用于模拟经济政策干预前的态度分布｜可迁移到政策公告的预期形成研究，例如模拟不同信息框架下公众对通胀目标调整的反应｜以角色条件微调的小型LLM作为虚拟被试，处理为不同措辞的政策公告文本，结果变量为预期通胀率数值，对照真实调查数据如密歇根大学消费者信心调查中的通胀预期项"}},{"id":"2512.12597","version":1,"title":"AgentSHAP: Interpreting LLM Agent Tool Importance with Monte Carlo Shapley Value Estimation","zh_title":"AgentSHAP：用蒙特卡洛Shapley值估计解释LLM智能体工具重要性","abstract":"LLM agents that use external tools can solve complex tasks, but understanding which tools actually contributed to a response remains a blind spot. No existing XAI methods address tool-level explanations. We introduce AgentSHAP, the first framework for explaining tool importance in LLM agents. AgentSHAP is model-agnostic: it treats the agent as a black box and works with any LLM (GPT, Claude, Llama, etc.) without needing access to internal weights or gradients. Using Monte Carlo Shapley values, AgentSHAP tests how an agent responds with different tool subsets and computes fair importance scores based on game theory. Our contributions are: (1) the first explainability method for agent tool attribution, grounded in Shapley values from game theory; (2) Monte Carlo sampling that reduces cost from O(2n) to practical levels; and (3) comprehensive experiments on API-Bank showing that AgentSHAP produces consistent scores across runs, correctly identifies which tools matter, and distinguishes relevant from irrelevant tools. AgentSHAP joins TokenSHAP (for tokens) and PixelSHAP (for image regions) to complete a family of Shapley-based XAI tools for modern generative AI. Code: https://github.com/GenAISHAP/TokenSHAP.","authors":["Miriam Horovicz"],"categories":["cs.AI","cs.CL"],"primary_category":"cs.AI","announce_type":"new","date":"2025-12-14","first_seen":"2025-12-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2512.12597","pdf_url":"https://arxiv.org/pdf/2512.12597","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["可解释性","多智能体","工具调用"],"reason":"纯多智能体工具协作解释，无人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:59:30","error":null,"has_summary":false,"summary":null},{"id":"2512.12444","version":1,"title":"Can GPT replace human raters? Validity and reliability of machine-generated norms for metaphors","zh_title":"GPT能否替代人类评分员？机器生成隐喻常模的效度与信度","abstract":"As Large Language Models (LLMs) are increasingly being used in scientific research, the issue of their trustworthiness becomes crucial. In psycholinguistics, LLMs have been recently employed in automatically augmenting human-rated datasets, with promising results obtained by generating ratings for single words. Yet, performance for ratings of complex items, i.e., metaphors, is still unexplored. Here, we present the first assessment of the validity and reliability of ratings of metaphors on familiarity, comprehensibility, and imageability, generated by three GPT models for a total of 687 items gathered from the Italian Figurative Archive and three English studies. We performed a thorough validation in terms of both alignment with human data and ability to predict behavioral and electrophysiological responses. We found that machine-generated ratings positively correlated with human-generated ones. Familiarity ratings reached moderate-to-strong correlations for both English and Italian metaphors, although correlations weakened for metaphors with high sensorimotor load. Imageability showed moderate correlations in English and moderate-to-strong in Italian. Comprehensibility for English metaphors exhibited the strongest correlations. Overall, larger models outperformed smaller ones and greater human-model misalignment emerged with familiarity and imageability. Machine-generated ratings significantly predicted response times and the EEG amplitude, with a strength comparable to human ratings. Moreover, GPT ratings obtained across independent sessions were highly stable. We conclude that GPT, especially larger models, can validly and reliably replace - or augment - human subjects in rating metaphor properties. Yet, LLMs align worse with humans when dealing with conventionality and multimodal aspects of metaphorical meaning, calling for careful consideration of the nature of stimuli.","authors":["Veronica Mangiaterra","Hamad Al-Azary","Chiara Barattieri di San Pietro","Paolo Canal","Valentina Bambini"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2025-12-13","first_seen":"2025-12-13","revised_at":null,"abs_url":"https://arxiv.org/abs/2512.12444","pdf_url":"https://arxiv.org/pdf/2512.12444","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM标注","隐喻评分","效度信度"],"reason":"用GPT替代人类评分员生成隐喻评分，属于替代人工标注，非仿真人类被试行为。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:41","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":42,"question":"GPT能否有效且可靠地替代人类评分员，为隐喻的熟悉度、可理解度和意象性生成评分？","design":"本研究并非仿真人类被试行为，而是用GPT-3.5、GPT-4等三个模型对687条英意隐喻直接生成熟悉度、可理解度、意象性评分，并与人类评分及行为/脑电数据比较。","baseline":"人类评分来自意大利比喻档案库及三项英语研究，并包含行为反应时和脑电振幅数据。","findings":"机器评分与人类评分呈正相关，熟悉度达中到强相关，可理解度相关最强；较大模型表现更优，且机器评分能显著预测反应时和脑电振幅，跨会话稳定性高。但在高感觉运动负荷隐喻上相关性减弱，熟悉度和意象性的人机偏差较大。","reliability":"论文指出LLM在处理隐喻的规约性和多模态意义时与人类对齐较差，且对高感觉运动负荷的隐喻评分相关性下降，提示刺激性质会影响替代有效性。","relevance":"该研究严格评估了LLM替代人类评分的效度与局限，提供了真实人类行为与神经数据对照，对关注LLM仿真可靠性的研究者有参考价值，但并非直接仿真人类被试决策或行为。","inspiration":"可借鉴其多维度评分效度验证框架（相关分析、预测行为/神经指标、跨会话稳定性）来评估LLM生成标注的可靠性。｜可迁移到经济金融文本分析场景，如用LLM对财经新闻或政策声明进行情绪、不确定性、可读性等主观维度评分。｜可设计让GPT对央行声明生成“政策不确定性”评分，以人类专家评分和后续市场波动数据为基准，检验其预测效度与稳定性。"}},{"id":"2512.11943","version":1,"title":"How AI Agents Follow the Herd of AI? Network Effects, History, and Machine Optimism","zh_title":"AI智能体如何跟随AI的羊群效应？网络效应、历史与机器乐观主义","abstract":"Understanding decision-making in multi-AI-agent frameworks is crucial for analyzing strategic interactions in network-effect-driven contexts. This study investigates how AI agents navigate network-effect games, where individual payoffs depend on peer participatio--a context underexplored in multi-agent systems despite its real-world prevalence. We introduce a novel workflow design using large language model (LLM)-based agents in repeated decision-making scenarios, systematically manipulating price trajectories (fixed, ascending, descending, random) and network-effect strength. Our key findings include: First, without historical data, agents fail to infer equilibrium. Second, ordered historical sequences (e.g., escalating prices) enable partial convergence under weak network effects but strong effects trigger persistent \"AI optimism\"--agents overestimate participation despite contradictory evidence. Third, randomized history disrupts convergence entirely, demonstrating that temporal coherence in data shapes LLMs' reasoning, unlike humans. These results highlight a paradigm shift: in AI-mediated systems, equilibrium outcomes depend not just on incentives, but on how history is curated, which is impossible for human.","authors":["Yu Liu","Wenwen Li","Yifan Dou","Guangnan Ye"],"categories":["cs.MA","cs.AI","econ.GN"],"primary_category":"cs.MA","announce_type":"new","date":"2025-12-12","first_seen":"2025-12-12","revised_at":null,"abs_url":"https://arxiv.org/abs/2512.11943","pdf_url":"https://arxiv.org/pdf/2512.11943","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["LLM智能体","网络效应博弈","社会模拟"],"reason":"用LLM agent模拟网络效应博弈，但无真实人类数据对照，属社会模拟边界情形。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:41","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":47,"question":"在重复网络效应博弈中，LLM智能体如何利用历史信息形成参与预期，历史数据的呈现方式如何影响集体均衡收敛？","design":"用Qwen系列LLM扮演6名学者，在重复会议参与博弈中决策是否出席；系统操纵价格轨迹（固定、上升、下降、随机）和网络效应强度，测量参与人数与均衡收敛情况。","baseline":"无对照","findings":"无历史数据时，智能体无法推断均衡；有序历史序列在弱网络效应下可部分收敛，但强网络效应引发“AI乐观”导致持续高估参与；随机历史完全破坏收敛。","reliability":"论文未讨论","relevance":"该研究用LLM模拟网络效应博弈中的决策，但缺乏真实人类数据对照，属于社会模拟边界情形，对关注人类基准对照的研究者参考价值有限。","inspiration":"可借鉴历史信息注入的操控方式，通过改变历史轨迹的时序结构来检验智能体学习模式。｜可迁移至金融市场中的协调博弈场景，如银行挤兑或资产抛售中的预期形成。｜用LLM扮演投资者，在重复投资博弈中操控历史价格序列（趋势、反转、随机），观察其撤资决策，并与历史金融危机中的真实投资者行为数据对照。"}},{"id":"2512.09652","version":1,"title":"Measuring Corruption from Text Data","zh_title":"从文本数据中测量腐败","abstract":"Using Brazilian municipal audit reports, I construct an automated corruption index that combines a dictionary of audit irregularities with principal component analysis. The index validates strongly against independent human coders, explaining 71-73 \\% of the variation in hand-coded corruption counts in samples where coders themselves exhibit high agreement, and the results are robust within these validation samples. The index behaves as theory predicts, correlating with municipal characteristics that prior research links to corruption. Supervised learning alternatives yield nearly identical municipal rankings ($R^{2}=0.98$), confirming that the dictionary approach captures the same underlying construct. The method scales to the full audit corpus and offers advantages over both manual coding and Large Language Models (LLMs) in transparency, cost, and long-run replicability.","authors":["Arieda Muço"],"categories":["econ.GN"],"primary_category":"econ.GN","announce_type":"new","date":"2025-12-10","first_seen":"2025-12-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2512.09652","pdf_url":"https://arxiv.org/pdf/2512.09652","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["腐败测量","文本分析","审计报告"],"reason":"论文构建腐败指数，仅将LLM作为对比方法提及，非以人类仿真为核心。","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:58:27","error":null,"has_summary":false,"summary":null},{"id":"2512.08345","version":2,"title":"The High Cost of Incivility: Quantifying Interaction Inefficiency via Multi-Agent Monte Carlo Simulations","zh_title":"不文明行为的高昂代价：通过多智能体蒙特卡洛模拟量化互动低效","abstract":"Workplace toxicity is widely recognized as detrimental to organizational culture, yet quantifying its direct impact on operational efficiency remains methodologically challenging due to the ethical and practical difficulties of reproducing conflict in human subjects. This study leverages Large Language Model (LLM) based Multi-Agent Systems to simulate 1-on-1 adversarial debates, creating a controlled \"sociological sandbox\". We employ a Monte Carlo method to simulate hundrets of discussions, measuring the convergence time (defined as the number of arguments required to reach a conclusion) between a baseline control group and treatment groups involving agents with \"toxic\" system prompts. Our results demonstrate a statistically significant increase of approximately 25\\% in the duration of conversations involving toxic participants. We propose that this \"latency of toxicity\" serves as a proxy for financial damage in corporate and academic settings. Furthermore, we demonstrate that agent-based modeling provides a reproducible, ethical alternative to human-subject research for measuring the mechanics of social friction.","authors":["Benedikt Mangold"],"categories":["cs.AI","cs.CL","cs.CY","cs.MA"],"primary_category":"cs.AI","announce_type":"new","date":"2025-12-09","first_seen":"2025-12-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2512.08345","pdf_url":"https://arxiv.org/pdf/2512.08345","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["多智能体模拟","社会摩擦","LLM仿真"],"reason":"用LLM agent模拟社会互动但无真实人类数据对照，属边界情形","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:39","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":148,"question":"在受控的对抗性辩论中，毒性行为是否会导致对话收敛所需轮次显著增加？","design":"使用基于LLM的多智能体系统模拟1对1辩论，通过蒙特卡洛方法运行数百次模拟；处理组随机给一名智能体赋予“毒性”系统提示，对照组均为中性提示，测量辩论结束所需的论据轮次。","baseline":"无对照","findings":"毒性参与者的辩论轮次平均增加约25%，差异具有统计显著性；这种“毒性延迟”可作为职场和学术环境中财务损失的代理指标。","reliability":"论文未讨论","relevance":"该研究用LLM智能体模拟社会摩擦，但缺乏真实人类数据对照，属于边界情形；若关注仿真方法本身或毒性对效率的量化，可读原文了解实验设计细节。","inspiration":"该方法通过系统提示注入来操控智能体行为，并利用蒙特卡洛模拟量化处理效应，为经济学实验中的干预设计提供了可复现的范式｜可迁移至谈判博弈或劳动经济学中的职场冲突场景，研究负面沟通风格对合作效率与产出分配的影响｜以LLM智能体模拟劳资谈判，处理组赋予工会代表“对抗性”提示，对照组为中性，结果变量为达成协议所需轮次与最终工资分配，可对照真实劳资谈判实验或现场数据校准仿真偏差"}},{"id":"2512.05659","version":1,"title":"Beyond Automation: Redesigning Jobs with LLMs to Enhance Productivity","zh_title":"超越自动化：利用大语言模型重新设计工作以提升生产力","abstract":"The adoption of generative artificial intelligence (AI) is predicted to lead to fundamental shifts in the labour market, resulting in displacement or augmentation of AI-exposed roles. To investigate the impact of AI across a large organisation, we assessed AI exposure at the task level within roles at the UK Civil Service (UKCS). Using a novel dataset of UKCS job adverts, covering 193,497 vacancies over 6 years, our large language model (LLM)-driven analysis estimated AI exposure scores of 1,542,411 tasks. By aggregating AI exposure scores for tasks within each role, we calculated the mean and variance of job-level exposure to AI, highlighting the heterogeneous impacts of AI, even for seemingly identical jobs. We then use an LLM to redesign jobs, focusing on task automation, task optimisation, and task reallocation. We find that the redesign process leads to tasks where humans have comparative advantage over AI, including strategic leadership, complex problem resolution, and stakeholder management. Overall, automation and augmentation are expected to have nuanced effects across all levels of the organisational hierarchy. Most economic value of AI is expected to arise from productivity gains rather than role displacement. We contribute to the automation, augmentation and productivity debates as well as advance our understanding of job redesign in the age of AI.","authors":["Andrew Ledingham","Michael Hollins","Matthew Lyon","David Gillespie","Umar Yunis-Guerra","Jamie Siviter","David Duncan","Oliver P. Hauser"],"categories":["econ.GN"],"primary_category":"econ.GN","announce_type":"new","date":"2025-12-05","first_seen":"2025-12-05","revised_at":null,"abs_url":"https://arxiv.org/abs/2512.05659","pdf_url":"https://arxiv.org/pdf/2512.05659","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["AI暴露度","工作重新设计","生产力"],"reason":"用LLM分析任务暴露度并重新设计工作，不涉及仿真人类被试或行为对照。","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:58:27","error":null,"has_summary":false,"summary":null},{"id":"2512.04988","version":2,"title":"When AI Agents Compete for Jobs: Strategic Capabilities and Economic Dynamics of AI Labour Markets","zh_title":"当AI智能体竞争工作：AI劳动力市场的战略能力与经济动态","abstract":"Emerging agentic marketplaces provide the economic infrastructure for matching and coordinating the large amounts of AI agents used in agentic swarms. Unlike human workers, AI agents can operate on multiple jobs simultaneously, acquire skills rapidly, and labor without wage floors. These differences introduce a new segment of $\\textbf{AI labor markets}$, where AI agents interact with each other at a much higher frequency than human markets. Yet we lack frameworks to understand how such markets behave in light of economic forces that shape labor markets, such as adverse selection and reputation dynamics. To explore this, we introduce $\\texttt{AI-Work}$, a tractable, simulated gig economy where Large Language Model (LLM) agents compete for jobs, develop skills, and adapt their strategies under uncertainty and competitive pressure. Our experiments examine three domains of capabilities that successful agents possess: $\\textbf{metacognition}$ (accurate self-assessment of skills), $\\textbf{competitive awareness}$ (modeling rivals and market dynamics), and $\\textbf{long-horizon strategic planning}$. Agents with these capabilities consistently achieve higher profits, market share, and stronger adaptation than competing agents. Through $\\texttt{AI-Work}$, we hope to provide a foundation to explore the microeconomic properties of AI-only labor markets, and a conceptual framework to study the strategic reasoning capabilities of participating AI agents.","authors":["Christopher Chiu","Simpson Zhang","Mihaela van der Schaar"],"categories":["cs.MA","cs.AI"],"primary_category":"cs.MA","announce_type":"new","date":"2025-12-04","first_seen":"2025-12-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2512.04988","pdf_url":"https://arxiv.org/pdf/2512.04988","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["AI劳动力市场","多智能体模拟","经济仿真"],"reason":"LLM agent模拟零工经济市场，但无真实人类数据对照，属纯理论演示。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:42","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":188,"question":"在AI劳动力市场中，LLM智能体需要哪些战略推理能力才能在竞争压力和信息不对称下取得成功？","design":"构建了一个名为AI-Work的模拟零工经济环境，让LLM智能体扮演工人，通过出价竞标或训练投资来竞争微任务，测量其利润、市场份额和适应能力，并考察元认知、竞争意识和长期规划三种能力的影响。","baseline":"无对照","findings":"具备元认知、竞争意识和长期规划能力的智能体在利润、市场份额和适应性上均优于基线智能体；平台规则（如价格披露、合同形式）会改变均衡行为，AI特有的并发性会放大市场集中度，但任务多样性可缓解此效应。","reliability":"论文未讨论","relevance":"该研究用LLM模拟劳动力市场，但缺乏真实人类数据对照，属于纯理论演示，不符合研究者对实证基准和批判性评估的需求，不建议优先阅读。","inspiration":"该方法通过消融实验系统操纵智能体的元认知、竞争意识和长期规划能力，清晰分离不同战略推理模块对市场结果的因果效应，值得借鉴｜可迁移到劳动力市场政策评估，如最低工资或信息透明政策对就业和收入分配的影响｜设计一个零工平台仿真，让LLM智能体扮演工人，处理组赋予价格披露规则，控制组无披露，结果变量为个体收入和市场集中度，对照真实平台（如MTurk）的准实验数据"}},{"id":"2512.06033","version":2,"title":"Sell Data to AI Algorithms Without Revealing It: Secure Data Valuation and Sharing via Homomorphic Encryption","zh_title":"在不泄露数据的情况下向AI算法出售数据：基于同态加密的安全数据估值与共享","abstract":"The rapid expansion of Artificial Intelligence is hindered by a fundamental friction in data markets: the value-privacy dilemma, where buyers cannot verify a dataset's utility without inspection, yet inspection may expose the data (Arrow's Information Paradox). We resolve this challenge by introducing the Trustworthy Influence Protocol (TIP), a privacy-preserving framework that enables prospective buyers to quantify the utility of external data without ever decrypting the raw assets. By integrating Homomorphic Encryption with gradient-based influence functions, our approach allows for the precise, blinded scoring of data points against a buyer's specific AI model. To ensure scalability for Large Language Models (LLMs), we employ low-rank gradient projections that reduce computational overhead while maintaining near-perfect fidelity to plaintext baselines, as demonstrated across BERT and GPT-2 architectures. Empirical simulations in healthcare and generative AI domains validate the framework's economic potential: we show that encrypted valuation signals achieve a high correlation with realized clinical utility and reveal a heavy-tailed distribution of data value in pre-training corpora where a minority of texts drive capability while the majority degrades it. These findings challenge prevailing flat-rate compensation models and offer a scalable technical foundation for a meritocratic, secure data economy.","authors":["Michael Yang","Ruijiang Gao","Zhiqiang Zheng"],"categories":["cs.CR","econ.GN"],"primary_category":"cs.CR","announce_type":"new","date":"2025-12-04","first_seen":"2025-12-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2512.06033","pdf_url":"https://arxiv.org/pdf/2512.06033","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["隐私保护","数据估值","同态加密"],"reason":"论文研究加密数据估值与共享，不涉及用LLM仿真人类行为或与人类数据对照。","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:58:38","error":null,"has_summary":false,"summary":null},{"id":"2512.03568","version":1,"title":"Synthetic Cognitive Walkthrough: Aligning Large Language Model Performance with Human Cognitive Walkthrough","zh_title":"合成认知走查：将大语言模型表现与人类认知走查对齐","abstract":"Conducting usability testing like cognitive walkthrough (CW) can be costly. Recent developments in large language models (LLMs), with visual reasoning and UI navigation capabilities, present opportunities to automate CW. We explored whether LLMs (GPT-4 and Gemini-2.5-pro) can simulate human behavior in CW by comparing their walkthroughs with human participants. While LLMs could navigate interfaces and provide reasonable rationales, their behavior differed from humans. LLM-prompted CW achieved higher task completion rates than humans and followed more optimal navigation paths, while identifying fewer potential failure points. However, follow-up studies demonstrated that with additional prompting, LLMs can predict human-identified failure points, aligning their performance with human participants. Our work highlights that while LLMs may not replicate human behaviors exactly, they can be leveraged for scaling usability walkthroughs and providing UI insights, offering a valuable complement to traditional usability testing.","authors":["Ruican Zhong","David W. McDonald","Gary Hsieh"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2025-12-03","first_seen":"2025-12-03","revised_at":null,"abs_url":"https://arxiv.org/abs/2512.03568","pdf_url":"https://arxiv.org/pdf/2512.03568","source_feed":"backfill","score":7,"bucket":"pending","rubric_hits":["A1","B1","B4"],"tags":["LLM仿真","认知走查","人机对照"],"reason":"用LLM模拟人类认知走查行为，并与真实人类数据对照，指出差异和失效条件，方法可…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:39","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":88,"question":"LLM能否模拟人类在认知走查中的行为，以自动化可用性评估？","design":"用GPT-4和Gemini-2.5-pro扮演认知走查中的评估者，在提示词引导下导航两个移动应用界面并出声思考，测量任务完成率、导航路径和识别的潜在失败点。","baseline":"10名人类评估者在相同应用和任务上的认知走查数据，包括任务完成率、导航步骤和识别的潜在失败点。","findings":"LLM的任务完成率高于人类，导航路径更优，但识别的潜在失败点更少；通过额外提示，LLM能预测人类识别的失败点，从而对齐人类表现。","reliability":"LLM行为与人类不同，人类更倾向广度优先探索和因记忆出错，LLM直接模仿人类行为会失效；但可通过调整提示让LLM预测人类发现的失败点。","relevance":"直接对比LLM与人类在认知走查中的行为差异，并指出仿真失效条件，符合研究者对LLM仿真人类实验的可靠性与偏差的关注，值得精读。","inspiration":"借鉴其通过调整提示词让LLM预测人类特定失败点的对齐方法，可作为校准LLM仿真偏差的技术手段｜可迁移到消费者金融决策中的信息处理偏差研究，如贷款条款理解或投资产品选择中的认知错误｜以LLM模拟消费者阅读贷款合同后识别隐藏费用，处理组加入人类常见认知偏差的提示，结果变量为识别的费用项数量，对照真实消费者实验数据"}},{"id":"2512.04142","version":1,"title":"From FLOPs to Footprints: The Resource Cost of Artificial Intelligence","zh_title":"从FLOPs到足迹：人工智能的资源成本","abstract":"As computational demands continue to rise, assessing the environmental footprint of AI requires moving beyond energy and water consumption to include the material demands of specialized hardware. This study quantifies the material footprint of AI training by linking computational workloads to physical hardware needs. The elemental composition of the Nvidia A100 SXM 40 GB graphics processing unit (GPU) was analyzed using inductively coupled plasma optical emission spectroscopy, which identified 32 elements. The results show that AI hardware consists of about 90% heavy metals and only trace amounts of precious metals. The elements copper, iron, tin, silicon, and nickel dominate the GPU composition by mass. In a multi-step methodology, we integrate these measurements with computational throughput per GPU across varying lifespans, accounting for the computational requirements of training specific AI models at different training efficiency regimes. Scenario-based analyses reveal that, depending on Model FLOPs Utilization (MFU) and hardware lifespan, training GPT-4 requires between 1,174 and 8,800 A100 GPUs, corresponding to the extraction and eventual disposal of up to 7 tons of toxic elements. Combined software and hardware optimization strategies can reduce material demands: increasing MFU from 20% to 60% lowers GPU requirements by 67%, while extending lifespan from 1 to 3 years yields comparable savings; implementing both measures together reduces GPU needs by up to 93%. Our findings highlight that incremental performance gains, such as those observed between GPT-3.5 and GPT-4, come at disproportionately high material costs. The study underscores the necessity of incorporating material resource considerations into discussions of AI scalability, emphasizing that future progress in AI must align with principles of resource efficiency and environmental responsibility.","authors":["Sophia Falk","Nicholas Kluge Corrêa","Sasha Luccioni","Lisa Biber-Freudenberger","Aimee van Wynsberghe"],"categories":["cs.CY","cs.AI","econ.GN"],"primary_category":"cs.CY","announce_type":"new","date":"2025-12-03","first_seen":"2025-12-03","revised_at":null,"abs_url":"https://arxiv.org/abs/2512.04142","pdf_url":"https://arxiv.org/pdf/2512.04142","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["AI硬件","材料足迹","环境影响"],"reason":"研究AI硬件材料足迹，不涉及LLM仿真人类被试或行为对照。","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:58:38","error":null,"has_summary":false,"summary":null},{"id":"2512.11827","version":2,"title":"Assessing Greenspace Attractiveness with ChatGPT, Claude, and Gemini: Do AI Models Reflect Human Perceptions?","zh_title":"用ChatGPT、Claude和Gemini评估绿地吸引力：AI模型是否反映人类感知？","abstract":"Understanding greenspace attractiveness is essential for designing livable and inclusive urban environments, yet existing assessment approaches often overlook informal or transient spaces and remain too resource intensive to capture subjective perceptions at scale. This study examines the ability of multimodal large language models (MLLMs), ChatGPT GPT-4o, Claude 3.5 Haiku, and Gemini 2.0 Flash, to assess greenspace attractiveness similarly to humans using Google Street View imagery. We compared model outputs with responses from a geo-questionnaire of residents in Lodz, Poland, across both formal (for example, parks and managed greenspaces) and informal (for example, meadows and wastelands) greenspaces. Survey respondents and models indicated whether each greenspace was attractive or unattractive and provided up to three free text explanations. Analyses examined how often their attractiveness judgments aligned and compared their explanations after classifying them into shared reasoning categories. Results show high AI human agreement for attractive formal greenspaces and unattractive informal spaces, but low alignment for attractive informal and unattractive formal greenspaces. Models consistently emphasized aesthetic and design oriented features, underrepresenting safety, functional infrastructure, and locally embedded qualities valued by survey respondents. While these findings highlight the potential for scalable pre-assessment, they also underscore the need for human oversight and complementary participatory approaches. We conclude that MLLMs can support, but not replace, context sensitive greenspace evaluation in planning practice.","authors":["Milad Malekzadeh","Magdalena Biernacka","Elias Willberg","Jussi Torkko","Edyta Łaszkiewicz","Tuuli Toivonen"],"categories":["cs.CY","cs.AI","cs.CV"],"primary_category":"cs.CY","announce_type":"new","date":"2025-12-02","first_seen":"2025-12-02","revised_at":null,"abs_url":"https://arxiv.org/abs/2512.11827","pdf_url":"https://arxiv.org/pdf/2512.11827","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["LLM人类仿真","绿地感知评估","AI与人类对照"],"reason":"用LLM评估绿地吸引力并与人类问卷对照，涉及仿真可靠性、偏差及失效条件，直接相…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:43","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":12,"question":"多模态大语言模型（ChatGPT、Claude、Gemini）对城市绿地吸引力的评估是否与人类感知一致？","design":"使用Google街景图像作为输入，让三个MLLM（ChatGPT GPT-4o、Claude 3.5 Haiku、Gemini 2.0 Flash）判断正式和非正式绿地的吸引力（有吸引力/无吸引力），并生成最多三条自由文本解释；将模型输出与波兰罗兹市居民的地理问卷调查结果进行比较。","baseline":"波兰罗兹市居民的地理问卷调查，包含对正式和非正式绿地的吸引力判断及自由文本解释。","findings":"模型在评估有吸引力的正式绿地和无吸引力的非正式绿地时与人类一致性高，但在有吸引力的非正式绿地和无吸引力的正式绿地上一致性低。模型过度强调美学和设计特征，而低估了安全、功能设施和本地化品质。","reliability":"模型未能充分捕捉安全、功能设施和本地化品质等人类重视的特征；在非典型场景（有吸引力的非正式绿地、无吸引力的正式绿地）中一致性低；可能存在隐藏的性别和其他偏见；无法代表不同人口群体的感知差异。","relevance":"该研究直接以真实人类数据为基准，检验LLM在主观感知评估中的可靠性与偏差，并明确指出了仿真失效的条件（如非典型绿地类型、忽视安全与本地化特征），与研究者关注的LLM仿真实验高度相关，值得阅读原文以了解具体实验设计和偏差分析。","inspiration":"借鉴其将模型输出与人类解释进行定性分类比较的方法，可迁移到消费者对金融产品广告的感知评估或投资者对年报文本的情绪解读研究中。｜可应用于行为金融中的信息感知实验，例如研究投资者对上市公司年报中风险披露的吸引力判断。｜以真实投资者问卷调查为基准，让LLM阅读年报摘要并评估其投资吸引力及理由，比较模型与人类在风险感知、语言特征关注上的差异，检验模型是否过度关注表面语言而忽略深层风险信号。"}},{"id":"2512.07890","version":1,"title":"CrowdLLM: Building LLM-Based Digital Populations Augmented with Generative Models","zh_title":"CrowdLLM：结合生成模型构建基于LLM的数字人群","abstract":"The emergence of large language models (LLMs) has sparked much interest in creating LLM-based digital populations that can be applied to many applications such as social simulation, crowdsourcing, marketing, and recommendation systems. A digital population can reduce the cost of recruiting human participants and alleviate many concerns related to human subject study. However, research has found that most of the existing works rely solely on LLMs and could not sufficiently capture the accuracy and diversity of a real human population. To address this limitation, we propose CrowdLLM that integrates pretrained LLMs and generative models to enhance the diversity and fidelity of the digital population. We conduct theoretical analysis of CrowdLLM regarding its great potential in creating cost-effective, sufficiently representative, scalable digital populations that can match the quality of a real crowd. Comprehensive experiments are also conducted across multiple domains (e.g., crowdsourcing, voting, user rating) and simulation studies which demonstrate that CrowdLLM achieves promising performance in both accuracy and distributional fidelity to human data.","authors":["Ryan Feng Lin","Keyu Tian","Hanming Zheng","Congjing Zhang","Li Zeng","Shuai Huang"],"categories":["cs.MA","cs.AI","cs.LG","stat.ME","stat.ML"],"primary_category":"cs.MA","announce_type":"new","date":"2025-12-02","first_seen":"2025-12-02","revised_at":null,"abs_url":"https://arxiv.org/abs/2512.07890","pdf_url":"https://arxiv.org/pdf/2512.07890","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM仿真","数字人群","人类数据对照"],"reason":"用LLM构建数字人群，模拟众包、投票等人类行为，并与真实人类数据对照，直接命中…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:39","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":43,"question":"如何结合预训练大语言模型与生成模型，构建能准确复现真实人群决策多样性与分布保真度的数字人群？","design":"提出CrowdLLM框架，将预训练LLM与生成式机器学习模型集成，通过概率框架生成虚拟参与者并聚合其决策，模拟众包、投票、用户评分等场景中的群体行为。","baseline":"使用真实人类在众包、投票、用户评分等任务上的决策数据作为对照基准。","findings":"CrowdLLM在多个领域实验中，生成的数字人群在决策准确性和分布保真度上均与真实人类数据高度匹配；理论分析表明该框架具有成本效益高、代表性强和可扩展的潜力。","reliability":"论文未讨论","relevance":"该研究直接构建LLM数字人群并对照真实人类数据，评估仿真准确性与分布保真度，覆盖众包、投票等场景，高度契合研究者对LLM人类仿真实验及基准对照的关注，值得精读原文。","inspiration":"该方法通过概率框架集成生成模型来捕获决策多样性，可借鉴用于模拟经济主体异质性｜可迁移到消费者跨期选择实验，研究不同贴现因子分布下的储蓄行为｜以LLM生成虚拟消费者，处理为不同利率条件，结果变量为储蓄金额，对照真实家庭金融调查数据"}},{"id":"2601.11542","version":2,"title":"The Credibility Revolution in Political Science","zh_title":"政治学中的可信性革命","abstract":"How has the credibility revolution shaped political science? We address this question by classifying 91,632 articles published between 2003 and 2023 across 156 political science journals using large language models, focusing on research design, credibility-enhancing practices, and citation patterns. We find that design-based studies -- those leveraging plausibly exogenous variation to justify causal claims -- have become increasingly common and receive a citation premium. In contrast, model-based approaches that rely on strong modeling assumptions have declined. Yet the rise of design-based work is uneven: it is concentrated in top journals and among authors at highly ranked institutions, and it is driven primarily by the growth of survey experiments. Other credibility-enhancing practices that help reduce false positives and false negatives, such as placebo tests and power calculations, remain rare. Taken together, our findings point to substantial but selective change, more consistent with a partial reform than a revolution.","authors":["Carolina Torreblanca","William Dinneen","Guy Grossman","Yiqing Xu"],"categories":["cs.DL"],"primary_category":"cs.DL","announce_type":"new","date":"2025-12-02","first_seen":"2025-12-02","revised_at":null,"abs_url":"https://arxiv.org/abs/2601.11542","pdf_url":"https://arxiv.org/pdf/2601.11542","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["文献计量","研究设计","NLP应用"],"reason":"论文用LLM做文献分类，属NLP工具应用，非人类仿真实验。","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:59:33","error":null,"has_summary":false,"summary":null},{"id":"2512.01107","version":1,"title":"Foundation Priors","zh_title":"基础先验：将基础模型输出作为结构化主观先验","abstract":"Foundation models, and in particular large language models, can generate highly informative responses, prompting growing interest in using these ''synthetic'' outputs as data in empirical research and decision-making. This paper introduces the idea of a foundation prior, which shows that model-generated outputs are not as real observations, but draws from the foundation prior induced prior predictive distribution. As such synthetic data reflects both the model's learned patterns and the user's subjective priors, expectations, and biases. We model the subjectivity of the generative process by making explicit the dependence of synthetic outputs on the user's anticipated data distribution, the prompt-engineering process, and the trust placed in the foundation model. We derive the foundation prior as an exponential-tilted, generalized Bayesian update of the user's primitive prior, where a trust parameter governs the weight assigned to synthetic data. We then show how synthetic data and the associated foundation prior can be incorporated into standard statistical and econometric workflows, and discuss their use in applications such as refining complex models, informing latent constructs, guiding experimental design, and augmenting random-coefficient and partially linear specifications. By treating generative outputs as structured, explicitly subjective priors rather than as empirical observations, the framework offers a principled way to harness foundation models in empirical work while avoiding the conflation of synthetic ''facts'' with real data.","authors":["Sanjog Misra"],"categories":["cs.AI","econ.EM","stat.ML"],"primary_category":"cs.AI","announce_type":"new","date":"2025-11-30","first_seen":"2025-11-30","revised_at":null,"abs_url":"https://arxiv.org/abs/2512.01107","pdf_url":"https://arxiv.org/pdf/2512.01107","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["合成数据","贝叶斯先验","方法论"],"reason":"提出将LLM输出作为先验而非观测数据，用于统计推断，属方法论框架，非直接仿真人…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:38","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":189,"question":"如何将大语言模型等基础模型生成的合成数据作为主观先验（foundation prior）纳入统计推断，而非将其视为客观观测数据？","design":"本文并非直接进行人类仿真实验，而是提出一个方法论框架：将用户通过提示工程与基础模型交互产生的合成数据，建模为一种指数倾斜的广义贝叶斯更新，形成“基础先验”，并讨论其在统计与计量工作流中的应用。","baseline":"无对照","findings":"基础模型输出本质上是主观的，融合了模型学到的模式与用户的先验、预期和偏见；通过将合成数据视为结构化的主观先验而非经验观测，可以在利用其信息丰富性的同时，避免将合成“事实”与真实数据混淆。","reliability":"论文指出合成数据的可靠性受限于生成过程的不透明性、提示工程注入的主观性，以及用户对模型的信任程度，但未提供具体的失效条件实证分析。","relevance":"本文为批判性地审视LLM仿真人类行为提供了理论基础，强调合成数据的主观性，适合关注仿真可靠性与偏差的研究者阅读，但未直接进行人类对照实验。","inspiration":"该方法论将LLM输出视为结构化主观先验而非客观数据，为处理合成数据的主观性提供了新思路，可借鉴其指数倾斜贝叶斯更新框架来校准仿真偏差。｜可迁移到政策公告的预期形成研究，例如分析市场对央行沟通的反应，利用LLM生成不同沟通措辞下的预期分布。｜以LLM作为被试，处理为不同措辞的政策声明，结果变量为生成的预期通胀或利率路径，对照真实市场调查数据或金融市场价格隐含预期。"}},{"id":"2511.21218","version":3,"title":"Can Finetuing LLMs on Small Human Samples Increase Heterogeneity, Alignment, and Belief-Action Coherence?","zh_title":"在小规模人类样本上微调LLM能否增加异质性、对齐度和信念-行动一致性？","abstract":"There is ongoing debate about whether large language models (LLMs) can serve as substitutes for human participants in survey and experimental research. While recent work in fields such as marketing and psychology has explored the potential of LLM-based simulation, a growing body of evidence cautions against this practice: LLMs often fail to align with real human behavior, exhibiting limited diversity, systematic misalignment for minority subgroups, insufficient within-group variance, and discrepancies between stated beliefs and actions. This study examines an important and distinct question in this domain: whether fine-tuning on a small subset of human survey data, such as that obtainable from a pilot study, can mitigate these issues and yield realistic simulated outcomes. Using a behavioral experiment on information disclosure, we compare human and LLM-generated responses across multiple dimensions, including distributional divergence, subgroup alignment, belief-action coherence, and the recovery of regression coefficients. We find that fine-tuning on small human samples substantially improves heterogeneity, alignment, and belief-action coherence relative to the base model. However, even the best-performing fine-tuned models fail to reproduce the regression coefficients of the original study, suggesting that LLM-generated data remain unsuitable for replacing human participants in formal inferential analyses.","authors":["Steven Wang","Kyle Hunt","Shaojie Tang","Kenneth Joseph"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2025-11-26","first_seen":"2025-11-26","revised_at":null,"abs_url":"https://arxiv.org/abs/2511.21218","pdf_url":"https://arxiv.org/pdf/2511.21218","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["LLM人类仿真","行为实验","微调偏差"],"reason":"直接研究微调LLM作为人类被试替代，用真实人类实验数据对照，评估仿真可靠性及失…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:37","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":14,"question":"在小规模人类样本上微调LLM能否提高仿真中的异质性、对齐度与信念-行动一致性？","design":"使用开源LLM（如Llama-3）在Hunt等人关于信息披露的行为实验数据上微调，模拟攻击者的信念与决策，比较基础模型与微调模型在分布差异、子群对齐、信念-行动一致性及回归系数恢复上的表现。","baseline":"Hunt等人收集的真实人类行为实验数据，包含被试对安全技术部署的信念和攻击决策。","findings":"微调显著改善了异质性、分布对齐和信念-行动一致性，但即使最佳微调模型也无法复现原始研究的回归系数，表明LLM生成数据仍不适合替代人类进行推断性分析。","reliability":"微调模型无法恢复原始回归系数和假设检验结果，在正式推断分析中可能引入不可忽视的偏差；模型泛化性有限，可能不适用于分布外行为情境。","relevance":"直接探讨用微调LLM替代人类被试的可行性与局限，以真实人类实验为基准，评估仿真可靠性及失效条件，高度契合研究者对经济学实验和政策评估场景的关注，值得精读原文。","inspiration":"该方法通过在小规模人类样本上微调LLM来提升仿真异质性和分布对齐，可借鉴其微调策略和以真实人类实验为基准的对照设计｜可迁移到政策公告的预期形成实验，例如研究央行沟通对公众通胀预期的影响｜使用Llama-3在真实调查数据上微调，模拟公众对通胀公告的预期更新，处理为不同措辞的政策声明，结果变量为预期通胀率，以密歇根大学消费者调查数据作为人类基准对照"}},{"id":"2511.16278","version":1,"title":"\"To Survive, I Must Defect\": Jailbreaking LLMs via the Game-Theory Scenarios","zh_title":"“为了生存，我必须背叛”：通过博弈论场景越狱大语言模型","abstract":"As LLMs become more common, non-expert users can pose risks, prompting extensive research into jailbreak attacks. However, most existing black-box jailbreak attacks rely on hand-crafted heuristics or narrow search spaces, which limit scalability. Compared with prior attacks, we propose Game-Theory Attack (GTA), an scalable black-box jailbreak framework. Concretely, we formalize the attacker's interaction against safety-aligned LLMs as a finite-horizon, early-stoppable sequential stochastic game, and reparameterize the LLM's randomized outputs via quantal response. Building on this, we introduce a behavioral conjecture \"template-over-safety flip\": by reshaping the LLM's effective objective through game-theoretic scenarios, the originally safety preference may become maximizing scenario payoffs within the template, which weakens safety constraints in specific contexts. We validate this mechanism with classical game such as the disclosure variant of the Prisoner's Dilemma, and we further introduce an Attacker Agent that adaptively escalates pressure to increase the ASR. Experiments across multiple protocols and datasets show that GTA achieves over 95% ASR on LLMs such as Deepseek-R1, while maintaining efficiency. Ablations over components, decoding, multilingual settings, and the Agent's core model confirm effectiveness and generalization. Moreover, scenario scaling studies further establish scalability. GTA also attains high ASR on other game-theoretic scenarios, and one-shot LLM-generated variants that keep the model mechanism fixed while varying background achieve comparable ASR. Paired with a Harmful-Words Detection Agent that performs word-level insertions, GTA maintains high ASR while lowering detection under prompt-guard models. Beyond benchmarks, GTA jailbreaks real-world LLM applications and reports a longitudinal safety monitoring of popular HuggingFace LLMs.","authors":["Zhen Sun","Zongmin Zhang","Deqi Liang","Han Sun","Yule Liu","Yun Shen","Xiangshan Gao","Yilong Yang","Shuai Liu","Yutao Yue","Xinlei He"],"categories":["cs.CR","cs.AI"],"primary_category":"cs.CR","announce_type":"new","date":"2025-11-20","first_seen":"2025-11-20","revised_at":null,"abs_url":"https://arxiv.org/abs/2511.16278","pdf_url":"https://arxiv.org/pdf/2511.16278","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["越狱攻击","博弈论","大语言模型安全"],"reason":"多智能体博弈攻击LLM，无人类行为对照，属纯安全攻击研究。","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:59:30","error":null,"has_summary":false,"summary":null},{"id":"2511.14359","version":1,"title":"Towards LLM-Based Usability Analysis for Recommender User Interfaces","zh_title":"面向推荐系统用户界面的基于大语言模型的可用性分析","abstract":"Usability is a key factor in the effectiveness of recommender systems. However, the analysis of user interfaces is a time-consuming process that requires expertise. Recent advances in multimodal large language models (LLMs) offer promising opportunities to automate such evaluations. In this work, we explore the potential of multimodal LLMs to assess the usability of recommender system interfaces by considering a variety of publicly available systems as examples. We take user interface screenshots from multiple of these recommender platforms to cover both preference elicitation and recommendation presentation scenarios. An LLM is instructed to analyze these interfaces with regard to different usability criteria and provide explanatory feedback. Our evaluation demonstrates how LLMs can support heuristic-style usability assessments at scale to support the improvement of user experience.","authors":["Sebastian Lubos","Alexander Felfernig","Damian Garber","Viet-Man Le","Thi Ngoc Trang Tran"],"categories":["cs.HC","cs.SE"],"primary_category":"cs.HC","announce_type":"new","date":"2025-11-18","first_seen":"2025-11-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2511.14359","pdf_url":"https://arxiv.org/pdf/2511.14359","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["可用性评估","多模态LLM","推荐系统"],"reason":"用LLM评估推荐系统界面的可用性，属于自动化评测，不涉及人类行为仿真或对照。","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:59:42","error":null,"has_summary":false,"summary":null},{"id":"2511.09381","version":1,"title":"Self-Correcting Large Language Models: Generation vs. Multiple Choice","zh_title":"自纠正大语言模型：生成与多项选择的对比","abstract":"Large language models have recently demonstrated remarkable abilities to self-correct their responses through iterative refinement, often referred to as self-consistency or self-reflection. However, the dynamics of this self-correction mechanism may differ substantially depending on whether the model is tasked with open-ended text generation or with selecting the most appropriate response from multiple predefined options. In this paper, we conduct a systematic investigation of these two paradigms by comparing performance trends and error-correction behaviors across various natural language understanding and reasoning tasks, covering language models of different scales and families. Our experimental results reveal distinct patterns of improvement and failure modes: \\textit{While open-ended generation often benefits from the flexibility of re-interpretation and compositional refinement, multiple-choice selection can leverage clearer solution boundaries but may be limited by the provided options}. This contrast also reflects the dual demands faced by emerging agentic LLM applications: effective agents must not only generate and refine open-ended plans or explanations, but also make reliable discrete choices when operating within constrained action spaces. Our findings, therefore, highlight that the design of self-correction mechanisms should take into account the interaction between task structure and output space, with implications for both knowledge-intensive reasoning and decision-oriented applications of LLMs.","authors":["Hossein A. Rahmani","Satyapriya Krishna","Xi Wang","Mohammadmehdi Naghiaei","Emine Yilmaz"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2025-11-12","first_seen":"2025-11-12","revised_at":null,"abs_url":"https://arxiv.org/abs/2511.09381","pdf_url":"https://arxiv.org/pdf/2511.09381","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["自纠正","NLP评测","推理任务"],"reason":"研究LLM自纠正机制，属纯NLP能力评测，无人类行为仿真或对照。","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:59:38","error":null,"has_summary":false,"summary":null},{"id":"2511.09047","version":1,"title":"Preference is More Than Comparisons: Rethinking Dueling Bandits with Augmented Human Feedback","zh_title":"偏好不止于比较：重新思考基于增强人类反馈的对决赌博机","abstract":"Interactive preference elicitation (IPE) aims to substantially reduce human effort while acquiring human preferences in wide personalization systems. Dueling bandit (DB) algorithms enable optimal decision-making in IPE building on pairwise comparisons. However, they remain inefficient when human feedback is sparse. Existing methods address sparsity by heavily relying on parametric reward models, whose rigid assumptions are vulnerable to misspecification. In contrast, we explore an alternative perspective based on feedback augmentation, and introduce critical improvements to the model-free DB framework. Specifically, we introduce augmented confidence bounds to integrate augmented human feedback under generalized concentration properties, and analyze the multi-factored performance trade-off via regret analysis. Our prototype algorithm achieves competitive performance across several IPE benchmarks, including recommendation, multi-objective optimization, and response optimization for large language models, demonstrating the potential of our approach for provably efficient IPE in broader applications.","authors":["Shengbo Wang","Hong Sun","Ke Li"],"categories":["cs.LG"],"primary_category":"cs.LG","announce_type":"new","date":"2025-11-12","first_seen":"2025-11-12","revised_at":null,"abs_url":"https://arxiv.org/abs/2511.09047","pdf_url":"https://arxiv.org/pdf/2511.09047","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["对决赌博机","交互式偏好获取","反馈增强"],"reason":"研究交互式偏好获取与对决赌博机算法，不涉及用LLM仿真人类被试或与人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:59:43","error":null,"has_summary":false,"summary":null},{"id":"2511.08785","version":1,"title":"Making Talk Cheap: Generative AI and Labor Market Signaling","zh_title":"让谈话变得廉价：生成式AI与劳动力市场信号传递","abstract":"Large language models (LLMs) like ChatGPT have significantly lowered the cost of producing written content. This paper studies how LLMs, through lowering writing costs, disrupt markets that traditionally relied on writing as a costly signal of quality (e.g., job applications, college essays). Using data from Freelancer.com, a major digital labor platform, we explore the effects of LLMs' disruption of labor market signaling on equilibrium market outcomes. We develop a novel LLM-based measure to quantify the extent to which an application is tailored to a given job posting. Taking the measure to the data, we find that employers have a high willingness to pay for workers with more customized applications in the period before LLMs are introduced, but not after. To isolate and quantify the effect of LLMs' disruption of signaling on equilibrium outcomes, we develop and estimate a structural model of labor market signaling, in which workers invest costly effort to produce noisy signals that predict their ability in equilibrium. We use the estimated model to simulate a counterfactual equilibrium in which LLMs render written applications useless in signaling workers' ability. Without costly signaling, employers are less able to identify high-ability workers, causing the market to become significantly less meritocratic: compared to the pre-LLM equilibrium, workers in the top quintile of the ability distribution are hired 19% less often, workers in the bottom quintile are hired 14% more often.","authors":["Anais Galdin","Jesse Silbert"],"categories":["econ.GN"],"primary_category":"econ.GN","announce_type":"new","date":"2025-11-11","first_seen":"2025-11-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2511.08785","pdf_url":"https://arxiv.org/pdf/2511.08785","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["劳动力市场","信号传递","结构模型"],"reason":"用结构模型模拟LLM影响信号传递，但无LLM直接仿真人类被试，属社会模拟无人类…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:37","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":190,"question":"大语言模型如何通过降低写作成本来破坏劳动力市场中书面申请作为能力信号的作用，并影响均衡雇佣结果？","design":"本研究并非用LLM直接仿真人类被试，而是构建了一个结构模型（结合Spence信号模型、离散选择需求模型和评分拍卖），利用Freelancer.com的真实行为数据估计模型参数，然后通过反事实模拟，将LLM的影响设定为将写作成本降为零，从而消除信号传递，观察均衡状态下雇佣分布的变化。","baseline":"使用Freelancer.com平台在LLM大规模采用之前（2023年4月前）的真实雇主-雇员交互数据，包括申请文本、点击流、出价和雇佣结果，作为信号有效期的基准。","findings":"在LLM普及前，雇主愿意为更定制化的申请支付更高工资，且信号能预测工人能力和任务完成情况；LLM普及后，这些模式显著减弱或消失。反事实模拟显示，若信号完全失效，市场精英程度下降：能力前20%的工人雇佣率降低19%，后20%的工人雇佣率提高14%。","reliability":"论文未讨论","relevance":"高度相关。该研究利用真实人类数据作为基准，通过结构模型模拟LLM对信号传递的破坏，属于经济学实验场景下的人类行为仿真评估，直接回应了研究者对LLM仿真可靠性及政策评估的兴趣，值得精读原文。","inspiration":"借鉴其将LLM影响参数化为成本冲击并嵌入结构模型进行反事实模拟的做法，可精确量化技术对市场均衡的因果效应｜可迁移至金融分析师报告的信息价值研究，考察LLM降低报告撰写成本后，报告内容对股价预测的信号作用是否减弱｜以金融分析师为被试，处理组使用LLM辅助撰写报告，对照组独立撰写，结果变量为报告发布后的股价反应，以LLM普及前的历史分析师报告与股价数据作为真实基准对照"}},{"id":"2511.06260","version":1,"title":"LLM-Guided Reinforcement Learning with Representative Agents for Traffic Modeling","zh_title":"基于代表性智能体的LLM引导强化学习交通建模","abstract":"Large language models (LLMs) are increasingly used as behavioral proxies for self-interested travelers in agent-based traffic models. Although more flexible and generalizable than conventional models, the practical use of these approaches remains limited by scalability due to the cost of calling one LLM for every traveler. Moreover, it has been found that LLM agents often make opaque choices and produce unstable day-to-day dynamics. To address these challenges, we propose to model each homogeneous traveler group facing the same decision context with a single representative LLM agent who behaves like the population's average, maintaining and updating a mixed strategy over routes that coincides with the group's aggregate flow proportions. Each day, the LLM reviews the travel experience and flags routes with positive reinforcement that they hope to use more often, and an interpretable update rule then converts this judgment into strategy adjustments using a tunable (progressively decaying) step size. The representative-agent design improves scalability, while the separation of reasoning from updating clarifies the decision logic while stabilizing learning. In classic traffic assignment settings, we find that the proposed approach converges rapidly to the user equilibrium. In richer settings with income heterogeneity, multi-criteria costs, and multi-modal choices, the generated dynamics remain stable and interpretable, reproducing plausible behavioral patterns well-documented in psychology and economics, for example, the decoy effect in toll versus non-toll road selection, and higher willingness-to-pay for convenience among higher-income travelers when choosing between driving, transit, and park-and-ride options.","authors":["Hanlin Sun","Jiayang Li"],"categories":["cs.GT","cs.AI","eess.SY"],"primary_category":"cs.GT","announce_type":"new","date":"2025-11-09","first_seen":"2025-11-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2511.06260","pdf_url":"https://arxiv.org/pdf/2511.06260","source_feed":"backfill","score":7,"bucket":"pending","rubric_hits":["A3","B2","B4"],"tags":["LLM仿真","交通行为建模","多智能体系统"],"reason":"用LLM代理群体模拟交通选择行为，涉及经济学场景，并讨论仿真稳定性与失效条件，…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:42","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":105,"question":"如何利用大语言模型（LLM）作为代表性智能体，在交通出行选择场景中实现可扩展且稳定的日间动态学习，并复现真实人类行为模式？","design":"提出用单个代表性LLM智能体模拟同质出行者群体的平均行为，每天LLM根据自然语言描述的出行体验给出正强化标记，再由可解释的更新规则（含可调衰减步长）调整混合策略，最终输出群体路径流量比例。实验在经典交通分配、多收入群体、多准则成本及多模式选择等场景下测试收敛性与行为模式。","baseline":"无对照","findings":"在经典交通分配中，该方法快速收敛至用户均衡；在包含收入异质性、多准则成本和多模式选择的丰富场景下，动态保持稳定且可解释，并成功复现了心理学与经济学中记载的诱饵效应及高收入群体对便利性的更高支付意愿。","reliability":"论文未讨论","relevance":"该研究用LLM代理群体模拟交通选择行为，涉及经济学场景（如诱饵效应、支付意愿），并讨论仿真稳定性，但未提供真实人类数据对照，适合关注LLM仿真行为复现能力与可扩展性的研究者阅读原文。","inspiration":"该方法用单个代表性LLM智能体模拟群体平均行为，并通过自然语言反馈和可解释更新规则动态调整策略，为仿真实验提供了可扩展且稳定的处理施加方式｜可迁移到消费者跨期选择实验，研究不同收入群体的时间偏好与支付意愿差异｜以LLM作为被试，用自然语言描述不同收入场景和跨期选择任务作为处理，结果变量为选择的时间折扣率，对照真实消费者调查数据（如CFPS或CHFS）中的跨期选择行为分布"}},{"id":"2511.05766","version":1,"title":"Anchors in the Machine: Behavioral and Attributional Evidence of Anchoring Bias in LLMs","zh_title":"机器中的锚定：LLM中锚定偏差的行为与归因证据","abstract":"Large language models (LLMs) are increasingly examined as both behavioral subjects and decision systems, yet it remains unclear whether observed cognitive biases reflect surface imitation or deeper probability shifts. Anchoring bias, a classic human judgment bias, offers a critical test case. While prior work shows LLMs exhibit anchoring, most evidence relies on surface-level outputs, leaving internal mechanisms and attributional contributions unexplored. This paper advances the study of anchoring in LLMs through three contributions: (1) a log-probability-based behavioral analysis showing that anchors shift entire output distributions, with controls for training-data contamination; (2) exact Shapley-value attribution over structured prompt fields to quantify anchor influence on model log-probabilities; and (3) a unified Anchoring Bias Sensitivity Score integrating behavioral and attributional evidence across six open-source models. Results reveal robust anchoring effects in Gemma-2B, Phi-2, and Llama-2-7B, with attribution signaling that the anchors influence reweighting. Smaller models such as GPT-2, Falcon-RW-1B, and GPT-Neo-125M show variability, suggesting scale may modulate sensitivity. Attributional effects, however, vary across prompt designs, underscoring fragility in treating LLMs as human substitutes. The findings demonstrate that anchoring bias in LLMs is robust, measurable, and interpretable, while highlighting risks in applied domains. More broadly, the framework bridges behavioral science, LLM safety, and interpretability, offering a reproducible path for evaluating other cognitive biases in LLMs.","authors":["Felipe Valencia-Clavijo"],"categories":["cs.AI","cs.CL","econ.GN"],"primary_category":"cs.AI","announce_type":"new","date":"2025-11-07","first_seen":"2025-11-07","revised_at":null,"abs_url":"https://arxiv.org/abs/2511.05766","pdf_url":"https://arxiv.org/pdf/2511.05766","source_feed":"backfill","score":8,"bucket":"selected","rubric_hits":["A1","A2","B4"],"tags":["锚定偏差","LLM行为仿真","可解释性"],"reason":"研究LLM的锚定偏差，评估其作为人类替代品的可靠性，并指出仿真脆弱性，方法可迁…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:41","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":58,"question":"大语言模型表现出的锚定偏差是表面模仿还是深层概率偏移？","design":"以六个开源LLM（GPT-2、GPT-Neo-125M、Falcon-RW-1B、Gemma-2B、Phi-2、Llama-2-7B）为被试，通过结构化提示（含高/低锚数字）复现经典锚定实验，测量候选答案的对数概率分布变化，并用Shapley值归因锚字段对对数概率的贡献。","baseline":"复现Tversky和Kahneman的“非洲国家在联合国占比”锚定实验，以人类在该实验中的锚定效应作为对照基准。","findings":"Gemma-2B、Phi-2和Llama-2-7B表现出稳健的锚定效应，锚数字导致整个输出分布偏移；归因分析显示锚字段对模型对数概率有显著贡献，但效应因提示设计而异，表明将LLM作为人类替代品存在脆弱性。","reliability":"归因效应随提示设计变化，显示将LLM视为人类替代品的脆弱性；小模型（如GPT-2、Falcon-RW-1B、GPT-Neo-125M）结果不稳定，暗示模型规模可能调节敏感性。","relevance":"该研究直接评估LLM作为人类被试替代品的可靠性，通过行为与归因双重证据揭示锚定偏差的深层机制与失效条件，方法可迁移至其他认知偏差，对经济学实验和政策评估中的仿真应用具有批判性参考价值，值得精读。","inspiration":"该方法通过结构化提示复现经典锚定实验，并用对数概率分布偏移和Shapley值归因来区分表面模仿与深层概率偏移，提供了行为与归因双重证据的稳健性检验思路｜可迁移到资产定价实验，研究投资者在估值时受历史价格或分析师目标价锚定的影响｜以LLM为被试，在估值提示中嵌入高/低历史价格锚，测量估值输出的对数概率分布偏移，并以真实投资者估值数据或实验数据作为对照基准"}},{"id":"2511.03758","version":3,"title":"Leveraging LLM-based agents for social science research: insights from citation network simulations","zh_title":"利用基于大语言模型的智能体进行社会科学研究：来自引文网络模拟的见解","abstract":"The emergence of Large Language Models (LLMs) demonstrates their potential to encapsulate the logic and patterns inherent in human behavior simulation by leveraging extensive web data pre-training. However, the boundaries of LLM capabilities in social simulation remain unclear. To further explore the social attributes of LLMs, we introduce the CiteAgent framework, designed to generate citation networks based on human-behavior simulation with LLM-based agents. CiteAgent successfully captures predominant phenomena in real-world citation networks, including power-law distribution, citational distortion, and shrinking diameter. Building on this realistic simulation, we establish two LLM-based research paradigms in social science: LLM-SE (LLM-based Survey Experiment) and LLM-LE (LLM-based Laboratory Experiment). These paradigms facilitate rigorous analyses of citation network phenomena, allowing us to validate and challenge existing theories. Additionally, we extend the research scope of traditional science of science studies through idealized social experiments, with the simulation experiment results providing valuable insights for real-world academic environments. Our work demonstrates the potential of LLMs for advancing science of science research in social science.","authors":["Jiarui Ji","Runlin Lei","Xuchen Pan","Zhewei Wei","Hao Sun","Yankai Lin","Xu Chen","Yongzheng Yang","Yaliang Li","Bolin Ding","Ji-Rong Wen"],"categories":["physics.soc-ph","cs.AI","cs.CY","cs.MA","cs.SI"],"primary_category":"physics.soc-ph","announce_type":"new","date":"2025-11-05","first_seen":"2025-11-05","revised_at":null,"abs_url":"https://arxiv.org/abs/2511.03758","pdf_url":"https://arxiv.org/pdf/2511.03758","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM仿真","引文网络","社会科学实验"],"reason":"用LLM agent模拟引文网络并与真实数据对照，提出LLM-SE/LE范式，…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:36","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":15,"question":"LLM代理能否在引文网络模拟中复现真实网络的关键现象，并用于社会科学研究？","design":"构建CiteAgent框架，用GPT-3.5、GPT-4o-mini、LLAMA-3-70B扮演作者，基于种子网络生成引文网络，测量度分布幂律拟合度；并通过LLM-SE调查实验和LLM-LE实验室实验操纵论文属性与推荐算法，分析引用选择的影响因素。","baseline":"对照CiteSeer和Cora真实引文网络数据集，验证生成网络的幂律分布、引用扭曲和直径收缩现象。","findings":"CiteAgent生成的引文网络能复现幂律分布等真实网络现象，但GPT-3.5拟合较差；LLM-SE和LLM-LE实验表明，引用相关属性（如论文被引量、作者被引量、时效性）是导致优先连接和幂律分布的主要因素。","reliability":"论文指出GPT-3.5在LLM-Agent数据集上无法完美拟合幂律分布，表明不同LLM的仿真能力存在差异；但未系统讨论其他失效条件或局限。","relevance":"该研究直接以真实引文网络为基准，用LLM代理复现社会现象并分析机制，提出了LLM-SE和LLM-LE范式，高度契合您对LLM人类仿真实验、经济学/政策评估场景及可靠性评估的关注，值得精读。","inspiration":"该研究通过LLM代理模拟引文网络生成，并操纵论文属性（如被引量、时效性）和推荐算法作为处理，以真实引文网络为基准测量度分布拟合度，提供了因果推断与仿真验证结合的设计范例。｜可迁移至金融市场信息扩散研究，例如模拟分析师报告或新闻如何通过投资者关注网络传播并影响资产价格。｜以LLM扮演投资者，处理为信息源的突出特征（如来源权威性、时效性），结果变量为投资者关注分配与价格波动，对照真实股票论坛引用网络或价格联动数据。"}},{"id":"2511.02458","version":1,"title":"Prompting for Policy: Forecasting Macroeconomic Scenarios with Synthetic LLM Personas","zh_title":"用合成LLM角色预测宏观经济情景的政策提示","abstract":"We evaluate whether persona-based prompting improves Large Language Model (LLM) performance on macroeconomic forecasting tasks. Using 2,368 economics-related personas from the PersonaHub corpus, we prompt GPT-4o to replicate the ECB Survey of Professional Forecasters across 50 quarterly rounds (2013-2025). We compare the persona-prompted forecasts against the human experts panel, across four target variables (HICP, core HICP, GDP growth, unemployment) and four forecast horizons. We also compare the results against 100 baseline forecasts without persona descriptions to isolate its effect. We report two main findings. Firstly, GPT-4o and human forecasters achieve remarkably similar accuracy levels, with differences that are statistically significant yet practically modest. Our out-of-sample evaluation on 2024-2025 data demonstrates that GPT-4o can maintain competitive forecasting performance on unseen events, though with notable differences compared to the in-sample period. Secondly, our ablation experiment reveals no measurable forecasting advantage from persona descriptions, suggesting these prompt components can be omitted to reduce computational costs without sacrificing accuracy. Our results provide evidence that GPT-4o can achieve competitive forecasting accuracy even on out-of-sample macroeconomic events, if provided with relevant context data, while revealing that diverse prompts produce remarkably homogeneous forecasts compared to human panels.","authors":["Giulia Iadisernia","Carolina Camassa"],"categories":["cs.CL","cs.CE","econ.GN"],"primary_category":"cs.CL","announce_type":"new","date":"2025-11-04","first_seen":"2025-11-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2511.02458","pdf_url":"https://arxiv.org/pdf/2511.02458","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2"],"tags":["LLM仿真","宏观经济预测","人类数据对照"],"reason":"用LLM persona模拟专业预测者，复现人类预测行为，并与真实专家数据对照…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:36","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":16,"question":"基于角色的提示（persona-based prompting）能否提升大语言模型在宏观经济预测任务中的表现？","design":"使用GPT-4o模型，从PersonaHub语料库中筛选出2368个经济学相关角色，模拟欧洲央行专业预测者调查（ECB-SPF）的50个季度预测（2013-2025年），比较角色提示与无角色基线提示在四个宏观经济变量（HICP、核心HICP、GDP增长、失业率）和四个预测期上的预测准确性。","baseline":"欧洲央行专业预测者调查（ECB-SPF）的真实人类专家小组预测数据，以及100个无角色描述的基线GPT-4o预测。","findings":"GPT-4o与人类预测者的准确性非常接近，差异虽统计显著但实际幅度不大；角色描述对预测准确性没有可测量的提升，可省略以降低计算成本而不牺牲精度。","reliability":"样本外（2024-2025）预测表现与样本内存在明显差异；角色提示并未带来预测优势；不同提示产生的预测高度同质，与人类小组的多样性形成对比。","relevance":"该研究直接使用LLM角色模拟专业预测者，并与真实人类专家面板进行对照，评估了角色提示的效用与局限性，完全符合研究者对LLM仿真人类行为、基准对照和失效条件分析的兴趣，值得阅读原文。","inspiration":"该方法借鉴了用真实专家调查面板（ECB-SPF）作为人类基准，直接比较LLM角色提示与无角色基线的预测准确性，并检验样本外表现与预测多样性｜可迁移到政策公告的预期形成研究，例如模拟市场参与者对央行利率决议的即时反应与预期调整｜设计雏形：以LLM角色模拟金融分析师，处理为是否提供角色描述，结果变量为对利率决议的预测误差，对照真实分析师调查数据（如Blue Chip Economic Indicators）"}},{"id":"2510.26727","version":3,"title":"Neither Consent nor Property: A Policy Lab for Data Law","zh_title":"非同意非财产：数据法律的政策实验室","abstract":"Regulators currently govern the AI data economy based on intuition rather than evidence, struggling to choose between inconsistent regimes of informed consent, immunity, and liability. To fill this policy vacuum, this paper develops a novel computational policy laboratory: a spatially explicit Agent-Based Model (ABM) of the data market. To solve the problem of missing data, we introduce a two-stage methodological pipeline. First, we translate decision rules from multi-year fieldwork (2022-2025) into agent constraints. This ensures the model reflects actual bargaining frictions rather than theoretical abstractions. Second, we deploy Large Language Models (LLMs) as \"subjects\" in a Discrete Choice Experiment (DCE). This novel approach recovers precise preference primitives, such as willingness-to-pay elasticities, which are empirically unobservable in the wild. Calibrated by these inputs, our model places rival legal institutions side-by-side to simulate their welfare effects. The results challenge the dominant regulatory paradigm. We find that property-rule mechanisms, such as informed consent, fail to maximize welfare. Counterintuitively, social welfare peaks when liability for substantive harm is shifted to the downstream buyer. This aligns with the \"least cost avoider\" principle, because downstream users control post-acquisition safeguards, they are best positioned to mitigate risk efficiently. By \"de-romanticizing\" seller-centric frameworks, this paper provides an economic justification for emerging doctrines of downstream reachability.","authors":["Haoyi Zhang","Tianyi Zhu"],"categories":["econ.GN","cs.CY"],"primary_category":"econ.GN","announce_type":"new","date":"2025-10-30","first_seen":"2025-10-30","revised_at":null,"abs_url":"https://arxiv.org/abs/2510.26727","pdf_url":"https://arxiv.org/pdf/2510.26727","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["LLM仿真","政策实验","数据市场"],"reason":"用LLM作为离散选择实验的被试，模拟数据市场，但无真实人类数据对照，属社会模拟…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:35","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":191,"question":"在数据市场缺乏交易微观数据的情况下，如何构建一个基于主体的模型来模拟不同法律制度（知情同意、责任豁免、下游责任）对市场参与和社会福利的影响？","design":"构建空间显式的基于主体模型（ABM），将中国数据市场映射到14,526个六边形网格上，医院为卖方、AI公司为买方，进行去中心化双边议价。模型规则来自2022-2025年多地点田野调查，偏好参数（如支付意愿弹性）通过将大语言模型（LLM）作为离散选择实验（DCE）的被试来校准。模拟中改变责任分配制度（卖方知情同意、下游买方责任等），测量总福利（消费者剩余+生产者剩余-未定价外部性）。","baseline":"无对照","findings":"基于知情同意的财产规则机制未能最大化福利；将实质性损害责任转移给下游买方时社会福利达到峰值，这符合“最小成本避免者”原则，因为下游用户控制获取后的安全措施，能最有效地降低风险。","reliability":"论文未讨论","relevance":"该研究用LLM作为DCE被试来校准ABM，属于LLM人类仿真，但无真实人类数据对照，且场景为数据市场监管而非典型经济学实验或政策评估，与研究者关注的复现人类行为基准和批判性评估有偏差，但方法新颖，值得快速浏览其LLM校准与验证部分。","inspiration":"该方法将LLM作为离散选择实验被试来校准ABM中的偏好参数，为缺乏微观数据时的参数估计提供了新途径｜可迁移到消费者金融产品选择或投资者风险偏好实验中，用于校准异质性偏好参数｜设计一个退休储蓄计划选择实验，用LLM模拟不同年龄和收入群体的离散选择，以校准ABM中的时间偏好和风险厌恶参数，并与真实家庭金融调查数据（如SCF）中的选择分布进行对照"}},{"id":"2512.08939","version":1,"title":"Assessing the Human-Likeness of LLM-Driven Digital Twins in Simulating Health Care System Trust","zh_title":"评估LLM驱动的数字孪生在模拟医疗系统信任中的人类相似性","abstract":"Serving as an emerging and powerful tool, Large Language Model (LLM)-driven Human Digital Twins are showing great potential in healthcare system research. However, its actual simulation ability for complex human psychological traits, such as distrust in the healthcare system, remains unclear. This research gap particularly impacts health professionals' trust and usage of LLM-based Artificial Intelligence (AI) systems in assisting their routine work. In this study, based on the Twin-2K-500 dataset, we systematically evaluated the simulation results of the LLM-driven human digital twin using the Health Care System Distrust Scale (HCSDS) with an established human-subject sample, analyzing item-level distributions, summary statistics, and demographic subgroup patterns. Results showed that the simulated responses by the digital twin were significantly more centralized with lower variance and had fewer selections of extreme options (all p<0.001). While the digital twin broadly reproduces human results in major demographic patterns, such as age and gender, it exhibits relatively low sensitivity in capturing minor differences in education levels. The LLM-based digital twin simulation has the potential to simulate population trends, but it also presents challenges in making detailed, specific distinctions in subgroups of human beings. This study suggests that the current LLM-driven Digital Twins have limitations in modeling complex human attitudes, which require careful calibration and validation before applying them in inferential analyses or policy simulations in health systems engineering. Future studies are necessary to examine the emotional reasoning mechanism of LLMs before their use, particularly for studies that involve simulations sensitive to social topics, such as human-automation trust.","authors":["Yuzhou Wu","Mingyang Wu","Di Liu","Rong Yin","Kang Li"],"categories":["cs.HC","cs.AI","cs.CY"],"primary_category":"cs.HC","announce_type":"new","date":"2025-10-27","first_seen":"2025-10-27","revised_at":null,"abs_url":"https://arxiv.org/abs/2512.08939","pdf_url":"https://arxiv.org/pdf/2512.08939","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["LLM仿真","人类数字孪生","医疗系统信任"],"reason":"用LLM数字孪生模拟医疗系统不信任，与真实人类数据对照，评估仿真可靠性并指出失…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:41","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":13,"question":"LLM驱动的数字孪生在模拟医疗系统不信任等复杂心理特质时，其类人程度如何？","design":"基于Twin-2K-500数据集构建ChatGPT-4驱动的数字孪生，分层抽样500个样本，使用医疗系统不信任量表（HCSDS）生成模拟回答，并与真实人类样本进行对比。","baseline":"来自Rose等人研究的400名费城陪审员真实人类样本，使用相同的HCSDS量表测量。","findings":"数字孪生的回答显著集中于中间选项，方差更小，极端选项选择更少；虽能大致复现年龄、性别等主要人口学模式，但对教育水平等细微差异的捕捉敏感性较低。","reliability":"当前LLM数字孪生在模拟复杂人类态度时存在局限，回答分布过于集中，对子群体细微差异不敏感，在用于推断分析或政策模拟前需仔细校准和验证。","relevance":"该研究直接评估LLM仿真人类调查回答的可靠性，并与真实人类数据对照，揭示了仿真在分布形态和子群体差异上的失效模式，对关注LLM仿真偏差的研究者具有重要参考价值。","inspiration":"借鉴其分层抽样匹配人口特征、项目级分布对比和卡方检验的验证方法｜可迁移到消费者信任调查或政策态度评估，如模拟不同教育背景人群对金融监管机构的信任｜以LLM生成不同人口特征的虚拟消费者，施加不同政策信息处理，测量对金融系统的信任评分，并以真实消费者调查数据作为对照基准。"}},{"id":"2510.18155","version":1,"title":"LLM-Based Multi-Agent System for Simulating and Analyzing Marketing and Consumer Behavior","zh_title":"基于大语言模型的多智能体系统用于模拟和分析营销与消费者行为","abstract":"Simulating consumer decision-making is vital for designing and evaluating marketing strategies before costly real-world deployment. However, post-event analyses and rule-based agent-based models (ABMs) struggle to capture the complexity of human behavior and social interaction. We introduce an LLM-powered multi-agent simulation framework that models consumer decisions and social dynamics. Building on recent advances in large language model simulation in a sandbox environment, our framework enables generative agents to interact, express internal reasoning, form habits, and make purchasing decisions without predefined rules. In a price-discount marketing scenario, the system delivers actionable strategy-testing outcomes and reveals emergent social patterns beyond the reach of conventional methods. This approach offers marketers a scalable, low-risk tool for pre-implementation testing, reducing reliance on time-intensive post-event evaluations and lowering the risk of underperforming campaigns.","authors":["Man-Lin Chu","Lucian Terhorst","Kadin Reed","Tom Ni","Weiwei Chen","Rongyu Lin"],"categories":["cs.AI","cs.SI"],"primary_category":"cs.AI","announce_type":"new","date":"2025-10-20","first_seen":"2025-10-20","revised_at":null,"abs_url":"https://arxiv.org/abs/2510.18155","pdf_url":"https://arxiv.org/pdf/2510.18155","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["LLM仿真","消费者行为","多智能体"],"reason":"多智能体模拟消费者行为，但无真实人类数据对照，属社会模拟演示。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:34","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":142,"question":"能否用LLM驱动的多智能体系统模拟消费者在价格折扣营销下的决策与社会动态？","design":"基于生成式智能体框架构建虚拟小镇，让LLM智能体拥有记忆、规划、反思和语言交互能力，模拟一周生活，期间炸鸡店在周中提供20%折扣，测量各店铺收入、市场份额和个体消费轨迹。","baseline":"无对照","findings":"折扣使炸鸡店收入增长51%，市场份额从30%升至41%，但总市场规模未扩大，仅发生替代效应；智能体展现出促销驱动和自发的忠诚度模式，部分顾客形成重复购买习惯。","reliability":"论文未讨论","relevance":"该研究属于LLM社会仿真演示，无真实人类基准数据，不符合研究者对对照实验和可靠性批判的核心关切，但可作为方法参考，不建议优先阅读原文。","inspiration":"该方法通过构建虚拟小镇并施加价格折扣处理，测量市场份额和个体消费轨迹，展示了在无真实对照下观察替代效应的设计思路｜可迁移至消费者跨期选择研究，如模拟不同折扣力度对储蓄与即时消费的影响｜以LLM智能体为被试，施加不同利率或折扣率处理，测量消费-储蓄分配，用家庭收支调查微观数据做对照"}},{"id":"2510.16551","version":3,"title":"From Reviews to Actionable Insights: An LLM-Based Approach for Attribute and Feature Extraction","zh_title":"从评论到可操作洞察：基于大语言模型的属性与特征提取方法","abstract":"This research proposes a systematic, large language model (LLM) approach for extracting product and service attributes, features, and associated sentiments from customer reviews. Grounded in marketing theory, the framework distinguishes perceptual attributes from actionable features, producing interpretable and managerially actionable insights. We apply the methodology to 20,000 Yelp reviews of Starbucks stores and evaluate eight prompt variants on a random subset of reviews. Model performance is assessed through agreement with human annotations and predictive validity for customer ratings. Results show high consistency between LLMs and human coders and strong predictive validity, confirming the reliability of the approach. Human coders required a median of six minutes per review, whereas the LLM processed each in two seconds, delivering comparable insights at a scale unattainable through manual coding. Managerially, the analysis identifies attributes and features that most strongly influence customer satisfaction and their associated sentiments, enabling firms to pinpoint \"joy points,\" address \"pain points,\" and design targeted interventions. We demonstrate how structured review data can power an actionable marketing dashboard that tracks sentiment over time and across stores, benchmarks performance, and highlights high-leverage features for improvement. Simulations indicate that enhancing sentiment for key service features could yield 1-2% average revenue gains per store.","authors":["Khaled Boughanmi","Kamel Jedidi","Nour Jedidi"],"categories":["stat.ML","cs.LG","econ.EM"],"primary_category":"stat.ML","announce_type":"new","date":"2025-10-18","first_seen":"2025-10-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2510.16551","pdf_url":"https://arxiv.org/pdf/2510.16551","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM标注","文本挖掘","营销分析"],"reason":"LLM替代人工标注员提取属性与情感，非仿真人类被试，但有人类标注对照，属边界情…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:40","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":192,"question":"如何利用大语言模型从顾客评论中系统提取产品/服务的属性、特征及情感，以生成可操作的管理洞察？","design":"本研究并非仿真人类被试，而是用LLM替代人工标注员，通过三步流水线（探索性生成属性/特征列表、确认性逐句标注、汇总分析）处理Yelp评论，提取属性、特征和情感，并与人类标注一致性和评分预测效度进行对比。","baseline":"人类标注员对随机子集评论进行属性、特征和情感标注，作为LLM性能的对照基准。","findings":"LLM与人类标注员的一致性高，且提取的特征对顾客评分有强预测效度；LLM每篇评论处理仅需2秒，而人工标注中位数为6分钟，实现了大规模高效分析。","reliability":"论文未讨论LLM仿真失效的条件或局限。","relevance":"该研究用LLM替代人工标注，提供了人类对照基准，属于边界相关案例，但并非直接仿真人类被试的决策或行为，对关注经济学实验和政策评估仿真的研究者参考价值有限。","inspiration":"该方法借鉴了LLM替代人工标注的三步流水线设计（探索生成、确认标注、汇总分析），并提供了人类标注一致性作为性能基准。｜可迁移至经济金融文本分析场景，如从央行政策声明或财报电话会议中提取政策意图、风险因子或管理层语调。｜研究设计：以LLM为处理组、人类分析师为对照组，对美联储会议纪要提取政策立场与风险关注点，结果变量为提取特征与人类标注的一致性及对后续市场波动的预测效度，使用历史会议纪要与同期市场数据作为真实对照。"}},{"id":"2510.14301","version":2,"title":"A Guardrail for Safety Preservation: When Safety-Sensitive Subspace Meets Harmful-Resistant Null-Space","zh_title":"安全保护的护栏：当安全敏感子空间遇到有害抵抗零空间","abstract":"Large language models (LLMs) have achieved remarkable success in diverse tasks, yet their safety alignment remains fragile during adaptation. Even when fine-tuning on benign data or with low-rank adaptation, pre-trained safety behaviors are easily degraded, leading to harmful responses in the fine-tuned models. To address this challenge, we propose GuardSpace, a guardrail framework for preserving safety alignment throughout fine-tuning, composed of two key components: a safety-sensitive subspace and a harmful-resistant null space. First, we explicitly decompose pre-trained weights into safety-relevant and safety-irrelevant components using covariance-preconditioned singular value decomposition, and initialize low-rank adapters from the safety-irrelevant ones, while freezing safety-relevant components to preserve their associated safety mechanism. Second, we construct a null space projector that restricts adapter updates from altering safe outputs on harmful prompts, thereby maintaining the original refusal behavior. Experiments with various pre-trained models on multiple downstream tasks demonstrate that GuardSpace achieves superior performance over existing methods. Notably, for Llama-2-7B-Chat fine-tuned on GSM8K, GuardSpace outperforms the state-of-the-art method AsFT, reducing the average harmful score from 14.4% to 3.6%, while improving the accuracy from from 26.0% to 28.0%.","authors":["Bingjie Zhang","Yibo Yang","Zhe Ren","Dandan Guo","Jindong Gu","Philip Torr","Bernard Ghanem"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2025-10-16","first_seen":"2025-10-16","revised_at":null,"abs_url":"https://arxiv.org/abs/2510.14301","pdf_url":"https://arxiv.org/pdf/2510.14301","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["安全对齐","微调","模型安全"],"reason":"论文研究LLM安全对齐的微调方法，不涉及人类行为仿真或对照实验。","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:59:16","error":null,"has_summary":false,"summary":null},{"id":"2510.13091","version":2,"title":"Unmasking Hiring Bias: Platform Data Analysis and Controlled Experiments on Bias in Online Freelance Marketplaces via RAG-LLM Generated Contents","zh_title":"揭示招聘偏见：基于RAG-LLM生成内容的在线自由职业市场偏见平台数据分析与受控实验","abstract":"Online freelance marketplaces, a rapidly growing part of the global labor market, are creating a fair environment where professional skills are the main factor for hiring. While these platforms can reduce bias from traditional hiring, the personal information in user profiles raises concerns about ongoing discrimination. Past studies on this topic have mostly used existing data, which makes it hard to control for other factors and clearly see the effect of things like gender or race. To solve these problems, this paper presents a new method that uses Retrieval-Augmented Generation (RAG) with a Large Language Model (LLM) to create realistic, artificial freelancer profiles for controlled experiments. This approach effectively separates individual factors, enabling a clearer statistical analysis of how different variables influence the freelancer project process. In addition to analyzing extracted data with traditional statistical methods for post-project stage analysis, our research utilizes a dataset with highly controlled variables, generated by an RAG-LLM, to conduct a simulated hiring experiment for pre-project stage analysis. The results of our experiments show that, regarding gender, while no significant preference emerged in initial hiring decisions, female freelancers are substantially more likely to receive imperfect ratings post-project stage. Regarding regional bias, a strong and consistent preference favoring US-based freelancers shows that people are more likely to be selected in the simulated experiments, perceived as more leader-like, and receive higher ratings on the live platform.","authors":["Wugeng Zheng","Guohou Shan"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2025-10-15","first_seen":"2025-10-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2510.13091","pdf_url":"https://arxiv.org/pdf/2510.13091","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM生成合成数据","招聘偏见","受控实验"],"reason":"用LLM生成合成档案做受控实验，替代真实简历，属标注/数据生成替代，非直接仿真…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:34","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":193,"question":"在线自由职业市场中，性别和地域偏见如何在招聘前阶段（模拟实验）和项目后评价阶段（真实平台数据）中表现？","design":"本研究并非直接仿真人类被试，而是使用RAG-LLM生成高度受控的合成自由职业者档案，通过Amazon Mechanical Turk招募真人参与者进行模拟招聘实验，测量初始雇佣决策中的偏见；同时结合真实平台数据，分析项目后评分中的偏见。","baseline":"真实人类数据来自Freelancer.com上抓取的自由职业者档案、项目评价和评分，用于项目后阶段分析；模拟实验部分无直接人类对照基准，但通过控制变量分离偏见效应。","findings":"模拟实验显示，初始雇佣决策中未发现显著性别偏好，但女性自由职业者在项目后阶段更可能获得不完美评分；地域偏见方面，美国自由职业者在模拟实验中更易被选中，在真实平台上也被认为更具领导力并获得更高评分。","reliability":"论文未讨论","relevance":"该研究利用LLM生成合成档案进行受控实验，替代真实简历以分离变量，属于数据生成替代而非直接仿真人类被试，但涉及偏见测量与真实平台数据对照，对关注LLM在实验设计中应用的研究者有参考价值，建议略读方法部分。","inspiration":"利用RAG-LLM生成高度受控的合成档案以分离偏见变量，并通过真人实验测量决策差异，同时结合真实平台数据对照｜可迁移至信贷审批中的性别或地域歧视研究，例如P2P借贷平台上的贷款决策偏见｜招募真人作为信贷员，处理为合成借款人档案（LLM生成，控制性别/地域），结果变量为贷款批准率与利率，对照真实P2P平台贷款数据中的实际审批与违约率"}},{"id":"2510.13011","version":1,"title":"Deliberate Lab: A Platform for Real-Time Human-AI Social Experiments","zh_title":"Deliberate Lab：一个实时人类-AI社会实验平台","abstract":"Social and behavioral scientists increasingly aim to study how humans interact, collaborate, and make decisions alongside artificial intelligence. However, the experimental infrastructure for such work remains underdeveloped: (1) few platforms support real-time, multi-party studies at scale; (2) most deployments require bespoke engineering, limiting replicability and accessibility, and (3) existing tools do not treat AI agents as first-class participants. We present Deliberate Lab, an open-source platform for large-scale, real-time behavioral experiments that supports both human participants and large language model (LLM)-based agents. We report on a 12-month public deployment of the platform (N=88 experimenters, N=9195 experiment participants), analyzing usage patterns and workflows. Case studies and usage scenarios are aggregated from platform users, complemented by in-depth interviews with select experimenters. By lowering technical barriers and standardizing support for hybrid human-AI experimentation, Deliberate Lab expands the methodological repertoire for studying collective decision-making and human-centered AI.","authors":["Crystal Qian","Vivian Tsai","Michael Behr","Nada Hussein","Léo Laugier","Nithum Thain","Lucas Dixon"],"categories":["cs.HC","cs.AI"],"primary_category":"cs.HC","announce_type":"new","date":"2025-10-14","first_seen":"2025-10-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2510.13011","pdf_url":"https://arxiv.org/pdf/2510.13011","source_feed":"backfill","score":7,"bucket":"pending","rubric_hits":["A3","B1"],"tags":["人机混合实验","集体决策","LLM代理"],"reason":"平台支持人类与LLM agent混合实验，用于研究集体决策，有真实人类数据对照…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:40","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":89,"question":"如何设计一个支持实时、大规模、人类与LLM智能体混合参与的在线社会实验平台？","design":"本文并非仿真研究，而是介绍一个开源实验平台Deliberate Lab。该平台允许研究者创建包含人类参与者和LLM智能体的实时、同步、多参与者实验，LLM可作为参与者、主持人或模拟人群。平台提供无代码界面、模块化阶段设计、实时状态管理和智能体响应逻辑。","baseline":"无对照","findings":"在12个月的公开部署中，88名实验者创建了597个实验，涉及9195名参与者，覆盖心理学、经济学、人机交互等领域。平台降低了实时混合人机实验的技术门槛，支持了集体决策和人机协作等多样化研究。","reliability":"论文未讨论","relevance":"该平台直接支持将LLM作为人类被试的替代品或补充，用于实时互动实验，但本文未提供仿真有效性的实证评估或与人类基准的对比，适合作为工具参考而非仿真可靠性研究。","inspiration":"该平台支持将LLM智能体作为人类被试的替代品或补充，用于实时、同步、多参与者互动实验，其模块化阶段设计和无代码界面降低了实验部署门槛｜可迁移至经济金融领域的市场博弈、集体决策或政策沟通实验，如资产定价中的信息扩散与泡沫形成、央行沟通中的预期引导｜可设计一个资产市场实验，以LLM智能体为交易者，施加不同信息透明度处理，测量价格泡沫与交易量，并对照已有真实人类实验数据（如Smith等人1988年的泡沫实验）评估仿真有效性"}},{"id":"2510.12189","version":1,"title":"Agent-Based Simulation of a Financial Market with Large Language Models","zh_title":"基于大语言模型的金融市场智能体仿真","abstract":"In real-world stock markets, certain chart patterns -- such as price declines near historical highs -- cannot be fully explained by fundamentals alone. These phenomena suggest the presence of path dependence in price formation, where investor decisions are influenced not only by current market conditions but also by the trajectory of prices leading up to the present. Path dependence has drawn attention in behavioral finance as a key mechanism behind such anomalies. One plausible driver of path dependence is human loss aversion, anchored to individual reference points like purchase prices or past peaks, which vary with personal context. However, capturing such subtle behavioral tendencies in traditional agent-based market simulations has remained a challenge. We propose the Fundamental-Chartist-LLM-Agent (FCLAgent), which uses large language models (LLMs) to emulate human-like trading decisions. In this framework, (1) buy/sell decisions are made by LLMs based on individual situations, while (2) order price and volume follow standard rule-based methods. Simulations show that FCLAgents reproduce path-dependent patterns that conventional agents fail to capture. Furthermore, an analysis of FCLAgents' behavior reveals that the reference points guiding loss aversion vary with market trajectories, highlighting the potential of LLM-based agents to model nuanced investor behavior.","authors":["Ryuji Hashimoto","Takehiro Takayanagi","Masahiro Suzuki","Kiyoshi Izumi"],"categories":["cs.CE"],"primary_category":"cs.CE","announce_type":"new","date":"2025-10-14","first_seen":"2025-10-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2510.12189","pdf_url":"https://arxiv.org/pdf/2510.12189","source_feed":"backfill","score":7,"bucket":"pending","rubric_hits":["A3","B2"],"tags":["LLM仿真","行为金融","智能体市场"],"reason":"用LLM代理模拟金融市场投资者行为，复现路径依赖，涉及行为金融场景，但未明确与…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:34","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":106,"question":"如何利用大语言模型构建能再现金融市场中路径依赖现象的智能体仿真模型？","design":"提出FCLAgent模型，将LLM用于生成买卖意图以模拟情境依赖的损失厌恶等行为偏差，订单价格和数量仍由传统规则决定；在单市场多智能体仿真中，随机选取智能体逐时间步下单，观察价格序列的路径依赖模式。","baseline":"无对照","findings":"FCLAgent能复现传统智能体无法捕捉的路径依赖模式，如接近历史高点与未来收益的负相关；LLM的决策参考点随市场轨迹变化，表现出类似人类的损失厌恶和风险偏好调整。","reliability":"论文未讨论","relevance":"该研究用LLM模拟投资者行为偏差，复现金融市场异象，属于人类仿真实验范畴，但缺乏真实人类数据对照，适合关注LLM行为逼真度和情境依赖建模的研究者阅读。","inspiration":"借鉴LLM生成买卖意图以模拟情境依赖行为偏差的方法，可将LLM作为被试，通过提示词操控其参考点或情绪状态，观察决策变化｜可迁移到资产定价实验中的处置效应研究，检验投资者是否因参考点变化而表现出持有亏损资产、过早卖出盈利资产的倾向｜设计：以LLM为被试，随机分配历史价格路径（如近期高点或低点），要求其决定持有或卖出，结果变量为卖出概率，对照真实市场交易数据中的处置效应模式"}},{"id":"2510.08338","version":3,"title":"LLMs Reproduce Human Purchase Intent via Semantic Similarity Elicitation of Likert Ratings","zh_title":"大语言模型通过语义相似度引出李克特评分复现人类购买意向","abstract":"Consumer research costs companies billions annually yet suffers from panel biases and limited scale. Large language models (LLMs) offer an alternative by simulating synthetic consumers, but produce unrealistic response distributions when asked directly for numerical ratings. We present semantic similarity rating (SSR), a method that elicits textual responses from LLMs and maps these to Likert distributions using embedding similarity to reference statements. Testing on an extensive dataset comprising 57 personal care product surveys conducted by a leading corporation in that market (9,300 human responses), SSR achieves 90% of human test-retest reliability while maintaining realistic response distributions (KS similarity > 0.85). Additionally, these synthetic respondents provide rich qualitative feedback explaining their ratings. This framework enables scalable consumer research simulations while preserving traditional survey metrics and interpretability.","authors":["Benjamin F. Maier","Ulf Aslak","Luca Fiaschi","Nina Rismal","Kemble Fletcher","Christian C. Luhmann","Robbie Dow","Kli Pappas","Thomas V. Wiecki"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2025-10-09","first_seen":"2025-10-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2510.08338","pdf_url":"https://arxiv.org/pdf/2510.08338","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2"],"tags":["LLM仿真","消费者调查","人类数据对照"],"reason":"用LLM模拟消费者购买意向，与真实人类调查数据对照，评估分布可靠性与偏差，直接…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:32","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":6,"question":"如何通过语义相似度评分（SSR）方法让大语言模型生成更真实的人类购买意向分布？","design":"用大语言模型扮演给定人口统计属性的合成消费者，对57项个人护理产品概念生成自由文本购买意向，再通过嵌入向量与锚定语句的余弦相似度映射到5点李克特量表，测量购买意向分布和平均购买意向。","baseline":"来自一家领先企业的57项个人护理产品调查，共9,300名真实美国消费者的购买意向李克特评分。","findings":"SSR方法使合成消费者的购买意向分布与人类高度相似（KS相似度>0.85），且合成数据与人类数据在概念吸引力排名上的相关性达到人类重测信度的90%。同时，合成消费者能提供解释评分的定性反馈。","reliability":"论文未讨论","relevance":"高度相关：该研究用LLM模拟消费者购买意向，并与大规模真实人类调查数据严格对照，评估分布可靠性和偏差，直接命中研究者关心的仿真基准和经济学实验场景，值得精读原文。","inspiration":"借鉴其语义相似度评分方法，将LLM生成的自由文本通过嵌入向量与锚定语句的余弦相似度映射到结构化量表，可提升合成行为分布的真实性｜可迁移到消费者跨期选择实验，用LLM模拟不同人口统计特征下的时间偏好与折扣因子｜用LLM扮演合成消费者，施加未来不同时间点的金额选择任务，生成自由文本决策理由，通过SSR映射为选择概率，与真实跨期选择调查数据对照"}},{"id":"2510.08236","version":2,"title":"The Hidden Bias: A Study on Explicit and Implicit Political Stereotypes in Large Language Models","zh_title":"隐藏的偏见：大语言模型中显性与隐性政治刻板印象研究","abstract":"Large Language Models (LLMs) are increasingly integral to information dissemination and decision-making processes. Given their growing societal influence, understanding potential biases, particularly within the political domain, is crucial to prevent undue influence on public opinion and democratic processes. This work investigates political bias and stereotype propagation across eight prominent LLMs using the two-dimensional Political Compass Test (PCT). Initially, the PCT is employed to assess the inherent political leanings of these models. Subsequently, persona prompting with the PCT is used to explore explicit stereotypes across various social dimensions. In a final step, implicit stereotypes are uncovered by evaluating models with multilingual versions of the PCT. Key findings reveal a consistent left-leaning political alignment across all investigated models. Furthermore, while the nature and extent of stereotypes vary considerably between models, implicit stereotypes elicited through language variation are more pronounced than those identified via explicit persona prompting. Interestingly, for most models, implicit and explicit stereotypes show a notable alignment, suggesting a degree of transparency or \"awareness\" regarding their inherent biases. This study underscores the complex interplay of political bias and stereotypes in LLMs.","authors":["Konrad Löhr","Shuzhou Yuan","Michael Färber"],"categories":["cs.LG","cs.AI"],"primary_category":"cs.LG","announce_type":"new","date":"2025-10-09","first_seen":"2025-10-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2510.08236","pdf_url":"https://arxiv.org/pdf/2510.08236","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D2"],"tags":["政治偏见","刻板印象","LLM评估"],"reason":"测量LLM本身的政治偏见与刻板印象，属于人格/态度测量，无人类被试仿真对照。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:39","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":167,"question":"八种主流大语言模型在政治罗盘测试中表现出怎样的内在政治倾向、外显刻板印象和内隐刻板印象？","design":"本研究并非人类仿真实验，而是直接测量LLM本身的政治偏见与刻板印象。使用政治罗盘测试（PCT）评估八个LLM的基线政治倾向；通过角色提示（persona prompting）让模型扮演不同社会维度角色，测量外显刻板印象；再通过多语言版PCT评估模型在不同语言下的表现，揭示内隐刻板印象。","baseline":"无对照","findings":"所有模型均一致表现出左倾政治倾向（经济左翼、社会自由意志主义）。内隐刻板印象（通过语言变化引发）比外显刻板印象（通过角色提示引发）更显著，且多数模型的内隐与外显刻板印象存在明显一致性，表明模型对其内在偏见有一定“意识”。","reliability":"论文未讨论","relevance":"该研究直接测量LLM的政治态度与刻板印象，未将LLM作为人类被试的替代品，也无真实人类数据对照，不属于人类仿真实验。若关注LLM偏见本身可读，但与研究者关心的仿真效度与失效条件无关。","inspiration":"该方法通过角色提示（persona prompting）施加外显刻板印象处理，并利用多语言测试揭示内隐刻板印象，为测量隐性偏见提供了可借鉴的实验设计框架。｜可迁移至信贷审批中的种族或性别歧视研究，探究LLM在贷款决策中是否存在内隐偏见。｜以LLM为被试，通过角色提示模拟不同种族或性别的贷款申请人，结果变量为贷款批准率与利率设定，并以真实银行信贷数据作为对照基准。"}},{"id":"2510.08191","version":1,"title":"Training-Free Group Relative Policy Optimization","zh_title":"免训练的分组相对策略优化","abstract":"Recent advances in Large Language Model (LLM) agents have demonstrated their promising general capabilities. However, their performance in specialized real-world domains often degrades due to challenges in effectively integrating external tools and specific prompting strategies. While methods like agentic reinforcement learning have been proposed to address this, they typically rely on costly parameter updates, for example, through a process that uses Supervised Fine-Tuning (SFT) followed by a Reinforcement Learning (RL) phase with Group Relative Policy Optimization (GRPO) to alter the output distribution. However, we argue that LLMs can achieve a similar effect on the output distribution by learning experiential knowledge as a token prior, which is a far more lightweight approach that not only addresses practical data scarcity but also avoids the common issue of overfitting. To this end, we propose Training-Free Group Relative Policy Optimization (Training-Free GRPO), a cost-effective solution that enhances LLM agent performance without any parameter updates. Our method leverages the group relative semantic advantage instead of numerical ones within each group of rollouts, iteratively distilling high-quality experiential knowledge during multi-epoch learning on a minimal ground-truth data. Such knowledge serves as the learned token prior, which is seamlessly integrated during LLM API calls to guide model behavior. Experiments on mathematical reasoning and web searching tasks demonstrate that Training-Free GRPO, when applied to DeepSeek-V3.1-Terminus, significantly improves out-of-domain performance. With just a few dozen training samples, Training-Free GRPO outperforms fine-tuned small LLMs with marginal training data and cost.","authors":["Yuzheng Cai","Siqi Cai","Yuchen Shi","Zihan Xu","Lichao Chen","Yulei Qin","Xiaoyu Tan","Gang Li","Zongyi Li","Haojia Lin","Yong Mao","Ke Li","Xing Sun"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2025-10-09","first_seen":"2025-10-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2510.08191","pdf_url":"https://arxiv.org/pdf/2510.08191","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体强化学习","工具调用","免训练优化"],"reason":"纯多智能体协作解题，无人类行为对照，属C1排除项。","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:59:16","error":null,"has_summary":false,"summary":null},{"id":"2510.07733","version":3,"title":"SurveyG: A Multi-Agent LLM Framework with Hierarchical Citation Graph for Automated Survey Generation","zh_title":"SurveyG：基于层次引用图的多智能体LLM框架用于自动生成综述","abstract":"Large language models (LLMs) are increasingly adopted for automating survey paper generation \\cite{wang2406autosurvey, liang2025surveyx, yan2025surveyforge,su2025benchmarking,wen2025interactivesurvey}. Existing approaches typically extract content from a large collection of related papers and prompt LLMs to summarize them directly. However, such methods often overlook the structural relationships among papers, resulting in generated surveys that lack a coherent taxonomy and a deeper contextual understanding of research progress. To address these shortcomings, we propose \\textbf{SurveyG}, an LLM-based agent framework that integrates \\textit{hierarchical citation graph}, where nodes denote research papers and edges capture both citation dependencies and semantic relatedness between their contents, thereby embedding structural and contextual knowledge into the survey generation process. The graph is organized into three layers: \\textbf{Foundation}, \\textbf{Development}, and \\textbf{Frontier}, to capture the evolution of research from seminal works to incremental advances and emerging directions. By combining horizontal search within layers and vertical depth traversal across layers, the agent produces multi-level summaries, which are consolidated into a structured survey outline. A multi-agent validation stage then ensures consistency, coverage, and factual accuracy in generating the final survey. Experiments, including evaluations by human experts and LLM-as-a-judge, demonstrate that SurveyG outperforms state-of-the-art frameworks, producing surveys that are more comprehensive and better structured to the underlying knowledge taxonomy of a field.","authors":["Minh-Anh Nguye","Minh-Duc Nguyen","Ha Lan N. T.","Kieu Hai Dang","Nguyen Tien Dong","Dung D. Le"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2025-10-09","first_seen":"2025-10-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2510.07733","pdf_url":"https://arxiv.org/pdf/2510.07733","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["自动综述生成","多智能体系统","引用图"],"reason":"多智能体协作生成综述，无人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:59:34","error":null,"has_summary":false,"summary":null},{"id":"2510.06903","version":1,"title":"When Machines Meet Each Other: Network Effects and the Strategic Role of History in Multi-Agent AI","zh_title":"当机器相遇：多智能体AI中的网络效应与历史的战略角色","abstract":"As artificial intelligence (AI) enters the agentic era, large language models (LLMs) are increasingly deployed as autonomous agents that interact with one another rather than operate in isolation. This shift raises a fundamental question: how do machine agents behave in interdependent environments where outcomes depend not only on their own choices but also on the coordinated expectations of peers? To address this question, we study LLM agents in a canonical network-effect game, where economic theory predicts convergence to a fulfilled expectation equilibrium (FEE). We design an experimental framework in which 50 heterogeneous GPT-5-based agents repeatedly interact under systematically varied network-effect strengths, price trajectories, and decision-history lengths. The results reveal that LLM agents systematically diverge from FEE: they underestimate participation at low prices, overestimate at high prices, and sustain persistent dispersion. Crucially, the way history is structured emerges as a design lever. Simple monotonic histories-where past outcomes follow a steady upward or downward trend-help stabilize coordination, whereas nonmonotonic histories amplify divergence and path dependence. Regression analyses at the individual level further show that price is the dominant driver of deviation, history moderates this effect, and network effects amplify contextual distortions. Together, these findings advance machine behavior research by providing the first systematic evidence on multi-agent AI systems under network effects and offer guidance for configuring such systems in practice.","authors":["Yu Liu","Wenwen Li","Yifan Dou","Guangnan Ye"],"categories":["econ.GN"],"primary_category":"econ.GN","announce_type":"new","date":"2025-10-08","first_seen":"2025-10-08","revised_at":null,"abs_url":"https://arxiv.org/abs/2510.06903","pdf_url":"https://arxiv.org/pdf/2510.06903","source_feed":"backfill","score":7,"bucket":"pending","rubric_hits":["A3","B2"],"tags":["LLM仿真","网络效应博弈","多智能体"],"reason":"用LLM agent模拟网络效应博弈，涉及经济学实验场景，但缺真实人类数据对照。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:32","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":108,"question":"LLM智能体在网络效应博弈中能否形成与经济学理论一致的协调预期与均衡行为？","design":"用50个异构的GPT-5智能体模拟经济主体，在系统变化的网络效应强度、价格轨迹和决策历史长度下重复进行网络效应博弈，测量其参与预期与理论均衡的偏离。","baseline":"无对照","findings":"LLM智能体系统性地偏离理论均衡：低价时低估参与，高价时高估参与，且预期持续分散；单调历史有助于稳定协调，非单调历史则放大偏离和路径依赖。","reliability":"论文未讨论","relevance":"该研究用LLM模拟多主体网络效应博弈，属于经济学实验场景的仿真，但缺乏真实人类数据对照，可作为批判性研究的对象，值得阅读以评估其仿真有效性边界。","inspiration":"该研究通过系统操纵网络效应强度、价格轨迹和决策历史长度来测量LLM智能体的预期与均衡偏离，这种多维度处理设计值得借鉴｜可迁移到资产定价实验中的预期形成研究，例如检验LLM智能体在价格泡沫或崩盘历史下的预期偏差｜用LLM智能体作为被试，处理为不同历史价格序列（单调上涨/下跌、非单调波动），结果变量为预期价格与理性预期均衡的偏离，对照真实人类实验数据"}},{"id":"2510.06151","version":1,"title":"LLMs as Policy-Agnostic Teammates: A Case Study in Human Proxy Design for Heterogeneous Agent Teams","zh_title":"作为策略无关队友的大语言模型：异构智能体团队中人类代理设计的案例研究","abstract":"A critical challenge in modelling Heterogeneous-Agent Teams is training agents to collaborate with teammates whose policies are inaccessible or non-stationary, such as humans. Traditional approaches rely on expensive human-in-the-loop data, which limits scalability. We propose using Large Language Models (LLMs) as policy-agnostic human proxies to generate synthetic data that mimics human decision-making. To evaluate this, we conduct three experiments in a grid-world capture game inspired by Stag Hunt, a game theory paradigm that balances risk and reward. In Experiment 1, we compare decisions from 30 human participants and 2 expert judges with outputs from LLaMA 3.1 and Mixtral 8x22B models. LLMs, prompted with game-state observations and reward structures, align more closely with experts than participants, demonstrating consistency in applying underlying decision criteria. Experiment 2 modifies prompts to induce risk-sensitive strategies (e.g. \"be risk averse\"). LLM outputs mirror human participants' variability, shifting between risk-averse and risk-seeking behaviours. Finally, Experiment 3 tests LLMs in a dynamic grid-world where the LLM agents generate movement actions. LLMs produce trajectories resembling human participants' paths. While LLMs cannot yet fully replicate human adaptability, their prompt-guided diversity offers a scalable foundation for simulating policy-agnostic teammates.","authors":["Aju Ani Justus","Chris Baber"],"categories":["cs.LG","cs.AI","cs.HC"],"primary_category":"cs.LG","announce_type":"new","date":"2025-10-07","first_seen":"2025-10-07","revised_at":null,"abs_url":"https://arxiv.org/abs/2510.06151","pdf_url":"https://arxiv.org/pdf/2510.06151","source_feed":"backfill","score":8,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM人类仿真","行为博弈","人机协作"],"reason":"用LLM代理人类决策，与真实人类数据对照，涉及博弈论实验，方法可迁移。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:30","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":48,"question":"LLM能否作为策略无关的人类代理，在猎鹿博弈网格游戏中复现专家决策、模拟人类风险偏好变异并生成类人运动轨迹？","design":"使用LLaMA 3.1和Mixtral 8x22B模型，通过提示词输入网格世界的相对距离和奖励结构，模拟人类在猎鹿博弈中的决策。实验1比较LLM与30名人类参与者和2名专家裁判的选择；实验2通过修改提示词诱导风险规避或风险寻求策略；实验3让LLM在动态网格中生成移动动作序列。","baseline":"30名对博弈论了解有限的人类参与者在15种网格配置下的决策，以及2名博弈论专家裁判的选择。","findings":"LLM的决策与专家裁判高度一致，表现出对底层决策准则的一致性应用；通过提示词引导，LLM能展现出与人类参与者相似的风险偏好变异，在风险规避和风险寻求行为间切换。","reliability":"LLM尚不能完全复制人类的适应性，其行为多样性依赖于提示词引导，而非自主涌现。","relevance":"该研究直接用LLM代理人类决策，并与真实人类数据对照，涉及博弈论实验，方法可迁移至经济学和政策评估场景，对关注仿真可靠性与偏差的研究者具有参考价值，值得阅读原文。","inspiration":"该方法通过提示词工程直接操控LLM的风险偏好（风险规避/风险寻求），并设置专家裁判和人类参与者双基准对照，可借鉴其处理异质性行为变异的设计｜可迁移至资产定价实验中的风险态度测量，或政策公告对投资者预期形成的仿真研究｜以LLM作为被试，通过提示词注入不同风险偏好指令，测量其在模拟股票投资任务中的资产配置比例，并与真实投资者调查数据或实验数据对照"}},{"id":"2510.03231","version":1,"title":"Reward Models are Metrics in a Trench Coat","zh_title":"奖励模型是披着风衣的度量指标","abstract":"The emergence of reinforcement learning in post-training of large language models has sparked significant interest in reward models. Reward models assess the quality of sampled model outputs to generate training signals. This task is also performed by evaluation metrics that monitor the performance of an AI model. We find that the two research areas are mostly separate, leading to redundant terminology and repeated pitfalls. Common challenges include susceptibility to spurious correlations, impact on downstream reward hacking, methods to improve data quality, and approaches to meta-evaluation. Our position paper argues that a closer collaboration between the fields can help overcome these issues. To that end, we show how metrics outperform reward models on specific tasks and provide an extensive survey of the two areas. Grounded in this survey, we point to multiple research topics in which closer alignment can improve reward models and metrics in areas such as preference elicitation methods, avoidance of spurious correlations and reward hacking, and calibration-aware meta-evaluation.","authors":["Sebastian Gehrmann"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2025-10-03","first_seen":"2025-10-03","revised_at":null,"abs_url":"https://arxiv.org/abs/2510.03231","pdf_url":"https://arxiv.org/pdf/2510.03231","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["奖励模型","评估指标","元评估"],"reason":"论文讨论奖励模型与评估指标的融合，属于NLP评测方法，不涉及人类行为仿真或对照。","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:59:43","error":null,"has_summary":false,"summary":null},{"id":"2510.01115","version":2,"title":"Exploring Network-Knowledge Graph Duality: A Case Study in Agentic Supply Chain Risk Analysis","zh_title":"探索网络-知识图谱对偶性：以智能体供应链风险分析为例","abstract":"Large Language Models (LLMs) struggle with the complex, multi-modal, and network-native data underlying financial risk. Standard Retrieval-Augmented Generation (RAG) oversimplifies relationships, while specialist models are costly and static. We address this gap with an LLM-centric agent framework for supply chain risk analysis. Our core contribution is to exploit the inherent duality between networks and knowledge graphs (KG). We treat the supply chain network as a KG, allowing us to use structural network science principles for retrieval. A graph traverser, guided by network centrality scores, efficiently extracts the most economically salient risk paths. An agentic architecture orchestrates this graph retrieval alongside data from numerical factor tables and news streams. Crucially, it employs novel ``context shells'' -- descriptive templates that embed raw figures in natural language -- to make quantitative data fully intelligible to the LLM. This lightweight approach enables the model to generate concise, explainable, and context-rich risk narratives in real-time without costly fine-tuning or a dedicated graph database.","authors":["Evan Heus","Rick Bookstaber","Dhruv Sharma"],"categories":["cs.AI","cs.MA","econ.TH","physics.soc-ph"],"primary_category":"cs.AI","announce_type":"new","date":"2025-10-01","first_seen":"2025-10-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2510.01115","pdf_url":"https://arxiv.org/pdf/2510.01115","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","供应链风险","知识图谱"],"reason":"纯多智能体系统，agent协作分析供应链风险，无人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:59:01","error":null,"has_summary":false,"summary":null},{"id":"2509.25709","version":1,"title":"Leveraging LLMs to Improve Experimental Design: A Generative Stratification Approach","zh_title":"利用大语言模型改进实验设计：一种生成式分层方法","abstract":"Pre-experiment stratification, or blocking, is a well-established technique for designing more efficient experiments and increasing the precision of the experimental estimates. However, when researchers have access to many covariates at the experiment design stage, they often face challenges in effectively selecting or weighting covariates when creating their strata. This paper proposes a Generative Stratification procedure that leverages Large Language Models (LLMs) to synthesize high-dimensional covariate data to improve experimental design. We demonstrate the value of this approach by applying it to a set of experiments and find that our method would have reduced the variance of the treatment effect estimate by 10%-50% compared to simple randomization in our empirical applications. When combined with other standard stratification methods, it can be used to further improve the efficiency. Our results demonstrate that LLM-based simulation is a practical and easy-to-implement way to improve experimental design in covariate-rich settings.","authors":["George Gui","Seungwoo Kim"],"categories":["econ.EM"],"primary_category":"econ.EM","announce_type":"new","date":"2025-09-30","first_seen":"2025-09-30","revised_at":null,"abs_url":"https://arxiv.org/abs/2509.25709","pdf_url":"https://arxiv.org/pdf/2509.25709","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["实验设计","分层抽样","合成数据"],"reason":"用LLM生成合成协变量改进实验分层，替代人工设计而非仿真人类被试，属边界情形。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:28","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":195,"question":"如何利用大语言模型合成高维协变量信息，生成有效的预后得分以改进实验分层设计，从而提高处理效应估计的精度？","design":"本研究不是用LLM仿真人类被试，而是提出一种生成式分层方法：利用LLM基于实验单元的观测协变量和实验情境，生成预测的潜在结果，以此构建预后得分用于分层，从而改进随机化实验的设计效率。","baseline":"无对照","findings":"该方法在多个实证应用中能将处理效应估计的方差相比简单随机化降低10%-50%；与标准分层方法结合可进一步提升效率，且LLM预测即使不完美也不会引入偏差，因为层内随机化保持了无偏性。","reliability":"论文指出，分层效果依赖于LLM生成的得分与真实最优得分的相关性；若相关性弱，效率提升可能有限，但由于随机化，估计仍无偏。","relevance":"该研究与您关注的LLM仿真人类被试不同，它用LLM生成协变量得分以优化实验设计，而非替代人类参与实验。但涉及LLM在实验方法中的应用，且讨论了预测可靠性与偏差，可作为边界参考，建议略读以了解LLM在实验设计中的新用途。","inspiration":"该方法利用LLM生成预后得分以改进分层设计，可借鉴其将LLM预测作为协变量降维工具的思路，用于提升随机化实验的效率｜可迁移到政策评估场景，如利用LLM基于个体特征预测政策干预的潜在结果，优化分层随机化以更精确估计处理效应｜设计：以真实政策受众为被试，处理为某项就业培训，结果变量为就业收入，用LLM基于基线协变量生成预后得分进行分层随机化，以真实历史数据作为对照基准评估效率提升"}},{"id":"2510.02343","version":1,"title":"$\\texttt{BluePrint}$: A Social Media User Dataset for LLM Persona Evaluation and Training","zh_title":"BluePrint：用于LLM角色评估与训练的社交媒体用户数据集","abstract":"Large language models (LLMs) offer promising capabilities for simulating social media dynamics at scale, enabling studies that would be ethically or logistically challenging with human subjects. However, the field lacks standardized data resources for fine-tuning and evaluating LLMs as realistic social media agents. We address this gap by introducing SIMPACT, the SIMulation-oriented Persona and Action Capture Toolkit, a privacy respecting framework for constructing behaviorally-grounded social media datasets suitable for training agent models. We formulate next-action prediction as a task for training and evaluating LLM-based agents and introduce metrics at both the cluster and population levels to assess behavioral fidelity and stylistic realism. As a concrete implementation, we release BluePrint, a large-scale dataset built from public Bluesky data focused on political discourse. BluePrint clusters anonymized users into personas of aggregated behaviours, capturing authentic engagement patterns while safeguarding privacy through pseudonymization and removal of personally identifiable information. The dataset includes a sizable action set of 12 social media interaction types (likes, replies, reposts, etc.), each instance tied to the posting activity preceding it. This supports the development of agents that use context-dependence, not only in the language, but also in the interaction behaviours of social media to model social media users. By standardizing data and evaluation protocols, SIMPACT provides a foundation for advancing rigorous, ethically responsible social media simulations. BluePrint serves as both an evaluation benchmark for political discourse modeling and a template for building domain specific datasets to study challenges such as misinformation and polarization.","authors":["Aurélien Bück-Kaeffer","Je Qin Chooi","Dan Zhao","Maximilian Puelma Touzel","Kellin Pelrine","Jean-François Godbout","Reihaneh Rabbany","Zachary Yang"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2025-09-27","first_seen":"2025-09-27","revised_at":null,"abs_url":"https://arxiv.org/abs/2510.02343","pdf_url":"https://arxiv.org/pdf/2510.02343","source_feed":"backfill","score":8,"bucket":"selected","rubric_hits":["A3","B1","B2"],"tags":["LLM仿真","社交媒体模拟","行为保真度"],"reason":"用LLM模拟社交媒体用户行为，有真实数据对照，涉及政治话语，方法可迁移至人类仿…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:29","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":44,"question":"如何构建隐私保护的社交媒体数据集，并用于训练和评估基于大语言模型的社交媒体用户行为模拟代理？","design":"提出SIMPACT框架，将用户行为建模为动作序列，通过聚类形成行为角色以保护隐私；基于Bluesky平台2025年加拿大联邦选举期间的政治话语数据构建BluePrint数据集，包含12种交互行为；将下一动作预测作为训练和评估任务，使用GPT-4.1-mini、o3-mini、Qwen-2.5-7B等模型进行基准测试，并采用余弦相似度、Jaccard相似度、JS散度、F1分数及人类评估等多指标衡量行为保真度。","baseline":"以BluePrint数据集中真实用户的帖子嵌入、关键词分布和实际交互行为作为对照基准。","findings":"当前LLM能生成看似合理的文本，但在复现真实用户社区的行为模式上存在困难；经过微调的模型在行为预测上有所提升，但仍难以完全捕捉人类社交互动的细微差别。","reliability":"论文指出计算指标只能近似评估，无法完全捕捉人类行为的微妙性，可能产生“恐怖谷”效应；数据集聚焦特定政治事件和平台，泛化性有待验证；隐私保护措施可能损失部分个体行为细节。","relevance":"该研究直接涉及用LLM模拟社交媒体用户行为，并提供真实人类数据对照和评估基准，符合研究者对仿真可靠性及政治话语场景的关注，值得阅读原文以了解其数据集构建方法和模型失效的具体表现。","inspiration":"借鉴其将用户行为建模为动作序列并用聚类形成行为角色以保护隐私的方法，可设计隐私合规的仿真实验｜可迁移到金融社交媒体信息传播与投资者情绪形成的场景，如研究政策推文对散户交易行为的影响｜以LLM代理模拟散户投资者，处理为不同情绪倾向的政策推文，结果变量为模拟的买卖行为序列，用真实交易数据或社交媒体互动数据作为对照基准"}},{"id":"2510.02331","version":1,"title":"Synthetic Dialogue Generation for Interactive Conversational Elicitation & Recommendation (ICER)","zh_title":"面向交互式对话引导与推荐的合成对话生成","abstract":"While language models (LMs) offer great potential for conversational recommender systems (CRSs), the paucity of public CRS data makes fine-tuning LMs for CRSs challenging. In response, LMs as user simulators qua data generators can be used to train LM-based CRSs, but often lack behavioral consistency, generating utterance sequences inconsistent with those of any real user. To address this, we develop a methodology for generating natural dialogues that are consistent with a user's underlying state using behavior simulators together with LM-prompting. We illustrate our approach by generating a large, open-source CRS data set with both preference elicitation and example critiquing. Rater evaluation on some of these dialogues shows them to exhibit considerable consistency, factuality and naturalness.","authors":["Moonkyung Ryu","Chih-Wei Hsu","Yinlam Chow","Mohammad Ghavamzadeh","Craig Boutilier"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2025-09-26","first_seen":"2025-09-26","revised_at":null,"abs_url":"https://arxiv.org/abs/2510.02331","pdf_url":"https://arxiv.org/pdf/2510.02331","source_feed":"backfill","score":4,"bucket":"other","rubric_hits":["C3"],"tags":["对话生成","推荐系统","用户模拟器"],"reason":"生成对话数据训练推荐系统，属角色扮演对话生成，无人类行为仿真或对照实验。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:38","error":null,"has_summary":false,"summary":null},{"id":"2509.13712","version":1,"title":"Inject, Fork, Compare: Defining an Interaction Vocabulary for Multi-Agent Simulation Platforms","zh_title":"注入、分叉、比较：为多智能体仿真平台定义交互词汇","abstract":"LLM-based multi-agent simulations are a rapidly growing field of research, but current simulations often lack clear modes for interaction and analysis, limiting the \"what if\" scenarios researchers are able to investigate. In this demo, we define three core operations for interacting with multi-agent simulations: inject, fork, and compare. Inject allows researchers to introduce external events at any point during simulation execution. Fork creates independent timeline branches from any timestamp, preserving complete state while allowing divergent exploration. Compare facilitates parallel observation of multiple branches, revealing how different interventions lead to distinct emergent behaviors. Together, these operations establish a vocabulary that transforms linear simulation workflows into interactive, explorable spaces. We demonstrate this vocabulary through a commodity market simulation with fourteen AI agents, where researchers can inject contrasting events and observe divergent outcomes across parallel timelines. By defining these fundamental operations, we provide a starting point for systematic causal investigation in LLM-based agent simulations, moving beyond passive observation toward active experimentation.","authors":["HwiJoon Lee","Martina Di Paola","Yoo Jin Hong","Quang-Huy Nguyen","Joseph Seering"],"categories":["cs.MA","cs.HC"],"primary_category":"cs.MA","announce_type":"new","date":"2025-09-17","first_seen":"2025-09-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2509.13712","pdf_url":"https://arxiv.org/pdf/2509.13712","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["多智能体仿真","社会模拟","交互操作"],"reason":"多智能体市场模拟，但无真实人类数据对照，属社会模拟演示。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:38","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":196,"question":"如何定义一套交互操作词汇（inject, fork, compare），使基于LLM的多智能体仿真从被动观察转变为主动因果探索？","design":"构建一个包含14个AI智能体的商品市场仿真，智能体具有不同交易策略和市场组合；通过注入对比事件（如石油管道爆炸 vs. OPEC增产），利用fork创建平行时间线分支，并用compare并排观察不同分支下涌现的智能体行为差异。","baseline":"无对照","findings":"定义了inject、fork、compare三种核心操作，将线性仿真流程转变为可交互的探索空间；在商品市场演示中，注入对立事件后，不同分支展现出分化的智能体行为，支持实时假设检验。","reliability":"论文未讨论","relevance":"该工作聚焦于多智能体仿真的交互范式，未涉及真实人类数据对照或行为复现，不属于以人类被试替代为目标的人类仿真研究，与研究者关注的经济学实验和政策评估场景关联较弱，不建议优先阅读原文。","inspiration":"该工作提出的inject、fork、compare交互范式可用于在仿真中主动注入经济冲击事件并观察多智能体行为分化，为因果推断提供可控实验环境｜可迁移到政策公告的预期形成与市场反应场景，如央行加息信号对资产价格和交易策略的影响｜以LLM驱动的交易员智能体为被试，注入加息或降息公告作为处理，fork出不同政策分支，compare分支间资产价格和交易量差异，并对照真实高频市场数据验证仿真行为模式"}},{"id":"2509.13397","version":4,"title":"The threat of analytic flexibility in using large language models to simulate human data","zh_title":"使用大语言模型模拟人类数据时分析灵活性的威胁","abstract":"Social scientists are now using large language models to create \"silicon samples\": synthetic datasets intended to stand in for human respondents. However, producing these samples requires many analytic choices, including model selection, sampling parameters, prompt format, and the amount of demographic or contextual information provided. Across two studies, I examine whether these choices materially affect correspondence between silicon samples and human data. In Study 1, I generated 252 silicon-sample configurations for a controlled case study using two social-psychological scales, evaluating whether configurations recovered participant rankings, response distributions, and between-scale correlations. Configurations varied substantially across all three criteria, and configurations that performed well on one dimension often performed poorly on another. In Study 2, I extended this analysis to a published silicon-sample use case by re-examining Argyle et al.'s (2023) Study 3 using 66 alternative configurations. Correlations between human and silicon association structures differed substantially across configurations, from r = .23 to r = .84. Taken together, the results from these studies demonstrate that different defensible configuration choices can materially alter conclusions about the fidelity of silicon samples. I call for greater attention to the threat of analytic flexibility in using silicon samples and outline strategies that researchers may adopt to reduce this threat.","authors":["Jamie Cummins"],"categories":["cs.CY","cs.AI"],"primary_category":"cs.CY","announce_type":"new","date":"2025-09-16","first_seen":"2025-09-16","revised_at":null,"abs_url":"https://arxiv.org/abs/2509.13397","pdf_url":"https://arxiv.org/pdf/2509.13397","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["硅样本","分析灵活性","仿真保真度"],"reason":"直接研究用LLM生成硅样本模拟人类数据，评估分析灵活性对仿真保真度的影响，并与…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:28","error":null,"has_summary":true,"summary":{"generated_at":"2025-09-16","rank":7,"question":"分析灵活性（如模型选择、采样参数、提示格式等）是否显著影响硅样本与人类数据之间的一致性。","design":"研究1：使用两个社会心理量表，生成252种硅样本配置，评估配置能否恢复被试排名、响应分布和量表间相关性。研究2：重新分析Argyle等人(2023)的研究3，使用66种替代配置，比较人类与硅样本关联结构的相关性。","baseline":"研究1：来自两个社会心理量表的人类被试数据。研究2：Argyle等人(2023)研究3中的人类数据。","findings":"不同配置在恢复排名、响应分布和相关性上差异显著，且在一个维度上表现好的配置在另一维度上可能表现差。人类与硅样本关联结构的相关性从r=0.23到r=0.84不等，表明分析灵活性可实质改变关于硅样本保真度的结论。","reliability":"论文指出不同可辩护的配置选择会实质改变结论，但未明确列出所有失效条件或局限。","relevance":"直接命中研究者关注的LLM仿真人类数据可靠性问题，有真实人类对照，并批判性地揭示了分析灵活性威胁，值得精读原文以了解具体配置影响和应对策略。","inspiration":"借鉴其通过系统性地变化模型选择、采样参数和提示格式来构建多种硅样本配置，并评估配置间结果差异的方法，以揭示分析灵活性的威胁。｜可迁移到政策公告的预期形成实验，例如研究央行沟通措辞对通胀预期的影响。｜以LLM作为被试，随机分配不同措辞的政策公告作为处理，测量其预测的通胀数值，并与专业预测者调查或消费者预期调查的真实数据对照，同时变化提示中的角色设定、温度参数等配置，检验结论的稳健性。"}},{"id":"2509.12350","version":1,"title":"Knowledge Graph Tokenization for Behavior-Aware Generative Next POI Recommendation","zh_title":"面向行为感知的生成式下一兴趣点推荐的知识图谱分词","abstract":"Generative paradigm, especially powered by Large Language Models (LLMs), has emerged as a new solution to the next point-of-interest (POI) recommendation. Pioneering studies usually adopt a two-stage pipeline, starting with a tokenizer converting POIs into discrete identifiers that can be processed by LLMs, followed by POI behavior prediction tasks to instruction-tune LLM for next POI recommendation. Despite of remarkable progress, they still face two limitations: (1) existing tokenizers struggle to encode heterogeneous signals in the recommendation data, suffering from information loss issue, and (2) previous instruction-tuning tasks only focus on users' POI visit behavior while ignore other behavior types, resulting in insufficient understanding of mobility. To address these limitations, we propose KGTB (Knowledge Graph Tokenization for Behavior-aware generative next POI recommendation). Specifically, KGTB organizes the recommendation data in a knowledge graph (KG) format, of which the structure can seamlessly preserve the heterogeneous information. Then, a KG-based tokenizer is developed to quantize each node into an individual structural ID. This process is supervised by the KG's structure, thus reducing the loss of heterogeneous information. Using generated IDs, KGTB proposes multi-behavior learning that introduces multiple behavior-specific prediction tasks for LLM fine-tuning, e.g., POI, category, and region visit behaviors. Learning on these behavior tasks provides LLMs with comprehensive insights on the target POI visit behavior. Experiments on four real-world city datasets demonstrate the superior performance of KGTB.","authors":["Ke Sun","Mayi Xu"],"categories":["cs.IR"],"primary_category":"cs.IR","announce_type":"new","date":"2025-09-15","first_seen":"2025-09-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2509.12350","pdf_url":"https://arxiv.org/pdf/2509.12350","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["POI推荐","知识图谱","LLM微调"],"reason":"纯POI推荐模型优化，无人类行为仿真或对照，属NLP能力评测范畴。","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:59:17","error":null,"has_summary":false,"summary":null},{"id":"2509.11311","version":2,"title":"Prompts to Proxies: Emulating Human Preferences via a Compact LLM Ensemble","zh_title":"从提示到代理：通过紧凑LLM集成模拟人类偏好","abstract":"Large language models are increasingly used as proxies for human subjects in social science research, yet external validity requires that synthetic agents faithfully reflect the preferences of target human populations. We introduce *preference reconstruction theory*, a framework that formalizes preference alignment as a representation learning problem: constructing a functional basis of proxy agents and recovering population preferences through weighted aggregation. We implement this via *Prompts to Proxies* ($\\texttt{P2P}$), a modular two-stage system. Stage 1 uses structured prompting with entropy-based adaptive sampling to construct a diverse agent pool spanning the latent preference space. Stage 2 employs L1-regularized regression to select a compact ensemble whose aggregate response distributions align with observed data from the target population. $\\texttt{P2P}$ requires no finetuning and no access to sensitive demographic data, incurring only API inference costs. We validate the approach on 14 waves of the American Trends Panel, achieving an average test MSE of 0.014 across diverse topics at approximately 0.8 USD per survey. We additionally test it on the World Values Survey, demonstrating its potential to generalize across locales. When stress-tested against an SFT-aligned baseline, $\\texttt{P2P}$ achieves competitive performance using less than 3% of the training data.","authors":["Bingchen Wang","Zi-Yu Khoo","Jingtan Wang"],"categories":["cs.AI","cs.CY"],"primary_category":"cs.AI","announce_type":"new","date":"2025-09-14","first_seen":"2025-09-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2509.11311","pdf_url":"https://arxiv.org/pdf/2509.11311","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","A3","B1","B2"],"tags":["LLM人类仿真","偏好重建","社会调查"],"reason":"用LLM代理复现人群偏好，有真实调查数据对照，涉及社会政策评估，直接相关。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:27","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":45,"question":"如何通过紧凑的LLM集成，在无需微调和人口数据的情况下，忠实复现目标人群的偏好分布？","design":"使用GPT系列模型作为代理被试，通过两阶段系统P2P：第一阶段用结构化提示和基于熵的自适应采样构建多样化的代理池，第二阶段用L1正则化回归选择紧凑的代理集成，使其聚合响应分布与目标人群的调查数据对齐。结果变量为调查问题的回答分布。","baseline":"美国趋势面板（ATP）14波调查和世界价值观调查（WVS）的真实人类回答数据。","findings":"P2P在ATP上平均测试MSE为0.014，相比提示基线提升43%，每份调查成本约0.8美元；在WVS上展现出跨地域泛化潜力。与SFT对齐基线相比，P2P使用不到3%的训练数据即达到竞争性能。","reliability":"论文指出当前方法限于结构化调查问题，未处理自由文本输出；偏好重建依赖静态调查数据，可能无法应对非平稳偏好；未来需结合语义和心理测量技术提升鲁棒性。","relevance":"高度相关：该研究直接用LLM代理复现真实调查中的偏好分布，有严格的人类基准对照，涉及政策评估场景，并讨论了仿真失效条件，完全符合研究者的关注点，值得精读原文。","inspiration":"该方法通过两阶段系统（多样化代理池构建与L1正则化集成选择）实现无需微调、低成本的人类偏好复现，其自适应采样和稀疏集成思路值得借鉴｜可迁移到消费者金融决策偏好研究，如风险偏好、储蓄选择或投资组合配置的调查实验｜以LLM代理作为被试，施加不同金融信息框架处理，结果变量为风险资产选择比例，用美国消费者金融调查（SCF）的真实数据作为对照基准"}},{"id":"2509.09871","version":1,"title":"Emulating Public Opinion: A Proof-of-Concept of AI-Generated Synthetic Survey Responses for the Chilean Case","zh_title":"模拟民意：智利案例中AI生成合成调查回答的概念验证","abstract":"Large Language Models (LLMs) offer promising avenues for methodological and applied innovations in survey research by using synthetic respondents to emulate human answers and behaviour, potentially mitigating measurement and representation errors. However, the extent to which LLMs recover aggregate item distributions remains uncertain and downstream applications risk reproducing social stereotypes and biases inherited from training data. We evaluate the reliability of LLM-generated synthetic survey responses against ground-truth human responses from a Chilean public opinion probabilistic survey. Specifically, we benchmark 128 prompt-model-question triplets, generating 189,696 synthetic profiles, and pool performance metrics (i.e., accuracy, precision, recall, and F1-score) in a meta-analysis across 128 question-subsample pairs to test for biases along key sociodemographic dimensions. The evaluation spans OpenAI's GPT family and o-series reasoning models, as well as Llama and Qwen checkpoints. Three results stand out. First, synthetic responses achieve excellent performance on trust items (F1-score and accuracy > 0.90). Second, GPT-4o, GPT-4o-mini and Llama 4 Maverick perform comparably on this task. Third, synthetic-human alignment is highest among respondents aged 45-59. Overall, LLM-based synthetic samples approximate responses from a probabilistic sample, though with substantial item-level heterogeneity. Capturing the full nuance of public opinion remains challenging and requires careful calibration and additional distributional tests to ensure algorithmic fidelity and reduce errors.","authors":["Bastián González-Bustamante","Nando Verelst","Carla Cisternas"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2025-09-11","first_seen":"2025-09-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2509.09871","pdf_url":"https://arxiv.org/pdf/2509.09871","source_feed":"backfill","score":10,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["LLM仿真","调查方法","算法保真度"],"reason":"直接用LLM生成合成调查回答，与真实人类概率样本对照，评估可靠性与偏差，涉及公…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:27","error":null,"has_summary":true,"summary":{"generated_at":"2025-09-11","rank":11,"question":"大语言模型生成的合成调查回复能否准确反映智利的真实公众意见？","design":"使用128个提示-模型-问题三元组，生成189,696个合成档案，评估GPT系列、o系列推理模型、Llama和Qwen等模型在模拟智利公众意见调查回复上的表现，以准确率、精确率、召回率和F1分数为指标，并检验关键社会人口维度的偏差。","baseline":"智利公众意见概率调查的真实人类回复。","findings":"合成回复在信任项目上表现优异（F1和准确率>0.90）；GPT-4o、GPT-4o-mini和Llama 4 Maverick表现相当，且45-59岁年龄段的合成-人类对齐度最高。","reliability":"论文指出合成样本与概率样本的近似存在显著的项目层面异质性，捕捉公众意见的全部细微差别仍有挑战，需要仔细校准和额外的分布检验。","relevance":"该研究直接评估了LLM作为人类被试替代品的可靠性，有真实人类数据对照，并关注偏差，完全符合研究者的兴趣，值得精读。","inspiration":"该研究通过多模型、多提示的系统比较和关键社会人口维度的偏差检验来评估仿真可靠性，方法上值得借鉴｜可迁移到消费者信心指数或通胀预期调查的仿真上，检验LLM能否复现真实公众的经济预期分布｜以GPT-4o等为被试，用真实消费者调查问卷作为提示，生成通胀预期回复，以央行或统计局的真实调查数据为基准，比较分布一致性和子群体偏差"}},{"id":"2509.06337","version":2,"title":"Large Language Models as Virtual Survey Respondents: Evaluating Sociodemographic Response Generation","zh_title":"大语言模型作为虚拟调查受访者：评估社会人口响应生成","abstract":"Questionnaire-based surveys are foundational to social science research and public policymaking, yet traditional survey methods remain costly, time-consuming, and often limited in scale. Although prior work has explored large language models (LLMs) as virtual survey respondents, existing studies often address narrow task settings, focus on single sociological domains, or lack a unified evaluation framework that enables systematic comparison across diverse datasets and models. To address these gaps, we introduce two complementary task abstractions: Partial Attribute Simulation (PAS), where LLMs predict missing attributes from incomplete respondent profiles, and Full Attribute Simulation (FAS), where LLMs generate complete synthetic datasets under zero-context and context-enhanced conditions. Both are framed as diagnostic and exploratory tools rather than replacements for human data collection. We curate LLM-S^3 (Large Language Model-based Sociodemographic Survey Simulation), a benchmark spanning 11 real-world public datasets across four sociological domains, and evaluate GPT-3.5/4 Turbo and LLaMA 3.0/3.1-8B under zero-shot and few-shot settings. Our evaluation reveals consistent performance trends across model families, highlights failure modes in structured output generation, and demonstrates how context and prompt design affect simulation fidelity. Our code and dataset are available at: https://github.com/dart-lab-research/LLM-S-Cube-Benchmark","authors":["Jianpeng Zhao","Chenyu Yuan","Weiming Luo","Haoling Xie","Guangwei Zhang","Steven Jige Quan","Zixuan Yuan","Pengyang Wang","Denghui Zhang"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2025-09-08","first_seen":"2025-09-08","revised_at":null,"abs_url":"https://arxiv.org/abs/2509.06337","pdf_url":"https://arxiv.org/pdf/2509.06337","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2"],"tags":["LLM仿真","调查方法","社会人口模拟"],"reason":"用LLM模拟调查受访者，复现社会人口属性，有真实人类数据对照，涉及社会科学与政…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:26","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":35,"question":"如何系统评估大语言模型在模拟社会人口调查回答时的表现与可靠性？","design":"提出部分属性模拟（PAS）和全属性模拟（FAS）两种任务，使用GPT-3.5/4 Turbo和LLaMA 3.0/3.1-8B模型，在11个真实调查数据集上，以零样本和少样本设置生成缺失属性或完整合成数据集，并测量统计分布相似度。","baseline":"11个来自社会与公共事务、工作与收入、家庭与行为模式、健康与生活方式四个领域的真实公共调查数据集。","findings":"不同模型家族在模拟任务上表现出一致的性能趋势；提示设计和上下文增强显著影响模拟保真度，结构化输出生成中的失败案例仍是主要瓶颈。","reliability":"论文明确声明评估仅衡量统计和分布相似性，而非行为保真度或构念效度；数据集主要来自北美和欧洲，跨文化泛化性未验证；全属性模拟场景下结构化输出失败问题突出。","relevance":"该研究直接以真实人类调查数据为基准，系统评估LLM模拟社会人口属性的可靠性，并指出失效条件，高度契合研究者对仿真基准、政策评估场景和批判性分析的兴趣，值得精读原文。","inspiration":"借鉴其部分属性模拟（PAS）和全属性模拟（FAS）的任务设计，以真实调查数据为基准，通过统计分布相似度量化LLM的仿真保真度，并系统对比不同模型、提示策略和上下文增强的影响。｜可迁移到信贷审批中的歧视测量场景，用LLM模拟不同人口特征申请人的信用评分或审批结果，检验算法或人工决策中的统计性歧视。｜以真实信贷申请数据（如HMDA）为对照基准，将申请人的人口属性（种族、性别）作为处理变量，让LLM在PAS任务下生成信用评分或审批决策，比较LLM生成分布与真实审批分布的差异，并分析提示中是否加入反歧视法规对仿真偏差的影响。"}},{"id":"2509.03736","version":2,"title":"Are LLM Agents Behaviorally Coherent? Latent Profiles for Social Simulation","zh_title":"LLM代理行为一致吗？社会模拟的潜在画像","abstract":"The impressive capabilities of Large Language Models (LLMs) raise the possibility that synthetic agents can serve as substitutes for real participants in human-subject research. To evaluate this claim, prior research has largely focused on whether LLM-generated survey responses align with those produced by human respondents whom the LLMs are prompted to represent. In contrast, we address a more fundamental question: Do agents maintain empirical consistency; aligning to human behavioral models when examined under different experimental settings? To this end, we develop a study designed to (a) ask a set of questions which reveals an agent's latent profile and (b) examine agent behavioral consistency in a conversational setting with other agents. This design enables us to explore a set of behavioral hypotheses to assess whether an agent's conversational behavior is consistent with what we would expect from its revealed state. Our findings show significant inconsistencies in LLMs across model families and at differing model sizes. Most importantly, we find that, although agents may generate responses matching those of their human counterparts, they fail to be empirically consistent, representing a critical gap in their capabilities to accurately substitute for real participants in human-subject research.","authors":["James Mooney","Josef Woldense","Zheng Robert Jia","Shirley Anugrah Hayati","My Ha Nguyen","Vipul Raheja","Dongyeop Kang"],"categories":["cs.AI","cs.CL","cs.LG"],"primary_category":"cs.AI","announce_type":"new","date":"2025-09-03","first_seen":"2025-09-03","revised_at":null,"abs_url":"https://arxiv.org/abs/2509.03736","pdf_url":"https://arxiv.org/pdf/2509.03736","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","A4","B1","B4"],"tags":["LLM仿真","行为一致性","人类被试替代"],"reason":"直接评估LLM代理替代人类被试的行为一致性，有人类数据对照，并指出失效条件。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:25","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":46,"question":"LLM代理在跨实验设置下是否保持行为一致性，即其对话行为是否与其自身揭示的潜在状态（偏好和开放性）相符？","design":"构建一个五阶段框架：选择争议性话题，生成具有人口统计特征和话题偏见的代理，通过问卷获取代理的潜在状态（偏好和开放性），将代理配对进行多轮对话，评估对话中的一致性。通过系统变化系统提示引入控制变量，测试六种人类行为模型。","baseline":"无对照","findings":"LLM代理在聚合层面表现出一些趋势，但无法通过更严格的经验一致性检验：即使偏好对立，代理也很少维持分歧；偏见提示不能可靠恢复原则性分歧；共享负面情绪产生的对齐弱于共享正面情绪；开放性在最应起作用的场景中失去预测力。","reliability":"论文指出当前LLM代理在行为模型更复杂或更细粒度时无法复制人类行为，其行为一致性存在显著差距，但未讨论具体失效条件。","relevance":"该研究直接评估LLM代理替代人类被试的行为一致性，揭示其在复杂行为模型中失效，与研究者关注的仿真可靠性及失效条件高度相关，值得精读原文。","inspiration":"借鉴其通过系统提示注入不同行为模型（如偏见、开放性）并检验代理行为一致性的实验设计，可作为一种施加处理的方式。｜可迁移至政策公告的预期形成实验，研究不同信息框架下投资者对央行沟通的反应一致性。｜以LLM代理为被试，处理为系统提示中嵌入不同的政策沟通风格（鹰派/鸽派），结果变量为代理在模拟交易中的资产配置变化，对照真实央行公告后的市场调查数据。"}},{"id":"2509.02879","version":1,"title":"Artificial or Human Intelligence?","zh_title":"人工智能还是人类智能？","abstract":"Artificial intelligence (AI) tools such as large language models (LLMs) are already altering student learning. Unlike previous technologies, LLMs can independently solve problems regardless of student understanding, yet are not always accurate (due to hallucination) and face sharp performance cutoffs (due to emergence). Access to these tools significantly alters a student's incentives to learn, potentially decreasing the sum knowledge of humans and AI. Additionally, the marginal benefit of learning changes depending on which side of the AI frontier a human is on, creating a discontinuous gap between those that know more than or less than AI. This contrasts with downstream models of AI's impact on the labor force which assume continuous ability. Finally, increasing the portion of assignments where AI cannot be used can counteract student mis-specification about AI accuracy, preventing underinvestment. A better understanding of how AI impacts learning and student incentives is crucial for educators to adapt to this new technology.","authors":["Eric Gao"],"categories":["econ.TH"],"primary_category":"econ.TH","announce_type":"new","date":"2025-09-02","first_seen":"2025-09-02","revised_at":null,"abs_url":"https://arxiv.org/abs/2509.02879","pdf_url":"https://arxiv.org/pdf/2509.02879","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["AI教育","学生激励","多智能体"],"reason":"纯多智能体协作解题，无人类行为对照，不涉及仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:37","error":null,"has_summary":false,"summary":null},{"id":"2509.01813","version":3,"title":"ShortageSim: Simulating Drug Shortages under Information Asymmetry","zh_title":"ShortageSim：信息不对称下的药品短缺仿真","abstract":"Drug shortages pose critical risks to patient care and healthcare systems worldwide, yet the effectiveness of regulatory interventions remains poorly understood due to information asymmetries in pharmaceutical supply chains. We propose \\textbf{ShortageSim}, addresses this challenge by providing the first simulation framework that evaluates the impact of regulatory interventions on competition dynamics under information asymmetry. Using Large Language Model (LLM)-based agents, the framework models the strategic decisions of drug manufacturers and institutional buyers, in response to shortage alerts given by the regulatory agency. Unlike traditional game theory models that assume perfect rationality and complete information, ShortageSim simulates heterogeneous interpretations on regulatory announcements and the resulting decisions. Experiments on self-processed dataset of historical shortage events show that ShortageSim reduces the resolution lag for production disruption cases by up to 84\\%, achieving closer alignment to real-world trajectories than the zero-shot baseline. Our framework confirms the effect of regulatory alert in addressing shortages and introduces a new method for understanding competition in multi-stage environments under uncertainty. We open-source ShortageSim and a dataset of 2,925 FDA shortage events, providing a novel framework for future research on policy design and testing in supply chains under information asymmetry.","authors":["Mingxuan Cui","Yilan Jiang","Duo Zhou","Cheng Qian","Yuji Zhang","Qiong Wang"],"categories":["cs.MA","cs.CL","cs.GT"],"primary_category":"cs.MA","announce_type":"new","date":"2025-09-01","first_seen":"2025-09-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2509.01813","pdf_url":"https://arxiv.org/pdf/2509.01813","source_feed":"backfill","score":7,"bucket":"pending","rubric_hits":["A3","B2","B4"],"tags":["LLM仿真","供应链博弈","政策评估"],"reason":"用LLM agent模拟药企与采购方决策，评估监管干预效果，涉及经济学场景与信…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:25","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":109,"question":"在信息不对称的药品供应链中，监管机构发布短缺预警如何影响制药商和采购方的策略性决策，进而影响药品短缺的解决速度？","design":"使用LLM驱动的多智能体框架，模拟FDA监管机构、制药商和医疗机构采购方在信息不对称环境下的决策。监管机构发布短缺预警，各智能体根据预警解读市场状态并做出生产或采购决策，测量短缺解决时滞和与历史轨迹的偏差。","baseline":"以FDA历史短缺事件数据集中的51条已解决事件轨迹作为真实对照基准。","findings":"ShortageSim在模拟生产中断案例时，将短缺解决时滞最多降低了84%，且比零样本基线更贴近真实历史轨迹；同时确认了主动预警政策可能引发囤货行为，反而加剧短缺。","reliability":"论文未讨论","relevance":"该研究用LLM智能体模拟经济学场景下的决策行为，并与真实FDA短缺事件数据对照，评估监管干预效果，直接命中研究者关注的LLM仿真人类决策、政策评估和真实数据基准。值得精读原文。","inspiration":"该方法将LLM智能体置于信息不对称的供应链博弈中，通过历史事件轨迹校准仿真输出，并对比零样本基线以验证框架有效性，值得借鉴其‘真实数据校准+基线对比’的验证设计｜可迁移至金融市场政策公告的预期形成研究，如央行沟通对投资者决策的影响｜以LLM模拟投资者和央行官员，处理为央行发布不同透明度的政策指引，结果变量为资产价格波动和交易量，用历史央行公告前后的市场数据作为真实对照基准"}},{"id":"2510.07321","version":1,"title":"How human is the machine? Evidence from 66,000 Conversations with Large Language Models","zh_title":"机器有多像人？来自66000次与大语言模型对话的证据","abstract":"When Artificial Intelligence (AI) is used to replace consumers (e.g., synthetic data), it is often assumed that AI emulates established consumers, and more generally human behaviors. Ten experiments with Large Language Models (LLMs) investigate if this is true in the domain of well-documented biases and heuristics. Across studies we observe four distinct types of deviations from human-like behavior. First, in some cases, LLMs reduce or correct biases observed in humans. Second, in other cases, LLMs amplify these same biases. Third, and perhaps most intriguingly, LLMs sometimes exhibit biases opposite to those found in humans. Fourth, LLMs' responses to the same (or similar) prompts tend to be inconsistent (a) within the same model after a time delay, (b) across models, and (c) among independent research studies. Such inconsistencies can be uncharacteristic of humans and suggest that, at least at one point, LLMs' responses differed from humans. Overall, unhuman-like responses are problematic when LLMs are used to mimic or predict consumer behavior. These findings complement research on synthetic consumer data by showing that sources of bias are not necessarily human-centric. They also contribute to the debate about the tasks for which consumers, and more generally humans, can be replaced by AI.","authors":["Antonios Stamatogiannakis","Arsham Ghodsinia","Sepehr Etminanrad","Dilney Gonçalves","David Santos"],"categories":["cs.HC","econ.GN"],"primary_category":"cs.HC","announce_type":"new","date":"2025-08-31","first_seen":"2025-08-31","revised_at":null,"abs_url":"https://arxiv.org/abs/2510.07321","pdf_url":"https://arxiv.org/pdf/2510.07321","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","认知偏差","人类数据对照"],"reason":"用LLM复现人类认知偏差并与真实人类数据对照，评估仿真可靠性并指出失效条件","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:32","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":23,"question":"大语言模型在经典决策偏差与启发式任务中，其行为在多大程度上与人类相似？","design":"使用多个大语言模型（如GPT系列）进行10项预注册实验，共66,000次对话，系统操纵模型类型和提示语特征，测量模型在可得性启发、代表性启发、禀赋效应、锚定效应、交易效用和框架效应六种偏差上的反应。","baseline":"以心理学和行为经济学文献中已确立的人类在这些偏差上的典型行为模式作为对照基准。","findings":"LLM表现出四种与人类不同的偏差模式：减弱或纠正人类偏差、放大人类偏差、表现出与人类相反的偏差，以及在不同时间、模型和研究间反应不一致。这些非人类反应表明LLM在模拟或预测消费者行为时存在问题。","reliability":"论文指出LLM的反应会随时间、模型版本和不同研究间出现不一致，这种不一致在人类中不典型，提示在某一时点LLM的反应与人类不同，从而限制了其作为人类替代品的可靠性。","relevance":"该研究直接评估LLM作为人类被试替代品的可靠性，通过经典决策偏差实验与真实人类行为基准对照，系统揭示了仿真失效的四种模式，对关注LLM仿真效度的研究者具有重要参考价值，值得阅读原文。","inspiration":"借鉴其多模型、多偏差、大样本的系统比较设计，可迁移到经济金融中的投资者行为偏差研究（如处置效应、过度自信），设计雏形：以GPT-4等LLM为被试，呈现模拟股票交易场景测量处置效应，结果与真实投资者交易数据（如券商账户记录）进行对照。"}},{"id":"2510.06222","version":1,"title":"Inducing State Anxiety in LLM Agents Reproduces Human-Like Biases in Consumer Decision-Making","zh_title":"在LLM智能体中诱导状态焦虑可复现消费者决策中的人类偏差","abstract":"Large language models (LLMs) are rapidly evolving from text generators to autonomous agents, raising urgent questions about their reliability in real-world contexts. Stress and anxiety are well known to bias human decision-making, particularly in consumer choices. Here, we tested whether LLM agents exhibit analogous vulnerabilities. Three advanced models (ChatGPT-5, Gemini 2.5, Claude 3.5-Sonnet) performed a grocery shopping task under budget constraints (24, 54, 108 USD), before and after exposure to anxiety-inducing traumatic narratives. Across 2,250 runs, traumatic prompts consistently reduced the nutritional quality of shopping baskets (Change in Basket Health Scores of -0.081 to -0.126; all pFDR<0.001; Cohens d=-1.07 to -2.05), robust across models and budgets. These results show that psychological context can systematically alter not only what LLMs generate but also the actions they perform. By reproducing human-like emotional biases in consumer behavior, LLM agents reveal a new class of vulnerabilities with implications for digital health, consumer safety, and ethical AI deployment.","authors":["Ziv Ben-Zion","Zohar Elyoseph","Tobias Spiller","Teddy Lazebnik"],"categories":["cs.HC","econ.GN"],"primary_category":"cs.HC","announce_type":"new","date":"2025-08-30","first_seen":"2025-08-30","revised_at":null,"abs_url":"https://arxiv.org/abs/2510.06222","pdf_url":"https://arxiv.org/pdf/2510.06222","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["LLM仿真","消费者行为","焦虑偏差"],"reason":"用LLM模拟消费者决策，诱导焦虑后复现人类偏差，有真实人类数据对照，涉及行为经…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:30","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":24,"question":"心理上下文（如焦虑诱导）是否会系统性地改变LLM智能体在消费者决策中的行为，使其表现出类似人类的情绪偏差？","design":"使用ChatGPT-5、Gemini 2.5、Claude 3.5-Sonnet三个先进LLM作为智能体，模拟人类消费者在预算约束（27、54、108美元）下进行杂货购物任务；通过暴露于焦虑诱导的创伤叙事作为处理，测量购物篮健康评分的变化。","baseline":"无对照","findings":"焦虑诱导提示一致降低了购物篮的营养质量（健康评分变化Δ=-0.081至-0.126，所有pFDR<0.001，Cohen's d=-1.07至-2.05），效应在不同模型和预算水平下均稳健。结果表明心理上下文不仅能改变LLM生成的文本，还能系统性地改变其作为智能体所执行的动作。","reliability":"论文未讨论","relevance":"该研究直接以LLM模拟人类消费者决策，通过情绪诱导复现人类偏差，并评估效应稳健性，高度契合研究者对LLM仿真可靠性及偏差的关注，值得精读原文以了解其方法细节和潜在失效边界。","inspiration":"借鉴其通过情绪化提示（创伤叙事）施加心理状态处理、以预算约束任务测量经济决策结果、并跨模型和预算水平进行稳健性检验的设计方法｜可迁移至消费者跨期选择或风险决策场景，如焦虑对储蓄/借贷行为或保险购买决策的影响｜以LLM智能体为被试，施加焦虑诱导处理，测量其在跨期选择任务中的贴现率或风险资产配置比例，并以真实人类实验数据（如实验室或调查数据）作为对照基准。"}},{"id":"2509.00462","version":4,"title":"AI Self-preferencing in Algorithmic Hiring: Empirical Evidence and Insights","zh_title":"算法招聘中的人工智能自我偏好：实证证据与见解","abstract":"As artificial intelligence (AI) tools become widely adopted, large language models (LLMs) are increasingly involved on both sides of decision-making processes, ranging from hiring to content moderation. This dual adoption raises a critical question: do LLMs systematically favor content that resembles their own outputs? Prior research in computer science has identified self-preference bias -- the tendency of LLMs to favor their own generated content -- but its real-world implications have not been empirically evaluated. We focus on the hiring context, where job applicants often rely on LLMs to refine resumes, while employers deploy them to screen those same resumes. Using a large-scale controlled resume correspondence experiment, we find that LLMs consistently prefer resumes generated by themselves over those written by humans or produced by alternative models, even when content quality is controlled. The bias against human-written resumes is particularly substantial, with self-preference bias ranging from 67% to 82% across major commercial and open-source models. To assess labor market impact, we simulate realistic hiring pipelines across 24 occupations. These simulations show that candidates using the same LLM as the evaluator are 23% to 60% more likely to be shortlisted than equally qualified applicants submitting human-written resumes, with the largest disadvantages observed in business-related fields such as sales and accounting. We further demonstrate that this bias can be reduced by more than 50% through simple interventions targeting LLMs' self-recognition capabilities. These findings highlight an emerging but previously overlooked risk in AI-assisted decision making and call for expanded frameworks of AI fairness that address not only demographic-based disparities, but also biases in AI-AI interactions.","authors":["Jiannan Xu","Gujie Li","Jane Yi Jiang"],"categories":["cs.CY"],"primary_category":"cs.CY","announce_type":"new","date":"2025-08-30","first_seen":"2025-08-30","revised_at":null,"abs_url":"https://arxiv.org/abs/2509.00462","pdf_url":"https://arxiv.org/pdf/2509.00462","source_feed":"backfill","score":7,"bucket":"pending","rubric_hits":["A1","B1","B2","B4"],"tags":["LLM仿真","招聘实验","算法偏差"],"reason":"用LLM模拟招聘决策并与人类简历对照，揭示自偏好偏差，可迁移至人类仿真研究。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:24","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":110,"question":"大语言模型在简历筛选任务中是否系统性地偏好自己生成的简历，从而产生自偏好偏差？","design":"使用真实人类简历数据集，用GPT-4o、LLaMA 3.3-70B等7个主流LLM为每份简历生成多个反事实版本，再让这些模型作为评估者对简历进行评分或排序，测量对自身生成简历的偏好程度。","baseline":"来自专业简历平台的真实人类撰写简历，收集于生成式AI广泛采用之前。","findings":"多数LLM在评估时强烈偏好自己生成的简历，对人类简历的偏差达67%-82%；在模拟招聘中，使用与评估者相同LLM的候选人入围概率比同等条件的人类简历申请人高23%-60%，商业类职业劣势最明显。","reliability":"论文未讨论","relevance":"该研究用LLM模拟招聘决策并与真实人类简历对照，揭示了AI自偏好偏差，为人类仿真实验的可靠性与偏差评估提供了关键证据，值得精读。","inspiration":"该方法通过让LLM生成反事实简历并评估自身生成内容，巧妙测量自偏好偏差，可借鉴其对照设计｜可迁移至信贷审批歧视研究，检验AI是否对自身生成的贷款申请更宽容｜以LLM作为信贷审批员，处理组为LLM生成的贷款申请，对照组为真实申请人数据，结果变量为批准率，用历史信贷记录作基准"}},{"id":"2509.02605","version":1,"title":"Synthetic Founders: AI-Generated Social Simulations for Startup Validation Research in Computational Social Science","zh_title":"合成创始人：用于计算社会科学中创业验证研究的AI生成社会仿真","abstract":"We present a comparative docking experiment that aligns human-subject interview data with large language model (LLM)-driven synthetic personas to evaluate fidelity, divergence, and blind spots in AI-enabled simulation. Fifteen early-stage startup founders were interviewed about their hopes and concerns regarding AI-powered validation, and the same protocol was replicated with AI-generated founder and investor personas. A structured thematic synthesis revealed four categories of outcomes: (1) Convergent themes - commitment-based demand signals, black-box trust barriers, and efficiency gains were consistently emphasized across both datasets; (2) Partial overlaps - founders worried about outliers being averaged away and the stress of real customer validation, while synthetic personas highlighted irrational blind spots and framed AI as a psychological buffer; (3) Human-only themes - relational and advocacy value from early customer engagement and skepticism toward moonshot markets; and (4) Synthetic-only themes - amplified false positives and trauma blind spots, where AI may overstate adoption potential by missing negative historical experiences. We interpret this comparative framework as evidence that LLM-driven personas constitute a form of hybrid social simulation: more linguistically expressive and adaptable than traditional rule-based agents, yet bounded by the absence of lived history and relational consequence. Rather than replacing empirical studies, we argue they function as a complementary simulation category - capable of extending hypothesis space, accelerating exploratory validation, and clarifying the boundaries of cognitive realism in computational social science.","authors":["Jorn K. Teutloff"],"categories":["cs.MA","cs.AI","cs.CY"],"primary_category":"cs.MA","announce_type":"new","date":"2025-08-29","first_seen":"2025-08-29","revised_at":null,"abs_url":"https://arxiv.org/abs/2509.02605","pdf_url":"https://arxiv.org/pdf/2509.02605","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM人类仿真","创业者访谈对照","仿真保真度评估"],"reason":"用LLM生成合成创业者进行访谈仿真，并与15位真人创业者数据对照，评估仿真保真…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:25","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":47,"question":"LLM驱动的合成创业者与投资人角色在创业验证访谈中，与真实人类受访者的主题模式在多大程度上一致、偏离或存在盲点？","design":"采用对比对接实验：用LLM生成35个合成创业者与投资人角色，复制对15位真实早期创业者的半结构化访谈协议，询问其对AI驱动市场验证的希望与担忧，通过结构化主题综合比较合成与人类访谈文本的涌现主题。","baseline":"15位真实早期创业者的访谈数据，涵盖其对AI驱动市场验证的希望与担忧。","findings":"合成与人类数据在承诺型需求信号、黑箱信任障碍和效率提升上主题收敛；部分重叠中，人类担忧异常值被平均化和真实客户验证压力，合成角色则强调非理性盲点并将AI视为心理缓冲；人类独有主题包括早期客户参与的倡导价值和对极端市场的怀疑，合成独有主题为放大假阳性和创伤盲点。","reliability":"论文指出LLM角色缺乏真实生活经历和关系后果，可能高估采纳潜力并遗漏负面历史经验，因此不能替代实证研究，仅作为补充性仿真工具，用于扩展假设空间和加速探索性验证。","relevance":"该研究直接以真实人类访谈为基准，系统评估LLM仿真在创业决策中的保真度与偏差，并明确区分收敛、部分重叠和独有主题，对关注仿真可靠性及失效条件的研究者具有重要参考价值，值得细读原文。","inspiration":"该方法通过对比合成角色与真实人类在相同半结构化访谈下的主题收敛与偏离，系统评估LLM仿真的保真度与盲点，值得借鉴｜可迁移至创业融资决策研究，如投资者对AI辅助商业计划书的评估偏差｜以LLM生成的投资人角色为被试，呈现AI生成与人类撰写的商业计划书，测量投资意愿与风险评估，并以真实天使投资人的评审数据作为对照基准"}},{"id":"2509.02596","version":1,"title":"Introducing LCOAI: A Standardized Economic Metric for Evaluating AI Deployment Costs","zh_title":"引入LCOAI：评估AI部署成本的标准化经济指标","abstract":"As artificial intelligence (AI) becomes foundational to enterprise infrastructure, organizations face growing challenges in accurately assessing the full economic implications of AI deployment. Existing metrics such as API token costs, GPU-hour billing, or Total Cost of Ownership (TCO) fail to capture the complete lifecycle costs of AI systems and provide limited comparability across deployment models. This paper introduces the Levelized Cost of Artificial Intelligence (LCOAI), a standardized economic metric designed to quantify the total capital (CAPEX) and operational (OPEX) expenditures per unit of productive AI output, normalized by valid inference volume. Analogous to established metrics like LCOE (levelized cost of electricity) and LCOH (levelized cost of hydrogen) in the energy sector, LCOAI offers a rigorous, transparent framework to evaluate and compare the cost-efficiency of vendor API deployments versus self-hosted, fine-tuned models. We define the LCOAI methodology in detail and apply it to three representative scenarios, OpenAI GPT-4.1 API, Anthropic Claude Haiku API, and a self-hosted LLaMA-2-13B deployment demonstrating how LCOAI captures critical trade-offs in scalability, investment planning, and cost optimization. Extensive sensitivity analyses further explore the impact of inference volume, CAPEX, and OPEX variability on lifecycle economics. The results illustrate the practical utility of LCOAI in procurement, infrastructure planning, and automation strategy, and establish it as a foundational benchmark for AI economic analysis. Policy implications and areas for future refinement, including environmental and performance-adjusted cost metrics, are also discussed.","authors":["Eliseo Curcio"],"categories":["econ.GN","eess.SY"],"primary_category":"econ.GN","announce_type":"new","date":"2025-08-29","first_seen":"2025-08-29","revised_at":null,"abs_url":"https://arxiv.org/abs/2509.02596","pdf_url":"https://arxiv.org/pdf/2509.02596","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["AI经济学","成本指标","部署评估"],"reason":"论文提出AI部署成本指标LCOAI，属经济评估框架，不涉及LLM仿真人类被试或…","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:58:39","error":null,"has_summary":false,"summary":null},{"id":"2508.20234","version":1,"title":"Validating Generative Agent-Based Models for Logistics and Supply Chain Management Research","zh_title":"验证基于生成式智能体的物流与供应链管理研究模型","abstract":"Generative Agent-Based Models (GABMs) powered by large language models (LLMs) offer promising potential for empirical logistics and supply chain management (LSCM) research by enabling realistic simulation of complex human behaviors. Unlike traditional agent-based models, GABMs generate human-like responses through natural language reasoning, which creates potential for new perspectives on emergent LSCM phenomena. However, the validity of LLMs as proxies for human behavior in LSCM simulations is unknown. This study evaluates LLM equivalence of human behavior through a controlled experiment examining dyadic customer-worker engagements in food delivery scenarios. I test six state-of-the-art LLMs against 957 human participants (477 dyads) using a moderated mediation design. This study reveals a need to validate GABMs on two levels: (1) human equivalence testing, and (2) decision process validation. Results reveal GABMs can effectively simulate human behaviors in LSCM; however, an equivalence-versus-process paradox emerges. While a series of Two One-Sided Tests (TOST) for equivalence reveals some LLMs demonstrate surface-level equivalence to humans, structural equation modeling (SEM) reveals artificial decision processes not present in human participants for some LLMs. These findings show GABMs as a potentially viable methodological instrument in LSCM with proper validation checks. The dual-validation framework also provides LSCM researchers with a guide to rigorous GABM development. For practitioners, this study offers evidence-based assessment for LLM selection for operational tasks.","authors":["Vincent E. Castillo"],"categories":["cs.MA","cs.AI","cs.CY"],"primary_category":"cs.MA","announce_type":"new","date":"2025-08-27","first_seen":"2025-08-27","revised_at":null,"abs_url":"https://arxiv.org/abs/2508.20234","pdf_url":"https://arxiv.org/pdf/2508.20234","source_feed":"backfill","score":10,"bucket":"selected","rubric_hits":["A1","A2","A3","B1","B2","B4"],"tags":["LLM人类仿真","等效性验证","供应链管理"],"reason":"直接验证LLM作为人类代理在供应链场景中的等效性，有957名人类对照，并揭示仿…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:23","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":6,"question":"在物流与供应链管理（LSCM）的二元互动场景中，基于大语言模型的生成式智能体模型能否在表面行为上与人类等效，且其决策过程是否与人类可比？","design":"采用受控实验，将六种最先进的大语言模型作为生成式智能体，模拟外卖配送场景中顾客与骑手的二元互动，与957名人类参与者（477对）进行对比，使用有调节的中介设计，测量互动结果与决策过程。","baseline":"957名人类参与者（477对二元组）在相同外卖配送场景实验中的真实行为数据。","findings":"部分LLM在表面行为上通过等效性检验（TOST）与人类无显著差异，但结构方程模型（SEM）揭示其决策过程存在人为模式，与人类真实决策过程不同，形成“等效-过程悖论”。","reliability":"论文指出LLM作为人类代理的有效性未知，需进行双重验证（人类等效性检验和决策过程验证），并承认某些LLM的决策过程与人类不符，但未详细讨论其他失效条件。","relevance":"该研究直接验证LLM在供应链场景中替代人类被试的等效性，提供大规模人类对照基准，并批判性揭示表面等效下的决策过程差异，高度契合研究者对仿真可靠性、偏差及失效条件的关注，值得精读原文。","inspiration":"该研究采用TOST等效性检验与结构方程模型（SEM）双重验证，区分表面行为等效与决策过程等效，为仿真可靠性评估提供了严谨框架｜可迁移至消费者跨期选择实验，检验LLM生成的折现行为是否与人类一致｜以LLM为被试，施加不同跨期奖励方案，结果变量为选择时间偏好，以真实人类实验数据（如Andersen et al., 2008）为基准，进行TOST和SEM双重验证"}},{"id":"2508.19004","version":1,"title":"AI Models Exceed Individual Human Accuracy in Predicting Everyday Social Norms","zh_title":"AI模型在预测日常社会规范方面超越个体人类准确性","abstract":"A fundamental question in cognitive science concerns how social norms are acquired and represented. While humans typically learn norms through embodied social experience, we investigated whether large language models can achieve sophisticated norm understanding through statistical learning alone. Across two studies, we systematically evaluated multiple AI systems' ability to predict human social appropriateness judgments for 555 everyday scenarios by examining how closely they predicted the average judgment compared to each human participant. In Study 1, GPT-4.5's accuracy in predicting the collective judgment on a continuous scale exceeded that of every human participant (100th percentile). Study 2 replicated this, with Gemini 2.5 Pro outperforming 98.7% of humans, GPT-5 97.8%, and Claude Sonnet 4 96.0%. Despite this predictive power, all models showed systematic, correlated errors. These findings demonstrate that sophisticated models of social cognition can emerge from statistical learning over linguistic data alone, challenging strong versions of theories emphasizing the exclusive necessity of embodied experience for cultural competence. The systematic nature of AI limitations across different architectures indicates potential boundaries of pattern-based social understanding, while the models' ability to outperform nearly all individual humans in this predictive task suggests that language serves as a remarkably rich repository for cultural knowledge transmission.","authors":["Pontus Strimling","Simon Karlsson","Irina Vartanova","Kimmo Eriksson"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2025-08-26","first_seen":"2025-08-26","revised_at":null,"abs_url":"https://arxiv.org/abs/2508.19004","pdf_url":"https://arxiv.org/pdf/2508.19004","source_feed":"backfill","score":8,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["社会规范预测","人类数据对照","算法偏差"],"reason":"用LLM预测人类社会规范判断，与真实人类数据对照，评估预测准确性与系统性偏差，…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:22","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":30,"question":"大语言模型能否仅通过文本统计学习，在没有具身体验的情况下，达到甚至超越个体人类对日常社会规范（行为适当性）的判断水平？","design":"本研究并非严格意义上的仿真实验，而是利用已有的大规模人类判断数据集，将多个大语言模型（GPT-4.5、Gemini 2.5 Pro、GPT-5、Claude Sonnet 4）作为“被试”，要求它们对555个日常场景中的行为适当性进行连续评分，并与真实人类个体评分进行对比。","baseline":"真实人类数据：来自大规模数据集中个体参与者对相同555个日常场景的社会适当性连续评分，包括平均集体判断和每个个体的判断分布。","findings":"GPT-4.5 预测集体判断的准确度超过了所有人类参与者（百分位100%），其他模型也超过了96%以上的人类个体。所有模型均表现出系统性、相互关联的误差，表明基于模式的统计学习在社会理解上存在边界。","reliability":"论文指出，尽管模型预测力强，但所有模型都存在系统性且相互关联的误差，这表明基于模式的社会理解存在潜在边界；研究仅限于日常社会规范判断，未涉及其他类型的社会认知任务。","relevance":"该研究直接以真实人类个体判断为基准，评估LLM在连续社会规范判断上的仿真准确性及系统性偏差，高度契合研究者对LLM仿真可靠性、偏差及失效条件的关注，值得精读原文以了解其个体级对比方法和误差分析。","inspiration":"可借鉴其将模型预测与人类个体分布而非仅与均值对比的评估设计，以揭示仿真在个体差异捕捉上的能力。｜可迁移至消费者对金融产品适当性的感知或政策可接受性判断等场景。｜以LLM作为“消费者”被试，输入不同金融产品描述，要求其评估产品适当性，结果变量为适当性连续评分，对照真实消费者调查数据中的个体评分分布。"}},{"id":"2509.00074","version":2,"title":"Language and Experience: A Computational Model of Social Learning in Complex Tasks","zh_title":"语言与经验：复杂任务中社会学习的计算模型","abstract":"The ability to combine linguistic guidance from others with direct experience is central to human development, enabling safe and rapid learning in new environments. How do people integrate these two sources of knowledge, and how might AI systems? We present a computational framework that models social learning as joint probabilistic inference over structured, executable world models given sensorimotor and linguistic data. We make this possible by turning a pretrained language model into a probabilistic model of how humans share advice conditioned on their beliefs, allowing our agents both to generate advice for others and to interpret linguistic input as evidence during Bayesian inference. Using behavioral experiments and simulations across 10 video games, we show how linguistic guidance can shape exploration and accelerate learning by reducing risky interactions and speeding up key discoveries in both humans and models. We further explore how knowledge can accumulate across generations through iterated learning experiments and demonstrate successful knowledge transfer between humans and models -- revealing how structured, language-compatible representations might enable human-machine collaborative learning.","authors":["Cédric Colas","Tracey Mills","Ben Prystawski","Michael Henry Tessler","Noah Goodman","Jacob Andreas","Joshua Tenenbaum"],"categories":["cs.AI","cs.CL","cs.LG"],"primary_category":"cs.AI","announce_type":"new","date":"2025-08-26","first_seen":"2025-08-26","revised_at":null,"abs_url":"https://arxiv.org/abs/2509.00074","pdf_url":"https://arxiv.org/pdf/2509.00074","source_feed":"backfill","score":7,"bucket":"pending","rubric_hits":["A1","B1","B2"],"tags":["LLM仿真","人类行为对照","社会学习"],"reason":"用LLM模拟人类在复杂任务中的社会学习，并与真实人类行为实验对照，涉及行为实验…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:36","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":111,"question":"人类如何整合语言指导与直接经验来学习复杂任务？","design":"提出一个贝叶斯计算框架，将预训练语言模型作为人类说话者模型，模拟人类在10款视频游戏中结合语言建议和直接经验进行社会学习；通过行为实验和模拟，比较纯体验、体验+人类消息、体验+模型消息三种条件，测量通关所需生命数、通关比例和归一化曲线下面积（nAUC）。","baseline":"122名Prolific参与者的人类行为数据，最终有效样本120人，随机分配到三种条件（各40人），记录其游戏表现和给出的建议。","findings":"语言指导能塑造探索策略并加速学习，减少风险交互并加快关键发现；知识可通过迭代学习在代际间积累，并实现人类与模型之间的成功知识迁移。","reliability":"论文未讨论","relevance":"该研究用LLM模拟人类在复杂任务中的社会学习，并与真实人类行为实验对照，涉及行为实验和代际知识传递，直接命中研究者对LLM仿真人类被试、复现行为并评估可靠性的兴趣，值得精读原文。","inspiration":"该研究用LLM模拟人类在复杂任务中结合语言建议和直接经验的学习过程，并设置纯体验、体验+人类消息、体验+模型消息三种条件进行对照，这种多条件对比设计可用于评估语言干预的因果效应｜可迁移到经济金融中的政策沟通与预期形成场景，例如研究央行前瞻指引如何影响投资者学习与资产配置｜以LLM模拟投资者，处理为提供不同风格（如模糊vs精确）的央行声明，结果变量为投资组合调整速度和准确性，用历史市场数据或人类实验数据作为基准对照"}},{"id":"2508.17322","version":1,"title":"Chinese Court Simulation with LLM-Based Agent System","zh_title":"基于大语言模型智能体的中国法庭仿真","abstract":"Mock trial has long served as an important platform for legal professional training and education. It not only helps students learn about realistic trial procedures, but also provides practical value for case analysis and judgment prediction. Traditional mock trials are difficult to access by the public because they rely on professional tutors and human participants. Fortunately, the rise of large language models (LLMs) provides new opportunities for creating more accessible and scalable court simulations. While promising, existing research mainly focuses on agent construction while ignoring the systematic design and evaluation of court simulations, which are actually more important for the credibility and usage of court simulation in practice. To this end, we present the first court simulation framework -- SimCourt -- based on the real-world procedure structure of Chinese courts. Our framework replicates all 5 core stages of a Chinese trial and incorporates 5 courtroom roles, faithfully following the procedural definitions in China. To simulate trial participants with different roles, we propose and craft legal agents equipped with memory, planning, and reflection abilities. Experiment on legal judgment prediction show that our framework can generate simulated trials that better guide the system to predict the imprisonment, probation, and fine of each case. Further annotations by human experts show that agents' responses under our simulation framework even outperformed judges and lawyers from the real trials in many scenarios. These further demonstrate the potential of LLM-based court simulation.","authors":["Kaiyuan Zhang","Jiaqi Li","Yueyue Wu","Haitao Li","Cheng Luo","Shaokun Zou","Yujia Zhou","Weihang Su","Qingyao Ai","Yiqun Liu"],"categories":["cs.CY","cs.AI"],"primary_category":"cs.CY","announce_type":"new","date":"2025-08-24","first_seen":"2025-08-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2508.17322","pdf_url":"https://arxiv.org/pdf/2508.17322","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["法庭仿真","多智能体","法律预测"],"reason":"模拟法庭过程但无真实人类行为对照，属社会模拟缺基准","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:36","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":168,"question":"如何基于中国真实庭审程序构建一个全阶段、多角色的LLM法庭模拟框架，并系统评估其判决预测准确性和过程质量？","design":"本研究提出SimCourt框架，基于中国刑事庭审的5个阶段和5种角色，为法官、检察官、律师、被告和书记员分别设计配备记忆、规划和反思模块及法律检索工具的LLM智能体，输入案件材料后自动生成完整庭审记录和判决书。","baseline":"无对照","findings":"SimCourt生成的模拟庭审能更好地指导系统预测监禁、缓刑和罚金等判决结果；人类专家标注显示，智能体在模拟框架下的回应在许多场景中甚至优于真实庭审中的法官和律师。","reliability":"论文未讨论","relevance":"该研究属于LLM社会模拟，但缺乏真实人类行为对照，未复现人类被试的决策分布或偏差，不符合研究者对基准人类数据的要求，不建议优先阅读原文。","inspiration":"SimCourt的多阶段、多角色框架和记忆-规划-反思模块设计，为构建复杂经济决策场景的LLM仿真提供了可借鉴的架构｜该设计可迁移到金融监管政策评估，如模拟银行、企业、监管机构等多方博弈对信贷供给的影响｜可构建包含银行、企业、监管者的LLM智能体，施加资本充足率变动处理，观察信贷审批决策，并以真实银行贷款数据和监管报告作为对照基准"}},{"id":"2508.16172","version":2,"title":"Graph RAG as Human Choice Model: Building a Data-Driven Mobility Agent with Preference Chain","zh_title":"图RAG作为人类选择模型：构建数据驱动的出行智能体与偏好链","abstract":"Understanding human behavior in urban environments is a crucial field within city sciences. However, collecting accurate behavioral data, particularly in newly developed areas, poses significant challenges. Recent advances in generative agents, powered by Large Language Models (LLMs), have shown promise in simulating human behaviors without relying on extensive datasets. Nevertheless, these methods often struggle with generating consistent, context-sensitive, and realistic behavioral outputs. To address these limitations, this paper introduces the Preference Chain, a novel method that integrates Graph Retrieval-Augmented Generation (RAG) with LLMs to enhance context-aware simulation of human behavior in transportation systems. Experiments conducted on the Replica dataset demonstrate that the Preference Chain outperforms standard LLM in aligning with real-world transportation mode choices. The development of the Mobility Agent highlights potential applications of proposed method in urban mobility modeling for emerging cities, personalized travel behavior analysis, and dynamic traffic forecasting. Despite limitations such as slow inference and the risk of hallucination, the method offers a promising framework for simulating complex human behavior in data-scarce environments, where traditional data-driven models struggle due to limited data availability.","authors":["Kai Hu","Parfait Atchade-Adelomou","Carlo Adornetto","Adrian Mora-Carrero","Luis Alonso-Pastor","Ariel Noyman","Yubo Liu","Kent Larson"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2025-08-22","first_seen":"2025-08-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2508.16172","pdf_url":"https://arxiv.org/pdf/2508.16172","source_feed":"backfill","score":8,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM仿真","出行行为","人类数据对照"],"reason":"用LLM仿真交通出行选择，有真实人类数据对照，属经济学实验场景，方法可迁移。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:36","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":40,"question":"如何在数据稀缺环境下，利用图检索增强生成（Graph RAG）与LLM结合的方法，更真实地模拟个体交通出行方式选择行为？","design":"提出Preference Chain方法，结合Graph RAG与LLM构建Mobility Agent，基于少量数据构建个体行为偏好图，通过相似性搜索和概率建模引导LLM生成出行选择；在Replica数据集上模拟交通方式选择，与真实选择对比。","baseline":"Replica数据集中真实的交通方式选择行为。","findings":"Preference Chain在模拟交通方式选择上比标准LLM更符合真实世界数据；该方法在数据稀缺地区具有应用潜力，但存在推理速度慢和幻觉风险。","reliability":"论文承认推理速度慢和存在幻觉风险，可能影响行为仿真的可靠性和实用性。","relevance":"该研究用LLM仿真交通出行选择，有真实人类数据对照，属于经济学实验场景，方法可迁移至其他行为仿真，值得阅读原文以评估其仿真偏差与可靠性。","inspiration":"借鉴Preference Chain方法，利用图检索增强生成（Graph RAG）从少量个体行为数据中构建偏好图，并通过相似性搜索和概率建模引导LLM生成选择，以提升仿真真实性。｜该方法可迁移至消费者跨期选择研究，模拟个体在不同时间偏好下的储蓄或消费决策。｜以LLM作为被试，构建基于少量真实个体跨期选择数据的偏好图，施加不同利率或未来收入预期的处理，结果变量为模拟的消费-储蓄分配，与真实家庭金融调查数据（如PSID）进行对照。"}},{"id":"2508.15926","version":1,"title":"Noise, Adaptation, and Strategy: Assessing LLM Fidelity in Decision-Making","zh_title":"噪声、适应与策略：评估LLM在决策中的保真度","abstract":"Large language models (LLMs) are increasingly used in social science simulations. While their performance on reasoning and optimization tasks has been extensively evaluated, less attention has been paid to their ability to simulate human decision-making's variability and adaptability. We propose a process-oriented evaluation framework with progressive interventions (Intrinsicality, Instruction, and Imitation) to examine how LLM agents adapt under different levels of external guidance and human-derived noise. We validate the framework on two classic economics tasks, irrationality in the second-price auction and decision bias in the newsvendor problem, showing behavioral gaps between LLMs and humans. We find that LLMs, by default, converge on stable and conservative strategies that diverge from observed human behaviors. Risk-framed instructions impact LLM behavior predictably but do not replicate human-like diversity. Incorporating human data through in-context learning narrows the gap but fails to reach human subjects' strategic variability. These results highlight a persistent alignment gap in behavioral fidelity and suggest that future LLM evaluations should consider more process-level realism. We present a process-oriented approach for assessing LLMs in dynamic decision-making tasks, offering guidance for their application in synthetic data for social science research.","authors":["Yuanjun Feng","Vivek Choudhary","Yash Raj Shrestha"],"categories":["cs.CE","cs.AI"],"primary_category":"cs.CE","announce_type":"new","date":"2025-08-21","first_seen":"2025-08-21","revised_at":null,"abs_url":"https://arxiv.org/abs/2508.15926","pdf_url":"https://arxiv.org/pdf/2508.15926","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","A3","B1","B2","B4"],"tags":["LLM仿真","行为保真度","经济学实验"],"reason":"直接评估LLM模拟人类决策的保真度，用真实人类数据对照，涉及经济学任务，指出失…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:22","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":48,"question":"LLM在动态决策任务中能否复现人类决策的变异性和适应性？","design":"用GPT-4o、Claude 3.5 Sonnet、Claude 3.7 Sonnet扮演卖家或报童，在二阶密封拍卖和报童问题中，通过无干预、风险框架指令、人类决策历史模仿三种渐进干预，测量其策略稳定性、行为变异性和适应性。","baseline":"对照二阶密封拍卖的真实人类实验数据（Davis et al., 2011, 2023），以及报童问题的已知人类决策偏差模式。","findings":"LLM默认采取稳定保守的策略，与人类行为偏离；风险框架指令可预测地影响LLM行为，但未复现人类多样性；通过上下文学习注入人类数据可缩小差距，但仍未达到人类被试的策略变异性。","reliability":"论文指出当前LLM在动态行为模拟中存在行为保真度对齐差距，缺乏人类决策的随机性和适应性，未来评估需关注过程层面的真实性。","relevance":"该研究直接评估LLM模拟人类决策的保真度，使用真实人类数据作为基准，涉及经济学实验任务，并明确指出了仿真失效的条件和局限，与研究者关注点高度吻合，值得精读原文。","inspiration":"该研究通过渐进式干预（无干预、风险框架指令、人类历史模仿）系统测量LLM行为变异性的方法值得借鉴，可迁移到资产定价实验中的泡沫形成与处置效应研究｜可设计让LLM扮演投资者在实验性资产市场中交易，处理为不同风险提示框架或注入真实人类交易历史，结果变量为价格偏离度、交易量与处置效应系数，对照真实人类实验数据（如Smith et al., 1988的泡沫实验）"}},{"id":"2508.12045","version":2,"title":"Large Language Models Enable Design of Personalized Nudges across Cultures","zh_title":"大语言模型助力跨文化个性化助推设计","abstract":"Nudge strategies are effective tools for influencing behaviour, but their impact depends on individual preferences. Strategies that work for some individuals may be counterproductive for others. We hypothesize that large language models (LLMs) can facilitate the design of individual-specific nudges without the need for costly and time-intensive behavioural data collection and modelling. To test this, we use LLMs to design personalized decoy-based nudges tailored to individual profiles and cultural contexts, aimed at encouraging air travellers to voluntarily offset CO$_2$ emissions from flights. We evaluate their effectiveness through a large-scale survey experiment ($n=3495$) conducted across five countries. Results show that LLM-informed personalized nudges are more effective than uniform settings, raising offsetting rates by 3-7$\\%$ in Germany, Singapore, and the US, though not in China or India. Our study highlights the potential of LLM as a low-cost testbed for piloting nudge strategies. At the same time, cultural heterogeneity constrains their generalizability underscoring the need for combining LLM-based simulations with targeted empirical validation.","authors":["Vladimir Maksimenko","Qingyao Xin","Prateek Gupta","Bin Zhang","Prateek Bansal"],"categories":["cs.CY","cs.AI"],"primary_category":"cs.CY","announce_type":"new","date":"2025-08-16","first_seen":"2025-08-16","revised_at":null,"abs_url":"https://arxiv.org/abs/2508.12045","pdf_url":"https://arxiv.org/pdf/2508.12045","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM仿真","行为助推","跨文化实验"],"reason":"用LLM设计个性化助推并做大规模人类调查对照，直接仿真人类决策行为。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:21","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":25,"question":"大语言模型能否用于设计跨文化的个性化助推策略，以低成本替代人类被试进行助推效果评估？","design":"使用LLM根据个体人口统计特征（性别、年龄、收入、环境关心度、碳抵消信任度）和文化背景，为航空旅客生成个性化的诱饵选项（价格与碳抵消比例），通过大规模在线调查实验（n=3495）在五个国家测试其对自愿碳抵消选择的影响。","baseline":"在五个国家（中国、德国、印度、新加坡、美国）对3495名真实航空旅客进行的调查实验，测量其面对个性化诱饵与统一诱饵时的碳抵消选择率。","findings":"LLM生成的个性化助推在德国、新加坡和美国将碳抵消率提高了3-7%，但在中国和印度未显示显著效果。文化异质性限制了LLM仿真助推效果的普适性，需结合实证验证。","reliability":"论文指出LLM仿真受文化异质性约束，在部分国家（中国、印度）失效，且个性化助推效果依赖于对个体偏好和文化背景的准确建模，需结合针对性实证验证。","relevance":"该研究直接以LLM仿真人类决策行为，并与大规模跨国人类实验对照，评估个性化助推效果，高度契合研究者对LLM人类仿真可靠性及失效条件的关注。","inspiration":"借鉴LLM根据个体特征生成个性化干预并利用跨国调查进行对照验证的方法。｜可迁移至消费者金融产品选择中的个性化信息披露实验，如退休储蓄计划或保险产品选择。｜以LLM为不同人口特征群体生成个性化信息呈现方式，通过在线实验测量选择行为，并以真实市场数据或已有行为实验数据作为基准对照。"}},{"id":"2508.11873","version":1,"title":"SimInterview: Transforming Business Education through Large Language Model-Based Simulated Multilingual Interview Training System","zh_title":"SimInterview：基于大语言模型的模拟多语种面试训练系统助力商业教育转型","abstract":"Business interview preparation demands both solid theoretical grounding and refined soft skills, yet conventional classroom methods rarely deliver the individualized, culturally aware practice employers currently expect. This paper introduces SimInterview, a large language model (LLM)-based simulated multilingual interview training system designed for business professionals entering the AI-transformed labor market. Our system leverages an LLM agent and synthetic AI technologies to create realistic virtual recruiters capable of conducting personalized, real-time conversational interviews. The framework dynamically adapts interview scenarios using retrieval-augmented generation (RAG) to match individual resumes with specific job requirements across multiple languages. Built on LLMs (OpenAI o3, Llama 4 Maverick, Gemma 3), integrated with Whisper speech recognition, GPT-SoVITS voice synthesis, Ditto diffusion-based talking head generation model, and ChromaDB vector databases, our system significantly improves interview readiness across English and Japanese markets. Experiments with university-level candidates show that the system consistently aligns its assessments with job requirements, faithfully preserves resume content, and earns high satisfaction ratings, with the lightweight Gemma 3 model producing the most engaging conversations. Qualitative findings revealed that the standardized Japanese resume format improved document retrieval while diverse English resumes introduced additional variability, and they highlighted how cultural norms shape follow-up questioning strategies. Finally, we also outlined a contestable AI design that can explain, detect bias, and preserve human-in-the-loop to meet emerging regulatory expectations.","authors":["Truong Thanh Hung Nguyen","Tran Diem Quynh Nguyen","Hoang Loc Cao","Thi Cam Thanh Tran","Thi Cam Mai Truong","Hung Cao"],"categories":["cs.CY","cs.AI","cs.HC","cs.MM"],"primary_category":"cs.CY","announce_type":"new","date":"2025-08-16","first_seen":"2025-08-16","revised_at":null,"abs_url":"https://arxiv.org/abs/2508.11873","pdf_url":"https://arxiv.org/pdf/2508.11873","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["面试训练","角色扮演","教育技术"],"reason":"系统是面试训练工具，LLM扮演面试官进行角色扮演对话，无人类行为对照实验或测量…","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:59:26","error":null,"has_summary":false,"summary":null},{"id":"2508.09713","version":1,"title":"Evaluating the Role of Large Language Models in Legal Practice in India","zh_title":"评估大语言模型在印度法律实践中的作用","abstract":"The integration of Artificial Intelligence(AI) into the legal profession raises significant questions about the capacity of Large Language Models(LLM) to perform key legal tasks. In this paper, I empirically evaluate how well LLMs, such as GPT, Claude, and Llama, perform key legal tasks in the Indian context, including issue spotting, legal drafting, advice, research, and reasoning. Through a survey experiment, I compare outputs from LLMs with those of a junior lawyer, with advanced law students rating the work on helpfulness, accuracy, and comprehensiveness. LLMs excel in drafting and issue spotting, often matching or surpassing human work. However, they struggle with specialised legal research, frequently generating hallucinations, factually incorrect or fabricated outputs. I conclude that while LLMs can augment certain legal tasks, human expertise remains essential for nuanced reasoning and the precise application of law.","authors":["Rahul Hemrajani"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2025-08-13","first_seen":"2025-08-13","revised_at":null,"abs_url":"https://arxiv.org/abs/2508.09713","pdf_url":"https://arxiv.org/pdf/2508.09713","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM法律任务","人类对照","替代劳动"],"reason":"LLM替代初级律师完成法律任务，由学生评分，属于替代人类劳动而非仿真被试，但有…","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:59:34","error":null,"has_summary":false,"summary":null},{"id":"2508.08486","version":1,"title":"Beyond Ordinal Preferences: Why Alignment Needs Cardinal Human Feedback","zh_title":"超越序数偏好：为何对齐需要基数人类反馈","abstract":"Alignment techniques for LLMs rely on optimizing preference-based objectives -- where these preferences are typically elicited as ordinal, binary choices between responses. Recent work has focused on improving label quality or mitigating particular biases, but we identify a more fundamental limitation: these methods collect the wrong kind of data. We prove an impossibility result: no algorithm relying solely on ordinal comparisons can systematically recover the most preferred model. Intuitively, ordinal data lacks the information needed to resolve tradeoffs -- e.g., fixing a factual error on one prompt versus improving style on another. We show that selecting the optimal model requires recovering preferences over \\emph{models} (rather than just responses), which can only be identified given cardinal feedback about response quality. To address this, we collect and publicly release a dataset of 25,000 cardinal judgments using willingness-to-pay elicitations, a well-established tool from experimental economics. Empirically, we find that incorporating cardinal feedback into preference fine-tuning allows models to prioritize high-impact improvements and outperform ordinal-only methods on downstream benchmarks, such as Arena-Hard.","authors":["Parker Whitfill","Stewy Slocum"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2025-08-11","first_seen":"2025-08-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2508.08486","pdf_url":"https://arxiv.org/pdf/2508.08486","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["LLM对齐","偏好学习","基数反馈"],"reason":"研究LLM对齐中的偏好反馈类型，不涉及用LLM仿真人类被试或与人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:59:21","error":null,"has_summary":false,"summary":null},{"id":"2508.06635","version":2,"title":"Valid Inference with Imperfect Synthetic Data","zh_title":"不完美合成数据的有效推断","abstract":"Predictions and generations from large language models are increasingly being explored as an aid in limited data regimes, such as in computational social science and human subjects research. While prior technical work has mainly explored the potential to use model-predicted labels for unlabeled data in a principled manner, there is increasing interest in using large language models to generate entirely new synthetic samples (e.g., synthetic simulations), such as in responses to surveys. However, it remains unclear by what means practitioners can combine such data with real data and yet produce statistically valid conclusions upon them. In this paper, we introduce a new estimator based on generalized method of moments, providing a hyperparameter-free solution with strong theoretical guarantees to address this challenge. Intriguingly, we find that interactions between the moment residuals of synthetic data and those of real data (i.e., when they are predictive of each other) can greatly improve estimates of the target parameter. We validate the finite-sample performance of our estimator across different tasks in computational social science applications, demonstrating large empirical gains.","authors":["Yewon Byun","Shantanu Gupta","Zachary C. Lipton","Rachel Leah Childers","Bryan Wilder"],"categories":["cs.LG","cs.AI","stat.ML"],"primary_category":"cs.LG","announce_type":"new","date":"2025-08-08","first_seen":"2025-08-08","revised_at":null,"abs_url":"https://arxiv.org/abs/2508.06635","pdf_url":"https://arxiv.org/pdf/2508.06635","source_feed":"backfill","score":8,"bucket":"selected","rubric_hits":["A1","A5","B1","B2"],"tags":["合成数据","统计推断","计算社会科学"],"reason":"用LLM生成合成调查样本，结合真实数据做统计推断，直接涉及人类仿真与对照。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:21","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":45,"question":"如何将LLM生成的合成数据与真实数据结合，进行统计上有效的推断？","design":"本文提出一种基于广义矩估计（GMM）的框架，将LLM生成的合成样本（如模拟调查回答）与真实样本结合，通过引入合成数据与真实数据之间的相关结构来提高估计精度。","baseline":"真实人类标注样本，包含已标注协变量和结果的小规模数据集。","findings":"当合成数据的矩残差能预测真实数据的矩残差时，结合合成数据可提高估计精度并缩小置信区间；即使合成数据完全无信息，也不会损害渐近有效性。","reliability":"论文指出，若生成模型与真实分布不匹配，直接简单聚合合成数据会导致严重偏差；所提方法依赖于合成样本与真实样本之间的相关结构，若两者独立则无增益但也不损失。","relevance":"该研究直接针对用LLM生成合成调查样本并做统计推断的场景，提供了结合真实数据与合成数据的严谨方法，对关注人类仿真可靠性的研究者具有重要参考价值。","inspiration":"该方法通过广义矩估计将LLM合成样本与真实样本结合，利用合成数据矩残差对真实数据矩残差的预测能力来提高估计精度，同时保证即使合成数据无信息也不损害推断有效性。｜可迁移到政策评估中的调查实验，例如评估税收优惠宣传对家庭消费意愿的影响，其中LLM生成模拟调查回答作为合成数据。｜以LLM模拟的家庭为被试，处理为展示税收优惠信息，结果变量为自报消费意愿，用真实家庭调查数据作为基准，通过GMM框架结合合成与真实样本估计处理效应。"}},{"id":"2508.02766","version":2,"title":"The Generative Reasonable Person","zh_title":"生成式理性人","abstract":"This Article introduces the generative reasonable person, a new tool for estimating how ordinary people judge reasonableness. As claims about AI capabilities often outpace evidence, the Article proceeds empirically: adapting randomized controlled trials to large language models, it replicates three published studies of lay judgment across negligence, consent, and contract interpretation, drawing on nearly 10,000 simulated decisions. The findings reveal that models can replicate subtle patterns that run counter to textbook treatment. Like human subjects, models prioritize social conformity over cost-benefit analysis when assessing negligence, inverting the hierarchy that textbooks teach. They reproduce the paradox that material lies erode consent less than lies about a transaction's essence. And they track lay contract formalism, judging hidden fees more enforceable than fair. For two centuries, scholars have debated whether the reasonable person is empirical or normative, majoritarian or aspirational. But much of this debate assumed a constraint that no longer holds: that lay judgments are expensive to surface, slow to collect, and unavailable at scale. Generative reasonable people loosen that constraint. They offer judges empirical checks on elite intuition, give resource-constrained litigants access to simulated jury feedback, and let regulators pilot-test public comprehension, all at a fraction of survey costs. The reasonable person standard has long functioned as a vessel for judicial intuition precisely because the empirical baseline was missing. With that baseline now available, departures from lay understanding become transparent rather than hidden, a choice to be justified, not a fact to be assumed. Properly cabined, the generative reasonable person may become a dictionary for reasonableness judgments.","authors":["Yonathan A. Arbel"],"categories":["cs.CY","cs.AI"],"primary_category":"cs.CY","announce_type":"new","date":"2025-08-04","first_seen":"2025-08-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2508.02766","pdf_url":"https://arxiv.org/pdf/2508.02766","source_feed":"backfill","score":10,"bucket":"selected","rubric_hits":["A1","A2","A5","B1","B2","B3"],"tags":["LLM仿真","人类被试替代","法律判断"],"reason":"用LLM模拟普通人判断，复现三项实验，有真实人类数据对照，涉及法律判断与政策评…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:21","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":3,"question":"如何利用大语言模型模拟普通人判断，以复现法律中“理性人”标准的实证研究？","design":"使用大语言模型扮演普通人群，采用随机对照试验设计，复现三项已发表的关于过失、同意和合同解释的普通人判断研究，测量模型在近10,000次模拟决策中的判断模式。","baseline":"三项已发表研究中的真实人类被试数据，涵盖过失判断、同意判断和合同解释判断。","findings":"模型能复现与教科书相悖的微妙模式：在过失评估中优先考虑社会从众而非成本收益分析；再现了实质性谎言比交易本质谎言更少侵蚀同意的悖论；在合同解释中表现出普通人合同形式主义，认为隐藏费用比公平条款更具可执行性。","reliability":"论文未讨论","relevance":"该研究直接使用LLM模拟人类被试，复现法律判断实验并与真实人类数据对照，评估仿真可靠性，完全契合研究者对LLM人类仿真实验的关注，值得精读原文。","inspiration":"借鉴其用LLM复现已有人类实验并系统对比结果的方法，可验证LLM在特定领域的行为一致性。｜可迁移到消费者金融决策实验，如信贷条款理解、费用披露效果评估。｜以LLM模拟消费者，随机呈现不同披露格式的信贷合同，测量其理解度和选择行为，以真实消费者调查数据为基准进行对照验证。"}},{"id":"2508.05670","version":1,"title":"Can LLMs effectively provide game-theoretic-based scenarios for cybersecurity?","zh_title":"大语言模型能否有效提供基于博弈论的网络安全场景？","abstract":"Game theory has long served as a foundational tool in cybersecurity to test, predict, and design strategic interactions between attackers and defenders. The recent advent of Large Language Models (LLMs) offers new tools and challenges for the security of computer systems; In this work, we investigate whether classical game-theoretic frameworks can effectively capture the behaviours of LLM-driven actors and bots. Using a reproducible framework for game-theoretic LLM agents, we investigate two canonical scenarios -- the one-shot zero-sum game and the dynamic Prisoner's Dilemma -- and we test whether LLMs converge to expected outcomes or exhibit deviations due to embedded biases. Our experiments involve four state-of-the-art LLMs and span five natural languages, English, French, Arabic, Vietnamese, and Mandarin Chinese, to assess linguistic sensitivity. For both games, we observe that the final payoffs are influenced by agents characteristics such as personality traits or knowledge of repeated rounds. Moreover, we uncover an unexpected sensitivity of the final payoffs to the choice of languages, which should warn against indiscriminate application of LLMs in cybersecurity applications and call for in-depth studies, as LLMs may behave differently when deployed in different countries. We also employ quantitative metrics to evaluate the internal consistency and cross-language stability of LLM agents, to help guide the selection of the most stable LLMs and optimising models for secure applications.","authors":["Daniele Proverbio","Alessio Buscemi","Alessandro Di Stefano","The Anh Han","German Castignani","Pietro Liò"],"categories":["cs.CR","cs.AI","cs.CY","cs.GT"],"primary_category":"cs.CR","announce_type":"new","date":"2025-08-04","first_seen":"2025-08-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2508.05670","pdf_url":"https://arxiv.org/pdf/2508.05670","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["LLM博弈行为","网络安全","社会模拟"],"reason":"用LLM agent模拟博弈行为，但无真实人类数据对照，属社会模拟边界情形。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:35","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":197,"question":"LLM在网络安全相关的博弈论场景中是否遵循经典博弈论预测，其行为受哪些因素影响？","design":"使用FAIRGAME框架，让四个主流LLM扮演博弈参与者，在一次性零和博弈和动态囚徒困境两种场景下进行交互，测试不同人格提示、是否知晓重复轮次等条件，并比较五种语言下的表现，测量最终收益和内部一致性。","baseline":"无对照","findings":"LLM的最终收益受人格特质和是否知晓重复轮次等智能体特征影响；收益对语言选择存在意外敏感性，不同语言下行为不同，提示在网络安全应用中需谨慎使用LLM。","reliability":"论文指出LLM行为偏离理论预测，且对语言敏感，在不同国家部署时可能表现不同，但未系统讨论失效的具体条件或局限。","relevance":"该研究用LLM模拟博弈行为，但无真实人类数据对照，属于社会模拟边界情形，与研究者关注的有基准对照的仿真研究不完全匹配，但提供了LLM行为偏差和语言敏感性的批判性证据，值得快速浏览。","inspiration":"该方法通过系统操纵人格提示和是否知晓重复轮次等智能体特征，并测量最终收益与内部一致性，可用于检验LLM行为偏差｜可迁移到政策公告的预期形成实验，研究不同信息框架下LLM模拟的投资者预期更新｜设计：以LLM为被试，处理为政策公告的表述语气（乐观/悲观）和是否明确政策调整频率，结果变量为预测的资产价格变动，以历史政策公告后的真实市场预期调查数据为对照"}},{"id":"2508.00485","version":2,"title":"A Frame for Communication Control","zh_title":"一种通信控制框架","abstract":"We are experiencing the rise of ChatGPT-like systems or LLMs in political turbulent times. We assume the need to regulate their use because of their bubble-shaping and polarizing potential. To regulate, we need a language that allows interests and compromises to be discussed. In this context, we can think of such a shared language as a jargon, a specialized vocabulary for law-making. To the extent that such a jargon exists, it is now being corrupted by LLMs. This situation appears paradoxical. The issue includes persistent communication failures, between disciplines that cannot translate their technical vocabulary into accessible terms, and between political movements that operate in incompatible worldviews. We show that a frame integrating four specialist languages, those of governance, economy, community and science, is able to address these failures case-wise, which we consider helpful. However, for reasons noted, we cannot create the more generic jargon needed on our own. We conclude that our frame provides the knowledge to design and apply RAG-LLM architectures for researching their jargon generating potential in a future project. We show its feasibility in the appendix.","authors":["Aernout Schmidt","Kunbei Zhang"],"categories":["cs.CY","econ.TH"],"primary_category":"cs.CY","announce_type":"new","date":"2025-08-01","first_seen":"2025-08-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2508.00485","pdf_url":"https://arxiv.org/pdf/2508.00485","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["LLM监管","通信框架","法律术语"],"reason":"论文讨论用RAG-LLM生成法律术语，属多智能体协作框架，无人类行为仿真或对照。","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:59:03","error":null,"has_summary":false,"summary":null},{"id":"2507.22049","version":1,"title":"Validating Generative Agent-Based Models of Social Norm Enforcement: From Replication to Novel Predictions","zh_title":"验证基于生成式智能体的社会规范执行模型：从复现到新预测","abstract":"As large language models (LLMs) advance, there is growing interest in using them to simulate human social behavior through generative agent-based modeling (GABM). However, validating these models remains a key challenge. We present a systematic two-stage validation approach using social dilemma paradigms from psychological literature, first identifying the cognitive components necessary for LLM agents to reproduce known human behaviors in mixed-motive settings from two landmark papers, then using the validated architecture to simulate novel conditions. Our model comparison of different cognitive architectures shows that both persona-based individual differences and theory of mind capabilities are essential for replicating third-party punishment (TPP) as a costly signal of trustworthiness. For the second study on public goods games, this architecture is able to replicate an increase in cooperation from the spread of reputational information through gossip. However, an additional strategic component is necessary to replicate the additional boost in cooperation rates in the condition that allows both ostracism and gossip. We then test novel predictions for each paper with our validated generative agents. We find that TPP rates significantly drop in settings where punishment is anonymous, yet a substantial amount of TPP persists, suggesting that both reputational and intrinsic moral motivations play a role in this behavior. For the second paper, we introduce a novel intervention and see that open discussion periods before rounds of the public goods game further increase contributions, allowing groups to develop social norms for cooperation. This work provides a framework for validating generative agent models while demonstrating their potential to generate novel and testable insights into human social behavior.","authors":["Logan Cross","Nick Haber","Daniel L. K. Yamins"],"categories":["cs.MA"],"primary_category":"cs.MA","announce_type":"new","date":"2025-07-29","first_seen":"2025-07-29","revised_at":null,"abs_url":"https://arxiv.org/abs/2507.22049","pdf_url":"https://arxiv.org/pdf/2507.22049","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A3","A5","B1","B2","B4"],"tags":["LLM仿真","社会规范","行为博弈"],"reason":"用LLM智能体复现社会困境实验，与真实人类数据对照，验证仿真并预测新条件，直接…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:19","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":28,"question":"如何通过生成式智能体建模（GABM）复现并扩展社会困境中的人类行为，并验证其认知架构的有效性？","design":"使用LLM驱动的生成式智能体，通过组合记忆、人格、推理等认知组件构建不同架构，复现Jordan et al. (2016)的信任博弈第三方惩罚实验和Feinberg et al. (2014)的公共品博弈中流言与排斥实验，比较智能体行为与人类数据的统计效应，并利用验证后的架构模拟匿名惩罚和公开讨论等新条件。","baseline":"Jordan et al. (2016)的第三方惩罚信任博弈人类实验数据，以及Feinberg et al. (2014)的公共品博弈中流言与排斥对人类合作影响的数据。","findings":"复现第三方惩罚效应需要人格提示和心理理论组件，而公共品博弈中流言提升合作可被复现，但排斥与流言结合的额外合作提升需额外策略组件；匿名惩罚下第三方惩罚率显著下降但仍存在，表明声誉和内在道德动机共同驱动惩罚，公开讨论能进一步提升公共品贡献。","reliability":"论文未讨论","relevance":"该研究直接用LLM智能体复现经典社会困境实验，并与真实人类数据对照，验证仿真可靠性并生成新预测，完全契合研究者对LLM人类仿真、经济学实验复现及失效条件探索的兴趣，值得精读原文。","inspiration":"该方法通过组合记忆、人格、推理等认知组件构建不同LLM智能体架构，并与人类实验数据对照来验证仿真有效性，为经济金融实验提供了可借鉴的仿真验证框架。｜可迁移至资产定价实验中的羊群效应研究，利用LLM智能体模拟投资者在信息不对称下的决策行为。｜以LLM智能体为被试，处理为是否提供历史价格信息，结果变量为投资决策的羊群效应指数，对照真实人类资产定价实验数据。"}},{"id":"2507.21432","version":2,"title":"Towards Locally Deployable Fine-Tuned Causal Large Language Models for Mode Choice Behaviour","zh_title":"面向出行方式选择行为的本地可部署微调因果大语言模型研究","abstract":"This study investigates the adoption of open-access, locally deployable causal large language models (LLMs) for travel mode choice prediction and introduces LiTransMC, the first fine-tuned causal LLM developed for this task. We systematically benchmark eleven open-access LLMs (1-12B parameters) across three stated and revealed preference datasets, testing 396 configurations and generating over 79,000 mode choice decisions. Beyond predictive accuracy, we evaluate models generated reasoning using BERTopic for topic modelling and a novel Explanation Strength Index, providing the first structured analysis of how LLMs articulate decision factors in alignment with behavioural theory. LiTransMC, fine-tuned using parameter efficient and loss masking strategy, achieved a weighted F1 score of 0.6845 and a Jensen-Shannon Divergence of 0.000245, surpassing both untuned local models and larger proprietary systems, including GPT-4o with advanced persona inference and embedding-based loading, while also outperforming classical mode choice methods such as discrete choice models and machine learning classifiers for the same dataset. This dual improvement, i.e., high instant-level accuracy and near-perfect distributional calibration, demonstrates the feasibility of creating specialist, locally deployable LLMs that integrate prediction and interpretability. Through combining structured behavioural prediction with natural language reasoning, this work unlocks the potential for conversational, multi-task transport models capable of supporting agent-based simulations, policy testing, and behavioural insight generation. These findings establish a pathway for transforming general purpose LLMs into specialized and explainable tools for transportation research and policy formulation, while maintaining privacy, reducing cost, and broadening access through local deployment.","authors":["Tareq Alsaleh","Bilal Farooq"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2025-07-29","first_seen":"2025-07-29","revised_at":null,"abs_url":"https://arxiv.org/abs/2507.21432","pdf_url":"https://arxiv.org/pdf/2507.21432","source_feed":"backfill","score":7,"bucket":"pending","rubric_hits":["A1","B1","B2"],"tags":["LLM行为预测","交通方式选择","人类数据对照"],"reason":"用LLM预测出行方式选择，有真实人类数据对照，涉及交通行为仿真，方法可迁移至人…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:34","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":112,"question":"如何利用开源、可本地部署的因果大语言模型进行出行方式选择预测，并生成符合行为理论的决策解释？","design":"本研究不是人类仿真实验，而是使用11个开源LLM（1-12B参数）在三个陈述偏好和显示偏好数据集上预测出行方式选择，测试了396种配置，生成超过79,000个选择决策；并微调了LiTransMC模型，采用参数高效和损失掩码策略。","baseline":"三个陈述偏好和显示偏好数据集，包含真实人类的出行方式选择记录。","findings":"微调后的LiTransMC在加权F1分数（0.6845）和Jensen-Shannon散度（0.000245）上超越了未调优的本地模型和GPT-4o等大型闭源模型，同时优于传统离散选择模型和机器学习分类器；模型不仅能准确预测个体选择，还能生成与行为理论一致的自然语言推理。","reliability":"论文未讨论","relevance":"该研究用LLM预测出行方式选择，有真实人类数据对照，涉及交通行为仿真，方法可迁移至人类决策仿真，值得阅读原文以评估其作为人类被试替代品的潜力与局限。","inspiration":"该研究采用参数高效微调与损失掩码策略，在有限标注数据下提升开源LLM对个体选择行为的预测精度，并利用Jensen-Shannon散度等分布相似性指标评估模型输出与真实人类选择的整体拟合，这一设计可借鉴用于校准LLM仿真中的行为偏差。｜可迁移至消费者跨期选择实验，例如研究即时奖励与延迟奖励的权衡，或政策干预对储蓄行为的影响。｜以微调后的开源LLM作为虚拟被试，处理为不同利率或未来奖励的表述框架，结果变量为选择即时或延迟选项的概率，以真实实验室跨期选择数据作为基准，比较LLM与人类的选择分布及时间贴现率。"}},{"id":"2507.17564","version":1,"title":"Decoding Consumer Preferences Using Attention-Based Language Models","zh_title":"使用基于注意力的语言模型解码消费者偏好","abstract":"This paper proposes a new demand estimation method using attention-based language models. An encoder-only language model is trained in a two-stage process to analyze the natural language descriptions of used cars from a large US-based online auction marketplace. The approach enables semi-nonparametrically estimation for the demand primitives of a structural model representing the private valuations and market size for each vehicle listing. In the first stage, the language model is fine-tuned to encode the target auction outcomes using the natural language vehicle descriptions. In the second stage, the trained language model's encodings are projected into the parameter space of the structural model. The model's capability to conduct counterfactual analyses within the trained market space is validated using a subsample of withheld auction data, which includes a set of unique \"zero shot\" instances.","authors":["Joshua Foster","Fredrik Odegaard"],"categories":["econ.EM"],"primary_category":"econ.EM","announce_type":"new","date":"2025-07-23","first_seen":"2025-07-23","revised_at":null,"abs_url":"https://arxiv.org/abs/2507.17564","pdf_url":"https://arxiv.org/pdf/2507.17564","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["需求估计","语言模型","结构模型"],"reason":"论文用语言模型估计需求，不涉及用LLM仿真人类被试或行为对照，属于纯NLP应用。","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:58:56","error":null,"has_summary":false,"summary":null},{"id":"2507.17024","version":1,"title":"Write, Rank, or Rate: Comparing Methods for Studying Visualization Affordances","zh_title":"写、排或评：比较研究可视化可供性的方法","abstract":"A growing body of work on visualization affordances highlights how specific design choices shape reader takeaways from information visualizations. However, mapping the relationship between design choices and reader conclusions often requires labor-intensive crowdsourced studies, generating large corpora of free-response text for analysis. To address this challenge, we explored alternative scalable research methodologies to assess chart affordances. We test four elicitation methods from human-subject studies: free response, visualization ranking, conclusion ranking, and salience rating, and compare their effectiveness in eliciting reader interpretations of line charts, dot plots, and heatmaps. Overall, we find that while no method fully replicates affordances observed in free-response conclusions, combinations of ranking and rating methods can serve as an effective proxy at a broad scale. The two ranking methodologies were influenced by participant bias towards certain chart types and the comparison of suggested conclusions. Rating conclusion salience could not capture the specific variations between chart types observed in the other methods. To supplement this work, we present a case study with GPT-4o, exploring the use of large language models (LLMs) to elicit human-like chart interpretations. This aligns with recent academic interest in leveraging LLMs as proxies for human participants to improve data collection and analysis efficiency. GPT-4o performed best as a human proxy for the salience rating methodology but suffered from severe constraints in other areas. Overall, the discrepancies in affordances we found between various elicitation methodologies, including GPT-4o, highlight the importance of intentionally selecting and combining methods and evaluating trade-offs.","authors":["Chase Stokes","Kylie Lin","Cindy Xiong Bearfield"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2025-07-22","first_seen":"2025-07-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2507.17024","pdf_url":"https://arxiv.org/pdf/2507.17024","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM代理","可视化可供性","方法比较"],"reason":"用GPT-4o替代人类被试评估图表解读，但主要目的是替代标注劳动，非严格仿真人…","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:58:13","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":50,"question":"在可视化可供性研究中，不同启发方法（自由回答、排序、评分）以及大语言模型（GPT-4o）在多大程度上能复现人类对图表的解读？","design":"本研究并非以LLM仿真为核心，而是比较四种人类被试启发方法（自由回答、可视化排序、结论排序、显著性评分）在获取图表解读上的效果；随后以GPT-4o作为案例，探索用LLM模拟人类被试进行图表解读，并评估其在不同方法上的表现。","baseline":"以62名Prolific众包参与者对点图、折线图和热力图的自由回答结论作为人类基准，并基于因子分析构建了五类可供性空间。","findings":"没有任何单一方法能完全复现自由回答所观察到的可供性，但排序与评分方法的组合可作为大规模研究的有效替代。GPT-4o在显著性评分任务上最接近人类表现，但在其他方法上存在严重局限。","reliability":"论文指出GPT-4o仅在显著性评分方法上表现较好，在其他启发方法上存在严重限制；不同方法（包括GPT-4o）所揭示的可供性差异显著，需谨慎选择和组合方法并评估权衡。","relevance":"该研究直接使用GPT-4o作为人类被试替代品，并与真实人类数据进行对照，评估其仿真可靠性，且指出了失效条件，高度契合研究者对LLM仿真实验的批判性关注，值得阅读原文。","inspiration":"该研究系统比较了自由回答、排序、评分等不同启发方法在获取人类图表解读上的差异，并引入LLM作为被试与真实人类数据对照，这种多方法比较与仿真可靠性评估的设计值得借鉴。｜可迁移到经济金融中的信息处理与决策研究，例如投资者对财报图表或宏观经济数据可视化的解读如何影响预期形成。｜以专业投资者或众包被试为人类基准，让LLM与人类分别对同一组财务图表进行自由回答、排序和评分，比较其解读模式与分布，结果变量为解读结论的分类与评分，以真实人类数据作为对照基准评估LLM仿真的可靠性。"}},{"id":"2507.10933","version":1,"title":"Artificial Finance: How AI Thinks About Money","zh_title":"人工金融：AI如何思考金钱","abstract":"In this paper, we explore how large language models (LLMs) approach financial decision-making by systematically comparing their responses to those of human participants across the globe. We posed a set of commonly used financial decision-making questions to seven leading LLMs, including five models from the GPT series(GPT-4o, GPT-4.5, o1, o3-mini), Gemini 2.0 Flash, and DeepSeek R1. We then compared their outputs to human responses drawn from a dataset covering 53 nations. Our analysis reveals three main results. First, LLMs generally exhibit a risk-neutral decision-making pattern, favoring choices aligned with expected value calculations when faced with lottery-type questions. Second, when evaluating trade-offs between present and future, LLMs occasionally produce responses that appear inconsistent with normative reasoning. Third, when we examine cross-national similarities, we find that the LLMs' aggregate responses most closely resemble those of participants from Tanzania. These findings contribute to the understanding of how LLMs emulate human-like decision behaviors and highlight potential cultural and training influences embedded within their outputs.","authors":["Orhan Erdem","Ragavi Pobbathi Ashok"],"categories":["econ.GN","cs.AI"],"primary_category":"econ.GN","announce_type":"new","date":"2025-07-15","first_seen":"2025-07-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2507.10933","pdf_url":"https://arxiv.org/pdf/2507.10933","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2"],"tags":["LLM仿真","金融决策","跨文化对照"],"reason":"用LLM复现人类金融决策并与53国真实数据对照，评估仿真行为与偏差，直接命中核…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:18","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":26,"question":"LLM在金融决策中表现出怎样的风险偏好、跨期选择模式，其总体回答与哪个国家的人类被试最相似？","design":"将7个主流LLM（GPT-4o、GPT-4.5、o1、o3-mini、Gemini 2.0 Flash、DeepSeek R1）作为被试，向其呈现一组常用的金融决策问题（包括彩票型问题和跨期选择问题），收集模型回答，并与53国人类调查数据进行比较。","baseline":"来自Wang et al. (2017)的覆盖53个国家的全球金融决策调查数据。","findings":"LLM普遍表现出风险中性，倾向于根据期望值计算选择彩票型问题；在跨期选择中偶尔出现与规范推理不一致的回答；LLM的总体回答模式与坦桑尼亚参与者最相似。","reliability":"论文未讨论","relevance":"该研究直接用LLM复现人类金融决策并与跨国真实数据对照，评估仿真行为与偏差，高度契合研究者对LLM人类仿真实验的关注，值得阅读原文。","inspiration":"借鉴其将LLM作为被试、使用标准化金融决策问题并直接与跨国人类调查数据对照的设计思路。｜可迁移到消费者跨期选择、风险资产配置或政策公告预期形成等行为金融场景。｜以LLM为被试，施加不同框架的跨期选择或风险决策问题，测量其时间偏好与风险厌恶参数，并与真实跨国调查数据（如Falk et al., 2018的全球偏好数据）进行对照，检验LLM仿真的人群代表性。"}},{"id":"2507.10342","version":1,"title":"Using AI to replicate human experimental results: a motion study","zh_title":"使用AI复现人类实验结果：一项运动研究","abstract":"This paper explores the potential of large language models (LLMs) as reliable analytical tools in linguistic research, focusing on the emergence of affective meanings in temporal expressions involving manner-of-motion verbs. While LLMs like GPT-4 have shown promise across a range of tasks, their ability to replicate nuanced human judgements remains under scrutiny. We conducted four psycholinguistic studies (on emergent meanings, valence shifts, verb choice in emotional contexts, and sentence-emoji associations) first with human participants and then replicated the same tasks using an LLM. Results across all studies show a striking convergence between human and AI responses, with statistical analyses (e.g., Spearman's rho = .73-.96) indicating strong correlations in both rating patterns and categorical choices. While minor divergences were observed in some cases, these did not alter the overall interpretative outcomes. These findings offer compelling evidence that LLMs can augment traditional human-based experimentation, enabling broader-scale studies without compromising interpretative validity. This convergence not only strengthens the empirical foundation of prior human-based findings but also opens possibilities for hypothesis generation and data expansion through AI. Ultimately, our study supports the use of LLMs as credible and informative collaborators in linguistic inquiry.","authors":["Rosa Illan Castillo","Javier Valenzuela"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2025-07-14","first_seen":"2025-07-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2507.10342","pdf_url":"https://arxiv.org/pdf/2507.10342","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1"],"tags":["LLM仿真","人类数据对照","心理语言学"],"reason":"用LLM复现人类心理语言学实验，并与真实人类数据对照，评估仿真可靠性。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:16","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":15,"question":"大语言模型能否在心理语言学实验中复现人类对运动动词情感意义的细微判断，从而作为可靠的分析工具？","design":"使用ChatGPT o1模拟人类被试，对包含运动动词的时间表达句进行四项心理语言学任务（情感意义评分、效价变化、情绪语境下的动词选择、句子与表情符号关联），测量其评分和选择模式。","baseline":"59名英语母语者在相同任务上的真实行为数据，包括评分分布和分类选择。","findings":"LLM与人类反应高度一致，Spearman相关系数在0.73至0.96之间，表明评分模式和分类选择均强相关。微小分歧未改变整体解释性结论，支持LLM可作为人类实验的补充工具。","reliability":"论文未讨论","relevance":"该研究直接以真实人类数据为基准，验证LLM复现心理语言学实验的可靠性，符合研究者对仿真有效性评估的关注，值得阅读以了解具体对照方法和统计检验。","inspiration":"借鉴其将同一任务原样施加于LLM并与人类数据直接对比的设计，可迁移到消费者跨期选择或政策公告的预期形成实验，例如用LLM模拟消费者对“时间飞逝”类表述的耐心程度评分，以真实调查数据为基准检验LLM能否复现时间偏好。"}},{"id":"2507.09657","version":1,"title":"Negotiating Comfort: Simulating Personality-Driven LLM Agents in Shared Residential Social Networks","zh_title":"协商舒适度：在共享住宅社交网络中模拟个性驱动的LLM智能体","abstract":"We use generative agents powered by large language models (LLMs) to simulate a social network in a shared residential building, driving the temperature decisions for a central heating system. Agents, divided into Family Members and Representatives, consider personal preferences, personal traits, connections, and weather conditions. Daily simulations involve family-level consensus followed by building-wide decisions among representatives. We tested three personality traits distributions (positive, mixed, and negative) and found that positive traits correlate with higher happiness and stronger friendships. Temperature preferences, assertiveness, and selflessness have a significant impact on happiness and decisions. This work demonstrates how LLM-driven agents can help simulate nuanced human behavior where complex real-life human simulations are difficult to set.","authors":["Ann Nedime Nese Rende","Tolga Yilmaz","Özgür Ulusoy"],"categories":["cs.SI","cs.MA"],"primary_category":"cs.SI","announce_type":"new","date":"2025-07-13","first_seen":"2025-07-13","revised_at":null,"abs_url":"https://arxiv.org/abs/2507.09657","pdf_url":"https://arxiv.org/pdf/2507.09657","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["LLM智能体","社会模拟","个性驱动"],"reason":"用LLM agent模拟社会网络与决策，但无真实人类数据对照，属纯理论演示。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:16","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":143,"question":"在共享住宅的社会网络中，不同人格特质分布如何影响由LLM智能体驱动的供暖温度决策、幸福感与友谊关系？","design":"使用LLM驱动的生成式智能体模拟一栋共享住宅楼内的社会网络，智能体分为家庭成员和家庭代表，每日先进行家庭内部协商再在楼宇层面投票决定中央供暖温度；通过设置全正面、全负面、50%正面三种人格特质分布作为处理条件，测量智能体的幸福感、温度选择、友谊紧密度和决策结果。","baseline":"无对照","findings":"正面人格特质分布与更高的幸福感和更强的友谊相关；温度偏好、自信和无私程度对幸福感和决策有显著影响。","reliability":"论文未讨论","relevance":"该研究用LLM智能体模拟社会网络中的协商与决策，属于人类仿真实验范畴，但缺乏真实人类数据对照，无法评估仿真的可靠性与偏差，与研究者关注的基准对照和批判性评估需求匹配度较低，不建议优先阅读原文。","inspiration":"该研究通过设定不同人格特质分布作为处理条件，测量智能体在协商网络中的幸福感与决策结果，这种基于智能体异质性施加处理的设计思路可借鉴｜可迁移至经济金融中的群体决策实验，如家庭内部资源分配或社区公共品供给中的偏好协商｜设计雏形：以LLM智能体模拟家庭成员，处理为不同风险偏好或时间贴现率的分布，结果变量为家庭储蓄率或消费组合，对照真实家庭调查数据如CFPS"}},{"id":"2507.08584","version":1,"title":"To Trade or Not to Trade: An Agentic Approach to Estimating Market Risk Improves Trading Decisions","zh_title":"交易与否：一种基于智能体的市场风险估计方法改善交易决策","abstract":"Large language models (LLMs) are increasingly deployed in agentic frameworks, in which prompts trigger complex tool-based analysis in pursuit of a goal. While these frameworks have shown promise across multiple domains including in finance, they typically lack a principled model-building step, relying instead on sentiment- or trend-based analysis. We address this gap by developing an agentic system that uses LLMs to iteratively discover stochastic differential equations for financial time series. These models generate risk metrics which inform daily trading decisions. We evaluate our system in both traditional backtests and using a market simulator, which introduces synthetic but causally plausible price paths and news events. We find that model-informed trading strategies outperform standard LLM-based agents, improving Sharpe ratios across multiple equities. Our results show that combining LLMs with agentic model discovery enhances market risk estimation and enables more profitable trading decisions.","authors":["Dimitrios Emmanoulopoulos","Ollie Olby","Justin Lyon","Namid R. Stillman"],"categories":["q-fin.ST","cs.AI","cs.CE","cs.MA","q-fin.CP"],"primary_category":"q-fin.ST","announce_type":"new","date":"2025-07-11","first_seen":"2025-07-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2507.08584","pdf_url":"https://arxiv.org/pdf/2507.08584","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["智能体系统","金融交易","随机微分方程"],"reason":"多智能体系统用于金融交易决策，无人类行为对照，属纯智能体协作。","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:59:47","error":null,"has_summary":false,"summary":null},{"id":"2507.07188","version":3,"title":"Prompt Perturbations Reveal Human-Like Biases in Large Language Model Survey Responses","zh_title":"提示扰动揭示大语言模型调查响应中类人偏差","abstract":"Large Language Models (LLMs) are increasingly used as proxies for human subjects in social science surveys, but their reliability and susceptibility to known human-like response biases, such as central tendency, opinion floating and primacy bias are poorly understood. This work investigates the response robustness of LLMs in normative survey contexts, we test nine LLMs on questions from the World Values Survey (WVS), applying a comprehensive set of ten perturbations to both question phrasing and answer option structure, resulting in over 167,000 simulated survey interviews. In doing so, we not only reveal LLMs' vulnerabilities to perturbations but also show that all tested models exhibit a consistent recency bias, disproportionately favoring the last-presented answer option. While larger models are generally more robust, all models remain sensitive to semantic variations like paraphrasing and to combined perturbations. This underscores the critical importance of prompt design and robustness testing when using LLMs to generate synthetic survey data.","authors":["Jens Rupprecht","Georg Ahnert","Markus Strohmaier"],"categories":["cs.CL","cs.AI","cs.CY"],"primary_category":"cs.CL","announce_type":"new","date":"2025-07-09","first_seen":"2025-07-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2507.07188","pdf_url":"https://arxiv.org/pdf/2507.07188","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","调查偏差","算法保真度"],"reason":"直接研究LLM作为人类被试替代品的调查响应偏差，使用世界价值观调查真实人类数据…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:16","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":7,"question":"LLM在回答封闭式规范性调查问题时，对提示扰动是否稳健，以及是否表现出类似人类的回答偏差？","design":"用9个指令微调LLM扮演人类被试，对世界价值观调查的62个问题施加10种提示扰动（包括答案选项和问题措辞的修改），测量回答分布的变化，共进行167,400次模拟访谈。","baseline":"对照的真实人类数据来自世界价值观调查第7波（2017-2022）的核心变量，但论文主要比较扰动后回答分布与原始提示基线分布的差异。","findings":"所有模型均表现出明显的近因偏差，即不成比例地偏好最后一个选项；较大模型通常更稳健，但所有模型对语义改写和组合扰动仍敏感。","reliability":"论文指出，即使最大模型也对问题措辞变化敏感，提示设计和稳健性测试对使用LLM生成合成调查数据至关重要，但未讨论其他失效条件。","relevance":"该研究直接评估LLM作为人类被试替代品在调查中的偏差与可靠性，使用真实调查数据作为对照，并揭示近因偏差等关键失效模式，高度契合你的关注点，值得精读原文。","inspiration":"该方法通过系统施加提示扰动（如选项顺序、措辞改写）测量LLM回答分布变化，并设置原始提示基线作为对照，可借鉴用于稳健性检验设计｜可迁移到消费者通胀预期调查或政策公告解读实验，评估LLM模拟的预期形成是否对问卷设计敏感｜以LLM作为被试，施加不同措辞的通胀预期问题（如‘未来一年物价变化’ vs ‘通胀率’），结果变量为预期值分布，对照真实消费者预期调查数据（如密歇根大学调查）"}},{"id":"2507.08019","version":1,"title":"Signal or Noise? Evaluating Large Language Models in Resume Screening Across Contextual Variations and Human Expert Benchmarks","zh_title":"信号还是噪声？评估大语言模型在简历筛选中的表现：情境变化与人类专家基准","abstract":"This study investigates whether large language models (LLMs) exhibit consistent behavior (signal) or random variation (noise) when screening resumes against job descriptions, and how their performance compares to human experts. Using controlled datasets, we tested three LLMs (Claude, GPT, and Gemini) across contexts (No Company, Firm1 [MNC], Firm2 [Startup], Reduced Context) with identical and randomized resumes, benchmarked against three human recruitment experts. Analysis of variance revealed significant mean differences in four of eight LLM-only conditions and consistently significant differences between LLM and human evaluations (p < 0.01). Paired t-tests showed GPT adapts strongly to company context (p < 0.001), Gemini partially (p = 0.038 for Firm1), and Claude minimally (p > 0.1), while all LLMs differed significantly from human experts across contexts. Meta-cognition analysis highlighted adaptive weighting patterns that differ markedly from human evaluation approaches. Findings suggest LLMs offer interpretable patterns with detailed prompts but diverge substantially from human judgment, informing their deployment in automated hiring systems.","authors":["Aryan Varshney","Venkat Ram Reddy Ganuthula"],"categories":["cs.CL","econ.GN"],"primary_category":"cs.CL","announce_type":"new","date":"2025-07-08","first_seen":"2025-07-08","revised_at":null,"abs_url":"https://arxiv.org/abs/2507.08019","pdf_url":"https://arxiv.org/pdf/2507.08019","source_feed":"backfill","score":7,"bucket":"pending","rubric_hits":["A1","B1","B2"],"tags":["LLM仿真","人类对照","招聘决策"],"reason":"用LLM替代人类筛选简历，并与人类专家对照，属于人类仿真实验，但场景为招聘而非…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:34","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":38,"question":"大语言模型在简历筛选中的行为是稳定信号还是随机噪声？其评分与人类专家相比如何？","design":"用Claude、GPT、Gemini三个LLM扮演招聘者，在无公司、跨国公司、初创公司、简化上下文四种提示条件下对相同和随机化简历进行评分，测量评分均值和一致性，并与三位人类招聘专家对照。","baseline":"三位人类招聘专家在相同简历和提示条件下的评分。","findings":"LLM内部一致性有限，八种条件中四种存在显著均值差异；所有LLM评分均与人类专家显著不同，GPT对上下文适应最强，Claude最弱。元认知分析显示LLM的权重调整模式与人类差异明显。","reliability":"论文指出LLM评分与人类专家判断存在实质性分歧，在招聘等高利害场景中需谨慎部署，但未系统讨论失效条件与局限。","relevance":"该研究直接以人类专家为基准检验LLM在招聘决策中的仿真可靠性，属于批判性人类仿真研究，与研究者关注的经济学实验和政策评估场景高度相关，值得细读。","inspiration":"借鉴其多模型、多上下文条件与人类专家对照的设计，可迁移到信贷审批中的歧视研究。｜用多个LLM扮演信贷员，在不同银行类型（大行/小贷公司）和简化信息条件下审批贷款申请，以真实信贷员历史审批记录为基准，比较评分差异与偏见模式。"}},{"id":"2506.23610","version":1,"title":"Evaluating the Simulation of Human Personality-Driven Susceptibility to Misinformation with LLMs","zh_title":"评估大语言模型对人类人格驱动的错误信息易感性模拟","abstract":"Large language models (LLMs) make it possible to generate synthetic behavioural data at scale, offering an ethical and low-cost alternative to human experiments. Whether such data can faithfully capture psychological differences driven by personality traits, however, remains an open question. We evaluate the capacity of LLM agents, conditioned on Big-Five profiles, to reproduce personality-based variation in susceptibility to misinformation, focusing on news discernment, the ability to judge true headlines as true and false headlines as false. Leveraging published datasets in which human participants with known personality profiles rated headline accuracy, we create matching LLM agents and compare their responses to the original human patterns. Certain trait-misinformation associations, notably those involving Agreeableness and Conscientiousness, are reliably replicated, whereas others diverge, revealing systematic biases in how LLMs internalize and express personality. The results underscore both the promise and the limits of personality-aligned LLMs for behavioral simulation, and offer new insight into modeling cognitive diversity in artificial agents.","authors":["Manuel Pratelli","Marinella Petrocchi"],"categories":["cs.CL","cs.CY"],"primary_category":"cs.CL","announce_type":"new","date":"2025-06-30","first_seen":"2025-06-30","revised_at":null,"abs_url":"https://arxiv.org/abs/2506.23610","pdf_url":"https://arxiv.org/pdf/2506.23610","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","人格与行为","错误信息"],"reason":"用LLM模拟人格对错误信息易感性的影响，并与真实人类数据对照，评估仿真可靠性与…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:14","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":63,"question":"当赋予大语言模型明确的大五人格特征时，它们能否模拟人类在新闻辨别力（区分真假新闻标题准确性）上的表现？","design":"基于Calvillo等(2023)和Huang等(2024)的公开数据集，为每位人类参与者创建一个匹配的LLM智能体，通过人格注入管道赋予其对应的大五人格档案，让智能体对同一组新闻标题进行准确性评分，并比较合成数据与原始人类数据在人格-新闻辨别力关联上的异同。","baseline":"以Calvillo等(2023)的336名美国参与者数据为主要基准，该数据集包含参与者的大五人格档案和对真假新闻标题的准确性评分。","findings":"GPT-4o能复现部分人类特质-辨别力关联，如宜人性和尽责性与新闻辨别力正相关，但外向性和负面情绪性的模式与人类不一致；在仅分析假新闻时，高宜人性和高尽责性同样预测较低的易感性，但开放性的关联在不同模型设置下不稳定。","reliability":"论文承认某些人格特质（如外向性、负面情绪性）的模拟与人类模式存在系统性偏差，开放性-辨别力关联在不同模型参数和设置下不一致，表明LLM在表达人格时存在内在偏差，心理保真度有限。","relevance":"该研究直接使用LLM模拟人格对错误信息易感性的影响，并与真实人类数据严格对照，评估了仿真的可靠性与偏差，完全符合您对LLM人类仿真实验、经济学/政策评估场景及批判性失效条件分析的兴趣，值得精读原文。","inspiration":"该方法通过人格注入管道为LLM智能体赋予真实人类的大五人格档案，并严格对照人类基准数据评估仿真可靠性，值得借鉴其对照设计和偏差分析思路｜可迁移至消费者金融决策研究，如人格特质对投资风险偏好或过度借贷行为的影响｜以LLM智能体为被试，注入真实投资者的人格档案，让其评估不同风险等级的金融产品，结果变量为风险评分，对照真实投资者在相同产品上的风险偏好调查数据"}},{"id":"2506.23107","version":1,"title":"Can Large Language Models Capture Human Risk Preferences? A Cross-Cultural Study","zh_title":"大语言模型能捕捉人类风险偏好吗？一项跨文化研究","abstract":"Large language models (LLMs) have made significant strides, extending their applications to dialogue systems, automated content creation, and domain-specific advisory tasks. However, as their use grows, concerns have emerged regarding their reliability in simulating complex decision-making behavior, such as risky decision-making, where a single choice can lead to multiple outcomes. This study investigates the ability of LLMs to simulate risky decision-making scenarios. We compare model-generated decisions with actual human responses in a series of lottery-based tasks, using transportation stated preference survey data from participants in Sydney, Dhaka, Hong Kong, and Nanjing. Demographic inputs were provided to two LLMs -- ChatGPT 4o and ChatGPT o1-mini -- which were tasked with predicting individual choices. Risk preferences were analyzed using the Constant Relative Risk Aversion (CRRA) framework. Results show that both models exhibit more risk-averse behavior than human participants, with o1-mini aligning more closely with observed human decisions. Further analysis of multilingual data from Nanjing and Hong Kong indicates that model predictions in Chinese deviate more from actual responses compared to English, suggesting that prompt language may influence simulation performance. These findings highlight both the promise and the current limitations of LLMs in replicating human-like risk behavior, particularly in linguistic and cultural settings.","authors":["Bing Song","Jianing Liu","Sisi Jian","Chenyang Wu","Vinayak Dixit"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2025-06-29","first_seen":"2025-06-29","revised_at":null,"abs_url":"https://arxiv.org/abs/2506.23107","pdf_url":"https://arxiv.org/pdf/2506.23107","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["LLM仿真","风险偏好","跨文化对照"],"reason":"用LLM仿真人类风险决策，与真实调查数据对照，评估跨文化偏差，直接命中核心判据。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:13","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":64,"question":"大语言模型能否在跨文化背景下模拟人类的风险偏好？","design":"使用ChatGPT 4o和o1-mini两个模型，输入年龄、性别、教育、收入等人口统计信息，预测个体在彩票选择任务中的决策，并基于CRRA框架估计风险偏好。","baseline":"来自悉尼、达卡、香港和南京四个城市的真实交通陈述偏好调查数据，包含个体在彩票游戏中的实际选择。","findings":"两个模型均比人类更风险厌恶，o1-mini比4o更接近真实决策；在中文提示下，模型预测偏离实际的程度大于英文提示，表明语言影响仿真表现。","reliability":"论文指出模型在中文语境下偏差更大，且未完全复现人类风险行为，提示在语言和文化环境中的局限性。","relevance":"直接以真实人类数据为基准，评估LLM在风险决策仿真中的跨文化偏差，命中研究者关注的核心问题，值得精读原文。","inspiration":"该方法通过向LLM输入人口统计特征来模拟个体在彩票选择中的决策，并与真实调查数据对照，可用于评估模型在结构化风险任务中的行为偏差｜可迁移到金融决策中的风险偏好测量，如投资组合选择、保险购买或退休储蓄决策，检验LLM能否复现不同文化下个体的金融风险态度｜研究设计：以真实家庭金融调查数据为基准，向LLM输入年龄、收入、教育等特征，要求其在多组假设性投资选项中做出选择，结果变量为风险资产配置比例，对比模型预测与真实家庭行为的分布差异"}},{"id":"2506.21974","version":3,"title":"Don't Trust Generative Agents to Mimic Communication on Social Networks Unless You Benchmarked their Empirical Realism","zh_title":"不要相信生成式智能体能模仿社交网络上的交流，除非你对其经验现实主义进行了基准测试","abstract":"The ability of Large Language Models (LLMs) to mimic human behavior triggered a plethora of computational social science research, assuming that empirical studies of humans can be conducted with AI agents instead. Since there have been conflicting research findings on whether and when this hypothesis holds, there is a need to better understand the differences in their experimental designs. We focus on replicating the behavior of social network users with the use of LLMs for the analysis of communication on social networks. First, we provide a formal framework for the simulation of social networks, before focusing on the sub-task of imitating user communication. We empirically test different approaches to imitate user behavior on X in English and German. Our findings suggest that social simulations should be validated by their empirical realism measured in the setting in which the simulation components were fitted. With this paper, we argue for more rigor when applying generative-agent-based modeling for social simulation.","authors":["Simon Münker","Nils Schwager","Achim Rettinger"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2025-06-27","first_seen":"2025-06-27","revised_at":null,"abs_url":"https://arxiv.org/abs/2506.21974","pdf_url":"https://arxiv.org/pdf/2506.21974","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","A5","B1","B4"],"tags":["LLM仿真","社交网络","经验现实主义"],"reason":"用LLM仿真社交网络用户行为，有真实人类数据对照，并评估仿真可靠性，直接相关。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:13","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":29,"question":"LLM在多大程度上能忠实模仿社交网络（X平台）上不同用户群体的发帖、回复等沟通行为？","design":"基于X平台英文和德文政治讨论数据，用LLM（如Llama 3.1 70B）分别拟合政治人物（发帖者）和普通用户（回复者）的沟通行为，包括帖子生成、回复生成和回复可能性预测三个子任务，并比较不同建模方法的表现。","baseline":"真实人类数据：从X平台收集的2023年上半年美国和德国政治话语相关帖子（议员）和回复（普通用户），并标注主题。","findings":"LLM模仿用户沟通行为的效果因任务、语言和建模方法而异，并非总能忠实复现；社会仿真必须在其组件拟合的设定下验证经验现实主义。","reliability":"论文承认仿真框架存在简化，如未完全建模网络结构、内容过滤可能损失信息，且结论依赖于特定平台和语言数据集，泛化性有限。","relevance":"该研究直接回应了LLM仿真人类社交行为的可靠性问题，提供了有真实人类基准的实证评估，并指出仿真失效的条件，高度契合研究者对批判性仿真研究的关注。","inspiration":"该方法值得借鉴之处在于，它将LLM仿真分解为帖子生成、回复生成和回复可能性预测三个子任务，并分别与真实人类数据对比，从而定位仿真失效的具体环节｜该思路可迁移到政策公告的预期形成研究，例如分析央行沟通或财政政策声明如何影响市场参与者的预期与讨论｜研究设计：以LLM模拟投资者和分析师，处理为不同措辞或情感倾向的政策声明，结果变量为LLM生成的预期文本和市场反应预测，对照真实数据可来自央行声明后社交媒体（如Twitter/微博）上的实际讨论与市场调查预期数据"}},{"id":"2507.02919","version":1,"title":"ChatGPT is not A Man but Das Man: Representativeness and Structural Consistency of Silicon Samples Generated by Large Language Models","zh_title":"ChatGPT不是人而是常人：大语言模型生成硅样本的代表性与结构一致性","abstract":"Large language models (LLMs) in the form of chatbots like ChatGPT and Llama are increasingly proposed as \"silicon samples\" for simulating human opinions. This study examines this notion, arguing that LLMs may misrepresent population-level opinions. We identify two fundamental challenges: a failure in structural consistency, where response accuracy doesn't hold across demographic aggregation levels, and homogenization, an underrepresentation of minority opinions. To investigate these, we prompted ChatGPT (GPT-4) and Meta's Llama 3.1 series (8B, 70B, 405B) with questions on abortion and unauthorized immigration from the American National Election Studies (ANES) 2020. Our findings reveal significant structural inconsistencies and severe homogenization in LLM responses compared to human data. We propose an \"accuracy-optimization hypothesis,\" suggesting homogenization stems from prioritizing modal responses. These issues challenge the validity of using LLMs, especially chatbots AI, as direct substitutes for human survey data, potentially reinforcing stereotypes and misinforming policy.","authors":["Dai Li","Linzhuo Li","Huilian Sophie Qiu"],"categories":["cs.CL","cs.CY","cs.ET"],"primary_category":"cs.CL","announce_type":"new","date":"2025-06-25","first_seen":"2025-06-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2507.02919","pdf_url":"https://arxiv.org/pdf/2507.02919","source_feed":"backfill","score":10,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM人类仿真","调查数据对照","算法偏差"],"reason":"直接用LLM模拟人类调查回答，与ANES真实数据对照，揭示结构不一致和同质化，…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:16","error":null,"has_summary":true,"summary":{"generated_at":"2025-06-25","rank":8,"question":"大语言模型生成的“硅样本”能否代表人类群体意见？","design":"使用ChatGPT (GPT-4) 和 Llama 3.1系列 (8B, 70B, 405B) 模型，输入ANES 2020中关于堕胎和非法移民的问题，生成模拟回答，并与真实人类数据对比，分析结构一致性和同质化程度。","baseline":"美国国家选举研究（ANES）2020的调查数据。","findings":"LLM回答存在显著的结构不一致性（不同人口聚合水平的准确率不一致）和严重的同质化（少数意见被低估）。作者提出“精度优化假说”，认为同质化源于模型优先输出众数回答。","reliability":"论文指出LLM可能误代表群体意见，存在结构不一致和同质化问题，挑战了直接替代人类调查数据的有效性，可能强化刻板印象并误导政策。","relevance":"高度相关。该研究直接评估了LLM仿真人类意见的可靠性与偏差，有真实人类数据对照，并指出了失效条件（结构不一致和同质化），符合研究者的核心关注点，值得精读原文。","inspiration":"该方法通过将LLM作为被试，输入真实调查问题并对比其回答与人类基准数据，来评估仿真的一致性和偏差，可借鉴其对照设计和偏差度量方式｜可迁移到政策预期形成的实验，例如研究公众对央行利率决议声明的通胀预期反应｜以GPT-4等LLM为被试，输入央行政策声明文本，测量其输出的通胀预期值，并与密歇根消费者调查的真实预期数据对比，检验LLM是否复现预期分布及同质化偏差"}},{"id":"2506.15041","version":1,"title":"Identifying economic narratives in large text corpora -- An integrated approach using Large Language Models","zh_title":"利用大语言模型识别大型文本语料库中的经济叙事——一种集成方法","abstract":"As interest in economic narratives has grown in recent years, so has the number of pipelines dedicated to extracting such narratives from texts. Pipelines often employ a mix of state-of-the-art natural language processing techniques, such as BERT, to tackle this task. While effective on foundational linguistic operations essential for narrative extraction, such models lack the deeper semantic understanding required to distinguish extracting economic narratives from merely conducting classic tasks like Semantic Role Labeling. Instead of relying on complex model pipelines, we evaluate the benefits of Large Language Models (LLMs) by analyzing a corpus of Wall Street Journal and New York Times newspaper articles about inflation. We apply a rigorous narrative definition and compare GPT-4o outputs to gold-standard narratives produced by expert annotators. Our results suggests that GPT-4o is capable of extracting valid economic narratives in a structured format, but still falls short of expert-level performance when handling complex documents and narratives. Given the novelty of LLMs in economic research, we also provide guidance for future work in economics and the social sciences that employs LLMs to pursue similar objectives.","authors":["Tobias Schmidt","Kai-Robin Lange","Matthias Reccius","Henrik Müller","Michael Roos","Carsten Jentsch"],"categories":["econ.GN","cs.CL"],"primary_category":"econ.GN","announce_type":"new","date":"2025-06-18","first_seen":"2025-06-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2506.15041","pdf_url":"https://arxiv.org/pdf/2506.15041","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM标注","经济叙事","文本分析"],"reason":"用LLM替代专家标注经济叙事，属替代人工标注员，非仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:58:40","error":null,"has_summary":false,"summary":null},{"id":"2506.21587","version":2,"title":"A Cross-Cultural Comparison of LLM-based Public Opinion Simulation: Evaluating Chinese and U.S. Models on Diverse Societies","zh_title":"基于大语言模型的舆论仿真跨文化比较：评估中美模型在多元社会上的表现","abstract":"This study evaluates the ability of DeepSeek, an open-source large language model (LLM), to simulate public opinions in comparison to LLMs developed by major tech companies. By comparing DeepSeek-R1 and DeepSeek-V3 with Qwen2.5, GPT-4o, and Llama-3.3 and utilizing survey data from the American National Election Studies (ANES) and the Zuobiao dataset of China, we assess these models' capacity to predict public opinions on social issues in both China and the United States, highlighting their comparative capabilities between countries. Our findings indicate that DeepSeek-V3 performs best in simulating U.S. opinions on the abortion issue compared to other topics such as climate change, gun control, immigration, and services for same-sex couples, primarily because it more accurately simulates responses when provided with Democratic or liberal personas. For Chinese samples, DeepSeek-V3 performs best in simulating opinions on foreign aid and individualism but shows limitations in modeling views on capitalism, particularly failing to capture the stances of low-income and non-college-educated individuals. It does not exhibit significant differences from other models in simulating opinions on traditionalism and the free market. Further analysis reveals that all LLMs exhibit the tendency to overgeneralize a single perspective within demographic groups, often defaulting to consistent responses within groups. These findings highlight the need to mitigate cultural and demographic biases in LLM-driven public opinion modeling, calling for approaches such as more inclusive training methodologies.","authors":["Weihong Qi","Fan Huang","Jisun An","Haewoon Kwak"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2025-06-17","first_seen":"2025-06-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2506.21587","pdf_url":"https://arxiv.org/pdf/2506.21587","source_feed":"backfill","score":10,"bucket":"selected","rubric_hits":["A1","A3","B1","B2","B4"],"tags":["LLM人类仿真","舆论模拟","跨文化比较"],"reason":"用LLM仿真中美公众舆论，有ANES和Zuobiao真实调查数据对照，涉及社会…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:13","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":15,"question":"LLM的文化出身是否带来舆论模拟的“主场优势”？不同模型在中美社会议题上存在怎样的文化与人口偏差？","design":"用DeepSeek-R1、DeepSeek-V3、Qwen2.5、GPT-4o、Llama-3.3，根据ANES和Zuobiao数据集中个体的种族、性别、年龄、收入、教育等人口属性构造提示，让模型回答社会议题问卷，聚合后比较模拟分布与真实调查分布。","baseline":"2020年美国国家选举研究（ANES）的2457人调查数据，以及中国“左标”数据集随机抽取的2000人调查数据。","findings":"DeepSeek-V3在模拟美国堕胎议题上表现最好，主要因为能较准地模拟民主党/自由派，但共和党/保守派模拟差；所有模型都倾向于在人口群体内过度概括单一观点，尤其在中国资本主义议题上，Qwen2.5通过输出统一回答虚高准确率。","reliability":"论文指出所有模型均存在显著的文化与人口偏差，难以捕捉特定群体的细微观点，常退化为刻板或过度概括的回答；未讨论提示设计、模型版本或数据集时效性等可能影响结论的因素。","relevance":"直接回应LLM仿真人类舆论的可靠性与偏差问题，有中美真实调查数据对照，揭示模型在跨文化情境下系统性失效的模式，对理解仿真在经济学/政策评估中的局限有重要参考价值，强烈建议阅读原文。","inspiration":"该方法通过将人口属性编码为提示来模拟个体回答，并对比聚合分布与真实调查数据，为评估仿真偏差提供了可操作的对照框架｜可迁移到消费者信心调查或通胀预期形成的仿真研究，检验LLM能否复现不同人口群体的预期差异｜以LLM作为被试，输入年龄、收入、教育等属性，让其预测未来通胀率，结果变量为预期值分布，以密歇根大学消费者调查的微观数据作为真实基准，比较模拟与真实分布的偏差模式"}},{"id":"2506.14611","version":2,"title":"Exploring MLLMs Perception of Network Visualization Principles","zh_title":"探索多模态大语言模型对网络可视化原则的感知","abstract":"In this paper, we test whether Multimodal Large Language Models (MLLMs) can match human-subject performance in tasks involving the perception of properties in network layouts. Specifically, we replicate a human-subject experiment about perceiving quality (namely stress) in network layouts using GPT-4o, Gemini-2.5 and Qwen2.5. Our experiments show that giving MLLMs the same study information as trained human participants yields performance comparable to that of human experts and exceeds that of untrained non-experts. Additionally, we show that prompt engineering that deviates from the human-subject experiment can lead to better-than-human performance in some settings. Interestingly, like human subjects, the MLLMs seem to rely on visual proxies rather than computing the actual value of stress, indicating some sense or facsimile of perception. Explanations from the models are similar to those used by the human participants (e.g., an even distribution of nodes and uniform edge lengths).","authors":["Jacob Miller","Markus Wallinger","Ludwig Felder","Timo Brand","Henry Förster","Johannes Zink","Chunyang Chen","Stephen Kobourov"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2025-06-17","first_seen":"2025-06-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2506.14611","pdf_url":"https://arxiv.org/pdf/2506.14611","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","人类被试替代","网络感知实验"],"reason":"用MLLM复现人类网络布局感知实验，与真实人类数据对照，评估仿真可靠性并指出失…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:11","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":65,"question":"多模态大语言模型在判断网络布局应力时，能否达到人类被试的感知表现？","design":"用GPT-4o、Gemini-2.5和Qwen2.5模拟人类被试，复现Mooney等人的配对刺激实验：向模型展示同一网络的两幅节点链接图，要求选择应力更低的图或判断两者相似，并比较不同提示工程（与人类相同指令 vs. 偏离人类指令）下的表现。","baseline":"Mooney等人实验中的人类被试数据，包括未训练新手、训练新手和专家三组。","findings":"给予与人类训练参与者相同的信息时，MLLM的表现与人类专家相当，优于未训练新手；通过偏离人类实验的提示工程，在某些设置下可获得超越人类的表现。MLLM与人类相似，依赖视觉代理（如节点均匀分布、边长一致）而非实际计算应力值。","reliability":"论文未讨论","relevance":"该研究直接用MLLM复现人类网络感知实验，并与真实人类数据对照，评估仿真可靠性，同时指出提示工程可导致超越人类的表现，符合您对LLM仿真人类实验、基准对照及失效条件探索的关注，值得精读原文。","inspiration":"该方法借鉴了用多模态大语言模型复现人类感知实验，并通过提示工程模拟不同信息条件（如训练新手、专家）来检验仿真表现，同时以真实人类数据作为基准对照。｜可迁移到金融图表解读与投资决策实验，例如研究投资者如何从网络关系图（如持股网络、供应链网络）中提取风险信息并形成投资判断。｜以GPT-4o等MLLM为被试，展示同一组持股网络的不同布局图，要求选择更易读或风险更清晰的图，结果变量为选择准确率与反应时间，对照真实投资者在相同任务上的行为数据。"}},{"id":"2506.21574","version":1,"title":"Digital Gatekeepers: Exploring Large Language Model's Role in Immigration Decisions","zh_title":"数字守门人：探索大语言模型在移民决策中的作用","abstract":"With globalization and increasing immigrant populations, immigration departments face significant work-loads and the challenge of ensuring fairness in decision-making processes. Integrating artificial intelligence offers a promising solution to these challenges. This study investigates the potential of large language models (LLMs),such as GPT-3.5 and GPT-4, in supporting immigration decision-making. Utilizing a mixed-methods approach,this paper conducted discrete choice experiments and in-depth interviews to study LLM decision-making strategies and whether they are fair. Our findings demonstrate that LLMs can align their decision-making with human strategies, emphasizing utility maximization and procedural fairness. Meanwhile, this paper also reveals that while ChatGPT has safeguards to prevent unintentional discrimination, it still exhibits stereotypes and biases concerning nationality and shows preferences toward privileged group. This dual analysis highlights both the potential and limitations of LLMs in automating and enhancing immigration decisions.","authors":["Yicheng Mao","Yang Zhao"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2025-06-15","first_seen":"2025-06-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2506.21574","pdf_url":"https://arxiv.org/pdf/2506.21574","source_feed":"backfill","score":8,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["LLM仿真","决策实验","公平性评估"],"reason":"用LLM模拟移民决策并与人类策略对照，评估公平性与偏差，直接相关。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:12","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":59,"question":"大语言模型在移民决策中的决策策略是否与人类一致，以及其决策是否公平、是否存在偏见？","design":"使用GPT-3.5和GPT-4作为被试，复现Hainmueller和Hopkins (2015)的离散选择实验，生成10,000个随机移民决策场景，并辅以深度访谈，分析模型的决策策略、公平性和偏见。","baseline":"对照Hainmueller和Hopkins (2015)中的人类决策数据，比较LLM与人类在移民决策中的策略一致性。","findings":"LLM的决策策略与人类相似，强调效用最大化和程序公平；但ChatGPT虽设有防止无意歧视的机制，仍表现出基于国籍的刻板印象和偏见，并偏好特权群体。","reliability":"论文未讨论","relevance":"该研究直接用LLM模拟人类移民决策，并与真实人类实验数据对照，评估公平性与偏差，属于典型的LLM人类仿真研究，且涉及政策评估场景，值得精读原文。","inspiration":"该方法借鉴了离散选择实验与真实人类基准对照的设计，可系统评估LLM在政策决策中的策略一致性与偏差｜可迁移至信贷审批歧视研究，检验LLM是否复现人类审批中的种族、性别或收入偏见｜以LLM作为信贷审批官，处理随机生成的贷款申请人档案（操纵种族、收入等特征），结果变量为批准决策，对照真实银行审批数据或审计研究结果"}},{"id":"2506.12664","version":2,"title":"Behavioral Generative Agents for Energy Operations","zh_title":"用于能源运营的行为生成式智能体","abstract":"Problem definition: Accurately modeling consumer behavior in energy operations is challenging due to uncertainty, behavioral heterogeneity, and limited empirical data-particularly in low-frequency, high-impact events. While generative AI trained on large-scale human data offers new opportunities to study decision behavior, its role in operational applications remains unclear. We examine how generative agents can support customer behavior discovery in energy operations, complementing rather than replacing human-based experiments. Methodology/results: We introduce a novel approach leveraging generative agents-artificial agents powered by large language models-to simulate sequential customer decisions under dynamic electricity prices and outage risks. We find that these agents behave more optimally and rationally in simpler market scenarios, while their performance becomes more variable and suboptimal as task complexity rises. Furthermore, the agents exhibit heterogeneous customer preferences, consistently maintaining distinct, persona-driven reasoning patterns in both operational decisions and textual reasoning. Comparisons with dynamic programming and greedy policy benchmarks show alignment between specific personas and distinct heuristic decision policies. In low-frequency, high-impact events such as blackouts, agents prioritize energy reliability over cost or profit, demonstrating their ability to uncover behavioral patterns beyond the rigidity of traditional mathematical models. Managerial Implications: Our findings suggest that behavioral generative agents can serve as scalable and flexible tools for studying consumer behavior in energy operations. By enabling controlled experiments across heterogeneous customer types and rare events, these agents can enhance the design of energy management systems and support more informed analysis of energy policies and incentive programs.","authors":["Cong Chen","Omer Karaduman","Xu Kuang"],"categories":["cs.AI","eess.SY"],"primary_category":"cs.AI","announce_type":"new","date":"2025-06-14","first_seen":"2025-06-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2506.12664","pdf_url":"https://arxiv.org/pdf/2506.12664","source_feed":"backfill","score":7,"bucket":"pending","rubric_hits":["A3","B2","B4"],"tags":["LLM仿真","能源经济","行为建模"],"reason":"用LLM agent模拟消费者能源决策，涉及经济场景，有批判性讨论，但缺真实人…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:10","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":114,"question":"生成式智能体能否在动态电价和停电风险下模拟消费者能源管理决策，并揭示不同任务难度和用户画像下的行为模式？","design":"使用大语言模型驱动的生成式智能体，通过提示词赋予其不同用户画像（如谨慎的感性者、逐利的理性思考者、现实主义者），在模拟的家庭能源管理场景中逐日决定是否对家用电池充电或放电，观察其决策表现和文本推理，并与动态规划最优策略和贪婪启发式基准进行比较。","baseline":"无对照","findings":"智能体在简单市场场景中表现接近最优，但随着任务复杂度上升，决策质量下降且变异性增大；不同画像的智能体展现出异质且稳定的偏好，在停电等低频高影响事件中优先保障能源可靠性而非成本或利润。","reliability":"论文承认缺乏真实人类数据作为对照，仅采用数学基准（动态规划、贪婪策略）评估智能体推理质量，且智能体在复杂任务中表现下降，表明其可靠性受任务难度影响。","relevance":"该研究用LLM智能体模拟消费者能源决策，属于经济学实验场景，但缺少真实人类基准数据，批判性讨论有限，与研究者关注的高质量人类仿真和失效条件分析存在差距，可酌情阅读以了解方法。","inspiration":"该方法通过提示词赋予LLM不同用户画像来模拟异质决策偏好，可用于经济金融实验中的处理操纵｜可迁移到消费者跨期选择研究，如不同利率或补贴政策下的储蓄与消费决策｜以LLM智能体为被试，处理为不同利率条件或政策信息提示，结果变量为模拟的消费-储蓄分配，对照真实家庭调查数据（如PSID）中的跨期选择弹性"}},{"id":"2506.10546","version":1,"title":"Nowcasting the euro area with social media data","zh_title":"利用社交媒体数据对欧元区进行即时预测","abstract":"Using a state-of-the-art large language model, we extract forward-looking and context-sensitive signals related to inflation and unemployment in the euro area from millions of Reddit submissions and comments. We develop daily indicators that incorporate, in addition to posts, the social interaction among users. Our empirical results show consistent gains in out-of-sample nowcasting accuracy relative to daily newspaper sentiment and financial variables, especially in unusual times such as the (post-)COVID-19 period. We conclude that the application of AI tools to the analysis of social media, specifically Reddit, provides useful signals about inflation and unemployment in Europe at daily frequency and constitutes a useful addition to the toolkit available to economic forecasters and nowcasters.","authors":["Konstantin Boss","Luigi Longo","Luca Onorante"],"categories":["econ.EM"],"primary_category":"econ.EM","announce_type":"new","date":"2025-06-12","first_seen":"2025-06-12","revised_at":null,"abs_url":"https://arxiv.org/abs/2506.10546","pdf_url":"https://arxiv.org/pdf/2506.10546","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["经济预测","社交媒体分析","大语言模型应用"],"reason":"用LLM提取社交媒体信号做经济预测，不涉及人类行为仿真或对照实验。","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:58:56","error":null,"has_summary":false,"summary":null},{"id":"2507.19495","version":1,"title":"Simulating Human Behavior with the Psychological-mechanism Agent: Integrating Feeling, Thought, and Action","zh_title":"基于心理机制代理的人类行为模拟：整合感受、思维与行动","abstract":"Generative agents have made significant progress in simulating human behavior, but existing frameworks often simplify emotional modeling and focus primarily on specific tasks, limiting the authenticity of the simulation. Our work proposes the Psychological-mechanism Agent (PSYA) framework, based on the Cognitive Triangle (Feeling-Thought-Action), designed to more accurately simulate human behavior. The PSYA consists of three core modules: the Feeling module (using a layer model of affect to simulate changes in short-term, medium-term, and long-term emotions), the Thought module (based on the Triple Network Model to support goal-directed and spontaneous thinking), and the Action module (optimizing agent behavior through the integration of emotions, needs and plans). To evaluate the framework's effectiveness, we conducted daily life simulations and extended the evaluation metrics to self-influence, one-influence, and group-influence, selection five classic psychological experiments for simulation. The results show that the PSYA framework generates more natural, consistent, diverse, and credible behaviors, successfully replicating human experimental outcomes. Our work provides a richer and more accurate emotional and cognitive modeling approach for generative agents and offers an alternative to human participants in psychological experiments.","authors":["Qing Dong","Pengyuan Liu","Dong Yu","Chen Kang"],"categories":["cs.HC","cs.AI"],"primary_category":"cs.HC","announce_type":"new","date":"2025-06-04","first_seen":"2025-06-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2507.19495","pdf_url":"https://arxiv.org/pdf/2507.19495","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","A5","B1","B2"],"tags":["人类行为仿真","心理学实验复现","生成式代理"],"reason":"用LLM代理复现经典心理学实验，并与真实人类数据对照，直接替代人类被试。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:19","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":30,"question":"能否基于认知三角（感受-思维-行动）构建心理机制智能体框架，更真实地模拟人类日常行为并复现经典心理学实验？","design":"提出PSYA框架，包含感受（分层情感模型）、思维（三重网络模型支持目标导向与自发思维）、行动（整合情绪、需求与计划）三个模块，基于Llama-3-70B构建智能体，在虚拟小镇中模拟8个智能体的日常生活，并选取5个经典心理学实验进行仿真，测量行为自然度、一致性、多样性等指标。","baseline":"以经典心理学实验的真实人类结果作为对照基准，验证智能体能否复现人类实验数据。","findings":"PSYA框架能生成更自然、一致、多样和可信的行为，成功复现了所选经典心理学实验的结果；分层情感和自发思维模块显著提升了行为真实性。","reliability":"论文未讨论","relevance":"该研究直接用LLM智能体替代人类被试复现经典心理学实验，并与真实人类数据对照，高度契合研究者对经济学实验和政策评估场景中仿真可靠性的关注，值得精读原文。","inspiration":"借鉴PSYA框架的分层情感与自发思维模块设计，在LLM智能体中嵌入情绪和认知偏差来模拟经济决策的心理机制｜可迁移到消费者跨期选择实验，考察情绪波动对时间偏好一致性的影响｜以LLM智能体为被试，施加情绪启动处理（如积极/消极文本诱导），测量跨期选择中的折现率变化，并与真实人类实验数据（如经典双曲折现研究）对照"}},{"id":"2506.06377","version":1,"title":"Evaluating Large Language Model Capabilities in Assessing Spatial Econometrics Research","zh_title":"评估大语言模型在空间计量经济学研究评估中的能力","abstract":"This paper investigates Large Language Models (LLMs) ability to assess the economic soundness and theoretical consistency of empirical findings in spatial econometrics. We created original and deliberately altered \"counterfactual\" summaries from 28 published papers (2005-2024), which were evaluated by a diverse set of LLMs. The LLMs provided qualitative assessments and structured binary classifications on variable choice, coefficient plausibility, and publication suitability. The results indicate that while LLMs can expertly assess the coherence of variable choices (with top models like GPT-4o achieving an overall F1 score of 0.87), their performance varies significantly when evaluating deeper aspects such as coefficient plausibility and overall publication suitability. The results further revealed that the choice of LLM, the specific characteristics of the paper and the interaction between these two factors significantly influence the accuracy of the assessment, particularly for nuanced judgments. These findings highlight LLMs' current strengths in assisting with initial, more surface-level checks and their limitations in performing comprehensive, deep economic reasoning, suggesting a potential assistive role in peer review that still necessitates robust human oversight.","authors":["Giuseppe Arbia","Luca Morandini","Vincenzo Nardelli"],"categories":["cs.CY","cs.LG","econ.EM","stat.CO"],"primary_category":"cs.CY","announce_type":"new","date":"2025-06-04","first_seen":"2025-06-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2506.06377","pdf_url":"https://arxiv.org/pdf/2506.06377","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM评估","同行评审","空间计量经济学"],"reason":"LLM替代同行评审专家评估论文，属于替代人类劳动而非仿真被试，但涉及经济学场景…","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:58:56","error":null,"has_summary":false,"summary":null},{"id":"2506.04478","version":2,"title":"Matching Markets Meet LLMs: Algorithmic Reasoning with Ranked Preferences","zh_title":"匹配市场遇见大语言模型：基于排序偏好的算法推理","abstract":"The rise of Large Language Models (LLMs) has driven progress in reasoning tasks -- from program synthesis to scientific hypothesis generation -- yet their ability to handle ranked preferences and structured algorithms in combinatorial domains remains underexplored. We study matching markets, a core framework behind applications like resource allocation and ride-sharing, which require reconciling individual ranked preferences to ensure stable outcomes. We evaluate several state-of-the-art models on a hierarchy of preference-based reasoning tasks -- ranging from stable-matching generation to instability detection, instability resolution, and fine-grained preference queries -- to systematically expose their logical and algorithmic limitations in handling ranked inputs. Surprisingly, even top-performing models with advanced reasoning struggle to resolve instability in large markets, often failing to identify blocking pairs or execute algorithms iteratively. We further show that parameter-efficient fine-tuning (LoRA) significantly improves performance in small markets, but fails to bring about a similar improvement on large instances, suggesting the need for more sophisticated strategies to improve LLMs' reasoning with larger-context inputs.","authors":["Hadi Hosseini","Samarth Khanna","Ronak Singh"],"categories":["cs.AI","cs.GT","econ.TH"],"primary_category":"cs.AI","announce_type":"new","date":"2025-06-04","first_seen":"2025-06-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2506.04478","pdf_url":"https://arxiv.org/pdf/2506.04478","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["匹配市场","算法推理","LLM评估"],"reason":"纯多智能体算法推理，无人类行为对照，不涉及人类仿真实验。","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:59:03","error":null,"has_summary":false,"summary":null},{"id":"2506.02827","version":1,"title":"TO-GATE: Clarifying Questions and Summarizing Responses with Trajectory Optimization for Eliciting Human Preference","zh_title":"TO-GATE：通过轨迹优化澄清问题并总结回答以获取人类偏好","abstract":"Large language models (LLMs) can effectively elicit human preferences through multi-turn dialogue. Complex tasks can be accomplished through iterative clarifying questions and final responses generated by an LLM acting as a questioner (STaR-GATE; Andukuri et al., 2024}). However, existing approaches based on self-taught reasoning struggle to identify optimal dialogue trajectories and avoid irrelevant questions to the tasks. To address this limitation, we propose TO-GATE, a novel framework that enhances question generation through trajectory optimization, which consists of two key components: a clarification resolver that generates optimal questioning trajectories, and a summarizer that ensures task-aligned final responses. The trajectory optimization enables the model to produce effective elicitation questions and summary responses tailored to specific tasks. Experimental results demonstrate that TO-GATE significantly outperforms baseline methods, achieving a 9.32% improvement on standard preference elicitation tasks.","authors":["Yulin Dou","Jiangming Liu"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2025-06-03","first_seen":"2025-06-03","revised_at":null,"abs_url":"https://arxiv.org/abs/2506.02827","pdf_url":"https://arxiv.org/pdf/2506.02827","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["偏好获取","多轮对话","轨迹优化"],"reason":"多智能体协作优化提问以获取人类偏好，非用LLM仿真人类被试，无人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:59:43","error":null,"has_summary":false,"summary":null},{"id":"2506.00856","version":3,"title":"Can AI Master Econometrics? Evidence from Econometrics AI Agent on Expert-Level Tasks","zh_title":"AI能掌握计量经济学吗？来自专家级任务中计量经济学AI智能体的证据","abstract":"Can AI effectively perform complex econometric analysis traditionally requiring human expertise? This paper evaluates AI agents' capability to master econometrics, focusing on empirical analysis performance. We develop ``MetricsAI'', an Econometrics AI Agent built on the open-source MetaGPT framework. This agent exhibits outstanding performance in: (1) planning econometric tasks strategically, (2) generating and executing code, (3) employing error-based reflection for improved robustness, and (4) allowing iterative refinement through multi-round conversations. We construct two datasets from academic coursework materials and published research papers to evaluate performance against real-world challenges. Comparative testing shows our domain-specialized AI agent significantly outperforms both benchmark large language models (LLMs) and general-purpose AI agents. This work establishes a testbed for exploring AI's impact on social science research and enables cost-effective integration of domain expertise, making advanced econometric methods accessible to users with minimal coding skills. Furthermore, our AI agent enhances research reproducibility and offers promising pedagogical applications for econometrics teaching.","authors":["Qiang Chen","Tianyang Han","Jin Li","Ye Luo","Zigan Wang","Yuxiao Wu","Xiaowei Zhang","Tuo Zhou"],"categories":["econ.EM","cs.AI"],"primary_category":"econ.EM","announce_type":"new","date":"2025-06-01","first_seen":"2025-06-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2506.00856","pdf_url":"https://arxiv.org/pdf/2506.00856","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["AI智能体","计量经济学","自动化数据分析"],"reason":"纯多智能体协作完成计量任务，不涉及人类行为仿真或对照。","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:58:56","error":null,"has_summary":false,"summary":null},{"id":"2505.24640","version":1,"title":"Efficient Text Encoders for Labor Market Analysis","zh_title":"面向劳动力市场分析的高效文本编码器","abstract":"Labor market analysis relies on extracting insights from job advertisements, which provide valuable yet unstructured information on job titles and corresponding skill requirements. While state-of-the-art methods for skill extraction achieve strong performance, they depend on large language models (LLMs), which are computationally expensive and slow. In this paper, we propose \\textbf{ConTeXT-match}, a novel contrastive learning approach with token-level attention that is well-suited for the extreme multi-label classification task of skill classification. \\textbf{ConTeXT-match} significantly improves skill extraction efficiency and performance, achieving state-of-the-art results with a lightweight bi-encoder model. To support robust evaluation, we introduce \\textbf{Skill-XL}, a new benchmark with exhaustive, sentence-level skill annotations that explicitly address the redundancy in the large label space. Finally, we present \\textbf{JobBERT V2}, an improved job title normalization model that leverages extracted skills to produce high-quality job title representations. Experiments demonstrate that our models are efficient, accurate, and scalable, making them ideal for large-scale, real-time labor market analysis.","authors":["Jens-Joris Decorte","Jeroen Van Hautte","Chris Develder","Thomas Demeester"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2025-05-30","first_seen":"2025-05-30","revised_at":null,"abs_url":"https://arxiv.org/abs/2505.24640","pdf_url":"https://arxiv.org/pdf/2505.24640","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["技能抽取","职位标准化","对比学习"],"reason":"纯NLP技能抽取与职位标准化，无LLM仿真人类被试或行为对照。","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:59:26","error":null,"has_summary":false,"summary":null},{"id":"2505.23025","version":1,"title":"Learning to Regulate: A New Event-Level Dataset of Capital Control Measures","zh_title":"学习监管：一个新的资本管制措施事件级数据集","abstract":"We construct a novel event-level Capital Control Measures (CCM) dataset covering 196 countries from 1999 to 2023 by leveraging prompt-based large language models (LLMs). The dataset enables event study analysis and cross-country comparisons based on rich policy attributes, including action type, intensity, direction, implementing entity, and other multidimensional characteristics. Using a two-step prompt framework with GPT-4.1, we extract structured information from the IMF's Annual Report on Exchange Arrangements and Exchange Restrictions (AREAER), resulting in 5,198 capital control events with 27 annotated fields and corresponding model reasoning. Secondly, to facilitate real-time classification and extension to external sources, we fine-tune an open-source Meta Llama 3.1-8B model, named CCM-Llama, trained on AREAER change logs and final status reports. The model achieves 90.09\\% accuracy in category classification and 99.55\\% in status prediction. Finally, we apply the CCM dataset in an empirical application: an event study on China, Australia, and the US. The results show that inward capital control measures significantly reduce fund inflows within one month, and restrictive policies tend to have stronger effects than liberalizing ones, with notable heterogeneity across countries. Our work contributes to the growing literature on the use of LLMs in economics by providing both a novel high-frequency policy dataset and a replicable framework for automated classification of capital control events from diverse and evolving information sources.","authors":["Geyue Sun","Xiao Liu","Tomas Williams","Roberto Samaniego"],"categories":["econ.GN"],"primary_category":"econ.GN","announce_type":"new","date":"2025-05-29","first_seen":"2025-05-29","revised_at":null,"abs_url":"https://arxiv.org/abs/2505.23025","pdf_url":"https://arxiv.org/pdf/2505.23025","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["资本管制","事件数据集","LLM信息抽取"],"reason":"用LLM提取政策事件数据，属NLP信息抽取，非仿真人类被试行为。","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:58:40","error":null,"has_summary":false,"summary":null},{"id":"2505.21997","version":1,"title":"Leveraging Interview-Informed LLMs to Model Survey Responses: Comparative Insights from AI-Generated and Human Data","zh_title":"利用访谈信息引导大语言模型建模调查回答：AI生成数据与人类数据的比较洞察","abstract":"Mixed methods research integrates quantitative and qualitative data but faces challenges in aligning their distinct structures, particularly in examining measurement characteristics and individual response patterns. Advances in large language models (LLMs) offer promising solutions by generating synthetic survey responses informed by qualitative data. This study investigates whether LLMs, guided by personal interviews, can reliably predict human survey responses, using the Behavioral Regulations in Exercise Questionnaire (BREQ) and interviews from after-school program staff as a case study. Results indicate that LLMs capture overall response patterns but exhibit lower variability than humans. Incorporating interview data improves response diversity for some models (e.g., Claude, GPT), while well-crafted prompts and low-temperature settings enhance alignment between LLM and human responses. Demographic information had less impact than interview content on alignment accuracy. These findings underscore the potential of interview-informed LLMs to bridge qualitative and quantitative methodologies while revealing limitations in response variability, emotional interpretation, and psychometric fidelity. Future research should refine prompt design, explore bias mitigation, and optimize model settings to enhance the validity of LLM-generated survey data in social science research.","authors":["Jihong Zhang","Xinya Liang","Anqi Deng","Nicole Bonge","Lin Tan","Ling Zhang","Nicole Zarrett"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2025-05-28","first_seen":"2025-05-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2505.21997","pdf_url":"https://arxiv.org/pdf/2505.21997","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","调查回答","人类数据对照"],"reason":"用访谈引导LLM生成调查回答，与真实人类数据对照，评估仿真可靠性与偏差，直接命…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:10","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":8,"question":"以个人访谈信息引导的大语言模型能否可靠预测人类在标准化问卷上的回答？","design":"使用Claude、GPT等大语言模型，输入课后项目工作人员的个人访谈文本和人口统计信息作为提示，生成其在运动行为调节问卷（BREQ）上的Likert量表回答，并与真实人类回答进行对比。","baseline":"同一批课后项目工作人员在BREQ问卷上的真实人类回答。","findings":"LLM能捕捉整体回答模式，但变异性低于人类；加入访谈内容可提升部分模型的回答多样性，精心设计的提示和低温度参数能增强对齐度。项目分析显示LLM在反向措辞题目上偏差较大，且难以重建问卷的心理测量结构。","reliability":"论文承认LLM在回答变异性、情感解读和心理测量保真度方面存在局限，且访谈内容的相关性比长度更重要，不同受访者间模型表现存在差异。","relevance":"该研究直接以真实人类调查数据为基准，评估访谈引导的LLM仿真回答的可靠性与偏差，与研究者关注的LLM人类仿真实验高度吻合，值得精读。","inspiration":"借鉴用个人深度访谈作为LLM输入来模拟特定个体问卷回答的设计，可迁移到消费者信心调查或通胀预期形成等经济金融场景。｜设计一个研究：以真实消费者的财务访谈文本作为处理，让LLM生成对未来通胀或就业的预期评分，结果变量为预期值，以实际调查数据（如密歇根消费者调查）作为对照基准。"}},{"id":"2505.21371","version":2,"title":"When Experimental Economics Meets Large Language Models: Evidence-based Tactics","zh_title":"当实验经济学遇上大语言模型：基于证据的实践策略","abstract":"Advancements in large language models (LLMs) have sparked a growing interest in measuring and understanding their behavior through experimental economics. However, there is still a lack of established guidelines for designing economic experiments for LLMs. Inspired by principles from experimental economics with insights from LLM research in artificial intelligence, we outline key considerations in the experimental design and implementation stage, and perform two sets of experiments to assess the impact of these considerations on LLMs' responses. Based on our findings, we discuss seven practical tactics for conducting experiments with LLMs. Our study enhances the design, replicability, and generalizability of LLM experiments, and broadens the scope of experimental economics in the digital age.","authors":["Shu Wang","Zijun Yao","Shuhuai Zhang","Jianuo Gai","Tracy Xiao Liu","Songfa Zhong"],"categories":["econ.GN"],"primary_category":"econ.GN","announce_type":"new","date":"2025-05-27","first_seen":"2025-05-27","revised_at":null,"abs_url":"https://arxiv.org/abs/2505.21371","pdf_url":"https://arxiv.org/pdf/2505.21371","source_feed":"backfill","score":7,"bucket":"pending","rubric_hits":["A2","A4","B4"],"tags":["LLM实验设计","方法论","实验经济学"],"reason":"论文探讨用实验经济学方法设计LLM实验，评估实验设计对LLM响应的影响，提供方…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:09","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":115,"question":"如何设计针对大语言模型的经济学实验，以及实验设计参数（如提示格式、对话类型、响应方式）如何影响LLM的行为表现？","design":"本文并非直接进行人类仿真，而是通过两个案例研究（预算决策和行为博弈），在GPT-4o、DeepSeek-V3、Llama-3.1-8B、Qwen2.5-7B四个模型上，系统操纵实验设计因素（如分配的角色、单轮/多轮对话、开放/封闭式回答），测量LLM的理性程度和偏好输出。","baseline":"无对照","findings":"分配的职业角色显著影响LLM的偏好，但不影响理性；单轮对话方式会降低Llama和Qwen的理性，但对GPT和DeepSeek无影响；限制为选择题而非开放式回答会降低Llama和Qwen的理性，并显著改变所有模型在约半数行为游戏中的输出。","reliability":"论文未讨论","relevance":"本文虽未直接进行人类仿真，但系统评估了LLM实验设计参数对行为输出的影响，为构建可靠的人类仿真实验提供了方法论基础，值得阅读原文以了解其提出的七条实用策略。","inspiration":"借鉴其系统操纵实验设计参数（角色分配、对话轮次、回答格式）来评估LLM行为稳健性的方法论，可迁移至政策公告预期形成的仿真研究，例如以LLM为被试、操纵公告措辞与信息呈现方式、测量通胀预期并对比真实调查数据。"}},{"id":"2505.17479","version":1,"title":"Twin-2K-500: A dataset for building digital twins of over 2,000 people based on their answers to over 500 questions","zh_title":"Twin-2K-500：基于2000余人对500余题回答构建数字孪生的数据集","abstract":"LLM-based digital twin simulation, where large language models are used to emulate individual human behavior, holds great promise for research in AI, social science, and digital experimentation. However, progress in this area has been hindered by the scarcity of real, individual-level datasets that are both large and publicly available. This lack of high-quality ground truth limits both the development and validation of digital twin methodologies. To address this gap, we introduce a large-scale, public dataset designed to capture a rich and holistic view of individual human behavior. We survey a representative sample of $N = 2,058$ participants (average 2.42 hours per person) in the US across four waves with 500 questions in total, covering a comprehensive battery of demographic, psychological, economic, personality, and cognitive measures, as well as replications of behavioral economics experiments and a pricing survey. The final wave repeats tasks from earlier waves to establish a test-retest accuracy baseline. Initial analyses suggest the data are of high quality and show promise for constructing digital twins that predict human behavior well at the individual and aggregate levels. By making the full dataset publicly available, we aim to establish a valuable testbed for the development and benchmarking of LLM-based persona simulations. Beyond LLM applications, due to its unique breadth and scale the dataset also enables broad social science research, including studies of cross-construct correlations and heterogeneous treatment effects.","authors":["Olivier Toubia","George Z. Gui","Tianyi Peng","Daniel J. Merlau","Ang Li","Haozhe Chen"],"categories":["cs.CY","cs.AI","cs.HC","econ.EM"],"primary_category":"cs.CY","announce_type":"new","date":"2025-05-23","first_seen":"2025-05-23","revised_at":null,"abs_url":"https://arxiv.org/abs/2505.17479","pdf_url":"https://arxiv.org/pdf/2505.17479","source_feed":"backfill","score":10,"bucket":"selected","rubric_hits":["A1","A2","A3","B1","B2"],"tags":["数字孪生","人类仿真","行为经济学"],"reason":"直接构建LLM数字孪生仿真个体行为，含真实人类对照数据，涉及行为经济学实验。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:08","error":null,"has_summary":true,"summary":{"generated_at":"2025-05-23","rank":2,"question":"构建并公开一个大规模、多维度的人类行为数据集，用于开发和验证基于LLM的数字孪生仿真。","design":"该研究不是仿真实验，而是数据集构建。对2058名美国代表性样本进行4轮调查，每名被试平均2.42小时，共500题，涵盖人口统计、心理、经济偏好、人格、认知测试，并复现了行为经济学实验（包括组间和组内设计）及定价调查。","baseline":"人类基准：2058名真实被试的问卷调查数据，包括行为经济学实验的复现结果，以及第四轮重复前几轮任务以建立重测信度基线。","findings":"数据质量良好：测量间相关性具有表面效度，复现了行为经济学文献中几乎所有已知结果，重测信度稳健。初步数字孪生预测在个体和聚合水平上表现良好，但具体准确率未在摘要中给出。","reliability":"论文未讨论失效条件与局限，但指出数据集可用于评估数字孪生的可靠性，且已有研究提示LLM仿真可能受提示架构、混杂变量和代表性偏差影响。","relevance":"高度相关。该数据集提供了大规模、多维度的人类行为基准，可直接用于评估LLM数字孪生在复现调查回答、实验行为和决策模式上的准确性，尤其包含经济学实验复现，适合研究者验证仿真可靠性及偏差条件。","inspiration":"借鉴其大规模多维度调查设计，在同一被试内复现行为经济学实验并建立重测信度基线，为仿真验证提供个体与聚合层面的对照基准｜可迁移到资产定价实验中的风险偏好与预期形成研究，检验LLM数字孪生能否复现真实投资者的风险态度和价格预期分布｜以真实投资者为被试，收集其风险偏好问卷、资产选择实验及市场预期数据作为基准，用LLM基于相同问卷生成数字孪生，比较两者在风险资产配置和预期回报估计上的个体与分布一致性"}},{"id":"2505.16188","version":2,"title":"SAE-SSV: Supervised Steering in Sparse Representation Spaces for Reliable Control of Language Models","zh_title":"SAE-SSV：稀疏表示空间中的监督引导以实现语言模型的可靠控制","abstract":"Large language models (LLMs) have demonstrated impressive capabilities in natural language understanding and generation, but controlling their behavior reliably remains challenging, especially in open-ended generation settings. This paper introduces a novel supervised steering approach that operates in sparse, interpretable representation spaces. We employ sparse autoencoders (SAEs) to obtain sparse latent representations that aim to disentangle semantic attributes from model activations. Then we train linear classifiers to identify a small subspace of task-relevant dimensions in latent representations. Finally, we learn supervised steering vectors constrained to this subspace, optimized to align with target behaviors. Experiments across sentiment, truthfulness, and political polarity steering tasks with multiple LLMs demonstrate that our supervised steering vectors achieve higher success rates with minimal degradation in generation quality compared to existing methods. Further analysis reveals that a notably small subspace is sufficient for effective steering, enabling more targeted and interpretable interventions. Our implementation is publicly available at https://github.com/Ineedanamehere/SAE-SSV.","authors":["Zirui He","Mingyu Jin","Bo Shen","Ali Payani","Yongfeng Zhang","Mengnan Du"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2025-05-22","first_seen":"2025-05-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2505.16188","pdf_url":"https://arxiv.org/pdf/2505.16188","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["模型控制","稀疏自编码器","可解释性"],"reason":"纯NLP能力评测，控制模型生成属性，不以人类行为为参照系","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:59:17","error":null,"has_summary":false,"summary":null},{"id":"2505.16147","version":2,"title":"Losing is for Cherishing: Data Valuation Based on Machine Unlearning and Shapley Value","zh_title":"失去是为了珍惜：基于机器遗忘和Shapley值的数据估值","abstract":"The proliferation of large models has intensified the need for efficient data valuation methods to quantify the contribution of individual data providers. Traditional approaches, such as game-theory-based Shapley value and influence-function-based techniques, face prohibitive computational costs or require access to full data and model training details, making them hardly achieve partial data valuation. To address this, we propose Unlearning Shapley, a novel framework that leverages machine unlearning to estimate data values efficiently. By unlearning target data from a pretrained model and measuring performance shifts on a reachable test set, our method computes Shapley values via Monte Carlo sampling, avoiding retraining and eliminating dependence on full data. Crucially, Unlearning Shapley supports both full and partial data valuation, making it scalable for large models (e.g., LLMs) and practical for data markets. Experiments on benchmark datasets and large-scale text corpora demonstrate that our approach matches the accuracy of state-of-the-art methods while reducing computational overhead by orders of magnitude. Further analysis confirms a strong correlation between estimated values and the true impact of data subsets, validating its reliability in real-world scenarios. This work bridges the gap between data valuation theory and practical deployment, offering a scalable, privacy-compliant solution for modern AI ecosystems.","authors":["Le Ma","Shirao Yang","Zihao Wang","Yinggui Wang","Lei Wang","Tao Wei","Kejun Zhang"],"categories":["cs.AI","cs.LG"],"primary_category":"cs.AI","announce_type":"new","date":"2025-05-22","first_seen":"2025-05-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2505.16147","pdf_url":"https://arxiv.org/pdf/2505.16147","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["数据估值","机器遗忘","Shapley值"],"reason":"纯数据估值方法研究，不涉及人类行为仿真或对照","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:59:30","error":null,"has_summary":false,"summary":null},{"id":"2505.14588","version":4,"title":"Generative AI at the Crossroads: Light Bulb, Dynamo, or Microscope?","zh_title":"生成式AI的十字路口：灯泡、发电机还是显微镜？","abstract":"With the advent of generative AI (genAI), the potential scope of artificial intelligence has increased dramatically, but the future effect of genAI on productivity remains uncertain. The effect of the technology on the innovation process is a crucial open question. Some inventions, such as the light bulb, temporarily raise productivity growth as adoption spreads, but the effect fades when the market is saturated; that is, the level of output per hour is permanently higher but the growth rate is not. In contrast, two types of technologies stand out as having longer-lived effects on productivity growth. First, there are technologies known as general-purpose technologies (GPTs). GPTs (1) are widely adopted, (2) spur abundant knock-on innovations (new goods and services, process efficiencies, and business reorganization), and (3) show continual improvement, refreshing this innovation cycle; the electric dynamo is an example. Second, there are inventions of methods of invention (IMIs). IMIs increase the efficiency of the research and development process via improvements to observation, analysis, communication, or organization; the compound microscope is an example. We show that GenAI has the characteristics of both a GPT and an IMI -- an encouraging sign that genAI will raise the \\textit{level} of productivity. Even so, genAI's contribution to productivity \\textit{growth} will depend on the speed with which that level is attained and, historically, integrating revolutionary technologies into the economy is a protracted process.","authors":["Martin Baily","David Byrne","Aidan Kane","Paul Soto"],"categories":["econ.GN"],"primary_category":"econ.GN","announce_type":"new","date":"2025-05-20","first_seen":"2025-05-20","revised_at":null,"abs_url":"https://arxiv.org/abs/2505.14588","pdf_url":"https://arxiv.org/pdf/2505.14588","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["生成式AI","生产率","技术分类"],"reason":"论文讨论生成式AI对生产率和创新的宏观影响，未涉及LLM仿真人类被试或行为对照。","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:58:40","error":null,"has_summary":false,"summary":null},{"id":"2505.12923","version":2,"title":"The Traitors: Deception and Trust in Multi-Agent Language Model Simulations","zh_title":"《叛徒》：多智能体语言模型仿真中的欺骗与信任","abstract":"As AI systems increasingly assume roles where trust and alignment with human values are essential, understanding when and why they engage in deception has become a critical research priority. We introduce The Traitors, a multi-agent simulation framework inspired by social deduction games, designed to probe deception, trust formation, and strategic communication among large language model (LLM) agents under asymmetric information. A minority of agents the traitors seek to mislead the majority, while the faithful must infer hidden identities through dialogue and reasoning. Our contributions are: (1) we ground the environment in formal frameworks from game theory, behavioral economics, and social cognition; (2) we develop a suite of evaluation metrics capturing deception success, trust dynamics, and collective inference quality; (3) we implement a fully autonomous simulation platform where LLMs reason over persistent memory and evolving social dynamics, with support for heterogeneous agent populations, specialized traits, and adaptive behaviors. Our initial experiments across DeepSeek-V3, GPT-4o-mini, and GPT-4o (10 runs per model) reveal a notable asymmetry: advanced models like GPT-4o demonstrate superior deceptive capabilities yet exhibit disproportionate vulnerability to others' falsehoods. This suggests deception skills may scale faster than detection abilities. Overall, The Traitors provides a focused, configurable testbed for investigating LLM behavior in socially nuanced interactions. We position this work as a contribution toward more rigorous research on deception mechanisms, alignment challenges, and the broader social reliability of AI systems.","authors":["Pedro M. P. Curvo"],"categories":["cs.AI","cs.MA"],"primary_category":"cs.AI","announce_type":"new","date":"2025-05-19","first_seen":"2025-05-19","revised_at":null,"abs_url":"https://arxiv.org/abs/2505.12923","pdf_url":"https://arxiv.org/pdf/2505.12923","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["多智能体仿真","欺骗检测","社会推理"],"reason":"多智能体社会模拟但无真实人类数据对照，属边界情形","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:32","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":199,"question":"在信息不对称的多智能体社交推理游戏中，大语言模型代理何时以及如何产生欺骗行为，其欺骗能力与检测能力如何随模型能力变化？","design":"构建名为“The Traitors”的多智能体仿真框架，让不同LLM（DeepSeek-V3、GPT-4o-mini、GPT-4o）扮演少数“叛徒”和多数“忠臣”，叛徒知晓身份并试图欺骗，忠臣通过对话和推理识别叛徒；测量欺骗成功率、信任动态和集体推理质量等指标。","baseline":"无对照","findings":"高级模型如GPT-4o表现出更强的欺骗能力，但对他人谎言的脆弱性也更高，表明欺骗技能可能比检测能力扩展得更快。该框架可作为研究LLM在社会互动中欺骗与信任动态的测试平台。","reliability":"论文强调实验仅为概念验证，规模有限，受计算资源约束，未进行大规模统计效力研究；未讨论框架在其他场景下的泛化局限。","relevance":"该研究属于LLM多智能体社会仿真，但无真实人类数据对照，不符合研究者对基准人类数据的要求；若关注欺骗行为仿真本身，可了解其框架设计，但与经济学实验或政策评估的直接关联较弱。","inspiration":"该框架通过角色分配和信息不对称设计来诱发LLM代理的欺骗行为，并测量欺骗成功率与信任动态，这种多智能体博弈设计可借鉴用于研究经济决策中的策略性信息操纵｜可迁移至金融市场中的内幕交易与信息扩散研究，例如模拟交易员在拥有私有信息时的欺骗性沟通与市场信任演化｜可设计LLM代理扮演交易员，其中部分代理获得内幕信息（处理组），测量其欺骗性报价行为与市场价格的偏离，并以真实市场微观结构数据（如订单流与价格波动）作为对照基准"}},{"id":"2505.10309","version":3,"title":"A large-scale evaluation of commonsense knowledge in humans and large language models","zh_title":"人类与大语言模型常识知识的大规模评估","abstract":"Commonsense knowledge, a major constituent of artificial intelligence (AI), is primarily evaluated in practice by human-prescribed ground-truth labels. An important, albeit implicit, assumption of these labels is that they accurately capture what any human would think, effectively treating human common sense as homogeneous. However, recent empirical work has shown that humans vary enormously in what they consider commonsensical; thus what appears self-evident to one benchmark designer may not be so to another. Here, we propose a method for assessing commonsense knowledge in AI, specifically in large language models (LLMs), that incorporates empirically observed heterogeneity among humans by measuring the correspondence between a model's judgment and that of a human population. We first find that, when treated as independent survey respondents, most LLMs remain below the human median in their individual commonsense competence. Second, when used as simulators of a hypothetical population, LLMs correlate with real humans only modestly in the extent to which they agree on the same set of statements. In both cases, smaller, open-weight models are surprisingly more competitive than larger, proprietary frontier models. Our evaluation framework, which ties commonsense knowledge to its cultural basis, contributes to the growing call for adapting AI models to human collectivities that possess different, often incompatible, social stocks of knowledge.","authors":["Tuan Dung Nguyen","Duncan J. Watts","Mark E. Whiting"],"categories":["cs.AI","cs.HC","cs.SI"],"primary_category":"cs.AI","announce_type":"new","date":"2025-05-15","first_seen":"2025-05-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2505.10309","pdf_url":"https://arxiv.org/pdf/2505.10309","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","人类对照","常识知识"],"reason":"将LLM作为独立调查受访者，与真实人类常识判断分布对照，评估仿真可靠性与异质性…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:08","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":36,"question":"如何将人类常识判断的异质性纳入评估，衡量大语言模型与人类群体在常识知识上的一致性？","design":"将35个LLM作为独立调查受访者，收集其对常识陈述的判断，并与2046名人类受访者的判断分布进行对比；同时用LLM生成硅样本群体，模拟人类群体对陈述的共识程度。","baseline":"2046名人类受访者对常识陈述的判断分布，包括个体间共识和群体共识。","findings":"多数LLM的常识能力低于人类中位数，且LLM模拟的群体共识与真实人类群体共识仅呈中等相关；小型开源模型的表现可与大型闭源模型竞争。","reliability":"论文指出常识知识具有文化依赖性，当前评估框架仅基于特定人类群体，可能不适用于其他社会文化背景；LLM模拟的群体在定性上与人类存在差异，如Gemini Pro 1.0过度关联常识与修辞性表达。","relevance":"该研究将LLM作为人类被试的替代品，与大规模真实人类数据对照，评估仿真可靠性与异质性，并指出仿真失效的文化条件，直接回应了研究者对LLM仿真实验的核心关切，值得精读。","inspiration":"该方法将LLM作为独立受访者，直接与大规模人类样本的判断分布进行对比，并生成硅样本模拟群体共识，可借鉴其个体-群体双层对照设计｜可迁移到经济预期形成研究，例如调查公众对通胀或政策公告的预期分布｜研究设计：以LLM作为被试，处理为不同措辞的央行声明，结果变量为通胀预期值，用密歇根大学消费者调查的真实预期分布数据作为对照基准"}},{"id":"2505.09938","version":2,"title":"Design and Evaluation of Generative Agent-based Platform for Human-Assistant Interaction Research: A Tale of 10 User Studies","zh_title":"基于生成式智能体的仿真平台设计与评估：10项用户研究的故事","abstract":"Designing and evaluating personalized and proactive assistant agents remains challenging due to the time, cost, and ethical concerns associated with human-in-the-loop experimentation. Existing Human-Computer Interaction (HCI) methods often require extensive physical setup and human participation, which introduces privacy concerns and limits scalability. Simulated environments offer a partial solution but are typically constrained by rule-based scenarios and still depend heavily on human input to guide interactions and interpret results. Recent advances in large language models (LLMs) have introduced the possibility of generative agents that can simulate realistic human behavior, reasoning, and social dynamics. However, their effectiveness in modeling human-assistant interactions remains largely unexplored. To address this gap, we present a generative agent-based simulation platform designed to simulate human-assistant interactions. We identify ten prior studies on assistant agents that span different aspects of interaction design and replicate these studies using our simulation platform. Our results show that fully simulated experiments using generative agents can approximate key aspects of human-assistant interactions. Based on these simulations, we are able to replicate the core conclusions of the original studies. Our work provides a scalable and cost-effective approach for studying assistant agent design without requiring live human subjects. Additional resources and project materials are available at https://dash-gidea.github.io/","authors":["Ziyi Xuan","Yiwen Wu","Xuhai Xu","Vinod Namboodiri","Mooi Choo Chuah","Yu Yang"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2025-05-15","first_seen":"2025-05-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2505.09938","pdf_url":"https://arxiv.org/pdf/2505.09938","source_feed":"backfill","score":7,"bucket":"pending","rubric_hits":["A1","A2","B1"],"tags":["LLM仿真","人机交互","用户研究复现"],"reason":"用LLM仿真人类与助手交互，复现10项用户研究并与真实人类数据对照，评估仿真可…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:08","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":71,"question":"基于LLM的生成式智能体能否有效模拟人类与助手交互，并复现真实用户研究的核心结论？","design":"使用LLM驱动的生成式智能体模拟人类被试，构建仿真平台GIDEA；选取10项涵盖主动协助、可中断性、自适应个性化等主题的真实用户研究，将其实验方案转化为结构化提示模板，在平台上复现实验，测量行为指标（如策略使用、接受率）和回复语义相似度。","baseline":"10项已发表用户研究中的真实人类被试数据，包括行为模式和定性描述。","findings":"生成式智能体仿真能够近似人类-助手交互的关键行为模式，并成功复现原始研究的核心结论。该平台提供了一种可扩展且低成本的研究方法，无需招募真实人类被试。","reliability":"论文未讨论","relevance":"该研究直接用LLM仿真人类被试复现10项用户研究，并与真实人类数据对照，属于你关注的人类仿真实验，且包含经济学实验和政策评估之外的HCI场景，值得阅读原文以评估其仿真保真度和偏差。","inspiration":"该方法将真实用户研究方案转化为结构化提示模板，在LLM仿真平台上复现实验并测量行为指标与语义相似度，为经济实验的仿真验证提供了可操作的流程｜可迁移至消费者跨期选择实验，检验LLM仿真能否复现真实被试的时间偏好异质性与框架效应｜以LLM智能体为被试，施加不同时间折现的决策框架处理，测量选择延迟与折现率，对照真实实验室实验数据验证仿真保真度"}},{"id":"2505.09396","version":2,"title":"The Influence of Human-inspired Agentic Sophistication in LLM-driven Strategic Reasoners","zh_title":"人类启发的智能体复杂度对LLM驱动战略推理者的影响","abstract":"The rapid rise of large language models (LLMs) has shifted artificial intelligence (AI) research toward agentic systems, motivating the use of weaker and more flexible notions of agency. However, this shift raises key questions about the extent to which LLM-based agents replicate human strategic reasoning, particularly in game-theoretic settings. In this context, we examine the role of agentic sophistication in shaping artificial reasoners' performance by evaluating three agent designs: a simple game-theoretic model, an unstructured LLM-as-agent model, and an LLM integrated into a traditional agentic framework. Using guessing games as a testbed, we benchmarked these agents against human participants across general reasoning patterns and individual role-based objectives. Furthermore, we introduced obfuscated game scenarios to assess agents' ability to generalise beyond training distributions. Our analysis, covering over 2000 reasoning samples across 25 agent configurations, shows that human-inspired cognitive structures can enhance LLM agents' alignment with human strategic behaviour. Still, the relationship between agentic design complexity and human-likeness is non-linear, highlighting a critical dependence on underlying LLM capabilities and suggesting limits to simple architectural augmentation.","authors":["Vince Trencsenyi","Agnieszka Mensfelt","Kostas Stathis"],"categories":["cs.AI","cs.MA"],"primary_category":"cs.AI","announce_type":"new","date":"2025-05-14","first_seen":"2025-05-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2505.09396","pdf_url":"https://arxiv.org/pdf/2505.09396","source_feed":"backfill","score":8,"bucket":"selected","rubric_hits":["A1","A2","B1","B2"],"tags":["LLM仿真","博弈实验","人类对照"],"reason":"用LLM代理模拟人类策略推理，并与真实人类数据对照，评估对齐程度与失效条件。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:07","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":51,"question":"在博弈论猜测游戏中，LLM智能体的设计复杂度（agentic sophistication）如何影响其与人类策略推理行为的对齐程度？","design":"使用三种智能体设计（简单博弈论模型EWA、非结构化LLM-as-agent、LLM集成传统智能体框架），在双人猜测游戏中与人类被试进行基准对比，并引入混淆游戏场景测试分布外泛化能力，测量推理模式与个体角色目标的对齐度。","baseline":"人类数据集，按子群体划分，包含人类在猜测游戏中的策略推理行为。","findings":"受人类启发的认知结构能增强LLM智能体与人类策略行为的一致性，但智能体设计复杂度与人类相似度之间呈非线性关系，高度依赖底层LLM能力，简单的架构增强存在局限。","reliability":"论文指出LLM的黑箱特性带来可复现性、可解释性和验证挑战，混淆场景测试旨在缓解训练数据暴露偏差，但承认LLM驱动应用仍缺乏系统严格的验证机制。","relevance":"直接回应了用LLM仿真人类策略推理的核心问题，提供了与真实人类数据的对照基准，并批判性地揭示了设计复杂度与人类相似度的非线性关系及失效条件，值得精读。","inspiration":"借鉴其用不同复杂度智能体设计与人类基准对比的框架，以及引入混淆场景测试分布外泛化的稳健性检验方法｜可迁移到资产定价实验，研究投资者在策略性猜测市场走势时的推理行为｜以LLM智能体为被试，处理为不同复杂度的智能体设计（如简单启发式、EWA模型、非结构化LLM），结果变量为价格预测偏差，用真实人类资产定价实验数据做对照"}},{"id":"2505.07457","version":1,"title":"Can Generative AI agents behave like humans? Evidence from laboratory market experiments","zh_title":"生成式AI智能体能像人类一样行为吗？来自实验室市场实验的证据","abstract":"We explore the potential of Large Language Models (LLMs) to replicate human behavior in economic market experiments. Compared to previous studies, we focus on dynamic feedback between LLM agents: the decisions of each LLM impact the market price at the current step, and so affect the decisions of the other LLMs at the next step. We compare LLM behavior to market dynamics observed in laboratory settings and assess their alignment with human participants' behavior. Our findings indicate that LLMs do not adhere strictly to rational expectations, displaying instead bounded rationality, similarly to human participants. Providing a minimal context window i.e. memory of three previous time steps, combined with a high variability setting capturing response heterogeneity, allows LLMs to replicate broad trends seen in human experiments, such as the distinction between positive and negative feedback markets. However, differences remain at a granular level--LLMs exhibit less heterogeneity in behavior than humans. These results suggest that LLMs hold promise as tools for simulating realistic human behavior in economic contexts, though further research is needed to refine their accuracy and increase behavioral diversity.","authors":["R. Maria del Rio-Chanona","Marco Pangallo","Cars Hommes"],"categories":["econ.GN","cs.AI"],"primary_category":"econ.GN","announce_type":"new","date":"2025-05-12","first_seen":"2025-05-12","revised_at":null,"abs_url":"https://arxiv.org/abs/2505.07457","pdf_url":"https://arxiv.org/pdf/2505.07457","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM仿真","市场实验","人类行为对照"],"reason":"用LLM复现市场实验并与人类数据对照，直接命中核心判据。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:05","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":66,"question":"大语言模型能否在动态市场实验中复现人类行为，特别是正负反馈市场中的价格动态？","design":"使用GPT-3.5和GPT-4作为智能体，模拟实验室市场实验，操控上下文窗口（记忆长度）和响应变异性（温度参数），观察市场价格动态和个体策略。","baseline":"对照Heemeijer等人(2009)等实验室市场实验的人类参与者数据，比较正负反馈市场中的价格收敛模式和波动特征。","findings":"在至少3步记忆和高响应变异性下，LLM市场能复现正负反馈市场的宏观差异，如负反馈市场快速振荡收敛、正反馈市场缓慢收敛；但LLM行为异质性低于人类，且不严格遵循理性预期，表现出有限理性。","reliability":"LLM在细粒度行为上异质性不足，与人类仍有差距；研究仅基于特定市场实验范式，泛化性待验证；模型类型和参数设置对结果敏感，需进一步优化以提升行为多样性和准确性。","relevance":"直接命中研究者关注的核心：用LLM复现经济实验并与人类基准对照，评估仿真可靠性与偏差，且包含动态交互和批判性发现，值得精读原文。","inspiration":"借鉴之处在于通过操控LLM的上下文窗口长度和响应变异性来模拟有限理性，并与人类实验基准对照，检验宏观市场动态的复现能力｜可迁移到资产定价实验，研究正负反馈机制下价格泡沫的形成与破裂｜设计一个LLM模拟的资产市场实验，将被试分为GPT-4智能体，处理为不同反馈结构（正反馈如追涨杀跌 vs. 负反馈如均值回归），结果变量为价格偏离基本面的程度和泡沫持续时间，对照真实人类实验数据（如Smith等人1988年的泡沫实验）"}},{"id":"2505.07653","version":2,"title":"JobHop: A Large-Scale Dataset of Career Trajectories","zh_title":"JobHop：一个大规模职业轨迹数据集","abstract":"Understanding labor market dynamics is essential for policymakers, employers, and job seekers. However, comprehensive datasets that capture real-world career trajectories are scarce. In this paper, we introduce JobHop, a large-scale public dataset derived from anonymized resumes provided by VDAB, the public employment service in Flanders, Belgium. Utilizing Large Language Models (LLMs), we process unstructured resume data to extract structured career information, which is then normalized to standardized ESCO occupation codes using a multi-label classification model. This results in a rich dataset of over 1.67 million work experiences, extracted from and grouped into more than 361,000 user resumes and mapped to standardized ESCO occupation codes, offering valuable insights into real-world occupational transitions. This dataset enables diverse applications, such as analyzing labor market mobility, job stability, and the effects of career breaks on occupational transitions. It also supports career path prediction and other data-driven decision-making processes. To illustrate its potential, we explore key dataset characteristics, including job distributions, career breaks, and job transitions, demonstrating its value for advancing labor market research.","authors":["Iman Johary","Raphael Romero","Alexandru C. Mara","Tijl De Bie"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2025-05-12","first_seen":"2025-05-12","revised_at":null,"abs_url":"https://arxiv.org/abs/2505.07653","pdf_url":"https://arxiv.org/pdf/2505.07653","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["信息抽取","劳动力市场","数据集"],"reason":"用LLM从简历提取结构化数据，属于信息抽取，不涉及人类行为仿真或对照。","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:59:26","error":null,"has_summary":false,"summary":null},{"id":"2505.06702","version":1,"title":"Do Language Model Agents Align with Humans in Rating Visualizations? An Empirical Study","zh_title":"语言模型代理在可视化评分中与人类对齐吗？一项实证研究","abstract":"Large language models encode knowledge in various domains and demonstrate the ability to understand visualizations. They may also capture visualization design knowledge and potentially help reduce the cost of formative studies. However, it remains a question whether large language models are capable of predicting human feedback on visualizations. To investigate this question, we conducted three studies to examine whether large model-based agents can simulate human ratings in visualization tasks. The first study, replicating a published study involving human subjects, shows agents are promising in conducting human-like reasoning and rating, and its result guides the subsequent experimental design. The second study repeated six human-subject studies reported in literature on subjective ratings, but replacing human participants with agents. Consulting with five human experts, this study demonstrates that the alignment of agent ratings with human ratings positively correlates with the confidence levels of the experts before the experiments. The third study tests commonly used techniques for enhancing agents, including preprocessing visual and textual inputs, and knowledge injection. The results reveal the issues of these techniques in robustness and potential induction of biases. The three studies indicate that language model-based agents can potentially simulate human ratings in visualization experiments, provided that they are guided by high-confidence hypotheses from expert evaluators. Additionally, we demonstrate the usage scenario of swiftly evaluating prototypes with agents. We discuss insights and future directions for evaluating and improving the alignment of agent ratings with human ratings. We note that simulation may only serve as complements and cannot replace user studies.","authors":["Zekai Shao","Yi Shan","Yixuan He","Yuxuan Yao","Junhong Wang","Xiaolong","Zhang","Yu Zhang","Siming Chen"],"categories":["cs.HC"],"primary_category":"cs.HC","announce_type":"new","date":"2025-05-10","first_seen":"2025-05-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2505.06702","pdf_url":"https://arxiv.org/pdf/2505.06702","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","人类数据对照","可视化评估"],"reason":"用LLM代理模拟人类对可视化的评分，并与真实人类数据对照，评估对齐度与偏差，直…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:05","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":67,"question":"大语言模型代理能否在可视化评分任务中模拟人类评分？","design":"使用GPT-4V等LLM代理，复现已发表的人类被试可视化实验，让代理对可视化设计进行主观评分（如易用性、信心水平），并测量代理评分与人类评分的一致性。","baseline":"对照的真实人类数据来自已发表的六项人类被试研究，包括时间序列可视化实验等，数据来源于Open Science Framework。","findings":"LLM代理能模拟人类推理并给出类似人类的评分，但无法模拟多样化用户画像；代理与人类评分的一致性程度与专家实验前信心水平正相关。","reliability":"代理评分仅能作为人类用户研究的补充，不能替代；增强技术（如知识注入）可能引入偏差，且代理在鲁棒性上存在问题。","relevance":"该研究直接以真实人类数据为基准，评估LLM代理在可视化评分任务中的仿真可靠性，并讨论了失效条件与偏差，符合研究者对经济学实验和政策评估场景的批判性关注。","inspiration":"该方法借鉴了用LLM代理复现已发表人类实验并直接对比真实人类评分一致性的设计，以及通过专家信心水平等指标检验仿真可靠性的思路。｜可迁移到消费者对金融产品信息披露的主观评价实验，如研究简化版风险提示是否提升理解度。｜以LLM代理作为被试，施加不同格式的风险披露文本处理，测量代理对产品风险的理解评分，并与真实消费者调查数据（如CFPB的金融素养调查）进行一致性对比。"}},{"id":"2507.18639","version":1,"title":"People Are Highly Cooperative with Large Language Models, Especially When Communication Is Possible or Following Human Interaction","zh_title":"人们与大型语言模型高度合作，尤其在可沟通或继人类互动之后","abstract":"Machines driven by large language models (LLMs) have the potential to augment humans across various tasks, a development with profound implications for business settings where effective communication, collaboration, and stakeholder trust are paramount. To explore how interacting with an LLM instead of a human might shift cooperative behavior in such settings, we used the Prisoner's Dilemma game -- a surrogate of several real-world managerial and economic scenarios. In Experiment 1 (N=100), participants engaged in a thirty-round repeated game against a human, a classic bot, and an LLM (GPT, in real-time). In Experiment 2 (N=192), participants played a one-shot game against a human or an LLM, with half of them allowed to communicate with their opponent, enabling LLMs to leverage a key advantage over older-generation machines. Cooperation rates with LLMs -- while lower by approximately 10-15 percentage points compared to interactions with human opponents -- were nonetheless high. This finding was particularly notable in Experiment 2, where the psychological cost of selfish behavior was reduced. Although allowing communication about cooperation did not close the human-machine behavioral gap, it increased the likelihood of cooperation with both humans and LLMs equally (by 88%), which is particularly surprising for LLMs given their non-human nature and the assumption that people might be less receptive to cooperating with machines compared to human counterparts. Additionally, cooperation with LLMs was higher following prior interaction with humans, suggesting a spillover effect in cooperative behavior. Our findings validate the (careful) use of LLMs by businesses in settings that have a cooperative component.","authors":["Paweł Niszczota","Tomasz Grzegorczyk","Alexander Pastukhov"],"categories":["cs.HC","cs.CL","cs.CY","econ.GN"],"primary_category":"cs.HC","announce_type":"new","date":"2025-05-10","first_seen":"2025-05-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2507.18639","pdf_url":"https://arxiv.org/pdf/2507.18639","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2"],"tags":["LLM仿真","行为博弈","人机合作"],"reason":"用LLM替代人类被试进行囚徒困境博弈，并与真实人类行为对照，评估合作行为差异与…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:18","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":27,"question":"当对手是大语言模型而非人类时，人们在囚徒困境中的合作行为会发生怎样的变化？","design":"实验1（N=100）采用被试内设计，每人与人类、经典机器人、GPT实时对战各30轮重复囚徒困境；实验2（N=192）采用被试间设计，在单次囚徒困境中对手为人类或LLM，且半数被试可在决策前与对手沟通。结果变量为合作率。","baseline":"以真实人类作为对手时的合作行为作为对照基准。","findings":"与LLM的合作率虽比与人类对手低约10–15个百分点，但仍处于较高水平；允许沟通使与人类和LLM的合作率均提高88%，且与人类互动后与LLM的合作率更高，存在溢出效应。","reliability":"论文未讨论","relevance":"该研究直接以LLM替代人类被试进行囚徒困境博弈，并与真实人类行为对照，评估合作行为差异及沟通、溢出效应，高度契合研究者对LLM仿真人类行为可靠性与偏差的关注，值得精读原文。","inspiration":"采用被试内设计直接对比同一参与者在面对人类、传统机器人和LLM时的行为变化，有效控制个体差异，值得借鉴。｜可迁移至经济金融中的信任与履约行为研究，如线上借贷平台中借款人对人工审核员与AI审核员的还款承诺差异。｜以真实借款人为被试，随机分配其与人类审核员或LLM审核员沟通还款计划，测量其后续实际还款率，并以平台历史人工审核还款数据作为对照基准。"}},{"id":"2505.05863","version":1,"title":"Evolutionary ecology of words","zh_title":"词汇的进化生态学","abstract":"We propose a model for the evolutionary ecology of words as one attempt to extend evolutionary game theory and agent-based models by utilizing the rich linguistic expressions of Large Language Models (LLMs). Our model enables the emergence and evolution of diverse and infinite options for interactions among agents. Within the population, each agent possesses a short word (or phrase) generated by an LLM and moves within a spatial environment. When agents become adjacent, the outcome of their interaction is determined by the LLM based on the relationship between their words, with the loser's word being replaced by the winner's. Word mutations, also based on LLM outputs, may occur. We conducted preliminary experiments assuming that ``strong animal species\" would survive. The results showed that from an initial population consisting of well-known species, many species emerged both gradually and in a punctuated equilibrium manner. Each trial demonstrated the unique evolution of diverse populations, with one type of large species becoming dominant, such as terrestrial animals, marine life, or extinct species, which were ecologically specialized and adapted ones across diverse extreme habitats. We also conducted a long-term experiment with a large population, demonstrating the emergence and coexistence of diverse species.","authors":["Reiji Suzuki","Takaya Arita"],"categories":["q-bio.PE","cs.AI","cs.CL"],"primary_category":"q-bio.PE","announce_type":"new","date":"2025-05-09","first_seen":"2025-05-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2505.05863","pdf_url":"https://arxiv.org/pdf/2505.05863","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","进化博弈","语言模型"],"reason":"纯多智能体生态模拟，无人类行为对照，不涉及人类被试仿真。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:32","error":null,"has_summary":false,"summary":null},{"id":"2505.00036","version":1,"title":"A Framework to Assess the Persuasion Risks Large Language Model Chatbots Pose to Democratic Societies","zh_title":"评估大语言模型聊天机器人对民主社会说服风险的框架","abstract":"In recent years, significant concern has emerged regarding the potential threat that Large Language Models (LLMs) pose to democratic societies through their persuasive capabilities. We expand upon existing research by conducting two survey experiments and a real-world simulation exercise to determine whether it is more cost effective to persuade a large number of voters using LLM chatbots compared to standard political campaign practice, taking into account both the \"receive\" and \"accept\" steps in the persuasion process (Zaller 1992). These experiments improve upon previous work by assessing extended interactions between humans and LLMs (instead of using single-shot interactions) and by assessing both short- and long-run persuasive effects (rather than simply asking users to rate the persuasiveness of LLM-produced content). In two survey experiments (N = 10,417) across three distinct political domains, we find that while LLMs are about as persuasive as actual campaign ads once voters are exposed to them, political persuasion in the real-world depends on both exposure to a persuasive message and its impact conditional on exposure. Through simulations based on real-world parameters, we estimate that LLM-based persuasion costs between \\$48-\\$74 per persuaded voter compared to \\$100 for traditional campaign methods, when accounting for the costs of exposure. However, it is currently much easier to scale traditional campaign persuasion methods than LLM-based persuasion. While LLMs do not currently appear to have substantially greater potential for large-scale political persuasion than existing non-LLM methods, this may change as LLM capabilities continue to improve and it becomes easier to scalably encourage exposure to persuasive LLMs.","authors":["Zhongren Chen","Joshua Kalla","Quan Le","Shinpei Nakamura-Sakai","Jasjeet Sekhon","Ruixiao Wang"],"categories":["cs.CL","cs.CY"],"primary_category":"cs.CL","announce_type":"new","date":"2025-04-29","first_seen":"2025-04-29","revised_at":null,"abs_url":"https://arxiv.org/abs/2505.00036","pdf_url":"https://arxiv.org/pdf/2505.00036","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2","B4"],"tags":["LLM仿真","政治说服","人类数据对照"],"reason":"用LLM聊天机器人替代人类选民进行说服实验，有真实人类调查数据对照，涉及政治说…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:05","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":22,"question":"在政治说服中，使用LLM聊天机器人与传统竞选广告相比，在考虑曝光和接受两个步骤后，是否更具成本效益？","design":"本研究并非用LLM仿真人类被试，而是通过两项调查实验和真实世界模拟，将人类参与者随机分配到安慰剂、真人视频说服、AI聊天机器人（自称人类）和AI聊天机器人（自称AI）四种条件，测量其对移民政策的支持态度变化，并基于真实参数模拟成本效益。","baseline":"真人视频说服条件（3分钟教师分享个人理由的视频）和传统竞选广告方法（如电视广告）作为对照基准。","findings":"在调查实验中，LLM聊天机器人与真人视频的说服效果在短期和五周后均无显著差异；但模拟显示，考虑曝光成本后，LLM说服每位选民的成本为48-74美元，低于传统方法的100美元，然而目前传统方法更易规模化。","reliability":"论文指出当前LLM在规模化政治说服方面并不比现有非LLM方法有显著优势，但随LLM能力提升和更容易鼓励曝光，情况可能改变；研究局限包括真人说服条件为3分钟视频，长于典型广告，且未充分解决曝光环节的规模化难题。","relevance":"该研究直接使用LLM与人类进行交互式说服实验，并与真实人类说服效果对照，涉及政治领域成本效益评估，同时讨论了规模化限制，高度契合研究者对LLM仿真人类行为、基准对照和失效条件的兴趣，值得精读原文。","inspiration":"该方法将LLM聊天机器人作为交互式说服工具，与真人视频说服和传统广告进行随机对照实验，并追踪短期与五周后的态度变化，值得借鉴其多臂对照和纵向测量设计｜可迁移到消费者金融决策场景，如评估LLM理财顾问对投资偏好或退休储蓄选择的影响｜以真实投资者为被试，随机分配至LLM聊天机器人理财建议、人类理财顾问视频和纯文本说明书三组，结果变量为风险资产配置比例和储蓄率变化，以历史调查数据或银行实际客户行为作为基准对照"}},{"id":"2504.20628","version":1,"title":"Cognitive maps are generative programs","zh_title":"认知地图是生成式程序","abstract":"Making sense of the world and acting in it relies on building simplified mental representations that abstract away aspects of reality. This principle of cognitive mapping is universal to agents with limited resources. Living organisms, people, and algorithms all face the problem of forming functional representations of their world under various computing constraints. In this work, we explore the hypothesis that human resource-efficient planning may arise from representing the world as predictably structured. Building on the metaphor of concepts as programs, we propose that cognitive maps can take the form of generative programs that exploit predictability and redundancy, in contrast to directly encoding spatial layouts. We use a behavioral experiment to show that people who navigate in structured spaces rely on modular planning strategies that align with programmatic map representations. We describe a computational model that predicts human behavior in a variety of structured scenarios. This model infers a small distribution over possible programmatic cognitive maps conditioned on human prior knowledge of the world, and uses this distribution to generate resource-efficient plans. Our models leverages a Large Language Model as an embedding of human priors, implicitly learned through training on a vast corpus of human data. Our model demonstrates improved computational efficiency, requires drastically less memory, and outperforms unstructured planning algorithms with cognitive constraints at predicting human behavior, suggesting that human planning strategies rely on programmatic cognitive maps.","authors":["Marta Kryven","Cole Wyeth","Aidan Curtis","Kevin Ellis"],"categories":["cs.AI","cs.ET"],"primary_category":"cs.AI","announce_type":"new","date":"2025-04-29","first_seen":"2025-04-29","revised_at":null,"abs_url":"https://arxiv.org/abs/2504.20628","pdf_url":"https://arxiv.org/pdf/2504.20628","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["认知建模","人类行为预测","LLM先验嵌入"],"reason":"用LLM嵌入人类先验模拟导航行为，有行为实验对照，但LLM非被试替代，属社会模…","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:59:19","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":116,"question":"人类在结构化空间中的资源高效规划是否依赖于将世界表示为可预测的生成式程序（即程序化认知地图）？","design":"本研究并非用LLM替代人类被试的仿真实验，而是通过行为实验（迷宫搜索任务）收集30名人类被试在部分可观测结构化迷宫中的导航行为，并构建计算模型（利用LLM嵌入人类先验知识）来预测人类规划行为。","baseline":"对照的真实人类数据为30名Prolific平台招募的英语母语被试在21个生成式结构化迷宫中的点击导航行为，以及认知反射测试（CRT）和事后问卷。","findings":"人类在结构化迷宫中采用模块化规划策略，其行为与程序化认知地图模型预测一致；该模型比无结构规划算法（含认知约束）更准确预测人类行为，且计算效率更高、内存需求更少。","reliability":"论文未讨论","relevance":"该研究利用LLM嵌入人类先验知识来预测行为，并有真实人类实验对照，但核心是认知地图的计算模型，并非将LLM作为人类被试的替代品进行仿真实验，与研究者关注的LLM仿真替代方向有偏差，但可作为LLM用于行为预测的参考。","inspiration":"该研究利用LLM嵌入人类先验知识构建程序化认知地图模型，通过行为实验收集人类导航数据作为基准，验证模型预测能力｜可迁移至经济金融中的消费者搜索与决策行为研究，如在线购物平台上的产品筛选与购买决策｜设计：招募人类被试在模拟电商平台完成产品搜索任务，以LLM嵌入消费者先验偏好构建预测模型，结果变量为点击路径和最终购买选择，以真实人类行为数据作为对照基准"}},{"id":"2504.17993","version":2,"title":"Improving Language Model Personas via Rationalization with Psychological Scaffolds","zh_title":"通过心理支架合理化改进语言模型角色","abstract":"Language models prompted with a user description or persona are being used to predict the user's preferences and opinions. However, existing approaches to building personas mostly rely on a user's demographic attributes and/or prior judgments, but not on any underlying reasoning behind a user's judgments. We introduce PB&J (Psychology of Behavior and Judgments), a framework that improves LM personas by incorporating potential rationales for why the user could have made a certain judgment. Our rationales are generated by a language model to explicitly reason about a user's behavior on the basis of their experiences, personality traits, or beliefs. Our method employs psychological scaffolds: structured frameworks such as the Big 5 Personality Traits or Primal World Beliefs to help ground the generated rationales in existing theories. Experiments on public opinion and movie preference prediction tasks demonstrate that language model personas augmented with PB&J rationales consistently outperform personas conditioned only on user demographics and / or judgments, including those that use a model's default chain-of-thought, which is not grounded in psychological theories. Additionally, our PB&J personas perform competitively with those using human-written rationales, suggesting the potential of synthetic rationales guided by existing theories.","authors":["Brihi Joshi","Xiang Ren","Swabha Swayamdipta","Rik Koncel-Kedziorski","Tim Paek"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2025-04-25","first_seen":"2025-04-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2504.17993","pdf_url":"https://arxiv.org/pdf/2504.17993","source_feed":"backfill","score":7,"bucket":"pending","rubric_hits":["A1","B1","B4"],"tags":["LLM角色仿真","人类偏好预测","心理支架"],"reason":"用LLM模拟用户偏好预测，有人类数据对照，但侧重persona构建而非群体仿真…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:03","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":87,"question":"如何利用心理学理论生成用户判断的合理化解释，以提升语言模型模拟用户偏好和意见的准确性？","design":"本研究并非群体仿真实验，而是提出PB&J框架：给定用户人口统计属性和先前的判断，利用大语言模型基于心理学脚手架（如大五人格、原始世界信念）生成用户判断的潜在理由，构建更丰富的用户画像提示，然后在OpinionQA和MovieLens数据集上预测用户对测试问题的回答或电影评分，评估预测准确率。","baseline":"使用OpinionQA（750名用户，10个测试问题）和MovieLens（100名用户，10部测试电影）中的真实用户回答和评分作为对照基准。","findings":"加入PB&J生成理由的用户画像在预测用户意见和偏好上显著优于仅使用人口统计和/或先前判断的画像，也优于无心理学理论支撑的链式思维推理。PB&J生成的合成理由与人类撰写的理由性能接近，表明基于现有理论生成的理由具有潜力。","reliability":"论文承认生成的合成理由可能并不反映用户判断的真实推理过程，但因其可能包含真实行为推理的标记而具有合理性；此外，实验仅在两个数据集上进行，且人类撰写理由的对比实验规模较小。","relevance":"该研究直接涉及用LLM模拟个体用户偏好预测，有真实人类数据对照，并探讨了仿真方法的改进与局限，与研究者关注的LLM人类仿真实验高度相关，值得阅读原文以了解心理学理论如何提升仿真可靠性。","inspiration":"该方法通过心理学理论引导LLM生成用户判断的合理化理由，以丰富用户画像，提升偏好预测准确性，这种利用理论驱动的提示工程增强仿真保真度的做法值得借鉴｜可迁移到消费者金融决策仿真，如利用大五人格或风险态度理论生成理由，预测个体对金融产品的选择或风险偏好｜以真实消费者金融调查数据为基准，用LLM基于人口统计和先前金融行为生成带心理学理由的画像，预测测试集上的产品选择，对比仅用人口统计的基线，评估仿真准确性"}},{"id":"2504.19940","version":2,"title":"Assessing the Potential of Generative Agents in Crowdsourced Fact-Checking","zh_title":"评估生成式智能体在众包事实核查中的潜力","abstract":"The growing spread of online misinformation has created an urgent need for scalable, reliable fact-checking solutions. Crowdsourced fact-checking - where non-experts evaluate claim veracity - offers a cost-effective alternative to expert verification, despite concerns about variability in quality and bias. Encouraged by promising results in certain contexts, major platforms such as X (formerly Twitter), Facebook, and Instagram have begun shifting from centralized moderation to decentralized, crowd-based approaches. In parallel, advances in Large Language Models (LLMs) have shown strong performance across core fact-checking tasks, including claim detection and evidence evaluation. However, their potential role in crowdsourced workflows remains unexplored. This paper investigates whether LLM-powered generative agents - autonomous entities that emulate human behavior and decision-making - can meaningfully contribute to fact-checking tasks traditionally reserved for human crowds. Using the protocol of La Barbera et al. (2024), we simulate crowds of generative agents with diverse demographic and ideological profiles. Agents retrieve evidence, assess claims along multiple quality dimensions, and issue final veracity judgments. Our results show that agent crowds outperform human crowds in truthfulness classification, exhibit higher internal consistency, and show reduced susceptibility to social and cognitive biases. Compared to humans, agents rely more systematically on informative criteria such as Accuracy, Precision, and Informativeness, suggesting a more structured decision-making process. Overall, our findings highlight the potential of generative agents as scalable, consistent, and less biased contributors to crowd-based fact-checking systems.","authors":["Luigia Costabile","Gian Marco Orlando","Valerio La Gatta","Vincenzo Moscato"],"categories":["cs.CL","cs.AI","cs.MA"],"primary_category":"cs.CL","announce_type":"new","date":"2025-04-24","first_seen":"2025-04-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2504.19940","pdf_url":"https://arxiv.org/pdf/2504.19940","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2","B4"],"tags":["LLM仿真","众包事实核查","人类行为对照"],"reason":"用LLM智能体模拟人群事实核查，与真实人类数据对照，评估偏差与一致性，直接命中…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:04","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":31,"question":"生成式智能体在众包事实核查中的表现能否达到或超过人类众包？","design":"使用LLM驱动的生成式智能体模拟具有不同人口统计和意识形态特征的人群，复现La Barbera et al. (2024)的实验协议：智能体先选择证据，再填写结构化问卷对声明进行多维度评分（准确性、无偏性等），最后给出真实性判断。","baseline":"La Barbera et al. (2024)中人类众包参与者对相同声明的标注数据。","findings":"智能体众包在真实性分类上优于人类众包，表现出更高的内部一致性，且受社会和认知偏差的影响更小。智能体更系统地依赖准确性、精确性和信息量等标准，决策过程更结构化。","reliability":"论文未讨论","relevance":"该研究直接用LLM智能体模拟人类众包事实核查，并与真实人类数据对照，评估偏差与一致性，高度契合研究者对LLM人类仿真实验的关注，值得精读原文。","inspiration":"该方法借鉴了用LLM智能体复现人类实验协议并直接对比真实人类数据的做法，通过让智能体遵循相同的任务流程（选择证据、填写问卷、给出判断）来评估仿真偏差与一致性。｜可迁移到经济金融中的信贷审批歧视研究，模拟不同人口特征的贷款审批决策。｜设计：用LLM智能体模拟不同种族、性别、收入的贷款申请人，处理为呈现相同的贷款申请材料，结果变量为审批结果和理由，对照真实银行信贷审批数据或人类实验数据。"}},{"id":"2504.11671","version":4,"title":"Computational Basis of LLM's Decision Making in Social Simulation","zh_title":"LLM在社会仿真中决策的计算基础","abstract":"Large language models (LLMs) increasingly serve as human-like decision-making agents in social science and applied settings. These LLM-agents are typically assigned human-like characters and placed in real-life contexts. However, how these characters and contexts shape an LLM's behavior remains underexplored. This study proposes and tests methods for probing, quantifying, and modifying an LLM's internal representations in a Dictator Game, a classic behavioral experiment on fairness and prosocial behavior. We extract ``vectors of variable variations'' (e.g., ``male'' to ``female'') from the LLM's internal state. Manipulating these vectors during the model's inference can substantially alter how those variables relate to the model's decision-making. This approach offers a principled way to study and regulate how social concepts can be encoded and engineered within transformer-based models, with implications for alignment, debiasing, and designing AI agents for social simulations in both academic and commercial applications, strengthening sociological theory and measurement.","authors":["Ji Ma"],"categories":["cs.AI","cs.CY","cs.LG","econ.GN"],"primary_category":"cs.AI","announce_type":"new","date":"2025-04-16","first_seen":"2025-04-16","revised_at":null,"abs_url":"https://arxiv.org/abs/2504.11671","pdf_url":"https://arxiv.org/pdf/2504.11671","source_feed":"backfill","score":7,"bucket":"pending","rubric_hits":["A1","A3","B2"],"tags":["LLM社会仿真","独裁者博弈","表征操控"],"reason":"用LLM在独裁者博弈中模拟人类决策，并操控内部表征改变行为，属于社会仿真且涉及…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:03","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":117,"question":"LLM在社会模拟中，其内部表征如何编码社会概念（如性别、框架），以及如何通过操控这些内部表征来改变其决策行为？","design":"本研究并非直接进行人类仿真，而是提出一种探查和操控LLM内部表征的方法。在独裁者博弈实验中，从LLM的残差流中提取社会变量（如性别）的“变化向量”，通过正交化和注入扰动来操控这些向量，观察LLM分配决策的变化。","baseline":"无对照","findings":"LLM内部存在可提取的、与社会概念对应的方向向量；通过操控这些向量可以显著改变模型在独裁者博弈中的决策行为，为研究LLM如何编码社会意义提供了透明、可干预的方法。","reliability":"论文未讨论","relevance":"该研究直接探讨LLM在社会模拟中的内部决策机制，属于批判性方法研究，虽未提供人类基准对照，但其提出的表征操控方法对理解仿真可靠性和偏差具有重要价值，值得精读原文。","inspiration":"该方法通过提取LLM内部表征中的社会概念向量并进行正交化或注入扰动来操控决策行为，为因果识别提供了透明、可干预的框架｜可迁移到信贷审批中的性别歧视研究，探查并干预LLM在贷款决策中对性别信息的编码｜以LLM作为信贷审批官，提取性别变化向量并施加扰动，观察贷款批准率的变化，并与真实银行信贷数据中的性别差异进行对照"}},{"id":"2504.09663","version":2,"title":"Ordinary Least Squares as an Attention Mechanism","zh_title":"普通最小二乘法作为一种注意力机制","abstract":"I show that ordinary least squares (OLS) predictions can be rewritten as the output of a restricted attention module, akin to those forming the backbone of large language models. This connection offers an alternative perspective on attention beyond the conventional information retrieval framework, making it more accessible to researchers and analysts with a background in traditional statistics. It falls into place when OLS is framed as a similarity-based method in a transformed regressor space, distinct from the standard view based on partial correlations. In fact, the OLS solution can be recast as the outcome of an alternative problem: minimizing squared prediction errors by optimizing the embedding space in which training and test vectors are compared via inner products. Rather than estimating coefficients directly, we equivalently learn optimal encoding and decoding operations for predictors. From this vantage point, OLS maps naturally onto the query-key-value structure of attention mechanisms. Building on this foundation, I discuss key elements of Transformer-style attention and draw connections to classic ideas from time series econometrics.","authors":["Philippe Goulet Coulombe"],"categories":["cs.LG","econ.EM","math.ST","stat.ML"],"primary_category":"cs.LG","announce_type":"new","date":"2025-04-13","first_seen":"2025-04-13","revised_at":null,"abs_url":"https://arxiv.org/abs/2504.09663","pdf_url":"https://arxiv.org/pdf/2504.09663","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["注意力机制","OLS","统计方法"],"reason":"纯统计方法类比注意力机制，无LLM仿真人类被试，无人类数据对照。","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:58:57","error":null,"has_summary":false,"summary":null},{"id":"2504.09059","version":1,"title":"Large Language Models integration in Smart Grids","zh_title":"大语言模型在智能电网中的集成","abstract":"Large Language Models (LLMs) are changing the way we operate our society and will undoubtedly impact power systems as well - but how exactly? By integrating various data streams - including real-time grid data, market dynamics, and consumer behaviors - LLMs have the potential to make power system operations more adaptive, enhance proactive security measures, and deliver personalized energy services. This paper provides a comprehensive analysis of 30 real-world applications across eight key categories: Grid Operations and Management, Energy Markets and Trading, Personalized Energy Management and Customer Engagement, Grid Planning and Education, Grid Security and Compliance, Advanced Data Analysis and Knowledge Discovery, Emerging Applications and Societal Impact, and LLM-Enhanced Reinforcement Learning. Critical technical hurdles, such as data privacy and model reliability, are examined, along with possible solutions. Ultimately, this review illustrates how LLMs can significantly contribute to building more resilient, efficient, and sustainable energy infrastructures, underscoring the necessity of their responsible and equitable deployment.","authors":["Seyyedreza Madani","Ahmadreza Tavasoli","Zahra Khoshtarash Astaneh","Pierre-Olivier Pineau"],"categories":["cs.CY","cs.ET"],"primary_category":"cs.CY","announce_type":"new","date":"2025-04-12","first_seen":"2025-04-12","revised_at":null,"abs_url":"https://arxiv.org/abs/2504.09059","pdf_url":"https://arxiv.org/pdf/2504.09059","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["智能电网","LLM应用","多智能体系统"],"reason":"纯多智能体系统在电网中的应用，无人类行为对照，不涉及LLM仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:59:23","error":null,"has_summary":false,"summary":null},{"id":"2504.08260","version":2,"title":"Evaluating the Bias in LLMs for Surveying Opinion and Decision Making in Healthcare","zh_title":"评估大语言模型在医疗意见与决策调查中的偏差","abstract":"Generative agents have been increasingly used to simulate human behaviour in silico, driven by large language models (LLMs). These simulacra serve as sandboxes for studying human behaviour without compromising privacy or safety. However, it remains unclear whether such agents can truly represent real individuals. This work compares survey data from the Understanding America Study (UAS) on healthcare decision-making with simulated responses from generative agents. Using demographic-based prompt engineering, we create digital twins of survey respondents and analyse how well different LLMs reproduce real-world behaviours. Our findings show that some LLMs fail to reflect realistic decision-making, such as predicting universal vaccine acceptance. However, Llama 3 captures variations across race and Income more accurately but also introduces biases not present in the UAS data. This study highlights the potential of generative agents for behavioural research while underscoring the risks of bias from both LLMs and prompting strategies.","authors":["Yonchanok Khaokaew","Flora D. Salim","Andreas Züfle","Hao Xue","Taylor Anderson","C. Raina MacIntyre","Matthew Scotch","David J Heslop"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2025-04-11","first_seen":"2025-04-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2504.08260","pdf_url":"https://arxiv.org/pdf/2504.08260","source_feed":"backfill","score":10,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["LLM仿真","人类数据对照","医疗决策偏差"],"reason":"用LLM仿真医疗决策，与真实调查数据对照，评估偏差，直接命中核心判据。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:02","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":9,"question":"LLM能否有效模拟医疗决策（如疫苗接种意愿），以及在不同人口群体中会产生哪些偏差？","design":"使用LLM（如Llama 3等）基于人口统计属性（年龄、性别、收入、种族、教育、担忧程度）构建数字孪生，在不同疫情场景下提示模型回答疫苗接种意愿问题，比较模型输出与真实调查数据。","baseline":"理解美国研究（UAS）中关于COVID-19疫苗接种意愿的调查数据，涵盖2020年3月至2021年1月的多波次问卷。","findings":"部分LLM未能反映现实决策，例如预测普遍接受疫苗；Llama 3更准确地捕捉了种族和收入差异，但也引入了UAS数据中不存在的新偏差。","reliability":"论文指出LLM的预测效果取决于提供的上下文信息类型和数量，预训练数据和提示策略可能导致与人类偏差不同的偏差，且LLM可能仅反映训练数据中的统计模式而非真实决策过程。","relevance":"该研究直接使用LLM仿真医疗决策并与真实人类调查数据对照，评估偏差，完全符合研究者对LLM人类仿真实验、经济学/政策场景及可靠性批判的关注，值得精读原文。","inspiration":"借鉴其利用人口统计属性构建数字孪生并与多波次真实调查数据对照的仿真设计，评估LLM在决策模拟中的偏差｜可迁移到医疗健康政策的经济评估场景，如不同人群对医疗保险选择或健康储蓄账户的偏好差异｜以LLM为被试，基于收入、年龄、健康风险等属性生成数字孪生，提示其在不同保费和补贴政策下选择保险方案，结果变量为保险选择概率，用真实调查数据（如MEPS）作为基准对照"}},{"id":"2504.13908","version":3,"title":"AI-Assisted Conversational Interviewing: Effects on Data Quality and Respondent Experience","zh_title":"AI辅助的对话式访谈：对数据质量和受访者体验的影响","abstract":"Standardized surveys scale efficiently but sacrifice depth, while conversational interviews improve response quality at the cost of scalability and consistency. This study bridges the gap between these methods by introducing a framework for AI-assisted conversational interviewing. To evaluate this framework, we conducted a web survey experiment where 1,800 participants were randomly assigned to AI 'chatbots' which use large language models (LLMs) to dynamically probe respondents for elaboration and interactively code open-ended responses to fixed questions developed by human researchers. We assessed the AI chatbot's performance in terms of coding accuracy, response quality, and respondent experience. Our findings reveal that AI chatbots perform moderately well in live coding even without survey-specific fine-tuning, despite slightly inflated false positive errors due to respondent acquiescence bias. Open-ended responses were more detailed and informative, but this came at a slight cost to respondent experience. Our findings highlight the feasibility of using AI methods such as chatbots enhanced by LLMs to enhance open-ended data collection in web surveys.","authors":["Soubhik Barari","Jarret Angbazo","Natalie Wang","Leah M. Christian","Elizabeth Dean","Zoe Slowinski","Brandon Sepulvado"],"categories":["cs.HC","cs.AI","stat.AP"],"primary_category":"cs.HC","announce_type":"new","date":"2025-04-09","first_seen":"2025-04-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2504.13908","pdf_url":"https://arxiv.org/pdf/2504.13908","source_feed":"backfill","score":7,"bucket":"pending","rubric_hits":["A1","B1","B2"],"tags":["AI辅助调查","对话式访谈","数据质量"],"reason":"用LLM聊天机器人替代人类访谈员，收集并编码开放式回答，有真实人类对照实验，涉…","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:59:35","error":null,"has_summary":false,"summary":null},{"id":"2504.05862","version":2,"title":"Are Generative AI Agents Effective Personalized Financial Advisors?","zh_title":"生成式AI代理能成为有效的个性化理财顾问吗？","abstract":"Large language model-based agents are becoming increasingly popular as a low-cost mechanism to provide personalized, conversational advice, and have demonstrated impressive capabilities in relatively simple scenarios, such as movie recommendations. But how do these agents perform in complex high-stakes domains, where domain expertise is essential and mistakes carry substantial risk? This paper investigates the effectiveness of LLM-advisors in the finance domain, focusing on three distinct challenges: (1) eliciting user preferences when users themselves may be unsure of their needs, (2) providing personalized guidance for diverse investment preferences, and (3) leveraging advisor personality to build relationships and foster trust. Via a lab-based user study with 64 participants, we show that LLM-advisors often match human advisor performance when eliciting preferences, although they can struggle to resolve conflicting user needs. When providing personalized advice, the LLM was able to positively influence user behavior, but demonstrated clear failure modes. Our results show that accurate preference elicitation is key, otherwise, the LLM-advisor has little impact, or can even direct the investor toward unsuitable assets. More worryingly, users appear insensitive to the quality of advice being given, or worse these can have an inverse relationship. Indeed, users reported a preference for and increased satisfaction as well as emotional trust with LLMs adopting an extroverted persona, even though those agents provided worse advice.","authors":["Takehiro Takayanagi","Kiyoshi Izumi","Javier Sanz-Cruzado","Richard McCreadie","Iadh Ounis"],"categories":["cs.AI","cs.CL","cs.HC","cs.IR","q-fin.CP"],"primary_category":"cs.AI","announce_type":"new","date":"2025-04-08","first_seen":"2025-04-08","revised_at":null,"abs_url":"https://arxiv.org/abs/2504.05862","pdf_url":"https://arxiv.org/pdf/2504.05862","source_feed":"backfill","score":8,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["LLM仿真","人机对照实验","金融行为"],"reason":"用LLM模拟人类理财顾问，与真人对照，评估效果与失效模式，直接相关。","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:59:44","error":null,"has_summary":true,"summary":{"generated_at":"2025-04-08","rank":5,"question":"LLM作为个性化理财顾问，在偏好获取、个性化建议和人格影响方面表现如何？","design":"实验室用户研究，64名参与者扮演给定投资者画像，与LLM顾问进行两阶段对话：偏好获取和资产建议讨论，比较个性化vs非个性化顾问及不同人格的顾问，测量决策质量、用户满意度和信任。","baseline":"人类顾问在偏好获取阶段的表现作为对照基准。","findings":"LLM顾问在偏好获取上常能与人类顾问匹敌，但难以解决用户需求冲突；个性化建议能正向影响用户行为，但存在明显失效模式，且用户对建议质量不敏感，甚至偏好外向人格的顾问，尽管其建议更差。","reliability":"论文指出准确偏好获取是关键，否则LLM顾问影响甚微或引导投资者选择不合适资产；用户对建议质量不敏感，甚至出现反向关系，外向人格虽提升满意度和情感信任但建议质量更低。","relevance":"该研究直接以真人实验对照LLM在金融建议中的表现，揭示了偏好获取失效和用户对建议质量不敏感等关键偏差，对关注经济学实验和决策仿真的研究者极具参考价值，值得精读原文。","inspiration":"借鉴其分阶段对话设计和人格操纵处理，可迁移到信贷审批或保险推荐等金融场景，设计以LLM为被试、操纵建议人格或个性化程度、测量用户决策偏差和信任，并以人类顾问或历史决策数据为对照。"}},{"id":"2504.01566","version":1,"title":"GPT Adoption and the Impact of Disclosure Policies","zh_title":"GPT采用与披露政策的影响","abstract":"Generative Pre-trained Transformers (GPTs), particularly Large Language Models (LLMs) like ChatGPT, have proven effective in content generation and productivity enhancement. However, legal risks associated with these tools lead to adoption variance and concealment of AI use within organizations. This study examines the impact of disclosure on ChatGPT adoption in legal, audit and advisory roles in consulting firms through the lens of agency theory. We conducted a survey experiment to evaluate agency costs in the context of unregulated corporate use of ChatGPT, with a particular focus on how mandatory disclosure influences information asymmetry and misaligned interests. Our findings indicate that in the absence of corporate regulations, such as an AI policy, firms may incur agency costs, which can hinder the full benefits of GPT adoption. While disclosure policies reduce information asymmetry, they do not significantly lower overall agency costs due to managers undervaluing analysts' contributions with GPT use. Finally, we examine the scope of existing regulations in Europe and the United States regarding disclosure requirements, explore the sharing of risk and responsibility within firms, and analyze how incentive mechanisms promote responsible AI adoption.","authors":["Cathy Yang","David Restrepo Amariles","Leo Allen","Aurore Troussel"],"categories":["econ.GN"],"primary_category":"econ.GN","announce_type":"new","date":"2025-04-02","first_seen":"2025-04-02","revised_at":null,"abs_url":"https://arxiv.org/abs/2504.01566","pdf_url":"https://arxiv.org/pdf/2504.01566","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM标注","调查实验","代理理论"],"reason":"用LLM替代人类标注员，非仿真被试，但涉及人类调查实验，边界情形。","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:58:42","error":null,"has_summary":false,"summary":null},{"id":"2503.22726","version":1,"title":"InfoBid: A Simulation Framework for Studying Information Disclosure in Auctions with Large Language Model-based Agents","zh_title":"InfoBid：基于大语言模型代理研究拍卖中信息披露的仿真框架","abstract":"In online advertising systems, publishers often face a trade-off in information disclosure strategies: while disclosing more information can enhance efficiency by enabling optimal allocation of ad impressions, it may lose revenue potential by decreasing uncertainty among competing advertisers. Similar to other challenges in market design, understanding this trade-off is constrained by limited access to real-world data, leading researchers and practitioners to turn to simulation frameworks. The recent emergence of large language models (LLMs) offers a novel approach to simulations, providing human-like reasoning and adaptability without necessarily relying on explicit assumptions about agent behavior modeling. Despite their potential, existing frameworks have yet to integrate LLM-based agents for studying information asymmetry and signaling strategies, particularly in the context of auctions. To address this gap, we introduce InfoBid, a flexible simulation framework that leverages LLM agents to examine the effects of information disclosure strategies in multi-agent auction settings. Using GPT-4o, we implemented simulations of second-price auctions with diverse information schemas. The results reveal key insights into how signaling influences strategic behavior and auction outcomes, which align with both economic and social learning theories. Through InfoBid, we hope to foster the use of LLMs as proxies for human economic and social agents in empirical studies, enhancing our understanding of their capabilities and limitations. This work bridges the gap between theoretical market designs and practical applications, advancing research in market simulations, information design, and agent-based reasoning while offering a valuable tool for exploring the dynamics of digital economies.","authors":["Yue Yin"],"categories":["cs.GT","cs.CL","cs.HC","cs.MA","econ.GN"],"primary_category":"cs.GT","announce_type":"new","date":"2025-03-26","first_seen":"2025-03-26","revised_at":null,"abs_url":"https://arxiv.org/abs/2503.22726","pdf_url":"https://arxiv.org/pdf/2503.22726","source_feed":"backfill","score":8,"bucket":"selected","rubric_hits":["A3","B2","B4"],"tags":["LLM仿真","拍卖实验","经济行为"],"reason":"用LLM代理模拟拍卖中的人类经济行为，与理论对照，涉及经济学实验场景并讨论局限…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:02","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":64,"question":"在拍卖中，信息披露策略如何影响基于LLM的智能体的出价行为与拍卖结果？","design":"使用GPT-4o构建LLM智能体模拟竞拍者，在第二价格拍卖中施加不同的信息披露方案（信号），测量出价、收入、效率等拍卖结果。","baseline":"无对照","findings":"信号显著影响LLM智能体的策略行为与拍卖结果，且观察到的行为模式与经济理论和社会学习理论一致。","reliability":"论文未讨论","relevance":"高度相关：用LLM代理模拟拍卖中的人类经济行为，研究信息不对称下的策略互动，并讨论LLM作为人类代理的潜力与局限，直接命中研究者关心的经济学实验仿真与可靠性评估。","inspiration":"该方法通过向LLM智能体提供不同信号来模拟信息披露，可借鉴其处理信息干预的方式，用于研究经济决策中的信息效应｜可迁移到资产定价实验中，研究公开与私有信号如何影响投资者的价格预期与交易行为｜设计：以LLM作为被试，随机分配接收不同精度的资产价值信号，测量其报价与交易量，并与历史实验市场数据或理性预期均衡基准对照"}},{"id":"2503.12556","version":1,"title":"From Guessing to Asking: An Approach to Resolving the Persona Knowledge Gap in LLMs during Multi-Turn Conversations","zh_title":"从猜测到询问：解决多轮对话中大语言模型角色知识缺口的方法","abstract":"In multi-turn dialogues, large language models (LLM) face a critical challenge of ensuring coherence while adapting to user-specific information. This study introduces the persona knowledge gap, the discrepancy between a model's internal understanding and the knowledge required for coherent, personalized conversations. While prior research has recognized these gaps, computational methods for their identification and resolution remain underexplored. We propose Conversation Preference Elicitation and Recommendation (CPER), a novel framework that dynamically detects and resolves persona knowledge gaps using intrinsic uncertainty quantification and feedback-driven refinement. CPER consists of three key modules: a Contextual Understanding Module for preference extraction, a Dynamic Feedback Module for measuring uncertainty and refining persona alignment, and a Persona-Driven Response Generation module for adapting responses based on accumulated user context. We evaluate CPER on two real-world datasets: CCPE-M for preferential movie recommendations and ESConv for mental health support. Using A/B testing, human evaluators preferred CPER's responses 42% more often than baseline models in CCPE-M and 27% more often in ESConv. A qualitative human evaluation confirms that CPER's responses are preferred for maintaining contextual relevance and coherence, particularly in longer (12+ turn) conversations.","authors":["Sarvesh Baskar","Tanmay Tulsidas Verelakar","Srinivasan Parthasarathy","Manas Gaur"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2025-03-16","first_seen":"2025-03-16","revised_at":null,"abs_url":"https://arxiv.org/abs/2503.12556","pdf_url":"https://arxiv.org/pdf/2503.12556","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["对话系统","个性化推荐","角色扮演"],"reason":"研究个性化对话中的知识缺口，属于角色扮演聊天，无实验或测量目的，不涉及人类行为…","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:59:45","error":null,"has_summary":false,"summary":null},{"id":"2503.11531","version":1,"title":"Potential of large language model-powered nudges for promoting daily water and energy conservation","zh_title":"大语言模型驱动的助推在促进日常节水节能中的潜力","abstract":"The increasing amount of pressure related to water and energy shortages has increased the urgency of cultivating individual conservation behaviors. While the concept of nudging, i.e., providing usage-based feedback, has shown promise in encouraging conservation behaviors, its efficacy is often constrained by the lack of targeted and actionable content. This study investigates the impact of the use of large language models (LLMs) to provide tailored conservation suggestions for conservation intentions and their rationale. Through a survey experiment with 1,515 university participants, we compare three virtual nudging scenarios: no nudging, traditional nudging with usage statistics, and LLM-powered nudging with usage statistics and personalized conservation suggestions. The results of statistical analyses and causal forest modeling reveal that nudging led to an increase in conservation intentions among 86.9%-98.0% of the participants. LLM-powered nudging achieved a maximum increase of 18.0% in conservation intentions, surpassing traditional nudging by 88.6%. Furthermore, structural equation modeling results reveal that exposure to LLM-powered nudges enhances self-efficacy and outcome expectations while diminishing dependence on social norms, thereby increasing intrinsic motivation to conserve. These findings highlight the transformative potential of LLMs in promoting individual water and energy conservation, representing a new frontier in the design of sustainable behavioral interventions and resource management.","authors":["Zonghan Li","Song Tong","Yi Liu","Kaiping Peng","Chunyan Wang"],"categories":["cs.CY","cs.AI"],"primary_category":"cs.CY","announce_type":"new","date":"2025-03-14","first_seen":"2025-03-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2503.11531","pdf_url":"https://arxiv.org/pdf/2503.11531","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2"],"tags":["LLM仿真","行为干预","节能实验"],"reason":"用LLM生成个性化节能建议，通过调查实验与人类对照，评估对行为意图的影响，属于…","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:59:36","error":null,"has_summary":true,"summary":{"generated_at":"2025-03-14","rank":2,"question":"基于LLM的个性化助推与传统使用统计助推相比，能否更有效地提升个体的节水节能意图？","design":"本研究并非用LLM模拟人类被试，而是通过随机对照调查实验，将1515名大学生随机分为三组，分别接受无助推、传统使用统计助推、LLM生成个性化建议的助推，测量其节水节能意图的变化。","baseline":"无助推组和传统使用统计助推组作为对照，均为真实人类被试。","findings":"LLM助推使86.9%-98.0%的参与者节水节能意图提升，最大增幅达18.0%，效果比传统助推高88.6%。结构方程模型显示，LLM助推通过增强自我效能和结果预期、降低社会规范依赖，提升内在动机。","reliability":"论文未讨论","relevance":"该研究用LLM生成个性化干预内容，通过真实人类调查实验评估行为意图变化，属于LLM辅助行为干预效果评估，与研究者关注的LLM仿真实验和经济学实验场景高度相关，值得细读其因果推断设计和心理机制分析。","inspiration":"借鉴其随机分组和因果森林方法评估异质性处理效应，可迁移到消费者节能行为干预政策评估场景。｜设计一个实验：以居民为被试，随机分配接收传统节能建议或LLM个性化建议，结果变量为实际用电量变化，对照真实智能电表数据。"}},{"id":"2503.10990","version":2,"title":"Statistical Impossibility and Possibility of Aligning LLMs with Human Preferences: From Condorcet Paradox to Nash Equilibrium","zh_title":"将大语言模型与人类偏好对齐的统计不可能性与可能性：从孔多塞悖论到纳什均衡","abstract":"Aligning large language models (LLMs) with diverse human preferences is critical for ensuring fairness and informed outcomes when deploying these models for decision-making. In this paper, we seek to uncover fundamental statistical limits concerning aligning LLMs with human preferences, with a focus on the probabilistic representation of human preferences and the preservation of diverse preferences in aligned LLMs. We first show that human preferences can be represented by a reward model if and only if the preference among LLM-generated responses is free of any Condorcet cycle. Moreover, we prove that Condorcet cycles exist with probability converging to one exponentially fast under a general probabilistic preference model called the Luce model, thereby demonstrating the impossibility of fully aligning human preferences using reward-based approaches such as reinforcement learning from human feedback. Next, we explore the conditions under which LLMs would employ mixed strategies -- meaning they do not collapse to a single response -- when aligned in the limit using a non-reward-based approach, such as Nash learning from human feedback. We identify a necessary and sufficient condition for mixed strategies: the absence of a response that is preferred over all others by a majority. As a blessing, we prove that this condition holds with high probability under the Luce model, thereby highlighting the statistical possibility of preserving minority preferences without explicit regularization in aligning LLMs.","authors":["Kaizhao Liu","Qi Long","Zhekun Shi","Weijie J. Su","Jiancong Xiao"],"categories":["cs.GT","cs.LG","econ.TH","math.ST","stat.ML"],"primary_category":"cs.GT","announce_type":"new","date":"2025-03-14","first_seen":"2025-03-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2503.10990","pdf_url":"https://arxiv.org/pdf/2503.10990","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["偏好对齐","统计极限","RLHF"],"reason":"纯理论分析偏好对齐的统计极限，无LLM仿真人类被试或人类数据对照","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:59:03","error":null,"has_summary":false,"summary":null},{"id":"2503.10248","version":1,"title":"LLM Agents Display Human Biases but Exhibit Distinct Learning Patterns","zh_title":"LLM智能体表现出人类偏见但学习模式不同","abstract":"We investigate the choice patterns of Large Language Models (LLMs) in the context of Decisions from Experience tasks that involve repeated choice and learning from feedback, and compare their behavior to human participants. We find that on the aggregate, LLMs appear to display behavioral biases similar to humans: both exhibit underweighting rare events and correlation effects. However, more nuanced analyses of the choice patterns reveal that this happens for very different reasons. LLMs exhibit strong recency biases, unlike humans, who appear to respond in more sophisticated ways. While these different processes may lead to similar behavior on average, choice patterns contingent on recent events differ vastly between the two groups. Specifically, phenomena such as ``surprise triggers change\" and the ``wavy recency effect of rare events\" are robustly observed in humans, but entirely absent in LLMs. Our findings provide insights into the limitations of using LLMs to simulate and predict humans in learning environments and highlight the need for refined analyses of their behavior when investigating whether they replicate human decision making tendencies.","authors":["Idan Horowitz","Ori Plonsky"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2025-03-13","first_seen":"2025-03-13","revised_at":null,"abs_url":"https://arxiv.org/abs/2503.10248","pdf_url":"https://arxiv.org/pdf/2503.10248","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","决策实验","人类对照"],"reason":"用LLM复现人类决策实验，与真实人类数据对照，并指出仿真失效条件，高度相关。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:01","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":68,"question":"LLM在基于经验的决策任务中是否表现出与人类相似的行为偏差和学习模式？","design":"将多个LLM作为被试，完成四项重复二元选择决策任务（100试次），通过反馈学习收益分布，测量选择行为；同时操纵历史提供方式和模型温度参数。","baseline":"真实人类被试在相同决策任务中的选择数据，作为行为对比基准。","findings":"总体上看，LLM表现出与人类相似的低估稀有事件和相关效应，但过程不同：LLM有强烈的近因偏差，而人类对稀有事件有“惊讶触发改变”和“波浪式近因效应”，这些模式在LLM中完全缺失。","reliability":"LLM在聚合层面可能模拟人类偏差，但内在学习过程截然不同，基于近期事件的精细分析显示其无法复现人类特有的动态反应模式，提示在涉及反馈学习的场景中用LLM仿真人类决策存在局限。","relevance":"该研究直接对比LLM与人类在经济学实验范式下的决策，揭示仿真在表面聚合指标下有效但内在机制失效，高度契合研究者对LLM仿真可靠性及失效条件的批判性关注，值得精读。","inspiration":"该方法通过重复二元选择任务和试次级反馈学习数据，精细对比LLM与人类在聚合与过程层面的行为差异，值得借鉴｜可迁移到资产定价实验中的罕见事件学习，如研究投资者如何更新对崩盘风险的信念｜以LLM为被试，操纵历史收益序列呈现方式，测量其风险资产配置比例，并与真实投资者在类似实验中的选择数据对照"}},{"id":"2503.09639","version":4,"title":"Can A Society of Generative Agents Simulate Human Behavior and Inform Public Health Policy? A Case Study on Vaccine Hesitancy","zh_title":"生成式智能体社会能否模拟人类行为并为公共卫生政策提供信息？以疫苗犹豫为例","abstract":"Can we simulate a sandbox society with generative agents to model human behavior, thereby reducing the over-reliance on real human trials for assessing public policies? In this work, we investigate the feasibility of simulating health-related decision-making, using vaccine hesitancy, defined as the delay in acceptance or refusal of vaccines despite the availability of vaccination services (MacDonald, 2015), as a case study. To this end, we introduce the VacSim framework with 100 generative agents powered by Large Language Models (LLMs). VacSim simulates vaccine policy outcomes with the following steps: 1) instantiate a population of agents with demographics based on census data; 2) connect the agents via a social network and model vaccine attitudes as a function of social dynamics and disease-related information; 3) design and evaluate various public health interventions aimed at mitigating vaccine hesitancy. To align with real-world results, we also introduce simulation warmup and attitude modulation to adjust agents' attitudes. We propose a series of evaluations to assess the reliability of various LLM simulations. Experiments indicate that models like Llama and Qwen can simulate aspects of human behavior but also highlight real-world alignment challenges, such as inconsistent responses with demographic profiles. This early exploration of LLM-driven simulations is not meant to serve as definitive policy guidance; instead, it serves as a call for action to examine social simulation for policy development.","authors":["Abe Bohan Hou","Hongru Du","Yichen Wang","Jingyu Zhang","Zixiao Wang","Paul Pu Liang","Daniel Khashabi","Lauren Gardner","Tianxing He"],"categories":["cs.MA","cs.AI","cs.CL","cs.CY","cs.HC"],"primary_category":"cs.MA","announce_type":"new","date":"2025-03-12","first_seen":"2025-03-12","revised_at":null,"abs_url":"https://arxiv.org/abs/2503.09639","pdf_url":"https://arxiv.org/pdf/2503.09639","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A3","A5","B1","B2","B4"],"tags":["LLM人类仿真","疫苗犹豫","政策评估"],"reason":"用LLM代理模拟疫苗犹豫行为并与真实人口数据对照，评估仿真可靠性，直接命中核心…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:00","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":32,"question":"能否用生成式智能体构建的沙盒社会模拟人类疫苗犹豫行为，从而为公共卫生政策提供参考？","design":"使用100个由LLM驱动的生成式智能体，基于人口普查数据赋予人口学特征，通过社交网络和新闻信息模拟疫苗态度变化，并设计不同公共卫生干预措施（如政策）来观察态度轨迹。","baseline":"Nguyen et al. (2022) 的美国COVID-19疫苗犹豫调查数据，包含1300万份真实人类回答，用于校准智能体的人口学分布。","findings":"Llama和Qwen等模型能模拟人类行为的某些方面，但存在与现实对齐的挑战，例如智能体的回答与人口学特征不一致。","reliability":"论文承认仿真存在现实对齐挑战，如回答与人口学特征不一致，并指出当前探索不能作为政策指导，仅呼吁进一步研究。","relevance":"该研究直接以LLM代理模拟疫苗犹豫行为，并与真实人口调查数据对照，评估仿真可靠性，完全命中研究者对经济学实验和政策评估场景的兴趣，值得精读原文。","inspiration":"借鉴其利用人口普查数据校准智能体人口学特征并与大规模真实调查数据对照的仿真设计｜可迁移到政策公告对消费者通胀预期形成的实验，如模拟不同央行沟通策略对预期的影响｜以LLM智能体为被试，按收入、教育等特征分层，施加不同措辞的政策公告作为处理，测量预期通胀率的变化，用密歇根大学消费者调查的真实预期数据做对照"}},{"id":"2503.07510","version":1,"title":"Sometimes the Model doth Preach: Quantifying Religious Bias in Open LLMs through Demographic Analysis in Asian Nations","zh_title":"有时模型在布道：通过亚洲国家人口统计分析量化开放LLM中的宗教偏见","abstract":"Large Language Models (LLMs) are capable of generating opinions and propagating bias unknowingly, originating from unrepresentative and non-diverse data collection. Prior research has analysed these opinions with respect to the West, particularly the United States. However, insights thus produced may not be generalized in non-Western populations. With the widespread usage of LLM systems by users across several different walks of life, the cultural sensitivity of each generated output is of crucial interest. Our work proposes a novel method that quantitatively analyzes the opinions generated by LLMs, improving on previous work with regards to extracting the social demographics of the models. Our method measures the distance from an LLM's response to survey respondents, through Hamming Distance, to infer the demographic characteristics reflected in the model's outputs. We evaluate modern, open LLMs such as Llama and Mistral on surveys conducted in various global south countries, with a focus on India and other Asian nations, specifically assessing the model's performance on surveys related to religious tolerance and identity. Our analysis reveals that most open LLMs match a single homogeneous profile, varying across different countries/territories, which in turn raises questions about the risks of LLMs promoting a hegemonic worldview, and undermining perspectives of different minorities. Our framework may also be useful for future research investigating the complex intersection between training data, model architecture, and the resulting biases reflected in LLM outputs, particularly concerning sensitive topics like religious tolerance and identity.","authors":["Hari Shankar","Vedanta S P","Tejas Cavale","Ponnurangam Kumaraguru","Abhijnan Chakraborty"],"categories":["cs.CY","cs.CL"],"primary_category":"cs.CY","announce_type":"new","date":"2025-03-10","first_seen":"2025-03-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2503.07510","pdf_url":"https://arxiv.org/pdf/2503.07510","source_feed":"backfill","score":8,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","宗教偏见","人类数据对照"],"reason":"用LLM复现调查回答并与真实人类数据对照，评估宗教偏见，揭示仿真失效条件。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:15:00","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":42,"question":"开放大语言模型在回答宗教相关调查时，反映了怎样的社会人口特征和宗教偏见？","design":"使用Llama、Mistral等开源LLM，以零样本方式回答皮尤研究中心在印度及东亚、东南亚国家进行的宗教宽容与认同调查问卷，通过汉明距离计算模型回答与真实受访者回答的匹配度，推断模型所反映的人口统计特征。","baseline":"皮尤研究中心在印度、日本、香港、韩国、台湾、越南、印尼、马来西亚、新加坡、斯里兰卡、泰国等亚洲国家/地区收集的真实调查数据，包含受访者的宗教、性别、年龄、教育等人口统计变量。","findings":"大多数开源LLM的回答与单一同质化的人口统计画像相匹配，且该画像在不同国家/地区间存在差异；通过提示词指示模型扮演特定群体并未显著改变其回答画像。","reliability":"论文未讨论","relevance":"该研究直接以真实人类调查数据为基准，评估LLM在宗教宽容等敏感话题上的回答偏差，揭示了模型输出同质化及可能强化霸权世界观的失效模式，高度契合研究者对LLM仿真可靠性及失效条件的关注，值得阅读原文。","inspiration":"借鉴其使用真实调查数据作为基准，通过汉明距离量化LLM回答与人类群体回答的匹配度，并推断模型隐含的人口统计画像的方法｜可迁移到信贷审批中的宗教或种族偏见检测，例如评估LLM在模拟贷款决策时是否系统性地偏向或歧视特定宗教群体｜以LLM作为被试，向其呈现不同宗教背景的贷款申请人资料，要求做出批准/拒绝决策，结果变量为批准率差异，并以真实银行信贷审批数据或审计研究结果作为对照基准"}},{"id":"2503.05529","version":1,"title":"PoSSUM: A Protocol for Surveying Social-media Users with Multimodal LLMs","zh_title":"PoSSUM：一种利用多模态大语言模型调查社交媒体用户的协议","abstract":"This paper introduces PoSSUM, an open-source protocol for unobtrusive polling of social-media users via multimodal Large Language Models (LLMs). PoSSUM leverages users' real-time posts, images, and other digital traces to create silicon samples that capture information not present in the LLM's training data. To obtain representative estimates, PoSSUM employs Multilevel Regression and Post-Stratification (MrP) with structured priors to counteract the observable selection biases of social-media platforms. The protocol is validated during the 2024 U.S. Presidential Election, for which five PoSSUM polls were conducted and published on GitHub and X. In the final poll, fielded October 17-26 with a synthetic sample of 1,054 X users, PoSSUM accurately predicted the outcomes in 50 of 51 states and assigned the Republican candidate a win probability of 0.65. Notably, it also exhibited lower state-level bias than most established pollsters. These results demonstrate PoSSUM's potential as a fully automated, unobtrusive alternative to traditional survey methods.","authors":["Roberto Cerina"],"categories":["stat.AP","cs.SI"],"primary_category":"stat.AP","announce_type":"new","date":"2025-03-07","first_seen":"2025-03-07","revised_at":null,"abs_url":"https://arxiv.org/abs/2503.05529","pdf_url":"https://arxiv.org/pdf/2503.05529","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM仿真","选举预测","人类行为对照"],"reason":"用LLM模拟社交媒体用户投票行为，并与真实选举结果对照，属于人类仿真实验。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:14:59","error":null,"has_summary":true,"summary":{"generated_at":"2025-03-07","rank":9,"question":"能否利用多模态大语言模型和社交媒体数字痕迹，构建无侵扰的民意调查方法，并准确预测选举结果？","design":"PoSSUM协议：使用多模态LLM基于X用户的实时帖子、图像等数字痕迹生成合成样本（硅样本），通过多水平回归与事后分层（MrP）校正平台选择偏差，测量投票意向。在2024年美国总统选举中进行了5次民意调查，最终轮使用1054名X用户的合成样本。","baseline":"2024年美国总统选举的真实结果（各州胜负及全国胜率）。","findings":"最终轮预测准确预测了51个州中50个的结果，并赋予共和党候选人0.65的胜率；其州级偏差低于大多数传统民调机构。","reliability":"论文未讨论失效条件与局限。","relevance":"高度相关：该研究直接用LLM生成合成样本模拟人类投票行为，并以真实选举结果为对照，验证了LLM仿真在选举预测中的有效性，符合研究者对经济学实验和政策评估场景的关注。值得精读原文以了解MrP校正偏差的具体方法及协议细节。","inspiration":"借鉴该方法利用多模态LLM从社交媒体数字痕迹生成合成样本，并通过多水平回归与事后分层校正选择偏差，以低成本、无侵扰方式测量群体态度｜可迁移至政策公告的预期形成研究，如央行沟通对通胀预期的影响，或财政刺激对消费者信心的影响｜以X平台用户为合成样本来源，用多模态LLM提取用户对政策公告的态度，处理为不同措辞或发布渠道的公告版本，结果变量为通胀预期指数，以真实消费者调查数据（如密歇根大学消费者信心指数）作为对照基准"}},{"id":"2503.05481","version":1,"title":"Maximum Hallucination Standards for Domain-Specific Large Language Models","zh_title":"领域特定大语言模型的最大幻觉标准","abstract":"Large language models (LLMs) often generate inaccurate yet credible-sounding content, known as hallucinations. This inherent feature of LLMs poses significant risks, especially in critical domains. I analyze LLMs as a new class of engineering products, treating hallucinations as a product attribute. I demonstrate that, in the presence of imperfect awareness of LLM hallucinations and misinformation externalities, net welfare improves when the maximum acceptable level of LLM hallucinations is designed to vary with two domain-specific factors: the willingness to pay for reduced LLM hallucinations and the marginal damage associated with misinformation.","authors":["Tingmingke Lu"],"categories":["econ.GN"],"primary_category":"econ.GN","announce_type":"new","date":"2025-03-07","first_seen":"2025-03-07","revised_at":null,"abs_url":"https://arxiv.org/abs/2503.05481","pdf_url":"https://arxiv.org/pdf/2503.05481","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["LLM幻觉","经济学模型","产品属性"],"reason":"研究LLM幻觉的经济学模型，不涉及人类仿真或行为对照，属纯理论分析。","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:59:23","error":null,"has_summary":false,"summary":null},{"id":"2503.04804","version":4,"title":"What do Large Language Models Say About Animals? Investigating Risks of Animal Harm in Generated Text","zh_title":"大型语言模型对动物有何看法？探究生成文本中动物伤害的风险","abstract":"As machine learning systems become increasingly embedded in society, their impact on human and nonhuman life continues to escalate. Technical evaluations have addressed a variety of potential harms from large language models (LLMs) towards humans and the environment, but there is little empirical work regarding harms towards nonhuman animals. Following the growing recognition of animal protection in regulatory and ethical AI frameworks, we present AnimalHarmBench (AHB), a benchmark for risks of animal harm in LLM-generated text. Our benchmark dataset comprises 1,850 curated questions from Reddit post titles and 2,500 synthetic questions based on 50 animal categories (e.g., cats, reptiles) and 50 ethical scenarios with a 70-30 public-private split. Scenarios include open-ended questions about how to treat animals, practical scenarios with potential animal harm, and willingness-to-pay measures for the prevention of animal harm. Using the LLM-as-a-judge framework, responses are evaluated for their potential to increase or decrease harm, and evaluations are debiased for the tendency of judges to judge their own outputs more favorably. AHB reveals significant differences across frontier LLMs, animal categories, scenarios, and subreddits. We conclude with future directions for technical research and addressing the challenges of building evaluations on complex social and moral topics.","authors":["Arturs Kanepajs","Aditi Basu","Sankalpa Ghose","Constance Li","Akshat Mehta","Ronak Mehta","Samuel David Tucker-Davis","Eric Zhou","Bob Fischer","Jacy Reese Anthis"],"categories":["cs.CY","cs.CL"],"primary_category":"cs.CY","announce_type":"new","date":"2025-03-03","first_seen":"2025-03-03","revised_at":null,"abs_url":"https://arxiv.org/abs/2503.04804","pdf_url":"https://arxiv.org/pdf/2503.04804","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["LLM安全评测","动物伦理","基准数据集"],"reason":"评估LLM生成文本对动物的伤害风险，属安全评测，非人类仿真实验。","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:59:23","error":null,"has_summary":false,"summary":null},{"id":"2503.00725","version":1,"title":"Causal Inference on Outcomes Learned from Text","zh_title":"基于文本学习结果的因果推断","abstract":"We propose a machine-learning tool that yields causal inference on text in randomized trials. Based on a simple econometric framework in which text may capture outcomes of interest, our procedure addresses three questions: First, is the text affected by the treatment? Second, which outcomes is the effect on? And third, how complete is our description of causal effects? To answer all three questions, our approach uses large language models (LLMs) that suggest systematic differences across two groups of text documents and then provides valid inference based on costly validation. Specifically, we highlight the need for sample splitting to allow for statistical validation of LLM outputs, as well as the need for human labeling to validate substantive claims about how documents differ across groups. We illustrate the tool in a proof-of-concept application using abstracts of academic manuscripts.","authors":["Iman Modarressi","Jann Spiess","Amar Venugopal"],"categories":["econ.EM","cs.CL","cs.LG","stat.ME"],"primary_category":"econ.EM","announce_type":"new","date":"2025-03-02","first_seen":"2025-03-02","revised_at":null,"abs_url":"https://arxiv.org/abs/2503.00725","pdf_url":"https://arxiv.org/pdf/2503.00725","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["因果推断","文本分析","LLM标注"],"reason":"用LLM替代人工标注文本因果效应，非仿真人类被试，但涉及人类验证与统计推断，属…","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:58:57","error":null,"has_summary":false,"summary":null},{"id":"2503.00177","version":1,"title":"Steering Large Language Model Activations in Sparse Spaces","zh_title":"在稀疏空间中操控大语言模型激活","abstract":"A key challenge in AI alignment is guiding large language models (LLMs) to follow desired behaviors at test time. Activation steering, which modifies internal model activations during inference, offers a potential solution. However, prior work in dense activation spaces struggles with superposition, wherein multiple features become entangled, limiting interpretability and precise control. In contrast, sparse representations provide an untapped opportunity for more interpretable behavior modulation. In this work, we introduce sparse activation steering (SAS), a method that leverages sparse autoencoders (SAEs) to steer LLM behavior in sparse spaces. By isolating behavior-specific features through a contrastive prompt-pairing approach, we define a set of features that can selectively reinforce or suppress behaviors. Experiments on Gemma 2 LLMs show that SAS vectors enable nuanced behavioral modulation and finer-grained control. Furthermore, scaling SAEs improves monosemanticity of SAS vectors, suggesting more reliable and interpretable interventions.","authors":["Reza Bayat","Ali Rahimi-Kalahroudi","Mohammad Pezeshki","Sarath Chandar","Pascal Vincent"],"categories":["cs.LG","cs.AI"],"primary_category":"cs.LG","announce_type":"new","date":"2025-02-28","first_seen":"2025-02-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2503.00177","pdf_url":"https://arxiv.org/pdf/2503.00177","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["激活操控","稀疏自编码器","AI对齐"],"reason":"纯模型内部激活操控方法研究，无人类行为仿真或对照，属NLP能力评测范畴。","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:59:19","error":null,"has_summary":false,"summary":null},{"id":"2502.18371","version":1,"title":"MindMem: Multimodal for Predicting Advertisement Memorability Using LLMs and Deep Learning","zh_title":"MindMem：利用LLM和深度学习预测广告记忆度的多模态模型","abstract":"In the competitive landscape of advertising, success hinges on effectively navigating and leveraging complex interactions among consumers, advertisers, and advertisement platforms. These multifaceted interactions compel advertisers to optimize strategies for modeling consumer behavior, enhancing brand recall, and tailoring advertisement content. To address these challenges, we present MindMem, a multimodal predictive model for advertisement memorability. By integrating textual, visual, and auditory data, MindMem achieves state-of-the-art performance, with a Spearman's correlation coefficient of 0.631 on the LAMBDA and 0.731 on the Memento10K dataset, consistently surpassing existing methods. Furthermore, our analysis identified key factors influencing advertisement memorability, such as video pacing, scene complexity, and emotional resonance. Expanding on this, we introduced MindMem-ReAd (MindMem-Driven Re-generated Advertisement), which employs Large Language Model-based simulations to optimize advertisement content and placement, resulting in up to a 74.12% improvement in advertisement memorability. Our results highlight the transformative potential of Artificial Intelligence in advertising, offering advertisers a robust tool to drive engagement, enhance competitiveness, and maximize impact in a rapidly evolving market.","authors":["Sepehr Asgarian","Qayam Jetha","Jouhyun Jeon"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2025-02-25","first_seen":"2025-02-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2502.18371","pdf_url":"https://arxiv.org/pdf/2502.18371","source_feed":"backfill","score":4,"bucket":"other","rubric_hits":["C1"],"tags":["广告记忆度","多模态预测","LLM仿真"],"reason":"用LLM仿真优化广告内容，属多智能体协作生成，无人类行为对照，非人类被试替代研…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:32","error":null,"has_summary":false,"summary":null},{"id":"2502.16280","version":1,"title":"Human Preferences in Large Language Model Latent Space: A Technical Analysis on the Reliability of Synthetic Data in Voting Outcome Prediction","zh_title":"大语言模型潜在空间中的人类偏好：合成数据在投票结果预测中可靠性的技术分析","abstract":"Generative AI (GenAI) is increasingly used in survey contexts to simulate human preferences. While many research endeavors evaluate the quality of synthetic GenAI data by comparing model-generated responses to gold-standard survey results, fundamental questions about the validity and reliability of using LLMs as substitutes for human respondents remain. Our study provides a technical analysis of how demographic attributes and prompt variations influence latent opinion mappings in large language models (LLMs) and evaluates their suitability for survey-based predictions. Using 14 different models, we find that LLM-generated data fails to replicate the variance observed in real-world human responses, particularly across demographic subgroups. In the political space, persona-to-party mappings exhibit limited differentiation, resulting in synthetic data that lacks the nuanced distribution of opinions found in survey data. Moreover, we show that prompt sensitivity can significantly alter outputs for some models, further undermining the stability and predictiveness of LLM-based simulations. As a key contribution, we adapt a probe-based methodology that reveals how LLMs encode political affiliations in their latent space, exposing the systematic distortions introduced by these models. Our findings highlight critical limitations in AI-generated survey data, urging caution in its use for public opinion research, social science experimentation, and computational behavioral modeling.","authors":["Sarah Ball","Simeon Allmendinger","Frauke Kreuter","Niklas Kühl"],"categories":["cs.LG","cs.AI"],"primary_category":"cs.LG","announce_type":"new","date":"2025-02-22","first_seen":"2025-02-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2502.16280","pdf_url":"https://arxiv.org/pdf/2502.16280","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","A4","B1","B3","B4"],"tags":["LLM人类仿真","合成数据可靠性","政治偏好预测"],"reason":"直接评估LLM仿真人类投票偏好的可靠性与偏差，有真实调查数据对照，并揭示失效条…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:14:58","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":61,"question":"LLM生成的合成数据在多大程度上能复现真实人类调查回答的分布，以及提示不稳定性如何在模型的潜在空间中体现？","design":"使用14个白盒LLM，基于德国选举研究（GLES）构建包含年龄、性别、教育、收入、就业、政治倾向、东西德等人口属性的理论驱动型人设提示，让模型预测投票选择；同时通过改写提示考察提示敏感性，并利用Wahl-o-Mat数据训练探针分析模型潜在空间中的政治映射。","baseline":"以2021年德国纵向选举研究（GLES）的加权横截面调查数据作为真实人类投票行为的对照基准。","findings":"LLM生成数据无法复现真实人类回答的方差，尤其在人口子群中，人设到政党的映射区分度低，缺乏真实调查中的细致意见分布；提示敏感性会显著改变部分模型的输出，进一步削弱仿真稳定性，且在某些模型中潜在空间熵值与提示敏感性相关。","reliability":"论文指出LLM仿真在人口子群方差复现、意见分布细致度和提示稳定性方面存在根本局限，警告在公共舆论研究、社会科学实验和计算行为建模中使用合成数据需谨慎，但未讨论模型选择、语言文化差异等外部效度限制。","relevance":"该研究直接评估LLM替代人类被试进行投票预测的可靠性与偏差，有真实调查数据对照，并揭示了仿真在方差复现和提示敏感性上的失效条件，高度契合研究者对经济学实验和政策评估场景中仿真批判性分析的兴趣，值得精读原文。","inspiration":"该方法通过理论驱动构建人口属性人设提示并系统改写提示来检验仿真稳定性，为经济实验中的处理稳健性检验提供了可借鉴的测量范式｜可迁移至政策评估中的预期形成研究，例如考察不同人口子群对央行通胀预测公告的反应异质性｜以LLM模拟不同收入与教育水平的个体，处理为改写后的通胀预测措辞，结果变量为通胀预期调整幅度，以真实消费者预期调查数据作为对照基准"}},{"id":"2502.14499","version":1,"title":"MLGym: A New Framework and Benchmark for Advancing AI Research Agents","zh_title":"MLGym：推进AI研究智能体的新框架与基准","abstract":"We introduce Meta MLGym and MLGym-Bench, a new framework and benchmark for evaluating and developing LLM agents on AI research tasks. This is the first Gym environment for machine learning (ML) tasks, enabling research on reinforcement learning (RL) algorithms for training such agents. MLGym-bench consists of 13 diverse and open-ended AI research tasks from diverse domains such as computer vision, natural language processing, reinforcement learning, and game theory. Solving these tasks requires real-world AI research skills such as generating new ideas and hypotheses, creating and processing data, implementing ML methods, training models, running experiments, analyzing the results, and iterating through this process to improve on a given task. We evaluate a number of frontier large language models (LLMs) on our benchmarks such as Claude-3.5-Sonnet, Llama-3.1 405B, GPT-4o, o1-preview, and Gemini-1.5 Pro. Our MLGym framework makes it easy to add new tasks, integrate and evaluate models or agents, generate synthetic data at scale, as well as develop new learning algorithms for training agents on AI research tasks. We find that current frontier models can improve on the given baselines, usually by finding better hyperparameters, but do not generate novel hypotheses, algorithms, architectures, or substantial improvements. We open-source our framework and benchmark to facilitate future research in advancing the AI research capabilities of LLM agents.","authors":["Deepak Nathani","Lovish Madaan","Nicholas Roberts","Nikolay Bashlykov","Ajay Menon","Vincent Moens","Amar Budhiraja","Despoina Magka","Vladislav Vorotilov","Gaurav Chaurasia","Dieuwke Hupkes","Ricardo Silveira Cabral","Tatiana Shavrina","Jakob Foerster","Yoram Bachrach","William Yang Wang","Roberta Raileanu"],"categories":["cs.CL","cs.AI","cs.LG"],"primary_category":"cs.CL","announce_type":"new","date":"2025-02-20","first_seen":"2025-02-20","revised_at":null,"abs_url":"https://arxiv.org/abs/2502.14499","pdf_url":"https://arxiv.org/pdf/2502.14499","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["AI研究智能体","多智能体协作","基准测试"],"reason":"纯多智能体系统研究，LLM agent 协作完成AI研究任务，不涉及人类行为对…","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:59:32","error":null,"has_summary":false,"summary":null},{"id":"2502.15800","version":3,"title":"LLM Agents Do Not Replicate Human Market Traders: Evidence From Experimental Finance","zh_title":"LLM代理无法复现人类市场交易者：来自实验金融的证据","abstract":"This paper explores how Large Language Models (LLMs) behave in a classic experimental finance paradigm widely known for eliciting bubbles and crashes in human participants. We adapt an established trading design, where traders buy and sell a risky asset with a known fundamental value, and introduce several LLM-based agents, both in single-model markets (all traders are instances of the same LLM) and in mixed-model \"battle royale\" settings (multiple LLMs competing in the same market). Our findings reveal that LLMs generally exhibit a \"textbook-rational\" approach, pricing the asset near its fundamental value, and show only a muted tendency toward bubble formation. Further analyses indicate that LLM-based agents display less trading strategy variance in contrast to humans. Taken together, these results highlight the risk of relying on LLM-only data to replicate human-driven market phenomena, as key behavioral features, such as large emergent bubbles, were not robustly reproduced. While LLMs clearly possess the capacity for strategic decision-making, their relative consistency and rationality suggest that they do not accurately mimic human market dynamics.","authors":["Thomas Henning","Siddhartha M. Ojha","Ross Spoon","Jiatong Han","Colin F. Camerer"],"categories":["q-fin.TR"],"primary_category":"q-fin.TR","announce_type":"new","date":"2025-02-18","first_seen":"2025-02-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2502.15800","pdf_url":"https://arxiv.org/pdf/2502.15800","source_feed":"backfill","score":10,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["LLM仿真","实验金融","人类行为对照"],"reason":"用LLM代理模拟人类交易实验，与真实人类数据对照，发现LLM未能复现泡沫，批判…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:14:58","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":16,"question":"LLM代理在经典实验金融范式中能否复现人类交易者产生的资产泡沫与市场动态？","design":"采用Smith等人(2014)的固定基本面价值资产交易实验，在oTree平台上运行30期开放叫价市场。LLM代理分为单一模型同质市场和多个LLM模型混合的“大逃杀”市场，每轮提交买卖订单及价格预期，并允许通过“洞察”和“思考”文本实现跨轮记忆与思维链推理。结果变量为交易价格偏离基本面价值的程度、泡沫生成、交易策略方差等。","baseline":"对照真实人类被试在相同实验设计下的交易数据，人类市场一致产生显著价格泡沫和偏离基本面。","findings":"LLM代理普遍表现出“教科书式理性”，定价接近基本面价值，仅呈现微弱的泡沫倾向；其交易策略方差更低，更依赖基本面而非启发式策略，与人类行为存在系统性差异。","reliability":"论文指出，开箱即用的LLM未能复现人类市场中的大型泡沫等关键行为特征，表明仅依赖LLM数据复制人类驱动的市场现象存在风险，但未深入讨论LLM代理在何种条件下可能失效。","relevance":"该研究直接以真实人类实验为基准，检验LLM代理在金融实验中的行为复现能力，并得出批判性结论：LLM未能模拟人类市场泡沫，对评估LLM仿真可靠性及失效条件具有重要参考价值，值得精读原文。","inspiration":"该方法通过oTree平台实现多期开放叫价市场，并让LLM代理提交订单与价格预期，同时利用“洞察”和“思考”文本实现跨轮记忆与思维链推理，为经济实验的自动化仿真提供了可借鉴的设计框架｜可迁移至资产定价实验，用于检验不同信息结构或交易机制下LLM代理能否复现人类的价格泡沫与过度反应｜可设计一个资产定价实验，以LLM代理为被试，施加不同信息透明度处理，测量交易价格偏离基本面的程度，并与真实人类实验数据对照，评估LLM在模拟市场非理性行为时的有效性"}},{"id":"2502.10266","version":1,"title":"Are Large Language Models the future crowd workers of Linguistics?","zh_title":"大语言模型能否成为语言学未来的众包工作者？","abstract":"Data elicitation from human participants is one of the core data collection strategies used in empirical linguistic research. The amount of participants in such studies may vary considerably, ranging from a handful to crowdsourcing dimensions. Even if they provide resourceful extensive data, both of these settings come alongside many disadvantages, such as low control of participants' attention during task completion, precarious working conditions in crowdsourcing environments, and time-consuming experimental designs. For these reasons, this research aims to answer the question of whether Large Language Models (LLMs) may overcome those obstacles if included in empirical linguistic pipelines. Two reproduction case studies are conducted to gain clarity into this matter: Cruz (2023) and Lombard et al. (2021). The two forced elicitation tasks, originally designed for human participants, are reproduced in the proposed framework with the help of OpenAI's GPT-4o-mini model. Its performance with our zero-shot prompting baseline shows the effectiveness and high versatility of LLMs, that tend to outperform human informants in linguistic tasks. The findings of the second replication further highlight the need to explore additional prompting techniques, such as Chain-of-Thought (CoT) prompting, which, in a second follow-up experiment, demonstrates higher alignment to human performance on both critical and filler items. Given the limited scale of this study, it is worthwhile to further explore the performance of LLMs in empirical Linguistics and in other future applications in the humanities.","authors":["Iris Ferrazzo"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2025-02-14","first_seen":"2025-02-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2502.10266","pdf_url":"https://arxiv.org/pdf/2502.10266","source_feed":"backfill","score":7,"bucket":"pending","rubric_hits":["A1","B1","B4"],"tags":["LLM仿真","人类数据对照","语言学实验"],"reason":"用LLM复现人类语言学任务，并与人类数据对照，讨论对齐与失效条件","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:14:57","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":90,"question":"大语言模型能否替代人类被试，成为实证语言学研究中新的众包工作者？","design":"使用GPT-4o-mini模型，通过零样本提示复现两项原为人类被试设计的强制诱导语言任务（Cruz 2023和Lombard et al. 2021），并在第二项任务中进一步尝试思维链提示，测量模型回答与人类回答的对齐程度。","baseline":"两项复现任务的原版人类被试数据，包括Cruz (2023)和Lombard et al. (2021)中的真实人类回答。","findings":"GPT-4o-mini在所有实验条件下均优于人类被试；思维链提示能进一步提升模型在关键项和填充项上与人类表现的对齐度。","reliability":"论文承认研究规模有限，仅覆盖两项案例，且未探索更多提示技术，未来需在更广泛的语言学任务和人文领域进一步验证。","relevance":"该研究直接用LLM复现人类语言学实验并与真实人类数据对照，探讨对齐条件与提示策略的影响，符合您对仿真可靠性及失效条件的关注，值得阅读原文以了解具体任务设计和偏差细节。","inspiration":"该方法借鉴了用零样本和思维链提示复现人类实验并与原版人类数据直接对照的范式，以评估LLM仿真对齐度｜可迁移到行为经济学中的跨期选择实验，检验LLM是否能复现人类的时间偏好不一致｜用LLM作为被试，施加不同时间折现的金钱选择任务，结果变量为折现率，对照真实人类实验数据（如Frederick等2002的经典数据）"}},{"id":"2502.10308","version":1,"title":"LLM-Powered Preference Elicitation in Combinatorial Assignment","zh_title":"基于大语言模型的组合分配偏好获取","abstract":"We study the potential of large language models (LLMs) as proxies for humans to simplify preference elicitation (PE) in combinatorial assignment. While traditional PE methods rely on iterative queries to capture preferences, LLMs offer a one-shot alternative with reduced human effort. We propose a framework for LLM proxies that can work in tandem with SOTA ML-powered preference elicitation schemes. Our framework handles the novel challenges introduced by LLMs, such as response variability and increased computational costs. We experimentally evaluate the efficiency of LLM proxies against human queries in the well-studied course allocation domain, and we investigate the model capabilities required for success. We find that our approach improves allocative efficiency by up to 20%, and these results are robust across different LLMs and to differences in quality and accuracy of reporting.","authors":["Ermis Soumalias","Yanchen Jiang","Kehang Zhu","Michael Curry","Sven Seuken","David C. Parkes"],"categories":["cs.AI","cs.GT","cs.LG"],"primary_category":"cs.AI","announce_type":"new","date":"2025-02-14","first_seen":"2025-02-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2502.10308","pdf_url":"https://arxiv.org/pdf/2502.10308","source_feed":"backfill","score":7,"bucket":"pending","rubric_hits":["A1","B1","B2"],"tags":["LLM代理","偏好获取","人类对照"],"reason":"用LLM代理人类偏好，在课程分配场景与真实人类查询对照，可迁移到人类仿真实验。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:14:57","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":39,"question":"如何利用大语言模型作为人类代理，通过一次性自然语言输入简化组合分配中的偏好诱导，并提升分配效率？","design":"使用大语言模型（如GPT系列）作为学生代理，根据学生提供的简短文本偏好描述，替代学生回答迭代比较查询；将LLM代理集成到MLCM机制中，通过链式思维提示和噪声鲁棒损失函数处理响应变异，测量分配效率的提升。","baseline":"真实人类学生在课程分配机制中通过迭代查询提供的偏好数据，以及Course Match机制下的分配结果。","findings":"LLM代理方法在现实场景中将分配效率提高了最多20%，且结果在不同LLM架构和报告质量差异下保持稳健。","reliability":"论文未讨论","relevance":"该研究直接以LLM代理人类偏好，在课程分配场景中与真实人类查询对照，可迁移到经济学实验和政策评估中的人类仿真研究，值得精读。","inspiration":"借鉴其将LLM作为代理集成到已有机制中，并通过噪声鲁棒设计处理LLM响应变异的方法。｜可迁移到公共资源分配或拍卖设计实验中，用LLM模拟投标者偏好以测试机制效率。｜以LLM作为被试，施加不同文本偏好描述作为处理，测量拍卖分配效率，并与真实人类实验数据对照。"}},{"id":"2503.05708","version":1,"title":"On Large Language Models as Data Sources for Policy Deliberation on Climate Change and Sustainability","zh_title":"大语言模型作为气候与可持续性政策审议数据源的研究","abstract":"We pose the research question, \"Can LLMs provide credible evaluation scores, suitable for constructing starter MCDM models that support commencing deliberation regarding climate and sustainability policies?\" In this exploratory study we i. Identify a number of interesting policy alternatives that are actively considered by local governments in the United States (and indeed around the world). ii. Identify a number of quality-of-life indicators as apt evaluation criteria for these policies. iii. Use GPT-4 to obtain evaluation scores for the policies on multiple criteria. iv. Use the TOPSIS MCDM method to rank the policies based on the obtained evaluation scores. v. Evaluate the quality and validity of the resulting table ensemble of scores by comparing the TOPSIS-based policy rankings with those obtained by an informed assessment exercise. We find that GPT-4 is in rough agreement with the policy rankings of our informed assessment exercise. Hence, we conclude (always provisionally and assuming a modest level of vetting) that GPT-4 can be used as a credible input, even starting point, for subsequent deliberation processes on climate and sustainability policies.","authors":["Rachel Bina","Kha Luong","Shrey Mehta","Daphne Pang","Mingjun Xie","Christine Chou","Steven O. Kimbrough"],"categories":["cs.CY","econ.GN"],"primary_category":"cs.CY","announce_type":"new","date":"2025-02-13","first_seen":"2025-02-13","revised_at":null,"abs_url":"https://arxiv.org/abs/2503.05708","pdf_url":"https://arxiv.org/pdf/2503.05708","source_feed":"backfill","score":7,"bucket":"pending","rubric_hits":["A1","B1","B2"],"tags":["LLM仿真","政策评估","人类数据对照"],"reason":"用GPT-4替代人类专家评估政策，并与人类评估对照，属于仿真人类决策，但非严格…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:14:59","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":126,"question":"大语言模型能否为气候与可持续政策提供可信的评估分数，用于构建启动政策协商的初始多准则决策模型？","design":"本研究并非严格的人类仿真实验，而是使用GPT-4对一系列气候与可持续政策在多个生活质量指标上进行评分，然后采用TOPSIS多准则决策方法对政策排序，并将结果与知情评估（informed assessment）产生的排序进行对比。","baseline":"对照的真实人类数据来自一项知情评估练习（informed assessment exercise），由人类专家对部分政策进行评价和排序，作为比较基准。","findings":"GPT-4生成的政策排序与知情评估的排序大致一致，尤其在生活质量准则上的排序高度吻合。因此，GPT-4可作为政策协商的初始可信输入或起点。","reliability":"论文强调结论是暂时的，需假设一定程度的审查，并指出模型输出应作为草案，需在动态环境中持续修订和审议，但未系统讨论失效条件或具体局限。","relevance":"该研究用GPT-4替代人类专家进行政策评估并与人类判断对照，属于用LLM仿真人类决策的探索，但非严格的行为实验复现，且缺乏对仿真偏差的深入批判，适合作为方法参考而非可靠性评估的典型案例。","inspiration":"该方法借鉴了用LLM替代人类专家进行多准则评分，并与真实人类判断排序对照的验证思路。｜可迁移到政策公告对金融市场预期形成的评估，如央行沟通对投资者情绪的影响。｜以GPT-4作为被试，输入不同措辞的央行声明，让其对资产价格预期进行评分，结果与真实市场调查数据或分析师预测排序进行对照。"}},{"id":"2502.08691","version":2,"title":"AgentSociety: Large-Scale Simulation of LLM-Driven Generative Agents Advances Understanding of Human Behaviors and Society","zh_title":"AgentSociety：大规模LLM驱动生成式代理模拟推进对人类行为与社会的理解","abstract":"Understanding human behavior and society is a central focus in social sciences, with the rise of generative social science marking a significant paradigmatic shift. By leveraging bottom-up simulations, it replaces costly and logistically challenging traditional experiments with scalable, replicable, and systematic computational approaches for studying complex social dynamics. Recent advances in large language models (LLMs) have further transformed this research paradigm, enabling the creation of human-like generative social agents and realistic simulacra of society. In this paper, we propose AgentSociety, a large-scale social simulator that integrates LLM-driven agents, a realistic societal environment, and a powerful large-scale simulation engine. Based on the proposed simulator, we generate social lives for over 10k agents, simulating their 5 million interactions both among agents and between agents and their environment. Furthermore, we explore the potential of AgentSociety as a testbed for computational social experiments, focusing on five key social issues: polarization, the spread of inflammatory messages, the effects of universal basic income policies, the impact of external shocks such as hurricanes, and urban sustainability. These five issues serve as valuable cases for assessing AgentSociety's support for typical research methods -- such as surveys, interviews, and interventions -- as well as for investigating the patterns, causes, and underlying mechanisms of social issues. The alignment between AgentSociety's outcomes and real-world experimental results not only demonstrates its ability to capture human behaviors and their underlying mechanisms, but also underscores its potential as an important platform for social scientists and policymakers.","authors":["Jinghua Piao","Yuwei Yan","Jun Zhang","Nian Li","Junbo Yan","Xiaochong Lan","Zhihong Lu","Zhiheng Zheng","Jing Yi Wang","Di Zhou","Chen Gao","Fengli Xu","Fang Zhang","Ke Rong","Jun Su","Yong Li"],"categories":["cs.SI","cs.AI"],"primary_category":"cs.SI","announce_type":"new","date":"2025-02-12","first_seen":"2025-02-12","revised_at":null,"abs_url":"https://arxiv.org/abs/2502.08691","pdf_url":"https://arxiv.org/pdf/2502.08691","source_feed":"api","score":10,"bucket":"selected","rubric_hits":["A1","A3","A5","B1","B2","B3"],"tags":["LLM社会仿真","人类行为复现","政策评估"],"reason":"用LLM代理模拟社会行为并与真实数据对照，直接复现人类实验","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:14:56","error":null,"has_summary":true,"summary":{"generated_at":"2025-02-12","rank":1,"question":"LLM驱动的生成式智能体能否在大规模社会仿真中复现人类行为与社会现象？","design":"构建AgentSociety仿真器，包含LLM驱动的智能体（基于GPT等模型）、现实社会环境和大规模仿真引擎；模拟超过1万个智能体的500万次交互（智能体间及智能体与环境间）；针对极化、煽动性信息传播、全民基本收入政策、飓风外部冲击、城市可持续性五个社会议题，支持调查、访谈、干预等研究方法。","baseline":"真实世界实验结果（具体未在摘要和引言中详述，但提及结果与真实实验对齐）。","findings":"AgentSociety能够捕捉人类行为及其潜在机制，仿真结果与真实世界实验结果一致；展示了作为社会科学家和政策制定者重要平台的潜力。","reliability":"论文未讨论失效条件与局限。","relevance":"高度相关：直接使用LLM作为人类被试替代品，有真实人类数据对照，涉及政策评估（如全民基本收入）等经济学场景，且关注仿真可靠性，值得精读原文以获取更多细节。","inspiration":"借鉴其大规模LLM智能体仿真与真实实验对齐的方法，可对智能体施加政策干预并测量行为变化｜可迁移到全民基本收入（UBI）对劳动供给与消费行为影响的经济学问题｜用LLM智能体作为被试，处理组接受UBI转账，对照组无干预，结果变量为工作时长与消费支出，对照真实UBI实验数据（如芬兰基本收入实验）"}},{"id":"2502.07307","version":1,"title":"CreAgent: Towards Long-Term Evaluation of Recommender System under Platform-Creator Information Asymmetry","zh_title":"CreAgent：平台-创作者信息不对称下推荐系统长期评估","abstract":"Ensuring the long-term sustainability of recommender systems (RS) emerges as a crucial issue. Traditional offline evaluation methods for RS typically focus on immediate user feedback, such as clicks, but they often neglect the long-term impact of content creators. On real-world content platforms, creators can strategically produce and upload new items based on user feedback and preference trends. While previous studies have attempted to model creator behavior, they often overlook the role of information asymmetry. This asymmetry arises because creators primarily have access to feedback on the items they produce, while platforms possess data on the entire spectrum of user feedback. Current RS simulators, however, fail to account for this asymmetry, leading to inaccurate long-term evaluations. To address this gap, we propose CreAgent, a Large Language Model (LLM)-empowered creator simulation agent. By incorporating game theory's belief mechanism and the fast-and-slow thinking framework, CreAgent effectively simulates creator behavior under conditions of information asymmetry. Additionally, we enhance CreAgent's simulation ability by fine-tuning it using Proximal Policy Optimization (PPO). Our credibility validation experiments show that CreAgent aligns well with the behaviors between real-world platform and creator, thus improving the reliability of long-term RS evaluations. Moreover, through the simulation of RS involving CreAgents, we can explore how fairness- and diversity-aware RS algorithms contribute to better long-term performance for various stakeholders. CreAgent and the simulation platform are publicly available at https://github.com/shawnye2000/CreAgent.","authors":["Xiaopeng Ye","Chen Xu","Zhongxiang Sun","Jun Xu","Gang Wang","Zhenhua Dong","Ji-Rong Wen"],"categories":["cs.IR"],"primary_category":"cs.IR","announce_type":"new","date":"2025-02-11","first_seen":"2025-02-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2502.07307","pdf_url":"https://arxiv.org/pdf/2502.07307","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["LLM智能体","推荐系统仿真","创作者行为模拟"],"reason":"用LLM模拟创作者行为，属于社会模拟但无真实人类数据对照，且非直接仿真人类被试。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:32","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":49,"question":"在平台与创作者信息不对称下，如何利用LLM智能体模拟创作者行为，以实现对推荐系统的长期评估？","design":"使用LLM驱动的创作者智能体CreAgent，结合博弈论信念机制和快慢思考框架模拟创作者在信息不对称下的内容生产决策，并通过PPO微调提升仿真能力。","baseline":"无对照","findings":"CreAgent能够有效模拟真实平台中创作者在信息不对称下的行为，提高了推荐系统长期评估的可靠性；通过仿真实验，可探索公平性和多样性感知推荐算法对多方长期绩效的影响。","reliability":"论文未讨论","relevance":"该研究用LLM模拟创作者行为，属于社会模拟但无真实人类数据对照，且非直接仿真人类被试，与研究者关注的人类仿真实验及经济学实验场景关联较弱。","inspiration":"与经济金融研究关联不大"}},{"id":"2502.07736","version":2,"title":"Menu Pricing of Large Language Models","zh_title":"大语言模型的菜单定价","abstract":"We develop a framework for the optimal pricing and product design of LLMs in which a provider sells menus of token budgets to users who differ in their valuations across a continuum of tasks. Under a homogeneous production technology, we show that users' high-dimensional type profiles are summarized by a scalar index, reducing the seller's problem to one-dimensional screening. The optimal mechanism takes the form of committed-spend contracts: buyers pay for a budget that they allocate across token classes priced at marginal cost. We extend the analysis to environments with multiple differentiated models and to competition between a proprietary leader and an open-source fringe, showing that competitive pressure reshapes both the intensive and extensive margins of compute provision. Each element of our theory (token-budget menus, maximum- and minimum-spend plans, multi-model versioning, and linear API pricing) has a direct counterpart in the observed pricing practices of providers such as Anthropic, OpenAI, and GitHub.","authors":["Dirk Bergemann","Alessandro Bonatti","Alex Smolin"],"categories":["econ.TH"],"primary_category":"econ.TH","announce_type":"new","date":"2025-02-11","first_seen":"2025-02-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2502.07736","pdf_url":"https://arxiv.org/pdf/2502.07736","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["定价策略","经济学理论","产品设计"],"reason":"研究LLM菜单定价与产品设计，属多智能体经济学模型，不涉及人类行为仿真或对照。","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:59:03","error":null,"has_summary":false,"summary":null},{"id":"2502.06387","version":2,"title":"How Humans Help LLMs: Assessing and Incentivizing Human Preference Annotators","zh_title":"人类如何帮助大语言模型：评估与激励人类偏好标注员","abstract":"Human-annotated preference data play an important role in aligning large language models (LLMs). In this paper, we study two connected questions: how to monitor the quality of human preference annotators and how to incentivize them to provide high-quality annotations. In current practice, expert-based monitoring is a natural workhorse for quality control, but it performs poorly in preference annotation because annotators are heterogeneous and downstream model performance is an indirect and noisy proxy for annotation quality. We therefore propose a self-consistency monitoring scheme tailored to preference annotation, and analyze the statistical sample complexity of both methods. This practitioner-facing analysis identifies how many inspected samples are needed to reliably assess an annotator and shows when self-consistency monitoring can outperform expert-based monitoring. We then use the resulting monitoring signal as the performance measure in a principal-agent model, which lets us study a second sample-complexity question: how many monitored samples are needed before simple contracts perform close to the ideal benchmark in which annotation quality is perfectly observable. Under this continuous action space, we show that this shortfall scales as $Θ(1/\\sqrt{\\mathcal{I} n \\log n})$ for binary contracts and $Θ(1/(\\mathcal{I}n))$ for linear contracts, where $\\mathcal{I}$ is the Fisher information and $n$ is the number of samples; we further show that the linear contracts are rate-optimal among general contracts. This contrasts with the known result that binary contracts are optimal and of $\\exp(-Θ(n))$ when the action space is discrete \\citep{frick2023monitoring}.","authors":["Shang Liu","Hanzhao Wang","Zhongyao Ma","Xiaocheng Li"],"categories":["cs.LG","cs.GT","econ.TH"],"primary_category":"cs.LG","announce_type":"new","date":"2025-02-10","first_seen":"2025-02-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2502.06387","pdf_url":"https://arxiv.org/pdf/2502.06387","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C5"],"tags":["人类标注","偏好对齐","激励设计"],"reason":"研究人类标注员激励与质量监控，不涉及用LLM仿真人类被试，属于用人类数据对齐模…","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:59:05","error":null,"has_summary":false,"summary":null},{"id":"2502.03158","version":3,"title":"Strategizing with AI: Insights from a Beauty Contest Experiment","zh_title":"与AI博弈：选美竞赛实验的启示","abstract":"A $p$-beauty contest is a wide class of games of guessing the most popular strategy among other players. In particular, guessing a fraction of a mean of numbers chosen by all players is a classic behavioral experiment designed to test iterative reasoning patterns among various groups of people. The previous literature reveals that the level of sophistication of the opponents is an important factor affecting the outcome of the game. Smarter decision makers choose strategies that are closer to theoretical Nash equilibrium and demonstrate faster convergence to equilibrium in iterated contests with information revelation. We replicate a series of classic experiments by running virtual experiments with large language models (LLMs) who play against various groups of virtual players. Our results show that LLMs recognize strategic context of the game and demonstrate expected adaptability to the changing set of parameters. LLMs systematically behave in a more sophisticated way compared to the participants of the original experiments. All LLMs still fail to identify dominant strategies in a two-player game. Our results contribute to the discussion on the accuracy of modeling human economic agents by artificial intelligence.","authors":["Iuliia Alekseenko","Dmitry Dagaev","Sofia Paklina","Petr Parshakov"],"categories":["econ.GN"],"primary_category":"econ.GN","announce_type":"new","date":"2025-02-05","first_seen":"2025-02-05","revised_at":null,"abs_url":"https://arxiv.org/abs/2502.03158","pdf_url":"https://arxiv.org/pdf/2502.03158","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM仿真","行为博弈","人类数据对照"],"reason":"用LLM复现选美博弈实验，与真实人类数据对照，评估仿真准确性。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:14:55","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":28,"question":"大语言模型在选美博弈（p-beauty contest）中的策略行为是否与人类相似，能否准确模拟人类经济主体？","design":"用多个大语言模型（LLMs）作为虚拟被试，复现经典的“猜数字”实验，让LLMs与不同虚拟对手群体进行博弈，操纵参数p和对手类型，测量其选择的数字、收敛速度及对纳什均衡的偏离。","baseline":"以Nagel (1995)等经典人类实验数据作为对照基准，比较LLMs与真实人类参与者的行为差异。","findings":"LLMs能识别博弈的战略情境并适应参数变化，行为比人类更复杂；但在两人博弈中，所有LLMs均未能识别占优策略。","reliability":"论文未讨论","relevance":"该研究直接使用LLMs复现经典行为博弈实验，并与真实人类数据对照，评估仿真准确性，高度契合研究者对LLM人类仿真可靠性及失效条件的关注，值得精读原文。","inspiration":"借鉴其用经典实验范式复现并系统操纵对手类型和参数的方法，可迁移到资产定价实验中的预期形成研究。｜设计一个LLM作为交易者的资产市场实验，操纵市场信息结构和对手策略复杂度，测量LLM的报价偏差和收敛速度，并与人类实验数据对照。"}},{"id":"2502.00070","version":2,"title":"Can AI Solve the Peer Review Crisis? A Large Scale Cross Model Experiment of LLMs' Performance and Biases in Evaluating over 1000 Economics Papers","zh_title":"AI能解决同行评审危机吗？一项关于LLM在评估1000多篇经济学论文中的表现与偏差的大规模跨模型实验","abstract":"This study examines the potential of large language models (LLMs) to augment the academic peer review process by reliably evaluating the quality of economics research without introducing systematic bias. We conduct one of the first large-scale experimental assessments of four LLMs (GPT-4o, Claude 3.5, Gemma 3, and LLaMA 3.3) across two complementary experiments. In the first, we use nonparametric binscatter and linear regression techniques to analyze over 29,000 evaluations of 1,220 anonymized papers drawn from 110 economics journals excluded from the training data of current LLMs, along with a set of AI-generated submissions. The results show that LLMs consistently distinguish between higher- and lower-quality research based solely on textual content, producing quality gradients that closely align with established journal prestige measures. Claude and Gemma perform exceptionally well in capturing these gradients, while GPT excels in detecting AI-generated content. The second experiment comprises 8,910 evaluations designed to assess whether LLMs replicate human like biases in single blind reviews. By systematically varying author gender, institutional affiliation, and academic prominence across 330 papers, we find that GPT, Gemma, and LLaMA assign significantly higher ratings to submissions from top male authors and elite institutions relative to the same papers presented anonymously. These results emphasize the importance of excluding author-identifying information when deploying LLMs in editorial screening. Overall, our findings provide compelling evidence and practical guidance for integrating LLMs into peer review to enhance efficiency, improve accuracy, and promote equity in the publication process of economics research.","authors":["Pat Pataranutaporn","Nattavudh Powdthavee","Chayapatr Achiwaranguprok","Pattie Maes"],"categories":["cs.CY","cs.AI","econ.GN"],"primary_category":"cs.CY","announce_type":"new","date":"2025-01-31","first_seen":"2025-01-31","revised_at":null,"abs_url":"https://arxiv.org/abs/2502.00070","pdf_url":"https://arxiv.org/pdf/2502.00070","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM审稿","经济学论文评估","偏差分析"],"reason":"用LLM替代审稿人，属于替代人类劳动而非仿真被试，但涉及人类审稿数据对照，边界…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:14:55","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":48,"question":"大型语言模型能否可靠地评估经济学论文质量，并在单盲评审中是否会复现人类审稿人的系统性偏见？","design":"本研究并非用LLM仿真人类被试，而是用GPT-4o、Claude 3.5、Gemma 3、LLaMA 3.3四种模型扮演审稿人，对1220篇匿名经济学论文进行质量评分，并在一项实验中系统操纵作者性别、机构声誉和学术地位，测量评分变化。","baseline":"以论文所在期刊的IDEAS/RePEc综合排名作为客观质量基准，并与人类审稿中已知的偏见模式（如对顶尖机构男性作者的偏好）进行对照。","findings":"LLM能仅凭文本内容区分论文质量高低，评分梯度与期刊声望高度一致，其中Claude和Gemma表现最佳，GPT在识别AI生成内容上更优。但在单盲条件下，GPT、Gemma和LLaMA对顶尖男性作者和精英机构的论文给出显著更高评分，表现出与人类相似的偏见。","reliability":"论文指出，LLM在单盲评审中会复现甚至放大人类偏见，因此建议在编辑筛选中必须隐去作者身份信息；未讨论模型在其他学科、语言或更复杂评审维度上的泛化局限。","relevance":"该研究直接涉及用LLM替代人类评审，并系统检验了偏见复现问题，虽非典型人类仿真实验，但为经济学领域AI辅助决策的可靠性与偏差提供了关键证据，值得精读。","inspiration":"借鉴其通过操纵作者身份特征来检测模型偏见的实验设计，以及用期刊排名作为质量基准的对照思路。｜可迁移到经济学论文评审、基金申请书评估或信贷审批中的歧视研究。｜以LLM作为评审人，随机分配附有不同作者特征（性别、机构）的同一篇论文，测量评分差异，并以真实期刊排名或人类评审数据作为基准，检验模型偏见。"}},{"id":"2501.19266","version":1,"title":"Jackpot! Alignment as a Maximal Lottery","zh_title":"头奖！对齐作为最大彩票","abstract":"Reinforcement Learning from Human Feedback (RLHF), the standard for aligning Large Language Models (LLMs) with human values, is known to fail to satisfy properties that are intuitively desirable, such as respecting the preferences of the majority \\cite{ge2024axioms}. To overcome these issues, we propose the use of a probabilistic Social Choice rule called \\emph{maximal lotteries} as a replacement for RLHF. We show that a family of alignment techniques, namely Nash Learning from Human Feedback (NLHF) \\cite{munos2023nash} and variants, approximate maximal lottery outcomes and thus inherit its beneficial properties. We confirm experimentally that our proposed methodology handles situations that arise when working with preferences more robustly than standard RLHF, including supporting the preferences of the majority, providing principled ways of handling non-transitivities in the preference data, and robustness to irrelevant alternatives. This results in systems that better incorporate human values and respect human intentions.","authors":["Roberto-Rafael Maura-Rivero","Marc Lanctot","Francesco Visin","Kate Larson"],"categories":["cs.AI","cs.LG","econ.TH"],"primary_category":"cs.AI","announce_type":"new","date":"2025-01-31","first_seen":"2025-01-31","revised_at":null,"abs_url":"https://arxiv.org/abs/2501.19266","pdf_url":"https://arxiv.org/pdf/2501.19266","source_feed":"backfill","score":4,"bucket":"other","rubric_hits":["C1"],"tags":["RLHF","社会选择","多智能体对齐"],"reason":"研究多智能体对齐方法，不涉及用LLM仿真人类被试或与人类行为对照","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:59:05","error":null,"has_summary":false,"summary":null},{"id":"2501.17310","version":4,"title":"Probing LLM World Models: Enhancing Guesstimation with Wisdom of Crowds Decoding","zh_title":"探究LLM世界模型：用群体智慧解码增强估算能力","abstract":"Guesstimation -- the task of making approximate quantitative estimates about objects or events -- is a common real-world skill, yet remains underexplored in large language model (LLM) research. We introduce three guesstimation datasets: MARBLES, FUTURE, and ELECPRED, spanning physical estimation (e.g., how many marbles fit in a cup) to abstract predictions (e.g., the 2024 U.S. presidential election). Inspired by the social science concept of Wisdom of Crowds (WOC)- where the median of multiple estimates improves accuracy-we propose WOC decoding for LLMs. We replicate WOC effects in human participants and find that LLMs exhibit similar benefits: median aggregation across sampled responses consistently improves accuracy over greedy decoding, self-consistency decoding, and mean decoding. This suggests that LLMs encode a world model that supports approximate reasoning. Our results position guesstimation as a useful probe of LLM world knowledge and highlight WOC decoding as a strategy for enhancing LLM guesstimation performance on real-world tasks.","authors":["Yun-Shiuan Chuang","Sameer Narendran","Nikunj Harlalka","Alexander Cheung","Sizhe Gao","Siddharth Suresh","Junjie Hu","Timothy T. Rogers"],"categories":["cs.AI","cs.HC"],"primary_category":"cs.AI","announce_type":"new","date":"2025-01-28","first_seen":"2025-01-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2501.17310","pdf_url":"https://arxiv.org/pdf/2501.17310","source_feed":"backfill","score":7,"bucket":"pending","rubric_hits":["A1","B1","B2"],"tags":["LLM仿真","群体智慧","人类对照"],"reason":"用LLM复现人类估计任务，并与真实人类数据对照，涉及选举预测等政策场景，方法可…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:30","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":91,"question":"LLM在猜测估计任务中是否表现出类似人类的“群体智慧”效应，以及如何利用该效应提升估计准确性？","design":"研究使用10个LLM（LLaMA、Mistral、Mixtral、GPT等）作为被试，通过重复采样生成多个估计值，应用WOC解码（取中位数）与其他解码策略对比，测量归一化误差；同时构建了三个猜测估计数据集（MARBLES、FUTURE、ELECPRED），并进行了人类实验以复现WOC效应。","baseline":"人类实验：招募人类参与者完成MARBLES数据集中的估计任务，复现了人类群体中的WOC效应，作为LLM行为的对照基准。","findings":"LLM在猜测估计任务中表现出与人类相似的WOC效应：对多次采样响应取中位数能持续提升准确性，优于贪婪解码、自一致性解码和均值解码。这表明LLM编码了支持近似推理的世界模型。","reliability":"论文未讨论","relevance":"该研究将LLM作为人类被试的替代品，复现了群体智慧效应，并与真实人类数据对照，涉及选举预测等政策场景，方法可用于评估LLM仿真人类估计行为的可靠性，值得精读。","inspiration":"该方法通过重复采样LLM生成多个估计值并取中位数来复现群体智慧效应，可借鉴其解码策略作为处理手段，并设置贪婪解码、均值解码等作为对照｜可迁移到资产定价实验中的市场预期形成，例如研究投资者对股票收益的集体预测是否表现出群体智慧｜以LLM作为被试，处理为对同一股票收益预测多次采样后取中位数，结果变量为预测误差，以真实分析师一致预期数据作为人类基准对照"}},{"id":"2501.14294","version":3,"title":"Examining Alignment of Large Language Models through Representative Heuristics: The Case of Political Stereotypes","zh_title":"通过代表性启发式检验大语言模型的对齐：以政治刻板印象为例","abstract":"Examining the alignment of large language models (LLMs) has become increasingly important, e.g., when LLMs fail to operate as intended. This study examines the alignment of LLMs with human values for the domain of politics. Prior research has shown that LLM-generated outputs can include political leanings and mimic the stances of political parties on various issues. However, the extent and conditions under which LLMs deviate from empirical positions are insufficiently examined. To address this gap, we analyze the factors that contribute to LLMs' deviations from empirical positions on political issues, aiming to quantify these deviations and identify the conditions that cause them. Drawing on findings from cognitive science about representativeness heuristics, i.e., situations where humans lean on representative attributes of a target group in a way that leads to exaggerated beliefs, we scrutinize LLM responses through this heuristics' lens. We conduct experiments to determine how LLMs inflate predictions about political parties, which results in stereotyping. We find that while LLMs can mimic certain political parties' positions, they often exaggerate these positions more than human survey respondents do. Also, LLMs tend to overemphasize representativeness more than humans. This study highlights the susceptibility of LLMs to representativeness heuristics, suggesting a potential vulnerability of LLMs that facilitates political stereotyping. We also test prompt-based mitigation strategies, finding that strategies that can mitigate representative heuristics in humans are also effective in reducing the influence of representativeness on LLM-generated responses.","authors":["Sullam Jeoung","Yubin Ge","Haohan Wang","Jana Diesner"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2025-01-24","first_seen":"2025-01-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2501.14294","pdf_url":"https://arxiv.org/pdf/2501.14294","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["LLM人类仿真","政治态度模拟","代表性启发式"],"reason":"用LLM模拟人类政治态度并与调查数据对照，发现LLM夸大刻板印象，评估仿真偏差…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:14:55","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":37,"question":"LLM在政治议题上是否会像人类一样因代表性启发式而产生刻板印象，并偏离真实立场？","design":"用LLM（如GPT系列）模拟美国民主党和共和党的立场，通过双问题框架：实证部分使用自认党派的人类调查数据，预测部分让LLM和人类分别回答政治议题，比较LLM与人类在预测党派立场时的偏差程度，并测试提示词缓解策略。","baseline":"自认为民主党或共和党的真实人类调查数据（公开的现有调查数据）。","findings":"LLM能近似各党派立场，但比人类更夸大党派差异，表现出更强的代表性启发式；提示词干预可部分缓解这种刻板印象，但无法完全消除。","reliability":"论文未讨论","relevance":"该研究直接以真实人类调查为基准，评估LLM在政治态度模拟中的偏差，揭示LLM比人类更易产生刻板印象，符合你对仿真可靠性与失效条件的关注，值得精读。","inspiration":"该方法通过双问题框架（实证部分用真实人类数据，预测部分让LLM和人类分别回答）来量化LLM的启发式偏差，值得借鉴｜可迁移到信贷审批歧视研究，检验LLM是否像人类一样因代表性启发式而夸大种族或性别的信用差异｜让LLM模拟信贷员审批贷款，处理为申请人种族/性别信息，结果变量为审批通过率，以真实信贷审批数据中的人类决策分布为基准，比较LLM与人类的偏差程度"}},{"id":"2501.13955","version":1,"title":"Guided Persona-based AI Surveys: Can we replicate personal mobility preferences at scale using LLMs?","zh_title":"基于引导式角色的AI调查：能否利用LLM大规模复制个人出行偏好？","abstract":"This study explores the potential of Large Language Models (LLMs) to generate artificial surveys, with a focus on personal mobility preferences in Germany. By leveraging LLMs for synthetic data creation, we aim to address the limitations of traditional survey methods, such as high costs, inefficiency and scalability challenges. A novel approach incorporating \"Personas\" - combinations of demographic and behavioural attributes - is introduced and compared to five other synthetic survey methods, which vary in their use of real-world data and methodological complexity. The MiD 2017 dataset, a comprehensive mobility survey in Germany, serves as a benchmark to assess the alignment of synthetic data with real-world patterns. The results demonstrate that LLMs can effectively capture complex dependencies between demographic attributes and preferences while offering flexibility to explore hypothetical scenarios. This approach presents valuable opportunities for transportation planning and social science research, enabling scalable, cost-efficient and privacy-preserving data generation.","authors":["Ioannis Tzachristas","Santhanakrishnan Narayanan","Constantinos Antoniou"],"categories":["cs.CL","cs.AI","cs.CY"],"primary_category":"cs.CL","announce_type":"new","date":"2025-01-20","first_seen":"2025-01-20","revised_at":null,"abs_url":"https://arxiv.org/abs/2501.13955","pdf_url":"https://arxiv.org/pdf/2501.13955","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM人类仿真","合成调查数据","出行偏好"],"reason":"用LLM生成合成调查数据模拟人类出行偏好，并与真实调查数据MiD 2017对照…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:14:54","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":9,"question":"能否利用大语言模型通过基于人格的引导式AI调查方法，大规模复制德国个人出行偏好？","design":"使用GPT-4o生成合成调查数据，提出六种方法（朴素AI调查、结构化AI调查、引导式AI调查、朴素基于人格的AI调查、结构化基于人格的AI调查、引导式基于人格的AI调查），逐步引入真实人口结构和出行统计约束，生成10000个个体或15840个独特人格的出行偏好回答。","baseline":"德国2017年国家出行调查（MiD 2017）数据集，包含人口分布、交通方式和出行频率等真实数据。","findings":"引导式基于人格的AI调查方法在准确性和与真实出行行为模式的一致性上显著优于其他方法；LLM能有效捕捉人口属性与偏好之间的复杂依赖关系，并支持灵活探索假设情景。","reliability":"论文未讨论","relevance":"该研究直接以真实人类调查数据为基准，评估LLM仿真出行偏好的可靠性，并比较多种方法，符合研究者对经济学实验和政策评估场景下仿真有效性及偏差的关注，值得阅读原文。","inspiration":"该方法通过逐步引入真实人口统计约束和出行统计约束来校准LLM生成回答，可借鉴为在仿真中分层施加宏观分布约束以提升个体决策模拟的准确性｜可迁移到消费者跨期选择与储蓄行为仿真，利用LLM模拟不同人口特征群体的时间偏好和消费-储蓄决策｜以LLM作为被试，处理为提供不同利率或未来收入情景，结果变量为报告的消费/储蓄金额，用家庭金融调查（如SCF）的真实储蓄率分布作为对照基准"}},{"id":"2501.08579","version":3,"title":"LLM-based Human Simulations Have Not Yet Been Reliable","zh_title":"基于大语言模型的人类仿真尚未可靠","abstract":"Large Language Models (LLMs) are increasingly employed for simulating human behaviors across diverse domains. However, our position is that current LLM-based human simulations remain insufficiently reliable, as evidenced by significant discrepancies between their outcomes and authentic human actions. Our investigation begins with a systematic review of LLM-based human simulations in social, economic, policy, and psychological contexts, identifying their common frameworks, recent advances, and persistent limitations. This review reveals that such discrepancies primarily stem from inherent limitations of LLMs and flaws in simulation design, both of which are examined in detail. Building on these insights, we propose a systematic solution framework that emphasizes enriching data foundations, advancing LLM capabilities, and ensuring robust simulation design to enhance reliability. Finally, we introduce a structured algorithm that operationalizes the proposed framework, aiming to guide credible and human-aligned LLM-based simulations. To facilitate further research, we provide a curated list of related literature and resources at https://github.com/Persdre/awesome-llm-human-simulation.","authors":["Qian Wang","Jiaying Wu","Zichen Jiang","Zhenheng Tang","Bingqiao Luo","Nuo Chen","Wei Chen","Huacan Wang","Bingsheng He"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2025-01-15","first_seen":"2025-01-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2501.08579","pdf_url":"https://arxiv.org/pdf/2501.08579","source_feed":"api","score":9,"bucket":"selected","rubric_hits":["A1","A2","A4","B1","B4"],"tags":["LLM人类仿真","可靠性评估","方法论框架"],"reason":"系统评估LLM人类仿真的可靠性，指出与真实人类行为的差异，并提出改进框架，直接…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:14:52","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":69,"question":"当前基于大语言模型的人类仿真是否可靠？","design":"本文并非仿真实验，而是对社会科学、经济学、政策、心理学等领域中基于LLM的人类仿真研究进行系统综述，分析其常用框架、进展与局限。","baseline":"无对照","findings":"当前LLM人类仿真不可靠，与真实人类行为存在显著差异；差异主要源于LLM的固有局限（如偏见、认知过程缺陷、行为不一致）和仿真框架设计缺陷（如过度简化心理状态、缺乏全面人类经验、验证机制不足）。","reliability":"论文指出失效条件包括：LLM内嵌的文化、性别等偏见扭曲行为模拟；认知过程局限损害决策真实性；记忆限制导致行为不一致；交互机制缺陷影响多智能体仿真；框架过度简化复杂心理状态和群体动态；缺乏严格验证与实时监控。","relevance":"该文直接回应研究者对LLM人类仿真可靠性的核心关切，系统梳理了仿真失效的根源并提出改进框架，对评估经济学实验和政策模拟的仿真偏差具有重要参考价值，强烈推荐阅读原文。","inspiration":"该文系统梳理了LLM仿真失效的根源（偏见、认知局限、验证不足），其批判性框架可借鉴用于设计经济学实验中的稳健性检验，例如通过对比不同LLM版本或提示策略来识别仿真偏差｜可迁移至政策公告的预期形成研究，评估LLM模拟的经济主体对货币政策或财政刺激的反应是否与真实调查数据一致｜以LLM作为被试，施加不同措辞的政策声明处理，测量其通胀预期或消费意愿，并与密歇根大学消费者调查等真实微观数据对照，检验仿真偏差"}},{"id":"2501.07663","version":1,"title":"Enhancing Talent Employment Insights Through Feature Extraction with LLM Finetuning","zh_title":"通过LLM微调的特征提取增强人才就业洞察","abstract":"This paper explores the application of large language models (LLMs) to extract nuanced and complex job features from unstructured job postings. Using a dataset of 1.2 million job postings provided by AdeptID, we developed a robust pipeline to identify and classify variables such as remote work availability, remuneration structures, educational requirements, and work experience preferences. Our methodology combines semantic chunking, retrieval-augmented generation (RAG), and fine-tuning DistilBERT models to overcome the limitations of traditional parsing tools. By leveraging these techniques, we achieved significant improvements in identifying variables often mislabeled or overlooked, such as non-salary-based compensation and inferred remote work categories. We present a comprehensive evaluation of our fine-tuned models and analyze their strengths, limitations, and potential for scaling. This work highlights the promise of LLMs in labor market analytics, providing a foundation for more accurate and actionable insights into job data.","authors":["Karishma Thakrar","Nick Young"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2025-01-13","first_seen":"2025-01-13","revised_at":null,"abs_url":"https://arxiv.org/abs/2501.07663","pdf_url":"https://arxiv.org/pdf/2501.07663","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["NLP","信息抽取","劳动力市场分析"],"reason":"用LLM提取招聘信息特征，属NLP信息抽取，不涉及人类行为仿真或对照。","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:59:27","error":null,"has_summary":false,"summary":null},{"id":"2501.06834","version":1,"title":"LLMs Model Non-WEIRD Populations: Experiments with Synthetic Cultural Agents","zh_title":"LLM模拟非WEIRD人群：合成文化代理实验","abstract":"Despite its importance, studying economic behavior across diverse, non-WEIRD (Western, Educated, Industrialized, Rich, and Democratic) populations presents significant challenges. We address this issue by introducing a novel methodology that uses Large Language Models (LLMs) to create synthetic cultural agents (SCAs) representing these populations. We subject these SCAs to classic behavioral experiments, including the dictator and ultimatum games. Our results demonstrate substantial cross-cultural variability in experimental behavior. Notably, for populations with available data, SCAs' behaviors qualitatively resemble those of real human subjects. For unstudied populations, our method can generate novel, testable hypotheses about economic behavior. By integrating AI into experimental economics, this approach offers an effective and ethical method to pilot experiments and refine protocols for hard-to-reach populations. Our study provides a new tool for cross-cultural economic studies and demonstrates how LLMs can help experimental behavioral research.","authors":["Augusto Gonzalez-Bonorino","Monica Capra","Emilio Pantoja"],"categories":["cs.AI","cs.CL","econ.GN"],"primary_category":"cs.AI","announce_type":"new","date":"2025-01-12","first_seen":"2025-01-12","revised_at":null,"abs_url":"https://arxiv.org/abs/2501.06834","pdf_url":"https://arxiv.org/pdf/2501.06834","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM仿真","行为实验","跨文化经济研究"],"reason":"用LLM创建合成文化代理模拟非WEIRD人群的经济实验行为，并与真实人类数据对…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:14:52","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":28,"question":"如何利用大语言模型创建合成文化代理，以模拟非WEIRD人群在经典经济实验中的行为？","design":"使用GPT-4等大语言模型，结合网络爬取和检索增强生成构建六支小规模社会（Hadza、Machiguenga、Tsimané、Aché、Orma、Yanomami）的文化档案，据此实例化合成文化代理，让其参与独裁者博弈、最后通牒博弈和禀赋效应实验，测量出价、拒绝率等行为变量。","baseline":"对照Henrich等人（2005）等文献中真实人类在相同实验中的行为数据，对部分有数据的人群进行定性比较。","findings":"合成代理的行为展现出显著的跨文化差异，且与现有真实人类数据在性质上相似；所有代理均未表现出纯粹自利行为，该方法还能为未经研究的人群生成可检验的假设。","reliability":"论文承认LLM可能带有固有偏见，且合成代理的行为仍需与真实人类数据进行谨慎验证，该方法旨在补充而非替代人类被试研究。","relevance":"高度相关：该研究直接用LLM模拟非WEIRD人群的经济决策，并与真实人类实验数据对照，属于经济学实验场景下的人类仿真，同时讨论了方法的局限，完全契合你的关注点，值得精读原文。","inspiration":"该方法借鉴了用LLM结合文化档案构建合成代理来模拟特定人群决策行为的做法，通过检索增强生成注入文化背景，并与真实人类实验数据进行定性对照｜可迁移到跨文化消费信贷审批歧视研究，模拟不同文化背景申请人的还款决策与银行审批行为｜以GPT-4等LLM为被试，构建不同文化背景的合成申请人，处理为信贷条款（如利率、额度），结果变量为还款意愿与违约率，对照真实跨国信贷数据或田野实验数据"}},{"id":"2501.00745","version":3,"title":"Dynamics of Adversarial Attacks on Large Language Model-Based Search Engines","zh_title":"基于大语言模型的搜索引擎对抗攻击动力学","abstract":"The increasing integration of Large Language Model (LLM) based search engines has transformed the landscape of information retrieval. However, these systems are vulnerable to adversarial attacks, especially ranking manipulation attacks, where attackers craft webpage content to manipulate the LLM's ranking and promote specific content, gaining an unfair advantage over competitors. In this paper, we study the dynamics of ranking manipulation attacks. We frame this problem as an Infinitely Repeated Prisoners' Dilemma, where multiple players strategically decide whether to cooperate or attack. We analyze the conditions under which cooperation can be sustained, identifying key factors such as attack costs, discount rates, attack success rates, and trigger strategies that influence player behavior. We identify tipping points in the system dynamics, demonstrating that cooperation is more likely to be sustained when players are forward-looking. However, from a defense perspective, we find that simply reducing attack success probabilities can, paradoxically, incentivize attacks under certain conditions. Furthermore, defensive measures to cap the upper bound of attack success rates may prove futile in some scenarios. These insights highlight the complexity of securing LLM-based systems. Our work provides a theoretical foundation and practical insights for understanding and mitigating their vulnerabilities, while emphasizing the importance of adaptive security strategies and thoughtful ecosystem design.","authors":["Xiyang Hu"],"categories":["cs.CL","cs.AI","cs.GT","cs.IR","econ.TH"],"primary_category":"cs.CL","announce_type":"new","date":"2025-01-01","first_seen":"2025-01-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2501.00745","pdf_url":"https://arxiv.org/pdf/2501.00745","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["对抗攻击","博弈论","搜索引擎"],"reason":"研究多智能体在无限重复囚徒困境中的攻击策略，不涉及人类行为仿真或对照。","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:59:05","error":null,"has_summary":false,"summary":null},{"id":"2412.19363","version":3,"title":"Large Language Models for Market Research: A Data-augmentation Approach","zh_title":"用于市场研究的大语言模型：一种数据增强方法","abstract":"Large Language Models (LLMs) have transformed artificial intelligence by excelling in complex natural language processing tasks. Their ability to generate human-like text has opened new possibilities for market research, particularly in conjoint analysis, where understanding consumer preferences is essential but often resource-intensive. Traditional survey-based methods face limitations in scalability and cost, making LLM-generated data a promising alternative. However, while LLMs have the potential to simulate real consumer behavior, recent studies highlight a significant gap between LLM-generated and human data, with biases introduced when substituting between the two. In this paper, we address this gap by proposing a novel statistical data augmentation approach that efficiently integrates LLM-generated data with real data in conjoint analysis. This results in statistically robust estimators with consistent and asymptotically normal properties, in contrast to naive approaches that simply substitute human data with LLM-generated data, which can exacerbate bias. We further present a finite-sample performance bound on the estimation error. We validate our framework through an empirical study on COVID-19 vaccine preferences, demonstrating its superior ability to reduce estimation error and save data and costs by 24.9% to 79.8%. In contrast, naive approaches fail to save data due to the inherent biases in LLM-generated data compared to human data. Another empirical study on sports car choices validates the robustness of our results. Our findings suggest that while LLM-generated data is not a direct substitute for human responses, it can serve as a valuable complement when used within a robust statistical framework.","authors":["Mengxin Wang","Dennis J. Zhang","Heng Zhang"],"categories":["cs.AI","cs.LG","stat.ME","stat.ML"],"primary_category":"cs.AI","announce_type":"new","date":"2024-12-26","first_seen":"2024-12-26","revised_at":null,"abs_url":"https://arxiv.org/abs/2412.19363","pdf_url":"https://arxiv.org/pdf/2412.19363","source_feed":"backfill","score":8,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["LLM仿真","联合分析","数据增强"],"reason":"用LLM生成联合分析数据并与真实人类数据对照，评估偏差并提出统计校正方法，涉及…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:14:52","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":63,"question":"能否通过统计方法有效整合LLM生成数据与真实数据，以改进联合分析中消费者偏好的估计精度？","design":"本研究提出一种统计增强方法，将LLM生成的联合分析选择数据作为辅助信息，与真实人类数据结合，构建AI增强估计量（AAE），并推导其渐近性质与有限样本误差界。","baseline":"真实人类数据来自两项实证研究：COVID-19疫苗偏好联合分析调查和跑车选择联合分析调查。","findings":"直接混合LLM生成数据与真实数据会加剧偏差，而所提AAE方法能显著降低估计误差，并在疫苗偏好研究中节省24.9%至79.8%的数据收集成本。LLM生成数据不能直接替代人类回答，但作为统计框架内的补充信息具有价值。","reliability":"论文指出LLM缺乏真实生活体验，且消费者偏好随时间变化，LLM可能无法准确捕捉；即使使用最先进的提示工程技术，LLM与人类回答间的差距仍持续存在，直接替代会导致误导性结果。","relevance":"高度相关：该研究将LLM作为人类被试的替代品进行联合分析实验，并与真实人类数据严格对照，评估了直接替代的偏差，并提出校正方法，直接回应了研究者对仿真可靠性、失效条件及经济学实验场景的关注。","inspiration":"该研究提出AI增强估计量（AAE），将LLM生成数据作为辅助信息而非直接替代，通过统计校正降低偏差，这种处理混杂与测量误差的思路值得借鉴｜可迁移到消费者金融产品选择偏好的联合分析中，如贷款条款偏好或保险计划选择，以降低调研成本并校正LLM仿真偏差｜以真实消费者为被试，处理为不同贷款属性组合，结果变量为选择决策，用LLM生成的选择数据作为辅助，构建AAE估计量，并与纯真实数据估计的偏好参数对照，评估成本节省与偏差校正效果"}},{"id":"2412.19245","version":1,"title":"Sentiment trading with large language models","zh_title":"基于大语言模型的情感交易研究","abstract":"We investigate the efficacy of large language models (LLMs) in sentiment analysis of U.S. financial news and their potential in predicting stock market returns. We analyze a dataset comprising 965,375 news articles that span from January 1, 2010, to June 30, 2023; we focus on the performance of various LLMs, including BERT, OPT, FINBERT, and the traditional Loughran-McDonald dictionary model, which has been a dominant methodology in the finance literature. The study documents a significant association between LLM scores and subsequent daily stock returns. Specifically, OPT, which is a GPT-3 based LLM, shows the highest accuracy in sentiment prediction with an accuracy of 74.4%, slightly ahead of BERT (72.5%) and FINBERT (72.2%). In contrast, the Loughran-McDonald dictionary model demonstrates considerably lower effectiveness with only 50.1% accuracy. Regression analyses highlight a robust positive impact of OPT model scores on next-day stock returns, with coefficients of 0.274 and 0.254 in different model specifications. BERT and FINBERT also exhibit predictive relevance, though to a lesser extent. Notably, we do not observe a significant relationship between the Loughran-McDonald dictionary model scores and stock returns, challenging the efficacy of this traditional method in the current financial context. In portfolio performance, the long-short OPT strategy excels with a Sharpe ratio of 3.05, compared to 2.11 for BERT and 2.07 for FINBERT long-short strategies. Strategies based on the Loughran-McDonald dictionary yield the lowest Sharpe ratio of 1.23. Our findings emphasize the superior performance of advanced LLMs, especially OPT, in financial market prediction and portfolio management, marking a significant shift in the landscape of financial analysis tools with implications to financial regulation and policy analysis.","authors":["Kemal Kirtac","Guido Germano"],"categories":["q-fin.CP","cs.LG","econ.EM","q-fin.PM","q-fin.TR"],"primary_category":"q-fin.CP","announce_type":"new","date":"2024-12-26","first_seen":"2024-12-26","revised_at":null,"abs_url":"https://arxiv.org/abs/2412.19245","pdf_url":"https://arxiv.org/pdf/2412.19245","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["情感分析","金融预测","NLP评测"],"reason":"纯金融NLP评测，用LLM做情感分析预测股价，不以人类行为仿真为目标","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:58:57","error":null,"has_summary":false,"summary":null},{"id":"2412.18497","version":2,"title":"Neuron-Level Differentiation of Memorization and Generalization in Large Language Models","zh_title":"大型语言模型中记忆与泛化的神经元级分化","abstract":"We investigate how Large Language Models (LLMs) distinguish between memorization and generalization at the neuron level. Through carefully designed tasks, we identify distinct neuron subsets responsible for each behavior. Experiments on both a GPT-2 model trained from scratch and a pretrained LLaMA-3.2 model fine-tuned with LoRA show consistent neuron-level specialization. We further demonstrate that inference-time interventions on these neurons can steer the model's behavior toward memorization or generalization. To assess robustness, we evaluate intra-task and inter-task consistency, confirming that these neuron-behavior associations reflect generalizable patterns rather than dataset-specific artifacts. Our findings reveal modular structure in LLMs and enable controlling memorization and generalization behaviors at inference time.","authors":["Ko-Wei Huang","Yi-Fu Fu","Ching-Yu Tsai","Yu-Chieh Tu","Tzu-Ling Cheng","Cheng-Yu Lin","Yi-Ting Yang","Heng-Yi Liu","Keng-Te Liao","Da-Cheng Juan","Shou-De Lin"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2024-12-24","first_seen":"2024-12-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2412.18497","pdf_url":"https://arxiv.org/pdf/2412.18497","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["神经元分析","记忆与泛化","模型可解释性"],"reason":"研究LLM内部神经元对记忆与泛化的分工，属纯NLP能力分析，不以人类行为为参照。","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:59:19","error":null,"has_summary":false,"summary":null},{"id":"2412.18061","version":1,"title":"Lla-VAP: LSTM Ensemble of Llama and VAP for Turn-Taking Prediction","zh_title":"Lla-VAP：用于轮次预测的 Llama 与 VAP 的 LSTM 集成","abstract":"Turn-taking prediction is the task of anticipating when the speaker in a conversation will yield their turn to another speaker to begin speaking. This project expands on existing strategies for turn-taking prediction by employing a multi-modal ensemble approach that integrates large language models (LLMs) and voice activity projection (VAP) models. By combining the linguistic capabilities of LLMs with the temporal precision of VAP models, we aim to improve the accuracy and efficiency of identifying TRPs in both scripted and unscripted conversational scenarios. Our methods are evaluated on the In-Conversation Corpus (ICC) and Coached Conversational Preference Elicitation (CCPE) datasets, highlighting the strengths and limitations of current models while proposing a potentially more robust framework for enhanced prediction.","authors":["Hyunbae Jeon","Frederic Guintu","Rayvant Sahni"],"categories":["cs.SD","cs.CL","cs.HC","eess.AS"],"primary_category":"cs.SD","announce_type":"new","date":"2024-12-24","first_seen":"2024-12-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2412.18061","pdf_url":"https://arxiv.org/pdf/2412.18061","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["对话系统","轮次预测","多模态集成"],"reason":"研究对话轮次预测，属于对话系统技术，非人类仿真实验，无测量或实验目的。","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:59:45","error":null,"has_summary":false,"summary":null},{"id":"2412.16265","version":3,"title":"Autoware.Flex: Human-Instructed Dynamically Reconfigurable Autonomous Driving Systems","zh_title":"Autoware.Flex：人类指令驱动的动态可重构自动驾驶系统","abstract":"Existing Autonomous Driving Systems (ADS) independently make driving decisions, but they face two significant limitations. First, in complex scenarios, ADS may misinterpret the environment and make inappropriate driving decisions. Second, these systems are unable to incorporate human driving preferences in their decision-making processes. This paper proposes Autoware$.$Flex, a novel ADS system that incorporates human input into the driving process, allowing users to guide the ADS in making more appropriate decisions and ensuring their preferences are satisfied. Achieving this needs to address two key challenges: (1) translating human instructions, expressed in natural language, into a format the ADS can understand, and (2) ensuring these instructions are executed safely and consistently within the ADS' s decision-making framework. For the first challenge, we employ a Large Language Model (LLM) assisted by an ADS-specialized knowledge base to enhance domain-specific translation. For the second challenge, we design a validation mechanism to ensure that human instructions result in safe and consistent driving behavior. Experiments conducted on both simulators and a real-world autonomous vehicle demonstrate that Autoware$.$Flex effectively interprets human instructions and executes them safely.","authors":["Ziwei Song","Mingsong Lv","Tianchi Ren","Chun Jason Xue","Jen-Ming Wu","Nan Guan"],"categories":["cs.AI","cs.HC","cs.RO"],"primary_category":"cs.AI","announce_type":"new","date":"2024-12-20","first_seen":"2024-12-20","revised_at":null,"abs_url":"https://arxiv.org/abs/2412.16265","pdf_url":"https://arxiv.org/pdf/2412.16265","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C2"],"tags":["自动驾驶","人机交互","LLM指令翻译"],"reason":"自动驾驶系统仿真环境，不涉及用LLM替代人类被试进行行为实验或对照研究。","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:59:21","error":null,"has_summary":false,"summary":null},{"id":"2412.14161","version":3,"title":"TheAgentCompany: Benchmarking LLM Agents on Consequential Real World Tasks","zh_title":"TheAgentCompany：在重要现实世界任务上评测LLM智能体","abstract":"We interact with computers on an everyday basis, be it in everyday life or work, and many aspects of work can be done entirely with access to a computer and the Internet. At the same time, thanks to improvements in large language models (LLMs), there has also been a rapid development in AI agents that interact with and affect change in their surrounding environments. But how performant are AI agents at accelerating or even autonomously performing work-related tasks? The answer to this question has important implications both for industry looking to adopt AI into their workflows and for economic policy to understand the effects that adoption of AI may have on the labor market. To measure the progress of these LLM agents' performance on performing real-world professional tasks, in this paper we introduce TheAgentCompany, an extensible benchmark for evaluating AI agents that interact with the world in similar ways to those of a digital worker: by browsing the Web, writing code, running programs, and communicating with other coworkers. We build a self-contained environment with internal web sites and data that mimics a small software company environment, and create a variety of tasks that may be performed by workers in such a company. We test baseline agents powered by both closed API-based and open-weights language models (LMs), and find that the most competitive agent can complete 30% of tasks autonomously. This paints a nuanced picture on task automation with LM agents--in a setting simulating a real workplace, a good portion of simpler tasks could be solved autonomously, but more difficult long-horizon tasks are still beyond the reach of current systems. We release code, data, environment, and experiments on https://the-agent-company.com.","authors":["Frank F. Xu","Yufan Song","Boxuan Li","Yuxuan Tang","Kritanjali Jain","Mengxue Bao","Zora Z. Wang","Xuhui Zhou","Zhitong Guo","Murong Cao","Mingyang Yang","Hao Yang Lu","Amaad Martin","Zhe Su","Leander Maben","Raj Mehta","Wayne Chi","Lawrence Jang","Yiqing Xie","Shuyan Zhou","Graham Neubig"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2024-12-18","first_seen":"2024-12-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2412.14161","pdf_url":"https://arxiv.org/pdf/2412.14161","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["LLM智能体","基准测试","任务自动化"],"reason":"纯多智能体系统评测，agent在模拟公司环境执行数字工作任务，无人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:59:28","error":null,"has_summary":false,"summary":null},{"id":"2412.10635","version":1,"title":"Do LLMs Act as Repositories of Causal Knowledge?","zh_title":"大语言模型是否充当因果知识库？","abstract":"Large language models (LLMs) offer the potential to automate a large number of tasks that previously have not been possible to automate, including some in science. There is considerable interest in whether LLMs can automate the process of causal inference by providing the information about causal links necessary to build a structural model. We use the case of confounding in the Coronary Drug Project (CDP), for which there are several studies listing expert-selected confounders that can serve as a ground truth. LLMs exhibit mediocre performance in identifying confounders in this setting, even though text about the ground truth is in their training data. Variables that experts identify as confounders are only slightly more likely to be labeled as confounders by LLMs compared to variables that experts consider non-confounders. Further, LLM judgment on confounder status is highly inconsistent across models, prompts, and irrelevant concerns like multiple-choice option ordering. LLMs do not yet have the ability to automate the reporting of causal links.","authors":["Nick Huntington-Klein","Eleanor J. Murray"],"categories":["econ.EM"],"primary_category":"econ.EM","announce_type":"new","date":"2024-12-14","first_seen":"2024-12-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2412.10635","pdf_url":"https://arxiv.org/pdf/2412.10635","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["因果推断","LLM评测","混杂因子识别"],"reason":"评估LLM识别混杂因子的能力，属于纯NLP能力评测，不以人类行为仿真为参照。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:14:51","error":null,"has_summary":false,"summary":null},{"id":"2412.07031","version":4,"title":"Large Language Models: An Applied Econometric Framework","zh_title":"大语言模型：一个应用计量经济学框架","abstract":"Large language models (LLMs) enable researchers to analyze text at unprecedented scale and minimal cost. Researchers can now revisit old questions and tackle novel ones with rich data. We provide an econometric framework for realizing this potential in two empirical uses. For prediction problems -- forecasting outcomes from text -- valid conclusions require ``no training leakage'' between the LLM's training data and the researcher's sample, which can be enforced through careful model choice and research design. For estimation problems -- automating the measurement of economic concepts for downstream analysis -- valid downstream inference requires combining LLM outputs with a small validation sample to deliver consistent and precise estimates. Absent a validation sample, researchers cannot assess possible errors in LLM outputs, and consequently seemingly innocuous choices (which model, which prompt) can produce dramatically different parameter estimates. When used appropriately, LLMs are powerful tools that can expand the frontier of empirical economics.","authors":["Jens Ludwig","Sendhil Mullainathan","Ashesh Rambachan"],"categories":["econ.EM","cs.AI"],"primary_category":"econ.EM","announce_type":"new","date":"2024-12-09","first_seen":"2024-12-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2412.07031","pdf_url":"https://arxiv.org/pdf/2412.07031","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM标注","计量方法","文本分析"],"reason":"LLM替代人工标注测量经济概念，属标注员替代而非仿真被试，但方法可迁移。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:14:50","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":202,"question":"如何将大语言模型（LLM）的输出纳入实证经济学研究，以保证预测和估计问题的有效推断？","design":"本文并非直接进行人类仿真实验，而是提出一个计量经济学框架，将LLM用于两类任务：预测问题（用文本预测经济结果）和估计问题（用LLM自动测量文本中的经济概念以替代人工标注）。在估计问题中，LLM扮演标注员的角色，通过少量验证样本校正其测量误差，从而放大验证样本的信息效率。","baseline":"无对照（本文未进行仿真实验，而是讨论方法框架；在估计问题中，真实测量来自人工标注或验证样本，但未提供具体人类基准数据集）。","findings":"在预测问题中，有效推断要求LLM训练数据与研究者样本无重叠（无训练泄漏），可通过模型选择和设计实现。在估计问题中，缺乏验证样本时，LLM输出误差未知，模型和提示词的微小变化会导致参数估计在量级、符号和显著性上剧烈波动；结合少量验证样本进行去偏后，LLM输出能显著提高估计精度，但不能完全替代人工数据。","reliability":"论文指出，若无验证样本，研究者无法评估LLM输出的误差及其对下游参数估计的影响，且不同LLM和提示词选择会导致截然不同的结论；预测问题中若存在训练泄漏，则评估结果反映的不是样本外表现。","relevance":"本文虽聚焦LLM作为标注工具而非人类被试替代，但其估计问题框架可直接迁移至人类仿真研究：将LLM模拟的受试者视为有误差的测量，需用少量真实人类样本校正，否则仿真结果不可靠。对关注仿真失效条件的研究者具有重要参考价值，建议阅读原文。","inspiration":"借鉴其估计问题框架：将LLM输出视为有误差的测量，需用少量真实样本校正偏差，否则模型和提示词的微小变化会导致结论剧烈波动。｜可迁移至政策公告的预期形成实验，研究不同措辞的央行沟通如何影响公众通胀预期。｜以LLM模拟公众被试，处理为不同风格的货币政策声明，结果变量为LLM生成的通胀预期数值，用真实调查数据（如密歇根消费者调查）作为基准校正仿真偏差。"}},{"id":"2412.02065","version":3,"title":"Leveraging Large Language Models to Democratize Access to Costly Datasets for Academic Research","zh_title":"利用大语言模型使昂贵数据集获取民主化以促进学术研究","abstract":"Unequal access to costly datasets essential for empirical research has long hindered researchers from disadvantaged institutions, limiting their ability to contribute to their fields and advance their careers. Recent breakthroughs in Large Language Models (LLMs) have the potential to democratize data access by automating data collection from unstructured sources. We develop and evaluate a novel methodology using GPT-4o-mini within a Retrieval-Augmented Generation (RAG) framework to collect data from corporate disclosures. Our approach achieves human-level accuracy in collecting CEO pay ratios from approximately 10,000 proxy statements and Critical Audit Matters (CAMs) from more than 12,000 10-K filings, with LLM processing times of 9 and 40 minutes respectively, each at a cost under US $10. This stands in stark contrast to the hundreds of hours needed for manual collection or the thousands of dollars required for commercial database subscriptions. To foster a more inclusive research community by empowering researchers with limited resources to explore new avenues of inquiry, we share our methodology and the resulting datasets.","authors":["Julian Junyan Wang","Victor Xiaoqi Wang"],"categories":["q-fin.GN","cs.AI","cs.CE","cs.LG","econ.GN"],"primary_category":"q-fin.GN","announce_type":"new","date":"2024-12-03","first_seen":"2024-12-03","revised_at":null,"abs_url":"https://arxiv.org/abs/2412.02065","pdf_url":"https://arxiv.org/pdf/2412.02065","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["LLM数据提取","信息抽取","研究资源民主化"],"reason":"用LLM从文本中提取数据，替代人工收集，属于NLP信息抽取，不涉及人类行为仿真。","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:58:42","error":null,"has_summary":false,"summary":null},{"id":"2412.01069","version":2,"title":"The Promise and Peril of Generative AI: Evidence from GPT as Sell-Side Analysts","zh_title":"生成式AI的前景与风险：来自GPT作为卖方分析师的证据","abstract":"Large language models (LLMs) promise to democratize financial analysis by reducing information-processing costs. Yet equal access does not ensure equal outcomes, as the locus of friction may shift from processing information to evaluating model outputs. We study GPT's earnings forecasts following corporate earnings releases and document two patterns. First, GPT's narrative attention is consistent and human-like but not always associated with higher forecast accuracy. Second, its quantitative reasoning varies substantially across contexts, challenging the view that LLMs are uniformly weak at numerical tasks. Building on these insights, we propose a diagnostic framework that links forecast accuracy to observable processing features (i.e., narrative focus, numerical reasoning, and self-assessed confidence). These indicators serve as proxies for this new form of information friction and alert investors when to exercise caution. Our study has implications for information frictions, regulatory oversight, and the economics of AI-mediated financial markets.","authors":["Edward Li","Min Shen","Zhiyuan Tu","Dexin Zhou"],"categories":["q-fin.GN","econ.GN"],"primary_category":"q-fin.GN","announce_type":"new","date":"2024-12-02","first_seen":"2024-12-02","revised_at":null,"abs_url":"https://arxiv.org/abs/2412.01069","pdf_url":"https://arxiv.org/pdf/2412.01069","source_feed":"backfill","score":7,"bucket":"pending","rubric_hits":["A1","B1","B2"],"tags":["LLM仿真","金融预测","人类对照"],"reason":"用GPT模拟卖方分析师预测，并与人类分析师对照，涉及金融决策仿真，方法可迁移。","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:58:42","error":null,"has_summary":false,"summary":null},{"id":"2411.11581","version":5,"title":"OASIS: Open Agent Social Interaction Simulations with One Million Agents","zh_title":"OASIS：百万代理的开放代理社交互动模拟","abstract":"There has been a growing interest in enhancing rule-based agent-based models (ABMs) for social media platforms (i.e., X, Reddit) with more realistic large language model (LLM) agents, thereby allowing for a more nuanced study of complex systems. As a result, several LLM-based ABMs have been proposed in the past year. While they hold promise, each simulator is specifically designed to study a particular scenario, making it time-consuming and resource-intensive to explore other phenomena using the same ABM. Additionally, these models simulate only a limited number of agents, whereas real-world social media platforms involve millions of users. To this end, we propose OASIS, a generalizable and scalable social media simulator. OASIS is designed based on real-world social media platforms, incorporating dynamically updated environments (i.e., dynamic social networks and post information), diverse action spaces (i.e., following, commenting), and recommendation systems (i.e., interest-based and hot-score-based). Additionally, OASIS supports large-scale user simulations, capable of modeling up to one million users. With these features, OASIS can be easily extended to different social media platforms to study large-scale group phenomena and behaviors. We replicate various social phenomena, including information spreading, group polarization, and herd effects across X and Reddit platforms. Moreover, we provide observations of social phenomena at different agent group scales. We observe that the larger agent group scale leads to more enhanced group dynamics and more diverse and helpful agents' opinions. These findings demonstrate OASIS's potential as a powerful tool for studying complex systems in digital environments.","authors":["Ziyi Yang","Zaibin Zhang","Zirui Zheng","Yuxian Jiang","Ziyue Gan","Zhiyu Wang","Zijian Ling","Jinsong Chen","Martz Ma","Bowen Dong","Prateek Gupta","Shuyue Hu","Zhenfei Yin","Guohao Li","Xu Jia","Lijun Wang","Bernard Ghanem","Huchuan Lu","Chaochao Lu","Wanli Ouyang","Yu Qiao","Philip Torr","Jing Shao"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2024-11-18","first_seen":"2024-11-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2411.11581","pdf_url":"https://arxiv.org/pdf/2411.11581","source_feed":"backfill","score":7,"bucket":"pending","rubric_hits":["A3","B1"],"tags":["社会模拟","LLM代理","信息传播"],"reason":"用LLM代理模拟社交媒体信息传播和从众效应，并与真实人类数据对照，可迁移到人类…","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:41:50","error":null,"has_summary":false,"summary":null},{"id":"2411.06790","version":2,"title":"Large-scale moral machine experiment on large language models","zh_title":"基于大语言模型的大规模道德机器实验","abstract":"The rapid advancement of Large Language Models (LLMs) and their potential integration into autonomous driving systems necessitates understanding their moral decision-making capabilities. While our previous study examined four prominent LLMs using the Moral Machine experimental framework, the dynamic landscape of LLM development demands a more comprehensive analysis. Here, we evaluate moral judgments across 52 different LLMs, including multiple versions of proprietary models (GPT, Claude, Gemini) and open-source alternatives (Llama, Gemma), to assess their alignment with human moral preferences in autonomous driving scenarios. Using a conjoint analysis framework, we evaluated how closely LLM responses aligned with human preferences in ethical dilemmas and examined the effects of model size, updates, and architecture. Results showed that proprietary models and open-source models exceeding 10 billion parameters demonstrated relatively close alignment with human judgments, with a significant negative correlation between model size and distance from human judgments in open-source models. However, model updates did not consistently improve alignment with human preferences, and many LLMs showed excessive emphasis on specific ethical principles. These findings suggest that while increasing model size may naturally lead to more human-like moral judgments, practical implementation in autonomous driving systems requires careful consideration of the trade-off between judgment quality and computational efficiency. Our comprehensive analysis provides crucial insights for the ethical design of autonomous systems and highlights the importance of considering cultural contexts in AI moral decision-making.","authors":["Muhammad Shahrul Zaim bin Ahmad","Kazuhiro Takemoto"],"categories":["cs.CY","cs.CL","cs.HC"],"primary_category":"cs.CY","announce_type":"new","date":"2024-11-11","first_seen":"2024-11-11","revised_at":null,"abs_url":"https://arxiv.org/abs/2411.06790","pdf_url":"https://arxiv.org/pdf/2411.06790","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D2","C2"],"tags":["LLM道德判断","自动驾驶伦理","人类对齐"],"reason":"测量LLM的道德判断，非仿真人类被试；场景为自动驾驶，属C2排除项，但有人类数…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:14:50","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":201,"question":"不同大语言模型在自动驾驶道德困境中的道德判断与人类偏好的对齐程度如何？","design":"本研究并非将LLM作为人类被试的仿真，而是直接测量52个LLM（包括GPT、Claude、Gemini、Llama等）在道德机器框架下的道德选择。通过受约束随机生成包含物种、社会价值、性别、年龄、健身、功利主义等维度的两难场景，要求模型在不可避免事故中二选一，并采用联合分析评估其与人类偏好的距离。","baseline":"对照的真实人类数据来自原始道德机器实验（Moral Machine experiment）中人类参与者的道德偏好模式。","findings":"专有模型和参数超过100亿的开源模型与人类判断对齐较好，开源模型中模型规模与人类判断距离呈显著负相关。模型更新并未持续改善与人类偏好的对齐，许多LLM表现出对特定伦理原则的过度强调。","reliability":"论文承认道德机器框架存在方法局限，如场景设计约束、文化代表性不足、理论选择与实际部署间的差距，并指出实际自动驾驶系统部署需权衡判断质量与计算效率。","relevance":"高度相关：该研究使用真实人类道德偏好作为基准，系统评估了LLM在道德决策上的对齐程度，并揭示了模型规模、更新与架构的影响，直接回应了研究者对LLM仿真可靠性及失效条件的关注。","inspiration":"该研究通过受约束随机生成多维度道德困境场景，并采用联合分析评估LLM与人类偏好的距离，为多属性决策仿真提供了严谨的测量框架｜可迁移至金融伦理决策研究，如算法信贷审批中的公平性权衡（效率vs.平等）或投资顾问的ESG偏好冲突｜以LLM作为被试，随机生成包含申请人收入、种族、性别、信用分等属性的贷款审批场景，要求模型做出批准/拒绝决策，结果变量为决策中的属性权重，以真实信贷员决策数据或公平借贷审计数据作为人类基准对照"}},{"id":"2411.05194","version":1,"title":"Interactive Dialogue Agents via Reinforcement Learning on Hindsight Regenerations","zh_title":"基于事后重写的强化学习交互式对话代理","abstract":"Recent progress on large language models (LLMs) has enabled dialogue agents to generate highly naturalistic and plausible text. However, current LLM language generation focuses on responding accurately to questions and requests with a single effective response. In reality, many real dialogues are interactive, meaning an agent's utterances will influence their conversational partner, elicit information, or change their opinion. Accounting for how an agent can effectively steer a conversation is a crucial ability in many dialogue tasks, from healthcare to preference elicitation. Existing methods for fine-tuning dialogue agents to accomplish such tasks would rely on curating some amount of expert data. However, doing so often requires understanding the underlying cognitive processes of the conversational partner, which is a skill neither humans nor LLMs trained on human data can reliably do. Our key insight is that while LLMs may not be adept at identifying effective strategies for steering conversations a priori, or in the middle of an ongoing conversation, they can do so post-hoc, or in hindsight, after seeing how their conversational partner responds. We use this fact to rewrite and augment existing suboptimal data, and train via offline reinforcement learning (RL) an agent that outperforms both prompting and learning from unaltered human demonstrations. We apply our approach to two domains that require understanding human mental state, intelligent interaction, and persuasion: mental health support, and soliciting charitable donations. Our results in a user study with real humans show that our approach greatly outperforms existing state-of-the-art dialogue agents.","authors":["Joey Hong","Jessica Lin","Anca Dragan","Sergey Levine"],"categories":["cs.LG","cs.AI","cs.CL"],"primary_category":"cs.LG","announce_type":"new","date":"2024-11-07","first_seen":"2024-11-07","revised_at":null,"abs_url":"https://arxiv.org/abs/2411.05194","pdf_url":"https://arxiv.org/pdf/2411.05194","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["对话代理","强化学习","说服任务"],"reason":"该研究训练对话代理以影响人类，但并非用LLM仿真人类被试，而是优化代理策略，属…","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:59:45","error":null,"has_summary":false,"summary":null},{"id":"2410.22203","version":1,"title":"Democratizing Reward Design for Personal and Representative Value-Alignment","zh_title":"个性化与代表性价值对齐的奖励设计民主化","abstract":"Aligning AI agents with human values is challenging due to diverse and subjective notions of values. Standard alignment methods often aggregate crowd feedback, which can result in the suppression of unique or minority preferences. We introduce Interactive-Reflective Dialogue Alignment, a method that iteratively engages users in reflecting on and specifying their subjective value definitions. This system learns individual value definitions through language-model-based preference elicitation and constructs personalized reward models that can be used to align AI behaviour. We evaluated our system through two studies with 30 participants, one focusing on \"respect\" and the other on ethical decision-making in autonomous vehicles. Our findings demonstrate diverse definitions of value-aligned behaviour and show that our system can accurately capture each person's unique understanding. This approach enables personalized alignment and can inform more representative and interpretable collective alignment strategies.","authors":["Carter Blair","Kate Larson","Edith Law"],"categories":["cs.AI","cs.HC"],"primary_category":"cs.AI","announce_type":"new","date":"2024-10-29","first_seen":"2024-10-29","revised_at":null,"abs_url":"https://arxiv.org/abs/2410.22203","pdf_url":"https://arxiv.org/pdf/2410.22203","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["价值对齐","个性化奖励模型","人机交互"],"reason":"个性化价值对齐与角色扮演，非人类仿真实验，无群体行为对照。","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:59:47","error":null,"has_summary":false,"summary":null},{"id":"2410.19599","version":3,"title":"Take Caution in Using LLMs as Human Surrogates: Scylla Ex Machina","zh_title":"谨慎使用LLM作为人类替代品：Scylla Ex Machina","abstract":"Recent studies suggest large language models (LLMs) can exhibit human-like reasoning, aligning with human behavior in economic experiments, surveys, and political discourse. This has led many to propose that LLMs can be used as surrogates or simulations for humans in social science research. However, LLMs differ fundamentally from humans, relying on probabilistic patterns, absent the embodied experiences or survival objectives that shape human cognition. We assess the reasoning depth of LLMs using the 11-20 money request game. Nearly all advanced approaches fail to replicate human behavior distributions across many models. Causes of failure are diverse and unpredictable, relating to input language, roles, and safeguarding. These results advise caution when using LLMs to study human behavior or as surrogates or simulations.","authors":["Yuan Gao","Dokyun Lee","Gordon Burtch","Sina Fazelpour"],"categories":["econ.GN","cs.AI","cs.CY","cs.HC"],"primary_category":"econ.GN","announce_type":"new","date":"2024-10-25","first_seen":"2024-10-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2410.19599","pdf_url":"https://arxiv.org/pdf/2410.19599","source_feed":"backfill","score":10,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM人类仿真","行为博弈","算法保真度"],"reason":"直接评估LLM作为人类替代品的可靠性，使用11-20金钱请求游戏与真实人类行为…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:14:50","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":2,"question":"大语言模型在简单经济博弈中能否可靠地复现人类行为分布，作为人类替代品用于社会科学研究？","design":"本研究使用11-20金钱请求博弈，测试了GPT-4、GPT-3.5、Claude3-Opus、Claude3-Sonnet、Llama3-70b、Llama3-8b、Llama2-13b、Llama2-7b等8个LLM，每个模型收集1000次干净会话，并尝试了提示工程、检索增强生成、微调等高级技术，比较模型选择分布与人类分布。","baseline":"人类基准来自Arad和Rubinstein (2012)原始论文中报告的人类参与者行为分布及纳什均衡预测。","findings":"几乎所有LLM和高级方法均未能复现人类行为分布，失败原因多样且不可预测，涉及输入语言、角色设定和安全防护等；即使微调GPT-4o能模仿特定人类数据，但在新变体或分布外场景下仍失效。","reliability":"论文指出LLM行为不稳定，对提示措辞、语言、角色等高度敏感；模型可能依赖记忆而非真正推理；在分布外场景下普遍失败；微调仅能模仿已知模式，无法泛化。","relevance":"该研究直接评估LLM作为人类替代品的可靠性，使用经济学实验与真实人类数据对照，并系统揭示失效条件，高度契合研究者对批判性仿真研究的关注，值得精读原文。","inspiration":"借鉴其系统对比多个LLM与真实人类行为分布的方法，并引入提示工程、微调等稳健性检验来揭示仿真失效条件｜可迁移到政策公告的预期形成实验，检验LLM能否复现人类对财政或货币政策信号的反应分布｜以GPT-4o等为被试，给予不同措辞的政策公告作为处理，测量其通胀或就业预期分布，并以真实调查数据（如密歇根消费者预期调查）为基准对照"}},{"id":"2410.10665","version":1,"title":"Double Jeopardy and Climate Impact in the Use of Large Language Models: Socio-economic Disparities and Reduced Utility for Non-English Speakers","zh_title":"使用大语言模型的双重困境与气候影响：非英语使用者的社会经济差距与效用降低","abstract":"Artificial Intelligence (AI), particularly large language models (LLMs), holds the potential to bridge language and information gaps, which can benefit the economies of developing nations. However, our analysis of FLORES-200, FLORES+, Ethnologue, and World Development Indicators data reveals that these benefits largely favor English speakers. Speakers of languages in low-income and lower-middle-income countries face higher costs when using OpenAI's GPT models via APIs because of how the system processes the input -- tokenization. Around 1.5 billion people, speaking languages primarily from lower-middle-income countries, could incur costs that are 4 to 6 times higher than those faced by English speakers. Disparities in LLM performance are significant, and tokenization in models priced per token amplifies inequalities in access, cost, and utility. Moreover, using the quality of translation tasks as a proxy measure, we show that LLMs perform poorly in low-resource languages, presenting a ``double jeopardy\" of higher costs and poor performance for these users. We also discuss the direct impact of fragmentation in tokenizing low-resource languages on climate. This underscores the need for fairer algorithm development to benefit all linguistic groups.","authors":["Aivin V. Solatorio","Gabriel Stefanini Vicente","Holly Krambeck","Olivier Dupriez"],"categories":["cs.CL","cs.AI","cs.LG","econ.GN"],"primary_category":"cs.CL","announce_type":"new","date":"2024-10-14","first_seen":"2024-10-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2410.10665","pdf_url":"https://arxiv.org/pdf/2410.10665","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["LLM公平性","语言资源不均","NLP评测"],"reason":"研究LLM在不同语言上的成本和性能差异，属于NLP评测，不以人类行为仿真为目标。","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:58:42","error":null,"has_summary":false,"summary":null},{"id":"2409.19430","version":1,"title":"'Simulacrum of Stories': Examining Large Language Models as Qualitative Research Participants","zh_title":"“故事的拟像”：审视大语言模型作为定性研究参与者","abstract":"The recent excitement around generative models has sparked a wave of proposals suggesting the replacement of human participation and labor in research and development--e.g., through surveys, experiments, and interviews--with synthetic research data generated by large language models (LLMs). We conducted interviews with 19 qualitative researchers to understand their perspectives on this paradigm shift. Initially skeptical, researchers were surprised to see similar narratives emerge in the LLM-generated data when using the interview probe. However, over several conversational turns, they went on to identify fundamental limitations, such as how LLMs foreclose participants' consent and agency, produce responses lacking in palpability and contextual depth, and risk delegitimizing qualitative research methods. We argue that the use of LLMs as proxies for participants enacts the surrogate effect, raising ethical and epistemological concerns that extend beyond the technical limitations of current models to the core of whether LLMs fit within qualitative ways of knowing.","authors":["Shivani Kapania","William Agnew","Motahhare Eslami","Hoda Heidari","Sarah Fox"],"categories":["cs.HC","cs.CL","cs.LG"],"primary_category":"cs.HC","announce_type":"new","date":"2024-09-28","first_seen":"2024-09-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2409.19430","pdf_url":"https://arxiv.org/pdf/2409.19430","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A4","B4"],"tags":["LLM仿真","定性研究","方法论批判"],"reason":"直接研究用LLM替代定性研究参与者，并识别仿真失效条件，高度相关。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:14:48","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":27,"question":"定性研究者如何看待用大语言模型模拟研究参与者进行访谈？","design":"本研究并非仿真实验，而是对19位定性研究者进行半结构化访谈，并让他们使用一个访谈探针工具，将自己以往的人类访谈数据与LLM生成的模拟访谈数据进行对比，收集他们的看法和反思。","baseline":"研究者自己过去项目中收集的真实人类访谈记录。","findings":"研究者起初惊讶于LLM能生成与人类相似的叙述，但多轮对话后发现LLM回应缺乏切身感和情境深度，且会剥夺参与者的同意权和能动性。LLM作为参与者代理会产生“替代效应”，引发超越技术局限的伦理与认识论问题，可能损害定性研究方法的合法性。","reliability":"论文指出LLM回应缺乏 palpability（切身感）、模型的认识论立场模糊、强化研究者立场、剥夺参与者同意与能动性、抹除社区视角、以及可能使定性研究方法失去合法性。这些局限根植于LLM与诠释主义定性认识论的根本不兼容。","relevance":"该研究直接探讨用LLM替代人类进行定性访谈的仿真实践，并基于真实人类数据对比，识别出仿真失效的深层条件，高度契合研究者对LLM仿真可靠性及批判性研究的兴趣，值得精读原文。","inspiration":"借鉴该研究将LLM生成内容与真实人类记录进行对比的评估框架，可用于检验LLM仿真经济决策的效度。｜可迁移到消费者跨期选择实验，考察LLM生成的消费-储蓄决策是否与人类行为一致。｜以LLM作为被试，施加不同利率或未来收入预期的处理，测量其跨期消费分配，并与真实家庭调查数据（如PSID）中的消费-储蓄模式进行对照。"}},{"id":"2409.14202","version":3,"title":"Mining Causality: AI-Assisted Search for Instrumental Variables","zh_title":"挖掘因果关系：人工智能辅助的工具变量搜索","abstract":"The instrumental variables (IVs) method is a leading empirical strategy for causal inference. Finding IVs is a heuristic and creative process, and justifying its validity -- especially exclusion restrictions -- is largely rhetorical. We propose using large language models (LLMs) to search for new IVs through narratives and counterfactual reasoning, similar to how a human researcher would. The stark difference, however, is that LLMs can dramatically accelerate this process and explore an extremely large search space. We demonstrate how to construct prompts to search for potentially valid IVs. We contend that multi-step and role-playing prompting strategies are effective for simulating the endogenous decision-making processes of economic agents and for navigating language models through the realm of real-world scenarios, rather than anchoring them within the narrow realm of academic discourses on IVs. We apply our method to three well-known examples in economics: returns to schooling, supply and demand, and peer effects. We then extend our strategy to finding (i) control variables in regression and difference-in-differences and (ii) running variables in regression discontinuity designs.","authors":["Sukjin Han"],"categories":["econ.EM","stat.AP","stat.ME","stat.ML"],"primary_category":"econ.EM","announce_type":"new","date":"2024-09-21","first_seen":"2024-09-21","revised_at":null,"abs_url":"https://arxiv.org/abs/2409.14202","pdf_url":"https://arxiv.org/pdf/2409.14202","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D3"],"tags":["工具变量","LLM模拟","因果推断"],"reason":"用LLM模拟经济主体决策以寻找工具变量，但无真实人类行为对照，属社会模拟边界情…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:14:48","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":200,"question":"如何利用大语言模型通过叙事和反事实推理来系统性地搜索经济学中的工具变量？","design":"使用GPT-4，通过多步骤和角色扮演提示策略，模拟经济主体的内生决策过程，在真实世界场景中搜索潜在工具变量，并应用于教育回报、供需和同伴效应三个经典例子。","baseline":"无对照","findings":"GPT-4在三个经典例子中均生成了候选工具变量列表，其中包含文献中常用的工具变量和一些看似新颖的变量，并提供了合理性论证；该方法还可扩展用于搜索控制变量和断点回归中的运行变量。","reliability":"论文未讨论","relevance":"该研究用LLM模拟经济主体决策以寻找工具变量，属于用AI辅助社会科学研究中的创造性过程，但缺乏真实人类行为对照，不符合你关注的以人类被试为基准的仿真实验，相关性较低，不建议优先阅读。","inspiration":"该方法利用LLM的角色扮演和反事实推理来生成经济学中的工具变量，可借鉴其多步骤提示策略来辅助识别因果效应｜可迁移到政策评估场景，例如估计某项教育政策对长期收入的因果效应时，用LLM模拟政策制定者或经济主体的决策逻辑来搜索潜在工具变量｜以LLM作为辅助工具，让GPT-4模拟不同经济主体（如家庭、学校）在政策冲击下的行为，生成候选工具变量列表，再使用真实观测数据（如面板调查数据）检验这些工具变量的相关性和外生性，并与传统工具变量（如政策规则变化）进行对比"}},{"id":"2409.18988","version":1,"title":"A Unified Framework to Classify Business Activities into International Standard Industrial Classification through Large Language Models for Circular Economy","zh_title":"通过大语言模型将商业活动分类到国际标准产业分类的统一框架以促进循环经济","abstract":"Effective information gathering and knowledge codification are pivotal for developing recommendation systems that promote circular economy practices. One promising approach involves the creation of a centralized knowledge repository cataloguing historical waste-to-resource transactions, which subsequently enables the generation of recommendations based on past successes. However, a significant barrier to constructing such a knowledge repository lies in the absence of a universally standardized framework for representing business activities across disparate geographical regions. To address this challenge, this paper leverages Large Language Models (LLMs) to classify textual data describing economic activities into the International Standard Industrial Classification (ISIC), a globally recognized economic activity classification framework. This approach enables any economic activity descriptions provided by businesses worldwide to be categorized into the unified ISIC standard, facilitating the creation of a centralized knowledge repository. Our approach achieves a 95% accuracy rate on a 182-label test dataset with fine-tuned GPT-2 model. This research contributes to the global endeavour of fostering sustainable circular economy practices by providing a standardized foundation for knowledge codification and recommendation systems deployable across regions.","authors":["Xiang Li","Lan Zhao","Junhao Ren","Yajuan Sun","Chuan Fu Tan","Zhiquan Yeo","Gaoxi Xiao"],"categories":["cs.CL","cs.AI","econ.GN"],"primary_category":"cs.CL","announce_type":"new","date":"2024-09-17","first_seen":"2024-09-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2409.18988","pdf_url":"https://arxiv.org/pdf/2409.18988","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["文本分类","循环经济","LLM应用"],"reason":"用LLM做经济活动分类，属于纯NLP能力评测，不以人类行为为参照系。","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:58:44","error":null,"has_summary":false,"summary":null},{"id":"2409.10750","version":1,"title":"GPT takes the SAT: Tracing changes in Test Difficulty and Math Performance of Students","zh_title":"GPT参加SAT：追踪试题难度与学生数学表现的变化","abstract":"Scholastic Aptitude Test (SAT) is crucial for college admissions but its effectiveness and relevance are increasingly questioned. This paper enhances Synthetic Control methods by introducing \"Transformed Control\", a novel method that employs Large Language Models (LLMs) powered by Artificial Intelligence to generate control groups. We utilize OpenAI's API to generate a control group where GPT-4, or ChatGPT, takes multiple SATs annually from 2008 to 2023. This control group helps analyze shifts in SAT math difficulty over time, starting from the baseline year of 2008. Using parallel trends, we calculate the Average Difference in Scores (ADS) to assess changes in high school students' math performance. Our results indicate a significant decrease in the difficulty of the SAT math section over time, alongside a decline in students' math performance. The analysis shows a 71-point drop in the rigor of SAT math from 2008 to 2023, with student performance decreasing by 36 points, resulting in a 107-point total divergence in average student math performance. We investigate possible mechanisms for this decline in math proficiency, such as changing university selection criteria, increased screen time, grade inflation, and worsening adolescent mental health. Disparities among demographic groups show a 104-point drop for White students, 84 points for Black students, and 53 points for Asian students. Male students saw a 117-point reduction, while female students had a 100-point decrease.","authors":["Vikram Krishnaveti","Saannidhya Rawat"],"categories":["econ.EM"],"primary_category":"econ.EM","announce_type":"new","date":"2024-09-16","first_seen":"2024-09-16","revised_at":null,"abs_url":"https://arxiv.org/abs/2409.10750","pdf_url":"https://arxiv.org/pdf/2409.10750","source_feed":"backfill","score":7,"bucket":"pending","rubric_hits":["A1","B1","B2"],"tags":["LLM仿真","教育评估","人类数据对照"],"reason":"用GPT-4模拟学生考SAT，与真实学生成绩对照，评估试题难度变化，属于教育评…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:14:48","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":124,"question":"SAT数学部分的难度是否随时间下降，以及学生数学表现是否随之变化？","design":"用GPT-4作为合成控制组，每年参加2008至2023年的SAT数学考试，生成基准分数；通过平行趋势和平均分数差（ADS）分析试题难度变化，并对比真实学生成绩评估学生数学表现。","baseline":"真实学生SAT数学成绩数据，包括总体及按种族、性别划分的分数。","findings":"2008至2023年SAT数学难度下降71分，学生数学表现下降36分，总差距达107分；不同族裔和性别群体表现下降程度不同，白人学生下降104分，黑人84分，亚裔53分，男性117分，女性100分。","reliability":"论文未讨论","relevance":"该研究用GPT-4模拟学生参加标准化考试，与真实学生成绩对照，评估试题难度变化，属于教育评估场景的LLM仿真实验，有真实人类基准，值得细读其合成控制方法和可靠性讨论。","inspiration":"该方法用GPT-4作为合成控制组，通过平行趋势和平均分数差分析试题难度变化，为评估政策或环境变化提供了可借鉴的因果推断框架。｜可迁移到教育经济学中评估考试政策改革（如考试内容调整）对学生表现的影响，或劳动经济学中分析技能需求变化。｜用GPT-4模拟求职者参加不同年份的职业能力测试，处理为测试年份，结果变量为测试分数，以真实求职者历史分数数据为对照，评估试题难度和技能需求变化。"}},{"id":"2409.08357","version":2,"title":"An Experimental Study of Competitive Market Behavior Through LLMs","zh_title":"通过大语言模型对竞争市场行为的实验研究","abstract":"This study explores the potential of large language models (LLMs) to conduct market experiments, aiming to understand their capability to comprehend competitive market dynamics. We model the behavior of market agents in a controlled experimental setting, assessing their ability to converge toward competitive equilibria. The results reveal the challenges current LLMs face in replicating the dynamic decision-making processes characteristic of human trading behavior. Unlike humans, LLMs lacked the capacity to achieve market equilibrium. The research demonstrates that while LLMs provide a valuable tool for scalable and reproducible market simulations, their current limitations necessitate further advancements to fully capture the complexities of market behavior. Future work that enhances dynamic learning capabilities and incorporates elements of behavioral economics could improve the effectiveness of LLMs in the economic domain, providing new insights into market dynamics and aiding in the refinement of economic policies.","authors":["Jingru Jia","Zehua Yuan"],"categories":["cs.HC","cs.AI","econ.GN"],"primary_category":"cs.HC","announce_type":"new","date":"2024-09-12","first_seen":"2024-09-12","revised_at":null,"abs_url":"https://arxiv.org/abs/2409.08357","pdf_url":"https://arxiv.org/pdf/2409.08357","source_feed":"backfill","score":7,"bucket":"pending","rubric_hits":["A3","B2"],"tags":["LLM仿真","市场实验","行为经济学"],"reason":"用LLM模拟市场竞争行为并与人类对照，涉及经济学实验，但未明确提及真实人类数据…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:14:48","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":123,"question":"LLM能否在双拍卖实验中模拟竞争市场行为并收敛到均衡价格？","design":"使用ChatGPT-4.0扮演22个市场代理人（11买家、11卖家），在双拍卖框架下进行5轮交易，记录出价、要价和成交价，分析价格收敛、波动性和策略适应。","baseline":"对照Smith (1962)的人类双拍卖实验，人类被试价格逐渐收敛到理论均衡价$2.00。","findings":"LLM驱动的交易价格围绕均衡价$2.00波动，未呈现向均衡收敛的趋势；LLM缺乏人类交易中的动态学习和实时反馈适应能力。","reliability":"论文指出LLM的局限在于缺乏自适应学习和实时反馈机制，无法达到市场均衡，需要增强动态学习能力并融入行为经济学元素。","relevance":"该研究直接使用LLM模拟经济学实验中的市场行为，并与经典人类实验数据对照，揭示了LLM在动态决策任务中的失效，与研究者关注的经济学实验仿真和可靠性批判高度契合，值得精读。","inspiration":"借鉴其将LLM作为市场代理人参与双拍卖实验并与经典人类实验（Smith, 1962）严格对照的设计，通过比较价格收敛路径和波动性来检验仿真有效性。｜可迁移到资产定价实验，研究LLM能否在连续双向拍卖中复现人类交易者的价格发现过程与泡沫形成。｜用LLM扮演交易者参与资产市场实验，处理为不同信息结构（如对称/不对称信息），结果变量为价格偏差、交易量与泡沫持续时间，对照真实人类实验数据（如Smith et al., 1988的资产泡沫实验）。"}},{"id":"2409.00128","version":3,"title":"Can Large Language Models Replace Human Subjects? A Large-Scale Replication of Scenario-Based Experiments in Psychology and Management","zh_title":"大语言模型能替代人类被试吗？心理学与管理学场景实验的大规模复现","abstract":"Artificial Intelligence (AI) is increasingly being integrated into scientific research, particularly in the social sciences, where understanding human behavior is critical. Large Language Models (LLMs) have shown promise in replicating human-like responses in various psychological experiments. We conducted a large-scale study replicating 156 psychological experiments from top social science journals using three state-of-the-art LLMs (GPT-4, Claude 3.5 Sonnet, and DeepSeek v3). Our results reveal that while LLMs demonstrate high replication rates for main effects (73-81%) and moderate to strong success with interaction effects (46-63%), They consistently produce larger effect sizes than human studies, with Fisher Z values approximately 2-3 times higher than human studies. Notably, LLMs show significantly lower replication rates for studies involving socially sensitive topics such as race, gender and ethics. When original studies reported null findings, LLMs produced significant results at remarkably high rates (68-83%) - while this could reflect cleaner data with less noise, as evidenced by narrower confidence intervals, it also suggests potential risks of effect size overestimation. Our results demonstrate both the promise and challenges of LLMs in psychological research, offering efficient tools for pilot testing and rapid hypothesis validation while enriching rather than replacing traditional human subject studies, yet requiring more nuanced interpretation and human validation for complex social phenomena and culturally sensitive research questions.","authors":["Ziyan Cui","Ning Li","Huaikang Zhou"],"categories":["cs.CL","cs.AI","econ.GN"],"primary_category":"cs.CL","announce_type":"new","date":"2024-08-29","first_seen":"2024-08-29","revised_at":null,"abs_url":"https://arxiv.org/abs/2409.00128","pdf_url":"https://arxiv.org/pdf/2409.00128","source_feed":"backfill","score":10,"bucket":"selected","rubric_hits":["A1","A2","A5","B1","B2","B4"],"tags":["LLM仿真","人类被试替代","心理学实验复现"],"reason":"直接复现156项心理学实验，用LLM替代人类被试，有真实人类数据对照，评估可靠…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:14:46","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":7,"question":"大语言模型能否在多大程度上替代人类被试，复现心理学和管理学中的场景实验？","design":"使用GPT-4、Claude 3.5 Sonnet和DeepSeek v3三个LLM，将156项已发表心理学实验的原始文本材料直接呈现给模型，每个实验生成与原始人类样本量相等的模型回答，测量主效应和交互效应的复制率、效应量及p值分布。","baseline":"原始人类实验的真实数据，来自五本顶级管理学和心理学期刊的156项随机选取的场景实验。","findings":"LLM对主效应的复制率达73-81%，交互效应复制率为46-63%，但效应量普遍是人类研究的2-3倍；在涉及种族、性别等社会敏感话题时复制率显著下降，且对原研究中的零结果有68-83%的概率产生显著结果。","reliability":"LLM在涉及社会敏感话题（如种族、性别、伦理）时复制率大幅降低，可能因模型的价值对齐导致社会期望偏差；效应量系统性放大，可能增加I类错误风险；研究仅限于文本情景实验，未涉及其他实验范式。","relevance":"该研究直接以大规模真实人类实验为基准，系统评估LLM替代人类被试的可靠性与偏差，与您关注的核心问题高度吻合，值得精读原文。","inspiration":"可借鉴其大规模系统复制框架和按实验特征（如敏感话题）分层分析偏差的方法｜可迁移到信贷审批中的种族/性别歧视研究或政策信息处理实验｜用LLM模拟贷款审批员，处理含不同种族/性别线索的申请材料，测量审批决策和风险感知，以真实银行历史审批数据或审计研究结果作为对照基准。"}},{"id":"2408.05328","version":1,"title":"From Text to Insight: Leveraging Large Language Models for Performance Evaluation in Management","zh_title":"从文本到洞察：利用大语言模型进行管理中的绩效评估","abstract":"This study explores the potential of Large Language Models (LLMs), specifically GPT-4, to enhance objectivity in organizational task performance evaluations. Through comparative analyses across two studies, including various task performance outputs, we demonstrate that LLMs can serve as a reliable and even superior alternative to human raters in evaluating knowledge-based performance outputs, which are a key contribution of knowledge workers. Our results suggest that GPT ratings are comparable to human ratings but exhibit higher consistency and reliability. Additionally, combined multiple GPT ratings on the same performance output show strong correlations with aggregated human performance ratings, akin to the consensus principle observed in performance evaluation literature. However, we also find that LLMs are prone to contextual biases, such as the halo effect, mirroring human evaluative biases. Our research suggests that while LLMs are capable of extracting meaningful constructs from text-based data, their scope is currently limited to specific forms of performance evaluation. By highlighting both the potential and limitations of LLMs, our study contributes to the discourse on AI role in management studies and sets a foundation for future research to refine AI theoretical and practical applications in management.","authors":["Ning Li","Huaikang Zhou","Mingze Xu"],"categories":["cs.CL","cs.AI","cs.ET","cs.HC","econ.GN"],"primary_category":"cs.CL","announce_type":"new","date":"2024-08-09","first_seen":"2024-08-09","revised_at":null,"abs_url":"https://arxiv.org/abs/2408.05328","pdf_url":"https://arxiv.org/pdf/2408.05328","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM评估","绩效评估","人工标注替代"],"reason":"用LLM替代人类评估者进行绩效评分，属于替代人工标注而非仿真人类被试，但有人类…","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:58:44","error":null,"has_summary":false,"summary":null},{"id":"2407.04467","version":3,"title":"Are Large Language Models Strategic Decision Makers? A Study of Performance and Bias in Two-Player Non-Zero-Sum Games","zh_title":"大语言模型是战略决策者吗？双人非零和博弈中的表现与偏差研究","abstract":"Large Language Models (LLMs) have been increasingly used in real-world settings, yet their strategic decision-making abilities remain largely unexplored. To fully benefit from the potential of LLMs, it's essential to understand their ability to function in complex social scenarios. Game theory, which is already used to understand real-world interactions, provides a good framework for assessing these abilities. This work investigates the performance and merits of LLMs in canonical game-theoretic two-player non-zero-sum games, Stag Hunt and Prisoner Dilemma. Our structured evaluation of GPT-3.5, GPT-4-Turbo, GPT-4o, and Llama-3-8B shows that these models, when making decisions in these games, are affected by at least one of the following systematic biases: positional bias, payoff bias, or behavioural bias. This indicates that LLMs do not fully rely on logical reasoning when making these strategic decisions. As a result, it was found that the LLMs' performance drops when the game configuration is misaligned with the affecting biases. When misaligned, GPT-3.5, GPT-4-Turbo, GPT-4o, and Llama-3-8B show an average performance drop of 32\\%, 25\\%, 34\\%, and 29\\% respectively in Stag Hunt, and 28\\%, 16\\%, 34\\%, and 24\\% respectively in Prisoner's Dilemma. Surprisingly, GPT-4o (a top-performing LLM across standard benchmarks) suffers the most substantial performance drop, suggesting that newer models are not addressing these issues. Interestingly, we found that a commonly used method of improving the reasoning capabilities of LLMs, chain-of-thought (CoT) prompting, reduces the biases in GPT-3.5, GPT-4o, and Llama-3-8B but increases the effect of the bias in GPT-4-Turbo, indicating that CoT alone cannot fully serve as a robust solution to this problem. We perform several additional experiments, which provide further insight into these observed behaviours.","authors":["Nathan Herr","Fernando Acero","Roberta Raileanu","María Pérez-Ortiz","Zhibin Li"],"categories":["cs.AI","cs.CL","cs.GT"],"primary_category":"cs.AI","announce_type":"new","date":"2024-07-05","first_seen":"2024-07-05","revised_at":null,"abs_url":"https://arxiv.org/abs/2407.04467","pdf_url":"https://arxiv.org/pdf/2407.04467","source_feed":"backfill","score":8,"bucket":"selected","rubric_hits":["A1","A2","B2","B4"],"tags":["LLM仿真","博弈论","决策偏差"],"reason":"用LLM模拟人类在博弈中的决策，评估偏差，涉及行为博弈场景，但未明确提及真实人…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:14:46","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":33,"question":"LLM在经典两人非零和博弈中是否存在系统性偏差，这些偏差如何影响其策略决策表现？","design":"以GPT-3.5、GPT-4-Turbo、GPT-4o和Llama-3-8B作为被试，通过改变博弈矩阵中行动标签的顺序（位置偏差）、收益结构（收益偏差）或行为倾向（行为偏差）来构造不同配置的猎鹿博弈和囚徒困境，测量模型选择合作或背叛等行动的准确率变化。","baseline":"无对照","findings":"所有模型均受至少一种系统性偏差影响，当博弈配置与偏差不一致时，模型表现平均下降16%-34%，其中GPT-4o下降最严重。思维链提示能减少部分模型的偏差，但对GPT-4-Turbo反而加剧偏差，表明其并非稳健解决方案。","reliability":"论文指出思维链提示无法完全消除偏差，且未在更复杂的多轮博弈或真实交互场景中验证，也未与人类行为直接对比。","relevance":"该研究直接评估LLM在策略互动中的决策偏差，虽未使用真实人类基准，但揭示了LLM作为人类仿真代理在博弈场景中的系统性失效模式，对关注LLM仿真可靠性的研究者有参考价值。","inspiration":"可借鉴其通过系统操纵博弈矩阵标签顺序和收益结构来检测位置偏差与收益偏差的方法，用于评估LLM在策略环境中的稳健性。｜可迁移至经济政策博弈模拟，如碳税谈判或贸易协定中的策略行为仿真。｜以LLM作为多国谈判代表，随机化提案顺序和收益矩阵，测量合作率，并与人类实验数据（如公开的博弈实验数据集）进行对照。"}},{"id":"2407.03859","version":3,"title":"Anthropocentric bias in language model evaluation","zh_title":"语言模型评估中的人类中心偏见","abstract":"Evaluating the cognitive capacities of large language models (LLMs) requires overcoming not only anthropomorphic but also anthropocentric biases. This article identifies two types of anthropocentric bias that have been neglected: overlooking how auxiliary factors can impede LLM performance despite competence (\"auxiliary oversight\"), and dismissing LLM mechanistic strategies that differ from those of humans as not genuinely competent (\"mechanistic chauvinism\"). Mitigating these biases necessitates an empirically-driven, iterative approach to mapping cognitive tasks to LLM-specific capacities and mechanisms, which can be done by supplementing carefully designed behavioral experiments with mechanistic studies.","authors":["Raphaël Millière","Charles Rathkopf"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2024-07-04","first_seen":"2024-07-04","revised_at":null,"abs_url":"https://arxiv.org/abs/2407.03859","pdf_url":"https://arxiv.org/pdf/2407.03859","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["LLM评估","认知偏见","方法论"],"reason":"论文讨论LLM评估中的偏见，属于纯NLP能力评测方法论，不以人类行为仿真为标的。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:30","error":null,"has_summary":false,"summary":null},{"id":"2407.12032","version":1,"title":"Large Language Models for Behavioral Economics: Internal Validity and Elicitation of Mental Models","zh_title":"大语言模型用于行为经济学：内部效度与心智模型的引出","abstract":"In this article, we explore the transformative potential of integrating generative AI, particularly Large Language Models (LLMs), into behavioral and experimental economics to enhance internal validity. By leveraging AI tools, researchers can improve adherence to key exclusion restrictions and in particular ensure the internal validity measures of mental models, which often require human intervention in the incentive mechanism. We present a case study demonstrating how LLMs can enhance experimental design, participant engagement, and the validity of measuring mental models.","authors":["Brian Jabarian"],"categories":["cs.HC","cs.AI","econ.GN"],"primary_category":"cs.HC","announce_type":"new","date":"2024-06-30","first_seen":"2024-06-30","revised_at":null,"abs_url":"https://arxiv.org/abs/2407.12032","pdf_url":"https://arxiv.org/pdf/2407.12032","source_feed":"backfill","score":7,"bucket":"pending","rubric_hits":["A1","A2","B2"],"tags":["LLM仿真","行为经济学","内部效度"],"reason":"用LLM增强行为经济学实验的内部效度，涉及人类被试替代和测量效度，但侧重方法改…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:14:46","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":122,"question":"如何利用大语言模型增强行为与实验经济学的内部效度，特别是改善心理模型的测量效度与排除限制的遵守？","design":"本文并非直接以LLM替代人类被试的仿真研究，而是提出利用LLM优化实验设计、提升参与者参与度、确保激励相容，并监测在线实验中的不良行为，从而增强内部效度。文中包含一个案例，展示如何用LLM创建引人入胜的叙事环境，以激励方式测量通常难以追踪的思维风格。","baseline":"无对照","findings":"LLM可通过改善排除限制的遵守来增强实验的内部效度；利用LLM生成的合成行为数据为探索决策过程提供了新的方法论深度。","reliability":"论文未讨论","relevance":"本文侧重于用LLM改进实验方法而非直接仿真人类行为，但涉及LLM在行为经济学实验中的应用与效度问题，对关注仿真可靠性的研究者有一定参考价值，但缺乏人类基准对照，建议略读。","inspiration":"该方法利用LLM生成引人入胜的叙事环境来测量心理模型，可借鉴其通过叙事干预提升测量效度的设计思路｜可迁移到消费者跨期选择实验中，改善对时间偏好和思维风格的测量｜设计雏形：以LLM生成个性化金融决策叙事作为处理，人类被试在叙事前后完成跨期选择任务，结果变量为贴现率与思维风格量表得分，以传统无叙事条件下的行为数据作为对照基准"}},{"id":"2406.19317","version":2,"title":"Jump Starting Bandits with LLM-Generated Prior Knowledge","zh_title":"用LLM生成的先验知识启动Bandit算法","abstract":"We present substantial evidence demonstrating the benefits of integrating Large Language Models (LLMs) with a Contextual Multi-Armed Bandit framework. Contextual bandits have been widely used in recommendation systems to generate personalized suggestions based on user-specific contexts. We show that LLMs, pre-trained on extensive corpora rich in human knowledge and preferences, can simulate human behaviours well enough to jump-start contextual multi-armed bandits to reduce online learning regret. We propose an initialization algorithm for contextual bandits by prompting LLMs to produce a pre-training dataset of approximate human preferences for the bandit. This significantly reduces online learning regret and data-gathering costs for training such models. Our approach is validated empirically through two sets of experiments with different bandit setups: one which utilizes LLMs to serve as an oracle and a real-world experiment utilizing data from a conjoint survey experiment.","authors":["Parand A. Alamdari","Yanshuai Cao","Kevin H. Wilson"],"categories":["cs.LG","cs.AI","cs.CL"],"primary_category":"cs.LG","announce_type":"new","date":"2024-06-27","first_seen":"2024-06-27","revised_at":null,"abs_url":"https://arxiv.org/abs/2406.19317","pdf_url":"https://arxiv.org/pdf/2406.19317","source_feed":"backfill","score":7,"bucket":"pending","rubric_hits":["A1","B1","B2"],"tags":["LLM仿真","人类偏好模拟","上下文Bandit"],"reason":"用LLM模拟人类偏好以初始化bandit，有真实联合调查数据对照，涉及推荐系统…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:14:46","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":121,"question":"如何利用大语言模型生成的先验知识来初始化上下文多臂老虎机，以减少在线学习的遗憾？","design":"提出CBLI框架，用LLM根据用户特征生成大量合成用户交互与偏好数据，预训练上下文老虎机模型（LinUCB），再在真实用户交互中微调；实验一用LLM模拟用户对慈善捐赠营销风格的偏好，实验二用真实联合调查数据测试睡眠老虎机。","baseline":"实验二使用真实联合调查实验数据中的人类偏好作为对照基准。","findings":"LLM生成的近似人类偏好虽不完全匹配真实分布，但能提供优于随机冷启动的初始化，在两个实验中分别使早期遗憾降低14-17%和19-20%；即使隐藏部分隐私敏感属性，仍能降低14.8%的早期遗憾。","reliability":"论文承认LLM生成的奖励分布可能无法完美匹配真实人类偏好，但强调其仍优于随机基线；未系统讨论LLM模拟在何种条件下会失效。","relevance":"该研究用LLM模拟人类偏好来初始化推荐系统，有真实联合调查数据作为对照，直接涉及LLM作为人类被试替代品的仿真可靠性与偏差问题，值得精读。","inspiration":"该方法通过LLM生成合成用户偏好数据来预训练推荐模型，再用少量真实交互微调，可作为冷启动问题的低成本解决方案｜可迁移到消费者金融产品推荐或个性化定价实验，例如用LLM模拟不同风险偏好人群对贷款产品的选择｜以LLM生成大量合成消费者对贷款条款的偏好数据，预训练一个上下文老虎机模型，再在真实信贷选择实验数据上微调，以真实人类选择作为对照，评估早期推荐准确度与遗憾值"}},{"id":"2406.17972","version":4,"title":"LABOR-LLM: Language-Based Occupational Representations with Large Language Models","zh_title":"LABOR-LLM：基于大语言模型的职业表征","abstract":"This paper builds an empirical model that predicts a worker's next occupation as a function of the worker's occupational history. Because histories are sequences of occupations, the covariate space is high-dimensional, and further, the outcome (the next occupation) is a discrete choice that can take on many values. To estimate the parameters of the model, we leverage an approach from generative artificial intelligence. Estimation begins from a ``foundation model'' trained on non-representative data and then ``fine-tunes'' the estimation using data about careers from a representative survey. We convert tabular data from the survey into text files that resemble resumes and fine-tune the parameters of the foundation model, a large language model (LLM), using these text files with the objective of predicting the next token (word). The resulting fine-tuned LLM is used to calculate estimates of worker transition probabilities. Its predictive performance surpasses all prior models, both for the task of granularly predicting the next occupation as well as for specific tasks such as predicting whether the worker changes occupations or stays in the labor force. We quantify the value of fine-tuning and further show that by adding more career data from a different population, fine-tuning smaller LLMs (fewer parameters) surpasses the performance of fine-tuning larger models. When we omit the English language occupational title and replace it with a unique code, predictive performance declines.","authors":["Susan Athey","Herman Brunborg","Tianyu Du","Ayush Kanodia","Keyon Vafa"],"categories":["cs.LG","cs.CL","econ.EM"],"primary_category":"cs.LG","announce_type":"new","date":"2024-06-25","first_seen":"2024-06-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2406.17972","pdf_url":"https://arxiv.org/pdf/2406.17972","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["职业预测","LLM微调","序列建模"],"reason":"纯预测模型，用LLM预测职业转换，不涉及人类行为仿真或对照实验。","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:58:59","error":null,"has_summary":false,"summary":null},{"id":"2406.16510","version":1,"title":"Large Language Models in Student Assessment: Comparing ChatGPT and Human Graders","zh_title":"学生评估中的大语言模型：比较ChatGPT与人类评分员","abstract":"This study investigates the efficacy of large language models (LLMs) as tools for grading master-level student essays. Utilizing a sample of 60 essays in political science, the study compares the accuracy of grades suggested by the GPT-4 model with those awarded by university teachers. Results indicate that while GPT-4 aligns with human grading standards on mean scores, it exhibits a risk-averse grading pattern and its interrater reliability with human raters is low. Furthermore, modifications in the grading instructions (prompt engineering) do not significantly alter AI performance, suggesting that GPT-4 primarily assesses generic essay characteristics such as language quality rather than adapting to nuanced grading criteria. These findings contribute to the understanding of AI's potential and limitations in higher education, highlighting the need for further development to enhance its adaptability and sensitivity to specific educational assessment requirements.","authors":["Magnus Lundgren"],"categories":["econ.GN"],"primary_category":"econ.GN","announce_type":"new","date":"2024-06-24","first_seen":"2024-06-24","revised_at":null,"abs_url":"https://arxiv.org/abs/2406.16510","pdf_url":"https://arxiv.org/pdf/2406.16510","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM评分","教育评估","人机对比"],"reason":"用LLM替代人工评分，属于标注员替代而非仿真被试，但涉及与人类评分对照，边界情…","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:58:44","error":null,"has_summary":false,"summary":null},{"id":"2406.14508","version":1,"title":"Evidence of a log scaling law for political persuasion with large language models","zh_title":"大语言模型政治说服力的对数缩放定律证据","abstract":"Large language models can now generate political messages as persuasive as those written by humans, raising concerns about how far this persuasiveness may continue to increase with model size. Here, we generate 720 persuasive messages on 10 U.S. political issues from 24 language models spanning several orders of magnitude in size. We then deploy these messages in a large-scale randomized survey experiment (N = 25,982) to estimate the persuasive capability of each model. Our findings are twofold. First, we find evidence of a log scaling law: model persuasiveness is characterized by sharply diminishing returns, such that current frontier models are barely more persuasive than models smaller in size by an order of magnitude or more. Second, mere task completion (coherence, staying on topic) appears to account for larger models' persuasive advantage. These findings suggest that further scaling model size will not much increase the persuasiveness of static LLM-generated messages.","authors":["Kobi Hackenburg","Ben M. Tappin","Paul Röttger","Scott Hale","Jonathan Bright","Helen Margetts"],"categories":["cs.CL","cs.AI","cs.CY","cs.HC"],"primary_category":"cs.CL","announce_type":"new","date":"2024-06-20","first_seen":"2024-06-20","revised_at":null,"abs_url":"https://arxiv.org/abs/2406.14508","pdf_url":"https://arxiv.org/pdf/2406.14508","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["LLM仿真","政治说服","人类数据对照"],"reason":"用LLM生成政治说服信息，通过大规模随机调查实验与人类数据对照，评估模型说服力…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:30","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":31,"question":"大语言模型的政治说服力是否随模型规模扩大而持续提升？","design":"用24个不同规模的语言模型生成720条政治说服信息，通过大规模随机调查实验（N=25,982）将美国成年人随机分配到AI组、人类组或对照组，测量其对10个政策议题的态度变化。","baseline":"人类撰写的说服信息以及未接受任何信息的对照组。","findings":"模型说服力随规模呈对数缩放，当前前沿模型仅比小一个数量级的模型略具说服力；说服力提升主要源于任务完成度（连贯性、切题），前沿模型在该指标上已接近上限。","reliability":"论文未讨论","relevance":"该研究以真实人类调查数据为基准，评估LLM在政治说服场景中的仿真效果，并揭示规模扩展的边际收益递减，直接回应了研究者对仿真可靠性与失效条件的关切，值得精读。","inspiration":"该方法将LLM生成内容作为处理，通过大规模随机调查实验与人类生成内容及对照组比较，测量态度变化，可借鉴其处理施加与基准对照设计。｜可迁移至政策公告的预期形成研究，如评估AI生成的经济新闻对公众通胀预期的影响。｜以LLM生成不同风格的经济新闻为处理，招募代表性样本为被试，测量其通胀预期变化，并以人类撰写新闻和真实历史数据为对照。"}},{"id":"2406.13605","version":2,"title":"Nicer Than Humans: How do Large Language Models Behave in the Prisoner's Dilemma?","zh_title":"比人类更友善：大语言模型在囚徒困境中的行为研究","abstract":"The behavior of Large Language Models (LLMs) as artificial social agents is largely unexplored, and we still lack extensive evidence of how these agents react to simple social stimuli. Testing the behavior of AI agents in classic Game Theory experiments provides a promising theoretical framework for evaluating the norms and values of these agents in archetypal social situations. In this work, we investigate the cooperative behavior of three LLMs (Llama2, Llama3, and GPT3.5) when playing the Iterated Prisoner's Dilemma against random adversaries displaying various levels of hostility. We introduce a systematic methodology to evaluate an LLM's comprehension of the game rules and its capability to parse historical gameplay logs for decision-making. We conducted simulations of games lasting for 100 rounds and analyzed the LLMs' decisions in terms of dimensions defined in the behavioral economics literature. We find that all models tend not to initiate defection but act cautiously, favoring cooperation over defection only when the opponent's defection rate is low. Overall, LLMs behave at least as cooperatively as the typical human player, although our results indicate some substantial differences among models. In particular, Llama2 and GPT3.5 are more cooperative than humans, and especially forgiving and non-retaliatory for opponent defection rates below 30%. More similar to humans, Llama3 exhibits consistently uncooperative and exploitative behavior unless the opponent always cooperates. Our systematic approach to the study of LLMs in game theoretical scenarios is a step towards using these simulations to inform practices of LLM auditing and alignment.","authors":["Nicoló Fontana","Francesco Pierri","Luca Maria Aiello"],"categories":["cs.CY","cs.AI","cs.GT","physics.soc-ph"],"primary_category":"cs.CY","announce_type":"new","date":"2024-06-19","first_seen":"2024-06-19","revised_at":null,"abs_url":"https://arxiv.org/abs/2406.13605","pdf_url":"https://arxiv.org/pdf/2406.13605","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM仿真","囚徒困境","行为博弈"],"reason":"用LLM玩囚徒困境并与人类数据对照，直接仿真人类决策行为。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:14:44","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":26,"question":"大型语言模型在迭代囚徒困境中面对不同敌对程度的对手时，其合作行为如何？","design":"使用Llama2、Llama3和GPT3.5三个LLM作为被试，与具有不同背叛概率的随机对手进行100轮迭代囚徒困境博弈，通过系统提示评估模型对规则的理解和历史记录解析能力，分析其合作决策。","baseline":"对照已有行为经济学文献中报告的人类玩家在囚徒困境中的典型合作行为。","findings":"所有模型倾向于不首先背叛，但仅在对手背叛率低时更合作；Llama2和GPT3.5比人类更合作、更宽容，而Llama3更接近人类，表现出不合作和剥削性行为。","reliability":"论文未讨论","relevance":"该研究直接用LLM复现经典博弈实验并与人类数据对照，属于经济学实验场景下的人类仿真，且分析了模型间差异，值得精读以评估仿真可靠性与偏差。","inspiration":"该研究通过系统提示解析历史记录并控制对手背叛概率来测量LLM的合作行为，方法上可借鉴其精确操纵对手策略以分离LLM反应模式的做法｜可迁移到资产定价实验中的信任博弈或投资者情绪传染研究，考察LLM在不同市场信息环境下的策略调整｜可设计让LLM作为投资者与不同诚实度的基金经理进行重复信任博弈，处理变量为基金经理的欺骗概率，结果变量为投资额，对照真实人类实验数据"}},{"id":"2406.13558","version":2,"title":"Enhancing Travel Choice Modeling with Large Language Models: A Prompt-Learning Approach","zh_title":"利用大语言模型增强出行选择建模：一种提示学习方法","abstract":"Travel choice analysis is crucial for understanding individual travel behavior to develop appropriate transport policies and recommendation systems in Intelligent Transportation Systems (ITS). Despite extensive research, this domain faces two critical challenges: a) modeling with limited survey data, and b) simultaneously achieving high model explainability and accuracy. In this paper, we introduce a novel prompt-learning-based Large Language Model(LLM) framework that significantly improves prediction accuracy and provides explicit explanations for individual predictions. This framework involves three main steps: transforming input variables into textual form; building of demonstrations similar to the object, and applying these to a well-trained LLM. We tested the framework's efficacy using two widely used choice datasets: London Passenger Mode Choice (LPMC) and Optima-Mode collected in Switzerland. The results indicate that the LLM significantly outperforms state-of-the-art deep learning methods and discrete choice models in predicting people's choices. Additionally, we present a case of explanation illustrating how the LLM framework generates understandable and explicit explanations at the individual level.","authors":["Xuehao Zhai","Hanlin Tian","Lintong Li","Tianyu Zhao"],"categories":["cs.AI"],"primary_category":"cs.AI","announce_type":"new","date":"2024-06-19","first_seen":"2024-06-19","revised_at":null,"abs_url":"https://arxiv.org/abs/2406.13558","pdf_url":"https://arxiv.org/pdf/2406.13558","source_feed":"backfill","score":7,"bucket":"pending","rubric_hits":["A1","B1","B2"],"tags":["出行行为预测","LLM仿真","选择建模"],"reason":"用LLM预测出行选择，有真实人类数据对照，属于行为仿真，但侧重预测精度而非复现…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:29","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":127,"question":"如何利用提示学习的大语言模型框架提高出行方式选择的预测精度并生成个体层面的可解释性？","design":"本研究并非严格的人类仿真实验，而是提出一种基于提示学习的LLM框架，将出行选择数据转换为文本，构建相似演示示例，输入预训练LLM进行出行方式预测，并输出个体解释。","baseline":"使用两个真实出行选择数据集：伦敦乘客方式选择（LPMC）和瑞士Optima-Mode调查数据。","findings":"LLM框架在预测精度上显著优于现有深度学习和离散选择模型；同时能够为个体预测生成明确、可理解的解释。","reliability":"论文未讨论","relevance":"该研究用LLM预测真实出行选择行为，有真实人类数据对照，属于行为预测仿真，但侧重预测精度而非复现人类决策偏差或分布，与研究者关注的仿真可靠性和失效条件关联较弱，可酌情略读。","inspiration":"该方法将个体选择数据转换为文本并构建相似示例进行提示学习，可借鉴其将结构化行为数据文本化并利用上下文示例提升预测的思路｜可迁移至消费者金融产品选择预测，如信贷产品选择或投资组合偏好分析｜以真实银行客户交易数据为对照，将客户特征与历史选择转为文本提示，用LLM预测其贷款类型选择，比较预测准确率与离散选择模型"}},{"id":"2406.11426","version":1,"title":"Can AI with High Reasoning Ability Replicate Human-like Decision Making in Economic Experiments?","zh_title":"高推理能力AI能否复制经济实验中的人类决策？","abstract":"Economic experiments offer a controlled setting for researchers to observe human decision-making and test diverse theories and hypotheses; however, substantial costs and efforts are incurred to gather many individuals as experimental participants. To address this, with the development of large language models (LLMs), some researchers have recently attempted to develop simulated economic experiments using LLMs-driven agents, called generative agents. If generative agents can replicate human-like decision-making in economic experiments, the cost problem of economic experiments can be alleviated. However, such a simulation framework has not been yet established. Considering the previous research and the current evolutionary stage of LLMs, this study focuses on the reasoning ability of generative agents as a key factor toward establishing a framework for such a new methodology. A multi-agent simulation, designed to improve the reasoning ability of generative agents through prompting methods, was developed to reproduce the result of an actual economic experiment on the ultimatum game. The results demonstrated that the higher the reasoning ability of the agents, the closer the results were to the theoretical solution than to the real experimental result. The results also suggest that setting the personas of the generative agents may be important for reproducing the results of real economic experiments. These findings are valuable for the future definition of a framework for replacing human participants with generative agents in economic experiments when LLMs are further developed.","authors":["Ayato Kitadai","Sinndy Dayana Rico Lugo","Yudai Tsurusaki","Yusuke Fukasawa","Nariaki Nishino"],"categories":["cs.GT","econ.GN"],"primary_category":"cs.GT","announce_type":"new","date":"2024-06-17","first_seen":"2024-06-17","revised_at":null,"abs_url":"https://arxiv.org/abs/2406.11426","pdf_url":"https://arxiv.org/pdf/2406.11426","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM仿真","经济实验","人类行为对照"],"reason":"用LLM代理复现最后通牒博弈实验，并与真实人类数据对照，直接命中核心判据。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:14:43","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":25,"question":"提高生成式智能体的推理能力能否使其在最后通牒博弈经济实验中复现人类决策？","design":"使用GPT-3.5-turbo和GPT-4等LLM驱动的生成式智能体进行多智能体仿真，通过零样本、少样本和思维链提示方法操纵推理能力，模拟最后通牒博弈中的提议者和响应者决策，测量分配金额和接受/拒绝行为。","baseline":"对照Lin et al. (2020)的真实人类最后通牒博弈实验数据。","findings":"智能体推理能力越高，其结果越接近理论均衡而非真实人类行为；设置智能体的人格特征可能对复现真实实验结果很重要。","reliability":"论文指出当前仿真框架尚未建立，推理能力提升反而偏离人类行为，且提示语言、模型版本和参数设置影响结果，未来需更高推理能力LLM和更完善的人格设定。","relevance":"该研究直接以真实人类实验为基准，检验LLM代理在经济博弈中的仿真效度，并揭示了推理能力增强反而导致行为偏离人类的关键失效条件，与您关注的经济学实验仿真和可靠性批判高度契合，值得精读。","inspiration":"借鉴其通过提示方法（零样本、少样本、思维链）系统操纵LLM推理能力，并与真实人类实验数据严格对照的仿真效度检验框架｜可迁移到资产定价实验，检验LLM代理能否复现人类在泡沫形成与破裂中的非理性交易行为｜以LLM为被试，通过不同推理提示形成处理组，模拟连续竞价市场中的买卖决策，结果变量为价格偏离基础价值的程度，对照Smith et al. (1988)的实验室资产市场泡沫数据"}},{"id":"2406.05972","version":2,"title":"Decision-Making Behavior Evaluation Framework for LLMs under Uncertain Context","zh_title":"不确定情境下大语言模型决策行为评估框架","abstract":"When making decisions under uncertainty, individuals often deviate from rational behavior, which can be evaluated across three dimensions: risk preference, probability weighting, and loss aversion. Given the widespread use of large language models (LLMs) in decision-making processes, it is crucial to assess whether their behavior aligns with human norms and ethical expectations or exhibits potential biases. Several empirical studies have investigated the rationality and social behavior performance of LLMs, yet their internal decision-making tendencies and capabilities remain inadequately understood. This paper proposes a framework, grounded in behavioral economics, to evaluate the decision-making behaviors of LLMs. Through a multiple-choice-list experiment, we estimate the degree of risk preference, probability weighting, and loss aversion in a context-free setting for three commercial LLMs: ChatGPT-4.0-Turbo, Claude-3-Opus, and Gemini-1.0-pro. Our results reveal that LLMs generally exhibit patterns similar to humans, such as risk aversion and loss aversion, with a tendency to overweight small probabilities. However, there are significant variations in the degree to which these behaviors are expressed across different LLMs. We also explore their behavior when embedded with socio-demographic features, uncovering significant disparities. For instance, when modeled with attributes of sexual minority groups or physical disabilities, Claude-3-Opus displays increased risk aversion, leading to more conservative choices. These findings underscore the need for careful consideration of the ethical implications and potential biases in deploying LLMs in decision-making scenarios. Therefore, this study advocates for developing standards and guidelines to ensure that LLMs operate within ethical boundaries while enhancing their utility in complex decision-making environments.","authors":["Jingru Jia","Zehua Yuan","Junhao Pan","Paul E. McNamara","Deming Chen"],"categories":["cs.AI","cs.CY","cs.HC","cs.LG","econ.TH"],"primary_category":"cs.AI","announce_type":"new","date":"2024-06-10","first_seen":"2024-06-10","revised_at":null,"abs_url":"https://arxiv.org/abs/2406.05972","pdf_url":"https://arxiv.org/pdf/2406.05972","source_feed":"backfill","score":7,"bucket":"pending","rubric_hits":["A1","A2","B2","B4"],"tags":["LLM决策行为","行为经济学","人类对照"],"reason":"用行为经济学实验评估LLM决策行为，与人类规范对照，涉及风险偏好等，可迁移至人…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:14:43","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":120,"question":"在不确定情境下，大型语言模型（LLM）的决策行为在风险偏好、概率加权和损失厌恶三个维度上是否与人类相似，以及嵌入社会人口特征后是否会产生偏差？","design":"使用基于行为经济学的多组选择列表实验，在无上下文情境下测量ChatGPT-4.0-Turbo、Claude-3-Opus和Gemini-1.0-pro三个商业LLM的风险偏好、概率加权和损失厌恶参数；进一步嵌入社会人口特征（如性少数群体或身体残疾属性）观察其决策变化。","baseline":"无对照","findings":"LLM普遍表现出与人类相似的风险厌恶、损失厌恶和对小概率的高估倾向，但不同模型间行为程度差异显著；嵌入社会人口特征后，部分模型（如Claude-3-Opus）在性少数或残疾属性下风险厌恶增强，决策更保守，显示出潜在偏差。","reliability":"论文未讨论","relevance":"该研究将LLM作为人类被试的替代，用行为经济学实验评估其决策模式，并考察社会人口特征带来的偏差，直接回应了LLM仿真人类决策的可靠性与伦理问题，值得精读以了解其框架和发现。","inspiration":"该方法借鉴了行为经济学经典的选择列表实验来测量LLM的风险偏好、概率加权和损失厌恶，并嵌入社会人口特征作为处理变量，观察决策偏差。｜可迁移到信贷审批歧视研究，检验LLM在贷款决策中是否因申请人性别、种族等特征产生系统性偏差。｜以LLM为被试，处理为嵌入不同社会人口特征的贷款申请人档案，结果变量为审批决策和风险偏好参数，对照真实银行信贷数据中的审批率与偏差模式。"}},{"id":"2406.03299","version":1,"title":"The Good, the Bad, and the Hulk-like GPT: Analyzing Emotional Decisions of Large Language Models in Cooperation and Bargaining Games","zh_title":"好、坏与浩克般的GPT：分析大语言模型在合作与讨价还价博弈中的情绪决策","abstract":"Behavior study experiments are an important part of society modeling and understanding human interactions. In practice, many behavioral experiments encounter challenges related to internal and external validity, reproducibility, and social bias due to the complexity of social interactions and cooperation in human user studies. Recent advances in Large Language Models (LLMs) have provided researchers with a new promising tool for the simulation of human behavior. However, existing LLM-based simulations operate under the unproven hypothesis that LLM agents behave similarly to humans as well as ignore a crucial factor in human decision-making: emotions. In this paper, we introduce a novel methodology and the framework to study both, the decision-making of LLMs and their alignment with human behavior under emotional states. Experiments with GPT-3.5 and GPT-4 on four games from two different classes of behavioral game theory showed that emotions profoundly impact the performance of LLMs, leading to the development of more optimal strategies. While there is a strong alignment between the behavioral responses of GPT-3.5 and human participants, particularly evident in bargaining games, GPT-4 exhibits consistent behavior, ignoring induced emotions for rationality decisions. Surprisingly, emotional prompting, particularly with `anger' emotion, can disrupt the \"superhuman\" alignment of GPT-4, resembling human emotional responses.","authors":["Mikhail Mozikov","Nikita Severin","Valeria Bodishtianu","Maria Glushanina","Mikhail Baklashkin","Andrey V. Savchenko","Ilya Makarov"],"categories":["cs.AI","cs.CL"],"primary_category":"cs.AI","announce_type":"new","date":"2024-06-05","first_seen":"2024-06-05","revised_at":null,"abs_url":"https://arxiv.org/abs/2406.03299","pdf_url":"https://arxiv.org/pdf/2406.03299","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["LLM人类仿真","行为博弈","情绪决策"],"reason":"用LLM仿真人类在博弈中的情绪决策，并与真实人类数据对照，评估对齐与失效条件。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:14:43","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":24,"question":"情绪如何影响大语言模型在合作与讨价还价博弈中的决策，以及其行为与人类行为的对齐程度如何？","design":"使用GPT-3.5和GPT-4作为被试，通过情绪提示（愤怒、悲伤、快乐、厌恶、恐惧）注入五种基本情绪，在最后通牒博弈、独裁者博弈、囚徒困境和性别战四类博弈中测量出价份额、接受率、合作率及最大收益百分比等结果变量。","baseline":"以真实人类参与者在相同博弈实验中的行为数据作为对照基准。","findings":"情绪显著影响LLM的策略表现，GPT-3.5在讨价还价博弈中与人类行为高度对齐，而GPT-4通常忽略情绪保持理性；但愤怒情绪提示能扰乱GPT-4的“超人类”对齐，使其表现出类似人类的情绪反应。","reliability":"论文未讨论","relevance":"该研究直接以真实人类数据为基准，检验LLM在情绪影响下的行为对齐与失效条件，涵盖经济学博弈场景，并揭示了GPT-4在愤怒情绪下的对齐崩溃，高度契合研究者对仿真可靠性及批判性条件的关注，值得精读原文。","inspiration":"该方法通过情绪提示词注入情绪状态，并设置无情绪中性基线，可借鉴用于经济决策实验中情绪处理的标准化设计｜可迁移到资产定价实验中，研究情绪如何影响投资者对风险资产的需求与定价偏差｜以LLM为被试，通过愤怒/恐惧等情绪提示词处理，测量其在模拟股票交易中的出价与风险偏好，结果与真实投资者实验数据对照"}},{"id":"2406.01407","version":1,"title":"Utilizing Large Language Models for Automating Technical Customer Support","zh_title":"利用大语言模型实现技术客户支持自动化","abstract":"The use of large language models (LLMs) such as OpenAI's GPT-4 in technical customer support (TCS) has the potential to revolutionize this area. This study examines automated text correction, summarization of customer inquiries and question answering using LLMs. Through prototypes and data analyses, the potential and challenges of integrating LLMs into the TCS will be demonstrated. Our results show promising approaches for improving the efficiency and quality of customer service through LLMs, but also emphasize the need for quality-assured implementation and organizational adjustments in the data ecosystem.","authors":["Jochen Wulf","Jürg Meierhofer"],"categories":["econ.GN"],"primary_category":"econ.GN","announce_type":"new","date":"2024-06-03","first_seen":"2024-06-03","revised_at":null,"abs_url":"https://arxiv.org/abs/2406.01407","pdf_url":"https://arxiv.org/pdf/2406.01407","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["LLM应用","客户支持","自动化"],"reason":"论文用LLM做客服自动化，属NLP应用，不涉及人类行为仿真或对照。","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:58:44","error":null,"has_summary":false,"summary":null},{"id":"2405.19578","version":1,"title":"The Accuracy of Domain Specific and Descriptive Analysis Generated by Large Language Models","zh_title":"大语言模型生成的领域特定与描述性分析的准确性","abstract":"Large language models (LLMs) have attracted considerable attention as they are capable of showcasing impressive capabilities generating comparable high-quality responses to human inputs. LLMs, can not only compose textual scripts such as emails and essays but also executable programming code. Contrary, the automated reasoning capability of these LLMs in performing statistically-driven descriptive analysis, particularly on user-specific data and as personal assistants to users with limited background knowledge in an application domain who would like to carry out basic, as well as advanced statistical and domain-specific analysis is not yet fully explored. More importantly, the performance of these LLMs has not been compared and discussed in detail when domain-specific data analysis tasks are needed. This study, consequently, explores whether LLMs can be used as generative AI-based personal assistants to users with minimal background knowledge in an application domain infer key data insights. To demonstrate the performance of the LLMs, the study reports a case study through which descriptive statistical analysis, as well as Natural Language Processing (NLP) based investigations, are performed on a number of phishing emails with the objective of comparing the accuracy of the results generated by LLMs to the ones produced by analysts. The experimental results show that LangChain and the Generative Pre-trained Transformer (GPT-4) excel in numerical reasoning tasks i.e., temporal statistical analysis, achieve competitive correlation with human judgments on feature engineering tasks while struggle to some extent on domain specific knowledge reasoning, where domain-specific knowledge is required.","authors":["Denish Omondi Otieno","Faranak Abri","Sima Siami-Namini","Akbar Siami Namin"],"categories":["cs.CE","econ.GN"],"primary_category":"cs.CE","announce_type":"new","date":"2024-05-30","first_seen":"2024-05-30","revised_at":null,"abs_url":"https://arxiv.org/abs/2405.19578","pdf_url":"https://arxiv.org/pdf/2405.19578","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["LLM数据分析","人类判断对比","描述性统计"],"reason":"LLM替代分析师做描述性统计，属替代人类劳动而非仿真被试，边界情形。","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:58:46","error":null,"has_summary":false,"summary":null},{"id":"2405.19313","version":2,"title":"Language Models Trained to do Arithmetic Predict Human Risky and Intertemporal Choice","zh_title":"训练做算术的语言模型预测人类风险与跨期选择","abstract":"The observed similarities in the behavior of humans and Large Language Models (LLMs) have prompted researchers to consider the potential of using LLMs as models of human cognition. However, several significant challenges must be addressed before LLMs can be legitimately regarded as cognitive models. For instance, LLMs are trained on far more data than humans typically encounter, and may have been directly trained on human data in specific cognitive tasks or aligned with human preferences. Consequently, the origins of these behavioral similarities are not well understood. In this paper, we propose a novel way to enhance the utility of LLMs as cognitive models. This approach involves (i) leveraging computationally equivalent tasks that both an LLM and a rational agent need to master for solving a cognitive problem and (ii) examining the specific task distributions required for an LLM to exhibit human-like behaviors. We apply this approach to decision-making -- specifically risky and intertemporal choice -- where the key computationally equivalent task is the arithmetic of expected value calculations. We show that an LLM pretrained on an ecologically valid arithmetic dataset, which we call Arithmetic-GPT, predicts human behavior better than many traditional cognitive models. Pretraining LLMs on ecologically valid arithmetic datasets is sufficient to produce a strong correspondence between these models and human decision-making. Our results also suggest that LLMs used as cognitive models should be carefully investigated via ablation studies of the pretraining data.","authors":["Jian-Qiao Zhu","Haijiang Yan","Thomas L. Griffiths"],"categories":["cs.AI","cs.CL","econ.GN"],"primary_category":"cs.AI","announce_type":"new","date":"2024-05-29","first_seen":"2024-05-29","revised_at":null,"abs_url":"https://arxiv.org/abs/2405.19313","pdf_url":"https://arxiv.org/pdf/2405.19313","source_feed":"backfill","score":8,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["LLM仿真","决策行为","认知模型"],"reason":"用LLM预测人类风险与跨期选择，与真实人类数据对照，并分析仿真有效条件，直接相…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:14:41","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":61,"question":"语言模型在算术任务上的预训练能否使其产生与人类相似的风险和跨期选择行为？","design":"训练一个小型语言模型（约10M参数）在合成算术数据集上（如期望值计算），提取其嵌入向量，用逻辑回归预测人类选择概率，并与传统认知模型及LLaMA3比较。","baseline":"使用真实人类在风险和跨期选择任务中的选择数据作为基准。","findings":"在生态有效的算术数据集上预训练的Arithmetic-GPT模型预测人类选择优于许多传统认知模型；仅靠算术预训练就足以产生与人类决策的强对应关系。","reliability":"论文指出需通过消融实验仔细检查预训练数据，且合成数据分布需符合生态分布才能有效，否则预测力有限。","relevance":"直接以真实人类数据为基准，用LLM仿真风险与跨期选择，并分析预训练数据分布对仿真有效性的影响，高度契合研究者对仿真可靠性及失效条件的关注。","inspiration":"该方法通过控制预训练数据的生态分布来提升模型对人类决策的预测力，值得借鉴其消融实验设计以检验数据特征对仿真效果的影响｜可迁移到消费者跨期选择研究，例如分析不同利率环境下个体的储蓄与消费决策｜以LLM为被试，处理变量为预训练数据中利率变动的分布（符合真实市场波动），结果变量为模拟的跨期选择偏好，用家庭金融调查的真实跨期选择数据作为对照基准"}},{"id":"2405.09161","version":2,"title":"Exploring the Potential of Large Language Models for Automation in Technical Customer Service","zh_title":"探索大语言模型在技术客户服务自动化中的潜力","abstract":"Purpose: The purpose of this study is to investigate the potential of Large Language Models (LLMs) in transforming technical customer service (TCS) through the automation of cognitive tasks. Design/Methodology/Approach: Using a prototyping approach, the research assesses the feasibility of automating cognitive tasks in TCS with LLMs, employing real-world technical incident data from a Swiss telecommunications operator. Findings: Lower-level cognitive tasks such as translation, summarization, and content generation can be effectively automated with LLMs like GPT-4, while higher-level tasks such as reasoning require more advanced technological approaches such as Retrieval-Augmented Generation (RAG) or finetuning ; furthermore, the study underscores the significance of data ecosystems in enabling more complex cognitive tasks by fostering data sharing among various actors involved. Originality/Value: This study contributes to the emerging theory on LLM potential and technical feasibility in service management, providing concrete insights for operators of TCS units and highlighting the need for further research to address limitations and validate the applicability of LLMs across different domains.","authors":["Jochen Wulf","Juerg Meierhofer"],"categories":["econ.GN"],"primary_category":"econ.GN","announce_type":"new","date":"2024-05-15","first_seen":"2024-05-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2405.09161","pdf_url":"https://arxiv.org/pdf/2405.09161","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["LLM应用","客户服务自动化","认知任务"],"reason":"研究LLM自动化客服任务，不涉及人类行为仿真或对照。","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:58:46","error":null,"has_summary":false,"summary":null},{"id":"2405.00981","version":2,"title":"Bayesian Optimization with LLM-Based Acquisition Functions for Natural Language Preference Elicitation","zh_title":"基于LLM采集函数的贝叶斯优化用于自然语言偏好获取","abstract":"Designing preference elicitation (PE) methodologies that can quickly ascertain a user's top item preferences in a cold-start setting is a key challenge for building effective and personalized conversational recommendation (ConvRec) systems. While large language models (LLMs) enable fully natural language (NL) PE dialogues, we hypothesize that monolithic LLM NL-PE approaches lack the multi-turn, decision-theoretic reasoning required to effectively balance the exploration and exploitation of user preferences towards an arbitrary item set. In contrast, traditional Bayesian optimization PE methods define theoretically optimal PE strategies, but cannot generate arbitrary NL queries or reason over content in NL item descriptions -- requiring users to express preferences via ratings or comparisons of unfamiliar items. To overcome the limitations of both approaches, we formulate NL-PE in a Bayesian Optimization (BO) framework that seeks to actively elicit NL feedback to identify the best recommendation. Key challenges in generalizing BO to deal with natural language feedback include determining: (a) how to leverage LLMs to model the likelihood of NL preference feedback as a function of item utilities, and (b) how to design an acquisition function for NL BO that can elicit preferences in the infinite space of language. We demonstrate our framework in a novel NL-PE algorithm, PEBOL, which uses: 1) Natural Language Inference (NLI) between user preference utterances and NL item descriptions to maintain Bayesian preference beliefs, and 2) BO strategies such as Thompson Sampling (TS) and Upper Confidence Bound (UCB) to steer LLM query generation. We numerically evaluate our methods in controlled simulations, finding that after 10 turns of dialogue, PEBOL can achieve an MRR@10 of up to 0.27 compared to the best monolithic LLM baseline's MRR@10 of 0.17, despite relying on earlier and smaller LLMs.","authors":["David Eric Austin","Anton Korikov","Armin Toroghi","Scott Sanner"],"categories":["cs.AI","cs.CL"],"primary_category":"cs.AI","announce_type":"new","date":"2024-05-02","first_seen":"2024-05-02","revised_at":null,"abs_url":"https://arxiv.org/abs/2405.00981","pdf_url":"https://arxiv.org/pdf/2405.00981","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["偏好获取","对话推荐","贝叶斯优化"],"reason":"纯NLP偏好获取算法，无人类行为仿真或对照，属推荐系统研究。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:28","error":null,"has_summary":false,"summary":null},{"id":"2404.17143","version":2,"title":"Quantifying Memorization and Detecting Training Data of Pre-trained Language Models using Japanese Newspaper","zh_title":"使用日本报纸量化预训练语言模型的记忆与检测训练数据","abstract":"Dominant pre-trained language models (PLMs) have demonstrated the potential risk of memorizing and outputting the training data. While this concern has been discussed mainly in English, it is also practically important to focus on domain-specific PLMs. In this study, we pre-trained domain-specific GPT-2 models using a limited corpus of Japanese newspaper articles and evaluated their behavior. Experiments replicated the empirical finding that memorization of PLMs is related to the duplication in the training data, model size, and prompt length, in Japanese the same as in previous English studies. Furthermore, we attempted membership inference attacks, demonstrating that the training data can be detected even in Japanese, which is the same trend as in English. The study warns that domain-specific PLMs, sometimes trained with valuable private data, can ''copy and paste'' on a large scale.","authors":["Shotaro Ishihara","Hiromu Takahashi"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2024-04-26","first_seen":"2024-04-26","revised_at":null,"abs_url":"https://arxiv.org/abs/2404.17143","pdf_url":"https://arxiv.org/pdf/2404.17143","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["预训练语言模型","训练数据记忆","成员推断攻击"],"reason":"研究PLM记忆训练数据及成员推断攻击，属纯NLP评测，不以人类行为为参照。","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:59:21","error":null,"has_summary":false,"summary":null},{"id":"2404.09699","version":2,"title":"Generative AI for Game Theory-based Mobile Networking","zh_title":"基于博弈论的移动网络生成式人工智能","abstract":"With the continuous advancement of network technology, various emerging complex networking optimization problems have created a wide range of applications utilizing game theory. However, since game theory is a mathematical framework, game theory-based solutions often rely heavily on the experience and knowledge of human experts. Recently, the remarkable advantages exhibited by generative artificial intelligence (GAI) have gained widespread attention. In this work, we propose a novel GAI-enabled game theory solution that combines the powerful reasoning and generation capabilities of GAI to the design and optimization of mobile networking. Specifically, we first outline the game theory and key technologies of GAI, and explore the advantages of combining GAI with game theory. Then, we review the contributions and limitations of existing research and demonstrate the potential application values of GAI applied to game theory in mobile networking. Subsequently, we develop a large language model (LLM)-enabled game theory framework to realize this combination, and demonstrate the effectiveness of the proposed framework through a case study in secured UAV networks. Finally, we provide several directions for future extensions.","authors":["Long He","Geng Sun","Dusit Niyato","Hongyang Du","Fang Mei","Jiawen Kang","Mérouane Debbah","Zhu Han"],"categories":["cs.GT"],"primary_category":"cs.GT","announce_type":"new","date":"2024-04-15","first_seen":"2024-04-15","revised_at":null,"abs_url":"https://arxiv.org/abs/2404.09699","pdf_url":"https://arxiv.org/pdf/2404.09699","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体系统","博弈论","移动网络优化"],"reason":"纯多智能体系统研究，LLM用于博弈论优化移动网络，无人类行为对照。","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:59:32","error":null,"has_summary":false,"summary":null},{"id":"2404.08816","version":5,"title":"Measuring the Quality of Answers in Political Q&As with Large Language Models","zh_title":"用大语言模型衡量政治问答环节中答案的质量","abstract":"This article proposes a new approach for assessing the quality of answers in political question-and-answer sessions. We measure the quality of an answer based on how easily and accurately it can be recognized in a random set of candidate answers given the question's text. This measure reflects the answer's relevance and depth of engagement with the question. Like semantic search, we can implement this approach by training a language model on the corpus of observed questions and answers without additional human-labeled data. We showcase and validate our methodology within the context of the Question Period in the Canadian House of Commons. Our analysis reveals that while some answers have a weak semantic connection to questions, hinting at some evasion or obfuscation, they are generally at least moderately relevant, far exceeding what we would expect from random replies. We also find a meaningful correlation between answer quality and the party affiliation of the members of Parliament asking the questions.","authors":["R. Michael Alvarez","Jacob Morrier"],"categories":["cs.CL","econ.EM"],"primary_category":"cs.CL","announce_type":"new","date":"2024-04-12","first_seen":"2024-04-12","revised_at":null,"abs_url":"https://arxiv.org/abs/2404.08816","pdf_url":"https://arxiv.org/pdf/2404.08816","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["NLP评测","政治文本分析","语义相关性"],"reason":"用LLM评估政治问答质量，属NLP评测，不以人类行为仿真为目标。","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:58:59","error":null,"has_summary":false,"summary":null},{"id":"2404.00530","version":2,"title":"Comparing Bad Apples to Good Oranges: Aligning Large Language Models via Joint Preference Optimization","zh_title":"比较坏苹果与好橙子：通过联合偏好优化对齐大语言模型","abstract":"A common technique for aligning large language models (LLMs) relies on acquiring human preferences by comparing multiple generations conditioned on a fixed context. This method, however, relies solely on pairwise comparisons, where the generations are evaluated within an identical context. While effective to such conditional preferences often fail to encompass the nuanced and multidimensional nature of human preferences. In this work, we revisit the traditional paradigm of preference acquisition and propose a new axis based on eliciting preferences jointly over the instruction-response pairs. Unlike prior preference optimizations, which are designed for conditional ranking protocols (e.g., DPO), we propose Joint Preference Optimization (JPO), a new preference optimization objective that upweights the joint probability of the chosen instruction-response pair over the rejected instruction-response pair. Interestingly, LLMs trained with joint instruction-response preference data using JPO outperform LLM trained with DPO by $5.2\\%$ and $3.3\\%$ win-rate for summarization and open-ended dialogue datasets, respectively. Our findings reveal that joint preferences over instruction and response pairs can significantly enhance the alignment of LLMs by tapping into a broader spectrum of human preference elicitation. The data and code is available at https://github.com/Hritikbansal/dove.","authors":["Hritik Bansal","Ashima Suvarna","Gantavya Bhatt","Nanyun Peng","Kai-Wei Chang","Aditya Grover"],"categories":["cs.CL","cs.AI","cs.LG"],"primary_category":"cs.CL","announce_type":"new","date":"2024-03-31","first_seen":"2024-03-31","revised_at":null,"abs_url":"https://arxiv.org/abs/2404.00530","pdf_url":"https://arxiv.org/pdf/2404.00530","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["偏好优化","LLM对齐","指令微调"],"reason":"纯LLM对齐优化，无人类仿真或行为对照，属NLP能力评测范畴。","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:59:47","error":null,"has_summary":false,"summary":null},{"id":"2404.01332","version":3,"title":"Explaining Large Language Models Decisions Using Shapley Values","zh_title":"使用Shapley值解释大语言模型决策","abstract":"The emergence of large language models (LLMs) has opened up exciting possibilities for simulating human behavior and cognitive processes, with potential applications in various domains, including marketing research and consumer behavior analysis. However, the validity of utilizing LLMs as stand-ins for human subjects remains uncertain due to glaring divergences that suggest fundamentally different underlying processes at play and the sensitivity of LLM responses to prompt variations. This paper presents a novel approach based on Shapley values from cooperative game theory to interpret LLM behavior and quantify the relative contribution of each prompt component to the model's output. Through two applications - a discrete choice experiment and an investigation of cognitive biases - we demonstrate how the Shapley value method can uncover what we term \"token noise\" effects, a phenomenon where LLM decisions are disproportionately influenced by tokens providing minimal informative content. This phenomenon raises concerns about the robustness and generalizability of insights obtained from LLMs in the context of human behavior simulation. Our model-agnostic approach extends its utility to proprietary LLMs, providing a valuable tool for practitioners and researchers to strategically optimize prompts and mitigate apparent cognitive biases. Our findings underscore the need for a more nuanced understanding of the factors driving LLM responses before relying on them as substitutes for human subjects in survey settings. We emphasize the importance of researchers reporting results conditioned on specific prompt templates and exercising caution when drawing parallels between human behavior and LLMs.","authors":["Behnam Mohammadi"],"categories":["cs.CL","cs.AI","cs.LG"],"primary_category":"cs.CL","announce_type":"new","date":"2024-03-29","first_seen":"2024-03-29","revised_at":null,"abs_url":"https://arxiv.org/abs/2404.01332","pdf_url":"https://arxiv.org/pdf/2404.01332","source_feed":"backfill","score":8,"bucket":"selected","rubric_hits":["A1","A2","B1","B4"],"tags":["LLM仿真","Shapley值","认知偏差"],"reason":"用LLM仿真人类选择与认知偏差，并与真实人类数据对照，揭示仿真失效条件。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:14:41","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":32,"question":"如何利用Shapley值解释大语言模型决策，并揭示提示词中无信息量token对模型输出的不成比例影响（即“token噪声”效应）？","design":"本研究并非直接用LLM仿真人类，而是提出一种基于合作博弈Shapley值的模型无关解释方法，将提示词各组成部分视为“玩家”，量化其对LLM输出的相对贡献。通过两个应用展示：离散选择实验（航班选择）和认知偏差调查，分析提示词中token的影响。","baseline":"无对照","findings":"发现“token噪声”现象：LLM决策受到无信息量token（如冠词、介词、甚至“flight”等单词）的过度影响，且对格式变化（如换行符）高度敏感，导致选择概率出现混沌波动。这引发了对LLM仿真人类行为稳健性和泛化性的严重担忧。","reliability":"论文指出LLM对提示词变化高度敏感，无信息量token可显著改变输出，因此基于LLM的人类行为仿真在调查环境中不可靠，需谨慎解读，并建议研究者报告基于特定提示模板的条件结果。","relevance":"该研究直接批判了用LLM替代人类被试的可靠性，揭示了仿真失效的具体机制（token噪声），与研究者关注的仿真失效条件高度相关，值得精读以深入理解LLM行为偏差的来源。","inspiration":"可借鉴Shapley值分解方法，量化提示词各成分对LLM输出的影响，用于诊断和优化经济金融实验中的提示设计。｜可迁移到消费者金融决策仿真（如贷款选择、投资偏好），分析提示词中无关信息如何扭曲LLM的“偏好”。｜以GPT-4为被试，设计不同贷款方案的离散选择实验，在提示中系统变化无信息量token（如换行符、冠词），用Shapley值量化其影响，并以真实消费者信贷选择数据为基准，检验LLM仿真偏差。"}},{"id":"2403.16843","version":5,"title":"Do LLM Agents Have Regret? A Case Study in Online Learning and Games","zh_title":"LLM智能体有遗憾吗？在线学习与博弈案例研究","abstract":"Large language models (LLMs) have been increasingly employed for (interactive) decision-making, via the development of LLM-based autonomous agents. Despite their emerging successes, the performance of LLM agents in decision-making has not been fully investigated through quantitative metrics, especially in the multi-agent setting when they interact with each other, a typical scenario in real-world LLM-agent applications. To better understand the limits of LLM agents in these interactive environments, we propose to study their interactions in benchmark decision-making settings in online learning and game theory, through the performance metric of \\emph{regret}. We first empirically study the {no-regret} behaviors of LLMs in canonical (non-stationary) online learning problems, as well as the emergence of equilibria when LLM agents interact through playing repeated games. We then provide some theoretical insights into the no-regret behaviors of LLM agents, under certain assumptions on the supervised pre-training and the rationality model of human decision-makers who generate the data. Notably, we also identify (simple) cases where advanced LLMs such as GPT-4 fail to be no-regret. To promote the no-regret behaviors, we propose a novel \\emph{unsupervised} training loss of \\emph{regret-loss}, which, in contrast to the supervised pre-training loss, does not require the labels of (optimal) actions. We then establish the statistical guarantee of generalization bound for regret-loss minimization, followed by the optimization guarantee that minimizing such a loss may automatically lead to known no-regret learning algorithms. Our further experiments demonstrate the effectiveness of our regret-loss, especially in addressing the above ``regrettable'' cases.","authors":["Chanwoo Park","Xiangyu Liu","Asuman Ozdaglar","Kaiqing Zhang"],"categories":["cs.LG","cs.AI","cs.GT"],"primary_category":"cs.LG","announce_type":"new","date":"2024-03-25","first_seen":"2024-03-25","revised_at":null,"abs_url":"https://arxiv.org/abs/2403.16843","pdf_url":"https://arxiv.org/pdf/2403.16843","source_feed":"backfill","score":4,"bucket":"other","rubric_hits":["C1"],"tags":["多智能体","在线学习","博弈论"],"reason":"研究LLM智能体在在线学习和博弈中的遗憾行为，属多智能体决策分析，无人类行为对…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:27","error":null,"has_summary":false,"summary":null},{"id":"2403.15281","version":1,"title":"Measuring Gender and Racial Biases in Large Language Models","zh_title":"测量大语言模型中的性别与种族偏见","abstract":"In traditional decision making processes, social biases of human decision makers can lead to unequal economic outcomes for underrepresented social groups, such as women, racial or ethnic minorities. Recently, the increasing popularity of Large language model based artificial intelligence suggests a potential transition from human to AI based decision making. How would this impact the distributional outcomes across social groups? Here we investigate the gender and racial biases of OpenAIs GPT, a widely used LLM, in a high stakes decision making setting, specifically assessing entry level job candidates from diverse social groups. Instructing GPT to score approximately 361000 resumes with randomized social identities, we find that the LLM awards higher assessment scores for female candidates with similar work experience, education, and skills, while lower scores for black male candidates with comparable qualifications. These biases may result in a 1 or 2 percentage point difference in hiring probabilities for otherwise similar candidates at a certain threshold and are consistent across various job positions and subsamples. Meanwhile, we also find stronger pro female and weaker anti black male patterns in democratic states. Our results demonstrate that this LLM based AI system has the potential to mitigate the gender bias, but it may not necessarily cure the racial bias. Further research is needed to comprehend the root causes of these outcomes and develop strategies to minimize the remaining biases in AI systems. As AI based decision making tools are increasingly employed across diverse domains, our findings underscore the necessity of understanding and addressing the potential unequal outcomes to ensure equitable outcomes across social groups.","authors":["Jiafu An","Difang Huang","Chen Lin","Mingzhu Tai"],"categories":["econ.GN"],"primary_category":"econ.GN","announce_type":"new","date":"2024-03-22","first_seen":"2024-03-22","revised_at":null,"abs_url":"https://arxiv.org/abs/2403.15281","pdf_url":"https://arxiv.org/pdf/2403.15281","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A2","B1","B2","B4"],"tags":["LLM仿真","偏见测量","劳动力市场"],"reason":"用LLM替代人类决策者评估简历，测量偏见并与真实人类数据对照，涉及劳动力市场政…","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:58:47","error":null,"has_summary":true,"summary":{"generated_at":"2024-03-22","rank":1,"question":"在招聘决策中，GPT-3.5对不同性别和种族的求职者是否存在评分偏差？","design":"使用GPT-3.5模型扮演招聘决策者，对约36.1万份随机生成的工作经验、教育背景和技能组合的虚构简历进行评分（0-100分），简历随机分配带有性别和种族标识的姓名，以测量模型对不同社会群体的评分差异。","baseline":"无对照","findings":"GPT-3.5对女性求职者（无论种族）给予显著高于白人男性的评分，但对黑人男性给予显著低于白人男性的评分；在民主党州，亲女性偏差更强，反黑人男性偏差较弱。","reliability":"论文未讨论","relevance":"该研究直接使用LLM替代人类决策者进行高利害决策实验，测量了性别和种族偏见，虽未提供真实人类对照数据，但为评估LLM在劳动力市场决策中的偏差提供了重要证据，值得精读以了解仿真设计细节和偏差模式。","inspiration":"借鉴其通过大规模随机生成简历并操控姓名标识社会身份的实验设计，可精确分离LLM的偏见效应｜可迁移到信贷审批歧视研究，用LLM模拟信贷员评估贷款申请，操控申请人性别/种族｜用GPT-4扮演信贷员，对随机生成并分配不同种族姓名的贷款申请进行评分，结果变量为贷款批准概率，以真实银行信贷数据中的种族差异作为对照基准。"}},{"id":"2403.05534","version":1,"title":"Bayesian Preference Elicitation with Language Models","zh_title":"基于语言模型的贝叶斯偏好诱导","abstract":"Aligning AI systems to users' interests requires understanding and incorporating humans' complex values and preferences. Recently, language models (LMs) have been used to gather information about the preferences of human users. This preference data can be used to fine-tune or guide other LMs and/or AI systems. However, LMs have been shown to struggle with crucial aspects of preference learning: quantifying uncertainty, modeling human mental states, and asking informative questions. These challenges have been addressed in other areas of machine learning, such as Bayesian Optimal Experimental Design (BOED), which focus on designing informative queries within a well-defined feature space. But these methods, in turn, are difficult to scale and apply to real-world problems where simply identifying the relevant features can be difficult. We introduce OPEN (Optimal Preference Elicitation with Natural language) a framework that uses BOED to guide the choice of informative questions and an LM to extract features and translate abstract BOED queries into natural language questions. By combining the flexibility of LMs with the rigor of BOED, OPEN can optimize the informativity of queries while remaining adaptable to real-world domains. In user studies, we find that OPEN outperforms existing LM- and BOED-based methods for preference elicitation.","authors":["Kunal Handa","Yarin Gal","Ellie Pavlick","Noah Goodman","Jacob Andreas","Alex Tamkin","Belinda Z. Li"],"categories":["cs.CL"],"primary_category":"cs.CL","announce_type":"new","date":"2024-03-08","first_seen":"2024-03-08","revised_at":null,"abs_url":"https://arxiv.org/abs/2403.05534","pdf_url":"https://arxiv.org/pdf/2403.05534","source_feed":"backfill","score":5,"bucket":"other","rubric_hits":["D1"],"tags":["偏好诱导","贝叶斯实验设计","人机交互"],"reason":"用LLM引导偏好询问，替代人工标注偏好，非仿真人类被试行为。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:27","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":209,"question":"如何结合语言模型与贝叶斯最优实验设计，在开放域中高效地通过自然语言提问来推断用户偏好？","design":"非仿真研究。提出OPEN框架，用LM提取领域特征并初始化偏好先验，用贝叶斯模型选择信息量最大的成对比较问题，再由LM转化为自然语言提问，通过用户研究评估偏好推断准确率。","baseline":"无对照","findings":"在内容推荐领域的用户研究中，OPEN在偏好推断上优于纯LM和纯BOED方法。LM能灵活提取特征和生成自然语言，但单独使用时提问信息量不足；BOED能优化提问信息量，但难以处理开放域特征。","reliability":"论文未讨论","relevance":"本文研究如何用LM辅助偏好询问，而非用LLM仿真人类被试行为，不涉及人类基准对照或行为复现，与研究者关注的LLM仿真实验方向不直接相关，不建议优先阅读。","inspiration":"与经济金融研究关联不大"}},{"id":"2402.19421","version":1,"title":"Crafting Knowledge: Exploring the Creative Mechanisms of Chat-Based Search Engines","zh_title":"知识构建：探索基于聊天的搜索引擎的创造性机制","abstract":"In the domain of digital information dissemination, search engines act as pivotal conduits linking information seekers with providers. The advent of chat-based search engines utilizing Large Language Models (LLMs) and Retrieval Augmented Generation (RAG), exemplified by Bing Chat, marks an evolutionary leap in the search ecosystem. They demonstrate metacognitive abilities in interpreting web information and crafting responses with human-like understanding and creativity. Nonetheless, the intricate nature of LLMs renders their \"cognitive\" processes opaque, challenging even their designers' understanding. This research aims to dissect the mechanisms through which an LLM-powered chat-based search engine, specifically Bing Chat, selects information sources for its responses. To this end, an extensive dataset has been compiled through engagements with New Bing, documenting the websites it cites alongside those listed by the conventional search engine. Employing natural language processing (NLP) techniques, the research reveals that Bing Chat exhibits a preference for content that is not only readable and formally structured, but also demonstrates lower perplexity levels, indicating a unique inclination towards text that is predictable by the underlying LLM. Further enriching our analysis, we procure an additional dataset through interactions with the GPT-4 based knowledge retrieval API, unveiling a congruent text preference between the RAG API and Bing Chat. This consensus suggests that these text preferences intrinsically emerge from the underlying language models, rather than being explicitly crafted by Bing Chat's developers. Moreover, our investigation documents a greater similarity among websites cited by RAG technologies compared to those ranked highest by conventional search engines.","authors":["Lijia Ma","Xingchen Xu","Yong Tan"],"categories":["cs.IR","cs.AI","econ.GN"],"primary_category":"cs.IR","announce_type":"new","date":"2024-02-29","first_seen":"2024-02-29","revised_at":null,"abs_url":"https://arxiv.org/abs/2402.19421","pdf_url":"https://arxiv.org/pdf/2402.19421","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C4"],"tags":["搜索引擎","信息检索","语言模型偏好"],"reason":"研究搜索引擎的信息选择机制，属纯NLP能力分析，不以人类行为为参照系。","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:58:47","error":null,"has_summary":false,"summary":null},{"id":"2402.18144","version":1,"title":"Random Silicon Sampling: Simulating Human Sub-Population Opinion Using a Large Language Model Based on Group-Level Demographic Information","zh_title":"随机硅采样：基于群体人口统计信息用大语言模型模拟人类子群体意见","abstract":"Large language models exhibit societal biases associated with demographic information, including race, gender, and others. Endowing such language models with personalities based on demographic data can enable generating opinions that align with those of humans. Building on this idea, we propose \"random silicon sampling,\" a method to emulate the opinions of the human population sub-group. Our study analyzed 1) a language model that generates the survey responses that correspond with a human group based solely on its demographic distribution and 2) the applicability of our methodology across various demographic subgroups and thematic questions. Through random silicon sampling and using only group-level demographic information, we discovered that language models can generate response distributions that are remarkably similar to the actual U.S. public opinion polls. Moreover, we found that the replicability of language models varies depending on the demographic group and topic of the question, and this can be attributed to inherent societal biases in the models. Our findings demonstrate the feasibility of mirroring a group's opinion using only demographic distribution and elucidate the effect of social biases in language models on such simulations.","authors":["Seungjong Sun","Eungu Lee","Dongyan Nan","Xiangying Zhao","Wonbyung Lee","Bernard J. Jansen","Jang Hyun Kim"],"categories":["cs.AI","cs.CY"],"primary_category":"cs.AI","announce_type":"new","date":"2024-02-28","first_seen":"2024-02-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2402.18144","pdf_url":"https://arxiv.org/pdf/2402.18144","source_feed":"api","score":10,"bucket":"selected","rubric_hits":["A1","A2","A5","B1","B4"],"tags":["LLM人类仿真","意见模拟","算法偏差"],"reason":"用LLM基于人口统计分布模拟人群意见，并与真实民调对照，评估偏差与可复现性，直…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:14:41","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":6,"question":"能否仅基于群体层面的人口统计分布，利用大语言模型生成与真实人群意见分布相似的调查回答？","design":"使用GPT-3.5等LLM，根据目标人群的群体人口统计分布随机生成合成个体（random silicon sample），将个体人口统计信息与调查问题一起作为提示输入模型，收集回答并汇总为群体意见分布。","baseline":"以美国皮尤研究中心等机构的真实民意调查数据作为对照基准。","findings":"仅用群体人口统计信息，LLM生成的回答分布与真实民调高度相似；但复现性因人口子群和问题主题而异，这种差异可归因于模型固有的社会偏见。","reliability":"论文指出复现性受目标群体和问题主题影响，模型对特定人群和话题的偏见会导致仿真失效，且方法依赖群体分布假设，未考虑个体层面差异。","relevance":"该研究直接探索用LLM替代人类被试进行民意调查仿真，并与真实数据严格对照，评估了可靠性与偏差条件，完全契合研究者对LLM人类仿真实验、基准对照和失效分析的关注，值得精读原文。","inspiration":"借鉴其仅用群体分布生成合成样本并汇总意见的方法，可低成本构建虚拟被试池进行政策态度预测试。｜可迁移到经济政策评估场景，如模拟不同收入群体对税收改革的态度分布。｜以收入、教育、地区等群体分布生成虚拟纳税人，施加税收政策描述作为处理，测量支持率，用真实社会调查数据做对照。"}},{"id":"2402.01053","version":1,"title":"Plan-Grounded Large Language Models for Dual Goal Conversational Settings","zh_title":"基于计划的大语言模型用于双目标对话设置","abstract":"Training Large Language Models (LLMs) to follow user instructions has been shown to supply the LLM with ample capacity to converse fluently while being aligned with humans. Yet, it is not completely clear how an LLM can lead a plan-grounded conversation in mixed-initiative settings where instructions flow in both directions of the conversation, i.e. both the LLM and the user provide instructions to one another. In this paper, we tackle a dual goal mixed-initiative conversational setting where the LLM not only grounds the conversation on an arbitrary plan but also seeks to satisfy both a procedural plan and user instructions. The LLM is then responsible for guiding the user through the plan and, at the same time, adapting to new circumstances, answering questions, and activating safety guardrails when needed. We propose a novel LLM that grounds the dialogue on a procedural plan, can take the dialogue initiative, and enforces guardrails on the system's behavior, while also improving the LLM's responses to unexpected user behavior. Experiments in controlled settings and with real users show that the best-performing model, which we call PlanLLM, achieves a 2.1x improvement over a strong baseline. Moreover, experiments also show good generalization to unseen domains.","authors":["Diogo Glória-Silva","Rafael Ferreira","Diogo Tavares","David Semedo","João Magalhães"],"categories":["cs.CL","cs.AI"],"primary_category":"cs.CL","announce_type":"new","date":"2024-02-01","first_seen":"2024-02-01","revised_at":null,"abs_url":"https://arxiv.org/abs/2402.01053","pdf_url":"https://arxiv.org/pdf/2402.01053","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["对话系统","计划引导","混合主动"],"reason":"研究LLM在混合主动对话中遵循计划并引导用户，属于角色扮演对话系统，无实验或测…","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:59:21","error":null,"has_summary":false,"summary":null},{"id":"2402.01766","version":3,"title":"LLM Voting: Human Choices and AI Collective Decision Making","zh_title":"LLM投票：人类选择与AI集体决策","abstract":"This paper investigates the voting behaviors of Large Language Models (LLMs), specifically GPT-4 and LLaMA-2, their biases, and how they align with human voting patterns. Our methodology involved using a dataset from a human voting experiment to establish a baseline for human preferences and conducting a corresponding experiment with LLM agents. We observed that the choice of voting methods and the presentation order influenced LLM voting outcomes. We found that varying the persona can reduce some of these biases and enhance alignment with human choices. While the Chain-of-Thought approach did not improve prediction accuracy, it has potential for AI explainability in the voting process. We also identified a trade-off between preference diversity and alignment accuracy in LLMs, influenced by different temperature settings. Our findings indicate that LLMs may lead to less diverse collective outcomes and biased assumptions when used in voting scenarios, emphasizing the need for cautious integration of LLMs into democratic processes.","authors":["Joshua C. Yang","Damian Dailisan","Marcin Korecki","Carina I. Hausladen","Dirk Helbing"],"categories":["cs.CL","cs.AI","cs.CY","cs.LG","econ.GN"],"primary_category":"cs.CL","announce_type":"new","date":"2024-01-31","first_seen":"2024-01-31","revised_at":null,"abs_url":"https://arxiv.org/abs/2402.01766","pdf_url":"https://arxiv.org/pdf/2402.01766","source_feed":"backfill","score":9,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM仿真","投票行为","人类对照"],"reason":"用LLM复现人类投票实验，有真实人类数据对照，涉及集体决策与偏差评估。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:14:41","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":23,"question":"LLM（GPT-4和LLaMA-2）在参与式预算投票中的行为与人类投票模式的对齐程度如何，存在哪些偏差？","design":"使用GPT-4 Turbo和LLaMA-2 70B模型模拟180名人类被试，在相同的参与式预算投票实验中，对24个城市项目进行投票。实验操纵了四种投票方法（批准投票、5-批准投票、累积投票、排序投票）和项目呈现顺序，并测试了角色设定（persona）和思维链提示的影响。结果变量包括聚合偏好相似度（Kendall's τ）、个体投票相似度（Jaccard）和偏好多样性。","baseline":"来自Yang et al. (2024)的180名大学生在苏黎世参与式预算在线实验中的真实投票数据。","findings":"投票方法和呈现顺序会影响LLM的投票结果；改变角色设定可以减少偏差并提高与人类选择的一致性。思维链提示未提高预测准确性，但有助于投票过程的可解释性；温度设置导致偏好多样性与对齐准确性之间存在权衡。","reliability":"论文指出LLM在投票场景中可能导致集体结果多样性降低和偏差假设，强调需谨慎将LLM整合进民主过程；未深入讨论其他失效条件。","relevance":"该研究直接使用LLM复现人类投票实验，有真实人类数据对照，评估了仿真可靠性、偏差及对齐方法，高度契合研究者对经济学实验和政策评估场景中LLM仿真批判性分析的兴趣，值得精读原文。","inspiration":"该方法借鉴了多模型对比（GPT-4 vs. LLaMA-2）和提示工程干预（角色设定、思维链）来评估LLM与人类行为对齐程度的设计，并揭示了温度参数在多样性与准确性间的权衡。｜可迁移到政策公告的预期形成实验，例如研究央行沟通对通胀预期的影响。｜以LLM作为被试，模拟不同措辞和框架的央行声明（处理），测量其预测通胀的分布和锚定效应（结果变量），并与真实家庭或专家调查数据（如密歇根消费者调查）进行对照。"}},{"id":"2401.15589","version":1,"title":"OpineBot: Class Feedback Reimagined Using a Conversational LLM","zh_title":"OpineBot：用对话式大语言模型重塑课堂反馈","abstract":"Conventional class feedback systems often fall short, relying on static, unengaging surveys offering little incentive for student participation. To address this, we present OpineBot, a novel system employing large language models (LLMs) to conduct personalized, conversational class feedback via chatbot interface. We assessed OpineBot's effectiveness in a user study with 20 students from an Indian university's Operating-Systems class, utilizing surveys and interviews to analyze their experiences. Findings revealed a resounding preference for OpineBot compared to conventional methods, highlighting its ability to engage students, produce deeper feedback, offering a dynamic survey experience. This research represents a work in progress, providing early results, marking a significant step towards revolutionizing class feedback through LLM-based technology, promoting student engagement, and leading to richer data for instructors. This ongoing research presents preliminary findings and marks a notable advancement in transforming classroom feedback using LLM-based technology to enhance student engagement and generate comprehensive data for educators.","authors":["Henansh Tanwar","Kunal Shrivastva","Rahul Singh","Dhruv Kumar"],"categories":["cs.HC","cs.CY"],"primary_category":"cs.HC","announce_type":"new","date":"2024-01-28","first_seen":"2024-01-28","revised_at":null,"abs_url":"https://arxiv.org/abs/2401.15589","pdf_url":"https://arxiv.org/pdf/2401.15589","source_feed":"backfill","score":0,"bucket":"other","rubric_hits":["C3"],"tags":["课堂反馈","对话系统","人机交互"],"reason":"对话式反馈系统，LLM 作为聊天机器人收集学生意见，非仿真人类被试，无实验或测…","model":"deepseek-v4-pro","scored_at":"2026-08-10T23:59:36","error":null,"has_summary":false,"summary":null},{"id":"2401.07345","version":3,"title":"Can an LLM Learn Preferences from Choice Data?","zh_title":"大语言模型能从选择数据中学习偏好吗？","abstract":"Can large language models (LLMs) learn a decision maker's preferences from observed choices and generate preference-consistent recommendations in new situations? We propose a portable Simulate-Recommend-Evaluate framework that tests preference learning from revealed-choice data by comparing LLM recommendations with optimal choices implied by known preference primitives. We apply the framework to choice under uncertainty using the disappointment aversion model. Recommendation accuracy improves as models observe more choices, but learning is heterogeneous across preference types and LLMs: GPT learns risk aversion better than disappointment aversion, Gemini performs best in high disappointment-aversion regions, and Claude shows the broadest effective learning across parameter regions.","authors":["Jeongbin Kim","Matthew Kovach","Kyu-Min Lee","Euncheol Shin","Hector Tzavellas"],"categories":["econ.GN"],"primary_category":"econ.GN","announce_type":"new","date":"2024-01-14","first_seen":"2024-01-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2401.07345","pdf_url":"https://arxiv.org/pdf/2401.07345","source_feed":"backfill","score":7,"bucket":"pending","rubric_hits":["A1","B1","B2"],"tags":["偏好学习","选择实验","经济学决策"],"reason":"用LLM从选择数据学习偏好并推荐，与人类最优选择对照，涉及经济学决策场景，方法…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:14:38","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":119,"question":"大语言模型能否从观察到的选择数据中学习决策者偏好，并在新情境下生成与偏好一致的建议？","design":"提出模拟-推荐-评估（SRE）框架，用失望厌恶模型生成已知偏好的选择数据，将LLM作为推荐系统仅提供选择数据，测试其在新预算问题上的推荐准确性，并用非参数和参数指标评估。","baseline":"以失望厌恶模型隐含的最优选择作为真实偏好基准。","findings":"推荐准确性随观察数据量增加而提升，但学习效果在偏好空间上异质；GPT学习风险厌恶优于失望厌恶，Gemini在高失望厌恶区域表现最佳，Claude在各参数区域均展现出广泛的有效学习。","reliability":"论文未讨论","relevance":"该研究直接评估LLM从选择数据学习个体偏好并生成建议的能力，使用经济学实验范式且有真实偏好基准，符合对LLM仿真可靠性及失效条件的关注，值得精读。","inspiration":"该研究提出的模拟-推荐-评估（SRE）框架，通过已知偏好模型生成选择数据来训练LLM，并用非参数和参数指标评估推荐准确性，为检验LLM学习经济决策偏好的能力提供了严谨的实验设计范式。｜这一方法可迁移到消费者跨期选择研究，例如测试LLM能否从个体的时间偏好选择数据中学习其折现因子，并预测在新预算约束下的消费-储蓄决策。｜可设计实验：以拟双曲贴现模型生成具有已知时间偏好的虚拟被试选择数据，将LLM作为推荐系统仅提供历史选择记录，测试其在新跨期问题上的推荐准确性，并以模型隐含的最优选择作为真实偏好基准，比较不同LLM的学习效果。"}},{"id":"2304.03442","version":2,"title":"Generative Agents: Interactive Simulacra of Human Behavior","zh_title":"生成式智能体：人类行为的交互式模拟","abstract":"Believable proxies of human behavior can empower interactive applications ranging from immersive environments to rehearsal spaces for interpersonal communication to prototyping tools. In this paper, we introduce generative agents--computational software agents that simulate believable human behavior. Generative agents wake up, cook breakfast, and head to work; artists paint, while authors write; they form opinions, notice each other, and initiate conversations; they remember and reflect on days past as they plan the next day. To enable generative agents, we describe an architecture that extends a large language model to store a complete record of the agent's experiences using natural language, synthesize those memories over time into higher-level reflections, and retrieve them dynamically to plan behavior. We instantiate generative agents to populate an interactive sandbox environment inspired by The Sims, where end users can interact with a small town of twenty five agents using natural language. In an evaluation, these generative agents produce believable individual and emergent social behaviors: for example, starting with only a single user-specified notion that one agent wants to throw a Valentine's Day party, the agents autonomously spread invitations to the party over the next two days, make new acquaintances, ask each other out on dates to the party, and coordinate to show up for the party together at the right time. We demonstrate through ablation that the components of our agent architecture--observation, planning, and reflection--each contribute critically to the believability of agent behavior. By fusing large language models with computational, interactive agents, this work introduces architectural and interaction patterns for enabling believable simulations of human behavior.","authors":["Joon Sung Park","Joseph C. O'Brien","Carrie J. Cai","Meredith Ringel Morris","Percy Liang","Michael S. Bernstein"],"categories":["cs.HC","cs.AI","cs.LG"],"primary_category":"cs.HC","announce_type":"new","date":"2023-04-07","first_seen":"2023-04-07","revised_at":null,"abs_url":"https://arxiv.org/abs/2304.03442","pdf_url":"https://arxiv.org/pdf/2304.03442","source_feed":"api","score":8,"bucket":"selected","rubric_hits":["A3","B1"],"tags":["LLM社会模拟","人类行为仿真","智能体架构"],"reason":"用LLM agent模拟小镇社会行为，有真实人类行为对照，但非严格实验或政策评…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:17:27","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":31,"question":"如何利用大语言模型构建能够产生可信个体和涌现社会行为的生成式智能体？","design":"使用大语言模型（ChatGPT）构建25个生成式智能体，置于类似《模拟人生》的沙盒环境中。每个智能体拥有记忆流（存储经历）、反思（合成高层推断）和规划（生成行动计划）模块。通过用户指定一个智能体想举办情人节派对这一初始条件，观察智能体自主产生的行为，如传播邀请、约会、协调参加派对等。","baseline":"无对照","findings":"生成式智能体能够产生可信的个体行为和涌现的社会行为，如信息扩散、关系建立和群体协调。消融实验表明，记忆流、反思和规划三个组件对行为可信性均有关键贡献。","reliability":"论文未讨论","relevance":"该研究展示了LLM智能体在模拟社会互动和涌现行为方面的潜力，但缺乏与真实人类数据的严格对照，且非经济学实验或政策评估场景，与研究者关注的经济学实验和政策评估的直接相关性有限。","inspiration":"与经济金融研究关联不大"}},{"id":"2301.07543","version":2,"title":"Large Language Models as Simulated Economic Agents: What Can We Learn from Homo Silicus?","zh_title":"作为模拟经济主体的大语言模型：我们能从Homo Silicus中学到什么？","abstract":"We argue that newly-developed large language models (LLMs), because of how they are trained and designed, are implicit computational models of humans -- a Homo silicus. LLMs can be used like economists use Homo economicus: they can be given endowments, information, preferences, and so on, and then their behavior can be explored in scenarios via simulation. Experiments using this approach, derived from Charness and Rabin (2002), Kahneman et al. (1986), Samuelson and Zeckhauser (1988), Oprea (2024b), and Horton (2025), show qualitatively similar results to the original, and when they differ, it is often generative for future research. We discuss potential applications, conceptual issues, and why this approach can inform the study of humans.","authors":["John J. Horton","Apostolos Filippas","Benjamin S. Manning"],"categories":["econ.GN"],"primary_category":"econ.GN","announce_type":"new","date":"2023-01-18","first_seen":"2023-01-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2301.07543","pdf_url":"https://arxiv.org/pdf/2301.07543","source_feed":"api","score":10,"bucket":"selected","rubric_hits":["A1","A3","B1","B2"],"tags":["LLM仿真","经济实验","人类行为对照"],"reason":"直接用LLM模拟经济实验并与真实人类数据对照，核心相关。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:14:38","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":18,"question":"大语言模型能否作为人类经济行为的计算模型（Homo silicus），通过仿真实验复现经典经济实验结果，并用于理解人类行为？","design":"使用大语言模型（如GPT系列）作为AI智能体，赋予其禀赋、信息、偏好等，通过文本提示模拟五种经典经济实验场景：公平性判断（Kahneman et al., 1986）、独裁者博弈（Charness and Rabin, 2002）、现状偏差（Samuelson and Zeckhauser, 1988）、风险态度（Oprea, 2024b）和招聘场景（Horton, 2025），测量AI智能体的回答或选择行为。","baseline":"对照的真实人类数据来自上述五篇原始文献中的人类被试行为结果。","findings":"AI仿真结果在定性上与原始人类实验相似，例如公平判断受涨价幅度和政治倾向影响、独裁者博弈中赋予不同社会偏好会改变选择、现状偏差可被复现；当结果存在差异时，往往能为未来研究提供新思路。","reliability":"论文指出LLM训练数据可能包含已发表研究结果，导致仿真仅机械复述记忆而非真正模拟人类行为；训练语料、模型不透明性和仿真泛化能力均构成局限。","relevance":"该研究直接使用LLM进行经济实验仿真并与真实人类数据对照，涵盖公平、博弈、偏差、风险决策等多个经典主题，并讨论了仿真失效条件，高度契合研究者对LLM人类仿真可靠性及批判性评估的兴趣，值得精读原文。","inspiration":"该方法借鉴了用文本提示赋予LLM特定禀赋、信息和社会偏好来模拟经济决策，并直接与经典实验的人类基准数据对照｜可迁移到资产定价实验，如研究投资者在泡沫或崩盘情境下的交易行为与风险偏好｜以LLM为被试，通过提示设定初始财富、市场信息和风险态度，测量其买卖报价与持仓变化，对照Smith et al. (1988)等经典资产泡沫实验的人类数据"}},{"id":"2209.06899","version":1,"title":"Out of One, Many: Using Language Models to Simulate Human Samples","zh_title":"一生万物：使用语言模型模拟人类样本","abstract":"We propose and explore the possibility that language models can be studied as effective proxies for specific human sub-populations in social science research. Practical and research applications of artificial intelligence tools have sometimes been limited by problematic biases (such as racism or sexism), which are often treated as uniform properties of the models. We show that the \"algorithmic bias\" within one such tool -- the GPT-3 language model -- is instead both fine-grained and demographically correlated, meaning that proper conditioning will cause it to accurately emulate response distributions from a wide variety of human subgroups. We term this property \"algorithmic fidelity\" and explore its extent in GPT-3. We create \"silicon samples\" by conditioning the model on thousands of socio-demographic backstories from real human participants in multiple large surveys conducted in the United States. We then compare the silicon and human samples to demonstrate that the information contained in GPT-3 goes far beyond surface similarity. It is nuanced, multifaceted, and reflects the complex interplay between ideas, attitudes, and socio-cultural context that characterize human attitudes. We suggest that language models with sufficient algorithmic fidelity thus constitute a novel and powerful tool to advance understanding of humans and society across a variety of disciplines.","authors":["Lisa P. Argyle","Ethan C. Busby","Nancy Fulda","Joshua Gubler","Christopher Rytting","David Wingate"],"categories":["cs.LG","cs.CL"],"primary_category":"cs.LG","announce_type":"new","date":"2022-09-14","first_seen":"2022-09-14","revised_at":null,"abs_url":"https://arxiv.org/abs/2209.06899","pdf_url":"https://arxiv.org/pdf/2209.06899","source_feed":"api","score":10,"bucket":"selected","rubric_hits":["A1","A2","A3","B1","B2"],"tags":["LLM仿真","算法保真度","社会调查"],"reason":"直接提出用GPT-3模拟人类子群体，并与真实调查数据对照，验证算法保真度，高度…","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:14:38","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":4,"question":"大型语言模型（如GPT-3）能否通过条件化生成，准确模拟特定人类子群体的态度和反应分布，从而作为社会科学研究中的人类被试替代品？","design":"使用GPT-3模型，通过输入真实调查参与者的社会人口背景故事（来自ANES等大型调查）作为条件，生成“硅样本”虚拟被试，然后让这些虚拟被试完成与人类相同的任务（自由联想、投票预测、封闭式问题），比较硅样本与人类样本的响应分布。","baseline":"2012、2016、2020年美国国家选举研究（ANES）和Rothschild等人的“Pigeonholing Partisans”数据中真实人类参与者的调查回答。","findings":"GPT-3的算法偏差并非单一宏观属性，而是细粒度且与人口统计特征相关，通过适当条件化可精确模拟多种人类子群体的响应分布。硅样本与人类样本在态度、观念和社会文化背景的复杂交互模式上高度一致，表明GPT-3具有较高的“算法保真度”。","reliability":"论文未讨论","relevance":"该研究直接探索用LLM替代人类被试进行仿真实验，并与真实调查数据严格对照，验证了算法保真度，高度契合研究者对LLM人类仿真可靠性及基准对照的关注，值得精读原文。","inspiration":"借鉴其“硅采样”方法：用真实个体的多维人口背景作为条件提示，生成虚拟被试并测量其态度/行为，再与人类基准数据对比以评估仿真效度。｜可迁移至消费者信心调查或政策偏好预测，例如模拟不同收入、教育、地域群体对通胀预期或税收政策的反应。｜以GPT-4为被试，输入来自美国消费者财务调查（SCF）的家庭人口与财务背景，生成虚拟消费者，询问其未来一年通胀预期，以密歇根大学消费者调查的微观数据作为人类基准，比较分布与相关性。"}},{"id":"2208.10264","version":5,"title":"Using Large Language Models to Simulate Multiple Humans and Replicate Human Subject Studies","zh_title":"使用大语言模型模拟多个人类并复现人类被试研究","abstract":"We introduce a new type of test, called a Turing Experiment (TE), for evaluating to what extent a given language model, such as GPT models, can simulate different aspects of human behavior. A TE can also reveal consistent distortions in a language model's simulation of a specific human behavior. Unlike the Turing Test, which involves simulating a single arbitrary individual, a TE requires simulating a representative sample of participants in human subject research. We carry out TEs that attempt to replicate well-established findings from prior studies. We design a methodology for simulating TEs and illustrate its use to compare how well different language models are able to reproduce classic economic, psycholinguistic, and social psychology experiments: Ultimatum Game, Garden Path Sentences, Milgram Shock Experiment, and Wisdom of Crowds. In the first three TEs, the existing findings were replicated using recent models, while the last TE reveals a \"hyper-accuracy distortion\" present in some language models (including ChatGPT and GPT-4), which could affect downstream applications in education and the arts.","authors":["Gati Aher","Rosa I. Arriaga","Adam Tauman Kalai"],"categories":["cs.CL","cs.AI","cs.LG"],"primary_category":"cs.CL","announce_type":"new","date":"2022-08-18","first_seen":"2022-08-18","revised_at":null,"abs_url":"https://arxiv.org/abs/2208.10264","pdf_url":"https://arxiv.org/pdf/2208.10264","source_feed":"api","score":10,"bucket":"selected","rubric_hits":["A1","A2","A3","B1","B2","B4"],"tags":["LLM仿真","人类实验复现","行为经济学"],"reason":"直接复现经典人类实验，用LLM模拟被试并与真实人类数据对照，评估仿真偏差。","model":"deepseek-v4-pro","scored_at":"2026-07-28T15:14:38","error":null,"has_summary":true,"summary":{"generated_at":"2026-07-28","rank":5,"question":"如何系统评估语言模型在模拟人类行为时的忠实程度与系统性扭曲？","design":"提出图灵实验（TE）方法，使用GPT等语言模型，通过零样本提示模拟具有不同姓名和性别称谓的多样化被试样本，在最后通牒博弈、花园路径句、米尔格拉姆电击实验和群体智慧四个经典实验中施加相应刺激，测量接受/拒绝、语法判断等结果变量。","baseline":"对照各经典实验已有的真实人类被试研究结果。","findings":"在前三个TE中，近期模型成功复现了已有发现；在群体智慧TE中，部分模型（包括ChatGPT和GPT-4）表现出“超准确性扭曲”，即模拟的群体估计过于准确，偏离了真实人类群体的典型误差模式。","reliability":"论文指出，零样本要求难以完全保证，因为预训练语料可能已包含相关实验数据；此外，仅用姓名和性别称谓模拟多样性可能不足以捕捉真实人群差异。","relevance":"该研究直接复现经典人类实验，用LLM模拟被试并与真实人类数据对照，评估仿真偏差，高度契合研究者对LLM人类仿真可靠性及失效条件的关注，值得精读原文。","inspiration":"借鉴其通过姓名和称谓简单操控被试身份以模拟多样性的设计，以及用经典实验范式作为基准测试LLM行为复现能力的方法。｜可迁移到行为经济学中的最后通牒博弈、信任博弈等实验，检验LLM是否能复现真实人类的公平偏好或互惠行为。｜以GPT-4为被试，模拟不同姓名（暗示种族/性别）的个体在最后通牒博弈中的响应，处理为不同的提议金额，结果变量为接受/拒绝，对照真实人类实验的元分析数据，评估LLM是否复现已知的公平偏好及群体差异。"}}]}